Can I run Qwen3 30B-A3B (2507) on the Apple M4 Pro?
Runs greatQ4_K_M fits in unified memory (19.9 of 33.6 GB usable by the GPU): Runs great, est. 29.1 tok/s (23.3–34.9, calibrated estimate ±20%).
48 GB unified memory · 273.0 GB/s · FP16 17.0 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Verdict by memory size
The page uses the 48 GB configuration; here is every size.
- 24 GBRuns slowlyest. 29.1 tok/s (20.4–37.8, theoretical estimate ±30%)
- 48 GBRuns greatest. 29.1 tok/s (23.3–34.9, calibrated estimate ±20%)
- 64 GBRuns greatest. 29.1 tok/s (23.3–34.9, calibrated estimate ±20%)
Every quant of Qwen3 30B-A3B (2507) on the Apple M4 Pro
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 11.26 GB | Runs great | est. 40.5 tok/s32.4–48.6calibrated estimate ±20% | 12.6 / 33.6 GB | Unified memory | — |
| IQ4_XS | 16.38 GB | Runs great | est. 31.8 tok/s25.4–38.1calibrated estimate ±20% | 17.8 / 33.6 GB | Unified memory | — |
| Q4_K_Mbaseline | 18.56 GB | Runs great | est. 29.1 tok/s23.3–34.9calibrated estimate ±20% | 19.9 / 33.6 GB | Unified memory | — |
| Q8_0 | 32.48 GB | Runs slowly | est. 19.0 tok/s13.3–24.6theoretical estimate ±30% | 33.3 GB RAM | CPU only | CPU inference — capped at Runs slowly |
Where the memory goes at Q4_K_M
- Weights
- 18.6 GB
- KV cache
- 0.8 GB
- Compute buffer
- 0.6 GB
- Free
- 13.7 GB
- Not usable by the GPU
- 14.4 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 19.5 GBRuns great est. 34.0 tok/s27.2–40.8calibrated estimate ±20% | 19.3 GBRuns great est. 36.8 tok/s29.5–44.2calibrated estimate ±20% | 19.2 GBRuns great est. 38.6 tok/s30.9–46.3calibrated estimate ±20% |
| 8k | 19.9 GBRuns great est. 29.1 tok/s23.3–34.9calibrated estimate ±20% | 19.6 GBRuns great est. 33.6 tok/s26.9–40.3calibrated estimate ±20% | 19.4 GBRuns great est. 36.6 tok/s29.3–43.9calibrated estimate ±20% |
| 16k | 20.8 GBRuns great est. 22.6 tok/s18.1–27.2calibrated estimate ±20% | 20.1 GBRuns great est. 28.5 tok/s22.8–34.2calibrated estimate ±20% | 19.7 GBRuns great est. 33.2 tok/s26.6–39.8calibrated estimate ±20% |
| 32k | 22.6 GBRuns well est. 15.7 tok/s12.5–18.8calibrated estimate ±20% | 21.1 GBRuns great est. 21.9 tok/s17.6–26.3calibrated estimate ±20% | 20.3 GBRuns great est. 28.0 tok/s22.4–33.6calibrated estimate ±20% |
| 64k | 26.1 GBRuns well est. 9.7 tok/s7.8–11.6calibrated estimate ±20% | 23.1 GBRuns well est. 15.0 tok/s12.0–18.0calibrated estimate ±20% | 21.5 GBRuns great est. 21.3 tok/s17.0–25.6calibrated estimate ±20% |
| 128k | 33.1 GBRuns slowly est. 5.5 tok/s4.4–6.6calibrated estimate ±20% | 27.2 GBRuns well est. 9.2 tok/s7.4–11.0calibrated estimate ±20% | 23.9 GBRuns well est. 14.4 tok/s11.5–17.3calibrated estimate ±20% |
| 256k | 47.2 GBdoes not fit | 32.3 GBRuns slowly est. 5.2 tok/s3.6–6.7theoretical estimate ±30% | 28.8 GBRuns well est. 8.8 tok/s7.0–10.5calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs great up to 16k context with an f16 KV cache, and up to 32k with q8_0.
Estimated speed
- Generation
- est. 29.1 tok/s (23.3–34.9, calibrated estimate ±20%)
- Prompt processing
- about 765 tok/s (459–1,071, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A larger model that also runs well on the Apple M4 Pro
Runs greatQwen3.5 35B-A3B: Runs great, est. 39.8 tok/s (31.9–47.8, calibrated estimate ±20%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf (18.56 GB) from unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF.
llama-server -hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.43 GB at 8k context.
Terminal:
OLLAMA_CONTEXT_LENGTH=8192 ollama serve
ollama run hf.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: Running local LLMs on a Mac (M4, M5): how much unified memory?
Other models on the Apple M4 Pro
Qwen3 30B-A3B (2507) on other hardware
Frequently asked questions
Can I run Qwen3 30B-A3B (2507) on the Apple M4 Pro?
Q4_K_M fits in unified memory (19.9 of 33.6 GB usable by the GPU): Runs great, est. 29.1 tok/s (23.3–34.9, calibrated estimate ±20%). Q4_K_M needs 19.9 GB at 8k context; the GPU can use 33.6 GB of the unified memory.
How much context can Qwen3 30B-A3B (2507) use on the Apple M4 Pro?
At Q4_K_M the verdict stays Runs great up to 16k context with an f16 KV cache (1.6 GB of KV) and up to 32k with a q8_0 KV cache (1.7 GB). The model supports up to 256k.
Which quant should I use, and how fast is it?
Q4_K_M (18.56 GB) is the recommended balance: Runs great, est. 29.1 tok/s (23.3–34.9, calibrated estimate ±20%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.