Can I run Qwen3 30B-A3B (2507) on the Apple M4?

Runs wellQ4_K_M fits in unified memory (19.9 of 22.4 GB usable by the GPU): Runs well, est. 12.8 tok/s (10.2–15.4, calibrated estimate ±20%).

Unified memory — by default the GPU can use about 70% of itOpen license · Apache-2.0Community GGUF · unsloth

32 GB unified memory · 120.0 GB/s · FP16 8.5 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Verdict by memory size

The page uses the 32 GB configuration; here is every size.

  • 16 GBWon't run
  • 24 GBRuns slowlyest. 12.8 tok/s (9.0–16.6, theoretical estimate ±30%)
  • 32 GBRuns wellest. 12.8 tok/s (10.2–15.4, calibrated estimate ±20%)

Every quant of Qwen3 30B-A3B (2507) on the Apple M4

4 tracked GGUF files, smallest first, at 8k context
QuantFile sizeVerdictSpeedMemoryRuns asNotes
Q2_K11.26 GBRuns well
est. 17.8 tok/s14.2–21.3calibrated estimate ±20%
12.6 / 22.4 GBUnified memory—
IQ4_XS16.38 GBRuns well
est. 14.0 tok/s11.2–16.8calibrated estimate ±20%
17.8 / 22.4 GBUnified memory—
Q4_K_Mbaseline18.56 GBRuns well
est. 12.8 tok/s10.2–15.4calibrated estimate ±20%
19.9 / 22.4 GBUnified memory—
Q8_032.48 GBWon't run—Needs about 33.9 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)——

Where the memory goes at Q4_K_M

Weights, KV cache and compute buffer are the model; the OS reservation is the display driver.
Unified memory
Weights
18.6 GB
KV cache
0.8 GB
Compute buffer
0.6 GB
Free
2.5 GB
Not usable by the GPU
9.6 GB

Context length vs. KV cache

KV cache

Memory needed and verdict at Q4_K_M for each context length and KV cache type, with estimated tokens per second.
ContextKV f16KV q8_0KV q4_0
4k
19.5 GBRuns well
est. 14.9 tok/s11.9–17.9calibrated estimate ±20%
19.3 GBRuns well
est. 16.2 tok/s13.0–19.4calibrated estimate ±20%
19.2 GBRuns well
est. 17.0 tok/s13.6–20.3calibrated estimate ±20%
8k
19.9 GBRuns well
est. 12.8 tok/s10.2–15.4calibrated estimate ±20%
19.6 GBRuns well
est. 14.8 tok/s11.8–17.7calibrated estimate ±20%
19.4 GBRuns well
est. 16.1 tok/s12.9–19.3calibrated estimate ±20%
16k
20.8 GBRuns well
est. 9.9 tok/s8.0–11.9calibrated estimate ±20%
20.1 GBRuns well
est. 12.5 tok/s10.0–15.1calibrated estimate ±20%
19.7 GBRuns well
est. 14.6 tok/s11.7–17.5calibrated estimate ±20%
32k
21.8 GBRuns slowly
est. 6.9 tok/s4.8–8.9theoretical estimate ±30%
21.1 GBRuns well
est. 9.6 tok/s7.7–11.6calibrated estimate ±20%
20.3 GBRuns well
est. 12.3 tok/s9.8–14.8calibrated estimate ±20%
64k
25.0 GBRuns slowly
est. 4.3 tok/s3.0–5.5theoretical estimate ±30%
22.0 GBRuns slowly
est. 6.6 tok/s4.6–8.6theoretical estimate ±30%
21.5 GBRuns well
est. 9.4 tok/s7.5–11.2calibrated estimate ±20%
128k
33.1 GBdoes not fit
25.5 GBRuns slowly
est. 4.0 tok/s2.8–5.3theoretical estimate ±30%
22.2 GBRuns slowly
est. 6.3 tok/s4.4–8.2theoretical estimate ±30%
256k
47.2 GBdoes not fit
35.2 GBdoes not fit
25.9 GBRuns slowly
est. 3.8 tok/s2.7–5.0theoretical estimate ±30%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At Q4_K_M the verdict stays Runs well up to 16k context with an f16 KV cache, and up to 32k with q8_0.

Estimated speed

Generation
est. 12.8 tok/s (10.2–15.4, calibrated estimate ±20%)
Prompt processing
about 383 tok/s (230–536, rough estimate ±40%)

Measured on this exact combination

No public measurement for this exact combination yet; the numbers above are estimates.

If this is not enough

How to run it

Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.

File: Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf (18.56 GB) from unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF.

llama.cpp
llama-server -hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M -c 8192 -ngl all -fa on
  • -ngl all loads every layer on the GPU.
  • -fa on enables flash attention (needed for KV cache quantization).
  • --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.43 GB at 8k context.
Ollama

Terminal:

OLLAMA_CONTEXT_LENGTH=8192 ollama serve
ollama run hf.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M
  • Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
  • /set parameter num_ctx 8192

Related pages

Frequently asked questions

Can I run Qwen3 30B-A3B (2507) on the Apple M4?

Q4_K_M fits in unified memory (19.9 of 22.4 GB usable by the GPU): Runs well, est. 12.8 tok/s (10.2–15.4, calibrated estimate ±20%). Q4_K_M needs 19.9 GB at 8k context; the GPU can use 22.4 GB of the unified memory.

How much context can Qwen3 30B-A3B (2507) use on the Apple M4?

At Q4_K_M the verdict stays Runs well up to 16k context with an f16 KV cache (1.6 GB of KV) and up to 32k with a q8_0 KV cache (1.7 GB). The model supports up to 256k.

Which quant should I use, and how fast is it?

Q4_K_M (18.56 GB) is the recommended balance: Runs well, est. 12.8 tok/s (10.2–15.4, calibrated estimate ±20%). It is also the largest tracked file that gets this verdict.

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.