Can I run gpt-oss-20b on the Apple M4 Pro?
Runs greatMXFP4 fits in unified memory (12.9 of 16.8 GB usable by the GPU): Runs great, est. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%).
24 GB unified memory · 273.0 GB/s · FP16 17.0 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Verdict by memory size
The page uses the 24 GB configuration; here is every size.
- 24 GBRuns greatest. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%)
- 48 GBRuns greatest. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%)
- 64 GBRuns greatest. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%)
Every quant of gpt-oss-20b on the Apple M4 Pro
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| MXFP4baseline | 12.11 GB | Runs great | est. 35.8 tok/s28.6–43.0calibrated estimate ±20% | 12.9 / 16.8 GB | Unified memory | — |
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Where the memory goes at MXFP4
- Weights
- 12.1 GB
- KV cache
- 0.2 GB
- Compute buffer
- 0.6 GB
- Free
- 3.9 GB
- Not usable by the GPU
- 7.2 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 12.7 GBRuns great est. 37.5 tok/s30.0–44.9calibrated estimate ±20% | 12.7 GBRuns great est. 38.3 tok/s30.6–45.9calibrated estimate ±20% | 12.7 GBRuns great est. 38.7 tok/s31.0–46.5calibrated estimate ±20% |
| 8k | 12.9 GBRuns great est. 35.8 tok/s28.6–43.0calibrated estimate ±20% | 12.8 GBRuns great est. 37.3 tok/s29.9–44.8calibrated estimate ±20% | 12.7 GBRuns great est. 38.2 tok/s30.6–45.9calibrated estimate ±20% |
| 16k | 13.2 GBRuns great est. 32.9 tok/s26.3–39.5calibrated estimate ±20% | 13.0 GBRuns great est. 35.6 tok/s28.5–42.7calibrated estimate ±20% | 12.9 GBRuns great est. 37.2 tok/s29.8–44.7calibrated estimate ±20% |
| 32k | 13.7 GBRuns well est. 28.3 tok/s22.7–34.0calibrated estimate ±20% | 13.3 GBRuns great est. 32.5 tok/s26.0–39.1calibrated estimate ±20% | 13.1 GBRuns great est. 35.4 tok/s28.3–42.4calibrated estimate ±20% |
| 64k | 14.8 GBRuns well est. 22.2 tok/s17.7–26.6calibrated estimate ±20% | 14.1 GBRuns well est. 27.8 tok/s22.2–33.3calibrated estimate ±20% | 13.7 GBRuns great est. 32.2 tok/s25.7–38.6calibrated estimate ±20% |
| 128k | 15.3 GBRuns slowly est. 15.4 tok/s10.8–20.1theoretical estimate ±30% | 15.5 GBRuns well est. 21.5 tok/s17.2–25.8calibrated estimate ±20% | 14.7 GBRuns well est. 27.3 tok/s21.8–32.7calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At MXFP4 the verdict stays Runs great up to 16k context with an f16 KV cache, and up to 32k with q8_0.
Estimated speed
- Generation
- est. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%)
- Prompt processing
- about 765 tok/s (459–1,071, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A larger model that also runs well on the Apple M4 Pro
Runs wellMistral Small 3.2 24B: Runs well, est. 10.5 tok/s (8.4–12.5, calibrated estimate ±20%)
How to run it
Commands for the MXFP4 file. Both tools download from Hugging Face on first run.
File: gpt-oss-20b-MXFP4.gguf (12.11 GB) from ggml-org/gpt-oss-20b-GGUF.
llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.11 GB at 8k context.
Terminal:
OLLAMA_CONTEXT_LENGTH=8192 ollama serve
ollama run hf.co/ggml-org/gpt-oss-20b-GGUF:MXFP4- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: Running local LLMs on a Mac (M4, M5): how much unified memory?
Other models on the Apple M4 Pro
gpt-oss-20b on other hardware
Frequently asked questions
Can I run gpt-oss-20b on the Apple M4 Pro?
MXFP4 fits in unified memory (12.9 of 16.8 GB usable by the GPU): Runs great, est. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%). MXFP4 needs 12.9 GB at 8k context; the GPU can use 16.8 GB of the unified memory.
How much context can gpt-oss-20b use on the Apple M4 Pro?
At MXFP4 the verdict stays Runs great up to 16k context with an f16 KV cache (0.4 GB of KV) and up to 32k with a q8_0 KV cache (0.4 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
MXFP4 (12.11 GB) is the only tracked file: Runs great, est. 35.8 tok/s (28.6–43.0, calibrated estimate ±20%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.