Can I run gpt-oss-120b on the Apple M4 Pro?
Won't runMXFP4 needs 64.3 GB, 19.5 GB more than the 44.8 GB of usable memory, and no tracked quant runs at a usable speed with 64 GB of RAM.
64 GB unified memory · 273.0 GB/s · FP16 17.0 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Verdict by memory size
The page uses the 64 GB configuration; here is every size.
- 24 GBWon't run
- 48 GBWon't run
- 64 GBWon't run
Every quant of gpt-oss-120b on the Apple M4 Pro
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| MXFP4baseline | 63.39 GB | Won't run | — | Needs about 64.3 GB; 60.0 GB of memory is free (44.8 GB usable by the GPU) | — | — |
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Where the memory goes at MXFP4
MXFP4 does not fit — it needs 64.3 GB against 44.8 GB of usable memory, so there is nothing to break down.
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 64.1 GBdoes not fit | 64.0 GBdoes not fit | 64.0 GBdoes not fit |
| 8k | 64.3 GBdoes not fit | 64.1 GBdoes not fit | 64.1 GBdoes not fit |
| 16k | 64.6 GBdoes not fit | 64.4 GBdoes not fit | 64.2 GBdoes not fit |
| 32k | 65.4 GBdoes not fit | 64.8 GBdoes not fit | 64.5 GBdoes not fit |
| 64k | 66.9 GBdoes not fit | 65.8 GBdoes not fit | 65.2 GBdoes not fit |
| 128k | 69.9 GBdoes not fit | 67.7 GBdoes not fit | 66.5 GBdoes not fit |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At MXFP4 no context length fits: even 4k needs 64.1 GB against 44.8 GB.
Estimated speed
- Generation
- —
- Prompt processing
- Prompt-processing estimates only apply when the whole model runs on the GPU.
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A smaller model that runs well on the Apple M4 Pro
Runs greatQwen3.5 35B-A3B: Runs great, est. 39.8 tok/s (31.9–47.8, calibrated estimate ±20%)
gpt-oss-120b on the Apple M4 Max (40-core GPU)
Runs greatMXFP4: Runs great, est. 53.4 tok/s (42.7–64.0, calibrated estimate ±20%)
How to run it
No run command: this file does not fit in the 44.8 GB of unified memory the GPU can use.
Related pages
Guide: Running local LLMs on a Mac (M4, M5): how much unified memory?
Other models on the Apple M4 Pro
gpt-oss-120b on other hardware
Frequently asked questions
Can I run gpt-oss-120b on the Apple M4 Pro?
MXFP4 needs 64.3 GB, 19.5 GB more than the 44.8 GB of usable memory, and no tracked quant runs at a usable speed with 64 GB of RAM. MXFP4 needs 64.3 GB at 8k context; the GPU can use 44.8 GB of the unified memory.
How much context can gpt-oss-120b use on the Apple M4 Pro?
No context length fits: even at 4k, MXFP4 needs 64.1 GB against 44.8 GB of usable memory.
Which quant should I use, and how fast is it?
No tracked quant runs here: the smallest file (MXFP4, 63.39 GB) needs 64.3 GB at 8k context, and the GPU can use 44.8 GB of the unified memory.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.