Can I run gpt-oss-120b on the Apple M4?

Won't runMXFP4 needs 64.3 GB, 41.9 GB more than the 22.4 GB of usable memory, and no tracked quant runs at a usable speed with 32 GB of RAM.

Unified memory — by default the GPU can use about 70% of itOpen license · Apache-2.0llama.cpp team · ggml-orgreasoning model

32 GB unified memory · 120.0 GB/s · FP16 8.5 TFLOPS. Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Verdict by memory size

The page uses the 32 GB configuration; here is every size.

  • 16 GBWon't run
  • 24 GBWon't run
  • 32 GBWon't run

Every quant of gpt-oss-120b on the Apple M4

1 tracked GGUF file, smallest first, at 8k context
QuantFile sizeVerdictSpeedMemoryRuns asNotes
MXFP4baseline63.39 GBWon't run—Needs about 64.3 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)——

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Where the memory goes at MXFP4

MXFP4 does not fit — it needs 64.3 GB against 22.4 GB of usable memory, so there is nothing to break down.

Context length vs. KV cache

KV cache

Memory needed and verdict at MXFP4 for each context length and KV cache type, with estimated tokens per second.
ContextKV f16KV q8_0KV q4_0
4k
64.1 GBdoes not fit
64.0 GBdoes not fit
64.0 GBdoes not fit
8k
64.3 GBdoes not fit
64.1 GBdoes not fit
64.1 GBdoes not fit
16k
64.6 GBdoes not fit
64.4 GBdoes not fit
64.2 GBdoes not fit
32k
65.4 GBdoes not fit
64.8 GBdoes not fit
64.5 GBdoes not fit
64k
66.9 GBdoes not fit
65.8 GBdoes not fit
65.2 GBdoes not fit
128k
69.9 GBdoes not fit
67.7 GBdoes not fit
66.5 GBdoes not fit

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At MXFP4 no context length fits: even 4k needs 64.1 GB against 22.4 GB.

Estimated speed

Generation
—
Prompt processing
Prompt-processing estimates only apply when the whole model runs on the GPU.

Measured on this exact combination

No public measurement for this exact combination yet; the numbers above are estimates.

If this is not enough

How to run it

No run command: this file does not fit in the 22.4 GB of unified memory the GPU can use.

Related pages

Frequently asked questions

Can I run gpt-oss-120b on the Apple M4?

MXFP4 needs 64.3 GB, 41.9 GB more than the 22.4 GB of usable memory, and no tracked quant runs at a usable speed with 32 GB of RAM. MXFP4 needs 64.3 GB at 8k context; the GPU can use 22.4 GB of the unified memory.

How much context can gpt-oss-120b use on the Apple M4?

No context length fits: even at 4k, MXFP4 needs 64.1 GB against 22.4 GB of usable memory.

Which quant should I use, and how fast is it?

No tracked quant runs here: the smallest file (MXFP4, 63.39 GB) needs 64.3 GB at 8k context, and the GPU can use 22.4 GB of the unified memory.

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.