Which local LLMs can the Apple M5 run?

Unified memory — by default the GPU can use about 70% of it

Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Specs

Memory options
16 GB · 24 GB · 32 GB
Memory bandwidth
153.6 GB/s
FP16 compute
10.0 TFLOPS
Launch year
2025

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • 16 GBRuns great2 modelsRuns well8 modelsRuns slowly3 modelsWon't run16 models3 more models run at a lower quant.
  • 24 GBRuns great2 modelsRuns well10 modelsRuns slowly7 modelsWon't run10 models3 more models run at a lower quant.
  • 32 GBRuns great2 modelsRuns well12 modelsRuns slowly9 modelsWon't run6 models1 more model runs at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the Apple M5

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 32 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs great
est. 70.2 tok/s56.1–84.2calibrated estimate ±20%
Q4_K_M1.9 / 22.4 GBUnified memoryup to 32k—
HyperCLOVA X SEED 1.5B
Runs great
est. 50.8 tok/s40.6–60.9calibrated estimate ±20%
Q4_K_M2.4 / 22.4 GBUnified memoryup to 16k—
gpt-oss-20b
Runs well
est. 20.1 tok/s16.1–24.2calibrated estimate ±20%
MXFP412.9 / 22.4 GBUnified memoryup to 64k
  • reasoning model
Qwen3 30B-A3B (2507)
Runs well
est. 16.4 tok/s13.1–19.7calibrated estimate ±20%
Q4_K_M19.9 / 22.4 GBUnified memoryup to 16k
  • Q2_K (heavy quality loss): Runs great
Gemma 4 26B-A4B
Runs well
est. 15.9 tok/s12.8–19.1calibrated estimate ±20%
Q4_K_M17.9 / 22.4 GBUnified memoryup to 64k
  • UD-Q2_K_XL (heavy quality loss): Runs great
DeepSeek R1 Distill Llama 8B
Runs well
est. 15.4 tok/s12.3–18.5calibrated estimate ±20%
Q4_K_M6.6 / 22.4 GBUnified memoryup to 16k
  • reasoning model
Kanana 1.5 8B
Runs well
est. 15.4 tok/s12.3–18.5calibrated estimate ±20%
Q4_K_M6.6 / 22.4 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs great
Llama 3.1 8B
Runs well
est. 15.4 tok/s12.3–18.5calibrated estimate ±20%
Q4_K_M6.6 / 22.4 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs great
Kanana 1.5 15.7B-A3B
Runs well
est. 15.0 tok/s12.0–18.0calibrated estimate ±20%
Q4_K_M12.1 / 22.4 GBUnified memoryup to 16k—
Qwen3.5 9B
Runs well
est. 15.0 tok/s12.0–18.0calibrated estimate ±20%
Q4_K_M6.7 / 22.4 GBUnified memoryup to 128k—
Qwen3 8B
Runs well
est. 14.8 tok/s11.8–17.7calibrated estimate ±20%
Q4_K_M6.8 / 22.4 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs great
Gemma 4 12B
Runs well
est. 12.0 tok/s9.6–14.4calibrated estimate ±20%
Q4_K_M8.2 / 22.4 GBUnified memoryup to 64k—
Gemma 3 12B
Runs well
est. 11.8 tok/s9.4–14.1calibrated estimate ±20%
Q4_K_M8.4 / 22.4 GBUnified memoryup to 32k—
Qwen3 14B
Runs well
est. 8.9 tok/s7.1–10.7calibrated estimate ±20%
Q4_K_M10.9 / 22.4 GBUnified memoryup to 8k—
Qwen3.5 35B-A3B
Runs slowly
est. 22.4 tok/s15.7–29.1theoretical estimate ±30%
Q4_K_M22.8 GB RAMCPU onlyup to 256k
  • CPU inference — capped at Runs slowly
  • Try IQ4_XS: Runs great
HyperCLOVA X SEED Think 14B
Runs slowly
est. 9.0 tok/s6.3–11.8theoretical estimate ±30%
Q4_K_M10.8 / 22.4 GBUnified memoryup to 64k
  • reasoning model
Phi-4
Runs slowly
est. 8.6 tok/s6.9–10.3calibrated estimate ±20%
Q4_K_M11.3 / 22.4 GBUnified memoryup to 16k
  • reasoning model
Mistral Small 3.2 24B
Runs slowly
est. 5.9 tok/s4.7–7.1calibrated estimate ±20%
Q4_K_M16.2 / 22.4 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs well
Gemma 3 27B
Runs slowly
est. 5.4 tok/s4.3–6.4calibrated estimate ±20%
Q4_K_M17.8 / 22.4 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs well
Qwen3.5 27B
Runs slowly
est. 5.2 tok/s4.2–6.3calibrated estimate ±20%
Q4_K_M18.2 / 22.4 GBUnified memoryup to 32k—
EXAONE 4.0 32B
Runs slowly
est. 4.6 tok/s3.7–5.6calibrated estimate ±20%
Q4_K_M20.5 / 22.4 GBUnified memoryup to 32k
  • reasoning model
EXAONE 4.5 33B
Runs slowly
est. 4.5 tok/s3.6–5.4calibrated estimate ±20%
Q4_K_M21.2 / 22.4 GBUnified memoryup to 16k
  • reasoning model
Qwen3 32B
Runs slowly
est. 2.4 tok/s1.7–3.2theoretical estimate ±30%
Q4_K_M21.9 GB RAMCPU onlyup to 16k—
DeepSeek R1 Distill Qwen 32B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 6.4 tok/s5.1–7.6calibrated estimate ±20%
Q2_K15.0 / 22.4 GBUnified memoryup to 32k
  • reasoning model
gpt-oss-120b
Won't run
—MXFP4needs 64.3 GB——
  • Needs about 64.3 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
  • reasoning model
Llama 3.3 70B
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
Qwen3.5 122B-A10B
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
Solar Open 100B
Won't run
—Q4_K_Mneeds 64.4 GB——
  • Needs about 64.4 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
Solar Open 2 250B
Won't run
—IQ4_XSneeds 137.2 GB——
  • Needs about 137.2 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Verdict by memory size

Q4_K_M verdict for each memory configuration — open the model page for speeds
Model16 GB24 GB32 GB
EXAONE 4.0 1.2BRuns greatRuns greatRuns great
HyperCLOVA X SEED 1.5BRuns greatRuns greatRuns great
gpt-oss-20bWon't runRuns wellRuns well
Qwen3 30B-A3B (2507)Won't runRuns slowlyRuns well
Gemma 4 26B-A4BWon't runRuns slowlyRuns well
DeepSeek R1 Distill Llama 8BRuns wellRuns wellRuns well
Kanana 1.5 8BRuns wellRuns wellRuns well
Llama 3.1 8BRuns wellRuns wellRuns well
Kanana 1.5 15.7B-A3BRuns slowlyRuns wellRuns well
Qwen3.5 9BRuns wellRuns wellRuns well
Qwen3 8BRuns wellRuns wellRuns well
Gemma 4 12BRuns wellRuns wellRuns well
Gemma 3 12BRuns wellRuns wellRuns well
Qwen3 14BRuns wellRuns wellRuns well
Qwen3.5 35B-A3BWon't runWon't runRuns slowly
HyperCLOVA X SEED Think 14BRuns slowlyRuns slowlyRuns slowly
Phi-4Runs slowlyRuns slowlyRuns slowly
Mistral Small 3.2 24BWon't runRuns slowlyRuns slowly
Gemma 3 27BWon't runRuns slowlyRuns slowly
Qwen3.5 27BWon't runRuns slowlyRuns slowly
EXAONE 4.0 32BWon't runWon't runRuns slowly
EXAONE 4.5 33BWon't runWon't runRuns slowly
Qwen3 32BWon't runWon't runRuns slowly
DeepSeek R1 Distill Qwen 32BWon't runWon't runWon't run
gpt-oss-120bWon't runWon't runWon't run
Llama 3.3 70BWon't runWon't runWon't run
Qwen3.5 122B-A10BWon't runWon't runWon't run
Solar Open 100BWon't runWon't runWon't run
Solar Open 2 250BWon't runWon't runWon't run

Measured results on the Apple M5

No public measurements for this device yet.

Frequently asked questions

What is the largest model that runs entirely on the Apple M5?

Qwen3.5 35B-A3B at IQ4_XS (a 18.17 GB file) fits entirely in 32 GB of unified memory with 8k context, at est. 27.4 tok/s (21.9–32.8, calibrated estimate ±20%).

How many local LLMs run well on the Apple M5?

At Q4_K_M with 8k context, with 32 GB of memory, out of 29 tracked models: 2 run great, 12 run well, 9 run slowly and 6 won't run.

Can the Apple M5 run a 70B model like Llama 3.3 70B?

No. Llama 3.3 70B at Q4_K_M needs about 45.8 GB, while this setup offers 22.4 GB of GPU memory and 28.0 GB of free system RAM.

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.