Which local LLMs can the Apple M5 Pro run?
Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Specs
- Memory options
- 24 GB · 48 GB · 64 GB
- Memory bandwidth
- 307.0 GB/s
- FP16 compute
- 19.0 TFLOPS
- Launch year
- 2026
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- 24 GBRuns great11 modelsRuns well4 modelsRuns slowly5 modelsWon't run9 models4 more models run at a lower quant.
- 48 GBRuns great14 modelsRuns well7 modelsRuns slowly3 modelsWon't run5 models3 more models run at a lower quant.
- 64 GBRuns great14 modelsRuns well7 modelsRuns slowly4 modelsWon't run4 models2 more models run at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the Apple M5 Pro
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 140.3 tok/s112.2–168.3calibrated estimate ±20% | Q4_K_M | 1.9 / 44.8 GB | Unified memory | up to 64k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 101.5 tok/s81.2–121.8calibrated estimate ±20% | Q4_K_M | 2.4 / 44.8 GB | Unified memory | up to 16k | — |
| Qwen3.5 35B-A3B | Runs great | est. 44.8 tok/s35.8–53.7calibrated estimate ±20% | Q4_K_M | 23.4 / 44.8 GB | Unified memory | up to 128k | — |
| gpt-oss-20b | Runs great | est. 40.3 tok/s32.2–48.3calibrated estimate ±20% | MXFP4 | 12.9 / 44.8 GB | Unified memory | up to 32k |
|
| Qwen3 30B-A3B (2507) | Runs great | est. 32.7 tok/s26.2–39.3calibrated estimate ±20% | Q4_K_M | 19.9 / 44.8 GB | Unified memory | up to 16k | — |
| Gemma 4 26B-A4B | Runs great | est. 31.9 tok/s25.5–38.2calibrated estimate ±20% | Q4_K_M | 17.9 / 44.8 GB | Unified memory | up to 32k | — |
| DeepSeek R1 Distill Llama 8B | Runs great | est. 30.7 tok/s24.6–36.9calibrated estimate ±20% | Q4_K_M | 6.6 / 44.8 GB | Unified memory | up to 8k |
|
| Kanana 1.5 8B | Runs great | est. 30.7 tok/s24.6–36.9calibrated estimate ±20% | Q4_K_M | 6.6 / 44.8 GB | Unified memory | up to 16k | — |
| Llama 3.1 8B | Runs great | est. 30.7 tok/s24.6–36.9calibrated estimate ±20% | Q4_K_M | 6.6 / 44.8 GB | Unified memory | up to 16k | — |
| Kanana 1.5 15.7B-A3B | Runs great | est. 30.0 tok/s24.0–36.0calibrated estimate ±20% | Q4_K_M | 12.1 / 44.8 GB | Unified memory | up to 16k | — |
| Qwen3.5 9B | Runs great | est. 30.0 tok/s24.0–36.0calibrated estimate ±20% | Q4_K_M | 6.7 / 44.8 GB | Unified memory | up to 64k | — |
| Qwen3 8B | Runs great | est. 29.5 tok/s23.6–35.4calibrated estimate ±20% | Q4_K_M | 6.8 / 44.8 GB | Unified memory | up to 16k | — |
| Gemma 4 12B | Runs great | est. 24.1 tok/s19.2–28.9calibrated estimate ±20% | Q4_K_M | 8.2 / 44.8 GB | Unified memory | up to 16k | — |
| Gemma 3 12B | Runs great | est. 23.5 tok/s18.8–28.2calibrated estimate ±20% | Q4_K_M | 8.4 / 44.8 GB | Unified memory | up to 16k | — |
| HyperCLOVA X SEED Think 14B | Runs well | est. 18.1 tok/s12.6–23.5theoretical estimate ±30% | Q4_K_M | 10.8 / 44.8 GB | Unified memory | up to 32k |
|
| Qwen3 14B | Runs well | est. 17.8 tok/s14.2–21.4calibrated estimate ±20% | Q4_K_M | 10.9 / 44.8 GB | Unified memory | up to 32k |
|
| Phi-4 | Runs well | est. 17.2 tok/s13.7–20.6calibrated estimate ±20% | Q4_K_M | 11.3 / 44.8 GB | Unified memory | up to 16k |
|
| Mistral Small 3.2 24B | Runs well | est. 11.8 tok/s9.4–14.1calibrated estimate ±20% | Q4_K_M | 16.2 / 44.8 GB | Unified memory | up to 32k | — |
| Gemma 3 27B | Runs well | est. 10.7 tok/s8.6–12.8calibrated estimate ±20% | Q4_K_M | 17.8 / 44.8 GB | Unified memory | up to 64k | — |
| Qwen3.5 27B | Runs well | est. 10.4 tok/s8.4–12.5calibrated estimate ±20% | Q4_K_M | 18.2 / 44.8 GB | Unified memory | up to 64k | — |
| Qwen3 32B | Runs well | est. 8.4 tok/s6.7–10.1calibrated estimate ±20% | Q4_K_M | 22.5 / 44.8 GB | Unified memory | up to 8k | — |
| EXAONE 4.0 32B | Runs slowly | est. 9.3 tok/s7.4–11.1calibrated estimate ±20% | Q4_K_M | 20.5 / 44.8 GB | Unified memory | up to 128k |
|
| EXAONE 4.5 33B | Runs slowly | est. 8.9 tok/s7.2–10.7calibrated estimate ±20% | Q4_K_M | 21.2 / 44.8 GB | Unified memory | up to 256k |
|
| DeepSeek R1 Distill Qwen 32B | Runs slowly | est. 8.4 tok/s6.7–10.0calibrated estimate ±20% | Q4_K_M | 22.6 / 44.8 GB | Unified memory | up to 64k |
|
| Llama 3.3 70B | Runs slowly | est. 2.4 tok/s1.7–3.1theoretical estimate ±30% | Q4_K_M | 45.2 GB RAM | CPU only | up to 32k | — |
| Qwen3.5 122B-A10B | Won't runTry UD-Q2_K_XL (heavy quality loss): Runs great | est. 25.4 tok/s20.3–30.5calibrated estimate ±20% | UD-Q2_K_XL | 43.6 / 44.8 GB | Unified memory | up to 32k | — |
| Solar Open 100B | Won't runTry IQ4_XS: Runs slowly | est. 11.4 tok/s8.0–14.8theoretical estimate ±30% | IQ4_XS | 57.1 GB RAM | CPU only | up to 16k |
|
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Verdict by memory size
| Model | 24 GB | 48 GB | 64 GB |
|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | Runs great | Runs great |
| HyperCLOVA X SEED 1.5B | Runs great | Runs great | Runs great |
| Qwen3.5 35B-A3B | Won't run | Runs great | Runs great |
| gpt-oss-20b | Runs great | Runs great | Runs great |
| Qwen3 30B-A3B (2507) | Runs slowly | Runs great | Runs great |
| Gemma 4 26B-A4B | Runs slowly | Runs great | Runs great |
| DeepSeek R1 Distill Llama 8B | Runs great | Runs great | Runs great |
| Kanana 1.5 8B | Runs great | Runs great | Runs great |
| Llama 3.1 8B | Runs great | Runs great | Runs great |
| Kanana 1.5 15.7B-A3B | Runs great | Runs great | Runs great |
| Qwen3.5 9B | Runs great | Runs great | Runs great |
| Qwen3 8B | Runs great | Runs great | Runs great |
| Gemma 4 12B | Runs great | Runs great | Runs great |
| Gemma 3 12B | Runs great | Runs great | Runs great |
| HyperCLOVA X SEED Think 14B | Runs well | Runs well | Runs well |
| Qwen3 14B | Runs well | Runs well | Runs well |
| Phi-4 | Runs well | Runs well | Runs well |
| Mistral Small 3.2 24B | Runs well | Runs well | Runs well |
| Gemma 3 27B | Runs slowly | Runs well | Runs well |
| Qwen3.5 27B | Runs slowly | Runs well | Runs well |
| Qwen3 32B | Won't run | Runs well | Runs well |
| EXAONE 4.0 32B | Runs slowly | Runs slowly | Runs slowly |
| EXAONE 4.5 33B | Won't run | Runs slowly | Runs slowly |
| DeepSeek R1 Distill Qwen 32B | Won't run | Runs slowly | Runs slowly |
| Llama 3.3 70B | Won't run | Won't run | Runs slowly |
| Qwen3.5 122B-A10B | Won't run | Won't run | Won't run |
| Solar Open 100B | Won't run | Won't run | Won't run |
| gpt-oss-120b | Won't run | Won't run | Won't run |
| Solar Open 2 250B | Won't run | Won't run | Won't run |
Measured results on the Apple M5 Pro
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the Apple M5 Pro?
Llama 3.3 70B at IQ4_XS (a 37.90 GB file) fits entirely in 64 GB of unified memory with 8k context, at est. 4.5 tok/s (3.6–5.4, calibrated estimate ±20%).
How many local LLMs run well on the Apple M5 Pro?
At Q4_K_M with 8k context, with 64 GB of memory, out of 29 tracked models: 14 run great, 7 run well, 4 run slowly and 4 won't run.
Can the Apple M5 Pro run a 70B model like Llama 3.3 70B?
Runs slowly — CPU only, Q4_K_M: est. 2.4 tok/s (1.7–3.1, theoretical estimate ±30%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.