Which local LLMs can the Apple M5 Max (32-core GPU) run?

Unified memory — by default the GPU can use about 70% of it

Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Specs

Memory options
36 GB · 48 GB
Memory bandwidth
460.0 GB/s
FP16 compute
30.0 TFLOPS
Launch year
2026

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • 36 GBRuns great15 modelsRuns well9 modelsRuns slowly0 modelsWon't run5 models1 more model runs at a lower quant.
  • 48 GBRuns great15 modelsRuns well9 modelsRuns slowly0 modelsWon't run5 models3 more models run at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the Apple M5 Max (32-core GPU)

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 48 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs great
est. 210.2 tok/s168.1–252.2calibrated estimate ±20%
Q4_K_M1.9 / 33.6 GBUnified memoryup to 64k—
HyperCLOVA X SEED 1.5B
Runs great
est. 152.0 tok/s121.6–182.4calibrated estimate ±20%
Q4_K_M2.4 / 33.6 GBUnified memoryup to 16k—
Qwen3.5 35B-A3B
Runs great
est. 67.1 tok/s53.7–80.5calibrated estimate ±20%
Q4_K_M23.4 / 33.6 GBUnified memoryup to 128k—
gpt-oss-20b
Runs great
est. 60.3 tok/s48.3–72.4calibrated estimate ±20%
MXFP412.9 / 33.6 GBUnified memoryup to 64k
  • reasoning model
Qwen3 30B-A3B (2507)
Runs great
est. 49.1 tok/s39.2–58.9calibrated estimate ±20%
Q4_K_M19.9 / 33.6 GBUnified memoryup to 32k—
Gemma 4 26B-A4B
Runs great
est. 47.7 tok/s38.2–57.3calibrated estimate ±20%
Q4_K_M17.9 / 33.6 GBUnified memoryup to 64k—
DeepSeek R1 Distill Llama 8B
Runs great
est. 46.0 tok/s36.8–55.3calibrated estimate ±20%
Q4_K_M6.6 / 33.6 GBUnified memoryup to 16k
  • reasoning model
Kanana 1.5 8B
Runs great
est. 46.0 tok/s36.8–55.3calibrated estimate ±20%
Q4_K_M6.6 / 33.6 GBUnified memoryup to 32k—
Llama 3.1 8B
Runs great
est. 46.0 tok/s36.8–55.3calibrated estimate ±20%
Q4_K_M6.6 / 33.6 GBUnified memoryup to 64k—
Kanana 1.5 15.7B-A3B
Runs great
est. 45.0 tok/s36.0–54.0calibrated estimate ±20%
Q4_K_M12.1 / 33.6 GBUnified memoryup to 32k—
Qwen3.5 9B
Runs great
est. 45.0 tok/s36.0–54.0calibrated estimate ±20%
Q4_K_M6.7 / 33.6 GBUnified memoryup to 128k—
Qwen3 8B
Runs great
est. 44.2 tok/s35.4–53.1calibrated estimate ±20%
Q4_K_M6.8 / 33.6 GBUnified memoryup to 32k—
Gemma 4 12B
Runs great
est. 36.0 tok/s28.8–43.3calibrated estimate ±20%
Q4_K_M8.2 / 33.6 GBUnified memoryup to 64k—
Gemma 3 12B
Runs great
est. 35.2 tok/s28.2–42.3calibrated estimate ±20%
Q4_K_M8.4 / 33.6 GBUnified memoryup to 64k—
Qwen3 14B
Runs great
est. 26.7 tok/s21.3–32.0calibrated estimate ±20%
Q4_K_M10.9 / 33.6 GBUnified memoryup to 16k—
HyperCLOVA X SEED Think 14B
Runs well
est. 27.1 tok/s19.0–35.2theoretical estimate ±30%
Q4_K_M10.8 / 33.6 GBUnified memoryup to 64k
  • reasoning model
Phi-4
Runs well
est. 25.7 tok/s20.6–30.9calibrated estimate ±20%
Q4_K_M11.3 / 33.6 GBUnified memoryup to 16k
  • reasoning model
Mistral Small 3.2 24B
Runs well
est. 17.6 tok/s14.1–21.1calibrated estimate ±20%
Q4_K_M16.2 / 33.6 GBUnified memoryup to 64k
  • Q2_K (heavy quality loss): Runs great
Gemma 3 27B
Runs well
est. 16.0 tok/s12.8–19.2calibrated estimate ±20%
Q4_K_M17.8 / 33.6 GBUnified memoryup to 128k
  • Q2_K (heavy quality loss): Runs great
Qwen3.5 27B
Runs well
est. 15.6 tok/s12.5–18.8calibrated estimate ±20%
Q4_K_M18.2 / 33.6 GBUnified memoryup to 128k
  • UD-Q2_K_XL (heavy quality loss): Runs great
EXAONE 4.0 32B
Runs well
est. 13.9 tok/s11.1–16.7calibrated estimate ±20%
Q4_K_M20.5 / 33.6 GBUnified memoryup to 32k
  • reasoning model
EXAONE 4.5 33B
Runs well
est. 13.4 tok/s10.7–16.1calibrated estimate ±20%
Q4_K_M21.2 / 33.6 GBUnified memoryup to 32k
  • reasoning model
Qwen3 32B
Runs well
est. 12.6 tok/s10.1–15.1calibrated estimate ±20%
Q4_K_M22.5 / 33.6 GBUnified memoryup to 32k—
DeepSeek R1 Distill Qwen 32B
Runs well
est. 12.5 tok/s10.0–15.1calibrated estimate ±20%
Q4_K_M22.6 / 33.6 GBUnified memoryup to 8k
  • reasoning model
Qwen3.5 122B-A10B
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 38.0 tok/s26.6–49.5theoretical estimate ±30%
UD-Q2_K_XL43.1 GB RAMCPU onlyup to 32k
  • CPU inference — capped at Runs slowly
Solar Open 100B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 22.9 tok/s16.1–29.8theoretical estimate ±30%
Q2_K39.3 GB RAMCPU onlyup to 16k
  • CPU inference — capped at Runs slowly
Llama 3.3 70B
Won't runTry IQ4_XS: Runs slowly
est. 4.0 tok/s2.8–5.1theoretical estimate ±30%
IQ4_XS40.6 GB RAMCPU onlyup to 16k
  • Q2_K (heavy quality loss): Runs well
gpt-oss-120b
Won't run
—MXFP4needs 64.3 GB——
  • Needs about 64.3 GB; 44.0 GB of memory is free (33.6 GB usable by the GPU)
  • reasoning model
Solar Open 2 250B
Won't run
—IQ4_XSneeds 137.2 GB——
  • Needs about 137.2 GB; 44.0 GB of memory is free (33.6 GB usable by the GPU)

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Verdict by memory size

Q4_K_M verdict for each memory configuration — open the model page for speeds
Model36 GB48 GB
EXAONE 4.0 1.2BRuns greatRuns great
HyperCLOVA X SEED 1.5BRuns greatRuns great
Qwen3.5 35B-A3BRuns greatRuns great
gpt-oss-20bRuns greatRuns great
Qwen3 30B-A3B (2507)Runs greatRuns great
Gemma 4 26B-A4BRuns greatRuns great
DeepSeek R1 Distill Llama 8BRuns greatRuns great
Kanana 1.5 8BRuns greatRuns great
Llama 3.1 8BRuns greatRuns great
Kanana 1.5 15.7B-A3BRuns greatRuns great
Qwen3.5 9BRuns greatRuns great
Qwen3 8BRuns greatRuns great
Gemma 4 12BRuns greatRuns great
Gemma 3 12BRuns greatRuns great
Qwen3 14BRuns greatRuns great
HyperCLOVA X SEED Think 14BRuns wellRuns well
Phi-4Runs wellRuns well
Mistral Small 3.2 24BRuns wellRuns well
Gemma 3 27BRuns wellRuns well
Qwen3.5 27BRuns wellRuns well
EXAONE 4.0 32BRuns wellRuns well
EXAONE 4.5 33BRuns wellRuns well
Qwen3 32BRuns wellRuns well
DeepSeek R1 Distill Qwen 32BRuns wellRuns well
Qwen3.5 122B-A10BWon't runWon't run
Solar Open 100BWon't runWon't run
Llama 3.3 70BWon't runWon't run
gpt-oss-120bWon't runWon't run
Solar Open 2 250BWon't runWon't run

Measured results on the Apple M5 Max (32-core GPU)

No public measurements for this device yet.

Frequently asked questions

What is the largest model that runs entirely on the Apple M5 Max (32-core GPU)?

Qwen3.5 35B-A3B at Q4_K_M (a 22.63 GB file) fits entirely in 48 GB of unified memory with 8k context, at est. 67.1 tok/s (53.7–80.5, calibrated estimate ±20%).

How many local LLMs run well on the Apple M5 Max (32-core GPU)?

At Q4_K_M with 8k context, with 48 GB of memory, out of 29 tracked models: 15 run great, 9 run well, 0 run slowly and 5 won't run.

Can the Apple M5 Max (32-core GPU) run a 70B model like Llama 3.3 70B?

Runs slowly — CPU only, IQ4_XS: est. 4.0 tok/s (2.8–5.1, theoretical estimate ±30%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.