Which local LLMs can the Apple M4 Max (32-core GPU) run?

Unified memory — by default the GPU can use about 70% of it

Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Specs

Memory options
36 GB · 48 GB
Memory bandwidth
410.0 GB/s
FP16 compute
27.2 TFLOPS
Launch year
2024

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • 36 GBRuns great15 modelsRuns well7 modelsRuns slowly2 modelsWon't run5 models1 more model runs at a lower quant.
  • 48 GBRuns great15 modelsRuns well7 modelsRuns slowly2 modelsWon't run5 models3 more models run at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the Apple M4 Max (32-core GPU)

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 48 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs great
est. 187.3 tok/s149.8–224.8calibrated estimate ±20%
Q4_K_M1.9 / 33.6 GBUnified memoryup to 64k—
HyperCLOVA X SEED 1.5B
Runs great
est. 135.5 tok/s108.4–162.6calibrated estimate ±20%
Q4_K_M2.4 / 33.6 GBUnified memoryup to 16k—
Qwen3.5 35B-A3B
Runs great
est. 59.8 tok/s47.9–71.8calibrated estimate ±20%
Q4_K_M23.4 / 33.6 GBUnified memoryup to 128k—
gpt-oss-20b
Runs great
est. 53.8 tok/s43.0–64.5calibrated estimate ±20%
MXFP412.9 / 33.6 GBUnified memoryup to 64k
  • reasoning model
Qwen3 30B-A3B (2507)
Runs great
est. 43.7 tok/s35.0–52.5calibrated estimate ±20%
Q4_K_M19.9 / 33.6 GBUnified memoryup to 32k—
Gemma 4 26B-A4B
Runs great
est. 42.6 tok/s34.0–51.1calibrated estimate ±20%
Q4_K_M17.9 / 33.6 GBUnified memoryup to 64k—
DeepSeek R1 Distill Llama 8B
Runs great
est. 41.0 tok/s32.8–49.3calibrated estimate ±20%
Q4_K_M6.6 / 33.6 GBUnified memoryup to 16k
  • reasoning model
Kanana 1.5 8B
Runs great
est. 41.0 tok/s32.8–49.3calibrated estimate ±20%
Q4_K_M6.6 / 33.6 GBUnified memoryup to 32k—
Llama 3.1 8B
Runs great
est. 41.0 tok/s32.8–49.3calibrated estimate ±20%
Q4_K_M6.6 / 33.6 GBUnified memoryup to 32k—
Kanana 1.5 15.7B-A3B
Runs great
est. 40.1 tok/s32.1–48.1calibrated estimate ±20%
Q4_K_M12.1 / 33.6 GBUnified memoryup to 16k—
Qwen3.5 9B
Runs great
est. 40.1 tok/s32.1–48.1calibrated estimate ±20%
Q4_K_M6.7 / 33.6 GBUnified memoryup to 128k—
Qwen3 8B
Runs great
est. 39.4 tok/s31.5–47.3calibrated estimate ±20%
Q4_K_M6.8 / 33.6 GBUnified memoryup to 32k—
Gemma 4 12B
Runs great
est. 32.1 tok/s25.7–38.6calibrated estimate ±20%
Q4_K_M8.2 / 33.6 GBUnified memoryup to 64k—
Gemma 3 12B
Runs great
est. 31.4 tok/s25.1–37.7calibrated estimate ±20%
Q4_K_M8.4 / 33.6 GBUnified memoryup to 64k—
Qwen3 14B
Runs great
est. 23.8 tok/s19.0–28.5calibrated estimate ±20%
Q4_K_M10.9 / 33.6 GBUnified memoryup to 16k—
HyperCLOVA X SEED Think 14B
Runs well
est. 24.1 tok/s16.9–31.4theoretical estimate ±30%
Q4_K_M10.8 / 33.6 GBUnified memoryup to 64k
  • reasoning model
Phi-4
Runs well
est. 22.9 tok/s18.3–27.5calibrated estimate ±20%
Q4_K_M11.3 / 33.6 GBUnified memoryup to 16k
  • reasoning model
Mistral Small 3.2 24B
Runs well
est. 15.7 tok/s12.6–18.8calibrated estimate ±20%
Q4_K_M16.2 / 33.6 GBUnified memoryup to 64k
  • Q2_K (heavy quality loss): Runs great
Gemma 3 27B
Runs well
est. 14.3 tok/s11.4–17.1calibrated estimate ±20%
Q4_K_M17.8 / 33.6 GBUnified memoryup to 128k
  • Q2_K (heavy quality loss): Runs great
Qwen3.5 27B
Runs well
est. 13.9 tok/s11.2–16.7calibrated estimate ±20%
Q4_K_M18.2 / 33.6 GBUnified memoryup to 128k—
EXAONE 4.0 32B
Runs well
est. 12.4 tok/s9.9–14.9calibrated estimate ±20%
Q4_K_M20.5 / 33.6 GBUnified memoryup to 16k
  • reasoning model
Qwen3 32B
Runs well
est. 11.2 tok/s9.0–13.5calibrated estimate ±20%
Q4_K_M22.5 / 33.6 GBUnified memoryup to 32k—
EXAONE 4.5 33B
Runs slowly
est. 11.9 tok/s9.6–14.3calibrated estimate ±20%
Q4_K_M21.2 / 33.6 GBUnified memoryup to 128k
  • Try IQ4_XS: Runs well
  • reasoning model
DeepSeek R1 Distill Qwen 32B
Runs slowly
est. 11.2 tok/s8.9–13.4calibrated estimate ±20%
Q4_K_M22.6 / 33.6 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs well
  • reasoning model
Qwen3.5 122B-A10B
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 33.9 tok/s23.7–44.1theoretical estimate ±30%
UD-Q2_K_XL43.1 GB RAMCPU onlyup to 32k
  • CPU inference — capped at Runs slowly
Solar Open 100B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 20.4 tok/s14.3–26.6theoretical estimate ±30%
Q2_K39.3 GB RAMCPU onlyup to 16k
  • CPU inference — capped at Runs slowly
Llama 3.3 70B
Won't runTry IQ4_XS: Runs slowly
est. 3.5 tok/s2.5–4.6theoretical estimate ±30%
IQ4_XS40.6 GB RAMCPU onlyup to 16k
  • Q2_K (heavy quality loss): Runs well
gpt-oss-120b
Won't run
—MXFP4needs 64.3 GB——
  • Needs about 64.3 GB; 44.0 GB of memory is free (33.6 GB usable by the GPU)
  • reasoning model
Solar Open 2 250B
Won't run
—IQ4_XSneeds 137.2 GB——
  • Needs about 137.2 GB; 44.0 GB of memory is free (33.6 GB usable by the GPU)

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Verdict by memory size

Q4_K_M verdict for each memory configuration — open the model page for speeds
Model36 GB48 GB
EXAONE 4.0 1.2BRuns greatRuns great
HyperCLOVA X SEED 1.5BRuns greatRuns great
Qwen3.5 35B-A3BRuns greatRuns great
gpt-oss-20bRuns greatRuns great
Qwen3 30B-A3B (2507)Runs greatRuns great
Gemma 4 26B-A4BRuns greatRuns great
DeepSeek R1 Distill Llama 8BRuns greatRuns great
Kanana 1.5 8BRuns greatRuns great
Llama 3.1 8BRuns greatRuns great
Kanana 1.5 15.7B-A3BRuns greatRuns great
Qwen3.5 9BRuns greatRuns great
Qwen3 8BRuns greatRuns great
Gemma 4 12BRuns greatRuns great
Gemma 3 12BRuns greatRuns great
Qwen3 14BRuns greatRuns great
HyperCLOVA X SEED Think 14BRuns wellRuns well
Phi-4Runs wellRuns well
Mistral Small 3.2 24BRuns wellRuns well
Gemma 3 27BRuns wellRuns well
Qwen3.5 27BRuns wellRuns well
EXAONE 4.0 32BRuns wellRuns well
Qwen3 32BRuns wellRuns well
EXAONE 4.5 33BRuns slowlyRuns slowly
DeepSeek R1 Distill Qwen 32BRuns slowlyRuns slowly
Qwen3.5 122B-A10BWon't runWon't run
Solar Open 100BWon't runWon't run
Llama 3.3 70BWon't runWon't run
gpt-oss-120bWon't runWon't run
Solar Open 2 250BWon't runWon't run

Measured results on the Apple M4 Max (32-core GPU)

No public measurements for this device yet.

Frequently asked questions

What is the largest model that runs entirely on the Apple M4 Max (32-core GPU)?

Qwen3.5 35B-A3B at Q4_K_M (a 22.63 GB file) fits entirely in 48 GB of unified memory with 8k context, at est. 59.8 tok/s (47.9–71.8, calibrated estimate ±20%).

How many local LLMs run well on the Apple M4 Max (32-core GPU)?

At Q4_K_M with 8k context, with 48 GB of memory, out of 29 tracked models: 15 run great, 7 run well, 2 run slowly and 5 won't run.

Can the Apple M4 Max (32-core GPU) run a 70B model like Llama 3.3 70B?

Runs slowly — CPU only, IQ4_XS: est. 3.5 tok/s (2.5–4.6, theoretical estimate ±30%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.