Which local LLMs can the Apple M3 Ultra run?

Unified memory — by default the GPU can use about 70% of it

Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Specs

Memory options
96 GB · 256 GB · 512 GB
Memory bandwidth
819.0 GB/s
FP16 compute
57.0 TFLOPS
Launch year
2025

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • 96 GBRuns great22 modelsRuns well4 modelsRuns slowly2 modelsWon't run1 model
  • 256 GBRuns great24 modelsRuns well4 modelsRuns slowly1 modelWon't run0 models
  • 512 GBRuns great24 modelsRuns well4 modelsRuns slowly1 modelWon't run0 models

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the Apple M3 Ultra

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 512 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs great
est. 274.4 tok/s219.5–329.3calibrated estimate ±20%
Q4_K_M1.9 / 358.4 GBUnified memoryup to 64k—
HyperCLOVA X SEED 1.5B
Runs great
est. 198.5 tok/s158.8–238.2calibrated estimate ±20%
Q4_K_M2.4 / 358.4 GBUnified memoryup to 16k—
Qwen3.5 35B-A3B
Runs great
est. 115.5 tok/s92.4–138.6calibrated estimate ±20%
Q4_K_M23.4 / 358.4 GBUnified memoryup to 256k—
gpt-oss-20b
Runs great
115.5 tok/s8k estimate 103.8 tok/s (83.1–124.6, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP412.9 / 358.4 GBUnified memoryup to 128k
  • reasoning model
Qwen3 30B-A3B (2507)
Runs great
est. 84.4 tok/s67.5–101.3calibrated estimate ±20%
Q4_K_M19.9 / 358.4 GBUnified memoryup to 64k—
Gemma 4 26B-A4B
Runs great
est. 82.2 tok/s65.7–98.6calibrated estimate ±20%
Q4_K_M17.9 / 358.4 GBUnified memoryup to 128k—
Kanana 1.5 15.7B-A3B
Runs great
est. 77.4 tok/s61.9–92.9calibrated estimate ±20%
Q4_K_M12.1 / 358.4 GBUnified memoryup to 32k—
gpt-oss-120b
Runs great
est. 77.4 tok/s61.9–92.8calibrated estimate ±20%
MXFP464.3 / 358.4 GBUnified memoryup to 128k
  • reasoning model
DeepSeek R1 Distill Llama 8B
Runs great
est. 60.1 tok/s48.1–72.1calibrated estimate ±20%
Q4_K_M6.6 / 358.4 GBUnified memoryup to 32k
  • reasoning model
Kanana 1.5 8B
Runs great
est. 60.1 tok/s48.1–72.1calibrated estimate ±20%
Q4_K_M6.6 / 358.4 GBUnified memoryup to 32k—
Llama 3.1 8B
Runs great
est. 60.1 tok/s48.1–72.1calibrated estimate ±20%
Q4_K_M6.6 / 358.4 GBUnified memoryup to 64k—
Qwen3.5 9B
Runs great
est. 58.7 tok/s47.0–70.4calibrated estimate ±20%
Q4_K_M6.7 / 358.4 GBUnified memoryup to 256k—
Qwen3 8B
Runs great
est. 57.8 tok/s46.2–69.3calibrated estimate ±20%
Q4_K_M6.8 / 358.4 GBUnified memoryup to 32k—
Gemma 4 12B
Runs great
est. 47.1 tok/s37.7–56.5calibrated estimate ±20%
Q4_K_M8.2 / 358.4 GBUnified memoryup to 128k—
Gemma 3 12B
Runs great
est. 46.0 tok/s36.8–55.2calibrated estimate ±20%
Q4_K_M8.4 / 358.4 GBUnified memoryup to 128k—
Qwen3.5 122B-A10B
Runs great
est. 36.8 tok/s29.4–44.1calibrated estimate ±20%
Q4_K_M79.0 / 358.4 GBUnified memoryup to 128k—
HyperCLOVA X SEED Think 14B
Runs great
est. 35.3 tok/s24.7–46.0theoretical estimate ±30%
Q4_K_M10.8 / 358.4 GBUnified memoryup to 16k
  • reasoning model
Qwen3 14B
Runs great
est. 34.8 tok/s27.9–41.8calibrated estimate ±20%
Q4_K_M10.9 / 358.4 GBUnified memoryup to 32k—
Phi-4
Runs great
est. 33.6 tok/s26.9–40.3calibrated estimate ±20%
Q4_K_M11.3 / 358.4 GBUnified memoryup to 8k
  • reasoning model
Solar Open 2 250B
Runs great
est. 27.7 tok/s22.2–33.3calibrated estimate ±20%
IQ4_XS137.2 / 358.4 GBUnified memoryup to 64k—
Solar Open 100B
Runs great
est. 26.7 tok/s21.4–32.1calibrated estimate ±20%
Q4_K_M64.4 / 358.4 GBUnified memoryup to 16k—
Mistral Small 3.2 24B
Runs great
est. 23.0 tok/s18.4–27.6calibrated estimate ±20%
Q4_K_M16.2 / 358.4 GBUnified memoryup to 16k—
Gemma 3 27B
Runs great
est. 20.9 tok/s16.7–25.1calibrated estimate ±20%
Q4_K_M17.8 / 358.4 GBUnified memoryup to 16k—
Qwen3.5 27B
Runs great
est. 20.4 tok/s16.3–24.5calibrated estimate ±20%
Q4_K_M18.2 / 358.4 GBUnified memoryup to 8k—
EXAONE 4.0 32B
Runs well
est. 18.1 tok/s14.5–21.8calibrated estimate ±20%
Q4_K_M20.5 / 358.4 GBUnified memoryup to 128k
  • reasoning model
EXAONE 4.5 33B
Runs well
est. 17.5 tok/s14.0–21.0calibrated estimate ±20%
Q4_K_M21.2 / 358.4 GBUnified memoryup to 128k
  • reasoning model
Qwen3 32B
Runs well
est. 16.4 tok/s13.2–19.7calibrated estimate ±20%
Q4_K_M22.5 / 358.4 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs great
DeepSeek R1 Distill Qwen 32B
Runs well
est. 16.4 tok/s13.1–19.7calibrated estimate ±20%
Q4_K_M22.6 / 358.4 GBUnified memoryup to 32k
  • reasoning model
Llama 3.3 70B
Runs slowly
est. 8.0 tok/s6.4–9.6calibrated estimate ±20%
Q4_K_M45.8 / 358.4 GBUnified memoryup to 128k
  • Try IQ4_XS: Runs well

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Verdict by memory size

Q4_K_M verdict for each memory configuration — open the model page for speeds
Model96 GB256 GB512 GB
EXAONE 4.0 1.2BRuns greatRuns greatRuns great
HyperCLOVA X SEED 1.5BRuns greatRuns greatRuns great
Qwen3.5 35B-A3BRuns greatRuns greatRuns great
gpt-oss-20bRuns greatRuns greatRuns great
Qwen3 30B-A3B (2507)Runs greatRuns greatRuns great
Gemma 4 26B-A4BRuns greatRuns greatRuns great
Kanana 1.5 15.7B-A3BRuns greatRuns greatRuns great
gpt-oss-120bRuns greatRuns greatRuns great
DeepSeek R1 Distill Llama 8BRuns greatRuns greatRuns great
Kanana 1.5 8BRuns greatRuns greatRuns great
Llama 3.1 8BRuns greatRuns greatRuns great
Qwen3.5 9BRuns greatRuns greatRuns great
Qwen3 8BRuns greatRuns greatRuns great
Gemma 4 12BRuns greatRuns greatRuns great
Gemma 3 12BRuns greatRuns greatRuns great
Qwen3.5 122B-A10BRuns slowlyRuns greatRuns great
HyperCLOVA X SEED Think 14BRuns greatRuns greatRuns great
Qwen3 14BRuns greatRuns greatRuns great
Phi-4Runs greatRuns greatRuns great
Solar Open 2 250BWon't runRuns greatRuns great
Solar Open 100BRuns greatRuns greatRuns great
Mistral Small 3.2 24BRuns greatRuns greatRuns great
Gemma 3 27BRuns greatRuns greatRuns great
Qwen3.5 27BRuns greatRuns greatRuns great
EXAONE 4.0 32BRuns wellRuns wellRuns well
EXAONE 4.5 33BRuns wellRuns wellRuns well
Qwen3 32BRuns wellRuns wellRuns well
DeepSeek R1 Distill Qwen 32BRuns wellRuns wellRuns well
Llama 3.3 70BRuns slowlyRuns slowlyRuns slowly

Measured results on the Apple M3 Ultra

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured results on the Apple M3 Ultra
ModelQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
gpt-oss-20bMXFP4llama.cpp2k2,816.0115.5—github.com2025-08-15
llama-2-7bQ4_0llama.cpp5121,471.092.1Metalgithub.com2025-03-15

Frequently asked questions

What is the largest model that runs entirely on the Apple M3 Ultra?

Solar Open 2 250B at IQ4_XS (a 136.24 GB file) fits entirely in 512 GB of unified memory with 8k context, at est. 27.7 tok/s (22.2–33.3, calibrated estimate ±20%).

How many local LLMs run well on the Apple M3 Ultra?

At Q4_K_M with 8k context, with 512 GB of memory, out of 29 tracked models: 24 run great, 4 run well, 1 runs slowly and 0 won't run.

Can the Apple M3 Ultra run a 70B model like Llama 3.3 70B?

Runs slowly — Unified memory, Q4_K_M: est. 8.0 tok/s (6.4–9.6, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.