Which local LLMs can the DDR4-3200 dual-channel CPU run?

No GPU — the whole model runs from system RAM

CPU-only inference with 4 GB of RAM kept for the OS, 8k context and an f16 KV cache. Verdicts are shown for each RAM size.

Specs

Memory options
16 GB · 32 GB · 64 GB
Memory bandwidth
51.2 GB/s
Launch year
2016

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • 16 GBRuns great0 modelsRuns well0 modelsRuns slowly11 modelsWon't run18 models3 more models run at a lower quant.
  • 32 GBRuns great0 modelsRuns well0 modelsRuns slowly15 modelsWon't run14 models3 more models run at a lower quant.
  • 64 GBRuns great0 modelsRuns well0 modelsRuns slowly15 modelsWon't run14 models5 more models run at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the DDR4-3200 dual-channel CPU

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 64 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs slowly
est. 19.5 tok/s15.6–23.4calibrated estimate ±20%
Q4_K_M1.3 GB RAMCPU onlyup to 64k
  • CPU inference — capped at Runs slowly
HyperCLOVA X SEED 1.5B
Runs slowly
est. 14.1 tok/s11.3–16.9calibrated estimate ±20%
Q4_K_M1.8 GB RAMCPU onlyup to 16k
  • CPU inference — capped at Runs slowly
Qwen3.5 35B-A3B
Runs slowly
est. 10.7 tok/s8.6–12.8calibrated estimate ±20%
Q4_K_M22.8 GB RAMCPU onlyup to 256k
  • CPU inference — capped at Runs slowly
gpt-oss-20b
Runs slowly
est. 9.6 tok/s7.7–11.6calibrated estimate ±20%
MXFP412.3 GB RAMCPU onlyup to 128k
  • reasoning model
Qwen3 30B-A3B (2507)
Runs slowly
est. 7.8 tok/s6.3–9.4calibrated estimate ±20%
Q4_K_M19.4 GB RAMCPU onlyup to 64k—
Gemma 4 26B-A4B
Runs slowly
est. 7.6 tok/s6.1–9.1calibrated estimate ±20%
Q4_K_M17.3 GB RAMCPU onlyup to 128k—
Kanana 1.5 15.7B-A3B
Runs slowly
est. 7.2 tok/s5.7–8.6calibrated estimate ±20%
Q4_K_M11.5 GB RAMCPU onlyup to 32k—
DeepSeek R1 Distill Llama 8B
Runs slowly
est. 4.3 tok/s3.4–5.1calibrated estimate ±20%
Q4_K_M6.0 GB RAMCPU onlyup to 16k
  • reasoning model
Kanana 1.5 8B
Runs slowly
est. 4.3 tok/s3.4–5.1calibrated estimate ±20%
Q4_K_M6.0 GB RAMCPU onlyup to 32k—
Llama 3.1 8B
Runs slowly
est. 4.3 tok/s3.4–5.1calibrated estimate ±20%
Q4_K_M6.0 GB RAMCPU onlyup to 32k—
Qwen3.5 9B
Runs slowly
est. 4.2 tok/s3.3–5.0calibrated estimate ±20%
Q4_K_M6.1 GB RAMCPU onlyup to 128k—
Qwen3 8B
Runs slowly
est. 4.1 tok/s3.3–4.9calibrated estimate ±20%
Q4_K_M6.2 GB RAMCPU onlyup to 32k—
Gemma 4 12B
Runs slowly
est. 3.3 tok/s2.7–4.0calibrated estimate ±20%
Q4_K_M7.7 GB RAMCPU onlyup to 64k—
Gemma 3 12B
Runs slowly
est. 3.3 tok/s2.6–3.9calibrated estimate ±20%
Q4_K_M7.8 GB RAMCPU onlyup to 64k—
Qwen3 14B
Runs slowly
est. 2.5 tok/s2.0–3.0calibrated estimate ±20%
Q4_K_M10.3 GB RAMCPU onlyup to 16k—
Qwen3.5 122B-A10B
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 6.1 tok/s4.9–7.3calibrated estimate ±20%
UD-Q2_K_XL43.1 GB RAMCPU onlyup to 256k—
Solar Open 100B
Won't runTry IQ4_XS: Runs slowly
est. 2.7 tok/s2.2–3.3calibrated estimate ±20%
IQ4_XS57.1 GB RAMCPU onlyup to 16k—
HyperCLOVA X SEED Think 14B
Won't run
est. 2.5 tok/s1.8–3.3theoretical estimate ±30%
Q4_K_M10.2 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
Mistral Small 3.2 24B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.5 tok/s2.0–3.0calibrated estimate ±20%
Q2_K10.2 GB RAMCPU onlyup to 16k—
Phi-4
Won't run
est. 2.4 tok/s1.9–2.9calibrated estimate ±20%
Q4_K_M10.7 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
Gemma 3 27B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.3 tok/s1.8–2.7calibrated estimate ±20%
Q2_K11.2 GB RAMCPU onlyup to 16k—
Qwen3.5 27B
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 2.0 tok/s1.6–2.4calibrated estimate ±20%
UD-Q2_K_XL12.6 GB RAMCPU onlyup to 8k—
EXAONE 4.0 32B
Won't run
est. 1.3 tok/s1.0–1.5calibrated estimate ±20%
Q4_K_M19.9 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
EXAONE 4.5 33B
Won't run
est. 1.2 tok/s1.0–1.5calibrated estimate ±20%
Q4_K_M20.6 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
Qwen3 32B
Won't run
est. 1.2 tok/s0.9–1.4calibrated estimate ±20%
Q4_K_M21.9 GB RAMCPU only—
  • Fits, but too slow to use
DeepSeek R1 Distill Qwen 32B
Won't run
est. 1.2 tok/s0.9–1.4calibrated estimate ±20%
Q4_K_M22.0 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
Llama 3.3 70B
Won't run
est. 0.6 tok/s0.5–0.7calibrated estimate ±20%
Q4_K_M45.2 GB RAMCPU only—
  • Fits, but too slow to use
gpt-oss-120b
Won't run
—MXFP4needs 63.7 GB——
  • Needs about 63.7 GB; 60.0 GB of RAM is free
  • reasoning model
Solar Open 2 250B
Won't run
—IQ4_XSneeds 136.6 GB——
  • Needs about 136.6 GB; 60.0 GB of RAM is free

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Verdict by memory size

Q4_K_M verdict for each memory configuration — open the model page for speeds
Model16 GB32 GB64 GB
EXAONE 4.0 1.2BRuns slowlyRuns slowlyRuns slowly
HyperCLOVA X SEED 1.5BRuns slowlyRuns slowlyRuns slowly
Qwen3.5 35B-A3BWon't runRuns slowlyRuns slowly
gpt-oss-20bWon't runRuns slowlyRuns slowly
Qwen3 30B-A3B (2507)Won't runRuns slowlyRuns slowly
Gemma 4 26B-A4BWon't runRuns slowlyRuns slowly
Kanana 1.5 15.7B-A3BRuns slowlyRuns slowlyRuns slowly
DeepSeek R1 Distill Llama 8BRuns slowlyRuns slowlyRuns slowly
Kanana 1.5 8BRuns slowlyRuns slowlyRuns slowly
Llama 3.1 8BRuns slowlyRuns slowlyRuns slowly
Qwen3.5 9BRuns slowlyRuns slowlyRuns slowly
Qwen3 8BRuns slowlyRuns slowlyRuns slowly
Gemma 4 12BRuns slowlyRuns slowlyRuns slowly
Gemma 3 12BRuns slowlyRuns slowlyRuns slowly
Qwen3 14BRuns slowlyRuns slowlyRuns slowly
Qwen3.5 122B-A10BWon't runWon't runWon't run
Solar Open 100BWon't runWon't runWon't run
HyperCLOVA X SEED Think 14BWon't runWon't runWon't run
Mistral Small 3.2 24BWon't runWon't runWon't run
Phi-4Won't runWon't runWon't run
Gemma 3 27BWon't runWon't runWon't run
Qwen3.5 27BWon't runWon't runWon't run
EXAONE 4.0 32BWon't runWon't runWon't run
EXAONE 4.5 33BWon't runWon't runWon't run
Qwen3 32BWon't runWon't runWon't run
DeepSeek R1 Distill Qwen 32BWon't runWon't runWon't run
Llama 3.3 70BWon't runWon't runWon't run
gpt-oss-120bWon't runWon't runWon't run
Solar Open 2 250BWon't runWon't runWon't run

Measured results on the DDR4-3200 dual-channel CPU

No public measurements for this device yet.

Frequently asked questions

What is the largest model this CPU setup can run?

Solar Open 100B at IQ4_XS (a 55.53 GB file) runs from 64 GB of RAM at est. 2.7 tok/s (2.2–3.3, calibrated estimate ±20%).

How many local LLMs run well on the DDR4-3200 dual-channel CPU?

At Q4_K_M with 8k context, with 64 GB of memory, out of 29 tracked models: 0 run great, 0 run well, 15 run slowly and 14 won't run.

Can the DDR4-3200 dual-channel CPU run a 70B model like Llama 3.3 70B?

It fits in memory but is too slow to use: est. 0.6 tok/s (0.5–0.7, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.