Which local LLMs can the NVIDIA DGX Spark run?

Unified memory — by default the GPU can use about 70% of it

Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.

Specs

Memory options
128 GB
Memory bandwidth
273.0 GB/s
FP16 compute
62.5 TFLOPS
Launch year
2025
Price
$3,999 launch MSRP

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context, with 128 GB of memory.

  • Runs great8 modelsRuns well10 modelsRuns slowly10 modelsWon't run1 model1 more model runs at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the NVIDIA DGX Spark

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 128 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs great
est. 83.1 tok/s66.5–99.8calibrated estimate ±20%
Q4_K_M1.9 / 89.6 GBUnified memoryup to 64k—
HyperCLOVA X SEED 1.5B
Runs great
est. 60.2 tok/s48.1–72.2calibrated estimate ±20%
Q4_K_M2.4 / 89.6 GBUnified memoryup to 16k—
Qwen3.5 35B-A3B
Runs great
est. 53.1 tok/s42.5–63.7calibrated estimate ±20%
Q4_K_M23.4 / 89.6 GBUnified memoryup to 128k—
gpt-oss-20b
Runs great
est. 47.7 tok/s38.2–57.3calibrated estimate ±20%
MXFP412.9 / 89.6 GBUnified memoryup to 32k
  • reasoning model
Qwen3 30B-A3B (2507)
Runs great
est. 38.8 tok/s31.1–46.6calibrated estimate ±20%
Q4_K_M19.9 / 89.6 GBUnified memoryup to 32k—
Gemma 4 26B-A4B
Runs great
est. 37.8 tok/s30.2–45.3calibrated estimate ±20%
Q4_K_M17.9 / 89.6 GBUnified memoryup to 64k—
Kanana 1.5 15.7B-A3B
Runs great
est. 35.6 tok/s28.5–42.7calibrated estimate ±20%
Q4_K_M12.1 / 89.6 GBUnified memoryup to 16k—
gpt-oss-120b
Runs great
35.0 tok/s8k estimate 35.6 tok/s (28.5–42.7, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP464.3 / 89.6 GBUnified memoryup to 16k
  • reasoning model
DeepSeek R1 Distill Llama 8B
Runs well
est. 18.2 tok/s14.6–21.9calibrated estimate ±20%
Q4_K_M6.6 / 89.6 GBUnified memoryup to 16k
  • reasoning model
Kanana 1.5 8B
Runs well
est. 18.2 tok/s14.6–21.9calibrated estimate ±20%
Q4_K_M6.6 / 89.6 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs great
Llama 3.1 8B
Runs well
est. 18.2 tok/s14.6–21.9calibrated estimate ±20%
Q4_K_M6.6 / 89.6 GBUnified memoryup to 64k
  • Q2_K (heavy quality loss): Runs great
Qwen3.5 9B
Runs well
est. 17.8 tok/s14.2–21.3calibrated estimate ±20%
Q4_K_M6.7 / 89.6 GBUnified memoryup to 128k—
Qwen3 8B
Runs well
est. 17.5 tok/s14.0–21.0calibrated estimate ±20%
Q4_K_M6.8 / 89.6 GBUnified memoryup to 32k
  • Q2_K (heavy quality loss): Runs great
Qwen3.5 122B-A10B
Runs well
est. 16.9 tok/s13.5–20.3calibrated estimate ±20%
Q4_K_M79.0 / 89.6 GBUnified memoryup to 256k
  • UD-Q2_K_XL (heavy quality loss): Runs great
Gemma 4 12B
Runs well
est. 14.3 tok/s11.4–17.1calibrated estimate ±20%
Q4_K_M8.2 / 89.6 GBUnified memoryup to 64k
  • UD-Q2_K_XL (heavy quality loss): Runs great
Gemma 3 12B
Runs well
est. 13.9 tok/s11.1–16.7calibrated estimate ±20%
Q4_K_M8.4 / 89.6 GBUnified memoryup to 64k
  • Q2_K (heavy quality loss): Runs great
Solar Open 100B
Runs well
est. 12.3 tok/s9.8–14.7calibrated estimate ±20%
Q4_K_M64.4 / 89.6 GBUnified memoryup to 16k—
Qwen3 14B
Runs well
est. 10.6 tok/s8.4–12.7calibrated estimate ±20%
Q4_K_M10.9 / 89.6 GBUnified memoryup to 16k—
HyperCLOVA X SEED Think 14B
Runs slowly
est. 10.7 tok/s7.5–13.9theoretical estimate ±30%
Q4_K_M10.8 / 89.6 GBUnified memoryup to 128k
  • reasoning model
Phi-4
Runs slowly
est. 10.2 tok/s8.1–12.2calibrated estimate ±20%
Q4_K_M11.3 / 89.6 GBUnified memoryup to 16k
  • reasoning model
Mistral Small 3.2 24B
Runs slowly
est. 7.0 tok/s5.6–8.4calibrated estimate ±20%
Q4_K_M16.2 / 89.6 GBUnified memoryup to 128k
  • Q2_K (heavy quality loss): Runs well
Gemma 3 27B
Runs slowly
est. 6.3 tok/s5.1–7.6calibrated estimate ±20%
Q4_K_M17.8 / 89.6 GBUnified memoryup to 128k
  • Q2_K (heavy quality loss): Runs well
Qwen3.5 27B
Runs slowly
est. 6.2 tok/s5.0–7.4calibrated estimate ±20%
Q4_K_M18.2 / 89.6 GBUnified memoryup to 256k
  • UD-Q2_K_XL (heavy quality loss): Runs well
EXAONE 4.0 32B
Runs slowly
est. 5.5 tok/s4.4–6.6calibrated estimate ±20%
Q4_K_M20.5 / 89.6 GBUnified memoryup to 128k
  • reasoning model
EXAONE 4.5 33B
Runs slowly
est. 5.3 tok/s4.2–6.4calibrated estimate ±20%
Q4_K_M21.2 / 89.6 GBUnified memoryup to 128k
  • reasoning model
Qwen3 32B
Runs slowly
est. 5.0 tok/s4.0–6.0calibrated estimate ±20%
Q4_K_M22.5 / 89.6 GBUnified memoryup to 32k—
DeepSeek R1 Distill Qwen 32B
Runs slowly
est. 5.0 tok/s4.0–6.0calibrated estimate ±20%
Q4_K_M22.6 / 89.6 GBUnified memoryup to 32k
  • reasoning model
Llama 3.3 70B
Runs slowly
est. 2.4 tok/s1.9–2.9calibrated estimate ±20%
Q4_K_M45.8 / 89.6 GBUnified memoryup to 32k—
Solar Open 2 250B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 15.3 tok/s10.7–19.9theoretical estimate ±30%
Q2_K95.8 GB RAMCPU onlyup to 256k
  • CPU inference — capped at Runs slowly

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Measured results on the NVIDIA DGX Spark

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured results on the NVIDIA DGX Spark
ModelQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
gpt-oss-120bMXFP4llama.cpp2k1,717.035.0—github.com2025-10-14

Frequently asked questions

What is the largest model that runs entirely on the NVIDIA DGX Spark?

Qwen3.5 122B-A10B at Q4_K_M (a 78.26 GB file) fits entirely in 128 GB of unified memory with 8k context, at est. 16.9 tok/s (13.5–20.3, calibrated estimate ±20%).

How many local LLMs run well on the NVIDIA DGX Spark?

At Q4_K_M with 8k context, with 128 GB of memory, out of 29 tracked models: 8 run great, 10 run well, 10 run slowly and 1 won't run.

Can the NVIDIA DGX Spark run a 70B model like Llama 3.3 70B?

Runs slowly — Unified memory, Q4_K_M: est. 2.4 tok/s (1.9–2.9, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.