Which local LLMs can the GeForce RTX 5070 Ti run?

Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.

Specs

VRAM
16 GB
Memory bandwidth
896.0 GB/s
FP16 compute
44.1 TFLOPS
Launch year
2025
Price
$949 street (as of 2026-08-09)

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • Runs great14 modelsRuns well2 modelsRuns slowly8 modelsWon't run5 models1 more model runs at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the GeForce RTX 5070 Ti

Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.

Scroll sideways to see every column.

29 models sorted by verdict and speed
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs great
est. 477.6 tok/s420.3–534.9calibrated estimate ±12%
Q4_K_M2.5 / 16.0 GBFull GPUup to 64k—
HyperCLOVA X SEED 1.5B
Runs great
est. 345.5 tok/s304.0–387.0calibrated estimate ±12%
Q4_K_M3.0 / 16.0 GBFull GPUup to 16k—
gpt-oss-20b
Runs great
189.5 tok/s8k estimate 176.3 tok/s (141.0–211.5, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 16.0 GBFull GPUup to 64k
  • reasoning model
Kanana 1.5 15.7B-A3B
Runs great
est. 131.4 tok/s105.1–157.7calibrated estimate ±20%
Q4_K_M12.7 / 16.0 GBFull GPUup to 16k—
DeepSeek R1 Distill Llama 8B
Runs great
est. 104.6 tok/s92.1–117.2calibrated estimate ±12%
Q4_K_M7.2 / 16.0 GBFull GPUup to 64k
  • reasoning model
Kanana 1.5 8B
Runs great
est. 104.6 tok/s92.1–117.2calibrated estimate ±12%
Q4_K_M7.2 / 16.0 GBFull GPUup to 32k—
Llama 3.1 8B
Runs great
est. 104.6 tok/s92.1–117.2calibrated estimate ±12%
Q4_K_M7.2 / 16.0 GBFull GPUup to 64k—
Qwen3.5 9B
Runs great
est. 102.2 tok/s89.9–114.4calibrated estimate ±12%
Q4_K_M7.3 / 16.0 GBFull GPUup to 128k—
Qwen3 8B
Runs great
est. 100.5 tok/s88.5–112.6calibrated estimate ±12%
Q4_K_M7.4 / 16.0 GBFull GPUup to 32k—
Gemma 4 12B
Runs great
est. 81.9 tok/s72.1–91.7calibrated estimate ±12%
Q4_K_M8.8 / 16.0 GBFull GPUup to 64k—
Gemma 3 12B
Runs great
est. 80.0 tok/s70.4–89.6calibrated estimate ±12%
Q4_K_M9.0 / 16.0 GBFull GPUup to 64k—
HyperCLOVA X SEED Think 14B
Runs great
est. 61.5 tok/s43.1–80.0theoretical estimate ±30%
Q4_K_M11.4 / 16.0 GBFull GPUup to 32k
  • reasoning model
Qwen3 14B
Runs great
est. 60.6 tok/s53.4–67.9calibrated estimate ±12%
Q4_K_M11.5 / 16.0 GBFull GPUup to 32k—
Phi-4
Runs great
est. 58.5 tok/s51.4–65.5calibrated estimate ±12%
Q4_K_M11.9 / 16.0 GBFull GPUup to 16k
  • reasoning model
Mistral Small 3.2 24B
Runs well
est. 22.6 tok/s18.1–27.2calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 0.8 GB RAMPartial offloadup to 8k
  • Not fully on the GPU — Runs great needs the whole model in GPU memory
  • Try IQ4_XS: Runs great
Qwen3.5 35B-A3B
Runs well
est. 20.4 tok/s14.3–26.5theoretical estimate ±30%
Q4_K_M2.6 / 16.0 GB + 21.4 GB RAMMoE experts in RAMup to 256k
  • Not fully on the GPU — Runs great needs the whole model in GPU memory
  • UD-Q2_K_XL (heavy quality loss): Runs great
Gemma 4 26B-A4B
Runs slowly
est. 58.8 tok/s41.1–76.4theoretical estimate ±30%
Q4_K_M16.0 / 16.0 GB + 2.5 GB RAMPartial offloadup to 256k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • UD-Q2_K_XL (heavy quality loss): Runs great
Qwen3 30B-A3B (2507)
Runs slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
Q4_K_M2.9 / 16.0 GB + 17.7 GB RAMMoE experts in RAMup to 64k
  • Experts run from system RAM — Runs well needs 20 tok/s or more on this path
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
Gemma 3 27B
Runs slowly
est. 12.6 tok/s10.1–15.2calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 2.4 GB RAMPartial offloadup to 64k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
Qwen3.5 27B
Runs slowly
est. 11.3 tok/s9.0–13.6calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 2.8 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • UD-Q2_K_XL (heavy quality loss): Runs great
EXAONE 4.0 32B
Runs slowly
est. 7.2 tok/s5.7–8.6calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 5.1 GB RAMPartial offloadup to 64k
  • reasoning model
EXAONE 4.5 33B
Runs slowly
est. 6.4 tok/s5.1–7.7calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 5.8 GB RAMPartial offloadup to 64k
  • reasoning model
Qwen3 32B
Runs slowly
est. 5.1 tok/s4.0–6.1calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 7.1 GB RAMPartial offloadup to 32k
  • Q2_K (heavy quality loss): Runs great
DeepSeek R1 Distill Qwen 32B
Runs slowly
est. 5.0 tok/s4.0–6.0calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 7.2 GB RAMPartial offloadup to 16k
  • Q2_K (heavy quality loss): Runs great
  • reasoning model
Llama 3.3 70B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.2–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—
gpt-oss-120b
Won't run
—MXFP4needs 64.3 GB——
  • Needs about 64.3 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
  • reasoning model
Qwen3.5 122B-A10B
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
Solar Open 100B
Won't run
—Q4_K_Mneeds 64.4 GB——
  • Needs about 64.4 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 64 GB RAM: Runs slowly
Solar Open 2 250B
Won't run
—IQ4_XSneeds 137.2 GB——
  • Needs about 137.2 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 128 GB RAM: Runs slowly

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Measured results on the GeForce RTX 5070 Ti

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured results on the GeForce RTX 5070 Ti
ModelQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
llama-2-7bQ4_0llama.cpp5126,952.0176.9—github.com2025-08-01
gpt-oss-20bMXFP4llama.cpp2k6,326.0189.5—github.com2025-08-15

Frequently asked questions

What is the largest model that runs entirely on the GeForce RTX 5070 Ti?

Mistral Small 3.2 24B at IQ4_XS (a 12.76 GB file) fits entirely in 16 GB with 8k context, at est. 44.5 tok/s (39.1–49.8, calibrated estimate ±12%).

How many local LLMs run well on the GeForce RTX 5070 Ti?

At Q4_K_M with 8k context, out of 29 tracked models: 14 run great, 2 run well, 8 run slowly and 5 won't run.

Can the GeForce RTX 5070 Ti run a 70B model like Llama 3.3 70B?

Runs slowly — Partial offload, Q2_K (heavy quality loss): est. 2.7 tok/s (2.2–3.2, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.