Which local LLMs can the GeForce RTX 4080 run?
Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Specs
- VRAM
- 16 GB
- Memory bandwidth
- 717.0 GB/s
- FP16 compute
- 48.7 TFLOPS
- Launch year
- 2022
- Price
- $1,199 launch MSRP
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- Runs great14 modelsRuns well2 modelsRuns slowly8 modelsWon't run5 models1 more model runs at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the GeForce RTX 4080
Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 382.2 tok/s336.3–428.0calibrated estimate ±12% | Q4_K_M | 2.5 / 16.0 GB | Full GPU | up to 64k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 276.5 tok/s243.3–309.7calibrated estimate ±12% | Q4_K_M | 3.0 / 16.0 GB | Full GPU | up to 16k | — |
| gpt-oss-20b | Runs great | est. 141.1 tok/s112.9–169.3calibrated estimate ±20% | MXFP4 | 13.5 / 16.0 GB | Full GPU | up to 64k |
|
| Kanana 1.5 15.7B-A3B | Runs great | est. 105.1 tok/s84.1–126.2calibrated estimate ±20% | Q4_K_M | 12.7 / 16.0 GB | Full GPU | up to 16k | — |
| DeepSeek R1 Distill Llama 8B | Runs great | est. 83.7 tok/s73.7–93.8calibrated estimate ±12% | Q4_K_M | 7.2 / 16.0 GB | Full GPU | up to 64k |
|
| Kanana 1.5 8B | Runs great | est. 83.7 tok/s73.7–93.8calibrated estimate ±12% | Q4_K_M | 7.2 / 16.0 GB | Full GPU | up to 32k | — |
| Llama 3.1 8B | Runs great | est. 83.7 tok/s73.7–93.8calibrated estimate ±12% | Q4_K_M | 7.2 / 16.0 GB | Full GPU | up to 64k | — |
| Qwen3.5 9B | Runs great | est. 81.8 tok/s72.0–91.6calibrated estimate ±12% | Q4_K_M | 7.3 / 16.0 GB | Full GPU | up to 128k | — |
| Qwen3 8B | Runs great | est. 80.5 tok/s70.8–90.1calibrated estimate ±12% | Q4_K_M | 7.4 / 16.0 GB | Full GPU | up to 32k | — |
| Gemma 4 12B | Runs great | est. 65.5 tok/s57.7–73.4calibrated estimate ±12% | Q4_K_M | 8.8 / 16.0 GB | Full GPU | up to 64k | — |
| Gemma 3 12B | Runs great | est. 64.0 tok/s56.4–71.7calibrated estimate ±12% | Q4_K_M | 9.0 / 16.0 GB | Full GPU | up to 64k | — |
| HyperCLOVA X SEED Think 14B | Runs great | est. 49.2 tok/s34.5–64.0theoretical estimate ±30% | Q4_K_M | 11.4 / 16.0 GB | Full GPU | up to 32k |
|
| Qwen3 14B | Runs great | est. 48.5 tok/s42.7–54.4calibrated estimate ±12% | Q4_K_M | 11.5 / 16.0 GB | Full GPU | up to 32k | — |
| Phi-4 | Runs great | est. 46.8 tok/s41.2–52.4calibrated estimate ±12% | Q4_K_M | 11.9 / 16.0 GB | Full GPU | up to 16k |
|
| Qwen3.5 35B-A3B | Runs well | est. 20.4 tok/s14.3–26.5theoretical estimate ±30% | Q4_K_M | 2.6 / 16.0 GB + 21.4 GB RAM | MoE experts in RAM | up to 256k |
|
| Mistral Small 3.2 24B | Runs well | est. 20.0 tok/s16.0–24.0calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 0.8 GB RAM | Partial offload | up to 8k |
|
| Gemma 4 26B-A4B | Runs slowly | est. 53.9 tok/s37.7–70.1theoretical estimate ±30% | Q4_K_M | 16.0 / 16.0 GB + 2.5 GB RAM | Partial offload | up to 256k |
|
| Qwen3 30B-A3B (2507) | Runs slowly | est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | Q4_K_M | 2.9 / 16.0 GB + 17.7 GB RAM | MoE experts in RAM | up to 64k |
|
| Gemma 3 27B | Runs slowly | est. 11.8 tok/s9.4–14.1calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 2.4 GB RAM | Partial offload | up to 64k |
|
| Qwen3.5 27B | Runs slowly | est. 10.6 tok/s8.5–12.7calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 2.8 GB RAM | Partial offload | up to 128k |
|
| EXAONE 4.0 32B | Runs slowly | est. 6.9 tok/s5.5–8.3calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 5.1 GB RAM | Partial offload | up to 64k |
|
| EXAONE 4.5 33B | Runs slowly | est. 6.2 tok/s5.0–7.4calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 5.8 GB RAM | Partial offload | up to 64k |
|
| Qwen3 32B | Runs slowly | est. 4.9 tok/s3.9–5.9calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 7.1 GB RAM | Partial offload | up to 32k |
|
| DeepSeek R1 Distill Qwen 32B | Runs slowly | est. 4.9 tok/s3.9–5.8calibrated estimate ±20% | Q4_K_M | 16.0 / 16.0 GB + 7.2 GB RAM | Partial offload | up to 16k |
|
| Llama 3.3 70B | Won't runTry Q2_K (heavy quality loss): Runs slowly | est. 2.7 tok/s2.1–3.2calibrated estimate ±20% | Q2_K | 16.0 / 16.0 GB + 14.2 GB RAM | Partial offload | up to 16k | — |
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Qwen3.5 122B-A10B | Won't run | — | Q4_K_M | needs 79.0 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Measured results on the GeForce RTX 4080
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the GeForce RTX 4080?
Mistral Small 3.2 24B at IQ4_XS (a 12.76 GB file) fits entirely in 16 GB with 8k context, at est. 35.6 tok/s (31.3–39.9, calibrated estimate ±12%).
How many local LLMs run well on the GeForce RTX 4080?
At Q4_K_M with 8k context, out of 29 tracked models: 14 run great, 2 run well, 8 run slowly and 5 won't run.
Can the GeForce RTX 4080 run a 70B model like Llama 3.3 70B?
Runs slowly — Partial offload, Q2_K (heavy quality loss): est. 2.7 tok/s (2.1–3.2, calibrated estimate ±20%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.