Which local LLMs can the GeForce RTX 4080 Laptop run?
Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Specs
- VRAM
- 12 GB
- Memory bandwidth
- 432.0 GB/s
- FP16 compute
- 24.7 TFLOPS
- Launch year
- 2023
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- Runs great10 modelsRuns well3 modelsRuns slowly11 modelsWon't run5 models1 more model runs at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the GeForce RTX 4080 Laptop
Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 230.3 tok/s202.6–257.9calibrated estimate ±12% | Q4_K_M | 2.5 / 12.0 GB | Full GPU | up to 64k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 166.6 tok/s146.6–186.6calibrated estimate ±12% | Q4_K_M | 3.0 / 12.0 GB | Full GPU | up to 16k | — |
| DeepSeek R1 Distill Llama 8B | Runs great | est. 50.5 tok/s44.4–56.5calibrated estimate ±12% | Q4_K_M | 7.2 / 12.0 GB | Full GPU | up to 32k |
|
| Kanana 1.5 8B | Runs great | est. 50.5 tok/s44.4–56.5calibrated estimate ±12% | Q4_K_M | 7.2 / 12.0 GB | Full GPU | up to 32k | — |
| Llama 3.1 8B | Runs great | est. 50.5 tok/s44.4–56.5calibrated estimate ±12% | Q4_K_M | 7.2 / 12.0 GB | Full GPU | up to 32k | — |
| Qwen3.5 9B | Runs great | est. 49.3 tok/s43.4–55.2calibrated estimate ±12% | Q4_K_M | 7.3 / 12.0 GB | Full GPU | up to 64k | — |
| Qwen3 8B | Runs great | est. 48.5 tok/s42.7–54.3calibrated estimate ±12% | Q4_K_M | 7.4 / 12.0 GB | Full GPU | up to 32k | — |
| Gemma 4 12B | Runs great | est. 39.5 tok/s34.8–44.2calibrated estimate ±12% | Q4_K_M | 8.8 / 12.0 GB | Full GPU | up to 32k | — |
| Gemma 3 12B | Runs great | est. 38.6 tok/s34.0–43.2calibrated estimate ±12% | Q4_K_M | 9.0 / 12.0 GB | Full GPU | up to 32k | — |
| Qwen3 14B | Runs great | est. 29.2 tok/s25.7–32.7calibrated estimate ±12% | Q4_K_M | 11.5 / 12.0 GB | Full GPU | up to 8k | — |
| HyperCLOVA X SEED Think 14B | Runs well | est. 29.7 tok/s20.8–38.6theoretical estimate ±30% | Q4_K_M | 11.4 / 12.0 GB | Full GPU | up to 8k |
|
| Phi-4 | Runs well | est. 28.2 tok/s24.8–31.6calibrated estimate ±12% | Q4_K_M | 11.9 / 12.0 GB | Full GPU | up to 8k |
|
| Qwen3.5 35B-A3B | Runs well | est. 20.4 tok/s14.3–26.5theoretical estimate ±30% | Q4_K_M | 2.6 / 12.0 GB + 21.4 GB RAM | MoE experts in RAM | up to 256k |
|
| Gemma 4 26B-A4B | Runs slowly | est. 26.5 tok/s18.5–34.4theoretical estimate ±30% | Q4_K_M | 12.0 / 12.0 GB + 6.5 GB RAM | Partial offload | up to 128k |
|
| Kanana 1.5 15.7B-A3B | Runs slowly | est. 19.3 tok/s13.5–25.1theoretical estimate ±30% | Q4_K_M | 3.0 / 12.0 GB + 9.6 GB RAM | MoE experts in RAM | up to 32k |
|
| Qwen3 30B-A3B (2507) | Runs slowly | est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | Q4_K_M | 2.9 / 12.0 GB + 17.7 GB RAM | MoE experts in RAM | up to 64k |
|
| gpt-oss-20b | Runs slowly | est. 18.5 tok/s12.9–24.0theoretical estimate ±30% | MXFP4 | 2.0 / 12.0 GB + 11.5 GB RAM | MoE experts in RAM | up to 128k |
|
| Mistral Small 3.2 24B | Runs slowly | est. 6.6 tok/s5.2–7.9calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 4.8 GB RAM | Partial offload | up to 32k |
|
| Gemma 3 27B | Runs slowly | est. 5.4 tok/s4.4–6.5calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 6.4 GB RAM | Partial offload | up to 64k |
|
| Qwen3.5 27B | Runs slowly | est. 5.2 tok/s4.2–6.2calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 6.8 GB RAM | Partial offload | up to 64k | — |
| EXAONE 4.0 32B | Runs slowly | est. 4.1 tok/s3.3–4.9calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 9.1 GB RAM | Partial offload | up to 32k |
|
| EXAONE 4.5 33B | Runs slowly | est. 3.9 tok/s3.1–4.6calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 9.8 GB RAM | Partial offload | up to 32k |
|
| Qwen3 32B | Runs slowly | est. 3.3 tok/s2.6–3.9calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 11.1 GB RAM | Partial offload | up to 16k | — |
| DeepSeek R1 Distill Qwen 32B | Runs slowly | est. 3.2 tok/s2.6–3.9calibrated estimate ±20% | Q4_K_M | 12.0 / 12.0 GB + 11.2 GB RAM | Partial offload | up to 8k |
|
| Llama 3.3 70B | Won't runTry Q2_K (heavy quality loss): Runs slowly | est. 2.1 tok/s1.7–2.5calibrated estimate ±20% | Q2_K | 12.0 / 12.0 GB + 18.2 GB RAM | Partial offload | up to 8k | — |
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Qwen3.5 122B-A10B | Won't run | — | Q4_K_M | needs 79.0 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Measured results on the GeForce RTX 4080 Laptop
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the GeForce RTX 4080 Laptop?
Kanana 1.5 15.7B-A3B at IQ4_XS (a 8.71 GB file) fits entirely in 12 GB with 8k context, at est. 71.0 tok/s (56.8–85.2, calibrated estimate ±20%).
How many local LLMs run well on the GeForce RTX 4080 Laptop?
At Q4_K_M with 8k context, out of 29 tracked models: 10 run great, 3 run well, 11 run slowly and 5 won't run.
Can the GeForce RTX 4080 Laptop run a 70B model like Llama 3.3 70B?
Runs slowly — Partial offload, Q2_K (heavy quality loss): est. 2.1 tok/s (1.7–2.5, calibrated estimate ±20%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.