Which local LLMs can the GeForce RTX 5090 Laptop run?
Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Specs
- VRAM
- 24 GB
- Memory bandwidth
- 896.0 GB/s
- FP16 compute
- 60.0 TFLOPS
- Launch year
- 2025
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- Runs great23 modelsRuns well1 modelRuns slowly0 modelsWon't run5 models2 more models run at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the GeForce RTX 5090 Laptop
Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 477.6 tok/s420.3–534.9calibrated estimate ±12% | Q4_K_M | 2.5 / 24.0 GB | Full GPU | up to 64k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 345.5 tok/s304.0–387.0calibrated estimate ±12% | Q4_K_M | 3.0 / 24.0 GB | Full GPU | up to 16k | — |
| Qwen3.5 35B-A3B | Runs great | est. 196.1 tok/s156.9–235.3calibrated estimate ±20% | Q4_K_M | 24.0 / 24.0 GB | Full GPU | up to 8k | — |
| gpt-oss-20b | Runs great | est. 176.3 tok/s141.0–211.5calibrated estimate ±20% | MXFP4 | 13.5 / 24.0 GB | Full GPU | up to 128k |
|
| Qwen3 30B-A3B (2507) | Runs great | est. 143.3 tok/s114.6–172.0calibrated estimate ±20% | Q4_K_M | 20.5 / 24.0 GB | Full GPU | up to 32k | — |
| Gemma 4 26B-A4B | Runs great | est. 139.5 tok/s111.6–167.4calibrated estimate ±20% | Q4_K_M | 18.5 / 24.0 GB | Full GPU | up to 64k | — |
| Kanana 1.5 15.7B-A3B | Runs great | est. 131.4 tok/s105.1–157.7calibrated estimate ±20% | Q4_K_M | 12.7 / 24.0 GB | Full GPU | up to 32k | — |
| DeepSeek R1 Distill Llama 8B | Runs great | est. 104.6 tok/s92.1–117.2calibrated estimate ±12% | Q4_K_M | 7.2 / 24.0 GB | Full GPU | up to 64k |
|
| Kanana 1.5 8B | Runs great | est. 104.6 tok/s92.1–117.2calibrated estimate ±12% | Q4_K_M | 7.2 / 24.0 GB | Full GPU | up to 32k | — |
| Llama 3.1 8B | Runs great | est. 104.6 tok/s92.1–117.2calibrated estimate ±12% | Q4_K_M | 7.2 / 24.0 GB | Full GPU | up to 64k | — |
| Qwen3.5 9B | Runs great | est. 102.2 tok/s89.9–114.4calibrated estimate ±12% | Q4_K_M | 7.3 / 24.0 GB | Full GPU | up to 256k | — |
| Qwen3 8B | Runs great | est. 100.5 tok/s88.5–112.6calibrated estimate ±12% | Q4_K_M | 7.4 / 24.0 GB | Full GPU | up to 32k | — |
| Gemma 4 12B | Runs great | est. 81.9 tok/s72.1–91.7calibrated estimate ±12% | Q4_K_M | 8.8 / 24.0 GB | Full GPU | up to 128k | — |
| Gemma 3 12B | Runs great | est. 80.0 tok/s70.4–89.6calibrated estimate ±12% | Q4_K_M | 9.0 / 24.0 GB | Full GPU | up to 128k | — |
| HyperCLOVA X SEED Think 14B | Runs great | est. 61.5 tok/s43.1–80.0theoretical estimate ±30% | Q4_K_M | 11.4 / 24.0 GB | Full GPU | up to 64k |
|
| Qwen3 14B | Runs great | est. 60.6 tok/s53.4–67.9calibrated estimate ±12% | Q4_K_M | 11.5 / 24.0 GB | Full GPU | up to 32k | — |
| Phi-4 | Runs great | est. 58.5 tok/s51.4–65.5calibrated estimate ±12% | Q4_K_M | 11.9 / 24.0 GB | Full GPU | up to 16k |
|
| Mistral Small 3.2 24B | Runs great | est. 40.0 tok/s35.2–44.8calibrated estimate ±12% | Q4_K_M | 16.8 / 24.0 GB | Full GPU | up to 32k | — |
| Gemma 3 27B | Runs great | est. 36.4 tok/s32.1–40.8calibrated estimate ±12% | Q4_K_M | 18.4 / 24.0 GB | Full GPU | up to 64k | — |
| Qwen3.5 27B | Runs great | est. 35.5 tok/s31.3–39.8calibrated estimate ±12% | Q4_K_M | 18.8 / 24.0 GB | Full GPU | up to 64k | — |
| EXAONE 4.0 32B | Runs great | est. 31.6 tok/s27.8–35.3calibrated estimate ±12% | Q4_K_M | 21.1 / 24.0 GB | Full GPU | up to 16k |
|
| EXAONE 4.5 33B | Runs great | est. 30.5 tok/s26.8–34.1calibrated estimate ±12% | Q4_K_M | 21.8 / 24.0 GB | Full GPU | up to 8k |
|
| Qwen3 32B | Runs great | est. 28.6 tok/s25.2–32.1calibrated estimate ±12% | Q4_K_M | 23.1 / 24.0 GB | Full GPU | up to 8k | — |
| DeepSeek R1 Distill Qwen 32B | Runs well | est. 28.5 tok/s25.1–31.9calibrated estimate ±12% | Q4_K_M | 23.2 / 24.0 GB | Full GPU | up to 8k |
|
| Qwen3.5 122B-A10B | Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly | est. 20.3 tok/s14.2–26.4theoretical estimate ±30% | UD-Q2_K_XL | 24.0 / 24.0 GB + 20.2 GB RAM | Partial offload | up to 128k |
|
| Llama 3.3 70B | Won't runTry IQ4_XS: Runs slowly | est. 2.2 tok/s1.7–2.6calibrated estimate ±20% | IQ4_XS | 24.0 / 24.0 GB + 17.8 GB RAM | Partial offload | up to 8k | — |
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Measured results on the GeForce RTX 5090 Laptop
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the GeForce RTX 5090 Laptop?
Qwen3.5 35B-A3B at Q4_K_M (a 22.63 GB file) fits entirely in 24 GB with 8k context, at est. 196.1 tok/s (156.9–235.3, calibrated estimate ±20%).
How many local LLMs run well on the GeForce RTX 5090 Laptop?
At Q4_K_M with 8k context, out of 29 tracked models: 23 run great, 1 runs well, 0 run slowly and 5 won't run.
Can the GeForce RTX 5090 Laptop run a 70B model like Llama 3.3 70B?
Runs slowly — Partial offload, IQ4_XS: est. 2.2 tok/s (1.7–2.6, calibrated estimate ±20%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.