Which local LLMs can the GeForce RTX 4060 Laptop run?
Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Specs
- VRAM
- 8 GB
- Memory bandwidth
- 256.0 GB/s
- FP16 compute
- 11.6 TFLOPS
- Launch year
- 2023
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- Runs great6 modelsRuns well2 modelsRuns slowly13 modelsWon't run8 models3 more models run at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the GeForce RTX 4060 Laptop
Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 136.4 tok/s120.1–152.8calibrated estimate ±12% | Q4_K_M | 2.5 / 8.0 GB | Full GPU | up to 64k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 98.7 tok/s86.9–110.6calibrated estimate ±12% | Q4_K_M | 3.0 / 8.0 GB | Full GPU | up to 16k | — |
| Kanana 1.5 8B | Runs great | est. 29.9 tok/s26.3–33.5calibrated estimate ±12% | Q4_K_M | 7.2 / 8.0 GB | Full GPU | up to 8k | — |
| Llama 3.1 8B | Runs great | est. 29.9 tok/s26.3–33.5calibrated estimate ±12% | Q4_K_M | 7.2 / 8.0 GB | Full GPU | up to 8k | — |
| Qwen3.5 9B | Runs great | est. 29.2 tok/s25.7–32.7calibrated estimate ±12% | Q4_K_M | 7.3 / 8.0 GB | Full GPU | up to 16k | — |
| Qwen3 8B | Runs great | est. 28.7 tok/s25.3–32.2calibrated estimate ±12% | Q4_K_M | 7.4 / 8.0 GB | Full GPU | up to 8k | — |
| DeepSeek R1 Distill Llama 8B | Runs well | est. 29.9 tok/s26.3–33.5calibrated estimate ±12% | Q4_K_M | 7.2 / 8.0 GB | Full GPU | up to 8k |
|
| Qwen3.5 35B-A3B | Runs well | est. 20.4 tok/s14.3–26.5theoretical estimate ±30% | Q4_K_M | 2.6 / 8.0 GB + 21.4 GB RAM | MoE experts in RAM | up to 128k |
|
| Kanana 1.5 15.7B-A3B | Runs slowly | est. 19.3 tok/s13.5–25.1theoretical estimate ±30% | Q4_K_M | 3.0 / 8.0 GB + 9.6 GB RAM | MoE experts in RAM | up to 32k |
|
| Qwen3 30B-A3B (2507) | Runs slowly | est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | Q4_K_M | 2.9 / 8.0 GB + 17.7 GB RAM | MoE experts in RAM | up to 32k |
|
| gpt-oss-20b | Runs slowly | est. 18.5 tok/s12.9–24.0theoretical estimate ±30% | MXFP4 | 2.0 / 8.0 GB + 11.5 GB RAM | MoE experts in RAM | up to 128k |
|
| Gemma 4 26B-A4B | Runs slowly | est. 17.9 tok/s12.5–23.3theoretical estimate ±30% | Q4_K_M | 8.0 / 8.0 GB + 10.5 GB RAM | Partial offload | up to 128k |
|
| Gemma 4 12B | Runs slowly | est. 17.3 tok/s13.9–20.8calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 0.8 GB RAM | Partial offload | up to 64k |
|
| Gemma 3 12B | Runs slowly | est. 16.2 tok/s12.9–19.4calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 1.0 GB RAM | Partial offload | up to 64k |
|
| HyperCLOVA X SEED Think 14B | Runs slowly | est. 8.2 tok/s5.8–10.7theoretical estimate ±30% | Q4_K_M | 8.0 / 8.0 GB + 3.4 GB RAM | Partial offload | up to 32k |
|
| Qwen3 14B | Runs slowly | est. 8.0 tok/s6.4–9.6calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 3.5 GB RAM | Partial offload | up to 32k |
|
| Phi-4 | Runs slowly | est. 7.3 tok/s5.8–8.7calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 3.9 GB RAM | Partial offload | up to 16k |
|
| Mistral Small 3.2 24B | Runs slowly | est. 4.0 tok/s3.2–4.8calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 8.8 GB RAM | Partial offload | up to 32k | — |
| Gemma 3 27B | Runs slowly | est. 3.6 tok/s2.9–4.3calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 10.4 GB RAM | Partial offload | up to 64k | — |
| Qwen3.5 27B | Runs slowly | est. 3.5 tok/s2.8–4.2calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 10.8 GB RAM | Partial offload | up to 64k | — |
| Qwen3 32B | Runs slowly | est. 2.5 tok/s2.0–3.0calibrated estimate ±20% | Q4_K_M | 8.0 / 8.0 GB + 15.1 GB RAM | Partial offload | up to 16k | — |
| DeepSeek R1 Distill Qwen 32B | Won't runTry Q2_K (heavy quality loss): Runs slowly | est. 4.3 tok/s3.5–5.2calibrated estimate ±20% | Q2_K | 8.0 / 8.0 GB + 7.6 GB RAM | Partial offload | up to 16k |
|
| EXAONE 4.0 32B | Won't runTry IQ4_XS: Runs slowly | est. 3.4 tok/s2.7–4.1calibrated estimate ±20% | IQ4_XS | 8.0 / 8.0 GB + 11.1 GB RAM | Partial offload | up to 16k |
|
| EXAONE 4.5 33B | Won't runTry IQ4_XS: Runs slowly | est. 3.3 tok/s2.6–3.9calibrated estimate ±20% | IQ4_XS | 8.0 / 8.0 GB + 11.7 GB RAM | Partial offload | up to 16k |
|
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Llama 3.3 70B | Won't run | — | Q4_K_M | needs 45.8 GB | — | — |
|
| Qwen3.5 122B-A10B | Won't run | — | Q4_K_M | needs 79.0 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Measured results on the GeForce RTX 4060 Laptop
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the GeForce RTX 4060 Laptop?
Qwen3.5 9B at Q4_K_M (a 5.87 GB file) fits entirely in 8 GB with 8k context, at est. 29.2 tok/s (25.7–32.7, calibrated estimate ±12%).
How many local LLMs run well on the GeForce RTX 4060 Laptop?
At Q4_K_M with 8k context, out of 29 tracked models: 6 run great, 2 run well, 13 run slowly and 8 won't run.
Can the GeForce RTX 4060 Laptop run a 70B model like Llama 3.3 70B?
No. Llama 3.3 70B at Q4_K_M needs about 45.8 GB, while this setup offers 7.4 GB of GPU memory and 28.0 GB of free system RAM.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.