Can I run Llama 3.1 8B on the GeForce RTX 4060?
Runs greatQ4_K_M fits in VRAM with 0.8 GB to spare: Runs great, 38.1 tok/s measured at 4k context (1 run; 8k estimate 31.8 tok/s, 28.0–35.6, calibrated estimate ±12%).
8 GB VRAM · 272.0 GB/s · FP16 15.1 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Every quant of Llama 3.1 8B on the GeForce RTX 4060
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 3.18 GB | Runs great | est. 44.8 tok/s39.4–50.1calibrated estimate ±12% | 5.4 / 8.0 GB | Full GPU | — |
| IQ4_XS | 4.45 GB | Runs great | est. 34.5 tok/s30.3–38.6calibrated estimate ±12% | 6.7 / 8.0 GB | Full GPU | — |
| Q4_K_Mbaseline | 4.92 GB | Runs great | 38.1 tok/s8k estimate 31.8 tok/s (28.0–35.6, calibrated estimate ±12%)measured (1 run, 4k context) | 7.2 / 8.0 GB | Full GPU | — |
| Q8_0 | 8.54 GB | Runs slowly | est. 9.6 tok/s7.7–11.5calibrated estimate ±20% | 8.0 / 8.0 GB + 2.8 GB RAM | Partial offload | Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly |
Where the memory goes at Q4_K_M
- Weights
- 4.9 GB
- KV cache
- 1.1 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 0.8 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 6.0 GBRuns great est. 34.9 tok/s30.7–39.1calibrated estimate ±12% | 5.7 GBRuns great est. 36.6 tok/s32.2–41.0calibrated estimate ±12% | 5.6 GBRuns great est. 37.5 tok/s33.0–42.0calibrated estimate ±12% |
| 8k | 6.6 GBRuns great 38.1 tok/s8k estimate 31.8 tok/s (28.0–35.6, calibrated estimate ±12%)measured (1 run, 4k context) | 6.1 GBRuns great est. 34.7 tok/s30.5–38.8calibrated estimate ±12% | 5.8 GBRuns great est. 36.4 tok/s32.1–40.8calibrated estimate ±12% |
| 16k | 7.7 GBRuns well est. 22.3 tok/s17.8–26.7calibrated estimate ±20% | 6.7 GBRuns great est. 31.4 tok/s27.6–35.1calibrated estimate ±12% | 6.2 GBRuns great est. 34.4 tok/s30.3–38.5calibrated estimate ±12% |
| 32k | 10.0 GBRuns slowly est. 7.6 tok/s6.1–9.1calibrated estimate ±20% | 8.0 GBRuns slowly est. 18.7 tok/s15.0–22.5calibrated estimate ±20% | 6.9 GBRuns great est. 31.0 tok/s27.3–34.7calibrated estimate ±12% |
| 64k | 13.5 GBRuns slowly est. 3.3 tok/s2.7–4.0calibrated estimate ±20% | 10.6 GBRuns slowly est. 6.4 tok/s5.1–7.7calibrated estimate ±20% | 8.5 GBRuns slowly est. 15.2 tok/s12.1–18.2calibrated estimate ±20% |
| 128k | 22.1 GBRuns slowly est. 2.0 tok/s1.6–2.4calibrated estimate ±20% | 14.1 GBRuns slowly est. 3.2 tok/s2.5–3.8calibrated estimate ±20% | 11.5 GBRuns slowly est. 5.2 tok/s4.2–6.3calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs great up to 8k context with an f16 KV cache, and up to 16k with q8_0.
Estimated speed
- Generation
- 38.1 tok/s measured at 4k context (1 run; 8k estimate 31.8 tok/s, 28.0–35.6, calibrated estimate ±12%)
- Prompt processing
- about 2,114 tok/s (1,268–2,960, rough estimate ±40%)
Measured on this exact combination
Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.
| Hardware | Quant | Backend | Context | Prompt (tok/s) | Generation (tok/s) | Flags | Source | Measured |
|---|---|---|---|---|---|---|---|---|
| GeForce RTX 4060 | Q4_K_M | llama.cpp | 4k | 1,528.0 | 38.1 | ctx-weighted | localscore.ai | 2026-08-09 |
If this is not enough
A larger model that also runs well on the GeForce RTX 4060
Runs wellQwen3.5 35B-A3B: Runs well, est. 20.4 tok/s (14.3–26.5, theoretical estimate ±30%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf (4.92 GB) from bartowski/Meta-Llama-3.1-8B-Instruct-GGUF.
llama-server -hf bartowski/Meta-Llama-3.1-8B-Instruct-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.57 GB at 8k context.
PowerShell (quit the Ollama tray app first):
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 4060
Llama 3.1 8B on other hardware
Frequently asked questions
Can I run Llama 3.1 8B on the GeForce RTX 4060?
Q4_K_M fits in VRAM with 0.8 GB to spare: Runs great, 38.1 tok/s measured at 4k context (1 run; 8k estimate 31.8 tok/s, 28.0–35.6, calibrated estimate ±12%). Q4_K_M needs 6.6 GB at 8k context; this setup has 7.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can Llama 3.1 8B use on the GeForce RTX 4060?
At Q4_K_M the verdict stays Runs great up to 8k context with an f16 KV cache (1.1 GB of KV) and up to 16k with a q8_0 KV cache (1.1 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
Q4_K_M (4.92 GB) is the recommended balance: Runs great, 38.1 tok/s measured at 4k context (1 run; 8k estimate 31.8 tok/s, 28.0–35.6, calibrated estimate ±12%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.