Can I run Gemma 3 12B on the GeForce RTX 4060?
Runs slowlyQ4_K_M is 1.0 GB short of fitting in VRAM, so 14% of the weights run from system RAM: Runs slowly, est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%). Q2_K (heavy quality loss) would reach Runs great, est. 35.9 tok/s (31.6–40.2, calibrated estimate ±12%).
8 GB VRAM · 272.0 GB/s · FP16 15.1 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Every quant of Gemma 3 12B on the GeForce RTX 4060
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 4.77 GB | Runs great | est. 35.9 tok/s31.6–40.2calibrated estimate ±12% | 6.5 / 8.0 GB | Full GPU | — |
| IQ4_XS | 6.55 GB | Runs well | est. 23.8 tok/s19.0–28.5calibrated estimate ±20% | 8.0 / 8.0 GB + 0.3 GB RAM | Partial offload | Not fully on the GPU — Runs great needs the whole model in GPU memory |
| Q4_K_Mbaseline | 7.30 GB | Runs slowly | est. 16.7 tok/s13.4–20.1calibrated estimate ±20% | 8.0 / 8.0 GB + 1.0 GB RAM | Partial offload | Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly |
| Q8_0 | 12.51 GB | Runs slowly | est. 5.6 tok/s4.5–6.7calibrated estimate ±20% | 8.0 / 8.0 GB + 6.2 GB RAM | Partial offload | — |
Where the memory goes at Q4_K_M
- Weights
- 6.3 GB
- KV cache
- 0.5 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Weights in system RAM
- 1.0 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 8.1 GBRuns well est. 19.1 tok/s15.3–23.0calibrated estimate ±20% | 8.0 GBRuns well est. 20.3 tok/s16.3–24.4calibrated estimate ±20% | 7.9 GBRuns well est. 21.0 tok/s16.8–25.2calibrated estimate ±20% |
| 8k | 8.4 GBRuns slowly est. 16.7 tok/s13.4–20.1calibrated estimate ±20% | 8.2 GBRuns slowly est. 18.7 tok/s15.0–22.5calibrated estimate ±20% | 8.0 GBRuns well est. 20.0 tok/s16.0–24.0calibrated estimate ±20% |
| 16k | 9.0 GBRuns slowly est. 13.2 tok/s10.6–15.8calibrated estimate ±20% | 8.5 GBRuns slowly est. 16.1 tok/s12.9–19.3calibrated estimate ±20% | 8.3 GBRuns slowly est. 18.1 tok/s14.5–21.8calibrated estimate ±20% |
| 32k | 10.2 GBRuns slowly est. 8.9 tok/s7.1–10.7calibrated estimate ±20% | 9.2 GBRuns slowly est. 12.4 tok/s9.9–14.8calibrated estimate ±20% | 8.7 GBRuns slowly est. 15.2 tok/s12.2–18.2calibrated estimate ±20% |
| 64k | 12.7 GBRuns slowly est. 4.9 tok/s3.9–5.9calibrated estimate ±20% | 10.7 GBRuns slowly est. 8.0 tok/s6.4–9.6calibrated estimate ±20% | 9.6 GBRuns slowly est. 11.2 tok/s9.0–13.5calibrated estimate ±20% |
| 128k | 15.9 GBRuns slowly est. 2.8 tok/s2.3–3.4calibrated estimate ±20% | 13.6 GBRuns slowly est. 4.3 tok/s3.4–5.1calibrated estimate ±20% | 11.4 GBRuns slowly est. 7.0 tok/s5.6–8.4calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs slowly up to 128k context with an f16 KV cache, and up to 128k with q8_0.
Estimated speed
- Generation
- est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%)
- Prompt processing
- Prompt-processing estimates only apply when the whole model runs on the GPU.
Why this verdict: Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A smaller model that runs well on the GeForce RTX 4060
Runs greatQwen3.5 9B: Runs great, est. 31.0 tok/s (27.3–34.7, calibrated estimate ±12%)
A lower quant of Gemma 3 12B
Runs wellIQ4_XS (6.55 GB): Runs well, est. 23.8 tok/s (19.0–28.5, calibrated estimate ±20%)
Cheapest GPU that runs Gemma 3 12B great
Runs greatGeForce RTX 3060 12GB — $250 street (as of 2026-08-09): est. 32.2 tok/s (28.3–36.0, calibrated estimate ±12%).
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: google_gemma-3-12b-it-Q4_K_M.gguf (7.30 GB) from bartowski/google_gemma-3-12b-it-GGUF.
llama-server -hf bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M -c 8192 -ngl 41 -fa on- -ngl 41 puts 41 of 48 layers on the GPU; if VRAM overflows, lower it by 1–2.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.29 GB at 8k context.
PowerShell (quit the Ollama tray app first):
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
- Ollama splits layers between GPU and CPU automatically.
/set parameter num_ctx 8192
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 4060
Gemma 3 12B on other hardware
Frequently asked questions
Can I run Gemma 3 12B on the GeForce RTX 4060?
Q4_K_M is 1.0 GB short of fitting in VRAM, so 14% of the weights run from system RAM: Runs slowly, est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%). Q2_K (heavy quality loss) would reach Runs great, est. 35.9 tok/s (31.6–40.2, calibrated estimate ±12%). Q4_K_M needs 8.4 GB at 8k context; this setup has 7.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can Gemma 3 12B use on the GeForce RTX 4060?
At Q4_K_M the verdict stays Runs slowly up to 128k context with an f16 KV cache (8.6 GB of KV) and up to 128k with a q8_0 KV cache (4.6 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
Q4_K_M (7.30 GB) is the recommended balance: Runs slowly, est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%). The largest tracked file that still gets Runs slowly is Q8_0 (12.51 GB), est. 5.6 tok/s (4.5–6.7, calibrated estimate ±20%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.