Can I run Gemma 3 12B on the GeForce RTX 5070?
Runs greatQ4_K_M fits in VRAM with 3.0 GB to spare: Runs great, est. 60.0 tok/s (52.8–67.2, calibrated estimate ±12%).
12 GB VRAM · 672.0 GB/s · FP16 30.9 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache. The GeForce RTX 5070 Ti Laptop has the same memory setup (12 GB, 672.0 GB/s), so its generation-speed estimate is the same; only prompt processing differs. Prompt processing: about 4,326 tok/s (2,596–6,056, rough estimate ±40%) here vs about 4,200 tok/s (2,520–5,880, rough estimate ±40%) on the GeForce RTX 5070 Ti Laptop.
Every quant of Gemma 3 12B on the GeForce RTX 5070
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 4.77 GB | Runs great | est. 88.6 tok/s78.0–99.3calibrated estimate ±12% | 6.5 / 12.0 GB | Full GPU | — |
| IQ4_XS | 6.55 GB | Runs great | est. 66.4 tok/s58.4–74.3calibrated estimate ±12% | 8.3 / 12.0 GB | Full GPU | — |
| Q4_K_Mbaseline | 7.30 GB | Runs great | est. 60.0 tok/s52.8–67.2calibrated estimate ±12% | 9.0 / 12.0 GB | Full GPU | — |
| Q8_0 | 12.51 GB | Runs slowly | est. 13.4 tok/s10.7–16.1calibrated estimate ±20% | 12.0 / 12.0 GB + 2.2 GB RAM | Partial offload | Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly |
Where the memory goes at Q4_K_M
- Weights
- 7.3 GB
- KV cache
- 0.5 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 3.0 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 8.1 GBRuns great est. 62.2 tok/s54.7–69.6calibrated estimate ±12% | 8.0 GBRuns great est. 63.2 tok/s55.6–70.8calibrated estimate ±12% | 7.9 GBRuns great est. 63.8 tok/s56.1–71.4calibrated estimate ±12% |
| 8k | 8.4 GBRuns great est. 60.0 tok/s52.8–67.2calibrated estimate ±12% | 8.2 GBRuns great est. 62.0 tok/s54.6–69.4calibrated estimate ±12% | 8.0 GBRuns great est. 63.1 tok/s55.5–70.7calibrated estimate ±12% |
| 16k | 9.0 GBRuns great est. 56.2 tok/s49.4–62.9calibrated estimate ±12% | 8.5 GBRuns great est. 59.7 tok/s52.6–66.9calibrated estimate ±12% | 8.3 GBRuns great est. 61.8 tok/s54.4–69.3calibrated estimate ±12% |
| 32k | 10.2 GBRuns great est. 49.8 tok/s43.8–55.8calibrated estimate ±12% | 9.2 GBRuns great est. 55.7 tok/s49.0–62.4calibrated estimate ±12% | 8.7 GBRuns great est. 59.5 tok/s52.3–66.6calibrated estimate ±12% |
| 64k | 12.7 GBRuns slowly est. 15.1 tok/s12.1–18.1calibrated estimate ±20% | 10.7 GBRuns great est. 49.0 tok/s43.1–54.9calibrated estimate ±12% | 9.6 GBRuns great est. 55.2 tok/s48.6–61.8calibrated estimate ±12% |
| 128k | 17.6 GBRuns slowly est. 3.3 tok/s2.6–3.9calibrated estimate ±20% | 13.6 GBRuns slowly est. 10.3 tok/s8.2–12.3calibrated estimate ±20% | 11.4 GBRuns well est. 45.4 tok/s36.3–54.5calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs great up to 32k context with an f16 KV cache, and up to 64k with q8_0.
Estimated speed
- Generation
- est. 60.0 tok/s (52.8–67.2, calibrated estimate ±12%)
- Prompt processing
- about 4,326 tok/s (2,596–6,056, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A larger model that also runs well on the GeForce RTX 5070
Runs wellQwen3.5 35B-A3B: Runs well, est. 20.4 tok/s (14.3–26.5, theoretical estimate ±30%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: google_gemma-3-12b-it-Q4_K_M.gguf (7.30 GB) from bartowski/google_gemma-3-12b-it-GGUF.
llama-server -hf bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.29 GB at 8k context.
PowerShell (quit the Ollama tray app first):
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 5070
Gemma 3 12B on other hardware
Frequently asked questions
Can I run Gemma 3 12B on the GeForce RTX 5070?
Q4_K_M fits in VRAM with 3.0 GB to spare: Runs great, est. 60.0 tok/s (52.8–67.2, calibrated estimate ±12%). Q4_K_M needs 8.4 GB at 8k context; this setup has 11.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can Gemma 3 12B use on the GeForce RTX 5070?
At Q4_K_M the verdict stays Runs great up to 32k context with an f16 KV cache (2.1 GB of KV) and up to 64k with a q8_0 KV cache (2.3 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
Q4_K_M (7.30 GB) is the recommended balance: Runs great, est. 60.0 tok/s (52.8–67.2, calibrated estimate ±12%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.