Can I run Gemma 3 12B on the GeForce RTX 5090?
Runs greatQ4_K_M fits in VRAM with 23.0 GB to spare: Runs great, est. 160.1 tok/s (140.9–179.3, calibrated estimate ±12%).
32 GB VRAM · 1,792.0 GB/s · FP16 104.8 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Every quant of Gemma 3 12B on the GeForce RTX 5090
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 4.77 GB | Runs great | est. 236.4 tok/s208.0–264.7calibrated estimate ±12% | 6.5 / 32.0 GB | Full GPU | — |
| IQ4_XS | 6.55 GB | Runs great | est. 177.0 tok/s155.8–198.2calibrated estimate ±12% | 8.3 / 32.0 GB | Full GPU | — |
| Q4_K_Mbaseline | 7.30 GB | Runs great | est. 160.1 tok/s140.9–179.3calibrated estimate ±12% | 9.0 / 32.0 GB | Full GPU | — |
| Q8_0 | 12.51 GB | Runs great | est. 96.1 tok/s84.6–107.7calibrated estimate ±12% | 14.2 / 32.0 GB | Full GPU | — |
Where the memory goes at Q4_K_M
- Weights
- 7.3 GB
- KV cache
- 0.5 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 23.0 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 8.1 GBRuns great est. 165.7 tok/s145.9–185.6calibrated estimate ±12% | 8.0 GBRuns great est. 168.5 tok/s148.3–188.7calibrated estimate ±12% | 7.9 GBRuns great est. 170.1 tok/s149.6–190.5calibrated estimate ±12% |
| 8k | 8.4 GBRuns great est. 160.1 tok/s140.9–179.3calibrated estimate ±12% | 8.2 GBRuns great est. 165.3 tok/s145.5–185.2calibrated estimate ±12% | 8.0 GBRuns great est. 168.3 tok/s148.1–188.5calibrated estimate ±12% |
| 16k | 9.0 GBRuns great est. 149.8 tok/s131.8–167.8calibrated estimate ±12% | 8.5 GBRuns great est. 159.3 tok/s140.2–178.4calibrated estimate ±12% | 8.3 GBRuns great est. 164.9 tok/s145.1–184.7calibrated estimate ±12% |
| 32k | 10.2 GBRuns great est. 132.8 tok/s116.8–148.7calibrated estimate ±12% | 9.2 GBRuns great est. 148.5 tok/s130.7–166.3calibrated estimate ±12% | 8.7 GBRuns great est. 158.5 tok/s139.5–177.6calibrated estimate ±12% |
| 64k | 12.7 GBRuns great est. 108.2 tok/s95.2–121.2calibrated estimate ±12% | 10.7 GBRuns great est. 130.7 tok/s115.0–146.4calibrated estimate ±12% | 9.6 GBRuns great est. 147.2 tok/s129.5–164.8calibrated estimate ±12% |
| 128k | 17.6 GBRuns great est. 78.9 tok/s69.5–88.4calibrated estimate ±12% | 13.6 GBRuns great est. 105.5 tok/s92.8–118.1calibrated estimate ±12% | 11.4 GBRuns great est. 128.7 tok/s113.2–144.1calibrated estimate ±12% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs great up to 128k context with an f16 KV cache, and up to 128k with q8_0.
Estimated speed
- Generation
- est. 160.1 tok/s (140.9–179.3, calibrated estimate ±12%)
- Prompt processing
- about 14,672 tok/s (8,803–20,541, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A larger model that also runs well on the GeForce RTX 5090
Runs greatQwen3.5 35B-A3B: Runs great, est. 235.0 tok/s (188.0–282.0, calibrated estimate ±20%)
A larger quant that still runs great
Runs greatQ8_0 (12.51 GB): Runs great, est. 96.1 tok/s (84.6–107.7, calibrated estimate ±12%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: google_gemma-3-12b-it-Q4_K_M.gguf (7.30 GB) from bartowski/google_gemma-3-12b-it-GGUF.
llama-server -hf bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.29 GB at 8k context.
PowerShell (quit the Ollama tray app first):
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 5090
Gemma 3 12B on other hardware
Frequently asked questions
Can I run Gemma 3 12B on the GeForce RTX 5090?
Q4_K_M fits in VRAM with 23.0 GB to spare: Runs great, est. 160.1 tok/s (140.9–179.3, calibrated estimate ±12%). Q4_K_M needs 8.4 GB at 8k context; this setup has 31.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can Gemma 3 12B use on the GeForce RTX 5090?
At Q4_K_M the verdict stays Runs great up to 128k context with an f16 KV cache (8.6 GB of KV) and up to 128k with a q8_0 KV cache (4.6 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
Q4_K_M (7.30 GB) is the recommended balance: Runs great, est. 160.1 tok/s (140.9–179.3, calibrated estimate ±12%). The largest tracked file that still gets Runs great is Q8_0 (12.51 GB), est. 96.1 tok/s (84.6–107.7, calibrated estimate ±12%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.