Can I run Gemma 3 12B on the GeForce RTX 4060?

Runs slowlyQ4_K_M is 1.0 GB short of fitting in VRAM, so 14% of the weights run from system RAM: Runs slowly, est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%). Q2_K (heavy quality loss) would reach Runs great, est. 35.9 tok/s (31.6–40.2, calibrated estimate ±12%).

8 GB VRAM · 272.0 GB/s · FP16 15.1 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.

Every quant of Gemma 3 12B on the GeForce RTX 4060

4 tracked GGUF files, smallest first, at 8k context
QuantFile sizeVerdictSpeedMemoryRuns asNotes
Q2_K4.77 GBRuns great
est. 35.9 tok/s31.6–40.2calibrated estimate ±12%
6.5 / 8.0 GBFull GPU—
IQ4_XS6.55 GBRuns well
est. 23.8 tok/s19.0–28.5calibrated estimate ±20%
8.0 / 8.0 GB + 0.3 GB RAMPartial offloadNot fully on the GPU — Runs great needs the whole model in GPU memory
Q4_K_Mbaseline7.30 GBRuns slowly
est. 16.7 tok/s13.4–20.1calibrated estimate ±20%
8.0 / 8.0 GB + 1.0 GB RAMPartial offloadLess than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
Q8_012.51 GBRuns slowly
est. 5.6 tok/s4.5–6.7calibrated estimate ±20%
8.0 / 8.0 GB + 6.2 GB RAMPartial offload—

Where the memory goes at Q4_K_M

Weights, KV cache and compute buffer are the model; the OS reservation is the display driver.
GPU memory
System RAM
Weights
6.3 GB
KV cache
0.5 GB
Compute buffer
0.6 GB
OS reserve
0.6 GB
Weights in system RAM
1.0 GB

Context length vs. KV cache

KV cache

Memory needed and verdict at Q4_K_M for each context length and KV cache type, with estimated tokens per second.
ContextKV f16KV q8_0KV q4_0
4k
8.1 GBRuns well
est. 19.1 tok/s15.3–23.0calibrated estimate ±20%
8.0 GBRuns well
est. 20.3 tok/s16.3–24.4calibrated estimate ±20%
7.9 GBRuns well
est. 21.0 tok/s16.8–25.2calibrated estimate ±20%
8k
8.4 GBRuns slowly
est. 16.7 tok/s13.4–20.1calibrated estimate ±20%
8.2 GBRuns slowly
est. 18.7 tok/s15.0–22.5calibrated estimate ±20%
8.0 GBRuns well
est. 20.0 tok/s16.0–24.0calibrated estimate ±20%
16k
9.0 GBRuns slowly
est. 13.2 tok/s10.6–15.8calibrated estimate ±20%
8.5 GBRuns slowly
est. 16.1 tok/s12.9–19.3calibrated estimate ±20%
8.3 GBRuns slowly
est. 18.1 tok/s14.5–21.8calibrated estimate ±20%
32k
10.2 GBRuns slowly
est. 8.9 tok/s7.1–10.7calibrated estimate ±20%
9.2 GBRuns slowly
est. 12.4 tok/s9.9–14.8calibrated estimate ±20%
8.7 GBRuns slowly
est. 15.2 tok/s12.2–18.2calibrated estimate ±20%
64k
12.7 GBRuns slowly
est. 4.9 tok/s3.9–5.9calibrated estimate ±20%
10.7 GBRuns slowly
est. 8.0 tok/s6.4–9.6calibrated estimate ±20%
9.6 GBRuns slowly
est. 11.2 tok/s9.0–13.5calibrated estimate ±20%
128k
15.9 GBRuns slowly
est. 2.8 tok/s2.3–3.4calibrated estimate ±20%
13.6 GBRuns slowly
est. 4.3 tok/s3.4–5.1calibrated estimate ±20%
11.4 GBRuns slowly
est. 7.0 tok/s5.6–8.4calibrated estimate ±20%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At Q4_K_M the verdict stays Runs slowly up to 128k context with an f16 KV cache, and up to 128k with q8_0.

Estimated speed

Generation
est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%)
Prompt processing
Prompt-processing estimates only apply when the whole model runs on the GPU.

Why this verdict: Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly

Measured on this exact combination

No public measurement for this exact combination yet; the numbers above are estimates.

If this is not enough

How to run it

Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.

File: google_gemma-3-12b-it-Q4_K_M.gguf (7.30 GB) from bartowski/google_gemma-3-12b-it-GGUF.

llama.cpp
llama-server -hf bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M -c 8192 -ngl 41 -fa on
  • -ngl 41 puts 41 of 48 layers on the GPU; if VRAM overflows, lower it by 1–2.
  • -fa on enables flash attention (needed for KV cache quantization).
  • --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.29 GB at 8k context.
Ollama

PowerShell (quit the Ollama tray app first):

$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/bartowski/google_gemma-3-12b-it-GGUF:Q4_K_M
  • Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
  • Ollama splits layers between GPU and CPU automatically.
  • /set parameter num_ctx 8192

Related pages

Frequently asked questions

Can I run Gemma 3 12B on the GeForce RTX 4060?

Q4_K_M is 1.0 GB short of fitting in VRAM, so 14% of the weights run from system RAM: Runs slowly, est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%). Q2_K (heavy quality loss) would reach Runs great, est. 35.9 tok/s (31.6–40.2, calibrated estimate ±12%). Q4_K_M needs 8.4 GB at 8k context; this setup has 7.4 GB of usable memory and 28.0 GB of free system RAM.

How much context can Gemma 3 12B use on the GeForce RTX 4060?

At Q4_K_M the verdict stays Runs slowly up to 128k context with an f16 KV cache (8.6 GB of KV) and up to 128k with a q8_0 KV cache (4.6 GB). The model supports up to 128k.

Which quant should I use, and how fast is it?

Q4_K_M (7.30 GB) is the recommended balance: Runs slowly, est. 16.7 tok/s (13.4–20.1, calibrated estimate ±20%). The largest tracked file that still gets Runs slowly is Q8_0 (12.51 GB), est. 5.6 tok/s (4.5–6.7, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.