Can I run Qwen3 14B on the GeForce RTX 3060 12GB?
Runs greatQ4_K_M fits in VRAM with 0.5 GB to spare: Runs great, est. 24.4 tok/s (21.4–27.3, calibrated estimate ±12%).
12 GB VRAM · 360.0 GB/s · FP16 12.7 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Every quant of Qwen3 14B on the GeForce RTX 3060 12GB
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 5.75 GB | Runs great | est. 35.5 tok/s31.3–39.8calibrated estimate ±12% | 8.3 / 12.0 GB | Full GPU | — |
| IQ4_XS | 8.14 GB | Runs great | est. 26.6 tok/s23.4–29.8calibrated estimate ±12% | 10.7 / 12.0 GB | Full GPU | — |
| Q4_K_Mbaseline | 9.00 GB | Runs great | est. 24.4 tok/s21.4–27.3calibrated estimate ±12% | 11.5 / 12.0 GB | Full GPU | — |
| Q8_0 | 15.70 GB | Runs slowly | est. 5.2 tok/s4.2–6.3calibrated estimate ±20% | 12.0 / 12.0 GB + 6.2 GB RAM | Partial offload | — |
Where the memory goes at Q4_K_M
- Weights
- 9.0 GB
- KV cache
- 1.3 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 0.5 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 10.2 GBRuns great est. 26.1 tok/s22.9–29.2calibrated estimate ±12% | 9.9 GBRuns great est. 26.9 tok/s23.7–30.2calibrated estimate ±12% | 9.7 GBRuns great est. 27.4 tok/s24.1–30.7calibrated estimate ±12% |
| 8k | 10.9 GBRuns great est. 24.4 tok/s21.4–27.3calibrated estimate ±12% | 10.3 GBRuns great est. 25.9 tok/s22.8–29.0calibrated estimate ±12% | 10.0 GBRuns great est. 26.9 tok/s23.6–30.1calibrated estimate ±12% |
| 16k | 12.3 GBRuns slowly est. 14.6 tok/s11.7–17.5calibrated estimate ±20% | 11.1 GBRuns great est. 24.1 tok/s21.2–27.0calibrated estimate ±12% | 10.4 GBRuns great est. 25.8 tok/s22.7–28.9calibrated estimate ±12% |
| 32k | 15.2 GBRuns slowly est. 6.0 tok/s4.8–7.2calibrated estimate ±20% | 12.7 GBRuns slowly est. 12.8 tok/s10.3–15.4calibrated estimate ±20% | 11.3 GBRuns great est. 23.9 tok/s21.1–26.8calibrated estimate ±12% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs great up to 8k context with an f16 KV cache, and up to 16k with q8_0.
Estimated speed
- Generation
- est. 24.4 tok/s (21.4–27.3, calibrated estimate ±12%)
- Prompt processing
- about 1,778 tok/s (1,067–2,489, rough estimate ±40%)
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A larger model that also runs well on the GeForce RTX 3060 12GB
Runs wellQwen3.5 35B-A3B: Runs well, est. 20.4 tok/s (14.3–26.5, theoretical estimate ±30%)
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: Qwen3-14B-Q4_K_M.gguf (9.00 GB) from unsloth/Qwen3-14B-GGUF.
llama-server -hf unsloth/Qwen3-14B-GGUF:Q4_K_M -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.72 GB at 8k context.
PowerShell (quit the Ollama tray app first):
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/unsloth/Qwen3-14B-GGUF:Q4_K_M- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 3060 12GB
Qwen3 14B on other hardware
Frequently asked questions
Can I run Qwen3 14B on the GeForce RTX 3060 12GB?
Q4_K_M fits in VRAM with 0.5 GB to spare: Runs great, est. 24.4 tok/s (21.4–27.3, calibrated estimate ±12%). Q4_K_M needs 10.9 GB at 8k context; this setup has 11.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can Qwen3 14B use on the GeForce RTX 3060 12GB?
At Q4_K_M the verdict stays Runs great up to 8k context with an f16 KV cache (1.3 GB of KV) and up to 16k with a q8_0 KV cache (1.4 GB). The model supports up to 32k.
Which quant should I use, and how fast is it?
Q4_K_M (9.00 GB) is the recommended balance: Runs great, est. 24.4 tok/s (21.4–27.3, calibrated estimate ±12%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.