Can I run Qwen3 30B-A3B (2507) on the GeForce RTX 5060 Ti 16GB?

Runs slowlyQ4_K_M keeps attention on the GPU (2.9 GB) and its experts in 17.7 GB of system RAM: Runs slowly, est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%). Q2_K (heavy quality loss) would reach Runs great, est. 99.6 tok/s (79.7–119.5, calibrated estimate ±20%).

16 GB VRAM · 448.0 GB/s · FP16 23.7 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.

Every quant of Qwen3 30B-A3B (2507) on the GeForce RTX 5060 Ti 16GB

4 tracked GGUF files, smallest first, at 8k context
QuantFile sizeVerdictSpeedMemoryRuns asNotes
Q2_K11.26 GBRuns great
est. 99.6 tok/s79.7–119.5calibrated estimate ±20%
13.2 / 16.0 GBFull GPU—
IQ4_XS16.38 GBRuns well
est. 21.7 tok/s15.2–28.3theoretical estimate ±30%
2.9 / 16.0 GB + 15.5 GB RAMMoE experts in RAMNot fully on the GPU — Runs great needs the whole model in GPU memory
Q4_K_Mbaseline18.56 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
2.9 / 16.0 GB + 17.7 GB RAMMoE experts in RAMExperts run from system RAM — Runs well needs 20 tok/s or more on this path
Q8_032.48 GBWon't run—Needs about 33.9 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free——

Where the memory goes at Q4_K_M

Weights, KV cache and compute buffer are the model; the OS reservation is the display driver.
GPU memory
System RAM
Weights
0.9 GB
KV cache
0.8 GB
Compute buffer
0.6 GB
OS reserve
0.6 GB
Free
13.1 GB
Weights in system RAM
17.7 GB

Context length vs. KV cache

KV cache

Memory needed and verdict at Q4_K_M for each context length and KV cache type, with estimated tokens per second.
ContextKV f16KV q8_0KV q4_0
4k
19.5 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
19.3 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
19.2 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
8k
19.9 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
19.6 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
19.4 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
16k
20.8 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
20.1 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
19.7 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
32k
22.6 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
21.1 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
20.3 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
64k
26.1 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
23.1 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
21.5 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
128k
33.1 GBdoes not fit
27.2 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
23.9 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%
256k
47.2 GBdoes not fit
35.2 GBdoes not fit
28.8 GBRuns slowly
est. 19.2 tok/s13.4–24.9theoretical estimate ±30%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At Q4_K_M the verdict stays Runs slowly up to 64k context with an f16 KV cache, and up to 128k with q8_0.

Estimated speed

Generation
est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%)
Prompt processing
Prompt-processing estimates only apply when the whole model runs on the GPU.

Why this verdict: Experts run from system RAM — Runs well needs 20 tok/s or more on this path

Measured on this exact combination

No public measurement for this exact combination yet; the numbers above are estimates.

If this is not enough

How to run it

Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.

File: Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf (18.56 GB) from unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF.

llama.cpp
llama-server -hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M -c 8192 -ngl all --n-cpu-moe 48 -fa on
  • -ngl all loads every layer on the GPU.
  • --n-cpu-moe 48 keeps the experts of all 48 layers in system RAM, which is what this estimate assumes.
  • With 13 the model still fits; fewer layers on the CPU means faster generation.
  • -fa on enables flash attention (needed for KV cache quantization).
  • --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.43 GB at 8k context.
Ollama

Ollama has no expert-offload flag, so this mode is llama.cpp only.

Related pages

Frequently asked questions

Can I run Qwen3 30B-A3B (2507) on the GeForce RTX 5060 Ti 16GB?

Q4_K_M keeps attention on the GPU (2.9 GB) and its experts in 17.7 GB of system RAM: Runs slowly, est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%). Q2_K (heavy quality loss) would reach Runs great, est. 99.6 tok/s (79.7–119.5, calibrated estimate ±20%). Q4_K_M needs 19.9 GB at 8k context; this setup has 15.4 GB of usable memory and 28.0 GB of free system RAM.

How much context can Qwen3 30B-A3B (2507) use on the GeForce RTX 5060 Ti 16GB?

At Q4_K_M the verdict stays Runs slowly up to 64k context with an f16 KV cache (6.4 GB of KV) and up to 128k with a q8_0 KV cache (6.9 GB). The model supports up to 256k.

Which quant should I use, and how fast is it?

Q4_K_M (18.56 GB) is the recommended balance: Runs slowly, est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%). It is also the largest tracked file that gets this verdict.

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.