Can I run Qwen3 30B-A3B (2507) on the GeForce RTX 5070?
Runs slowlyQ4_K_M keeps attention on the GPU (2.9 GB) and its experts in 17.7 GB of system RAM: Runs slowly, est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%).
12 GB VRAM · 672.0 GB/s · FP16 30.9 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache. The GeForce RTX 5070 Ti Laptop has the same memory setup (12 GB, 672.0 GB/s), so its generation-speed estimate is the same; only prompt processing differs.
Every quant of Qwen3 30B-A3B (2507) on the GeForce RTX 5070
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| Q2_K | 11.26 GB | Runs well | est. 31.6 tok/s22.1–41.1theoretical estimate ±30% | 2.9 / 12.0 GB + 10.4 GB RAM | MoE experts in RAM | Not fully on the GPU — Runs great needs the whole model in GPU memory |
| IQ4_XS | 16.38 GB | Runs well | est. 21.7 tok/s15.2–28.3theoretical estimate ±30% | 2.9 / 12.0 GB + 15.5 GB RAM | MoE experts in RAM | Not fully on the GPU — Runs great needs the whole model in GPU memory |
| Q4_K_Mbaseline | 18.56 GB | Runs slowly | est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 2.9 / 12.0 GB + 17.7 GB RAM | MoE experts in RAM | Experts run from system RAM — Runs well needs 20 tok/s or more on this path |
| Q8_0 | 32.48 GB | Won't run | — | Needs about 33.9 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free | — | — |
Where the memory goes at Q4_K_M
- Weights
- 0.9 GB
- KV cache
- 0.8 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 9.1 GB
- Weights in system RAM
- 17.7 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 19.5 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 19.3 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 19.2 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
| 8k | 19.9 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 19.6 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 19.4 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
| 16k | 20.8 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 20.1 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 19.7 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
| 32k | 22.6 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 21.1 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 20.3 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
| 64k | 26.1 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 23.1 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 21.5 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
| 128k | 33.1 GBdoes not fit | 27.2 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | 23.9 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
| 256k | 47.2 GBdoes not fit | 35.2 GBdoes not fit | 28.8 GBRuns slowly est. 19.2 tok/s13.4–24.9theoretical estimate ±30% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At Q4_K_M the verdict stays Runs slowly up to 64k context with an f16 KV cache, and up to 128k with q8_0.
Estimated speed
- Generation
- est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%)
- Prompt processing
- Prompt-processing estimates only apply when the whole model runs on the GPU.
Why this verdict: Experts run from system RAM — Runs well needs 20 tok/s or more on this path
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A smaller model that runs well on the GeForce RTX 5070
Runs greatQwen3 14B: Runs great, est. 45.5 tok/s (40.0–50.9, calibrated estimate ±12%)
A lower quant of Qwen3 30B-A3B (2507)
Runs wellIQ4_XS (16.38 GB): Runs well, est. 21.7 tok/s (15.2–28.3, theoretical estimate ±30%)
Keep only the experts in system RAM
--n-cpu-moe 48 puts the experts of every layer in RAM (this estimate); 24 is the smallest value that still fits, and fewer layers on the CPU means faster.
How to run it
Commands for the Q4_K_M file. Both tools download from Hugging Face on first run.
File: Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf (18.56 GB) from unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF.
llama-server -hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M -c 8192 -ngl all --n-cpu-moe 48 -fa on- -ngl all loads every layer on the GPU.
- --n-cpu-moe 48 keeps the experts of all 48 layers in system RAM, which is what this estimate assumes.
- With 24 the model still fits; fewer layers on the CPU means faster generation.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.43 GB at 8k context.
Ollama has no expert-offload flag, so this mode is llama.cpp only.
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 5070
Qwen3 30B-A3B (2507) on other hardware
Frequently asked questions
Can I run Qwen3 30B-A3B (2507) on the GeForce RTX 5070?
Q4_K_M keeps attention on the GPU (2.9 GB) and its experts in 17.7 GB of system RAM: Runs slowly, est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%). Q4_K_M needs 19.9 GB at 8k context; this setup has 11.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can Qwen3 30B-A3B (2507) use on the GeForce RTX 5070?
At Q4_K_M the verdict stays Runs slowly up to 64k context with an f16 KV cache (6.4 GB of KV) and up to 128k with a q8_0 KV cache (6.9 GB). The model supports up to 256k.
Which quant should I use, and how fast is it?
Q4_K_M (18.56 GB) is the recommended balance: Runs slowly, est. 19.2 tok/s (13.4–24.9, theoretical estimate ±30%). It is also the largest tracked file that gets this verdict.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.