Can I run gpt-oss-20b on the GeForce RTX 5070?

Runs slowlyMXFP4 keeps attention on the GPU (2.0 GB) and its experts in 11.5 GB of system RAM: Runs slowly, est. 18.5 tok/s (12.9–24.0, theoretical estimate ±30%).

12 GB VRAM · 672.0 GB/s · FP16 30.9 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache. The GeForce RTX 5070 Ti Laptop has the same memory setup (12 GB, 672.0 GB/s), so its generation-speed estimate is the same; only prompt processing differs.

Every quant of gpt-oss-20b on the GeForce RTX 5070

1 tracked GGUF file, smallest first, at 8k context
QuantFile sizeVerdictSpeedMemoryRuns asNotes
MXFP4baseline12.11 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
2.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMExperts run from system RAM — Runs well needs 30 tok/s or more on this path

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Where the memory goes at MXFP4

Weights, KV cache and compute buffer are the model; the OS reservation is the display driver.
GPU memory
System RAM
Weights
0.7 GB
KV cache
0.2 GB
Compute buffer
0.6 GB
OS reserve
0.6 GB
Free
10.0 GB
Weights in system RAM
11.5 GB

Context length vs. KV cache

KV cache

Memory needed and verdict at MXFP4 for each context length and KV cache type, with estimated tokens per second.
ContextKV f16KV q8_0KV q4_0
4k
12.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
12.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
12.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
8k
12.9 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
12.8 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
12.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
16k
13.2 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
13.0 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
12.9 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
32k
13.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
13.3 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
13.1 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
64k
14.8 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
14.1 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
13.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
128k
17.0 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
15.5 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
14.7 GBRuns slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At MXFP4 the verdict stays Runs slowly up to 128k context with an f16 KV cache, and up to 128k with q8_0.

Estimated speed

Generation
est. 18.5 tok/s (12.9–24.0, theoretical estimate ±30%)
Prompt processing
Prompt-processing estimates only apply when the whole model runs on the GPU.

Why this verdict: Experts run from system RAM — Runs well needs 30 tok/s or more on this path

Measured on this exact combination

No public measurement for this exact combination yet; the numbers above are estimates.

If this is not enough

How to run it

Commands for the MXFP4 file. Both tools download from Hugging Face on first run.

File: gpt-oss-20b-MXFP4.gguf (12.11 GB) from ggml-org/gpt-oss-20b-GGUF.

llama.cpp
llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 8192 -ngl all --n-cpu-moe 24 -fa on
  • -ngl all loads every layer on the GPU.
  • --n-cpu-moe 24 keeps the experts of all 24 layers in system RAM, which is what this estimate assumes.
  • With 5 the model still fits; fewer layers on the CPU means faster generation.
  • -fa on enables flash attention (needed for KV cache quantization).
  • --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.11 GB at 8k context.
Ollama

Ollama has no expert-offload flag, so this mode is llama.cpp only.

Related pages

Frequently asked questions

Can I run gpt-oss-20b on the GeForce RTX 5070?

MXFP4 keeps attention on the GPU (2.0 GB) and its experts in 11.5 GB of system RAM: Runs slowly, est. 18.5 tok/s (12.9–24.0, theoretical estimate ±30%). MXFP4 needs 12.9 GB at 8k context; this setup has 11.4 GB of usable memory and 28.0 GB of free system RAM.

How much context can gpt-oss-20b use on the GeForce RTX 5070?

At MXFP4 the verdict stays Runs slowly up to 128k context with an f16 KV cache (3.2 GB of KV) and up to 128k with a q8_0 KV cache (1.7 GB). The model supports up to 128k.

Which quant should I use, and how fast is it?

MXFP4 (12.11 GB) is the only tracked file: Runs slowly, est. 18.5 tok/s (12.9–24.0, theoretical estimate ±30%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.