Can I run gpt-oss-20b on the GeForce RTX 3090?

Runs greatMXFP4 fits in VRAM with 10.5 GB to spare: Runs great, 162.0 tok/s measured at 2k context (1 run; 8k estimate 184.2 tok/s, 147.3–221.0, calibrated estimate ±20%).

24 GB VRAM · 936.0 GB/s · FP16 35.6 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.

Every quant of gpt-oss-20b on the GeForce RTX 3090

1 tracked GGUF file, smallest first, at 8k context
QuantFile sizeVerdictSpeedMemoryRuns asNotes
MXFP4baseline12.11 GBRuns great
162.0 tok/s8k estimate 184.2 tok/s (147.3–221.0, calibrated estimate ±20%)measured (1 run, 2k context)
13.5 / 24.0 GBFull GPU—

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Where the memory goes at MXFP4

Weights, KV cache and compute buffer are the model; the OS reservation is the display driver.
GPU memory
Weights
12.1 GB
KV cache
0.2 GB
Compute buffer
0.6 GB
OS reserve
0.6 GB
Free
10.5 GB

Context length vs. KV cache

KV cache

Memory needed and verdict at MXFP4 for each context length and KV cache type, with estimated tokens per second.
ContextKV f16KV q8_0KV q4_0
4k
12.7 GBRuns great
est. 192.6 tok/s154.1–231.2calibrated estimate ±20%
12.7 GBRuns great
est. 196.8 tok/s157.5–236.2calibrated estimate ±20%
12.7 GBRuns great
est. 199.2 tok/s159.3–239.0calibrated estimate ±20%
8k
12.9 GBRuns great
162.0 tok/s8k estimate 184.2 tok/s (147.3–221.0, calibrated estimate ±20%)measured (1 run, 2k context)
12.8 GBRuns great
est. 192.0 tok/s153.6–230.4calibrated estimate ±20%
12.7 GBRuns great
est. 196.5 tok/s157.2–235.8calibrated estimate ±20%
16k
13.2 GBRuns great
est. 169.3 tok/s135.4–203.1calibrated estimate ±20%
13.0 GBRuns great
est. 183.0 tok/s146.4–219.6calibrated estimate ±20%
12.9 GBRuns great
est. 191.4 tok/s153.1–229.7calibrated estimate ±20%
32k
13.7 GBRuns great
est. 145.7 tok/s116.5–174.8calibrated estimate ±20%
13.3 GBRuns great
est. 167.4 tok/s133.9–200.8calibrated estimate ±20%
13.1 GBRuns great
est. 181.9 tok/s145.5–218.3calibrated estimate ±20%
64k
14.8 GBRuns great
est. 113.9 tok/s91.2–136.7calibrated estimate ±20%
14.1 GBRuns great
est. 142.9 tok/s114.3–171.5calibrated estimate ±20%
13.7 GBRuns great
est. 165.5 tok/s132.4–198.6calibrated estimate ±20%
128k
17.0 GBRuns great
est. 79.4 tok/s63.5–95.2calibrated estimate ±20%
15.5 GBRuns great
est. 110.6 tok/s88.5–132.7calibrated estimate ±20%
14.7 GBRuns great
est. 140.2 tok/s112.2–168.3calibrated estimate ±20%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

At MXFP4 the verdict stays Runs great up to 128k context with an f16 KV cache, and up to 128k with q8_0.

Estimated speed

Generation
162.0 tok/s measured at 2k context (1 run; 8k estimate 184.2 tok/s, 147.3–221.0, calibrated estimate ±20%)
Prompt processing
about 4,984 tok/s (2,990–6,978, rough estimate ±40%)

Measured on this exact combination

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured on this exact combination
HardwareQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
GeForce RTX 3090MXFP4llama.cpp2k5,143.0162.0—github.com2025-08-15

If this is not enough

How to run it

Commands for the MXFP4 file. Both tools download from Hugging Face on first run.

File: gpt-oss-20b-MXFP4.gguf (12.11 GB) from ggml-org/gpt-oss-20b-GGUF.

llama.cpp
llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 8192 -ngl all -fa on
  • -ngl all loads every layer on the GPU.
  • -fa on enables flash attention (needed for KV cache quantization).
  • --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.11 GB at 8k context.
Ollama

PowerShell (quit the Ollama tray app first):

$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/ggml-org/gpt-oss-20b-GGUF:MXFP4
  • Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
  • /set parameter num_ctx 8192

Related pages

Frequently asked questions

Can I run gpt-oss-20b on the GeForce RTX 3090?

MXFP4 fits in VRAM with 10.5 GB to spare: Runs great, 162.0 tok/s measured at 2k context (1 run; 8k estimate 184.2 tok/s, 147.3–221.0, calibrated estimate ±20%). MXFP4 needs 12.9 GB at 8k context; this setup has 23.4 GB of usable memory and 28.0 GB of free system RAM.

How much context can gpt-oss-20b use on the GeForce RTX 3090?

At MXFP4 the verdict stays Runs great up to 128k context with an f16 KV cache (3.2 GB of KV) and up to 128k with a q8_0 KV cache (1.7 GB). The model supports up to 128k.

Which quant should I use, and how fast is it?

MXFP4 (12.11 GB) is the only tracked file: Runs great, 162.0 tok/s measured at 2k context (1 run; 8k estimate 184.2 tok/s, 147.3–221.0, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.