Can I run gpt-oss-20b on the GeForce RTX 5090?
Runs greatMXFP4 fits in VRAM with 18.5 GB to spare: Runs great, 282.3 tok/s measured at 2k context (1 run; 8k estimate 235.0 tok/s, 188.0–282.0, calibrated estimate ±20%).
32 GB VRAM · 1,792.0 GB/s · FP16 104.8 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Every quant of gpt-oss-20b on the GeForce RTX 5090
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| MXFP4baseline | 12.11 GB | Runs great | 282.3 tok/s8k estimate 235.0 tok/s (188.0–282.0, calibrated estimate ±20%)measured (1 run, 2k context) | 13.5 / 32.0 GB | Full GPU | — |
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Where the memory goes at MXFP4
- Weights
- 12.1 GB
- KV cache
- 0.2 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 18.5 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 12.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 12.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 12.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% |
| 8k | 12.9 GBRuns great 282.3 tok/s8k estimate 235.0 tok/s (188.0–282.0, calibrated estimate ±20%)measured (1 run, 2k context) | 12.8 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 12.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% |
| 16k | 13.2 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 13.0 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 12.9 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% |
| 32k | 13.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 13.3 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 13.1 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% |
| 64k | 14.8 GBRuns great est. 218.1 tok/s174.5–261.8calibrated estimate ±20% | 14.1 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% | 13.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% |
| 128k | 17.0 GBRuns great est. 151.9 tok/s121.6–182.3calibrated estimate ±20% | 15.5 GBRuns great est. 211.7 tok/s169.4–254.0calibrated estimate ±20% | 14.7 GBRuns great est. 235.0 tok/s188.0–282.0calibrated estimate ±20% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At MXFP4 the verdict stays Runs great up to 128k context with an f16 KV cache, and up to 128k with q8_0.
Estimated speed
- Generation
- 282.3 tok/s measured at 2k context (1 run; 8k estimate 235.0 tok/s, 188.0–282.0, calibrated estimate ±20%)
- Prompt processing
- about 14,672 tok/s (8,803–20,541, rough estimate ±40%)
Measured on this exact combination
Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.
| Hardware | Quant | Backend | Context | Prompt (tok/s) | Generation (tok/s) | Flags | Source | Measured |
|---|---|---|---|---|---|---|---|---|
| GeForce RTX 5090 | MXFP4 | llama.cpp | 2k | 9,841.0 | 282.3 | — | github.com | 2025-08-15 |
If this is not enough
A larger model that also runs well on the GeForce RTX 5090
Runs greatQwen3.5 35B-A3B: Runs great, est. 235.0 tok/s (188.0–282.0, calibrated estimate ±20%)
How to run it
Commands for the MXFP4 file. Both tools download from Hugging Face on first run.
File: gpt-oss-20b-MXFP4.gguf (12.11 GB) from ggml-org/gpt-oss-20b-GGUF.
llama-server -hf ggml-org/gpt-oss-20b-GGUF:MXFP4 -c 8192 -ngl all -fa on- -ngl all loads every layer on the GPU.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.11 GB at 8k context.
PowerShell (quit the Ollama tray app first):
$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/ggml-org/gpt-oss-20b-GGUF:MXFP4- Ollama defaults to 4k context on GPUs under 24 GB. Set OLLAMA_CONTEXT_LENGTH when starting the server, or /set parameter num_ctx inside the chat.
/set parameter num_ctx 8192
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 5090
gpt-oss-20b on other hardware
Frequently asked questions
Can I run gpt-oss-20b on the GeForce RTX 5090?
MXFP4 fits in VRAM with 18.5 GB to spare: Runs great, 282.3 tok/s measured at 2k context (1 run; 8k estimate 235.0 tok/s, 188.0–282.0, calibrated estimate ±20%). MXFP4 needs 12.9 GB at 8k context; this setup has 31.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can gpt-oss-20b use on the GeForce RTX 5090?
At MXFP4 the verdict stays Runs great up to 128k context with an f16 KV cache (3.2 GB of KV) and up to 128k with a q8_0 KV cache (1.7 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
MXFP4 (12.11 GB) is the only tracked file: Runs great, 282.3 tok/s measured at 2k context (1 run; 8k estimate 235.0 tok/s, 188.0–282.0, calibrated estimate ±20%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.