gpt-oss-20b hardware requirements

About the model

Parameters
20.9B total · 3.6B active (MoE)
Attention
Sliding-window attention
Max context
131,072 tokens
Released
2025-08-05

GGUF files

Tracked quantizations of gpt-oss-20b, smallest first
QuantFile sizeQualityPublished by
MXFP4baseline12.11 GBNear-losslessllama.cpp team · ggml-org

Memory needed by quant and context

Weights + f16 KV cache + compute buffer, in GB. Add about 0.6 GB if the GPU also drives your display on Windows.
Quant4k8k16k32k64k128k
MXFP412.712.913.213.714.817.0

Which GPUs can run gpt-oss-20b?

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Scroll sideways to see every column.

54 GPUs, Macs and CPU setups at MXFP4 with 8k context, sorted by verdict and speed
HardwareVerdictSpeedQuantMemoryRuns asContextNotesPrice
GeForce RTX 5090
Runs greatDetails
282.3 tok/s8k estimate 235.0 tok/s (188.0–282.0, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 32.0 GBFull GPUup to 128k—$4,200 street (as of 2026-08-09)
GeForce RTX 4090
Runs greatDetails
225.2 tok/s8k estimate 198.3 tok/s (158.7–238.0, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 24.0 GBFull GPUup to 128k—$2,755 street (as of 2026-08-09)
Radeon RX 7900 XTX
Runs great
est. 209.9 tok/s167.9–251.8calibrated estimate ±20%
MXFP413.5 / 24.0 GBFull GPUup to 128k—$999 launch MSRP
GeForce RTX 5080
Runs great
204.9 tok/s8k estimate 188.9 tok/s (151.1–226.6, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 16.0 GBFull GPUup to 64k—$1,256 street (as of 2026-08-09)
GeForce RTX 3090 Ti
Runs great
est. 198.3 tok/s158.7–238.0calibrated estimate ±20%
MXFP413.5 / 24.0 GBFull GPUup to 128k—$1,999 launch MSRP
GeForce RTX 5070 Ti
Runs great
189.5 tok/s8k estimate 176.3 tok/s (141.0–211.5, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 16.0 GBFull GPUup to 64k—$949 street (as of 2026-08-09)
GeForce RTX 4080 Super
Runs great
186.4 tok/s8k estimate 144.8 tok/s (115.8–173.8, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 16.0 GBFull GPUup to 64k—$999 launch MSRP
GeForce RTX 5080 Laptop
Runs great
est. 176.3 tok/s141.0–211.5calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k——
GeForce RTX 5090 Laptop
Runs great
est. 176.3 tok/s141.0–211.5calibrated estimate ±20%
MXFP413.5 / 24.0 GBFull GPUup to 128k——
GeForce RTX 3090
Runs greatDetails
162.0 tok/s8k estimate 184.2 tok/s (147.3–221.0, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 24.0 GBFull GPUup to 128k—$1,050 street (as of 2026-08-09)
GeForce RTX 4080
Runs great
est. 141.1 tok/s112.9–169.3calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$1,199 launch MSRP
Radeon RX 9070
Runs great
est. 141.0 tok/s112.8–169.2calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$639 street (as of 2026-08-09)
Radeon RX 9070 XT
Runs great
est. 141.0 tok/s112.8–169.2calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$689 street (as of 2026-08-09)
Radeon RX 7800 XT
Runs great
est. 136.4 tok/s109.1–163.7calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$499 launch MSRP
GeForce RTX 4070 Ti Super
Runs great
est. 132.2 tok/s105.8–158.7calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$799 launch MSRP
Apple M3 Ultra · 96 GB
Runs great
115.5 tok/s8k estimate 103.8 tok/s (83.1–124.6, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP412.9 / 67.2 GBUnified memoryup to 128k——
GeForce RTX 4090 Laptop
Runs great
est. 113.3 tok/s90.7–136.0calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k——
GeForce RTX 5060 Ti 16GB
Runs greatDetails
111.6 tok/s8k estimate 88.1 tok/s (70.5–105.8, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 16.0 GBFull GPUup to 64k—$569 street (as of 2026-08-09)
Radeon RX 7900 XT
Runs great
101.9 tok/s8k estimate 174.9 tok/s (139.9–209.9, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP413.5 / 20.0 GBFull GPUup to 128k—$899 launch MSRP
Apple M2 Ultra · 64 GB
Runs great
est. 101.4 tok/s81.1–121.7calibrated estimate ±20%
MXFP412.9 / 44.8 GBUnified memoryup to 128k——
Apple M4 Max (40-core GPU) · 48 GB
Runs great
92.4 tok/s8k estimate 71.6 tok/s (57.3–85.9, calibrated estimate ±20%)measured (1 run, 2k context)
MXFP412.9 / 33.6 GBUnified memoryup to 128k——
Apple M5 Max (40-core GPU) · 48 GB
Runs great
est. 80.5 tok/s64.4–96.6calibrated estimate ±20%
MXFP412.9 / 33.6 GBUnified memoryup to 128k——
Radeon RX 9060 XT 16GB
Runs great
est. 70.4 tok/s56.3–84.5calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$449 street (as of 2026-08-09)
Apple M5 Max (32-core GPU) · 36 GB
Runs great
est. 60.3 tok/s48.3–72.4calibrated estimate ±20%
MXFP412.9 / 25.2 GBUnified memoryup to 64k——
Ryzen AI Max+ 395 (Strix Halo) · 32 GB
Runs great
est. 59.3 tok/s47.5–71.2calibrated estimate ±20%
MXFP412.9 / 22.4 GBUnified memoryup to 64k—$1,999 launch MSRP
GeForce RTX 4060 Ti 16GB
Runs great
est. 56.7 tok/s45.3–68.0calibrated estimate ±20%
MXFP413.5 / 16.0 GBFull GPUup to 64k—$499 launch MSRP
Apple M4 Max (32-core GPU) · 36 GB
Runs great
est. 53.8 tok/s43.0–64.5calibrated estimate ±20%
MXFP412.9 / 25.2 GBUnified memoryup to 64k——
NVIDIA DGX Spark · 128 GB
Runs great
est. 47.7 tok/s38.2–57.3calibrated estimate ±20%
MXFP412.9 / 89.6 GBUnified memoryup to 32k—$3,999 launch MSRP
Apple M5 Pro · 24 GB
Runs great
est. 40.3 tok/s32.2–48.3calibrated estimate ±20%
MXFP412.9 / 16.8 GBUnified memoryup to 32k——
Apple M4 Pro · 24 GB
Runs greatDetails
est. 35.8 tok/s28.6–43.0calibrated estimate ±20%
MXFP412.9 / 16.8 GBUnified memoryup to 16k——
Apple M5 · 24 GB
Runs well
est. 20.1 tok/s16.1–24.2calibrated estimate ±20%
MXFP412.9 / 16.8 GBUnified memoryup to 64k——
Apple M4 · 24 GB
Runs wellDetails
est. 15.7 tok/s12.6–18.9calibrated estimate ±20%
MXFP412.9 / 16.8 GBUnified memoryup to 32k——
Intel Arc B580
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$290 street (as of 2026-08-09)
GeForce GTX 1080 Ti
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 11.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$699 launch MSRP
GeForce RTX 2080 Ti
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 11.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$999 launch MSRP
GeForce RTX 3050 8GB
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$249 launch MSRP
GeForce RTX 3060 12GB
Runs slowlyDetails
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$250 street (as of 2026-08-09)
GeForce RTX 3070
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$499 launch MSRP
GeForce RTX 3080 10GB
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 10.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$699 launch MSRP
GeForce RTX 3080 12GB
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$799 launch MSRP
GeForce RTX 4060
Runs slowlyDetails
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$299 launch MSRP
GeForce RTX 4060 Laptop
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
—
GeForce RTX 4060 Ti 8GB
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$399 launch MSRP
GeForce RTX 4070
Runs slowlyDetails
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$599 launch MSRP
GeForce RTX 4070 Laptop
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
—
GeForce RTX 4070 Super
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$599 launch MSRP
GeForce RTX 4080 Laptop
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
—
GeForce RTX 5060
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$339 street (as of 2026-08-09)
GeForce RTX 5060 Ti 8GB
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 8.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$429 street (as of 2026-08-09)
GeForce RTX 5070
Runs slowlyDetails
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
$629 street (as of 2026-08-09)
GeForce RTX 5070 Ti Laptop
Runs slowly
est. 18.5 tok/s12.9–24.0theoretical estimate ±30%
MXFP42.0 / 12.0 GB + 11.5 GB RAMMoE experts in RAMup to 128k
  • Experts run from system RAM — Runs well needs 30 tok/s or more on this path
—
DDR5-6000 dual-channel CPU · 32 GB
Runs slowly
est. 18.0 tok/s14.4–21.7calibrated estimate ±20%
MXFP412.3 GB RAMCPU onlyup to 128k
  • CPU inference — capped at Runs slowly
—
DDR5-5600 dual-channel CPU · 32 GB
Runs slowly
est. 16.8 tok/s13.5–20.2calibrated estimate ±20%
MXFP412.3 GB RAMCPU onlyup to 128k
  • CPU inference — capped at Runs slowly
—
DDR4-3200 dual-channel CPU · 32 GB
Runs slowly
est. 9.6 tok/s7.7–11.6calibrated estimate ±20%
MXFP412.3 GB RAMCPU onlyup to 128k——

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Cheapest GPUs that run it

  • Runs greatRadeon RX 9060 XT 16GB$449 street (as of 2026-08-09), est. 70.4 tok/s (56.3–84.5, calibrated estimate ±20%).

Measured results for gpt-oss-20b

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured results for gpt-oss-20b
HardwareQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
GeForce RTX 5090MXFP4llama.cpp2k9,841.0282.3—github.com2025-08-15
GeForce RTX 4090MXFP4llama.cpp2k8,078.0225.2—github.com2025-08-15
GeForce RTX 5080MXFP4llama.cpp2k7,469.0204.9—github.com2025-08-15
GeForce RTX 5070 TiMXFP4llama.cpp2k6,326.0189.5—github.com2025-08-15
GeForce RTX 4080 SuperMXFP4llama.cpp2k8,111.0186.4—github.com2025-08-15
GeForce RTX 3090MXFP4llama.cpp2k5,143.0162.0—github.com2025-08-15
GeForce RTX 5060 Ti 16GBMXFP4llama.cpp2k3,821.0111.6—github.com2025-08-15
Radeon RX 7900 XTMXFP4llama.cpp2k4,252.0101.9ROCmgithub.com2025-08-15
Apple M3 UltraMXFP4llama.cpp2k2,816.0115.5—github.com2025-08-15
Apple M4 Max (40-core GPU)MXFP4llama.cpp2k1,277.092.4—github.com2025-08-15

Frequently asked questions

How much VRAM does gpt-oss-20b need?

At MXFP4 the weights are 12.11 GB; with 8k context the total is about 12.9 GB (KV cache 0.2 GB, compute buffer 0.6 GB).

What is the cheapest GPU that runs gpt-oss-20b well?

Cheapest that runs it great: Radeon RX 9060 XT 16GB at $449 (street price as of 2026-08-09), est. 70.4 tok/s (56.3–84.5, calibrated estimate ±20%).

Can gpt-oss-20b run on an 8 GB, 16 GB or 24 GB GPU?

At MXFP4 with 8k context: GeForce RTX 4060 — Runs slowly, est. 18.5 tok/s (12.9–24.0, theoretical estimate ±30%); GeForce RTX 5060 Ti 16GB — Runs great, 111.6 tok/s measured at 2k context (1 run; 8k estimate 88.1 tok/s, 70.5–105.8, calibrated estimate ±20%); GeForce RTX 3090 — Runs great, 162.0 tok/s measured at 2k context (1 run; 8k estimate 184.2 tok/s, 147.3–221.0, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.