Can I run gpt-oss-120b on the GeForce RTX 4060?
Won't runMXFP4 needs 64.3 GB, 56.9 GB more than the 7.4 GB of usable memory. With 96 GB of system RAM: Runs slowly, est. 13.9 tok/s (9.7–18.1, theoretical estimate ±30%).
8 GB VRAM · 272.0 GB/s · FP16 15.1 TFLOPS. Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Every quant of gpt-oss-120b on the GeForce RTX 4060
| Quant | File size | Verdict | Speed | Memory | Runs as | Notes |
|---|---|---|---|---|---|---|
| MXFP4baseline | 63.39 GB | Won't run | — | Needs about 64.3 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free | — | — |
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Where the memory goes at MXFP4
The memory, context and speed figures below assume 96 GB of system RAM (the configuration from the RAM hint above); with 32 GB nothing fits.
- Weights
- 0.8 GB
- KV cache
- 0.3 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 5.7 GB
- Weights in system RAM
- 62.6 GB
Context length vs. KV cache
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 64.1 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.0 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.0 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% |
| 8k | 64.3 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.1 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.1 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% |
| 16k | 64.6 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.4 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.2 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% |
| 32k | 65.4 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.8 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 64.5 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% |
| 64k | 66.9 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 65.8 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 65.2 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% |
| 128k | 69.9 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 67.7 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% | 66.5 GBRuns slowly est. 13.9 tok/s9.7–18.1theoretical estimate ±30% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
With 96 GB of system RAM: At MXFP4 the verdict stays Runs slowly up to 128k context with an f16 KV cache, and up to 128k with q8_0.
Estimated speed
- Generation
- With 96 GB of system RAM: est. 13.9 tok/s (9.7–18.1, theoretical estimate ±30%)
- Prompt processing
- Prompt-processing estimates only apply when the whole model runs on the GPU.
Why this verdict: Experts run from system RAM — Runs well needs 30 tok/s or more on this path
Measured on this exact combination
No public measurement for this exact combination yet; the numbers above are estimates.
If this is not enough
A smaller model that runs well on the GeForce RTX 4060
Runs wellQwen3.5 35B-A3B: Runs well, est. 20.4 tok/s (14.3–26.5, theoretical estimate ±30%)
Same GPU with 96 GB of system RAM
Runs slowlyMXFP4: Runs slowly, est. 13.9 tok/s (9.7–18.1, theoretical estimate ±30%)
How to run it
Commands for the MXFP4 file. Both tools download from Hugging Face on first run.
These commands assume 96 GB of system RAM (the configuration from the RAM hint above).
File: gpt-oss-120b-MXFP4.gguf (63.39 GB) from ggml-org/gpt-oss-120b-GGUF.
llama-server -hf ggml-org/gpt-oss-120b-GGUF:MXFP4 -c 8192 -ngl all --n-cpu-moe 36 -fa on- -ngl all loads every layer on the GPU.
- --n-cpu-moe 36 keeps the experts of all 36 layers in system RAM, which is what this estimate assumes.
- With 34 the model still fits; fewer layers on the CPU means faster generation.
- -fa on enables flash attention (needed for KV cache quantization).
- --cache-type-k q8_0 --cache-type-v q8_0 shrinks the KV cache to 0.16 GB at 8k context.
Ollama has no expert-offload flag, so this mode is llama.cpp only.
Related pages
Guide: How much VRAM do local LLMs need? 8–32 GB tiers (2026)
Other models on the GeForce RTX 4060
gpt-oss-120b on other hardware
Frequently asked questions
Can I run gpt-oss-120b on the GeForce RTX 4060?
MXFP4 needs 64.3 GB, 56.9 GB more than the 7.4 GB of usable memory. With 96 GB of system RAM: Runs slowly, est. 13.9 tok/s (9.7–18.1, theoretical estimate ±30%). MXFP4 needs 64.3 GB at 8k context; this setup has 7.4 GB of usable memory and 28.0 GB of free system RAM.
How much context can gpt-oss-120b use on the GeForce RTX 4060?
With 96 GB of system RAM: At MXFP4 the verdict stays Runs slowly up to 128k context with an f16 KV cache (4.8 GB of KV) and up to 128k with a q8_0 KV cache (2.6 GB). The model supports up to 128k.
Which quant should I use, and how fast is it?
No tracked quant runs with 32 GB of RAM; with 96 GB, MXFP4 reaches Runs slowly, est. 13.9 tok/s (9.7–18.1, theoretical estimate ±30%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.