How much VRAM do local LLMs need? 8–32 GB tiers (2026)
By CanRun · Updated September 30, 2026
A model runs fully on the GPU when its file, its KV cache (the memory that holds the context) and a small compute buffer (the runtime's working memory) fit in the VRAM your card can use. Llama 3.1 8B at Q4_K_M, a 4-bit quant, needs 6.6 GB with an 8k-token context, just inside the 7.4 GB an 8GB RTX 4060 can use (little margin, so check ollama ps), and it Runs great. At 8k, 12GB holds 12B to 14B models (Qwen3 14B with little margin), 16GB holds gpt-oss-20b, 24GB holds 30B-class models such as Qwen3 30B-A3B and 32GB mainly adds room for longer context. With 32GB of system RAM, gpt-oss-120b fits on none of these cards; in CanRun's calculation it needs 96GB of system RAM.
How much VRAM a model needs: file, KV cache and buffer
Three things share a graphics card's memory while a model runs: the model file at the quant you downloaded (Q4_K_M, a 4-bit GGUF quant, is the usual default), the KV cache, which holds keys and values for every token in the context, and a compute buffer, the runtime's working memory, which grows slightly with the context. CanRun's verdicts use Q4_K_M (MXFP4 for gpt-oss, its native format), an 8k context and an f16 (unquantized 16-bit) KV cache. These are the totals for the models this guide follows:
| Model | Parameters | Default quant file | KV cache at 8k | Total at 8k |
|---|---|---|---|---|
| Llama 3.1 8B | 8.03B | Q4_K_M · 4.9 GB | 1.1 GB | 6.6 GB |
| Gemma 3 12B | 12.2B | Q4_K_M · 7.3 GB | 0.5 GB | 8.4 GB |
| Qwen3 14B | 14.8B | Q4_K_M · 9.0 GB | 1.3 GB | 10.9 GB |
| gpt-oss-20b | 20.9B (3.6B active) | MXFP4 · 12.1 GB | 0.2 GB | 12.9 GB |
| Gemma 3 27B | 27.4B | Q4_K_M · 16.6 GB | 0.7 GB | 17.8 GB |
| Qwen3 30B-A3B (2507) | 30.5B (3.3B active) | Q4_K_M · 18.6 GB | 0.8 GB | 19.9 GB |
| Qwen3 32B | 32.8B | Q4_K_M · 19.8 GB | 2.1 GB | 22.5 GB |
| Llama 3.3 70B | 70.6B | Q4_K_M · 42.5 GB | 2.7 GB | 45.8 GB |
| gpt-oss-120b | 116.8B (5.1B active) | MXFP4 · 63.4 GB | 0.3 GB | 64.3 GB |
On a graphics card the OS also keeps some VRAM (about 0.6 GB on a Windows display GPU); on a Mac the GPU can use about 70% of unified memory by default.
Compare the Total column with the VRAM your card can use, not the number on the box. When the card also drives a Windows display, CanRun sets some VRAM aside for the desktop (the note under the table gives the amount). That leaves an 8GB card such as the RTX 4060 with 7.4 GB and a 12GB card such as the RTX 3060 12GB with 11.4 GB.
The file is the largest part. CanRun takes file sizes from the real GGUF files on Hugging Face. The quantize table in llama.cpp gives a rule of thumb: Q4_K_M averages a little under five bits per weight, so a Q4_K_M file in gigabytes is about three-fifths of the parameter count in billions. Llama 3.1 8B comes to 4.9 GB and Qwen3 14B to 9.0 GB. Q8_0 is roughly one gigabyte per billion parameters, so Qwen3 14B at Q8_0 is 15.7 GB.
At 8k the KV cache is small next to the file (1.1 GB for Llama 3.1 8B, 1.3 GB for Qwen3 14B), but it grows with the context, as the section on context length shows.
Qwen3 14B on the RTX 3060 12GB shows how the parts fill a 12GB card: it needs 10.9 GB at 8k against 11.4 GB usable, so in CanRun's calculation it Runs great with an f16 KV cache, with less than a gigabyte free. Check ollama ps after loading it: PROCESSOR should read 100% GPU. Our page for Qwen3 14B on the GeForce RTX 3060 12GB has the commands.
Where the memory goes: Qwen3 14B (Q4_K_M) on the GeForce RTX 3060 12GB
- Weights
- 9.0 GB
- KV cache
- 1.3 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 0.5 GB
Active parameters do not set the memory need: mixture-of-experts (MoE) models such as gpt-oss-20b use only a few billion parameters per token, but the whole file must be in memory (the Parameters column shows both counts). Macs use unified memory instead, of which the GPU gets about 70% by default in CanRun's calculation; see Running local LLMs on a Mac (M4, M5): how much unified memory?.
What fits on 8, 12, 16, 24 and 32 GB graphics cards
The grid crosses six popular models with eight NVIDIA cards. Every verdict assumes an 8k context, an f16 KV cache, 32GB of DDR5 system RAM and Windows. Each cell shows a verdict, an estimated speed band and a confidence label; cells labelled theoretical get an extra note under the grid.
| Hardware | Llama 3.1 8B | Gemma 3 12B | Qwen3 14B | gpt-oss-20b | Qwen3 30B-A3B (2507) | gpt-oss-120b |
|---|---|---|---|---|---|---|
| GeForce RTX 40608 GB | Runs great 38.1 tok/s8k estimate 31.8 tok/s (28.0–35.6, calibrated estimate ±12%)measured (1 run, 4k context) Details | |||||
| GeForce RTX 3060 12GB12 GB | Runs great 51.6 tok/s8k estimate 42.0 tok/s (37.0–47.1, calibrated estimate ±12%)measured (1 run, 4k context) Details | |||||
| GeForce RTX 407012 GB | ||||||
| GeForce RTX 507012 GB | ||||||
| GeForce RTX 5060 Ti 16GB16 GB | Runs great 111.6 tok/s8k estimate 88.1 tok/s (70.5–105.8, calibrated estimate ±20%)measured (1 run, 2k context) Details | |||||
| GeForce RTX 309024 GB | Runs great 162.0 tok/s8k estimate 184.2 tok/s (147.3–221.0, calibrated estimate ±20%)measured (1 run, 2k context) Details | |||||
| GeForce RTX 409024 GB | Runs great 91.0 tok/s8k estimate 117.7 tok/s (103.6–131.8, calibrated estimate ±12%)measured (1 run, 4k context) Details | Runs great 225.2 tok/s8k estimate 198.3 tok/s (158.7–238.0, calibrated estimate ±20%)measured (1 run, 2k context) Details | ||||
| GeForce RTX 509032 GB | Runs great 282.3 tok/s8k estimate 235.0 tok/s (188.0–282.0, calibrated estimate ±20%)measured (1 run, 2k context) Details |
Theoretical estimates (±30%) mark setups we have not validated against measurements yet, such as MoE experts kept in system RAM. They are likely on the conservative side.
8GB: RTX 4060
On the GeForce RTX 4060, Llama 3.1 8B fits fully with less than a gigabyte to spare (6.6 GB against 7.4 GB usable; confirm 100% GPU in ollama ps) and Runs great, a measured cell. Our page for Llama 3.1 8B on the GeForce RTX 4060 covers longer contexts.
The next size up does not fit. Gemma 3 12B needs 8.4 GB and Qwen3 14B needs 10.9 GB, both more than the card can use, so part of their weights go to system RAM (partial offload). Both are rated Runs slowly, as calibrated estimates. Our page for Gemma 3 12B on the GeForce RTX 4060 shows what a smaller quant changes.
Two of the MoE models, gpt-oss-20b and Qwen3 30B-A3B, run on this card only with their expert weights in system RAM. They are rated Runs slowly too, but those cells are theoretical estimates and likely conservative (see the MoE section). gpt-oss-120b is rated Won't run.
The GeForce RTX 5060 Ti 8GB shares its name with a 16GB card, so check the memory on the box. It gets the same fit results as the GeForce RTX 4060: Llama 3.1 8B Runs great and Qwen3 14B Runs slowly.
12GB: RTX 3060 12GB, RTX 4070 and RTX 5070
The grid has three 12GB cards: the GeForce RTX 3060 12GB, the GeForce RTX 4070 and the GeForce RTX 5070. Each can use 11.4 GB, so in CanRun's calculation they fit the same models.
Gemma 3 12B needs 8.4 GB, leaving room to spare, and Runs great on all three. Qwen3 14B fits at 8k with little margin (see the bar in the first section, and check ollama ps) and Runs great there too. At 16k with an f16 KV cache it spills over; see the context table further down. gpt-oss-20b needs 12.9 GB, more than a 12GB card can use, so its expert weights go to system RAM and the cell reads Runs slowly as a theoretical estimate.
Memory bandwidth largely sets the speed. The RTX 3060 12GB has 360 GB/s, the RTX 4070 504 GB/s and the RTX 5070 672 GB/s. Phi-4 needs 11.3 GB at 8k and fits on all three with almost no margin (check ollama ps), yet it is rated Runs well on the RTX 3060 12GB and Runs great on the RTX 4070. The gap comes from the reasoning rule: CanRun counts Phi-4 as a reasoning model and holds it to 1.5× stricter speed thresholds, which the faster RTX 4070 clears and the RTX 3060 12GB does not.
16GB: RTX 5060 Ti 16GB
On the GeForce RTX 5060 Ti 16GB, gpt-oss-20b needs 12.9 GB against 15.4 GB usable, so it fits fully on the GPU, in line with its model card, which says it runs within 16GB of memory. It Runs great, a measured cell. Our page for gpt-oss-20b on the GeForce RTX 5060 Ti 16GB has the full breakdown.
Qwen3 14B Runs great here as well, and the extra memory goes to context: at 32k with an f16 KV cache it needs 15.2 GB, still inside the 15.4 GB this card can use but with very little margin (check ollama ps).
At Q4_K_M, models larger than gpt-oss-20b do not fit fully; the table further down names the largest one that does at IQ4_XS. Qwen3 30B-A3B needs 19.9 GB, so its experts go to system RAM and the cell reads Runs slowly as a theoretical estimate. Gemma 3 27B, outside the grid, needs 17.8 GB; it runs with part of its weights in system RAM and is rated Runs slowly, a calibrated estimate.
24GB: RTX 3090 and RTX 4090
The GeForce RTX 3090 and the GeForce RTX 4090 can each use 23.4 GB. That is enough for Qwen3 30B-A3B, which needs 19.9 GB: it runs fully on the GPU and Runs great on both. Our page for Qwen3 30B-A3B (2507) on the GeForce RTX 3090 shows how it holds up at longer contexts.
The dense Qwen3 32B needs 22.5 GB at 8k, so it fits on the RTX 3090 with less than a gigabyte to spare (confirm 100% GPU in ollama ps) and Runs great there. A longer context with an f16 KV cache pushes it partly off the GPU, and Ollama's default on 24GB cards is already 32k; for that case, see Ollama slow or forgetting your prompt? Context length defaults.
Speed can still separate two cards with the same memory. DeepSeek R1 Distill Qwen 32B, a reasoning model, is rated Runs well on the RTX 3090 but Runs great on the RTX 4090. gpt-oss-120b is rated Won't run on both.
32GB: RTX 5090
For the six grid models, the RTX 5090 gets the same 8k verdicts as the 24GB cards. Its extra memory goes to context: Qwen3 32B at 32k with an f16 KV cache needs 29.1 GB, inside the 31.4 GB this card can use but over the 23.4 GB of a 24GB card. At the 8k baseline it Runs great on the RTX 5090.
At Q4_K_M, 32GB still does not hold a dense 70B model. Llama 3.3 70B needs 45.8 GB at that quant, so it runs with part of its weights in system RAM and is rated Runs slowly on the RTX 5090. On the RTX 4090 it is rated Won't run at Q4_K_M because partial offload is too slow. Only Heavy loss quants fit Llama 3.3 70B fully on the RTX 5090. gpt-oss-120b is still rated Won't run with 32GB of system RAM, even here.
The whole catalog, card by card
The table below counts every model in CanRun's catalog by verdict for each of the eight cards (the caption gives the total) and names the largest model that fits on the GPU alone, with its quant.
| Hardware | Memory | Runs great | Runs well | Runs slowly | Won't run | Largest model that fits on the GPU alone |
|---|---|---|---|---|---|---|
| GeForce RTX 4060 | 8 GB | 7 | 1 | 13 | 8 | Qwen3.5 9B Q4_K_M |
| GeForce RTX 3060 12GB | 12 GB | 10 | 3 | 11 | 5 | Kanana 1.5 15.7B-A3B IQ4_XS |
| GeForce RTX 4070 | 12 GB | 12 | 1 | 11 | 5 | Kanana 1.5 15.7B-A3B IQ4_XS |
| GeForce RTX 5070 | 12 GB | 12 | 1 | 11 | 5 | Kanana 1.5 15.7B-A3B IQ4_XS |
| GeForce RTX 5060 Ti 16GB | 16 GB | 13 | 3 | 8 | 5 | Mistral Small 3.2 24B IQ4_XS |
| GeForce RTX 3090 | 24 GB | 23 | 1 | 0 | 5 | Qwen3.5 35B-A3B Q4_K_M |
| GeForce RTX 4090 | 24 GB | 24 | 0 | 0 | 5 | Qwen3.5 35B-A3B Q4_K_M |
| GeForce RTX 5090 | 32 GB | 24 | 0 | 1 | 4 | Qwen3.5 35B-A3B Q4_K_M |
At 8k, the set of models that fit fully on the GPU grows at each step from 8GB to 24GB; the 32GB RTX 5090 adds none at the baseline. Its extra memory goes to longer context and to running Llama 3.3 70B at Q4_K_M partly from system RAM, as the 32GB section explains. Between cards of the same size, only speed moves models from one verdict to another. The counts change as models join the catalog.
Other cards with the same memory
In CanRun's calculation, fit depends only on the memory a card can use. Another 16GB card such as the Radeon RX 9060 XT 16GB gets the same fit result for gpt-oss-20b as the GeForce RTX 5060 Ti 16GB and Runs great there too; only the speeds differ.
A laptop GPU can carry a desktop card's name with less VRAM: the GeForce RTX 4090 Laptop has 16 GB and the GeForce RTX 5090 Laptop has 24 GB. Look up the laptop chip itself.
Memory targets by model size
- 8B-class models: 8GB of VRAM.
- 12B to 14B models: 12GB, or 16GB for longer context.
- gpt-oss-20b fully on the GPU: 16GB.
- 30B-class models: 24GB, or 32GB for longer context.
MoE models: the whole file counts, but experts can sit in system RAM
In a mixture-of-experts (MoE) model, the bulk of each layer is split into many expert blocks, and a router picks a few for every token. Qwen3 30B-A3B has 128 experts with 8 active per token, according to its model card; gpt-oss-20b and gpt-oss-120b work the same way. Fit needs the whole file, because the router can pick any expert; speed depends mostly on the active part.
When the model does not fit fully in VRAM (file, KV cache and buffer together), CanRun's calculation keeps the shared and attention weights, the KV cache and the buffer on the GPU and moves the expert weights to system RAM, as llama.cpp's --cpu-moe and --n-cpu-moe N options do. Qwen3 30B-A3B (2507) needs 19.9 GB, against 11.4 GB usable on the RTX 3060 12GB, so the bar shows a small GPU part and most of the weights in system RAM:
Where the memory goes: Qwen3 30B-A3B (2507) (Q4_K_M) on the GeForce RTX 3060 12GB
- Weights
- 0.9 GB
- KV cache
- 0.8 GB
- Compute buffer
- 0.6 GB
- OS reserve
- 0.6 GB
- Free
- 9.1 GB
- Weights in system RAM
- 17.7 GB
In this mode the verdict follows a stricter rule, shown in the verdict legend further down: the best verdict is Runs well, which needs the speed Runs great needs on a full GPU; slower than that is Runs slowly, and very slow is Won't run. The gpt-oss models are reasoning models, so their thresholds are 1.5× stricter again. Qwen3 30B-A3B (2507) on the GeForce RTX 3060 12GB is rated Runs slowly by that rule, as a theoretical estimate.
Treat these speeds with care: they are theoretical estimates (±30%) and likely conservative. Against CanRun's one public data point for this setup, an RTX 3060 12GB with DDR5 running a 35B-A3B MoE model, the estimate is about 2.5 times too low. Read the Runs slowly cells for gpt-oss-20b and Qwen3 30B-A3B on the 8GB to 16GB cards as "experts in system RAM, speed not yet validated", not as proof that MoE models are slow on small cards.
For the largest MoE model in the grid, system RAM becomes the limit. The gpt-oss-120b file is 63.4 GB. With 32GB of system RAM the experts have nowhere to go, so it is rated Won't run on every card in the grid, from the 8GB RTX 4060 to the 32GB RTX 5090. In CanRun's calculation, which keeps every expert weight in system RAM, 64GB is still short; 96GB is the first size checked where it fits, and the grid cells already show this as "With 96 GB RAM: Runs slowly", a theoretical estimate. Its model card targets a single 80GB GPU.
Runtimes differ here. Ollama picks its own GPU/CPU split and documents no switch for keeping experts in RAM; ollama ps shows what it chose under PROCESSOR. Combo pages print the full llama.cpp command; for the setup above, see our page for Qwen3 30B-A3B (2507) on the GeForce RTX 3060 12GB.
How quantization and context length change the answer
On a given card, two settings move a model across the lines above: the quant you download and the context you run it at.
Quantization: a smaller file, at a cost in quality
A quant stores each weight in fewer bits, which makes the file smaller. CanRun sorts quants into four quality groups: Near-lossless (Q5_K_M, Q6_K, Q8_0, and MXFP4, gpt-oss's native format), Good balance (Q4_K_M and IQ4_XS), Noticeable loss (3-bit quants such as Q3_K_M and IQ3_XS) and Heavy loss (2-bit quants such as Q2_K and IQ2_M). A Q8_0 file is well over one and a half times the Q4_K_M file: Gemma 3 12B is 12.5 GB at Q8_0 against 7.3 GB at Q4_K_M. The ladder below shows each step for Gemma 3 12B on the GeForce RTX 4060; every cell is calibrated.
| Quant | Quality group | Verdict | Speed | Memory |
|---|---|---|---|---|
| Q8_0 | Near-lossless | Runs slowly | est. 5.6 tok/s4.5–6.7calibrated estimate ±20% | 8.0 / 8.0 GB + 6.2 GB RAM |
| Q4_K_M | Good balance | Runs slowly | est. 16.7 tok/s13.4–20.1calibrated estimate ±20% | 8.0 / 8.0 GB + 1.0 GB RAM |
| IQ4_XS | Good balance | Runs well | est. 23.8 tok/s19.0–28.5calibrated estimate ±20% | 8.0 / 8.0 GB + 0.3 GB RAM |
| Q2_K | Heavy loss | Runs great | est. 35.9 tok/s31.6–40.2calibrated estimate ±12% | 6.5 / 8.0 GB |
At Q8_0 and Q4_K_M the model spills into system RAM; Q4_K_M, the baseline, is rated Runs slowly. IQ4_XS needs 7.7 GB, still over the 7.4 GB the card can use, so it remains partial offload, but with at least 90% on the GPU the table shows Runs well. Q2_K fits fully and the table shows Runs great, but it belongs to the Heavy loss group. CanRun suggests a Heavy loss quant only when every other quant is rated Won't run; the better first move is one level down within Good balance, from Q4_K_M to IQ4_XS. For a closer comparison of the levels, see GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download.
More VRAM lets you move up in quality instead, at a cost: a bigger file is slower per token and leaves less room for context, so CanRun does not suggest it on its own. Qwen3 14B at Q8_0 needs 17.6 GB, more than the 15.4 GB of a 16GB card but inside the 23.4 GB of a 24GB card, where in CanRun's calculation it stays fully on the GPU.
Context length: the KV cache grows
In a model with full attention on every layer, such as Llama 3.1 8B, Qwen3 14B, Qwen3 32B and Qwen3 30B-A3B, the KV cache grows in step with the context: double the context, double the cache. Gemma 3 and gpt-oss mix in sliding-window layers that only see a short stretch of recent tokens (as the Gemma 3 technical report and the gpt-oss configuration files describe), so their cache stays much smaller. At 32k with an f16 cache, Llama 3.1 8B needs 4.3 GB for the cache alone, gpt-oss-20b only 0.8 GB.
| Model | KV per token | at 8k | at 32k | at 128k |
|---|---|---|---|---|
| Qwen3 32B | 256 KiB | 2.1 GB | 8.6 GB | — |
| Qwen3 14B | 160 KiB | 1.3 GB | 5.4 GB | — |
| Llama 3.1 8B | 128 KiB | 1.1 GB | 4.3 GB | 17.2 GB |
| Qwen3 30B-A3B (2507) | 96 KiB | 0.8 GB | 3.2 GB | 12.9 GB |
| Gemma 3 12B | 64 KiB | 0.5 GB | 2.1 GB | 8.6 GB |
| gpt-oss-120b | 36 KiB | 0.3 GB | 1.2 GB | 4.8 GB |
| gpt-oss-20b | 24 KiB | 0.2 GB | 0.8 GB | 3.2 GB |
A dash means the context is longer than the model supports. q8_0 and q4_0 KV caches take about 54% and 29% of the f16 size.
Qwen3 14B and Qwen3 32B support 32k natively, so their 128k column shows a dash. The context table for Qwen3 14B on a 12GB card shows where the line falls; every cell is calibrated.
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 10.2 GBRuns great est. 26.1 tok/s22.9–29.2calibrated estimate ±12% | 9.9 GBRuns great est. 26.9 tok/s23.7–30.2calibrated estimate ±12% | 9.7 GBRuns great est. 27.4 tok/s24.1–30.7calibrated estimate ±12% |
| 8k | 10.9 GBRuns great est. 24.4 tok/s21.4–27.3calibrated estimate ±12% | 10.3 GBRuns great est. 25.9 tok/s22.8–29.0calibrated estimate ±12% | 10.0 GBRuns great est. 26.9 tok/s23.6–30.1calibrated estimate ±12% |
| 16k | 12.3 GBRuns slowly est. 14.6 tok/s11.7–17.5calibrated estimate ±20% | 11.1 GBRuns great est. 24.1 tok/s21.2–27.0calibrated estimate ±12% | 10.4 GBRuns great est. 25.8 tok/s22.7–28.9calibrated estimate ±12% |
| 32k | 15.2 GBRuns slowly est. 6.0 tok/s4.8–7.2calibrated estimate ±20% | 12.7 GBRuns slowly est. 12.8 tok/s10.3–15.4calibrated estimate ±20% | 11.3 GBRuns great est. 23.9 tok/s21.1–26.8calibrated estimate ±12% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
At 16k with an f16 KV cache, Qwen3 14B needs 12.3 GB, over the 11.4 GB the card can use, so part of the model moves to system RAM. A q8_0 (8-bit) KV cache brings 16k back onto the GPU, at 11.1 GB, with little margin. At 32k only a q4_0 (4-bit) cache keeps it on the GPU, at 11.3 GB with very little margin; in both cases confirm 100% GPU in ollama ps. The 32k f16 and q8_0 cells are partial offload; the table gives each cell's verdict.
Ollama adds its own default. Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). So 8GB to 16GB cards start at 4k and 24GB and 32GB cards at 32k, while CanRun's verdicts use 8k. Ollama quantizes the KV cache only when the server starts with OLLAMA_KV_CACHE_TYPE (f16, q8_0 or q4_0) and flash attention on; the setting applies to every model the server runs (Ollama docs, checked 2026-09-29). The Ollama FAQ puts q8_0 at about half and q4_0 at about a quarter of f16; CanRun's tables use about 54% and 29%. Parallel requests multiply context memory (OLLAMA_NUM_PARALLEL); the tables assume one slot. For the settings, see Ollama slow or forgetting your prompt? Context length defaults and KV cache explained: how context length eats VRAM.
Where to go next
- KV cache explained: how context length eats VRAM: how the KV cache works and what q8_0 and q4_0 save.
- GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download: quant levels, file sizes and quality groups.
- Ollama slow or forgetting your prompt? Context length defaults: Ollama's default context and how to change it.
- Running local LLMs on a Mac (M4, M5): how much unified memory?: the same questions for unified memory.
- Card and model pages, such as the GeForce RTX 4060, the GeForce RTX 3090 and Qwen3 14B, list every combination.
How sure are these numbers?
The memory figures are calculated by formula from real GGUF file sizes, each model's published configuration and the compute buffer. Where a real file size is missing, CanRun falls back to a weight-size formula, which matched eleven real GGUF files with a mean error of 0.7%. CanRun's KV formulas matched the published values on seven architectures.
The speeds are estimates with a band and a label. Measured means a public benchmark for that card, model and quant. Calibrated means ±12% for dense models fully on an NVIDIA or AMD card and ±20% for MoE models, partial offload and other hardware. Theoretical means ±30%, for setups no measurement backs yet; for MoE experts in system RAM these estimates are likely conservative. Against 12 public measurements, the speed formula has a mean error of 6.3% and a maximum of 17.1%.
In the grid, about half the cells are calibrated and several are measured; the cells with MoE experts in system RAM, including the 96GB hint for gpt-oss-120b, are theoretical.
Below is one public measurement, for gpt-oss-20b on the GeForce RTX 5060 Ti 16GB, from llama.cpp's gpt-oss guide thread (llama-bench, llama.cpp's benchmark tool, at a 2k context; not a CanRun run), next to CanRun's own 8k estimate band.
All of this assumes Windows with the card driving the display, 32GB of DDR5-5600 dual-channel system RAM, an 8k context, an f16 KV cache, one model loaded, one request slot and a llama.cpp-class runtime. Ollama and llama.cpp make their own estimates, so on your PC the PROCESSOR column of ollama ps is the ground truth: 100% GPU, or a CPU/GPU split.
Verdict names in the text of this guide refer to the 8k, f16 KV baseline; for other contexts and quants, read the tables. Models CanRun tags as reasoning models, such as gpt-oss, Phi-4 and the DeepSeek R1 distills, face 1.5× stricter speed thresholds, which the legend does not show. The thresholds:
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
FAQ
Is 8GB of VRAM enough to run a local LLM?
Yes, for 8B-class models at Q4_K_M with an 8k context. Llama 3.1 8B needs 6.6 GB against the 7.4 GB an 8GB card can use, and it Runs great on the RTX 4060; the margin is small, so check ollama ps. At Q4_K_M, the 12B to 14B dense models in the catalog do not fit fully: Qwen3 14B needs 10.9 GB, so part of it runs from system RAM and it Runs slowly. Longer context makes it tighter: at 16k with an f16 KV cache, Llama 3.1 8B needs 7.7 GB, already over, and a q8_0 (8-bit) KV cache or a shorter context keeps it on the GPU.
How do I work out how much VRAM a model needs?
Add the GGUF file size at your quant, the KV cache for your context and a small compute buffer, then compare the total with the VRAM your card can use: the number on the box minus what the OS keeps for the display. At Q4_K_M a file in gigabytes is about three-fifths of the parameter count in billions, at Q8_0 roughly one gigabyte per billion; the real file size is better still. For Qwen3 14B, a 9.0 GB file plus 1.3 GB of KV cache at 8k, plus the buffer, comes to 10.9 GB.
Does a longer context need more VRAM?
Yes. The KV cache stores every token in the context, so for models with full attention on every layer, doubling the context doubles the cache. How much depends on the architecture: at 32k with an f16 cache it is 4.3 GB for Llama 3.1 8B but 0.8 GB for gpt-oss-20b, whose sliding-window layers keep only recent tokens. Mind the runtime default as well: Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). A q8_0 or q4_0 (8-bit or 4-bit) KV cache shrinks it; see KV cache explained: how context length eats VRAM.
Do MoE models like gpt-oss-20b or Qwen3 30B-A3B need less VRAM than dense models of the same size?
Not to fit. The whole file has to be in memory; the small number of active parameters reduces the work per token, not the file size. gpt-oss-20b needs 12.9 GB at 8k, so it fits fully on a 16GB card such as the RTX 5060 Ti 16GB and Runs great there, in line with its model card. What MoE does allow is keeping the expert weights in system RAM (llama.cpp's --cpu-moe); CanRun marks that setup's speed as a theoretical estimate for now.
Can system RAM make up for too little VRAM?
Partly. Runtimes put whatever does not fit into system RAM, so the model still runs, but slower, and for dense models the drop is steep: Qwen3 14B on an 8GB RTX 4060 Runs slowly. For large MoE models, system RAM itself becomes the limit. The gpt-oss-120b file is 63.4 GB, and with 32GB of RAM it is rated Won't run even on the 32GB RTX 5090; in CanRun's calculation it needs 96GB of RAM, and 64GB is not enough. ollama ps or the llama.cpp log shows what your runtime did.
Sources
- llama.cpp: quantize README: bits per weight and file sizes for Q4_K_M, IQ4_XS, Q8_0 and other quants of Llama 3.1 8B, the basis of the file-size rule of thumb.
- llama.cpp: server README:
-ngl/--n-gpu-layers,--cpu-moe(all MoE weights on the CPU),--n-cpu-moe N(MoE weights of the first N layers on the CPU),-c/--ctx-sizeand--cache-type-k(f16 by default, q8_0, q4_0). - Ollama docs: context length: the default context by VRAM tier, the advice to avoid offloading to the CPU, and checking with
ollama ps. - Ollama FAQ: the PROCESSOR column of
ollama ps, KV cache quantization (q8_0 about half and q4_0 about a quarter of f16, with flash attention), and memory scaling withOLLAMA_NUM_PARALLEL. - gpt-oss-20b model card: total and active parameters, MoE weights in MXFP4, runs within 16GB of memory.
- gpt-oss-120b model card: total and active parameters, MXFP4, fits a single 80GB GPU.
- gpt-oss-20b configuration and gpt-oss-120b configuration:
layer_typesalternating sliding-window and full attention, with a shortsliding_window. - Qwen3-30B-A3B-Instruct-2507 model card: total and activated parameters, 128 experts with 8 activated, 48 layers, native context.
- Qwen3-14B model card: parameters, layers, KV heads and native context, with the YaRN extension.
- Qwen3-32B model card: parameters, layers, KV heads and native context.
- Llama 3.1 8B Instruct model card: parameters, 128k context, grouped-query attention.
- Llama 3.3 70B Instruct model card: parameters and 128k context, used for the 70B example in the 32GB section.
- Gemma 3 12B model card: 128k input context (downloading the files requires accepting Google's license; the card text is public).
- Gemma 3 Technical Report: more local than global attention layers with a short local span, to limit KV cache growth at long context.
- llama.cpp: guide to running gpt-oss: llama-bench results on RTX cards, the source of the measured gpt-oss-20b point on the RTX 5060 Ti 16GB.
Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.