KV cache explained: how context length eats VRAM
By CanRun · Updated September 30, 2026
A longer context needs more memory because, on top of the fixed-size model file, the KV cache stores a key and a value for every token in every attention layer. For Llama 3.1 8B at Q4_K_M, a 4-bit quant, the file is 4.9 GB, while a 16-bit (f16) KV cache at 128k takes 17.2 GB, more than three times the size of the file. How steeply the cache grows depends on the attention design: sliding-window and hybrid models need a fraction of that. Storing the cache as q8_0 or q4_0 shrinks it to roughly a half or a quarter, at some cost in quality.
Why a longer context costs memory
The memory a model needs has three parts. The model file is fixed by the weight quant you download, Q4_K_M in the examples here unless noted. The KV cache grows with the context you set. The compute buffer is the working memory the runtime needs during inference; CanRun approximates it from llama.cpp logs, and it grows with context too, but more slowly than the KV cache.
In the KV cache, each attention layer keeps a key and a value per KV head (the heads that store keys and values; several query heads can share one) for every token already in the context, so new tokens can look back at them without recomputing them. The size per token follows from the model's configuration:
2 (key + value) × layers × KV heads × head size × bytes per value
With an f16 cache, the default in llama.cpp and Ollama, each value takes 16 bits. The whole cache is this per-token size times the context length, so doubling the context, as each column below does, doubles the cache.
| Model | KV per token | at 4k | at 8k | at 16k | at 32k |
|---|---|---|---|---|---|
| Qwen3 32B | 256 KiB | 1.1 GB | 2.1 GB | 4.3 GB | 8.6 GB |
| Qwen3 14B | 160 KiB | 0.7 GB | 1.3 GB | 2.7 GB | 5.4 GB |
| Llama 3.1 8B | 128 KiB | 0.5 GB | 1.1 GB | 2.1 GB | 4.3 GB |
A dash means the context is longer than the model supports. q8_0 and q4_0 KV caches take about 54% and 29% of the f16 size.
In their published configurations, these three models share the same KV-head layout (the same number of KV heads, each the same size), so their per-token cost differs only by the number of layers: 32 for Llama 3.1 8B, 40 for Qwen3 14B and 64 for Qwen3 32B. At 32k with an f16 cache, that works out to 4.3 GB, 5.4 GB and 8.6 GB of KV cache. By 128k, the f16 cache of Llama 3.1 8B is more than three times the size of its Q4_K_M file.
A model's native limit caps what is useful: Qwen3 14B and Qwen3 32B handle 32k natively and Llama 3.1 8B handles 128k (the Qwen model cards describe extending this with YaRN, which is outside this guide).
Why the KV cost differs so much between models
The table below covers the whole catalog at 8k, 32k and 128k (a dash means the model does not support that context). The order does not follow parameter count; three design choices explain most of it.
| Model | KV per token | at 8k | at 32k | at 128k |
|---|---|---|---|---|
| Llama 3.3 70B | 320 KiB | 2.7 GB | 10.7 GB | 42.9 GB |
| DeepSeek R1 Distill Qwen 32B | 256 KiB | 2.1 GB | 8.6 GB | 34.4 GB |
| Qwen3 32B | 256 KiB | 2.1 GB | 8.6 GB | — |
| Phi-4 | 200 KiB | 1.7 GB | — | — |
| Solar Open 100B | 192 KiB | 1.6 GB | 6.4 GB | 25.8 GB |
| Mistral Small 3.2 24B | 160 KiB | 1.3 GB | 5.4 GB | 21.5 GB |
| Qwen3 14B | 160 KiB | 1.3 GB | 5.4 GB | — |
| HyperCLOVA X SEED Think 14B | 152 KiB | 1.3 GB | 5.1 GB | 20.4 GB |
| Qwen3 8B | 144 KiB | 1.2 GB | 4.8 GB | — |
| DeepSeek R1 Distill Llama 8B | 128 KiB | 1.1 GB | 4.3 GB | 17.2 GB |
| Kanana 1.5 15.7B-A3B | 128 KiB | 1.1 GB | 4.3 GB | — |
| Kanana 1.5 8B | 128 KiB | 1.1 GB | 4.3 GB | — |
| Llama 3.1 8B | 128 KiB | 1.1 GB | 4.3 GB | 17.2 GB |
| HyperCLOVA X SEED 1.5B | 96 KiB | 0.8 GB | — | — |
| Qwen3 30B-A3B (2507) | 96 KiB | 0.8 GB | 3.2 GB | 12.9 GB |
| Gemma 3 27B | 80 KiB | 0.7 GB | 2.7 GB | 10.7 GB |
| EXAONE 4.0 32B | 64 KiB | 0.5 GB | 2.1 GB | 8.6 GB |
| EXAONE 4.5 33B | 64 KiB | 0.5 GB | 2.1 GB | 8.6 GB |
| Gemma 3 12B | 64 KiB | 0.5 GB | 2.1 GB | 8.6 GB |
| Gemma 4 12B | 64 KiB | 0.5 GB | 2.1 GB | 8.6 GB |
| Qwen3.5 27B | 64 KiB | 0.5 GB | 2.1 GB | 8.6 GB |
| EXAONE 4.0 1.2B | 60 KiB | 0.5 GB | 2.0 GB | — |
| Solar Open 2 250B | 48 KiB | 0.4 GB | 1.6 GB | 6.4 GB |
| Gemma 4 26B-A4B | 40 KiB | 0.3 GB | 1.3 GB | 5.4 GB |
| gpt-oss-120b | 36 KiB | 0.3 GB | 1.2 GB | 4.8 GB |
| Qwen3.5 9B | 32 KiB | 0.3 GB | 1.1 GB | 4.3 GB |
| gpt-oss-20b | 24 KiB | 0.2 GB | 0.8 GB | 3.2 GB |
| Qwen3.5 122B-A10B | 24 KiB | 0.2 GB | 0.8 GB | 3.2 GB |
| Qwen3.5 35B-A3B | 20 KiB | 0.2 GB | 0.7 GB | 2.7 GB |
A dash means the context is longer than the model supports. q8_0 and q4_0 KV caches take about 54% and 29% of the f16 size.
Grouped-query attention
In grouped-query attention (GQA), many query heads share one KV head, so the cache stores fewer keys and values per token; the Llama 3.1 model card says all its versions use GQA. For a model with full attention on every layer, the cost follows layers × KV heads × head size, not the total parameter count, which is why Phi-4, a 14B model, costs more per token than Qwen3 14B.
Sliding-window attention
Gemma 3, Gemma 4, gpt-oss, EXAONE 4.0 32B and EXAONE 4.5 33B mix local layers, which attend only to a short window of recent tokens, with global (full-attention) layers; only the global layers keep growing with the context. The Gemma 3 technical report uses more local layers per global one and a short local span to cut KV-cache memory, the gpt-oss configuration alternates the two with a 128-token window, and the EXAONE 4.0 model card gives the 32B model a 3:1 ratio of local to global layers. In the table, Gemma 3 12B costs less than half as much per token as Qwen3 14B.
The effect is easiest to see at the same size. In CanRun's figures, EXAONE 4.0 32B, a Korean model from LG AI Research, needs 2.1 GB of f16 KV cache at 32k, a quarter of the 8.6 GB Qwen3 32B needs; the fixed sliding-window part, left out of that figure (see below), narrows the gap somewhat.
Hybrid linear attention
The Qwen3.5 9B model card lists eight blocks of three Gated DeltaNet layers followed by one Gated Attention layer, so only one layer in four keeps a KV cache; the linear layers carry a fixed-size state instead. At 32k with f16, Qwen3.5 9B needs 1.1 GB of KV cache, about a quarter of the 4.8 GB Qwen3 8B needs.
MoE and model size are poor guides
A mixture-of-experts (MoE) design does not shrink the KV cache by itself: the cache depends on the attention layout, not on the parameters active per token. The MoE Kanana 1.5 15.7B-A3B costs the same per token as the dense Llama 3.1 8B, and Qwen3 30B-A3B (2507) costs four times as much per token as gpt-oss-20b. Total size does not predict it either: at 32k with f16, gpt-oss-120b needs 1.2 GB of KV cache, well under the 4.3 GB Llama 3.1 8B needs.
How CanRun counts these models
For sliding-window and hybrid models, CanRun's figure includes only the full-attention layers. It leaves out the part whose size does not grow with the context: the sliding windows and the linear layers' recurrent state. These figures therefore run low, most visibly at short contexts and for models with wider windows (EXAONE 4.0 32B's window is far wider than gpt-oss's); next to the model file the gap is small. The counted part still grows linearly with the context, so the cache is not flat: Gemma 3 12B at 128k with f16 still needs 8.6 GB of KV cache.
Will a longer context fit on your card?
To see whether a context fits, compare the total need (file, KV cache and compute buffer) with the memory your card can actually use, not the VRAM on the box. A card that drives a Windows display keeps some VRAM back (see the note under the table below), so in CanRun's calculation the GeForce RTX 5060 Ti 16GB has 15.4 GB to work with.
At 128k, the smallest file does not mean the smallest need.
| Model | Parameters | Default quant file | KV cache at 128k | Total at 128k |
|---|---|---|---|---|
| Llama 3.1 8B | 8.03B | Q4_K_M · 4.9 GB | 17.2 GB | 23.8 GB |
| Qwen3.5 9B | 9.65B | Q4_K_M · 5.9 GB | 4.3 GB | 11.9 GB |
| Gemma 3 12B | 12.2B | Q4_K_M · 7.3 GB | 8.6 GB | 17.6 GB |
| gpt-oss-20b | 20.9B (3.6B active) | MXFP4 · 12.1 GB | 3.2 GB | 17.0 GB |
| Qwen3 30B-A3B (2507) | 30.5B (3.3B active) | Q4_K_M · 18.6 GB | 12.9 GB | 33.1 GB |
On a graphics card the OS also keeps some VRAM (about 0.6 GB on a Windows display GPU); on a Mac the GPU can use about 70% of unified memory by default.
Llama 3.1 8B needs 23.8 GB at 128k with f16, more than Gemma 3 12B at 17.6 GB and gpt-oss-20b at 17.0 GB. Yet gpt-oss-20b has roughly two and a half times as many parameters as Llama 3.1 8B (its default file is MXFP4, the 4-bit format gpt-oss is released in). Qwen3.5 9B needs 11.9 GB, about half of what Llama 3.1 8B needs. Qwen3 30B-A3B (2507) needs 33.1 GB, more than even the 31.4 GB an RTX 5090 can use.
A 16GB card: sliding window against full attention
Take Gemma 3 12B on the GeForce RTX 5060 Ti 16GB. At 8k with an f16 KV cache it Runs great. At 64k with f16 it needs 12.7 GB, inside the card's usable memory with room left. At 128k with f16 it needs 17.6 GB, over the limit, so in CanRun's calculation part of the weights moves to system RAM (partial offload) and the estimated speed in that cell drops sharply. A q8_0 KV cache brings 128k down to 13.6 GB, which fits (see the next section).
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 8.1 GBRuns great est. 41.4 tok/s36.5–46.4calibrated estimate ±12% | 8.0 GBRuns great est. 42.1 tok/s37.1–47.2calibrated estimate ±12% | 7.9 GBRuns great est. 42.5 tok/s37.4–47.6calibrated estimate ±12% |
| 8k | 8.4 GBRuns great est. 40.0 tok/s35.2–44.8calibrated estimate ±12% | 8.2 GBRuns great est. 41.3 tok/s36.4–46.3calibrated estimate ±12% | 8.0 GBRuns great est. 42.1 tok/s37.0–47.1calibrated estimate ±12% |
| 16k | 9.0 GBRuns great est. 37.5 tok/s33.0–41.9calibrated estimate ±12% | 8.5 GBRuns great est. 39.8 tok/s35.0–44.6calibrated estimate ±12% | 8.3 GBRuns great est. 41.2 tok/s36.3–46.2calibrated estimate ±12% |
| 32k | 10.2 GBRuns great est. 33.2 tok/s29.2–37.2calibrated estimate ±12% | 9.2 GBRuns great est. 37.1 tok/s32.7–41.6calibrated estimate ±12% | 8.7 GBRuns great est. 39.6 tok/s34.9–44.4calibrated estimate ±12% |
| 64k | 12.7 GBRuns great est. 27.0 tok/s23.8–30.3calibrated estimate ±12% | 10.7 GBRuns great est. 32.7 tok/s28.8–36.6calibrated estimate ±12% | 9.6 GBRuns great est. 36.8 tok/s32.4–41.2calibrated estimate ±12% |
| 128k | 17.6 GBRuns slowly est. 7.0 tok/s5.6–8.5calibrated estimate ±20% | 13.6 GBRuns great est. 26.4 tok/s23.2–29.5calibrated estimate ±12% | 11.4 GBRuns great est. 32.2 tok/s28.3–36.0calibrated estimate ±12% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
A full-attention model on the same card tells a different story. Qwen3 14B on the GeForce RTX 5060 Ti 16GB also Runs great at the 8k baseline. At its native maximum of 32k with an f16 cache it needs 15.2 GB against the same 15.4 GB: in CanRun's calculation it fits with little margin, so check that PROCESSOR in ollama ps shows 100% GPU. At the same 32k, Gemma 3 12B needs 10.2 GB, and it reaches 64k on this card with memory to spare.
What context does to speed
While everything stays in VRAM, the estimated speed falls steadily row by row (by about a third from 4k to 64k in the f16 column of the Gemma 3 12B table), because in CanRun's speed model each generated token reads the whole KV cache, and each row assumes the context is filled to that length. The sharp drop comes where the need passes the card's usable memory: in CanRun's calculation the 128k f16 cell of that table becomes a partial offload.
In Ollama, the default context follows your GPU memory, not the model's native maximum. Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). For how to check and change it, see our guide: Ollama slow or forgetting your prompt? Context length defaults.
KV cache quantization: what q8_0 and q4_0 save and cost
KV cache quantization stores the keys and values in blocks of 8-bit (q8_0) or 4-bit (q4_0) numbers instead of 16-bit ones. In CanRun's tables a q8_0 cache takes about 54% of the f16 size and a q4_0 cache about 29%, a little more than a half and a quarter, because each block also stores a scale. The Ollama FAQ gives approximately a half and a quarter.
This is separate from weight quantization. Q4_K_M or Q8_0, in capitals, name the model file you download, whose size is fixed; a q8_0 or q4_0 KV cache, in lower case, is a runtime setting that shrinks the memory the context takes, so its savings grow with the context. The two combine. For choosing the file itself, see our guide: GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download.
On quality, the Ollama FAQ describes the loss as very small for q8_0 and small-medium for q4_0 (more in this guide's FAQ below). CanRun does not measure quality, so the rule here is based on memory alone: when f16 does not fit, try q8_0 first, and q4_0 only if q8_0 still spills over, or shorten the context.
A 12GB card where the KV type decides
Take Qwen3 14B on the GeForce RTX 3060 12GB. At 8k with an f16 KV cache it Runs great. At 16k with f16 it needs 12.3 GB against 11.4 GB of usable memory, so in CanRun's calculation it becomes a partial offload; see the drop in that cell. A q8_0 cache brings 16k down to 11.1 GB, which fits with little margin. At the model's native 32k, q8_0 needs 12.7 GB, still over, and q4_0 needs 11.3 GB, just under with little margin. In both close cases, check that ollama ps shows 100% GPU.
| Context | KV f16 | KV q8_0 | KV q4_0 |
|---|---|---|---|
| 4k | 10.2 GBRuns great est. 26.1 tok/s22.9–29.2calibrated estimate ±12% | 9.9 GBRuns great est. 26.9 tok/s23.7–30.2calibrated estimate ±12% | 9.7 GBRuns great est. 27.4 tok/s24.1–30.7calibrated estimate ±12% |
| 8k | 10.9 GBRuns great est. 24.4 tok/s21.4–27.3calibrated estimate ±12% | 10.3 GBRuns great est. 25.9 tok/s22.8–29.0calibrated estimate ±12% | 10.0 GBRuns great est. 26.9 tok/s23.6–30.1calibrated estimate ±12% |
| 16k | 12.3 GBRuns slowly est. 14.6 tok/s11.7–17.5calibrated estimate ±20% | 11.1 GBRuns great est. 24.1 tok/s21.2–27.0calibrated estimate ±12% | 10.4 GBRuns great est. 25.8 tok/s22.7–28.9calibrated estimate ±12% |
| 32k | 15.2 GBRuns slowly est. 6.0 tok/s4.8–7.2calibrated estimate ±20% | 12.7 GBRuns slowly est. 12.8 tok/s10.3–15.4calibrated estimate ±20% | 11.3 GBRuns great est. 23.9 tok/s21.1–26.8calibrated estimate ±12% |
S = Runs great · A = Runs well · B = Runs slowly · F = Won't run
Compare other cards on the Qwen3 14B page and other models on the GeForce RTX 3060 12GB page.
Turning it on
Ollama quantizes the KV cache only when the server starts with OLLAMA_KV_CACHE_TYPE (f16, q8_0 or q4_0) and flash attention on; the setting applies to every model the server runs (Ollama docs, checked 2026-09-29). On Windows, quit the Ollama app from its taskbar icon, then run this in PowerShell for a one-off start that matches the 16k q8_0 cell above:
$env:OLLAMA_CONTEXT_LENGTH="16384"; $env:OLLAMA_FLASH_ATTENTION="1"; $env:OLLAMA_KV_CACHE_TYPE="q8_0"; ollama serve
Then, in a second PowerShell window, load the model with ollama run hf.co/unsloth/Qwen3-14B-GGUF:Q4_K_M. To keep the KV cache settings permanently, add OLLAMA_FLASH_ATTENTION and OLLAMA_KV_CACHE_TYPE (and OLLAMA_CONTEXT_LENGTH if you want 16k by default) under Edit environment variables for your account (no admin rights needed), then restart Ollama. On macOS, quit the Ollama app first; on Linux, stop the service with sudo systemctl stop ollama. The one-off form on both is OLLAMA_CONTEXT_LENGTH=16384 OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve.
In llama.cpp, add -fa on --cache-type-k q8_0 --cache-type-v q8_0 to llama-server; the server README lists the allowed cache types and gives f16 as the default. For the 16k q8_0 cell above, the combo page gives this command when you choose 16k and q8_0:
llama-server -hf unsloth/Qwen3-14B-GGUF:Q4_K_M -c 16384 -ngl all -fa on --cache-type-k q8_0 --cache-type-v q8_0
Treat KV cache quantization as a memory tool, not as a speed-up: where f16 also fits in VRAM, the extra speed the q8_0 and q4_0 cells show comes from CanRun's formula alone, since no measurement behind CanRun's calibration uses a quantized KV cache.
How sure are these numbers?
The file sizes are those of the real GGUF files, and the KV figures are calculated with a formula from each model's published configuration; the formula reproduces the known per-token sizes of the seven full-attention architectures checked in CanRun's engine tests. The compute buffer is an approximation drawn from llama.cpp logs. For sliding-window and hybrid models the KV figure is a simplification: CanRun leaves out the fixed window part and the recurrent state, so these figures run low, most visibly at short contexts; next to the model file the gap is small. llama.cpp's server documents a --swa-full switch, off by default, that keeps a full-size cache for the sliding-window layers.
Ollama makes its own memory estimate and GPU/CPU split, so on your machine ollama ps is the ground truth: SIZE for the memory in use, PROCESSOR for the GPU/CPU split, CONTEXT for the context allocated. CanRun's tables assume Windows with the card driving the display, 32GB of DDR5 system RAM, one model loaded and one request slot; the Ollama FAQ notes that the memory needed scales with OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH.
Speeds are estimates, each with a band and a confidence label. Calibrated means ±12% for dense models running fully on an NVIDIA or AMD card, and ±20% for partial offload, MoE models and other hardware; theoretical means ±30%. Every speed cell in this guide's two context tables is calibrated. The q8_0 and q4_0 cells carry the same label, but the effect of the smaller KV cache on their speed rests on the formula alone, since no public measurement CanRun calibrates against records a quantized KV cache. On combo pages for MoE models at long contexts, cells that keep the experts in system RAM carry the theoretical label, and their estimates are likely lower than real speeds, so do not read them as results.
Verdict names in the text of this guide refer to the 8k context, f16 KV baseline; for other contexts, read the tables. The thresholds:
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Where to go next:
- Ollama slow or forgetting your prompt? Context length defaults: setting the context length in Ollama.
- How much VRAM do local LLMs need? 8–32 GB tiers (2026): which models fit each memory tier.
- GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download: how a smaller model file frees room for context.
- Running local LLMs on a Mac (M4, M5): how much unified memory?: on a Mac, the KV cache shares unified memory with everything else.
FAQ
Does a longer context use more memory even before the chat gets long?
Yes. Memory follows the context length you set, not how much of it you have used: Ollama's context-length documentation says a larger context length increases the memory required. With an f16 cache, Llama 3.1 8B's KV cache is 1.1 GB at 8k and 17.2 GB at 128k. Set the context you actually need, then check CONTEXT and SIZE in ollama ps. While everything stays on the GPU, an unfilled context slows each token less than its table row shows, since the rows assume a full context; a spill to system RAM slows it regardless.
Why can a bigger model need less memory for its context than a smaller one?
Because the KV cache depends on the attention design (the layers, the KV heads, and whether some layers use sliding windows or linear attention), not on the parameter count. At 32k with f16, gpt-oss-120b needs 1.2 GB of KV cache while Llama 3.1 8B needs 4.3 GB. MoE alone does not help: the MoE Kanana 1.5 15.7B-A3B costs the same per token as the dense Llama 3.1 8B.
Is KV cache quantization the same as choosing a Q4_K_M model file?
No. Q4_K_M is weight quantization and fixes the size of the model file; a q8_0 or q4_0 KV cache shrinks the memory the context takes at run time, so its savings grow with the context. The two are independent and combine. Turn it on with OLLAMA_KV_CACHE_TYPE in Ollama (server-wide, with flash attention on) or with -fa on --cache-type-k q8_0 --cache-type-v q8_0 in llama.cpp.
How much quality do q8_0 and q4_0 KV caches cost?
According to the Ollama FAQ, q8_0 has a very small loss in precision, and q4_0 a small-medium loss that may be more noticeable at higher context sizes; the impact depends on the model and the task, and models with a high GQA count (its example is Qwen2) may be affected more. CanRun does not measure quality, so test your own prompts at your target context.
Sources
- Ollama FAQ: K/V cache quantization (f16, q8_0 and q4_0 memory use and precision notes, a server-wide option that needs flash attention, impact that depends on the model and task, the high-GQA note with Qwen2),
OLLAMA_FLASH_ATTENTION, and memory scaling withOLLAMA_NUM_PARALLELtimesOLLAMA_CONTEXT_LENGTH. - Ollama docs: context length: a larger context length increases the memory required, the default context by VRAM, the
ollama pscolumns (SIZE, PROCESSOR, CONTEXT), and the advice to avoid CPU offload. - llama.cpp server README:
--cache-type-kand--cache-type-vwith the allowed types and the f16 default,-fa,-c, and--swa-full(a full-size sliding-window cache, off by default). - Llama 3.1 8B Instruct model card: grouped-query attention in all Llama 3.1 versions, 128k context.
- Llama 3.1 8B Instruct configuration: layers, attention and KV heads, and the hidden size the head size follows from (a gated repository: the file opens once you accept the license on Hugging Face).
- Qwen3-8B model card: layers, query and KV heads, native context and YaRN extension.
- Qwen3-14B model card: layers, query and KV heads, native context and YaRN extension.
- Qwen3-14B configuration: layers, KV heads, head size.
- Qwen3-32B model card: layers, query and KV heads, native context.
- Qwen3-32B configuration: layers, KV heads, head size.
- Qwen3-30B-A3B-Instruct-2507 model card: layers, query and KV heads, experts, native context.
- Qwen3.5-9B model card: the hybrid layout of Gated DeltaNet and Gated Attention layers, KV heads of the attention layers, native context.
- Gemma 3 Technical Report: cutting long-context KV-cache memory by raising the ratio of local to global attention layers and keeping the local span short.
- Gemma 4 12B configuration: sliding-window and full-attention layers, 1024-token window.
- gpt-oss-20b configuration: sliding-attention layers with a 128-token window alternating with full-attention layers, KV heads, head size, maximum context.
- EXAONE 4.0 32B model card: hybrid attention for the 32B model, with local (sliding-window) and global layers in a 3:1 ratio, layers, GQA, maximum context.
- EXAONE 4.5 33B configuration: local and global layers in a 3:1 pattern, 4096-token window.
Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.