Ollama slow or forgetting your prompt? Context length defaults

By CanRun · Updated September 30, 2026

Out of the box, Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). That puts 8GB, 12GB and 16GB graphics cards on a 4k window, and when a chat grows past it Ollama quietly drops the oldest messages, so the model seems to forget what you told it. On 24GB cards the 32k default can instead push a dense 32B model out of VRAM: Qwen3 32B at Q4_K_M, a 4-bit quant, needs 22.5 GB at 8k, which fits in the 23.4 GB a 24GB card can use. At 32k the KV cache (the memory that holds the conversation so far) is four times larger and the total reaches 29.1 GB, so the overflow goes to system RAM, and that is where the slowdown comes from. The fix in both cases is to set the context yourself with OLLAMA_CONTEXT_LENGTH or num_ctx, choose the largest size that still fits in VRAM, and check the result with ollama ps.

Ollama's default context depends on your GPU memory

Ollama has no single fixed context length. Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). The help text for OLLAMA_CONTEXT_LENGTH lists the same three tiers, and nothing in a normal chat tells you which one you got.

Among the graphics cards in our catalog, the GeForce RTX 4060 (8GB), the GeForce RTX 3060 12GB and the GeForce RTX 5060 Ti 16GB land in the 4k tier. The GeForce RTX 3090 and the GeForce RTX 4090 (24GB each) and the GeForce RTX 5090 (32GB) land in the 32k tier. No card we list reaches the 256k tier; the largest has 32GB. We have not verified how Ollama detects memory on a Mac with unified memory, so Macs are left out of this mapping.

A 4k window is small next to what current models support. Llama 3.1 8B handles 128k and Qwen3 14B handles 32k natively (Qwen's model card describes going further with YaRN).

What happens when a chat outgrows the window

Ollama does not stop with an error. By default its server drops the oldest messages until the rest fits, and it always keeps the system messages and your latest message. If the reply itself fills the window, Ollama shifts the window forward and keeps generating, again by default. Nothing shows up in the chat: dropped messages are logged only at debug level, so instructions you gave early on, or a document you pasted a few turns back, quietly leave the model's view. A single message longer than the window is cut from the start as well, with only a warning in the server log.

Reasoning models fill the window quickly. Their thinking takes up the same window as your messages, and Qwen3's model card recommends an output length of 32k tokens for most queries, far more than a 4k window holds.

Some older guides still give a flat 2k or 4k default (see the FAQ below). To see what you actually have, run ollama ps while a model is loaded: the CONTEXT column shows the context allocated, and PROCESSOR shows the split between GPU and CPU.

What a longer context costs in memory

A model's memory need has three parts: the model file at the quant our tables use as a baseline (Q4_K_M for every model below except gpt-oss-20b, which uses MXFP4), the KV cache, and a small compute buffer, the working memory the runtime needs during inference. The file does not change with context. The KV cache holds keys and values for every token in the window, so for a model with full attention on every layer it grows in step with the context: double the context, double the cache.

How fast it grows depends on the architecture. With an f16 KV cache, Ollama's default type, 32k takes 4.3 GB for Llama 3.1 8B but only 0.8 GB for gpt-oss-20b, whose sliding-window layers keep only a short stretch of recent tokens. EXAONE 4.0 32B mixes sliding-window (local) and global attention layers in a 3:1 ratio, according to its model card, so its 32k total sits closer to Gemma 3 27B's than to Qwen3 32B's. The mechanics are explained in KV cache explained: how context length eats VRAM.

Here are the totals at 32k, the context Ollama picks by default on 24GB and 32GB cards.

Memory a model needs at 32k context: file of the default quant + f16 KV cache + compute buffer
ModelParametersDefault quant fileKV cache at 32kTotal at 32k
Llama 3.1 8B8.03BQ4_K_M · 4.9 GB4.3 GB10.0 GB
Gemma 3 12B12.2BQ4_K_M · 7.3 GB2.1 GB10.2 GB
Qwen3 14B14.8BQ4_K_M · 9.0 GB5.4 GB15.2 GB
gpt-oss-20b20.9B (3.6B active)MXFP4 · 12.1 GB0.8 GB13.7 GB
Gemma 3 27B27.4BQ4_K_M · 16.6 GB2.7 GB20.0 GB
Qwen3 30B-A3B (2507)30.5B (3.3B active)Q4_K_M · 18.6 GB3.2 GB22.6 GB
EXAONE 4.0 32B32BQ4_K_M · 19.3 GB2.1 GB22.3 GB
Qwen3 32B32.8BQ4_K_M · 19.8 GB8.6 GB29.1 GB

On a graphics card the OS also keeps some VRAM (about 0.6 GB on a Windows display GPU); on a Mac the GPU can use about 70% of unified memory by default.

Compare each total with what your card can use, not the number on the box: a card driving a Windows display keeps some VRAM back (see the note under the table), so an 8GB card has 7.4 GB in our model.

An 8GB card from 4k to 32k

Take Llama 3.1 8B on the GeForce RTX 4060, a common first setup. At Ollama's 4k default it needs 6.0 GB, and at 8k it needs 6.6 GB. Both fit in the 7.4 GB the card can use, and at the 8k baseline it Runs great.

At 16k with an f16 KV cache the need rises to 7.7 GB, just over the usable amount, and 32k takes it to 10.0 GB. Past that line our model moves part of the weights to system RAM (partial offload). In the 64k f16 cell and the 128k f16 and q8_0 cells, the KV cache alone is larger than the card's usable memory, so our model runs those on the CPU only. The table shows what each does to the verdict and the speed band, and its q8_0 and q4_0 columns show the other way out, a smaller KV cache.

Llama 3.1 8B on the GeForce RTX 4060 (Q4_K_M): memory and verdict by context length and KV cache type
ContextKV f16KV q8_0KV q4_0
4k
6.0 GBRuns great
est. 34.9 tok/s30.7–39.1calibrated estimate ±12%
5.7 GBRuns great
est. 36.6 tok/s32.2–41.0calibrated estimate ±12%
5.6 GBRuns great
est. 37.5 tok/s33.0–42.0calibrated estimate ±12%
8k
6.6 GBRuns great
38.1 tok/s8k estimate 31.8 tok/s (28.0–35.6, calibrated estimate ±12%)measured (1 run, 4k context)
6.1 GBRuns great
est. 34.7 tok/s30.5–38.8calibrated estimate ±12%
5.8 GBRuns great
est. 36.4 tok/s32.1–40.8calibrated estimate ±12%
16k
7.7 GBRuns well
est. 22.3 tok/s17.8–26.7calibrated estimate ±20%
6.7 GBRuns great
est. 31.4 tok/s27.6–35.1calibrated estimate ±12%
6.2 GBRuns great
est. 34.4 tok/s30.3–38.5calibrated estimate ±12%
32k
10.0 GBRuns slowly
est. 7.6 tok/s6.1–9.1calibrated estimate ±20%
8.0 GBRuns slowly
est. 18.7 tok/s15.0–22.5calibrated estimate ±20%
6.9 GBRuns great
est. 31.0 tok/s27.3–34.7calibrated estimate ±12%
64k
13.5 GBRuns slowly
est. 3.3 tok/s2.7–4.0calibrated estimate ±20%
10.6 GBRuns slowly
est. 6.4 tok/s5.1–7.7calibrated estimate ±20%
8.5 GBRuns slowly
est. 15.2 tok/s12.1–18.2calibrated estimate ±20%
128k
22.1 GBRuns slowly
est. 2.0 tok/s1.6–2.4calibrated estimate ±20%
14.1 GBRuns slowly
est. 3.2 tok/s2.5–3.8calibrated estimate ±20%
11.5 GBRuns slowly
est. 5.2 tok/s4.2–6.3calibrated estimate ±20%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

A 12GB card has more headroom. For Gemma 3 12B on the GeForce RTX 3060 12GB, 32k with an f16 KV cache needs 10.2 GB, still inside the 11.4 GB the card can use, and at the 8k baseline it Runs great. Every card we cover for Llama 3.1 8B is listed on the Llama 3.1 8B page.

Why Ollama slows down: the 32k default on 24GB cards

A 24GB card already gets 32k by default, and for a dense 32B model with full attention on every layer, such as Qwen3 32B, that is more than the card can hold. At Q4_K_M and 8k, Qwen3 32B needs 22.5 GB, which fits in the 23.4 GB an RTX 3090 can use, and at that baseline it Runs great. At 32k with an f16 KV cache, the KV cache alone is 8.6 GB and the total is 29.1 GB, so part of the weights have to sit in system RAM.

Ollama makes that split on its own, loading what fits onto the GPU and running the remaining layers on the CPU; ollama ps then shows a CPU/GPU percentage under PROCESSOR instead of 100% GPU. Ollama's documentation advises the largest context that avoids offloading to the CPU. The table below shows where that line falls on the RTX 3090, and the Qwen3 32B page compares other cards.

Qwen3 32B on the GeForce RTX 3090 (Q4_K_M): memory and verdict by context length and KV cache type
ContextKV f16KV q8_0KV q4_0
4k
21.4 GBRuns great
est. 31.4 tok/s27.7–35.2calibrated estimate ±12%
20.9 GBRuns great
est. 32.2 tok/s28.4–36.1calibrated estimate ±12%
20.6 GBRuns great
est. 32.7 tok/s28.7–36.6calibrated estimate ±12%
8k
22.5 GBRuns great
est. 29.9 tok/s26.3–33.5calibrated estimate ±12%
21.5 GBRuns great
est. 31.3 tok/s27.6–35.1calibrated estimate ±12%
20.9 GBRuns great
est. 32.2 tok/s28.3–36.0calibrated estimate ±12%
16k
24.7 GBRuns well
est. 14.3 tok/s11.5–17.2calibrated estimate ±20%
22.7 GBRuns great
est. 29.7 tok/s26.1–33.3calibrated estimate ±12%
21.6 GBRuns great
est. 31.2 tok/s27.5–35.0calibrated estimate ±12%
32k
29.1 GBRuns slowly
est. 4.7 tok/s3.7–5.6calibrated estimate ±20%
25.2 GBRuns well
est. 12.2 tok/s9.7–14.6calibrated estimate ±20%
23.0 GBRuns great
est. 29.5 tok/s26.0–33.0calibrated estimate ±12%

S = Runs great · A = Runs well · B = Runs slowly · F = Won't run

Not every model hits this wall. At 32k with an f16 KV cache, Gemma 3 27B needs 20.0 GB, Qwen3 30B-A3B (2507) needs 22.6 GB and EXAONE 4.0 32B needs 22.3 GB, all under the RTX 3090's usable memory, and each Runs great at the 8k baseline. Our page for Qwen3 30B-A3B (2507) on the GeForce RTX 3090 shows how that MoE model fares at longer contexts. The GeForce RTX 4090 has the same 24GB and the same default, so the same arithmetic applies. The RTX 5090 is in the same tier, but its 31.4 GB of usable memory covers the 29.1 GB Qwen3 32B needs at 32k, so in our model Qwen3 32B stays fully on the GPU there at 32k (check PROCESSOR in ollama ps); its 8k baseline also Runs great.

Two fixes that keep the model on the GPU

The first fix is a smaller context for that model. At 16k, Qwen3 32B needs 24.7 GB, still over the RTX 3090's usable memory; 8k fits.

The second fix is a smaller KV cache. Ollama quantizes the KV cache only when the server starts with OLLAMA_KV_CACHE_TYPE (f16, q8_0 or q4_0) and flash attention on; the setting applies to every model the server runs (Ollama docs, checked 2026-09-29). According to the Ollama FAQ, q8_0 costs very little precision, while q4_0 costs more and the loss may be more noticeable at long contexts; how much it matters depends on the model and the task, and the FAQ adds that models with many query heads per KV head, such as Qwen2, may be more affected. In our tables q8_0 and q4_0 KV caches take about 54% and 29% of the f16 size, slightly more than a half and a quarter, because the block scales are stored too. For Qwen3 32B at 32k, q8_0 brings the need to 25.2 GB, still over, and q4_0 brings it to 23.0 GB, just under in our model, with little margin, so confirm 100% GPU in ollama ps after starting. The two fixes also combine: 16k with a q8_0 KV cache needs 22.7 GB, which fits in our model. A one-off start in PowerShell looks like this:

$env:OLLAMA_CONTEXT_LENGTH="32768"; $env:OLLAMA_FLASH_ATTENTION="1"; $env:OLLAMA_KV_CACHE_TYPE="q4_0"; ollama serve

Inside VRAM, the estimated speed eases down row by row as the context grows, because each new token reads the whole KV cache. The rows assume the window is filled to that length. A larger window that you never fill still reserves the memory, but it does not slow each token as much as its row shows. The sharp drop comes where the need passes the card's usable memory.

How to change the context length in Ollama

Ollama documents five ways to set the context:

  • Ollama app: the context length slider in Settings.
  • Server-wide default: the environment variable OLLAMA_CONTEXT_LENGTH, for example OLLAMA_CONTEXT_LENGTH=16384 ollama serve on macOS or Linux.
  • One interactive session: type /set parameter num_ctx 16384 inside ollama run.
  • One API request: put num_ctx inside options, as below.
  • One model, every time: add PARAMETER num_ctx 16384 to a Modelfile and build it with ollama create.

On macOS or Linux, a request looks like this (on Windows, run it in WSL or Git Bash; Windows PowerShell treats curl as a different command):

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "messages": [{ "role": "user", "content": "Why is the sky blue?" }],
  "options": { "num_ctx": 16384 }
}'

OLLAMA_CONTEXT_LENGTH only replaces the automatic default. A num_ctx set in a Modelfile or sent with a request overrides it, and a request overrides the Modelfile. If ollama ps shows a smaller CONTEXT than you set on the server, look for a num_ctx in the Modelfile or the client sending requests.

Setting the variable on Windows, macOS and Linux

Ollama reads the variable when the server starts, so restart Ollama after any change.

  • Windows: quit Ollama from its taskbar icon, open Settings and search for environment variables, edit the variables for your account, add OLLAMA_CONTEXT_LENGTH with a value such as 16384, save, and start Ollama again from the Start menu. For a one-off test with the app closed, run $env:OLLAMA_CONTEXT_LENGTH="16384"; ollama serve in PowerShell, the form our combo pages print.
  • macOS app: run launchctl setenv OLLAMA_CONTEXT_LENGTH 16384, then restart the Ollama app.
  • Linux with systemd: run systemctl edit ollama.service, add the lines below, then run systemctl daemon-reload and systemctl restart ollama.
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=16384"

Picking a size that fits

Aim for the largest context whose total need stays under your card's usable memory, using the tables above or each combo page. Ollama's docs suggest at least 64k for agents, web search and coding tools. For scale, Llama 3.1 8B at 64k with an f16 KV cache needs 14.6 GB.

The 4k default can also leave memory unused. Running Qwen3 14B on the GeForce RTX 5060 Ti 16GB at its full 32k with an f16 KV cache needs 15.2 GB against 15.4 GB usable, so in our model it only just fits, with little margin; check ollama ps for 100% GPU. At the 8k baseline it Runs great. Other cards are compared on the Qwen3 14B page.

Going past a model's own limit does not help. Qwen3 32B's native limit is 32k; its model card describes extending it with YaRN, which is outside this guide. Parallel requests multiply the context memory: the Ollama FAQ notes that the memory needed scales with OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH, while our tables assume one request slot.

Each combo page prints an Ollama command whose OLLAMA_CONTEXT_LENGTH follows the context you pick in its table, and for q8_0 or q4_0 cells it adds OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE. Our page for Llama 3.1 8B on the GeForce RTX 4060 is an 8GB example.

Where to go next:

How sure are these numbers?

The memory figures (file, KV cache, compute buffer) are deterministic: they are calculated from real GGUF file sizes and each model's published configuration. The speeds in the tables are estimates, each with a band and a confidence label. Calibrated means ±12% for dense models running fully on an NVIDIA or AMD card, and ±20% for partial offload, MoE models and other hardware; theoretical means ±30%. Every cell in this guide's two context tables is calibrated (±12% when fully on the GPU, ±20% for partial-offload or CPU-only cells), except the 8k f16 cell for Llama 3.1 8B on the RTX 4060, which is measured.

That measurement is one public data point for the GeForce RTX 4060. It comes from LocalScore at a 4k context, not from an Ollama run, and it is shown next to our 8k estimate band.

Measured: Llama 3.1 8B (Q4_K_M) on the GeForce RTX 4060
38.1 tok/s8k estimate 31.8 tok/s (28.0–35.6, calibrated estimate ±12%)measured (1 run, 4k context)
Source: localscore.ai

Ollama makes its own memory estimate and picks its own GPU/CPU split, so the SIZE and PROCESSOR values in ollama ps will not match our figures exactly. On your machine, ollama ps is the ground truth. Our tables assume Windows with the card driving the display, 32GB of DDR5 system RAM, one model loaded, one request slot (Ollama's default for OLLAMA_NUM_PARALLEL, per its FAQ) and flash attention for the quantized KV columns. Every combo page, including our page for Llama 3.1 8B on the GeForce RTX 4060, lists the same assumptions.

Verdict names in the text of this guide refer to the 8k context, f16 KV baseline; for other context lengths, read the tables. The thresholds:

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

FAQ

Why does Ollama forget what I said earlier in a long chat?

The context window is full. Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29), so 8GB, 12GB and 16GB cards start at only 4k. Once a chat outgrows it, Ollama by default drops the oldest messages, keeping the system prompt and your latest message, with no error. Raise the context with OLLAMA_CONTEXT_LENGTH, the app slider or num_ctx as far as memory allows, then check CONTEXT in ollama ps.

Does a longer context make Ollama slower?

The big slowdown comes when the larger KV cache no longer fits in VRAM and Ollama moves part of the model to system RAM; ollama ps then shows a CPU/GPU split. Qwen3 32B at Q4_K_M, for example, needs 29.1 GB at 32k with an f16 KV cache, against 23.4 GB usable on a 24GB card, where 32k is already the default. On the GPU alone, the estimated speed falls only gradually as the window fills.

How much extra memory does a 32k context need?

It depends on the model's KV cache per token. With an f16 KV cache, 32k costs 4.3 GB for Llama 3.1 8B and 8.6 GB for Qwen3 32B, but only 0.8 GB for gpt-oss-20b with its sliding-window layers. Add the model file and a compute buffer for the total, as in the 32k memory table, and compare it with your card's usable memory.

Should I turn on q8_0 or q4_0 KV cache in Ollama?

Only when you need the memory: OLLAMA_KV_CACHE_TYPE is server-wide, needs flash attention and applies to every model the server runs. q8_0 roughly halves the KV cache with a very small precision loss; q4_0 saves more, but the Ollama FAQ notes its loss may show at long contexts. A memory-based rule: try q8_0 first when f16 does not fit, and use q4_0 only if q8_0 still spills. It can decide whether a model fits: on a 24GB card such as the RTX 3090, Qwen3 32B at 32k needs 29.1 GB with f16 and 25.2 GB with q8_0, both over the 23.4 GB the card can use, and 23.0 GB with q4_0, just under.

Why do other guides say Ollama's default is 2k or 4k tokens?

Older Ollama versions and some reference pages list one flat default: the Modelfile parameter table shows 2k and the FAQ mentions 4k. The context-length documentation now ties it to VRAM. Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). The CONTEXT column of ollama ps shows what your machine actually allocated.

Sources

Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.