GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download

By CanRun · Updated September 30, 2026

Start with Q4_K_M: llama.cpp's -hf option and Hugging Face's Ollama integration pick it when you name no quant, and our verdicts use it as the baseline. For the dense models in this guide, IQ4_XS is about a tenth smaller and in the same quality group as Q4_K_M (Good balance); Q8_0 is roughly three quarters larger, worth it only when it fits with room for your context. One step down can decide the fit: on the RTX 5060 Ti 16GB with an 8k context, Mistral Small 3.2 24B at Q4_K_M needs 16.2 GB, more than the 15.4 GB the card can use. At IQ4_XS it needs 14.7 GB, which in our calculation stays on the GPU, with little margin — check ollama ps.

What the parts of a GGUF quant name mean

A quantized GGUF file stores a model's weights at reduced precision, and the quant name says which scheme was used. The table shows the four files we track for Qwen3 14B, from the unsloth repository.

Qwen3 14B: GGUF files by quantization
QuantBits per weightFile sizeQuality groupGGUF source
Q2_K3.165.75 GBHeavy lossCommunity GGUF · unsloth
IQ4_XS4.468.14 GBGood balanceCommunity GGUF · unsloth
Q4_K_M4.899.00 GBGood balanceCommunity GGUF · unsloth
Q8_08.5015.70 GBNear-losslessCommunity GGUF · unsloth

Taking Q4_K_M apart:

  • Q4: the digit is the nominal bit width.
  • K: a k-quant. Weights are stored in super-blocks that contain smaller blocks, and each block has its own scale (llama.cpp PR #1684).
  • M: the size mix, one of S, M or L. M uses a higher-bit type for some of the more sensitive tensors, and which tensors get it has changed since the original PR.

Other names follow different schemes:

  • IQ4_XS is an i-quant, a type whose weights are derived using an importance matrix. Its PR (#5747) describes it as IQ4_NL in super-blocks of 256 weights with 6-bit block scales; XS is the smallest variant.
  • Q8_0 is plain 8-bit quantization in blocks of 32 weights, with one scale per block.
  • MXFP4 is a 4-bit microscaling block floating-point format. Per the gpt-oss-20b model card, the gpt-oss models were post-trained with MXFP4 quantization of their mixture-of-experts (MoE) weights, so it is the only file we list for gpt-oss-20b (12.1 GB) and sits in our Near-lossless group as the release precision.
  • UD- files, such as UD-Q2_K_XL, are Unsloth Dynamic quants, which choose a quant type per layer, so their bits per weight cannot be read from the name; as for every file, we use the real file size.

The bits-per-weight column gives the file-level values llama.cpp's quantize README lists for Llama 3.1 8B. They are higher than the digit in the name because block scales and the higher-bit tensors add bits, and real files for another model differ a little. As a rough rule, file size in gigabytes is billions of parameters times bits per weight, divided by eight, except for MXFP4 and UD files.

In the Qwen3 14B table, IQ4_XS is about a tenth smaller than Q4_K_M, Q8_0 roughly three quarters larger, and Q2_K, a 2-bit file in the Heavy loss group, about two thirds the size. The same ratios hold for Mistral Small 3.2 24B and Gemma 3 12B in our data, but not for every model, including some MoE models and models whose Q4_K_M row is a UD file, so check each model's own table.

Which file gets downloaded

llama-server -hf picks Q4_K_M when you give no quant (or the first file in the repo if it has no Q4_K_M file), and Hugging Face's Ollama integration uses Q4_K_M when the repo has it.

To get a different file, append the quant as a case-insensitive tag. Both commands run unchanged in PowerShell on Windows, and -c 8192 matches the 8k context our numbers assume:

ollama run hf.co/unsloth/Qwen3-14B-GGUF:IQ4_XS
llama-server -hf unsloth/Qwen3-14B-GGUF:IQ4_XS -c 8192

The ollama run line does not set a context. Ollama picks its default context length from the GPU memory it finds: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more (Ollama docs, checked 2026-09-29). To match our 8k, quit the Ollama tray app, start the server with $env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve as our combo pages do, and run the command from a second PowerShell window.

Quality groups, and why 2-bit files exist

We sort every quant type into one of four quality groups:

  • Near-lossless: Q8_0, Q6_K, Q5_K_M, F16, and MXFP4 where it is the model's native format.
  • Good balance: Q4_K_M and IQ4_XS.
  • Noticeable loss: the 3-bit Q3_K_M and IQ3_XS.
  • Heavy loss: the 2-bit Q2_K, IQ2_M, UD-Q2_K_XL and UD-IQ2_M.

The grouping is our policy by quant type, not a quality measurement for each model; real loss depends on the model and the task. No model in this guide's tables has a 3-bit file in our data, so Noticeable loss does not appear below.

The ordering follows llama.cpp's evidence. Its quantize README says quantization may cost accuracy, measured with perplexity or KL divergence (scores of how far the quantized model's predictions drift from the original), and that an importance-matrix (imatrix) file can reduce the loss. The 2023 PR that introduced k-quants, tested on LLaMA models from 7B to 65B, found perplexity a fairly smooth function of file size, with the 6-bit type very close to fp16. Relative error also stopped shrinking as models grew, back to about the 7B level at 30B and 65B, so a large model is not automatically safe at 2-bit.

Why 2-bit quants exist

The reason is file size. Llama 3.3 70B at Q4_K_M is a 42.5 GB file, and even its 2-bit IQ2_M file, labeled Heavy loss, is 24.1 GB, more than the 23.4 GB a 24GB card can use.

Llama 3.3 70B: GGUF files by quantization
QuantBits per weightFile sizeQuality groupGGUF source
IQ2_M2.9324.12 GBHeavy lossCommunity GGUF · bartowski
Q2_K3.1626.38 GBHeavy lossCommunity GGUF · bartowski
IQ4_XS4.4637.90 GBGood balanceCommunity GGUF · bartowski
Q4_K_M4.8942.52 GBGood balanceCommunity GGUF · bartowski
Q8_08.5074.98 GBNear-losslessCommunity GGUF · bartowski

At the Q4_K_M baseline, Llama 3.3 70B Won't run on the GeForce RTX 3090. The reason is speed, not a failure to load: its need of 45.8 GB is about twice the card's usable memory, so in our calculation roughly half the model would sit in system RAM, and that part sets the pace. One step down does not bring it onto the card either: IQ4_XS still needs 41.2 GB, so in our calculation it stays a partial offload, though at IQ4_XS the pair moves up to Runs slowly, which is why our pages suggest IQ4_XS for this pair and print their commands with it.

How we suggest a different quant

Our pages suggest trying a smaller quant only when it raises the verdict, and a 2-bit file only when every non-2-bit file Won't run. When a 2-bit file would reach a higher verdict than that suggestion, or than the baseline when there is none, the page mentions it separately as a last resort, marked as heavy quality loss. We never suggest a larger quant that gets the same verdict, since it costs speed and context room for the same outcome; a combo page may still name one as an option.

When one step down changes the result

A rule of thumb follows from the file sizes: dropping from Q4_K_M to IQ4_XS brings a model fully onto the GPU only when the Q4_K_M need is over your usable memory by less than the gap between the two files, about a tenth of the file here.

Mistral Small 3.2 24B on the RTX 5060 Ti 16GB

Mistral Small 3.2 24B on the GeForce RTX 5060 Ti 16GB is the case where one step is enough. At Q4_K_M and 8k it needs 16.2 GB against 15.4 GB usable, so in our calculation a small part is offloaded to system RAM (a partial offload), and at that baseline it Runs well. At IQ4_XS it needs 14.7 GB, which in our calculation fits fully on the GPU, with little margin — check ollama ps. Mistral Small 3.2 is also a vision model, and our figures leave out its vision projector (mmproj) file, which llama-server -hf fetches and loads when the repo has one (add --no-mmproj to skip it); at this margin that can decide the fit. In the ladder, the IQ4_XS row moves up a verdict and the Q8_0 row goes the other way.

Mistral Small 3.2 24B on the GeForce RTX 5060 Ti 16GB: verdict by quantization (8k context, f16 KV cache)
QuantQuality groupVerdictSpeedMemory
Q8_0Near-losslessRuns slowly
est. 3.2 tok/s2.5–3.8calibrated estimate ±20%
16.0 / 16.0 GB + 11.6 GB RAM
Q4_K_MGood balanceRuns well
est. 14.8 tok/s11.8–17.7calibrated estimate ±20%
16.0 / 16.0 GB + 0.8 GB RAM
IQ4_XSGood balanceRuns great
est. 22.2 tok/s19.6–24.9calibrated estimate ±12%
15.3 / 16.0 GB
Q2_KHeavy lossRuns great
est. 30.6 tok/s27.0–34.3calibrated estimate ±12%
11.4 / 16.0 GB

The fit is the same on every 16GB card in our catalog, but whether the IQ4_XS row also moves up a verdict depends on the card's speed, so check your own card in the calculator.

Context changes the picture: at 16k, IQ4_XS needs 16.1 GB, over the usable memory again. A q8_0 KV cache (the memory that holds the context, stored at 8-bit instead of f16) brings that to 14.8 GB, which fits in CanRun's calculation, again with little margin — check ollama ps.

For this pair, the commands CanRun prints use the Q4_K_M tag (they switch to a smaller quant only when Q4_K_M Won't run), so swap in IQ4_XS yourself. For our 8k context, quit the Ollama tray app, run the first line, then the second in another PowerShell window:

$env:OLLAMA_CONTEXT_LENGTH="8192"; ollama serve
ollama run hf.co/bartowski/mistralai_Mistral-Small-3.2-24B-Instruct-2506-GGUF:IQ4_XS

Gemma 3 12B on the RTX 4060

For Gemma 3 12B on the GeForce RTX 4060, Q4_K_M needs 8.4 GB against 7.4 GB usable. In our calculation, less than 90% of the model stays on the GPU, so at the baseline it Runs slowly.

IQ4_XS needs 7.7 GB, still over the usable memory, so it is still a partial offload. What changes in our calculation is the split: more than 90% of the model now stays on the GPU, and the ladder row moves up one verdict. The Q2_K row needs 5.9 GB and fits entirely, but as a 2-bit Heavy loss file it is only a last resort. Our pages suggest trying IQ4_XS instead.

Gemma 3 12B on the GeForce RTX 4060: verdict by quantization (8k context, f16 KV cache)
QuantQuality groupVerdictSpeedMemory
Q8_0Near-losslessRuns slowly
est. 5.6 tok/s4.5–6.7calibrated estimate ±20%
8.0 / 8.0 GB + 6.2 GB RAM
Q4_K_MGood balanceRuns slowly
est. 16.7 tok/s13.4–20.1calibrated estimate ±20%
8.0 / 8.0 GB + 1.0 GB RAM
IQ4_XSGood balanceRuns well
est. 23.8 tok/s19.0–28.5calibrated estimate ±20%
8.0 / 8.0 GB + 0.3 GB RAM
Q2_KHeavy lossRuns great
est. 35.9 tok/s31.6–40.2calibrated estimate ±12%
6.5 / 8.0 GB

When one step down is not enough

For Qwen3 14B on the GeForce RTX 4060, the gap is too wide. At the baseline it Runs slowly, and IQ4_XS still needs 10.1 GB, far over the 7.4 GB the card can use, so one step down does not change the verdict. Only the Heavy loss Q2_K comes close, and even it needs 7.7 GB, still over.

One caveat for MoE models: where a ladder row keeps the MoE experts in system RAM, as the Q4_K_M and IQ4_XS rows do for Qwen3 30B-A3B (2507) on the RTX 5060 Ti 16GB, those rows are theoretical (±30%) and likely conservative, since that path has no measurements to check against yet.

When a larger file is worth the memory

A Near-lossless file such as Q8_0 or Q6_K keeps more precision, but only pays off when it fits with room for your context.

Qwen3 14B on the GeForce RTX 5060 Ti 16GB shows the limit. At the Q4_K_M baseline it Runs great. Q8_0 needs 17.6 GB at 8k, over the 15.4 GB the card can use, so the ladder's Q8_0 row becomes a partial offload.

Qwen3 14B on the GeForce RTX 5060 Ti 16GB: verdict by quantization (8k context, f16 KV cache)
QuantQuality groupVerdictSpeedMemory
Q8_0Near-losslessRuns slowly
est. 10.0 tok/s8.0–12.0calibrated estimate ±20%
16.0 / 16.0 GB + 2.2 GB RAM
Q4_K_MGood balanceRuns great
est. 30.3 tok/s26.7–34.0calibrated estimate ±12%
11.5 / 16.0 GB
IQ4_XSGood balanceRuns great
est. 33.1 tok/s29.1–37.0calibrated estimate ±12%
10.7 / 16.0 GB
Q2_KHeavy lossRuns great
est. 44.2 tok/s38.9–49.5calibrated estimate ±12%
8.3 / 16.0 GB

On a 24GB card the same file has room. For Qwen3 14B on the GeForce RTX 3090, the Q8_0 need of 17.6 GB is under the 23.4 GB the card can use, and the Q4_K_M baseline Runs great.

What it costs

The first cost is speed. For a dense model, generating each token reads all of the weights, so a file roughly three quarters larger is estimated to be clearly slower. The quantize README's table for Llama 3.1 8B also shows Q8_0 slower than Q4_K_M, and the ladder rows show our band for each quant.

The second cost is context room. For Llama 3.1 8B on the GeForce RTX 3060 12GB, the Q4_K_M baseline Runs great. Q8_0 needs 10.2 GB at 8k, under the 11.4 GB the card can use. At 16k it needs 11.3 GB, which fits with little margin — check ollama ps. At 32k it needs 13.6 GB, over the usable memory, while Q4_K_M at 32k needs 10.0 GB and fits.

A memory-based rule: take a Near-lossless file only when its need at the context you actually use stays under your usable memory. Otherwise use Q4_K_M, and try IQ4_XS when Q4_K_M is just over.

Where to go next:

How sure are these numbers?

File sizes are the real GGUF sizes from the Hugging Face repos, in decimal GB. The memory needs add the file, the KV cache and a compute buffer (working memory for inference), and they are calculated with a formula, not estimated.

Speeds are estimates, each with a band and a confidence label, and the verdicts follow from those estimates and from how much of the model stays on the GPU. Calibrated means ±12% for dense models running fully on an NVIDIA or AMD card, and ±20% for partial offload, MoE models and other hardware. Theoretical means ±30%; it marks setups such as MoE experts kept in system RAM, where our estimates are likely conservative. Every ladder row in this guide is calibrated.

Measurements are joined only when the quant matches exactly, so a run of a Q4_K_XL file is never used for a Q4_K_M cell. The one measured point in this guide, shown below, comes from a llama.cpp GitHub discussion at a 2k context, not from an Ollama run, and sits next to our 8k estimate band.

Measured: gpt-oss-20b (MXFP4) on the GeForce RTX 5060 Ti 16GB
111.6 tok/s8k estimate 88.1 tok/s (70.5–105.8, calibrated estimate ±20%)measured (1 run, 2k context)
Source: github.com

At its MXFP4 baseline, gpt-oss-20b on the GeForce RTX 5060 Ti 16GB Runs great. It is a reasoning model, so its speed thresholds are 1.5× stricter.

Our figures assume Windows with the card driving the display, 32GB of DDR5 system RAM, an 8k context, an f16 KV cache, one model loaded and one request slot, and they cover the language-model file only, not a vision projector (mmproj) file. Ollama and llama.cpp make their own memory estimates, so check the real GPU/CPU split on your machine with ollama ps (or the load log in llama.cpp).

Unless a sentence names another quant, verdict names in the text refer to the baseline quant (Q4_K_M, or MXFP4 for gpt-oss) at 8k with an f16 KV cache; for other quants, read the ladder rows. The thresholds:

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

FAQ

Which GGUF quant should I download?

Start with Q4_K_M, the default for llama-server -hf and Hugging Face's Ollama integration, and switch only when memory says so. If its need at your context is over your usable memory, try IQ4_XS, in the same quality group (Good balance). At an 8k context, Mistral Small 3.2 24B needs 16.2 GB at Q4_K_M and 14.7 GB at IQ4_XS, against 15.4 GB usable on an RTX 5060 Ti 16GB, so in our calculation IQ4_XS fits, with little margin — check ollama ps. Take Q8_0 or Q6_K only if it fits with room for your context, and choose the file with a quant tag, as in ollama run hf.co/{user}/{repo}:IQ4_XS.

Is IQ4_XS worse than Q4_K_M?

It keeps a little less precision, but we put both in the same Good balance group. For Qwen3 14B the IQ4_XS file is 8.1 GB against 9.0 GB for Q4_K_M. The llama.cpp PR that added IQ4_XS reported that it sits well on the error-versus-size curve; we have no per-model quality measurements, and the difference depends on the model and the task. The practical reason to pick it is fit, when Q4_K_M is just over.

Is Q8_0 worth it over Q4_K_M?

Only when it fits with room for your context. For the dense models in this guide, Q8_0 is roughly three quarters larger than Q4_K_M. Qwen3 14B at Q8_0 needs 17.6 GB at 8k: over the 15.4 GB an RTX 5060 Ti 16GB can use, under the 23.4 GB of an RTX 3090. Even where it fits, each token reads a larger file, so the estimated speed is lower, and less memory is left for context.

Are 2-bit quants like Q2_K and IQ2_M usable?

Only as a last resort. Our pages suggest a 2-bit file only when every non-2-bit file Won't run. Otherwise a combo page mentions one separately when it would reach a higher verdict than the suggested quant (or the baseline), always marked as heavy quality loss. They exist mainly for large models, yet even Llama 3.3 70B's IQ2_M file (24.1 GB) is larger than the 23.4 GB a 24GB card can use. llama.cpp's 2023 k-quant measurements found that relative error did not keep shrinking in larger models, so size alone does not make 2-bit safe.

Is a q8_0 KV cache the same as a Q8_0 model file?

No. The model quant (Q8_0, Q4_K_M and so on) sets the size of the weights file you download. The KV cache type (f16, q8_0 or q4_0) is a separate runtime setting for the memory that holds the context, so you can lower both, and the savings add up. Ollama quantizes the KV cache only when the server starts with OLLAMA_KV_CACHE_TYPE (f16, q8_0 or q4_0) and flash attention on; the setting applies to every model the server runs (Ollama docs, checked 2026-09-29).

Sources

Speeds on this page are estimates with an error band and a confidence label, or public measurements with their source. Verdicts assume the setup stated with each table.