Guides
Practical guides to running LLMs on your own PC or Mac. Every number in them comes from the same data and estimates as our GPU, model and combo pages.
- How much VRAM do local LLMs need? 8–32 GB tiers (2026)
What fits on 8, 12, 16, 24 and 32 GB graphics cards: the memory formula, model-by-model verdicts, and how context length and quantization change the answer.
September 30, 2026
- Ollama slow or forgetting your prompt? Context length defaults
Why Ollama trims long conversations or slows down: its default context length by VRAM tier, how to raise it, and what a longer context costs in memory.
September 30, 2026
- GGUF quantization: Q4_K_M vs IQ4_XS vs Q8_0 — which to download
What each GGUF quant level costs in gigabytes, how we group them by quality, and when dropping one level turns a model that does not fit into one that runs.
September 30, 2026
- Korean open LLMs (EXAONE, Kanana, HyperCLOVA X SEED, Solar): hardware guide
Specs, licenses and what it takes to run LG EXAONE, Kakao Kanana, NAVER HyperCLOVA X SEED and Upstage Solar locally, with verdicts on common GPUs and Macs.
September 30, 2026
- Running local LLMs on a Mac (M4, M5): how much unified memory?
How much of a Mac’s unified memory a local model can use, which M4 and M5 configurations fit which models, and why memory bandwidth sets the speed.
September 30, 2026
- KV cache explained: how context length eats VRAM
Why a longer context needs more memory, how much each model’s KV cache costs per token, and what q8_0 and q4_0 KV cache quantization buy you.
September 30, 2026