Which local LLMs can your PC run?

Pick your hardware; every model gets a verdict, a recommended quant and an estimated speed band.

Results

Pick a GPU or chip above to see verdicts.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Frequently asked questions

How accurate are the speed estimates?

Every speed is a band with a confidence label: measured (a public benchmark for this exact GPU, model and quant), calibrated (a bandwidth model fitted to public benchmarks — ±12% for NVIDIA and AMD cards running a dense model fully on the GPU, ±20% otherwise) or theoretical (±30%, outside our calibration set). Drivers, backend, context and thermals move real results within that band.

What do the four verdicts mean?

Runs great: the whole model sits on the GPU and generates faster than you read. Runs well: a comfortable wait, including MoE models that keep their experts in system RAM. Runs slowly: it works but crawls (CPU or a GPU/CPU split). Won't run: not enough memory, or too slow to be usable. Reasoning models are judged 1.5× more strictly because they spend tokens thinking.

Is my hardware information sent anywhere?

No. Detection and every calculation run inside your browser; there is no server. A share link only carries the GPU, RAM and settings you picked as URL parameters.

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.