Which local LLMs can the GeForce GTX 1080 Ti run?
Assumes 32 GB of DDR5-5600 system RAM, Windows with this GPU driving the display, 8k context and an f16 KV cache.
Specs
- VRAM
- 11 GB
- Memory bandwidth
- 484.0 GB/s
- FP16 compute
- 0.2 TFLOPS
- Launch year
- 2017
- Price
- $699 launch MSRP
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- Runs great9 modelsRuns well4 modelsRuns slowly11 modelsWon't run5 models1 more model runs at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the GeForce GTX 1080 Ti
Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 258.0 tok/s180.6–335.4theoretical estimate ±30% | Q4_K_M | 2.5 / 11.0 GB | Full GPU | up to 64k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 186.6 tok/s130.6–242.6theoretical estimate ±30% | Q4_K_M | 3.0 / 11.0 GB | Full GPU | up to 16k | — |
| DeepSeek R1 Distill Llama 8B | Runs great | est. 56.5 tok/s39.6–73.5theoretical estimate ±30% | Q4_K_M | 7.2 / 11.0 GB | Full GPU | up to 32k |
|
| Kanana 1.5 8B | Runs great | est. 56.5 tok/s39.6–73.5theoretical estimate ±30% | Q4_K_M | 7.2 / 11.0 GB | Full GPU | up to 32k | — |
| Llama 3.1 8B | Runs great | est. 56.5 tok/s39.6–73.5theoretical estimate ±30% | Q4_K_M | 7.2 / 11.0 GB | Full GPU | up to 32k | — |
| Qwen3.5 9B | Runs great | est. 55.2 tok/s38.6–71.8theoretical estimate ±30% | Q4_K_M | 7.3 / 11.0 GB | Full GPU | up to 64k | — |
| Qwen3 8B | Runs great | est. 54.3 tok/s38.0–70.6theoretical estimate ±30% | Q4_K_M | 7.4 / 11.0 GB | Full GPU | up to 16k | — |
| Gemma 4 12B | Runs great | est. 44.2 tok/s31.0–57.5theoretical estimate ±30% | Q4_K_M | 8.8 / 11.0 GB | Full GPU | up to 32k | — |
| Gemma 3 12B | Runs great | est. 43.2 tok/s30.3–56.2theoretical estimate ±30% | Q4_K_M | 9.0 / 11.0 GB | Full GPU | up to 32k | — |
| HyperCLOVA X SEED Think 14B | Runs well | est. 26.1 tok/s18.3–34.0theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 0.4 GB RAM | Partial offload | up to 8k |
|
| Qwen3 14B | Runs well | est. 23.8 tok/s16.7–30.9theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 0.5 GB RAM | Partial offload | up to 8k |
|
| Qwen3.5 35B-A3B | Runs well | est. 20.4 tok/s14.3–26.5theoretical estimate ±30% | Q4_K_M | 2.6 / 11.0 GB + 21.4 GB RAM | MoE experts in RAM | up to 256k |
|
| Phi-4 | Runs well | est. 19.1 tok/s13.4–24.8theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 0.9 GB RAM | Partial offload | up to 8k |
|
| Gemma 4 26B-A4B | Runs slowly | est. 24.7 tok/s17.3–32.1theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 7.5 GB RAM | Partial offload | up to 128k |
|
| Kanana 1.5 15.7B-A3B | Runs slowly | est. 19.3 tok/s13.5–25.1theoretical estimate ±30% | Q4_K_M | 3.0 / 11.0 GB + 9.6 GB RAM | MoE experts in RAM | up to 32k |
|
| Qwen3 30B-A3B (2507) | Runs slowly | est. 19.2 tok/s13.4–24.9theoretical estimate ±30% | Q4_K_M | 2.9 / 11.0 GB + 17.7 GB RAM | MoE experts in RAM | up to 64k |
|
| gpt-oss-20b | Runs slowly | est. 18.5 tok/s12.9–24.0theoretical estimate ±30% | MXFP4 | 2.0 / 11.0 GB + 11.5 GB RAM | MoE experts in RAM | up to 128k |
|
| Mistral Small 3.2 24B | Runs slowly | est. 5.9 tok/s4.1–7.6theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 5.8 GB RAM | Partial offload | up to 32k |
|
| Gemma 3 27B | Runs slowly | est. 5.0 tok/s3.5–6.5theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 7.4 GB RAM | Partial offload | up to 64k | — |
| Qwen3.5 27B | Runs slowly | est. 4.8 tok/s3.4–6.2theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 7.8 GB RAM | Partial offload | up to 64k | — |
| EXAONE 4.0 32B | Runs slowly | est. 3.9 tok/s2.7–5.0theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 10.1 GB RAM | Partial offload | up to 32k |
|
| EXAONE 4.5 33B | Runs slowly | est. 3.6 tok/s2.5–4.7theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 10.8 GB RAM | Partial offload | up to 16k |
|
| Qwen3 32B | Runs slowly | est. 3.1 tok/s2.2–4.0theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 12.1 GB RAM | Partial offload | up to 16k | — |
| DeepSeek R1 Distill Qwen 32B | Runs slowly | est. 3.1 tok/s2.1–4.0theoretical estimate ±30% | Q4_K_M | 11.0 / 11.0 GB + 12.2 GB RAM | Partial offload | up to 8k |
|
| Llama 3.3 70B | Won't runTry Q2_K (heavy quality loss): Runs slowly | est. 2.0 tok/s1.4–2.6theoretical estimate ±30% | Q2_K | 11.0 / 11.0 GB + 19.2 GB RAM | Partial offload | up to 8k | — |
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Qwen3.5 122B-A10B | Won't run | — | Q4_K_M | needs 79.0 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Measured results on the GeForce GTX 1080 Ti
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the GeForce GTX 1080 Ti?
Kanana 1.5 15.7B-A3B at IQ4_XS (a 8.71 GB file) fits entirely in 11 GB with 8k context, at est. 79.5 tok/s (55.7–103.4, theoretical estimate ±30%).
How many local LLMs run well on the GeForce GTX 1080 Ti?
At Q4_K_M with 8k context, out of 29 tracked models: 9 run great, 4 run well, 11 run slowly and 5 won't run.
Can the GeForce GTX 1080 Ti run a 70B model like Llama 3.3 70B?
Runs slowly — Partial offload, Q2_K (heavy quality loss): est. 2.0 tok/s (1.4–2.6, theoretical estimate ±30%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.