MiniCPM5 2B hardware requirements
About the model
- Parameters
- 2.52B
- Attention
- Full attention
- Max context
- 131,072 tokens
- Released
- 2026-09-06
- Model card
- MiniCPM5 2B on Hugging Face
GGUF files
| Quant | File size | Quality | Published by |
|---|---|---|---|
| Q4_K_Mbaseline | 1.56 GB | Good balance | Official GGUF · openbmb |
| Q8_0 | 2.68 GB | Near-lossless | Official GGUF · openbmb |
Memory needed by quant and context
| Quant | 4k | 8k | 16k | 32k | 64k | 128k |
|---|---|---|---|---|---|---|
| Q4_K_M | 2.3 | 2.5 | 2.9 | 3.8 | 5.5 | 8.9 |
| Q8_0 | 3.4 | 3.6 | 4.0 | 4.9 | 6.6 | 10.0 |
Which GPUs can run MiniCPM5 2B?
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Scroll sideways to see every column.
| Hardware | Verdict | Speed | Quant | Memory | Runs as | Context | Notes | Price |
|---|---|---|---|---|---|---|---|---|
| GeForce RTX 5090 | Runs great | est. 656.0 tok/s577.2–734.7calibrated estimate ±12% | Q4_K_M | 3.1 / 32.0 GB | Full GPU | up to 128k | — | $4,200 street (as of 2026-08-09) |
| Radeon RX 7900 XTX | Runs great | est. 391.6 tok/s344.6–438.6calibrated estimate ±12% | Q4_K_M | 3.1 / 24.0 GB | Full GPU | up to 128k | — | $999 launch MSRP |
| GeForce RTX 3090 Ti | Runs great | est. 369.0 tok/s324.7–413.3calibrated estimate ±12% | Q4_K_M | 3.1 / 24.0 GB | Full GPU | up to 128k | — | $1,999 launch MSRP |
| GeForce RTX 4090 | Runs great | est. 369.0 tok/s324.7–413.3calibrated estimate ±12% | Q4_K_M | 3.1 / 24.0 GB | Full GPU | up to 128k | — | $2,755 street (as of 2026-08-09) |
| GeForce RTX 5080 | Runs great | est. 351.4 tok/s309.2–393.6calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $1,256 street (as of 2026-08-09) |
| GeForce RTX 3090 | Runs great | est. 342.6 tok/s301.5–383.7calibrated estimate ±12% | Q4_K_M | 3.1 / 24.0 GB | Full GPU | up to 128k | — | $1,050 street (as of 2026-08-09) |
| GeForce RTX 3080 12GB | Runs great | est. 333.8 tok/s293.8–373.9calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | $799 launch MSRP |
| GeForce RTX 5070 Ti | Runs great | est. 328.0 tok/s288.6–367.3calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $949 street (as of 2026-08-09) |
| GeForce RTX 5080 Laptop | Runs great | est. 328.0 tok/s288.6–367.3calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | — |
| GeForce RTX 5090 Laptop | Runs great | est. 328.0 tok/s288.6–367.3calibrated estimate ±12% | Q4_K_M | 3.1 / 24.0 GB | Full GPU | up to 128k | — | — |
| Radeon RX 7900 XT | Runs great | est. 326.3 tok/s287.1–365.5calibrated estimate ±12% | Q4_K_M | 3.1 / 20.0 GB | Full GPU | up to 128k | — | $899 launch MSRP |
| GeForce RTX 3080 10GB | Runs great | est. 278.2 tok/s244.8–311.6calibrated estimate ±12% | Q4_K_M | 3.1 / 10.0 GB | Full GPU | up to 128k | — | $699 launch MSRP |
| GeForce RTX 4080 Super | Runs great | est. 269.4 tok/s237.1–301.7calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $999 launch MSRP |
| Radeon RX 9070 | Runs great | est. 263.1 tok/s231.5–294.7calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $639 street (as of 2026-08-09) |
| Radeon RX 9070 XT | Runs great | est. 263.1 tok/s231.5–294.7calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $689 street (as of 2026-08-09) |
| GeForce RTX 4080 | Runs great | est. 262.5 tok/s231.0–294.0calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $1,199 launch MSRP |
| Radeon RX 7800 XT | Runs great | est. 254.5 tok/s224.0–285.1calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $499 launch MSRP |
| GeForce RTX 4070 Ti Super | Runs great | est. 246.0 tok/s216.5–275.5calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $799 launch MSRP |
| GeForce RTX 5070 | Runs great | est. 246.0 tok/s216.5–275.5calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | $629 street (as of 2026-08-09) |
| GeForce RTX 5070 Ti Laptop | Runs great | est. 246.0 tok/s216.5–275.5calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | — |
| GeForce RTX 2080 Ti | Runs great | est. 225.5 tok/s198.4–252.5calibrated estimate ±12% | Q4_K_M | 3.1 / 11.0 GB | Full GPU | up to 128k | — | $999 launch MSRP |
| GeForce RTX 4090 Laptop | Runs great | est. 210.8 tok/s185.5–236.1calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | — |
| Apple M5 Max (40-core GPU) · 48 GB | Runs great | est. 192.6 tok/s154.1–231.2calibrated estimate ±20% | Q4_K_M | 2.5 / 33.6 GB | Unified memory | up to 128k | — | — |
| Apple M3 Ultra · 96 GB | Runs great | est. 188.4 tok/s150.8–226.1calibrated estimate ±20% | Q4_K_M | 2.5 / 67.2 GB | Unified memory | up to 128k | — | — |
| GeForce RTX 4070 | Runs great | est. 184.5 tok/s162.3–206.6calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | $599 launch MSRP |
| GeForce RTX 4070 Super | Runs great | est. 184.5 tok/s162.3–206.6calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | $599 launch MSRP |
| Apple M2 Ultra · 64 GB | Runs great | est. 184.1 tok/s147.3–220.9calibrated estimate ±20% | Q4_K_M | 2.5 / 44.8 GB | Unified memory | up to 128k | — | — |
| GeForce GTX 1080 Ti | Runs great | est. 177.2 tok/s124.0–230.3theoretical estimate ±30% | Q4_K_M | 3.1 / 11.0 GB | Full GPU | up to 128k | — | $699 launch MSRP |
| Apple M4 Max (40-core GPU) · 48 GB | Runs great | est. 171.3 tok/s137.0–205.6calibrated estimate ±20% | Q4_K_M | 2.5 / 33.6 GB | Unified memory | up to 128k | — | — |
| GeForce RTX 3070 | Runs great | est. 164.0 tok/s144.3–183.7calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | $499 launch MSRP |
| GeForce RTX 5060 | Runs great | est. 164.0 tok/s144.3–183.7calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | $339 street (as of 2026-08-09) |
| GeForce RTX 5060 Ti 16GB | Runs great | est. 164.0 tok/s144.3–183.7calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $569 street (as of 2026-08-09) |
| GeForce RTX 5060 Ti 8GB | Runs great | est. 164.0 tok/s144.3–183.7calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | $429 street (as of 2026-08-09) |
| GeForce RTX 4080 Laptop | Runs great | est. 158.1 tok/s139.2–177.1calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | — |
| Apple M5 Max (32-core GPU) · 36 GB | Runs great | est. 144.3 tok/s115.5–173.2calibrated estimate ±20% | Q4_K_M | 2.5 / 25.2 GB | Unified memory | up to 128k | — | — |
| Intel Arc B580 | Runs great | est. 140.7 tok/s112.6–168.8calibrated estimate ±20% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | $290 street (as of 2026-08-09) |
| GeForce RTX 3060 12GB | Runs great | est. 131.8 tok/s116.0–147.6calibrated estimate ±12% | Q4_K_M | 3.1 / 12.0 GB | Full GPU | up to 128k | — | $250 street (as of 2026-08-09) |
| Radeon RX 9060 XT 16GB | Runs great | est. 131.3 tok/s115.6–147.1calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $449 street (as of 2026-08-09) |
| Apple M4 Max (32-core GPU) · 36 GB | Runs great | est. 128.6 tok/s102.9–154.4calibrated estimate ±20% | Q4_K_M | 2.5 / 25.2 GB | Unified memory | up to 128k | — | — |
| GeForce RTX 4060 Ti 16GB | Runs great | est. 105.4 tok/s92.8–118.1calibrated estimate ±12% | Q4_K_M | 3.1 / 16.0 GB | Full GPU | up to 128k | — | $499 launch MSRP |
| GeForce RTX 4060 Ti 8GB | Runs great | est. 105.4 tok/s92.8–118.1calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | $399 launch MSRP |
| GeForce RTX 4060 | Runs great | est. 99.6 tok/s87.6–111.5calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | $299 launch MSRP |
| Apple M5 Pro · 24 GB | Runs great | est. 96.3 tok/s77.1–115.6calibrated estimate ±20% | Q4_K_M | 2.5 / 16.8 GB | Unified memory | up to 128k | — | — |
| GeForce RTX 4060 Laptop | Runs great | est. 93.7 tok/s82.5–105.0calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | — |
| GeForce RTX 4070 Laptop | Runs great | est. 93.7 tok/s82.5–105.0calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | — |
| Ryzen AI Max+ 395 (Strix Halo) · 32 GB | Runs great | est. 87.0 tok/s69.6–104.4calibrated estimate ±20% | Q4_K_M | 2.5 / 22.4 GB | Unified memory | up to 128k | — | $1,999 launch MSRP |
| Apple M4 Pro · 24 GB | Runs great | est. 85.7 tok/s68.5–102.8calibrated estimate ±20% | Q4_K_M | 2.5 / 16.8 GB | Unified memory | up to 128k | — | — |
| GeForce RTX 3050 8GB | Runs great | est. 82.0 tok/s72.2–91.8calibrated estimate ±12% | Q4_K_M | 3.1 / 8.0 GB | Full GPU | up to 64k | — | $249 launch MSRP |
| NVIDIA DGX Spark · 128 GB | Runs great | est. 57.1 tok/s45.7–68.5calibrated estimate ±20% | Q4_K_M | 2.5 / 89.6 GB | Unified memory | up to 64k | — | $3,999 launch MSRP |
| Apple M5 · 16 GB | Runs great | est. 48.2 tok/s38.6–57.8calibrated estimate ±20% | Q4_K_M | 2.5 / 11.2 GB | Unified memory | up to 64k | — | — |
| Apple M4 · 16 GB | Runs great | est. 37.7 tok/s30.1–45.2calibrated estimate ±20% | Q4_K_M | 2.5 / 11.2 GB | Unified memory | up to 32k | — | — |
| DDR5-6000 dual-channel CPU · 16 GB | Runs slowly | est. 25.1 tok/s20.1–30.1calibrated estimate ±20% | Q4_K_M | 1.9 GB RAM | CPU only | up to 128k |
| — |
| DDR5-5600 dual-channel CPU · 16 GB | Runs slowly | est. 23.4 tok/s18.7–28.1calibrated estimate ±20% | Q4_K_M | 1.9 GB RAM | CPU only | up to 128k |
| — |
| DDR4-3200 dual-channel CPU · 16 GB | Runs slowly | est. 13.4 tok/s10.7–16.1calibrated estimate ±20% | Q4_K_M | 1.9 GB RAM | CPU only | up to 128k |
| — |
Cheapest GPUs that run it
- Runs greatGeForce RTX 3060 12GB$250 street (as of 2026-08-09), est. 131.8 tok/s (116.0–147.6, calibrated estimate ±12%).
Measured results for MiniCPM5 2B
No public measurements for this model yet.
Frequently asked questions
How much VRAM does MiniCPM5 2B need?
At Q4_K_M the weights are 1.56 GB; with 8k context the total is about 2.5 GB (KV cache 0.4 GB, compute buffer 0.6 GB).
What is the cheapest GPU that runs MiniCPM5 2B well?
Cheapest that runs it great: GeForce RTX 3060 12GB at $250 (street price as of 2026-08-09), est. 131.8 tok/s (116.0–147.6, calibrated estimate ±12%).
Can MiniCPM5 2B run on an 8 GB, 16 GB or 24 GB GPU?
At Q4_K_M with 8k context: GeForce RTX 4060 — Runs great, est. 99.6 tok/s (87.6–111.5, calibrated estimate ±12%); GeForce RTX 5060 Ti 16GB — Runs great, est. 164.0 tok/s (144.3–183.7, calibrated estimate ±12%); GeForce RTX 3090 — Runs great, est. 342.6 tok/s (301.5–383.7, calibrated estimate ±12%).
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.