Qwen3.8 27B hardware requirements

About the model

Parameters
27.78B
Attention
Hybrid linear attention
Max context
262,144 tokens
Released
2026-08-05

GGUF files

Tracked quantizations of Qwen3.8 27B, smallest first
QuantFile sizeQualityPublished by
Q2_K10.82 GBHeavy lossCommunity GGUF · bartowski
IQ4_XS15.48 GBGood balanceCommunity GGUF · bartowski
Q4_K_Mbaseline17.44 GBGood balanceCommunity GGUF · bartowski
Q8_029.12 GBNear-losslessCommunity GGUF · bartowski

Memory needed by quant and context

Weights + f16 KV cache + compute buffer, in GB. Add about 0.6 GB if the GPU also drives your display on Windows.
Quant4k8k16k32k64k128k256k
Q2_K11.611.912.513.816.221.130.9
IQ4_XS16.316.617.218.420.925.835.6
Q4_K_M18.218.619.220.422.827.737.5
Q8_029.930.230.832.134.539.449.2

Which GPUs can run Qwen3.8 27B?

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.

Scroll sideways to see every column.

54 GPUs, Macs and CPU setups at Q4_K_M with 8k context, sorted by verdict and speed
HardwareVerdictSpeedQuantMemoryRuns asContextNotesPrice
GeForce RTX 5090
Runs great
est. 69.8 tok/s61.4–78.2calibrated estimate ±12%
Q4_K_M19.2 / 32.0 GBFull GPUup to 128k—$4,200 street (as of 2026-08-09)
Radeon RX 7900 XTX
Runs great
est. 41.7 tok/s36.7–46.7calibrated estimate ±12%
Q4_K_M19.2 / 24.0 GBFull GPUup to 64k—$999 launch MSRP
GeForce RTX 3090 Ti
Runs great
est. 39.3 tok/s34.5–44.0calibrated estimate ±12%
Q4_K_M19.2 / 24.0 GBFull GPUup to 64k—$1,999 launch MSRP
GeForce RTX 4090
Runs great
est. 39.3 tok/s34.5–44.0calibrated estimate ±12%
Q4_K_M19.2 / 24.0 GBFull GPUup to 64k—$2,755 street (as of 2026-08-09)
GeForce RTX 3090
Runs great
est. 36.4 tok/s32.1–40.8calibrated estimate ±12%
Q4_K_M19.2 / 24.0 GBFull GPUup to 64k—$1,050 street (as of 2026-08-09)
GeForce RTX 5090 Laptop
Runs great
est. 34.9 tok/s30.7–39.1calibrated estimate ±12%
Q4_K_M19.2 / 24.0 GBFull GPUup to 64k——
Radeon RX 7900 XT
Runs great
est. 34.7 tok/s30.5–38.9calibrated estimate ±12%
Q4_K_M19.2 / 20.0 GBFull GPUup to 16k—$899 launch MSRP
Apple M5 Max (40-core GPU) · 48 GB
Runs great
est. 20.5 tok/s16.4–24.6calibrated estimate ±20%
Q4_K_M18.6 / 33.6 GBUnified memoryup to 8k——
Apple M3 Ultra · 96 GB
Runs great
est. 20.0 tok/s16.0–24.1calibrated estimate ±20%
Q4_K_M18.6 / 67.2 GBUnified memoryup to 8k——
Apple M2 Ultra · 64 GB
Runs well
est. 19.6 tok/s15.7–23.5calibrated estimate ±20%
Q4_K_M18.6 / 44.8 GBUnified memoryup to 256k
  • Try IQ4_XS: Runs great
—
Apple M4 Max (40-core GPU) · 48 GB
Runs well
est. 18.2 tok/s14.6–21.9calibrated estimate ±20%
Q4_K_M18.6 / 33.6 GBUnified memoryup to 128k
  • Try IQ4_XS: Runs great
—
Apple M5 Max (32-core GPU) · 36 GB
Runs well
est. 15.4 tok/s12.3–18.4calibrated estimate ±20%
Q4_K_M18.6 / 25.2 GBUnified memoryup to 64k
  • Q2_K (heavy quality loss): Runs great
—
Apple M4 Max (32-core GPU) · 36 GB
Runs well
est. 13.7 tok/s10.9–16.4calibrated estimate ±20%
Q4_K_M18.6 / 25.2 GBUnified memoryup to 64k
  • Q2_K (heavy quality loss): Runs great
—
Apple M5 Pro · 48 GB
Runs well
est. 10.2 tok/s8.2–12.3calibrated estimate ±20%
Q4_K_M18.6 / 33.6 GBUnified memoryup to 64k——
Ryzen AI Max+ 395 (Strix Halo) · 32 GB
Runs well
est. 9.3 tok/s7.4–11.1calibrated estimate ±20%
Q4_K_M18.6 / 22.4 GBUnified memoryup to 32k—$1,999 launch MSRP
Apple M4 Pro · 48 GB
Runs well
est. 9.1 tok/s7.3–10.9calibrated estimate ±20%
Q4_K_M18.6 / 33.6 GBUnified memoryup to 32k——
GeForce RTX 5080
Runs slowly
est. 10.6 tok/s8.5–12.7calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$1,256 street (as of 2026-08-09)
GeForce RTX 5070 Ti
Runs slowly
est. 10.4 tok/s8.3–12.5calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$949 street (as of 2026-08-09)
GeForce RTX 5080 Laptop
Runs slowly
est. 10.4 tok/s8.3–12.5calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
—
GeForce RTX 4080 Super
Runs slowly
est. 9.9 tok/s7.9–11.9calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$999 launch MSRP
Radeon RX 9070
Runs slowly
est. 9.8 tok/s7.9–11.8calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$639 street (as of 2026-08-09)
Radeon RX 9070 XT
Runs slowly
est. 9.8 tok/s7.9–11.8calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$689 street (as of 2026-08-09)
GeForce RTX 4080
Runs slowly
est. 9.8 tok/s7.9–11.8calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$1,199 launch MSRP
Radeon RX 7800 XT
Runs slowly
est. 9.7 tok/s7.8–11.7calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$499 launch MSRP
GeForce RTX 4070 Ti Super
Runs slowly
est. 9.6 tok/s7.7–11.6calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$799 launch MSRP
GeForce RTX 4090 Laptop
Runs slowly
est. 9.2 tok/s7.3–11.0calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
—
GeForce RTX 5060 Ti 16GB
Runs slowly
est. 8.4 tok/s6.7–10.0calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$569 street (as of 2026-08-09)
Radeon RX 9060 XT 16GB
Runs slowly
est. 7.6 tok/s6.1–9.1calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Try IQ4_XS: Runs well
  • Q2_K (heavy quality loss): Runs great
$449 street (as of 2026-08-09)
GeForce RTX 4060 Ti 16GB
Runs slowly
est. 6.9 tok/s5.5–8.2calibrated estimate ±20%
Q4_K_M16.0 / 16.0 GB + 3.2 GB RAMPartial offloadup to 128k
  • Try IQ4_XS: Runs well
$499 launch MSRP
NVIDIA DGX Spark · 128 GB
Runs slowly
est. 6.1 tok/s4.9–7.3calibrated estimate ±20%
Q4_K_M18.6 / 89.6 GBUnified memoryup to 256k
  • Q2_K (heavy quality loss): Runs well
$3,999 launch MSRP
GeForce RTX 3080 12GB
Runs slowly
est. 5.5 tok/s4.4–6.6calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
$799 launch MSRP
GeForce RTX 5070
Runs slowly
est. 5.3 tok/s4.3–6.4calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
$629 street (as of 2026-08-09)
GeForce RTX 5070 Ti Laptop
Runs slowly
est. 5.3 tok/s4.3–6.4calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
—
GeForce RTX 4070
Runs slowly
est. 5.1 tok/s4.1–6.2calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
$599 launch MSRP
GeForce RTX 4070 Super
Runs slowly
est. 5.1 tok/s4.1–6.2calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
$599 launch MSRP
GeForce RTX 4080 Laptop
Runs slowly
est. 5.0 tok/s4.0–6.0calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
—
Intel Arc B580
Runs slowly
est. 4.9 tok/s3.9–5.9calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
$290 street (as of 2026-08-09)
GeForce RTX 3060 12GB
Runs slowly
est. 4.8 tok/s3.9–5.8calibrated estimate ±20%
Q4_K_M12.0 / 12.0 GB + 7.2 GB RAMPartial offloadup to 64k
  • Q2_K (heavy quality loss): Runs well
$250 street (as of 2026-08-09)
GeForce RTX 2080 Ti
Runs slowly
est. 4.8 tok/s3.8–5.7calibrated estimate ±20%
Q4_K_M11.0 / 11.0 GB + 8.2 GB RAMPartial offloadup to 64k—$999 launch MSRP
GeForce GTX 1080 Ti
Runs slowly
est. 4.6 tok/s3.2–6.0theoretical estimate ±30%
Q4_K_M11.0 / 11.0 GB + 8.2 GB RAMPartial offloadup to 64k—$699 launch MSRP
GeForce RTX 3080 10GB
Runs slowly
est. 4.4 tok/s3.5–5.3calibrated estimate ±20%
Q4_K_M10.0 / 10.0 GB + 9.2 GB RAMPartial offloadup to 64k—$699 launch MSRP
GeForce RTX 3070
Runs slowly
est. 3.6 tok/s2.9–4.3calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k—$499 launch MSRP
GeForce RTX 5060
Runs slowly
est. 3.6 tok/s2.9–4.3calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k—$339 street (as of 2026-08-09)
GeForce RTX 5060 Ti 8GB
Runs slowly
est. 3.6 tok/s2.9–4.3calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k—$429 street (as of 2026-08-09)
GeForce RTX 4060 Ti 8GB
Runs slowly
est. 3.5 tok/s2.8–4.2calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k—$399 launch MSRP
GeForce RTX 4060
Runs slowly
est. 3.4 tok/s2.8–4.1calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k—$299 launch MSRP
GeForce RTX 4060 Laptop
Runs slowly
est. 3.4 tok/s2.7–4.1calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k——
GeForce RTX 4070 Laptop
Runs slowly
est. 3.4 tok/s2.7–4.1calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k——
GeForce RTX 3050 8GB
Runs slowly
est. 3.4 tok/s2.7–4.0calibrated estimate ±20%
Q4_K_M8.0 / 8.0 GB + 11.2 GB RAMPartial offloadup to 64k—$249 launch MSRP
Apple M5 · 24 GB
Runs slowly
est. 3.0 tok/s2.1–3.9theoretical estimate ±30%
Q4_K_M18.0 GB RAMCPU onlyup to 32k
  • Q2_K (heavy quality loss): Runs well
—
DDR5-6000 dual-channel CPU · 32 GB
Runs slowly
est. 2.7 tok/s2.1–3.2calibrated estimate ±20%
Q4_K_M18.0 GB RAMCPU onlyup to 64k——
DDR5-5600 dual-channel CPU · 32 GB
Runs slowly
est. 2.5 tok/s2.0–3.0calibrated estimate ±20%
Q4_K_M18.0 GB RAMCPU onlyup to 64k——
Apple M4 · 24 GB
Runs slowly
est. 2.3 tok/s1.6–3.0theoretical estimate ±30%
Q4_K_M18.0 GB RAMCPU onlyup to 32k——
DDR4-3200 dual-channel CPU · 64 GB
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.3 tok/s1.8–2.7calibrated estimate ±20%
Q2_K11.4 GB RAMCPU onlyup to 16k——

Cheapest GPUs that run it

  • Runs greatGeForce RTX 3090$1,050 street (as of 2026-08-09), est. 36.4 tok/s (32.1–40.8, calibrated estimate ±12%).

Measured results for Qwen3.8 27B

No public measurements for this model yet.

Frequently asked questions

How much VRAM does Qwen3.8 27B need?

At Q4_K_M the weights are 17.44 GB; with 8k context the total is about 18.6 GB (KV cache 0.5 GB, compute buffer 0.6 GB). The smallest tracked file (Q2_K (heavy quality loss)) needs about 11.9 GB.

What is the cheapest GPU that runs Qwen3.8 27B well?

Cheapest that runs it great: GeForce RTX 3090 at $1,050 (street price as of 2026-08-09), est. 36.4 tok/s (32.1–40.8, calibrated estimate ±12%).

Can Qwen3.8 27B run on an 8 GB, 16 GB or 24 GB GPU?

At Q4_K_M with 8k context: GeForce RTX 4060 — Runs slowly, est. 3.4 tok/s (2.8–4.1, calibrated estimate ±20%); GeForce RTX 5060 Ti 16GB — Runs slowly, est. 8.4 tok/s (6.7–10.0, calibrated estimate ±20%); GeForce RTX 3090 — Runs great, est. 36.4 tok/s (32.1–40.8, calibrated estimate ±12%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.