Qwen3.5 122B-A10B hardware requirements

About the model

Parameters
125.09B total · 10B active (MoE)
Attention
Hybrid linear attention
Max context
262,144 tokens
Released
2026-05-01

GGUF files

Tracked quantizations of Qwen3.5 122B-A10B, smallest first
QuantFile sizeQualityPublished by
UD-Q2_K_XL42.85 GBHeavy lossCommunity GGUF · unsloth
Q4_K_Mbaseline78.26 GBGood balanceCommunity GGUF · unsloth

Some repositories publish only Unsloth Dynamic (UD) variants; sizes are those of the file you would download.

Memory needed by quant and context

Weights + f16 KV cache + compute buffer, in GB. Add about 0.6 GB if the GPU also drives your display on Windows.
Quant4k8k16k32k64k128k256k
UD-Q2_K_XL43.543.643.944.545.647.852.2
Q4_K_M78.979.079.379.981.083.287.6

Which GPUs can run Qwen3.5 122B-A10B?

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.

Scroll sideways to see every column.

54 GPUs, Macs and CPU setups at Q4_K_M with 8k context, sorted by verdict and speed
HardwareVerdictSpeedQuantMemoryRuns asContextNotesPrice
Apple M3 Ultra · 256 GB
Runs great
est. 36.8 tok/s29.4–44.1calibrated estimate ±20%
Q4_K_M79.0 / 179.2 GBUnified memoryup to 128k——
Apple M2 Ultra · 128 GB
Runs great
est. 35.9 tok/s28.7–43.1calibrated estimate ±20%
Q4_K_M79.0 / 89.6 GBUnified memoryup to 128k——
Apple M5 Max (40-core GPU) · 128 GB
Runs great
est. 28.5 tok/s22.8–34.2calibrated estimate ±20%
Q4_K_M79.0 / 89.6 GBUnified memoryup to 64k——
Apple M4 Max (40-core GPU) · 128 GB
Runs great
est. 25.4 tok/s20.3–30.4calibrated estimate ±20%
Q4_K_M79.0 / 89.6 GBUnified memoryup to 64k——
Ryzen AI Max+ 395 (Strix Halo) · 128 GB
Runs great
est. 21.0 tok/s16.8–25.2calibrated estimate ±20%
Q4_K_M79.0 / 89.6 GBUnified memoryup to 16k—$1,999 launch MSRP
NVIDIA DGX Spark · 128 GB
Runs well
est. 16.9 tok/s13.5–20.3calibrated estimate ±20%
Q4_K_M79.0 / 89.6 GBUnified memoryup to 256k
  • UD-Q2_K_XL (heavy quality loss): Runs great
$3,999 launch MSRP
DDR5-6000 dual-channel CPU · 96 GB
Runs slowly
est. 6.4 tok/s5.1–7.7calibrated estimate ±20%
Q4_K_M78.5 GB RAMCPU onlyup to 256k——
DDR5-5600 dual-channel CPU · 96 GB
Runs slowly
est. 6.0 tok/s4.8–7.2calibrated estimate ±20%
Q4_K_M78.5 GB RAMCPU onlyup to 256k——
Apple M5 Max (32-core GPU) · 48 GB
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 38.0 tok/s26.6–49.5theoretical estimate ±30%
UD-Q2_K_XL43.1 GB RAMCPU onlyup to 32k
  • CPU inference — capped at Runs slowly
—
Apple M4 Max (32-core GPU) · 48 GB
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 33.9 tok/s23.7–44.1theoretical estimate ±30%
UD-Q2_K_XL43.1 GB RAMCPU onlyup to 32k
  • CPU inference — capped at Runs slowly
—
GeForce RTX 5090
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 33.3 tok/s23.3–43.2theoretical estimate ±30%
UD-Q2_K_XL32.0 / 32.0 GB + 12.2 GB RAMPartial offloadup to 256k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
$4,200 street (as of 2026-08-09)
Apple M5 Pro · 64 GB
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs great
est. 25.4 tok/s20.3–30.5calibrated estimate ±20%
UD-Q2_K_XL43.6 / 44.8 GBUnified memoryup to 32k——
Apple M4 Pro · 64 GB
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs great
est. 22.6 tok/s18.1–27.1calibrated estimate ±20%
UD-Q2_K_XL43.6 / 44.8 GBUnified memoryup to 16k——
Radeon RX 7900 XTX
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 20.7 tok/s14.5–26.8theoretical estimate ±30%
UD-Q2_K_XL24.0 / 24.0 GB + 20.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
$999 launch MSRP
GeForce RTX 3090 Ti
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 20.6 tok/s14.4–26.7theoretical estimate ±30%
UD-Q2_K_XL24.0 / 24.0 GB + 20.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
$1,999 launch MSRP
GeForce RTX 4090
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 20.6 tok/s14.4–26.7theoretical estimate ±30%
UD-Q2_K_XL24.0 / 24.0 GB + 20.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
$2,755 street (as of 2026-08-09)
GeForce RTX 3090
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 20.4 tok/s14.3–26.5theoretical estimate ±30%
UD-Q2_K_XL24.0 / 24.0 GB + 20.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
$1,050 street (as of 2026-08-09)
GeForce RTX 5090 Laptop
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 20.3 tok/s14.2–26.4theoretical estimate ±30%
UD-Q2_K_XL24.0 / 24.0 GB + 20.2 GB RAMPartial offloadup to 128k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
—
Radeon RX 7900 XT
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 17.5 tok/s12.2–22.7theoretical estimate ±30%
UD-Q2_K_XL20.0 / 20.0 GB + 24.2 GB RAMPartial offloadup to 64k
  • Less than 90% on the GPU — a GPU/CPU split is capped at Runs slowly
  • With 64 GB RAM: Runs slowly
$899 launch MSRP
DDR4-3200 dual-channel CPU · 64 GB
Won't runTry UD-Q2_K_XL (heavy quality loss): Runs slowly
est. 6.1 tok/s4.9–7.3calibrated estimate ±20%
UD-Q2_K_XL43.1 GB RAMCPU onlyup to 256k——
Apple M4 · 32 GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
—
Apple M5 · 32 GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
—
Intel Arc B580
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$290 street (as of 2026-08-09)
GeForce GTX 1080 Ti
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 10.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$699 launch MSRP
GeForce RTX 2080 Ti
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 10.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$999 launch MSRP
GeForce RTX 3050 8GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$249 launch MSRP
GeForce RTX 3060 12GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$250 street (as of 2026-08-09)
GeForce RTX 3070
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$499 launch MSRP
GeForce RTX 3080 10GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 9.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$699 launch MSRP
GeForce RTX 3080 12GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$799 launch MSRP
GeForce RTX 4060
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$299 launch MSRP
GeForce RTX 4060 Laptop
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
—
GeForce RTX 4060 Ti 16GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$499 launch MSRP
GeForce RTX 4060 Ti 8GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$399 launch MSRP
GeForce RTX 4070
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$599 launch MSRP
GeForce RTX 4070 Laptop
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
—
GeForce RTX 4070 Super
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$599 launch MSRP
GeForce RTX 4070 Ti Super
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$799 launch MSRP
GeForce RTX 4080
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$1,199 launch MSRP
GeForce RTX 4080 Laptop
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
—
GeForce RTX 4080 Super
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$999 launch MSRP
GeForce RTX 4090 Laptop
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
—
GeForce RTX 5060
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$339 street (as of 2026-08-09)
GeForce RTX 5060 Ti 16GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$569 street (as of 2026-08-09)
GeForce RTX 5060 Ti 8GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$429 street (as of 2026-08-09)
GeForce RTX 5070
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$629 street (as of 2026-08-09)
GeForce RTX 5070 Ti
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$949 street (as of 2026-08-09)
GeForce RTX 5070 Ti Laptop
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 11.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
—
GeForce RTX 5080
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$1,256 street (as of 2026-08-09)
GeForce RTX 5080 Laptop
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
—
Radeon RX 7800 XT
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$499 launch MSRP
Radeon RX 9060 XT 16GB
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$449 street (as of 2026-08-09)
Radeon RX 9070
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$639 street (as of 2026-08-09)
Radeon RX 9070 XT
Won't run
—Q4_K_Mneeds 79.0 GB——
  • Needs about 79.0 GB; 15.4 GB of VRAM and 28.0 GB of RAM are free
  • With 96 GB RAM: Runs slowly
$689 street (as of 2026-08-09)

Cheapest GPUs that run it

No GPU with a current street price reaches Runs great or Runs well at Q4_K_M. The table above lists Macs, unified-memory PCs and CPU setups.

Measured results for Qwen3.5 122B-A10B

No public measurements for this model yet.

Frequently asked questions

How much VRAM does Qwen3.5 122B-A10B need?

At Q4_K_M the weights are 78.26 GB; with 8k context the total is about 79.0 GB (KV cache 0.2 GB, compute buffer 0.6 GB). The smallest tracked file (UD-Q2_K_XL (heavy quality loss)) needs about 43.6 GB.

What is the cheapest GPU that runs Qwen3.5 122B-A10B well?

No GPU with a current street price reaches Runs great or Runs well at Q4_K_M. The best tracked option is Apple M3 Ultra · 256 GB: Runs great.

Can Qwen3.5 122B-A10B run on an 8 GB, 16 GB or 24 GB GPU?

At Q4_K_M with 8k context: GeForce RTX 4060 — Won't run; GeForce RTX 5060 Ti 16GB — Won't run; GeForce RTX 3090 — Won't run.

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.