GPUFits

Models / Qwen3.8 Flash-Next

Qwen3.8 Flash-Next VRAM Requirements & GPU Pairing

Qwen3.8-Flash-Next is the September 2026 preview of the Qwen4 architecture: 180B total parameters (including a 51.2B n-gram lookup table and MTP layers) with only ~6B active per token (512 routed experts, top-10 + 1 shared). Of its 48 layers, 36 are Gated DeltaNet linear attention and just 12 are QSA full attention (24 Q heads / 2 KV heads), giving it a native 256K context. Note the license: Qwen Community License 1.0, not Apache-2.0. Q4_K_M weighs ~110.3GB, ~112.6GB total — a 96GB Mac (~72GB usable) cannot hold it, and every GPU in our database fails at Q4.

Realistic paths: a 128GB M5 Max (~96GB usable) runs Q2_K (~73.6GB) comfortably and Q3_K_M (~92.3GB) tight; only the 256GB M5 Ultra runs Q4 comfortably. At 6B active, whatever fits is fast — an M5 Max at 614GB/s theoretically reaches ~125 tok/s, and the n-gram table can be offloaded to cut pressure further. Unsloth GGUFs, Ollama support and Mac deployment write-ups all landed within September — this is the largest new model consumer hardware can almost reach.

Architecture Specs

Total parameters180B
Active parameters(MoE: VRAM follows total params, speed follows active params) 6B
Layers48
KV heads 2
Head dim256
Native context256K
Licenseqwen-community
KV bytes per token (fp16)96.0 KB

Hybrid architecture: only 12 of 48 layers are full attention; the 36 Gated DeltaNet layers' KV does not grow with context. The KV column above conservatively applies the standard GQA formula to all 48 layers — real KV is roughly 1/4 of the shown value.

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 191.5 GB 193.8 GB
Q6_K 147.8 GB 150.1 GB
Q4_K_M 110.3 GB 112.6 GB
Q3_K_M 90.0 GB 92.3 GB
Q2_K 71.3 GB 73.6 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Not feasible — — Try it →
RTX 3090 24.0 GB Not feasible — — Try it →
RTX 4070 Ti Super 16.0 GB Not feasible — — Try it →
RTX 4090 24.0 GB Not feasible — — Try it →
RTX 5090 32.0 GB Not feasible needs 4 cards — Try it →
RTX A6000 48.0 GB Not feasible needs 3 cards — Try it →
A100 80GB 80.0 GB Not feasible Q2_K ≈643 tok/s Try it →
H100 80GB 80.0 GB Not feasible Q2_K ≈1057 tok/s Try it →
RX 7900 XTX 24.0 GB Not feasible — — Try it →
Mac mini M4 Pro (48GB) 36.0 GB Not feasible needs 4 cards — Try it →
Mac Studio M4 Max (64GB) 48.0 GB Not feasible needs 3 cards — Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Not feasible needs 2 cards — Try it →
Mac Studio M5 Ultra (96GB) 72.0 GB Not feasible needs 2 cards — Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

Why does a 180B model run at 6B speed?
Sparse MoE activation: only 10 of 512 routed experts plus 1 shared expert fire per token — ~6B active. Generation speed depends on the weights read per token (~3.7GB at Q4_K_M), not the total parameter count.
Can a 96GB Mac Studio run Qwen3.8-Flash-Next?
Q4_K_M (~112.6GB) is impossible; even the smallest Q2_K (~73.6GB) misses 72GB usable by 1.6GB, landing in 'needs multi-GPU' territory. The single-machine answers are a 128GB M5 Max (Q2 comfortable / Q3 tight) or a 256GB M5 Ultra (Q4 comfortable).
What should I check before commercial use?
It ships under the Qwen Community License 1.0 (not Apache-2.0) — read the commercial and redistribution terms. Also, the 51.2B n-gram table can be offloaded: slightly lower quality, much lower memory pressure.

Related guides

Try Qwen3.8 Flash-Next in the GPU compatibility checker →

Data verified 2026-10-01