GPUFits

Models / GLM-5.3-Flash

GLM-5.3-Flash VRAM Requirements & GPU Pairing

GLM-5.3-Flash (Z.ai, released Aug 25, 2026, MIT) is September's dark horse in local deployment: a 321.3B-total, 18B-active MoE where 11 of 45 layers use DSA sparse attention (MLA compressed latents, 512+64 dims, same family as DeepSeek) and 34 use KDA linear attention, with a native 1M context. It topped OpenRouter's daily charts in September, with Unsloth 3-bit GGUFs and Mac/DGX Spark deployment guides following. Q4_K_M weighs ~196.8GB — like DeepSeek-R1, no consumer single card comes close.

Realistic paths: the Unsloth 3-bit quant (~129GB) fits a 256GB M5 Ultra comfortably or 3×A100 80GB tightly (at Q4); otherwise use Z.ai's official API. At 18B active, whatever holds it is fast — an M5 Ultra at 1.2TB/s theoretically reaches ~82 tok/s. The KV cache is tiny: only 11 layers actually hold KV, ~0.1GB measured at 8K and ~13GB at full 1M context; this site conservatively counts all 45 layers, showing roughly 4× the real value.

Architecture Specs

Total parameters321.3B
Active parameters(MoE: VRAM follows total params, speed follows active params) 18B
Layers45
KV heads(MLA latent compression: KV = layers × (kvLoraRank + qkRopeHeadDim) × bytes — the standard GQA formula does not apply) —
Head dim—
Native context1024K
Licensemit
KV bytes per token (fp16)50.6 KB

Hybrid architecture: only 11 of 45 layers are MLA sparse attention (DSA); the 34 KDA linear layers' KV does not grow with context. The KV column above conservatively applies the MLA formula to all 45 layers — real KV is roughly 1/4 of the shown value.

VRAM by Quant Tier (@8K context, fp16 KV)

Quant Weights Total (incl. 1.5GB runtime overhead)
Q8_0 341.8 GB 343.7 GB
Q6_K 263.9 GB 265.8 GB
Q4_K_M 196.8 GB 198.7 GB
Q3_K_M 160.7 GB 162.6 GB
Q2_K 127.3 GB 129.2 GB

KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.

GPU Verdict Matrix (Q4_K_M @8K)

GPU Usable VRAM Verdict Recommended quant Theoretical speed
RTX 3060 12GB 12.0 GB Not feasible — — Try it →
RTX 3090 24.0 GB Not feasible — — Try it →
RTX 4070 Ti Super 16.0 GB Not feasible — — Try it →
RTX 4090 24.0 GB Not feasible — — Try it →
RTX 5090 32.0 GB Not feasible — — Try it →
RTX A6000 48.0 GB Not feasible — — Try it →
A100 80GB 80.0 GB Not feasible needs 3 cards — Try it →
H100 80GB 80.0 GB Not feasible needs 3 cards — Try it →
RX 7900 XTX 24.0 GB Not feasible — — Try it →
Mac mini M4 Pro (48GB) 36.0 GB Not feasible — — Try it →
Mac Studio M4 Max (64GB) 48.0 GB Not feasible — — Try it →
Mac Studio M3 Ultra (96GB) 72.0 GB Not feasible needs 3 cards — Try it →
Mac Studio M5 Ultra (96GB) 72.0 GB Not feasible needs 3 cards — Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

What hardware does GLM-5.3-Flash actually need?
Q4_K_M totals ~198.7GB: tight on 3×A100 80GB, impossible on 2×A100. A 256GB M5 Ultra (~192GB usable) just misses Q4, runs Q3_K_M (~162.6GB) tight and 3-bit (~129GB) comfortably. No consumer single/dual-card path exists.
Why is the KV shown here so much larger than measured?
Only 11 of its 45 layers hold KV (MLA, 512+64 dims each); the other 34 KDA linear-attention layers' KV does not grow with context. Our MLA formula conservatively counts all 45 layers — a ~4× overestimate: real KV at 8K is ~0.1GB, not 0.4GB.
GLM-5.3-Flash or DeepSeek-R1?
GLM-5.3-Flash is under half R1's size (321B vs 671B) with 18B vs 37B active, also MIT, plus a native 1M context. Neither fits consumer single-card setups, but the 3-bit + 256GB Mac path is far more attainable than R1's. Pick R1 for the absolute reasoning benchmark, GLM-5.3-Flash for long context and reachability.

Related guides

Try GLM-5.3-Flash in the GPU compatibility checker →

Data verified 2026-10-01