Models / GLM-5.3-Flash
GLM-5.3-Flash VRAM Requirements & GPU Pairing
GLM-5.3-Flash (Z.ai, released Aug 25, 2026, MIT) is September's dark horse in local deployment: a 321.3B-total, 18B-active MoE where 11 of 45 layers use DSA sparse attention (MLA compressed latents, 512+64 dims, same family as DeepSeek) and 34 use KDA linear attention, with a native 1M context. It topped OpenRouter's daily charts in September, with Unsloth 3-bit GGUFs and Mac/DGX Spark deployment guides following. Q4_K_M weighs ~196.8GB — like DeepSeek-R1, no consumer single card comes close.
Realistic paths: the Unsloth 3-bit quant (~129GB) fits a 256GB M5 Ultra comfortably or 3×A100 80GB tightly (at Q4); otherwise use Z.ai's official API. At 18B active, whatever holds it is fast — an M5 Ultra at 1.2TB/s theoretically reaches ~82 tok/s. The KV cache is tiny: only 11 layers actually hold KV, ~0.1GB measured at 8K and ~13GB at full 1M context; this site conservatively counts all 45 layers, showing roughly 4× the real value.
Architecture Specs
| Total parameters | 321.3B |
| Active parameters(MoE: VRAM follows total params, speed follows active params) | 18B |
| Layers | 45 |
| KV heads(MLA latent compression: KV = layers × (kvLoraRank + qkRopeHeadDim) × bytes — the standard GQA formula does not apply) | — |
| Head dim | — |
| Native context | 1024K |
| License | mit |
| KV bytes per token (fp16) | 50.6 KB |
Hybrid architecture: only 11 of 45 layers are MLA sparse attention (DSA); the 34 KDA linear layers' KV does not grow with context. The KV column above conservatively applies the MLA formula to all 45 layers — real KV is roughly 1/4 of the shown value.
VRAM by Quant Tier (@8K context, fp16 KV)
| Quant | Weights | Total (incl. 1.5GB runtime overhead) |
|---|---|---|
| Q8_0 | 341.8 GB | 343.7 GB |
| Q6_K | 263.9 GB | 265.8 GB |
| Q4_K_M | 196.8 GB | 198.7 GB |
| Q3_K_M | 160.7 GB | 162.6 GB |
| Q2_K | 127.3 GB | 129.2 GB |
KV cache and runtime overhead do not change across quants; longer contexts grow the KV part linearly — adjust context length in the GPU checker tool.
GPU Verdict Matrix (Q4_K_M @8K)
| GPU | Usable VRAM | Verdict | Recommended quant | Theoretical speed | |
|---|---|---|---|---|---|
| RTX 3060 12GB | 12.0 GB | Not feasible | — | — | Try it → |
| RTX 3090 | 24.0 GB | Not feasible | — | — | Try it → |
| RTX 4070 Ti Super | 16.0 GB | Not feasible | — | — | Try it → |
| RTX 4090 | 24.0 GB | Not feasible | — | — | Try it → |
| RTX 5090 | 32.0 GB | Not feasible | — | — | Try it → |
| RTX A6000 | 48.0 GB | Not feasible | — | — | Try it → |
| A100 80GB | 80.0 GB | Not feasible | needs 3 cards | — | Try it → |
| H100 80GB | 80.0 GB | Not feasible | needs 3 cards | — | Try it → |
| RX 7900 XTX | 24.0 GB | Not feasible | — | — | Try it → |
| Mac mini M4 Pro (48GB) | 36.0 GB | Not feasible | — | — | Try it → |
| Mac Studio M4 Max (64GB) | 48.0 GB | Not feasible | — | — | Try it → |
| Mac Studio M3 Ultra (96GB) | 72.0 GB | Not feasible | needs 3 cards | — | Try it → |
| Mac Studio M5 Ultra (96GB) | 72.0 GB | Not feasible | needs 3 cards | — | Try it → |
Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.
FAQ
- What hardware does GLM-5.3-Flash actually need?
- Q4_K_M totals ~198.7GB: tight on 3×A100 80GB, impossible on 2×A100. A 256GB M5 Ultra (~192GB usable) just misses Q4, runs Q3_K_M (~162.6GB) tight and 3-bit (~129GB) comfortably. No consumer single/dual-card path exists.
- Why is the KV shown here so much larger than measured?
- Only 11 of its 45 layers hold KV (MLA, 512+64 dims each); the other 34 KDA linear-attention layers' KV does not grow with context. Our MLA formula conservatively counts all 45 layers — a ~4× overestimate: real KV at 8K is ~0.1GB, not 0.4GB.
- GLM-5.3-Flash or DeepSeek-R1?
- GLM-5.3-Flash is under half R1's size (321B vs 671B) with 18B vs 37B active, also MIT, plus a native 1M context. Neither fits consumer single-card setups, but the 3-bit + 256GB Mac path is far more attainable than R1's. Pick R1 for the absolute reasoning benchmark, GLM-5.3-Flash for long context and reachability.
Related guides
- Running DeepSeek-R1 Locally: The Complete Hardware Guide
- MoE Model Hardware Requirements: Total vs Active Parameters
- KV Cache Explained: Why Long Contexts Eat Your VRAM
Try GLM-5.3-Flash in the GPU compatibility checker →
Data verified 2026-10-01