GPUFits

Can Mac Studio M5 Ultra (96GB) run Mistral Small 3.2 24B?

✅ Comfortable
Mistral Small 3.2 24B @ Q4_K_M · 8K context
Needed
17.5 GB
Usable
72.0 GB

VRAM breakdown

Weights (Q4_K_M)14.7 GB
KV cache (8K context)1.3 GB
Runtime overhead1.5 GB
Total17.5 GB

Computed at 8K context with fp16 KV cache. Longer contexts need more VRAM — use the VRAM calculator for other settings.

Recommended quantization

FP16

Estimated generation speed

~19 tok/s @ FP16

Theoretical estimates based on published architecture data and measured GGUF sizes; real-world speed varies ±30%.

Cheaper GPUs that run it

Smaller models this GPU runs well

Other GPUs that run Mistral Small 3.2 24B

FAQ

How much VRAM does Mistral Small 3.2 24B need?
At Q4_K_M with 8K context: 14.7GB weights + 1.3GB KV cache + 1.5GB runtime overhead = 17.5GB total. Mac Studio M5 Ultra (96GB) offers 72.0GB usable VRAM, so the verdict is: Comfortable.
What is the best quantization for Mistral Small 3.2 24B on Mac Studio M5 Ultra (96GB)?
FP16 — the highest tier that still fits within 72.0GB usable VRAM at 8K context. Lower tiers (Q3/Q2) fit too but cost noticeable quality.
How fast does Mistral Small 3.2 24B run on Mac Studio M5 Ultra (96GB)?
About 19 tokens/s at FP16 (theoretical estimate, ±30% in real-world use).

Check another combination in the GPU Checker →

Data verified 2026-10-01