GPUFits

GPUs / Mac Studio M5 Ultra (96GB)

Mac Studio M5 Ultra (96GB) for Local LLMs: What It Runs and How Fast

The Mac Studio M5 Ultra 96GB is Apple's new flagship, announced August 25, 2026 and shipping September 22: 1.2TB/s of unified-memory bandwidth — the highest Apple has ever shipped, 50% above the M3 Ultra — at a $5,499 starting price, just $200 above the repriced M3 Ultra. Usable memory is the same ~72GB: gpt-oss-120b Q4 (~73.6GB) still misses by a hair, Q3_K_M (~60GB) is tight, 70B Q6_K (~62GB) is comfortable; but everything that fits runs about half again as fast — gpt-oss-120b theoretically ~288 tok/s, dense 70B ~21 tok/s.

The real choice is 96GB Ultra versus 128GB M5 Max (614GB/s): if your working set fits in 72GB, take the Ultra's bandwidth; only 85-100GB working sets justify the 128GB Max. The 256GB Ultra (+~$4,000) is the tier that genuinely changes the big-model map — Qwen3.8-Flash-Next Q4 (~112.6GB) and GLM-5.3-Flash 3-bit (~129GB) both fit comfortably; the 512GB version ships in late October. Up to four Mac Studios can pool memory over Thunderbolt 5. tdpW is a whole-system estimate.

Specs

Nominal VRAM96 GB
Usable VRAM (for models)72.0 GB
Memory bandwidth1200 GB/s
TypeUnified memory
Release year2026
MSRP$5,499
TDP140 W

Apple unified memory is counted at 75% usable for models (the rest serves the system/display).

Value Metrics

Bandwidth per dollar
0.22 GB/s/$
Usable VRAM per dollar
0.013 GB/$

Based on MSRP (no used price yet); used prices are market estimates, see commentary for volatility.

Model Verdict Matrix (Q4_K_M @8K)

Model VRAM needed Verdict Recommended quant Theoretical speed
Llama 3.2 3B 4.4 GB Comfortable FP16 ≈140 tok/s Try it →
Llama 3.1 8B 7.5 GB Comfortable FP16 ≈56 tok/s Try it →
Qwen3 8B 7.7 GB Comfortable FP16 ≈55 tok/s Try it →
Phi-4 14B 12.2 GB Comfortable FP16 ≈31 tok/s Try it →
Mistral Small 3.2 24B 17.5 GB Comfortable FP16 ≈19 tok/s Try it →
Gemma 3 27B 22.4 GB Comfortable FP16 ≈16 tok/s Try it →
Qwen3.8 27B 20.7 GB Comfortable FP16 ≈16 tok/s Try it →
Muse Glimmer 30B 20.1 GB Comfortable FP16 ≈15 tok/s Try it →
Qwen3 30B-A3B 21.0 GB Comfortable FP16 ≈136 tok/s Try it →
Qwen3 32B 23.7 GB Comfortable FP16 ≈14 tok/s Try it →
gpt-oss-20b 14.7 GB Comfortable FP16 ≈125 tok/s Try it →
Llama 3.3 70B 47.4 GB Comfortable Q6_K ≈16 tok/s Try it →
gpt-oss-120b 73.6 GB Needs multi-GPU Q3_K_M ≈353 tok/s Try it →
Qwen3.8 Flash-Next 112.6 GB Not feasible needs 2 cards — Try it →
GLM-5.3-Flash 198.7 GB Not feasible needs 3 cards — Try it →
DeepSeek-R1 671B 413.1 GB Not feasible — — Try it →

Theoretical speed = bandwidth × 0.75 ÷ per-token weight bytes (active params for MoE); real-world results vary with framework/driver/CPU, ±30%.

FAQ

What can the Mac Studio M5 Ultra 96GB run?
With ~72GB usable, capacity conclusions match the M3 Ultra: gpt-oss-120b Q3_K_M (~60GB) tight, 70B Q6_K (~62GB) comfortable, everything else comfortable; at Q4, gpt-oss-120b (73.6GB) misses by 1.6GB. The difference is speed: 1.2TB/s makes whatever fits run ~50% faster.
96GB M5 Ultra or 128GB M5 Max?
If the target model fits within 72GB usable: take the Ultra — 1.2TB/s vs 614GB/s is nearly 2×. If the real working set is 85-100GB (e.g. an 80-90GB model file): the 128GB Max — bandwidth cannot accelerate weights that never fit. ~119GB-class Q4 models fit neither; that needs the 256GB Ultra.
Is it worth $200 over the M3 Ultra?
Yes. Local LLM generation is bandwidth-bound, so 1.2TB/s vs 819GB/s converts directly into ~+50% tok/s, and the M5 Ultra's per-GPU-core Neural Accelerators speed up long-prompt processing too. The only reason to pick an M3 Ultra is a clearly cheaper used/discounted channel.

Related guides

Try Mac Studio M5 Ultra (96GB) in the GPU compatibility checker →

Data verified 2026-10-01