LLM VRAM Calculator
Calculate exact VRAM requirements for any LLM: weights, KV cache, and overhead by quantization level and context length. Free, data-sourced, no signup.
Total
3.9GB
Llama 3.2 3.21B · Q4_K_M · 4K
Weights
2.0 GB 50%
KV cache
0.5 GB 12%
Overhead
1.5 GB 38%
Many runtimes default to a smaller context than the model maximum (e.g. Ollama defaults to 2048 tokens). Actual usage depends on your configuration.
GPUs that can hold it
- rtx-3060-12gb✅ Comfortable
- rtx-4070-ti-super✅ Comfortable
- rtx-3090✅ Comfortable
- rtx-4090✅ Comfortable
- rx-7900-xtx✅ Comfortable
- rtx-5090✅ Comfortable
- mac-mini-m4-pro-48gb✅ Comfortable
- rtx-a6000✅ Comfortable
- mac-studio-m4-max-64gb✅ Comfortable
- mac-studio-m3-ultra-96gb✅ Comfortable
- a100-80gb✅ Comfortable
- h100-80gb✅ Comfortable
Figures are estimates computed from published model architectures and measured GGUF file sizes; real usage varies with runtime and drivers.