GPUFits

LLM VRAM Calculator

Calculate exact VRAM requirements for any LLM: weights, KV cache, and overhead by quantization level and context length. Free, data-sourced, no signup.

Total
3.9GB
Llama 3.2 3.21B · Q4_K_M · 4K
Weights
2.0 GB 50%
KV cache
0.5 GB 12%
Overhead
1.5 GB 38%

Many runtimes default to a smaller context than the model maximum (e.g. Ollama defaults to 2048 tokens). Actual usage depends on your configuration.