Estimate GPU memory requirements for inference. Adjust parameters to see real-time VRAM breakdown.
| Precision | Weights | KV Cache | Activations | Total |
|---|---|---|---|---|
| FP16 / BF16 | 140 GB | 2.68 GB | 21 GB | 163.68 GB |
| INT8 | 70 GB | 1.34 GB | 10.5 GB | 81.84 GB |
| INT4 | 35 GB | 0.67 GB | 5.25 GB | 40.92 GB |
Estimates account for model weights, KV cache (2 × layers × kv_heads × head_dim × seq_len × batch × bytes), and activations (~15% of weights). Actual usage varies by framework, kernel efficiency, and fragmentation.