LLM VRAM Calculator

172 LLM architectures · 130 GPUs · inference & fine-tuning

8B paramsGQA128K ctx32L · 4096dMeta

Lower precision = less VRAM, usually faster inference

4K
1K4K8K16K32K64K128K

Max for this model: 128K

1
12481632
1
12481632

Global avg ≈ 0.47 · US ≈ 0.38 · EU ≈ 0.25 · France (nuclear) ≈ 0.06

81%
VRAM
High
19.54 GB
of 24 GB available
Tok/sec
93.3
TFTT
716ms
Total
93/s
Power
172W
CO₂/hr
0.07kg
ms/tok
10.7
Rent on GPUniq marketplace
RTX 4090
Loading live price…
Rent now

Memory breakdown

Model Weights
16.00 GB
KV Cache
0.50 GB
Activations
2.00 GB
Framework Overhead
1.04 GB

Carbon footprint

Per hour
0.07 kg
Per day
1.7 kg
Per month
0.05 t
Per year
0.60 t
Try it live

Chat with Llama 3.1 8B — no setup

Skip the numbers and just try it. Our chat opens Llama 3.1 8B (and any close matches) so you can send a prompt and see how it responds right away.

Open in chat

Popular combinations

Pre-computed VRAM, throughput and hourly price for the most searched LLM × GPU pairs.

Looking for a different model? Use the calculator above — it covers 172 architectures and 130 GPU SKUs.