LLM VRAM Calculator
172 LLM architectures · 130 GPUs · inference & fine-tuning
8B paramsGQA128K ctx32L · 4096dMeta
Lower precision = less VRAM, usually faster inference
4K
1K4K8K16K32K64K128K
Max for this model: 128K
1
12481632
1
12481632
Global avg ≈ 0.47 · US ≈ 0.38 · EU ≈ 0.25 · France (nuclear) ≈ 0.06
81%
VRAM
High
19.54 GB
of 24 GB available
Tok/sec
93.3
TFTT
716ms
Total
93/s
Power
172W
CO₂/hr
0.07kg
ms/tok
10.7
Rent on GPUniq marketplace
RTX 4090
Loading live price…
Memory breakdown
Model Weights
16.00 GBKV Cache
0.50 GBActivations
2.00 GBFramework Overhead
1.04 GBCarbon footprint
Per hour
0.07 kg
Per day
1.7 kg
Per month
0.05 t
Per year
0.60 t
Try it live
Chat with Llama 3.1 8B — no setup
Skip the numbers and just try it. Our chat opens Llama 3.1 8B (and any close matches) so you can send a prompt and see how it responds right away.
Popular combinations
Pre-computed VRAM, throughput and hourly price for the most searched LLM × GPU pairs.
Llama 3.1 8B
Llama 3.1 70B
Looking for a different model? Use the calculator above — it covers 172 architectures and 130 GPU SKUs.