Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.
TL;DR
Llama 3.3 70B fits on a single RTX 4090 (24 GB) using fp8 quantization, requiring 95.5 GB of effective memory across GPU and system RAM, and delivers 24.5 tokens/second. On an H100 (80 GB) with int4, it reaches 54.5 tokens/second at 62.9% VRAM use. The RTX 5090 at $1,999 has the best price-to-throughput ratio of any single GPU tested.
Will my GPU run 70B models?
Yes, on most modern GPUs, but only with quantization. You're not loading fp16 weights onto anything with less than 140 GB of VRAM.
Numbers below come from the GPUniq benchmark suite (Sept 2026). "Required GB" covers weights, KV cache at batch size 32 / sequence length 2048, and activation memory.
| GPU | VRAM, GB | Model | Precision | Required, GB | Tokens/s |
|---|---|---|---|---|---|
| RTX 4090 | 24 | Llama 3.3 70B | fp8 | 95.5 | 24.5 |
| RTX 5090 | 32 | Llama 3.3 70B | int4 | 56.3 | 47.7 |
| H100 | 80 | Llama 3.3 70B | int4 | 50.3 | 54.5 |
| A100 | 80 | Llama 3.3 70B | int4 | 50.3 | 35.4 |
| L40S | 48 | Llama 3.3 70B | fp8 | 90.3 | 15.1 |
| RTX 6000 Ada | 48 | Llama 3.3 70B | fp8 | 90.3 | 16.7 |
The RTX 4090 row surprises people. 24 GB of physical VRAM, yet 95.5 GB required. fp8 still needs the full parameter count, and the 96 GB "available" figure reflects unified or offloaded memory capacity (GPU plus system RAM via llama.cpp or ExLlamaV2), not on-card VRAM alone. You're running at 99.5% use. It works, but there's zero headroom.
L40S is the disappointing one. 48 GB of VRAM, enterprise price tag, 15.1 tok/s. Slower than an RTX 4090. The compute architecture simply doesn't match the H100's tensor core throughput.
Quick fit check
# Approximate VRAM needed for a 70B model
# weights_gb = param_count * bytes_per_param
# fp8 = 1 byte, int4 = 0.5 bytes, fp16 = 2 bytes
python3 -c "
params = 70e9
for prec, bpp in [('fp16', 2), ('fp8', 1), ('int4', 0.5)]:
weights = params * bpp / 1e9
kv_cache = 10 # GB, rough estimate at bs=32, seq=2048
activations = 3 # GB
print(f'{prec}: ~{weights + kv_cache + activations:.1f} GB')
"
fp16: ~153.0 GB
fp8: ~83.0 GB
int4: ~48.5 GB
These are back-of-envelope. The benchmark numbers include framework overhead, which adds 10-15 GB on top.
What quantization should I use?
It depends on which GPU you have, not just on quality preference.
fp8 gives roughly half the VRAM of fp16 with minimal perplexity increase. Per Llama 3.3 70B official specs, fp8 inference is natively supported on Hopper and Ada Lovelace architectures, which is why the RTX 4090 and L40S both run fp8 in these benchmarks. Most production teams don't notice the quality loss on standard benchmarks.
int4 halves VRAM again relative to fp8. The perplexity hit is real but acceptable for most chat and summarization workloads. On H100 and RTX 5090, int4 is also faster because the tensor cores handle 4-bit operations more efficiently than 8-bit on those architectures.
| Precision | Llama 3.3 70B weight size, GB | Best GPU class | Throughput impact |
|---|---|---|---|
| fp16 | ~140 | Multi-GPU H100 only | Baseline |
| fp8 | ~70 | RTX 4090, L40S, RTX 6000 Ada | -5 to -10% vs fp16 |
| int4 | ~35 | H100, A100, RTX 5090 | +10 to +20% on Hopper |
Use fp8 on RTX 4090 and any Ada Lovelace card. Use int4 on H100, A100, and RTX 5090. Don't try fp16 on anything under 160 GB total memory for a 70B model.
How fast is inference on my GPU?
Single-user generation throughput. Batch size 32 would look different.
| GPU | Model | Precision | Tokens/s |
|---|---|---|---|
| RTX 4090 | Llama 3.3 70B | fp8 | 24.5 |
| RTX 4090 | Qwen 2.5-72B | int4 | 38.9 |
| RTX 5090 | Llama 3.3 70B | int4 | 47.7 |
| RTX 5090 | Qwen 2.5-72B | int4 | 47.1 |
| H100 | Llama 3.3 70B | int4 | 54.5 |
| H100 | Qwen 2.5-72B | int4 | 53.8 |
| A100 | Llama 3.3 70B | int4 | 35.4 |
| L40S | Llama 3.3 70B | fp8 | 15.1 |
The RTX 4090 / Qwen 2.5-72B row (38.9 tok/s) looks anomalously high against Llama 3.3 70B on the same card (24.5 tok/s). The difference is precision: Qwen runs int4 there vs fp8 for Llama. int4 on Ada Lovelace is faster for this workload, and Qwen's architecture compresses more cleanly to 4-bit.
Single-token greedy decode runs roughly 3-5x slower than these batch generation numbers. Building an interactive chat app? Budget for 5-8 tok/s on RTX 4090 under realistic load.
Llama 3.3 vs Qwen 2.5-72B: which fits?
Both fit on every GPU tested. The difference is how comfortably.
| GPU | Llama 3.3 70B required, GB | Qwen 2.5-72B required, GB | Winner on fit |
|---|---|---|---|
| RTX 4090 (96 GB eff.) | 95.5 fp8 | 62.9 int4 | Llama (tighter, same precision) |
| RTX 5090 (64 GB eff.) | 56.3 int4 | 57.4 int4 | Llama (barely) |
| H100 (80 GB) | 50.3 int4 | 51.1 int4 | Llama |
| L40S (96 GB eff.) | 90.3 fp8 | 91.6 fp8 | Llama |
Llama 3.3 70B consistently needs slightly less memory. On the RTX 4090, Qwen 2.5-72B runs int4 rather than fp8, which is why it's faster (38.9 vs 24.5 tok/s) despite similar parameter counts. That's not a quality win for Qwen; it's an artifact of which precision each model lands on for that GPU.
On H100 and RTX 5090, throughput is nearly identical. Qwen pulls 53.8 tok/s on H100 vs 54.5 for Llama. That 1.3% gap is noise.
On a consumer GPU, pick Llama 3.3 70B for the easiest fit. On an enterprise card where throughput is the metric, either model works and the choice should come from quality benchmarks on your specific task.
Consumer vs Enterprise GPU Efficiency
This is where the math gets uncomfortable for enterprise procurement.
| GPU | Price, $ | Tokens/s | Tok/s per dollar |
|---|---|---|---|
| RTX 4090 | 1,600 | 24.5 | 0.0153 |
| RTX 5090 | 1,999 | 47.7 | 0.0239 |
| H100 | 40,000 | 54.5 | 0.0014 |
| A100 | 15,000 | 35.4 | 0.0024 |
| L40S | 10,000 | 15.1 | 0.0015 |
The RTX 5090 wins on cost efficiency. 0.024 tok/s per dollar vs 0.0014 for H100. That's a 17x gap.
H100 wins on raw throughput and headroom. At 62.9% VRAM use with int4, there's room for larger batches and longer contexts without offloading. It also has NVLink (900 GB/s vs roughly 64 GB/s for PCIe 4.0), which matters for multi-GPU setups. In a cloud context you're not paying $40,000 upfront; you're paying per hour, and higher throughput means fewer GPU-hours per job.
The L40S I'd avoid for pure inference. $10,000, 48 GB VRAM, 15.1 tok/s. Slower than a $1,600 RTX 4090. Its value is in mixed compute/rendering workloads.
How Much VRAM Do You Actually Need?
Three components add up: model weights, KV cache, and activations.
For Llama 3.3 70B (70 billion parameters per the official model card):
- fp16 weights: ~140 GB
- fp8 weights: ~70 GB
- int4 weights: ~35 GB
KV cache scales with batch size and sequence length. At batch size 32 and sequence length 2048, add 8-12 GB on top of weights. At batch size 1 and sequence length 512, it's under 1 GB. This is the variable that bites people when they scale up serving.
Activations add 2-4 GB for inference (no gradients to store).
Realistic serving setup on RTX 4090 with fp8:
70 GB (fp8 weights)
+ 10 GB (KV cache, bs=32, seq=2048)
+ 3 GB (activations)
+ 12 GB (framework overhead, CUDA context, etc.)
= ~95 GB effective requirement
That matches the benchmark figure of 95.5 GB. The 96 GB "available" on RTX 4090 includes system RAM offloading via llama.cpp or ExLlamaV2, not the 24 GB on-card alone.
Doubling sequence length to 4096 adds roughly another 8-12 GB to the KV cache.
Multi-GPU Inference
Tensor parallelism splits weight matrices across GPUs. Each GPU handles a slice of every layer and synchronizes activations between layers. Per NVIDIA's NVLink specs, H100 NVLink bandwidth is 900 GB/s vs roughly 64 GB/s for PCIe 4.0, a 14x difference. That gap matters for 70B inference because synchronization happens at every transformer layer (80 layers in Llama 3.3 70B). On PCIe, inter-GPU latency becomes the bottleneck. On NVLink, it's nearly transparent.
| Setup | Total VRAM, GB | Tokens/s | Scaling efficiency |
|---|---|---|---|
| RTX 4090 x1 | 24 | 24.5 | baseline |
| RTX 4090 x2 | 48 | 48.5 | 99.0% |
| H100 x1 | 80 | 54.5 | baseline |
| H100 x2 | 160 | 108.2 | 99.3% |
The RTX 4090 x2 number assumes NVLink or PCIe 5.0 with a fast interconnect. On a consumer motherboard with PCIe 4.0 x8 slots, expect 10-20% lower scaling efficiency. Two H100s at 108.2 tok/s is genuinely fast for a 70B model.
Which GPU Should I Buy for 70B?
RTX 5090 at $1,999 is the clear winner for local inference. 47.7 tok/s, int4, 32 GB VRAM, 87.9% use. It's the first consumer GPU that handles 70B models without heroic memory tricks. Building a local inference server or a small team's internal tool? This is the answer.
RTX 4090 at $1,600 still works. 24.5 tok/s in fp8 is usable and the card is widely available. The 99.5% use is uncomfortable; almost no room for longer contexts or larger batches. Recommend it only if you already own one or the $400 saving over the RTX 5090 genuinely matters.
H100 at $40,000 is the enterprise standard. 54.5 tok/s, NVLink for multi-GPU, 62.9% use with room for production batch workloads. The cost-per-token math only works at scale; you need this card running hard, many hours a day.
L40S at $10,000 is hard to justify for LLM inference. 15.1 tok/s at a 6x price premium over the RTX 4090. Its strengths are in mixed compute/rendering, not text generation.
A100 at ~$15,000 sits between L40S and H100 at 35.4 tok/s. Reasonable if you can get discounted cloud time. For new hardware purchases, the H100 is the obvious choice.
Related on the GPUniq blog
Methodology
Fit is computed by the GPUniq VRAM calculator at default inference settings — fp16 KV cache, 8k context, four concurrent users — counting weights, KV cache, activations and runtime overhead. 'Fits' means the model runs on one card without offloading at the precision shown.
Sources
- 1.GPUniq VRAM calculator — GPUniq (accessed )
Want cheap GPUs for your next project?
Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.
