Won't fitfp16 · 4K ctx · batch 1

Does RTX 5090 run DeepSeek-V3 671B? No at fp16.

At fp16 precision, DeepSeek-V3 671B (671 B MoE parameters) needs roughly 1428.9 GB of VRAM. RTX 5090 only has 32 GB, short by 1396.9 GB. INT4 quantization brings it down to 359.2 GB.

>999%
VRAM
Won't fit
Estimated VRAM (fp16)
1428.9 GB
of 32 GB on RTX 5090· 1396.9 GB short
Tok/sec
16
TFTT
5830ms
Power
575W
Rent on GPUniq marketplace
RTX 5090
Live prices from verified providers · billed hourly, no commitment.
Rent now

Memory breakdown (fp16)

Shared Weights
74.00 GB
Expert Weights
1342.00 GB
KV Cache
1.67 GB
Activations
6.67 GB
Framework Overhead
4.54 GB

Quantization comparison

Lower precision = less VRAM with a small quality trade-off. Quality order: FP16 > INT8 > INT4.

PrecisionVRAMUtilisationFits on RTX 5090?
FP16 (full precision)1428.9 GB4465%Won't fit
INT8 (8-bit)715.8 GB2237%Won't fit
INT4 / Q4 (4-bit)359.2 GB1123%Won't fit

Frequently asked

Can RTX 5090 run DeepSeek-V3 671B?
Not at fp16 — DeepSeek-V3 671B needs about 1428.9 GB while RTX 5090 has 32 GB. It fits at INT4 quantization (359.2 GB) with some accuracy trade-off.
How much VRAM does DeepSeek-V3 671B use?
About 1428.9 GB at fp16, 359.2 GB at INT4 (for a 4K context, batch size 1). Longer contexts add to KV cache size; larger batches increase activations.
What's the fastest way to run DeepSeek-V3 671B on RTX 5090?
Use a production inference engine like vLLM, SGLang, or TensorRT-LLM with paged attention — they cut KV cache 2–3× vs the naive estimate and batch multiple requests efficiently. For RTX 5090, enable Flash Attention 2.
Where can I rent a RTX 5090?
GPUniq aggregates live RTX 5090 offers from verified providers. You can deploy an instance in about a minute and pay hourly with no commitment.
Try it live

Chat with DeepSeek-V3 671B — no setup

Send a prompt and see how DeepSeek-V3 671B responds directly in our chat. No installation, no GPU required to test.

Open in chat
Want to tweak sequence length, batch size, or fine-tuning? Open the full calculator →
None of the common quantization levels fit this model on RTX 5090. Consider multi-GPU deployment or a larger card.