Runs wellfp16 · 4K ctx · batch 1

Run Mixtral-8x7B-v0.1 on H200

At fp16 precision, Mixtral-8x7B-v0.1 (47 B MoE parameters) needs roughly 111.2 GB of VRAM. H200 has 141 GB, leaving 29.8 GB of headroom — plenty of room for longer contexts and larger batch sizes. Expected throughput ≈ 132 tokens/sec on a single card.

79%
VRAM
Runs well
Estimated VRAM (fp16)
111.2 GB
of 141 GB on H200· 29.8 GB free
Tok/sec
132
TFTT
923ms
Power
395W
Rent on GPUniq marketplace
H200
Live prices from verified providers · billed hourly, no commitment.
Rent now

Memory breakdown (fp16)

Shared Weights
14.00 GB
Expert Weights
93.40 GB
KV Cache
0.50 GB
Activations
2.00 GB
Framework Overhead
1.27 GB

Quantization comparison

Lower precision = less VRAM with a small quality trade-off. Quality order: FP16 > INT8 > INT4.

PrecisionVRAMUtilisationFits on H200?
FP16 (full precision)111.2 GB79%Runs well
INT8 (8-bit)56.3 GB40%Runs easily
INT4 / Q4 (4-bit)28.9 GB21%Runs easily

Frequently asked

Can H200 run Mixtral-8x7B-v0.1?
Yes. Mixtral-8x7B-v0.1 needs ≈ 111.2 GB VRAM at fp16 and H200 provides 141 GB. Expected throughput is 132 tokens/sec per GPU.
How much VRAM does Mixtral-8x7B-v0.1 use?
About 111.2 GB at fp16, 28.9 GB at INT4 (for a 4K context, batch size 1). Longer contexts add to KV cache size; larger batches increase activations.
What's the fastest way to run Mixtral-8x7B-v0.1 on H200?
Use a production inference engine like vLLM, SGLang, or TensorRT-LLM with paged attention — they cut KV cache 2–3× vs the naive estimate and batch multiple requests efficiently. For H200, enable Flash Attention 2.
Where can I rent a H200?
GPUniq aggregates live H200 offers from verified providers. You can deploy an instance in about a minute and pay hourly with no commitment.
Try it live

Chat with Mixtral-8x7B-v0.1 — no setup

Send a prompt and see how Mixtral-8x7B-v0.1 responds directly in our chat. No installation, no GPU required to test.

Open in chat
Want to tweak sequence length, batch size, or fine-tuning? Open the full calculator →