All articles
LLM API Comparison & Cost/

Mistral Large 4 API Pricing and Specs

At $0.68 input and $2.09 output per million tokens, the 1M context MoE model redefines scale.

4 minutes read

Mistral Large 4 API Pricing and Specs. A macro close-up of a high-end GPU heat sink fins glowing with a soft, cool blue ambient light, with a single, tightly coiled fiber optic cable resting across the metallic surface in a dark, — GPUniq
Data as of 4 sources

Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.

TL;DR

Mistral Large 4 costs $0.68 per million input tokens and $2.09 per million output tokens. It features a 1,048,576 token context window and native vision processing. Built on a sparse mixture-of-experts architecture with roughly 52B active parameters out of 1.05T total, it scored 93.3% on Lakera's B3 AI Security Benchmark.

How Much Does Mistral Large 4 Cost?

API rates for Mistral Large 4 are $0.68 per 1M input tokens and $2.09 per 1M output tokens. These rates apply uniformly across standard text and visual context inputs.

Input costs stay low even on large payloads. A prompt carrying 500,000 input tokens costs $0.34 per API call. Maxing out the 1,048,576 token input context runs $0.71 per request. For outputs, generating a full 262,144 token response costs $0.55.

Operational PayloadInput TokensOutput TokensTotal Request Cost
Short Query10,0001,000$0.0089
Medium Document100,00010,000$0.0889
Deep Analysis500,00050,000$0.4445
Full Context Payload1,048,576262,144$1.2610

What Are the Model Specs?

Mistral Large 4 uses a sparse Mixture-of-Experts (MoE) design. It routes incoming tokens dynamically to use ~52B active parameters out of ~1.05T total parameters.

SpecificationValue
ArchitectureSparse Mixture-of-Experts (MoE)
Total Parameters~1.05T
Active Parameters per Token~52B
Vision Encoder Parameters1.6B
Maximum Context Window1,048,576 tokens
Maximum Output Tokens262,144 tokens

The model takes context up to 1,048,576 tokens. Single API requests can output up to 262,144 tokens.

Self-hosting requires massive hardware capacity. Storing the parameters requires about 1,050 GB of VRAM in FP8 precision or roughly 525 GB in a 4-bit footprint.

Does It Support Images?

Mistral Large 4 processes image inputs natively alongside text prompts. The architecture includes a dedicated 1.6B parameter vision encoder.

You can feed the model text documents, charts, visual diagrams, and code screenshots. It handles visual grounding, structured table parsing, and document question-answering directly within standard API calls.

How Does It Compare to Mistral Large 3?

Mistral Large 4 scales capacity to ~1.05T total parameters with a sparse MoE layout, moving past the architecture of Mistral Large 3. It routes tokens to ~52B active parameters to balance execution speed with model capacity.

Feature or MetricMistral Large 3Mistral Large 4
ArchitectureNot specified in dataSparse MoE (~1.05T total)
Active Parameters-~52B per token
Input Price per 1M Tokens-$0.68
Output Price per 1M Tokens-$2.09
Maximum Context Window-1,048,576 tokens
Vision Support-Text + Image (1.6B vision encoder)

How Does It Score on Benchmarks?

Mistral Large 4 hit 93.3% on Lakera's public B3 AI Security Benchmark. That is the highest score recorded for an open-weight model on that evaluation.

BenchmarkScoreRanking / Context
Lakera B3 AI Security93.3%Highest among open-weight models
Vals AI (Harvey's Legal Agent)Rank #6 of 75Evaluated across 75 models

Performance depends heavily on your provider configuration. Some deployment endpoints cap the context limit at 512K tokens rather than the full 1M token ceiling.

How Do You Access the API?

Mistral AI released the model in public preview on October 6, 2026, via Mistral Studio and the Mistral API. The company promised open weights by the end of October 2026 under a proprietary license.

Use mistral-large-4 or mistral-large-4-0 as your model identifier.

curl https://api.mistral.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_API_KEY" \
  -d '{
    "model": "mistral-large-4",
    "messages": [
      {
        "role": "user",
        "content": "Analyze this contract for risk vectors."
      }
    ],
    "max_tokens": 1000
  }'

Integration support for frameworks like vLLM or llama.cpp was unconfirmed at launch.

What Are the Key Use Cases?

  • Legal Discovery: Pass full agreements into the 1M token context to locate specific clauses without chunking.
  • Secure Agents: Build agentic loops and tool-use workflows that demand defense against injection attacks.
  • Financial Analysis: Extract information from financial reports containing inline tables, images, and unstructured text using the 1.6B vision encoder.
  • Software Maintenance: Analyze large code bases in a single prompt and output long refactored completions up to 262,144 tokens.

Methodology

Facts are compiled from the vendor's own announcement and documentation, linked below, at the time of writing; where the vendor hasn't published a number it is marked as such. GPUniq prices, when the model is already available here, come from our live catalog.

Sources

  1. 1.Introducing Mistral Large 4 | Mistral — Mistral AI (accessed )
  2. 2.Mistral Large 4 - Mistral AI | Mistral Docs — Mistral AI (accessed )
  3. 3.What Is Mistral Large 4? Architecture, Context, Access, and Cost — Hugging Face (accessed )
  4. 4.Mistral Large 4 benchmarks & pricing · The Model Gap — The Model Gap (accessed )

Want cheap GPUs for your next project?

Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.

More from the GPUniq blog