Mistral Large 4 Released: 1T Parameters and API Pricing
Mistral Large 4 debuts a one-trillion-parameter MoE architecture trained on 3,800 Grace Blackwell GPUs, with preview API access starting at $1.36 per million input tokens.
Mistral Large 4 launched preview API access on October 6, 2026, introducing a Mixture-of-Experts architecture with one trillion total parameters and 49 billion active parameters. The preview API sets rates at $1.36 per million input tokens and $4.18 per million output tokens, ahead of an open-weights release scheduled for download and self-hosting at the end of October 2026.12
Mistral trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs housed across its own datacenters in Europe. The model is too large to run on consumer laptops and desktop computers, targeting corporate datacenters and private clouds for enterprise deployment.12
API rates and prompt caching
The $1.36 per million input token price reflects a notable shift from earlier releases. According to our model pricing data, Mistral Large and Mistral Large 3 both billed at $0.50 input and $1.50 output per million tokens, while Mistral Medium 3.5 billed at $1.50 input and $7.50 output.13
Cached tokens in the Mistral Chat Completion API are billed at 10% of standard input token pricing, making repeated prompt context cost $0.136 per million tokens. The API interface supports two reasoning effort levels: none and high.45

Enterprises planning on-premise deployment ahead of the open weights release can review hardware constraints via our self-hosted AI agents GPU sizing guide.1
Agentic coding and enterprise benchmarks
Mistral Large 4 scores 38 on Artificial Analysis, a marked improvement over the 9 scored by Mistral Large 3, placing it directly behind DeepSeek 4.1 Flash at 552 billion parameters. On the Coding Agent Index, the model achieved a 49.8% combined score.25
Evaluation results across software development benchmarks show strong performance on agentic workflows:12
| Benchmark | Domain | Mistral Large 4 Score | Key Competitor Comparison |
|---|---|---|---|
| DeepSWE v1.1 | Agentic coding | 61.7% | GLM-5.3: 61%, DeepSeek V4 Pro: 57% |
| Surge AI Blind Coding | Coding quality (1 to 5) | 3.74 | Claude Opus 5: 4.22, GLM-5.3: 3.60 |
| Terminal-Bench 4.0 | Terminal execution | 28.3% | Not reported |
| SWE-Atlas-QnA | Software engineering QA | 59.4% | Not reported |
| AutomationBench | Business workflows across Gmail, Sheets, Slack, and Salesforce | 59.9% | Not reported |
| Cybench | Security exercises (40 tests) | 93% | Top-tier open-weight benchmark |
| AA-Briefcase | Long-horizon knowledge work | 1,393 Elo | Ahead of DeepSeek V4 Pro |
| FinWorkBench | Finance domain tasks | 67% | DeepSeek V4 Pro: 67%, GLM-5.3: 65% |
| Harvey Legal Agent | Legal domain reasoning | 15% | GPT-6 Astra: 5%, Kimi K3: 13% |
| DIOR-RSVG | Visual grounding | 73% | GPT-6 Astra: 68%, Kimi K3: 55% |
| Dense 200 | Visual grounding | 42% | GPT-6 Astra: 41% |
Datacenter deployment and sovereign architecture
The training architecture reflects a substantial scale-up from Mistral Large 3, which utilized 675 billion total parameters and 41 billion active parameters. Mistral Large 4 spans more than 160 languages, including every official language of the European Union, trained following Mistral's €3 billion Series D funding round at a post-money valuation exceeding €21 billion.267
Teams requiring sovereign on-premise execution can test workloads immediately using API reasoning endpoints before committing datacenter capacity to the late-October weights.15