Mistral Large 4 Opens API Preview With 1 Trillion Parameters and European Infrastructure
Sparse MoE activates 52B of 1.05T parameters per token; open weights and license terms due end of October

Mistral AI opened public preview access to Mistral Large 4 on October 6, 2026, making a 1.05 trillion-parameter mixture-of-experts model available via its La Plateforme API — and promising to release the full model weights by the end of the month. If that release arrives as scheduled, it would make Large 4 by far the largest open-weight language model ever publicly released, surpassing Meta's Llama 3.1 405B by more than a factor of two on total parameter count.
The model, internally nicknamed "Le Chonk," was trained at Mistral's own European data centers using 3,800 NVIDIA Grace Blackwell GPUs — a substantial infrastructure investment that reflects both the €3 billion the company raised in its September 2026 Series D round, led by Samsung, and its sustained argument that AI capability can be built on European soil without routing data through US-controlled cloud infrastructure.
A Sparse Architecture That Punches Above Its Compute Weight
The defining engineering choice in Large 4 is sparse activation. Despite totaling 1.05 trillion parameters, the model routes each token through roughly 52 billion of them — around five percent of the full network — rather than activating every weight on every forward pass. This is the core tradeoff of mixture-of-experts design: total parameter count (and therefore the memory the model requires) is enormous, but per-token inference compute is comparable to a dense model roughly one-twentieth the size.
That distinction matters for pricing and scalability. During the current preview period, Mistral is charging $0.68 per million input tokens and $2.09 per million output tokens — rates that become more economically defensible once you understand that inference is not bearing the cost of 1 trillion active weights but of roughly 52 billion. Standard pricing after the preview period is listed at $1.36 and $4.18 per million tokens, respectively.
The model also processes a context window of one million tokens and incorporates a separate 1.6-billion-parameter vision encoder for multimodal inputs. Mistral trained Large 4's vision component as a distinct module attached to the LLM backbone rather than tightly integrating multimodal routing into the core MoE structure.
Reinforcement Learning at Industrial Scale With Mistral Forge
Post-training reinforcement learning is now the dominant lever for extracting capability from pretrained models, and Large 4 represents Mistral's most ambitious deployment of it. The company built and used an internal RL infrastructure called Mistral Forge, which runs a composable training environment: rather than applying sequential fine-tuning phases for different behavior domains, Forge combines reward signals for conversational quality, scientific reasoning, safety, and long-horizon tool use within a single training run.
The scale of that environment is notable — 33 billion tokens per day of RL throughput, with approximately 16 billion of those being trainable completion tokens. For reference, most published academic RL fine-tuning setups operate at a small fraction of that throughput; Mistral is effectively running industrial-scale preference optimization in its own European compute stack.
European Infrastructure as a Product Feature
Mistral's infrastructure story is inseparable from its market positioning. CEO Arthur Mensch has testified before the French parliament on AI sovereignty and has consistently argued that European enterprises should not be compelled to process sensitive data on US-hyperscaler infrastructure subject to the CLOUD Act.
Large 4 is trained and served from Mistral's own European facilities. The practical implication for enterprises operating under GDPR or sector-specific data-residency requirements — healthcare, finance, legal, government — is that they can use a frontier-class model without routing data through AWS, Azure, or Google Cloud. That distinction is not merely rhetorical: it affects enterprise procurement decisions in regulated industries across the EU.
The €3 billion Series D (led by Samsung, valuing Mistral at more than €21 billion) provides the capital base for exactly this kind of infrastructure investment. Samsung's involvement introduces questions about data governance that Mistral has not publicly addressed in detail, though the reported deal structure preserves the company's operational independence.
Where Large 4 Leads — and What the Benchmark Numbers Actually Mean
Mistral claims Large 4 leads across a range of evaluations: the DeepSWE v1.1 agentic coding benchmark, the Artificial Analysis Intelligence Index, the Harvey Legal Agent Benchmark, KORA, SciCode-Verified, and Dense 200. The Artificial Analysis Intelligence Index result carries meaningful weight because Artificial Analysis is an independent third-party evaluator rather than Mistral itself.
The Harvey Legal Agent Benchmark result is noteworthy but requires context: Harvey is a legal AI company that uses Mistral's models, making it a partner rather than a disinterested party. Its benchmark results are credible in the sense that Harvey is measuring its own domain expertise, but readers should understand the relationship when evaluating the claim.
Read more: StepFun's Step 5 Preview, a 600B MoE model with open weights promised October 15
Mistral's most dramatic-sounding claims involve cybersecurity: CyberGym-E2E (82%), Cybench, and the B3 AI Security Benchmark, where Large 4 reportedly outperforms Claude Opus 5.5 and GPT-6 Astra by very wide margins — in some cases by more than 80 percentage points. The explanation for those gaps is not primarily that competing models lack the technical capability to complete cybersecurity tasks. Claude Opus 5.5 and GPT-6 Astra are designed to refuse offensive security requests as a matter of safety policy, producing near-zero completion rates by deliberate abstention. DeepSeek V4 Pro 0813, a Chinese open-weight model with less aggressive safety filtering, scores substantially higher than the closed US models on the same benchmarks — consistent with the policy-refusal explanation.
This is a real product differentiation, not a fabricated one: Mistral explicitly markets Large 4 to security researchers and red teams who need a model that will complete offensive tasks. But framing the cybersecurity benchmark results as pure capability leadership over Anthropic and OpenAI is misleading without that context.
Deployment Constraints and What Remains Unknown
At 1.05 trillion total parameters, deploying Large 4 outside of Mistral's API will require substantial hardware. A full-precision (BF16) deployment would demand roughly two terabytes of GPU memory, requiring dozens of high-end accelerators. Practical open-weight deployment will depend almost entirely on quantized variants, which Mistral has not yet published. The compute and memory requirements mean "open weights" for Large 4 means something different than it does for a 7B or 70B model — it will be within reach of well-resourced organizations and universities, not individual developers on consumer hardware.
The license for the open-weight release is also unknown. Mistral's previous Large 3 shipped under Apache 2.0, the most permissive major open-source license; Large 4's terms have not yet been published. Commercial use conditions, fine-tuning restrictions, and deployment requirements will all depend on what Mistral announces alongside the weights.
Independent validation of the 1 million-token context window's quality at limit has not yet been published. Long-context degradation — where models lose coherent recall of early-context information as the window fills — is a documented failure mode in frontier models even when they nominally support million-token contexts. Mistral's claims about context quality are plausible given the model's scale but have not been confirmed by external evaluators.
What the Open-Weight Release Would Change
The approaching open-weight release is the larger event for the AI research community and for enterprise AI buyers. A freely deployable 1 trillion-parameter MoE at frontier capability — assuming Mistral's benchmark claims hold up to independent scrutiny — would fundamentally alter the open versus closed model calculus. Organizations that have been forced to choose between capability (closed US models) and control (smaller open-weight models) would gain a third option: frontier capability on their own infrastructure, subject to whatever license terms Mistral publishes.
For US-China AI competition dynamics, an open-weight European frontier model changes the geometry. European technology policy — which has pushed for its own AI capability — would gain a concrete proof of concept. The question is whether the late-October release arrives on time, whether the license is genuinely open, and whether the quantized variants make self-hosting tractable for organizations without data-center-scale GPU budgets. The window for those answers to arrive is now weeks, not months.