source preview
source preview Open source material ↗
Source figure
来自一手来源:web_source Open source material ↗
Mistral Large 4
Mistral Large 4 Open source material ↗

Evidence at a glance

parameters1.05T; token 49BEvidence
1M tokensEvidence
1.6B parametersEvidence
use 3,800 Grace Blackwell GPUEvidence
160Evidence
APIInput$1.36, $4.18/ tokensEvidence

The mechanism in one line

InputReduce the input to a workable scale

Compress the visual or contextual input before the main reasoning path.

MechanismSpend compute where it matters

Route or verify the expensive step instead of repeating the full path.

OutcomeEnd with a measurable workflow result

Translate the mechanism into a bounded deployment or evaluation check.

The Service Comes Before the Download

On October 6, 2026, Mistral AI released a public preview of Mistral Large 4, or ML4. It is a multimodal mixture-of-experts model aimed at coding, agent workflows, image understanding, and enterprise tasks. Its API is available now, while the weights are scheduled for release at the end of October. For now, “open” means teams can call the hosted service, not that they can download the model into their own infrastructure.

That gap matters to technical leaders because it separates two capabilities that are often conflated: trying a model through a hosted API and taking responsibility for deployment. The preview lets teams test function calling, long context, and image input, but it cannot yet answer how the model will fit on their own hardware, how its routing is configured, or what inference costs look like there. ML4 is therefore notable not just for its parameter count, but because access to the service and access to self-hosting arrive on different timelines.

Sparse Computation Does Not Erase Model Size

The easiest way to misread ML4’s specifications is to see “1.05 trillion parameters” and “49 billion active parameters per token” as competing figures. Mistral describes it as a granular mixture-of-experts model: the full parameter set represents the model’s capacity, while only part of it is used for each token. Active parameters are therefore a clue to computation, not a measure of the storage needed for the whole model. The full model still has to reside in memory so that different experts can be selected as needed.

The figures are clearer when read together: about 1.05 trillion total parameters, 49 billion active per token in the model documentation, or 52 billion when embeddings and the output layer are included; a one-million-token context window; and a 1.6-billion-parameter vision encoder. Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European data centers. That GPU count describes the scale of the training infrastructure. It does not tell customers how many GPUs they will need to deploy the model, and the active-parameter count alone cannot determine local operating cost.

A Million-Token Window Is an Option, Not a Quality Guarantee

A one-million-token context window gives ML4 room to support long-document question answering, multi-stage agent tasks, and workflows that need to bring large amounts of material into a single request. The model accepts images natively, and Mistral’s demonstrations cover charts, documents, natural images, and visual grounding. The available information establishes these input capabilities, but does not disclose how the vision encoder connects to the rest of the network or how reliable the model is on each task.

The preview API supports function calling, structured outputs, document question answering, batching, and the Agents and Conversations endpoints. Published prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. The cache rate could change the bill for agent loops that repeatedly use the same long context, but teams will need to model that against their own request patterns. How much information fits in a context window and how effectively the model uses it are separate questions.

Strong Cybersecurity Scores Need Their Evaluation Context

Mistral’s most striking reported results are in cybersecurity: 93% on Cybench and 82% on CyberGym-E2E. The company also says several closed frontier models score near zero on CyberGym-E2E because they refuse the task. That comparison raises an important question: when a benchmark requires reproducing a vulnerability to establish that it is real, refusal behavior affects the score. A low score may not mean a model lacks the relevant capability, but a lower refusal rate does not automatically make a model safer either.

For coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0. Some of these evaluations were conducted privately before the harness became public, so external teams cannot yet reproduce them under the same procedure. A separate blind human evaluation by Surge AI gave ML4 Preview a score of 3.74 out of 5, placing it second among five models and behind Claude Opus 5 at 4.22. This adds evidence beyond benchmark scores, but it does not replace testing against a team’s own production tasks.

Deployment Decisions Come After the Weights

Mistral links ML4 to its own European data centers, the 3,800 GPUs used for training, and training data spanning more than 160 languages, including every official EU language. These details support its positioning around European infrastructure and multilingual use, but they should not be treated as proof that customers already have full control over deployment. During the preview, requests still go through Mistral’s API. Teams that need data residency, auditability, or offline operation will need clearer information about weights and regional service options.

When the weights are scheduled to arrive at the end of October, the important questions will go beyond whether they can be downloaded. The number of experts, top-k routing, layer layout, and post-training methods have not yet been disclosed, and they will help determine how the 49-billion active-parameter figure translates into inference and deployment choices. For now, the practical approach is to use the API to test task quality, context use, and cached-call costs, while treating self-hosting as unverified. A large total parameter count does not guarantee a stronger model, and fewer active parameters per token do not guarantee a cheaper or easier deployment.