Evidence at a glance

2026 10 6Evidence
parameters1T, parameters49BEvidence
parameters1.05T, parameters52BEvidence
1.6B parametersEvidence
use 3,800 Grace Blackwell GPUEvidence
API none highEvidence

A Preview Now, Open Deployment Later

Mistral AI released a public preview of Mistral Large 4 on October 6, 2026. The model handles text and image inputs and is positioned for reasoning and agent tasks. It is currently available through the Mistral API, while the company says it plans to release the weights at the end of October. In other words, access to a hosted service today does not mean the model can be downloaded and deployed today.

That distinction should shape how technical leaders read the launch. Large 4 is not only a capability showcase; it also signals Mistral’s attempt to regain ground through its own compute infrastructure and a sparse-model strategy. But an API preview alone cannot establish the costs of serving it, its suitability for private deployment, or whether results can be reproduced. The weights, license, and further training details remain pending.

A Huge Parameter Count Does Not Mean Every Request Uses It

Mistral describes Large 4 as a mixture-of-experts model: it has a large total parameter count, but only a portion is activated for a given computation. That is why “one trillion parameters” should not be treated as a direct proxy for either per-request cost or capability. The press release gives one trillion total parameters and 49 billion active parameters, while the model card lists 1.05 trillion and 52 billion. Mistral has not explained the difference, so there is no basis for forcing the two figures into a single accounting convention.

The training scale calls for similar restraint. Mistral says the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European data center. The model card also lists a one-million-token context window and a 1.6-billion-parameter vision encoder. The company brings text, image understanding, instruction following, reasoning, and agent capabilities together in one model, but has not disclosed key architectural details such as the number of experts or how routing works. The GPU count indicates the scale of the investment, not the training cost, efficiency, or serving latency.

What One Pelican Image Can—and Cannot—Tell Us

Simon Willison compared the API’s two reasoning settings, none and high, on a pelican-drawing task. In his observation, the high setting produced the better-looking image while using 2,717 output tokens, fewer than the 3,275 used by none. It is a useful small example: a higher reasoning setting does not necessarily produce more output tokens, and teams should not infer workload from token counts alone.

But one image is not a systematic evaluation, nor does it show that high is more effective across tasks. A broader point of reference is Artificial Analysis, where Large 4 scores 38, just behind DeepSeek 4.1 Flash, a 552-billion-parameter model. The same source gives the previous Mistral Large 3 a score of 9. That jump suggests a substantial improvement over its predecessor, but one aggregate score cannot represent a team’s workload or replace measurements of latency, reliability, and call costs.

Useful Benchmark Numbers, Not Yet Reproducible Conclusions

Mistral reports coding results of 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and 49.8% on the Coding Agent Index. On the Dense 200 visual-localization test, the company reports 42% for Large 4 and 41% for GPT-6 Astra. These figures can help teams choose evaluation tasks, but the announcement does not provide complete configurations for independently reproducing all the results. They should therefore be treated as figures disclosed by the publisher, not as performance guarantees a team has already verified.

Post-training is also still in motion. Mistral describes a reinforcement-learning pipeline that combines task environments and uses reward models, unit tests, LLM judges, and static checks for verification. Parallel actors generate training trajectories while model training proceeds asynchronously. The company also says reinforcement learning for the preview is ongoing. The reported performance is therefore a snapshot, and later weights or API behavior may change; the preview should not be mistaken for a final specification.

Put It on the Shortlist Before You Plan a Migration

For teams choosing a model, Large 4 belongs on the shortlist for now, not yet at the center of a migration plan. Use the API to compare it on your own multimodal, coding, long-context, and agent tasks, recording quality, latency, and token use under both reasoning settings. The API offers only none and high, so teams cannot yet tune the quality-resource tradeoff with more granular reasoning controls.

If a decision depends on self-hosting, license review, or reproducibility, wait for the weights and fuller configuration details. Then check the final license, clarify the difference between the 49-billion and 52-billion active-parameter figures, and measure throughput and cost on your own hardware. Large 4 shows how quickly Mistral is catching up, but the practical test remains straightforward: establish whether it outperforms your current option on your workload before treating parameter scale or publisher-reported benchmarks as a conclusion.