🔍 Read the full analysis: Why Mistral Large 4 Has More Ground To Cover In AI on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral released Large 4 as a public preview API on October 6, 2026, and Artificial Analysis scored it 38 on its Intelligence Index. That places it below leading U.S. models and some Chinese competitors, though the score is not a direct measure of performance on every task. The model’s weights are expected later in October, and its current suitability for long agentic workflows remains an open question.
Mistral AI launched Mistral Large 4 as a public preview API on October 6, but an October 7 benchmark snapshot places it behind leading U.S. models and several Chinese rivals. The result gives developers a reason to test the model against their own workloads before relying on it for long, multi-step agentic tasks; it does not establish that the model will fail at any particular task.
Mistral describes Large 4 as its largest model to date: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it. At the time of the October 7 report, the model was available through a preview API, while the weights had not yet been released for download.
Artificial Analysis gave Mistral Large 4 Preview an Intelligence Index score of 38. In the same dated comparison, Anthropic’s Claude Opus 5.5 scored 58, Google’s Gemini 4 Argon scored 53, and OpenAI’s GPT-6.1 Sol scored 52. Chinese models GLM-5.3 and Kimi K3 scored 45 and 44, respectively; DeepSeek V4.1 Flash scored 39. These are benchmark index points, not percentages or direct predictions of success on a specific job.
The benchmark settings were not identical: the comparison includes models evaluated at different reasoning levels, and the listed scores are a snapshot that can change. Artificial Analysis also reports roughly 512,000 tokens of context capacity for Large 4. That describes the amount of input the model can handle, not whether it can reliably reason over all of it.
Benchmark Gaps Shape Developer Choices
For developers choosing a model for autonomous work, the issue is not just whether Large 4 can produce a good answer. An agent may need to plan, use tools, interpret results and carry decisions across steps. Errors or unsupported assumptions early in that sequence can affect later actions, while a fluent final response may not reveal where the process went wrong.
The benchmark gap is a reason for caution, not a verdict on every use case. Artificial Analysis’s aggregate score does not directly test a particular company’s coding, research or operations workflow. Mistral advertises strengths in agentic coding and specialized professional tasks, but those claims need to be checked against the work developers intend to delegate. The source report’s author said the current evidence would lead them to start with higher-scoring alternatives for demanding, long-running tasks.
The result also matters to Mistral’s position as a European AI provider. Training the model on European infrastructure is relevant to regional AI capacity, but it does not by itself demonstrate that Large 4 matches the strongest products available elsewhere. The score puts the discussion on a more specific footing: Mistral is ahead of Cohere’s Command A+ in this comparison, but remains below the cited leading U.S. models and stronger-scoring Chinese models.
As an affiliate, we earn on qualifying purchases.
A Preview, Not a Weight Release
The October 6 announcement marked the public preview stage, rather than a completed public release of model weights. That distinction affects what users can evaluate and how they can deploy the system: at the time covered, access was through Mistral’s API, and the weights were scheduled for release later in October. The schedule is not confirmation that the release had already happened.
The figures cited here come from Artificial Analysis’s release analysis and model profile, as reported on October 7. Its table lists developer locations, not the place where an individual API request is processed. It also compares models at the settings named in the table, rather than under a single, matched compute budget. Scores therefore provide a dated comparative signal, not a universal ranking across all tasks or operating conditions.
The report’s author also described encountering hallucinations while using the preview and said this reduced their confidence in assigning it longer tasks. That is an individual account, not a controlled comparison of hallucination rates. It should be treated as an experience that may guide further testing, rather than evidence that Large 4 hallucinates more often than every competitor.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, author of the October 7 report
As an affiliate, we earn on qualifying purchases.
Task Reliability Still Needs Testing
The available information does not establish how Large 4 performs across specific coding, research or professional workflows, or how reliably it completes long tasks without supervision. The Intelligence Index is an aggregate benchmark, and its results do not prove that the model will fail on a particular assignment. The report also provides no controlled hallucination-rate comparison, so the author’s experience cannot establish relative error rates.
It is also not clear from the material available whether the scheduled weight release will occur as planned, what changes may come before or after it, or how the model’s performance may shift as Mistral continues development. The benchmark snapshot may change, and the listed reasoning settings do not represent identical evaluation budgets.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Workload Tests
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October 2026. Developers and evaluators can then assess the release available to them and compare it with the preview, while watching for updated benchmark results. Until then, the clearest basis for a decision is to test the API on representative tasks and measure accuracy, unsupported claims, tool use and the amount of human review required.
For teams considering long-running agents, the key question is whether Large 4 performs reliably enough on their own work to justify delegation. The present benchmark and personal-use account raise doubts for the report’s author, but they do not settle that question for all users or workloads.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral release on October 6?
Mistral announced Mistral Large 4 as a public preview API. The model accepts text and images and has one trillion total parameters, with 49 billion active parameters, according to the announcement as described in the source report.
How did Mistral Large 4 score?
Artificial Analysis gave the preview an Intelligence Index score of 38 in a snapshot dated October 7, 2026. The comparison includes models tested at different reasoning settings, so the score is not a direct prediction of performance on every task.
Does the score prove Large 4 cannot handle agentic work?
No. The index is an aggregate benchmark, not a direct test of every multi-step workflow. The source report’s author recommends stronger-scoring alternatives for demanding, long tasks, but developers need workload-specific testing to judge the model for their own use.
Are Mistral Large 4’s weights available?
Not at the time of the October 7 report. Mistral had made the preview available through an API and scheduled the weights for release later in October; the source does not confirm that the release has occurred.
What remains to be established?
Independent, workload-specific evidence is still needed on long-task reliability, tool use and the frequency of unsupported output. The report’s account of hallucinations is personal experience, not a controlled comparison with competing models.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
