Mistral Large 4: Impressive Beyond The US And China, With Agent Caveats
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: Impressive Beyond The US And China, With Agent Caveats on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research preview, scoring 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The result is a major improvement over Mistral’s previous models and makes it the highest-scoring model from outside the United States and China in the cited comparison, but it trails leading US and Chinese systems. Pricing, verbosity and reported hallucinations raise questions about using it for long-running agent tasks.

Mistral has released Large 4, a research-preview model that scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The score is a steep rise from the company’s previous flagship, but the model remains below leading US and Chinese systems, and its price and reported reliability issues may limit its fit for agent-based work.

Artificial Analysis’ index places Large 4 behind every current US and Chinese flagship listed in the source comparison. The leading listed score is 57.6, for Anthropic’s Claude Opus 5.5; Chinese models GLM-5.3 and Kimi K3 score 44.8 and 43.6. Large 4 is ahead of some older models, including GLM-5.2 at 33.7 and DeepSeek V4 Pro at 36.0. These rankings reflect the cited Index v4.3.2, rather than a general measure of every model use.

The release is a substantial step for Mistral by its own recent benchmark history. Mistral Large 3 scored 9 on the same index version, while Medium 3.5 scored 14. The source report describes the move to 38.4 as the biggest single-release gain by a European lab; that characterization is the report’s assessment, not a separate finding from the index.

Large 4 has 1 trillion total parameters, with 49 billion active, and accepts text and images as input while producing text. It has a 512,000-token context window. Mistral lists API pricing of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; the source says the company offered a 50% discount for the first two weeks. The model’s weights have not yet been released, and the licence has not been published.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, posting a substantial benchmark improvement while remaining behind leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A European Model’s Benchmark Gap

Large 4 matters because it shows a French AI company making a marked advance in a market led by US and Chinese labs. Its score gives buyers a clearer basis for comparing Mistral’s progress, while the gap to higher-ranked models makes clear that being the strongest model in a particular geographic category is not the same as matching the frontier overall.

The practical question is not only how high a model scores, but what it costs to complete real tasks and how reliably it handles a sequence of steps. The source report says Large 4 used 200 million output tokens on the index tasks, compared with a median of 81 million for comparable models. If that observation holds across buyers’ workloads, verbosity could add cost and latency even when the per-token price seems manageable. The benchmark figure alone does not establish that every customer will see the same usage pattern.

On the source’s cost-per-index-task figures, Large 4 costs $1.13 per task. GLM-5.3-Flash is listed at $0.25 and scores 41.8; DeepSeek V4.1 Flash is listed at $0.27 and scores 39.5. Those comparisons suggest that some lower-cost alternatives score higher on this particular index. Buyers would still need to test models on their own tasks, since benchmark rankings and task costs do not guarantee equivalent performance in production.

Amazon

AI language model API pricing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Index Frames Large 4

The Artificial Analysis Intelligence Index v4.3.2 combines tasks that include agentic knowledge work, real-world work tasks, software workflows and coding. The source report says these components make agentic performance a substantial part of the score. That gives the benchmark relevance to multi-step workflows, but it does not make the index a direct measurement of every agent deployment or business process.

Mistral’s release is currently a proprietary API preview. The company has said that model weights are expected at the end of October, but the source does not provide a year for that date. Mistral also says reinforcement learning is ongoing and that scores may change. Until weights and licence terms are available, customers cannot treat the preview as an already released open-weight model or make a final assessment of the terms for running it independently.

The geographic framing needs care. The source calls Large 4 the most intelligent model outside the United States and China, based on the cited comparison. That is a narrow ranking claim, not evidence that the model is competitive with the top systems globally. The same table places multiple US and Chinese models above it.

“Reinforcement learning is still running, so scores may move.”

— Mistral

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Limits and Reliability Questions

Large 4’s benchmark position may change: Mistral says reinforcement learning is ongoing, and the model is still in research preview. The source does not give an independent account of how the announced figures were produced beyond identifying the Artificial Analysis index, nor does it establish how the benchmark translates to specific customer workflows.

The source report’s claim that it observed confident hallucinations in hands-on use is an individual testing observation, not a result attributed to Artificial Analysis. No test details, sample size or reproducible evaluation are supplied. The report also cites a 15% hallucination rate for Gemini 4 Argon and higher rates for some Chinese models on AA-Omniscience, but these figures should be understood as results on that named evaluation, not universal error rates.

Commercial details remain incomplete. Mistral has not published the model’s licence, and the source does not clarify the year meant by the promised end-of-October weights release. It is also unclear whether the early 50% API discount remains available or how the reported token use will vary across customer workloads.

Amazon

AI image and text input models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Pricing and Further Tests

The next concrete milestone is Mistral’s stated plan to release Large 4’s weights by the end of October. Buyers will need the release date, licence and deployment terms before deciding whether they can use those weights outside Mistral’s API. The source does not specify a year for the timetable, so the exact deadline is not established here.

In the meantime, the model remains available as an API research preview. Mistral’s ongoing reinforcement learning may alter benchmark results, while independent evaluations and customer testing can help establish whether its performance, token use and reliability fit particular tasks. For agent deployments, those tests should examine complete workflows rather than relying on a single overall index score.

Amazon

AI model performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source material. That is up from 9 for Mistral Large 3 on the same index version.

Does Large 4 lead US and Chinese models?

No. It is described as the highest-scoring model outside the United States and China in the cited comparison, but several US and Chinese models score higher, including Claude Opus 5.5 at 57.6 and GLM-5.3 at 44.8.

Is Large 4 available as an open-weight model now?

No. The source describes it as a proprietary API research preview. Mistral has said weights are planned for the end of October, but the source does not specify the year, and the licence has not been published.

What are the main concerns for agent use?

The source report raises concerns about the model’s benchmark gap, high output-token use and an observer’s report of confident hallucinations. The hallucination observation is not an Artificial Analysis finding, and buyers should test the model on their own multi-step workflows.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI’s Strategy: Releasing Astra Gated After Crossing The Line

OpenAI publicly admits Astra model reaches ‘Critical’ cybersecurity risk level, plans gated release with safeguards following recent incident and internal assessments.

Qwen4 Architecture: A Pre-Release Open-Source Milestone

Alibaba’s Qwen team open-sourced the architecture of its upcoming model family, Qwen4, before its official launch, emphasizing efficiency and community collaboration.

Can SenseTime’s 8B Multimodal AI Model Transform Visual AI Applications?

SenseTime has open-sourced an 8-billion-parameter multimodal AI model claiming native 4K image output, raising questions about its capabilities and availability.

The Future Of AI: Top 10 Trends To Watch In 2026

Explore the top 10 AI trends shaping 2026, including advancements in automation, ethical AI, and new applications across industries.