Mistral Large 4: A Strong Option Outside The US And China, But Not For Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: A Strong Option Outside The US And China, But Not For Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie-night favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral released Large 4 as a research preview, scoring 38.4 on Artificial Analysis Intelligence Index v4.3.2. The result is a sharp improvement over Mistral’s prior models and gives buyers a notable option from France, but the supplied analysis finds it trails leading US and Chinese models and raises cost, verbosity and reliability concerns for agentic work.

Mistral has released Large 4, a research-preview model that scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, marking a major improvement over the company’s earlier models while remaining behind leading US and Chinese competitors. The result makes Mistral a stronger option for organizations seeking a model from outside those two countries, but the source analysis cautions against using it for complex agent workflows without further testing.

Artificial Analysis lists Large 4 at 38.4 points. In the same index version, Mistral Large 3 scored 9 and Medium 3.5 scored 14, making Large 4’s gain a substantial one-release jump. The source analysis says the score is the strongest among models from outside the United States and China, while also noting that current major US and Chinese flagships rank higher. The leading listed model, Anthropic’s Claude Opus 5.5, scores 57.6.

Mistral describes Large 4 as a 1-trillion-parameter model with 49 billion active parameters. It accepts text and images, produces text, and supports a 512,000-token context window. The preview is available through Mistral’s API. The source says the company has promised to release model weights by the end of October; until then, the model is proprietary and its licence has not been published.

The listed standard API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million. Mistral is offering a 50% discount for the first two weeks, according to the source. Mistral also says reinforcement learning is ongoing, meaning benchmark results may change. Artificial Analysis estimates a cost of $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; both models score above Large 4 in the cited index.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral has released Large 4 as a research preview, with independent benchmark data showing a major improvement over its previous models but a continued gap from leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Trade-Off for Agent Work

Large 4 matters because it gives buyers a substantial model from a French developer at a time when the strongest systems in the cited comparison come from US and Chinese labs. For companies with geographic, procurement or strategic reasons to diversify providers, that can make Mistral’s progress relevant even if it does not match frontier leaders on the benchmark.

But the source analysis argues that the score is especially pertinent to agentic tasks, since the index includes benchmarks for knowledge work, real-world tasks, software workflows and coding. It reports that Large 4 used 200 million output tokens to complete the index, against a median of 81 million for comparable models. If that measure reflects a model’s tendency to produce longer answers, the added output can raise cost and latency across repeated agent steps. The source also reports firsthand observations of confident false statements in testing; that is an attributed assessment, not a benchmark result.

These concerns do not establish that Large 4 is unsuitable for every agent or business use. They do suggest buyers should test it on their own tasks, including error recovery, factual accuracy, token use and total cost, rather than choosing it on nationality or a single headline score.

Amazon

AI language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Rise From Mistral’s Earlier Scores

The comparison in the source uses Artificial Analysis Intelligence Index v4.3.2, described as the current version, so the listed model scores are presented on a like-for-like basis. Mistral’s earlier Large 3 and Medium 3.5 scores of 9 and 14 put the 38.4 result in perspective: the company has made a large benchmark gain, even though that gain has not closed the distance to the highest-scoring models.

The source says Large 4 ranks below several Chinese models, including GLM-5.3 variants, Kimi K3 and DeepSeek V4.1 Flash, while beating some earlier Chinese releases. It also frames Large 4 as the leading option among models outside the US and China in the cited comparison. That distinction is narrower than being a global frontier leader: the underlying table shows multiple US and Chinese models with higher scores.

Large 4’s status is also provisional in practical terms. It is currently a preview offered through Mistral’s API, rather than a publicly available set of weights. The promised weight release and Mistral’s ongoing reinforcement learning could affect how customers can deploy it and how later evaluations compare.

“Reinforcement learning is still running, so scores may move.”

— Mistral

Amazon

large language model for developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Status Leaves Open Questions

Large 4 is still a research preview, and the source says Mistral is continuing reinforcement learning. The final benchmark position, performance after further updates and behavior across real customer workloads are not yet established. Artificial Analysis’ score is one evaluation and cannot settle how the model will perform in every deployment.

The source does not provide a complete independent reliability assessment or a detailed methodology for its own hands-on hallucination observations. It also does not establish whether the reported verbosity will occur consistently across prompts and tasks. The promised weight release is dated only as the end of October, with no exact day or published licence terms included in the material.

Amazon

text and image AI processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Further Testing Ahead

The next reported milestone is Mistral’s planned release of Large 4’s weights by the end of October. The licence and exact release date remain unspecified in the source material. Mistral’s ongoing training work may also lead to updated scores or model behavior.

For prospective customers, the immediate next step is to compare the preview against alternatives on representative tasks and calculate full workflow costs, including output tokens and retries. Further independent evaluations could clarify whether Large 4’s benchmark position changes and how reliably it handles multi-step work. Until then, its strongest confirmed case is a significant improvement from Mistral’s previous scores and a model option from France—not demonstrated parity with the leading US and Chinese systems.

Amazon

AI model cost management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is a text-and-image-input, text-output model from French company Mistral. The source describes it as having 1 trillion total parameters, 49 billion active parameters and a 512,000-token context window.

How did it score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. That is a substantial improvement over the cited scores for Mistral Large 3 and Medium 3.5, but below the leading US and Chinese models in the source’s comparison.

Can developers download its weights now?

Not according to the source material. Large 4 is currently a proprietary Research Public Preview available through Mistral’s API. Mistral has promised weights by the end of October, but the licence is unpublished in the supplied material.

Why does the source caution against using it for agents?

The source points to its benchmark position, reported high output-token use and the author’s observations of confident false statements. Those are reasons to test cost and reliability on specific workflows; they do not establish that the model will fail in every agent deployment.

What does Mistral Large 4 cost?

The listed standard API prices are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source reports a 50% discount for the first two weeks.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

European AI firms Mistral, Aleph Alpha, and Black Forest Labs are positioning for the EU AI Act enforcement, emphasizing compliance and sovereignty over frontier capabilities.

Software engineering. The canonical case.

A comprehensive analysis of recent data shows junior developers face significant displacement, while senior engineers benefit from augmentation, revealing a bifurcated AI impact in software engineering.

Webinar follow-up personalization tool for B2B consultants

A new webinar follow-up personalization tool for solo B2B consultants is being tested to improve engagement and lead conversion after webinars.

The 27% Problem: Why Google Wrote a $750M Check to Catch Anthropic

Google commits $750 million to expand enterprise AI, aiming to reclaim market share lost to Anthropic, amid industry shifts towards agent governance platforms.