🔍 Read the full analysis: Mistral Large 4 Still Trails The AI Frontier on ThorstenMeyerAI.com
Get movie-night favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral released Large 4 in public API preview on October 6, but an Artificial Analysis benchmark snapshot puts it behind leading US models and two higher-scoring Chinese models. The model has one trillion total parameters and a reported context window of about 512,000 tokens; its weights are scheduled for release later in October. Benchmark scores and one reviewer’s experience do not establish how it will perform across all workloads.
Mistral AI released Mistral Large 4 in public API preview on October 6, but a benchmark snapshot published the next day places it behind leading US models and two Chinese competitors. The result adds evidence to the debate over whether the French company’s latest model can compete at the AI frontier, while leaving its performance on specific developer workloads unresolved—a question explored in discussions of whether Mistral is truly leading AI innovation.
Artificial Analysis gave Mistral Large 4 Preview a score of 38 on its Intelligence Index, according to its release analysis. The score matched OpenAI’s GPT-6 Luna at maximum reasoning effort and was one point below DeepSeek V4.1 Flash at maximum effort. In the same dated comparison, Anthropic’s Claude Opus 5.5 scored 58, Google’s Gemini 4 Argon scored 53, and OpenAI’s GPT-6.1 Sol scored 52. China’s GLM-5.3 and Kimi K3 scored 45 and 44, respectively.
The scores are index points, not percentages, and the models were evaluated at different reasoning settings, so the comparison does not represent identical compute budgets. Artificial Analysis’s index is an aggregate benchmark measure, not a direct test of how reliably a model completes a particular coding, research or agentic task. The scores are also a snapshot; they can change as models and evaluations are updated.
Mistral describes Large 4 as its largest model so far, with a mixture-of-experts design, one trillion total parameters and 49 billion active parameters. It accepts text and images. Mistral says it trained the model on its own European infrastructure and is continuing to improve it. The model is currently available through a preview API; the company has scheduled a release of its weights for later in October, meaning they were not publicly downloadable at the time of the report.
AI FRONTIER · OCTOBER 7, 2026
Mistral Large 4 Still Trails The AI Frontier
A public API preview brings a trillion-parameter model into view. In one dated benchmark snapshot, its score trails several leading US and Chinese models—while leaving real-world workload performance an open question.
The gap in this comparison
Artificial Analysis index points, as reported on October 7. Scores are not percentages; reasoning settings and compute budgets differ.
| Rank | Model | Developer region | Score |
|---|---|---|---|
| 01 | Claude Opus 5.5 | United States | 58 |
| 02 | Gemini 4 Argon | United States | 53 |
| 03 | GPT-6.1 Sol | United States | 52 |
| 04 | GLM-5.3 | China | 45 |
| 05 | Kimi K3 | China | 44 |
| 06 | DeepSeek V4.1 Flash · max effort | China | 39 |
| 07 | Mistral Large 4 Preview | France | 38 |
| 07 | GPT-6 Luna · max effort | United States | 38 |
| — | Cohere Command A+ | Canada | 13 |
How to read it: this is an aggregate benchmark snapshot, not a direct measure of coding reliability, research quality, hallucination rate, or performance on every agentic task. Developer locations identify companies, not request-processing locations.
Where 38 sits
A visual comparison of selected scores from the same snapshot. The chart shows index points, not percentages.
One reviewer’s experience and this benchmark do not establish how the model will perform across all workloads.
Scale, access, and claims
Mistral describes Large 4 as its largest model so far, built with a mixture-of-experts design and trained on its own European infrastructure.
1 trillion total
Mistral reports 49 billion active parameters in a mixture-of-experts model. Model size alone does not establish task quality.
Preview first
Public API access began October 6. Weights were scheduled for later in October and were not downloadable at the reporting date.
About 512,000 tokens
Context capacity describes how much material a request can contain; it does not show accurate reasoning across all of it.
Benchmarks narrow a shortlist
Agentic systems plan, call tools, interpret results, and carry decisions forward. A weak early assumption can shape every later step.
Regional infrastructure is a separate signal
Mistral says it trained Large 4 on its own infrastructure in Europe and continues to improve it. That matters to European AI capacity, but does not settle its standing on benchmarks or production workloads.
Test the work you actually need
Compare models with the same prompts, tools, supervision, and success criteria. Headline scores can inform an initial shortlist; workload-specific evidence should guide deployment.
The preview cannot settle everything
The available material leaves important questions unanswered. Keep observations about the current preview separate from expectations about later releases.
“The model was trained on the company’s own infrastructure in Europe and continues to be improved.”
MistralMistral’s claims about agentic coding and specialized professional tasks need evaluation on those workloads. Preview performance and index scores may change as the model and evaluations are updated.
What developers should take away
A dated comparison is evidence about one snapshot—not a final verdict on future versions or every use case.
Is Large 4 publicly available?
It was accessible through a public preview API after the October 6 announcement. Its weights were scheduled for later in October and were not publicly downloadable on October 7.
How did it score?
Artificial Analysis gave the preview 38 points. That matched GPT-6 Luna at maximum reasoning effort and was one point below DeepSeek V4.1 Flash at maximum effort.
Does 38 prove it is unreliable?
No. The Intelligence Index is an aggregate benchmark, not a direct measure of performance or reliability on every task.
What is the next useful evidence?
Independent workload evaluations and tests using consistent prompts, tools, supervision, and success criteria—alongside details of the weights release.
What the Scores Mean for Developers
For developers choosing a model for long, multi-step agentic work, benchmark position can inform an initial shortlist, but it cannot settle the choice by itself. Agentic systems plan, call tools, interpret results and carry decisions forward. A mistaken assumption early in that sequence can shape later steps, so sustained accuracy and verification matter alongside a model’s capacity or headline capabilities.
The source review’s author, Thorsten Meyer, said he would not select the current preview for demanding agentic tasks when higher-scoring alternatives are available. That is his assessment of the preview, based on benchmarks and personal use, rather than a controlled comparison of every model on the same workloads. The Intelligence Index score of 38 does not prove that Large 4 will fail a particular task; it does give developers a reason to test it against alternatives before relying on it for autonomous work.
Mistral’s launch also matters beyond model selection. The company says Large 4 was trained on its own infrastructure in Europe, a development relevant to European AI capacity. That regional significance is distinct from whether the preview matches the strongest available models on aggregate benchmarks or in production use.
As an affiliate, we earn on qualifying purchases.
The Preview Before the Weights
The October 6 announcement introduced a preview API, not a finished public release of model weights. Mistral’s planned later-October weights release could give developers another way to evaluate or use the model, but it had not happened by the October 7 reporting date. Any assessment of the current preview should be kept separate from expectations about future versions or the eventual release.
The benchmark comparison also includes a useful limit on broad conclusions: Canada’s Cohere Command A+ scored 13 on the same index, below Mistral Large 4. Cohere emphasizes enterprise applications, private deployment and grounded workflows, according to the source material. Its lower score means the data does not support saying every major competitor is ahead. The narrower finding is that several leading US models and two stronger-scoring Chinese models outranked Mistral in this snapshot. The locations in the comparison refer to developers, not where API requests are processed.
Artificial Analysis reports a context capacity of about 512,000 tokens for Large 4. That figure describes how much material a request can contain; it does not show that the model can reason accurately over all of it. Similarly, Mistral’s claims about agentic coding and specialized professional tasks require evaluations on those workloads rather than inference from model size or context length alone.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
What the Preview Cannot Establish
The available comparison does not show how Large 4 performs on a shared set of long-running coding, research or tool-use tasks, nor does it establish a controlled hallucination rate. Meyer reported encountering hallucinations in his own use and said this reduced his confidence, but he explicitly described that as personal experience, not a controlled study. It cannot establish that competing models hallucinate less across users or workloads.
The supplied source material ends before completing its discussion of model costs, so it does not support a cost comparison. It also does not provide a confirmed date for the weights release beyond saying it is scheduled for later in October. Mistral’s preview may improve, and benchmark scores may change; the extent and timing of those changes remain unclear.
As an affiliate, we earn on qualifying purchases.
Weights Release and Workload Tests
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October. Developers will then be able to assess the release in the form Mistral makes available, although the source does not specify the release date or terms. Mistral says it is continuing to improve the model.
For now, the most useful next step for prospective users is to test the preview on their own tasks and compare it with alternatives using the same prompts, tools, supervision and success criteria. Further benchmark results and independent workload evaluations could clarify whether the model’s advertised coding and professional-task strengths compensate for its lower aggregate score. Until those results arrive, the October 7 index should be treated as a dated comparison, not a final verdict on future versions.
As an affiliate, we earn on qualifying purchases.
Key Questions
Is Mistral Large 4 publicly available?
It was available through a public preview API following its October 6 announcement. Its weights were scheduled for release later in October and were not publicly downloadable at the time of the report.
How did Mistral Large 4 score against leading models?
Artificial Analysis gave the preview 38 on its Intelligence Index. That matched GPT-6 Luna at maximum reasoning effort, sat one point below DeepSeek V4.1 Flash at maximum effort, and trailed the cited scores for Claude Opus 5.5, Gemini 4 Argon, GPT-6.1 Sol, GLM-5.3 and Kimi K3.
Does the score prove Mistral Large 4 is unreliable?
No. The index is an aggregate benchmark, not a direct measure of performance on every task or a controlled reliability test. Developers need workload-specific evaluations to judge whether it suits their use case.
What is known about the model’s size and context?
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. Artificial Analysis reports a context capacity of about 512,000 tokens; that capacity does not guarantee accurate reasoning across a long input.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
