🔍 Read the full analysis: Mistral Large 4’S Standing In The Race For The AI Frontier on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral launched Mistral Large 4 in public API preview on October 6, 2026. Artificial Analysis scores it at 38, below several leading U.S. and Chinese models; the model’s weights are scheduled for release later in October, and its current preview score does not establish how it will perform on specific workloads.
Mistral AI launched Mistral Large 4 in public API preview on October 6, but an Artificial Analysis Intelligence Index score of 38 places it behind several leading U.S. and Chinese models in the October 7 snapshot. The release is a step for European AI development, but the available evidence does not yet establish the preview as a top choice for demanding agentic work; model weights are scheduled for release later this month.
Mistral describes Large 4 as its largest model so far: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. For now, developers can access it through a preview API; despite the planned weights release, the weights are not publicly downloadable as of October 7.
Artificial Analysis gives Mistral Large 4 Preview an Intelligence Index score of 38. In the same dated snapshot, Claude Opus 5.5 at maximum effort with default fallback scored 58, Gemini 4 Argon at high effort scored 53, and GPT-6.1 Sol at maximum effort scored 52. Chinese models GLM-5.3 and Kimi K3 scored 45 and 44, while DeepSeek V4.1 Flash at maximum effort scored 39. These are index points, not percentages or direct predictions of success on a particular task.
The comparison is not a controlled contest under identical settings: the listed models use different reasoning configurations, and the scores may change. GPT-6 Luna at maximum effort also scored 38, while Cohere Command A+ scored 13. Developer locations in the comparison refer to the companies, not the locations where API requests are processed.
Frontier AI · October 7, 2026 snapshot
Mistral Large 4’s Standing in the Race for the AI Frontier
A new European model enters public API preview with a trillion total parameters. Its first benchmark snapshot puts it behind several leading U.S. and Chinese models, while weights and workload-specific evidence are still ahead.
Artificial Analysis score for Mistral Large 4 Preview
Weights were not publicly downloadable as of October 7
Mistral has scheduled a weights release; no exact date is stated
PREVIEW STATUS · DATED BENCHMARK SNAPSHOT
Mistral’s largest model so far
Mixture-of-experts architecture
Tokens reported by Artificial Analysis
Available through the preview API
01 / Benchmark snapshot
A score of 38 in a crowded field
Index points are not percentages or direct predictions of success on a particular task. The listed models use different reasoning configurations, so this is not a controlled contest under identical settings.
Bar lengths scaled to the highest score shown. Developer locations refer to the companies, not where API requests are processed.
The snapshot may change as models and evaluation settings evolve. Large 4’s score is close to DeepSeek V4.1 Flash; the source analysis reports much lower measured cost per task for DeepSeek, but the material does not support a full price comparison across models.
02 / What the gap signals
Capability claims still need workload evidence
The index can inform a model choice, but it does not measure every part of sustained execution or predict reliability on a specific coding, research, or business workflow.
Behind several leaders
In this snapshot, Large 4 is well below the highest-scoring U.S. models and below two listed Chinese models. A score of 38 does not prove it will fail a particular task.
Many steps, more ways to drift
Long workflows may require planning, tool use, interpretation, and follow-through. An early mistake can affect later actions; aggregate scores alone cannot establish sustained reliability.
One view, not a consensus
The author advises against choosing this preview for demanding, long-running work when stronger models are available. That is a judgment, not a controlled finding about every task.
03 / Release status
Preview today, weights later
Two separate milestones matter for developers: API access is available now, while the planned public weights release had not happened as of October 7, 2026.
Public API preview
Mistral introduced Large 4 in public API preview. The model accepts text and images.
Weights not yet downloadable
Calling it an already available open-weight model would be inaccurate at this point in the reporting.
Weights release scheduled
No exact release date, terms, or confirmation of a completed release are provided in the source material.
04 / What remains open
Test the work that matters to you
Mistral says it trained Large 4 on its own infrastructure in Europe and continues to improve it. That signals European AI capacity; it does not settle how the model performs across use cases.
Will the release match?
The source does not specify the final configuration or whether public weights will match the API preview in capability.
How does it hold up over time?
Independent, workload-specific tests are still needed for long agentic tasks, coding, professional work, constraint following, and tool use.
Capacity is not accuracy
A context window of roughly 512,000 tokens describes how much material can fit in a request, not whether the model will use it correctly.
“The company says it trained the model on its own infrastructure in Europe and continues to improve it.”Mistral · Company statement as reported
05 / Key questions
What developers should know
Is Mistral Large 4 publicly available?
It is available through a public preview API. Its weights were not publicly downloadable as of October 7, 2026.
How does it score against frontier models?
Artificial Analysis gave the preview 38 points in its October 7 snapshot. Several listed U.S. and Chinese models scored higher, with differing reasoning settings.
Does a score of 38 mean it will fail agentic tasks?
No. The index is an aggregate benchmark, not a direct prediction for a specific workflow.
When are the weights due?
Mistral scheduled the release for later in October 2026. The source gives no exact date and does not confirm the release has occurred.
Are hallucination reports a formal comparison?
No. The author describes personal experience and explicitly says it was not a controlled comparative study.
What evidence should come next?
Test constraint following, result verification, consistent tool use, and supervision needs on your own tasks as new evaluations become available.
What the Benchmark Gap Signals
The score matters because Mistral is seeking a place among frontier AI providers, and developers choosing a model for complex work need evidence about capability as well as access. In this snapshot, Large 4 sits well below the highest-scoring U.S. models and below two listed Chinese models. It is close to DeepSeek V4.1 Flash on index score, though the source analysis says DeepSeek has a much lower measured cost per task.
For agentic workflows, a model may need to plan, use tools, interpret results, and carry decisions across several steps. Mistakes early in that chain can affect later actions, so strong aggregate performance can inform a choice. Still, the index is not a direct test of reliability on every coding, research, or business workflow. A score of 38 does not prove that Large 4 will fail a given task, just as a larger context window does not prove that it can reason accurately over everything supplied.
The article’s assessment that it would not select the current preview for demanding, long-running work is the author’s judgment, not a measured industry consensus. Mistral’s claims about agentic coding and specialist professional work need testing on those workloads. The practical question for users is whether it performs reliably enough, at an acceptable cost and supervision level, for their own tasks.
As an affiliate, we earn on qualifying purchases.
Preview Today, Weights Later
The release has two separate milestones. On October 6, Mistral introduced the preview API; it also said the weights are scheduled to be released later in October. Those weights had not been made publicly downloadable at the time of the October 7 reporting. Describing the current product as an already available open-weight model would therefore be inaccurate.
Mistral says it trained Large 4 on its own infrastructure in Europe and is continuing to improve it. That is relevant to European AI capacity, but it does not answer how the model compares on every use case. Artificial Analysis reports a context capacity of roughly 512,000 tokens; capacity describes how much material can fit into a request, not whether the model will use it correctly.
The source article also reports that its author encountered hallucinations while using the preview. That is personal experience, not a controlled comparative study. It does not show how often hallucinations occur across users or establish that competing models do not make similar errors.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
AI model training and testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Preview Results May Change
The preview may change as Mistral continues development, and the planned weights release could enable broader evaluation. The available source does not specify the exact release date, final model configuration, or whether the public weights will match the API preview in capability.
It is also unclear how Large 4 performs across independent, workload-specific tests of long agentic tasks, coding, and professional applications. The index offers an aggregate comparison, not a direct measure of sustained execution or hallucination rates. The author’s reported hallucinations are anecdotal; no controlled rate or comparable test is supplied. The source material’s discussion of cost is incomplete, so it does not support a full price comparison across models.
machine learning model evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights and Workload Tests Ahead
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October 2026. Whether that schedule is met, and what terms or technical details accompany the release, were not confirmed in the source material. New benchmark results may also change the current comparison as model versions and evaluation settings develop.
For developers, the next useful evidence will come from testing the model on their own tasks: whether it follows constraints, verifies results, uses tools consistently, and needs more supervision than alternatives. Until those results and the weights release are available, the current assessment applies to the API preview and the dated benchmark snapshot, not necessarily to future versions.
AI image and text processing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Is Mistral Large 4 publicly available?
It is available through a public preview API, according to the source. Its weights were not publicly downloadable as of October 7, 2026.
How does Mistral Large 4 score against frontier models?
Artificial Analysis gave the preview an Intelligence Index score of 38 in its October 7 snapshot. Several listed U.S. and Chinese models scored higher, though the comparison uses different reasoning settings.
Does a score of 38 mean the model will fail agentic tasks?
No. The index is an aggregate benchmark, not a direct prediction of performance on a specific workflow. The source author advises against choosing the preview for demanding long tasks, but that is an assessment rather than a controlled finding about every task.
When are Mistral Large 4’s weights due?
Mistral scheduled the weights release for later in October 2026. The source does not give an exact date or confirm that the release has occurred.
Are reports of hallucinations based on a formal comparison?
No. The source author describes hallucinations encountered in personal use and explicitly presents that experience as not a controlled comparative study.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
