AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4’S Standing In The Race For The AI Frontier on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral launched Mistral Large 4 in public API preview on October 6, 2026. Artificial Analysis scores it at 38, below several leading U.S. and Chinese models; the model’s weights are scheduled for release later in October, and its current preview score does not establish how it will perform on specific workloads.

Mistral AI launched Mistral Large 4 in public API preview on October 6, but an Artificial Analysis Intelligence Index score of 38 places it behind several leading U.S. and Chinese models in the October 7 snapshot. The release is a step for European AI development, but the available evidence does not yet establish the preview as a top choice for demanding agentic work; model weights are scheduled for release later this month.

Mistral describes Large 4 as its largest model so far: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. For now, developers can access it through a preview API; despite the planned weights release, the weights are not publicly downloadable as of October 7.

Artificial Analysis gives Mistral Large 4 Preview an Intelligence Index score of 38. In the same dated snapshot, Claude Opus 5.5 at maximum effort with default fallback scored 58, Gemini 4 Argon at high effort scored 53, and GPT-6.1 Sol at maximum effort scored 52. Chinese models GLM-5.3 and Kimi K3 scored 45 and 44, while DeepSeek V4.1 Flash at maximum effort scored 39. These are index points, not percentages or direct predictions of success on a particular task.

The comparison is not a controlled contest under identical settings: the listed models use different reasoning configurations, and the scores may change. GPT-6 Luna at maximum effort also scored 38, while Cohere Command A+ scored 13. Developer locations in the comparison refer to the companies, not the locations where API requests are processed.

At a glance
analysisWhen: Announced October 6, 2026; status as of…
The developmentMistral has released Mistral Large 4 in public API preview, with an independent benchmark snapshot placing it behind several leading models.
Mistral Large 4’s Standing in the Race for the AI Frontier

Frontier AI · October 7, 2026 snapshot

Mistral Large 4’s Standing in the Race for the AI Frontier

A new European model enters public API preview with a trillion total parameters. Its first benchmark snapshot puts it behind several leading U.S. and Chinese models, while weights and workload-specific evidence are still ahead.

Intelligence Index 38 points

Artificial Analysis score for Mistral Large 4 Preview

Availability today API preview

Weights were not publicly downloadable as of October 7

Next milestone Later in October

Mistral has scheduled a weights release; no exact date is stated

PREVIEW STATUS · DATED BENCHMARK SNAPSHOT

Total parameters1 trillion

Mistral’s largest model so far

Active parameters49 billion

Mixture-of-experts architecture

Context capacity~512K

Tokens reported by Artificial Analysis

Input typesText + images

Available through the preview API

01 / Benchmark snapshot

A score of 38 in a crowded field

Index points are not percentages or direct predictions of success on a particular task. The listed models use different reasoning configurations, so this is not a controlled contest under identical settings.

Model · developer locationIndex
Claude Opus 5.5 · maximum effort, default fallbackU.S.
58
Gemini 4 Argon · high effortU.S.
53
GPT-6.1 Sol · maximum effortU.S.
52
GLM-5.3China
45
Kimi K3China
44
DeepSeek V4.1 Flash · maximum effortChina
39
Mistral Large 4 PreviewFrance
38
GPT-6 Luna · maximum effortU.S.
38
Cohere Command A+Canada
13

Bar lengths scaled to the highest score shown. Developer locations refer to the companies, not where API requests are processed.

The snapshot may change as models and evaluation settings evolve. Large 4’s score is close to DeepSeek V4.1 Flash; the source analysis reports much lower measured cost per task for DeepSeek, but the material does not support a full price comparison across models.

02 / What the gap signals

Capability claims still need workload evidence

The index can inform a model choice, but it does not measure every part of sustained execution or predict reliability on a specific coding, research, or business workflow.

Aggregate result

Behind several leaders

In this snapshot, Large 4 is well below the highest-scoring U.S. models and below two listed Chinese models. A score of 38 does not prove it will fail a particular task.

Agentic work

Many steps, more ways to drift

Long workflows may require planning, tool use, interpretation, and follow-through. An early mistake can affect later actions; aggregate scores alone cannot establish sustained reliability.

Author assessment

One view, not a consensus

The author advises against choosing this preview for demanding, long-running work when stronger models are available. That is a judgment, not a controlled finding about every task.

03 / Release status

Preview today, weights later

Two separate milestones matter for developers: API access is available now, while the planned public weights release had not happened as of October 7, 2026.

01October 6

Public API preview

Mistral introduced Large 4 in public API preview. The model accepts text and images.

02As of October 7

Weights not yet downloadable

Calling it an already available open-weight model would be inaccurate at this point in the reporting.

03Planned for later October

Weights release scheduled

No exact release date, terms, or confirmation of a completed release are provided in the source material.

04 / What remains open

Test the work that matters to you

Mistral says it trained Large 4 on its own infrastructure in Europe and continues to improve it. That signals European AI capacity; it does not settle how the model performs across use cases.

Version

Will the release match?

The source does not specify the final configuration or whether public weights will match the API preview in capability.

Evaluation

How does it hold up over time?

Independent, workload-specific tests are still needed for long agentic tasks, coding, professional work, constraint following, and tool use.

Interpretation

Capacity is not accuracy

A context window of roughly 512,000 tokens describes how much material can fit in a request, not whether the model will use it correctly.

“The company says it trained the model on its own infrastructure in Europe and continues to improve it.”
Mistral · Company statement as reported

05 / Key questions

What developers should know

Is Mistral Large 4 publicly available?

It is available through a public preview API. Its weights were not publicly downloadable as of October 7, 2026.

How does it score against frontier models?

Artificial Analysis gave the preview 38 points in its October 7 snapshot. Several listed U.S. and Chinese models scored higher, with differing reasoning settings.

Does a score of 38 mean it will fail agentic tasks?

No. The index is an aggregate benchmark, not a direct prediction for a specific workflow.

When are the weights due?

Mistral scheduled the release for later in October 2026. The source gives no exact date and does not confirm the release has occurred.

Are hallucination reports a formal comparison?

No. The author describes personal experience and explicitly says it was not a controlled comparative study.

What evidence should come next?

Test constraint following, result verification, consistent tool use, and supervision needs on your own tasks as new evaluations become available.

What the Benchmark Gap Signals

The score matters because Mistral is seeking a place among frontier AI providers, and developers choosing a model for complex work need evidence about capability as well as access. In this snapshot, Large 4 sits well below the highest-scoring U.S. models and below two listed Chinese models. It is close to DeepSeek V4.1 Flash on index score, though the source analysis says DeepSeek has a much lower measured cost per task.

For agentic workflows, a model may need to plan, use tools, interpret results, and carry decisions across several steps. Mistakes early in that chain can affect later actions, so strong aggregate performance can inform a choice. Still, the index is not a direct test of reliability on every coding, research, or business workflow. A score of 38 does not prove that Large 4 will fail a given task, just as a larger context window does not prove that it can reason accurately over everything supplied.

The article’s assessment that it would not select the current preview for demanding, long-running work is the author’s judgment, not a measured industry consensus. Mistral’s claims about agentic coding and specialist professional work need testing on those workloads. The practical question for users is whether it performs reliably enough, at an acceptable cost and supervision level, for their own tasks.

Amazon

AI development API access tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Today, Weights Later

The release has two separate milestones. On October 6, Mistral introduced the preview API; it also said the weights are scheduled to be released later in October. Those weights had not been made publicly downloadable at the time of the October 7 reporting. Describing the current product as an already available open-weight model would therefore be inaccurate.

Mistral says it trained Large 4 on its own infrastructure in Europe and is continuing to improve it. That is relevant to European AI capacity, but it does not answer how the model compares on every use case. Artificial Analysis reports a context capacity of roughly 512,000 tokens; capacity describes how much material can fit into a request, not whether the model will use it correctly.

The source article also reports that its author encountered hallucinations while using the preview. That is personal experience, not a controlled comparative study. It does not show how often hallucinations occur across users or establish that competing models do not make similar errors.

“I would not choose it for demanding agentic work or long tasks when stronger models are available.”

— Thorsten Meyer, ThorstenMeyerAI.com

Amazon

AI model training and testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Preview Results May Change

The preview may change as Mistral continues development, and the planned weights release could enable broader evaluation. The available source does not specify the exact release date, final model configuration, or whether the public weights will match the API preview in capability.

It is also unclear how Large 4 performs across independent, workload-specific tests of long agentic tasks, coding, and professional applications. The index offers an aggregate comparison, not a direct measure of sustained execution or hallucination rates. The author’s reported hallucinations are anecdotal; no controlled rate or comparable test is supplied. The source material’s discussion of cost is incomplete, so it does not support a full price comparison across models.

Amazon

machine learning model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Workload Tests Ahead

The next stated milestone is Mistral’s planned release of Large 4’s weights later in October 2026. Whether that schedule is met, and what terms or technical details accompany the release, were not confirmed in the source material. New benchmark results may also change the current comparison as model versions and evaluation settings develop.

For developers, the next useful evidence will come from testing the model on their own tasks: whether it follows constraints, verifies results, uses tools consistently, and needs more supervision than alternatives. Until those results and the weights release are available, the current assessment applies to the API preview and the dated benchmark snapshot, not necessarily to future versions.

Amazon

AI image and text processing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Is Mistral Large 4 publicly available?

It is available through a public preview API, according to the source. Its weights were not publicly downloadable as of October 7, 2026.

How does Mistral Large 4 score against frontier models?

Artificial Analysis gave the preview an Intelligence Index score of 38 in its October 7 snapshot. Several listed U.S. and Chinese models scored higher, though the comparison uses different reasoning settings.

Does a score of 38 mean the model will fail agentic tasks?

No. The index is an aggregate benchmark, not a direct prediction of performance on a specific workflow. The source author advises against choosing the preview for demanding long tasks, but that is an assessment rather than a controlled finding about every task.

When are Mistral Large 4’s weights due?

Mistral scheduled the weights release for later in October 2026. The source does not give an exact date or confirm that the release has occurred.

Are reports of hallucinations based on a formal comparison?

No. The source author describes hallucinations encountered in personal use and explicitly presents that experience as not a controlled comparative study.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Electric Bus Registrations in Europe Up 41% in First Half of 2025

Sparking a green revolution, Europe’s electric bus registrations surged 41% in early 2025—discover what’s fueling this rapid shift toward sustainable transit.

ByteDance Prioritizes Other AI Strategies Over Distillation—Here’s Why

ByteDance’s Seed team commits to avoiding AI distillation, even if it slows development, signaling a focus on original training methods over industry shortcuts.

CTOs Are Escaping

Senior tech leaders are shifting from CTO roles to hands-on roles at Anthropic, prioritizing model development over organizational authority amid AI industry shifts.

Apple Has Added Persistent ‘Ads’ To iOS, And It’s Driving Users Crazy

Apple’s latest iOS update includes persistent ads, causing frustration among users. The change is confirmed but the full scope and impact remain unclear.