AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: How Its Global Strength Meets Agent Limitations on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, a sharp rise from its predecessor but below leading US and Chinese models. The release is a notable step for Mistral, while its preview status, high output-token use, pricing and reported hallucinations complicate its case for agent workflows.

Mistral released Large 4 as a Research Public Preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2. That result marks a steep improvement over Mistral’s earlier models, but remains below leading US and Chinese systems, raising questions about its fit for demanding agent tasks and its value at the announced price.

Artificial Analysis describes Large 4 as the most intelligent model from outside the United States and China. The source report notes that this ranking is accurate within the benchmark’s comparisons, but says it should not be read as a claim that the model matches the leading US or Chinese frontier systems. The report lists US models as high as 57.6 and Chinese models as high as 44.8, against Large 4’s 38.4.

Mistral says the model has one trillion parameters, with 49 billion active, accepts text and images, produces text, and supports a 512,000-token context window. It is currently available through Mistral’s API as a research preview. Mistral has promised to release the weights by the end of October; until then, the model is proprietary and its licence has not been published, according to the supplied report.

The reported standard API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. The report says Mistral offered a 50% discount for the first two weeks and that the company says reinforcement learning is still under way, so benchmark results could change. Artificial Analysis data cited in the report put Large 4’s cost at $1.13 per Intelligence Index task.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral has released Large 4 as a research preview, with independent benchmark data showing substantial progress but a remaining gap to leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Agent Work

The result matters to organisations considering models for multi-step automation, not just question-and-answer chat. The Artificial Analysis Index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Its score therefore offers evidence about performance on work-like tasks, though it is not a guarantee of how a particular deployed workflow will perform.

The source report highlights a cost and capability mismatch: it says GLM-5.3-Flash scored 41.8 at a reported $0.25 per Index task, while DeepSeek V4.1 Flash scored 39.5 at $0.27. Those figures are below Large 4’s reported $1.13 task cost. Such comparisons could matter in high-volume deployments, but buyers should check that pricing, task mix and benchmark conditions match their intended use.

The report also says Large 4 generated 200 million output tokens across the Index, compared with a median of 81 million for comparable models. That is a benchmark observation, not a universal estimate of token use in every application. If reproduced in practice, higher output volume could add latency and cost, especially when an agent makes repeated model calls.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Rise From Large 3

On the same Artificial Analysis Index version, Mistral Large 3 scored 9 and Medium 3.5 scored 14, according to the source report. Large 4’s 38.4 is therefore a substantial increase within the company’s own lineup. The report characterises it as a major step for a European lab, while stressing that the improvement does not place Mistral among the highest-scoring models in the table.

The release sits between two stages of availability. Customers can access the API research preview now, while Mistral has said model weights are due at the end of October. The report says the licence remains unpublished pending that release. This matters to teams weighing whether they can inspect, host or adapt the model, rather than use it only through a provider’s API.

Amazon

large language model API subscription

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results Still May Change

Large 4 is a research preview, and Mistral says reinforcement learning continues. It is not yet clear how much the Index score may change, when the weights will be released beyond the stated end-of-October target, or what licence will apply. The supplied source does not provide a later confirmation of those milestones.

The source report’s concern about confident hallucinations is explicitly an author’s hands-on observation, not a result established by the cited Index. It does not specify a test set, sample size or rate for Large 4. The report also cites hallucination figures for other models, but those comparisons do not establish how Large 4 would behave in a particular business workflow. More information is needed to judge reliability across tasks and deployment settings.

Benchmark scores and task-cost estimates do not settle a procurement decision on their own. Actual results can depend on prompts, tools, workload, context length, output volume and provider pricing. It also remains unclear whether the 50% introductory discount applies to every customer and use case described in the report.

Amazon

AI token usage monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Pricing Will Test Demand

The next stated milestone is Mistral’s planned release of Large 4’s weights by the end of October. The licence and final availability terms will help determine whether developers can run the model outside Mistral’s API and how broadly they can adapt it. Mistral may also publish updated results as reinforcement learning continues.

For buyers, the practical next step is to test the preview on representative tasks and measure accuracy, unsupported claims, token consumption and total cost, rather than relying on a single overall score. A model that is cheaper or more capable on one benchmark may not perform the same way in a company’s tools and data environment. Until more deployment evidence and final release terms are available, Large 4’s case is a clear performance gain for Mistral, but not a settled choice for autonomous workflows.

Amazon

AI model cost management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

It scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the source report. The score is higher than the reported results for Mistral Large 3 and Medium 3.5, but below the leading US and Chinese models listed.

Can developers download Mistral Large 4 now?

Not according to the supplied report. Large 4 is available through Mistral’s API as a Research Public Preview. Mistral has promised the weights for the end of October, but the licence has not yet been published.

Is Large 4 suitable for AI agents?

The benchmark includes agent-oriented tasks, but that alone does not establish suitability for a specific workflow. The report raises concerns about its relative score, output-token use, cost and observed hallucinations; teams would need to test it on their own tasks before relying on it.

What does Mistral Large 4 cost?

The source report lists standard API rates of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks; current eligibility and pricing should be checked with Mistral.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Every Top-Tier Frontier AI Model Now Leverages Mixture-of-Experts

Explains why MoE models dominate 2026 AI, enabling massive capacity without proportional costs, and how this transforms model scalability and efficiency.

The Role of AI in Optimizing Electric Bus Routes

Journey into how AI transforms electric bus routing, unlocking smarter, more efficient transit—discover the innovations that could change transportation forever.

Jack Clark Says It Out Loud — Reading the Co-Founder’s 60%/2028 Estimate on Automated AI R&D

Anthropic’s co-founder Jack Clark states there’s a 60%+ probability that AI systems can autonomously develop their own successors by the end of 2028, marking a significant policy forecast.

Buried Apple Feature Turns An iPhone Into The Perfect Kids’ Dumb Phone

A secret Apple feature enables iPhones to function as simplified devices, ideal for children, by disabling advanced functions and access.