🔍 Read the full analysis: Mistral Large 4: How Its Global Strength Meets Agent Limitations on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, a sharp rise from its predecessor but below leading US and Chinese models. The release is a notable step for Mistral, while its preview status, high output-token use, pricing and reported hallucinations complicate its case for agent workflows.
Mistral released Large 4 as a Research Public Preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2. That result marks a steep improvement over Mistral’s earlier models, but remains below leading US and Chinese systems, raising questions about its fit for demanding agent tasks and its value at the announced price.
Artificial Analysis describes Large 4 as the most intelligent model from outside the United States and China. The source report notes that this ranking is accurate within the benchmark’s comparisons, but says it should not be read as a claim that the model matches the leading US or Chinese frontier systems. The report lists US models as high as 57.6 and Chinese models as high as 44.8, against Large 4’s 38.4.
Mistral says the model has one trillion parameters, with 49 billion active, accepts text and images, produces text, and supports a 512,000-token context window. It is currently available through Mistral’s API as a research preview. Mistral has promised to release the weights by the end of October; until then, the model is proprietary and its licence has not been published, according to the supplied report.
The reported standard API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. The report says Mistral offered a 50% discount for the first two weeks and that the company says reinforcement learning is still under way, so benchmark results could change. Artificial Analysis data cited in the report put Large 4’s cost at $1.13 per Intelligence Index task.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Cost of Agent Work
The result matters to organisations considering models for multi-step automation, not just question-and-answer chat. The Artificial Analysis Index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Its score therefore offers evidence about performance on work-like tasks, though it is not a guarantee of how a particular deployed workflow will perform.
The source report highlights a cost and capability mismatch: it says GLM-5.3-Flash scored 41.8 at a reported $0.25 per Index task, while DeepSeek V4.1 Flash scored 39.5 at $0.27. Those figures are below Large 4’s reported $1.13 task cost. Such comparisons could matter in high-volume deployments, but buyers should check that pricing, task mix and benchmark conditions match their intended use.
The report also says Large 4 generated 200 million output tokens across the Index, compared with a median of 81 million for comparable models. That is a benchmark observation, not a universal estimate of token use in every application. If reproduced in practice, higher output volume could add latency and cost, especially when an agent makes repeated model calls.
As an affiliate, we earn on qualifying purchases.
A Sharp Rise From Large 3
On the same Artificial Analysis Index version, Mistral Large 3 scored 9 and Medium 3.5 scored 14, according to the source report. Large 4’s 38.4 is therefore a substantial increase within the company’s own lineup. The report characterises it as a major step for a European lab, while stressing that the improvement does not place Mistral among the highest-scoring models in the table.
The release sits between two stages of availability. Customers can access the API research preview now, while Mistral has said model weights are due at the end of October. The report says the licence remains unpublished pending that release. This matters to teams weighing whether they can inspect, host or adapt the model, rather than use it only through a provider’s API.
large language model API subscription
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Preview Results Still May Change
Large 4 is a research preview, and Mistral says reinforcement learning continues. It is not yet clear how much the Index score may change, when the weights will be released beyond the stated end-of-October target, or what licence will apply. The supplied source does not provide a later confirmation of those milestones.
The source report’s concern about confident hallucinations is explicitly an author’s hands-on observation, not a result established by the cited Index. It does not specify a test set, sample size or rate for Large 4. The report also cites hallucination figures for other models, but those comparisons do not establish how Large 4 would behave in a particular business workflow. More information is needed to judge reliability across tasks and deployment settings.
Benchmark scores and task-cost estimates do not settle a procurement decision on their own. Actual results can depend on prompts, tools, workload, context length, output volume and provider pricing. It also remains unclear whether the 50% introductory discount applies to every customer and use case described in the report.
As an affiliate, we earn on qualifying purchases.
Weights and Pricing Will Test Demand
The next stated milestone is Mistral’s planned release of Large 4’s weights by the end of October. The licence and final availability terms will help determine whether developers can run the model outside Mistral’s API and how broadly they can adapt it. Mistral may also publish updated results as reinforcement learning continues.
For buyers, the practical next step is to test the preview on representative tasks and measure accuracy, unsupported claims, token consumption and total cost, rather than relying on a single overall score. A model that is cheaper or more capable on one benchmark may not perform the same way in a company’s tools and data environment. Until more deployment evidence and final release terms are available, Large 4’s case is a clear performance gain for Mistral, but not a settled choice for autonomous workflows.
As an affiliate, we earn on qualifying purchases.
Key Questions
How did Mistral Large 4 score?
It scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the source report. The score is higher than the reported results for Mistral Large 3 and Medium 3.5, but below the leading US and Chinese models listed.
Can developers download Mistral Large 4 now?
Not according to the supplied report. Large 4 is available through Mistral’s API as a Research Public Preview. Mistral has promised the weights for the end of October, but the licence has not yet been published.
Is Large 4 suitable for AI agents?
The benchmark includes agent-oriented tasks, but that alone does not establish suitability for a specific workflow. The report raises concerns about its relative score, output-token use, cost and observed hallucinations; teams would need to test it on their own tasks before relying on it.
What does Mistral Large 4 cost?
The source report lists standard API rates of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks; current eligibility and pricing should be checked with Mistral.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
