🔍 Read the full analysis: Beyond US And China, Mistral Large 4 Has Appeal—but Agents Are A Concern on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, making it the highest-scoring model in the source’s comparison from outside the United States and China. The score is a major improvement over Mistral’s earlier models, but remains below leading US and Chinese systems; the source also raises concerns about agent performance, output volume, cost and hallucinations.
Mistral has released Large 4, a research preview that scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, putting it ahead of models from outside the United States and China in the comparison cited by the source. The result marks a sharp improvement for the French company, but the same index places Large 4 below major US and Chinese rivals, while the source report questions its cost and reliability for long-running agent tasks.
Artificial Analysis’s ranking puts Large 4 below the US models listed in the source, whose scores range from 52.6 to 57.6, and below several Chinese models, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The source characterizes Large 4 as the highest-scoring model outside the US and China in its comparison. That framing describes a limited competitive field; it does not put Mistral at the top of the overall rankings.
The release is a substantial step up from Mistral Large 3, which scored 9 on the same index version, and Medium 3.5, which scored 14. Large 4 has a trillion parameters, with 49 billion active, and accepts text and images while producing text. It has a 512,000-token context window. Mistral offers it through an API as a Research Public Preview; the source says Mistral promised model weights by the end of October, but the licence had not been published at the time of the report.
The reported standard API price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. Mistral offered a 50% discount for the first two weeks. The source says Large 4 cost $1.13 per Artificial Analysis Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models scored 41.8 and 39.5 respectively, according to the cited data. Mistral says reinforcement learning is still underway, so the model’s scores may change.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Cost of Using Large 4 for Agents
The ranking matters to companies choosing models for software agents and other work that requires multiple steps. Artificial Analysis’s index includes agent-focused tests such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. The source argues that performance gaps can become more consequential across a sequence of actions: an early mistake may shape later decisions, making a model’s score relevant beyond short chat exchanges.
Cost and output volume add practical concerns. In the index evaluation, Large 4 generated 200 million output tokens, compared with a median of 81 million for the peer group cited by the source. More output can mean higher token costs and longer waits in workflows that make repeated model calls. The source’s task-cost comparison also shows two cited Chinese models scoring higher at less than one-quarter of Large 4’s reported cost per task. These figures are tied to the report’s benchmark and pricing assumptions; they do not establish what every customer will pay in production.
The report’s concern about hallucinations is a hands-on observation by its author, not a finding from the Artificial Analysis index. The author says they saw Large 4 confidently state false information. If that behavior occurs within an agent workflow, later steps could rely on a false premise. That risk warrants testing, but the source does not provide a controlled hallucination rate for Large 4 that would quantify it.
As an affiliate, we earn on qualifying purchases.
A Sharp Rise From Mistral’s Earlier Scores
Mistral’s previous scores help explain why Large 4 drew attention. Its move from 9 for Large 3 to 38.4 for Large 4, measured on the same Artificial Analysis index version according to the source, is a marked improvement. It does not erase the gap to the top-ranked models, but it shows the company has moved substantially closer to the current field than its earlier results suggested.
The description of Large 4 as the most intelligent model outside the US and China comes from the Artificial Analysis result as presented by the source. The report cautions that this does not mean it competes evenly with the leading US and Chinese labs: the cited leaderboard includes multiple models from both countries scoring above it. The weights are also not yet available, so buyers cannot treat the preview as an already released open-weights model. Mistral’s planned weights release and still-unpublished licence leave important questions for developers evaluating deployment and reuse.
“Reinforcement learning is still running.”
— Mistral
As an affiliate, we earn on qualifying purchases.
Weights, Licensing and Reliability
Several details remain unsettled. The source says Mistral had promised to release Large 4’s weights by the end of October, but the weights were not yet available and the licence had not been published. The source does not establish whether that timetable was later met. Mistral also says reinforcement learning is ongoing, so the preview’s performance may change.
The reported hallucination concern is based on the source author’s hands-on testing, and no Large 4 hallucination rate is supplied. The benchmark scores and task-cost estimates come from Artificial Analysis’s specified index and the report’s pricing calculations; actual results may differ by workload, prompt design and usage pattern. The source also does not provide enough detail to determine how the model performs across different production agent setups.
AI image and text processing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Updated Scores
The next reported milestones are Mistral’s planned weights release by the end of October and further model evaluation as reinforcement learning continues. Developers and buyers will be able to assess licensing and deployment options once the weights and licence are available. Updated independent benchmark results could also show whether Large 4’s preview score changes.
For now, the source’s evidence points to a model with a large score improvement and a competitive position outside the US and China, alongside gaps in the cited leaderboard and unanswered questions about cost, hallucination behavior and agent reliability. Organizations considering it for multi-step work would need to test it against their own tasks rather than treating its geographic ranking as proof of suitability.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral release?
Mistral released Large 4 as a Research Public Preview through its API. The source describes it as a trillion-parameter, natively multimodal model with a 512,000-token context window; model weights were promised for the end of October.
How did Mistral Large 4 score?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. That was above Mistral Large 3’s score of 9 on the same index version, but below the US and Chinese models listed ahead of it.
Why does the report raise concerns about agents?
The report points to Large 4’s lower benchmark score than several leading rivals, its high output-token use in the index evaluation, and the author’s observation of confident false statements. These concerns may matter more in multi-step workflows, where later actions can rely on earlier outputs. The source does not give a controlled hallucination rate for Large 4.
Is Mistral Large 4 open-weight?
Not at the time described in the source. It was available as a proprietary API preview; Mistral had promised weights for the end of October, and the licence had not been published.
How does its reported task cost compare with some rivals?
The source reports an estimated $1.13 per index task for Large 4, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those models scored higher in the cited index, but the figures reflect the report’s benchmark and pricing assumptions rather than every real-world workload.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
