🔍 Read the full analysis: Mistral Large 4 Trails As The AI Frontier Pushes Forward on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral launched a public API preview of Large 4 on October 6, 2026. Artificial Analysis gives it an Intelligence Index score of 38, below leading US models and several Chinese alternatives; the model’s weights are not yet publicly available. Claims about agentic coding and professional tasks, as well as concerns based on one reviewer’s experience, still need workload-specific testing.
Mistral AI launched a public API preview of Mistral Large 4 on October 6, introducing its largest model to date as the company competes with leading US and Chinese AI developers. In a benchmark snapshot published by Artificial Analysis and available October 7, the preview scored 38 on the Intelligence Index—below several leading models, though comparable to OpenAI’s GPT-6 Luna at maximum reasoning effort. The model’s weights are not yet publicly downloadable.
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. The preview accepts text and images through an API. Mistral says the model was trained on its own infrastructure in Europe and is still being improved. The company has scheduled a release of the weights for later in October, but that release had not happened as of October 7.
Artificial Analysis’s dated comparison gives the preview a score of 38. In the same snapshot, Anthropic’s Claude Opus 5.5 scored 58, Google’s Gemini 4 Argon scored 53, and OpenAI’s GPT-6.1 Sol scored 52. China’s GLM-5.3 scored 45, Kimi K3 scored 44, and DeepSeek V4.1 Flash scored 39. The comparison includes different model reasoning settings, so it is not an evaluation under identical compute budgets.
The source review argues that the current preview is not its author’s preferred option for demanding agentic work or long tasks. That is an assessment, not a benchmark finding. The reviewer says their own use surfaced hallucinations, while acknowledging this was not a controlled comparison. Artificial Analysis reports a context capacity of roughly 512,000 tokens; that figure describes input capacity, not accuracy or reliable performance across long tasks.
The Gap in Frontier Benchmarks
The scores place Mistral’s preview below several competitors on one aggregate benchmark suite at a moment when developers are weighing models for increasingly complex work. In the source’s comparison, Claude Opus 5.5 leads by 20 index points, Gemini 4 Argon by 15, and GPT-6.1 Sol by 14. These are index-point differences, not percentages or direct predictions of success on a user’s task.
For developers considering autonomous workflows, aggregate performance is relevant but not decisive. Such workflows can involve planning, tool use, interpreting results, and carrying decisions across steps. A weakness at one stage may affect later work, yet an index score does not establish how reliably a model handles a particular coding, research, or business process. Mistral’s claims about agentic coding and specialized professional tasks need to be checked against those workloads.
The release also matters beyond the leaderboard. Mistral says it trained Large 4 on its own European infrastructure, and the model’s scale marks a substantial product-line development. But the announcement and European training claims do not by themselves establish benchmark leadership, dependable long-task performance, or the availability of downloadable weights today.
As an affiliate, we earn on qualifying purchases.
Preview Access Before Weights
The current release is an API preview, not a completed public-weight release. That distinction affects what developers can do now: they can test the model through Mistral’s API, but the source reports that its weights remain unavailable for public download. Mistral has said the weights are due later in October; the schedule is a future plan, not a completed release.
The Intelligence Index figures are a dated snapshot from October 7, 2026, and may change as models or evaluations are updated. The source comparison identifies developers’ locations, not where a particular API request is processed. It also notes that the models were tested at different reasoning settings, which limits direct comparisons.
The comparison is not a claim that every competitor scores higher. Canada’s Cohere Command A+ scored 13 in the same table, below Mistral. That result makes the narrower conclusion more appropriate: Mistral’s preview trails the named leading US models and several Chinese alternatives on this index, while the index does not capture every product strength or use case.
“Mistral says it trained Large 4 on its own infrastructure in Europe and continues to improve the model.”
— Mistral AI, as described in its announcement
As an affiliate, we earn on qualifying purchases.
Performance Beyond the Index
It remains unclear how Large 4 performs on specific coding, research, or professional workflows compared with alternatives. The available score is an aggregate benchmark measure, not a direct test of long-task reliability or a guarantee of success or failure on any individual assignment. Mistral’s advertised strengths in agentic coding and specialized work have not been established here by workload-specific results.
The reviewer’s report of hallucinations reflects personal experience, not a controlled comparative study, and does not show how often unsupported answers occur across users or tasks. The available source material also does not provide a complete cost comparison or sufficient pricing detail to determine whether the preview offers an advantage for a particular workload. Benchmark scores and model behavior may change as the preview is updated.
As an affiliate, we earn on qualifying purchases.
Weights and Further Testing
The next announced milestone is Mistral’s planned release of Large 4’s weights later in October. Until then, the public access described in the source is through the preview API. The company says it is continuing to improve the model, but the source does not specify what changes are planned or give a confirmed date beyond the month.
For developers, the practical next step is to test the preview against representative tasks and compare results under clearly described settings, including cost, error rates, and the amount of human review needed. Once the weights are released, their availability will add a new option for evaluation; it will not, on its own, settle how the model performs on complex agentic work. Updated benchmarks and direct tests will be needed to assess whether the preview’s current ranking changes.
AI model performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral announced a public API preview of Large 4 on October 6, 2026. It describes the model as a one-trillion-parameter mixture-of-experts system with 49 billion active parameters that accepts text and images.
Can the public download Mistral Large 4’s weights now?
No. As of October 7, the weights were not publicly downloadable. Mistral has scheduled their release for later in October, according to the source.
How does Large 4 rank in the cited comparison?
Artificial Analysis gives the preview an Intelligence Index score of 38. That is below several named US and Chinese models in the October 7 snapshot, and matches GPT-6 Luna at maximum reasoning effort. The comparison uses different reasoning settings.
Does the score show that Large 4 will fail at agentic tasks?
No. The score is an aggregate benchmark result, not a direct measure of success on a specific agentic workflow. The source reviewer’s recommendation against using it for demanding long tasks is an opinion informed by benchmarks and personal experience, not proof that it will fail every such task.
What is the next development to watch?
The main announced milestone is the planned release of the model weights later in October. Further benchmark results and workload-specific testing may clarify how the model performs as Mistral continues to update the preview.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
