🔍 Read the full analysis: Mistral Large 4 Trails The Leading Edge Of AI on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral launched Mistral Large 4 in public preview on October 6, 2026, with API access and a planned release of model weights later in the month. Artificial Analysis gives the preview an Intelligence Index score of 38, below several leading US and Chinese models; the score is a dated benchmark comparison, not a prediction of success on every task.
Mistral launched Mistral Large 4 in public preview on October 6, offering developers API access to its largest model yet, but benchmark results cited by ThorstenMeyerAI.com put it behind several leading US and Chinese systems. The model’s weights are scheduled for release later in October and were not publicly downloadable as of October 7.
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. It accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it. Those statements come from Mistral’s announcement; they do not by themselves establish how well the preview performs on specific tasks.
Artificial Analysis gives Mistral Large 4 Preview an Intelligence Index score of 38. In the comparison reported by ThorstenMeyerAI.com on October 7, that matched OpenAI’s GPT-6 Luna at maximum reasoning effort and sat just below DeepSeek V4.1 Flash at maximum effort, which scored 39. Higher scores indicate stronger aggregate performance on the benchmark suite, but the models were not evaluated under identical reasoning settings or compute budgets.
The same snapshot lists Anthropic’s Claude Opus 5.5 at 58, Google’s Gemini 4 Argon at 53 and OpenAI’s GPT-6.1 Sol at 52. China’s Z.ai GLM-5.3 scored 45 and Moonshot AI’s Kimi K3 scored 44. Cohere’s Command A+ scored 13. The figures are index points, not percentages, and the developer locations identify the companies, not where an individual API request is processed.
AI MODEL WATCH · OCTOBER 7, 2026
Mistral Large 4 Trails The Leading Edge Of AI
Mistral’s new preview opens API access to its largest model yet. An October 7 benchmark snapshot places it behind several leading systems, while its promised weight release remains ahead.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
Thorsten Meyer · Current preview assessment01 / BENCHMARK SNAPSHOT
A measured gap at launch
Artificial Analysis index points, as reported October 7. Higher scores indicate stronger aggregate results across its benchmark suite.
The comparison used differing reasoning settings and compute budgets. Scores are not percentages and do not predict performance on every task.
Announced October 6, 2026, with developer access through Mistral’s API.
As of October 7, users could not yet download the model weights.
Mistral says it trained Large 4 on its own infrastructure in Europe and continues to improve it.
02 / HOW TO READ THE RESULT
Benchmark position is one signal
The score is a dated aggregate comparison. It cannot settle how well the preview handles a particular project.
Test the actual work
Planning, tool use, coding, and research can expose different strengths and weaknesses. Run representative tasks before assigning consequential work.
Long tasks need scrutiny
An early error can steer later steps. A polished final answer may hide a process that went off course, so inspect intermediate results too.
No like-for-like trial
The source provides no controlled comparison of long workflows. Meyer reports hallucinations in personal use, not a measured competitor-wide rate.
03 / ACCESS TIMELINE
Preview first, weights to follow
Access format matters: an API preview is not yet a downloadable open-weight release.
04 / WHAT THE SNAPSHOT CANNOT SAY
Keep the measures in context
Large scale, context capacity, and infrastructure are useful facts, but they answer different questions from task accuracy.
Models were not evaluated at identical reasoning effort or compute budgets, so the ranking is not a perfectly matched trial.
Roughly 512,000 tokens can fit in a request; that capacity does not show whether the model reasons correctly over all the material.
Developer locations identify the companies. They do not reveal where an individual API request is routed or processed.
The source omits the cost figures and offers no controlled reliability study, so it supports no price or comparative hallucination conclusion.
05 / QUICK QUESTIONS
What developers should know
What did Mistral announce?
Large 4 entered public preview on October 6, 2026. API access was available; weights were scheduled for later in October.
What does a score of 38 mean?
It is an aggregate Intelligence Index result in the cited October 7 comparison, not a percentage or a guarantee for any task.
Does the score prove unreliability?
No. It places the preview below several models in that benchmark snapshot. Reliability on a specific workflow requires direct testing.
What remains unknown?
Weight-release details, updated scores, task-specific accuracy, long-workflow reliability, and the performance of future model updates.
The Gap for Complex AI Work
The results matter to developers choosing a model for tasks that require planning, tool use and multiple steps. In an agentic workflow, an early mistake can shape later decisions, while a polished final answer may not reveal that the process went off course. For those uses, aggregate benchmark standing can be one piece of evidence—but it does not substitute for testing a model on the developer’s own workload.
Thorsten Meyer, writing on ThorstenMeyerAI.com, concludes that he would not select the current preview for demanding agentic work or long tasks when stronger alternatives are available. That is his assessment, informed by the benchmark comparison and personal use, rather than a controlled study of all models or workflows. The reported score does not prove that Large 4 will fail a particular coding, research or business task.
Mistral promotes the model’s capabilities in agentic coding and specialized professional work, but the source material does not provide workload-specific results establishing those claims. The practical question is whether it performs reliably enough, with little enough supervision, on a real task. For now, the comparison gives developers reason to test before assigning consequential or lengthy work to the preview.
As an affiliate, we earn on qualifying purchases.
Preview Access Before Weight Release
The distinction between a preview API and a public weight release is central to the launch. As of October 7, developers could access Large 4 through Mistral’s API, but the weights were not yet available to download. Mistral said they were scheduled to arrive later in October. Until that release, users should not treat the model as an already available open-weight system.
Mistral’s European training infrastructure is relevant to the company’s effort to build AI capacity in Europe. But the report separates that development from the question of comparative capability: European training and a large parameter count do not demonstrate parity with higher-scoring models. The available figures are a dated snapshot from Artificial Analysis, and the scores may change as systems and evaluations are updated.
The benchmark comparison also has limits. The reasoning settings differ, and locations in the table refer to developers rather than the physical routing or processing of API requests. A large context window is another separate measure: Artificial Analysis reports roughly 512,000 tokens for Large 4, but context capacity indicates how much material can fit into a request, not whether the model can reason correctly over all of it.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
large language model training infrastructure
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Reliability Beyond the Index
The benchmark score does not establish how Large 4 will perform across particular coding, research or professional tasks, nor does it directly measure reliability across long agentic workflows. The source gives no controlled, like-for-like evaluation of those uses. Meyer reports encountering hallucinations in his own use, but says this was not a controlled comparison; it cannot establish how often the model hallucinates relative to competitors.
The source material also begins a discussion of model costs but does not include the cost figures or the remainder of that comparison. It therefore does not support a conclusion here about Large 4’s price or cost-effectiveness. The model’s scheduled weight release, further improvements, and performance after updates remain to be seen.
AI image and text processing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weight Release and Further Testing
Mistral said the model weights were scheduled for release later in October. The timing and details of that release were not confirmed further in the source material. Developers can assess the preview through the API, but should distinguish their own task results from broad claims about model capability.
Relevant next evidence would include the weight release, updated benchmark results and independent, workload-specific tests of accuracy and reliability on longer tasks. Until that evidence is available, the October 7 comparison remains a snapshot of the preview—not a final measure of how the model may perform after Mistral’s stated improvements.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral introduced Mistral Large 4 in public preview on October 6, 2026. Developers could access it through an API; the weights were scheduled for release later in October.
How did Mistral Large 4 score?
Artificial Analysis gave the preview an Intelligence Index score of 38 in the October 7 comparison cited by ThorstenMeyerAI.com. The score is an aggregate benchmark result, not a percentage or a guarantee of task performance.
Does the score prove the model is unreliable?
No. The score places the preview below several models in that benchmark comparison, but it does not prove how it will perform on a specific workflow. Long-task reliability requires testing on the intended work.
Are Mistral Large 4’s weights available?
Not according to the source material as of October 7, 2026. Mistral said the weights were scheduled for release later in October; the source does not confirm that they have since been released.
What remains unknown about the model?
Independent, like-for-like evidence about its performance on demanding agentic tasks, hallucination rates and cost-effectiveness is not provided in the source material. Its results may also change as Mistral updates the preview.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
