🔍 Read the full analysis: Mistral Large 4: A Leading Option Outside The US And China, But Not For Agents on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, making it a notable European model but leaving it behind leading US and Chinese systems. The source argues that its benchmark standing, output volume, and cost make it a poor choice for agent workflows; the model remains in research preview, with weights and licensing details still pending.
Mistral AI released Large 4 in research public preview, and the model scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. That marks a sharp rise from earlier Mistral models, but leaves Large 4 below major US and Chinese systems in the cited rankings, complicating the claim that it is a compelling alternative for demanding AI work.
The source describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, multimodal text and image input, text output, and a 512,000-token context window. Mistral has made it available through its API as a research preview. The company says it expects to release the weights at the end of October; until then, the model is proprietary, and the source says its licence has not been published.
On the cited Artificial Analysis index, Large 4 scored 38.4 points. The source compares that result with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version. It ranks below the listed US frontier models, whose scores range from 51.8 to 57.6, and below several Chinese models, including GLM-5.3 at 44.8 and Kimi K3 at 43.6. It is ahead of some older or smaller models in the table, including DeepSeek V4 Pro 0813 at 36.0.
The source lists API pricing at $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. It also reports a 50% discount for the first two weeks. Mistral says reinforcement learning is still underway, so the model’s scores could change. The source’s author additionally reports observing confident false statements in hands-on tests; that is an individual observation, not a benchmark finding.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
A European Model Faces a Cost Test
Large 4’s release matters because it gives buyers a new European-developed option at a time when the highest-scoring systems in the cited table come from US and Chinese labs. The result is also a substantial improvement over Mistral’s previous models on the same index. But the numbers do not support treating it as a peer to the leaders: its 38.4 score trails the listed top US models by a wide margin, and several Chinese systems also rank above it.
For companies choosing models for automated, multi-step work, raw capability is only part of the decision. The source says Large 4 generated 200 million output tokens while completing the index, compared with a median of 81 million for comparable models. More output can add cost and latency in repeated workflows. The source also puts Large 4 at $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, both of which scored higher in the cited results. Those figures make the model’s operational economics a central question, not a minor detail.
The label “outside the US and China” may be useful for buyers seeking geographic diversity, but it does not mean Large 4 leads the broader field. The source characterizes that comparison as a narrow one, with few other labs outside those countries competing at the same level. For procurement teams, the practical distinction is between having a European option and having a model that matches leading alternatives on performance, reliability, and cost.
As an affiliate, we earn on qualifying purchases.
How Large 4 Ranks Today
The ranking figures cited in the source come from Artificial Analysis Intelligence Index v4.3.2, which includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench, and Terminal-Bench 4.0. That means the score reflects performance across a set of work and coding tasks, rather than only general knowledge questions. The source argues this makes the index relevant to buyers considering agents, though any single composite score cannot settle how a model will perform in a particular company’s workflow.
Mistral’s earlier scores provide a measure of the change: Large 3 recorded 9, while Medium 3.5 recorded 14 on the same index version. The source calls Large 4’s move to 38.4 a major step for the company. Its ranking still falls short of the newer Chinese models listed in the source, while its announced weights have not yet been released. The comparison may change as new models arrive or as Mistral updates Large 4.
“Large 4 shows it has closed a lot of ground — and that it is still not one.”
— ThorstenMeyerAI.com source author
multimodal AI image text input device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights, Licensing and Reliability
Several details remain unsettled. Mistral has said the weights are expected at the end of October, but the source does not give a more precise date. It also says the model’s licence has not been published, so the terms under which organizations could use or modify the weights are not yet clear. The reported preview pricing includes a two-week discount, but the source does not specify the discount’s start and end dates.
Performance may also change while reinforcement learning continues. The source does not provide a complete account of the hands-on tests behind its hallucination observations, and those observations should not be confused with independent benchmark measurements. Nor do the cited aggregate scores establish how Large 4 will perform on every agent task, with every tool setup, or against a buyer’s own evaluation set.
As an affiliate, we earn on qualifying purchases.
October Weights Could Change Access
The next expected milestone is Mistral’s planned release of Large 4’s weights at the end of October. Publication of the weights and licence would clarify how broadly developers and businesses can deploy the model outside Mistral’s API. Buyers can also watch for updated benchmark results as reinforcement learning continues and for further detail on preview pricing after the initial discount period.
For now, organizations weighing Large 4 against alternatives can compare the cited benchmarks with their own tasks, token use, latency, and reliability requirements. The available figures support calling it a major improvement for Mistral, but do not establish that it is the best option for agent workloads or that the current ranking will hold after updates and new releases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4?
Mistral Large 4 is a model released by Mistral AI in research public preview. The source describes it as a one-trillion-parameter multimodal model with 49 billion active parameters and a 512,000-token context window.
How did Large 4 score against leading models?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. Several listed US and Chinese models scored higher, including top US entries above 50 and GLM-5.3 at 44.8.
Can developers download its weights now?
Not yet, according to the source. Mistral is offering Large 4 through its API as a research preview and has said the weights are expected at the end of October. The source says the licence has not been published.
Is Large 4 a good choice for AI agents?
The source argues against choosing it for demanding agent workflows based on its benchmark position, output volume, and per-task cost. Those are reported comparisons, not a universal verdict; buyers should test the model on their own tasks and compare reliability and total usage costs.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
