Mistral Large 4: A Leading Option Outside The US And China, But Not For Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: A Leading Option Outside The US And China, But Not For Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, making it a notable European model but leaving it behind leading US and Chinese systems. The source argues that its benchmark standing, output volume, and cost make it a poor choice for agent workflows; the model remains in research preview, with weights and licensing details still pending.

Mistral AI released Large 4 in research public preview, and the model scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. That marks a sharp rise from earlier Mistral models, but leaves Large 4 below major US and Chinese systems in the cited rankings, complicating the claim that it is a compelling alternative for demanding AI work.

The source describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, multimodal text and image input, text output, and a 512,000-token context window. Mistral has made it available through its API as a research preview. The company says it expects to release the weights at the end of October; until then, the model is proprietary, and the source says its licence has not been published.

On the cited Artificial Analysis index, Large 4 scored 38.4 points. The source compares that result with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version. It ranks below the listed US frontier models, whose scores range from 51.8 to 57.6, and below several Chinese models, including GLM-5.3 at 44.8 and Kimi K3 at 43.6. It is ahead of some older or smaller models in the table, including DeepSeek V4 Pro 0813 at 36.0.

The source lists API pricing at $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. It also reports a 50% discount for the first two weeks. Mistral says reinforcement learning is still underway, so the model’s scores could change. The source’s author additionally reports observing confident false statements in hands-on tests; that is an individual observation, not a benchmark finding.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 in research preview, presenting a major benchmark improvement while independent figures cited in the source show it trailing leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A European Model Faces a Cost Test

Large 4’s release matters because it gives buyers a new European-developed option at a time when the highest-scoring systems in the cited table come from US and Chinese labs. The result is also a substantial improvement over Mistral’s previous models on the same index. But the numbers do not support treating it as a peer to the leaders: its 38.4 score trails the listed top US models by a wide margin, and several Chinese systems also rank above it.

For companies choosing models for automated, multi-step work, raw capability is only part of the decision. The source says Large 4 generated 200 million output tokens while completing the index, compared with a median of 81 million for comparable models. More output can add cost and latency in repeated workflows. The source also puts Large 4 at $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, both of which scored higher in the cited results. Those figures make the model’s operational economics a central question, not a minor detail.

The label “outside the US and China” may be useful for buyers seeking geographic diversity, but it does not mean Large 4 leads the broader field. The source characterizes that comparison as a narrow one, with few other labs outside those countries competing at the same level. For procurement teams, the practical distinction is between having a European option and having a model that matches leading alternatives on performance, reliability, and cost.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Large 4 Ranks Today

The ranking figures cited in the source come from Artificial Analysis Intelligence Index v4.3.2, which includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench, and Terminal-Bench 4.0. That means the score reflects performance across a set of work and coding tasks, rather than only general knowledge questions. The source argues this makes the index relevant to buyers considering agents, though any single composite score cannot settle how a model will perform in a particular company’s workflow.

Mistral’s earlier scores provide a measure of the change: Large 3 recorded 9, while Medium 3.5 recorded 14 on the same index version. The source calls Large 4’s move to 38.4 a major step for the company. Its ranking still falls short of the newer Chinese models listed in the source, while its announced weights have not yet been released. The comparison may change as new models arrive or as Mistral updates Large 4.

“Large 4 shows it has closed a lot of ground — and that it is still not one.”

— ThorstenMeyerAI.com source author

Amazon

multimodal AI image text input device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Reliability

Several details remain unsettled. Mistral has said the weights are expected at the end of October, but the source does not give a more precise date. It also says the model’s licence has not been published, so the terms under which organizations could use or modify the weights are not yet clear. The reported preview pricing includes a two-week discount, but the source does not specify the discount’s start and end dates.

Performance may also change while reinforcement learning continues. The source does not provide a complete account of the hands-on tests behind its hallucination observations, and those observations should not be confused with independent benchmark measurements. Nor do the cited aggregate scores establish how Large 4 will perform on every agent task, with every tool setup, or against a buyer’s own evaluation set.

Amazon

large context window AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

October Weights Could Change Access

The next expected milestone is Mistral’s planned release of Large 4’s weights at the end of October. Publication of the weights and licence would clarify how broadly developers and businesses can deploy the model outside Mistral’s API. Buyers can also watch for updated benchmark results as reinforcement learning continues and for further detail on preview pricing after the initial discount period.

For now, organizations weighing Large 4 against alternatives can compare the cited benchmarks with their own tasks, token use, latency, and reliability requirements. The available figures support calling it a major improvement for Mistral, but do not establish that it is the best option for agent workloads or that the current ranking will hold after updates and new releases.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is a model released by Mistral AI in research public preview. The source describes it as a one-trillion-parameter multimodal model with 49 billion active parameters and a 512,000-token context window.

How did Large 4 score against leading models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. Several listed US and Chinese models scored higher, including top US entries above 50 and GLM-5.3 at 44.8.

Can developers download its weights now?

Not yet, according to the source. Mistral is offering Large 4 through its API as a research preview and has said the weights are expected at the end of October. The source says the licence has not been published.

Is Large 4 a good choice for AI agents?

The source argues against choosing it for demanding agent workflows based on its benchmark position, output volume, and per-task cost. Those are reported comparisons, not a universal verdict; buyers should test the model on their own tasks and compare reliability and total usage costs.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The 90-Day Window Closed. Nobody Sent a Notice.

The 90-day window for responsible vulnerability disclosure has effectively ended without any notices from vendors, raising concerns about security risks.

Two Channels: How the Pentagon Just Split Frontier-AI Procurement in Half

The Pentagon announced a split in its AI procurement, placing Anthropic in a separate cybersecurity channel from the multi-vendor classified network, affecting strategic relationships.

Using Management Tests To Understand AI’s Authentic Behavior

Exploring how management-style tests reveal true AI capabilities in decision-making, trust, and execution through live business simulations.

How to Reduce Heat and Noise in a High-Power AI Workstation

Learn effective strategies to lower heat and noise in high-power AI workstations, including undervolting, airflow improvements, and component choices.