Should You Choose Mistral Large 4 For Agents? Its Global Strengths And Limits
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Should You Choose Mistral Large 4 For Agents? Its Global Strengths And Limits on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research preview, scoring 38.4 on Artificial Analysis Intelligence Index v4.3.2. The score marks a substantial gain over earlier Mistral models, but the source’s benchmark comparisons show leading US and several Chinese models scoring higher, while its reported task cost and output volume raise concerns for agent use.

Mistral has released Large 4, a 1-trillion-parameter multimodal model available as an API research preview, with weights promised for the end of October. On Artificial Analysis Intelligence Index v4.3.2, it scores 38.4, a large jump from earlier Mistral models but below the leading US and several Chinese systems—evidence that its progress does not by itself establish it as a strong choice for long-running agents.

The source report says Large 4 has 49 billion active parameters, accepts text and images, produces text, and supports a 512,000-token context window. It is currently available through Mistral’s API as a research preview. Mistral has said model weights are expected at the end of October; the report says the licence has not been published. API pricing is listed as $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks.

Artificial Analysis gives Large 4 a score of 38.4. The report compares this with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, describing the new result as a substantial improvement. Its table lists several US models above 52 and Chinese models including GLM-5.3 at 44.8, Kimi K3 at 43.6, and DeepSeek V4.1 Flash at 39.5. These are benchmark results, not a guarantee of how a model will perform in a particular customer’s workflow.

The report also cites Artificial Analysis figures putting Large 4’s cost at $1.13 per index task and its total output at 200 million tokens to complete the index. It compares that output with a median of 81 million tokens for comparable models. Its account says GLM-5.3-Flash costs $0.25 per task and scores 41.8, while DeepSeek V4.1 Flash costs $0.27 and scores 39.5. Those figures suggest buyers should compare completed-work costs, not just advertised token rates.

At a glance
analysisWhen: Released the day before the source repo…
The developmentMistral released Large 4 as an API research preview, prompting fresh comparisons of its agent-task performance, price and standing against US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Workloads Face Cost and Reliability Tests

Artificial Analysis’ index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0, according to the source. That gives the overall score relevance for buyers considering agents, but it does not settle whether Large 4 will work well for a specific task, toolset or deployment. The source argues that lower benchmark performance could compound across long sequences of actions: a mistaken intermediate result can shape later steps rather than remain an isolated answer.

Output volume and cost add practical considerations. An agent that uses many tokens at each step can increase both latency and API spending, even when its per-token price seems manageable. Buyers evaluating the model should test completion rates, correction needs, tool-call behavior and total cost on representative workflows. The source’s benchmark and hands-on observations are reasons to test carefully, not substitutes for a controlled evaluation in the intended environment.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Major Jump From Mistral’s Prior Scores

The source places Large 4’s 38.4 score against earlier Mistral results of 9 for Large 3 and 14 for Medium 3.5, using the same Artificial Analysis index version. That points to a marked change in measured capability. The index version matters: comparisons are most useful when models are assessed on the same version and task set, rather than across changing evaluations.

The source characterizes the launch framing—Mistral as home to the strongest model outside the US and China—as technically accurate but dependent on which countries and competitors are included. Its own table shows Large 4 behind several US and Chinese models. It also says Mistral reports reinforcement learning is still underway, so scores may change. Until the weights and licence are available, customers cannot assess the released model as an open-weight option on the terms anticipated by the report.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— ThorstenMeyerAI.com report author

Amazon

multimodal AI models for agents

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Real-World Results

Several practical questions remain open. The source says model weights are promised for the end of October but does not provide a year, and it says the licence is unpublished. Mistral’s API preview also means buyers are assessing a product whose capabilities could shift while reinforcement learning continues. The report does not specify how much scores might change or when a final model will be available.

The source author reports seeing confident false claims during hands-on use, but gives no test protocol, sample size or measured error rate. That observation should not be treated as a general hallucination statistic. Nor do index scores alone establish performance on an individual company’s workflows. The source’s pricing and cost comparisons are tied to its cited evaluation; actual spending can vary with prompt design, caching, task length and the number of retries needed.

Amazon

large language model API pricing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for Weights and Updated Evaluations

The next reported milestone is the release of model weights, which Mistral has promised for the end of October. Their availability and licence terms will clarify whether customers can run or adapt Large 4 outside the API and under what conditions. Mistral may also publish updated results as its reinforcement-learning work continues.

For now, prospective users can assess the API preview against their own agent tasks, comparing successful completion, factual errors, tool use, latency and total token cost with alternatives. The source material does not identify a final-release date, a settled licence or independent evidence that Large 4 has already reached production-level reliability across agent workflows.

Amazon

AI token management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Is Mistral Large 4 available now?

According to the source report, it is available through Mistral’s API as a research preview. Model weights are promised for the end of October, but the report does not specify a year or publish licence terms.

How does Large 4 score against leading models?

It scores 38.4 on Artificial Analysis Intelligence Index v4.3.2. The source’s comparison table puts several US models and several Chinese models above it, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.

Does the benchmark prove Large 4 is unsuitable for agents?

No. The score and the report’s cost and output comparisons are relevant signals, but they do not predict every deployment. Buyers should test the preview on their own tasks and measure reliability, completion rates and total cost.

What remains unknown about its open weights?

The source says weights are expected at the end of October, but the licence has not been published. Until both are available, users cannot confirm the terms for downloading, modifying or deploying the model themselves.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Frontier AI Model Just Went Dark For 18 Days. The Kill-Switch Is Real Now.

A leading AI model was abruptly taken offline for 18 days due to government orders, signaling a new era of national security vetting for AI releases.

Why RAG Evaluation Is Harder Than Most Teams Expect

Probing the true challenges of RAG evaluation reveals hidden biases, inconsistent metrics, and complex human judgments that most teams overlook.

Why Playco Relies On GPT-6 Astra For Manual Fixes In Game Prototyping

Playco reports a 50% reduction in manual fixes during game prototyping with GPT-6 Astra, according to a case study by OpenAI; independent verification pending.

How Generative AI Infrastructure Differs From Traditional ML Stacks

Narrowing the gap between traditional ML and generative AI infrastructure reveals key differences that are essential for building advanced, scalable AI systems.