🔍 Read the full analysis: Should You Choose Mistral Large 4 For Agents? Its Global Strengths And Limits on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral has released Large 4 as a research preview, scoring 38.4 on Artificial Analysis Intelligence Index v4.3.2. The score marks a substantial gain over earlier Mistral models, but the source’s benchmark comparisons show leading US and several Chinese models scoring higher, while its reported task cost and output volume raise concerns for agent use.
Mistral has released Large 4, a 1-trillion-parameter multimodal model available as an API research preview, with weights promised for the end of October. On Artificial Analysis Intelligence Index v4.3.2, it scores 38.4, a large jump from earlier Mistral models but below the leading US and several Chinese systems—evidence that its progress does not by itself establish it as a strong choice for long-running agents.
The source report says Large 4 has 49 billion active parameters, accepts text and images, produces text, and supports a 512,000-token context window. It is currently available through Mistral’s API as a research preview. Mistral has said model weights are expected at the end of October; the report says the licence has not been published. API pricing is listed as $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks.
Artificial Analysis gives Large 4 a score of 38.4. The report compares this with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, describing the new result as a substantial improvement. Its table lists several US models above 52 and Chinese models including GLM-5.3 at 44.8, Kimi K3 at 43.6, and DeepSeek V4.1 Flash at 39.5. These are benchmark results, not a guarantee of how a model will perform in a particular customer’s workflow.
The report also cites Artificial Analysis figures putting Large 4’s cost at $1.13 per index task and its total output at 200 million tokens to complete the index. It compares that output with a median of 81 million tokens for comparable models. Its account says GLM-5.3-Flash costs $0.25 per task and scores 41.8, while DeepSeek V4.1 Flash costs $0.27 and scores 39.5. Those figures suggest buyers should compare completed-work costs, not just advertised token rates.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
Agent Workloads Face Cost and Reliability Tests
Artificial Analysis’ index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0, according to the source. That gives the overall score relevance for buyers considering agents, but it does not settle whether Large 4 will work well for a specific task, toolset or deployment. The source argues that lower benchmark performance could compound across long sequences of actions: a mistaken intermediate result can shape later steps rather than remain an isolated answer.
Output volume and cost add practical considerations. An agent that uses many tokens at each step can increase both latency and API spending, even when its per-token price seems manageable. Buyers evaluating the model should test completion rates, correction needs, tool-call behavior and total cost on representative workflows. The source’s benchmark and hands-on observations are reasons to test carefully, not substitutes for a controlled evaluation in the intended environment.
As an affiliate, we earn on qualifying purchases.
A Major Jump From Mistral’s Prior Scores
The source places Large 4’s 38.4 score against earlier Mistral results of 9 for Large 3 and 14 for Medium 3.5, using the same Artificial Analysis index version. That points to a marked change in measured capability. The index version matters: comparisons are most useful when models are assessed on the same version and task set, rather than across changing evaluations.
The source characterizes the launch framing—Mistral as home to the strongest model outside the US and China—as technically accurate but dependent on which countries and competitors are included. Its own table shows Large 4 behind several US and Chinese models. It also says Mistral reports reinforcement learning is still underway, so scores may change. Until the weights and licence are available, customers cannot assess the released model as an open-weight option on the terms anticipated by the report.
“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”
— ThorstenMeyerAI.com report author
As an affiliate, we earn on qualifying purchases.
Weights, Licensing and Real-World Results
Several practical questions remain open. The source says model weights are promised for the end of October but does not provide a year, and it says the licence is unpublished. Mistral’s API preview also means buyers are assessing a product whose capabilities could shift while reinforcement learning continues. The report does not specify how much scores might change or when a final model will be available.
The source author reports seeing confident false claims during hands-on use, but gives no test protocol, sample size or measured error rate. That observation should not be treated as a general hallucination statistic. Nor do index scores alone establish performance on an individual company’s workflows. The source’s pricing and cost comparisons are tied to its cited evaluation; actual spending can vary with prompt design, caching, task length and the number of retries needed.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Updated Evaluations
The next reported milestone is the release of model weights, which Mistral has promised for the end of October. Their availability and licence terms will clarify whether customers can run or adapt Large 4 outside the API and under what conditions. Mistral may also publish updated results as its reinforcement-learning work continues.
For now, prospective users can assess the API preview against their own agent tasks, comparing successful completion, factual errors, tool use, latency and total token cost with alternatives. The source material does not identify a final-release date, a settled licence or independent evidence that Large 4 has already reached production-level reliability across agent workflows.
As an affiliate, we earn on qualifying purchases.
Key Questions
Is Mistral Large 4 available now?
According to the source report, it is available through Mistral’s API as a research preview. Model weights are promised for the end of October, but the report does not specify a year or publish licence terms.
How does Large 4 score against leading models?
It scores 38.4 on Artificial Analysis Intelligence Index v4.3.2. The source’s comparison table puts several US models and several Chinese models above it, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.
Does the benchmark prove Large 4 is unsuitable for agents?
No. The score and the report’s cost and output comparisons are relevant signals, but they do not predict every deployment. Buyers should test the preview on their own tasks and measure reliability, completion rates and total cost.
What remains unknown about its open weights?
The source says weights are expected at the end of October, but the licence has not been published. Until both are available, users cannot confirm the terms for downloading, modifying or deploying the model themselves.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
