🔍 Read the full analysis: Mistral Large 4 Remains A Step Behind The AI Frontier on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral released an API preview of Large 4 on October 6, 2026, but the available benchmark snapshot places it below leading U.S. models and two stronger Chinese models. A reviewer also reported hallucinations in personal use, though that experience is not a controlled comparison. The model’s weights are scheduled for release later in October.
Mistral Large 4 entered public preview through an API on October 6, but an October 7 snapshot from Artificial Analysis gives it an Intelligence Index score of 38—below leading U.S. models and two higher-scoring Chinese alternatives. The release is a step forward for Mistral’s model lineup, but the results do not establish it as a leading choice for demanding, long-running AI agent tasks.
The preview is Mistral’s largest model to date, built as a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. Mistral said it trained the model on its own infrastructure in Europe and is continuing to improve it. The company’s weights are scheduled for release later in October; they were not publicly downloadable at the time of the October 7 report.
On the cited Intelligence Index snapshot, Large 4 Preview scored 38 points. That matched OpenAI’s GPT-6 Luna at maximum reasoning effort, fell one point below DeepSeek V4.1 Flash at maximum effort, and trailed Z.ai’s GLM-5.3 (45) and Moonshot AI’s Kimi K3 (44). The leading listed U.S. results were Anthropic’s Claude Opus 5.5 at 58, Google’s Gemini 4 Argon at 53, and OpenAI’s GPT-6.1 Sol at 52.
The scores are index points, not percentages or direct predictions of success on a particular task. The comparison also uses named reasoning settings that do not represent identical compute budgets. The source report’s author said personal use of the preview produced hallucinations and led them to avoid it for demanding agentic work. That account is an individual observation, not a controlled study of hallucination rates across models.
What the Benchmark Gap Means
For developers choosing a model for multi-step work, the preview’s score is a reason to test alternatives before assigning it complex autonomous tasks. Agentic systems may plan, call tools, interpret results and make later decisions based on earlier output. An unsupported assumption early in that chain can undermine the final result, even if the response reads smoothly. The source report’s author argues that this kind of work calls for reliable constraint-following, evidence checks and verification.
The benchmark does not settle how Large 4 will perform on every coding, research or business workflow. Nor does a higher aggregate score guarantee better results for an individual developer. But the gap from several listed models makes it harder to justify choosing the preview on the basis of frontier-level general performance alone. Mistral’s claims about agentic coding and specialized professional tasks need workload-specific evidence to show where it is competitive.
The result matters beyond a single model choice because Mistral is a prominent European AI developer. Training a large model on European infrastructure is a development in regional AI capacity; it is separate from whether this preview matches the strongest available systems. The benchmark comparison also has limits: it places Mistral ahead of Canada’s Cohere Command A+ score of 13, so the evidence does not support saying every competing lab scored higher.
As an affiliate, we earn on qualifying purchases.
Preview Now, Weights Later
Mistral announced Large 4 on October 6, 2026 as an API preview, not as a completed public-weight release. According to the company, it is still improving the model, with the weights scheduled to follow later in October. Any assessment based on the preview should be treated as a dated snapshot rather than a final verdict on the eventual release.
The source report cited Artificial Analysis for both the Intelligence Index comparison and a roughly 512,000-token context capacity. A large context window describes how much material can be supplied; it does not by itself show how accurately a model reasons across that material. The report also says DeepSeek V4.1 Flash has approximately comparable benchmark intelligence at a much lower measured cost per task, but the supplied material does not include the cost figures or details of the measurement.
Developer locations in the comparison refer to where the organizations are based, not necessarily where an API request is processed. The reported scores and model settings are specific to the October 7 snapshot and may change as providers update models or evaluations.
“The model was trained on Mistral’s own infrastructure in Europe and is continuing to improve.”
— Mistral, in its announcement as summarized by ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
What the Preview Cannot Yet Show
The available comparison does not show how Large 4 performs across specific customer workloads, how often it makes unsupported claims under controlled conditions, or whether it can match higher-scoring models on specialized tasks. The author’s reported hallucinations came from personal use, and no testing method, sample size or comparative rate is provided. The Intelligence Index is an aggregate benchmark, not a direct measure of reliability on every long-running agent workflow.
The company’s weights had not been released at the time covered, and Mistral said it was still improving the model. It remains unclear whether later versions will materially change the benchmark picture or what their results will be. The supplied source material also gives no full cost table, so the reported cost advantage for DeepSeek V4.1 Flash cannot be independently assessed from these details.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Retests
The next stated milestone is the planned release of Large 4’s weights later in October 2026. Until then, the immediate option described in the source is access through Mistral’s preview API. Mistral says development is continuing, but no specific date for a final version or further benchmark results is provided.
For developers considering the model, the practical next step is to test the available preview on representative tasks, including longer workflows where errors can compound. Future benchmark results, independent evaluations and the public weights—if released as scheduled—could clarify how the model performs and whether its standing changes. The October 7 comparison remains the current dated evidence in the supplied material, not a conclusion about later versions.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral introduced an API preview of Large 4 on October 6, 2026. It is a text-and-image mixture-of-experts model with one trillion total parameters and 49 billion active parameters.
How did Mistral Large 4 score?
Artificial Analysis gave the preview an Intelligence Index score of 38 in the snapshot cited on October 7, 2026. That is an aggregate benchmark score, not a percentage or a guarantee of performance on a particular task.
Is Mistral Large 4 open-weight now?
No. At the time of the report, it was available as an API preview, and its weights were not publicly downloadable. Mistral said the weights were scheduled for release later in October.
Does the report prove Large 4 is unreliable?
No. The source author reported hallucinations in personal use, but described that experience as not a controlled comparative study. The benchmark score also does not establish how the model will perform on every workflow.
When could the assessment change?
The source says Mistral planned to release the model’s weights later in October and was continuing development. New model versions or independent evaluations could change the assessment, but the source gives no confirmed date for further results.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
