📊 Full opportunity report: AI Performance Debates: Qwen3.8-Max’s Numbers Under The Microscope on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, revealing its benchmark scores and confirming its 2.4 trillion parameters. The model’s performance and openness are now under detailed scrutiny, raising questions about its actual capabilities and deployment potential.

Alibaba has officially published detailed benchmark scores for its Qwen3.8-Max model, confirming it as a 2.4 trillion-parameter, multimodal AI system. The company announced the model’s availability alongside a smaller, open-weight variant, Qwen3.8-27B, which is designed for deployment on individual hardware. This development marks a significant step in transparency for one of the largest AI models to date, with performance data now publicly accessible for the first time.

Alibaba’s Qwen3.8-Max, previously known only as a stealth preview, has been confirmed to contain approximately 2.4 trillion parameters with about 95 billion active parameters per query, using a sparse mixture-of-experts architecture built on the Qwen3.5 foundation. The model is multimodal, capable of processing text, images, and videos, and generating text output. Its benchmark scores, achieved on Alibaba’s own testing harness, include a top score of 93.0 on PaperBench and 86.6 on Terminal-Bench 2.1, surpassing several competitors such as Claude Opus 4.8 and Fable 5, and only trailing GPT-5.6 Sol at 88.8.

While the model demonstrates strong performance on multimodal and agentic benchmarks, it underperforms significantly on deep software engineering tasks like SWE-bench Pro, where it scores 67.7 against Fable 5’s 80.0. Notably, the model has shown substantial improvement over its predecessor, DeepSWE, jumping from 21.6 to 56.6 in agentic execution scores, indicating a meaningful advancement in autonomous reasoning capabilities.

At a glance
reportWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba officially released benchmark data and confirmed the open weights for Qwen3.8-Max, marking a significant milestone in large-scale AI models.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Benchmark Release and Open Weights

This release confirms Alibaba’s position as a major player in large-scale AI with the largest open-weight model publicly available, setting new benchmarks for transparency and performance. The detailed scores provide clarity on the model’s strengths in multimodal and agentic tasks, but also reveal persistent gaps in software engineering benchmarks, highlighting the ongoing challenges in scaling AI reasoning abilities. The open-weight release, scheduled for next week, could influence deployment practices and competitive dynamics in AI development, especially for organizations aiming to run large models on local hardware.

Compiler Engineering for AI Hardware: MLIR, TVM, XLA, and Custom Backends for Neural Network Accelerators (AI Infrastructure, Hardware & Compiler Engineering Series)

Compiler Engineering for AI Hardware: MLIR, TVM, XLA, and Custom Backends for Neural Network Accelerators (AI Infrastructure, Hardware & Compiler Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Recent Developments in Alibaba’s AI Strategy

Over the past two weeks, Alibaba’s AI initiatives have been shrouded in secrecy, with the company teasing its latest model through cryptic hints and stealth previews. The model, initially called kaleb, was revealed during the World AI Conference in Shanghai, where Alibaba confirmed it was Qwen3.8-Max. Prior to this, models like Kimi K3 and other large-scale models had stirred market interest, but detailed benchmarks and open weights had not been publicly available until now. Alibaba’s approach has combined strategic timing—announcing on a Sunday, releasing detailed data two weeks later—and selective disclosure, emphasizing multimodal and agentic capabilities.

"Qwen3.8-Max sets a new standard in multimodal AI, and our open weights will enable broader research and deployment opportunities."

— Alibaba spokesperson

Mastering Large Language Models from First Principles: A Practical Guide to Building Transformers, Attention Mechanisms, Tokenizers, and Intelligent AI Applications

Mastering Large Language Models from First Principles: A Practical Guide to Building Transformers, Attention Mechanisms, Tokenizers, and Intelligent AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Capabilities and Licensing

It remains unclear what the exact licensing terms will be for the 2.4 trillion-parameter weights, and whether they will be fully open-source under permissive licenses like Apache 2.0. Additionally, the performance on certain benchmarks, especially software engineering tasks, indicates significant gaps that could limit practical deployment. The impact of the open weights on real-world applications and how much the agentic improvements will hold up under compression or in diverse environments are still uncertain.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Model Deployment and Benchmark Validation

Alibaba plans to release the 2.4 trillion-parameter open weights next week, allowing organizations to evaluate the model independently. Meanwhile, the community will scrutinize the benchmark scores, test the model’s capabilities across diverse tasks, and assess its licensing terms once officially published. Further updates are expected as Alibaba clarifies licensing details and demonstrates the model’s performance in practical settings, especially on local hardware with the 27B variant.

Autel MaxiSYS Ultra S2 AI Scanner, Intelligent Topology 3, Multi-Point DVI

Autel MaxiSYS Ultra S2 AI Scanner, Intelligent Topology 3, Multi-Point DVI

🔥🔥🔥【2026 Autel Ultra S2 AI Scanner with 2 Years Update, V2.0 of MS919 S2/ MS909 S2】Autel unveil the...

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main capabilities of Alibaba’s Qwen3.8-Max?

It is a multimodal, 2.4 trillion-parameter model capable of processing text, images, and videos, with strong performance in multimodal and agentic benchmarks, but weaker results in deep software engineering tasks.

Will the open weights be available for commercial use?

Alibaba has announced the weights will ship next week, but licensing terms are still unpublished, so it is unclear whether they will be fully open-source or have restrictions.

How does Qwen3.8-Max compare to other large models?

It outperforms several competitors on multimodal and agentic benchmarks, ranking just below GPT-5.6 in some tests, but lags significantly on deep software engineering benchmarks.

What does this mean for AI development and deployment?

This release could accelerate research and deployment of large models on local hardware, but the true capabilities and licensing restrictions will influence its practical impact.

What are the limitations of the current benchmark data?

Benchmarks do not cover all aspects of real-world performance, especially in software engineering and long-horizon reasoning, and some results may vary in deployment scenarios.

Source: ThorstenMeyerAI.com

You May Also Like

Meta Launches Muse Spark 1.2 To Lead The AI Coding Revolution

Meta releases Muse Spark 1.2 and Muse Code, advancing AI coding tools with co-training, long-horizon capabilities, and improved safety features amid competitive benchmarks.

Cricket diplomacy is helping India and Pakistan play on a sticky wicket

India and Pakistan are set to play in the upcoming T20 World Cup match in Colombo, highlighting cricket diplomacy’s role in easing political tensions.

How 3D Printing Could Let You “Print” Your Own Gadgets at Home

How 3D printing could let you “print” your own gadgets at home and revolutionize your DIY projects—discover the possibilities that await you.

The gigawatt gap. Why China is structurally positioned for AI power and the US is engineering around its grid.

China leverages centralized planning and renewable infrastructure to close the gigawatt gap in AI deployment, challenging US dominance at the power layer.