📊 Full opportunity report: AI Performance Debates: Qwen3.8-Max’s Numbers Under The Microscope on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba announced the broad availability of Qwen3.8-Max, revealing its benchmark scores and confirming its 2.4 trillion parameters. The model’s performance and openness are now under detailed scrutiny, raising questions about its actual capabilities and deployment potential.
Alibaba has officially published detailed benchmark scores for its Qwen3.8-Max model, confirming it as a 2.4 trillion-parameter, multimodal AI system. The company announced the model’s availability alongside a smaller, open-weight variant, Qwen3.8-27B, which is designed for deployment on individual hardware. This development marks a significant step in transparency for one of the largest AI models to date, with performance data now publicly accessible for the first time.
Alibaba’s Qwen3.8-Max, previously known only as a stealth preview, has been confirmed to contain approximately 2.4 trillion parameters with about 95 billion active parameters per query, using a sparse mixture-of-experts architecture built on the Qwen3.5 foundation. The model is multimodal, capable of processing text, images, and videos, and generating text output. Its benchmark scores, achieved on Alibaba’s own testing harness, include a top score of 93.0 on PaperBench and 86.6 on Terminal-Bench 2.1, surpassing several competitors such as Claude Opus 4.8 and Fable 5, and only trailing GPT-5.6 Sol at 88.8.
While the model demonstrates strong performance on multimodal and agentic benchmarks, it underperforms significantly on deep software engineering tasks like SWE-bench Pro, where it scores 67.7 against Fable 5’s 80.0. Notably, the model has shown substantial improvement over its predecessor, DeepSWE, jumping from 21.6 to 56.6 in agentic execution scores, indicating a meaningful advancement in autonomous reasoning capabilities.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba’s Benchmark Release and Open Weights
This release confirms Alibaba’s position as a major player in large-scale AI with the largest open-weight model publicly available, setting new benchmarks for transparency and performance. The detailed scores provide clarity on the model’s strengths in multimodal and agentic tasks, but also reveal persistent gaps in software engineering benchmarks, highlighting the ongoing challenges in scaling AI reasoning abilities. The open-weight release, scheduled for next week, could influence deployment practices and competitive dynamics in AI development, especially for organizations aiming to run large models on local hardware.

Compiler Engineering for AI Hardware: MLIR, TVM, XLA, and Custom Backends for Neural Network Accelerators (AI Infrastructure, Hardware & Compiler Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Recent Developments in Alibaba’s AI Strategy
Over the past two weeks, Alibaba’s AI initiatives have been shrouded in secrecy, with the company teasing its latest model through cryptic hints and stealth previews. The model, initially called kaleb, was revealed during the World AI Conference in Shanghai, where Alibaba confirmed it was Qwen3.8-Max. Prior to this, models like Kimi K3 and other large-scale models had stirred market interest, but detailed benchmarks and open weights had not been publicly available until now. Alibaba’s approach has combined strategic timing—announcing on a Sunday, releasing detailed data two weeks later—and selective disclosure, emphasizing multimodal and agentic capabilities.
"Qwen3.8-Max sets a new standard in multimodal AI, and our open weights will enable broader research and deployment opportunities."
— Alibaba spokesperson

Mastering Large Language Models from First Principles: A Practical Guide to Building Transformers, Attention Mechanisms, Tokenizers, and Intelligent AI Applications
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Model Capabilities and Licensing
It remains unclear what the exact licensing terms will be for the 2.4 trillion-parameter weights, and whether they will be fully open-source under permissive licenses like Apache 2.0. Additionally, the performance on certain benchmarks, especially software engineering tasks, indicates significant gaps that could limit practical deployment. The impact of the open weights on real-world applications and how much the agentic improvements will hold up under compression or in diverse environments are still uncertain.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Model Deployment and Benchmark Validation
Alibaba plans to release the 2.4 trillion-parameter open weights next week, allowing organizations to evaluate the model independently. Meanwhile, the community will scrutinize the benchmark scores, test the model’s capabilities across diverse tasks, and assess its licensing terms once officially published. Further updates are expected as Alibaba clarifies licensing details and demonstrates the model’s performance in practical settings, especially on local hardware with the 27B variant.

Autel MaxiSYS Ultra S2 AI Scanner, Intelligent Topology 3, Multi-Point DVI
🔥🔥🔥【2026 Autel Ultra S2 AI Scanner with 2 Years Update, V2.0 of MS919 S2/ MS909 S2】Autel unveil the...
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the main capabilities of Alibaba’s Qwen3.8-Max?
It is a multimodal, 2.4 trillion-parameter model capable of processing text, images, and videos, with strong performance in multimodal and agentic benchmarks, but weaker results in deep software engineering tasks.
Will the open weights be available for commercial use?
Alibaba has announced the weights will ship next week, but licensing terms are still unpublished, so it is unclear whether they will be fully open-source or have restrictions.
How does Qwen3.8-Max compare to other large models?
It outperforms several competitors on multimodal and agentic benchmarks, ranking just below GPT-5.6 in some tests, but lags significantly on deep software engineering benchmarks.
What does this mean for AI development and deployment?
This release could accelerate research and deployment of large models on local hardware, but the true capabilities and licensing restrictions will influence its practical impact.
What are the limitations of the current benchmark data?
Benchmarks do not cover all aspects of real-world performance, especially in software engineering and long-horizon reasoning, and some results may vary in deployment scenarios.
Source: ThorstenMeyerAI.com