📊 Full opportunity report: Qwen3.8-Max's AI Results: Closer Than Ever To The Top? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba officially released Qwen3.8-Max, a 2.4 trillion-parameter AI model with top benchmark scores, and confirmed open weights will be available soon. The model shows promising performance, especially in multimodal and agentic tasks, though some limitations remain.
Alibaba has officially released the benchmark results for Qwen3.8-Max, a 2.4 trillion-parameter AI model that is now positioned among the top performers globally. The company confirmed that open weights will be available next week, marking a significant step in accessible large-scale AI models, and provided detailed performance metrics that substantiate its claims of high capability.
Alibaba’s Qwen3.8-Max features a sparse mixture-of-experts architecture built on the Qwen3.5 foundation, with approximately 95 billion active parameters per query. The model demonstrates impressive results on several benchmark tests, including a Terminal-Bench score of 86.6, surpassing models like Claude Opus 4.8 and Fable 5, and approaching GPT-5.6 Sol at the top end. It also leads in multimodal and agentic benchmarks, such as OSWorld-Verified and Parametric CAD Bench.
Alibaba also showcased the model’s ability to reproduce research results and outperform previous iterations in long-horizon agent tasks, notably improving DeepSWE scores from 21.6 to 56.6. The company confirmed that the full benchmark table was withheld until now but is now publicly available, providing transparency about the model’s strengths and limitations. The open weights, expected next week, will include a 2.4 trillion-parameter checkpoint, though practical deployment will likely rely on the 95B active parameter configuration due to hardware constraints.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Qwen3.8-Max’s Benchmark Results
The release of Qwen3.8-Max and its benchmark results mark a major milestone in AI development, demonstrating that large-scale models can achieve top-tier performance while maintaining open access. This could accelerate research and deployment in AI applications, especially in multimodal and agentic tasks. However, the model's limitations in deep software-engineering benchmarks highlight ongoing challenges in scaling AI capabilities across all domains, and the availability of open weights next week will be critical for developers and researchers aiming to leverage this technology.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Development
Alibaba has been gradually building its AI capabilities, with the Qwen series leading the charge. Its previous models, such as Qwen3.7-Max, laid the groundwork for the current release. The company’s strategy involved stealthy previews, culminating in a high-profile reveal during the World AI Conference in Shanghai on July 19. The announcement of Qwen3.8-Max followed a period of speculation, fueled by a mysterious model called kaleb appearing on leaderboards, which was later confirmed as the previewed model. Alibaba’s approach has combined selective disclosures with detailed benchmark releases, aiming to position itself as a top contender in large-language models.
"We are committed to advancing AI capabilities and providing the community with open weights that enable broader innovation."
— Alibaba spokesperson
large scale AI model deployment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Model Deployment and Capabilities
While benchmark results are promising, it is still unclear how Qwen3.8-Max will perform in real-world applications outside testing environments, especially in complex software engineering tasks where it trails behind Fable 5. The impact of compression on agentic performance in the open-weight version remains untested, and the licensing terms for the open weights are yet to be announced, which could influence adoption.
As an affiliate, we earn on qualifying purchases.
Upcoming Availability and Practical Use of Open Weights
The open weights for Qwen3.8-Max are scheduled for release next week, allowing developers to test and deploy the model on their own hardware. The community will closely evaluate how well the model's agentic and multimodal capabilities translate outside benchmark settings. Additionally, Alibaba may release further details on licensing terms, which will influence how broadly the model can be used and integrated into various applications.
AI research and development software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the main strengths of Qwen3.8-Max according to Alibaba?
Qwen3.8-Max excels in multimodal understanding, agentic reasoning, and benchmark performance, especially in tasks like research reproduction and long-horizon planning.
When will the open weights for Qwen3.8-Max be available?
Alibaba has announced that the open weights will be released next week, enabling broader access for developers and researchers.
How does Qwen3.8-Max compare to other leading models?
It outperforms models like Claude Opus 4.8 and Fable 5 on several benchmarks, and approaches GPT-5.6 Sol in top scores, though it still trails in some software-engineering benchmarks.
What limitations does Qwen3.8-Max have?
Despite strong benchmark results, it underperforms in deep software engineering tasks and its agentic capabilities may be affected by compression in the open-weight version.
Will the licensing terms be open source?
The licensing terms for the open weights have not yet been disclosed, which will influence how the model can be used in commercial and research settings.
Source: ThorstenMeyerAI.com