📊 Full opportunity report: Qwen3.8-Max's AI Results: Closer Than Ever To The Top? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba officially released Qwen3.8-Max, a 2.4 trillion-parameter AI model with top benchmark scores, and confirmed open weights will be available soon. The model shows promising performance, especially in multimodal and agentic tasks, though some limitations remain.

Alibaba has officially released the benchmark results for Qwen3.8-Max, a 2.4 trillion-parameter AI model that is now positioned among the top performers globally. The company confirmed that open weights will be available next week, marking a significant step in accessible large-scale AI models, and provided detailed performance metrics that substantiate its claims of high capability.

Alibaba’s Qwen3.8-Max features a sparse mixture-of-experts architecture built on the Qwen3.5 foundation, with approximately 95 billion active parameters per query. The model demonstrates impressive results on several benchmark tests, including a Terminal-Bench score of 86.6, surpassing models like Claude Opus 4.8 and Fable 5, and approaching GPT-5.6 Sol at the top end. It also leads in multimodal and agentic benchmarks, such as OSWorld-Verified and Parametric CAD Bench.

Alibaba also showcased the model’s ability to reproduce research results and outperform previous iterations in long-horizon agent tasks, notably improving DeepSWE scores from 21.6 to 56.6. The company confirmed that the full benchmark table was withheld until now but is now publicly available, providing transparency about the model’s strengths and limitations. The open weights, expected next week, will include a 2.4 trillion-parameter checkpoint, though practical deployment will likely rely on the 95B active parameter configuration due to hardware constraints.

At a glance
updateWhen: announced August 3, 2023, with open wei…
The developmentAlibaba announced the full benchmark results of its Qwen3.8-Max model, confirming it as one of the largest and most capable open-weight AI models to date.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Qwen3.8-Max’s Benchmark Results

The release of Qwen3.8-Max and its benchmark results mark a major milestone in AI development, demonstrating that large-scale models can achieve top-tier performance while maintaining open access. This could accelerate research and deployment in AI applications, especially in multimodal and agentic tasks. However, the model's limitations in deep software-engineering benchmarks highlight ongoing challenges in scaling AI capabilities across all domains, and the availability of open weights next week will be critical for developers and researchers aiming to leverage this technology.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Development

Alibaba has been gradually building its AI capabilities, with the Qwen series leading the charge. Its previous models, such as Qwen3.7-Max, laid the groundwork for the current release. The company’s strategy involved stealthy previews, culminating in a high-profile reveal during the World AI Conference in Shanghai on July 19. The announcement of Qwen3.8-Max followed a period of speculation, fueled by a mysterious model called kaleb appearing on leaderboards, which was later confirmed as the previewed model. Alibaba’s approach has combined selective disclosures with detailed benchmark releases, aiming to position itself as a top contender in large-language models.

"We are committed to advancing AI capabilities and providing the community with open weights that enable broader innovation."

— Alibaba spokesperson

Amazon

large scale AI model deployment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Deployment and Capabilities

While benchmark results are promising, it is still unclear how Qwen3.8-Max will perform in real-world applications outside testing environments, especially in complex software engineering tasks where it trails behind Fable 5. The impact of compression on agentic performance in the open-weight version remains untested, and the licensing terms for the open weights are yet to be announced, which could influence adoption.

Amazon

multimodal AI development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Availability and Practical Use of Open Weights

The open weights for Qwen3.8-Max are scheduled for release next week, allowing developers to test and deploy the model on their own hardware. The community will closely evaluate how well the model's agentic and multimodal capabilities translate outside benchmark settings. Additionally, Alibaba may release further details on licensing terms, which will influence how broadly the model can be used and integrated into various applications.

Amazon

AI research and development software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main strengths of Qwen3.8-Max according to Alibaba?

Qwen3.8-Max excels in multimodal understanding, agentic reasoning, and benchmark performance, especially in tasks like research reproduction and long-horizon planning.

When will the open weights for Qwen3.8-Max be available?

Alibaba has announced that the open weights will be released next week, enabling broader access for developers and researchers.

How does Qwen3.8-Max compare to other leading models?

It outperforms models like Claude Opus 4.8 and Fable 5 on several benchmarks, and approaches GPT-5.6 Sol in top scores, though it still trails in some software-engineering benchmarks.

What limitations does Qwen3.8-Max have?

Despite strong benchmark results, it underperforms in deep software engineering tasks and its agentic capabilities may be affected by compression in the open-weight version.

Will the licensing terms be open source?

The licensing terms for the open weights have not yet been disclosed, which will influence how the model can be used in commercial and research settings.

Source: ThorstenMeyerAI.com

You May Also Like

How AI Is Reshaping Urban Oversight And Civic Trust

Exploring how artificial intelligence and digital twins reshape city governance, privacy, and public trust amid evolving urban surveillance.

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Analysis of Mistral’s strategic shift to full-stack AI, its enterprise focus, and the debate over its future in the AI industry.

OpenEuroLLM. The third path.

European consortium OpenEuroLLM faces compute resource challenges amid progress, highlighting limits of pan-European AI development efforts.

13 Best Guides To AI-Powered Marketing Automation Tools For Smarter Campaigns In 2026

Explore the 13 best books and guides on AI marketing automation, helping marketers choose strategies and tools for smarter campaigns.