The Complexity Of Ranking AI: Qwen3.8-Max's Recent Performance Revealed

📊 Full opportunity report: The Complexity Of Ranking AI: Qwen3.8-Max's Recent Performance Revealed on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has officially released details on its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and demonstrating top benchmark performance. The open weights will be available next week, but the model’s true capabilities and limitations remain under evaluation.

Alibaba has confirmed the full specifications and benchmark results for Qwen3.8-Max, its largest-ever AI model with 2.4 trillion parameters. The company announced that open weights will be released next week, marking a significant step in making high-parameter models more accessible, though the model’s full capabilities and licensing details remain under wraps.

On August 3, Alibaba made Qwen3.8-Max broadly available, revealing it has 2.4 trillion parameters built on a sparse mixture-of-experts architecture. The model is multimodal, supporting text, image, and video inputs, with a focus on text output. Benchmark results on Alibaba’s internal tests show it outperforms many competitors on key tasks, including a top score of 93.0 on PaperBench and strong results on multimodal and agentic benchmarks. The company also announced a second checkpoint, Qwen3.8-27B, optimized for deployment on individual high-memory machines, which will be available next week.

The model’s active parameters are approximately 95 billion, with the full 2.4 trillion parameters representing a sparse network that activates only a small subset during inference. The release strategy included a preview endpoint at discounted pricing, generating significant media and market attention, with Alibaba’s shares rising up to 5.4 percent following the announcement.

At a glance
updateWhen: announced August 3, 2023; details unfol…
The developmentAlibaba announced the broad availability of Qwen3.8-Max, revealing its benchmark performance and confirming the open weights will be released next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's High-Parameter AI Model Release

This development marks a notable milestone in AI model scaling, demonstrating Alibaba’s capability to build and benchmark models with trillions of parameters. The release of the open weights next week could influence the landscape of accessible large models, especially with the availability of the smaller, deployable 27B checkpoint. However, the performance gaps on some benchmarks reveal limits in current scaling techniques, particularly in deep software engineering tasks. The move also underscores ongoing debates about transparency, licensing, and the practical deployment of enormous models in real-world applications.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's AI Model Strategy

Alibaba has been gradually revealing its large language model capabilities since July, with the initial preview of Qwen3.8-Max surfacing through an anonymous community leak and official confirmation at the World AI Conference in Shanghai. Prior to this, the company’s models, such as Kimi K3 and smaller Qwen variants, had garnered attention for their performance and scaling efforts. The recent announcement follows a pattern of strategic disclosures designed to generate market interest and demonstrate technical leadership, culminating in the full benchmark reveal and upcoming open-weight release.

"Next week, we will release the open weights for Qwen3.8-27B, enabling broader research and deployment on high-memory hardware."

— Alibaba spokesperson

Amazon

high memory GPU for AI development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Capabilities and Licensing

It remains unclear what the licensing terms will be for the open weights, as Alibaba has not yet published a license. The practical usability of the 2.4 trillion-parameter model outside Alibaba’s infrastructure is also uncertain, given the immense hardware requirements. Additionally, the performance on some benchmarks, particularly software engineering tasks, shows notable gaps, raising questions about the model's generalization and real-world applicability.

Amazon

multimodal AI development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s Model Deployment and Community Testing

Next week, Alibaba plans to release the open weights of Qwen3.8-27B, which will allow researchers and developers to test the model on their own hardware. The company is expected to publish licensing details and further benchmarks, providing clarity on the model’s deployment potential. Monitoring how the community adopts and adapts the open weights will be critical to understanding the model’s impact and limitations.

Amazon

large scale AI model deployment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main specifications of Alibaba's Qwen3.8-Max?

Qwen3.8-Max has 2.4 trillion parameters, built on a sparse mixture-of-experts architecture, with around 95 billion active parameters per query. It supports multimodal inputs and has demonstrated top benchmark performance in Alibaba's internal tests.

When will the open weights for Qwen3.8-27B be available?

Alibaba has announced that the open weights for the Qwen3.8-27B checkpoint will be released next week, enabling broader access for deployment on high-memory hardware.

What benchmarks did Qwen3.8-Max excel in?

The model scored 93.0 on PaperBench, outperformed many competitors on multimodal and agentic benchmarks, and demonstrated strong long-horizon reasoning capabilities. However, it lagged on some deep software engineering benchmarks.

What are the licensing implications for the open weights?

Alibaba has not yet published licensing details for the open weights, leaving uncertainty about usage rights and restrictions once they are released.

How does Qwen3.8-Max compare to other large models like GPT-5.6?

In Alibaba’s internal benchmarks, Qwen3.8-Max performs well, but GPT-5.6 still leads on some measures, especially in deep reasoning tasks. The model's true competitive edge may depend on future deployment and community adaptation.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

As Bitcoin’s Supply in Profit Drops, Are We Near a Market Bottom?

Potential market shifts arise as Bitcoin’s supply in profit dwindles, leaving investors to ponder whether a buying opportunity or further corrections lie ahead.

Every Benchmark Launched 2023-2024 Has Fallen — The METR / SWE-Bench / CORE-Bench / MLE-Bench / PostTrainBench Sequence

Every major benchmark launched in 2023-2024 measuring AI research capability has reached saturation or is nearing it, signaling rapid progress in AI development.

The Sec’s Latest Move Includes a New Unit Focused on Fighting Blockchain Fraud.

Blockchain fraud is facing a formidable challenge with the SEC’s new unit, but what innovative strategies will they implement to protect investors?

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers release a detailed framework outlining pathways from human-level AI to superintelligence, emphasizing scalability and potential hurdles.