Qwen3.8-Max Reveals Its AI Performance: How Does It Stand Up To Fable 5?

📊 Full opportunity report: Qwen3.8-Max Reveals Its AI Performance: How Does It Stand Up To Fable 5? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has officially released detailed benchmark data for its Qwen3.8-Max model, confirming it as the second-largest open-weight AI model. It outperforms several competitors in key benchmarks but lags behind Fable 5 in some software engineering tasks. The release highlights Alibaba’s advances in multimodal and agentic AI capabilities.

Alibaba has officially published comprehensive benchmark data for its Qwen3.8-Max model, confirming it as the second-largest open-weight AI model with 2.4 trillion parameters. This release highlights Alibaba’s advances in multimodal and agentic AI capabilities. This marks a significant milestone in AI development, as the model demonstrates strong performance across multiple benchmarks, especially in multimodal and agentic tasks, and will be available for open-weight download next week.

On August 3, Alibaba revealed the full benchmark table for Qwen3.8-Max, a model built on the Qwen3.5 architecture with 2.4 trillion parameters, of which approximately 95 billion are active per query. The model uses sparse mixture-of-experts technology and supports multimodal inputs—text, images, video—with text output. The benchmarks show it surpasses some competitors, such as Claude Opus 4.8 and Claude Fable 5, on Terminal-Bench 2.1, PaperBench, and other specialized tests, but trails behind GPT-5.6 Sol in overall maximum effort scores. For more context on recent AI model benchmarks, see this analysis.

Alibaba also demonstrated the model’s capabilities in reproducing research results and outperforming its own previous models in long-horizon agentic tasks, notably improving from unusable to competitive levels in RL-environment scaling benchmarks. Learn more about how these models are shaping the future in this article. The open weights are set to be released next week, with a smaller 27B checkpoint designed for local deployment, which is expected to perform well in inference tasks.

At a glance
updateWhen: announced August 3, 2023; benchmark dat…
The developmentAlibaba announced the full benchmark results for its Qwen3.8-Max model, confirming its performance and open-weight availability, positioning it as a major player in AI development.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Benchmark Release for AI Competition

The full disclosure of benchmark results confirms Alibaba's position as a major AI player, especially in multimodal and agentic tasks. The release of open weights for such a large model signals increased accessibility and potential for innovation in AI development. However, the model's performance gaps in software engineering benchmarks highlight ongoing challenges in achieving comprehensive AI capabilities. This development may influence industry standards, open-source contributions, and competitive dynamics in AI research.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Launches and Benchmarking

Over the past two weeks, Alibaba's AI models have generated significant attention through stealth previews and strategic announcements. On July 17, the Kimi K3 model with 2.8 trillion parameters was launched, briefly impacting US tech stocks. The following day, an anonymous model named 'kaleb' was revealed to be Alibaba’s Qwen3.8-Max, initially introduced at the World AI Conference in Shanghai. Prior to today’s full benchmark release, Alibaba had withheld detailed data, only hinting at its competitive position relative to Fable 5. The recent disclosure confirms the model’s specifications, capabilities, and open-weight plans, marking a major step in Alibaba’s AI strategy.

"Alibaba’s full benchmark disclosure confirms its model as a major contender, especially in multimodal and agentic tasks, but performance gaps remain in software engineering benchmarks."

— Thorsten Meyer, ThorstenMeyerAI.com

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

  • High Performance Dual-Core CPU: Equipped with dual-core processor and Type-C USB
  • Rich Peripheral Support: Includes SPI, LCD, Camera, UART, I2C, and more
  • Wireless Connectivity: Supports Wi-Fi 2.4 GHz and Bluetooth 5 (LE)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Licensing and Deployment

Details about the licensing terms for the 2.4 trillion-parameter weights remain unpublished, raising questions about usage rights and restrictions. It is also unclear whether the open weights will be fully open-source under permissive licenses or have revenue and attribution conditions. Additionally, the performance of the 27B checkpoint in real-world deployment and its agentic capabilities after compression are still unverified.

Amazon

AI inference acceleration hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s AI Model Rollout and Industry Impact

Alibaba will release the open weights next week, enabling researchers and developers to evaluate the model’s performance firsthand. The company is expected to publish licensing details and further benchmarks, especially for the 27B checkpoint. Industry observers will closely monitor how the model’s agentic and multimodal capabilities evolve in practical applications, and whether Alibaba’s claims influence broader AI development and competition.

Amazon

large language model deployment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main strengths of Qwen3.8-Max?

Qwen3.8-Max excels in multimodal tasks, agentic reasoning, and long-horizon research reproduction, outperforming some competitors in specific benchmarks like Terminal-Bench and Parametric CAD Bench.

How does Qwen3.8-Max compare to Fable 5?

In benchmarks such as Terminal-Bench 2.1 and PaperBench, Qwen3.8-Max performs close to or slightly better than Fable 5. However, it trails significantly in deep software-engineering benchmarks like SWE-bench Pro and FrontierSWE, where Fable 5 scores much higher.

Will the open weights be usable for individual developers?

The 2.4 trillion-parameter weights are a multi-node datacenter artifact, making them impractical for individual use. The smaller 27B checkpoint is designed for local deployment on high-memory machines and is expected to be more accessible for developers.

What licensing model will Alibaba use for the open weights?

As of now, licensing details remain unpublished. Historically, Alibaba’s open models have used Apache 2.0, but the upcoming release may include different conditions, especially given the large size of the checkpoint and potential revenue considerations.

What are the implications of this release for AI competition?

This release positions Alibaba as a serious contender in large-scale AI, especially in multimodal and agentic domains. It may accelerate industry-wide efforts toward openness and push competitors to disclose more detailed benchmark data.

Source: ThorstenMeyerAI.com

You May Also Like

Understanding AI’s Role in NATO: Friend or Foe at the Alliance Level

Analyzing the risks of Chinese technology within NATO’s infrastructure and the implications for alliance security and AI’s influence.

AMÁLIA · The Three Hard Questions.

Portugal’s €5.5M AMÁLIA LLM, launched in 2025, outperforms many models in Portuguese tasks but prompts key questions about openness, native data, and goals.

QAtrial: Compliance That Shows Its Work

QAtrial introduces an open-source platform ensuring AI assistance in life sciences complies with regulatory traceability and audit requirements.

Forezai · Polybot: When the AI Disagrees With the Odds

Polybot, an open-source AI trading experiment, tests when and how an AI can legitimately diverge from prediction market prices, highlighting risks and insights.