📊 Full opportunity report: Qwen3.8-Max Reveals Its AI Performance: How Does It Stand Up To Fable 5? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba has officially released detailed benchmark data for its Qwen3.8-Max model, confirming it as the second-largest open-weight AI model. It outperforms several competitors in key benchmarks but lags behind Fable 5 in some software engineering tasks. The release highlights Alibaba’s advances in multimodal and agentic AI capabilities.
Alibaba has officially published comprehensive benchmark data for its Qwen3.8-Max model, confirming it as the second-largest open-weight AI model with 2.4 trillion parameters. This release highlights Alibaba’s advances in multimodal and agentic AI capabilities. This marks a significant milestone in AI development, as the model demonstrates strong performance across multiple benchmarks, especially in multimodal and agentic tasks, and will be available for open-weight download next week.
On August 3, Alibaba revealed the full benchmark table for Qwen3.8-Max, a model built on the Qwen3.5 architecture with 2.4 trillion parameters, of which approximately 95 billion are active per query. The model uses sparse mixture-of-experts technology and supports multimodal inputs—text, images, video—with text output. The benchmarks show it surpasses some competitors, such as Claude Opus 4.8 and Claude Fable 5, on Terminal-Bench 2.1, PaperBench, and other specialized tests, but trails behind GPT-5.6 Sol in overall maximum effort scores. For more context on recent AI model benchmarks, see this analysis.
Alibaba also demonstrated the model’s capabilities in reproducing research results and outperforming its own previous models in long-horizon agentic tasks, notably improving from unusable to competitive levels in RL-environment scaling benchmarks. Learn more about how these models are shaping the future in this article. The open weights are set to be released next week, with a smaller 27B checkpoint designed for local deployment, which is expected to perform well in inference tasks.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba’s Benchmark Release for AI Competition
The full disclosure of benchmark results confirms Alibaba's position as a major AI player, especially in multimodal and agentic tasks. The release of open weights for such a large model signals increased accessibility and potential for innovation in AI development. However, the model's performance gaps in software engineering benchmarks highlight ongoing challenges in achieving comprehensive AI capabilities. This development may influence industry standards, open-source contributions, and competitive dynamics in AI research.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Launches and Benchmarking
Over the past two weeks, Alibaba's AI models have generated significant attention through stealth previews and strategic announcements. On July 17, the Kimi K3 model with 2.8 trillion parameters was launched, briefly impacting US tech stocks. The following day, an anonymous model named 'kaleb' was revealed to be Alibaba’s Qwen3.8-Max, initially introduced at the World AI Conference in Shanghai. Prior to today’s full benchmark release, Alibaba had withheld detailed data, only hinting at its competitive position relative to Fable 5. The recent disclosure confirms the model’s specifications, capabilities, and open-weight plans, marking a major step in Alibaba’s AI strategy.
"Alibaba’s full benchmark disclosure confirms its model as a major contender, especially in multimodal and agentic tasks, but performance gaps remain in software engineering benchmarks."
— Thorsten Meyer, ThorstenMeyerAI.com

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers
- High Performance Dual-Core CPU: Equipped with dual-core processor and Type-C USB
- Rich Peripheral Support: Includes SPI, LCD, Camera, UART, I2C, and more
- Wireless Connectivity: Supports Wi-Fi 2.4 GHz and Bluetooth 5 (LE)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Model Licensing and Deployment
Details about the licensing terms for the 2.4 trillion-parameter weights remain unpublished, raising questions about usage rights and restrictions. It is also unclear whether the open weights will be fully open-source under permissive licenses or have revenue and attribution conditions. Additionally, the performance of the 27B checkpoint in real-world deployment and its agentic capabilities after compression are still unverified.
AI inference acceleration hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s AI Model Rollout and Industry Impact
Alibaba will release the open weights next week, enabling researchers and developers to evaluate the model’s performance firsthand. The company is expected to publish licensing details and further benchmarks, especially for the 27B checkpoint. Industry observers will closely monitor how the model’s agentic and multimodal capabilities evolve in practical applications, and whether Alibaba’s claims influence broader AI development and competition.
large language model deployment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the main strengths of Qwen3.8-Max?
Qwen3.8-Max excels in multimodal tasks, agentic reasoning, and long-horizon research reproduction, outperforming some competitors in specific benchmarks like Terminal-Bench and Parametric CAD Bench.
How does Qwen3.8-Max compare to Fable 5?
In benchmarks such as Terminal-Bench 2.1 and PaperBench, Qwen3.8-Max performs close to or slightly better than Fable 5. However, it trails significantly in deep software-engineering benchmarks like SWE-bench Pro and FrontierSWE, where Fable 5 scores much higher.
Will the open weights be usable for individual developers?
The 2.4 trillion-parameter weights are a multi-node datacenter artifact, making them impractical for individual use. The smaller 27B checkpoint is designed for local deployment on high-memory machines and is expected to be more accessible for developers.
What licensing model will Alibaba use for the open weights?
As of now, licensing details remain unpublished. Historically, Alibaba’s open models have used Apache 2.0, but the upcoming release may include different conditions, especially given the large size of the checkpoint and potential revenue considerations.
What are the implications of this release for AI competition?
This release positions Alibaba as a serious contender in large-scale AI, especially in multimodal and agentic domains. It may accelerate industry-wide efforts toward openness and push competitors to disclose more detailed benchmark data.
Source: ThorstenMeyerAI.com