Tooling
Alibaba's Qwen3.8-Max Debuts at 2.4 Trillion Parameters, First Qwen This Big to See and Read
Alibaba announced Qwen3.8-Max on August 3, 2026, its largest model yet at 2.4 trillion total parameters on a sparse Mixture-of-Experts design. It is also the first Qwen model above 1 trillion parameters built to handle images, video, and documents alongside text.
Alibaba announced Qwen3.8-Max on August 3, 2026, and it is the biggest model the Qwen team has ever shipped. Total parameter count: 2.4 trillion, spread across a sparse Mixture-of-Experts architecture that activates 95 billion of them on any given token, according to TechNode Global and MLQ News.
The number that matters more than the parameter count is a first. Qwen developer Shuai Bai described it as the team’s first multimodal model above 1 trillion parameters, per MarkTechPost’s July 19 preview coverage from Alibaba’s showing at Shanghai’s World AI Conference. Every prior Qwen model that size handled text only. This one reads documents, watches video, and looks at images in the same request.
TL;DR
Alibaba shipped Qwen3.8-Max on August 3, 2026: 2.4 trillion total parameters, roughly 95 billion active per token through sparse MoE routing, and a 1 million token context window that Alibaba says can hold 200-plus pages of text or around 100 hours of video in a single request. It is the first Qwen above 1 trillion parameters built natively multimodal, taking text, images, video, and documents as input. The API went live immediately on Alibaba Cloud Model Studio at $2 per million input tokens and $6 per million output tokens, with open weights promised for the week of August 10. On launch benchmarks it ranked 5th in Text Arena, 2nd in Vision Arena, 4th in Frontend Code Arena, and posted 86.1 on OSWorld-Verified, ahead of the reported GPT-5.6 Sol Max score of 83.2, according to MarkTechPost and SiliconANGLE. It arrives days after Moonshot AI’s Kimi K3, part of a crowded summer for Chinese open-weight frontier models.
What was announced
Qwen3.8-Max is built on the Qwen 3.5 foundation with a sparse MoE layer and, per TechNode Global’s reporting, a hybrid attention mechanism. The gap between total and active parameters is the whole point of the design: you get a 2.4 trillion parameter model’s knowledge without paying 2.4 trillion parameters of compute on every token. 95 billion active parameters is still a lot, but it is roughly 4% of the total.
Specs confirmed across the launch coverage:
- 2.4 trillion total parameters, up from Qwen3-235B, Alibaba’s previous flagship.
- 95 billion active parameters per forward pass, per TechNode Global and MLQ News.
- 1 million token context window, with a maximum input around 991,000 tokens and maximum output near 131,000 tokens, per MarkTechPost.
- Multimodal input: text, images, video, and documents. Output is text.
- Five built-in tools exposed through the API: code_interpreter, web_search, web_extractor, t2i_search, and i2i_search, alongside function calling, structured outputs, and fine-tuning support.
What that multimodal window actually does in practice, per TechNode Global’s writeup: process hundred-page documents or a full television series in one pass, edit raw personal footage into an edited cut, reconstruct a front-end web project from a screenshot, and turn a 2D floor plan into a 3D visualization. Alibaba also says the model ran a 16-day autonomous coding project end to end, producing an open-sourced framework called oh-my-cli, and worked through a 500-plus step chip design optimization task without a human in the loop, according to both SiliconANGLE and TechNode Global.
Pricing on the hosted API: $2.00 per million input tokens, $6.00 per million output tokens, and $0.25 per million cached tokens, confirmed by MLQ News and MarkTechPost. Open weights, plus a smaller Qwen3.8-27B checkpoint sized for single-server deployment, are due the week of August 10. As of the August 3 launch, nothing had appeared yet on Qwen’s Hugging Face page and no license had been published.
Why it matters
China’s frontier labs are not waiting for anyone. Kimi K3 from Moonshot AI, at 2.8 trillion parameters, launched two days before Qwen3.8-Max’s July preview. DeepSeek shipped DeepSeek-V4-Flash around the same window. Add GLM-5.2 and MiniMax M3 to the list and you get a summer where five separate Chinese labs each put out a frontier-class model, most of them open-weight, within weeks of each other.
That pace is the story. A FourWeekMBA analysis makes the point directly: when open-weight models reach frontier performance, the moat stops being the model itself. It moves to distribution, the fine-tuning ecosystem around a model family, and who controls the inference infrastructure it runs on. Alibaba is not trying to out-ChatGPT OpenAI. It is trying to be the layer every other Chinese AI company builds on top of, the way AWS became the layer everyone else’s app ran on.
MoE is the technical trend underneath all five of those releases. Every one of them is sparse, and every one of them activates a fraction of its total parameters per token. That is partly an engineering choice and partly a response to hardware limits. Export controls have constrained the supply of top-tier Nvidia chips into China, and MoE lets a lab train and serve a model with strong benchmark numbers without needing the raw GPU count a dense model of the same size would demand. It is compute-efficiency as a geopolitical workaround, not just an architecture preference.
Native multimodality at this scale is the other half of the shift. A trillion-plus-parameter model that reads a document, watches a video clip, and looks at a screenshot in the same request is not a chatbot with a vision plugin bolted on. It changes what “ask a question” means. Reconstructing a working front-end from a screenshot, or watching 100 hours of footage and cutting it into a summary, requires the model to reason across modalities in one context window rather than stitching together separate calls to separate models.
One consequence worth flagging for anyone thinking about AI search visibility: as multimodal models this large become the default way people query information, and as they get folded into agent workflows that browse, extract, and synthesize on their own, the same content-structure fundamentals that drive citations in ChatGPT and Perplexity apply to a much wider set of assistants. A model that can ingest a hundred-page PDF in one pass is also a model that can cite, summarize, or ignore your site’s documentation just as easily.
Where Qwen3.8-Max actually lands competitively is mixed, and Alibaba’s own framing admits it. The July preview claim was that it ranks “second only to Fable 5” among benchmarked systems, but MarkTechPost’s preview coverage noted at the time that benchmark tables and model cards were still unpublished, so the claim carried a real caveat. By the August 3 launch, published numbers put it 5th in Text Arena and 4th in Frontend Code Arena, trailing Claude Opus 5 by 37 points in the latter, per SiliconANGLE. Where it does lead is Vision Arena, at 2nd place, and computer-use tasks: 86.1 on OSWorld-Verified against a reported 83.2 for GPT-5.6 Sol Max.
Frequently asked questions
How many parameters does Qwen3.8-Max have? 2.4 trillion total, using a sparse Mixture-of-Experts design that activates roughly 95 billion of them per token, according to TechNode Global and MLQ News.
Is Qwen3.8-Max open source? The API launched immediately on Alibaba Cloud Model Studio. Open weights, along with a smaller Qwen3.8-27B model sized for on-premise deployment, were scheduled for the week of August 10, roughly a week after the August 3 launch.
What can Qwen3.8-Max do that earlier Qwen models could not? It is the first Qwen model above 1 trillion parameters built natively multimodal, taking text, images, video, and documents in the same request rather than routing vision or document tasks to a separate model.
How does Qwen3.8-Max compare to GPT-5.6 and Claude? It trails Claude Opus 5 in Frontend Code Arena by 37 Elo points and ranks 5th overall in Text Arena, but it leads on computer-use benchmarks, scoring 86.1 on OSWorld-Verified versus a reported 83.2 for GPT-5.6 Sol Max, and it ranks 2nd in Vision Arena.
What does Qwen3.8-Max cost to use? Alibaba’s launch API pricing is $2.00 per million input tokens and $6.00 per million output tokens, with cached tokens at $0.25 per million.
Primary sources and further reading
- Alibaba Qwen Releases Qwen3.8-Max (MarkTechPost, Aug 3) - the most detailed technical breakdown, with full context window, pricing, and benchmark figures
- Alibaba Previews Qwen3.8-Max (MarkTechPost, Jul 19) - the World AI Conference preview, including the “first Qwen above 1 trillion parameters, multimodal” framing
- China’s Alibaba launches Qwen3.8-Max AI model (TechNode Global) - architecture details and multimodal use cases
- Alibaba debuts Qwen3.8-Max model with 2.4T parameters (SiliconANGLE) - benchmark rankings and the autonomous coding and chip-design claims
- Alibaba Launches Qwen3.8-Max, a 2.4 Trillion Parameter Open-Weight AI Model (MLQ News) - pricing and open-weight release timeline
- Alibaba’s Qwen3.8 Max and the 2.4-Trillion-Parameter Bet on Open-Weight Infrastructure (FourWeekMBA) - the strategic read on why Alibaba is racing to open-weight frontier scale