Key takeaways
- Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers…
- 8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips autonomously.
- 5 architecture, and the team says the focus is on completing complex tasks independently over extended periods rather than just answering…
What happened
8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips autonomously. The team plans to release the weights next week. 8-Max, its most capable language model to date. 4 trillion total parameters, with 95 billion active per query.
The model started with a working but bloated design using 8,298 gates and whittled it down to 678 gates over roughly 500 iterations. After an automated layout pass with the open-source tool OpenROAD, the physical chip area shrank from 106x106 to 46x46 micrometers, an 81 percent reduction. The Qwen team says the model kept making deep structural changes even after hundreds of iterations instead of settling for surface-level tweaks.
The second case study is E-Commerce-Bench, a simulation of an entire fiscal year in online retail based on anonymized data from Taobao and Tmall. The model starts with 100,000 yuan in capital and has to run multiple online stores in parallel for a full year. That means buying products, negotiating with suppliers in natural language, adjusting prices, managing returns, and dealing with crises like typhoons or supply chain disruptions.
Hidden in the supplier pool are 152 scammers that the model has to spot. 8-Max ended up with a balance of 416,252 yuan, quadrupling its starting capital. 7-Max managed. The model invested aggressively early in the year and pulled in a net profit of over 100,000 yuan during the holiday season. 8-Max can process documents with over 200 pages and videos longer than 100 hours, the team says.
They're also introducing RecreationBench, a new benchmark that requires the model to rebuild running applications without access to source code. The model can only observe the target app through interaction, meaning clicks and keyboard input. Testing covers Ubuntu, macOS, Windows, Android, and the web. Alongside this, the team is releasing Qwen-MM-Plugins, an extension library that adds image and video processing, visual tool use, and multimodal memory to existing agent systems.
Training no longer focused on single tasks alone but also covered multi-day workflows, nested directory structures instead of individual files, and a variety of agent harnesses. 725.
Why it matters
5 architecture, and the team says the focus is on completing complex tasks independently over extended periods rather than just answering one-off prompts. Alibaba first announced the model in mid-July as a preview version available through Alibaba's Token Plan, Qoder, and QoderWork at ten percent of the standard price. 4 trillion parameters and ranked the model just behind Fable 5, but didn't share benchmarks.
8-Max is the first model in the Qwen-Max class whose weights will be made publicly available. 8-Max's coding chops, the team presented three case studies in which the model worked without any human help. 8-Max spent 16 days building the command-line tool oh-my-cli. The model took incoming user requests, turned them into GitHub issues, assigned them to itself, wrote the code, ran tests, and improved the results iteratively.
By July 30, 2026, it had racked up 265 commits, 127 pull requests, and 151 issues, all without a single human touch. In the second case, the model received the research paper "Unified Data Selection for LLM Reasoning" but no starter code. Its job was to reproduce the paper's results and then improve on them.
8-Max wrote 7,600 lines of code and ran 33 GPU training jobs, according to the team. It first reproduced all six of the paper's main results. 7 points. The third case involved the WWW2025 Multimodal Dialogue Intent Recognition Challenge on Alibaba's Tianchi platform, where 526 human teams competed. 5-VL-7B for product screenshots and combined them into a voting system. 853. 8-Max ahead of 458 of the 526 human teams.
Two more case studies target tasks that stretch across hundreds of interaction rounds. 8-Max had to design a cryptographic building block for encryption schemes. The key efficiency metric for such a circuit is the number of logic gates it needs, the basic elements on a chip. Fewer gates mean a smaller, more efficient chip.
What to watch
6 Sol across many categories. 8-Max hits 93, the highest score in the comparison. 8. As is typical with self-reported numbers from model makers, these results come from internal runs. Independent verification is still pending. The Qwen team attributes the model's ability to sustain such long-running tasks to a major expansion of training environments during reinforcement learning.



