Key takeaways
- Alibaba's Qwen team is introducing Qwen3.8-Flash-Next, a multimodal mixture-of-experts model that serves as an architecture preview of…
- 8-Flash-Next, a multimodal mixture-of-experts model that serves as an architecture preview of Qwen4.
- " The model natively supports a 262,144-token context window and can scale to one million tokens using YaRN.
What happened
8-Flash-Next, a multimodal mixture-of-experts model that serves as an architecture preview of Qwen4. It aims to match much larger models at a fraction of the training cost. The model has 125 billion total parameters but only activates 6 billion per token. It also includes 51 billion parameters in a novel N-gram embedding layer, one of the architecture innovations slated for Qwen4.
That kind of pricing keeps ratcheting up the pressure on OpenAI and Anthropic. 8-27B has also been popular lately since it can run locally and delivers strong performance for almost no money, assuming you have the hardware. 6 model line. That's good for users but bad for the rapid revenue growth AI providers need to keep the investment narrative alive.
Why it matters
" The model natively supports a 262,144-token context window and can scale to one million tokens using YaRN. The technical report is on GitHub, and weights are available on Hugging Face and ModelScope. 47 per million output tokens. The API should go live shortly, according to Qwen. 7-Plus at roughly one-ninth the training cost, according to the Qwen team, with the biggest gains in coding and office tasks.
What to watch
7-Plus has 397 billion parameters with 17 billion activated per token, nearly three times Flash-Next's active count. 6 (Max). Despite both being much larger or more expensive, Flash-Next leads in the majority of tested tasks. The model seems optimized for agentic coding benchmarks, which require an AI to independently find and fix bugs in real software projects. 6. The gap on office and productivity tasks is even wider. 1. 6.
9), the models are closely matched. 6 only comes out ahead on Humanity's Last Exam, a test built around extremely hard multidisciplinary problems, but it's also an "older" Anthropic model from February 2026. As always, benchmark scores and real-world performance can differ. 6 Sol. Flash-Next performs just below the flagship but costs about one-twelfth as much, with a roughly 12x price gap on both input and output tokens.



