Key takeaways

  • Qwen-Image-3.0 can render multi-panel infographics in one pass, produce legible text as small as ten pixels, and write in twelve languages.
  • 0 can render multi-panel infographics in one pass, produce legible text as small as ten pixels, and write in twelve languages.
  • 0 accepts prompts of up to 4,500 tokens.

What happened

0 can render multi-panel infographics in one pass, produce legible text as small as ten pixels, and write in twelve languages. Whether AI-generated academic papers and newspaper pages are useful as static images remains an open question. 0, the third version of its image generator. " The model is meant to handle practical work such as newspaper layouts, storyboards, and exam sheets, not just produce attractive images.

It fills in the missing areas while matching the original brushwork and ink shading. " The model supports twelve languages natively, including Japanese, Korean, and Spanish. The published examples also show it recreating interfaces from websites, games, and livestreams. In one editing demo, the model turns an insect photo into a full identification plate with taxonomy, labels for physical features, enlarged detail views, and a scale bar.

The model can also pull in live internet data, according to Qwen, and uses it to generate things like a weather forecast for Hangzhou. Another example places Chinese ink painter Qi Baishi and Vincent van Gogh together in a simulated livestream studio. 0, just this past May.

Why it matters

0 accepts prompts of up to 4,500 tokens. According to the team, that gives the model enough room to create dense layouts in one pass rather than assemble them from several images. One demo packs nine separate infographics into a 3 x 3 grid, each with its own text, formulas, and illustrations.

The panels cover safe following distances near tunnels, perpendicular lines, a Confucian lesson about emotion and reason, and the detachment speed of a projectile from a rotating cylinder. Other panels explain the liver fluke life cycle, right-sided chest pain, Sylow theorems for groups of order 72, internal controls at banks, and DNA in animal and plant cells. The team also shows how the model handles nested interfaces.

One example starts with a VSCode window containing a Qwen Chat screen. Inside that screen is a WeChat conversation, which includes a poster explaining how to make pour-over coffee. Qwen says the model can produce legible text as small as ten pixels. Its examples include a whale shark infographic packed with text and a full page from a fictional algebraic geometry paper.

The paper contains multi-line LaTeX equations with subscripts, superscripts, braces, fractions, sums, and products. Other demos show a simulated newspaper page and red handwritten comments that resemble notes from a teacher. 0 also aims for photographic detail in portraits and objects, including visible pores, skin texture, and individual strands of hair. In another editing demo, the model repairs a damaged traditional ink painting of fighting eagles.

What to watch

The technical report focused on training and inference efficiency gains, including a faster variant that needed only four instead of 40 steps per image. 0 landed just behind OpenAI's GPT-Image-2 and Google's Nano Banana Pro. 0 appears to be available only through invite-only API access. The model should show up in first-party apps like Qwen Chat soon.

It's unlikely that the model weights will ship under an open license, as they did for the original Qwen-Image. The practical value of some demos remains unclear. Researchers typically write and typeset papers in LaTeX rather than render them as images, so AI-generated pages with formulas may be better suited to mockups and visual drafts than final papers. A similar question applies to newspaper pages and complex infographics. Modern image models can edit individual text ele