Qwen-Image-3.0 drops with a “Real” focus: 4.5k token prompts, legible 10px text
通义千问发布 Qwen-Image-3.0 图像生成模型,核心关键词为"实"
Qwen released Qwen-Image-3.0, its third-gen image foundation model, with “Real” as the one-word pitch. It accepts up to 4.5k tokens of instruction and can render a full 3×3 grid of dense infographics in one shot—formulas, charts, and mixed Chinese-English text all stay in place. At the detail level, 10px text remains legible, and LaTeX-heavy academic paper pages come out with correct superscripts, subscripts, and alignments. The model natively handles 12 languages and can simulate UIs like web pages, games, and livestreams. The post shows off an algebraic geometry paper page, a newspaper layout, a book page with red handwritten annotations, and portrait shots with pore-level texture. The post does not disclose parameter count, inference cost, or open-weight plans, so I’d hold off on the “productivity tool” label until real API latency and pricing land.
Why it matters: Alibaba's Qwen team ships its third-gen image foundation model, betting on information density and text fidelity — 4.5k token prompts, legible 10px text, and native 12-language rendering are hard claims. Scored as a domestic flagship release per policy. Held below 85 because w...