Qwen-Image-2.1: A version rollback that packs text rendering, editing, and native RGBA into one open-weight model
Qwen-Image-2.1:版本号断裂背后的双线策略,与一个模型接住三件事的野心
Qwen released Qwen-Image-2.1 on Sep 20, a 7B open-weight image model that unifies text-to-image, local editing, and native RGBA output in a single pipeline. The version number rolled back from 3.0 to 2.1 reflects a 2026 split: 3.0 is a closed-source commercial API, while 2.1 continues the open research branch. A built-in RGBA VAE outputs PNGs with transparency, skipping external matting. The interface supports up to 10 reference images and three mask types at native 2K. The license shifted from Apache 2.0 to a research-only agreement; commercial use requires a separate license. Community tests show ~25s per megapixel image on RTX 5070/5080 at 25 steps, ~15.6GB VRAM with Q8 quantization. Text rendering remains a strength, but multi-subject consistency shows facial generalization on well-known public figures—official demos don't guarantee universal performance.
Why it matters: Qwen open-sourced a model that combines image generation, editing, and native transparency output into one pipeline — a clear engineering increment, not a reskin. The backward version jump is inherently clickable, and it resonates with both designers and developers. Not scorin...