This one's worth opening because it collapses three separate tools into one pipeline: generate the image, knock out the transparent background, and edit local text—all in a single run. 7B parameters, native 2K, up to 10 reference images and three mask types. Community tests show ~25s per megapixel on an RTX 5070/5080, ~15.6GB VRAM with Q8 quantization.
The version number jumping from 3.0 back to 2.1—Qwen never explained it, but the release pattern tells the story. July's 3.0 is API-only, no weights, no tech report, focused on 4.5K-token prompts and newspaper-grade layout. September's 2.1 ships open weights but swaps Apache 2.0 for a research-only license; commercial use requires a separate agreement. Two tracks, and 2.1 continues the open branch.
Compared to peers: Z-Image Turbo is fast but explicitly doesn't do in-image text. FLUX.2 has solid local editing but no native transparency channel. 2.1 isn't chasing the highest single-dimension image quality—it's the only open-weight model right now that bundles text rendering, RGBA output, and multi-image editing in one package.
Two discounts I'd apply. First, the license: research and evaluation only, commercial deployment means emailing for a separate deal—don't treat this like Apache 2.0. Second, multi-subject consistency: community feedback notes facial generalization works well on well-known public figures, but the article itself says official demos don't guarantee universal performance. If you need an open-weight tool that spits out transparent assets and handles text edits, this is worth a test run. If you're banking on general-purpose face consistency, try it on your own images first.