Qwen-Image-2.1 open-sourced: unified 7B model for image generation, editing, and transparency
What happened
Qwen 放出了一个叫 Qwen-Image-2.1 的开源模型,参数量 7B,把文生图和图像编辑做进了一个轻量系统里。它最特别的地方是原生支持透明图像:你可以直接用文字让它生成带透明通道的 PNG,也能编辑已有的透明图层,甚至从普通照片里把主体抠出来变成 RGBA 素材。编辑能力上,它一次最多能参考 10 张图,支持局部修改和保持人物、产品不走样。为...
From AI HOT 精选
Coverage
Follow the reports to see the story from different sides.
- Computing Life · Share · YagePickQwen-Image-2.1: A version rollback that packs text rendering, editing, and native RGBA into one open-weight model
Qwen released Qwen-Image-2.1 on Sep 20, a 7B open-weight image model that unifies text-to-image, local editing, and native RGBA output in a single pipeline. The version number rolled back from 3.0 to 2.1 reflects a 2026 split: 3.0 is a closed-source commercial API, while 2.1 continues the open research branch. A built-in RGBA VAE outputs PNGs with transparency, skipping external matting. The interface supports up to 10 reference images and three mask types at native 2K. The license shifted from Apache 2.0 to a research-only agreement; commercial use requires a separate license. Community tests show ~25s per megapixel image on RTX 5070/5080 at 25 steps, ~15.6GB VRAM with Q8 quantization. Text rendering remains a strength, but multi-subject consistency shows facial generalization on well-known public figures—official demos don't guarantee universal performance.
- AI HOT (Curated Pool)PickQwen-Image-2.1 Tops Arena's Open-Source Leaderboards for Image Editing and Text-to-Image
Alibaba's Qwen-Image-2.1 ranks first among open-source models on Arena's Image Edit and Text-to-Image leaderboards. It scored 1367 in Image Edit Arena, placing 16th overall, just 3 points behind GPT-Image-1.5-high-fidelity at #15. The post doesn't disclose parameter count, architecture, or release timeline.
- AI HOT (Curated Pool)PickQwen-Image-2.1 released as open weights, tops Image Edit Arena among open-source models
Qwen-Image-2.1 is out with open weights. It scored 1367 on the Arena Image Edit Arena, ranking #1 among open-source models and #16 overall — just 3 points behind GPT-Image-1.5-high-fidelity at #15. It also landed #1 open-source on the Text-to-Image Arena. The post doesn't disclose parameter count, architecture details, or the exact open license.
- AI HOT (Curated Pool)PickQwen-Image-2.1: A 7B Single-Checkpoint Model for Both Image Generation and Editing
Qwen-Image-2.1 is a 7B native image generation and editing model. It uses a single checkpoint for both tasks and supports up to 10 reference images. The model includes a built-in prompt-enhancement LLM, integrates with diffusers and ComfyUI, and offers a no-install browser demo on Hugging Face Spaces. The post doesn't disclose training data, inference latency, or benchmark comparisons.
- AI HOT (Curated Pool)PickQwen-Image-2.1 now works with ComfyUI, open weights available
Alibaba Qwen released open weights for Qwen-Image-2.1 with native ComfyUI support. A single 7B checkpoint handles both image generation and editing, outputs up to 2K natively, accepts up to 10 reference images per instruction, and supports RGBA with alpha channel. The post doesn't spell out license terms or hardware requirements.
- AI HOT (Curated Pool)PickQwen open-sources Qwen-Image-2.1: a 7B model unifying generation and editing with native transparency
Qwen released Qwen-Image-2.1, a 7B model that merges text-to-image generation and image editing into one lightweight system. It natively handles transparent images—generating them from prompts, editing layers, and extracting subjects from photos as RGBA assets. Editing supports up to 10 reference images, local edits, and identity preservation. A mixed-granularity attention design with KV cache reuse cuts inference cost for multi-image tasks. The model is open-sourced on GitHub, Hugging Face, and ModelScope.