Microsoft releases Mage-Flow: a 4B native-resolution model for image generation and editing
Mage-Flow - An Efficient Native-Resolution Foundation Model for Image Generation and Editing - Microsoft
Microsoft open-sourced Mage-Flow on HuggingFace, a 4B-parameter foundation model for text-to-image generation and instruction-based editing. It matches or beats much larger systems like Qwen-Image 20B and FLUX.2 32B by co-designing a lightweight VAE and a native-resolution diffusion transformer. Mage-VAE uses ~12× fewer encode MACs and ~22× fewer decode MACs per pixel than FLUX.2-VAE. The model natively handles 512–2048 resolutions at any aspect ratio, including extreme 4:1. Turbo variants hit 0.59 s/image for generation and 1.02 s/edit at 1024² on a single A100, with 18–20 GB peak memory. Training throughput improved ~2.5× via native-resolution packing and fused CUDA kernels. The family ships Base, RL-aligned, and 4-step Turbo checkpoints for both generation and editing. The post does not disclose 3090/4090 benchmarks, so it's unclear how well A100 kernel optimizations transfer to consumer 24 GB GPUs.
Why it matters: Microsoft open-sources a 4B image gen model that uses a lightweight VAE and native-resolution approach to compete with Qwen-Image 20B and FLUX.2 32B — a solid efficiency story. Score held at 78 because we only have the Reddit post so far; no formal paper or third-party reprodu...