Skip to content
Hacker News front page

Krea releases Krea 2 technical report: open-source text-to-image models built for aesthetic diversity and creative control

Krea 2 Technical Report

Krea 2 is a series of open-source text-to-image foundation models released under a permissive license. Instead of optimizing for a single polished default look, it aims to cover a broad range of visual styles and give users ways to explore them via text or reference images. The pretraining data deliberately excludes AI-generated images and avoids aesthetic-score filters—only duplicates, samples VLMs can't describe well, harmful biases, and overly complex images are removed. The architecture is a diffusion transformer (DiT) with iREPA, improved VAEs, Qwen3-VL text encoder, and components like GQA and sigmoid-gated attention to speed up convergence. Training runs through pretraining, midtraining, SFT, preference optimization, and RL. To bridge the gap between short user prompts and the model's rich conditioning space, Krea 2 adds a prompt expander (two-stage SFT+RL on open-source LLMs) and a style-reference system that lets users control style and mood from uploaded images, with adjustable strength and weighted mixing. It ranks in the top 10 on the Artificial Analysis text-to-image leaderboard and second among independent labs.

Why it matters: Krea 2 ships open-source with a detailed technical report and a clear data curation stance (no AI-generated images, no aesthetic scoring). Useful for model builders, but the image-gen space is crowded and Krea isn't a tier-1 lab, so it lands at the 78 featured threshold.

Read the original ↗Export Markdown