Skip to content
AI HOT (Curated Pool)

Boogu-Image-0.1: An open-source unified model for multimodal understanding and generation, trained for ~$400K

Boogu-Image-0.1 发布:开源统一多模态理解与生成模型,训练成本仅约 40 万美元

Boogu-Image-0.1 is an open-source unified multimodal model family that handles image understanding, text-to-image generation, instruction-based editing, and bilingual text rendering. It ships in four variants: Base, Turbo, Edit, and Edit-Turbo. The team used only 208.62 million unique images, and the base model's theoretical training cost is around $400K. The paper argues that better model understanding, data quality, training pipelines, and agentic inference-time scaling can push generation and editing performance even on a tight compute budget. Benchmarks show it matches or beats other open-source models and approaches closed-source systems like Nano-Banana Pro and GPT-Image-2. Weights, code, and recipes are released under Apache 2.0.

Why it matters: Open-source unified multimodal model with a $400K training cost hook and clear variant breakdown — substantive. But unknown team, paper just dropped, no community validation yet, so score sits right at the featured threshold.

Read the original ↗Export Markdown