MiniMax releases Music 3.0: open-weights music model that generates full 5-minute songs in one pass
MiniMax Music 3.0 发布:新一代开源权重、生产级全能音乐模型
MiniMax today released Music 3.0, an open-weights music generation model. Given a creative concept and optional lyrics, it outputs a complete song with arrangement, performance, and vocals in one pass, up to 5 minutes long. The upgrade targets three pain points: accurately interpreting creative intent, maintaining that intent across a full song, and making vocals and instruments sound performed rather than synthesized. The new Hybrid-LM architecture uses an 8B global model for song-level structure and a 0.6B local model for per-frame acoustic detail, with flow matching and a Flow-VAE converting discrete predictions to continuous audio. The post does not disclose training data scale, inference latency, the specific open-source license, or quantitative benchmark comparisons.
Why it matters: MiniMax released Music 3.0 with open weights, generating full songs up to 5 minutes with vocals in one pass. The technical paper details Hybrid-LM architecture and 8B parameters. First open-weights music model from a major Chinese lab — directly actionable for audio product te...