MiniMax launches H3, an open multimodal model that generates 2K video with native stereo sound
MiniMax H3
MiniMax launched H3 on Product Hunt, an open multimodal model for unified video generation. It takes text, image, and audio inputs to produce 2K video with native stereo sound. The model focuses on accurate text rendering, visual packaging, and complex instruction following for commercial content like motion design and branding. The post doesn't disclose parameter count, inference latency, or the specific open license. I'd hold off on the 'world-leading' claim until a technical report is out.
Why it matters: MiniMax drops H3, a multimodal video model that takes text, image, and audio and outputs 2K stereo video, targeting commercial content with text rendering and visual packaging as explicit strengths. But no parameter count, latency, or license disclosed — the 'world-leading' cl...