Skip to content
AI HOT (Curated Pool)

MiniMax launches open-source H3, a multimodal model that generates 2K video with native stereo audio

MiniMax H3 发布:开源全能多模态生成模型,支持 2K 原生立体声视频

MiniMax H3 unifies text, image, video, and audio understanding into one model. It generates up to 15-second videos at 2K resolution with native stereo sound. The company claims per-second pricing at 2K is under one-third of mainstream models, and 768p is under half the price of mainstream 720p. Model weights will be open-sourced in the coming days, subject to legal review. The post does not specify the open-source license or exact release date.

Why it matters: MiniMax dropped H3, a multimodal generation model that handles text, image, video, and audio in one model, outputting 15-second 2K video with native stereo sound. Pricing is aggressive — 2K per second at less than one-third of mainstream models — and weights are planned for op...

Read the original ↗Export Markdown