Google's Gemini Omni turns images, audio, and text into video
Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start
Google's Gemini Omni generates and edits video through conversation, using text, images, audio, and video as inputs, with Omni Flash named as the starting version; the RSS snippet says the model reasons across modalities, but the post does not disclose launch date, pricing, context limits, benchmarks, or API availability.
Why it matters: Google-scale Gemini multimodal video update clears HKR-H/K/R: Omni Flash, chat-based editing, and four input types are concrete. Pricing and rollout are not disclosed, so it sits in the lower must-write band.