Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

261–280 of 514

May 20Wednesday

AI HOT (Curated Pool)

Gemini Omni: A new model for creating content from any input

Google AI published a Gemini Omni post, saying the model creates content from any input starting with video; the post does not disclose parameters, pricing, availability timing, or benchmark results.

Why it matters: Google AI’s post gives Gemini Omni and “content from any input” as hooks, so HKR-H/R pass; HKR-K fails because price, timing, parameters, and benchmarks are absent, keeping it at low featured.

AI HOT (Curated Pool)

Google releases Gemini Omni multimodal generation model

Google released Gemini Omni, a multimodal generation model that combines image, video, and text inputs to generate videos grounded in Gemini’s real-world knowledge; the post does not disclose model size, pricing, or availability.

Why it matters: HKR-H/K/R all pass: Google’s Gemini Omni adds combined image, video, and text inputs for video generation. Missing parameters, pricing, and rollout timing keep it at the low end of the must-write band.

AI HOT (Curated Pool)

Google I/O Announces Multiple Gemini Updates

Google announced multiple Gemini updates at Google I/O, including a new experience design using neural expression technology, upcoming agent features with Daily Brief and Gemini Spark, plus Gemini Omni and 3.5 Flash; the RSS snippet does not disclose release dates, model parameters, pricing, or benchmark results.

Why it matters: HKR-H/K/R all pass: Google I/O brings concrete Gemini hooks in Omni, 3.5 Flash, and agents, with clear competitive resonance. Missing launch timing, specs, and pricing keep it in the 78–84 band.

r/LocalLLaMA

Nemotron-Labs-Diffusion from NVIDIA

NVIDIA released the Nemotron-Labs-Diffusion 3B, 8B, and 14B dense model family with AR decoding, diffusion parallel decoding, and self-speculation; the 8B model reaches 850 tok/s on GB200 at concurrency 1, compared with 253 tok/s for AR and 360 tok/s for Eagle3.

Why it matters: HKR-H/K/R all pass: NVIDIA diffusion LLMs, concrete sizes/mechanisms, and an 850 tok/s GB200 claim. Single-source Reddit sourcing keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Google Search gets its biggest redesign in 25 years with AI-driven interaction changes

Google announced at I/O 2026 the biggest Search redesign in 25 years, powered by Gemini 3.5 Flash, with AI Mode exceeding 1 billion monthly active users and query volume doubling each quarter.

Why it matters: Google Search is an internet entry-point product; 1B+ AI Mode MAU and quarterly query doubling put this beyond a routine feature update. HKR-H/K/R all pass, so it lands in p1.

The Verge · AI

Gemini will use Volvo’s external cameras to interpret parking signs

Google and Volvo announced at I/O that Gemini will access external cameras on the upcoming EX60 SUV, with the first stated use case translating hard-to-understand parking signs for vehicle owners.

Why it matters: HKR-H and HKR-K pass: Gemini tying into Volvo EX60 exterior cameras is a concrete multimodal in-car use case. HKR-R is weak because rollout scope, privacy, and safety mechanisms are not disclosed.

The Verge · AI

The 13 Biggest Announcements at Google I/O 2026

Google announced Gemini 3.5 Flash at I/O 2026. It becomes the default model today for the Gemini app and AI Mode in Search, while Gemini 3.5 Pro follows next month; the RSS snippet also mentions Search, Gmail, and Project Aura smart glasses updates but does not disclose the full list of 13 announcements.

Why it matters: HKR-H/K/R all pass, but the text only gives Gemini 3.5 Flash default rollout and Pro timing; it lacks the full 13 items, benchmarks, or pricing, so this stays featured below p1.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash with 55 intelligence score

Google released Gemini 3.5 Flash, raising its intelligence score by 9 points to 55, exceeding 280 output tokens per second, and increasing operating cost by 5.5 times versus the previous generation.

Why it matters: A Google Gemini 3.5 Flash release is a top-lab model update, backed by Artificial Analysis numbers for speed, intelligence, and cost. HKR-H/K/R all pass, with the 5.5x cost jump making it more than a routine launch.

TechCrunch · AI

Google’s Genie world model can now simulate real streets with Street View

Google DeepMind is integrating Street View with Project Genie for interactive street-level simulations in robotics, gaming, and travel; the post does not disclose model parameters, launch timing, or evaluation results.

Why it matters: Google DeepMind connecting Street View to Genie gives HKR-H/K/R: a novel hook, a concrete mechanism, and robotics/data-moat resonance. Missing params, launch timing, and evals keep it in the 78–84 band.

TechCrunch · AI

OpenAI is making it easier to check if an image was made by its models

OpenAI announced two measures for detecting AI-generated images: it joined the open C2PA standard and added Google’s SynthID to its products. The RSS snippet does not disclose which OpenAI products include SynthID, whether the checks cover legacy images, or when the measures become available to users.

Why it matters: HKR-H/K/R all pass: the OpenAI-Google provenance tie-up is clickable, and C2PA plus SynthID are concrete mechanisms. Coverage, launch timing, and verification flow are not disclosed, so this stays at the featured threshold.

TechCrunch · AI

Google's Gemini Omni turns images, audio, and text into video

Google's Gemini Omni generates and edits video through conversation, using text, images, audio, and video as inputs, with Omni Flash named as the starting version; the RSS snippet says the model reasons across modalities, but the post does not disclose launch date, pricing, context limits, benchmarks, or API availability.

Why it matters: Google-scale Gemini multimodal video update clears HKR-H/K/R: Omni Flash, chat-based editing, and four input types are concrete. Pricing and rollout are not disclosed, so it sits in the lower must-write band.

AI HOT (Curated Pool)

Google releases Gemini Omni for any-input-to-any-output generation and conversational video editing

Google released Gemini Omni and Omni Flash at I/O 2026, supporting text, image, audio, and video inputs and outputs, with conversational video editing; Omni Flash is available in Gemini App, Google Flow, and YouTube Shorts, while the post does not disclose the API launch date.

Why it matters: HKR-H/K/R all pass: Google announced Gemini Omni and Omni Flash as a major multimodal update at I/O. API timing, pricing, and benchmarks are not disclosed, so it stays below 90.

AI HOT (Curated Pool)

Gemini Omni Released: New Progress in Multimodal Generation

Google DeepMind released Gemini Omni, starting with video generation; the post says it combines Gemini with its generative media systems, but does not disclose parameters, pricing, or availability.

Why it matters: HKR-H/K/R pass: a DeepMind Gemini-branded video-generation launch has a clear hook, a stated mechanism, and competitive resonance. Missing parameters, pricing, and access timing keep it in the 72–77 band.

AI HOT (Curated Pool)

Google processes over 3,200 trillion tokens per month, up 7x year over year

Google said at I/O 2026 that it processed over 3,200 trillion tokens per month in May, while Gemini App exceeded 900 million monthly active users and the Nano Banana model generated more than 50 billion images cumulatively.

Why it matters: Google I/O disclosed usage scale, not a new model or major capability. HKR-H/K/R pass via 7x growth, 900M MAU, and 3.2Q tokens/month, but without a launch-level update it stays in the 78–84 band.

AI HOT (Curated Pool)

NVIDIA open-sources first 4-bit infrastructure for ultra-long video generation

NVIDIA researchers open-sourced LongLive 2.0, an end-to-end long-video generation infrastructure covering training and inference with 4-bit quantization, FP4 quantization, parallel acceleration, KV-cache optimization, and 45.7 FPS generation on a 5B model.

Why it matters: HKR-H/K/R all pass: NVIDIA researcher open-sources LongLive 2.0 with 4-bit long-video train/inference and 45.7 FPS on a 5B model. This is strong open-source infra, not a flagship model launch, so it fits the 78–84 band.

May 19Tuesday

r/LocalLLaMA

ByteDance released an open-source model that attempts broad multimodal tasks with 3B parameters

ByteDance released Lance, an open-source unified multimodal model with 3B active parameters that supports image and video understanding, generation, and editing, and the post says it was trained from scratch with a staged multi-task recipe under a 128-A100-GPU budget.

Why it matters: HKR-H/K/R all pass: ByteDance’s open Lance has a compact multimodal hook, concrete 3B/128-A100 facts, and clear cost/deployment resonance. Reddit-sourced details lack benchmarks, license terms, and official context, so it stays featured, not P1.

AI HOT (Curated Pool)

Advancing content provenance for a safer, more transparent AI ecosystem

OpenAI launched an AI content provenance system that combines Content Credentials and SynthID with a verification tool; the post does not disclose supported media formats, rollout scope, or detection accuracy.

Why it matters: HKR-H/K/R pass: the OpenAI provenance stack has a concrete cross-standard mechanism and trust/compliance relevance. Missing format coverage, rollout scope, and accuracy keep it in the low featured band.

QbitAI · WeChat

Chinese GPU vendor Moore Threads releases MT Lambda for embodied AI simulation

Moore Threads released MT Lambda, an embodied AI simulation platform that combines physics, rendering, and AI engines, and demonstrated the robot dog “Xiaofei” executing a Sim-to-Real policy trained 100% in simulation on domestic hardware.

Why it matters: HKR-H/K/R pass: the story has a concrete domestic-GPU simulation hook, a three-engine mechanism, and a clear NVIDIA/robotics-cost nerve. Importance stays in the low featured band because performance, pricing, access, and third-party validation are not disclosed.

QbitAI · WeChat

World model supports multiplayer FPS gameplay before Fei-Fei Li

Odyssey released Agora-1, a world model that supports up to four human and AI players fighting in the same generated FPS world in real time. The system decouples simulation from rendering and trains on GoldenEye internal game states.

Why it matters: HKR-H/K/R all pass: Agora-1 moves world models from solo demos to up to 4-player real-time FPS, with decoupled simulation/rendering and training-data clues. The lab is not a top-tier foundation-model vendor, so this stays in the 78–84 band.

New York Times Chinese

China’s AI Microdrama Boom Brings Job Anxiety and Tech Enthusiasm

Chinese companies are producing AI-generated microdramas for about $30 per minute without cameras, crews, or human actors; DataEye says nearly 50,000 new AI microdramas were uploaded to Douyin in March, almost matching the platform’s total uploads for all of 2025.

Why it matters: HKR-H/K/R all pass: the backlash angle is clickable, the story adds $30-per-minute production and nearly 50,000 March uploads, and it hits labor anxiety. It is strong industry reporting, not a core model or product release.