Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

201–220 of 514

May 30Saturday

Synced · WeChat

NVIDIA and Tsinghua Team's Gamma-World Tops Hugging Face Daily Chart

NVIDIA, Tsinghua, University of Toronto, and Vector Institute released Gamma-World, a multi-agent world model using simplex-based positional encoding and hub tokens to cut interaction cost from quadratic to linear, with 8-player latency dropping from 17.6 ms to 4.5 ms.

Why it matters: HKR-H/K/R all pass: Gamma-World has a concrete mechanism and latency claim from NVIDIA/Tsinghua. Scope remains multi-agent world-model research, so it sits in the 78–84 good-quality band rather than must-write.

AI HOT (Curated Pool)

OpenAI launches real-time translation model with 70+ input languages

OpenAI launched gpt-realtime-translate, a speech translation model that accepts 70+ input languages and outputs speech in 13 target languages; the post says the feature is running on smart glasses.

Why it matters: HKR-H/K/R all pass: OpenAI has a concrete realtime-translation model with numbers and a wearable demo. Missing latency, pricing, and API availability keep it below P1.

The Verge · AI

Tech companies desperately want to film you doing chores

Shift said it would clean New Yorkers’ homes for free if it can film cleaners doing chores such as washing dishes, wiping counters, dusting tables, and mopping floors, creating domestic robot training data; the snippet does not disclose consent terms, data retention, pricing, or expansion timing for cities such as London.

Why it matters: HKR-H/K/R all pass: the odd trade is clickable, the post gives a concrete chore-video collection mechanism, and it hits robotics data plus privacy nerves. This is a strong industry feature, not a major model or platform release, so it sits in 72–77.

May 29Friday

The Verge · AI

This AI startup will clean your home for free to train future robots

Shift offers free home cleaning and records cleaners scrubbing, vacuuming, dusting, tidying, and washing to collect robot training footage; the RSS snippet does not disclose service cities, privacy terms, consent mechanics, or dataset scale.

Why it matters: HKR-H and HKR-R are strong: Shift turns home cleaning into robot-training data collection. HKR-K has a clear mechanism, but city scope, privacy terms, and dataset scale are not disclosed, so this stays low-featured.

The Verge · AI

Adobe’s Conversational AI Agent Is a Mediocre Design Intern

The Verge tested Adobe Firefly AI Assistant in beta. It can operate Adobe design apps as a conversational middleman, rather than only generating images or video. The post says it explains edit steps clearly, but the results were not impressive. The RSS snippet does not disclose pricing, release timing, or the full list of supported apps.

Why it matters: HKR-H/K/R all pass because this is a Verge hands-on of Adobe’s Firefly AI Assistant beta with a clear negative usability hook. Missing pricing, launch timing, and supported-app details keep it in the 72–77 featured-threshold band.

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.

AI HOT (Curated Pool)

Google DeepMind CEO Demis Hassabis Says AGI Could Arrive Within Three Years

Demis Hassabis predicts AGI could arrive around 2029 to 2030, with mature multimodal capabilities and autonomous decision-making as key conditions, while warning that society remains underprepared and needs rules and safeguards before deployment.

Why it matters: HKR-H/K/R all pass: Hassabis gives a 2029-2030 AGI window and names multimodal plus autonomous decision-making as conditions. High-interest commentary, but thinner than a model release or major product update.

r/LocalLLaMA

StepFun 3.7 Flash

StepFun released Step 3.7 Flash with 196B total parameters, 11B active MoE, a built-in 1.8B ViT, and local execution on 128GB RAM.

Why it matters: HKR-H/K/R pass via the 196B/11B MoE specs and 128GB local-run claim. Sparse Reddit sourcing leaves license, eval method, and access conditions undisclosed, so it stays in the lower featured band.

AI HOT (Curated Pool)

StepFun Releases Step 3.7 Flash, Focused on Agent Efficiency

StepFun released the open-source Step 3.7 Flash model with a 198B-parameter MoE architecture, about 11B active parameters, a 256K context window, and a 67.1 score on ClawEval-1.1.

Why it matters: HKR-H/K/R all pass: the release has a clear sparse-model hook, concrete context and benchmark numbers, and practitioner resonance around open agent efficiency. Official-post sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

Nano Banana Pro and Nano Banana 2 officially released

Google AI Developers released Nano Banana Pro and Nano Banana 2, two image models available for production use through the Gemini API; the post names gemini-3-pro-image and gemini-3.1-flash-image but does not disclose pricing, benchmarks, or rate limits.

Why it matters: HKR-H/K/R all pass: Google shipped two production image models via Gemini API. The post gives no benchmarks, pricing, or safety mechanism, so this stays in the 78–84 band rather than p1.

The Verge · AI

A $2,000 AI-generated film will debut at Tribeca

Tribeca Festival will premiere the 75-minute AI-generated film Dreams of Violets, which cost $2,000 to make and uses people and images fully created by AI.

Why it matters: HKR-H/K/R all pass: a $2,000, 75-minute AI film entering Tribeca has novelty, concrete numbers, and labor resonance. It is a film-industry application, not a model or platform launch, so it stays in the low featured band.

May 28Thursday

NVIDIA Blog

NVIDIA Research Advances Robotics From Simulation to the Real World

NVIDIA Research presented 8 ICRA papers on sim-to-real robotics: ScheduleStream delivered a 3x speedup for multi-arm planning, COMPASS reached about 80% success across 20 real-world navigation trials, and Grasp-MPC achieved about 75% real-robot grasping success.

Why it matters: HKR-K and HKR-R are strong: the post gives concrete sim-to-real numbers from ICRA and addresses robot deployment reliability. HKR-H is moderate but passes on the real-world success-rate hook.

r/LocalLLaMA

Nvidia LocateAnything: Fast Vision-Language Grounding with Parallel Box Decoding

The title says Nvidia LocateAnything-3B performs vision-language grounding with parallel box decoding and runs 10x faster than Qwen3-VL; the post body only provides Hugging Face, GitHub, demo, and project links, and does not disclose benchmark setup or accuracy numbers.

Why it matters: HKR-H/K/R all pass, but the body is mostly links and title-level facts, with no full eval setup or quality metrics. NVIDIA open vision grounding is useful enough for featured, not same-day must-write.

Synced · WeChat

ICML 2026: AutoMoT reaches SOTA on Bench2Drive and nuScenes

NTU AutoMan Lab, Harvard, and Xiaomi Auto proposed AutoMoT, a unified VLA driving model using a 4B Qwen3-VL Understanding Expert and a 1.6B Action Expert with asynchronous inference, reaching 89.42 DS and 74.09% SR on Bench2Drive with AutoMoT+.

Why it matters: HKR-H/K pass via the async VLM-driving setup and concrete Bench2Drive numbers. The autonomy focus narrows HKR-R, so this sits at the featured threshold rather than the 78+ research tier.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

AI HOT (Curated Pool)

Open-source FastVideo Dreamverse real-time video generation tool

Hao AI Lab open-sourced FastVideo Dreamverse, a real-time video generation tool that generates a 30-second 1080p video in 7 seconds under the stated setup of one NVIDIA B200 GPU and LTX-2.

Why it matters: HKR-H/K/R all pass: the 7s-for-30s-1080p claim is concrete and practitioner-relevant. Single-source X sourcing and missing independent benchmarks keep it in the 78–84 band.

May 27Wednesday

AI HOT (Curated Pool)

Runway launches Model Context Protocol server

Runway launched an MCP server that lets compatible agents such as Claude, ChatGPT, and Cursor generate images and videos inside chat interfaces, with access to Gen-4.5, Seedance 2.0, GPT Image 2, Kling 3.0, and Nano Banana Pro.

Why it matters: HKR-H/K/R all pass, but this is a Runway product integration, not an MCP protocol change or model release. It clears featured, with the score kept in the 72–77 band.

TechCrunch · AI

YouTube will now automatically label AI videos

YouTube will automatically label videos using significant photorealistic AI, no longer relying only on creator self-disclosure. The RSS snippet says AI labels will become more prominent, but the post does not disclose rollout timing, detection thresholds, appeal rules, or whether the system covers shorts and livestreams.

Why it matters: HKR-H/K/R pass: YouTube shifts AI-video labels from creator self-reporting to platform detection. The article gives the mechanism, but not accuracy, appeals, or rollout scope, so it sits at the featured threshold.

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

r/LocalLLaMA

PrismML Released Binary and Ternary Bonsai Image 4B

PrismML released Binary and Ternary Bonsai Image 4B, 1-bit and ternary text-to-image diffusion transformers around 3GB, compared with FLUX.2 Klein 4B at about 16GB, with browser-local WebGPU demo links and an Apache-2.0 license disclosed in the Reddit snippet.

Why it matters: HKR-H/K/R all pass: low-bit image DiT plus local browser inference is a strong hook, backed by 4B, ~3GB, WebGPU, and license details. Reddit sourcing and limited lab weight keep it in the low featured band.