Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

241–260 of 514

May 22Friday

AI HOT (Curated Pool)

Plastic Interfaces: The Future Shape of AI-Driven Software

Salesforce has adopted a headless architecture that lets salespeople update data through AI; the post says MCPs, HTML, audio, and web interfaces can be generated dynamically by context, but it does not disclose implementation metrics or adoption numbers.

Why it matters: HKR-H/K/R all pass, but this is a software-form thesis without user metrics, launch timing, or a reproducible test. It fits the insightful-commentary band, not a must-write release.

AI HOT (Curated Pool)

Aleph 2.0 and Edit Studio

Runway released Aleph 2.0 and Edit Studio, combining generation, editing, and post-production into one platform; the post does not disclose pricing, technical parameters, or rollout scope.

Why it matters: Runway is a major AI video vendor, and Aleph 2.0 plus Edit Studio is a mid-weight product update. HKR-H/K/R pass, but missing price, specs, and rollout keep it at the featured threshold.

May 21Thursday

TechCrunch · AI

Hark raises $700M Series A for its secretive ‘universal’ AI interface

Hark raised a $700 million Series A and plans to release its first multimodal models this summer; the post does not disclose investors, valuation, model specifications, or a hardware launch schedule.

Why it matters: HKR-H/K/R all pass: the $700M Series A makes Hark a serious AI-interface contender. Investors, valuation, model specs, and hardware timing are not disclosed, so this stays featured rather than must-write.

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

Synced · WeChat

VAST and Tsinghua propose density-controlled 3D Gaussian generation for SIGGRAPH 2026

VAST and Tsinghua propose DeG, a 3D Gaussian generation method that samples Gaussian centers from a learned density distribution and trains density control with a render loss contribution gradient; in some settings, it reaches TRELLIS-like visual quality with less than half the Gaussian count.

Why it matters: HKR-H/K/R pass: DeG offers a concrete mechanism and a testable efficiency claim, reaching TRELLIS-like quality with under half the Gaussians in some scenes. SIGGRAPH research has some technical depth, but no hard-exclusion rule applies.

Synced · WeChat

Xie Saining’s Team Releases Second-Generation Representation Autoencoder RAEv2

Xie Saining’s team, Adobe Research, and the Australian National University released RAEv2, which reaches gFID 1.06 after 80 epochs on ImageNet-256 and reduces EPFID@2 from 177 epochs to 35 epochs while keeping compute at 189 GFLOPs.

Why it matters: HKR-K and HKR-R pass with concrete benchmark and training-efficiency claims. HKR-H is weak because the angle is a normal research release, so it lands at the featured threshold rather than a must-write item.

AI HOT (Curated Pool)

Tencent Launches OS-Level AI Assistant Mavis on Windows, Mac, and Android

Tencent launched the OS-level AI assistant Mavis on May 21 across Windows, Mac, and Android, with document parsing, image recognition, system maintenance, partial offline use, model dispatching, and desktop control of mobile apps listed as supported functions.

Why it matters: HKR-H/K/R all pass: Tencent’s OS-level assistant spans Windows, Mac, and Android with concrete tool abilities. Model, pricing, and permission design are not disclosed, so it stays at the lower featured band.

The Verge · AI

You can now remix other people’s YouTube Shorts with AI

Google added a “reimagine” option to YouTube Shorts Remix, letting users use Gemini Omni to restyle clips, alter contents, or insert themselves into other people’s videos, while creators can enable or disable reimagining for their uploads.

Why it matters: HKR-H/K/R all pass: Google is adding Gemini Omni to the Shorts remix flow with a named reimagine mode and creator controls. It is a meaningful platform feature, not a model release, so it sits in low featured.

May 20Wednesday

Hacker News front page

Show HN: Lance – Image/video generation and understanding in one model

ByteDance released Lance as a research project for image and video generation and understanding in one model; the RSS snippet states 3B active parameters, fewer than 128 GPUs used for training, and links to a homepage, arXiv paper, and Hugging Face model, while the post does not disclose benchmark results or licensing terms.

Why it matters: ByteDance’s Lance puts image/video generation and understanding in one model, with 3B active parameters and <128 GPUs for training. HKR-H/K/R all pass, but benchmarks, license details, and real outputs are not disclosed, keeping it below P1.

The Verge · AI

It’s Make-or-Break Time for AI Labeling Systems

Google announced at I/O an expanded ability to verify SynthID markers on AI-generated images, while C2PA Content Credentials also targets origin metadata for image, video, and audio files; the RSS snippet does not disclose the full rollout scope or verification limits.

Why it matters: HKR-H/K/R all pass, but the post lacks full coverage scope, rollout terms, and adoption data. This is a mid-weight Google/C2PA provenance update, not a must-write release.

Xinzhiyuan · WeChat

UISEE Lists in Hong Kong as a Full-Scenario L4 Autonomous Driving Stock

UISEE listed on the Hong Kong Stock Exchange at HK$60.30 per share, with its public offering oversubscribed 6,777.29 times and a 90.5% share of the Greater China airport L4 commercial vehicle market in 2025.

Why it matters: HKR-H/K/R all pass: the IPO hook is concrete, with subscription and market-share numbers, and it ties to AV commercialization. It stays below 85 because this is not a foundation-model company IPO.

AI HOT (Curated Pool)

Kling AI Launches the First Native 4K Video Generation Model

Kling AI launched a native 4K video generation model on April 23, supporting one-click true 4K video generation; the post says Hollywood teams and Wonder Studios have adopted it, but does not disclose pricing, inference cost, or access limits.

Why it matters: HKR-H and HKR-K pass: Kling AI’s native 4K video model has a concrete capability and named adoption. Source is product-side, with no benchmark, pricing, or clip-duration data, so it sits at the featured threshold.

Synced · WeChat

After I/O, Google turns the search box into an agent entry point

Google announced Gemini 3.5 Flash at I/O and added AI Mode directly to Search; the company said its AI services now process over 3.2 quadrillion tokens per month, with more than 8.5 million developers using Gemini.

Why it matters: HKR-H/K/R all pass: Google I/O combines a model update, Search distribution, and concrete usage numbers. AI Mode inside the search box is heavier than a routine feature release, so it clears the same-day must-write band.

Latent Space

Google I/O 2026: Gemini 3.5 Flash, Omni, Spark, and Antigravity 2.0

Google announced Gemini 3.5 Flash at I/O 2026 with a 1M-token context window, 65k max output, four thinking levels, and pricing of $1.50 per 1M input tokens and $9.00 per 1M output tokens.

Why it matters: HKR-H/K/R all pass: this is a Google I/O model-and-product bundle with concrete context, output, thinking-tier, and pricing facts. It has same-day relevance for Claude, OpenAI, and coding-agent competition, so it clears P1.

AI HOT (Curated Pool)

Qwen3.7: Agent Frontier

Qwen Studio released Qwen3.7 with chatbots, image and video understanding, and image generation. It also covers document processing, web search integration, tool calling, and artifact generation. The RSS snippet frames it as an agent-focused model, but the post does not disclose context length. It also omits benchmark scores, pricing, API limits, release schedule, and reproducible evaluation conditions.

Why it matters: HKR-H/K/R all pass: this is a Qwen flagship-model update with concrete capability coverage. Lack of benchmarks, pricing, and context-window details keeps it at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini Omni Supports Video Creation With Personal Likeness and Voice

Gemini Omni lets users create digital-avatar videos using their personal likeness and voice, and the avatar can generate videos without uploading an image each time; the post does not disclose pricing, regions, or launch timing.

Why it matters: HKR-H/K/R pass: personal avatar video is clicky, reusable identity is a concrete mechanism, and voice/likeness raises creator and safety stakes. Price, regions, and launch timing are not disclosed, keeping it near the featured floor.

AI HOT (Curated Pool)

ChatGPT Image Generation Surpasses 1.5 Billion Uses Per Week

OpenAI says users generate more than 1.5 billion images per week in ChatGPT, and the post discusses new use cases and trends since the release of Images 2.0.

Why it matters: HKR-H/K/R all pass because OpenAI disclosed a concrete 1.5B-per-week image-generation usage figure. This is a strong adoption signal, but no new capability, pricing, or technical mechanism keeps it in the lower 78–84 band.

AI HOT (Curated Pool)

Google launches new AI search box with multimodal interactions

Google launched an AI search box based on Gemini 3.5, combining AI Overviews and AI Mode into one AI search experience that supports multimodal multi-turn queries across text, images, files, and video, with global availability on desktop and mobile.

Why it matters: HKR-H/K/R all pass: a Google Search entry-point update with Gemini 3.5, multimodal file/video queries, and AI Overviews/AI Mode integration. The source is thin, so it lands at the lower end of must-write.

AI HOT (Curated Pool)

Google Tensor ML SDK Beta Released

Google released the Tensor ML SDK beta, letting developers convert, compile, and run PyTorch or TFLite models on Pixel 10 TPUs through LiteRT, with a model library containing more than 100 classic and generative AI models, including Gemma 3.

Why it matters: HKR-K is strong: the post gives a concrete Pixel 10 TPU workflow and a 100+ model library. HKR-H/R clear the featured bar, but this is a beta developer SDK rather than a flagship model or major consumer launch.

AI HOT (Curated Pool)

Gemini Omni launches with physical reasoning and multimodal generation

Google launched Gemini Omni video generation for global AI Plus, Pro, and Ultra subscribers, integrating it with Gemini app, Google Flow, and YouTube Shorts, while the RSS snippet says the model combines intuitive physical reasoning with Gemini’s historical, scientific, and cultural knowledge.

Why it matters: HKR-H/K/R all pass: this is an official Google Gemini video-generation launch with named tiers and product surfaces. Details are thin—no benchmarks, pricing, or physical-reasoning tests—so it stays at the low end of the 85–94 band.