Skip to content

#多模态

1 today

May 21Thursday

The Verge · AI

You can now remix other people’s YouTube Shorts with AI

Google added a “reimagine” option to YouTube Shorts Remix, letting users use Gemini Omni to restyle clips, alter contents, or insert themselves into other people’s videos, while creators can enable or disable reimagining for their uploads.

Why it matters: HKR-H/K/R all pass: Google is adding Gemini Omni to the Shorts remix flow with a named reimagine mode and creator controls. It is a meaningful platform feature, not a model release, so it sits in low featured.

May 20Wednesday

Hacker News front page

Show HN: Lance – Image/video generation and understanding in one model

ByteDance released Lance as a research project for image and video generation and understanding in one model; the RSS snippet states 3B active parameters, fewer than 128 GPUs used for training, and links to a homepage, arXiv paper, and Hugging Face model, while the post does not disclose benchmark results or licensing terms.

Why it matters: ByteDance’s Lance puts image/video generation and understanding in one model, with 3B active parameters and <128 GPUs for training. HKR-H/K/R all pass, but benchmarks, license details, and real outputs are not disclosed, keeping it below P1.

The Verge · AI

It’s Make-or-Break Time for AI Labeling Systems

Google announced at I/O an expanded ability to verify SynthID markers on AI-generated images, while C2PA Content Credentials also targets origin metadata for image, video, and audio files; the RSS snippet does not disclose the full rollout scope or verification limits.

Why it matters: HKR-H/K/R all pass, but the post lacks full coverage scope, rollout terms, and adoption data. This is a mid-weight Google/C2PA provenance update, not a must-write release.

Xinzhiyuan · WeChat

UISEE Lists in Hong Kong as a Full-Scenario L4 Autonomous Driving Stock

UISEE listed on the Hong Kong Stock Exchange at HK$60.30 per share, with its public offering oversubscribed 6,777.29 times and a 90.5% share of the Greater China airport L4 commercial vehicle market in 2025.

Why it matters: HKR-H/K/R all pass: the IPO hook is concrete, with subscription and market-share numbers, and it ties to AV commercialization. It stays below 85 because this is not a foundation-model company IPO.

AI HOT (Curated Pool)

Kling AI Launches the First Native 4K Video Generation Model

Kling AI launched a native 4K video generation model on April 23, supporting one-click true 4K video generation; the post says Hollywood teams and Wonder Studios have adopted it, but does not disclose pricing, inference cost, or access limits.

Why it matters: HKR-H and HKR-K pass: Kling AI’s native 4K video model has a concrete capability and named adoption. Source is product-side, with no benchmark, pricing, or clip-duration data, so it sits at the featured threshold.

Synced · WeChat

After I/O, Google turns the search box into an agent entry point

Google announced Gemini 3.5 Flash at I/O and added AI Mode directly to Search; the company said its AI services now process over 3.2 quadrillion tokens per month, with more than 8.5 million developers using Gemini.

Why it matters: HKR-H/K/R all pass: Google I/O combines a model update, Search distribution, and concrete usage numbers. AI Mode inside the search box is heavier than a routine feature release, so it clears the same-day must-write band.

Latent Space

Google I/O 2026: Gemini 3.5 Flash, Omni, Spark, and Antigravity 2.0

Google announced Gemini 3.5 Flash at I/O 2026 with a 1M-token context window, 65k max output, four thinking levels, and pricing of $1.50 per 1M input tokens and $9.00 per 1M output tokens.

Why it matters: HKR-H/K/R all pass: this is a Google I/O model-and-product bundle with concrete context, output, thinking-tier, and pricing facts. It has same-day relevance for Claude, OpenAI, and coding-agent competition, so it clears P1.

AI HOT (Curated Pool)

Qwen3.7: Agent Frontier

Qwen Studio released Qwen3.7 with chatbots, image and video understanding, and image generation. It also covers document processing, web search integration, tool calling, and artifact generation. The RSS snippet frames it as an agent-focused model, but the post does not disclose context length. It also omits benchmark scores, pricing, API limits, release schedule, and reproducible evaluation conditions.

Why it matters: HKR-H/K/R all pass: this is a Qwen flagship-model update with concrete capability coverage. Lack of benchmarks, pricing, and context-window details keeps it at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini Omni Supports Video Creation With Personal Likeness and Voice

Gemini Omni lets users create digital-avatar videos using their personal likeness and voice, and the avatar can generate videos without uploading an image each time; the post does not disclose pricing, regions, or launch timing.

Why it matters: HKR-H/K/R pass: personal avatar video is clicky, reusable identity is a concrete mechanism, and voice/likeness raises creator and safety stakes. Price, regions, and launch timing are not disclosed, keeping it near the featured floor.

AI HOT (Curated Pool)

ChatGPT Image Generation Surpasses 1.5 Billion Uses Per Week

OpenAI says users generate more than 1.5 billion images per week in ChatGPT, and the post discusses new use cases and trends since the release of Images 2.0.

Why it matters: HKR-H/K/R all pass because OpenAI disclosed a concrete 1.5B-per-week image-generation usage figure. This is a strong adoption signal, but no new capability, pricing, or technical mechanism keeps it in the lower 78–84 band.

AI HOT (Curated Pool)

Google launches new AI search box with multimodal interactions

Google launched an AI search box based on Gemini 3.5, combining AI Overviews and AI Mode into one AI search experience that supports multimodal multi-turn queries across text, images, files, and video, with global availability on desktop and mobile.

Why it matters: HKR-H/K/R all pass: a Google Search entry-point update with Gemini 3.5, multimodal file/video queries, and AI Overviews/AI Mode integration. The source is thin, so it lands at the lower end of must-write.

AI HOT (Curated Pool)

Google Tensor ML SDK Beta Released

Google released the Tensor ML SDK beta, letting developers convert, compile, and run PyTorch or TFLite models on Pixel 10 TPUs through LiteRT, with a model library containing more than 100 classic and generative AI models, including Gemma 3.

Why it matters: HKR-K is strong: the post gives a concrete Pixel 10 TPU workflow and a 100+ model library. HKR-H/R clear the featured bar, but this is a beta developer SDK rather than a flagship model or major consumer launch.

AI HOT (Curated Pool)

Gemini Omni launches with physical reasoning and multimodal generation

Google launched Gemini Omni video generation for global AI Plus, Pro, and Ultra subscribers, integrating it with Gemini app, Google Flow, and YouTube Shorts, while the RSS snippet says the model combines intuitive physical reasoning with Gemini’s historical, scientific, and cultural knowledge.

Why it matters: HKR-H/K/R all pass: this is an official Google Gemini video-generation launch with named tiers and product surfaces. Details are thin—no benchmarks, pricing, or physical-reasoning tests—so it stays at the low end of the 85–94 band.

AI HOT (Curated Pool)

Gemini Omni: A new model for creating content from any input

Google AI published a Gemini Omni post, saying the model creates content from any input starting with video; the post does not disclose parameters, pricing, availability timing, or benchmark results.

Why it matters: Google AI’s post gives Gemini Omni and “content from any input” as hooks, so HKR-H/R pass; HKR-K fails because price, timing, parameters, and benchmarks are absent, keeping it at low featured.

AI HOT (Curated Pool)

Google releases Gemini Omni multimodal generation model

Google released Gemini Omni, a multimodal generation model that combines image, video, and text inputs to generate videos grounded in Gemini’s real-world knowledge; the post does not disclose model size, pricing, or availability.

Why it matters: HKR-H/K/R all pass: Google’s Gemini Omni adds combined image, video, and text inputs for video generation. Missing parameters, pricing, and rollout timing keep it at the low end of the must-write band.

AI HOT (Curated Pool)

Google I/O Announces Multiple Gemini Updates

Google announced multiple Gemini updates at Google I/O, including a new experience design using neural expression technology, upcoming agent features with Daily Brief and Gemini Spark, plus Gemini Omni and 3.5 Flash; the RSS snippet does not disclose release dates, model parameters, pricing, or benchmark results.

Why it matters: HKR-H/K/R all pass: Google I/O brings concrete Gemini hooks in Omni, 3.5 Flash, and agents, with clear competitive resonance. Missing launch timing, specs, and pricing keep it in the 78–84 band.

r/LocalLLaMA

Nemotron-Labs-Diffusion from NVIDIA

NVIDIA released the Nemotron-Labs-Diffusion 3B, 8B, and 14B dense model family with AR decoding, diffusion parallel decoding, and self-speculation; the 8B model reaches 850 tok/s on GB200 at concurrency 1, compared with 253 tok/s for AR and 360 tok/s for Eagle3.

Why it matters: HKR-H/K/R all pass: NVIDIA diffusion LLMs, concrete sizes/mechanisms, and an 850 tok/s GB200 claim. Single-source Reddit sourcing keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Google Search gets its biggest redesign in 25 years with AI-driven interaction changes

Google announced at I/O 2026 the biggest Search redesign in 25 years, powered by Gemini 3.5 Flash, with AI Mode exceeding 1 billion monthly active users and query volume doubling each quarter.

Why it matters: Google Search is an internet entry-point product; 1B+ AI Mode MAU and quarterly query doubling put this beyond a routine feature update. HKR-H/K/R all pass, so it lands in p1.

The Verge · AI

Gemini will use Volvo’s external cameras to interpret parking signs

Google and Volvo announced at I/O that Gemini will access external cameras on the upcoming EX60 SUV, with the first stated use case translating hard-to-understand parking signs for vehicle owners.

Why it matters: HKR-H and HKR-K pass: Gemini tying into Volvo EX60 exterior cameras is a concrete multimodal in-car use case. HKR-R is weak because rollout scope, privacy, and safety mechanisms are not disclosed.

The Verge · AI

The 13 Biggest Announcements at Google I/O 2026

Google announced Gemini 3.5 Flash at I/O 2026. It becomes the default model today for the Gemini app and AI Mode in Search, while Gemini 3.5 Pro follows next month; the RSS snippet also mentions Search, Gmail, and Project Aura smart glasses updates but does not disclose the full list of 13 announcements.

Why it matters: HKR-H/K/R all pass, but the text only gives Gemini 3.5 Flash default rollout and Pro timing; it lacks the full 13 items, benchmarks, or pricing, so this stays featured below p1.

AI HOT (Curated Pool)

Google releases Gemini 3.5 Flash with 55 intelligence score

Google released Gemini 3.5 Flash, raising its intelligence score by 9 points to 55, exceeding 280 output tokens per second, and increasing operating cost by 5.5 times versus the previous generation.

Why it matters: A Google Gemini 3.5 Flash release is a top-lab model update, backed by Artificial Analysis numbers for speed, intelligence, and cost. HKR-H/K/R all pass, with the 5.5x cost jump making it more than a routine launch.

TechCrunch · AI

Google’s Genie world model can now simulate real streets with Street View

Google DeepMind is integrating Street View with Project Genie for interactive street-level simulations in robotics, gaming, and travel; the post does not disclose model parameters, launch timing, or evaluation results.

Why it matters: Google DeepMind connecting Street View to Genie gives HKR-H/K/R: a novel hook, a concrete mechanism, and robotics/data-moat resonance. Missing params, launch timing, and evals keep it in the 78–84 band.

TechCrunch · AI

OpenAI is making it easier to check if an image was made by its models

OpenAI announced two measures for detecting AI-generated images: it joined the open C2PA standard and added Google’s SynthID to its products. The RSS snippet does not disclose which OpenAI products include SynthID, whether the checks cover legacy images, or when the measures become available to users.

Why it matters: HKR-H/K/R all pass: the OpenAI-Google provenance tie-up is clickable, and C2PA plus SynthID are concrete mechanisms. Coverage, launch timing, and verification flow are not disclosed, so this stays at the featured threshold.

TechCrunch · AI

Google's Gemini Omni turns images, audio, and text into video

Google's Gemini Omni generates and edits video through conversation, using text, images, audio, and video as inputs, with Omni Flash named as the starting version; the RSS snippet says the model reasons across modalities, but the post does not disclose launch date, pricing, context limits, benchmarks, or API availability.

Why it matters: Google-scale Gemini multimodal video update clears HKR-H/K/R: Omni Flash, chat-based editing, and four input types are concrete. Pricing and rollout are not disclosed, so it sits in the lower must-write band.

AI HOT (Curated Pool)

Google releases Gemini Omni for any-input-to-any-output generation and conversational video editing

Google released Gemini Omni and Omni Flash at I/O 2026, supporting text, image, audio, and video inputs and outputs, with conversational video editing; Omni Flash is available in Gemini App, Google Flow, and YouTube Shorts, while the post does not disclose the API launch date.

Why it matters: HKR-H/K/R all pass: Google announced Gemini Omni and Omni Flash as a major multimodal update at I/O. API timing, pricing, and benchmarks are not disclosed, so it stays below 90.

AI HOT (Curated Pool)

Gemini Omni Released: New Progress in Multimodal Generation

Google DeepMind released Gemini Omni, starting with video generation; the post says it combines Gemini with its generative media systems, but does not disclose parameters, pricing, or availability.

Why it matters: HKR-H/K/R pass: a DeepMind Gemini-branded video-generation launch has a clear hook, a stated mechanism, and competitive resonance. Missing parameters, pricing, and access timing keep it in the 72–77 band.

AI HOT (Curated Pool)

Google processes over 3,200 trillion tokens per month, up 7x year over year

Google said at I/O 2026 that it processed over 3,200 trillion tokens per month in May, while Gemini App exceeded 900 million monthly active users and the Nano Banana model generated more than 50 billion images cumulatively.

Why it matters: Google I/O disclosed usage scale, not a new model or major capability. HKR-H/K/R pass via 7x growth, 900M MAU, and 3.2Q tokens/month, but without a launch-level update it stays in the 78–84 band.

AI HOT (Curated Pool)

NVIDIA open-sources first 4-bit infrastructure for ultra-long video generation

NVIDIA researchers open-sourced LongLive 2.0, an end-to-end long-video generation infrastructure covering training and inference with 4-bit quantization, FP4 quantization, parallel acceleration, KV-cache optimization, and 45.7 FPS generation on a 5B model.

Why it matters: HKR-H/K/R all pass: NVIDIA researcher open-sources LongLive 2.0 with 4-bit long-video train/inference and 45.7 FPS on a 5B model. This is strong open-source infra, not a flagship model launch, so it fits the 78–84 band.

May 19Tuesday

r/LocalLLaMA

ByteDance released an open-source model that attempts broad multimodal tasks with 3B parameters

ByteDance released Lance, an open-source unified multimodal model with 3B active parameters that supports image and video understanding, generation, and editing, and the post says it was trained from scratch with a staged multi-task recipe under a 128-A100-GPU budget.

Why it matters: HKR-H/K/R all pass: ByteDance’s open Lance has a compact multimodal hook, concrete 3B/128-A100 facts, and clear cost/deployment resonance. Reddit-sourced details lack benchmarks, license terms, and official context, so it stays featured, not P1.

AI HOT (Curated Pool)

Advancing content provenance for a safer, more transparent AI ecosystem

OpenAI launched an AI content provenance system that combines Content Credentials and SynthID with a verification tool; the post does not disclose supported media formats, rollout scope, or detection accuracy.

Why it matters: HKR-H/K/R pass: the OpenAI provenance stack has a concrete cross-standard mechanism and trust/compliance relevance. Missing format coverage, rollout scope, and accuracy keep it in the low featured band.

QbitAI · WeChat

Chinese GPU vendor Moore Threads releases MT Lambda for embodied AI simulation

Moore Threads released MT Lambda, an embodied AI simulation platform that combines physics, rendering, and AI engines, and demonstrated the robot dog “Xiaofei” executing a Sim-to-Real policy trained 100% in simulation on domestic hardware.

Why it matters: HKR-H/K/R pass: the story has a concrete domestic-GPU simulation hook, a three-engine mechanism, and a clear NVIDIA/robotics-cost nerve. Importance stays in the low featured band because performance, pricing, access, and third-party validation are not disclosed.

QbitAI · WeChat

World model supports multiplayer FPS gameplay before Fei-Fei Li

Odyssey released Agora-1, a world model that supports up to four human and AI players fighting in the same generated FPS world in real time. The system decouples simulation from rendering and trains on GoldenEye internal game states.

Why it matters: HKR-H/K/R all pass: Agora-1 moves world models from solo demos to up to 4-player real-time FPS, with decoupled simulation/rendering and training-data clues. The lab is not a top-tier foundation-model vendor, so this stays in the 78–84 band.

New York Times Chinese

China’s AI Microdrama Boom Brings Job Anxiety and Tech Enthusiasm

Chinese companies are producing AI-generated microdramas for about $30 per minute without cameras, crews, or human actors; DataEye says nearly 50,000 new AI microdramas were uploaded to Douyin in March, almost matching the platform’s total uploads for all of 2025.

Why it matters: HKR-H/K/R all pass: the backlash angle is clickable, the story adds $30-per-minute production and nearly 50,000 March uploads, and it hits labor anxiety. It is strong industry reporting, not a core model or product release.

AI HOT (Curated Pool)

Qwen3.7 Preview lands on Arena; Alibaba rises to fifth in vision ranking

Alibaba says Qwen3.7-Plus-Preview has landed on Arena and that Alibaba now ranks fifth in vision; the post does not disclose benchmark scores, the number of competing models, or a release timeline for the Qwen3.7 series.

Why it matters: HKR-H/K/R pass: Qwen3.7-Plus-Preview appears on Arena with a #5 vision rank. Score stays in the low featured band because the vendor post omits scores, model count, access, and timeline.

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 18Monday

Latent Space

The Autonomous Drone Tech Stack and Economics of Drones — Yaroslav Azhnyuk

Latent Space interviewed The Fourth Law founder Yaroslav Azhnyuk for a two-hour episode covering FPV drones, five levels of autonomy, eight dimensions of the autonomous battlefield, and China’s manufacturing advantage; the transcript claims Ukraine produced 4 million FPV drones last year and discusses a hypothetical Chinese capacity of 4 billion.

Why it matters: HKR-H/K/R all pass: the Latent Space interview offers concrete autonomy and battlefield frameworks. It is still commentary, not a model release, product update, or research artifact, so it stays just above the featured threshold.

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

AI HOT (Curated Pool)

Grok Now Supports Video Understanding and Analysis

Grok now supports full-video uploads for real-time analysis, summarization, translation, scene explanation, and context extraction; the post does not disclose duration limits, supported formats, or rollout scope.

Why it matters: HKR-H/K/R all pass, but duration limits, formats, and rollout scope are not disclosed, so this stays at the featured threshold for a mid-weight product update.

AI HOT (Curated Pool)

Alibaba Cloud launches HappyHorse video generation model

Alibaba Cloud launched HappyHorse on Model Studio, with prompt-to-1080p multi-shot video generation in one workflow; the post lists a limited-time 20% discount but does not disclose pricing, model parameters, or availability terms.

Why it matters: HKR-H/K/R pass on the named model, 1080p multi-shot capability, and cost/competition angle. Thin disclosure on price, parameters, and benchmarks keeps it near the featured threshold.

Google DeepMind

Google DeepMind releases Gemini Omni Flash video model

Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family. It combines image, audio, video and text inputs to generate high-quality video, and supports multi-turn editing in natural language.

Why it matters: Gemini Omni Flash folds video generation and conversational editing into one model, a shift in how multimodal creation gets accessed.