Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

281–300 of 514

May 19Tuesday

AI HOT (Curated Pool)

Qwen3.7 Preview lands on Arena; Alibaba rises to fifth in vision ranking

Alibaba says Qwen3.7-Plus-Preview has landed on Arena and that Alibaba now ranks fifth in vision; the post does not disclose benchmark scores, the number of competing models, or a release timeline for the Qwen3.7 series.

Why it matters: HKR-H/K/R pass: Qwen3.7-Plus-Preview appears on Arena with a #5 vision rank. Score stays in the low featured band because the vendor post omits scores, model count, access, and timeline.

AI HOT (Curated Pool)

NVIDIA fine-tunes Cosmos Predict 2.5 with LoRA/DoRA for robot video generation

NVIDIA published a Hugging Face post on fine-tuning Cosmos Predict 2.5 with LoRA and DoRA to generate robot first-person videos from text prompts; the post does not disclose dataset size, training cost, or evaluation results.

Why it matters: HKR-H/K/R pass: the robot POV video angle is clickable, and LoRA/DoRA on Cosmos Predict 2.5 is a concrete mechanism. Missing dataset scale and metrics keep it in the low featured band.

May 18Monday

Latent Space

The Autonomous Drone Tech Stack and Economics of Drones — Yaroslav Azhnyuk

Latent Space interviewed The Fourth Law founder Yaroslav Azhnyuk for a two-hour episode covering FPV drones, five levels of autonomy, eight dimensions of the autonomous battlefield, and China’s manufacturing advantage; the transcript claims Ukraine produced 4 million FPV drones last year and discusses a hypothetical Chinese capacity of 4 billion.

Why it matters: HKR-H/K/R all pass: the Latent Space interview offers concrete autonomy and battlefield frameworks. It is still commentary, not a model release, product update, or research artifact, so it stays just above the featured threshold.

r/LocalLLaMA

Qwen 3.6 27B on 24GB VRAM: backend comparisons, quant choice, and settings

The author tested Qwen 3.6 27B on an RTX 3090 24GB and kept ik_llama.cpp with Qwen3.6-27B-MTP-IQ4_KS.gguf; at 156k context with q8_0 KV and MTP, a ~5.9k-token prompt plus 1024-token output reached about 1261 tok/s prefill and 72.9 tok/s decode, while vLLM lacked a clean single-card long-context run.

Why it matters: HKR-H/K/R all pass: this is a first-person local inference benchmark with concrete VRAM, quant, context, and speed numbers. Its reach is narrower than a model release, so it sits at the featured threshold.

AI HOT (Curated Pool)

Grok Now Supports Video Understanding and Analysis

Grok now supports full-video uploads for real-time analysis, summarization, translation, scene explanation, and context extraction; the post does not disclose duration limits, supported formats, or rollout scope.

Why it matters: HKR-H/K/R all pass, but duration limits, formats, and rollout scope are not disclosed, so this stays at the featured threshold for a mid-weight product update.

AI HOT (Curated Pool)

Alibaba Cloud launches HappyHorse video generation model

Alibaba Cloud launched HappyHorse on Model Studio, with prompt-to-1080p multi-shot video generation in one workflow; the post lists a limited-time 20% discount but does not disclose pricing, model parameters, or availability terms.

Why it matters: HKR-H/K/R pass on the named model, 1080p multi-shot capability, and cost/competition angle. Thin disclosure on price, parameters, and benchmarks keeps it near the featured threshold.

Google DeepMind

Google DeepMind releases Gemini Omni Flash video model

Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family. It combines image, audio, video and text inputs to generate high-quality video, and supports multi-turn editing in natural language.

Why it matters: Gemini Omni Flash folds video generation and conversational editing into one model, a shift in how multimodal creation gets accessed.

May 17Sunday

Bloomberg Technology

Apple’s New ChatGPT-Like Siri App Will Have Auto-Deleting Chats

The title says Apple’s ChatGPT-like Siri app will support auto-deleting chats; the RSS snippet only adds that iOS 27 will include a Genmoji upgrade, and the post does not disclose retention periods, release timing, or feature details.

Why it matters: HKR-H and HKR-R pass because Bloomberg frames a specific Apple Siri privacy angle; HKR-K fails since retention and feature mechanics are missing, so this stays at the low featured threshold.

Google DeepMind

Google expands content provenance and verification tools across Search, Gemini, Chrome and Pixel

Google is widening its content transparency and verification tools across Search, Gemini, Chrome, Pixel and Cloud, and deepening industry partnerships. SynthID has watermarked over 100 billion images and videos plus 60,000 years of audio. SynthID verification in the Gemini app has been used 50 million times, and the capability reaches Search today, with Chrome in the coming weeks.

Why it matters: The post lays out where SynthID and C2PA land across Search, Gemini, Chrome and Pixel, which shows the current limits of content provenance tools.

QbitAI · WeChat

TGO Aligns Visual Generative Models with Scalar Feedback Without Preference Pairs | ICML 2026

NUS proposed Threshold-Guided Optimization, which converts scalar feedback into positive or negative updates through a score-distribution threshold and was accepted by ICML 2026; experiments cover Stable Diffusion v1.5, FLUX, Wan 1.3B, and Meissonic across image and video generation settings.

Why it matters: HKR-H/K/R pass: the paper has a concrete mechanism and tests across SD v1.5, FLUX, Wan 1.3B, and Meissonic. Impact is research-heavy, so it lands in featured, not must-write.

QbitAI · WeChat

A Robot Dog Challenges Nvidia's Compute Lead

Weilan Technology unveiled BabyAlpha A3, a consumer quadruped robot using a six-chip heterogeneous cluster that runs a 7B-parameter model on-device at 280 TPS; the article says it has 66MP vision, 2.232 million point-cloud samples per second, and a planned Q3 launch.

Why it matters: HKR-H/K/R pass: the robot-dog-versus-Nvidia angle is clickable, and 280 TPS on a local 7B model is concrete. Single-source summary lacks price, power draw, and benchmark setup, so it stays near the featured floor.

AI HOT (Curated Pool)

Grok Imagine image generation is officially released

Grok Imagine is now available on X for all users, with text-to-image generation for realistic images and multiple aspect ratios; the post does not disclose model parameters, pricing, or regional limits.

Why it matters: HKR-H/K/R pass, but the post only discloses availability and basic image features; model details, pricing, and regions are absent, so this lands at the featured threshold.

Financial Times · Technology

Chinese AI Groups Pull Ahead of US Rivals in Video Generation Race

FT says Chinese AI groups have moved ahead of US rivals in video generation; the RSS snippet names ByteDance and Kuaishou and says they outshine western competitors in advertising and entertainment quality, but the post does not disclose benchmark metrics or model details.

Why it matters: FT authority plus a China-vs-US video-generation lead claim clears HKR-H and HKR-R. HKR-K fails because the body lacks metrics, samples, and eval method, so it sits at the low featured threshold.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

May 16Saturday

Hacker News front page

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

SANA-WM’s title says the project is a 2.6B open-source world model for 1-minute 720p video; the RSS body only lists the project URL, Hacker News comments URL, 9 points, and 8 comments, and the post does not disclose training data, license terms, inference cost, evaluation setup, or benchmark results.

Why it matters: HKR-H/K/R pass on the concrete open-source world-model hook, 2.6B size, and video-model competition angle. Sparse body details keep it at the lower good-quality band.

Synced · WeChat

Why Robots Need World Models: Top Institutions Release Joint Survey

NTU MARS Lab and collaborators released a 43-page survey on robot world models, covering definitions, architectures, applications, benchmarks, and challenges around action-conditioned consistency, inference efficiency, and physical grounding.

Why it matters: HKR-H and HKR-K pass: the hook is robot world models, and the post cites a 43-page survey with benchmarks and action-consistency framing. HKR-R is weak, so this stays at the featured threshold.

QbitAI · WeChat

Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1

Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.

Why it matters: HKR-H/K/R all pass: the hook is unusual, and the post gives 3,000 text samples, Flow-GRPO, and a +10.23 dB PSNR gain. Strong multimodal research, but not a foundation-model launch, so 78.

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

AI HOT (Curated Pool)

Runway Agent Generates Complete Ads in One Session

Runway Agent turns product photos and ideas into fully produced ads in one session; the post does not disclose the model, pricing, generation length, or regional availability.

Why it matters: Runway’s ad-generation Agent clears HKR-H/K/R as a mid-weight product update. Missing model, pricing, duration, and region details keep it at the featured threshold, not a must-write release.

May 15Friday

r/LocalLLaMA

Fully Offline Suitcase Robot Built Around Jetson Orin NX SUPER 16GB

CreativelyBankrupt built Sparky as a fully offline suitcase robot on Jetson Orin NX SUPER 16GB, running Gemma 4 E4B Q4_K_M via llama.cpp with q8_0 KV cache, about 200 ms cached TTFT, 14-15 tok/s sustained output, 12K context, 30+ sensors, and no WiFi, Bluetooth, or cellular interface.

Why it matters: HKR-H/K/R all pass, with a named hands-on build and concrete latency/sensor numbers. It stays in low featured because this is a Reddit project post, not a product launch or research release.