Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

161–180 of 514

Jun 4Thursday

Synced · WeChat

Google releases Gemma 4 12B for 16GB laptops

Google released Gemma 4 12B, a medium-size model that runs locally with 16GB VRAM or unified memory. It uses an encoder-free multimodal architecture, supports native audio input, ships under Apache 2.0, and includes an MTP draft model for lower latency.

Why it matters: Google’s Gemma 4 12B has clear HKR-H/K/R: 16GB local running, 12B scale, and Apache 2.0 licensing. It is a strong open-model update, not a must-write foundation-model launch.

QbitAI · WeChat

CVPR 2026: NVIDIA, Tesla, and Waymo hear Xpeng present physical AI

Xpeng presented its world-model stack at CVPR 2026, covering X-World, X-Foresight, and X-Cache; the article says X-Cache cuts about 70% of repeated computation, the second-generation VLA used over 4 trillion training tokens, and the in-car stack reduced inference latency to 80 ms.

Why it matters: HKR-H comes from the CVPR stage contrast, HKR-K has X-Cache, 4T+ tokens, and 80 ms latency, and HKR-R fits autonomy competition. It is still a company tech showcase, below the 85 must-write band.

AI HOT (Curated Pool)

Ideogram 4.0 Open-Source Text-to-Image Model Released

Ideogram released Ideogram 4.0, an open-source text-to-image model with a 9.3B-parameter core, a single-stream DiT architecture, Qwen3-VL-8B-Instruct text encoder, and a No. 4 ranking in DesignArena human evaluation.

Why it matters: HKR-H/K/R all pass: Ideogram 4.0 brings open weights, 9.3B parameters, single-stream DiT, and a No. 4 human-eval rank. It is strong open image-model signal, not a top-tier general-model launch.

Financial Times · Technology

MP sues Musk’s xAI in UK test case over fake sexual images

UK MP Jess Asato sued Musk’s xAI over fake sexual images, using the claim to test whether AI model makers are liable for system outputs; the post does not disclose the model, generation mechanism, damages sought, or court timetable.

Why it matters: HKR-H/K/R all pass: FT ties xAI, Musk, fake sexual images, and a UK liability test. The article does not disclose the model, generation mechanism, or damages, so it stays in the 78–84 band.

Hacker News front page

Gemma 4 12B: A Unified, Encoder-Free Multimodal Model

Google’s title introduces Gemma 4 12B as a unified, encoder-free multimodal model; the RSS snippet only lists 137 Hacker News points and 48 comments, and the post does not disclose architecture details, training setup, pricing, release terms, or benchmark results.

Why it matters: HKR-H/K/R pass: Google names Gemma 4 12B and an encoder-free multimodal design, a strong hook for open-model practitioners. The post lacks training details, pricing, and benchmarks, so it stays in the low 78–84 band, not P1.

Jun 3Wednesday

r/LocalLLaMA

google/gemma-4-12B on Hugging Face

Google DeepMind released Gemma 4 open-weight models in five sizes, with the 12B variant supporting text, image, and audio input, instruction-tuned and pre-trained variants, native system prompts, function calling, and a context window of up to 256K tokens.

Why it matters: Gemma 4 clears HKR-H/K/R: open weights, multimodal input, and 256K context make it more than a routine update. Missing benchmarks, license detail, and fuller official context keep it in the 78–84 band.

MIT Technology Review · AI

The Download: Trump’s New AI Order, and Smart Glasses for Warfare

President Donald Trump signed a new AI order asking companies to voluntarily submit frontier models for government review 30 days before release, without mandatory licensing; the newsletter also says Anduril and Meta are prototyping a military AR headset that envisions drone-strike orders through eye tracking and voice commands.

Why it matters: HKR-H/K/R all pass: the article gives a concrete 30-day frontier-model review mechanism and a Meta/Anduril AR warfare prototype. A presidential AI order affecting release compliance clears the must-write band.

Synced · WeChat

RSS 2026: Ant Lingbo Proposes Autoregressive Causal World Model for Robot Manipulation with 50 Demos

Ant Lingbo and HKUST introduced LingBot-VA, an autoregressive video-action world model that unifies visual dynamics prediction and action inference, and the paper reports fine-tuning with 50 real-world demonstrations per task plus 92.0% and 91.1% success on RoboTwin 2.0 Easy and Hard settings.

Why it matters: HKR-H/K/R all pass: the hook is 50-demo robot control, with a concrete video-action world-model mechanism. Single-source coverage lacks code, benchmark detail, and deployment evidence, so it lands at 78.

Latent Space

[AINews] Microsoft Build: MAI-Thinking-1 and MAI Family Models

Microsoft announced seven MAI models at Build, with MAI-Thinking-1 described as a 35B-active-parameter MoE with a 256K context window, and released a 109-page technical report covering training, data lineage, and performance claims.

Why it matters: All HKR axes pass: Microsoft’s MAI family has concrete specs, a long technical report, and clear competitive stakes around its model stack. This clears the 85+ same-day bar, but no weights, pricing, or external evals are disclosed, so it lands at 87.

QbitAI · WeChat

Daxiao Robot and NTU Release PhysX-Omni for Unified Physical 3D Generation

Daxiao Robot and NTU introduced PhysX-Omni, a unified simulation-ready physical 3D generation framework for rigid, deformable, and articulated objects, with PhysXVerse covering 8.7K assets across 2.9K categories and PhysX-Bench evaluating six dimensions including geometry, scale, material, affordance, kinematics, and description.

Why it matters: HKR-H/K/R all pass: unified physical 3D generation is a clear hook, the dataset and benchmark numbers add substance, and robotics simulation data is a real practitioner pain. No open-source or product adoption is disclosed, so it stays at 78.

AI HOT (Curated Pool)

xAI releases Grok Imagine 1.5 preview image-to-video model

xAI released grok-imagine-video-1.5-preview via its API, letting users turn one still image into 720p video while controlling camera movement, pacing, and sound effects with natural-language prompts.

Why it matters: HKR-H/K/R all pass: xAI ships a named image-to-video API preview with 720p output and sound controls. It stays below 85 because this is a preview product update, not a flagship foundation-model release.

AI HOT (Curated Pool)

Runway API adds Aleph 2.0 video editing

Runway API now provides Aleph 2.0 video editing for integration into apps, products, and platforms, supporting precise edits on multi-shot videos up to 30 seconds at 1080p while changing only selected portions; the post does not disclose pricing, rate limits, latency, or model availability by region.

Why it matters: Runway is a core AI-video player, and Aleph 2.0 exposes partial video editing via API with 30s and 1080p limits. HKR-H/K/R all pass, but this is a mid-weight product update, not a model-class release.

r/LocalLLaMA

Using Gemma 4 E4B with LiteRT: about 2.4× faster text generation than Q4 GGUF

The author tested Gemma 4 E4B on an RTX 4060 Ti 16GB, where LiteRT averaged 157.2 tok/s for text generation versus 66.3 tok/s for llama.cpp Q4 GGUF; image captioning on 111 full-resolution images improved only 1.1×, at about 72 seconds versus 80 seconds.

Why it matters: HKR-H/K/R all pass, with a first-person benchmark including hardware, throughput, and sample count. Source authority is limited to one Reddit test, so it sits at the featured threshold rather than the 78+ band.

The Verge · AI

Microsoft’s Project Solara is an OS for AI agent gadgets

Microsoft announced Project Solara at Build 2026 as an Android-based OS for AI agent gadgets, not Windows, and the post discloses two concept devices: a desk device with facial recognition and a wearable badge with a camera and fingerprint scanner.

Why it matters: HKR-H/K/R all pass: Project Solara ties Microsoft, Android, and agent gadgets together, with two concrete hardware concepts. Score stays below P1 because shipping date, developer APIs, and pricing are not disclosed.

Jun 2Tuesday

AI HOT (Curated Pool)

StepFun releases Step 3.7 Flash as an open-weight model for agentic coding

StepFun released the open-weight Step 3.7 Flash model for fast agentic coding, with tool calling and multimodal understanding, and the model is already available in Kilo alongside MiniMax M3.

Why it matters: HKR-H/K/R pass on the open-weight agentic-coding angle and Kilo availability. Missing benchmarks, size, license, and pricing keep it at the lower featured threshold.

r/LocalLLaMA

NVIDIA releases Cosmos 3 Omnimodal world models on Hugging Face

NVIDIA released Cosmos 3 on Hugging Face with Nano at 16B parameters and Super at 64B parameters; the post says the models generate video, images, audio, and action commands from text, image, video, and action-trajectory inputs.

Why it matters: HKR-H/K/R all pass: NVIDIA world models on HF, concrete 16B/64B variants, and multimodal robotics relevance. Missing benchmarks, license, and training details keep it in the 78–84 band.

QbitAI · WeChat

Jensen Huang Brings NVIDIA CPUs Into the PC Market

NVIDIA RTX Spark will ship in Windows PCs this fall with 1 petaflop of AI compute and 128GB unified memory. The platform combines a Blackwell RTX GPU, a 20-core Arm-based Grace CPU, and NVLink-C2C, and NVIDIA says it can run 1-million-token-context, 120B-parameter language models locally.

Why it matters: HKR-H/K/R all pass: NVIDIA is moving RTX Spark into Windows PCs with concrete specs: 1 petaflop, 128GB unified memory, 1M context, and 120B local models. This is a strong hardware product update, not a foundation-model release, so it lands in 78–84.

Latent Space

[AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark

NVIDIA released Cosmos 3 and Nemotron 3 Ultra; Cosmos 3 uses a Mixture-of-Transformers design with 16B Nano and 64B Super variants, while Nemotron 3 Ultra is described as a 550B-A55B open-weight model.

Why it matters: HKR-H/K/R all pass: NVIDIA ships Cosmos 3, Nemotron 3 Ultra, and RTX Spark with concrete MoT, 16B/64B, and 550B-A55B open-weight details. Impact is broad, but below a frontier-lab model release.

AI HOT (Curated Pool)

NVIDIA Cosmos 3 Tops Open-Weight Image and Video Generation Rankings

NVIDIA Cosmos 3 ranked first in Artificial Analysis’s open-weight text-to-image and image-to-video categories, with 16B Nano and 64B Super variants, and the release includes weights, code, curated datasets, and fine-tuning recipes under the OpenMDW 1.1 license.

Why it matters: HKR-H/K/R all pass: Cosmos 3 leads both Artificial Analysis open-weight image and video charts, with 16B/64B variants and OpenMDW 1.1 artifacts disclosed. Single-source benchmark news keeps it in the 78–84 featured band.

AI HOT (Curated Pool)

Gemini Omni Supports Creating Personal Digital Avatars

Gemini App says Gemini Omni can add users to video creation by generating a digital avatar that resembles their appearance and voice; the post does not disclose rollout scope, pricing, or safety mechanisms.

Why it matters: HKR-H/K/R all pass: the official Gemini App post has a strong multimodal avatar hook. Scope, pricing, consent, and safety controls are not disclosed, keeping it in the mid-weight product-update band.