Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

181–200 of 514

Jun 1Monday

AI HOT (Curated Pool)

MiniMax Releases Open-Source M3 with Coding, Long-Context, and Multimodal Capabilities

MiniMax released the open-source M3 model with coding, a 1M-token context window, and native multimodal support; M3 scores 59.0% on SWE-Bench Pro, 83.5% on BrowseComp, and costs about one-twelfth per token versus GPT-5.5.

Why it matters: HKR-H/K/R all pass: M3 has open source, 1M context, multimodal support, and 59.0% on SWE-Bench Pro. A single X post without official docs or third-party tests keeps it in the 78–84 band.

Latent Space

Why Video Agent Models Are Next — Ethan He on xAI Grok Imagine

Ethan He says a small xAI team built Grok Imagine from zero to one in 3 months, and the episode discusses video agents, audio-video alignment, inference speedups, and the storage, egress, and GPU-hour costs behind large video datasets.

Why it matters: HKR-H/K/R all pass, but the body is interview-level signal: beyond the 3-month build and mechanism themes, it gives no benchmarks, cost figures, or reproducible test. Strong xAI video-agent context, not same-day must-write.

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

400 tokens/s: StepFun Step 3.7 Flash cuts Agent task costs

StepFun released Step 3.7 Flash, a sparse MoE model with 196B parameters plus a 1.8B ViT, activating 11B parameters per inference and reaching up to 400 tokens per second.

Why it matters: HKR-H/K/R all pass with concrete speed and parameter numbers. The feed does not disclose pricing, benchmark setup, or open-source terms, so this stays in the 78–84 quality update band.

Synced · WeChat

OpenAI recruits for robotics team led by Sora creator Aditya Ramesh

OpenAI has listed more than a dozen San Francisco robotics roles for OpenAI Robotics, a team that evolved from Aditya Ramesh’s Worldsim work, with the actuator design engineer role offering $342,000 to $445,000 in base cash pay plus PPU incentives.

Why it matters: HKR-H/K/R all pass: OpenAI robotics hiring adds a strong hook, plus concrete roles, leader, and salary range. This is still a hiring signal, not a model or product release, so it stays in the featured-threshold band.

Synced · WeChat

World models get a “save state”: VAST releases Project Eden

VAST released Project Eden, a three-layer world-model architecture that separates persistent state evolution from visual rendering, and disclosed nearly $200 million across its A+ and A++ funding rounds.

Why it matters: HKR-H/K/R all pass: Project Eden has a product hook, architecture detail, and funding scale. VAST is not a top foundation-model lab, and benchmarks or access terms are not disclosed, so this lands in 78–84.

QbitAI · WeChat

How Cloud Models Reach the Physical World: CMG Lion Rock AI Lab Uses LiOS for Embodied AI

CMG Lion Rock AI Lab released the LiOS edge-cloud architecture for embodied robotics, reporting about 30 ms one-way latency from local camera to cloud GPU memory in cross-machine tests, and open-sourced the low-latency video transmission module plus the LeFold laundry-folding dataset.

Why it matters: HKR-H/K/R pass: LiOS offers a concrete latency claim and open artifacts for embodied AI. Impact stays mid-tier because the lab is not a top platform vendor and no cross-source cluster is shown.

QbitAI · WeChat

VAST Raises Nearly $200M and Discloses Its Project Eden World Model Roadmap

VAST raised nearly $200 million in A+ and A++ rounds and disclosed Project Eden, a world model architecture that separates state evolution from visual rendering through a structured state layer, a conditional interface layer, and a generative rendering layer.

Why it matters: HKR-H/K/R all pass: the $200M A+/A++ financing is sizable, and Project Eden gives a concrete three-layer world-model mechanism. VAST is not a top-tier foundation-model lab and no metrics or release details are disclosed, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Introducing Cosmos Coalition

Runway joined Cosmos Coalition as a founding member and will co-develop the first open world-model foundation model for physical AI with NVIDIA.

Why it matters: HKR-H/K/R all pass: Runway plus NVIDIA and an open physical-AI world model is strong. Details are thin—no params, license, or benchmarks—so it stays in the 78–84 band.

AI HOT (Curated Pool)

Cosmos 3 Released: First Open Physical AI Generalist Model

NVIDIA released Cosmos 3 as an open physical AI generalist model with native visual reasoning, world generation, and action generation, offering two variants: Super at 32B parameters and Nano at 8B parameters.

Why it matters: HKR-H/K/R all pass: NVIDIA names two Cosmos 3 variants and concrete physical-AI capabilities. Source is a single launch post with no benchmark or license detail, so it stays in the 78–84 band.

AI HOT (Curated Pool)

MiniMax M3: Frontier coding, 1M-token context, and native multimodal model

MiniMax released M3 as an open-source unified model with coding, agent, and native multimodal capabilities, supporting a 1M-token context window and using MiniMax Sparse Attention to cut per-token compute at 1M context to 1/20 of its predecessor, with over 9x faster prefill and over 15x faster decoding.

Why it matters: HKR-H/K/R all pass: MiniMax M3 has a 1M-token context hook, MSA with a claimed 20x cost cut, and open-source China-model resonance. Single official-source release keeps it in the 78–84 band, not P1.

AI HOT (Curated Pool)

Qwen3.7-Plus: Multimodal Agent Intelligence

Qwen Studio lists seven capability areas: chatbots, image and video understanding, image generation, document processing, web search integration, tool use, and artifact generation; the post does not disclose Qwen3.7-Plus parameters, pricing, or release timing.

Why it matters: HKR-H/K/R pass, but the facts are thin: 7 capability categories, no params, pricing, benchmarks, or launch terms. A Qwen flagship update clears featured, not p1.

Bloomberg Technology

Gen-Z Gamer’s 3D-Model Startup Becomes China’s Latest AI Unicorn

Vast raised nearly $200 million and reached a $1 billion valuation; the RSS snippet says the 3D-modeling startup was founded by a 29-year-old gamer but does not disclose investors or product specifications.

Why it matters: HKR-H/K/R pass: the founder angle is clickable, the funding and valuation are concrete, and China AI funding has resonance. Missing investors and product specs keep it in the 72–77 band.

OpenAI News

OpenAI bans PRC-linked accounts using ChatGPT to generate comments on US tech policy and tariffs

OpenAI's June 2026 threat report details a banned cluster of ChatGPT accounts likely originating in China. The operators used Simplified Chinese prompts and VPNs to generate English comments and political cartoons criticizing US tariffs, rare earths, AI, and 5G policy. They instructed the model to depict only Trump, not Xi Jinping or China. The same cluster produced Chinese-language water army content attacking the US and Israel, amplifying anti-Jewish tropes, and harassing dissidents. OpenAI also linked these accounts to a separate X network that falsely claimed ChatGPT user data was compromised. The post does not disclose the exact number of banned accounts or the operators' specific institutional affiliation.

Why it matters: OpenAI's official threat report names a PRC-origin AI influence operation with concrete details and high topic sensitivity. Hits all three HKR axes, but it's a security incident report rather than a product/tech breakthrough, placing it in the 78-84 band per policy.

May 31Sunday

Xinzhiyuan · WeChat

Fudan-Linked Team Releases STI-WM Spatiotemporally Integrated World Model

MouShen Intelligence released STI-WM, a spatiotemporally integrated world-action model for robotics, claiming support for RGB, point-cloud, and proprioceptive inputs, hundred-second task planning, and disclosing five funding rounds in six months plus a RMB 300 million Pre-A round.

Why it matters: HKR-H/K/R pass: STI-WM combines RGB, point clouds, and proprioception for 100-second planning, plus 5 funding rounds and a RMB300m Pre-A. Company-claim framing lacks public benchmarks or reproducible access, so it stays near the featured threshold.

QbitAI · WeChat

Robot-Native World Action Model Debuts With Spatiotemporal Architecture From Fudan-Linked Team

Moushen Intelligence released STI-WM, a spatiotemporally integrated world action model for robotics, with RGB, depth point cloud, and proprioceptive inputs; the post says it supports hundred-second-scale long-horizon task rollout and closed-loop replanning, but does not disclose benchmark scores or deployment costs.

Why it matters: HKR-H/K/R all pass: the STI-WM angle is novel, with concrete input modalities and hundred-second rollouts. Kept near the featured floor because public weights, benchmark results, and reproducible tests are not disclosed.

May 30Saturday

AI HOT (Curated Pool)

AI Scammers Create Fake Black Personas to Sell Low-Quality Shein Goods

Sellers use AI-generated Black personas on TikTok, Facebook, and Instagram to pose as handmade creators and sell mass-produced dropshipped goods; the snippet cites one fake persona, “Aliyah,” selling fictional handmade belt buckles.

Why it matters: HKR-H/K/R all pass: the story has a sharp synthetic-identity scam hook, a concrete cross-platform dropshipping mechanism, and clear safety resonance. No scale numbers are disclosed, so it stays in the 72–77 band.

AI HOT (Curated Pool)

Nano Banana Pro and Nano Banana 2 officially released

Google AI Developers released Nano Banana Pro and Nano Banana 2, mapped to gemini-3-pro-image and gemini-3.1-flash-image. The post says both are production-ready through the Gemini API, but does not disclose pricing, benchmarks, or runtime limits.

Why it matters: HKR-H/K/R all pass: Google names two image models and production Gemini API access. Missing pricing, benchmarks, and invocation limits keep it in the mid product-update band rather than a must-write release.

QbitAI · WeChat

Key Gemini IMO Gold Contributor Nearly Became a Professional Pianist

Yi Tay served as a modeling co-captain for Gemini Deep Think when it reached IMO gold-medal level, co-founded Reka AI in 2023, and returned to Google DeepMind after 639 days, while the article also notes his 2012 Trinity classical piano associate diploma.

Why it matters: HKR-H/K/R all pass, but this is a profile, not a Gemini capability launch. The concrete value is Yi Tay's role, Reka history, and 639-day return, so it sits in the 72–77 featured band.

Synced · WeChat

Apple Uses AI to Rework Image Compression: Same Visual Quality at One-Third the File Size

Apple’s team published PICO, a perceptual image codec that uses 57%-70% fewer bits than AV1, VVC, and JPEG AI at the same subjective visual quality, while encoding a 12MP photo in 230 ms and decoding it in 150 ms on an iPhone 17 Pro Max.

Why it matters: HKR-H/K/R all pass: Apple PICO has concrete 57%-70% bitrate savings and 230 ms on-device encoding data. It remains a research release, not a shipped platform feature, so it sits in the 78-84 band.