Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

121–140 of 514

Jun 12Friday

AI HOT (Curated Pool)

MiniMax open-sources M3: 428B total params, 23B active, 1M-token context window

MiniMax uploaded M3 weights to HuggingFace, with the tech report and full weights expected in about 10 days. It's a 428B-total-param, 23B-active-param hybrid model using MiniMax sparse attention to push the context window to 1M tokens, plus native multimodal support. Coding and agent scores: SWE-Bench Pro 59.0%, Terminal Bench 2.1 66.0%, SWE-fficiency 34.8%, KernelBench Hard 28.8%, MCP Atlas 74.2%. MiniMax Code tool and API platform launched alongside. The post doesn't disclose training data, inference cost, or license terms — I'd hold off on usability judgments until the report drops.

Why it matters: MiniMax's first open-weight flagship release: 428B MoE with 23B active params and 1M context, with benchmark scores directly competing against DeepSeek and Qwen on agent/code tasks. Tech report still pending and weights just landed — clear info gaps — but the open-source move ...

AI HOT (Curated Pool)

WSJ: OpenAI weighs steep price cuts and plans biggest ChatGPT overhaul ahead of IPO

WSJ reports OpenAI is weighing steep price cuts as Anthropic gains ground with Claude Code, which enterprise teams are already weaving into daily coding workflows and burning through tokens. OpenAI has the bigger consumer brand, but enterprise pays the bills, so the price move targets developers. At the same time, OpenAI is preparing its biggest ChatGPT overhaul yet ahead of an IPO, aiming to turn it into a super-app spanning coding, AI agents, image generation, and business software. The rollout starts in the coming weeks. OpenAI is also pouring more resources into Codex, with its engineering lead talking about building a 'personal agent.' The post does not disclose specific price cuts or a timeline.

Why it matters: WSJ exclusive: OpenAI is weighing a major price cut because Claude Code is eating into its enterprise developer base, while also prepping ChatGPT's biggest overhaul ahead of IPO. The competitive dynamic is shifting materially, and the pricing response is a direct countermove. ...

Jun 11Thursday

AI HOT (Curated Pool)

Runway and Lionsgate expand partnership with equity stake and joint IP development

Lionsgate has taken an equity interest in Runway and the two will co-develop new IP, starting with a short-form episodic series that blends Lionsgate's existing IP with Runway's generative models. Lionsgate will also be a presenting partner at the Runway AI Festival. The deal builds on their first-of-its-kind partnership from September 2024, where Runway's tools were used for pre-visualization, storyboarding, and final-frame production. Lionsgate was the first Hollywood studio to partner with an applied AI research company, hire a Chief AI Officer, and build out AI infrastructure. Runway co-CEO Cristóbal Valenzuela stressed that studios serious about AI see it as a creative resource, not a cost-cutting tool.

Why it matters: Lionsgate taking equity in Runway and launching a co-developed short series is the most concrete Hollywood bet on AI video generation yet. Score capped because only one project is disclosed—no data on production scale or audience reception.

AI HOT (Curated Pool)

Xiaomi open-sources MiMo Code terminal AI coding assistant, beats Claude Code on SWE-Bench Pro

Xiaomi open-sourced MiMo Code V0.1.0 under MIT license. The built-in MiMo-V2.5 multimodal model is free for a limited time and claims performance on par with Claude Sonnet 4.6; it also supports DeepSeek, Kimi, and GLM. Two standout features: a persistent memory system (project memory, session checkpoints, task progress) to avoid forgetting in long sessions, and a Compose mode for model-agent collaboration that hits 62% on SWE-Bench Pro (Claude Code scored 57%) and 73% on Terminal Bench 2. The post doesn't disclose how long the free period lasts or MiMo-V2.5's parameter count. Type `mimo` in the terminal to start; the UI is fully localized in Chinese.

Why it matters: Xiaomi open-sourcing a terminal coding assistant with MIT license and a free model is a concrete draw for developers. The MiMo-V2.5 claims parity with Claude Sonnet 4.6 but omits parameter count and free-tier cutoff; the persistent memory sub-agent design is more substantive t...

Jun 10Wednesday

OpenAI News

OpenAI banned PRC-linked ChatGPT accounts running covert influence ops on US AI debates

OpenAI published a threat report on June 10 detailing two clusters of ChatGPT accounts likely originating from China, both banned for covert influence operations. One cluster, named 'Data Center Bandwagon,' generated posts claiming AI data centers were raising household electricity prices. The other, 'Tech and Tariffs,' criticized US tariffs as tech competition tactics and instructed outputs to mention only President Trump, not Xi Jinping. That second cluster also spread false claims of a ChatGPT user data breach, which OpenAI calls entirely fabricated. OpenAI found no evidence the operations shifted public opinion, but sees them as testing narratives against US AI infrastructure. The post does not disclose account counts, target platforms, or reach metrics.

Why it matters: OpenAI's official threat report with concrete operational details and account clusters. Hits all three HKR axes, but as a security incident disclosure rather than a product/tech breakthrough, it lands in the 78-84 'good quality' band. Not scored higher because it doesn't resha...

AI Chat-Group Daily (群聊日报)

Anthropic drops Claude Fable 5 / Mythos 5, hits 80.3% on SWE-bench Pro, but safety classifier misfires badly

Anthropic launched two models: Fable 5 for everyone and the full Mythos 5 for trusted partners only. SWE-bench Pro hit 80.3%, well above Opus 4.8's 69.2% and GPT 5.5's 58.6%. It beat Pokémon FireRed using only screenshots. Pricing is double Opus 4.8 at $10/M input and $50/M output. Early testers burned through quota 2–3x faster than Opus; one user drained 73% of a 5-hour allowance in under two hours. The safety classifier became the day's biggest complaint—asking '9.9−9.11=?' triggered a downgrade, and writing an analysis of Anthropic's own safety report got the request blocked entirely. The article had to be finished by DeepSeek V4 Pro. One member pegged the $200 Coding Plan as roughly $5K–10K in API value, calling it a short-lived arbitrage. GitHub Copilot added Fable 5 the same day but requires dropping zero data retention, a dealbreaker for some enterprises. Anthropic's April advisor tool—where a cheap model calls an expensive one for advice—turns out to be the right cost fix for Fable 5. A rice-blast experiment in the safety report also surfaced a shift: AI is flattening domain expertise, but the people who can spot when its answers are wrong are becoming more valuable.

Why it matters: Anthropic flagship model launch with SWE-bench Pro at 80.3%, far ahead of GPT 5.5's 58.6%. Pricing doubled but the Coding Plan may offer a short-term cost arbitrage. Cross-source cluster confirmed, all three HKR axes hit. Minus 1 point because the post doesn't disclose Mythos ...

AI HOT (Curated Pool)

Google Gemini 3.5 Live Translate enters public preview with 70+ languages

Google released Gemini 3.5 Live Translate in public preview through the Gemini API, offering low-latency speech-to-speech translation across 70+ languages and 2,000 language pairs.

Why it matters: HKR-H/K/R all pass: Google’s speech-to-speech translation API has a clear developer hook and concrete scale numbers. Single X-source detail and missing price, latency benchmarks, and regions keep it at 78.

AI HOT (Curated Pool)

Anthropic launches safety-treated Mythos-class model Claude Fable 5

Anthropic released Claude Fable 5, a safety-treated Mythos-class model; in high-risk cyber, biochemistry, and distillation domains, it automatically falls back to Opus 4.8, with one trigger per 20 conversations on average.

Why it matters: Anthropic model launches sit in the 85–94 band; HKR-H/K/R all pass via the safety fallback hook, named mechanism, and Claude-user relevance. X-only sourcing limits confidence, so it stays below the top band.

AI HOT (Curated Pool)

Claude Fable 5 and Claude Mythos 5

Anthropic launched Claude Fable 5 and Claude Mythos 5 at $10 per million input tokens and $50 per million output tokens. Fable 5 leads FrontierCode among frontier models, while Mythos 5 reports about 10x acceleration in drug design and about 80% scientist preference in blinded molecular biology hypothesis tests.

Why it matters: HKR-H/K/R all pass: this is an official Anthropic dual-model release with pricing, coding benchmark, and drug-design speed claims. As a major Claude model update plus Anthropic substantive-update bump, it sits in the 85–94 band.

The Verge · AI

Anthropic releases its first Mythos-class model Claude Fable 5

Anthropic launched Claude Fable 5, calling it the most capable model it has ever released widely. It excels at software engineering, knowledge work, and vision, with its lead growing on longer, more complex tasks. This is the first broad release from the Mythos family, previously deemed too dangerous because of cybersecurity capabilities. New safeguards that block responses in specific high-risk areas made the release possible. The post doesn't disclose benchmarks, pricing, or a launch date.

Why it matters: First public Mythos-class model from Anthropic, previously withheld for safety reasons. Long-horizon task gains and a new safety interceptor are concrete new info. HKR all hit; this is an industry-shaking release.

TechCrunch · AI

Anthropic releases Claude Fable 5, a public version of its top Mythos model with hard safety limits

Anthropic opened its most powerful model family to the public for the first time with Claude Fable 5, a version of Mythos. It excels at software engineering, knowledge work, and vision, but blocks responses in high-risk areas like cybersecurity, biology, chemistry, and distillation, falling back to Claude Opus 4.8. Mythos was previewed in April for select partners only due to cybersecurity concerns. The post doesn't spell out Fable 5's parameter count, pricing, or regional availability.

Why it matters: Anthropic's first public release of a Mythos-tier model is a substantive product launch with explicit safety mechanisms. Cross-source cluster confirmed, all three HKR axes hit. Not scoring higher because the post doesn't disclose benchmark comparisons or pricing details.

Jun 9Tuesday

AI HOT (Curated Pool)

Google Releases Gemini 3.5 Live Translate for Real-Time Speech Translation

Google released Gemini 3.5 Live Translate, a speech-to-speech translation model that supports more than 70 languages, starts translating before the speaker finishes, uses streaming updates, and runs through Gemini Live API, Google Meet preview, and Google Translate apps on iOS and Android.

Why it matters: HKR-H/K/R all pass: Google ties real-time speech translation to 70+ languages and streaming output before the speaker finishes. It stays at 82 because rollout scope, pricing, and benchmarks are not disclosed.

AI HOT (Curated Pool)

Google DeepMind Releases Gemma 4 12B, a Unified Encoder-Free Multimodal Model

Google DeepMind released Gemma 4 12B, a multimodal model with a unified encoder-free architecture, native audio input, Apache 2.0 licensing, and local laptop runtime with 16GB of VRAM or unified memory.

Why it matters: HKR-H/K/R all pass: the hook is local multimodal audio in 16GB VRAM, and the new architecture is concrete. It is a strong Google DeepMind open-model release, but not a frontier-model launch, so it stays below p1.

AI HOT (Curated Pool)

GPT-5.5 Replaces OCR as ChinaRxiv Papers Become Freely Available

A developer replaced a complex OCR pipeline with GPT-5.5, making 23,000+ ChinaRxiv papers freely available with more complete English translations.

Why it matters: HKR-H/K/R all pass, but this is a developer use case rather than an OpenAI model launch. The 23,000+ paper corpus and OCR-pipeline replacement put it in the 78–84 recommendation band.

AI HOT (Curated Pool)

Tencent Hunyuan Releases UniRL, a Unified Multimodal RL Infrastructure

Tencent Hunyuan released UniRL, using one post-training loop to cover diffusion and flow-matching models, LLM/VLM systems, and unified multimodal models, while open-sourcing two algorithms, DRPO and Flow-DPPO.

Why it matters: HKR-H/K/R all pass: Tencent Hunyuan names a unified multimodal RL loop and two open-source algorithms. This fits a strong research/open-source infrastructure release, not a flagship model launch, so it stays in the 78–84 band.

AI HOT (Curated Pool)

How an Agent Chains Two HuggingFace Spaces to Build a 3D Paris Gallery

A coding agent chained ideogram-ai/ideogram4 and VAST-AI/TripoSplat to generate Paris monument images, reconstruct single-image 3D Gaussian splats as .ply files, convert them to .ksplat with about 3× smaller size, and deploy a static Three.js Space using APIs exposed through agents.md.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face Spaces tutorial-style build, not a model or platform release. The concrete chain and ~3x compression place it in the 72-77 featured band.

Jun 8Monday

AI HOT (Curated Pool)

Runway Aleph 2.0 Editing Model Adapts Videos to Any Format

Runway introduced the Aleph 2.0 video editing model, letting users upload an existing video in its desktop web app, choose an aspect ratio, and have the model fill the remaining scene area for the selected format.

Why it matters: Runway Aleph 2.0 is a mid-weight video product update with a concrete mechanism, but no pricing, quality evals, or rollout scope. HKR-H/K/R pass, placing it at the low featured threshold.

AI HOT (Curated Pool)

Microsoft AI CEO: Superintelligence Is Coming, but It Won’t Replace Your Job

Mustafa Suleyman said superintelligence is coming without causing mass unemployment; Microsoft signed a new OpenAI contract last October and released seven omnimodal models at Build this week.

Why it matters: HKR-H/K/R all pass: the job-safety claim creates tension, the piece gives an Oct contract and 7-model Build detail, and it hits automation plus Microsoft-OpenAI nerves. As a CEO interview, not a release, it stays in the 78-84 band.

r/LocalLLaMA

Been Watching Real Adversarial Input Hit My Detection API for Six Months

Bordair’s author says six months of detection-API traffic showed three recurring attack patterns: multi-turn setup, forward-momentum exploitation, and role redefinition; the public adversarial game produced roughly 6,700 attacks last month.

Why it matters: HKR-H/K/R all pass: the post offers real-world adversarial traffic, 3 named tactics, and a 6,700-attack sample. Reddit sourcing keeps it in the high-70s rather than must-write territory.

AI HOT (Curated Pool)

Amap Releases 3D-Native City World Model ABot-Earth0.5

Amap released ABot-Earth0.5, a 3D-native city world model covering more than 190 countries and regions, generating kilometer-scale 3D cities from satellite images or text within 10 minutes on consumer GPUs.

Why it matters: HKR-H/K/R all pass: Amap’s ABot-Earth0.5 has concrete claims, including 190+ countries and 10-minute km-scale 3D city generation. Strong world-model product signal, but below a major foundation-model release.