Skip to content

#多模态

1 today

Jun 10Wednesday

AI HOT (Curated Pool)

Anthropic launches safety-treated Mythos-class model Claude Fable 5

Anthropic released Claude Fable 5, a safety-treated Mythos-class model; in high-risk cyber, biochemistry, and distillation domains, it automatically falls back to Opus 4.8, with one trigger per 20 conversations on average.

Why it matters: Anthropic model launches sit in the 85–94 band; HKR-H/K/R all pass via the safety fallback hook, named mechanism, and Claude-user relevance. X-only sourcing limits confidence, so it stays below the top band.

AI HOT (Curated Pool)

Claude Fable 5 and Claude Mythos 5

Anthropic launched Claude Fable 5 and Claude Mythos 5 at $10 per million input tokens and $50 per million output tokens. Fable 5 leads FrontierCode among frontier models, while Mythos 5 reports about 10x acceleration in drug design and about 80% scientist preference in blinded molecular biology hypothesis tests.

Why it matters: HKR-H/K/R all pass: this is an official Anthropic dual-model release with pricing, coding benchmark, and drug-design speed claims. As a major Claude model update plus Anthropic substantive-update bump, it sits in the 85–94 band.

The Verge · AI

Anthropic releases its first Mythos-class model Claude Fable 5

Anthropic launched Claude Fable 5, calling it the most capable model it has ever released widely. It excels at software engineering, knowledge work, and vision, with its lead growing on longer, more complex tasks. This is the first broad release from the Mythos family, previously deemed too dangerous because of cybersecurity capabilities. New safeguards that block responses in specific high-risk areas made the release possible. The post doesn't disclose benchmarks, pricing, or a launch date.

Why it matters: First public Mythos-class model from Anthropic, previously withheld for safety reasons. Long-horizon task gains and a new safety interceptor are concrete new info. HKR all hit; this is an industry-shaking release.

TechCrunch · AI

Anthropic releases Claude Fable 5, a public version of its top Mythos model with hard safety limits

Anthropic opened its most powerful model family to the public for the first time with Claude Fable 5, a version of Mythos. It excels at software engineering, knowledge work, and vision, but blocks responses in high-risk areas like cybersecurity, biology, chemistry, and distillation, falling back to Claude Opus 4.8. Mythos was previewed in April for select partners only due to cybersecurity concerns. The post doesn't spell out Fable 5's parameter count, pricing, or regional availability.

Why it matters: Anthropic's first public release of a Mythos-tier model is a substantive product launch with explicit safety mechanisms. Cross-source cluster confirmed, all three HKR axes hit. Not scoring higher because the post doesn't disclose benchmark comparisons or pricing details.

Jun 9Tuesday

AI HOT (Curated Pool)

Google Releases Gemini 3.5 Live Translate for Real-Time Speech Translation

Google released Gemini 3.5 Live Translate, a speech-to-speech translation model that supports more than 70 languages, starts translating before the speaker finishes, uses streaming updates, and runs through Gemini Live API, Google Meet preview, and Google Translate apps on iOS and Android.

Why it matters: HKR-H/K/R all pass: Google ties real-time speech translation to 70+ languages and streaming output before the speaker finishes. It stays at 82 because rollout scope, pricing, and benchmarks are not disclosed.

AI HOT (Curated Pool)

Google DeepMind Releases Gemma 4 12B, a Unified Encoder-Free Multimodal Model

Google DeepMind released Gemma 4 12B, a multimodal model with a unified encoder-free architecture, native audio input, Apache 2.0 licensing, and local laptop runtime with 16GB of VRAM or unified memory.

Why it matters: HKR-H/K/R all pass: the hook is local multimodal audio in 16GB VRAM, and the new architecture is concrete. It is a strong Google DeepMind open-model release, but not a frontier-model launch, so it stays below p1.

AI HOT (Curated Pool)

GPT-5.5 Replaces OCR as ChinaRxiv Papers Become Freely Available

A developer replaced a complex OCR pipeline with GPT-5.5, making 23,000+ ChinaRxiv papers freely available with more complete English translations.

Why it matters: HKR-H/K/R all pass, but this is a developer use case rather than an OpenAI model launch. The 23,000+ paper corpus and OCR-pipeline replacement put it in the 78–84 recommendation band.

AI HOT (Curated Pool)

Tencent Hunyuan Releases UniRL, a Unified Multimodal RL Infrastructure

Tencent Hunyuan released UniRL, using one post-training loop to cover diffusion and flow-matching models, LLM/VLM systems, and unified multimodal models, while open-sourcing two algorithms, DRPO and Flow-DPPO.

Why it matters: HKR-H/K/R all pass: Tencent Hunyuan names a unified multimodal RL loop and two open-source algorithms. This fits a strong research/open-source infrastructure release, not a flagship model launch, so it stays in the 78–84 band.

AI HOT (Curated Pool)

How an Agent Chains Two HuggingFace Spaces to Build a 3D Paris Gallery

A coding agent chained ideogram-ai/ideogram4 and VAST-AI/TripoSplat to generate Paris monument images, reconstruct single-image 3D Gaussian splats as .ply files, convert them to .ksplat with about 3× smaller size, and deploy a static Three.js Space using APIs exposed through agents.md.

Why it matters: HKR-H/K/R all pass, but this is a Hugging Face Spaces tutorial-style build, not a model or platform release. The concrete chain and ~3x compression place it in the 72-77 featured band.

Jun 8Monday

AI HOT (Curated Pool)

Runway Aleph 2.0 Editing Model Adapts Videos to Any Format

Runway introduced the Aleph 2.0 video editing model, letting users upload an existing video in its desktop web app, choose an aspect ratio, and have the model fill the remaining scene area for the selected format.

Why it matters: Runway Aleph 2.0 is a mid-weight video product update with a concrete mechanism, but no pricing, quality evals, or rollout scope. HKR-H/K/R pass, placing it at the low featured threshold.

AI HOT (Curated Pool)

Microsoft AI CEO: Superintelligence Is Coming, but It Won’t Replace Your Job

Mustafa Suleyman said superintelligence is coming without causing mass unemployment; Microsoft signed a new OpenAI contract last October and released seven omnimodal models at Build this week.

Why it matters: HKR-H/K/R all pass: the job-safety claim creates tension, the piece gives an Oct contract and 7-model Build detail, and it hits automation plus Microsoft-OpenAI nerves. As a CEO interview, not a release, it stays in the 78-84 band.

r/LocalLLaMA

Been Watching Real Adversarial Input Hit My Detection API for Six Months

Bordair’s author says six months of detection-API traffic showed three recurring attack patterns: multi-turn setup, forward-momentum exploitation, and role redefinition; the public adversarial game produced roughly 6,700 attacks last month.

Why it matters: HKR-H/K/R all pass: the post offers real-world adversarial traffic, 3 named tactics, and a 6,700-attack sample. Reddit sourcing keeps it in the high-70s rather than must-write territory.

AI HOT (Curated Pool)

Amap Releases 3D-Native City World Model ABot-Earth0.5

Amap released ABot-Earth0.5, a 3D-native city world model covering more than 190 countries and regions, generating kilometer-scale 3D cities from satellite images or text within 10 minutes on consumer GPUs.

Why it matters: HKR-H/K/R all pass: Amap’s ABot-Earth0.5 has concrete claims, including 190+ countries and 10-minute km-scale 3D city generation. Strong world-model product signal, but below a major foundation-model release.

Financial Times · Technology

The AI Spying Breakthrough That Spooked the Kremlin

FT says AI can use CCTV data to identify targets, and Russia paused a surveillance system after the killing of Iran’s Supreme Leader; the RSS snippet does not disclose the system name, model mechanism, vendor, or timeline.

Why it matters: FT authority plus HKR-H and HKR-R make this a featured-threshold item: CCTV targeting and Russia pausing surveillance create a strong security hook. HKR-K is weak because system name, model mechanics, and timeline are not disclosed.

Jun 7Sunday

AI HOT (Curated Pool)

A Hokkaido Broccoli Farmer’s 8 Real AI Uses with ChatGPT and Codex

Hokkaido farmer Hiroki Tomiyasu uses ChatGPT and Codex for 8 farm tasks, including broccoli disease recognition, NDVI monitoring, ESP32 greenhouse control, LINE chatbots, sowing-count tracking, RTK-GPS steering study, and an Airtable farm database.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the post names 8 farm workflows, and Codex moving into physical operations will travel among practitioners. Single-X sourcing and missing outcome metrics keep it near the featured floor.

QbitAI · WeChat

Kuaishou Kling Proposes VLM-as-Teacher for Test-Time Video Reasoning Optimization

City University and Kuaishou Kling proposed VLM-as-Teacher, which uses VLM feedback to optimize VGM LoRA at test time, reporting a 16.7-point average gain and raising VBVR-Bench from 0.666 to 0.781.

Why it matters: HKR-H/K/R all pass: the story has a novel test-time VLM-teacher hook, concrete VBVR-Bench gains from 0.666 to 0.781, and clear resonance around controllable video generation. It is a strong research release, not a must-write model launch.

QbitAI · WeChat

Chinese open-source framework targets stable 5-minute AI long-video generation

JD open-sourced JoyAI-Echo, a long audio-video generation framework for 5-minute consistent videos, using cross-modal memory, DMD post-training for about 7.5x faster inference, and real-time upscaling from 720P to 1K or 2K output.

Why it matters: HKR-H/K/R all pass: the story has a clear 5-minute video hook, concrete speed and SR claims, and open-source competition resonance. Missing third-party evaluation keeps it in the lower 78–84 band.

Jun 6Saturday

AI HOT (Curated Pool)

OpenCV 5 Released with New DNN Engine and Native LLM Support

OpenCV 5 introduces a graph-based DNN engine, raising ONNX operator coverage from under 23% in 4.x to over 80%, with native support for Transformer, VLM, and LLM workloads.

Why it matters: HKR-H/K/R all pass for a substantive OpenCV major release: graph DNN engine, ONNX coverage jump, and native Transformer/VLM/LLM support. Strong featured item, but below must-write model-lab release territory.

r/LocalLLaMA

Big week for open AI, with 25+ notable open-weight drops across every modality

Victor M summarized 25+ open-weight model releases in one week, including NVIDIA Nemotron 3 Ultra, a 550B hybrid Mamba-MoE with 55B active parameters and a 1M-token context window.

Why it matters: HKR-H/K/R all pass: the story combines a 25+ open-weight wave with NVIDIA’s 550B, 1M-context Nemotron. Reddit/X sourcing keeps it in the 78-84 band, not p1.

Synced · WeChat

Daxiao Robotics and NTU Release PhysX-Omni for Simulation-Ready Physical 3D Generation

PhysX-Omni models rigid, deformable, and articulated objects in one simulation-ready 3D generation framework, while PhysXVerse contains over 8.7K physical 3D assets across more than 2.9K categories.

Why it matters: HKR-H and HKR-K pass: unified physical modeling plus 8.7K/2.9K+ dataset figures add substance. Source authority and entity weight are mid-tier, and the headline carries promo language, so it stays near the featured threshold.

Synced · WeChat

Video AI Moves to 5 Minutes: Fully Open Source, One-Pass Generation, No Blind-Box Sampling

JD open-sourced JoyAI-Echo, a long audio-video generation framework that supports up to 5 minutes of cross-shot audiovisual consistency, local repainting, 8-step DMD distillation, and output up to 1472×2560 resolution.

Why it matters: JoyAI-Echo clears HKR-H/K/R with a concrete open-source long-video claim: 5-minute output, cross-shot audio-video consistency, and 8-step DMD distillation. Single-source coverage and no independent evals keep it in the 78–84 band.

AI HOT (Curated Pool)

Google AI weekly product updates: Nano Banana 2, Co-Scientist, dreambeans, Gemma 4, and more

Google AI announced six updates: Nano Banana 2 is generally available, Gemma 4 12B can run fully offline on laptops, and Magenta RealTime 2 is open source.

Why it matters: HKR-H/K/R all pass: the post bundles six Google AI updates with concrete local and open-source hooks. Lacking benchmarks, licensing, and pricing keeps it below the 78+ good-quality band.

AI HOT (Curated Pool)

Gemini Live supports real-time image creation and editing

Gemini App adds real-time image creation and editing inside Live; users must open Live, share the camera, and tell Gemini what they want to see.

Why it matters: HKR-H/K/R pass: the real-time Gemini Live image workflow is clickable, concrete, and competitive. Scope is limited: the post gives entry and interaction conditions, not model, pricing, or rollout regions.

Hacker News front page

Launch HN: General Instinct (YC P26) – Frontier Models on Edge Devices

General Instinct open-sourced InstinctRazor, compressing Qwen3.5-122B-A10B from a roughly 245GB BF16 MoE model into a 48GiB GGUF, with a small-GPU mode that streams experts from system RAM and uses about 7.6–8GB peak VRAM at an 8k context window.

Why it matters: HKR-H/K/R all pass: the 122B-to-8GB edge claim is clickable and backed by memory figures. Source authority is still a YC Launch HN, so it fits featured, not must-write.

Jun 5Friday

AI HOT (Curated Pool)

Meta Smart Glasses App Contains Face Recognition Code, NameTag Pushed to Over 50 Million Devices

Meta pushed face-recognition code named NameTag into its smart-glasses companion app, which has more than 50 million downloads; the feature uses three AI models to convert faces into local face templates and match them against a phone database.

Why it matters: HKR-H/K/R all pass: hidden face recognition, 50M-device scale, and a concrete 3-model local-template mechanism. The story stays in the 78–84 band because the post does not confirm user-facing activation.

r/LocalLLaMA

Microsoft released MAI models instead of something like Qwen3.6-27B or Gemma-4-31B

Microsoft AI released seven MAI models, with MAI-Thinking-1 listed as 1T A35B with a 256K context window and MAI-Code-1-Flash listed as 137B A5B with a 256K context window.

Why it matters: Microsoft shipping 7 MAI models with reasoning/code variants and 256K context clears HKR-K/R, and the Qwen/Gemma catch-up angle clears HKR-H. Reddit sourcing and missing benchmarks, license, and pricing keep it below P1.

Synced · WeChat

MetaFine proposes a diagnostic meta-evaluation framework for fine-grained robot manipulation

Southeast University and Peking University researchers introduced MetaFine, a diagnostic meta-evaluation framework that tests fine-grained robot manipulation across understanding, perception, and behavior, and the article says traditional binary success metrics can overestimate fine-manipulation capability by up to 70%.

Why it matters: HKR-H comes from the success-rate illusion hook; HKR-K adds MetaFine’s three-axis diagnostic and a 70% overestimation claim; HKR-R fits robotics eval trust. Research scope keeps it at the low end of 78-84.

QbitAI · WeChat

Yao Shunyu Responds to Whether Tencent Is Behind in AI

Yao Shunyu said at Tencent Cloud’s AI industry application conference that Hunyuan 3 rebuilt pretraining and reinforcement-learning infrastructure, changed data and evaluation, and assigned its strongest post-training staff to improve Yuanbao first; he named coding agents, multimodality, and embodied AI as Tencent’s next focus areas.

Why it matters: HKR-H/K/R all pass, but the facts are conference remarks and roadmap signals, not a new model release with specs, benchmarks, or launch date. This fits the lower featured band for a major Chinese tech AI strategy update.

AI HOT (Curated Pool)

Google Magenta RealTime 2 (MRT2) real-time music model released

Google AI for Developers released the open-weight Magenta RealTime 2 music model, supporting MIDI, live text prompts, and gestures, with native MacBook latency under 200 ms.

Why it matters: HKR-H/K/R all pass: Google Magenta MRT2 has a concrete real-time audio hook, open weights, and sub-200ms local latency. It is strong for creative-AI builders, but narrower than a general foundation-model release.

AI HOT (Curated Pool)

Boson AI and LMSYS Release Higgs Audio v3 TTS End-to-End Service Based on SGLang-Omni

Boson AI and LMSYS released the Higgs Audio v3 TTS service with about 4B parameters, a Qwen3-4B backbone, support for 100 languages, streaming synthesis, and text tags for controlling 20+ emotions plus style, rhythm, and sound effects.

Why it matters: HKR-H and HKR-K pass via the 4B/100-language/streaming TTS hook. HKR-R is weaker because the post lacks latency, pricing, and release-form details, so this sits at the lower featured band.

Jun 4Thursday

AI HOT (Curated Pool)

Nex-N2-Pro launches as a 397B MoE reasoning model based on Qwen3.5

neolab released Nex-N2-Pro, a 397B-parameter MoE reasoning model based on Qwen3.5-397B-A17B, with 262K context, VLM support, claimed GPT-5.5 and Claude Opus 4.7-level performance, 30–50% fewer thinking tokens, SOTA results on Terminal Bench 2.1, GDPVal, and SWE-Verified, plus free access for the first two weeks via SiliconFlow.

Why it matters: HKR-H/K/R pass: the title has a strong benchmark hook and the post gives size, context, and token-reduction claims. Kept in 72-77 because it is a single X source and evaluation conditions are not disclosed.

Xinzhiyuan · WeChat

MoleculeMind releases MMDesign, claims over 90% target hit rate

MoleculeMind released MMDesign, an AI platform for de novo biologics design. In tests across 12 therapeutic targets, it validated specific binding on 11 targets, sending only 14 to 50 molecules per target into wet-lab assays and reporting a target success rate above 90%.

Why it matters: HKR-H/K/R all pass: MMDesign has concrete wet-lab numbers for de novo biologic design. The claim is vertical and partly promotional, so it stays in the 72–77 featured band rather than a broader must-write item.

Xinzhiyuan · WeChat

Silicon Valley CEO backs MiniMax M3 as it tops open-source rankings amid Chinese community debate

MiniMax M3 ranks first among open-source models on Artificial Analysis, and the article says it supports a 1M-token context window, used 100T-scale pretraining, and will open-source its weights and full technical report within 10 days.

Why it matters: HKR-H/K/R all pass: the hook is an open-source No.1 claim amid debate, with 1M context, 100T pretraining, and weights promised in 10 days. Since weights and full report are not out, this stays in 78–84, not P1.

Synced · WeChat

Google releases Gemma 4 12B for 16GB laptops

Google released Gemma 4 12B, a medium-size model that runs locally with 16GB VRAM or unified memory. It uses an encoder-free multimodal architecture, supports native audio input, ships under Apache 2.0, and includes an MTP draft model for lower latency.

Why it matters: Google’s Gemma 4 12B has clear HKR-H/K/R: 16GB local running, 12B scale, and Apache 2.0 licensing. It is a strong open-model update, not a must-write foundation-model launch.

QbitAI · WeChat

CVPR 2026: NVIDIA, Tesla, and Waymo hear Xpeng present physical AI

Xpeng presented its world-model stack at CVPR 2026, covering X-World, X-Foresight, and X-Cache; the article says X-Cache cuts about 70% of repeated computation, the second-generation VLA used over 4 trillion training tokens, and the in-car stack reduced inference latency to 80 ms.

Why it matters: HKR-H comes from the CVPR stage contrast, HKR-K has X-Cache, 4T+ tokens, and 80 ms latency, and HKR-R fits autonomy competition. It is still a company tech showcase, below the 85 must-write band.

AI HOT (Curated Pool)

Ideogram 4.0 Open-Source Text-to-Image Model Released

Ideogram released Ideogram 4.0, an open-source text-to-image model with a 9.3B-parameter core, a single-stream DiT architecture, Qwen3-VL-8B-Instruct text encoder, and a No. 4 ranking in DesignArena human evaluation.

Why it matters: HKR-H/K/R all pass: Ideogram 4.0 brings open weights, 9.3B parameters, single-stream DiT, and a No. 4 human-eval rank. It is strong open image-model signal, not a top-tier general-model launch.

Financial Times · Technology

MP sues Musk’s xAI in UK test case over fake sexual images

UK MP Jess Asato sued Musk’s xAI over fake sexual images, using the claim to test whether AI model makers are liable for system outputs; the post does not disclose the model, generation mechanism, damages sought, or court timetable.

Why it matters: HKR-H/K/R all pass: FT ties xAI, Musk, fake sexual images, and a UK liability test. The article does not disclose the model, generation mechanism, or damages, so it stays in the 78–84 band.

Hacker News front page

Gemma 4 12B: A Unified, Encoder-Free Multimodal Model

Google’s title introduces Gemma 4 12B as a unified, encoder-free multimodal model; the RSS snippet only lists 137 Hacker News points and 48 comments, and the post does not disclose architecture details, training setup, pricing, release terms, or benchmark results.

Why it matters: HKR-H/K/R pass: Google names Gemma 4 12B and an encoder-free multimodal design, a strong hook for open-model practitioners. The post lacks training details, pricing, and benchmarks, so it stays in the low 78–84 band, not P1.

Jun 3Wednesday

r/LocalLLaMA

google/gemma-4-12B on Hugging Face

Google DeepMind released Gemma 4 open-weight models in five sizes, with the 12B variant supporting text, image, and audio input, instruction-tuned and pre-trained variants, native system prompts, function calling, and a context window of up to 256K tokens.

Why it matters: Gemma 4 clears HKR-H/K/R: open weights, multimodal input, and 256K context make it more than a routine update. Missing benchmarks, license detail, and fuller official context keep it in the 78–84 band.

MIT Technology Review · AI

The Download: Trump’s New AI Order, and Smart Glasses for Warfare

President Donald Trump signed a new AI order asking companies to voluntarily submit frontier models for government review 30 days before release, without mandatory licensing; the newsletter also says Anduril and Meta are prototyping a military AR headset that envisions drone-strike orders through eye tracking and voice commands.

Why it matters: HKR-H/K/R all pass: the article gives a concrete 30-day frontier-model review mechanism and a Meta/Anduril AR warfare prototype. A presidential AI order affecting release compliance clears the must-write band.