Skip to content

#多模态

1 today

May 17Sunday

Bloomberg Technology

Apple’s New ChatGPT-Like Siri App Will Have Auto-Deleting Chats

The title says Apple’s ChatGPT-like Siri app will support auto-deleting chats; the RSS snippet only adds that iOS 27 will include a Genmoji upgrade, and the post does not disclose retention periods, release timing, or feature details.

Why it matters: HKR-H and HKR-R pass because Bloomberg frames a specific Apple Siri privacy angle; HKR-K fails since retention and feature mechanics are missing, so this stays at the low featured threshold.

Google DeepMind

Google expands content provenance and verification tools across Search, Gemini, Chrome and Pixel

Google is widening its content transparency and verification tools across Search, Gemini, Chrome, Pixel and Cloud, and deepening industry partnerships. SynthID has watermarked over 100 billion images and videos plus 60,000 years of audio. SynthID verification in the Gemini app has been used 50 million times, and the capability reaches Search today, with Chrome in the coming weeks.

Why it matters: The post lays out where SynthID and C2PA land across Search, Gemini, Chrome and Pixel, which shows the current limits of content provenance tools.

QbitAI · WeChat

TGO Aligns Visual Generative Models with Scalar Feedback Without Preference Pairs | ICML 2026

NUS proposed Threshold-Guided Optimization, which converts scalar feedback into positive or negative updates through a score-distribution threshold and was accepted by ICML 2026; experiments cover Stable Diffusion v1.5, FLUX, Wan 1.3B, and Meissonic across image and video generation settings.

Why it matters: HKR-H/K/R pass: the paper has a concrete mechanism and tests across SD v1.5, FLUX, Wan 1.3B, and Meissonic. Impact is research-heavy, so it lands in featured, not must-write.

QbitAI · WeChat

A Robot Dog Challenges Nvidia's Compute Lead

Weilan Technology unveiled BabyAlpha A3, a consumer quadruped robot using a six-chip heterogeneous cluster that runs a 7B-parameter model on-device at 280 TPS; the article says it has 66MP vision, 2.232 million point-cloud samples per second, and a planned Q3 launch.

Why it matters: HKR-H/K/R pass: the robot-dog-versus-Nvidia angle is clickable, and 280 TPS on a local 7B model is concrete. Single-source summary lacks price, power draw, and benchmark setup, so it stays near the featured floor.

AI HOT (Curated Pool)

Grok Imagine image generation is officially released

Grok Imagine is now available on X for all users, with text-to-image generation for realistic images and multiple aspect ratios; the post does not disclose model parameters, pricing, or regional limits.

Why it matters: HKR-H/K/R pass, but the post only discloses availability and basic image features; model details, pricing, and regions are absent, so this lands at the featured threshold.

Financial Times · Technology

Chinese AI Groups Pull Ahead of US Rivals in Video Generation Race

FT says Chinese AI groups have moved ahead of US rivals in video generation; the RSS snippet names ByteDance and Kuaishou and says they outshine western competitors in advertising and entertainment quality, but the post does not disclose benchmark metrics or model details.

Why it matters: FT authority plus a China-vs-US video-generation lead claim clears HKR-H and HKR-R. HKR-K fails because the body lacks metrics, samples, and eval method, so it sits at the low featured threshold.

Synced · WeChat

What Are World Models? Their History and the $10 Billion Bet

Jiqizhixin translated a MoE Capital blog tracing two world-model lineages. The article says more than $10 billion entered the category over 18 months, and cites DreamDojo as using 44,711 hours of first-person video pretraining to reach r=0.995 correlation with real-world robot policy outcomes.

Why it matters: HKR-H/K/R all pass: the hook is strong and the article gives concrete figures, but it is a compiled explainer rather than a new release. It fits the featured-threshold band for a strong commentary/tutorial.

May 16Saturday

Hacker News front page

SANA-WM, a 2.6B open-source world model for 1-minute 720p video

SANA-WM’s title says the project is a 2.6B open-source world model for 1-minute 720p video; the RSS body only lists the project URL, Hacker News comments URL, 9 points, and 8 comments, and the post does not disclose training data, license terms, inference cost, evaluation setup, or benchmark results.

Why it matters: HKR-H/K/R pass on the concrete open-source world-model hook, 2.6B size, and video-model competition angle. Sparse body details keep it at the lower good-quality band.

Synced · WeChat

Why Robots Need World Models: Top Institutions Release Joint Survey

NTU MARS Lab and collaborators released a 43-page survey on robot world models, covering definitions, architectures, applications, benchmarks, and challenges around action-conditioned consistency, inference efficiency, and physical grounding.

Why it matters: HKR-H and HKR-K pass: the hook is robot world models, and the post cites a 43-page survey with benchmarks and action-consistency framing. HKR-R is weak, so this stays at the featured threshold.

QbitAI · WeChat

Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1

Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.

Why it matters: HKR-H/K/R all pass: the hook is unusual, and the post gives 3,000 text samples, Flow-GRPO, and a +10.23 dB PSNR gain. Strong multimodal research, but not a foundation-model launch, so 78.

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

AI HOT (Curated Pool)

Runway Agent Generates Complete Ads in One Session

Runway Agent turns product photos and ideas into fully produced ads in one session; the post does not disclose the model, pricing, generation length, or regional availability.

Why it matters: Runway’s ad-generation Agent clears HKR-H/K/R as a mid-weight product update. Missing model, pricing, duration, and region details keep it at the featured threshold, not a must-write release.

May 15Friday

r/LocalLLaMA

Fully Offline Suitcase Robot Built Around Jetson Orin NX SUPER 16GB

CreativelyBankrupt built Sparky as a fully offline suitcase robot on Jetson Orin NX SUPER 16GB, running Gemma 4 E4B Q4_K_M via llama.cpp with q8_0 KV cache, about 200 ms cached TTFT, 14-15 tok/s sustained output, 12K context, 30+ sensors, and no WiFi, Bluetooth, or cellular interface.

Why it matters: HKR-H/K/R all pass, with a named hands-on build and concrete latency/sensor numbers. It stays in low featured because this is a Reddit project post, not a product launch or research release.

MIT Technology Review · AI

The Download: China’s AI Drama Factory and the WHO’s Missing Health Targets

China’s short-drama industry released an average of 470 AI-generated short dramas per day in January, while production timelines fell from months to weeks and costs dropped by up to 90%.

Why it matters: MIT Technology Review provides concrete output, cycle-time, and cost figures for China’s AI short-drama pipeline, clearing HKR-H/K/R. The story is application-layer, not a core model or product release, so it sits at the featured threshold.

MIT Technology Review · AI

How Chinese Short Dramas Became AI Content Machines

Chinese short-drama companies are using AI for full-series production, with DataEye counting an average of 470 AI-generated short dramas released per day in January 2026, while FlexTV says production time fell from three to four months to under one month and North American per-series costs can drop by 80% to 90%.

Why it matters: HKR-H/K/R all pass: the story has a strong content-factory hook, concrete production metrics, and clear labor/cost resonance. It is a quality industry feature, not a model or platform release, so 80 fits the 78-84 band.

QbitAI · WeChat

Understand LeCun’s JEPA World Model in 160 Lines of Code

A developer released the keon/jepa teaching repository with five JEPA variants implemented as standalone PyTorch files, ranging from 160 to 278 lines, depending only on PyTorch and torchvision; the post reports iJEPA runs on CIFAR-10 for 100 epochs and reaches 52.7% linear-probe accuracy, while V-JEPA, C-JEPA, and LeWorldModel use toy or synthetic datasets.

Why it matters: HKR-H/K/R pass via the 160-line JEPA hook, reproducible repo, and non-LLM world-model angle. It is a tutorial artifact, not a model or paper release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

Hassabis Praises Google DeepMind's AI-enabled Pointer Powered by Gemini

Google DeepMind released a Gemini-powered AI-enabled pointer and opened two demos in Google AI Studio: image editing and place finding on maps, while the post says Chrome pointer selection and a Googlebook Magic Pointer are planned product paths.

Why it matters: HKR-H/K/R all pass: the prompt-free pointer is clickable, the two AI Studio demos add concrete facts, and UI replacement resonates. Scope is still demo-level, with no metrics or API details, so 78 not 85+.

May 14Thursday

r/LocalLLaMA

Open-source one-prompt-to-cinematic-reel pipeline on one GPU with FLUX.2 and Wan2.2-I2V

The developer open-sourced StudioMI300, an 8-stage sequential pipeline that turns one English sentence into a 720p MP4 on a single AMD Instinct MI300X, cutting end-to-end time from 25.9 minutes to 10.4 minutes per clip.

Why it matters: HKR-H/K/R all pass: the post has a concrete one-GPU video pipeline, runtime numbers, and a local-build cost/control hook. Reddit single-source status and no third-party replication keep it below the 78+ band.

MIT Technology Review · AI

The Shock of Seeing Your Body Used in Deepfake Porn

MIT Technology Review documents Jennifer and other adult content creators whose bodies were used in NCII deepfakes, with examples spanning Jennifer’s circa-2013 video and the 2017 Reddit “deepfakes” uploads involving celebrity face swaps.

Why it matters: HKR-H and HKR-R are strong, with HKR-K from named cases and the 2013-to-2017 deepfake lineage. This is a high-quality safety/policy feature, not a model or product release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

Anthropic Overtakes OpenAI in Enterprise AI Adoption After Three Years

Ramp says Anthropic reached 34.4% enterprise adoption, surpassing OpenAI at 32.3% for the first time; the index is based on credit-card and invoice spending from more than 50,000 companies.

Why it matters: HKR-H/K/R all pass: a reversal hook, concrete 34.4%/32.3% figures, and a strong enterprise-AI rivalry angle. Score stays at 80 because Ramp spending data is not global market share.

QbitAI · WeChat

Alexandr Wang Responds to LeCun, Manus, and Meta AI Rebuild

Alexandr Wang said Meta rebuilt its pretraining, reinforcement learning, and data stacks in nine months, while Muse Spark remains closed because it triggered safety checks in areas including biosecurity, cyber capability, and loss of control.

Why it matters: HKR-H/K/R all pass: the named conflict draws clicks, the 9-month Meta stack rebuild and Muse Spark safety hold add facts, and open-source safety hits a real practitioner nerve. This is an interview, not a model launch, so it sits in the 78-84 band.

Synced · WeChat

China in Focus: PsiBot Uses 100,000 Hours of Human Data for Embodied AI

PsiBot says it uses 100,000 hours of human operation data to train robot policies, with the W0 world model acting only as a training-time transfer module while deployment runs R2 alone.

Why it matters: HKR-H/K/R all pass, but the facts come mainly from company framing and lack an artifact link, benchmark, or third-party replication. This fits a solid robotics research/product story, not the 78+ band.

AI HOT (Curated Pool)

Best Practices for Computer and Browser Use with Claude

Anthropic published guidance for Claude computer and browser use, with Claude 4.6 API screenshots capped at a 1,568-pixel long edge and 1.15 million total pixels, while Opus 4.7 raises the limits to 2,576 pixels and 3.75 million total pixels.

Why it matters: Anthropic’s first-party Claude computer/browser guide has actionable screenshot limits, not just promo copy. HKR-H/K/R all pass, but this is a practice guide rather than a major model or capability launch, so it sits in the 72–77 band.

AI HOT (Curated Pool)

Introducing Runway Agent

Runway launched Runway Agent, a video creation agent that turns one natural-language conversation into multi-scene videos with narration, dialogue, and music; new free-plan users receive 1,500 credits for their first video.

Why it matters: HKR-H/K/R pass: a notable AI-video vendor ships an agentic multi-scene workflow with a 1,500-credit free plan. Score stays in the 72–77 band because the post is still a vendor announcement without pricing, limits, or independent tests.

r/LocalLLaMA

sensenova/SenseNova-U1-A3B-MoT · Hugging Face

SenseNova published SenseNova-U1-A3B-MoT on Hugging Face; the post lists A3B MoT, 8B MoT, and 0.4B LoRA weight links, and says the NEO-unify architecture unifies multimodal understanding, reasoning, and generation in one model family.

Why it matters: HKR-H/K/R all pass: an open multimodal model release with multiple weight sizes and a named NEO-unify mechanism. Source authority and missing benchmarks/license details keep it in the lower featured band.

May 13Wednesday

r/LocalLLaMA

AIDC-AI/Ovis2.6-80B-A3B on Hugging Face

AIDC-AI released Ovis2.6-80B-A3B, a multimodal MoE model with 80B total parameters and about 3B active parameters at inference, supporting a 64K-token context window and images up to 2880×2880 resolution.

Why it matters: HKR-H/K/R pass: the open multimodal MoE has concrete specs and a real efficiency hook. Score stays near the featured floor because the post gives no benchmarks, license details, or hands-on results.

QbitAI · WeChat

ByteDance Proposes Generative Refinement Networks as a Third Route for Visual Generation

ByteDance’s commercial technology team proposed GRN, a visual generation architecture using HBQ, global refinement, and complexity-aware sampling to address quantization loss, error accumulation, and fixed-step inference; on a 130M model, adaptive sampling reduced inference from 50 steps to an average of 24, while gFID changed from 3.56 to 3.79.

Why it matters: HKR-H/K/R all pass: ByteDance’s GRN has a concrete hook plus 130M, 24-step inference and gFID 3.79. It is a strong research release, not a flagship model launch, so it stays in the 78–84 band.

Synced · WeChat

Lin Junyang Reportedly Starts New AI Lab Seeking $2 Billion Valuation

The Information says Lin Junyang is raising several hundred million dollars for a new AI Lab at a potential $2 billion post-financing valuation, while the lab’s research direction and final valuation remain undisclosed.

Why it matters: HKR-H/K/R all pass, but the article only gives The Information’s funding rumor and valuation; research focus, team, and product plan are not disclosed. This fits the 72–77 featured band.

AI HOT (Curated Pool)

SenseNova-U1 Technical Report Released: Guide to Native Multimodal Model Building

SenseTime released the SenseNova-U1 technical report, covering six-stage training, RL post-training, and distillation; the open-source SenseNova-U1-A3B-MoT uses an MoE architecture and activates only 3 billion parameters.

Why it matters: HKR-H/K/R all pass: A3B-MoT’s 3B active parameters and six-stage training recipe give concrete signal. The score stays near the featured floor because this is a vendor post with no benchmarks, license terms, or reproduction details disclosed.

Xinzhiyuan · WeChat

Tsinghua-affiliated team open-sources MiniCPM-V 4.6, a 1.3B model tunable on one RTX 4090

ModelBest, Tsinghua University, and OpenBMB open-sourced MiniCPM-V 4.6, a 1.3B multimodal model that supports full fine-tuning on one RTX 4090 and offers 4x/16x visual token compression for accuracy or speed trade-offs.

Why it matters: HKR-H/K/R all pass: the story gives a concrete open-source multimodal release with size, hardware condition, and token-compression details. It lowers local fine-tuning cost, but it is not a frontier-lab flagship release, so 78–84 fits.

AI HOT (Curated Pool)

Step Image Edit 2 image model released with leading performance and efficiency

StepFun released the 3.5B-parameter Step Image Edit 2 model, which ranks first in KRIS-Bench overall, factual, and conceptual categories, and is now available on the Stepfun Open Platform.

Why it matters: HKR-H/K/R all pass: the hook is a 3.5B image-editing model topping KRIS-Bench, with concrete launch details. Vendor-only sourcing and no independent test or pricing keep it at the low featured band.

May 12Tuesday

Synced · WeChat

ByteDance Open-Sources DreamLite for Offline Mobile Image Generation and Editing

ByteDance open-sourced DreamLite, a 0.39B-parameter unified diffusion model that generates or edits a 1024×1024 image on an iPhone 17 Pro in about 3 seconds, using 4-step DMD2 distillation and on-device offline inference without cloud dependency.

Why it matters: HKR-H/K/R all pass: 3-second on-device 1024×1024 generation is a strong hook, with 0.39B params and 4-step DMD2 as concrete claims. As a ByteDance open-source vision model, it sits below a general foundation-model release.

Xinzhiyuan · WeChat

The Largest Single Industrial Product in History Is Entering Mass Production in China

AgiBot says it had shipped 10,000 general-purpose embodied robots by the end of March, and its humanoid robots worked eight continuous hours on a Nanchang 3C production line, completing 2,283 tasks with zero errors under formal line-cycle requirements.

Why it matters: HKR-H/K/R all pass: AgiBot gives unit and factory-run numbers with clear robotics deployment resonance. The score stays in 78-84 because the key claims are company-sourced, with no third-party validation or cost data disclosed.

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

QbitAI · WeChat

OpenClaw quietly updates with Peekaboo v3 for Mac computer use

OpenClaw-related Peekaboo v3 adds Mac agent capabilities for pixel-level screenshots, UI position reading, clicks, text input, hotkeys, scrolling, and drag-and-drop, with MCP server integration for Cursor, Claude Code, and Codex.

Why it matters: HKR-H/K/R all pass: Peekaboo v3 adds Mac GUI perception and action primitives plus MCP access for Cursor, Claude Code, and Codex. This is a useful open-source agent-tooling update, not a model-level event, so it sits in the low featured band.

AI HOT (Curated Pool)

Thinking Machines Releases Native Multimodal Interaction Model for Real-Time Human-AI Collaboration

Thinking Machines released an interaction model that natively receives audio, video, and text input, processes foreground interaction at 200-millisecond intervals, and uses a background reasoning model for long-horizon planning and tool calls.

Why it matters: HKR-H/K/R all pass: this is more than a model notice, with a two-layer foreground/background interaction design. Pricing, access scope, and benchmarks are missing, so it sits at the lower end of 85-94.

The Verge · AI

Here’s What Mira Murati’s AI Company Is Up To

Thinking Machines announced work on “interaction models” that continuously take in audio, video, and text and respond or act in real time; the post does not disclose model size, release timing, pricing, or the final product format.

Why it matters: HKR-H/K/R all pass, but the body lacks parameters, launch timing, and product form. This is a high-interest startup direction reveal, not a usable model release, so it stays at the top of the 72–77 band.

AI HOT (Curated Pool)

The Evolution of Human-Computer Interfaces: From Text to Interactive Neural Video

Karpathy argues that LLM output is moving from Markdown toward richer HTML, while interactive neural video still has an open problem: how to combine neural generation with precise traditional software.

Why it matters: HKR-H/K/R pass: Karpathy gives a fresh UI frame, a concrete Markdown→HTML→neural-video path, and a builder-facing product question. Single X post with no data keeps it at the featured floor.

May 11Monday

AI HOT (Curated Pool)

Qwen-Image-2.0 Technical Report

Qwen-Image-2.0 uses a Qwen3-VL condition encoder and multimodal diffusion transformer for image generation and precise editing, with instruction inputs up to 1K tokens and reported gains in multilingual text rendering, layout quality, and human-rated generation and editing tasks.

Why it matters: HKR-H/K/R all pass: Qwen’s flagship image model report gives concrete architecture, 1K-token instruction input, and editing claims. The domestic flagship-model signal lifts it into the must-write band.

May 10Sunday

Synced · WeChat

Ted Xiao Reviews Three Eras of Robot Learning, from RT-1/RT-2 to Scaling

Ted Xiao divides nearly a decade of robot learning into three eras: Google’s team trained RT-1 on 87,000 teleoperation trajectories, then adapted 5B to 55B VLMs into VLA policies for RT-2.

Why it matters: HKR-H/K/R all pass: a named Google robotics insider, concrete RT-1/RT-2 numbers, and strong embodied-AI resonance. It is retrospective commentary, not a launch, so it stays in the 72–77 featured band.