Skip to content

#多模态

0 today

Aug 21Friday

Jun 7Sunday

QbitAI · WeChat

Kuaishou Kling Proposes VLM-as-Teacher for Test-Time Video Reasoning Optimization

City University and Kuaishou Kling proposed VLM-as-Teacher, which uses VLM feedback to optimize VGM LoRA at test time, reporting a 16.7-point average gain and raising VBVR-Bench from 0.666 to 0.781.

Why it matters: HKR-H/K/R all pass: the story has a novel test-time VLM-teacher hook, concrete VBVR-Bench gains from 0.666 to 0.781, and clear resonance around controllable video generation. It is a strong research release, not a must-write model launch.

Jun 6Saturday

Synced · WeChat

Daxiao Robotics and NTU Release PhysX-Omni for Simulation-Ready Physical 3D Generation

PhysX-Omni models rigid, deformable, and articulated objects in one simulation-ready 3D generation framework, while PhysXVerse contains over 8.7K physical 3D assets across more than 2.9K categories.

Why it matters: HKR-H and HKR-K pass: unified physical modeling plus 8.7K/2.9K+ dataset figures add substance. Source authority and entity weight are mid-tier, and the headline carries promo language, so it stays near the featured threshold.

Jun 5Friday

Synced · WeChat

MetaFine proposes a diagnostic meta-evaluation framework for fine-grained robot manipulation

Southeast University and Peking University researchers introduced MetaFine, a diagnostic meta-evaluation framework that tests fine-grained robot manipulation across understanding, perception, and behavior, and the article says traditional binary success metrics can overestimate fine-manipulation capability by up to 70%.

Why it matters: HKR-H comes from the success-rate illusion hook; HKR-K adds MetaFine’s three-axis diagnostic and a 70% overestimation claim; HKR-R fits robotics eval trust. Research scope keeps it at the low end of 78-84.

Jun 4Thursday

QbitAI · WeChat

CVPR 2026: NVIDIA, Tesla, and Waymo hear Xpeng present physical AI

Xpeng presented its world-model stack at CVPR 2026, covering X-World, X-Foresight, and X-Cache; the article says X-Cache cuts about 70% of repeated computation, the second-generation VLA used over 4 trillion training tokens, and the in-car stack reduced inference latency to 80 ms.

Why it matters: HKR-H comes from the CVPR stage contrast, HKR-K has X-Cache, 4T+ tokens, and 80 ms latency, and HKR-R fits autonomy competition. It is still a company tech showcase, below the 85 must-write band.

Jun 3Wednesday

Synced · WeChat

RSS 2026: Ant Lingbo Proposes Autoregressive Causal World Model for Robot Manipulation with 50 Demos

Ant Lingbo and HKUST introduced LingBot-VA, an autoregressive video-action world model that unifies visual dynamics prediction and action inference, and the paper reports fine-tuning with 50 real-world demonstrations per task plus 92.0% and 91.1% success on RoboTwin 2.0 Easy and Hard settings.

Why it matters: HKR-H/K/R all pass: the hook is 50-demo robot control, with a concrete video-action world-model mechanism. Single-source coverage lacks code, benchmark detail, and deployment evidence, so it lands at 78.

QbitAI · WeChat

Daxiao Robot and NTU Release PhysX-Omni for Unified Physical 3D Generation

Daxiao Robot and NTU introduced PhysX-Omni, a unified simulation-ready physical 3D generation framework for rigid, deformable, and articulated objects, with PhysXVerse covering 8.7K assets across 2.9K categories and PhysX-Bench evaluating six dimensions including geometry, scale, material, affordance, kinematics, and description.

Why it matters: HKR-H/K/R all pass: unified physical 3D generation is a clear hook, the dataset and benchmark numbers add substance, and robotics simulation data is a real practitioner pain. No open-source or product adoption is disclosed, so it stays at 78.

May 30Saturday

Synced · WeChat

Apple Uses AI to Rework Image Compression: Same Visual Quality at One-Third the File Size

Apple’s team published PICO, a perceptual image codec that uses 57%-70% fewer bits than AV1, VVC, and JPEG AI at the same subjective visual quality, while encoding a 12MP photo in 230 ms and decoding it in 150 ms on an iPhone 17 Pro Max.

Why it matters: HKR-H/K/R all pass: Apple PICO has concrete 57%-70% bitrate savings and 230 ms on-device encoding data. It remains a research release, not a shipped platform feature, so it sits in the 78-84 band.

Synced · WeChat

NVIDIA and Tsinghua Team's Gamma-World Tops Hugging Face Daily Chart

NVIDIA, Tsinghua, University of Toronto, and Vector Institute released Gamma-World, a multi-agent world model using simplex-based positional encoding and hub tokens to cut interaction cost from quadratic to linear, with 8-player latency dropping from 17.6 ms to 4.5 ms.

Why it matters: HKR-H/K/R all pass: Gamma-World has a concrete mechanism and latency claim from NVIDIA/Tsinghua. Scope remains multi-agent world-model research, so it sits in the 78–84 good-quality band rather than must-write.

May 28Thursday

NVIDIA Blog

NVIDIA Research Advances Robotics From Simulation to the Real World

NVIDIA Research presented 8 ICRA papers on sim-to-real robotics: ScheduleStream delivered a 3x speedup for multi-arm planning, COMPASS reached about 80% success across 20 real-world navigation trials, and Grasp-MPC achieved about 75% real-robot grasping success.

Why it matters: HKR-K and HKR-R are strong: the post gives concrete sim-to-real numbers from ICRA and addresses robot deployment reliability. HKR-H is moderate but passes on the real-world success-rate hook.

r/LocalLLaMA

Nvidia LocateAnything: Fast Vision-Language Grounding with Parallel Box Decoding

The title says Nvidia LocateAnything-3B performs vision-language grounding with parallel box decoding and runs 10x faster than Qwen3-VL; the post body only provides Hugging Face, GitHub, demo, and project links, and does not disclose benchmark setup or accuracy numbers.

Why it matters: HKR-H/K/R all pass, but the body is mostly links and title-level facts, with no full eval setup or quality metrics. NVIDIA open vision grounding is useful enough for featured, not same-day must-write.

Synced · WeChat

ICML 2026: AutoMoT reaches SOTA on Bench2Drive and nuScenes

NTU AutoMan Lab, Harvard, and Xiaomi Auto proposed AutoMoT, a unified VLA driving model using a 4B Qwen3-VL Understanding Expert and a 1.6B Action Expert with asynchronous inference, reaching 89.42 DS and 74.09% SR on Bench2Drive with AutoMoT+.

Why it matters: HKR-H/K pass via the async VLM-driving setup and concrete Bench2Drive numbers. The autonomy focus narrows HKR-R, so this sits at the featured threshold rather than the 78+ research tier.

May 27Wednesday

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

May 26Tuesday

AI HOT (Curated Pool)

Project Luxo: Crossing the Uncanny Valley of AI Media

Runway released Project Luxo, showing AI shorts and ad samples including The Rogue; each work was made by a single-person team, with production times ranging from three weeks to four hours.

Why it matters: HKR-H/K/R all pass, but this is a Runway research showcase with samples, not a new model or shipped product capability. It lands at the lower end of the good-quality band.

Alibaba Technology · WeChat

Nearly 9x training speedup: residual streams in DiT are becoming a convergence bottleneck

Nanjing University LAMDA and Alibaba Intelligent Engine proposed DAR, a timestep-aware cross-layer routing method that replaces fixed residual accumulation in DiT; on ImageNet 256x256, it reduced SiT-XL/2 FID from 9.67 to 7.56 and reached baseline convergence quality with 8.75x fewer training iterations.

Why it matters: HKR-H/K/R all pass, but the topic is a narrow DiT training method rather than a broad model or product launch. Concrete ImageNet metrics and the Alibaba/LAMDA mechanism clear the featured bar, not the 78+ band.

QbitAI · WeChat

Zhejiang University and Alibaba Make AI Think Before Drawing Sudoku or Burning Candles | ACL 2026

Zhejiang University and Alibaba introduced Unified Thinker, an independent planning module trained with 40,000 HieraReason-40K samples and a two-stage GRPO reinforcement-learning setup that turns structured reasoning traces into executable visual instructions for image generation and editing.

Why it matters: HKR-H/K/R all pass: the paper has a concrete visual-failure hook, a 40k-sample planning/RL mechanism, and relevance to multimodal-agent reliability. It remains a paper-level advance, not a product or flagship model release.

May 24Sunday

Synced · WeChat

ICML 2026: First Parallel Thinking Framework for Vision-Language Models

Visual Para-Thinker introduces a parallel thinking framework for vision-language models, using Pa-Attention and LPRoPE to isolate four visual reasoning paths and training on 163,000 question-answer pairs.

Why it matters: HKR-H/K/R pass: the ICML 2026 paper offers a concrete parallel-thinking mechanism, four isolated paths, and 163K training pairs. It remains a single research release without broad replication or product impact, so it fits 78–84.

May 23Saturday

Synced · WeChat

FlashAR speeds up pretrained autoregressive image models by 22.9x using 0.05% data

Zhejiang University and the University of Adelaide introduced FlashAR, using 0.05% of the original training data to reduce Emu3.5-Image-34B 512×512 generation latency from 130.10 seconds to 5.68 seconds, while GenEval changed from 80.48 to 80.29.

Why it matters: HKR-H/K/R all pass: FlashAR gives speedup, data ratio, latency, and GenEval deltas for AR image inference. It is a strong research item, but not a top-lab model release, so 80 featured rather than P1.

May 22Friday

Synced · WeChat

CVPR 2026 | HiF-VLA: A Motion-Centric World Action Model

Westlake University and collaborators introduced HiF-VLA, a motion-centric VLA framework that extracts compact Motion vectors with codecs such as H.264 and uses a joint expert to predict future visual motion and generate action sequences, reporting 31.4GB peak memory and 117.7ms latency under the cited history-window setting.

Why it matters: HKR-H/K/R all pass: the H.264-motion angle, concrete VRAM/latency numbers, and robotics deployment pressure are clear. It remains a single research item without adoption or cross-source heat, so it sits in the lower featured band.

Synced · WeChat

Meta Chinese Researcher Releases ATLAS for Generalizable Visual Reasoning with One Word

Meta AI and the Chinese University of Hong Kong proposed ATLAS, a visual reasoning method that uses one Functional Token to connect Agentic and Latent Visual Reasoning, with ATLAS-178K, a two-stage SFT+RL pipeline, and LA-GRPO to train sparse visual-operation tokens.

Why it matters: HKR-H/K/R pass: the one-token angle is clickable, and the post gives dataset and training details. As a Meta AI/CUHK research release rather than a flagship model or product launch, it fits the 78–84 band.

May 21Thursday

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

Synced · WeChat

VAST and Tsinghua propose density-controlled 3D Gaussian generation for SIGGRAPH 2026

VAST and Tsinghua propose DeG, a 3D Gaussian generation method that samples Gaussian centers from a learned density distribution and trains density control with a render loss contribution gradient; in some settings, it reaches TRELLIS-like visual quality with less than half the Gaussian count.

Why it matters: HKR-H/K/R pass: DeG offers a concrete mechanism and a testable efficiency claim, reaching TRELLIS-like quality with under half the Gaussians in some scenes. SIGGRAPH research has some technical depth, but no hard-exclusion rule applies.

Synced · WeChat

Xie Saining’s Team Releases Second-Generation Representation Autoencoder RAEv2

Xie Saining’s team, Adobe Research, and the Australian National University released RAEv2, which reaches gFID 1.06 after 80 epochs on ImageNet-256 and reduces EPFID@2 from 177 epochs to 35 epochs while keeping compute at 189 GFLOPs.

Why it matters: HKR-K and HKR-R pass with concrete benchmark and training-efficiency claims. HKR-H is weak because the angle is a normal research release, so it lands at the featured threshold rather than a must-write item.

May 20Wednesday

Hacker News front page

Show HN: Lance – Image/video generation and understanding in one model

ByteDance released Lance as a research project for image and video generation and understanding in one model; the RSS snippet states 3B active parameters, fewer than 128 GPUs used for training, and links to a homepage, arXiv paper, and Hugging Face model, while the post does not disclose benchmark results or licensing terms.

Why it matters: ByteDance’s Lance puts image/video generation and understanding in one model, with 3B active parameters and <128 GPUs for training. HKR-H/K/R all pass, but benchmarks, license details, and real outputs are not disclosed, keeping it below P1.

r/LocalLLaMA

Nemotron-Labs-Diffusion from NVIDIA

NVIDIA released the Nemotron-Labs-Diffusion 3B, 8B, and 14B dense model family with AR decoding, diffusion parallel decoding, and self-speculation; the 8B model reaches 850 tok/s on GB200 at concurrency 1, compared with 253 tok/s for AR and 360 tok/s for Eagle3.

Why it matters: HKR-H/K/R all pass: NVIDIA diffusion LLMs, concrete sizes/mechanisms, and an 850 tok/s GB200 claim. Single-source Reddit sourcing keeps it in the 78–84 band, not P1.

May 17Sunday

QbitAI · WeChat

TGO Aligns Visual Generative Models with Scalar Feedback Without Preference Pairs | ICML 2026

NUS proposed Threshold-Guided Optimization, which converts scalar feedback into positive or negative updates through a score-distribution threshold and was accepted by ICML 2026; experiments cover Stable Diffusion v1.5, FLUX, Wan 1.3B, and Meissonic across image and video generation settings.

Why it matters: HKR-H/K/R pass: the paper has a concrete mechanism and tests across SD v1.5, FLUX, Wan 1.3B, and Meissonic. Impact is research-heavy, so it lands in featured, not must-write.

May 16Saturday

Synced · WeChat

Why Robots Need World Models: Top Institutions Release Joint Survey

NTU MARS Lab and collaborators released a 43-page survey on robot world models, covering definitions, architectures, applications, benchmarks, and challenges around action-conditioned consistency, inference efficiency, and physical grounding.

Why it matters: HKR-H and HKR-K pass: the hook is robot world models, and the post cites a 43-page survey with benchmarks and action-consistency framing. HKR-R is weak, so this stays at the featured threshold.

QbitAI · WeChat

Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1

Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.

Why it matters: HKR-H/K/R all pass: the hook is unusual, and the post gives 3,000 text samples, Flow-GRPO, and a +10.23 dB PSNR gain. Strong multimodal research, but not a foundation-model launch, so 78.

May 15Friday

Xinzhiyuan · WeChat

Hassabis Praises Google DeepMind's AI-enabled Pointer Powered by Gemini

Google DeepMind released a Gemini-powered AI-enabled pointer and opened two demos in Google AI Studio: image editing and place finding on maps, while the post says Chrome pointer selection and a Googlebook Magic Pointer are planned product paths.

Why it matters: HKR-H/K/R all pass: the prompt-free pointer is clickable, the two AI Studio demos add concrete facts, and UI replacement resonates. Scope is still demo-level, with no metrics or API details, so 78 not 85+.

May 14Thursday

Synced · WeChat

China in Focus: PsiBot Uses 100,000 Hours of Human Data for Embodied AI

PsiBot says it uses 100,000 hours of human operation data to train robot policies, with the W0 world model acting only as a training-time transfer module while deployment runs R2 alone.

Why it matters: HKR-H/K/R all pass, but the facts come mainly from company framing and lack an artifact link, benchmark, or third-party replication. This fits a solid robotics research/product story, not the 78+ band.

May 13Wednesday

QbitAI · WeChat

ByteDance Proposes Generative Refinement Networks as a Third Route for Visual Generation

ByteDance’s commercial technology team proposed GRN, a visual generation architecture using HBQ, global refinement, and complexity-aware sampling to address quantization loss, error accumulation, and fixed-step inference; on a 130M model, adaptive sampling reduced inference from 50 steps to an average of 24, while gFID changed from 3.56 to 3.79.

Why it matters: HKR-H/K/R all pass: ByteDance’s GRN has a concrete hook plus 130M, 24-step inference and gFID 3.79. It is a strong research release, not a flagship model launch, so it stays in the 78–84 band.

AI HOT (Curated Pool)

SenseNova-U1 Technical Report Released: Guide to Native Multimodal Model Building

SenseTime released the SenseNova-U1 technical report, covering six-stage training, RL post-training, and distillation; the open-source SenseNova-U1-A3B-MoT uses an MoE architecture and activates only 3 billion parameters.

Why it matters: HKR-H/K/R all pass: A3B-MoT’s 3B active parameters and six-stage training recipe give concrete signal. The score stays near the featured floor because this is a vendor post with no benchmarks, license terms, or reproduction details disclosed.

May 12Tuesday

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

May 11Monday

AI HOT (Curated Pool)

Qwen-Image-2.0 Technical Report

Qwen-Image-2.0 uses a Qwen3-VL condition encoder and multimodal diffusion transformer for image generation and precise editing, with instruction inputs up to 1K tokens and reported gains in multilingual text rendering, layout quality, and human-rated generation and editing tasks.

Why it matters: HKR-H/K/R all pass: Qwen’s flagship image model report gives concrete architecture, 1K-token instruction input, and editing claims. The domestic flagship-model signal lifts it into the must-write band.

May 8Friday

Synced · WeChat

ICLR 2026: NVIDIA and Purdue Use an Agentic Loop for Text-to-3D Scene Generation

NVIDIA Cosmos Lab and Purdue University proposed Scenethesis, a language-and-vision agentic framework for text-to-3D scene generation that uses visual grounding, SDF-based physical constraints, and a judge module; experiments report about 72% first-pass success, 91% after self-checking, and collision rate reduction from 6.1% to 0.8%.

Why it matters: HKR-H/K/R all pass: NVIDIA/Purdue plus an agent loop is clickable, and the post gives SDF constraints, a judge module, and 72%→91% results. Strong research signal, but not a product release, so it stays in 78–84.

May 7Thursday

QbitAI · WeChat

Zhejiang University and Alibaba MetaCompress reaches 90% token compression for multi-turn VQA

Zhejiang University and Alibaba proposed MetaCompress, a learned token-compression framework that generates a compression mapping from the input image alone for multi-turn VQA. The article says it can remove 90% of visual tokens while preserving accuracy, and reports only 1.71% overlap between optimally retained tokens and high-attention tokens.

Why it matters: HKR-H/K/R all pass: 90% visual-token compression, no accuracy loss, and image-conditioned mapping give builders a testable cost-cutting mechanism. Zhejiang/Alibaba plus CVPR 2026 is strong research signal, not a platform-level product release.

Xinzhiyuan · WeChat

Zhejiang University and Harvard open-source UniGeo for geometry-guided camera-controllable editing

Zhejiang University and Harvard released UniGeo with code, a report, a project page, and an HF Space. UniGeo injects geometry guidance into representation, architecture, and loss layers; it reports SOTA on DL3DV, RE10K, and Tanks against five methods. The key is video priors plus geometry-anchor attention, not just using a video model.

Why it matters: HKR-H and HKR-K pass: open code, HF Space, and three geometry-guidance layers make it testable. HKR-R is weak because it is specialized vision-generation research, so this sits near the featured floor.

May 4Monday

Synced · WeChat

ACL 2026: PolyU Open-Sources SignThought for Gloss-Free Sign Language Translation

PolyU and Sichuan University introduced SignThought, accepted to ACL 2026 Main and slated for oral recommendation. It uses latent thoughts, plan-then-ground, and dual-stream decoding, reaching top gloss-free BLEU-4 on five SLT benchmarks. The team also built LC-HKSLT with 1,311 hours, 432K clips, and 14 signers.

Why it matters: ACL 2026 Main, an open model, and a new dataset satisfy HKR-H/K/R, with concrete mechanisms and five benchmarks. The niche sign-language focus keeps it below broader model or developer-tool releases.

Apr 30Thursday

r/LocalLLaMA

DeepSeek released Thinking with Visual Primitives framework

DeepSeek, Peking University, and Tsinghua released the Thinking with Visual Primitives paper and repository. The framework inserts coordinate points and bounding boxes into chain-of-thought; the post does not disclose benchmark scores.

Why it matters: HKR-H/K/R all pass: the hook is visual primitives inside reasoning, the new fact is point/box CoT plus an open repo, and the audience cares about grounded VLMs. No benchmark scores are disclosed, so it stays at 80, not P1.

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.