Skip to content

#多模态

1 today

May 29Friday

r/LocalLLaMA

StepFun 3.7 Flash

StepFun released Step 3.7 Flash with 196B total parameters, 11B active MoE, a built-in 1.8B ViT, and local execution on 128GB RAM.

Why it matters: HKR-H/K/R pass via the 196B/11B MoE specs and 128GB local-run claim. Sparse Reddit sourcing leaves license, eval method, and access conditions undisclosed, so it stays in the lower featured band.

AI HOT (Curated Pool)

StepFun Releases Step 3.7 Flash, Focused on Agent Efficiency

StepFun released the open-source Step 3.7 Flash model with a 198B-parameter MoE architecture, about 11B active parameters, a 256K context window, and a 67.1 score on ClawEval-1.1.

Why it matters: HKR-H/K/R all pass: the release has a clear sparse-model hook, concrete context and benchmark numbers, and practitioner resonance around open agent efficiency. Official-post sourcing and no independent eval keep it in the 78–84 band.

AI HOT (Curated Pool)

Nano Banana Pro and Nano Banana 2 officially released

Google AI Developers released Nano Banana Pro and Nano Banana 2, two image models available for production use through the Gemini API; the post names gemini-3-pro-image and gemini-3.1-flash-image but does not disclose pricing, benchmarks, or rate limits.

Why it matters: HKR-H/K/R all pass: Google shipped two production image models via Gemini API. The post gives no benchmarks, pricing, or safety mechanism, so this stays in the 78–84 band rather than p1.

The Verge · AI

A $2,000 AI-generated film will debut at Tribeca

Tribeca Festival will premiere the 75-minute AI-generated film Dreams of Violets, which cost $2,000 to make and uses people and images fully created by AI.

Why it matters: HKR-H/K/R all pass: a $2,000, 75-minute AI film entering Tribeca has novelty, concrete numbers, and labor resonance. It is a film-industry application, not a model or platform launch, so it stays in the low featured band.

May 28Thursday

NVIDIA Blog

NVIDIA Research Advances Robotics From Simulation to the Real World

NVIDIA Research presented 8 ICRA papers on sim-to-real robotics: ScheduleStream delivered a 3x speedup for multi-arm planning, COMPASS reached about 80% success across 20 real-world navigation trials, and Grasp-MPC achieved about 75% real-robot grasping success.

Why it matters: HKR-K and HKR-R are strong: the post gives concrete sim-to-real numbers from ICRA and addresses robot deployment reliability. HKR-H is moderate but passes on the real-world success-rate hook.

r/LocalLLaMA

Nvidia LocateAnything: Fast Vision-Language Grounding with Parallel Box Decoding

The title says Nvidia LocateAnything-3B performs vision-language grounding with parallel box decoding and runs 10x faster than Qwen3-VL; the post body only provides Hugging Face, GitHub, demo, and project links, and does not disclose benchmark setup or accuracy numbers.

Why it matters: HKR-H/K/R all pass, but the body is mostly links and title-level facts, with no full eval setup or quality metrics. NVIDIA open vision grounding is useful enough for featured, not same-day must-write.

Synced · WeChat

ICML 2026: AutoMoT reaches SOTA on Bench2Drive and nuScenes

NTU AutoMan Lab, Harvard, and Xiaomi Auto proposed AutoMoT, a unified VLA driving model using a 4B Qwen3-VL Understanding Expert and a 1.6B Action Expert with asynchronous inference, reaching 89.42 DS and 74.09% SR on Bench2Drive with AutoMoT+.

Why it matters: HKR-H/K pass via the async VLM-driving setup and concrete Bench2Drive numbers. The autonomy focus narrows HKR-R, so this sits at the featured threshold rather than the 78+ research tier.

Synced · WeChat

Chinese pretrained embodied model Wall-OSS-0.5 is open sourced

X Square Robot open sourced Wall-OSS-0.5, a VLA model whose 400k pretraining checkpoint scored above 80 on 4 of 17 real-robot zero-shot tasks, with weights, code, training recipe, ablations, and a DMuon optimizer implementation released.

Why it matters: Clear HKR-H/K/R: a 400k checkpoint and 17 real-robot zero-shot tasks add substance, while “post-training not required” is a sharp hook. X Square Robot is not a top foundation-model lab, so this stays at 79.

AI HOT (Curated Pool)

Open-source FastVideo Dreamverse real-time video generation tool

Hao AI Lab open-sourced FastVideo Dreamverse, a real-time video generation tool that generates a 30-second 1080p video in 7 seconds under the stated setup of one NVIDIA B200 GPU and LTX-2.

Why it matters: HKR-H/K/R all pass: the 7s-for-30s-1080p claim is concrete and practitioner-relevant. Single-source X sourcing and missing independent benchmarks keep it in the 78–84 band.

May 27Wednesday

AI HOT (Curated Pool)

Runway launches Model Context Protocol server

Runway launched an MCP server that lets compatible agents such as Claude, ChatGPT, and Cursor generate images and videos inside chat interfaces, with access to Gen-4.5, Seedance 2.0, GPT Image 2, Kling 3.0, and Nano Banana Pro.

Why it matters: HKR-H/K/R all pass, but this is a Runway product integration, not an MCP protocol change or model release. It clears featured, with the score kept in the 72–77 band.

TechCrunch · AI

YouTube will now automatically label AI videos

YouTube will automatically label videos using significant photorealistic AI, no longer relying only on creator self-disclosure. The RSS snippet says AI labels will become more prominent, but the post does not disclose rollout timing, detection thresholds, appeal rules, or whether the system covers shorts and livestreams.

Why it matters: HKR-H/K/R pass: YouTube shifts AI-video labels from creator self-reporting to platform detection. The article gives the mechanism, but not accuracy, appeals, or rollout scope, so it sits at the featured threshold.

QbitAI · WeChat

7B Medical AI Agent Beats o3 and GPT-5 by Learning Where and How to Look

Shanghai Innovation Institute’s LeapQuest and three universities released Ophiuchus and MedScope, applying Think with Images and Think with Videos to medical AI; Ophiuchus-7B scored 68.0 on eight VQA benchmarks, above OpenAI-o3 at 62.2, Gemini 2.5 Pro at 61.8, and GPT-5 at 59.9.

Why it matters: HKR-H/K/R all pass: a 7B model beating o3/GPT-5 is a strong hook, 8 VQA benchmarks with 68.0 vs 62.2 add a testable claim, and medical specialist evaluation will trigger debate. Not a frontier-lab general model release, so it stays in 78–84.

r/LocalLLaMA

PrismML Released Binary and Ternary Bonsai Image 4B

PrismML released Binary and Ternary Bonsai Image 4B, 1-bit and ternary text-to-image diffusion transformers around 3GB, compared with FLUX.2 Klein 4B at about 16GB, with browser-local WebGPU demo links and an Apache-2.0 license disclosed in the Reddit snippet.

Why it matters: HKR-H/K/R all pass: low-bit image DiT plus local browser inference is a strong hook, backed by 4B, ~3GB, WebGPU, and license details. Reddit sourcing and limited lab weight keep it in the low featured band.

May 26Tuesday

AI HOT (Curated Pool)

SenseNova-U1 full training code open-sourced for multimodal multitask training

OpenSenseNova released the full SenseNova-U1 training code on GitHub under Apache-2.0, supporting an 8B dense model, an A3B MoE architecture, and multimodal tasks such as text-to-image generation, image editing, interleaved generation, and text-visual understanding.

Why it matters: HKR-H/K/R all pass, but the source is a short official post with no dataset, training budget, or eval results disclosed. The practical value of full training code puts it in the featured band.

AI HOT (Curated Pool)

Project Luxo: Crossing the Uncanny Valley of AI Media

Runway released Project Luxo, showing AI shorts and ad samples including The Rogue; each work was made by a single-person team, with production times ranging from three weeks to four hours.

Why it matters: HKR-H/K/R all pass, but this is a Runway research showcase with samples, not a new model or shipped product capability. It lands at the lower end of the good-quality band.

Alibaba Technology · WeChat

Nearly 9x training speedup: residual streams in DiT are becoming a convergence bottleneck

Nanjing University LAMDA and Alibaba Intelligent Engine proposed DAR, a timestep-aware cross-layer routing method that replaces fixed residual accumulation in DiT; on ImageNet 256x256, it reduced SiT-XL/2 FID from 9.67 to 7.56 and reached baseline convergence quality with 8.75x fewer training iterations.

Why it matters: HKR-H/K/R all pass, but the topic is a narrow DiT training method rather than a broad model or product launch. Concrete ImageNet metrics and the Alibaba/LAMDA mechanism clear the featured bar, not the 78+ band.

QbitAI · WeChat

Zhejiang University and Alibaba Make AI Think Before Drawing Sudoku or Burning Candles | ACL 2026

Zhejiang University and Alibaba introduced Unified Thinker, an independent planning module trained with 40,000 HieraReason-40K samples and a two-stage GRPO reinforcement-learning setup that turns structured reasoning traces into executable visual instructions for image generation and editing.

Why it matters: HKR-H/K/R all pass: the paper has a concrete visual-failure hook, a 40k-sample planning/RL mechanism, and relevance to multimodal-agent reliability. It remains a paper-level advance, not a product or flagship model release.

AI HOT (Curated Pool)

Grok Build Beta Opens to SuperGrok Users

xAI opened Grok Build Beta to all SuperGrok and X Premium+ users, with Plan Mode, Imagine-based image and video creation, and a CLI for automation or orchestrator workflows at x.ai/cli.

Why it matters: HKR-H/K/R all pass: xAI opened a paid beta with named workflow features. The score stays at the featured floor because the post lacks capability limits, pricing detail, and test results.

May 25Monday

r/LocalLLaMA

NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction

Numind released NuExtract3, a 4B open-weight VLM based on Qwen3.5-4B under Apache-2.0, supporting image and text to Markdown, OCR, and JSON-template extraction, with self-hosting from 4GB VRAM and weights in Safetensors, GGUF, and MLX formats.

Why it matters: HKR-H/K/R all pass: NuExtract3 packages OCR, Markdown, and structured extraction into a 4B open-weight VLM with a 4GB self-hosting condition. Source and lab reach keep it in the low featured band.

Synced · WeChat

A 1B-Gaussian 3D World Runs in the Browser, Outperforming Fei-Fei Li’s Spark

Manycore Tech open-sourced Aholo Viewer, a browser-based 3D Gaussian Splatting viewer that used half the memory of Spark 2.0 in a 300M-Gaussian test, loaded 2x faster, rendered 3x faster, and supports scenes with up to 1B Gaussian points.

Why it matters: HKR-H/K/R all pass: the hook is vivid, the post gives 300M-point benchmarks and a 1B-point ceiling, and browser-side 3D deployment matters to practitioners. Score stays at 80 because this is a strong tool release, not a foundation-model event.

May 24Sunday

Synced · WeChat

ICML 2026: First Parallel Thinking Framework for Vision-Language Models

Visual Para-Thinker introduces a parallel thinking framework for vision-language models, using Pa-Attention and LPRoPE to isolate four visual reasoning paths and training on 163,000 question-answer pairs.

Why it matters: HKR-H/K/R pass: the ICML 2026 paper offers a concrete parallel-thinking mechanism, four isolated paths, and 163K training pairs. It remains a single research release without broad replication or product impact, so it fits 78–84.

Xinzhiyuan · WeChat

Anthropic’s Three Cards Surface: Mythos 1 Appears, Opus 4.8 Spotted

Xinzhiyuan says Anthropic’s claude-opus-4.8 appeared in Google Vertex AI, while a 59.8MB Claude Code source-map leak with 512,000 TypeScript lines exposed Sonnet 4.8 references and Mythos 1 clues tied to Claude Code and Claude Security.

Why it matters: HKR-H/K/R all pass, but this is a leak plus Vertex listing, not an Anthropic launch. No capability numbers, pricing, context window, or reproducible evals, so it stays in the 78–84 band.

r/LocalLLaMA

Vision-capable LLMs vs. OCR for long-document QA with charts, images, and tables

The author tested Claude Sonnet 4.5 on 171 questions from 30 image-heavy MMLongBench-Doc PDFs, comparing native PDF vision use with OCR pipelines. Native PDF ranked fifth of six at 52.0% accuracy and cost $0.2552 per query, while LlamaCloud premium with full context reached 59.6% at $0.1885 per query.

Why it matters: HKR-H/K/R pass: the post gives 30 PDFs, 171 questions, accuracy, and per-question cost for long-document QA. Limited sample and Reddit sourcing keep it in the featured-threshold band.

May 23Saturday

Synced · WeChat

FlashAR speeds up pretrained autoregressive image models by 22.9x using 0.05% data

Zhejiang University and the University of Adelaide introduced FlashAR, using 0.05% of the original training data to reduce Emu3.5-Image-34B 512×512 generation latency from 130.10 seconds to 5.68 seconds, while GenEval changed from 80.48 to 80.29.

Why it matters: HKR-H/K/R all pass: FlashAR gives speedup, data ratio, latency, and GenEval deltas for AR image inference. It is a strong research item, but not a top-lab model release, so 80 featured rather than P1.

The Verge · AI

Google’s New Anything-to-Anything AI Model Is Wild

The Verge tried Google’s new Gemini anything-to-anything model for a stuffed-deer deepfake video, but the RSS snippet discloses only one example and does not disclose model parameters, pricing, release timing, or safety controls.

Why it matters: HKR-H/R pass: a Google/Gemini multimodal hands-on has a strong deepfake hook and safety resonance. HKR-K fails because the feed discloses one example only, with no params, pricing, or launch timing.

r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

AI HOT (Curated Pool)

Gemini update: over 900 million users and new agent features

Google announced that the Gemini app has surpassed 900 million monthly active users and introduced two agent features: Daily Brief for personalized daily summaries and Gemini Spark, a 24/7 personal agent that manages tasks under user authorization.

Why it matters: HKR-H/K/R all pass: Google gives a 900M MAU number and two agent features for Gemini. This is an entry-point product update with competitive weight, not a routine small feature.

May 22Friday

TechCrunch · AI

We tried Google’s AI glasses and they’re almost there

Google demonstrated prototype Android XR glasses that overlay Gemini-powered translation, navigation, and other information into the user’s field of view; the post does not disclose pricing, launch timing, battery life, or hardware specifications.

Why it matters: HKR-H/K/R all pass: TechCrunch tested Google’s Android XR glasses and identified Gemini overlays for translation and navigation. Price, launch timing, and battery life are not disclosed, keeping it in the lower featured band.

AI HOT (Curated Pool)

Project Genie and Google Maps Street View launch interactive worlds

Project Genie partnered with Google Maps Street View to turn real U.S. locations into interactive worlds; the post does not disclose supported cities, generation mechanics, pricing, or access scope.

Why it matters: Google DeepMind’s official post says Genie × Street View turns real US locations into interactive worlds, so HKR-H and HKR-R pass. HKR-K fails because cities, generation method, and access are not disclosed.

AI HOT (Curated Pool)

NetEase Youdao Open-Sources Ziyue 4 Multimodal and Text-to-Speech Models

NetEase Youdao open-sourced its Ziyue 4.0 multimodal and text-to-speech models, with the 27B multimodal model reporting 81.4% accuracy on Chinese math reasoning tasks and the speech model supporting 14 languages.

Why it matters: HKR-H/K/R pass: the story has a concrete open-source hook, specific model numbers, and practitioner relevance. NetEase Youdao is not a frontier lab, so it stays below the 78+ good-quality band.

Hacker News front page

Deepfakes Tore a High School Apart

404 Media reports that five girls at Radnor Township High School were targeted with AI-generated CSAM, and a freshman allegedly spent $250 on a Movely subscription from Apple’s App Store; the visible article does not disclose the police outcome.

Why it matters: HKR-H/K/R all pass: 404 Media reports a concrete AI CSAM school incident with victim count and tool cost. It is strong safety-policy signal, not a model or platform launch, so it stays in the 78–84 band.

Synced · WeChat

CVPR 2026 | HiF-VLA: A Motion-Centric World Action Model

Westlake University and collaborators introduced HiF-VLA, a motion-centric VLA framework that extracts compact Motion vectors with codecs such as H.264 and uses a joint expert to predict future visual motion and generate action sequences, reporting 31.4GB peak memory and 117.7ms latency under the cited history-window setting.

Why it matters: HKR-H/K/R all pass: the H.264-motion angle, concrete VRAM/latency numbers, and robotics deployment pressure are clear. It remains a single research item without adoption or cross-source heat, so it sits in the lower featured band.

Synced · WeChat

Meta Chinese Researcher Releases ATLAS for Generalizable Visual Reasoning with One Word

Meta AI and the Chinese University of Hong Kong proposed ATLAS, a visual reasoning method that uses one Functional Token to connect Agentic and Latent Visual Reasoning, with ATLAS-178K, a two-stage SFT+RL pipeline, and LA-GRPO to train sparse visual-operation tokens.

Why it matters: HKR-H/K/R pass: the one-token angle is clickable, and the post gives dataset and training details. As a Meta AI/CUHK research release rather than a flagship model or product launch, it fits the 78–84 band.

AI HOT (Curated Pool)

Plastic Interfaces: The Future Shape of AI-Driven Software

Salesforce has adopted a headless architecture that lets salespeople update data through AI; the post says MCPs, HTML, audio, and web interfaces can be generated dynamically by context, but it does not disclose implementation metrics or adoption numbers.

Why it matters: HKR-H/K/R all pass, but this is a software-form thesis without user metrics, launch timing, or a reproducible test. It fits the insightful-commentary band, not a must-write release.

AI HOT (Curated Pool)

Aleph 2.0 and Edit Studio

Runway released Aleph 2.0 and Edit Studio, combining generation, editing, and post-production into one platform; the post does not disclose pricing, technical parameters, or rollout scope.

Why it matters: Runway is a major AI video vendor, and Aleph 2.0 plus Edit Studio is a mid-weight product update. HKR-H/K/R pass, but missing price, specs, and rollout keep it at the featured threshold.

May 21Thursday

TechCrunch · AI

Hark raises $700M Series A for its secretive ‘universal’ AI interface

Hark raised a $700 million Series A and plans to release its first multimodal models this summer; the post does not disclose investors, valuation, model specifications, or a hardware launch schedule.

Why it matters: HKR-H/K/R all pass: the $700M Series A makes Hark a serious AI-interface contender. Investors, valuation, model specs, and hardware timing are not disclosed, so this stays featured rather than must-write.

Xinzhiyuan · WeChat

USTC Papers Study Lifelong Learning for LMMs via Multimodal Knowledge Injection

USTC researchers released MMEVOKE and KORE: MMEVOKE contains 9,422 samples across 159 subcategories, while KORE uses knowledge-tree augmentation and null-space constrained fine-tuning to reduce catastrophic forgetting during multimodal knowledge injection.

Why it matters: HKR-K and HKR-R are solid: the post gives dataset size plus a concrete fine-tuning mechanism. It stays in the low featured band because this is paper-level knowledge injection without production evidence or full reproducibility details.

Synced · WeChat

VAST and Tsinghua propose density-controlled 3D Gaussian generation for SIGGRAPH 2026

VAST and Tsinghua propose DeG, a 3D Gaussian generation method that samples Gaussian centers from a learned density distribution and trains density control with a render loss contribution gradient; in some settings, it reaches TRELLIS-like visual quality with less than half the Gaussian count.

Why it matters: HKR-H/K/R pass: DeG offers a concrete mechanism and a testable efficiency claim, reaching TRELLIS-like quality with under half the Gaussians in some scenes. SIGGRAPH research has some technical depth, but no hard-exclusion rule applies.

Synced · WeChat

Xie Saining’s Team Releases Second-Generation Representation Autoencoder RAEv2

Xie Saining’s team, Adobe Research, and the Australian National University released RAEv2, which reaches gFID 1.06 after 80 epochs on ImageNet-256 and reduces EPFID@2 from 177 epochs to 35 epochs while keeping compute at 189 GFLOPs.

Why it matters: HKR-K and HKR-R pass with concrete benchmark and training-efficiency claims. HKR-H is weak because the angle is a normal research release, so it lands at the featured threshold rather than a must-write item.

AI HOT (Curated Pool)

Tencent Launches OS-Level AI Assistant Mavis on Windows, Mac, and Android

Tencent launched the OS-level AI assistant Mavis on May 21 across Windows, Mac, and Android, with document parsing, image recognition, system maintenance, partial offline use, model dispatching, and desktop control of mobile apps listed as supported functions.

Why it matters: HKR-H/K/R all pass: Tencent’s OS-level assistant spans Windows, Mac, and Android with concrete tool abilities. Model, pricing, and permission design are not disclosed, so it stays at the lower featured band.