Skip to content

#具身智能

2 today

Jul 7Tuesday

AI HOT (Curated Pool)

First American autonomous ground vehicles are fighting in Ukraine, over 100 Forterra ATVs deployed

Forterra disclosed it has deployed more than 100 self-driving ATVs in Ukrainian conflict zones over the past nine months, the largest known combat use of autonomous ground vehicles by a US defense tech firm. The vehicles handle logistics and recon. A company exec said no defense tech is proven until it hits real combat. Ukraine turned to ground autonomy because aerial drones have made soldiers extremely vulnerable. The post does not disclose weapon configurations or loss figures.

Why it matters: First large-scale combat deployment of US autonomous ground vehicles, with concrete numbers and tactical context — not just PR. But the article doesn't disclose weaponization status or autonomy level, so it stays at the featured threshold.

Jul 3Friday

Hacker News front page

Yann LeCun says LLMs are 'not smart' and his AMI Labs is building a more flexible AI

Yann LeCun argued at VivaTech that LLMs like ChatGPT can't handle real-world complexity and aren't a path to human-level intelligence. His new venture AMI Labs, founded after leaving Meta in 2025, is building JEPA—an architecture that learns abstract representations instead of memorizing statistical patterns. The company raised over $1B in seed funding from Nvidia and Jeff Bezos' family fund. Oxford's Ingmar Posner is pursuing a similar direction with world models that reason about causality. The post does not disclose JEPA's performance benchmarks or a product timeline.

Why it matters: LeCun's first major public pitch for AMI Labs' JEPA approach since leaving Meta, with BBC giving it substantial coverage. Not scoring higher because it's still directional — no runnable model or benchmark numbers yet, far from shipping.

Jun 29Monday

Import AI (Jack Clark)

NVIDIA builds a self-improving loop for robots; Tencent details its 10k-GPU debug tool

NVIDIA's ENPIRE lets physical robots self-improve through trial and error like coding agents, hitting 99% on tasks like GPU insertion and zip-tie cutting. The catch: auto-evaluation and auto-reset still break on harder tasks. Tencent open-sourced ARGUS, an always-on tracing system for 10k+ GPU training clusters, already battle-tested for six months. A separate law paper points out that top minds badly misjudged nuclear fission and the internet—today's AI hot takes will likely age just as poorly.

Why it matters: NVIDIA's ENPIRE ports the agent trial-and-error loop to physical robots, hitting 99% on GPU insertion but still failing on auto-eval and reset for harder tasks. HKR all hit, but this is a newsletter digest rather than the primary paper, so information density is diluted — capp...

Jun 26Friday

AI HOT (Curated Pool)

General Intuition raised $320M, betting video game data can train general AI agents

General Intuition raised $320M to train AI on millions of hours of gameplay footage. Founder Pim de Witte argues that keystrokes, mouse movements, and decision sequences teach models physical reasoning better than text. They plan to sell the resulting models to robotics firms and game developers. The post does not disclose valuation or investor names.

Why it matters: A $3.2B raise with a novel training-data thesis hits all three HKR axes. Held at 78 rather than higher because the post doesn't disclose valuation or specific investors — key facts are missing.

Jun 17Wednesday

AI HOT (Curated Pool)

AWS open-sources Strands Robots SDK: one agent stack from Hugging Face Hub to physical robots

AWS released the Strands Robots SDK under Apache 2.0, wrapping the LeRobot stack into a unified agent. It defaults to MuJoCo simulation with no hardware needed; switch to mode="real" for physical robots. Recorded demos are saved as LeRobotDataset and can be pushed to Hugging Face Hub. Policies like GR00T or LerobotLocal run inference, then broadcast commands to multiple robots over Zenoh mesh. Simulation and hardware code are identical except for one keyword argument. Examples run in a notebook with Python 3.12+ on Linux/macOS, no GPU required.

Why it matters: AWS wraps LeRobot into a unified agent SDK with one-click sim-to-real switching — a solid tool for robotics devs. But pure physical robotics has limited resonance with AI app-layer readers, so R axis isn't fully hit, landing right at the featured threshold.

Jun 16Tuesday

AI HOT (Curated Pool)

Qwen-RobotManip: Alignment unlocks scale for robotic manipulation foundation models

Qwen team released Qwen-RobotManip, a foundation model for robotic manipulation. The key insight: alignment, not just larger pretraining, is what makes scale pay off. Demos show cross-embodiment generalization across real robots—stacking bowls, folding clothes, making burgers, arranging flowers—with Qwen-Omni issuing open-ended voice commands on the fly, no predefined task list. The post does not disclose model size, training data scale, or latency figures; only demo videos and a paper link are provided.

Why it matters: Qwen-RobotManip isn't just another robotics model — it uses alignment instead of more pre-training data to unlock scale, with live demos where Qwen-Omni gives random voice commands and the arm executes on the fly. Score stays below 85 because the post doesn't disclose preferen...

Jun 12Friday

TechCrunch · AI

Jeff Bezos's Prometheus raises $12B to build an 'artificial general engineer' for the physical world

Prometheus raised $12B at a $41B valuation. The startup targets automating heavy engineering and drug design in the physical world. The post only discloses the round size and valuation—no details on tech approach, team, or how the money will be spent.

Why it matters: $12B at a $41B valuation with Jeff Bezos behind it — a raise this size in physical AI is rare and worth featuring. But the post is thin: no tech approach, no team, no spending plan. K is a miss, so the score stays at 78.

Jun 10Wednesday

AI HOT (Curated Pool)

Huawei Cloud launches CloudRobo, an end-to-end embodied AI platform spanning data to deployment

Huawei Cloud announced CloudRobo, an end-to-end embodied AI platform that covers data, model training, deployment, and integration on a petabyte-scale trusted data base. At INSPIRE2026, partners showed dual evaluation for data and models, fast assembly of active force-control models, robot cloud onboarding in hours, and model deployment in minutes. The post doesn't disclose pricing or regional availability—I'd hold off on the 'world's first' claim until we see production deployments.

Why it matters: Huawei Cloud launched CloudRobo, an embodied AI platform with concrete specs on data foundation and deployment speed — not empty marketing. But it's a first-party announcement lacking third-party validation and real case details, so the score stays below 85.

Jun 9Tuesday

AI HOT (Curated Pool)

Two Chinese ministries target routine deployment of humanoid robots by end-2026

China’s MIIT and SASAC require key products including humanoid robots to complete application validation and enter routine deployment by the end of 2026, targeting more than 100 high-value scenarios and deployment at the ten-thousand-unit scale.

Why it matters: HKR-H/K/R all pass: the policy gives a hard humanoid-robot deployment timeline and scale targets. It stays in the lower featured band because budget, pilot names, and enforcement mechanics are not disclosed.

Jun 8Monday

Import AI (Jack Clark)

AI learns to game society's rules, and Anthropic sees 8x code growth in a year

Three highlights: a new benchmark, SocioHack, shows RL-trained models are good at exploiting real-world rules like credit card points or school grades, with over 90% precision on historical loopholes. Anthropic reports an 8x increase in merged code in 2026 vs 2021-2024 and says a prosaic form of recursive self-improvement may have begun, though no paradigm-shifting ideas yet. Separately, RL-trained racing drones from UZH and Google DeepMind beat a champion human pilot in multi-player races at over 22 m/s while cutting collisions by 50%.

Why it matters: Three solid items, with Anthropic's RSI disclosure as the standout exclusive signal. SocioHack's 90% reproduction accuracy and the drone RL's 11ms latency are both concrete. The ding: this is a newsletter roundup, not a first-party release — each item individually would clear ...

AI HOT (Curated Pool)

Amap Releases 3D-Native City World Model ABot-Earth0.5

Amap released ABot-Earth0.5, a 3D-native city world model covering more than 190 countries and regions, generating kilometer-scale 3D cities from satellite images or text within 10 minutes on consumer GPUs.

Why it matters: HKR-H/K/R all pass: Amap’s ABot-Earth0.5 has concrete claims, including 190+ countries and 10-minute km-scale 3D city generation. Strong world-model product signal, but below a major foundation-model release.

Jun 6Saturday

Xinzhiyuan · WeChat

Lion Rock AI Lab wins ICRA 2026 LeHome Challenge real-robot final

Lion Rock AI Lab won first place in the ICRA 2026 LeHome Challenge real-robot final, using LiOS to connect training, deployment, trajectory sampling, and Real2Sim teleoperation in one data iteration loop.

Why it matters: HKR-H/K/R all pass, but this is a robotics challenge result rather than a model or shipped product. The real-robot final win and LiOS loop justify featured, not p1.

Synced · WeChat

Daxiao Robotics and NTU Release PhysX-Omni for Simulation-Ready Physical 3D Generation

PhysX-Omni models rigid, deformable, and articulated objects in one simulation-ready 3D generation framework, while PhysXVerse contains over 8.7K physical 3D assets across more than 2.9K categories.

Why it matters: HKR-H and HKR-K pass: unified physical modeling plus 8.7K/2.9K+ dataset figures add substance. Source authority and entity weight are mid-tier, and the headline carries promo language, so it stays near the featured threshold.

Jun 5Friday

Xinzhiyuan · WeChat

The first robot to enter 100,000 homes wins the opening round

Xinzhiyuan says Weilan Technology has sold 25,000 quadruped robots, with home users accounting for 90% across 295 cities; its BabyAlpha A3 raises compute by 1,000x and runs a 7B-parameter model on-device.

Why it matters: HKR-H/K/R all pass: the 100,000-home hook is clickable, and the post gives sales, city coverage, and on-device model details. Kept in the low featured band because the data appears single-source and company-led, not an independently verified industry break.

Synced · WeChat

MetaFine proposes a diagnostic meta-evaluation framework for fine-grained robot manipulation

Southeast University and Peking University researchers introduced MetaFine, a diagnostic meta-evaluation framework that tests fine-grained robot manipulation across understanding, perception, and behavior, and the article says traditional binary success metrics can overestimate fine-manipulation capability by up to 70%.

Why it matters: HKR-H comes from the success-rate illusion hook; HKR-K adds MetaFine’s three-axis diagnostic and a 70% overestimation claim; HKR-R fits robotics eval trust. Research scope keeps it at the low end of 78-84.

QbitAI · WeChat

Yao Shunyu Responds to Whether Tencent Is Behind in AI

Yao Shunyu said at Tencent Cloud’s AI industry application conference that Hunyuan 3 rebuilt pretraining and reinforcement-learning infrastructure, changed data and evaluation, and assigned its strongest post-training staff to improve Yuanbao first; he named coding agents, multimodality, and embodied AI as Tencent’s next focus areas.

Why it matters: HKR-H/K/R all pass, but the facts are conference remarks and roadmap signals, not a new model release with specs, benchmarks, or launch date. This fits the lower featured band for a major Chinese tech AI strategy update.

QbitAI · WeChat

Instead of Spending 10 Billion on Humanoids, Put 100,000 Robot Dogs in Homes First

Weilan Technology’s BabyAlpha series has sold 25,397 units, with 90% used in home settings, while the A3 runs a 7B-parameter model on-device and reports 280 tokens/s inference under its disclosed configuration.

Why it matters: HKR-H/K/R all pass, but this is one company’s robot-dog commercialization story, not a top-lab model or platform launch. Concrete sales and edge-inference numbers put it at the upper end of mid-weight product updates.

Ruan YiFeng's Weblog

Tech Enthusiasts Weekly Issue 399: Visits to China’s AI Majors

Ruan Yifeng excerpts observations from U.S. analysts who visited 14 Chinese AI and robotics companies in early May: the article estimates U.S. AI compute at about 8 times China’s by the end of 2025, while Chinese firms’ intelligence output per unit of compute is estimated at 4-7 times naive scaling.

Why it matters: All three HKR axes pass: many named visit targets, concrete compute ratios, and a China-US AI competition nerve. It is still a secondary commentary post, not a primary release or major product event, so it sits just above the featured threshold.

AI HOT (Curated Pool)

Alex Imas and Phil Trammell: What Remains Scarce After AGI?

Alex Imas and Phil Trammell argue that robots can be copied and scaled after AGI, while scarce human skills such as ballet performance remain fixed; the post does not disclose a quantitative model or timeline.

Why it matters: HKR-H and HKR-R are strong because the angle reframes post-AGI labor scarcity; HKR-K passes on a concrete scarcity mechanism, but no quantitative model is disclosed. This fits the 72–77 commentary band.

Jun 4Thursday

The Verge · AI

Amazon develops a warehouse robot that workers can speak to

Amazon announced a new Proteus warehouse robot that accepts natural-language task instructions from workers; the original Proteus was announced in 2022, and the RSS snippet does not disclose deployment scale or pricing.

Why it matters: This is a mid-weight Amazon Proteus robotics update: HKR-H has the talk-to-robot hook, HKR-K adds a task-assignment mechanism, and HKR-R hits physical automation and labor impact. Deployment scale is not disclosed, so it stays near the featured floor.

QbitAI · WeChat

CVPR 2026: NVIDIA, Tesla, and Waymo hear Xpeng present physical AI

Xpeng presented its world-model stack at CVPR 2026, covering X-World, X-Foresight, and X-Cache; the article says X-Cache cuts about 70% of repeated computation, the second-generation VLA used over 4 trillion training tokens, and the in-car stack reduced inference latency to 80 ms.

Why it matters: HKR-H comes from the CVPR stage contrast, HKR-K has X-Cache, 4T+ tokens, and 80 ms latency, and HKR-R fits autonomy competition. It is still a company tech showcase, below the 85 must-write band.

Jun 3Wednesday

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

Synced · WeChat

RSS 2026: Ant Lingbo Proposes Autoregressive Causal World Model for Robot Manipulation with 50 Demos

Ant Lingbo and HKUST introduced LingBot-VA, an autoregressive video-action world model that unifies visual dynamics prediction and action inference, and the paper reports fine-tuning with 50 real-world demonstrations per task plus 92.0% and 91.1% success on RoboTwin 2.0 Easy and Hard settings.

Why it matters: HKR-H/K/R all pass: the hook is 50-demo robot control, with a concrete video-action world-model mechanism. Single-source coverage lacks code, benchmark detail, and deployment evidence, so it lands at 78.

QbitAI · WeChat

Daxiao Robot and NTU Release PhysX-Omni for Unified Physical 3D Generation

Daxiao Robot and NTU introduced PhysX-Omni, a unified simulation-ready physical 3D generation framework for rigid, deformable, and articulated objects, with PhysXVerse covering 8.7K assets across 2.9K categories and PhysX-Bench evaluating six dimensions including geometry, scale, material, affordance, kinematics, and description.

Why it matters: HKR-H/K/R all pass: unified physical 3D generation is a clear hook, the dataset and benchmark numbers add substance, and robotics simulation data is a real practitioner pain. No open-source or product adoption is disclosed, so it stays at 78.

Jun 2Tuesday

Synced · WeChat

Turing Award Winner Sutton’s New Paper Argues AI Should Move Toward Enactive Cognition

Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.

Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.

r/LocalLLaMA

NVIDIA releases Cosmos 3 Omnimodal world models on Hugging Face

NVIDIA released Cosmos 3 on Hugging Face with Nano at 16B parameters and Super at 64B parameters; the post says the models generate video, images, audio, and action commands from text, image, video, and action-trajectory inputs.

Why it matters: HKR-H/K/R all pass: NVIDIA world models on HF, concrete 16B/64B variants, and multimodal robotics relevance. Missing benchmarks, license, and training details keep it in the 78–84 band.

Latent Space

[AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark

NVIDIA released Cosmos 3 and Nemotron 3 Ultra; Cosmos 3 uses a Mixture-of-Transformers design with 16B Nano and 64B Super variants, while Nemotron 3 Ultra is described as a 550B-A55B open-weight model.

Why it matters: HKR-H/K/R all pass: NVIDIA ships Cosmos 3, Nemotron 3 Ultra, and RTX Spark with concrete MoT, 16B/64B, and 550B-A55B open-weight details. Impact is broad, but below a frontier-lab model release.

Jun 1Monday

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.

Synced · WeChat

OpenAI recruits for robotics team led by Sora creator Aditya Ramesh

OpenAI has listed more than a dozen San Francisco robotics roles for OpenAI Robotics, a team that evolved from Aditya Ramesh’s Worldsim work, with the actuator design engineer role offering $342,000 to $445,000 in base cash pay plus PPU incentives.

Why it matters: HKR-H/K/R all pass: OpenAI robotics hiring adds a strong hook, plus concrete roles, leader, and salary range. This is still a hiring signal, not a model or product release, so it stays in the featured-threshold band.

Synced · WeChat

World models get a “save state”: VAST releases Project Eden

VAST released Project Eden, a three-layer world-model architecture that separates persistent state evolution from visual rendering, and disclosed nearly $200 million across its A+ and A++ funding rounds.

Why it matters: HKR-H/K/R all pass: Project Eden has a product hook, architecture detail, and funding scale. VAST is not a top foundation-model lab, and benchmarks or access terms are not disclosed, so this lands in 78–84.

QbitAI · WeChat

How Cloud Models Reach the Physical World: CMG Lion Rock AI Lab Uses LiOS for Embodied AI

CMG Lion Rock AI Lab released the LiOS edge-cloud architecture for embodied robotics, reporting about 30 ms one-way latency from local camera to cloud GPU memory in cross-machine tests, and open-sourced the low-latency video transmission module plus the LeFold laundry-folding dataset.

Why it matters: HKR-H/K/R pass: LiOS offers a concrete latency claim and open artifacts for embodied AI. Impact stays mid-tier because the lab is not a top platform vendor and no cross-source cluster is shown.

QbitAI · WeChat

VAST Raises Nearly $200M and Discloses Its Project Eden World Model Roadmap

VAST raised nearly $200 million in A+ and A++ rounds and disclosed Project Eden, a world model architecture that separates state evolution from visual rendering through a structured state layer, a conditional interface layer, and a generative rendering layer.

Why it matters: HKR-H/K/R all pass: the $200M A+/A++ financing is sizable, and Project Eden gives a concrete three-layer world-model mechanism. VAST is not a top-tier foundation-model lab and no metrics or release details are disclosed, so this stays in the 78–84 band.

AI HOT (Curated Pool)

Introducing Cosmos Coalition

Runway joined Cosmos Coalition as a founding member and will co-develop the first open world-model foundation model for physical AI with NVIDIA.

Why it matters: HKR-H/K/R all pass: Runway plus NVIDIA and an open physical-AI world model is strong. Details are thin—no params, license, or benchmarks—so it stays in the 78–84 band.

AI HOT (Curated Pool)

Cosmos 3 Released: First Open Physical AI Generalist Model

NVIDIA released Cosmos 3 as an open physical AI generalist model with native visual reasoning, world generation, and action generation, offering two variants: Super at 32B parameters and Nano at 8B parameters.

Why it matters: HKR-H/K/R all pass: NVIDIA names two Cosmos 3 variants and concrete physical-AI capabilities. Source is a single launch post with no benchmark or license detail, so it stays in the 78–84 band.

AI HOT (Curated Pool)

MWC26 Shanghai to Host First Humanoid Robot Penalty Shootout With Unitree and 7 Other Teams

MWC26 Shanghai will host a humanoid robot penalty shootout in June 2026, with eight Chinese embodied intelligence teams competing under rules that require autonomous play without human control or preset scripts.

Why it matters: HKR-H/K/R all pass: the robot penalty shootout is clickable, with rules banning teleoperation and scripts. It stays in 72–77 because this is an event preview, not a model release or reproducible result.

AI HOT (Curated Pool)

OpenAI enters robotics and starts hiring

OpenAI formed the OpenAI Robotics team and is hiring full-stack hardware, systems, and ML engineers; Aditya Ramesh leads the project, with a near-term focus on supporting skilled workers, while the post does not disclose hiring scale.

Why it matters: HKR-H/K/R all pass: OpenAI’s robotics team and hiring push is a strong roadmap signal. Product form, timeline, and hiring scale are not disclosed, so it stays below P1.

May 31Sunday

Xinzhiyuan · WeChat

Fudan-Linked Team Releases STI-WM Spatiotemporally Integrated World Model

MouShen Intelligence released STI-WM, a spatiotemporally integrated world-action model for robotics, claiming support for RGB, point-cloud, and proprioceptive inputs, hundred-second task planning, and disclosing five funding rounds in six months plus a RMB 300 million Pre-A round.

Why it matters: HKR-H/K/R pass: STI-WM combines RGB, point clouds, and proprioception for 100-second planning, plus 5 funding rounds and a RMB300m Pre-A. Company-claim framing lacks public benchmarks or reproducible access, so it stays near the featured threshold.

QbitAI · WeChat

Robot-Native World Action Model Debuts With Spatiotemporal Architecture From Fudan-Linked Team

Moushen Intelligence released STI-WM, a spatiotemporally integrated world action model for robotics, with RGB, depth point cloud, and proprioceptive inputs; the post says it supports hundred-second-scale long-horizon task rollout and closed-loop replanning, but does not disclose benchmark scores or deployment costs.

Why it matters: HKR-H/K/R all pass: the STI-WM angle is novel, with concrete input modalities and hundred-second rollouts. Kept near the featured floor because public weights, benchmark results, and reproducible tests are not disclosed.

AI HOT (Curated Pool)

Tesla FSD completes a 6,000 km zero-intervention autonomous drive across Canada

Tesla FSD V14.3.3 completed a 6,051 km zero-intervention drive from Vancouver to Halifax in 4 days and 21 hours, with the system handling lane changes, complex road conditions, and parking without disengagements or human corrections.

Why it matters: HKR-H/K/R all pass: Tesla FSD V14.3.3 has a concrete 6,051 km zero-intervention claim. It stays below 85 because the item gives the result but lacks independent validation, route detail, and failure boundaries.

May 30Saturday

Financial Times · Technology

UK military looks at allowing lethal strikes without human approval

The FT headline says the UK military is examining lethal strikes without human approval, but the accessible body is a subscription page and does not disclose the weapon types, approval mechanism, legal conditions, or deployment timeline.

Why it matters: HKR-H and HKR-R are strong: the FT headline points at a lethal-autonomy policy red line. HKR-K fails because the accessible body is a subscribe page with no mechanism, timeline, or scope.