Skip to content

Embodied AI

AI in the physical world: humanoid robots, embodied foundation models and real-world manipulation.

Latest picks

41–60 of 167

Jun 9Tuesday

AI HOT (Curated Pool)

Two Chinese ministries target routine deployment of humanoid robots by end-2026

China’s MIIT and SASAC require key products including humanoid robots to complete application validation and enter routine deployment by the end of 2026, targeting more than 100 high-value scenarios and deployment at the ten-thousand-unit scale.

Why it matters: HKR-H/K/R all pass: the policy gives a hard humanoid-robot deployment timeline and scale targets. It stays in the lower featured band because budget, pilot names, and enforcement mechanics are not disclosed.

Jun 8Monday

Import AI (Jack Clark)

AI learns to game society's rules, and Anthropic sees 8x code growth in a year

Three highlights: a new benchmark, SocioHack, shows RL-trained models are good at exploiting real-world rules like credit card points or school grades, with over 90% precision on historical loopholes. Anthropic reports an 8x increase in merged code in 2026 vs 2021-2024 and says a prosaic form of recursive self-improvement may have begun, though no paradigm-shifting ideas yet. Separately, RL-trained racing drones from UZH and Google DeepMind beat a champion human pilot in multi-player races at over 22 m/s while cutting collisions by 50%.

Why it matters: Three solid items, with Anthropic's RSI disclosure as the standout exclusive signal. SocioHack's 90% reproduction accuracy and the drone RL's 11ms latency are both concrete. The ding: this is a newsletter roundup, not a first-party release — each item individually would clear ...

AI HOT (Curated Pool)

Amap Releases 3D-Native City World Model ABot-Earth0.5

Amap released ABot-Earth0.5, a 3D-native city world model covering more than 190 countries and regions, generating kilometer-scale 3D cities from satellite images or text within 10 minutes on consumer GPUs.

Why it matters: HKR-H/K/R all pass: Amap’s ABot-Earth0.5 has concrete claims, including 190+ countries and 10-minute km-scale 3D city generation. Strong world-model product signal, but below a major foundation-model release.

Jun 6Saturday

Xinzhiyuan · WeChat

Lion Rock AI Lab wins ICRA 2026 LeHome Challenge real-robot final

Lion Rock AI Lab won first place in the ICRA 2026 LeHome Challenge real-robot final, using LiOS to connect training, deployment, trajectory sampling, and Real2Sim teleoperation in one data iteration loop.

Why it matters: HKR-H/K/R all pass, but this is a robotics challenge result rather than a model or shipped product. The real-robot final win and LiOS loop justify featured, not p1.

Synced · WeChat

Daxiao Robotics and NTU Release PhysX-Omni for Simulation-Ready Physical 3D Generation

PhysX-Omni models rigid, deformable, and articulated objects in one simulation-ready 3D generation framework, while PhysXVerse contains over 8.7K physical 3D assets across more than 2.9K categories.

Why it matters: HKR-H and HKR-K pass: unified physical modeling plus 8.7K/2.9K+ dataset figures add substance. Source authority and entity weight are mid-tier, and the headline carries promo language, so it stays near the featured threshold.

Jun 5Friday

Xinzhiyuan · WeChat

The first robot to enter 100,000 homes wins the opening round

Xinzhiyuan says Weilan Technology has sold 25,000 quadruped robots, with home users accounting for 90% across 295 cities; its BabyAlpha A3 raises compute by 1,000x and runs a 7B-parameter model on-device.

Why it matters: HKR-H/K/R all pass: the 100,000-home hook is clickable, and the post gives sales, city coverage, and on-device model details. Kept in the low featured band because the data appears single-source and company-led, not an independently verified industry break.

Synced · WeChat

MetaFine proposes a diagnostic meta-evaluation framework for fine-grained robot manipulation

Southeast University and Peking University researchers introduced MetaFine, a diagnostic meta-evaluation framework that tests fine-grained robot manipulation across understanding, perception, and behavior, and the article says traditional binary success metrics can overestimate fine-manipulation capability by up to 70%.

Why it matters: HKR-H comes from the success-rate illusion hook; HKR-K adds MetaFine’s three-axis diagnostic and a 70% overestimation claim; HKR-R fits robotics eval trust. Research scope keeps it at the low end of 78-84.

QbitAI · WeChat

Yao Shunyu Responds to Whether Tencent Is Behind in AI

Yao Shunyu said at Tencent Cloud’s AI industry application conference that Hunyuan 3 rebuilt pretraining and reinforcement-learning infrastructure, changed data and evaluation, and assigned its strongest post-training staff to improve Yuanbao first; he named coding agents, multimodality, and embodied AI as Tencent’s next focus areas.

Why it matters: HKR-H/K/R all pass, but the facts are conference remarks and roadmap signals, not a new model release with specs, benchmarks, or launch date. This fits the lower featured band for a major Chinese tech AI strategy update.

QbitAI · WeChat

Instead of Spending 10 Billion on Humanoids, Put 100,000 Robot Dogs in Homes First

Weilan Technology’s BabyAlpha series has sold 25,397 units, with 90% used in home settings, while the A3 runs a 7B-parameter model on-device and reports 280 tokens/s inference under its disclosed configuration.

Why it matters: HKR-H/K/R all pass, but this is one company’s robot-dog commercialization story, not a top-lab model or platform launch. Concrete sales and edge-inference numbers put it at the upper end of mid-weight product updates.

Ruan YiFeng's Weblog

Tech Enthusiasts Weekly Issue 399: Visits to China’s AI Majors

Ruan Yifeng excerpts observations from U.S. analysts who visited 14 Chinese AI and robotics companies in early May: the article estimates U.S. AI compute at about 8 times China’s by the end of 2025, while Chinese firms’ intelligence output per unit of compute is estimated at 4-7 times naive scaling.

Why it matters: All three HKR axes pass: many named visit targets, concrete compute ratios, and a China-US AI competition nerve. It is still a secondary commentary post, not a primary release or major product event, so it sits just above the featured threshold.

AI HOT (Curated Pool)

Alex Imas and Phil Trammell: What Remains Scarce After AGI?

Alex Imas and Phil Trammell argue that robots can be copied and scaled after AGI, while scarce human skills such as ballet performance remain fixed; the post does not disclose a quantitative model or timeline.

Why it matters: HKR-H and HKR-R are strong because the angle reframes post-AGI labor scarcity; HKR-K passes on a concrete scarcity mechanism, but no quantitative model is disclosed. This fits the 72–77 commentary band.

Jun 4Thursday

The Verge · AI

Amazon develops a warehouse robot that workers can speak to

Amazon announced a new Proteus warehouse robot that accepts natural-language task instructions from workers; the original Proteus was announced in 2022, and the RSS snippet does not disclose deployment scale or pricing.

Why it matters: This is a mid-weight Amazon Proteus robotics update: HKR-H has the talk-to-robot hook, HKR-K adds a task-assignment mechanism, and HKR-R hits physical automation and labor impact. Deployment scale is not disclosed, so it stays near the featured floor.

QbitAI · WeChat

CVPR 2026: NVIDIA, Tesla, and Waymo hear Xpeng present physical AI

Xpeng presented its world-model stack at CVPR 2026, covering X-World, X-Foresight, and X-Cache; the article says X-Cache cuts about 70% of repeated computation, the second-generation VLA used over 4 trillion training tokens, and the in-car stack reduced inference latency to 80 ms.

Why it matters: HKR-H comes from the CVPR stage contrast, HKR-K has X-Cache, 4T+ tokens, and 80 ms latency, and HKR-R fits autonomy competition. It is still a company tech showcase, below the 85 must-write band.

Jun 3Wednesday

NVIDIA Blog

NVIDIA Research Presents Grasping, Autonomous Driving and Agent Training Work at CVPR

NVIDIA Research presented three physical AI papers at CVPR: GraspGen-X was trained on 2 billion simulated grasps, LCDrive cuts reasoning tokens by about half versus text-based reasoning, and NitroGen trains embodied agents across more than 1,000 games and 40,000 hours of interaction.

Why it matters: HKR-H/K/R all pass: NVIDIA’s CVPR bundle gives concrete mechanisms and scale numbers. It stays in the low 78–84 band because it is a vendor research roundup, not a major model or product launch.

Synced · WeChat

RSS 2026: Ant Lingbo Proposes Autoregressive Causal World Model for Robot Manipulation with 50 Demos

Ant Lingbo and HKUST introduced LingBot-VA, an autoregressive video-action world model that unifies visual dynamics prediction and action inference, and the paper reports fine-tuning with 50 real-world demonstrations per task plus 92.0% and 91.1% success on RoboTwin 2.0 Easy and Hard settings.

Why it matters: HKR-H/K/R all pass: the hook is 50-demo robot control, with a concrete video-action world-model mechanism. Single-source coverage lacks code, benchmark detail, and deployment evidence, so it lands at 78.

QbitAI · WeChat

Daxiao Robot and NTU Release PhysX-Omni for Unified Physical 3D Generation

Daxiao Robot and NTU introduced PhysX-Omni, a unified simulation-ready physical 3D generation framework for rigid, deformable, and articulated objects, with PhysXVerse covering 8.7K assets across 2.9K categories and PhysX-Bench evaluating six dimensions including geometry, scale, material, affordance, kinematics, and description.

Why it matters: HKR-H/K/R all pass: unified physical 3D generation is a clear hook, the dataset and benchmark numbers add substance, and robotics simulation data is a real practitioner pain. No open-source or product adoption is disclosed, so it stays at 78.

Jun 2Tuesday

Synced · WeChat

Turing Award Winner Sutton’s New Paper Argues AI Should Move Toward Enactive Cognition

Banafsheh Rafiee and Richard S. Sutton propose an enactive cognition framework for AI, naming four pillars: experience, perception-action inseparability, autonomy, and embodiment.

Why it matters: HKR-H/K/R all pass, but the article centers on a conceptual framework and does not disclose experiments, code, or reproducible tests. Sutton’s name and the four pillars put it in the 78–84 research-commentary band.

r/LocalLLaMA

NVIDIA releases Cosmos 3 Omnimodal world models on Hugging Face

NVIDIA released Cosmos 3 on Hugging Face with Nano at 16B parameters and Super at 64B parameters; the post says the models generate video, images, audio, and action commands from text, image, video, and action-trajectory inputs.

Why it matters: HKR-H/K/R all pass: NVIDIA world models on HF, concrete 16B/64B variants, and multimodal robotics relevance. Missing benchmarks, license, and training details keep it in the 78–84 band.

Latent Space

[AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark

NVIDIA released Cosmos 3 and Nemotron 3 Ultra; Cosmos 3 uses a Mixture-of-Transformers design with 16B Nano and 64B Super variants, while Nemotron 3 Ultra is described as a 550B-A55B open-weight model.

Why it matters: HKR-H/K/R all pass: NVIDIA ships Cosmos 3, Nemotron 3 Ultra, and RTX Spark with concrete MoT, 16B/64B, and 550B-A55B open-weight details. Impact is broad, but below a frontier-lab model release.

Jun 1Monday

AI HOT (Curated Pool)

NVIDIA Open-Sources Cosmos 3, Its First Generalist Model for Physical AI

NVIDIA open-sourced Cosmos 3 at GTC Taipei, releasing two variants, Super 32B and Nano 8B, with model weights, code, and datasets made available.

Why it matters: HKR-H/K/R all pass: the concrete hook is NVIDIA opening Cosmos 3 with 32B/8B variants and released artifacts. The post is sparse and single-source, with no benchmarks or license details, so it stays in the 78–84 band.