Skip to content
AI HOT (Curated Pool)

Qwen-RobotNav: One model, five navigation domains, and a tool-call primitive for agentic systems

Qwen-RobotNav:面向智能体导航系统的可扩展导航模型

Qwen released Qwen-RobotNav, a single set of weights built on Qwen3-VL and trained on 15.6M samples that handles instruction following, object search, tracking, driving, and embodied QA. It exposes visual context as tunable inference-time parameters—token budget, temporal decay, per-camera weights—so an upper-level planner (Qwen3.7-Plus) can reconfigure it per call without retraining. On EXPRESS-Bench it beats the prior best by 15.4% while using 77% fewer navigation steps. Zero-shot deployment on a Unitree Go2 with a single low-res camera works in unseen outdoor environments.

Why it matters: Qwen ships a robot navigation model built on Qwen3-VL with 15.6M samples across five tasks. The core pitch is a parameterized visual memory interface configurable at inference time—frame count, attention weights, no retraining needed. Paper and GitHub are available, but no rea...

Read the original ↗Export Markdown