TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, et al.
arXiv:2609.09158v1Today’s standout paper is TANGO, a new vision-language navigation system for humanoid robots. The problem is harder than ordinary robot navigation: a humanoid moving through a cluttered room can’t just plan a 2D path, it has to coordinate its whole body in real time, shifting its torso, placing its arms, and adjusting its gait to avoid collisions. TANGO learns this directly from natural-language instructions and egocentric camera views, predicting low-level joint actions for 29 degrees of freedom. The authors train it entirely in simulation, but they don’t just collect random motion; they synthesize realistic, collision-free traversal behaviors using path planning, whole-body motion generation, obstacle-aware edits, and reinforcement learning. What makes this especially compelling is that the same policy transfers zero-shot to a real Unitree G1 humanoid and still works in cluttered indoor scenes. That suggests a practical route toward robots that can understand instructions and physically navigate human environments with much more flexibility.
Also spotted that day
Previous daily papers