World's Fair 2025
Robotics: why now? - Quan Vuong and Jost Tobias Springberg, Physical Intelligence
Quan Vuong , Jost Tobias Springberg , Physical Intelligence
Overview
This talk explores the advancements in robotics driven by the emergence of vision-language action (VLA) models. It highlights the transition from robots operating in highly constrained environments to performing complex tasks in semistructured real-world settings. The core thesis is that significant progress in robotics is now enabled by general AI developments, particularly VLA models, and that the primary bottleneck is shifting from hardware to software and model intelligence.
Who should watch
- AI engineers and researchers interested in robotics and embodied AI.
- Product Managers and builders exploring applications for advanced AI in physical systems.
- Developers working on robot control, simulation, and data collection pipelines.
- Anyone curious about the future of robots performing complex, real-world tasks autonomously.
Key takeaways
- Vision-language action (VLA) models adapt vision-language models (VLMs) by taking robot state inputs and producing actions for robot control, rather than text outputs.
- Training VLAs presents unique challenges compared to VLMs, particularly in identifying analytical data sources and adapting model architectures for high-frequency robot control.
- Physical Intelligence (PI) is building a data engine from scratch, using human operators to teleoperate robots and collect extensive, high-quality demonstration data for training.
- Their data collection efforts have scaled significantly, moving from static scenes to diverse mobile manipulation setups, enabling more complex and generalized robot behaviors.
- PI has developed models like PI Zero and PIO5, with PIO5 demonstrating open-world generalization by performing long-horizon tasks in entirely unseen environments.
- The ability of VLAs to generalize appears to improve with increased diversity in training data, specifically by incorporating data from a wider range of environments.
- The success of VLA models in controlling diverse hardware, even without prior specific knowledge of the robot, suggests that software and model intelligence are becoming the primary drivers of robotic advancement.
- The field is still facing significant scientific, engineering, and operational challenges, indicating a need for continued research and talent in this area.
Notable quotes
*Our mission is to make a model that can control any robot to do any task.*
*It's our belief that actually one of the main bottleneck is software and just model intelligence.*
Unofficial community note. Prefer the recording for nuance.