π0 (pi0)
Physical Intelligence generalist robot foundation model with flow-matching action generation.
Google DeepMind VLA that transfers internet vision-language knowledge into robot control.
RT-2 is Google DeepMind’s seminal Vision-Language-Action work that casts robot actions as text-like tokens produced by a VLM backbone. It established the modern VLA narrative: web-scale semantic knowledge transferred into robot control.
Public reproducibility is limited compared with OpenVLA; treat RT-2 primarily as a conceptual and paper reference unless you have access to internal stacks.
Readers studying VLA history and design; survey authors.
Physical Intelligence generalist robot foundation model with flow-matching action generation.
NVIDIA foundation model series for humanoid and general robots combining sim and real data.
Content reviewed 2026-07-23