PerAct
Perceiver-Actor — voxel-based language-conditioned manipulation widely used as a 3D baseline.
Robot View Transformer — multi-view transformers for language-conditioned 3D manipulation.
Why: Strong multi-view 3D manipulation open implementation
RVT / RVT-2 is catalogued on Embodied AI Hub as a curated project for embodied AI builders.
Robot View Transformer — multi-view transformers for language-conditioned 3D manipulation.
Why it matters: Strong multi-view 3D manipulation open implementation
Use the official repository and docs as the source of truth for install pins, licenses, and hardware requirements.
Researchers, students, and builders comparing open embodied stacks.
Content reviewed 2026-07-28