RVT / RVT-2
Robot View Transformer — multi-view transformers for language-conditioned 3D manipulation.
Perceiver-Actor — voxel-based language-conditioned manipulation widely used as a 3D baseline.
Why: Standard 3D language-conditioned manipulation baseline
PerAct is catalogued on Embodied AI Hub as a curated project for embodied AI builders.
Perceiver-Actor — voxel-based language-conditioned manipulation widely used as a 3D baseline.
Why it matters: Standard 3D language-conditioned manipulation baseline
Use the official repository and docs as the source of truth for install pins, licenses, and hardware requirements.
Researchers, students, and builders comparing open embodied stacks.
Content reviewed 2026-07-28