OpenVLA
Open 7B Vision-Language-Action model pretrained on Open X-Embodiment (~970k demos). Canonical paper is arXiv:2406.09246 (2024); for 2025 finetuning SOTA see OpenVLA-OFT (arXiv:2502.19645).
VLA that fuses 3D perception for stronger spatial understanding and manipulation planning.
3D-VLA explores bringing 3D perception into vision-language-action models, targeting richer spatial reasoning for manipulation. It is a research-oriented line: useful when 2D-only VLAs fail on geometry-heavy tasks.
Researchers combining 3D vision with language-conditioned control.
Content reviewed 2026-07-23