OpenVLA
Open 7B Vision-Language-Action model pretrained on Open X-Embodiment (~970k demos). Canonical paper is arXiv:2406.09246 (2024); for 2025 finetuning SOTA see OpenVLA-OFT (arXiv:2502.19645).
Low-cost imitation approach built on open VLMs for language-conditioned manipulation.
RoboFlamingo adapts Flamingo-style vision-language models to robot manipulation, emphasizing open-ended language-conditioned control with relatively accessible open components for research. It is part of the broader “VLM → robot” transition that led to modern VLAs.
Researchers studying VLM-based manipulation policies.
Content reviewed 2026-07-23