Octo
Generalist robot policy Transformer with multi-modal inputs and multi-embodiment finetuning.
Hugging Face / LeRobot lightweight VLA (~0.5B) for fast prototyping and low-GPU finetune loops.
SmolVLA aims at a smaller footprint for learning the VLA loop on limited GPUs. Prefer it when your goal is a closed finetune loop, not the largest leaderboard model. Still verify licenses and expect domain gap on real tasks. Beginner-to-intermediate.
SmolVLA refers to small VLA efforts in the LeRobot / Hugging Face ecosystem designed for accessible finetuning. The point is not SOTA leaderboard scores — it is fitting vision-language-action training on modest GPUs.
Follow LeRobot docs for dataset format, training entrypoints, and eval. Combine with OXE subsets or your own teleop data.
Learning Path (VLA): first finetune target before full OpenVLA / π0-scale models.
Researchers, students, and builders comparing open embodied stacks.
via LeRobot ecosystem
pip install lerobot
Generalist robot policy Transformer with multi-modal inputs and multi-embodiment finetuning.
Content reviewed 2026-07-28