Octo
Generalist robot policy Transformer with multi-modal inputs and multi-embodiment finetuning.
Optimized Fine-Tuning (OFT) for VLAs (RSS 2025). Boosts OpenVLA on LIBERO (~76.5%→97.1%) with parallel decoding, action chunking, and continuous actions; strong real ALOHA results.
Why: 2025 OpenVLA follow-up: OFT finetune recipe + paper arXiv:2502.19645
OpenVLA-OFT is the 2025 follow-up to the base OpenVLA model paper. Paper: Fine-Tuning Vision-Language-Action Models (arXiv:2502.19645, RSS 2025). It studies action decoding, continuous actions, and learning objectives, and proposes an Optimized Fine-Tuning (OFT) recipe.
Reported gains include lifting OpenVLA’s LIBERO average success from about 76.5% to 97.1% while increasing action throughput substantially, plus strong real bimanual ALOHA results versus π0 / RDT-style finetunes. If you already use OpenVLA weights, this is usually the first paper/code to read after the 2024 base paper.
Teams finetuning OpenVLA on limited demos; readers wanting the 2025 OpenVLA-line paper.
Content reviewed 2026-07-23