VIMA-Bench
Multimodal promptable manipulation benchmark accompanying VIMA models.
Simulation benchmark for lifelong learning and VLA eval with many manipulation tasks.
Why: Among the most used VLA evaluation benchmarks
LIBERO is a lifelong / language-conditioned robot manipulation benchmark with procedural suites (Spatial, Object, Goal, Long). It is widely used to stress compositional generalization of IL and VLA policies.
Not a “starter toy”: set up the env carefully and track success rates per suite. Datasets and eval protocols are part of the package.
Learning Path: imitation evaluation milestone after you can train a single-task policy.
VLA / lifelong learning authors needing standard sim eval.
git clone https://github.com/Lifelong-Robot-Learning/LIBERO.git
Multimodal promptable manipulation benchmark accompanying VIMA models.
Content reviewed 2026-07-28