DROID
Highly diverse real Franka manipulation data (~76K trajectories / 350h) for real-world finetune.
Large real manipulation corpus across 22 embodiments (1M+ trajectories) — key VLA pretrain base.
Why: Currently the largest cross-embodiment real robot manipulation dataset
Open X-Embodiment (OXE) is the multi-lab, multi-robot dataset collection that powered the open VLA wave. It aggregates trajectories across many embodiments and institutions, enabling cross-embodiment pretraining that single-lab datasets cannot support.
OXE is not one homogeneous corpus: licenses, sensors, and action spaces differ by constituent dataset. Successful users define explicit mixtures (sometimes called “magic soup” recipes), normalize actions carefully, and track which subsets dominate training.
If you train or finetune generalist policies, understanding OXE composition is as important as understanding model architecture.
Anyone pretraining or studying generalist VLAs; dataset mixture researchers.
Meta-dataset mixture, not a single robot corpus — check per-subset licenses.
Highly diverse real Franka manipulation data (~76K trajectories / 350h) for real-world finetune.
Large-scale real bimanual dataset from AgiBot across diverse scenes and tasks.
Content reviewed 2026-07-23