BEHAVIOR-1K
1000 everyday household activities benchmark for long-horizon embodied agents (OmniGibson).
Long-horizon language-conditioned manipulation benchmark with chain-of-tasks evaluation.
Why: Standard long-horizon language-conditioned eval
CALVIN benchmark is a curated project on Embodied AI Hub.
Long-horizon language-conditioned manipulation benchmark with chain-of-tasks evaluation.
Why it matters: Standard long-horizon language-conditioned eval
Always verify install pins, licenses, and hardware requirements on the official site before production use.
Builders mapping open embodied stacks; researchers comparing SOTA baselines.
Content reviewed 2026-07-28