EEmbodied AI Hub
DatasetFeatured

CALVIN

Language-conditioned long-horizon manipulation sim benchmark emphasizing compositional generalization.

Why: Classic language-conditioned long-horizon benchmark

Overview

CALVIN is a language-conditioned long-horizon manipulation benchmark in simulation, emphasizing compositional generalization across instructions and environments. It stresses multi-step behaviors more than single short skills.

Use CALVIN when your method claims long-horizon language following, not only single-step pick-and-place accuracy.

Who it is for

Long-horizon language-conditioned policy researchers.

Key highlights

  • Long-horizon language tasks
  • Compositional generalization focus
  • Established sim benchmark

When to use

  • Multi-step instruction following eval

When not to use

  • Only short-horizon single skills

Getting started

  1. 1Install CALVIN
  2. 2Run sequence evaluation protocols
  3. 3Report full multi-step success

Papers & reading

How it compares

DatasetFeatured

RoboCasa

Large kitchen sim suite with generative scenes and imitation/RL data.

100+ tasks, large synthetic demosCheck officialBest for pretrainBest for evalBest for finetuneSingle-armBimanualUpdated 2025-10-01

Caveats & pitfalls

  • Long-horizon metrics are strict — ablate carefully.

Content reviewed 2026-07-23