EEmbodied AI Hub
ProjectFeaturedAdvanced

CALVIN benchmark

Long-horizon language-conditioned manipulation benchmark with chain-of-tasks evaluation.

Why: Standard long-horizon language-conditioned eval

Overview

CALVIN benchmark is a curated project on Embodied AI Hub.

Long-horizon language-conditioned manipulation benchmark with chain-of-tasks evaluation.

Why it matters: Standard long-horizon language-conditioned eval

Always verify install pins, licenses, and hardware requirements on the official site before production use.

Who it is for

Builders mapping open embodied stacks; researchers comparing SOTA baselines.

Key highlights

  • Long-horizon language-conditioned manipulation benchmark with chain-of-tasks evaluation.
  • Type: project; tags: 基准, 语言条件, 长时程, 开源
  • Hub recommended

When to use

  • You need this category of open resource in your stack map
  • You are shortlisting baselines before deep evaluation

When not to use

  • License or hardware constraints block adoption
  • You only need a closed commercial stack with vendor SLA

Getting started

  1. 1Open the official URL and read the README / model card.
  2. 2Check license and citation requirements.
  3. 3Run the smallest official example before scaling.
  4. 4Cross-link related hub resources for a full pipeline.

Papers & reading

Caveats & pitfalls

  • Hub cards are curated summaries — verify upstream docs.
  • Stars and release names drift; treat metadata as approximate.

Content reviewed 2026-07-28