EEmbodied AI Hub
ModelPaper only

RT-2

Google DeepMind VLA that transfers internet vision-language knowledge into robot control.

Overview

RT-2 is Google DeepMind’s seminal Vision-Language-Action work that casts robot actions as text-like tokens produced by a VLM backbone. It established the modern VLA narrative: web-scale semantic knowledge transferred into robot control.

Public reproducibility is limited compared with OpenVLA; treat RT-2 primarily as a conceptual and paper reference unless you have access to internal stacks.

Who it is for

Readers studying VLA history and design; survey authors.

Key highlights

  • Defined the VLA paradigm for many follow-ups
  • Web-scale semantics → control story
  • Influential evaluation narrative

When to use

  • Literature review, design inspiration, citation baseline

When not to use

  • You need fully open weights for local finetune today

Getting started

  1. 1Read the paper carefully
  2. 2Compare claims with OpenVLA/Octo open stacks
  3. 3Reproduce ideas, not proprietary weights

Papers & reading

How it compares

Caveats & pitfalls

  • Not a drop-in open checkpoint for most labs.

Content reviewed 2026-07-23