EEmbodied AI Hub
ModelFully open

3D-VLA

VLA that fuses 3D perception for stronger spatial understanding and manipulation planning.

Overview

3D-VLA explores bringing 3D perception into vision-language-action models, targeting richer spatial reasoning for manipulation. It is a research-oriented line: useful when 2D-only VLAs fail on geometry-heavy tasks.

Who it is for

Researchers combining 3D vision with language-conditioned control.

Key highlights

  • 3D-aware VLA research direction
  • Spatial reasoning focus
  • Open research artifacts (check repo)

When to use

  • Geometry-critical manipulation
  • Depth/point-cloud fusion studies

When not to use

  • You only need a production 2D VLA baseline

Getting started

  1. 1Read paper + repo README
  2. 2Prepare 3D observations in the expected format
  3. 3Compare against 2D OpenVLA-style baselines

Papers & reading

Caveats & pitfalls

  • Ecosystem smaller than OpenVLA; expect more DIY.

Content reviewed 2026-07-23