V-JEPA
Meta video JEPA world-model style representations — predictive features for physical video understanding.
Reusable visual representations for robot manipulation pretrained on egocentric video.
Why: Classic reusable visual encoder for manipulation IL
R3M is catalogued on Embodied AI Hub as a curated model for embodied AI builders.
Reusable visual representations for robot manipulation pretrained on egocentric video.
Why it matters: Classic reusable visual encoder for manipulation IL
Use the official repository and docs as the source of truth for install pins, licenses, and hardware requirements.
Researchers, students, and builders comparing open embodied stacks.
Content reviewed 2026-07-28