R3M
Reusable visual representations for robot manipulation pretrained on egocentric video.
Meta video JEPA world-model style representations — predictive features for physical video understanding.
Why: Influential open predictive video representation model
V-JEPA is a curated model on Embodied AI Hub.
Meta video JEPA world-model style representations — predictive features for physical video understanding.
Why it matters: Influential open predictive video representation model
Always verify install pins, licenses, and hardware requirements on the official site before production use.
Builders mapping open embodied stacks; researchers comparing SOTA baselines.
Content reviewed 2026-07-28