Magma
Microsoft Magma multimodal foundation model for UI + robot action grounding research.
Embodied multimodal language model connecting vision, language, and robot control at scale.
Why: Landmark embodied multimodal LLM result
PaLM-E is catalogued on Embodied AI Hub as a curated model for embodied AI builders.
Embodied multimodal language model connecting vision, language, and robot control at scale.
Why it matters: Landmark embodied multimodal LLM result
Use the official repository and docs as the source of truth for install pins, licenses, and hardware requirements.
Researchers, students, and builders comparing open embodied stacks.
Content reviewed 2026-07-28