EEmbodied AI Hub
PapersEditor’s pick

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

Why it mattersThis work bridges simulation and real-world deployment for VLA models, leveraging AMD ROCm for cost-effective, scalable physical AI.

A pipeline for vision-language-action (VLA) manipulation using Real2Sim2Real transfer, built on AMD ROCm for efficient simulation-to-real deployment.

vision-language-actionReal2Sim2Realmanipulationphysical AIsim-to-realAMD ROCmresearchpaper

Open paper on arXiv (arXiv:2607.22997)

Source: arXiv · July 29, 2026

Share this article so more people can see it

This paper presents a Real2Sim2Real pipeline for vision-language-action (VLA) manipulation, built on AMD ROCm. The approach enables efficient transfer of policies learned in simulation to real-world robotic tasks.

The pipeline integrates large VLA models with embodied agents, addressing the challenge of sim-to-real gap. By using AMD ROCm, the system achieves high-performance computation for training and inference.

Key components include a simulation environment for data generation and policy learning, followed by real-world deployment with minimal fine-tuning. The Real2Sim2Real loop ensures robust performance across domains.

Experiments demonstrate successful manipulation tasks such as grasping and object rearrangement. The pipeline is designed to be modular and scalable for various robotic platforms.

This work contributes to the emerging field of physical AI, where models interact with the physical world. The use of AMD ROCm highlights an alternative to NVIDIA-based solutions.

Source: arXiv (2607.22997).

Related resources on this hub

Jump to projects, models, or datasets mentioned or closely related.

Discussion

Tell us what you think — comments make stories more useful for builders and founders.

Tell us what you think!

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

Have an account? Log in to use your display name and avatar.

Email is optional and never shown on the page.

More insights

PapersarXiv

A Replay-Constrained Simulation Framework for Personalization of Powered Knee-Ankle Prosthesis Controllers

Enables efficient personalization of prosthetic controllers, reducing reliance on time-intensive human-in-the-loop tuning and expanding optimization to high-dimensional parameter spaces.

A simulation framework that uses replay constraints to personalize impedance controllers for powered knee-ankle prostheses, enabling high-dimensional optimization without human-in-the-loop.

Read
PapersarXiv

Pose-Aware Modeling to Mitigate Pose-Related Artifacts in Tactile Gloves

Tactile gloves are crucial for dexterous manipulation and teleoperation, but pose artifacts limit their utility. This work directly addresses a key sensor limitation, enabling more accurate data collection for learning and control.

Tactile gloves digitize contact and force during hand-object interactions, but pose-related artifacts degrade data quality. This work proposes pose-aware modeling to mitigate such artifacts, improving tactile sensing reliability for robotics applications.

Read
PapersarXiv

OAT Tokenization: A New Interface for Visuomotor Policies That Balances Compression and Decodability

Enables more efficient and structured action representation for robot learning, potentially improving policy performance and scalability.

Ordered Action Tokenization (OAT) introduces a learned tokenizer that maps continuous robot actions into ordered discrete tokens, achieving high compression, total decodability, and an ordered token space. Tested across 60+ tasks and multiple policy backbones, OAT enables anytime inference tradeoffs and outperforms existing discretization methods. For builders, this means more flexible, efficient visuomotor policies without sacrificing performance.

Read