This paper presents a Real2Sim2Real pipeline for vision-language-action (VLA) manipulation, built on AMD ROCm. The approach enables efficient transfer of policies learned in simulation to real-world robotic tasks.
The pipeline integrates large VLA models with embodied agents, addressing the challenge of sim-to-real gap. By using AMD ROCm, the system achieves high-performance computation for training and inference.
Key components include a simulation environment for data generation and policy learning, followed by real-world deployment with minimal fine-tuning. The Real2Sim2Real loop ensures robust performance across domains.
Experiments demonstrate successful manipulation tasks such as grasping and object rearrangement. The pipeline is designed to be modular and scalable for various robotic platforms.
This work contributes to the emerging field of physical AI, where models interact with the physical world. The use of AMD ROCm highlights an alternative to NVIDIA-based solutions.
Source: arXiv (2607.22997).