EEmbodied AI Hub
IndustryDeploymentEditor’s pick

ACE-Data-0: A New Benchmark for Embodied AI Data in Real Homes

Daxiao Robotics and NTU S-Lab release ACE-Data-0, an L5-level multimodal dataset with 200 tasks, 17 million frames, and 150 hours of real-home interaction data. It captures synchronized video, motion, touch, and audio, aiming to train more robust embodied models. The dataset includes a three-tier benchmark revealing gaps in current models under occlusion and long-horizon tasks.

Daxiao Roboticsfeed:leiphoneACE-Data-0multimodaldeploymentleiphonedatasetchina

Source: 雷峰网 · August 1, 2026

Share this article so more people can see it

On July 31, 2026, Daxiao Robotics, in collaboration with NTU S-Lab, released ACE-Data-0, an open-source dataset designed to push embodied AI beyond controlled lab settings. The dataset is built on two novel concepts: the 'information density law' for embodied models and a human-centric environmental data collection scheme. With 200 task categories, 17 million video frames, 150 hours of data, 50 participants, and 75,000 interaction segments across two real home environments, ACE-Data-0 aims to provide a high-fidelity data foundation for training general-purpose physical intelligence.

What sets ACE-Data-0 apart is its focus on capturing the full physical process, not just isolated actions. Traditional datasets often record short clips of grasping or placing objects, missing the long-term dependencies inherent in real tasks. ACE-Data-0 records continuous sequences that include perception, action, contact, and state changes, all synchronized across multiple modalities. This includes first-person and third-person video, full-body and hand motion, object 6-DoF trajectories, multi-channel audio, and tactile signals, all aligned to a common timeline and spatial coordinate system.

The dataset is structured around two complementary capture systems: a tabletop scale for fine-grained hand-object interactions and a room scale for whole-home activities. The tabletop system uses multi-view cameras, optical motion capture, and tactile sensors to record tasks like pouring, wiping, cutting, and folding. The room-scale system covers entire living spaces with wide-baseline cameras and full-space motion capture, tracking participants as they move between kitchen, living room, and bedroom. Both systems share unified calibration and synchronization processes, ensuring that data from different devices can be accurately correlated.

A key innovation is the use of goal-level instructions rather than scripted actions. Participants are told the final task but are free to choose their own methods, sequences, and paths. This captures natural variations in human behavior, such as hesitation, plan changes, and error recovery, which are often treated as noise in standardized datasets but are crucial for robots operating in real homes. The dataset includes three task types: atomic interactions (about 3 minutes each), long-horizon activity chains (20-30 minutes, e.g., preparing a meal), and human-scene interactions (about 5 minutes).

ACE-Data-0 also addresses the challenge of multimodal synchronization. Different devices often have independent clocks and sampling rates, leading to misalignments. By implementing rigorous time and space calibration, the dataset ensures that video frames, motion data, object states, tactile readings, and audio are all locked to the same timeline and world coordinate system. This allows researchers to query, for example, the exact hand pose and contact pressure corresponding to a specific video frame.

Beyond the data itself, Daxiao has built a three-tier evaluation benchmark to assess model capabilities. The first tier tests contact and tactile prediction from visual input. The second evaluates recovery of human and hand motion under occlusion and complex poses. The third assesses full hand-object interaction understanding from first- and third-person video. The team evaluated over 30 representative methods and found significant gaps in existing models, particularly under occlusion, first-person egomotion, extreme viewpoints, and long-horizon tasks.

For startups and operators in embodied AI, ACE-Data-0 represents a valuable resource. It provides a more realistic training ground for imitation learning, world models, VLA models, and robot policy learning. The dataset's emphasis on long-horizon tasks and natural variability addresses a critical bottleneck: robots often fail when tasks require sustained memory and adaptation. By open-sourcing this data, Daxiao and NTU are enabling the broader community to train more robust models that can generalize to real-world home environments.

However, it's important to note that ACE-Data-0 is a starting point, not a final solution. The '0' in its name signifies the beginning of a data ecosystem. While the dataset is impressive in scale and fidelity, questions remain about its transferability to different robot embodiments and its coverage of edge cases. Startups should consider how to integrate such data with their specific hardware and task requirements.

In conclusion, ACE-Data-0 is a significant step toward making embodied AI more physically grounded. It moves beyond 'seeing actions' to 'understanding actions' by capturing the continuous physical process. For founders and operators, this dataset offers a new benchmark for evaluating and improving their models, potentially accelerating the path from research to real-world deployment.

Source: 雷峰网.

Related resources on this hub

Jump to projects, models, or datasets mentioned or closely related.

Discussion

Tell us what you think — comments make stories more useful for builders and founders.

Tell us what you think!

ACE-Data-0: A New Benchmark for Embodied AI Data in Real Homes

Have an account? Log in to use your display name and avatar.

Email is optional and never shown on the page.

More insights

IndustryDeploymentFeatured雷峰网

Video Generation Is Not Planning: Georgia Tech's Danfei Xu Proposes Factor Graphs to Bridge the Action Gap

Danfei Xu argues that high-fidelity video generation does not equal physical planning. He proposes compositional world models using factor graphs and temporal attention to enable test-time skill composition, improving success on out-of-distribution tasks. The approach reveals a video-action gap where action models lag behind visual generalization.

Read
IndustryDeployment雷峰网

Xingchi Power and PKU Healthcare: A Pragmatic Path for Embodied AI in Hospitals

Xingchi Power and PKU Healthcare are collaborating on a three-step pipeline—simulation training, scenario validation, and hospital deployment—to bring embodied AI into medical settings safely. The partnership leverages Xingchi's world model engine and data headset to create high-fidelity virtual hospitals, addressing the lack of training environments and expert data. This pragmatic approach prioritizes safety over speed, focusing on tasks like medication sorting and delivery.

Read
IndustryDeployment雷峰网

RSS 2026 Debate: Data Source, Control Modality Split Robotics Community

At RSS 2026, five experts debated whether world models and policy models should be unified or separate, whether control should rely on pixels or actions, and whether training data should come from the internet or embodied robots. The 'separate' camp won the first round, 'pixel' vs 'action' ended in a draw, and 'internet data' won the third round 61% to 39%. Key takeaways: modular design is preferred for debugging, video alone is insufficient for contact-rich tasks, and hybrid data strategies are

Read