Ropedia’s $22 million pre-Series A round underscores how data infrastructure has become a distinct bottleneck in scaling physical AI. The company, founded in 2025, now has $30 million total and is using the fresh capital to industrialize collection of egocentric human demonstrations rather than waiting for robot fleets to generate their own experience.
Its HOMIE device is a lightweight, head-mounted unit with four cameras delivering 360-degree coverage, plus audio sensors that localize sound sources and an onboard motion unit that tracks head movement. A swappable battery module supports extended capture sessions in real homes and workplaces.
Unlike teleoperation setups that tie data collection to specific robot hardware, HOMIE records natural human behavior from a first-person perspective. This approach lets Ropedia generate its own structured datasets and align them directly to model training pipelines, bypassing the fragmented ownership typical of third-party labeling services.
CEO Zhaoxi Chen said the round will fund team expansion across hardware, software, data infrastructure, and model development. A second priority is building a stronger North American presence, given that most existing clients are already based in the United States.
The company also intends to strengthen its annotation tooling, quality analytics, and compliance layers so that data volumes can move from thousands to millions of hours while meeting industry-grade standards. Parallel research work will target foundation models and world models trained on the resulting egocentric streams.
Ropedia’s full-stack model differentiates it from both conventional data-labeling vendors and hardware-centric teleoperation providers. By controlling capture hardware, processing, and downstream fine-tuning, the firm claims it can deliver consistent, embodiment-agnostic datasets at the scale required for generalist robot policies.
For builders evaluating data strategies, the round highlights two concrete bets: custom sensing hardware can measurably improve data quality over off-the-shelf alternatives, and owning the end-to-end pipeline reduces friction when moving from pilot datasets to production-scale training runs.