EEmbodied AI Hub
IndustryFundingEditor’s pick

Delta Intelligence Raises $70M in Six Months: A Data-First Bet on Full-Body Humanoid Models

Delta Intelligence, a Beijing-based humanoid foundation model startup founded in January 2026, closed a ~$70M (¥500M) angel++ round, its sixth in six months. The company is building a native 3D world engine and a head-mounted data capture device (Delta D1) to generate full-body interaction data at scale. This signals a strategic shift from simulation-dependent approaches to real-world data-driven full-body intelligence, with implications for founders building in the embodied AI stack.

humanoid foundation modelfull-body intelligencedata capture deviceDelta Intelligence3D world enginefeed:leiphoneleiphonechina

Source: 雷峰网 · July 31, 2026

Share this article so more people can see it

Delta Intelligence, a humanoid foundation model startup based in Beijing, has closed a nearly ¥500 million (~$70M) angel++ round, bringing its total funding to six rounds within six months of founding. The company was established in January 2026.

The round was backed by a mix of listed company strategic investors and top-tier financial institutions. Previous investors include humanoid OEMs such as Zhiyuan Robot, Leju Robot, and Xinghaitu, as well as financial investors like Yuanhe Holdings, Fosun Rui Zheng, Huaying Capital, and Lenovo Capital.

Delta Intelligence is building what it calls a "native general-purpose humanoid foundation model." The key differentiator is its approach to spatial representation. Most current embodied AI models rely on 2D visual representations—monocular images that lack true depth. This leads to cumulative errors in real-world tasks like estimating door handle distance or body clearance. Delta's solution is a native 3D world engine that directly processes point clouds and Gaussian splatting, enabling inherent 3D understanding and reasoning.

Under the hood, the company employs a three-layer architecture: "brain + cerebellum + force-position hybrid." The brain handles environment perception, long-horizon task planning, and manipulation decisions—trained almost exclusively on real robot interaction data, not simulation, because simulators struggle with soft-body deformation and multi-contact forces. The cerebellum manages full-body balance and low-level control, trained via massive reinforcement learning in simulation to convert sparse brain commands into high-frequency motor signals. The force-position hybrid layer then outputs compliant control.

The bottleneck for full-body intelligence is data. Existing data collection solutions capture only partial body motions (e.g., hands), which is insufficient for tasks requiring whole-body coordination like climbing ladders or pushing heavy doors. Delta's answer is the Delta D1, a head-mounted full-body panoramic data capture device launched alongside this funding round.

The D1 is worn by a human operator performing normal tasks. It simultaneously records first-person panoramic vision and full-body joint motion trajectories, with global positioning accuracy within 2 cm. No robot body, motion capture studio, or site modification is required. The collected data can be directly used for training. For fine bimanual manipulation, a handheld accessory (D1-Gripper) captures end-effector details.

Crucially, Delta is opening this capture system to the industry. Data collected can be retargeted across different humanoid platforms, including Unitree G1/H2, Zhiyuan Lingxi X2/A3, Leju Kuafu 4/5, and Xinghaitu Kengo. The company has also partnered with the China Academy of Information and Communications Technology to release an Industrial Dataset 2.0, collected by workers wearing D1 on real production lines without interrupting operations.

The foundation model is already deployed in real-world scenarios: power grid full-body inspection (moving and operating equipment in substations), automotive parts handling, and SMT bin sorting. Delta's standardized cross-platform adaptation pipeline—motion capture, retargeting, cerebellar RL simulation, and brain fine-tuning—enables rapid deployment across different hardware without starting from scratch.

Founder and CEO Ma Xiaojian holds a B.S. from Tsinghua University and a Ph.D. from UCLA, with prior stints at Google Robotics and NVIDIA Research. Co-founder Liu Hangxin earned degrees from Virginia Tech and UCLA, and is now an assistant professor at Peking University. Co-founder and Chief Scientist Huang Siyuan, a Tsinghua and UCLA alum, leads the Embodied Robotics Center at Beijing Institute for General Artificial Intelligence (BIGAI) and previously worked at DeepMind and Meta.

For founders and operators: Delta's rapid fundraising and product-market fit signal that the embodied AI race is shifting from locomotion demos to physical task execution. The critical bottleneck is no longer hardware but the data pipeline for full-body intelligence. Startups that can provide scalable data collection and cross-platform model adaptation may find themselves at the center of the ecosystem.

Source: 雷峰网.

Related resources on this hub

Jump to projects, models, or datasets mentioned or closely related.

Discussion

Tell us what you think — comments make stories more useful for builders and founders.

Tell us what you think!

Delta Intelligence Raises $70M in Six Months: A Data-First Bet on Full-Body Humanoid Models

Have an account? Log in to use your display name and avatar.

Email is optional and never shown on the page.

More insights

IndustryFunding雷峰网

Embodied AI Funding Update

A funding development in embodied AI and robotics worth watching. Source tip was non-English; full bilingual rewrite preferred when LLM is available.

Read
IndustryDeploymentFeatured雷峰网

Video Generation Is Not Planning: Georgia Tech's Danfei Xu Proposes Factor Graphs to Bridge the Action Gap

Danfei Xu argues that high-fidelity video generation does not equal physical planning. He proposes compositional world models using factor graphs and temporal attention to enable test-time skill composition, improving success on out-of-distribution tasks. The approach reveals a video-action gap where action models lag behind visual generalization.

Read