Omni-modal humanoid motion
OmniHM-ZeroScaling Humanoid Motion Foundation Models with Large-Scale Human Videos
01 / Full film
See OmniHM-Zero in motion.
A five-minute real-robot film spanning live and Internet video, spoken instructions, music, and text prompts.
02 / Why OmniHM-Zero
Human behavior, translated for humanoids.
Motion capture is precise but difficult to scale. Internet video is diverse, yet reconstructed motion can be noisy, physically invalid, or impossible for a humanoid to execute. OmniHM-Zero bridges that gap.
Web-scale data engine
Recover, repair, retarget, and verify motion from raw video.
Omni-modal supervision
Text, RGB video, speech, and music share one motion space.
Closed-loop generation
Replan from the latest condition and robot state in real time.

03 / Dataset
Human behavior at humanoid scale.
The OmniHM dataset is built from raw Internet video. Its data engine filters and annotates clips, reconstructs human motion, repairs physical artifacts, retargets behavior, and verifies execution before training.
Data engine
From raw video to executable motion
A five-stage data engine starts from raw Internet video, curates coherent clips, filters motion and visual quality, aligns multimodal semantics, recovers and retargets human motion, and finally refines it through closed-loop execution. Each stage narrows the gap between visually plausible human behavior and motion a humanoid can reliably track.
- 01Clip generationCoherent clips from segmented and filtered web video
- 02Quality filteringHuman-motion and visual-quality screening
- 03Semantic annotationVideo, text, speech, and music alignment
- 04Motion recoveryHuman reconstruction, repair, and retargeting
- 05Physical refinementClosed-loop optimization and execution verification

Inside one OmniHM record
From a human video to executable humanoid data.
Each showcase example pairs the source clip with recovered human motion, a retargeted humanoid reference, and a tracking rollout. Together, the eight records illustrate diverse motions, scenes, and objects.
Dataset profile
Scale is visible in the distribution.
Four views summarize temporal scale, behavior taxonomy, scene coverage, and the relationship between motions and their contexts.
Duration at scale

A median duration of 16.0 seconds preserves meaningful temporal behavior while supporting web-scale processing.
A hierarchical motion vocabulary

Performance, dance, daily activity, locomotion, fitness, gesture, and sport form a broad long-tail taxonomy.
Behavior in context

Studio, home, outdoor, stage, fitness, and public settings keep motion grounded in diverse visual contexts.
Diversity beyond raw counts

The same motion families appear across multiple environments, reducing dependence on a single visual setting.
Public release
Explore OmniHM-5M.
The public OmniHM-5M release is available on Hugging Face. The examples above illustrate the source-to-execution data structure independently of any dataset viewer row.
Open dataset on Hugging Face04 / Method
One model, many ways to express intent.
A frozen omni-modal context encoder conditions a shared flow-based action expert. The latest robot state keeps generation grounded in what the hardware is actually doing.

Task-mixed flow matching
Every modality trains one reusable motion prior instead of an isolated policy.
Multimodal real-time chunking
Overlapping predictions enable continuous replanning without abrupt transitions.
State-conditioned generation
Each chunk uses the newest robot state before a high-frequency controller tracks it.
05 / Capability showcase
Five ways to express intent. One shared motion model.
Explore a curated set of real-robot demonstrations across live video, Internet video, speech, music, and text.
Video / Live
Real-robot demonstrationsFollow a performer as the motion unfolds.
OmniHM-Zero uses the current video stream and the robot’s latest state to continuously replan executable whole-body motion.
Video / In the wild
Real-robot demonstrationsBring human motion from the Internet onto the robot.
In-the-wild references provide locomotion, gesture, rhythm, and direction changes that the model translates into continuous humanoid behavior.
Audio / Spoken instruction
Real-robot demonstrationsSay what to do—and how to do it.
Spoken instructions specify actions, body side, repetition, order, and duration, from left-leg lunges to alternating waves.
Audio / Music
Real-robot demonstrationsLet sound shape the movement.
Music conditions the timing and character of coordinated footwork, balance shifts, turns, and upper-body motion.
Language / Text
Real-robot demonstrationsTurn a prompt into a whole-body behavior.
Natural-language prompts can describe an action, a path, or a style—from counted squats and counterclockwise walking to salsa and zombie motion.
06 / Scaling study
Scaling data, models, and execution together.
The study varies data from 1% to 100% of OmniHM and the action expert from 95.6M to 1.30B parameters under a shared execution protocol.
Quantitative tables in the current technical report are still being finalized; no unreleased accuracy claims are shown here.
07 / Citation
Build on OmniHM-Zero.
Paper and model links will be added when the technical report is released. Code and the public dataset are available now.
@article{anonymous2026omnihmzero,
title = {OmniHM-Zero: Scaling Humanoid Motion Foundation Models with
Large-Scale Human Videos},
author = {Anonymous Authors},
journal = {Technical Report},
year = {2026}
}