A training-free framework that enables robots to acquire new manipulation skills in seconds from a single human video while retaining previously mastered capabilities.
MoreA training-free framework that enables robots to acquire new manipulation skills in seconds from a single human video while retaining previously mastered capabilities.
MoreA multimodal action tokenizer that doubles as a semantic interface between vision-language reasoning and continuous robot control.
MoreXRZero-G0 is an embodied data acquisition and strategy learning system designed through deep collaboration between hardware and software. It aims to overcome the fundamental bottleneck of acquiring high-quality, motion-aligned demonstration data in the field of dexterous robot operation.
MoreCarving World Action Modeling at the Event Joints
Morefully open-source, trained with gradient-bridged co-training, and deployable on real robots straight from pretraining.
MoreAn end-to-end embodied foundation model that leverages large-scale multimodal pretraining to achieve (1) embodiment-aware vision--language understanding, (2) strong language--action association, and (3) robust manipulation capability.
More