‹ 返回 2026-06-29

Qwen-RobotManip技术报告:对齐机制为机器人操作基础模型的实现提供了更大的规模支持

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

▲ 1 💬 1 2026-06-29

Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen

摘要

语言与多模态领域的基础模型通过以统一的方式对异构数据进行整合以及大规模训练,从而实现了良好的泛化能力。在本报告中,我们探讨了这种扩展方法是否可以用于机器人操作领域,以实现真正的泛化效果。这确实具有挑战性:与文本数据不同,机器人操作数据本身具有异构性,收集起来也成本较高,且多样性有限,因此同时实现数据的统一与大规模训练十分困难。我们提出了Qwen-RobotManip这一基于Qwen-VL构建的可泛化视觉-语言-动作基础模型。Qwen-RobotManip在操作的表示、运动和行为维度上采用了统一的对齐框架,使得大规模多源数据的训练能够保持一致性而非产生冲突。这种对齐能力使得Qwen-RobotManip能够处理大量操作数据,而之前的训练方法则无法做到这一点。一个从人类到机器人的合成流程可以将以人类为中心的手部动作转化为15种平台上的机器人轨迹;而严格的数据筛选流程则能够协调各种异构数据集。通过使用开源数据集和人类视频数据,而不依赖专有数据收集方式,Qwen-RobotManip能够构建出约38,100小时的预训练数据,并展现出多种泛化能力,包括零样本指令遵循能力、对干扰的鲁棒性、对错误的快速恢复能力以及跨载体迁移能力。我们发现传统的基准测试无法准确反映预训练质量,因此我们使用了包括RoboCasa365、LIBERO-Plus、EBench、RoboTwin-Clean2Rand、RoboTwin-IF和RoboTwin-XE等不同的测试环境。在所有OOD环境中,Qwen-RobotManip都显著优于现有的最先进模型,包括π0.5模型。在RoboChallenge测试中,它的排名位居第一,相对提升了20%。该模型还在包括AgileX ALOHA、Franka、UR和ARX等真实机器人平台上得到了验证。

English Abstract

Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including π0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.