SimFoundry:用于政策学习与评估的模块化与自动化场景生成
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
摘要
在现实世界中对机器人策略进行训练与评估既成本高昂,又难以实现规模化应用。我们推出了SimFoundry这一模块化、自动化的系统,它能够从视频数据出发,快速构建出适用于模拟环境的场景。SimFoundry能够生成适合模拟的数字孪生模型,同时还能对物体、场景和任务进行编辑操作,从而自动生成各种不同的数字版本——这些数字版本保留了真实场景的特征。在SimFoundry上训练的策略可以轻松应用于那些需要多步骤操作、复杂物体交互以及双手协同操作的现实场景。这些数字版本有助于让策略能够适应新的现实环境条件。在7种操作任务和5种策略架构中,SimFoundry的模拟评估结果显示,其性能与真实环境中的表现高度相关,平均Pearson相关系数达到了0.911,而平均最大排名偏差则仅为0.018。当在现实环境中对在模拟环境中训练出的策略进行评估时,使用模拟中的物体、场景和任务作为训练数据的策略,其任务成功率分别提高了17%、21%和40%。更多详细信息请访问:https://research.nvidia.com/labs/gear/simfoundry/
English Abstract
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .