‹ 返回 2026-06-24

PoLAR:机器人策略学习中潜在动作的因子分解与模式识别

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

▲ 7 💬 1 2026-06-24

Youngjoon Jeong, Jihwan Yu, Minsoo Jo, Junha Chun, Taesup Kim

摘要

潜式动作预训练能够从成对观测数据中学习视觉变化的表示方式。但现有方法通常将每次转换都表示为一个无结构的特征向量,从而使得转换的幅度与方式被混合在一起。我们提出了“极坐标式潜式动作”模型(PoLAR),该模型为潜式动作施加了极坐标结构,使得半径能够表示转换的幅度,而方向则用于反映转换的方式。PoLAR利用两个观测点之间的时间差作为转换幅度的近似指标,从而促使那些在时间上相距较远的观测对产生的潜式动作占据更大的半径范围。我们在双曲空间中实现了这种结构,随着半径的增加,空间体积也会扩大,这非常适合处理各种不同形式的转换模式。在任务内预训练和大规模预训练场景中,PoLAR都能提升模拟和现实机器人实验中的策略性能,其效果优于其他潜式动作模型以及经过强预训练的VLA模型。这些结果表明,潜式动作空间的几何结构是将视觉预训练技术应用到机器人策略学习中的关键设计因素。

English Abstract

Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured representation that entangles transition extent and transition mode. We introduce Polar Latent Actions with Radial structure (PoLAR), which imposes a radial-direction structure on latent actions, encouraging radius to encode transition extent and direction to retain transition mode. PoLAR uses temporal offset between two observations as a weak proxy for transition extent, encouraging latent action from observation pairs separated by larger temporal gaps to occupy larger radii. We instantiate this structure in hyperbolic space, whose expanding volume with radius offers a natural fit for more diverse transition modes at larger extents. Across in-task and large-scale pretraining settings, PoLAR improves downstream policy performance in simulation and real-world robot experiments, outperforming latent action baselines and strong pretrained VLAs. These results suggest that the geometry of the latent action space is an important design choice for transferring visual pretraining to downstream robot policy learning.