‹ 返回 2026-06-25

FlowR2A:用于多模态驾驶规划的学习奖励-动作分布方法

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

▲ 1 💬 1 2026-06-25

Xirui Li, Zhe Liu, Xiaoqing Ye, Wenhua Han, Yifeng Pan, Junyu Han, Hengshuang Zhao

摘要

多模态驾驶规划面临着两种方法理念之间的长期矛盾:基于评分的方法能够利用丰富的奖励信息,但受到固定动作词汇的限制;而基于锚点的方法则能够动态生成建议方案,但受到有限监督信息的约束,只能使用单一的真实轨迹作为参考。在本研究中,我们提出了FlowR2A这一方法,它通过将基于模拟的奖励从“区分性目标”转化为“生成性条件”,从而解决了这种矛盾。通过利用流匹配解码器从丰富的轨迹-奖励对中学习奖励相关的动作分布,FlowR2A将基于评分方法的丰富监督机制与基于锚点的方法中的建议生成功能整合到了一个统一的生成模型中。这样,模型就能在安全性、进展性、舒适性以及规则遵守等方面,充分理解某个动作与其结果之间的关系。为了平衡严格的安全性约束与灵活的进展目标,我们引入了细粒度的每个时间步的奖励条件化机制以及奖励噪声增强技术。这种生成式方法自然支持通过奖励引导和锚点采样来实现可控的测试时采样过程,从而生成高质量的建议方案。在NAVSIM v1和v2基准测试中,FlowR2A取得了最先进的性能,其多模态建议方案的质量也远远高于以往的方法。

English Abstract

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.