MVTrack4Gen:作为几何监督的多视图点跟踪技术,用于4D视频生成
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
摘要
为了根据目标相机轨迹从单目参考视频中生成新的视图视频,需要实现几何一致性与运动一致性与参考视频的匹配。基于显式3D表示的现有方法受到现成重建模块精度的限制,这些模块往往无法准确描述单目视频中动态物体的几何特征。相比之下,仅依赖相机条件的模型虽然能够获得良好的视觉质量,但往往难以保持几何和运动的一致性。在本研究中,我们提出了MVTrack4Gen——一种基于多视图点跟踪的训练框架,它利用多视图点跟踪作为额外的几何和运动监督信号,用于训练仅依赖相机条件的新视图视频生成模型。我们的关键发现是:特定的注意力层能够编码出强烈的对应关系信息;查询特征会关注不同视图中几何上对应的位置的关键特征,而这些对应关系的偏差则会导致运动不一致性。基于这一发现,我们将这些特征引入辅助的多视图跟踪模块中,然后与点跟踪目标一起训练扩散模型。通过明确强化这些与运动相关的对应关系,MVTrack4Gen能够提升现有模型的性能,使其能够更好地跟随参考视图中的运动趋势,并维持不同视图之间的几何一致性。在多种基准测试中,我们的方法实现了最先进的几何一致性水平以及具有竞争力的相机精度。
English Abstract
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of off-the-shelf reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos. In contrast, camera-conditioning-only methods can achieve high visual quality but often struggle to preserve geometric and motion consistency. In this work, we introduce MVTrack4Gen (Multi-View point Tracking for Novel-View Generation), a motion-aware training framework that leverages multi-view point tracking as an additional geometric and motion supervision signal for camera-conditioning-only novel-view video diffusion models. Our key finding is that specific attention layers encode strong correspondence cues, where query features attend to key features at geometrically corresponding locations across views and over time, and the misalignment of these correspondences causes motion inconsistency. Based on this observation, we route these features into an auxiliary multi-view tracking head and jointly train the diffusion model with a point-tracking objective. By explicitly strengthening these motion-aware correspondences, MVTrack4Gen improves existing models to better follow the motion in the reference view and maintain cross-view geometric consistency. Across diverse benchmarks, our method achieves state-of-the-art geometric consistency and competitive camera accuracy.