‹ 返回 2026-06-21

LooseControlVideo:利用空间分割技术进行导演级视频控制

LooseControlVideo: Directorial Video Control using Spatial Blocking

▲ 1 💬 1 2026-06-21

Shariq Farooq Bhat, Niloy J. Mitra, Kalyan Sunkavalli

摘要

在文本到视频生成过程中实现精确的3D空间调度仍然是一个重大挑战,尤其是在多物体场景中,语义布局与时间动态往往相互关联。虽然现有的深度条件模型能够保持良好的结构精度,但它们需要密集且精确到帧级的指导信息,而这对于包含可变形物体的动态场景来说非常困难。我们提出了LooseControlVideo框架,该框架利用稀疏、有方向的3D框作为“遮挡”代理,从而实现直观且富有表现力的控制效果。这样,用户就可以在利用视频生成模型来生成真实的遮挡、动态和互动效果的同时,设定出合理的布局与运动轨迹。我们通过对Wan 2.2主干网络进行微调,使其能够在带有DNOCS标注的视频数据集上进行训练,DNOCS是一种用于描述3D尺寸、方向和深度排序遮挡的新颖编码方式。此外,我们的方法还允许进行局部优化,比如调整跳跃轨迹或添加互动元素,而不会对整体场景情境造成太大影响。在nuScenes、HO-3D和BEHAVE等基准测试中的大量实验表明,LooseControlVideo的性能显著优于现有的2D框法和基于流法的基线方法。我们的研究结果显示,轨迹误差降低了1.2倍至3倍;刚性运动一致性提高了2倍;遮挡精度则提升了1.5倍至2倍。这表明,有方向的3D基本元素能够为复杂的多主体视频生成提供良好的几何约束条件。

English Abstract

Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-conditioned models achieve good structural fidelity, they necessitate dense, frame-accurate guidance that is labor-intensive to author for dynamic events involving deformable objects. We present LooseControlVideo, a framework that enables intuitive and expressive control by using sparse, oriented 3D boxes as a "blocking" proxy. This allows users to author high-level layout and trajectory while leveraging a video generative model to generate realistic occlusions, dynamics and interactions. We achieve this by fine-tuning a Wan 2.2 backbone on a video dataset annotated with DNOCS, a novel encoding for 3D size, orientation and depth-ordered occlusions. Furthermore, our method allows for localized refinement, such as adjusting a jump trajectory or adding an interaction, with minimal disruption to the global scene context. Extensive evaluations on the nuScenes, HO-3D, and BEHAVE benchmarks demonstrate that LooseControlVideo significantly outperforms existing 2D-box and flow-based baselines. Our findings indicate a 1.2x to 3x improvement in Trajectory Error; 2x improvement in Rigid Motion Consistency; and a 1.5x to 2x increase in Occlusion Accuracy over current state-of-the-art layout-conditioned models, demonstrating that oriented 3D primitives provide good geometric prior for complex, multi-agent video authoring.