‹ 返回 2026-06-16

BadWorld:对世界模型的对抗性攻击

BadWorld: Adversarial Attacks on World Models

▲ 14 💬 1 2026-06-16

Linghui Shen, Mingyue Cui, Xingyi Yang

摘要

视觉世界模型(Visual World Models,VWMs)能够从单个上下文图像中生成具有交互性的、由用户行为驱动的动画效果。然而,这些模型在面对对抗性干扰时的鲁棒性仍然是一个悬而未决的问题。传统的对抗性攻击无法评估这种脆弱性,因为攻击者缺乏真实的未来视频数据,也无法预测用户的后续行为。我们提出了BadWorld这一无标签的对抗性框架,该框架专门针对自回归型VWMs设计,能够系统地克服上述两个限制。首先,为了绕过对未来数据的依赖,我们提出了一种自监督的速度攻击方法,直接破坏模型早期的无噪声动态过程。其次,为了确保攻击能够在不可预测的用户行为下依然有效,我们提出了一种基于轨迹的自适应双层优化算法,该算法能够主动寻找难以控制的动作序列,从而生成与用户控制无关的抗干扰效果。在具有连续和离散控制方式的代表性VWMs上进行了测试,结果表明BadWorld存在严重的结构脆弱性。看起来几乎相同的对抗性图像实际上会导致动画效果的严重退化,从而导致去噪不完全、结构崩溃以及控制不一致等问题。这些发现揭示了在安全关键系统中使用VWMs时所面临的关键风险,同时也提供了一种实用的隐私保护机制。

English Abstract

Visual world models (VWMs) synthesize interactive, action-conditioned rollouts from a single context image. However, it remains an open question how robust these models are to adversarial perturbations. Standard adversarial attacks fail to assess this vulnerability because attackers lack ground-truth future videos and cannot predict subsequent user controls. We introduce BadWorld, a label-free adversarial framework tailored for autoregressive VWMs that systematically overcomes both constraints. First, to bypass the need for future supervision, we propose a self-supervised velocity attack that directly disrupts the early denoising dynamics of the model. Second, to ensure the attack generalizes across unpredictable user actions, we formulate a trajectory-adaptive bi-level optimization that actively mines hard control sequences to forge control-agnostic perturbations. Evaluated on representative VWMs with continuous and discrete controls, BadWorld exposes severe structural fragility. Visually indistinguishable adversarial images reliably trigger catastrophic degradation in future rollouts, leading to incomplete denoising, structural collapse, and control inconsistency. These findings reveal critical risks for deploying VWMs in safety-critical systems while highlighting a practical mechanism for privacy protection.