SAE干预不可靠:干预后被抑制行为的恢复情况
SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
摘要
稀疏自编码器能够将残差激活值分解为可理解的特征。近年来,许多潜在空间防御方法都依赖于这种分解方式,认为那些被识别出的“不安全”的SAE特征可以作为用于监控和干预的线索。在这种模式下,限制某个有害特征的值可以有效防止模型出现错误行为。然而,我们发现这种成功可能隐藏着一个可恢复的故障模式:限制某个特征可能会阻止某种行为发生,但并不能完全消除该行为本身。我们将这种漏洞视为一种“干预后恢复”问题,即一种受约束的残差空间优化问题。从干预后的残差状态出发,我们优化残差扰动,以恢复干预前的行为,同时保留目标SAE特征的干预后值。即使在干预过程持续进行的情况下,恢复仍然 가능。为了排除恢复只是抵消了干预的情况,我们使用单层干预时的编码器正交更新方法,以及跨层设置下的相关特征映射雅可比矩阵。在TPP实验中,即使经过去学习、IOI和拒绝转向操作,仍然有可能恢复行为,尽管特征层面的干预取得了成功。尤其是在安全至关重要的拒绝转向场景中,我们在有效样本上实现了95.8%的恢复率,同时保持被保护特征的相对漂移值仅为0.131,远远低于基于后缀的基线值。进一步分析表明,这种恢复现象主要发生在SAE重建残差中,也就是SAE无法解释的部分。这些结果揭示了特征层面控制与行为完整性之间的差距:SAE特征可以支持因果性干预,但控制它们并不能保证对底层行为的控制。
English Abstract
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that identified "unsafe" SAE features serve as actionable handles for monitoring and intervention. In this paradigm, clamping a specific harmful feature is expected to reliably prevent model misbehavior. However, we show that this success may hide a recoverable failure mode: the clamp may block one visible route to a behavior without eliminating the behavior itself. We formulate this vulnerability as post-intervention recovery, a constrained residual-space optimization problem. Starting from the post-intervention residual state, we optimize residual perturbations to recover the pre-intervention behavior while preserving the post-intervention values of the targeted SAE features. Even under a strong threat model where the intervention remains active throughout optimization and generation, recovery remains possible. To rule out that recovery simply undoes the intervention, we use encoder-orthogonal updates for single-layer interventions and the corresponding feature-map Jacobian in the cross-layer setting. Across TPP, unlearning, IOI, and refusal steering experiments, this stress test reveals recoverable behavior despite successful feature-level intervention. Especially in the safety-critical refusal-steering setting, we achieve a 95.8% recovery rate on valid samples while keeping defended-feature relative drift to 0.131, substantially below suffix-based baselines. A recovery-path attribution analysis further localizes this recovery to the SAE reconstruction residual, the component left unexplained by the SAE. These results expose a gap between feature-level control and behavioral completeness: SAE features can support causal intervention, but controlling them does not guarantee control over the underlying behavior.