‹ 返回 2026-06-18

显示信号,隐藏噪声:用于像素空间扩散的谱激励方法

Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion

▲ 15 💬 1 2026-06-18

Weichen Fan, Haiwen Diao, Penghao Wu, Ziwei Liu

摘要

像素空间扩散模型是在全带宽的噪声图像上进行训练的,不过,可供去噪器使用的有效信号其实具有很强的频率依赖性。在修正流扩散算法和自然图像幂律光谱的情况下,每个频带的信噪比阈值k*(t) = (1-t)^{-2/α},能够将低频区域与以噪声为主的高频区域区分开来。我们发现,这种从粗到细的结构不仅仅是一种描述性结构,它还会引发一种容量分配问题。标准的像素空间去噪器必须自行确定移动的频率带宽边界,并且可以将计算资源集中在那些最优预测可以简化为确定性基线的情况上,而不是用于数据分布建模。为了明确这一边界,我们引入了“光谱强制”技术——这是一种无需参数的、与时间相关的二维DCT低通算子,它被应用于噪声输入数据之后,再经过块嵌入处理。该算子的截止频率会随着扩散时间的推移而单调增加,最终在数据终点处变为恒等算子。通过可控的合成实验,我们确定了该算子适用的场景:当高频内容主要是噪声而非重要信号时,使用该算子能够改善效果。在ImageNet-256数据集上,使用JiT-700M/32模型时,“光谱强制”技术能够持续改善FID和Inception Score指标,显示出在训练过程中持续的提升效果;而在更精细的块级分割情况下,该技术的竞争力仍然很强。此外,我们将该算子整合到了SenseNova-U1模型中,该模型是一个统一的文本到图像生成模型。结果表明,该算子能够在不同训练阶段持续改善DPG-Bench和GenEval指标,表明其效果可以超越基于类别生成的限制。这些结果表明,通过揭示信号并隐藏噪声,我们可以实现高效利用容量的像素空间扩散技术。

English Abstract

Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffusion and natural-image power-law spectra, the per-band data-to-noise contour k^{*}(t) = (1-t)^{-2/α} separates a signal-bearing low-frequency region from a noise-dominated high-frequency region at each time t. We show that this implicit coarse-to-fine structure is not merely descriptive: it induces a capacity-allocation problem. A standard pixel-space denoiser must discover the moving bandwidth boundary internally and can spend computation on frequency-time regions where the optimal prediction collapses to deterministic baselines rather than data-distribution modeling. To make this boundary explicit, we introduce Spectral Forcing, a parameter-free, time-conditional 2D-DCT low-pass operator applied to the noisy input before the patch embedder. Its cutoff expands monotonically with the diffusion time and becomes the identity at the data endpoint. Through controlled synthetic experiments, we identify the regime in which the operator is beneficial: coarse patch tokenization and data whose high-frequency content is predominantly noise rather than essential signal. On ImageNet-256 with JiT-700M/32, Spectral Forcing consistently improves both FID and Inception Score across different training epochs, demonstrating robust gains throughout training; at finer tokenization, the spectral forcing is still competitive. We further insert the unchanged operator into SenseNova-U1, a unified text-to-image model, where it improves DPG-Bench and GenEval, showing that the input-side spectral prior transfers beyond class-conditional generation. These results suggest a route to capacity-efficient pixel-space diffusion by showing the signal and hiding the noise.