‹ 返回 2026-06-18

RefGC-SR^2:基于参考信息的生成内容超分辨率处理与优化技术

RefGC-SR^2: Reference-guided Generated Content Super-Resolution and Refinement

▲ 1 💬 1 2026-06-18

Jeahun Sung, Dahyeon Kye, Soo Ye Kim, Jihyong Oh

摘要

基于参考信息的生成技术(例如对象合成、个性化定制等)已经取得了显著进展。不过,当前的生成流程仍然存在一个根本性的缺陷:用户提供的以对象为中心的高分辨率参考图像在输入模型之前会被降采样到较低的分辨率,这样一来,精细的细节信息在输出结果生成之前就被丢失了。此外,生成过程还会引入额外的瑕疵,比如身份扭曲等问题。现有的基于参考信息的生成内容优化方法能够修复部分瑕疵,但它们仍然只能处理低分辨率的图像;而基于参考信息的超分辨率方法则能够恢复图像的分辨率,但会忽略生成过程中产生的瑕疵。为了弥补这两种方法的不足,我们提出了一种新的任务:基于参考信息的生成内容超分辨率优化(RefGC-SR^2)。在这种方法中,原始的高分辨率参考图像会在后处理阶段被重新使用,从而恢复丢失的细节信息,同时消除生成过程中产生的瑕疵。我们已经为这一任务构建了一个真实世界的数据生成流程,并训练了一个能够根据特定条件生成低质量样本的生成器,这些样本是现有预训练模型无法提供的。我们还提出了一个频率感知的扩散变换模型,该模型能够选择性地从高分辨率参考图像中提取细节信息,同时消除生成过程中产生的瑕疵。大量的实验表明,我们的RefGC-SR^2模型能够成功实现以下目标:(i) 根据参考图像精确地还原对象的身份特征;(ii) 恢复高分辨率的细节信息,使得最终生成的结果质量远高于现有的RefGCR和RefSR基线方法。

English Abstract

Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation: the object-centric high-resolution reference image (HRRI) provided by users is downsampled to a fixed low-resolution (LR) before being fed into the model, so the fine-grained details are discarded before the output is even produced. In addition, the generation step then introduces its own artifacts (e.g., identity distortion) on top of this loss. Existing reference-guided generated content refinement (RefGCR) methods can correct some of these artifacts but still operate in the LR domain; reference-guided super-resolution (RefSR) methods recover resolution but assume natural-image degradations and ignore the artifact distribution of generative pipelines. To address both gaps in a single formulation, we introduce a new task: reference-guided generated content super-resolution-refinement (RefGC-SR^2), where the original HRRI is reused at the post-processing stage to recover lost details, refine generative artifacts, and upscale the output simultaneously. We construct the first real-world triplet data generation pipeline for this RefGC-SR^2 task, training a diptych-conditioned generator to synthesize paired low-quality anchors that public pretrained models cannot provide. We further present a frequency-aware diffusion transformer model for RefGC-SR^2 that selectively injects fine details from the HRRI while removing generative artifacts. Extensive experiments demonstrate that our RefGC-SR^2 model successfully (i) refines the object identity faithfully with respect to the reference, and (ii) recovers high-resolution details, so that the final result is significantly higher quality and practically more usable compared to existing RefGCR and RefSR baselines.