‹ 返回 2026-06-25

FLUX3D:具有扩散对齐的稀疏表示的高保真3D高斯分布生成技术

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

▲ 0 2026-06-25

Haorui Ji, Weizhe Liu, Hongdong Li, Hengkai Guo

摘要

稀疏体素表示方式已成为从图像生成3D高斯云图的技术基础,但当前方法在保留输入图像的高频视觉细节方面存在两处结构瓶颈。首先,这些方法采用用于语义抽象的判别性2D特征来构建稀疏体素 latent 向量,这会削弱重建过程中的线索,导致表示层面的瓶颈。其次,在生成阶段,标准的扩散Transformer缺乏有效的机制来将密集的2D图像元素与稀疏的3D体素 latent 向量进行匹配,从而造成跨模态匹配的瓶颈。为了解决这些问题,我们提出了FLUX3D这一可扩展的从图像生成3D高斯云图的技术框架,该框架在生成过程中同时提升了表示学习能力和跨模态匹配能力。我们首先重新审视了基于稀疏体素的3D表示学习的2D特征选择方法,提出了Diffusion-Aligned Structured Latents(DA-SLAT)模型,并将其与仅包含解码器的架构相结合,以提高3D高斯云图的重建质量。我们还设计了一个注重稀疏结构的扩散框架,该框架整合了Sparse-structure Multimodal Diffusion Transformer(SMDiT)和Modal-Aware Rotary Positional Embedding(MARoPE)技术,从而实现与几何结构无关的2D-3D匹配。大量的基准测试表明,FLUX3D在外观精度方面取得了显著提升,其生成的3D高斯云图质量明显优于所有当前最先进的方法。

English Abstract

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods struggle to preserve high-frequency visual details of input images due to two structural bottlenecks. First, they adopt discriminative 2D features optimized for semantic abstraction to construct sparse voxel latents, which suppress reconstructive cues and induce a representation bottleneck. Second, in the generation stage, standard diffusion transformers lack effective mechanisms to align dense 2D image tokens with sparse 3D voxel latents, resulting in a cross-modal correspondence bottleneck. To address these issues, we propose FLUX3D, a scalable image-to-3DGS framework that boosts both representation learning and cross-modal alignment during generation. We first revisit 2D feature selection for sparse-voxel-based 3D representation learning, propose Diffusion-Aligned Structured Latents (DA-SLAT) and couple it with a decoder-only architecture to improve 3DGS reconstruction fidelity. We also design a sparse-structure-aware diffusion framework, which integrates the Sparse-structure Multimodal Diffusion Transformer (SMDiT) and Modal-Aware Rotary Positional Embedding (MARoPE) to achieve geometry-agnostic 2D-3D alignment. Extensive benchmark experiments demonstrate that FLUX3D yields substantial improvements in appearance fidelity and significantly outperforms all state-of-the-art (SOTA) methods in generating high-quality 3DGS assets.