‹ 返回 2026-06-23

SpatialAvatar-0:高质量4D头部虚拟形象,采用多阶段重建技术

SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction

▲ 2 💬 1 2026-06-23

Yiran Wang, Zeyu Zhang, Yuanming Li, Ziming Wang, Yang Zhao

摘要

高质量的4D头部虚拟形象,源自一个或多个源图像,是远程交互、AR/VR以及数字人类交互技术中的核心要素。3D高斯分形技术已成为主要的表示方式,其中两种互补的模型机制——可泛化的前馈预测器与针对每个个体的优化器——正在同时发展完善。不过,现有的前馈预测器是在单一数据集上训练的,且存在固定的源数量限制,从而带有相应的领域偏见。针对每个个体的优化器则需要30万到60万次迭代,并且依赖自适应加密算法,这种算法会破坏原有的高斯分布结构,导致这两种模型机制无法实现端到端的统一表示。为了整合这两种模型机制,我们提出了SpatialAvatar-0:该模型基于共享的FLAME网格结构的高斯表示方式,包含无参数参数的K源均值池化机制,以及单目时间信息与多视图空间信息的双阶段处理流程;此外,还引入了10万次迭代的个体优化循环,该循环能够保持FLAME结构和高斯分布的数量不变,并将加密过程替换为三组分抗尖峰正则化方法。在VFHQ/HDTF跨领域零样本测试中,尽管从未在测试领域进行训练,但SpatialAvatar-0的PSNR成绩仍比领域的领先模型GAGAvatar高出1.5 dB;在SplattingAvatar单目基准测试中,其各项指标均优于300万次迭代的GeoAvatar,其个体优化所需的迭代次数仅为常见最优模型的1/60。网站地址:https://spatialwalk.github.io/SpatialAvatar-0

English Abstract

High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human interaction. 3D Gaussian Splatting (3DGS) has emerged as the dominant representation, with two complementary regimes (generalizable feed-forward predictors and per-subject refiners) maturing in parallel. However, existing feed-forward predictors are trained on a single dataset family with a hard-coded source count, inheriting the corresponding domain bias. Per-subject refiners require 300K--600K iterations and rely on adaptive densification that destroys upstream Gaussian layouts, preventing the two regimes from sharing a representation end-to-end. To bridge both regimes we propose SpatialAvatar-0 on a shared FLAME-mesh-bound Gaussian representation: a feed-forward generator with a parameter-free K-source mean-pool and a monocular-temporal to multi-view-spatial two-phase schedule that anchors against identity-prior collapse onto the smaller multi-view set. We further introduce a 10K-iter layout-preserving per-subject refinement loop that freezes the FLAME-binding and Gaussian count and replaces densification with a three-component anti-spike regularization. On VFHQ/HDTF cross-domain zero-shot we surpass the in-domain leader GAGAvatar by +1.5 dB PSNR despite never training on either test domain, and on the SplattingAvatar monocular benchmark we lead every reported metric, surpassing the 300K-iter GeoAvatar by +1.3 dB PSNR at up to 60x shorter per-subject schedule than common SOTA baselines. Website: https://spatialwalk.github.io/SpatialAvatar-0.