跳到正文
arXiv:cs.LG· Ali Benlalah, Sepehr Johari, Patricia Vitoria, Armin Kappeler, Artem Sevastopolsky, Alexander Jung, Gabriele Fanelli, Kevin Mader, Manuel Breitenstein, Claudia Pl\"uss, Jan R\"uegg, Simon Biland, Thomas Etterlin, Dmitry Kostiaev, Mathias Deschler, Brian Amberg, Sebastian Martin·· 4 小时前

GHARP:基于大规模重建先验的实时高斯头部动画

GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior

AI 导读

GHARP 通过将问题解耦为离线身份建模与运行时表情残差预测两阶段,可基于少量输入图像和驱动表情信号实时驱动 3D 人头动画,动画网络轻量并针对移动设备优化。该方法在 Ava-256 基准上达到 SOTA 质量,在 A100 GPU 上运行速度最高提升 13 倍,高斯数量减少 8 倍。针对表情编码无法描述身体姿态与衣物位置导致的模糊与闪烁,研究提出身体对齐网络以消除训练信号歧义。

正文

Authors:Ali Benlalah, Sepehr Johari, Patricia Vitoria, Armin Kappeler, Artem Sevastopolsky, Alexander Jung, Gabriele Fanelli, Kevin Mader, Manuel Breitenstein, Claudia Plüss, Jan Rüegg, Simon Biland, Thomas Etterlin, Dmitry Kostiaev, Mathias Deschler, Brian Amberg, Sebastian Martin

View PDF HTML (experimental)

Abstract:We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
Comments: Accepted to ACCV 2026. 35 pages (14 main + references + 15 pages supplementary), 14 figures, 17 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.10945 [cs.CV]
  (or arXiv:2610.10945v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.10945

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Armin Kappeler [view email]
[v1] Wed, 7 Oct 2026 21:50:43 UTC (20,103 KB)

来源:arXiv:cs.LG · arxiv.org