Hi3D 3.0(Twinkle3D):高分辨率对象特定 3D 资产生成
HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution
Hi3D 3.0 图像到 3D 生成系统发布,其几何模型 Twinkle3D 可生成 2048³ 分辨率的水密三角网格。它通过重新设计的 DiT 架构将扩散生成扩展至最多 300K 几何 token 序列,单步训练时间从约十分钟降至十秒。在铭文文字还原上,Hi3D 3.0 取得 82.1% 召回率和 98.2% 精度,而最强竞品召回率仅 21.7%,并在所有报告指标上超越四款商业系统。
Authors:Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, Jie Li, Ruiyang Liu, Yibo Luo, Tengjiao Sun, Pei Tang, Shiwen Wang, Jiaqi Wu, Kang Wu, Kaiqiao Yang, Zherui Yang, Hu Zhang, Xuezhi Zhao, Xinhe Zheng, Yukun Li, Heliang Zheng, Rongfei Jia
Abstract:Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.
| Comments: | Hi3D 3.0 (Twinkle3D) Technical Report |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11685 [cs.CV] |
| (or arXiv:2610.11685v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11685 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ziying Li [view email]
[v1]
Thu, 8 Oct 2026 10:57:06 UTC (24,015 KB)
来源:arXiv:cs.AI · arxiv.org