HuggingFace Daily Papers·· 1 天前AI 评分42
从单张图像构建罗马:面向室内外场景的单图 3D 场景生成
Building Rome from a Single Image
AI 导读
研究者提出一种将物体中心生成器 Trellis 2 改造为室内外通用单图 3D 场景生成的方法:按距相机远近自适应切分场景,近处用小块保留细节、远处建筑用大块覆盖,并让模型显式感知自由空间、已观测表面与未观测区域。
正文
Abstract:Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b) making the generator capture explicit 2D-3D correspondence by lifting image features and making the model aware of the free space, observed surface, and unobserved region; (c) synthesizing around 4,000 outdoor scenes to broaden the training data, as existing scene datasets are largely indoor. Experiments on Tanks and Temples, ScanNet++, and in-the-wild images show that our method outperforms all baselines in geometric accuracy and perceptual quality across both indoor and outdoor scenes.
| Comments: | Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08790 [cs.CV] |
| (or arXiv:2610.08790v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08790 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiraphon Yenphraphai [view email]
[v1]
Tue, 6 Oct 2026 17:59:50 UTC (46,319 KB)
来源:HuggingFace Daily Papers · arxiv.org