arXiv:cs.LG· Chedly Ben Azizi, Claire Guilloteau, Gilles Roussel, Matthieu Puigt·· 5 小时前AI 评分29
面向不确定性感知光谱图像仿真的变分潜空间框架
A Variational Latent-Space Framework for Uncertainty-Aware Spectral Image Emulation
AI 导读
研究者提出一种基于 VAE 的变分潜空间框架,将光谱图像仿真建模为参数条件潜变量问题,同时给出逐像素不确定性估计。该框架在 PROSAIL 模拟的 211 波段高光谱植被立方体和 Sentinel-3 OLCI 21 波段真实海洋水色影像上评测,结果显示没有单一架构最优:像素级模型在受控高光谱仿真中表现最好,全卷积 VAE 在含缺失或污染像素的真实噪声观测上更稳健。
正文
Abstract:Synthetic spectral image generation is essential for remote sensing simulation and mission design, yet physically based radiative transfer models (RTMs) remain computationally expensive. Existing learning-based emulators reduce this cost, but are mostly deterministic parameter-to-spectrum regressors with limited spatial modeling and uncertainty information. We formulate spectral image emulation as a parameter-conditioned latent-variable problem and propose a variational autoencoder (VAE)-based framework combining nonlinear spectral-image representations, fast inference, and per-pixel uncertainty estimates. The framework is instantiated at spectrum and spatial--spectral levels through two-step VAE pretraining and latent mapping. We evaluate it on PROSAIL-simulated hyperspectral vegetation cubes (211 bands) and real Sentinel-3 OLCI multispectral ocean-colour imagery (21 bands) against classical regression emulators and a deep CNN baseline. Results show that no single architecture is optimal: pixel-to-pixel models perform best on controlled hyperspectral simulations, whereas a fully convolutional VAE is more robust on noisy real observations with missing or contaminated pixels. VAE-based emulators also achieve high throughput for large-scale generation. On Sentinel-3, the spatial--spectral VAE provides predictive intervals closer to empirical errors than pixel-wise neural and classical emulators, although absolute calibration remains incomplete. A look-up-table-based retrieval experiment further shows that reconstruction fidelity alone does not ensure reliable leaf area index and chlorophyll retrieval. Emulators should therefore be evaluated in representative remote-sensing end-use scenarios, not by reconstruction metrics alone.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV) |
| Cite as: | arXiv:2603.21911 [cs.CV] |
| (or arXiv:2603.21911v3 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2603.21911 arXiv-issued DOI via DataCite |
Submission history
From: Chedly Ben Azizi [view email]
[v1]
Mon, 23 Mar 2026 12:32:09 UTC (6,718 KB)
[v2]
Mon, 22 Jun 2026 08:42:22 UTC (9,547 KB)
[v3]
Tue, 6 Oct 2026 20:50:33 UTC (5,026 KB)
来源:arXiv:cs.LG · arxiv.org