arXiv:cs.LG(机器学习,全量分类)· Ruizhe Huang, Qidong Yang, Jonathan Giezendanner, Sherrie Wang·· 14 小时前AI 评分41
生成式模型在真实气象站观测数据同化上的基准测试
Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations
AI 导读
一项研究首次在真实气象站观测上对生成式天气数据同化进行受控基准测试,使用美国本土 11,849 个 NOAA MADIS 站点和四种气象变量,固定数据集、观测算子与深度学习架构。
正文
Abstract:Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00728 [cs.LG] |
| (or arXiv:2610.00728v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00728 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruizhe Huang [view email]
[v1]
Wed, 30 Sep 2026 21:18:39 UTC (14,964 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org