跳到正文
arXiv:cs.AI· Makoto Kelp, Amirhossein Arzani, Patricia Castellanos, Paul Griffiths, Ivan Higuera-Mendieta, Manuel Perez-Carrasco, Viral Shah, Patrick Obin Sturm, James Weber·· 3 小时前

地球科学中 AI 模型的战略治理

Strategic Governance of AI Models in Earth Science

AI 导读

基于天气与气候数据预训练的 AI 基础模型正被微调用于远超天气预报的地球科学任务,其发展速度已超出科学界的评估能力。这些模型几乎只靠 benchmark 技能指标评判,而技能与物理可靠性是两种不同属性,在气候变化带来的非平稳条件下差异尤为关键。

正文

View PDF

Abstract:AI foundation models pretrained on weather and climate data are increasingly fine-tuned to Earth science tasks well beyond weather forecasting. Their development and adoption are outpacing the scientific community's ability to evaluate them. These models are judged almost entirely by benchmark skill metrics, which measure how closely a forecast reproduces a reference product but not whether a model represents the physical processes governing the system it predicts. Forecast skill and physical reliability are therefore distinct properties. The distinction is most consequential under the nonstationary conditions of a changing climate for which these models were never trained. We identify five priorities for the physical evaluation of AI models in Earth science from task-specific emulators to foundation models, spanning training data, fine-tuning, behavioral testing, mechanistic interpretability, and output validation. We recommend three activities for the coming decade: 1) open AI-ready evaluation datasets, 2) a shared reporting standard for physics-based evaluation, and 3) a dedicated research program on the safety of these models.
Subjects: Physics and Society (physics.soc-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Atmospheric and Oceanic Physics (physics.ao-ph)
Cite as: arXiv:2610.10560 [physics.soc-ph]
  (or arXiv:2610.10560v1 [physics.soc-ph] for this version)
  https://doi.org/10.48550/arXiv.2610.10560

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Makoto Kelp [view email]
[v1] Thu, 1 Oct 2026 00:15:04 UTC (3,956 KB)

来源:arXiv:cs.AI · arxiv.org