arXiv:cs.LG· Madelyn Mathai, Timothy B. Higgins, Kevin M. Grise, Chirag Agarwal, Antonios Mamalakis·· 4 小时前AI 评分47
GraphCast 中大气河现象的机制可解释性研究
Mechanistic Interpretability of Atmospheric Rivers in GraphCast
AI 导读
研究者用稀疏自编码器(SAE)在 GraphCast 中提取所学概念,发现尽管积分水汽输送(IVT)既非输入也非目标,GraphCast 仍将其作为稳定的内部变量来计算大气河强度。标准 SAE 与 Matryoshka SAE 均能识别该变量,后者还按重要性排序概念并揭示其关联;大气河概念在模型各深度层持续存在,干预实验证实了其因果作用。
正文
Abstract:While AI weather models now rival operational forecasts, how they represent the atmosphere internally remains an open question: feature attribution reveals which input patterns matter, not what the model computes or how it combines information internally. We train sparse autoencoders (SAEs) on GraphCast to uncover its learned concepts, using atmospheric rivers as our phenomenon of focus. Both standard and Matryoshka SAEs show GraphCast computes atmospheric river intensity, measured by integrated vapor transport (IVT), as a stable internal variable, despite IVT being neither an input nor a target. In contrast to the unstructured concept retrieval of the standard SAE, the Matryoshka SAE orders concepts by importance and exposes their relations. Atmospheric river concepts persist across depth and direct interventions confirm causality. This method offers a way to find internal variables and determine which of them the model actually relies on, which is a prerequisite for asking whether those variables remain meaningful as the phenomenon changes under a warming climate.
| Comments: | Accepted to TCCML NeurIPS workshop 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07583 [cs.LG] |
| (or arXiv:2610.07583v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07583 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Madelyn Mathai [view email]
[v1]
Tue, 6 Oct 2026 01:22:10 UTC (2,519 KB)
来源:arXiv:cs.LG · arxiv.org