跳到正文
arXiv:cs.LG· Venkatesh Kolluru, Rajat Shinde, Abdelhak Marouane, Caden Helbling, Deepak Shah, Othneil Drew, Srinivas Kolluru, Iksha Gurung, Manil Maskey, Rahul Ramachandran·· 4 小时前AI 评分44

Prithvi-EO-2.0 作物分类基础模型在三大洲物候与地理迁移下的可迁移性与运行可靠性

Transferability and operational reliability of a Prithvi crop classification foundation model under phenological and geographic shift across three continents

AI 导读

研究评估了 Prithvi-EO-2.0 在三大洲 12 国 37 个事件上的分布外表现,平均总体精度从美国 0.65 降至欧洲 0.40。观测窗口与当地作物物候错位时精度崩塌,而模型置信度仍保持高位,8 组配对事件中 7 组期望校准误差上升。两项无需重训的调整可恢复精度:将 13 类合并为 10 类使平均 OA 提升 8.4 个百分点,将窗口压缩至约 75 天可在 45-90 天区间保持精度。

正文

View PDF

Abstract:Fine-tuned geospatial foundation models (GeoFMs) pretrained on large satellite archives have been shown to improve crop classification accuracy and geographic transferability. However, their operational performance beyond the training distribution remains poorly characterized. We evaluated the out-of-distribution performance of a widely adopted GeoFM [Prithvi-EO-2.0] across 37 events in 12 countries on three continents and validated against regional reference products. Results indicated that the mean overall accuracy (OA) declined from 0.65 in the United States to 0.40 in Europe. Beyond accuracy metrics, we assessed five key aspects of model performance: whether model confidence indicates signal failure, sensitivity to observation windows, the effect of coarsening class schemes, and robustness to both band loss and cloud- and shadow-contamination. Accuracy collapsed when the observation window misaligned with local crop phenology, while deterministic confidence remained high. Expected calibration error increased for seven of eight paired events, and 12-51% of each affected scene was confidently mislabeled at near-zero precision. Monte Carlo dropout entropy registered the shift in all eight, indicating that much of the apparent cross-continent decline reflected phenological misalignment rather than spatial transfer. Two adjustments recovered accuracy without retraining. Consolidating 13 classes into 10, based on the model's dominant confusions, raised the mean OA by 8.4 percentage points. Compressing the window toward near-real-time use preserved accuracy across a 45- to 90-day plateau, peaking near 75 days, though arms tighter than 30 days fell about 0.11 below that plateau. Fine-tuned crop GeoFMs therefore transfer usefully only where observation windows match local growing seasons. We translate these findings into operational guidance for the reliable deployment of the released model.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08810 [cs.LG]
  (or arXiv:2610.08810v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08810

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Venkatesh Kolluru Dr [view email]
[v1] Sun, 20 Sep 2026 03:02:50 UTC (11,620 KB)

来源:arXiv:cs.LG · arxiv.org