跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Fernando Dupin da Cunha Mello (Stricto Sensu Department, SENAI CIMATEC University, Salvador, Bahia, Brazil), Prashant Kumar (Global Centre for Clean Air Research), Erick G. Sperandio Nascimento (Stricto Sensu Department, SENAI CIMATEC University, Salvador, Bahia, Brazil)·· 5 小时前AI 评分33

仅用天气数据预测巴西大豆产量:一种输入精简的深度学习框架

An Input-Frugal Deep Learning Framework for Weather-Driven National Crop-Yield Forecasting: A Case Study of Brazilian Soybean

AI 导读

研究者提出一种架构无关的深度学习框架,仅以常规天气为唯一时变输入,外加作物年份和农业环境标签两个轻量静态输入,预测巴西大豆产量。基于 2001/02 至 2020/21 共 20 季数据、留一年交叉验证,Transformer 取得最佳全国精度(RMSE 149 kg ha^-1,rRMSE 5.3%,R^2 = 0.784),较五年移动平均基线误差降低 47.6%。

正文

Authors:Fernando Dupin da Cunha Mello (1), Prashant Kumar (2, 3, 4), Erick G. Sperandio Nascimento (1, 2, 3, 4) ((1) Stricto Sensu Department, SENAI CIMATEC University, Salvador, Bahia, Brazil, (2) Global Centre for Clean Air Research (GCARE), School of Engineering, Civil and Environmental Engineering, Faculty of Engineering and Physical Sciences, University of Surrey, Guildford GU2 7XH, United Kingdom, (3) Institute for Sustainability, University of Surrey, Guildford, GU2 7XH, United Kingdom, (4) Surrey Institute for People-Centred AI, Faculty of Engineering and Physical Sciences, University of Surrey, Guildford GU2 7XH, UK)

View PDF

Abstract:Reliable, timely crop-yield forecasts are essential for market stability and risk management, yet many approaches rely on costly or hard-to-scale inputs. We present a frugal, transferable, and architecture-agnostic deep learning framework that uses routine weather as the only time-varying input plus two lightweight static context inputs (crop year and an agro-environmental label) to capture long-run change and regional heterogeneity, while supporting multiple sequence encoders under identical data requirements. Using a 20-season Brazilian soybean case study (2001/02-2020/21) with leave-one-year-out cross-validation, we benchmark MLP, CNN, LSTM, CNN-LSTM, a Transformer encoder and the Mamba state-space model against linear ridge regression and a five-year moving-average "farmer" baseline. All deep learning variants outperform ridge, and all sequential encoders surpass the non-sequential MLP. The Transformer achieves the best national accuracy (RMSE 149 kg ha^-1; rRMSE 5.3%; R^2 = 0.784), reducing error by 47.6% relative to the farmer baseline. In-season forecasts improve monotonically from early- to late-season issuance, reaching approximately 50% lower error than the baseline at the latest forecast point. Ablations indicate that the agro-environmental label and spatial instance expansion (multiple grid-node weather sequences per municipality-year) contribute positively without increasing input complexity. SHAP diagnostics suggest crop year explains most of the long-run trajectory, whereas within-season weather and agro-environmental context primarily drive interannual deviations, with moisture/cloud and thermal-demand variables dominating. Overall, the framework is straightforward to deploy across other crops and geographic regions and is naturally compatible with operational weather forecasts for routine monitoring.
Subjects: Machine Learning (cs.LG); Applications (stat.AP)
Cite as: arXiv:2609.38447 [cs.LG]
  (or arXiv:2609.38447v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38447

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fernando Mello [view email]
[v1] Tue, 29 Sep 2026 19:40:09 UTC (2,339 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org