跳到正文
arXiv:cs.LG· Nathaniel T. Hindman, Fabricio Murai·· 4 小时前AI 评分28

LLM 辅助正则化能否提升低数据场景下的移民流预测精度?

Can LLM-assisted regularization increase forecast accuracy for migration flows in low data regimes?

AI 导读

研究将新闻文章中的推拉信号经 LLM 层级推理流水线分类后,以特征专属正则化惩罚融入加权 Lasso 预测框架,在 2021 年 11 月至 2022 年 11 月的墨西哥—美国、乌克兰—波兰、叙利亚—土耳其等移民走廊上评估。

正文

View PDF HTML (experimental)

Abstract:Predicting migration flows remains a significant challenge for traditional gravity-based forecasting models, which primarily rely on structured socio-economic indicators such as economic disparity, political stability, and geographic distance. This work investigates whether Large Language Models (LLMs) can improve migration forecasting by extracting contextual migration-related signals from news articles and incorporating them into a weighted Lasso forecasting framework through feature-specific regularization penalties. The proposed framework uses hierarchical LLM inference pipelines to classify migration-related push--pull signals from news data and evaluates the resulting forecasting performance across multiple migration corridors between November 2021 and November 2022, including Mexico--United States, Ukraine--Poland, and Syria--Turkey. Experimental results showed mixed performance across migration corridors and modeling strategies, and no single regularization approach consistently outperformed the others across all experiments. The best-performing Mexico configuration, which consisted of a gravity-based model augmented with the proposed push--pull ratios, achieved a Mean Absolute Percentage Error (MAPE) of 17.15%, while the strongest Syria configuration achieved a MAPE of 29.29% using Direct LLM-Lasso. For Ukraine, the best-performing configuration used LLM-Assisted Regularization (AR) and achieved a MAPE of 41.05%. Overall, the results suggest that contextual article-derived features and LLM-guided regularization can improve migration forecasting under certain conditions, although migration corridor characteristics, article volume, and hyperparameter configuration strongly influenced performance.
Subjects: Machine Learning (cs.LG); Computers and Society (cs.CY)
ACM classes: I.2.7; H.2.8; I.2.6
Cite as: arXiv:2610.07208 [cs.LG]
  (or arXiv:2610.07208v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07208

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fabricio Murai [view email]
[v1] Mon, 5 Oct 2026 18:22:57 UTC (254 KB)

来源:arXiv:cs.LG · arxiv.org