跳到正文
arXiv:cs.LG· Mingtian Tan, Mihir Parmar, Palash Goyal, Chun-Liang Li, Nanyun Peng, Thomas Hartvigsen, Jinsung Yoon, Tomas Pfister·· 5 小时前AI 评分40

LEAF:面向事件增强预测的活体基准

LEAF: A Living Benchmark for Event-Augmented Forecasting

AI 导读

研究者提出 LEAF,首个面向事件增强预测任务的活体基准,覆盖趋势、事件与时间序列预测,通过递归检索智能体系统与双智能体交叉验证收集辅助上下文。经 47 位领域专家对 500 项任务的审计,该流程将未来信息泄漏率从 8.6% 降至 1.6%。在 16 个前沿闭源与开源权重模型上的评测显示,LLM 能有效从已验证事件中提取信号以提升趋势与事件预测。

正文

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) are increasingly applied to real-world forecasting tasks, yet evaluating their true predictive capability remains compromised by pre-training data contamination and look-ahead leakage in automated search. Existing benchmarks either rely on static contexts, restrict evaluations to narrow environments, or fail to audit auxiliary textual events for future information leakage. To establish a rigorous evaluation paradigm, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including trend, event, and time series forecasting. LEAF couples a recursive retrieval agent system with dual-agent cross-validation to gather comprehensive, relevant, and temporally aligned auxiliary context. A comprehensive audit across 500 tasks by 47 domain specialists demonstrates that our pipeline suppresses future information leakage from 8.6% to 1.6%. Across extensive evaluations of 16 frontier proprietary and open-weight models on our benchmark, we show that LLMs effectively extract signals from verified events to boost trend and event forecasting.
Comments: 12 tables, 6 figures, 39 pages
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.16358 [cs.LG]
  (or arXiv:2605.16358v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2605.16358

arXiv-issued DOI via DataCite

Submission history

From: Mingtian Tan [view email]
[v1] Sat, 9 May 2026 03:17:59 UTC (27,843 KB)
[v2] Thu, 1 Oct 2026 22:35:36 UTC (28,033 KB)

来源:arXiv:cs.LG · arxiv.org