arXiv:cs.LG· Mingtian Tan, Mihir Parmar, Palash Goyal, Chun-Liang Li, Nanyun Peng, Thomas Hartvigsen, Jinsung Yoon, Tomas Pfister·· 5 小时前AI 评分40
LEAF:面向事件增强预测的活体基准
LEAF: A Living Benchmark for Event-Augmented Forecasting
AI 导读
研究者提出 LEAF,首个面向事件增强预测任务的活体基准,覆盖趋势、事件与时间序列预测,通过递归检索智能体系统与双智能体交叉验证收集辅助上下文。经 47 位领域专家对 500 项任务的审计,该流程将未来信息泄漏率从 8.6% 降至 1.6%。在 16 个前沿闭源与开源权重模型上的评测显示,LLM 能有效从已验证事件中提取信号以提升趋势与事件预测。
正文
Abstract:Large Language Models (LLMs) are increasingly applied to real-world forecasting tasks, yet evaluating their true predictive capability remains compromised by pre-training data contamination and look-ahead leakage in automated search. Existing benchmarks either rely on static contexts, restrict evaluations to narrow environments, or fail to audit auxiliary textual events for future information leakage. To establish a rigorous evaluation paradigm, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including trend, event, and time series forecasting. LEAF couples a recursive retrieval agent system with dual-agent cross-validation to gather comprehensive, relevant, and temporally aligned auxiliary context. A comprehensive audit across 500 tasks by 47 domain specialists demonstrates that our pipeline suppresses future information leakage from 8.6% to 1.6%. Across extensive evaluations of 16 frontier proprietary and open-weight models on our benchmark, we show that LLMs effectively extract signals from verified events to boost trend and event forecasting.
| Comments: | 12 tables, 6 figures, 39 pages |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.16358 [cs.LG] |
| (or arXiv:2605.16358v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.16358 arXiv-issued DOI via DataCite |
Submission history
From: Mingtian Tan [view email]
[v1]
Sat, 9 May 2026 03:17:59 UTC (27,843 KB)
[v2]
Thu, 1 Oct 2026 22:35:36 UTC (28,033 KB)
来源:arXiv:cs.LG · arxiv.org