arXiv:cs.LG(机器学习,全量分类)· Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha·· 14 小时前AI 评分51
arXiv 论文:用 2100+ Polymarket 题目训练 Qwen3.5-35B-A3B 学习边搜索边预测
Do Your Own Research: Learning to Forecast by Learning to Search
AI 导读
论文提出基于 2100+ 已结算 Polymarket 题目的智能体预测环境与数据集,Agent 在 rollout 时自行获取上下文(网页搜索、页面阅读、金融时间序列,并做分层泄漏过滤),用单轮 GRPO 和 Brier 分数奖励训练 Qwen3.5-35B-A3B(3B 激活参数)。
正文
Abstract:Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.
| Comments: | Accepted at the NeurIPS 2026 Workshop on Foundation Models for Temporal Systems (FMTS). 9 pages, 4 figures. Code and data: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01955 [cs.LG] |
| (or arXiv:2610.01955v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01955 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yusuf Afifi [view email]
[v1]
Thu, 1 Oct 2026 16:12:58 UTC (227 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org