跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Aditya Dubey, Namah Gupta, Vinti Agarwal·· 9 小时前AI 评分34

将大语言模型接入 DSGE 模拟器进行政策生成与预测

Grounding Large Language Models in DSGE Simulators for Policy Generation and Forecasting

AI 导读

研究将指令微调的大语言模型放入六个 Snowdrop 支持的 DSGE 模拟器中,模型每轮观察经济状态与话语变化、选择有界政策动作,并获得下一模拟状态和经济奖励。实验以 PPO 为主方法、GRPO 作为无价值函数的匹配基线,测试方向性语义信号、奖励视野、轨迹热启动、跨模拟器迁移及疫情与货币政策冲击。目标是按模拟经济后果而非语言合理性来评判政策动作。

正文

View PDF HTML (experimental)

Abstract:Large language models can produce economic policy responses that sound reasonable, but this does not show that their actions are consistent with economic dynamics. We test this by placing an instruction-tuned language model inside six Snowdrop-backed dynamic stochastic general equilibrium (DSGE) simulators. At each turn, the model observes the economy and a change in economic discourse, selects a bounded policy action, and receives the next simulated state and an economic reward. We implement a common Python interface for repeated rollouts, persistent shocks, state cloning, and rolling-horizon simulation.
This setting creates a long-horizon credit-assignment problem. Policy effects may appear several quarters after an action is taken. PPO has a learned value function that can propagate delayed reward to earlier tokens through generalized advantage estimation. GRPO has no learned value function and instead assigns a group-relative advantage from complete rollout returns. It therefore cannot distinguish which earlier turn caused the outcome; if every rollout receives the same return, the normalized advantage is zero. We use PPO as the primary method and GRPO as a matched critic-free baseline. The experiments also test directional semantic signals, reward horizon, trajectory warm starts, cross-simulator transfer, and historically anchored pandemic and monetary-policy shocks. The objective is to judge policy actions by their simulated economic consequences rather than by plausible language alone.
Subjects: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
Cite as: arXiv:2610.01128 [cs.AI]
  (or arXiv:2610.01128v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01128

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Vinti Agarwal [view email]
[v1] Thu, 1 Oct 2026 06:11:48 UTC (479 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org