跳到正文
arXiv:cs.AI· Siu Tung Wong (Institute of Finance and Technology, University College London), Carlo Campajola (Institute of Finance and Technology, University College London, UZH Blockchain Center)·· 4 小时前AI 评分32

当正确奖励仍不够:在解析可解的经纪商-交易者博弈中诊断与引导 PPO

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

AI 导读

研究将 PPO 智能体放入解析可解的连续时间经纪商-交易者博弈中,用有限步奖励替代经纪商决策。在无噪声订单流下 PPO-FFNN 接近参考动作,但在随机订单流下 PPO-FFNN 与 PPO-LSTM 均不准确,其 critic 无法可靠排序邻近动作,基于势能的奖励塑形也无稳定改善。冻结解析策略后微调,成本减半可稳定缩小 2.22% 的差距。

正文

Authors:Siu Tung Wong (1), Carlo Campajola (1 and 2) ((1) Institute of Finance and Technology, University College London, (2) UZH Blockchain Center)

View PDF HTML (experimental)

Abstract:Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results.
We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs.
Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
Comments: 8 pages; accepted for publication at ICAIF 2026
Subjects: Trading and Market Microstructure (q-fin.TR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03598 [q-fin.TR]
  (or arXiv:2610.03598v1 [q-fin.TR] for this version)
  https://doi.org/10.48550/arXiv.2610.03598

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Siu Tung Wong [view email]
[v1] Fri, 2 Oct 2026 17:03:29 UTC (539 KB)

来源:arXiv:cs.AI · arxiv.org