arXiv:cs.LG· Fernando Martinez, Abhishek Satyam, Tao Li, Junaid Farooq, Ying Wang, Juntao Chen·· 4 小时前AI 评分33
Ask the Expert:LLM 引导的强化学习框架实现自主网络防御
Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense
AI 导读
研究提出 Ask the Expert 训练期引导框架,先用 LLM 总结困难网络防御状态,再通过受限动作接口间歇查询主机级防御建议,并将其转化为分层奖励塑形以配合 PPO 训练,训练后丢弃 LLM,部署时仅为纯 RL 策略。
正文
Abstract:Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.
| Comments: | Accepted for publication at IEEE GLOBECOM 2026. Proceedings forthcoming. 6 pages, 4 figures |
| Subjects: | Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09337 [cs.CR] |
| (or arXiv:2610.09337v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09337 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fernando Martinez [view email]
[v1]
Wed, 7 Oct 2026 02:53:26 UTC (931 KB)
来源:arXiv:cs.LG · arxiv.org