跳到正文
arXiv:cs.AI· XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng·· 5 小时前AI 评分53

arXiv 论文提出长程智能体跨轮次安全失效模式 GHOST 及防御方法 STAR-Guard

A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

AI 导读

arXiv 论文(arXiv:2610.02664)提出长程智能体的新失效模式 GHOST,即在良性交互条件下智能体可能执行违反多轮之前指定安全约束的动作,在 GPT-5.5 上发生率为 11.5%。

正文

View PDF HTML (experimental)

Abstract:Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety concern. Under benign interaction conditions, an agent may execute an action that violates a safety constraint specified many turns earlier. We term this failure mode Governance Hazard from Overlooked Safety Constraints across Turns (GHOST), which may cause irreversible damage. Our experiments reveal that GHOST events are not isolated cases: this failure mode, occurring precisely under benign interaction conditions, yields an occurrence rate of 11.5% on GPT-5.5. Furthermore, we theoretically show that if the residual conditional violation hazard along each safe prefix is bounded below by a non-summable sequence, the execution enters the hazard region almost surely. Leveraging this theoretical insight, we further propose STAR-Guard, a two-layer defense coupling historical semantic safety constraint restoration with pre-execution audit. STAR-Guard restores applicable safety constraints to reduce unsafe proposals, while its deterministic audit layer prevents residual violations from reaching the environment. Consistent with this two-layer design, we observe no GHOST events in our experiments under the GPT-5.5 setup.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02664 [cs.AI]
  (or arXiv:2610.02664v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02664

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: XinPeng Shen [view email]
[v1] Fri, 2 Oct 2026 01:33:04 UTC (1,690 KB)

来源:arXiv:cs.AI · arxiv.org