arXiv:cs.LG· Leo Zeitler, Jack Richings, Victoria Nockles·· 4 小时前AI 评分49
论文探讨时间受限智能体的 AI 安全:部分可观测环境下无法保证安全行为
AI Safety Considerations for Agents With Limited Time to Act
AI 导读
一项理论研究指出,在只能部分观测且必须在有限时间内行动的环境中,即便是完美的 AI 智能体也无法保证安全行为。作者构造了无限状态空间和信号混合两种场景,证明智能体无关的安全保证存在理论边界,并提出任何 AI 安全或对齐证明都必须将环境与对应安全动作同智能体一并考虑。
正文
Abstract:In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theoretical bounds for agent-agnostic safety guarantees in environments that can only be partially observed and within which an action is required within limited time. We introduce two realistic scenarios, one with an infinite state space and one with signal mixture. In these scenarios, we prove that even a perfect agent cannot guarantee safe behaviour. It will be argued that for any proof of AI safety or alignment, the environment and associated safe actions need to be specifically considered together with the agent.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10285 [cs.AI] |
| (or arXiv:2610.10285v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10285 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Leo Zeitler [view email]
[v1]
Wed, 7 Oct 2026 15:49:06 UTC (204 KB)
来源:arXiv:cs.LG · arxiv.org