arXiv:cs.AI· Yunju Kang, Seonghyeon Cho, Irene Li, Yeo-Chan Yoon, Chanjun Park·· 6 小时前AI 评分32
POLAR:面向工具调用 LLM 智能体的本体引导风险预防框架
POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
AI 导读
POLAR 是一个面向小型工具调用智能体的护栏框架,通过结构化双层本体评估动作可逆性,为每个动作给出分级可逆性分数,并在执行前剪除未达阈值的调用。该框架在 τ²-bench 上跨六个智能体模型评测,使六个智能体中四个在 airline 域的平均任务奖励提升 0.11 至 0.18 分,但十八个模型—领域组合中仅八个整体改善,retail 域及更强智能体常出现性能回退。
正文
Abstract:LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on $\tau^2$-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.
| Comments: | Accepted Findings of AACL-IJCNLP 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08082 [cs.AI] |
| (or arXiv:2610.08082v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08082 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Seonghyeon Cho [view email]
[v1]
Tue, 6 Oct 2026 10:14:16 UTC (1,431 KB)
来源:arXiv:cs.AI · arxiv.org