跳到正文
arXiv:cs.AI· Saimon Amanuel Tsegai (Daphne), Alex Kantchelian (Daphne), Danfeng (Daphne), Yao, Peng Gao·· 3 小时前

从调查失败到可靠 SOC 智能体:理解并改进基于 LLM 的告警分诊

From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage

AI 导读

研究构建 ALERT-BENCH 交互式基准,通过实时 SIEM 回放企业遥测数据,评测单次工具调用、迭代检索、采样调查、自我审查与显式验证五类 LLM 智能体告警分诊方法,在 1,247 条多阶段攻击告警中每种方法至少漏报 40.4%。

正文

View PDF HTML (experimental)

Abstract:Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated. Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert. We study five representative approaches spanning single-pass tool use, iterative retrieval, sampled investigations, self-review, and explicit verification. To support this study, we build ALERT-BENCH, an interactive benchmark that replays enterprise telemetry through a live SIEM and requires each system to retrieve evidence. Across 1,247 alerts from a multi-stage attack scenario, every approach missed at least 40.4% of attack-related alerts. Trace analysis shows that attack alerts are more likely to be dismissed when searches return no records, same-context review has negative net correction, and dismissal receives no consistently stronger investigation than escalation. Based on these findings, we further design AIDA (Adversarial Investigation and Dialectical Analysis), a multi-agent framework that requires an explicit proposed decision before independent challenge and stronger evidentiary requirements before dismissal. AIDA preserves investigation history in an append-only Investigation Ledger and keeps the challenge in a separate reasoning context. A separate Judge adjudicates the proposed decision and challenge against evidence, resolving the alert or requesting another round when evidence is missing. On the same alerts, AIDA achieves an F1 score of 0.958, compared with 0.371-0.744 for the studied approaches, and reduces the false-negative rate from 40.4% to 3.1% while escalating 18.4% of alerts to analysts. These results show that structuring evidence retrieval and decision review can substantially improve agentic SOC triage.
Comments: Preprint
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.10608 [cs.CR]
  (or arXiv:2610.10608v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.10608

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saimon Amanuel Tsegai [view email]
[v1] Wed, 7 Oct 2026 03:20:51 UTC (991 KB)

来源:arXiv:cs.AI · arxiv.org