arXiv:cs.AI· Mubashar Iqbal, Asifullah Khan, Hifsa Asif, Saddam Hussain Khan, Umme Zahoora, Asifullah Khan·· 5 小时前AI 评分37
成本感知的分层多智能体勒索软件检测与分析预算下的家族归因
Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution under Analysis Budgets
AI 导读
研究者提出成本感知分层多智能体系统(HMAS),将证据获取建模为预算约束下的序贯决策问题,静态证据充分时不做动态与内存分析。在 12439 个样本(覆盖 16 个勒索软件家族)上,确定性 HMAS 取得 F1 0.93、macro F1 0.97,57.95% 的样本仅靠静态证据即可判定;平均分析成本 6.65 单位,相比穷举分析的 12 单位降低 44.6%。
正文
Abstract:Sandbox execution and memory forensics are among the most constrained resources in malware triage. Static analysis can scale to millions of files, whereas dynamic and memory analysis require minutes of analyst controlled infrastructure for each sample. Despite this difference, multimodal ransomware detectors often apply every modality to every sample, causing analysis cost and time to verdict to increase linearly with sample volume even when static evidence is already sufficient for a decision. We present a cost aware Hierarchical Multi-Agent System that formulates evidence acquisition as a budgeted sequential decision problem. Specialist agents generate schema validated risk signals for each modality, domain controllers aggregate these signals, and a Meta-Orchestrator begins with static evidence and escalates to dynamic and memory evidence only when confidence is insufficient or agents within a controller disagree. An optional, bounded, locally hosted large language model reviewer can adjust a verdict by at most one tier but cannot replace the deterministic pipeline. Each decision is recorded with a complete provenance trace. In multiple runs over 12439 samples from 16 ransomware families and benign samples, the deterministic HMAS achieves F1 0.93 and macro F1 0.97, resolving 57.95% of cases using static evidence alone, 35.83% after adding dynamic evidence, and only 6.21% through the full pipeline. The average internal analysis cost is 6.65 units, compared with 12 for exhaustive analysis, representing a 44.6% reduction. Standalone leave-one-family-out testing further shows that accuracy on families held out during tuning falls to 0.26 to 0.64 outside the Benign and high support classes. We report these results alongside a cost sensitivity analysis, a partial leave one component out ablation, and a full scale comparison with learned early and late fusion and cascade baselines.
| Comments: | 19 Pages |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.04820 [cs.CR] |
| (or arXiv:2609.04820v2 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2609.04820 arXiv-issued DOI via DataCite |
Submission history
From: Asifullah Khan [view email]
[v1]
Fri, 4 Sep 2026 07:17:09 UTC (1,011 KB)
[v2]
Tue, 6 Oct 2026 13:30:24 UTC (702 KB)
来源:arXiv:cs.AI · arxiv.org