arXiv:cs.AI· Mandana Ghadamian, David Mohaisen·· 6 小时前AI 评分34
从失败中学习:面向 LLM 漏洞分析的失败驱动提示词优化方法
Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
AI 导读
研究提出失败驱动提示词优化方法 FDPR,通过分析 LLM 在漏洞分析中的反复失败模式来指导提示词改进。在 DVJA 上识别出假阳性、假阴性、无依据推理和 CWE 误分类等失败模式,并据此优化提示词,随后在 Juliet Test Suite 上评估并做跨模型验证。结果显示该优化提升了 LLM 漏洞分析的可靠性,并给出可复用的提示词设计原则。
正文
Abstract:Large Language Models have emerged as promising tools for software vulnerability analysis, but their effectiveness depends heavily on prompt design. Existing research primarily compares prompting strategies using aggregate performance metrics, providing limited insight into why models fail or how prompts can be improved systematically. We propose Failure-Driven Prompt Refinement (FDPR), a methodology that analyzes recurring model failures to guide evidence-based prompt refinement. Using the Damn Vulnerable Java Application (DVJA), we identify recurring failure modes, including false positives, false negatives, unsupported reasoning, and CWE misclassification, and translate them into targeted prompt refinements. We then evaluate the resulting prompt on the Juliet Test Suite and perform cross-model validation to assess generalizability. The results show that failure-driven refinement improves the reliability of LLM-based vulnerability analysis while yielding reusable prompt design principles. More broadly, this work demonstrates that recurring model failures provide a principled foundation for prompt engineering, enabling the systematic development of more reliable LLM-based vulnerability analysis systems.
| Comments: | Accepted at CSoNet 2026. 15 pages, 6 figures/tables combined |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.08405 [cs.SE] |
| (or arXiv:2610.08405v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08405 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: David Mohaisen [view email]
[v1]
Tue, 6 Oct 2026 14:15:44 UTC (21 KB)
来源:arXiv:cs.AI · arxiv.org