arXiv:cs.LG· Marsalis Gibson, Claire Tomlin, Shankar Sastry·· 4 小时前AI 评分34
ATLAS:面向潜空间对抗搜索的自适应信赖域与主动学习框架
ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning
AI 导读
研究者提出 ATLAS(Adaptive Trust-Regions for Latent Adversarial Searches),一个基于查询的黑盒对抗输入集发现框架,将攻击生成建模为主动学习的水平集估计问题,结合校准近似与局部-全局采样定位对抗区域。
正文
Abstract:Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textit{sets} of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), which is a query-based framework that discovers adversarial input sets for black-box learning systems. ATLAS casts attack generation as an active learning level set estimation problem then combines calibrated approximations with a local-global sampling architecture to find regions of the input space that contain adversarial examples. Once discovered, ATLAS is designed to sample points within these adversarial regions to build adversarial sets that accurately represent the state of robustness of the target model. When applied on toy experiments, we find that ATLAS is able to recover more of the adversarial region under a limited query budget than does previous work. When applied to standard and adversarially trained MNIST, CIFAR, and ImageNet model targets, ATLAS produces better representative attacks than other query-based black-box attacks (NES, SignHunter, BayesOpt). ATLAS represents an automated red-teaming framework that can be used for both analyzing the robustness of learning-based systems under development and continuous auditing to see how the robustness of a system changes over time.
| Comments: | 9 pages main body, plus 10 additional pages for references and appendix |
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.07323 [cs.LG] |
| (or arXiv:2610.07323v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07323 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Marsalis Gibson [view email]
[v1]
Mon, 5 Oct 2026 19:56:17 UTC (15,602 KB)
来源:arXiv:cs.LG · arxiv.org