跳到正文
arXiv:cs.LG· Srini Ramaswamy, Deveeshree Nayak·· 4 小时前AI 评分40

BRaVeS:面向智能体 AI 自动化的有界自主与可验证安全框架

Bounded Autonomy and Verifiable Safety for Agentic AI Enabled Automation

AI 导读

论文提出 BRaVeS 有界推理与安全治理框架(Defensible Next-Gen Reasoning System,DNRS),将领域专家约束编码为不变锚点,并提出 MoDA-Style 深度感知访问机制在推理中保持这些锚点可见,同时用 SMARtAutonomy 状态层级随认知风险升高而降低自主度。

正文

View PDF

Abstract:Agentic AI-enabled automation cannot be safely deployed in high-stakes environments on probabilistic reasoning alone. A recurring risk is epistemic drift: as reasoning deepens, system behavior may move away from subject-matter-expert constraints for safe operation. This paper presents BRaVeS, a bounded reasoning and safety-governance framework termed the Defensible Next-Gen Reasoning System (DNRS). BRaVeS encodes SME-defined constraints as invariant anchors, proposes MoDA-Style (Mixture of Depths Attention) depth-aware access as a candidate mechanism for keeping these anchors visible during inference, and uses a state hierarchy (SMARtAutonomy) to reduce autonomy as epistemic risk increases. To formalize bounded recovery, we introduce the Lyapunov-Bounded Consensus Framework (LBCF), which maps continuous epistemic-risk signals into a finite K-bag abstraction and applies shielded state transitions that enforce Lyapunov-style energy descent or route the system to a human-mediated terminal state. The formal convergence result applies to the finite LBCF abstraction under fixed thresholds and feasible-shield assumptions; it does not prove safety of the full continuous neural activation space. We evaluate the framework through a discrete event Monte Carlo simulation using HAI 22.04 industrial-control-system time-series data with synthetic noise and sensor-degradation regimes. Across the tested parameter-grouping strategies and thresholds, the LBCF process achieved finite-step convergence and no safety-guard violations. These results provide simulation-based evidence that bounded governance behavior can be enforced under the stated abstraction, while motivating future work on deployed transformer implementations, live human-in-the-loop validation, and broader adversarial settings.
Comments: This paper has been accepted and will appear in the Journal of Intelligent and Robotic Systems (https://doi.org/10.1007/s10846-026-02467-w)
Subjects: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Systems and Control (eess.SY)
Cite as: arXiv:2610.08815 [cs.LG]
  (or arXiv:2610.08815v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08815

arXiv-issued DOI via DataCite

Submission history

From: Srini Ramaswamy [view email]
[v1] Thu, 24 Sep 2026 02:09:49 UTC (1,577 KB)

来源:arXiv:cs.LG · arxiv.org