跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Elias Malomgr\'e, Pieter Simoens·· 19 小时前AI 评分35

Alignment Flywheel:面向架构无关安全的治理中心混合 MAS

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

AI 导读

研究者提出 Alignment Flywheel,一种将决策生成与安全治理解耦的治理中心混合 MAS 架构,由 Proposer 生成候选轨迹、经 Safety Oracle 返回安全分数与审计覆盖不确定性、再由 Enforcement 层在运行时执行风险策略。

正文

View PDF HTML (experimental)

Abstract:Multi-agent systems provide mature abstractions for role decomposition, coordination, and normative governance, but increasingly capable learned components make post-deployment safety harder to inspect, audit, and update. When safety behavior is absorbed into a decision component, narrow failures may require retraining or rollback of the full component. This instantiates our vision of the Alignment Flywheel as a governance-centric hybrid MAS architecture that decouples decision generation from safety governance. We denote the agent or policy that generates candidate trajectories as the Proposer; it passes its output to a governed Safety Oracle stack, which returns safety scores, prediction uncertainty, audit coverage uncertainty, and evidence hooks through a stable interface. An Enforcement layer applies explicit risk policy at runtime. Around this loop, a governance MAS performs monitoring, red-teaming, verification, triage, refinement, and versioned release management. The central engineering principle is patch locality: many newly observed safety failures can be mitigated through small governance batches for the Oracle stack and its audit state rather than by retraining or retracting the Proposer. The architecture is implementation-agnostic with respect to both Proposer and Oracle. It defines the roles, artifacts, protocols, and release semantics needed for runtime gating, audit intake, signed updates, staged rollout, and rollback. We demonstrate executability in two scenarios: a learned spatial Oracle patched through regression-checked governance updates, and a clinical GenAI proxy setting illustrating structured norms, escalation, and audit coverage. Our implementation code and documentation are available open source at this https URL.
Comments: Accepted for the EMAS workshop at AAMAS 2026
Subjects: Multiagent Systems (cs.MA); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as: arXiv:2603.02259 [cs.MA]
  (or arXiv:2603.02259v3 [cs.MA] for this version)
  https://doi.org/10.48550/arXiv.2603.02259

arXiv-issued DOI via DataCite

Submission history

From: Elias Malomgré [view email]
[v1] Sat, 28 Feb 2026 00:48:06 UTC (143 KB)
[v2] Wed, 29 Apr 2026 14:17:45 UTC (2,429 KB)
[v3] Thu, 1 Oct 2026 10:59:38 UTC (2,567 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org