跳到正文
arXiv:cs.LG· Sagar Srinivas Sakhinana, Venkataramana Runkana·· 6 小时前AI 评分33

面向自主云 MLOps 的前置部署全栈工程框架

Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps

AI 导读

研究者提出一种证据门控多智能体框架,可将自然语言描述的 MLOps 云工程任务转化为经过验证的代码仓库与云端部署。框架在 Google Cloud Platform 上实现,由有状态 Graph Orchestrator 协调仓库生成、审查、执行、验证、发布与监控等专用智能体,并管理证据门控、重试边界与回滚路径。

正文

View PDF HTML (experimental)

Abstract:Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipelines, cloud infrastructure, security, deployment, monitoring, retraining, recovery, and rollback. We present an evidence-gated multi-agent framework for transforming a natural-language MLOps cloud engineering task into a verified repository and operational cloud deployment. The framework combines graph engineering, loop engineering, and agent harness engineering. A stateful Graph Orchestrator coordinates specialized agents for repository generation, review, execution, verification, release, and monitoring while governing workflow dependencies, evidence gates, retry bounds, recovery paths, and termination. Consequential lifecycle transitions proceed only when their required predicates are supported by verifiable execution or runtime evidence. Verification failures activate bounded reflection, repair, and re-verification, while runtime evidence of failure, drift, degradation, or policy violation can trigger bounded adaptation, recovery, or rollback. Agent harness engineering constrains repository generation, review, and repair, artifact execution, and cloud operations through controlled capabilities and isolated execution environments. We realize the framework on Google Cloud Platform and evaluate repository completeness, controlled execution, evidence-gated transitions, cloud promotion, and bounded recovery. Our experimental results show that the framework prevents unsupported lifecycle transitions and drives each run toward either a verified operational deployment or an auditable terminal failure.
Comments: Nill
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2608.29615 [cs.MA]
  (or arXiv:2608.29615v2 [cs.MA] for this version)
  https://doi.org/10.48550/arXiv.2608.29615

arXiv-issued DOI via DataCite

Submission history

From: Sagar Srinivas Sakhinana [view email]
[v1] Sun, 30 Aug 2026 07:13:10 UTC (40 KB)
[v2] Wed, 7 Oct 2026 05:58:49 UTC (228 KB)

来源:arXiv:cs.LG · arxiv.org