跳到正文
原文
Google AI:DEV 作者专属(RSS)· Benny·· 5 小时前AI 评分42

服务器自己提交工单,等我签字后自愈:基于 Sanity 的 Ops Brain 项目

My Server Files Its Own Tickets, Then Waits for My Signature to Heal Itself

AI 导读

开发者用 Sanity 构建了 Ops Brain:一个 Python 智能体监控 Ubuntu 上的 nginx,服务挂掉时自动抓取日志并在 Sanity 中生成包含诊断、建议操作和风险等级的工单,人工在 Studio 中把状态改为 approved 并 Publish 后,智能体才重启服务并验证结果。

正文

Cover image for My Server Files Its Own Tickets, Then Waits for My Signature to Heal Itself

Benny

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange

What I Built

Servers crash at 3 AM. A human wakes up, SSHs in, types systemctl restart, and goes back to sleep. I wondered if the ticket could do that work instead.

Ops Brain is a server that files its own incident reports in Sanity and heals itself, but only after a human approves.

The loop:

  1. A small Python agent watches my Ubuntu machine.
  2. nginx dies. The agent grabs the recent logs and creates an incident document in Sanity with a title, service, diagnosis, proposed action, risk level and log excerpt.
  3. I open Sanity Studio and find a ticket from a server that reported its own death.
  4. I flip the status to approved and hit Publish.
  5. The agent sees my approval, restarts nginx, checks that it's actually healthy, and stamps the ticket verified with a resolved time.

The server files the ticket, I sign it, and the server fixes itself.

Demo

{https://drive.google.com/file/d/1JbS2bEXw86HEF122E8LWS6lPCNezKAyf/view?usp=drivesdk}

Live Studio: https://opsbrain-benny.sanity.studio/

Code

{https://github.com/Bencoy09/ops-brain }

How I Used Sanity

The whole project rests on one decision: the workflow is data, not code.

An incident is a document that moves through states: awaiting_approval → approved → verified (or failed). The agent and a human push the same document forward. There's no separate approvals app and no webhook maze. The document is the process.

The schema is a single incident type:

  • service, title, logExcerpt: what broke, plus the evidence
  • diagnosis, action, risk: what the agent proposes
  • status: the workflow state, shown as radio buttons in Studio so approving takes one click
  • detectedAt, resolvedAt: a built-in timeline of how long each outage lasted

Because incidents are structured content, I can ask GROQ questions like "show me everything that failed this week" without grepping a single log file.

My Build Process

The one rule I wouldn't bend: the agent never runs arbitrary commands. It executes only actions from a hard-coded allowlist (restart_service on explicitly watched services). Anything else is blocked and marked failed. An agent that obeys whatever text lands in a database is a remote-code-execution bug waiting to happen, so the allowlist is part of the design.

What went wrong (the honest part):

  • npm kept timing out. My connection dropped create sanity mid-install. The project and dataset had already been created, so I reran only npm install with --maxsockets=3 and it went through.
  • A half-written package broke sanity deploy. An interrupted install left zod corrupted. Deleting it and reinstalling that one package fixed it.
  • I pasted a shell command into my Python file. Line 1 was cat > ..., and Python was not amused. I deleted the stray line.
  • The classic gotcha: the agent only sees published documents. I approved an incident, nothing happened, and the cause was that I hadn't clicked Publish.
  • I leaked my own token while testing. I typed it on the command line, so I revoked it and switched to a hidden prompt (read -s) for the new one.

Tools: I used Claude to plan the architecture and debug each error as it appeared. The rest is deliberately plain: WSL Ubuntu, nginx, Python's standard library and Sanity Studio.

What I'd build next:

  • An Approve button inside Studio using the App SDK, instead of changing a dropdown.
  • A learning step where, after a successful fix, the agent drafts a new runbook entry for me to approve.
  • More watched services, with risk levels that decide what needs a human and what doesn't.

Sanity Project Details

来源:Google AI:DEV 作者专属(RSS) · dev.to