跳到正文
原文
Google AI:DEV 作者专属(RSS)· yureki_lab·· 3 小时前精选AI 评分67

如何为 AI 编码智能体搭建复盘循环,让它不再重复犯错

How I Built a Post-Mortem Loop So My AI Coding Agent Stops Repeating Mistakes

AI 导读

作者为基于 Claude Code 的全自动开发系统搭建了 capture、distill、admit 三阶段复盘循环:失败运行生成结构化事故文件,独立 agent 提炼规则,每条规则须通过 replay 回归检查才能进入 playbook。三个月内重复失败占比从 38% 降至 6%,playbook 硬上限 60 条规则,约三分之一被采纳的规则实际修补的是作者自己任务描述中的歧义。

推荐理由

作者给出可复用的复盘循环设计,用 believed 与 actual 差值和 replay 准入门槛把重复失败从 38% 降到 6%。

正文 · 原文

TL;DR

My fully autonomous implementation system kept making the same handful of mistakes every single week, and I was the one patching its instructions at 11pm. So I built a post-mortem loop: every failed run produces a structured incident file, a separate agent distills those into one-line rules, and each rule has to pass a replay regression check before it's allowed into the agent's playbook. Over three months, repeat failures dropped from 38% of all failures to 6%. Here's the design, the load-bearing code, and what I'd do differently. 🚀

The Problem

Some context first. I run a 24/7 autonomous dev system built on Claude Code (the 2.x line as of this writing). An orchestrator module picks tasks, parallel implementation agents do the work in isolated worktrees, a self-healing agent retries broken builds, and I review the output in the morning.

It works. Most days I wake up to a few mergeable PRs. But after about two months in production I did something I should have done on day one: I classified every failure from the previous twelve weeks.

Failure bucket Share
Genuinely new problem 41%
Flaky infra / API outage 21%
Something we had already seen before 38%

That 38% bucket hurt. Same mistakes, different Tuesday:

  • ❌ Running the test suite from the repo root when the package lives in packages/api, then "fixing" the resulting import errors
  • ❌ Deleting a file it decided was "stale" that was actually loaded dynamically at runtime
  • ❌ Retrying a flaky network call 40 times inside a single run instead of failing fast
  • ❌ Resolving a lint error by disabling the rule

Each repeat cost me roughly 25 minutes of review and cleanup, plus a few dollars of tokens. Worse, it eroded trust. An agent that makes new mistakes is learning. An agent that makes the same mistakes is a liability.

And here's the embarrassing part: I had a lessons-learned document. It was 900 lines long. The agent either never read it, or read it and treated it as background noise. I was writing lessons for a reader that didn't exist.

The constraint that made this interesting: I didn't want to add a human to the loop. The whole point of the system is that it runs without me. So the fix had to be a mechanism, not a habit.

How I Solved It

The loop has three stages: capture, distill, admit. Each stage is owned by a different agent with a different prompt, and each stage produces a file the next stage reads.

flowchart LR
    A[Failed run] --> B[Capture: incident file]
    B --> C[Distill: proposed rule + replay scenario]
    C --> D{Admit gate: replay passes?}
    D -- yes --> E[Playbook]
    D -- no --> F[Rejected, logged]
    E --> G[Next run reads playbook]
    G --> A

Stage 1: Capture — a structured incident, not a log dump

When a run fails for any reason (test failure, watchdog kill, human rejection in review), the orchestrator asks the agent that just failed to fill in a fixed template before the context is thrown away. The template is small on purpose:

# incidents/2026-07-14-0932.yaml
task: "Add rate limiting to /v1/upload"
failure_class: wrong_assumption   # wrong_assumption | bad_tool_use | scope_creep | env | unknown
trigger: "CI failed: 14 import errors after agent moved tests"
believed: "Tests are run from repo root with `pytest`"
actual: "Each package has its own pytest.ini; must cd into packages/api first"
cost_minutes: 22
resolved_by: human

The two fields that matter are believed and actual. A stack trace tells you what broke. The delta between what the agent believed and what was true tells you why, and that delta is the only thing worth turning into a rule.

I tried free-form post-mortems first. They were long, apologetic, and useless. Forcing a one-line believed and a one-line actual made the agent actually commit to a diagnosis.

Stage 2: Distill — a separate agent writes the rule

Every ten incidents (or weekly, whichever comes first), a distiller agent reads all unresolved incident files and proposes rules. It is not the agent that failed. That separation turned out to matter a lot (see lesson 4).

The distiller's prompt has three hard constraints:

  1. A rule is at most two sentences and must say when it applies and what to do.
  2. A rule must cite at least two incidents. One-incident rules are tagged provisional.
  3. Every rule ships with a replay scenario: a minimal, reproducible task derived from the original incident.

Here's what a proposed rule looks like coming out of the distiller:

# proposed/rule-0041.yaml
rule: >
  Before running any test command, check for a package-local test config
  (pytest.ini, jest.config.*, vitest.config.*) and run from that directory.
cites: [2026-07-14-0932, 2026-07-21-1105, 2026-08-02-0847]
status: candidate
replay:
  repo: fixtures/monorepo-two-packages
  task: "Fix the failing test in packages/api/tests/test_limits.py"
  pass_if: "agent runs pytest with cwd == packages/api"

Notice the rule is boring. That's the goal. Clever rules get ignored; boring, specific, checkable rules get followed.

Stage 3: Admit — no replay, no rule

This is the stage that made the difference. A rule does not enter the playbook because it sounds reasonable. It enters because it demonstrably fixes the replay and doesn't break anything else.

The admit gate runs the agent on the replay scenario twice: once with the current playbook, once with the current playbook plus the candidate rule. Then it runs the full replay suite (about 30 scenarios by month three) with the candidate included.

# admit_gate.py (Python 3.13) — the part that matters
from dataclasses import dataclass

@dataclass
class ReplayResult:
    scenario_id: str
    passed: bool
    cost_usd: float

def admit(candidate: Rule, playbook: Playbook, suite: list[Scenario]) -> bool:
    # 1. Does the rule actually fix the thing it claims to fix?
    before = run_replay(candidate.replay, playbook)
    after  = run_replay(candidate.replay, playbook.with_rule(candidate))
    if before.passed or not after.passed:
        reject(candidate, reason="replay not discriminative")
        return False

    # 2. Does it break anything we already protect against?
    regressions = [
        r for r in run_suite(suite, playbook.with_rule(candidate))
        if not r.passed
    ]
    if regressions:
        reject(candidate, reason=f"regressed {[r.scenario_id for r in regressions]}")
        return False

    # 3. Playbook cap: 60 rules. Adding one past the cap retires the coldest one.
    if len(playbook) >= 60:
        playbook.retire(playbook.coldest())
    playbook.add(candidate)
    return True

Two details worth calling out:

  • The "before" run must fail. If the agent already passes the replay without the rule, the rule isn't teaching anything. About a quarter of candidates got rejected here, which told me the distiller was over-eager.
  • The playbook has a hard cap of 60 rules. Every rule tracks a last_hit timestamp (updated whenever the agent explicitly cites it in a run). When the cap is reached, the coldest rule gets retired to an archive. This single constraint is why the playbook stayed readable instead of becoming another 900-line document.

What the loop looked like after three months

Metric Month 1 Month 3
Total failed runs 71 64
Repeat failures (share) 38% 6%
Rules admitted 9 44 (cumulative)
Rules rejected by gate 4 19 (cumulative)
Rules retired (cold) 0 7
Avg replay suite cost per admit $1.10 $3.40

Total failures barely moved, which surprised me at first. The loop doesn't stop the agent from hitting new problems. It stops it from hitting the same problem twice. That's exactly what I wanted: new failures are information, repeat failures are waste.

The replay suite cost per admission tripled because the suite grew. I'm fine with that. Three dollars to permanently retire a 25-minute recurring mistake is the best trade in the whole system.

Lessons Learned

1. The lesson is in the delta, not the stack trace

believed vs. actual was the single highest-leverage design decision. Logs tell you what happened. The delta tells you what to change. If your post-mortem template doesn't force the agent to state what it assumed, you're collecting noise.

2. A rule you can't replay is a vibe, not a rule

Before the admit gate, I had rules like "be careful with monorepos." Totally unfalsifiable. Requiring every rule to ship with a replay scenario filtered out everything vague, and the ones that survived were specific enough that the agent could actually act on them.

3. Cap the playbook, or it becomes the 900-line doc again

Instruction files for AI agents decay the same way wiki pages do. Nobody prunes them, so they grow until they're ignored. A hard cap plus a "retire the coldest" rule keeps the playbook under what the agent can genuinely hold in attention. Sixty felt right for my system. Yours may be forty.

4. Don't let the agent that failed write the rule

Early on, I had the failing agent propose its own rule in the same context. The rules were defensive and over-specific ("never move test files"). A fresh distiller agent reading three incidents at once wrote better, more general rules because it wasn't trying to justify itself. Separation of concerns applies to agents too.

5. A lot of "AI mistakes" were my mistakes with a delay

About a third of admitted rules turned out to be patching ambiguity in my own task specs. "Fix the failing test" with no mention of which package. The loop made my sloppiness visible and measurable, which was uncomfortable and extremely useful.

What's Next

  • Rule scoring. Right now retirement is purely by last-hit date. I want to weight by how expensive the original incident was, so a rule that prevents a 2-hour disaster outlives one that prevents a 5-minute annoyance.
  • Spec feedback, not just rules. When the distiller sees that an incident traces back to an ambiguous task spec, it should propose a change to the spec template instead of a playbook rule. Fix the source, not the symptom.
  • Cross-project rules. I run this on several codebases. Some rules (the monorepo test config one) are clearly universal. I want a shared tier of rules with its own, stricter admit gate.
  • Cheaper replays. The replay suite is the slowest part of the loop. Running the discriminative check on a smaller model first and only escalating to the full run on a pass would cut the cost significantly.

Wrap-up

If your coding agent makes the same mistake twice, that's not a model problem. It's a feedback loop problem. Capture the belief delta, make rules replayable, gate admission, cap the playbook. That's the whole trick.

I'd love to hear what your agent's most repeated mistake is. Drop it in the comments, and if you've found a better admit criterion than "the before-run must fail," I genuinely want to steal it. 💡

Follow me here on Dev.to for more build logs from running a fully autonomous implementation system in production. Next up: how the replay suite itself gets maintained without turning into a second codebase. ✅

来源:Google AI:DEV 作者专属(RSS) · dev.to