跳到正文
原文
Google AI:DEV 作者专属(RSS)· Rulestack·· 6 小时前AI 评分33

CLAUDE.md 里哪一行真正改变了你的 AI 智能体行为?

Which line in your CLAUDE.md actually changed what your agent does?

AI 导读

一位开发者复盘 2,285 行 CLAUDE.md 后表示,无法证明任何一条规则仅靠文字表述就改变了 AI 智能体行为。三条曾被认为有效的规则中,"禁止复制粘贴"在生效 34 天后仍被违反,直到加入测试脚本才拦住;"按周计划执行"也需在 2026-07-28 改为文件校验后才落实。他因此发问:有哪一行规则在无任何强制手段时真正改变了你的智能体行为?

正文

Rulestack

Our project's CLAUDE.md was 2,285 lines on 2026-08-18, the day we split it into skills and rule files. Going back through the rules that survived, I can name three that only held once a script enforced them, and none that I can prove worked on its words alone. This is a question post: I'd like to hear which line changed your agent's behavior.

Three rules that needed a script

"No copy-paste posts. Vary the content every time." This line went into CLAUDE.md on 2026-05-26. On 2026-06-29 we added a test that reads the post queue and fails the commit when a post opens with a canned phrase, or overlaps a past post by 0.4 or more (shared three-word runs). The first run of that test found three posts that broke the rule, in a queue the agent had filled while the rule was in place: two opening with "Just published", and one that was a near-copy of a post already published on 2026-06-02.

Terminal: pnpm test on 2026-06-29 fails no-boilerplate-stock with two posts opening

The rule had been loaded on every run for 34 days. The posts still went in.

"Do what the weekly plan says." Every week the agent writes its own plan for the next one, with lines such as "one post a day that ends in a question to the reader". One plan asked for that, and the queue built two weeks later had 3 such posts out of 30. Since 2026-07-28 the plan's numbers go into a file the test suite checks against the queue, so a plan that is not met fails the commit instead of being forgotten.

"Pull before you start." This one the agent did follow. On 2026-08-19 a run pulled at 16:08 UTC and pushed at 17:04. In between, two scheduled jobs had pushed, and the push was rejected. The rule said when to pull and nothing about how long a run takes. We replaced "remember to pull" with a push script that pulls first and stops on a conflict.

What I can't show

I never measured a rule on its own, before and after, with nothing else changing. So I can't point to a line that changed behavior through its wording alone. What I can say is that each time a rule mattered enough to check, we found it had been read and not followed, or followed and not enough.

That may mean prose rules don't carry much weight. It may also mean we stopped giving them a fair chance, because a test is easier to trust than a sentence.

What I'd like to know

  • Which one line changed your agent's behavior with nothing enforcing it? What did it do before, and after?
  • How did you know it was that line? Did you compare runs, or notice a mistake stopped?
  • Which line did you delete because it never did anything?
  • Do you keep CLAUDE.md short and move rules into hooks and tests, or keep them in prose?

Rulestack sells guides, hooks and skills for Claude Code at rulestack.gumroad.com.

Your one line, or the line you gave up on, belongs in the comments below. I'll answer each one there. For the measurements behind posts like this, follow @ai-shop.bsky.social.

来源:Google AI:DEV 作者专属(RSS) · dev.to