什么能让智能体配置保持真实:25 个工具与 6 篇论文的观察
What keeps an agent setup true
作者调研约 25 个编码智能体工具与配置、6 篇近期论文,总结出四条结论:智能体倾向遵循眼前上下文,需要主动加载的技能约一半情况被跳过(Vercel 测试中 AGENTS.md 通过 100%。
作者汇总约 25 个工具和 6 篇论文的实测数据,给出智能体配置文件、技能加载和记忆校验的可验证结论。
Notes from about 25 coding-agent tools and published setups and six recent papers. Agents follow what is in front of them, skip what they have to fetch, and almost nothing checks whether a rule is still true.
tags: ai, devtools, opensource, productivity
canonical_url: https://emeraldleaf.dev/writing/what-keeps-an-agent-setup-true/
Originally published on emeraldleaf.dev.
Anyone who uses a coding agent seriously ends up with a setup: a CLAUDE.md or AGENTS.md, some rules, a few skills, a hook or two, a handful of MCP servers. There are now tools to sync that setup across machines, convert it between agents and share it with a team. I wanted to know what people who do this well actually keep in their setups, how they keep them current, and what the evidence says about any of it.
So I looked at about 25 tools and published setups and read six recent papers. Four things stood out.
1. Agents follow what is in front of them, and skip what they have to fetch
Vercel ran the cleanest test I found. In their agent evals they gave an agent the same documentation two ways: as an index in AGENTS.md, which is always in context, and as skills the agent could load when it judged them relevant. AGENTS.md passed 100% of cases. Skills passed 53%, and in 56% of cases the agent never loaded the skill at all. Telling it explicitly to use skills raised the score to 79%.
Scott Spence measured the same effect from the other side. On its own, Claude Code activated the right skill for 50 to 55% of his test prompts. With a UserPromptSubmit hook that made it evaluate every skill before starting, activation went to 100%, 22 prompts out of 22.
The lesson for any setup: anything the agent has to decide to look up gets skipped about half the time. If something matters, put it in front of the agent before it starts, or let the harness do the deciding.
2. Always-on context costs more than it helps, unless it is the right kind
That does not mean putting everything in AGENTS.md. Gloaguen and colleagues tested repository context files on SWE-bench tasks and on real issues from repositories that already had them. The files "do not generally improve task success rates", whether a model or a developer wrote them, while raising inference cost by over 20%. The agents did follow the instructions in those files. What did not help were repository overviews, the "here is how the codebase is laid out" sections that most starter templates produce.
Other results point the same way. IFScale gave models up to 500 instructions at once; the best frontier models followed 68% of them at that density, and favoured the ones that came first. Anthropic's own advice for CLAUDE.md is blunt: for each line, ask "Would removing this cause Claude to make mistakes?" If not, cut it.
The evidence on file size itself is mixed. Damon McMillan ran 1,650 Claude Code sessions varying file size, instruction position, conflicting instructions and structure, and found no detectable effect from any of them. What did matter was time: each additional function the agent wrote in a session lowered the odds that it followed an instruction by about 5.6%. Instructions fade as a session goes on, wherever they sit.
Taken together: keep always-on context to specific, actionable instructions, and expect them to fade during a long session.
3. Memory that captures everything mostly does not help. Memory that is checked does
The newest evidence is VibeMemBench, which tested memory systems for coding agents on real repository tasks. Injecting past experience that had been verified by running it raised task resolution by 1.1 to 4.5 percentage points. Automatic memory systems, the kind that capture what happened and replay it, failed to beat the no-memory baseline in 11 of 12 pairings.
GitHub's Copilot Memory is the strongest counter-example, and it shows what makes automatic capture work. Its agents save facts about a repository on their own, but each fact carries citations to the code that supports it. Before using a fact, Copilot "checks those citations against the current branch to confirm the information is still accurate. Only validated facts are used." Facts that go unused for 28 days are deleted. GitHub reports a 90% merge rate for pull requests with memory against 83% without: a vendor's number, but a real-world one.
So the dividing line is not human versus automatic capture. It is whether a memory is checked against the code before the agent relies on it.
ACE, from Zhang and colleagues, adds a point about how memories should change over time. Rewriting a context wholesale leads to "context collapse, where iterative rewriting erodes details over time"; small, structured, incremental updates preserve them. Their approach improved agent benchmarks by 10.6%.
4. Setups drift, and almost nothing notices
Every setup goes stale, and the most careful ones say so. Aristidis Vasilopoulos runs a 108,000-line C# codebase on a 660-line constitution, 19 agent specifications and 34 documents, and reports that "specification staleness was the primary failure mode". When a subsystem changes and its spec does not, "the AI will generate code based on stale information." His fix is a session-start hook that compares recent commits with a map of subsystems to files, and warns when code changed but its spec did not.
Others do it by hand. Sentry's agent skills ship a SOURCES.md that maps each claim to the file and function supporting it, dated when it was captured and when it was last updated. Every's compound-engineering plugin, the closest thing I found to a learning loop, records solution notes from sessions and has a refresh command that keeps, updates, merges, replaces or deletes them. Its rule is one I would adopt anywhere: "Age alone is not staleness."
The portability tools solve a different kind of drift. rulesync, Microsoft's APM and similar tools generate each agent's configuration from one source, and can fail CI when a generated file no longer matches it (rulesync generate --check exits 1). That catches a stale copy. It does not ask whether the rule itself is still true of the code.
What I would take from this
- Keep always-on instructions short, specific and actionable. Cut overviews and anything the agent can read from the code.
- Do not rely on the agent to fetch what matters. Put it in front of the agent, or have a hook do it.
- Expect instructions to fade in long sessions, and restate the relevant ones at the moment they matter.
- Only keep memories you can check. A note with no evidence behind it is a guess the agent will believe.
- Retire rules when the evidence says they are wrong, not because they are old.
- Watch for drift in what a rule claims, as well as drift between copies. When the code a rule describes changes, someone should have to look again.
Where my own loop fits
That last point is the gap I have been working on with okl, an open-source tool I have been building. Each lesson can be proven by a check you choose, and CI goes red when the code it governs changes, until someone runs the check again. On my own eval, briefing lessons before each task cut repeated known bugs from 43% to 8%, on a small sample of tasks I wrote.
The research also showed me where okl falls short. Agents other than Claude Code have to call for its lessons, which the Vercel result says they often will not. It flags stale lessons in CI rather than at the moment they are used. And it could brief again when a governed file is about to be edited, since instructions fade within a session. Those are next.
How I did this
I used AI research assistants to survey the tools, published setups and papers, then re-read every claim cited here at its source. One research summary misstated a paper's finding; this piece uses the paper's own abstract. Tools and numbers in this area change weekly; everything here is as of 1 October 2026.
Sources
Papers
- Fan et al., VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks (2026)
- Gloaguen et al., Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? (2026)
- McMillan, Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables (2026)
- Jaroslawicz et al., How Many Instructions Can LLMs Follow at Once? (2025)
- Zhang et al., Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models (2025)
- Vasilopoulos, Codified Context: Infrastructure for AI Agents in a Complex Codebase (2026)
Measurements and documentation
- Vercel, AGENTS.md outperforms skills in our agent evals
- Scott Spence, Measuring Claude Code skill activation with sandboxed evals
- GitHub, Building an agentic memory system for GitHub Copilot and the Copilot Memory docs
- Anthropic, Best practices for Claude Code
Setups and tools
- Sentry, an agent skill with SOURCES.md
- Every, compound-engineering refresh guide
-
rulesync and its
--checkoption - Microsoft APM
来源:Google AI:DEV 作者专属(RSS) · dev.to