个人智能体的记忆笔记并非越多越好:用 Claude Haiku 4.5 测试,违规率从无记忆时的 77% 降到 10 行记忆时的 20%,再拉长记忆反而回升到约 25%。需要计数或追踪的任务应交给代码,例如记录累计消费,写规则仍有 44% 失败率,而代码实现为 0%。让智能体根据用户抱怨自行改写记忆的方法初期有改善随后停滞,最佳方法违规率仍近 48%,而直接告知全部偏好仅 7.1%。
More memory does not keep helping personal agents, because relevant notes help up to a point and then extra lines start making the agent miss rules.
A personal agent can't learn every user preference through memory notes, and more notes eventually make it worse, so use code for anything it must count or track and keep memory short.
Written rules work for style, like signing texts with the user's first name. For a running spending total, a stated rule still failed 44% of the time, while code that kept the total failed 0%.
Memory size has a sweet spot. With Claude Haiku 4.5, violations dropped from 77% with no memory to 20% at 10 lines, then rose to about 25% with longer memories.
Agents that rewrite their own memory from user complaints improve early, then stall. The best methods ended near 48% violations, against 7.1% when the agent was simply told every preference.
来源:Rohan Paul · x.com