跳到正文
原文
Google AI:DEV 作者专属(RSS)· qianqiuwanzi·· 3 小时前AI 评分34

你的 AI 会忘记你说过的话:一个记忆层能解决大部分问题

Your AI forgets what you told it. A memory layer fixes most of it

AI 导读

开发者推出本地优先的智能体记忆层 HyperMarrow,主张在压缩上下文之前先把偏好、决策和结论写入本地存储,并保留原始措辞而非有损摘要。作者称记忆常已被记录,缺的是可寻址的返回路径,扩大上下文窗口只是移动而非消除这道墙。该设计将数据边界设为可配置项,存储在本机、向托管服务发送上下文需显式操作。

正文

You 官网 hm.qianshi.cooll your agent something on Monday. On Tuesday it asks you again. Not because it did not listen. The transcript is still there. The problem is that nothing which survived the session is addressable.

cover

A recent post on Dev.to put a number on that feeling: your AI ignores what you told it most of the time. I do not have a better measurement and I am not going to invent one. What I can do is point at public issue trackers, where the same failure was filed by people who hit it in production.

One report: assistant text emitted before a tool call is dropped from the UI, even though the text is persisted in the transcript. Another: conversations become unreachable, plus ghost and duplicate project entries. A third is a request rather than a bug: when resuming a large and stale session, suggest starting a new conversation instead.

Read those three together and the shape is clear. The memory is often already written down. What is missing is a way back to it.

The fix people reach for first

Rewriting the prompt. "Remember that I prefer X." It works inside one session and evaporates with the context window, because a prompt is an input, not a store. You cannot index it, version it, or cite it three sessions later.

The second instinct is a bigger context window. That moves the wall, it does not remove it. Compaction still fires, and it fires on its own schedule rather than on yours.

Four things a memory layer has to do instead

I build HyperMarrow, a local-first memory layer for agents. All four of these are write-time decisions, not query-time tricks.

1. Write before you compact

Preferences, decisions and conclusions go to the local store first, each with a source and a timestamp. Only then is the working context allowed to be summarized or dropped. If the summarizer throws something away, the durable copy is already on disk.

2. Separate the kinds of memory at write time

Not every sentence deserves the same treatment. A stated preference, like "we use pnpm here", has to survive verbatim. A conclusion is durable. Raw chatter is noise. Sorting them at query time is too late, because by then the original phrasing is gone.

3. Recall returns your words, not a summary of your words

This is where the reported bugs bite hardest. If what got stored was already a lossy summary, no ranker can recover the sentence you needed. Keep the exact phrasing and recall can hand it back. That is the part which makes an agent stop re-asking.

4. The boundary is a setting, not a promise

On a local-first design the store lives on your machine, and sending context to a hosted service is an explicit action rather than the default. "You own your data" and "your data never leaves your machine" are different statements, and only the second one is checkable from your side.

What changes in practice

The agent stops starting from zero. It knows you said pnpm. It knows which approach you already ruled out. It can point at the paragraph where you said it. The difference is not that the model got smarter. The difference is that it can read something which outlived the session.

If you run agents

The useful question is not whether your agent has memory. It is where that memory lives, and whether it is still addressable tomorrow. If the answer is "somewhere in the context window", you are one compaction away from starting over.

I am building HyperMarrow, the local-first memory layer described above. Docs and the client are here: HyperMarrow.

If you run agents: what is the thing you have to repeat most often?

(Disclosure: I build HyperMarrow, the local-first memory layer described above.)

来源:Google AI:DEV 作者专属(RSS) · dev.to