开发者自建跨客户端 AI 智能体记忆中枢 MemTether 并分享设计经验
I Built a Cross-Client Memory Hub for AI Agents — Here's What I Learned
开发者 lanbass869-cell 开源了 MemTether,一个本地优先的记忆中枢,让 23+ 个 AI 客户端共享同一个 SQLite 数据库,无需云端 API。
I use Claude Code for coding, Cursor for refactoring, and Windsurf for exploration. Each has its own memory. Switch tools and my AI forgets everything.
So I built MemTether — a local-first memory hub that lets 23+ AI clients share one physical SQLite database.
Here's what I learned building it.
The Problem Is Simpler Than You Think
Existing memory solutions (mem0, cognee, zep) all treat memory as a service. You send memories to their cloud API, they store and retrieve them. This works, but it means:
- Your memories live on someone else's server
- You pay per API call
- You need an API key even for a local project
- You can't inspect the raw data
My insight: if all your AI tools run on the same machine, you don't need a service — you need a file. Just make them all point to the same SQLite database.
No cloud. No API fees. No abstraction layer. One memory.db, 23 clients reading and writing to it.
Design Decisions That Matter
1. Supersession, Not Deletion
When a memory needs updating, I don't delete the old one. I mark it superseded and create a new version. This means you can always trace "what did we believe before we learned X?"
This sounds simple, but it changes everything about how the system works. Search must filter out superseded entries. FTS5 needs triggers to auto-sync. The projection system needs to prefer the latest version.
2. Bi-Temporal: Two Clocks, Not One
Every memory has two timestamps:
- T (valid time): when the fact was true in the real world
- T′ (recorded time): when the system learned about it
Example: an API key expired on September 15th, but I didn't notice until September 18th. Querying "what did we know on September 15th?" returns "the key is valid" (which is what the system believed). Querying "what was actually true?" returns "expired."
Without this distinction, you get retroactive truth bias — using today's knowledge to judge yesterday's decisions.
3. Q-Value: Memories That Get Used Should Rank Higher
Most memory systems rank by recency or semantic similarity. But a memory from 6 months ago that gets used every day is more valuable than one from yesterday that's never been retrieved.
So I added a Q-Value (inspired by reinforcement learning): every time a memory is searched and actually useful, its Q-Value increases. Next search, it ranks higher. Simple, effective, and nobody else does it.
4. Four-Factor Re-Ranking
Raw semantic similarity isn't enough. I blend four factors:
- Semantic similarity (45%)
- Recency (25%)
- Usage frequency (5%)
- Type importance (10%)
Then normalize with z-score + sigmoid and blend 70/30 with the RRF score. This fixed asset lookup queries that pure semantic search missed.
5. SQLite Triggers for FTS Sync (Not Python)
My first version had Python code to sync the FTS5 full-text index. It had except Exception: pass around the sync calls. Result: 254 stale entries and 260 missing entries.
The fix: SQLite triggers. Three triggers (INSERT, DELETE, UPDATE) at the SQL level. No Python code can accidentally skip them. The inconsistency went from 514 entries to zero.
Lesson: if SQLite can do it at the SQL level, don't do it in Python.
The Hard Part Isn't Code — It's Packaging
Writing the memory engine took 3 weeks. Making it installable took 2 months:
- Wheel building with correct
py-modules(flat layout, not packages) -
check_packaging.pyto catch version drift and missing modules - FTS5 triggers in the right place (inside
init_db, not floating in the schema string) - Console entry points (
memtether,memtether-connect) - Optional dependencies (
[vector]for chromadb,[server]for fastapi) - A Quick Start that actually works (I initially wrote commands that didn't exist)
I'm a solo developer. Every hour spent on packaging is an hour not spent on features. But without packaging, nobody can use your features.
Honest Benchmark Numbers
I ran LongMemEval (500 questions, full run):
| Metric | Score |
|---|---|
| Strict match | 62.6% |
| LLM judge | 54.6% |
| Multi-session strict | 48.8% |
| Multi-session LLM judge | 60.0% |
The multi-session gap (strict vs judge) is the most interesting finding. When an answer is computed (like "3 weeks" from multiple data points), strict substring matching fails because the computed answer doesn't appear verbatim in any single memory. The LLM judge is more forgiving.
I also built an E-Hybrid method (session summaries + flat evidence) that improved multi-session strict from 48.8% to 71.4% on a 15-question test. Small sample, but promising.
I'm not going to claim these are better than mem0's 94.4% — different harness, different methodology. But 48.8% strict is above the industry average of 27.9%.
What's Next
- MCP Registry submission (done: PR #4946)
- Community building (this blog post is part of that)
- NLPCC 2027 paper submission (CCF C, deadline ~April 2027)
- More examples and better docs
Try It
pip install memtether
memtether init
memtether connect --all
GitHub: MemTether/MemTether
PyPI: memtether
来源:Google AI:DEV 作者专属(RSS) · dev.to