微软等机构提出 CorpusMap,预先解析文档集合中反复出现的实体,为每个实体生成链接所有相关文档的页面,原文保持原位。在 7 个模型和 3 个基准上,答案质量提升 6.4 至 11.7 分,输入 token 减少 34% 至 57%,并优于 LLM Wiki 层及另外三种导航层。该地图无需 LLM 调用即可构建,并支持随新文档更新。
Banger paper from Microsoft and colleagues.
If you run agents that search a large document collection, this one is worth your time.
(bookmark it)
They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it.
The original documents stay in place.
The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again.
Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%.
It also beats an LLM Wiki layer and three other navigation layers.
The map can be built without LLM calls and updated as new documents arrive.
Paper: https://arxiv.org/abs/2609.37226
Chat with Paper: https://academy.dair.ai/papers/follow-the-entities-a-corpus-map-for-agentic-search-2609.37226
来源:DAIR.AI · x.com