跳到正文
DAIR.AI· @dair_ai · X·· 3 小时前AI 评分46
AI 导读

微软等机构提出 CorpusMap,预先解析文档集合中反复出现的实体,为每个实体生成链接所有相关文档的页面,原文保持原位。在 7 个模型和 3 个基准上,答案质量提升 6.4 至 11.7 分,输入 token 减少 34% 至 57%,并优于 LLM Wiki 层及另外三种导航层。该地图无需 LLM 调用即可构建,并支持随新文档更新。

正文

Banger paper from Microsoft and colleagues.

If you run agents that search a large document collection, this one is worth your time.

(bookmark it)

They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it.

The original documents stay in place.

The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again.

Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%.

It also beats an LLM Wiki layer and three other navigation layers.

The map can be built without LLM calls and updated as new documents arrive.

Paper: https://arxiv.org/abs/2609.37226

Chat with Paper: https://academy.dair.ai/papers/follow-the-entities-a-corpus-map-for-agentic-search-2609.37226

来源:DAIR.AI · x.com