arXiv:cs.CL· Theodore O. Cochran·· 3 小时前AI 评分43
LLM 维护 Wiki 知识库的渐进式披露:一项预注册消融研究
Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation
AI 导读
一项预注册消融研究在 LLM 维护的 709 页 markdown 知识库上测试渐进式披露,结果发现能力强的 Agent 根本不会加载大型索引,而是直接从问题推断页面位置并读取,预期中的节省并不存在,于是研究将答案质量设为主要指标,质量得以保持。
正文
Abstract:LLM agents now often answer questions from knowledge bases they help maintain. A common intuition says progressive disclosure should make this cheaper. Instead of loading one large index, the agent reads a compact catalog and one-line page summaries, then opens only the pages it needs. We tested that intuition in a preregistered study on a real 709-page markdown knowledge base maintained by an LLM. We retrofitted it for progressive disclosure and built four versions that differ only in how the agent reaches the pages. The pages themselves are identical in every version, so any difference comes from the access structure alone. Each version was tested three ways, with the agent following a set protocol, choosing its own path, or made to load the catalog first. A judge from a different model family graded the answers blind against verified reference answers.
A preparatory pilot changed the question. A capable agent never loaded the large index at all. It worked out from the question where a page was and read it directly. The saving we set out to measure did not exist for such an agent, so we made answer quality the primary outcome. Quality held. Answers from the retrofitted knowledge base were as good as answers from the original, within a margin we set in advance. Two limits apply. Our human rater and the model judge agreed far less than the plan required, so the quality result rests on the judge, backed by sensitivity checks. Quality was also not shown to hold when the agent was forced to load the catalog first, or on the two most reliably graded criteria under a stricter test. Cost fell clearly in every condition we tested, and the retrofitted version cited fewer pages and took fewer tool turns per answer.
| Comments: | 15 pages, 3 figures, 6 tables. v2 states its two limits in the abstract. Our human rater and the model judge agreed far less than the plan required. Quality was not shown to hold when the agent had to load the catalog first, or on the two most reliably graded criteria. Preregistered on OSF at this https URL, DOI https://doi.org/10.17605/OSF.IO/FEKA7 |
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2607.04576 [cs.CL] |
| (or arXiv:2607.04576v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.04576 arXiv-issued DOI via DataCite |
Submission history
From: Theodore Cochran [view email]
[v1]
Mon, 6 Jul 2026 01:03:27 UTC (59 KB)
[v2]
Wed, 7 Oct 2026 00:50:51 UTC (121 KB)
来源:arXiv:cs.CL · arxiv.org