跳到正文
arXiv:cs.AI· Viraj Bagal, Raviraja Ganta, Prabhath Chellingi·· 6 小时前AI 评分41

同一反馈,不同答案:衡量前沿模型客户反馈分析中的运行间不稳定性

Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis

AI 导读

研究提出重复运行评估框架,用主题流失率和数量分歧两项指标衡量 AI 智能体在客户反馈分析中的稳定性。在 8 个前沿模型、100 至 5000 条语料上测试三种执行方案,固定 Claude Opus 4.8 和 1000 条语料时,基于分类体系的智能体(TGA)相比原始生成和分层分解将主题流失率降低 86–88%,匹配主题的数量分歧为零。

正文

View PDF HTML (experimental)

Abstract:AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data. Such automation requires repeatability: when the underlying evidence is unchanged, the agent's categories, priorities, and counts should not shift materially between runs, even if each individual answer appears plausible. We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist. We evaluate three recurring customer-feedback tasks across eight frontier models, corpus sizes from 100 to 5,000 records, multiple prompts, and three execution designs: raw generation, taxonomy-free hierarchical decomposition, and a taxonomy-grounded agent (TGA) using persistent themes, subthemes, and record-level predictions. With Claude Opus 4.8 and the 1,000-record corpus fixed, TGA reduces theme churn by 86--88% relative to both raw generation and hierarchical decomposition, while matched-theme volumes have zero disagreement. The taxonomy-grounded agent is more stable than every raw model in the screen, remains more stable at each corpus size, and keeps this advantage when theme matching is made stricter or looser. Although evaluated on customer feedback, the framework targets repeated synthesis of unstructured corpora more broadly, including financial reports, legal documents, incident records, and scientific literature. Overall, these results show that taxonomy grounding produces more consistent and repeatable outputs for recurring knowledge work.
Comments: 9 pages
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08036 [cs.AI]
  (or arXiv:2610.08036v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08036

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Viraj Bagal [view email]
[v1] Tue, 6 Oct 2026 09:29:25 UTC (35 KB)

来源:arXiv:cs.AI · arxiv.org