arXiv:cs.CL· Mikhail L. Arbuzov (Independent researcher), Karan Dave (Independent researcher), Evgeniya Dontsova (Independent researcher), Yaodong Hu (Independent researcher), Vincent Lao (Independent researcher), Navita Jain (Independent researcher), Sisong Bei (Independent researcher), Dmitry Dimov (Independent researcher)·· 3 小时前
先澄清、再聚焦:面向大规模对话分析的陈述归一化方法
Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
AI 导读
研究提出"陈述归一化"方法,将对话转化为带说话人归属、来源引用和语义标签的短陈述,使含义更明确并支持按问题筛选证据。在客服通话的优惠抑制任务中,归一化提升了无筛选的监督分类器效果,而较弱的提示词读者同时受益于归一化与筛选。该流程可完全由小模型搭建,小模型学习归一化契约,轻量编码器负责打标签与下游决策,从而大幅降低百万级对话的分析成本。
正文
Abstract:Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.
| Comments: | 20 pages, 1 figure |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.10758 [cs.CL] |
| (or arXiv:2610.10758v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10758 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mikhail Arbuzov [view email]
[v1]
Wed, 7 Oct 2026 18:22:55 UTC (161 KB)
来源:arXiv:cs.CL · arxiv.org