arXiv:cs.CL· Nafiseh Ghoroghchian, Luis Scoccola, Tina Sedaghat, Omid Vaheb, Hannah Chen, Dino D'Agostino, Keyvan Golestan·· 3 小时前
表格数据上的对话式任务消歧:泄漏感知建模、基准套件与训练
Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training
AI 导读
研究者提出“模糊可验证任务”框架,将表格数据上的对话式任务消歧分解为提问策略与解题策略,并形式化 oracle 泄漏。基于该框架构建的 AmbiTab 基准套件统一了六个模糊数据集,用强化学习训练的提问策略在全部六个数据集上提升消歧指标、五个数据集上提升任务成功率。
正文
Abstract:Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.
| Comments: | 39 pages (9 main, 30 appendix), 12 figures (5 main, 7 appendix) |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.10740 [cs.LG] |
| (or arXiv:2610.10740v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10740 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nafiseh Ghoroghchian [view email]
[v1]
Wed, 7 Oct 2026 18:09:55 UTC (404 KB)
来源:arXiv:cs.CL · arxiv.org