arXiv:cs.LG· Mengzhe Geng·· 4 小时前AI 评分33
VoxReason:在语音合成前审计有源可依的语音计划
VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis
AI 导读
VoxReason 提出一个 100 案例的验证器基准,通过固定同一句话、只修改一个指定的来源标签线索,来评分引用证据、八个计划字段和允许的回应。在 24 案例的来源键不相交测试中,来源情感先验的计划槽准确率达 0.958,但所需变更准确率为 0.000;在 32 案例的情感不相交测试中,其计划槽准确率为 0.219,所需变更准确率为 1.000。
正文
Abstract:Plan accuracy alone cannot show whether a speech-delivery decision follows its source: a fixed prior may match the original label yet fail to respond appropriately when a cue changes. VoxReason provides a 100-case verifier benchmark that holds each utterance fixed, edits one designated source-label cue, and scores cited evidence, eight plan fields, and the permitted response. On a source-key-disjoint test of 24 cases, a source-emotion prior reaches plan-slot accuracy 0.958, but none of the 24 edited neutral targets appears in its training labels; its required-change accuracy is 0.000. This diagnoses the support boundary of this prior, not its performance on supported edits. In a complementary 32-case emotion-disjoint test, the prior has seen all edited neutral targets but neither original test emotion; its plan-slot accuracy is 0.219 and required-change accuracy is 1.000. The partitions reuse and overlap the same 100 cases, so these deterministic diagnostics are not independent cohorts or learned-planner results. The benchmark evaluates derived labels and structured plans, not audio input, generated speech, or listener judgments.
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2609.03203 [cs.SD] |
| (or arXiv:2609.03203v4 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2609.03203 arXiv-issued DOI via DataCite |
Submission history
From: Mengzhe Geng [view email]
[v1]
Wed, 2 Sep 2026 22:35:02 UTC (174 KB)
[v2]
Mon, 7 Sep 2026 22:11:56 UTC (136 KB)
[v3]
Sun, 20 Sep 2026 14:18:33 UTC (111 KB)
[v4]
Tue, 6 Oct 2026 06:24:04 UTC (111 KB)
来源:arXiv:cs.LG · arxiv.org