arXiv:cs.LG(机器学习,全量分类)· Esther Xin·· 14 小时前AI 评分40
合成 RLVR 语料中的数据痕迹是否被利用:对 GooseReason-0.7M 的因果审计
When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
AI 导读
研究对 GooseReason-0.7M 的 RLVR 合成语料做因果审计,检验策略是否学到"正确选项是人类原文、干扰项是合成"这一来源痕迹而非正确性本身。仅用五项表层统计的分类器在 315,499 个选项上 AUROC 仅 0.562,但代码域低至 0.416,原因是代码干扰项由金标准答案单算子变异而来。
正文
Abstract:Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principle learn the second instead of the first. We audit that possibility in GooseReason-0.7M. First we ask whether the asymmetry is visible at all: a classifier reading only five surface statistics (never the meaning) reaches AUROC 0.562 over 315,499 options, barely above chance. The aggregate hides something, though. Code sits at 0.416, below chance, and manual inspection explains why: code distractors turn out to be single-operator mutations of the gold answer rather than freely written alternatives, so the two classes are nearly identical by construction. Detecting a signal is not the same as showing a model uses it, so we then run an intervention. We build a paraphrase-matched control corpus, hold training-set size identical across arms, and train two policies under one fixed budget. The exploitation gap does not favour the unmodified-data arm: 0.021 against 0.027 for the control. Under our budget, in other words, a detectable artifact went unexploited. We think that dissociation, along with the domain-specific construction finding, is worth knowing for anyone curating corpora of this kind, and we release the audit as a mostly CPU-only protocol.
| Comments: | 9 pages, 2 figures,4 tables;Code and data this https URL |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7; I.2.6 |
| Cite as: | arXiv:2610.00202 [cs.CL] |
| (or arXiv:2610.00202v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00202 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Esther Xin [view email]
[v1]
Sun, 20 Sep 2026 11:31:26 UTC (41 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org