arXiv:cs.CL· Claudiu Creanga, Liviu Dinu·· 4 小时前AI 评分39
让 LLM 文本"更像人"反而更易被检测:RoBERTa 检测器的本体不稳定性与统计放大效应
Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text
AI 导读
研究分析 RoBERTa 系 AI 文本检测器在语义、结构和 tokenizer 层面扰动下的表现,发现让 Mistral-7B-Instruct 把机器文本改得更像人时,动词多样性从 0.77 升至 0.92,输出反而更易被检测。
正文
Abstract:Supervised AI-text detectors report high benchmark accuracy, but it is not clear what their decisions are based on. We analyze a RoBERTa-based detector under semantic, structural, and tokenizer-level perturbations, using the M4 dataset (N = 10,000) and controlled generations (N = 300). When Mistral-7B-Instruct was asked to make machine text sound more human, Verb Diversity rose from 0.77 to 0.92 and the outputs became easier to detect. Detection scores appear to track statistical complexity, which also leads to a 76.3% false-positive rate on formal human writing. As a control, we evaluate event-based Latent Space detection. Paraphrasing changed 87% of its event sequences (Jaccard = 0.067), and homoglyphs altered 70% of the extracted verbs even though extraction still ran (Jaccard = 0.30). Its best domain AUC was 0.577. RoBERTa's robustness seems specific to the features it uses, and structural abstraction did not make detection more robust.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03110 [cs.CL] |
| (or arXiv:2610.03110v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03110 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Claudiu Creanga [view email]
[v1]
Fri, 2 Oct 2026 10:28:25 UTC (703 KB)
来源:arXiv:cs.CL · arxiv.org