arXiv:cs.AI· Hamed Khosravi, Xiaoming Huo·· 3 小时前
面向自适应操纵的随时有效排行榜认证腐败预算
Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
AI 导读
研究者提出"认证腐败预算":在发布 t 条记录后为每个两两对比结论附带容忍度 B̂_t,若结论错误则说明被篡改记录超过该值。伪造记录与事后篡改需要不同证书,前者在攻击者翻票时以趋近于 1 的概率失效。在 180 万条 Chatbot Arena 投票回放中,数百条操纵票即可让标准置信区间给出错误排序,而该方法仍保持有效。
正文
Abstract:Public leaderboards for AI models are read continuously, and attackers can see every published standing. Vote rigging, selective disclosure of private variants, and benchmark contamination can each move a ranking. Existing guarantees assume genuine records or bound the corruption per step, which an attacker who corrupts in bursts evades. We introduce the certified corruption budget, a tolerance $\widehat{B}_t$ computed after $t$ records and published with each pairwise claim. With probability at least $1-\alpha$, simultaneously at all times, the claim is correct or more than $\widehat{B}_t$ records were corrupted. It holds against attackers who watch every certificate, with no bound on their budget. Forged records and records altered once seen require different certificates: the certificate for forgeries fails, with probability approaching one, against an attacker who flips votes it has seen, while one that charges roughly twice as much per record remains valid, with constant bets even against attackers who see the future, and no smaller charge is valid at every level. The certified budget grows nearly as fast as any valid method allows: with a win fraction $\frac{1}{2}+\delta$, each new record adds close to $2\delta$ to the number of forged records the claim can withstand ($\delta$ flipped). Publishing the best of $V$ private variants costs only an amount growing like $\log V$. In replays on 1.8 million Chatbot Arena votes, a few hundred rigged votes make standard confidence intervals certify false orderings, while ours stays valid. On real votes, our certificate shows that clearly separated models withstand about 2,000 forged votes.
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Methodology (stat.ME) |
| Cite as: | arXiv:2610.10597 [cs.CR] |
| (or arXiv:2610.10597v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10597 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hamed Khosravi [view email]
[v1]
Tue, 6 Oct 2026 21:20:14 UTC (435 KB)
来源:arXiv:cs.AI · arxiv.org