arXiv:cs.AI· Oliver Jaffe, Dane Sherburn·· 5 小时前AI 评分58
TasteVal 基准:衡量 AI 系统相对人类专家的实验研究品味
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
AI 导读
Oliver Jaffe 和 Dane Sherburn 发布 TasteVal 基准,以算力效率衡量前沿模型的实验研究品味,包含 8 个开放式前沿 AI 研发任务,招募 24 名人类专家并评测 2023 至 2026 年间发布的 20 个模型。
正文
Abstract:We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15-4.37), at roughly 1/30 of our baseliners' average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7-5.0), up from every 14 months between 2023 and December 2025. Measured by final normalized performance, frontier models show no trend break, doubling every 14.6 months. To keep TasteVal uncontaminated, we do not release the tasks.
| Comments: | 38 pages, 21 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06824 [cs.AI] |
| (or arXiv:2610.06824v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06824 arXiv-issued DOI via DataCite |
Submission history
From: Oliver Jaffe [view email]
[v1]
Mon, 5 Oct 2026 17:57:15 UTC (1,084 KB)
[v2]
Tue, 6 Oct 2026 07:00:00 UTC (1,085 KB)
来源:arXiv:cs.AI · arxiv.org