arXiv:cs.AI· Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling·· 5 小时前AI 评分44
TasteBench:从分子到可持续食品的多模态感官预测基准
TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods
AI 导读
研究团队推出 TasteBench,一个面向感官预测的多模态基准与隐私保护竞赛,包含基于 215 种植物基食品、21K+ 人类评估构建的食品级排序任务(935 个同类别排序对),以及覆盖 15K 风味分子的分子级味觉分类任务。
正文
Abstract:Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials discovery. Evaluating whether a novel food tastes like its animal-based target requires expensive human sensory panels, bottlenecking the design-build-test loop. We introduce TasteBench, a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning two tasks: a food-level ranking task built on 21K+ human evaluations across 215 plant-based foods in 24 product categories, yielding 935 within-category ranking pairs, and a supporting molecular-level taste classification task over 15K flavor molecules. To enable rigorous interpretation of model performance, we characterize the ground truth: inter-rater agreement among panelists is low (Krippendorff's $\alpha = .077$), and the split-half reliability ceiling of panel-aggregated rankings is .825, establishing the range within which ML systems on this benchmark should be assessed. We evaluate baselines across four input modalities; on the same pairs panelists rated, the best model achieves .661 pairwise accuracy, competitive with the median individual panelist (.650), and .683 across all within-category pairs. TasteBench provides the evaluation infrastructure and baselines for measuring progress on computational screening for sustainable protein discovery.
| Comments: | First two authors contributed equally. Accepted to NeurIPS 2026, Evaluations & Datasets track. Code available at this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02599 [cs.AI] |
| (or arXiv:2610.02599v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02599 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Anna Thomas [view email]
[v1]
Thu, 1 Oct 2026 23:52:39 UTC (2,768 KB)
来源:arXiv:cs.AI · arxiv.org