BEANS-Next 与 ROOTS:拓展生物声学中的音频-语言能力
BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics
研究者提出 BEANS-Next 基准与 ROOTS 训练资源,以拓展音频-语言模型在生物声学中的能力。BEANS-Next 覆盖声学感知、生物类别识别、场景理解和上下文学习四类任务,现有模型在既有评测重点之外表现有限。ROOTS 整合真实数据、行为与声学元数据及合成生成数据,训练后在各任务组均取得显著提升,基准、数据集与数据管线已全部开源。
Authors:Christos Plachouras, David Robinson, Marius Miron, Gagan Narula, Paul Laisné, Anthony L. T. Fine, Benno Weck, Ellen Gilsenan-McMahon, Diane Kim, Laura Hay Mack, Maddie Cusimano, Sara Keen, Lukas Rauch, Benjamin Hoffman, Emmanuel Chemla, Emmanouil Benetos, Johan Pauwels, Milad Alizadeh, Matthieu Geist, Olivier Pietquin
Abstract:Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models. However, progress in the field has so far been assessed on a narrow set of tasks, primarily centered on label-centric biological category recognition, such as species and call-type classification. In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning. Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications. To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient. We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology. To accelerate progress in the field, we open-source our benchmark, dataset, and data pipelines.
| Comments: | 36 pages, 7 figures. Project page: this https URL |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10663 [cs.SD] |
| (or arXiv:2610.10663v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10663 arXiv-issued DOI via DataCite |
Submission history
From: Christos Plachouras [view email]
[v1]
Wed, 7 Oct 2026 17:41:45 UTC (2,594 KB)
来源:arXiv:cs.LG · arxiv.org