跳到正文
arXiv:cs.LG· Kwanghee Choi, Suwon Shon, Dmitriy Serdyuk, Guitang Lan, Chao-Wei Huang, Mohammad Sadegh Rasooli, Sangeeta Srivastava, Zhaojiang Lin, Saurabh Adya, Ming Sun·· 4 小时前AI 评分46

Logbook:小时级长音频事件理解基准

Logbook: Extremely Long-form Audio Event Understanding

AI 导读

研究者推出 Logbook 基准,用于小时级音频理解,录音时长从十分钟到六天。系统需在给定连续音频和事件标签词表下,预测无间隙分段,并为每段标注事件标签与描述。研究对比 52 个端到端与级联系统,发现任务可解但最佳系统仍低于人类参考,过分割普遍,微调可部分缓解,端到端系统通常优于级联,但上下文越长性能越差。

正文

View PDF HTML (experimental)

Abstract:Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.
Comments: Submitted to ICASSP 2027. Source code available at this https URL
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
Cite as: arXiv:2610.07338 [eess.AS]
  (or arXiv:2610.07338v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2610.07338

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kwanghee Choi [view email]
[v1] Mon, 5 Oct 2026 20:13:46 UTC (128 KB)

来源:arXiv:cs.LG · arxiv.org