跳到正文
arXiv:cs.LG· Welmoed R. Eversteijn, Burooj Ghani, A. Leonie Baier, Dan Stowell·· 3 小时前AI 评分30

脉冲间隔变化对深度学习蝙蝠叫声分类的影响研究

Effects of interpulse-interval variation on deep-learning classification of bat vocalizations

AI 导读

研究用欧洲蝙蝠录音构建自然 IPI 与 50-ms 归一化 IPI 两组数据集,微调 EfficientNet-B0 和 PaSST 后发现,IPI 归一化影响因模型而异:PaSST 准确率从 71% 微降至 70%,EfficientNet 则从 47% 升至 57%。

正文

View PDF HTML (experimental)

Abstract:Temporal context may aid automated bat-species classification, but the contribution of specific features remains unclear. We investigated whether variation in the interpulse interval (IPI)-the time between consecutive call onsets-provides species-discriminative information and whether transformer-based models are more sensitive to this information than convolutional neural networks. We created two matched datasets from European bat recordings: a natural-IPI condition retaining the original call timing and a normalized-IPI condition in which call onsets were spaced at 50-ms intervals. EfficientNet-B0 and PaSST were fine-tuned and evaluated within each condition. In an additional experiment, each architecture was trained separately on natural-IPI and normalized-IPI recordings, and evaluated on the same natural-IPI test set. Finally, the pretrained classifiers BatDetect2 and BAT were evaluated on both conditions. Within-condition IPI normalization had model-dependent effects. PaSST accuracy differed little between the natural-IPI ($71 \pm 2.3\%$) and normalized-IPI ($70 \pm 6.3\%$) conditions, whereas EfficientNet accuracy increased from $47 \pm 4.7\%$ to $57 \pm 3.9\%$. PaSST exceeded EfficientNet under both conditions. In the cross-condition evaluation, models trained on natural-IPI recordings outperformed those trained on normalized-IPI recordings on the natural-IPI test set: accuracy decreased from 54% to 50% for EfficientNet and from 65% to 57% for PaSST. BatDetect2 and BAT differed little between IPI conditions. Overall, we found limited support for the hypotheses that natural IPI variation contributes substantially to bat-species classification and that it is used more effectively by transformer-based than CNN-based models. Nevertheless, the cross-condition performance decrease shows that results obtained under normalized conditions may not transfer fully to natural recordings.
Comments: 10 pages, 3 figures
Subjects: Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.02284 [cs.LG]
  (or arXiv:2610.02284v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02284

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Leonie Baier [view email]
[v1] Thu, 1 Oct 2026 14:01:53 UTC (932 KB)

来源:arXiv:cs.LG · arxiv.org