arXiv:cs.AI· Wonbin Son, Gyumum Choi, Junil Seo, Hyungjoon Kim·· 6 小时前AI 评分27
帧级标签能为热成像视频中的小型无人机点检测做什么、不能做什么
What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video
AI 导读
研究分析了仅用帧级目标有无标签训练的小型无人机点检测架构在热红外数据集 CST Anti-UAV 和 Anti-UAV410 上的表现,该架构冻结分类学到的空间特征,用同帧标签训练读出层生成空间得分图与点检测,无需外部检测器。
正文
Abstract:The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07705 [cs.CV] |
| (or arXiv:2610.07705v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07705 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hyungjoon Kim [view email]
[v1]
Tue, 6 Oct 2026 03:58:26 UTC (1,588 KB)
来源:arXiv:cs.AI · arxiv.org