跳到正文
arXiv:cs.LG· Utku \c{S}irin, Cathy Hou, Despina-Ekaterini Argiropoulos, David Alvarez-Melis, Stratos Idreos·· 5 小时前AI 评分43

纹理驱动视觉学习中的低频捷径问题研究

Low-Frequency Shortcuts in Texture-Driven Visual Learning

AI 导读

研究发现纹理驱动领域存在低频捷径:模型主要依赖少数频谱偏斜的低频分量(LFCs)做决策,尽管高频分量(HFCs)预测力更强。从训练和测试集中剪除 LFCs 可缓解捷径,在算法和真实域偏移下 ID 准确率最高提升 10%、OOD 准确率最高提升 40%。通用及领域专用基础模型同样存在低频捷径,大模型虽能缓解但计算成本高,准确率可能低于剪枝后从头训练的小模型。

正文

View PDF HTML (experimental)

Abstract:Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets. Existing studies are all based on a few standard benchmarks, which are shape-driven. Numerous application domains, however, are texture-driven. In this work, we present shortcut learning analysis for texture-driven domains and compare it with that of a standard benchmark. We show that texture-driven domains suffer from low-frequency shortcuts. They make the majority of their decisions based on a few low-frequency components (LFCs) with a skewed spectral behavior, despite that higher-frequency components (HFCs) have higher predictive power. Pruning LFCs from training and test sets mitigates the shortcut and provides a more balanced spectral behavior, improving the ID accuracy by up to 10% and OOD accuracy by up to 40% under algorithmic and real-world domain shifts. We show that general-purpose and domain-specific foundation models can also suffer from low-frequency shortcuts. While large models can mitigate the shortcuts, they incur a high computational cost and may result in a significantly lower accuracy than shortcut-pruned from-scratch trained small models. We show that reduced image resolutions amplify the degree of shortcuts; large frequency-transformation block sizes capture low-frequency shortcuts better than small block sizes; and, low-frequency shortcuts persist across different color spaces. Our findings provide valuable insights, which we hope will be useful for practitioners working on new, understudied domains.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2606.03493 [cs.CV]
  (or arXiv:2606.03493v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2606.03493

arXiv-issued DOI via DataCite

Submission history

From: Utku Sirin [view email]
[v1] Tue, 2 Jun 2026 11:11:20 UTC (6,509 KB)
[v2] Fri, 2 Oct 2026 13:10:26 UTC (7,431 KB)

来源:arXiv:cs.LG · arxiv.org