arXiv:cs.CL· Xi Xuan, Davide Carbone, Wenxin Zhang, Tomi H. Kinnunen·· 4 小时前AI 评分37
WaveScat:结合小波散射前端与自监督特征的语音深度伪造检测
WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
AI 导读
WaveScat 提出一类结合小波散射变换(WST)的特征提取器,将手工滤波器组特征的可解释性与 SSL 特征的高层信息结合,用于语音深度伪造检测。在 Deepfake-Eval-2024 基准及 SpoofCeleb、In-the-Wild、ASVspoof 5 跨数据集评测中,WaveScat 大幅优于现有前端。分析表明,小平均尺度配合高频与方向分辨率对捕捉细微伪造痕迹至关重要。
正文
Abstract:Existing front-ends for speech deepfake detection are primarily categorized into two types. Hand-crafted filterbank features are transparent but limited in capturing higher-level information. SSL features, in turn, lack interpretability and may overlook fine-grained spectral anomalies. We propose WaveScat, a novel family of feature extractors that combines the best of both worlds via the wavelet scattering transform (WST), which cascades wavelet convolutions with modulus nonlinearities to produce deformation-stable, multi-scale features. Experiments on the recent Deepfake-Eval-2024 benchmark, together with cross-dataset evaluations on SpoofCeleb, In-the-Wild, and ASVspoof 5, show that WaveScat outperforms existing front-ends by a wide margin. Our analysis reveals that a small averaging scale combined with high-frequency and directional resolutions is critical for capturing subtle artifacts. This underscores the value of stable and translation-invariant features for speech deepfake detection. The code and supplementary materials are available at this https URL.
| Comments: | Submitted to ICASSP 2027 |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Signal Processing (eess.SP) |
| Cite as: | arXiv:2602.02980 [eess.AS] |
| (or arXiv:2602.02980v3 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2602.02980 arXiv-issued DOI via DataCite |
Submission history
From: Xi Xuan [view email]
[v1]
Tue, 3 Feb 2026 01:39:28 UTC (2,690 KB)
[v2]
Thu, 30 Apr 2026 13:42:03 UTC (2,692 KB)
[v3]
Wed, 7 Oct 2026 11:39:47 UTC (2,277 KB)
来源:arXiv:cs.CL · arxiv.org