跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Lisan Al Amin, Lei Zhang, Vandana P. Janeja·· 13 小时前AI 评分29

量子核方法在低资源跨语料音频深伪检测中的鲁棒性评估

On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection

AI 导读

研究在仅 200 个训练样本的严格预算下,比较了量子支持向量机(QSVM)、经典 SVM 和 MLP 在跨语料音频深伪检测中的表现,三者均基于冻结的 wav2vec 2.0 嵌入并降至四维特征。

正文

View PDF HTML (experimental)

Abstract:Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited training data. We compare a Quantum Support Vector Machine (QSVM), a classical support vector machine (SVM), and a multilayer perceptron (MLP), all trained on frozen wav2vec 2.0 embeddings using a strict budget of 200 training samples. To match the qubit budget of near-term quantum hardware, embeddings are reduced to four dimensions using principal component analysis, and all models use the same reduced features. Experiments on ASVspoof 2019, ASVspoof 5, the ADD 2023 Challenge, and the In-the-Wild dataset show that under severe domain shift from ASVspoof 2019 to ADD 2023, the MLP degrades to near-random performance, with an area under the curve of approximately 50% and an equal error rate of 50.0%. In contrast, the QSVM maintains meaningful discrimination, achieving an area under the curve of 76.0% and an equal error rate of 27.0%. This advantage is not consistent across transfer directions. When trained on ADD 2023, the QSVM falls below chance on two of three transfers, while the MLP performs better. These results suggest that quantum kernel methods can be competitive under severe cross-corpus shifts and strict low-resource constraints, but do not provide a consistent advantage under near-domain transfer. We interpret these findings as an empirical characterization of quantum kernel inductive bias under distribution shift, rather than evidence of quantum advantage, since the four-qubit kernel can be simulated exactly on classical hardware.
Subjects: Sound (cs.SD); Machine Learning (cs.LG)
Cite as: arXiv:2610.00649 [cs.SD]
  (or arXiv:2610.00649v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.00649

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lisan Al Amin [view email]
[v1] Wed, 30 Sep 2026 19:51:25 UTC (5,073 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org