Task-Optimized Blind Channel Magnitude Response Estimation with Variational Autoencoders for Replay Attack Detection


BEKİRYAZICI Ş., Hanilci C., ÖZCAN SEMERCİ N.

IEEE Access, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1109/access.2026.3708257
  • Dergi Adı: IEEE Access
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
  • Anahtar Kelimeler: ASVspoof, Blind Channel Estimation, Replay Attack Detection, ReplayDF, Spoofing Countermeasures, Variational Autoencoders
  • Bursa Uludağ Üniversitesi Adresli: Evet

Özet

Replay attacks pose a serious threat to automatic speaker verification (ASV) systems due to distortions introduced by playback and recording channels. Conventional countermeasures typically rely on time–frequency representations of the observed signal, which may capture replay artifacts only implicitly and remain sensitive to dataset-specific biases. In this study, we propose a channel-focused neural framework for blind estimation of replay-induced channel log-magnitude responses. A variational autoencoder (VAE) is trained to reconstruct a clean time–frequency representation from an observed one, allowing replay-related channel characteristics to be extracted via the residual between observation and reconstruction. The resulting VAE-based channel magnitude representation (VAE-CMR) is used as input to replay countermeasure models. Extensive experiments on ASVspoof 2019, ASVspoof 2021, and ReplayDF using Res2Net50 and SE-Res2Net50 back-ends demonstrate that VAE-CMR consistently outperforms conventional acoustic features (spectrogram, Mel-spectrogram, LFCC) and self-supervised representations (WavLM, Wav2Vec 2.0). Joint optimization of the VAE and classifier further aligns channel estimation with the replay detection objective and yields substantial and consistent reductions in equal error rate (EER) and minimum tandem detection (t-DCF) cost function, including under cross-dataset conditions. For instance, while the best-performing spectrogram baseline achieves EERs of 2.53%, 35.06%, and 46.38% on ASVspoof 2019, ASVspoof 2021, and ReplayDF, respectively, the proposed method reduces these to 1.18%, 17.95%, and 26.34%. Additional analyses confirm robustness across operating points and reduced reliance on silence-related artifacts. The results establish blind neural channel magnitude estimation as an effective and robust approach for replay attack detection.