Lin Zhang 0054

dblp:37/1629-54 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-7826-2850ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HPQ: A Hybrid Framework for Joint Pruning and Quantization of Self-Supervised Speech Models
Junyi Peng, Lin Zhang 0054, Jiangyu Han, Oldrich Plchot, Shuai Wang 0016, Jan Cernocký
IEEE Signal Process. Lett.2
2025 Towards Generalized Source Tracing for Codec-Based Deepfake Speech
abstract
Recent attempts at source tracing for codecbased deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-theart performance on the CoSG test set of CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing.
I-Ming Lin, Xuanjun Chen, Lin Zhang 0054, Hung-yi Lee, Jyh-Shing Roger Jang
ASRU3
2025 Continual Unsupervised Domain Adaptation for Audio Deepfake Detection
abstract
Audio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models evolve, existing UDA methods struggle with catastrophic forgetting when facing continuously emerging spoofing methods. To address this challenge, we introduce continual UDA for ADD, which involves sequentially training across multiple target domains with continual learning. We propose a causality-distillation-based continual domain adversarial training framework for continual UDA, called CD-DAT. Specifically, we employ the domain adversarial training (DAT) framework to learn both spoofing-discriminative and domain-invariant deep features. In addition, we design a continual learning algorithm utilizing causality distillation to capture the mapping between utterances and classes, effectively mitigating forgetting and maintaining generalization. Experiments demonstrated that CD-DAT improved detection performance across all domains, confirming its memory stability and learning plasticity.
Xiaohuan Chen, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Lin Zhang 0054, Jianguo Wei
ICASSP6
2025 LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
abstract
Previous fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining valuable information for both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that the current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.49% equal error rate.
Hieu-Thi Luong, Haoyang Li 0018, Lin Zhang 0054, Kong-Aik Lee, Chng Eng Siong
ICASSP3
2025 CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
abstract
Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42%, 0.48%, and 0.96% on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.1
Junyi Peng, Ladislav Mosner, Lin Zhang 0054, Oldrich Plchot, Themos Stafylakis, Lukás Burget, Jan Cernocký
ICASSP3
2025 Analysis of ABC Frontend Audio Systems for the NIST-SRE24
abstract
Section: Speaker Recognition
Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang 0054, Jiangyu Han, Petr Pálka, Federico Landini, Lukás Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavácek, Martin Kodovsky, Tomás Pavlícek
INTERSPEECH7
2025 Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
Xuanjun Chen, I-Ming Lin, Lin Zhang 0054, Jiawei Du 0003, Hung-yi Lee, Jyh-Shing Roger Jang
INTERSPEECH3
2025 PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
You Zhang 0001, Baotong Tian, Lin Zhang 0054, Zhiyao Duan
INTERSPEECH3
2025 Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
Xinlei Ma, Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Lin Zhang 0054, Wenhuan Lu
Speech Commun.6
2025 SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-Spoofing
abstract
Audio anti-spoofing algorithms struggle with fake samples from unseen spoofing techniques, even when trained with diverse data sets or data augmentation strategies. Unsupervised domain adaptation (UDA) algorithms have the potential to mitigate this challenge. Typically, UDA assumes that the source and target domains are distinct distributions with clear boundaries and seeks to align model representations between them. However, in anti-spoofing, various spoofing algorithms could cause the distributions of the generated samples to overlap, resulting in unclear domain boundaries. This hinders UDA algorithms from effectively measuring and aligning domain discrepancies. Moreover, forcibly aligning samples with significant discrepancies could diminish the model’s discriminative capability. To solve this problem, we propose a domain attention algorithm with optimal transport (OT), termed Sinkhorn Domain Attention (SHDA). Unlike traditional attention mechanisms, SHDA identifies the optimal transfer plan by analyzing the global probability differences among cross-domain samples. Specifically, we first extract audio representations from various domains to compute the overall cost matrix between the source and target domains. Next, we employ Sinkhorn’s iteration to calculate the OT coupling matrix, where cross-domain samples with minor differences receive higher transfer weights, while those with substantial differences receive lower weights. Finally, we use the coupling and cost matrices to compute the adaptation loss, effectively transferring the anti-spoofing model from multiple sources to the target domain. We conducted eight cross-domain experiments using eleven well-known anti-spoofing corpora. The results indicate that our label-free SHDA surpassed the state-of-the-art model by 40%.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Lin Zhang 0054, Di Jin 0001, Junhai Xu, Wenhuan Lu
IEEE Trans. Inf. Forensics Secur.4
2024 CPAUG: Refining Copy-Paste Augmentation for Speech Anti-Spoofing
abstract
Conventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmentation for speech anti-spoofing, dubbed CpAug, to generate more training data with rich intra-class diversity. The CpAug employs two policies: concatenation to merge utterances with identical labels, and substitution to replace segments in an anchor utterance. Besides, considering the impacts of speakers and spoofing attack types, we craft four blending strategies for the CpAug. Furthermore, we explore how CpAug complements the Rawboost augmentation method. Experimental results reveal that the proposed CpAug significantly improves the performance of speech anti-spoofing. Particularly, CpAug with substitution policy leads to relative improvements of 43% and 38% on the ASVspoof’ 19LA and 21LA, respectively. Notably, the CpAug and Rawboost synergize effectively, achieving an EER of 2.91% on ASVspoof’ 21LA.
Linjuan Zhang, Kong-Aik Lee, Lin Zhang 0054, Longbiao Wang, Baoning Niu
ICASSP3
2024 How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
Tianchi Liu 0004, Lin Zhang 0054, Rohan Kumar Das, Ruijie Tao, Haizhou Li 0001
INTERSPEECH2
2024 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Mireia Díez, Federico Landini, Nicholas W. D. Evans, Junichi Yamagishi
INTERSPEECH1
2024 Unsupervised Adaptive Speaker Recognition by Coupling-Regularized Optimal Transport
abstract
Cross-domain speaker recognition (SR) can be improved by unsupervised domain adaptation (UDA) algorithms. UDA algorithms often reduce domain mismatch at the cost of decreasing the discrimination of speaker features. In contrast, optimal transport (OT) has the potential to achieve domain alignment while preserving the speaker discrimination capability in UDA applications; however, naively applying OT to measure global probability distribution discrepancies between the source and target domains may induce negative transports where samples belonging to different speakers are coupled in transportation. These negative transports reduce the SR model's discriminative power, degrading the SR performance. This paper proposes a coupling-regularized optimal transport (CROT) algorithm for cross-domain SR to reduce the negative transport during UDA. In the proposed CROT, two consecutive processing modules regularize the coupling paths for the OT solution: a progressive inter-speaker constraint (PISC) module and a coupling-smoothed regularization (CSR) module. The PISC, designed as a pseudo-label memory bank with curriculum learning, is first applied to select valid samples to guarantee that coupling samples are from the same speaker. The CSR, designed to control the information entropy of the coupling paths further, reduces the effect of negative transport in UDA. To evaluate the effectiveness of the proposed algorithm, cross-domain SR experiments were conducted under different target domains, speaker encoders, corpora, and acoustic features. Experimental results showed that CROT achieved a 50% relative reduction in equal error rates compared to conventional OT-based UDAs, outperforming the state-of-the-art UDAs.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification
abstract
Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has difficulty computing effective transports. To address this challenge, we propose an OT-based unsupervised domain adaptation (UDA) framework for SV, OT with a diversified memory bank, called DMB-OT, which ensures the accuracy of transfers by two strategies: (1) It regularizes the solution space of OT, which attempts to plan transformations between audio samples from the same speaker with high confidence; (2) it integrates a dynamic curriculum learning algorithm, preventing OT from calculating transport couplings based on hard-discriminative samples in the early stage of UDA. Experiments under different target domains showed that our unsupervised DMB-OT could significantly improve the performance of OT-based UDA and could even match the performance of the supervised PLDA-based adaptation.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu
ICASSP6
2023 Range-Based Equal Error Rate for Spoof Localization
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi
INTERSPEECH1
2023 TMS: Temporal multi-scale in time-delay neural network for speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu, Jianwu Dang 0001
Appl. Intell.6
2023 Self-supervised learning based domain regularization for mask-wearing speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Yantao Ji, Junhai Xu
Speech Commun.6
2023 The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance
abstract
Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial Spoof" (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively.
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization
abstract
Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-sequential re-parameterization (CS-Rep), a novel topology re-parameterization strategy for multi-type networks, to increase the inference speed and verification accuracy of models. CS-Rep solves the problem that existing re-parameterization methods are not suitable for typical ASV backbones. When a model applies CS-Rep, the training-period network utilizes a multi-branch topology to capture speaker information, whereas the inference-period model converts to a time-delay neural network (TDNN)-like plain backbone with stacked TDNN layers to achieve the fast inference speed. Based on CS-Rep, an improved TDNN with friendly test and deployment called Rep-TDNN is proposed. Compared with the state-of-the-art model ECAPA-TDNN, Rep-TDNN increases the actual inference speed by about 50% and reduces the EER by 10%. The code and trained models are available at https://github.com/zrtlemontree/CS-Rep.
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Lin Zhang 0054, Yantao Ji, Junhai Xu, Xugang Lu
ICASSP4
2022 Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio Detection
abstract
Fake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%.
Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki
INTERSPEECH6
2022 Spoofing-Aware Attention based ASV Back-end with Multiple Enrollment Utterances and a Sampling Strategy for the SASV Challenge 2022
abstract
Current state-of-the-art automatic speaker verification (ASV) systems are vulnerable to presentation attacks, and several countermeasures (CMs), which distinguish bona fide trials from spoofing ones, have been explored to protect ASV. However, ASV systems and CMs are generally developed and optimized independently without considering their inter-relationship. In this paper, we propose a new spoofing-aware ASV back-end module that efficiently computes a combined ASV score based on speaker similarity and CM score. In addition to the learnable fusion function of the two scores, the proposed back-end module has two types of attention components, scaled-dot and feed-forward self-attention, so that intra-relationship information of multiple enrollment utterances can also be learned at the same time. Moreover, a new effective trials-sampling strategy is designed for simulating new spoofing-aware verification scenarios introduced in the Spoof-Aware Speaker Verification (SASV) challenge 2022.
Chang Zeng, Lin Zhang 0054, Meng Liu 0017, Junichi Yamagishi
INTERSPEECH2
2021 An Initial Investigation for Detecting Partially Spoofed Audio
abstract
International audience
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Jose Patino 0001, Nicholas W. D. Evans
Interspeech1
2020 Regional Resonance of the Lower Vocal Tract and its Contribution to Speaker Characteristics
abstract
S.1391-1395
Lin Zhang 0054, Kiyoshi Honda, Jianguo Wei, Seiji Adachi
INTERSPEECH1
2020 ARET: Aggregated Residual Extended Time-Delay Neural Networks for Speaker Verification
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Longbiao Wang, Meng Liu 0017, Lin Zhang 0054, Jiayu Jin, Junhai Xu
INTERSPEECH6