EDBT 2026 Demo / reviewers in the wild / expert
Lin Zhang 0054
dblp:37/1629-54
· DBLP profile ↗
25ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-7826-2850ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HPQ: A Hybrid Framework for Joint Pruning and Quantization of Self-Supervised Speech Models
Junyi Peng, Lin Zhang 0054, Jiangyu Han, Oldrich Plchot, Shuai Wang 0016, Jan Cernocký |
IEEE Signal Process. Lett. | 2 |
| 2025 | Towards Generalized Source Tracing for Codec-Based Deepfake SpeechabstractRecent attempts at source tracing for codecbased deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-theart performance on the CoSG test set of CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing. I-Ming Lin, Xuanjun Chen, Lin Zhang 0054, Hung-yi Lee, Jyh-Shing Roger Jang |
ASRU | 3 |
| 2025 | Continual Unsupervised Domain Adaptation for Audio Deepfake DetectionabstractAudio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models evolve, existing UDA methods struggle with catastrophic forgetting when facing continuously emerging spoofing methods. To address this challenge, we introduce continual UDA for ADD, which involves sequentially training across multiple target domains with continual learning. We propose a causality-distillation-based continual domain adversarial training framework for continual UDA, called CD-DAT. Specifically, we employ the domain adversarial training (DAT) framework to learn both spoofing-discriminative and domain-invariant deep features. In addition, we design a continual learning algorithm utilizing causality distillation to capture the mapping between utterances and classes, effectively mitigating forgetting and maintaining generalization. Experiments demonstrated that CD-DAT improved detection performance across all domains, confirming its memory stability and learning plasticity. Xiaohuan Chen, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Lin Zhang 0054, Jianguo Wei |
ICASSP | 6 |
| 2025 | LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation GenerationabstractPrevious fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining valuable information for both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that the current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.49% equal error rate. Hieu-Thi Luong, Haoyang Li 0018, Lin Zhang 0054, Kong-Aik Lee, Chng Eng Siong |
ICASSP | 3 |
| 2025 | CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker VerificationabstractSelf-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42%, 0.48%, and 0.96% on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.1 Junyi Peng, Ladislav Mosner, Lin Zhang 0054, Oldrich Plchot, Themos Stafylakis, Lukás Burget, Jan Cernocký |
ICASSP | 3 |
| 2025 | Analysis of ABC Frontend Audio Systems for the NIST-SRE24abstractSection: Speaker Recognition Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang 0054, Jiangyu Han, Petr Pálka, Federico Landini, Lukás Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavácek, Martin Kodovsky, Tomás Pavlícek |
INTERSPEECH | 7 |
| 2025 | Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
Xuanjun Chen, I-Ming Lin, Lin Zhang 0054, Jiawei Du 0003, Hung-yi Lee, Jyh-Shing Roger Jang |
INTERSPEECH | 3 |
| 2025 | PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
You Zhang 0001, Baotong Tian, Lin Zhang 0054, Zhiyao Duan |
INTERSPEECH | 3 |
| 2025 | Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
Xinlei Ma, Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Lin Zhang 0054, Wenhuan Lu |
Speech Commun. | 6 |
| 2025 | SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-SpoofingabstractAudio anti-spoofing algorithms struggle with fake samples from unseen spoofing techniques, even when trained with diverse data sets or data augmentation strategies. Unsupervised domain adaptation (UDA) algorithms have the potential to mitigate this challenge. Typically, UDA assumes that the source and target domains are distinct distributions with clear boundaries and seeks to align model representations between them. However, in anti-spoofing, various spoofing algorithms could cause the distributions of the generated samples to overlap, resulting in unclear domain boundaries. This hinders UDA algorithms from effectively measuring and aligning domain discrepancies. Moreover, forcibly aligning samples with significant discrepancies could diminish the model’s discriminative capability. To solve this problem, we propose a domain attention algorithm with optimal transport (OT), termed Sinkhorn Domain Attention (SHDA). Unlike traditional attention mechanisms, SHDA identifies the optimal transfer plan by analyzing the global probability differences among cross-domain samples. Specifically, we first extract audio representations from various domains to compute the overall cost matrix between the source and target domains. Next, we employ Sinkhorn’s iteration to calculate the OT coupling matrix, where cross-domain samples with minor differences receive higher transfer weights, while those with substantial differences receive lower weights. Finally, we use the coupling and cost matrices to compute the adaptation loss, effectively transferring the anti-spoofing model from multiple sources to the target domain. We conducted eight cross-domain experiments using eleven well-known anti-spoofing corpora. The results indicate that our label-free SHDA surpassed the state-of-the-art model by 40%. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Lin Zhang 0054, Di Jin 0001, Junhai Xu, Wenhuan Lu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | CPAUG: Refining Copy-Paste Augmentation for Speech Anti-SpoofingabstractConventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmentation for speech anti-spoofing, dubbed CpAug, to generate more training data with rich intra-class diversity. The CpAug employs two policies: concatenation to merge utterances with identical labels, and substitution to replace segments in an anchor utterance. Besides, considering the impacts of speakers and spoofing attack types, we craft four blending strategies for the CpAug. Furthermore, we explore how CpAug complements the Rawboost augmentation method. Experimental results reveal that the proposed CpAug significantly improves the performance of speech anti-spoofing. Particularly, CpAug with substitution policy leads to relative improvements of 43% and 38% on the ASVspoof’ 19LA and 21LA, respectively. Notably, the CpAug and Rawboost synergize effectively, achieving an EER of 2.91% on ASVspoof’ 21LA. Linjuan Zhang, Kong-Aik Lee, Lin Zhang 0054, Longbiao Wang, Baoning Niu |
ICASSP | 3 |
| 2024 | How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
Tianchi Liu 0004, Lin Zhang 0054, Rohan Kumar Das, Ruijie Tao, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2024 | Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Mireia Díez, Federico Landini, Nicholas W. D. Evans, Junichi Yamagishi |
INTERSPEECH | 1 |
| 2024 | Unsupervised Adaptive Speaker Recognition by Coupling-Regularized Optimal TransportabstractCross-domain speaker recognition (SR) can be improved by unsupervised domain adaptation (UDA) algorithms. UDA algorithms often reduce domain mismatch at the cost of decreasing the discrimination of speaker features. In contrast, optimal transport (OT) has the potential to achieve domain alignment while preserving the speaker discrimination capability in UDA applications; however, naively applying OT to measure global probability distribution discrepancies between the source and target domains may induce negative transports where samples belonging to different speakers are coupled in transportation. These negative transports reduce the SR model's discriminative power, degrading the SR performance. This paper proposes a coupling-regularized optimal transport (CROT) algorithm for cross-domain SR to reduce the negative transport during UDA. In the proposed CROT, two consecutive processing modules regularize the coupling paths for the OT solution: a progressive inter-speaker constraint (PISC) module and a coupling-smoothed regularization (CSR) module. The PISC, designed as a pseudo-label memory bank with curriculum learning, is first applied to select valid samples to guarantee that coupling samples are from the same speaker. The CSR, designed to control the information entropy of the coupling paths further, reduces the effect of negative transport in UDA. To evaluate the effectiveness of the proposed algorithm, cross-domain SR experiments were conducted under different target domains, speaker encoders, corpora, and acoustic features. Experimental results showed that CROT achieved a 50% relative reduction in equal error rates compared to conventional OT-based UDAs, outperforming the state-of-the-art UDAs. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker VerificationabstractOptimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has difficulty computing effective transports. To address this challenge, we propose an OT-based unsupervised domain adaptation (UDA) framework for SV, OT with a diversified memory bank, called DMB-OT, which ensures the accuracy of transfers by two strategies: (1) It regularizes the solution space of OT, which attempts to plan transformations between audio samples from the same speaker with high confidence; (2) it integrates a dynamic curriculum learning algorithm, preventing OT from calculating transport couplings based on hard-discriminative samples in the early stage of UDA. Experiments under different target domains showed that our unsupervised DMB-OT could significantly improve the performance of OT-based UDA and could even match the performance of the supervised PLDA-based adaptation. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu |
ICASSP | 6 |
| 2023 | Range-Based Equal Error Rate for Spoof Localization
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi |
INTERSPEECH | 1 |
| 2023 | TMS: Temporal multi-scale in time-delay neural network for speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu, Jianwu Dang 0001 |
Appl. Intell. | 6 |
| 2023 | Self-supervised learning based domain regularization for mask-wearing speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Yantao Ji, Junhai Xu |
Speech Commun. | 6 |
| 2023 | The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an UtteranceabstractAutomatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial Spoof" (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively. Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | CS-REP: Making Speaker Verification Networks Embracing Re-ParameterizationabstractAutomatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-sequential re-parameterization (CS-Rep), a novel topology re-parameterization strategy for multi-type networks, to increase the inference speed and verification accuracy of models. CS-Rep solves the problem that existing re-parameterization methods are not suitable for typical ASV backbones. When a model applies CS-Rep, the training-period network utilizes a multi-branch topology to capture speaker information, whereas the inference-period model converts to a time-delay neural network (TDNN)-like plain backbone with stacked TDNN layers to achieve the fast inference speed. Based on CS-Rep, an improved TDNN with friendly test and deployment called Rep-TDNN is proposed. Compared with the state-of-the-art model ECAPA-TDNN, Rep-TDNN increases the actual inference speed by about 50% and reduces the EER by 10%. The code and trained models are available at https://github.com/zrtlemontree/CS-Rep. Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Lin Zhang 0054, Yantao Ji, Junhai Xu, Xugang Lu |
ICASSP | 4 |
| 2022 | Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio DetectionabstractFake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%. Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki |
INTERSPEECH | 6 |
| 2022 | Spoofing-Aware Attention based ASV Back-end with Multiple Enrollment Utterances and a Sampling Strategy for the SASV Challenge 2022abstractCurrent state-of-the-art automatic speaker verification (ASV) systems are vulnerable to presentation attacks, and several countermeasures (CMs), which distinguish bona fide trials from spoofing ones, have been explored to protect ASV. However, ASV systems and CMs are generally developed and optimized independently without considering their inter-relationship. In this paper, we propose a new spoofing-aware ASV back-end module that efficiently computes a combined ASV score based on speaker similarity and CM score. In addition to the learnable fusion function of the two scores, the proposed back-end module has two types of attention components, scaled-dot and feed-forward self-attention, so that intra-relationship information of multiple enrollment utterances can also be learned at the same time. Moreover, a new effective trials-sampling strategy is designed for simulating new spoofing-aware verification scenarios introduced in the Spoof-Aware Speaker Verification (SASV) challenge 2022. Chang Zeng, Lin Zhang 0054, Meng Liu 0017, Junichi Yamagishi |
INTERSPEECH | 2 |
| 2021 | An Initial Investigation for Detecting Partially Spoofed AudioabstractInternational audience Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Jose Patino 0001, Nicholas W. D. Evans |
Interspeech | 1 |
| 2020 | Regional Resonance of the Lower Vocal Tract and its Contribution to Speaker CharacteristicsabstractS.1391-1395 Lin Zhang 0054, Kiyoshi Honda, Jianguo Wei, Seiji Adachi |
INTERSPEECH | 1 |
| 2020 | ARET: Aggregated Residual Extended Time-Delay Neural Networks for Speaker Verification
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Longbiao Wang, Meng Liu 0017, Lin Zhang 0054, Jiayu Jin, Junhai Xu |
INTERSPEECH | 6 |