VLDB 2026 Research / reviewers in the wild / expert
Zongkun Sun
dblp:252/2019
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-6771-7215ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UnVC: Protecting Your Voiceprint by Generative Adversarial Speech
Zongkun Sun, Yihuan Huang, Yanzhen Ren, Wuyang Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Lombard-VLD: Voice Liveness Detection Based on Human Auditory FeedbackabstractVoice Liveness Detection (VLD) aims to protect speaker authentication from speech spoofing by determining whether speeches come from live speakers or loudspeakers. Previous methods mainly focus on their differences at the signal level. In this paper, we propose the first VLD that uses the human auditory feedback mechanism (i.e., the Lombard effect), called Lombard-VLD. The key idea is that live speakers can physiologically and involuntarily adjust their speaking patterns in a noisy background but loudspeakers cannot. Moreover, we design a reference-based dual input mode and a differential SE-ResBlock to model the acoustic differences caused by the Lombard effect. Experimental results show that Lombard-VLD achieves 0% and 0.24% EER in two datasets, outperforming the state-of-the-art methods. It is robust to various environmental factors, including different distances, postures of the speaker, and environmental noise, with an average accuracy of over 98.51%. It also has a good generalization to unseen speakers, genders, and datasets, with EER lower than 2.68%, 3.44%, and 7.32%, respectively. This work shows the advantages of the Lombard effect in VLD, which has fewer user limitations and better detection performance. Hongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 0008, Yongpeng Yan, Wuyang Liu, Yuhong Yang 0001, Weiping Tu |
SP | 2 |
| 2025 | APFT: Adaptive Phoneme Filter Template to Generate Anti-Compression Speech Adversarial Example in Real-TimeabstractAutomatic Speech Recognition (ASR) systems are widely used for speech censoring. Speech Adversarial Example (AE) offers a novel approach to protect speech privacy by forcing ASR to mistranscribe. However, existing speech AE faces two challenges in real-time voice communication scenarios, such as IP telephone, voice chat, or video conference, it cannot be generated in real-time, and its defensive capability is significantly reduced after the essential audio compression for network transmission. In this paper, we proposeAdaptive Phoneme Filter Template (APFT)to address these issues. The key features of APFT include: 1)Phoneme-level Templatesfor universal AE generation in real-time, 2)Filter, which eliminates redundant signals to improve compression robustness. 3)Adaptive Band Filtering, which limits the attack area from the frequency band without affecting the attack effectiveness and improves speech quality. The comprehensive experimental results show that APFT has four advantages: 1) Real-time Generation, with AE generation time below 1.1ms for 1s speech; 2) Compression Robustness, achieving a WER of 0.64 under AAC and Opus codecs; 3) Transferability, with an average WER of 0.72 across datasets and ASR systems; 4) Stealthiness, achieving a MOS of 4.07 for high-quality speech. In addition, the experiment on Telegram voice calls further proves the practical applicability of APFT. The demo of APFT can be obtained in https://yihuan-qaq.github.io/APFT.github.io/. Yihuan Huang, Yanzhen Ren, Zongkun Sun, Liming Zhai, Jingmin Wang, Wuyang Liu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | VFD-Net: Vocoder Fingerprints Detection for Fake AudioabstractWith the rapid development of audio deepfake technology, the credibility and authenticity of public opinion is facing a formidable challenge. Since vocoder is the key component of audio deepfake and leaves distinctive fingerprint features, we propose VFD-Net (Vocoder Fingerprints Detection Net), a new vocoder architectures attribution scheme, which is based on patch-wise supervised contrastive learning (PCL) to capture the global consistency of the vocoder fingerprints and to improve the detection performance in cross-set testing and audio compression scenario. PCL brings patches belonging to the same vocoder class closer together in the representation space, while pushing patches from different vocoder classes further apart. Comparative experimental results show that the average accuracy of our proposed outperforms state-of-the-art 30%-45% under cross-set testing and AAC compression circumstances. Furthermore, our proposed approach achieves a 83.67% average accuracy in short-term fake audio detection within one second. It can be used to detect partially fake audio by analyzing the consistency of vocoder fingerprints. Junlong Deng, Yanzhen Ren, Hongcheng Zhu, Zongkun Sun |
ICASSP | 5 |
| 2024 | FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder
Rubing Shen, Yanzhen Ren, Zongkun Sun |
INTERSPEECH | 3 |
| 2024 | AFPM: A Low-Cost and Universal Adversarial Defense for Speaker Recognition SystemsabstractSpeaker recognition systems (SRSs) are commonly used for biometric identification. However, these systems are vulnerable to adversarial attacks. Several defenses have been proposed but they require high costs in terms of additional data and computational resources to ensure robustness. To address these issues, this paper proposes a low-cost input reconstruction defense method called adaptive F-ratio-based partial masking (AFPM), which utilizes a robust feature extraction process to guarantee high defensibility. The underlying distribution of non-robust features is explored and filtered out by partial masking (PM), which helps maintain a low defense construction cost. An F-ratio-based PM (FPM) defense strategy is proposed by integrating the F-ratio, which reflects the weight of each frequency band for distinguishing between speakers, to balance classification accuracy and defensiveness. AFPM, which introduces an adaptive threshold calculation algorithm to FPM, is proposed to achieve further improved defensiveness and flexibility. Comparative experimental results show that AFPM is low-cost, highly defensive and universal. The construction process of AFPM does not involve training and its implementation does not require the protected SRSs to be retrained, only fine-tuned. While maintaining the classification accuracy at 99.42%, the average defense capability of AFPM against five white-box adaptive attacks is 90.89%, which is 9.23% better than that of the low-cost input reconstruction defense method and 3.77% better than that of the high-cost Parallel WaveGAN (PWG) defense approach. Against grey- and black-box adaptive attacks, FAKEBOB and Kenansville, AFPM reaches maximum defense effects of 96.01% and 74.49%, respectively, surpassing PWG by 4.5% and 65.82%. Furthermore, AFPM is universal and capable of protecting various SRSs against different attack strengths. Zongkun Sun, Yanzhen Ren, Yihuan Huang, Wuyang Liu, Hongcheng Zhu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice ConversionabstractVoice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use. Proactively tracing the origins of VC-generated speeches, i.e., speaker traceability, can prevent the misuse of VC, but unfortunately has not been extensively studied. In this paper, we are the first to investigate the speaker traceability for VC and propose a traceable VC framework named VoxTracer. Our VoxTracer is similar to but beyond the paradigm of audio watermarking. We first use unique speaker embedding to represent speaker identity. Then we design a VAE-Glow structure, in which the hiding process imperceptibly integrates the source speaker identity into the VC, and the tracing process accurately recovers the source speaker identity and even the source speech in spite of severe speech quality degradation. To address the speech mismatch between the hiding and tracing processes affected by different distortions, we also adopt an asynchronous training strategy to optimize the VAE-Glow models. The VoxTracer is versatile enough to be applied to arbitrary VC methods and popular audio coding standards. Extensive experiments demonstrate that the VoxTracer achieves not only high imperceptibility in hiding, but also nearly 100% tracing accuracy against various types of audio lossy compressions (AAC, MP3, Opus and SILK) with a broad range of bitrates (16 kbps - 128 kbps) even in a very short time duration (0.74s). Our source code is available at https://github.com/hongchengzhu/VoxTracer. Yanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun, Rubing Shen, Lina Wang 0001 |
ACM Multimedia | 4 |