EDBT 2026 Demo / reviewers in the wild / expert
Wuyang Liu
dblp:312/7464
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-2790-3069ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Security and privacy · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UnVC: Protecting Your Voiceprint by Generative Adversarial Speech
Zongkun Sun, Yihuan Huang, Yanzhen Ren, Wuyang Liu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Lombard-VLD: Voice Liveness Detection Based on Human Auditory FeedbackabstractVoice Liveness Detection (VLD) aims to protect speaker authentication from speech spoofing by determining whether speeches come from live speakers or loudspeakers. Previous methods mainly focus on their differences at the signal level. In this paper, we propose the first VLD that uses the human auditory feedback mechanism (i.e., the Lombard effect), called Lombard-VLD. The key idea is that live speakers can physiologically and involuntarily adjust their speaking patterns in a noisy background but loudspeakers cannot. Moreover, we design a reference-based dual input mode and a differential SE-ResBlock to model the acoustic differences caused by the Lombard effect. Experimental results show that Lombard-VLD achieves 0% and 0.24% EER in two datasets, outperforming the state-of-the-art methods. It is robust to various environmental factors, including different distances, postures of the speaker, and environmental noise, with an average accuracy of over 98.51%. It also has a good generalization to unseen speakers, genders, and datasets, with EER lower than 2.68%, 3.44%, and 7.32%, respectively. This work shows the advantages of the Lombard effect in VLD, which has fewer user limitations and better detection performance. Hongcheng Zhu, Zongkun Sun, Yanzhen Ren, Kun He 0008, Yongpeng Yan, Wuyang Liu, Yuhong Yang 0001, Weiping Tu |
SP | 7 |
| 2025 | APFT: Adaptive Phoneme Filter Template to Generate Anti-Compression Speech Adversarial Example in Real-TimeabstractAutomatic Speech Recognition (ASR) systems are widely used for speech censoring. Speech Adversarial Example (AE) offers a novel approach to protect speech privacy by forcing ASR to mistranscribe. However, existing speech AE faces two challenges in real-time voice communication scenarios, such as IP telephone, voice chat, or video conference, it cannot be generated in real-time, and its defensive capability is significantly reduced after the essential audio compression for network transmission. In this paper, we proposeAdaptive Phoneme Filter Template (APFT)to address these issues. The key features of APFT include: 1)Phoneme-level Templatesfor universal AE generation in real-time, 2)Filter, which eliminates redundant signals to improve compression robustness. 3)Adaptive Band Filtering, which limits the attack area from the frequency band without affecting the attack effectiveness and improves speech quality. The comprehensive experimental results show that APFT has four advantages: 1) Real-time Generation, with AE generation time below 1.1ms for 1s speech; 2) Compression Robustness, achieving a WER of 0.64 under AAC and Opus codecs; 3) Transferability, with an average WER of 0.72 across datasets and ASR systems; 4) Stealthiness, achieving a MOS of 4.07 for high-quality speech. In addition, the experiment on Telegram voice calls further proves the practical applicability of APFT. The demo of APFT can be obtained in https://yihuan-qaq.github.io/APFT.github.io/. Yihuan Huang, Yanzhen Ren, Zongkun Sun, Liming Zhai, Jingmin Wang, Wuyang Liu |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | FCC-MF: Detecting Violence in Audio-Visual Context with Frame-Wise Cluster Contrast and Modality-Stage FloodingabstractThis paper explores the detection of frame-wise instances of violence in both audio and visual modalities, where only clip-level labels are available. Previous works selected fixed value of frames for objective optimization to model frame-level features, and applied straightforward fusion strategy to aggregate audio and visual information. However, these two issues, namely Constant Frames Selection and Vulnerable Fusion, significantly impair the network’s detection performance. To address these issues, we present a novel framework called Frame-wise Cluster Contrast with Modality-stage Flooding (FCC-MF). Our contributions include: 1) We propose Frame-wise Cluster Contrast, which leverages unsupervised clustering for pseudo-labeling frames and triplet loss for contrastive learning to allow for dynamic frame-wise discrimination. 2) We propose Modality-stage Flooding, a two-stage flooding approach with the higher loss flooding level assigned to uni-modal features, which prevents over-memorization of redundant uni-modal data and promotes effective aggregation of multi-modal information. Our FCC-MF framework yields a promising average precision of 84.24% on the XD-Violence dataset, which performs favorably against previous SOTA methods. Extensive ablation studies exhibit that our FCC-MF framework produces finer frame-level violence discrimination ability and generalizable audio-visual fusion features. Jiaqing He, Yanzhen Ren, Liming Zhai, Wuyang Liu |
ICASSP | 4 |
| 2024 | Semantic Proximity Alignment: Towards Human Perception-Consistent Audio Tagging by Aligning with Label Text DescriptionabstractMost audio tagging models are trained with one-hot labels as supervised information. However, one-hot labels treat all sound events equally, ignoring the semantic hierarchy and proximity relationships between sound events. In contrast, the event descriptions contains richer information, describing the distance between different sound events with semantic proximity. In this paper, we explore the impact of training audio tagging models with auxiliary text descriptions of sound events. By aligning the audio features with the text features of corresponding labels, we inject the hierarchy and proximity information of sound events into audio encoders, improving the performance while making the prediction more consistent with human perception. We refer to this approach as Semantic Proximity Alignment (SPA). We use Ontology-aware mean Average Precision (OmAP) as the main evaluation metric for the models. OmAP reweights the false positives based on Audioset ontology distance and is more consistent with human perception compared to mAP. Experimental results show that the audio tagging models trained with SPA achieve higher OmAP compared to models trained with one-hot labels solely (+1.8 OmAP). Human evaluations also demonstrate that the predictions of SPA models are more consistent with human perception. Wuyang Liu, Yanzhen Ren |
ICASSP | 1 |
| 2024 | Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image FusionabstractMulti-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in real MMIF applications. This is because of the limited depth of the focus of visible optical lenses, which impedes the simultaneous capture of the focal information within the same scene. To address this issue, in this paper, we propose a MMIF framework for joint focused integration and modalities information extraction. Specifically, a semi-sparsity-based smoothing filter is introduced to decompose the images into structure and texture components. Subsequently, a novel multi-scale operator is proposed to fuse the texture components, capable of detecting significant information by considering the pixel focus attributes and relevant data from various modal images. Additionally, to achieve an effective capture of scene luminance and reasonable contrast maintenance, we consider the distribution of energy information in the structural components in terms of multi-directional frequency variance and information entropy. Extensive experiments on existing MMIF datasets, as well as the object detection and depth estimation tasks, consistently demonstrate that the proposed algorithm can surpass the state-of-the-art methods in visual perception and quantitative evaluation. The code is available at https://github.com/ixilai/MFIF-MMIF. Xilai Li, Xiaosong Li 0004, Tao Ye 0002, Xiaoqi Cheng, Wuyang Liu, Haishu Tan |
WACV | 5 |
| 2024 | AFPM: A Low-Cost and Universal Adversarial Defense for Speaker Recognition SystemsabstractSpeaker recognition systems (SRSs) are commonly used for biometric identification. However, these systems are vulnerable to adversarial attacks. Several defenses have been proposed but they require high costs in terms of additional data and computational resources to ensure robustness. To address these issues, this paper proposes a low-cost input reconstruction defense method called adaptive F-ratio-based partial masking (AFPM), which utilizes a robust feature extraction process to guarantee high defensibility. The underlying distribution of non-robust features is explored and filtered out by partial masking (PM), which helps maintain a low defense construction cost. An F-ratio-based PM (FPM) defense strategy is proposed by integrating the F-ratio, which reflects the weight of each frequency band for distinguishing between speakers, to balance classification accuracy and defensiveness. AFPM, which introduces an adaptive threshold calculation algorithm to FPM, is proposed to achieve further improved defensiveness and flexibility. Comparative experimental results show that AFPM is low-cost, highly defensive and universal. The construction process of AFPM does not involve training and its implementation does not require the protected SRSs to be retrained, only fine-tuned. While maintaining the classification accuracy at 99.42%, the average defense capability of AFPM against five white-box adaptive attacks is 90.89%, which is 9.23% better than that of the low-cost input reconstruction defense method and 3.77% better than that of the high-cost Parallel WaveGAN (PWG) defense approach. Against grey- and black-box adaptive attacks, FAKEBOB and Kenansville, AFPM reaches maximum defense effects of 96.01% and 74.49%, respectively, surpassing PWG by 4.5% and 65.82%. Furthermore, AFPM is universal and capable of protecting various SRSs against different attack strengths. Zongkun Sun, Yanzhen Ren, Yihuan Huang, Wuyang Liu, Hongcheng Zhu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Attention Mixup: An Accurate Mixup Scheme Based On Interpretable Attention Mechanism for Multi-Label Audio ClassificationabstractMixup proves to be an efficient data augmentation method on audio classification tasks. Original mixup scheme directly mixes the waveform of two random samples, which not only ignores the temporal distribution of the sound events but may also interfere with the original sound events in another sample. This paper proposes Attention MixUp (AMU), which only selects those segments that contain sound events for mixup, rather than simply mixing the entire sample. AMU utilizes the attention maps of pretrained audio classification Vision Transformer (ViT) to filter out the patches on the spectrogram that are useful for classification and then selects the regions for mixup according to three different strategies. Experimental results show a remarkable improvement (+1.9 mAP) on state-of-the-art Audioset classification methods with either CNN or ViT backbone. Further experiments show that AMU achieves the performance gain by improving the accuracy on short events (0.1s to 2s) by an average of 6.8% while keeping the accuracy on longer events. Wuyang Liu, Yanzhen Ren |
ICASSP | 1 |
| 2021 | Recalibrated Bandpass Filtering On Temporal Waveform For Audio Spoof DetectionabstractDeepfake techniques mislead people’s cognition with high-quality fake videos and audios, and speech synthesis is an important tool to implement cognitive attacks, which mainly include Text-To-Speech (TTS) and Voice Conversion (VC). In this paper, we propose a method for audio spoof detection based on frequency band recalibration via sinc convolution and squeeze-excitation module, extracting features from the temporal waveform and emphasizing the frequency bands that are more useful on this task. Experimental results show that the proposed method outperforms other similar methods by 18.6% with an average EER of 7.23%, and achieve better generalizability on the detection of unseen spoofing methods, while the size of the model is reduced by 30.8%. Yanzhen Ren, Wuyang Liu, Dengkai Liu, Lina Wang 0001 |
ICIP | 2 |