Wupeng Wang

dblp:270/4658 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0000-5360-2054ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 REAL-T: Real Conversational Mixtures for Target Speaker Extraction
Shaole Li, Shuai Wang 0016, Jiangyu Han, Wupeng Wang, Haizhou Li 0001
INTERSPEECH5
2025 Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang 0003, Kun Zhou 0003, Bin Ma 0001
INTERSPEECH2
2024 Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech
abstract
Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-labels, the resulting representations are only effective for the type of pre-train data used, either clean or mixture speech. With the idea of selective auditory attention, we propose a novel pre-training solution called Selective-HuBERT, or SHuBERT, which learns the selective extraction of target speech representations from either clean or mixture speech. Specifically, SHuBERT is trained to predict pseudo labels of a target speaker, conditioned on an enrolment speech from the target speaker. By doing so, SHuBERT is expected to selectively attend to the target speaker in a complex acoustic environment, thus benefiting various downstream tasks. We further introduce a dual-path training strategy and use the cross-correlation constraint between the two branches to encourage the model to generate noise-invariant representation. Experiments on SUPERB benchmark and LibriMix dataset demonstrate the universality and noise-robustness of SHuBERT. Furthermore, we find that our high-quality representation can be easily integrated with conventional supervised learning methods to achieve significant performance, even under extremely low-resource labeled data.
Jingru Lin, Meng Ge, Wupeng Wang, Haizhou Li 0001, Mengling Feng
IEEE Signal Process. Lett.3
2024 Speech Separation With Pretrained Frontend to Minimize Domain Mismatch
abstract
Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail party scenarios. As a result, there exists a domain gap between real and synthetic data when deploying speech separation models in real-world applications. In this paper, we propose a self-supervised domain-invariant pretrained (DIP) frontend that is exposed to mixture data without the need for target reference speech. The DIP frontend utilizes a Siamese network with two innovative pretext tasks, mixture predictive coding (MPC) and mixture invariant coding (MIC), to capture shared contextual cues between real and synthetic unlabeled mixtures. Subsequently, we freeze the DIP frontend as a feature extractor when training the downstream speech separation models on synthetic data. By pretraining the DIP frontend with the contextual cues, we expect that the speech separation skills learned from synthetic data can be effectively transferred to real data. To benefit from the DIP frontend, we introduce a novel separation pipeline to align the feature resolution of the separation models. We evaluate the speech separation quality on standard benchmarks and real-world datasets. The results confirm the superiority of our DIP frontend over existing speech separation models. This study underscores the potential of large-scale pretraining to enhance the quality and intelligibility of speech separation in real-world applications.
Wupeng Wang, Zexu Pan, Shuai Wang 0016, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting
abstract
The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video clip can serve as the auxiliary reference. The use of visual cue is not only feasible, but also effective due to its noise robustness, and becoming popular. However, it is difficult to guarantee that such parallel visual cue is always available in real-world applications where visual occlusion or intermittent communication can occur. In this paper, we study the audio-visual speaker extraction algorithms with intermittent visual cue. We propose a joint speaker extraction and visual embedding inpainting framework to explore the mutual benefits. To encourage the interaction between the two tasks, they are performed alternately with an interlacing structure and optimized jointly. We also propose two types of visual inpainting losses and study our proposed method with two types of popularly used visual embeddings. The experimental results show that we outperform the baseline in terms of signal quality, perceptual quality, and intelligibility.
Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li 0001
ICASSP2
2021 Vset: A Multimodal Transformer for Visual Speech Enhancement
abstract
The transformer architecture has shown great capability in learning long-term dependency and works well in multiple domains. However, transformer has been less considered in audio-visual speech enhancement (AVSE) research, partly due to the convention that treats speech enhancement as a short-time signal processing task. In this paper, we challenge this common belief and show that an audio-visual transformer can significantly improve AVSE performance, by learning the long-term dependency of both intra-modality and inter-modality. We test this new transformer-based AVSE model on the GRID and AVSpeech datasets, and show that it beats several state-of-the-art models by a large margin.
Karthik Ramesh, Wupeng Wang
ICASSP3
2021 Neural Speaker Extraction with Speaker-Speech Cross-Attention Network
Wupeng Wang, Chenglin Xu, Meng Ge, Haizhou Li 0001
Interspeech1
2020 A Robust Audio-Visual Speech Enhancement Model
abstract
Most existing audio-visual speech enhancement (AVSE) methods work well in conditions with strong noise, however when applied to conditions with a medium SNR, serious performance degradations are often observed. These degradations can be partly attributed to the feature-fusion(early fusion etc.) architecture that tightly couples the audio information that is very strong and the visual information that is relatively weak. In this paper, we present a safe AVSE approach that can make the visual stream contribute to audio speech enhancment(ASE) safely in conditions of various SNRs by late fusion.The key novelty is two-fold: Firstly, we define power binary masks (PBMs) as a rough representation of speech signals. This rough representation admits the weakness of the visual information and so can be easily predicted from the visual stream. Secondly, we design a posterior augmentation architecture that integrate the visual-derived PBMs to the audio-derived masks via a gating network. By this architecture, the entire performance is lower-bounded by the audio-based component. Our experiments on the Grid dataset demonstrated that this new approach consistently outperforms the audio-based system in all noise conditions, confirming that it is a safe way to incorporate visual knowledge in speech enhancement.
Wupeng Wang, Dong Wang 0013, Xiao Chen 0012, Fengyu Sun
ICASSP1