EDBT 2026 Demo / reviewers in the wild / expert
Lianwu Chen
dblp:223/9843
· DBLP profile ↗
25ranked-venue papers
5as first author
11since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech EnhancementabstractDeep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, personalized speech enhancement. As shown in the evaluation results, this is achieved by employing a multi-stage and multi-loss training architecture that incorporates the recently proposed two-step structure, ASR loss produced by a back-end ASR encoder, and the speaker extraction network. Lianwu Chen, Chenglin Xu, Xinlei Ren, Xiguang Zheng |
ICASSP | 1 |
| 2022 | A Two-Step Backward Compatible Fullband Speech Enhancement SystemabstractSpeech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system is proposed in this paper. Compared to the existing full-band systems that utilize perceptually motivated features to train the fullband speech enhancement with a single network structure, the proposed system is a two-step system ensuring good fullband speech enhancement quality while backward compatible to the existing wideband systems. Lianwu Chen, Xiguang Zheng, Xinlei Ren |
ICASSP | 2 |
| 2022 | A Deep Hierarchical Fusion Network for Fullband Acoustic Echo CancellationabstractDeep learning based wideband (16kHz) acoustic echo cancellation (AEC) approaches have surpassed traditional methods. This work proposes a deep hierarchical fusion (DHF) network with intra-network and inter-network fusion to further improve the wideband AEC performance. Meanwhile, this work extends the existing wideband systems to enable fullband (48kHz) AEC while simultaneously ensuring automatic speech recognition compatibility by incorporating with an ASR loss. The proposed system has ranked 2nd place in ICASSP 2022’s AEC Challenge. Runqiang Han, Lianwu Chen, Xiguang Zheng |
ICASSP | 4 |
| 2022 | Multi-Scale Temporal-Frequency Attention for Music Source SeparationabstractIn recent years, deep neural networks (DNNs) based approaches have achieved the start-of-the-art performance for music source separation (MSS). Although previous methods have addressed the large receptive field modeling using various methods, the temporal and frequency correlations of the music spectrogram with repeated patterns have not been explicitly explored for the MSS task. In this paper, a temporal-frequency attention module is proposed to model the spectrogram correlations along both temporal and frequency dimensions. Moreover, a multi-scale attention is proposed to effectively capture the correlations for music signal. The experimental results on MUSDB18 dataset show that the proposed method outperforms the existing state-of-the-art systems with 9.51 dB signal-to-distortion ratio (SDR) on separating the vocal stems, which is the primary practical application of MSS. Lianwu Chen, Xiguang Zheng |
ICME | 1 |
| 2022 | Impairment Representation Learning for Speech Quality Assessment
Lianwu Chen, Xinlei Ren, Xiguang Zheng |
INTERSPEECH | 1 |
| 2021 | ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech SeparationabstractSpeech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech recognition (ASR) systems. The conventional mask-based minimum variance distortionless response (MVDR) beamformer can be used to minimize the distortion, but comes with high level of residual noise. Furthermore, the matrix operations (e.g., matrix inversion) involved in the conventional MVDR solution are sometimes numerically unstable when jointly trained with neural networks. In this paper, we propose a novel all deep learning MVDR framework, where the matrix inversion and eigenvalue decomposition are replaced by two recurrent neural networks (RNNs), to resolve both issues at the same time. The proposed method can greatly reduce the residual noise while keeping the target speech undistorted by leveraging on the RNN-predicted frame-wise beamforming weights. The system is evaluated on a Mandarin audio-visual corpus and compared against several state-of-the-art (SOTA) speech separation systems. Experimental results demonstrate the superiority of the proposed method across several objective metrics and ASR accuracy. Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001 |
ICASSP | 5 |
| 2021 | A Causal U-Net Based Neural Beamforming Network for Real-Time Multi-Channel Speech Enhancement
Xinlei Ren, Lianwu Chen, Xiguang Zheng |
Interspeech | 3 |
| 2021 | TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech DereverberationabstractIn this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach on the deep neural network (DNN) for environment-aware speech dereverberation, which can adaptively attend to the contextual information. More specifically, a FullBand based Temporal Attention approach (FTA) is proposed, which models the correlations between the fullband information of the context frames. In addition, considering the difference between the attenuation of high frequency bands and low frequency bands (high frequency bands attenuate faster than low frequency bands) in the room impulse response (RIR), we also propose a SubBand based Temporal Attention approach (STA). In order to guide the network to be more aware of the reverberant environments, we jointly optimize the dereverberation network and the reverberation time (RT60) estimator in a multi-task manner. Our experimental results indicate that the proposed method outperforms our previously proposed reverberation-time-aware DNN and the learned attention weights are fully physical consistent. We also report a preliminary yet promising dereverberation and recognition experiment on real test data. Helin Wang, Bo Wu 0011, Lianwu Chen, Meng Yu 0003, Jianwei Yu 0001, Yong Xu 0004, Shixiong Zhang 0001, Chao Weng, Dan Su 0002, Dong Yu 0001 |
Interspeech | 3 |
| 2021 | Low-Delay Speech Enhancement Using Perceptually Motivated Target and Loss
Xinlei Ren, Xiguang Zheng, Lianwu Chen |
Interspeech | 4 |
| 2021 | Neural Mask based Multi-channel Convolutional Beamforming for Joint Dereverberation, Echo Cancellation and DenoisingabstractThis paper proposes a new joint optimization framework for simultaneous dereverberation, acoustic echo cancellation, and denoising, which is motivated by the recently proposed con-volutional beamformer for simultaneous denoising and dereverberation. Using the echo aware mask based beamforming framework, the proposed algorithm could effectively deal with double-talk case and local inference, etc. The evaluations based on ERLE for echo only, and PESQ for double-talk demonstrate that the proposed algorithm could significantly improve the performance. Meng Yu 0003, Yong Xu 0004, Chao Weng, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001 |
SLT | 6 |
| 2021 | Multi-Channel Multi-Frame ADL-MVDR for Target Speech SeparationabstractMany purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems. Minimum variance distortionless response (MVDR) filters are often adopted to remove nonlinear distortions, however, conventional neural mask-based MVDR systems still result in relatively high levels of residual noise. Moreover, the matrix inverse involved in the MVDR solution is sometimes numerically unstable during joint training with neural networks. In this study, we propose a multi-channel multi-frame (MCMF) all deep learning (ADL)-MVDR approach for target speech separation, which extends our preliminary multi-channel ADL-MVDR approach. The proposed MCMF ADL-MVDR system addresses linear and nonlinear distortions. Spatio-temporal cross correlations are also fully utilized in the proposed approach. The proposed systems are evaluated using a Mandarin audio-visual corpus and are compared with several state-of-the-art approaches. Experimental results demonstrate the superiority of our proposed systems under different scenarios and across several objective evaluation metrics, including ASR performance. Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Donald S. Williamson, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature LearningabstractHand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In this work, we propose an integrated architecture for learning spatial features directly from the multi-channel speech waveforms within an end-to-end speech separation framework. In this architecture, time-domain filters spanning signal channels are trained to perform adaptive spatial filtering. These filters are implemented by a 2d convolution (conv2d) layer and their parameters are optimized using a speech separation objective function in a purely data-driven fashion. Furthermore, inspired by the IPD formulation, we design a conv2d kernel to compute the inter-channel convolution differences (ICDs), which are expected to provide the spatial cues that help to distinguish the directional sources. Evaluation results on simulated multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed ICD based MCSS model improves the overall signal-to-distortion ratio by 10.4% over the IPD based MCSS model. Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001 |
ICASSP | 3 |
| 2020 | Improving Reverberant Speech Training Using Diffuse Acoustic SimulationabstractWe present an efficient and realistic geometric acoustic simulation approach for generating and augmenting training data in speech-related machine learning tasks. Our physically-based acoustic simulation method is capable of modeling occlusion, specular and diffuse reflections of sound in complicated acoustic environments, whereas the classical image method can only model specular reflections in simple room settings. We show that by using our synthetic training data, the same neural networks gain significant performance improvement on real test sets in far-field speech recognition by 1.58% and keyword spotting by 21%, without fine-tuning using real impulse responses. Zhenyu Tang 0001, Lianwu Chen, Bo Wu 0011, Dong Yu 0001, Dinesh Manocha |
ICASSP | 2 |
| 2020 | Neural Spatio-Temporal Beamformer for Target Speech SeparationabstractPurely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR).On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability.In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement.Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the intermicrophone correlation that is already utilized in prior arts.Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN.The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts. Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Chao Weng, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2020 | Audio-Visual Multi-Channel Recognition of Overlapped SpeechabstractAutomatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption, this paper presents an audio-visual multi-channel overlapped speech recognition system featuring tightly integrated separation front-end and recognition back-end. A series of audio-visual multi-channel speech separation front-end components based on \textit{TF masking}, \textit{filter\&sum} and \textit{mask-based MVDR} beamforming approaches were developed. To reduce the error cost mismatch between the separation and recognition components, they were jointly fine-tuned using the connectionist temporal classification (CTC) loss function, or a multi-task criterion interpolation with scale-invariant signal to noise ratio (Si-SNR) error cost. Experiments suggest that the proposed multi-channel AVSR system outperforms the baseline audio-only ASR system by up to 6.81\% (26.83\% relative) and 22.22\% (56.87\% relative) absolute word error rate (WER) reduction on overlapped speech constructed using either simulation or replaying of the lipreading sentence 2 (LRS2) dataset respectively. Jianwei Yu 0001, Bo Wu 0011, Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Dong Yu 0001, Xunying Liu, Helen M. Meng |
INTERSPEECH | 5 |
| 2019 | Time Domain Audio Visual Speech SeparationabstractAudio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker extraction from monaural mixtures. The architecture generalizes the previous TasNet (time-domain speech separation network) to enable multi-modal learning and at meanwhile it extends the classical audio-visual speech separation from frequency-domain to time-domain. The main components of proposed architecture include an audio encoder, a video encoder that extracts lip embedding from video streams, a multi-modal separation network and an audio decoder. Experiments on simulated mixtures based on recently released LRS2 dataset show that our method can bring 3dB+ and 4dB+ Si-SNR improvements on two- and three-speaker cases respectively, compared to audio-only TasNet and frequency-domain audio-visual networks. Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001 |
ASRU | 4 |
| 2019 | Improving Speech Enhancement with Phonetic Embedding FeaturesabstractIn this paper, we present a speech enhancement framework that leverages phonetic information obtained from the acoustic model. It consists of two separate components: (i) a long short-term memory recurrent neural network (LSTM-RNN) based speech enhancement model that takes the combination of log-power spectra (LPS) and phonetic embedding features as input to predict the complex ideal ratio mask (cIRM); and (ii) a convolutional, long short-term memory and fully connected deep neural network (CLDNN) based acoustic model that extracts the phonetic feature vector in the hidden units of its LSTM layer. Our experimental results show that the proposed framework outperforms both the conventional and phoneme-dependent speech enhancement systems under various noisy conditions, generalizes well to unseen conditions, and performs robustly to the speech interference. We further demonstrate its superior enhancement performance on unvoiced speech and report a preliminary yet promising recognition experiment on real test data. Bo Wu 0011, Meng Yu 0003, Lianwu Chen, Mingjie Jin, Dan Su 0002, Dong Yu 0001 |
ASRU | 3 |
| 2019 | Multi-band PIT and Model Integration for Improved Multi-channel Speech SeparationabstractThe recent exploration of deep learning for supervised speech separation has significantly accelerated the progress on the multi-talker speech separation problem. Multi-channel extension has attracted much research attention due to the benefit of spatial information in far-field acoustic environments. In this paper, We review the most recent models of multi-channel permutation invariant training (PIT), investigate spatial features formed by microphone pairs and their underlying impact and issue, present a multi-band architecture for effective feature encoding, and conduct a model integration between single-channel and multi-channel PIT for resolving the spatial overlapping problem in the conventional multi-channel PIT framework. The evaluation confirms the significant improvement achieved with the proposed model and training approach for the multi-channel speech separation. Lianwu Chen, Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 1 |
| 2019 | Neural Spatial Filter: Target Speaker Speech Separation Assisted with Directional Information
Rongzhi Gu, Lianwu Chen, Shixiong Zhang 0001, Jimeng Zheng, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2019 | Direction-Aware Speaker Beam for Multi-Channel Speaker Extraction
Guanjun Li, Shan Liang 0007, Shuai Nie 0001, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li |
INTERSPEECH | 6 |
| 2019 | Jointly Adversarial Enhancement Training for Robust End-to-End Speech Recognition
Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li |
INTERSPEECH | 6 |
| 2019 | Improved Speaker-Dependent Separation for CHiME-5 ChallengeabstractThis paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-channel, highly-overlapped conversational speech recognition in a dinner party scenario with reverberations and nonstationary noises.We adopt a speaker-aware training method by using i-vector as the target speaker information for multi-talker speech separation.With only one unified separation model for all speakers, we achieve a 10% absolute improvement in terms of word error rate (WER) over the previous baseline of 80.28% on the development set by leveraging our newly proposed data processing techniques and beamforming approach.With our improved back-end acoustic model, we further reduce WER to 60.15% which surpasses the result of our submitted CHiME-5 challenge system without applying any fusion techniques. Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2018 | Permutation Invariant Training of Generative Adversarial Network for Monaural Speech Separation
Lianwu Chen, Meng Yu 0003, Yanmin Qian, Dan Su 0002, Dong Yu 0001 |
INTERSPEECH | 1 |
| 2018 | Deep Extractor Network for Target Speaker Recovery from Single Channel Speech MixturesabstractSpeaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources.However, it remains challenging to achieve satisfying performance provided a very short available target speaker utterance (anchor).Here we present a novel "deep extractor network" which creates an extractor point for the target speaker in a canonical high dimensional embedding space, and pulls together the time-frequency bins corresponding to the target speaker.The proposed model is different from prior works in that the canonical embedding space encodes knowledges of both the anchor and the mixture during an end-to-end training phase: First, embeddings for the anchor and mixture speech are separately constructed in a primary embedding space, and then combined as an input to feed-forward layers to transform to a canonical embedding space which we discover more stable than the primary one.Experimental results show that given a very short utterance, the proposed model can efficiently recover high quality target speech from a mixture, which outperforms various baseline models, with 5.2% and 6.6% relative improvements in SDR and PESQ respectively compared with a baseline oracle deep attracor model.Meanwhile, we show it can be generalized well to more than one interfering speaker. Jun Wang 0091, Jie Chen 0057, Dan Su 0002, Lianwu Chen, Meng Yu 0003, Yanmin Qian, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2018 | Text-Dependent Speech Enhancement for Small-Footprint Robust Keyword Detection
Meng Yu 0003, Lianwu Chen, Jie Chen 0057, Jimeng Zheng, Dan Su 0002, Dong Yu 0001 |
INTERSPEECH | 4 |