EDBT 2026 Demo / reviewers in the wild / expert
Xueliang Zhang 0001
dblp:60/8053-1
· DBLP profile ↗
57ranked-venue papers
6as first author
24since 2021 · last 2025
0000-0002-0406-1105ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attention-Based Beamformer For Multi-Channel Speech EnhancementabstractMinimum Variance Distortionless Response (MVDR) is a classical adaptive beamformer that theoretically ensures the distortionless transmission of signals in the target direction, which makes it popular in real applications. Its noise reduction performance actually depends on the accuracy of the noise and speech spatial covariance matrices (SCMs) estimation. Time-frequency masks are often used to compute these SCMs. However, most mask-based beamforming methods typically assume that the sources are stationary, ignoring the case of moving sources, which leads to performance degradation. In this paper, we propose an attention-based mechanism to calculate the speech and noise SCMs and then apply MVDR to obtain the enhanced speech. To fully incorporate spatial information, the inplace convolution operator and frequency-independent LSTM are applied to facilitate SCMs estimation. The model is optimized in an end-to-end manner. Experiments demonstrate that the proposed method outperforms baselines with reduced computation and fewer parameters under various conditions. Jinglin Bai, Hao Li 0046, Xueliang Zhang 0001, Fei Chen 0011 |
ICASSP | 3 |
| 2025 | Vector Quantized Diffusion Model Based Speech Bandwidth ExtensionabstractRecent advancements in neural audio codec (NAC) unlock new potential in audio signal processing. Studies have increasingly explored leveraging the latent features of NAC for various speech signal processing tasks. This paper introduces the first approach to speech bandwidth extension (BWE) that utilizes the discrete features obtained from NAC. By restoring high-frequency details within highly compressed discrete tokens, this approach enhances speech intelligibility and naturalness. Based on Vector Quantized Diffusion, the proposed framework combines the strengths of advanced NAC, diffusion models, and Mamba-2 to reconstruct high-frequency speech components. Extensive experiments demonstrate that this method exhibits superior performance across both log-spectral distance and ViSQOL, significantly improving speech quality. Jinglin Bai, Xueliang Zhang 0001 |
ICASSP | 4 |
| 2025 | Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic TransformabstractThe performance of traditional beamforming algorithms is influenced by the number of microphones, with performance improving as the number increases. However, in practice, the number of microphones is often limited. In this paper, we propose a novel virtual microphone estimation method that combines the strengths of both traditional and neural network-based approaches using the spherical harmonic transform (SHT), effectively addressing their respective limitations. Our method predicts the SHT coefficients at virtual positions and inversely transforms them into virtual speech signals, leveraging spatial information in the spherical harmonic domain for more accurate and effective virtual microphone estimation. Evaluations on the open MS-SNSD dataset demonstrate that the proposed method outperforms established baselines. Hui Zhang 0031, Xueliang Zhang 0001 |
ICASSP | 3 |
| 2025 | Multi-Channel Acoustic Echo Cancellation Based on Direction-of-Arrival Estimation
Xueliang Zhang 0001, Zhongqiu Wang 0001 |
INTERSPEECH | 2 |
| 2024 | 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource ApplicationsabstractTarget speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction. Shulin He, Hao Li 0046, Yang Yang 0121, Fei Chen 0011, Xueliang Zhang 0001 |
ICASSP | 6 |
| 2024 | Hierarchical Speaker Representation for Target Speaker ExtractionabstractTarget speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the voice of the target speaker. However, the representation of the speaker embedding is too simplistic, often being merely a 1×1024 vector. This dense information makes it difficult for the separation network to harness effectively. To address this limitation, we introduce a pioneering methodology called Hierarchical Representation (HR) that seamlessly fuses anchor data across granular and overarching 5 layers of the separation network, enhancing the precision of target extraction. HR amplifies the efficacy of anchors to improve target speaker isolation. On the Libri-2talker dataset, HR substantially outperforms state-of-the-art time-frequency domain techniques. Further demonstrating HR’s capabilities, we achieved first place in the prestigious ICASSP 2023 Deep Noise Suppression Challenge. The proposed HR methodology shows great promise for advancing target speaker extraction through enhanced anchor utilization. Shulin He, Huaiwen Zhang, Wei Rao 0002, Kanghao Zhang, Yukai Jv, Yang Yang 0121, Xueliang Zhang 0001 |
ICASSP | 7 |
| 2024 | Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional EncodingabstractMulti-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is therefore key. Deep learning shows great potential on multi-channel speech enhancement and often takes short-time Fourier Transform (STFT) as inputs directly. To fully leverage the spatial information, we introduce a method using spherical harmonics transform (SHT) coefficients as auxiliary model inputs. These coefficients concisely represent spatial distributions. Specifically, our model has two encoders, one for the STFT and another for the SHT. By fusing both encoders in the decoder to estimate the enhanced STFT, we effectively incorporate spatial context. Evaluations on TIMIT under varying noise and reverberation show our model outperforms established benchmarks. Remarkably, this is achieved with fewer computations and parameters. By leveraging spherical harmonics to incorporate directional cues, our model efficiently improves the performance of the multi-channel speech enhancement. Pengjie Shen, Hui Zhang 0031, Xueliang Zhang 0001 |
ICASSP | 4 |
| 2024 | Innovative Directional Encoding in Speech Processing: Leveraging Spherical Harmonics Injection for Multi-Channel Speech Enhancement
Pengjie Shen, Hui Zhang 0031, Xueliang Zhang 0001 |
IJCAI | 4 |
| 2024 | Cross-Attention-Guided WaveNet for EEG-to-MEL Spectrogram Reconstruction
Hao Li 0046, Xueliang Zhang 0001, Fei Chen 0011, Guanglai Gao |
INTERSPEECH | 3 |
| 2023 | Speech Enhancement with Intelligent Neural Homomorphic SynthesisabstractMost neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral analysis to obtain noisy speech’s excitation and vocal tract. Unlike traditional signal processing, we use an attentive recurrent network (ARN) model predicted ratio mask to replace the liftering separation function. Then two convolutional attentive recurrent network (CARN) networks are used to predict the excitation and vocal tract of clean speech, respectively. The system’s output is synthesized from the estimated excitation and vocal. Experiments prove that our proposed method performs better, with SI-SNR improving by 1.363dB compared to FullSubNet. Shulin He, Wei Rao 0002, Jun Chen 0024, Yukai Jv, Xueliang Zhang 0001, Yannan Wang, Shidong Shang |
ICASSP | 6 |
| 2023 | Neural Multi-Channel and Multi-Microphone Acoustic Echo CancellationabstractDeep learning is introduced in multi-channel (MC) and multi-microphone (MM) acoustic echo cancellation (AEC) without decorrelation to the loudspeaker signals and achieves remarkable performance. In this paper, we propose a complex spectral mapping framework with inplace convolution and frequency-wise temporal modeling for MCAEC problem, which efficiently models the echo paths and spatial information. The proposed method is a multi-input and multi-output (MIMO) scheme, which filters out echoes from all microphone signals simultaneously, so the computational cost is greatly reduced. In addition, a cross-domain loss function with a multi-task learning strategy is designed for better generalization capability. Experiments are conducted on various unmatched scenarios and results show that the proposed method significantly outperforms previous methods. Moreover, a lightweight version of the proposed model with 0.29 million trainable parameters also shows good performance, which is essential for resource-limited and real-time applications. Chenggang Zhang, Hao Li 0046, Xueliang Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | DRC-NET: Densely Connected Recurrent Convolutional Neural Network for Speech DereverberationabstractUnder our previous work on frequency bin-wise independent processing, a dramatic reduction of the computational complexity for recurrent neural networks (RNN) is achieved. So that a massive deployment of RNN in time dimension is realized in this paper, by using the channel-wise long short-term memory neural network. Based on this approach, the processing of RNN on frequency dimension and time dimension in the time-frequency domain are unified. This allows us to combine convolutional neural network (CNN) and RNN as a basic neural operator, which finally leads to the Densely Connected Recurrent Convolutional Neural Network (DRC-NET). The DRC-NET sufficiently exploits the infinite response of RNN, and the finite response of CNN. Its balanced response characteristics significantly improve the system performance. Experimental result shows that both non-causal and causal version of DRC-NET outperforms the state-of-the-art (STOA) model for speech dereverberation task. Xueliang Zhang 0001 |
ICASSP | 2 |
| 2022 | Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex DomainabstractBone-conduction (BC) microphones capture speech signals by converting the vibrations of the human skull into electrical signals. BC sensors are insensitive to acoustic noise, but limited in bandwidth. On the other hand, conventional or air-conduction (AC) microphones are capable of capturing full-band speech, but are susceptible to background noise. We propose to combine the strengths of AC and BC microphones by employing a convolutional recurrent network that performs complex spectral mapping. To better utilize signals from both kinds of microphone, we employ attention-based fusion with early-fusion and late-fusion strategies. Experiments demonstrate the superiority of the proposed method over other recent speech enhancement methods combining BC and AC signals. In addition, our enhancement performance is significantly better than conventional speech enhancement counterparts, especially in low signal-to-noise ratio scenarios. Heming Wang, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 2 |
| 2022 | Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech EnhancementabstractIn this paper, we study the loss-metric mismatch problem of supervised single-channel speech enhancement system. Most of the existing speech enhancement systems achieve unsatisfying performance since their empirically selected loss functions have semantic gaps with the non-differentiable evaluation metrics, a.k.a., the loss-metric mismatch problem. In this work, we propose a simple yet efficient method to generate suitable loss functions for the real front-end speech enhancement scenarios to alleviate the loss-metric mismatch problem. Specifically, we adopt the function smoothing technique and approximate the non-differentiable evaluation metrics by a set of basis functions and their linear combination. Experimental results demonstrate that the loss function generated by our method helps the speech enhancement system achieve remarkable performance in most evaluation metrics than the traditional empirically selected ones. Yang Yang 0121, Hui Zhang 0031, Xueliang Zhang 0001, Huaiwen Zhang |
ICASSP | 3 |
| 2022 | A Robust Deep Audio Splicing Detection Method via Singularity Detection FeatureabstractThere are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose a high-frequency singularity detection feature obtained by wavelet transform. The proposed feature can explicitly show the location of the tampering operation on the waveform. Moreover, the long short-term memory (LSTM) is introduced to the CNN-architecture LCNN to ensure that the sequence information can be fully learned. The proposed feature is sent to the improved RNN-architecture LCNN together with the widely used linear frequency cepstral coefficients (LFCC) to learn forgery characteristics where the LFCC is used as a supplement. Systematic evaluation and comparison show that the proposed method has greatly improved the accuracy and generalization. Kanghao Zhang, Shan Liang 0007, Shuai Nie 0001, Shulin He, Xueliang Zhang 0001, Haoxin Ma, Jiangyan Yi |
ICASSP | 6 |
| 2022 | A Complex Spectral Mapping with Inplace Convolution Recurrent Neural Networks For Acoustic Echo CancellationabstractRecently, deep learning is introduced in acoustic echo cancellation (AEC) and achieves remarkable performance. For deep learning-based AEC, the most important problem is generalization ability in diversity scenarios. Different from most methods which process the entire frequency band, we propose inplace convolution recurrent neural networks (ICRN) for end-to-end AEC, which utilizes inplace convolution and channel-wise temporal modeling to ensure the near-end signal information being preserved. In addition, we employ complex spectral mapping with a multi-task learning strategy for better generalization capability. Experiments conducted on various unmatched scenarios show that the proposed method outperforms previous methods. Moreover, the system has 210K parameters and 1.76G MACs, which is suitable for real-time applications. Chenggang Zhang, Xueliang Zhang 0001 |
ICASSP | 3 |
| 2022 | Speaker recognition-assisted robust audio deepfake detection
Shuai Nie 0001, Hui Zhang 0031, Shulin He, Kanghao Zhang, Shan Liang 0007, Xueliang Zhang 0001, Jianhua Tao 0001 |
INTERSPEECH | 7 |
| 2022 | LCSM: A Lightweight Complex Spectral Mapping Framework for Stereophonic Acoustic Echo Cancellation
Chenggang Zhang, Xueliang Zhang 0001 |
INTERSPEECH | 3 |
| 2022 | Fusing Bone-Conduction and Air-Conduction Sensors for Complex-Domain Speech EnhancementabstractSpeech enhancement aims to improve the listening quality and intelligibility of noisy speech in adverse environments. It proves to be challenging to perform speech enhancement in very low signal-to-noise ratio (SNR) conditions. Conventional speech enhancement utilizes air-conduction (AC) microphones, which are sensitive to background noise but capable of capturing full-band signals. On the other hand, bone-conduction (BC) sensors are unaffected by acoustic noise, but recorded speech has limited bandwidth. This study proposes an attention-based fusion method to combine the strengths of AC and BC signals and perform complex spectral mapping for speech enhancement. Experiments on the EMSB dataset demonstrate that the proposed approach effectively leverages the advantages of AC and BC sensors, and outperforms a recent time-domain baseline in all conditions. We also show that the sensor fusion method is superior to single-sensor counterparts, especially in low SNR conditions. As the amount of BC data is very limited, we additionally propose a semi-supervised technique to utilize both parallelly and unparallely recorded AC and BC speech signals. With additional AC speech from the AISHELL-1 dataset, we achieve similar performance to supervised learning with only 50% parallel data. Heming Wang, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral MappingabstractSpeech quality and intelligibility can be severely degraded by back-ground noise in mobile communication. In order to attenuate back-ground noise, speech enhancement systems have been integrated into mobile phones, and a microphone array is typically deployed to improve the enhancement performance. This paper proposes a novel approach to real-time speech enhancement for dual-microphone mobile phones. Our approach employs a causal densely-connected convolutional recurrent network to perform dual-channel complex spectral mapping. We apply a structured pruning technique for compressing the model without significantly affecting the enhancement performance. This leads to a real-time enhancement system for on-device processing. Evaluation results show that the pro-posed approach substantially advances the performance of an earlier approach to dual-channel speech enhancement for mobile communication. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 2 |
| 2021 | Inplace Gated Convolutional Recurrent Neural Network for Dual-Channel Speech EnhancementabstractFor dual-channel speech enhancement, it is a promising idea to design an end-to-end model based on the traditional array signal processing guideline and the manifold space of multi-channel signals. We found that the idea above can be effectively implemented by the classical convolutional recurrent neural networks (CRN) architecture. We propose a very compact in place gated convolutional recurrent neural network (inplace GCRN) for end-to-end multi-channel speech enhancement, which utilizes inplace-convolution for frequency pattern extraction and reconstruction. The inplace characteristics efficiently preserve spatial cues in each frequency bin for channel-wise long short-term memory neural networks (LSTM) tracing the spatial source. In addition, we come up with a new spectrum recovery method by predict amplitude mask, mapping, and phase, which effectively improves the speech quality. Xueliang Zhang 0001 |
Interspeech | 2 |
| 2021 | DBNet: A Dual-Branch Network Architecture Processing on Spectrum and Waveform for Single-Channel Speech EnhancementabstractIn real acoustic environment, speech enhancement is an arduous task to improve the quality and intelligibility of speech interfered by background noise and reverberation.Over the past years, deep learning has shown great potential on speech enhancement.In this paper, we propose a novel real-time framework called DBNet which is a dual-branch structure with alternate interconnection.Each branch incorporates an encoderdecoder architecture with skip connections.The two branches are responsible for spectrum and waveform modeling, respectively.A bridge layer is adopted to exchange information between the two branches.Systematic evaluation and comparison show that the proposed system substantially outperforms related algorithms under very challenging environments.And in INTERSPEECH 2021 Deep Noise Suppression (DNS) challenge, the proposed system ranks the top 8 in real-time track 1 in terms of the Mean Opinion Score (MOS) of the ITU-T P.835 framework. Kanghao Zhang, Shulin He, Hao Li 0046, Xueliang Zhang 0001 |
Interspeech | 4 |
| 2021 | Recurrent Neural Networks and Acoustic Features for Frame-Level Signal-to-Noise Ratio EstimationabstractIt is important to know the presence and the relative level of background noise for many speech processing tasks. Frame-level signal-to-noise ratio (SNR) provides a measure of instantaneous noise level of a noisy signal, and its estimation has been researched for decades. This problem can be approached from a supervised learning perspective by predicting SNR from features of noisy speech. In this study, we introduce a deep learning algorithm for frame-level SNR estimation. The proposed algorithm employs recurrent neural networks (RNNs) with long short-term memory (LSTM) to leverage contextual information. We also systematically examine a range of acoustic features and investigate feature combinations using Group Lasso and sequential floating forward selection (SFFS). The proposed algorithm naturally leads to an utterance-level SNR estimator. Systematical evaluations show that the proposed algorithm provides an accurate estimate of frame-level SNR, as well as utterance-level SNR, under different noise conditions, outperforming other estimators. Hao Li 0046, DeLiang Wang, Xueliang Zhang 0001, Guanglai Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Deep Learning Based Real-Time Speech Enhancement for Dual-Microphone Mobile PhonesabstractIn mobile speech communication, speech signals can be severely corrupted by background noise when the far-end talker is in a noisy acoustic environment. To suppress background noise, speech enhancement systems are typically integrated into mobile phones, in which one or more microphones are deployed. In this study, we propose a novel deep learning based approach to real-time speech enhancement for dual-microphone mobile phones. The proposed approach employs a new densely-connected convolutional recurrent network to perform dual-channel complex spectral mapping. We utilize a structured pruning technique to compress the model without significantly degrading the enhancement performance, which yields a low-latency and memory-efficient enhancement system for real-time processing. Experimental results suggest that the proposed approach consistently outperforms an earlier approach to dual-channel speech enhancement for mobile phone communication, as well as a deep learning based beamformer. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Speakerfilter: Deep Learning-Based Target Speaker Extraction Using Anchor SpeechabstractSpeaker extraction aims to separate a target speaker from multiple voices which is useful for applications, e.g. teleconference. In many practical cases, it has an opportunity to get a piece voice of the target speaker in advance, which provides useful information for speaker extraction. This paper addresses the problem of extracting the target speaker from the mixture using a short piece of anchor speech. To effectively utilize anchor speech, we propose a multi-level feature extraction and seamlessly integrate the features into a speech separation model. Experiments are conducted on the two-speaker dataset (WSJ0-mix2) which is widely used for speaker extraction. The systematic evaluation shows that the proposed method significantly outperforms the previous methods and achieves a signal-to-distortion ratio (SDR) improvement of 11.3 dB on the unprocessed mixture. Shulin He, Hao Li 0046, Xueliang Zhang 0001 |
ICASSP | 3 |
| 2020 | Beamformed Feature for Learning-based Dual-channel Speech SeparationabstractThis paper deals with the problem of separating target speech signal from reverberant and noisy environment with dual microphones, where the target speech comes from a predefined direction range. First, we apply two differential beamformers with opposite directions to dual-channel inputs. Then, the power spectra of beamforming outputs are used as input feature of deep learning architecture. As input features, the beamformer outputs reflect not only spectral information but also directional information by their power level difference. And the calculation is very simple. Systematic evaluation and comparison show that the proposed system achieves very good separation performance and substantially outperforms related algorithms under very challenging environments where both interfering speaker, noise and reverberations are present. Hao Li 0046, Xueliang Zhang 0001, Guanglai Gao |
ICASSP | 2 |
| 2020 | An Efficient Joint Training Framework for Robust Small-Footprint Keyword Spotting
Zhihao Du, Hui Zhang 0031, Xueliang Zhang 0001 |
ICONIP (1) | 4 |
| 2020 | Double Adversarial Network Based Monaural Speech Enhancement for Robust Speech Recognition
Zhihao Du, Jiqing Han 0001, Xueliang Zhang 0001 |
INTERSPEECH | 3 |
| 2020 | Frame-Level Signal-to-Noise Ratio Estimation Using Deep Learning
Hao Li 0046, DeLiang Wang, Xueliang Zhang 0001, Guanglai Gao |
INTERSPEECH | 3 |
| 2020 | Polishing the Classical Likelihood Ratio Test by Supervised Learning for Voice Activity Detection
Tianjiao Xu, Hui Zhang 0031, Xueliang Zhang 0001 |
INTERSPEECH | 3 |
| 2020 | A Robust and Cascaded Acoustic Echo Cancellation Based on Deep Learning
Chenggang Zhang, Xueliang Zhang 0001 |
INTERSPEECH | 2 |
| 2020 | A Joint Framework of Denoising Autoencoder and Generative Vocoder for Monaural Speech EnhancementabstractConventional monaural speech enhancement methods usually enhance the magnitude spectrum of noisy speech and leave the phase unchanged. Recent studies suggest that phase is also important for both speech intelligibility and perceptual quality. Although deep learning exhibits great potential on enhancing the magnitude and phase spectra in complex spectrogram domain and waveform domain, complex spectrogram and waveform are always more difficult to predict than the magnitude spectrum due to lack of clear structure in them. In this study, a Mel-domain denoising autoencoder and a deep generative vocoder are stacked to form a joint framework for monaural speech enhancement, in which the clean speech waveform is reconstructed without using the phase. Specifically, a convolutional recurrent network (CRN) is employed as the denoising autoencoder to enhance the Mel power spectrum of noisy speech. Then, the enhanced Mel power spectrum is fed to a deep generative vocoder to synthesize the speech waveform. Furthermore, the denoising autoencoder and generative vocoder are jointly fine-tuned. Experimental results show that the proposed method significantly improves speech intelligibility and perceptual quality. More importantly, our method achieves much better generalization ability for untrained noises than previous methods. Zhihao Du, Xueliang Zhang 0001, Jiqing Han 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Supervised Speech Enhancement with Real Spectrum ApproximationabstractSpeech enhancement aims to separate a target speech from background noise. Recently, speech enhancement has been formulated as a supervised learning problem, in which a learning machine is trained to estimate the target spectrum denoted as mapping-based method or a time-frequency mask denoted as masking-based method. Signal approximation methods indirectly estimate the target spectrum via the mask estimation, which combines the advantages of both mapping based and masking based methods. Moreover, conventional methods usually ignore the phase which is also important to the speech quality. To consider the phase, the complex number spectrum needs to be modeled. However, modeling may be difficult. In this work, a pure real number spectrum is used as an alternative representation of the complex number spectrum, and a signal approximation method is used for speech enhancement. Experimental results show that the proposed method outperforms other commonly used methods. Hui Zhang 0031, Xueliang Zhang 0001, Linju Yang |
ICASSP | 3 |
| 2019 | Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk ScenariosabstractIn mobile speech communication, the quality and intelligibility of the received speech can be severely degraded by background noise if the far-end talker is in an adverse acoustic environment. Therefore, speech enhancement algorithms are typically integrated into mobile phones to remove background noise. In this paper, we propose a novel deep learning based framework for real-time speech enhancement on dual-microphone mobile phones in a close-talk scenario. It incorporates a convolutional recurrent network (CRN) with high computational efficiency. In addition, the framework amounts to a causal system, which is necessary for real-time processing on mobile phones. We find that the proposed approach consistently outperforms a deep neural network (DNN) based method, as well as two traditional methods for speech enhancement. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 2 |
| 2019 | A Robust Text-independent Speaker Verification Method Based on Speech Separation and Deep SpeakerabstractRecently, deep neural networks (DNNs) have achieved incredible performance in speaker verification. However, most of which remains sensitive to environment noise. In this paper, we propose an end-to-end speaker verification framework to enhance the robustness against background noise. The proposed framework first utilizes convolutional recurrent network (CRN) to address speech separation. Then the output of the middle layer of the CRN is used as the auxiliary feature, and together with the robust Filter banks (Fbanks) feature of noisy speech are fed to the speaker verification system. The speech separation and speaker verification are jointly optimized. Compared with deep speaker and DNN/i-vector, systematic evaluation indicates that the proposed algorithm can obtain a better performance in noisy conditions. Hao Li 0046, Xueliang Zhang 0001 |
ICASSP | 3 |
| 2019 | Investigation of Cost Function for Supervised Monaural Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001, Yuhang Cao |
INTERSPEECH | 3 |
| 2019 | Robust Speaker Localization Guided by Deep Learning-Based Time-Frequency MaskingabstractDeep learning-based time-frequency (T-F) masking has dramatically advanced monaural (single-channel) speech separation and enhancement. This study investigates its potential for direction of arrival (DOA) estimation in noisy and reverberant environments. We explore ways of combining T-F masking and conventional localization algorithms, such as generalized cross correlation with phase transform, as well as newly proposed algorithms based on steered-response SNR and steering vectors. The key idea is to utilize deep neural networks (DNNs) to identify speech dominant T-F units containing relatively clean phase for DOA estimation. Our DNN is trained using only monaural spectral information, and this makes the trained model directly applicable to arrays with various numbers of microphones arranged in diverse geometries. Although only monaural information is used for training, experimental results show strong robustness of the proposed approach in new environments with intense noise and room reverberation, outperforming traditional DOA estimation methods by large margins. Our study also suggests that the ideal ratio mask and its variants remain effective training targets for robust speaker localization. Zhongqiu Wang 0001, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Online Direction of Arrival Estimation Based on Deep LearningabstractDirection of arrival (DOA) estimation is an important topic in microphone array processing. Conventional methods work well in relatively clean conditions but suffer from noise and reverberation distortions. Recently, deep learning-based methods show the robustness to noise and reverberation. However, the performance is degraded rapidly or even model cannot work when microphone array structure changes. So it has to retrain the model with new data, which is a huge work. In this paper, we propose a supervised learning algorithm for DOA estimation combining convolutional neural network (CNN) and long short term memory (LSTM). Experimental results show that the proposed method can improve the accuracy significantly. In addition, due to an input feature design, the proposed method can adapt to a new microphone array conveniently only use a very small amount of data. Xueliang Zhang 0001, Hao Li 0046 |
ICASSP | 2 |
| 2018 | Training Supervised Speech Separation System to Improve STOI and PESQ DirectlyabstractSupervised speech separation methods train learning machine to cast the noisy speech to the target clean speech. Most of them use mean-square error (MSE) as loss function. However, MSE is not the perfect choice because it doesn't match the human auditory perception. Short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ) are closely related to the human auditory perception and widely used in speech separation research as evaluation criteria. Therefore, STOI and PESQ may be better choices for the loss function. However, they are nondifferentiable functions which cannot be optimized by the conventional gradient descent algorithm. In this work, a gradient approximation method is used to calculate the gradients of the STOI and PESQ. Then the calculated gradients are used in the gradient descent algorithm to optimize the STOI and PESQ directly. Experimental results show the speech separation performance can be improved by the proposed method. Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao |
ICASSP | 2 |
| 2018 | Using Shifted Real Spectrum Mask as Training Target for Supervised Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001 |
INTERSPEECH | 3 |
| 2018 | Robust TDOA Estimation Based on Time-Frequency Masking and Deep Neural Networks
Zhongqiu Wang 0001, Xueliang Zhang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | Deep Learning Based Speech Separation via NMF-Style ReconstructionsabstractDeep learning based speech separation usually uses a supervised algorithm to learn a mapping function from noisy features to separation targets. These separation targets, either ideal masks or magnitude spectrograms, have prominent spectro-temporal structures. Nonnegative matrix factorization (NMF) is a well-known representation learning technique that is capable of capturing the basic spectral structures. Therefore, the combination of deep learning and NMF as an organic whole is a smart strategy. However, previous methods typically use deep neural networks (DNN) and NMF for speech separation in a separate manner. In this paper, we propose a jointly combinatorial scheme to concentrate the strengths of both DNN and NMF for speech separation. NMF is used to learn the basis spectra that then are integrated into a DNN to directly reconstruct the magnitude spectrograms of speech and noise. Instead of predicting activation coefficients inferred by NMF, which is used as an intermediate target by the previous methods, DNN directly optimizes an actual separation objective in our system, so that the accumulated errors could be alleviated. Moreover, we explore a discriminative training objective with sparsity constraints to suppress noise and preserve more speech components further. Systematic experiments show that the proposed models are competitive with the previous methods. Shuai Nie 0001, Shan Liang 0001, Xueliang Zhang 0001, Jianhua Tao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | A speech enhancement algorithm by iterating single- and multi-microphone processing and its application to robust ASRabstractWe propose a speech enhancement algorithm based on single- and multi-microphone processing techniques. The core of the algorithm estimates a time-frequency mask which represents the target speech and use masking-based beamforming to enhance corrupted speech. Specifically, in single-microphone processing, the received signals of a microphone array are treated as individual signals and we estimate a mask for the signal of each microphone using a deep neural network (DNN). With these masks, in multi-microphone processing, we calculate a spatial covariance matrix of noise and steering vector for beamforming. In addition, we propose a masking-based post-filter to further suppress the noise in the output of beamforming. Then, the enhanced speech is sent back to DNN for mask re-estimation. When these steps are iterated for a few times, we obtain the final enhanced speech. The proposed algorithm is evaluated as a frontend for automatic speech recognition (ASR) and achieves a 5.05% average word error rate (WER) on the real environment test set of CHiME-3, outperforming the current best algorithm by 13.34%. Xueliang Zhang 0001, Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 1 |
| 2017 | Binaural Reverberant Speech Separation Based on Deep Neural Networks
Xueliang Zhang 0001, DeLiang Wang |
INTERSPEECH | 1 |
| 2017 | Multi-Target Ensemble Learning for Monaural Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao |
INTERSPEECH | 2 |
| 2017 | Deep Learning Based Binaural Speech Separation in Reverberant EnvironmentsabstractSpeech signal is usually degraded by room reverberation and additive noises in real environments. This paper focuses on separating target speech signal in reverberant conditions from binaural inputs. Binaural separation is formulated as a supervised learning problem, and we employ deep learning to map from both spatial and spectral features to a training target. With binaural inputs, we first apply a fixed beamformer and then extract several spectral features. A new spatial feature is proposed and extracted to complement the spectral features. The training target is the recently suggested ideal ratio mask. Systematic evaluations and comparisons show that the proposed system achieves very good separation performance and substantially outperforms related algorithms under challenging multi-source and reverberant environments. Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separationabstractThe targets of speech separation, whether ideal masks or magnitude spectrograms of interest, have prominent spectro-temporal structures. These characteristics are very worthy to be exploited for speech separation, however, they are usually ignored in previous works. In this paper, we use nonnegative matrix factorization (NMF) to exploit the spectro-temporal structures of magnitude spectrograms. With nonnegative constrains, NMF can capture the basis spectra patterns of speech and noise. Then the learned basis spectra are integrated into a deep neural network (DNN) to reconstruct the magnitude spectrograms of speech and noise with their nonnegative linear combination. Using the reconstructed spectrograms, we further explore a discriminative training objective and a joint optimization framework for the proposed model. Systematic experiments show that the proposed model is competitive with the previous methods in monaural speech separation tasks. Shuai Nie 0001, Hao Li 0046, Xueliang Zhang 0001, Zhanlei Yang, Like Dong |
ICASSP | 4 |
| 2016 | Convolutional neural network for robust pitch determinationabstractPitch is an important characteristic of speech and is useful for many applications. However, pitch determination in noisy conditions is difficult. In this paper, we propose a supervised learning algorithm to estimate pitch using a convolutional neural network (CNN). Specifically, we use a CNN for pitch candidate selection, and dynamic programming for pitch tracking. Our experimental results show that the proposed method can obtain accurate pitch estimation and they show good generalization ability to new speakers and noisy conditions. We credit the success to the use of CNN, which is suitable for modeling the shift-invariant spectral feature for pitch detection. Hong Su, Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao |
ICASSP | 3 |
| 2016 | Jointly Optimizing Activation Coefficients of Convolutive NMF Using DNN for Speech Separation
Hao Li 0046, Shuai Nie 0001, Xueliang Zhang 0001, Hui Zhang 0031 |
INTERSPEECH | 3 |
| 2016 | A Pairwise Algorithm Using the Deep Stacking Network for Speech Separation and Pitch EstimationabstractSpeech separation and pitch estimation in noisy conditions are considered to be a “chicken-and-egg” problem. On one hand, pitch information is an important cue for speech separation. On the other hand, speech separation makes pitch estimation easier when background noise is removed. In this paper, we propose a supervised learning architecture to solve these two problems iteratively. The proposed algorithm is based on the deep stacking network (DSN), which provides a method for stacking simple processing modules to build deep architectures. Each module is a classifier whose target is the ideal binary mask (IBM), and the input vector includes spectral features, pitch-based features and the output from the previous module. During the testing stage, we estimate the pitch using the separation results and update the pitch-based features to the next module. When embedded into the DSN, pitch estimation and speech separation each run several times. We obtain the final results from the last module. Systematic evaluations show that the proposed system results in both a high quality estimated binary mask and accurate pitch estimation and outperforms recent systems in its generalization ability. Xueliang Zhang 0001, Hui Zhang 0031, Shuai Nie 0001, Guanglai Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | A pairwise algorithm for pitch estimation and speech separation using deep stacking networkabstractPitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on deep stacking network (DSN) which provides a method of stacking simple processing modules in building deep architecture. In the training stage, an ideal binary mask is used as target. The input vector includes the outputs of lower module and frame-level features which consist of spectral and pitch-based features. In the testing stage, each module provides an estimated binary mask which is employed to re-estimate pitch. Then we update the pitch-based features to the next module. This procedure is embedded iteratively in DSN, and we obtain the final separation results from the last module of DSN. Systematic evaluations show that the proposed approach produces high quality estimated binary mask and outperforms recent systems in generalization. Hui Zhang 0031, Xueliang Zhang 0001, Shuai Nie 0001, Guanglai Gao |
ICASSP | 2 |
| 2015 | Two-stage multi-target joint learning for monaural speech separationabstractRecently, supervised speech separation has been extensively studied and shown considerable promise. Due to the temporal continuity of speech, speech auditory features and separation targets present prominent spectro-temporal structures and strong correlations over the time-frequency (T-F) domain, which can be exploited for speech separation. However, many supervised speech separation methods independently model each T-F unit with only one target and much ignore these useful information. In this paper, we propose a two-stage multi-target joint learning method to jointly model the related speech separation targets at the frame level. Systematic experiments show that the proposed approach consistently achieves better separation and generalization performances in the low signal-to-noise ratio(SNR) conditions. Shuai Nie 0001, Xueliang Zhang 0001, Like Dong |
INTERSPEECH | 4 |
| 2015 | Joint optimization of recurrent networks exploiting source auto-regression for source separationabstractIn music interferences condition, source separation is very difficult. In this paper, we propose a novel recurrent network exploiting the auto-regressions of speech and music interference for source separation. An auto-regression can capture the shortterm temporal dependencies in data to help the source separation. For the separation, we independently separate the magnitude spectra of speech and interference from the mixture spectra by including an extra masking layer in the recurrent network. Compared to directly evaluating the ideal mask, the extra masking layer relaxes the assumption of independence between speech and interference which is more suitable for the realworld environments. Using the separated spectra of speech and interference, we further explore a discriminative training objective and joint optimization framework for the proposed network, which incorporates the correlations and spectral dependencies of speech and interference into the separation. Systematic experiments show that the proposed model is competitive with the state-of-the-art method in singing-voice separations. Shuai Nie 0001, Xueliang Zhang 0001, Liwei Qiao |
INTERSPEECH | 4 |
| 2014 | Deep stacking networks with time series for speech separationabstractIn many present speech separation approaches, the separation task is formulated as a binary classification problem. Several classification-based approaches have been proposed and performed satisfactorily. However, they do not explicitly model the correlation in time and each time-frequency (T-F) unit is still classified individually. As we know, the speech signal has a very rich time series and temporal dynamic information that can be exploited for speech separation. In this study, we incorporate the correlation in time into classification. Compared with the previous approaches, the proposed approach achieves better separation and generalization performance by using deep stacking networks (DSN) with time series and re-threshold method. Shuai Nie 0001, Hui Zhang 0031, Xueliang Zhang 0001 |
ICASSP | 3 |
| 2011 | Monaural Voiced Speech Segregation Based on Pitch and Comb FilterabstractThe correlogram is an important mid-level representation for periodic sounds which is widely used in sound source separation and pitch detection. However, it is very time consuming. In this paper, we presented a novel scheme for monaural voiced speech separation without computing correlograms. The noisy speech is firstly decomposing into time-frequency units. Pitch contour of the target speech is extracted according to the zero crossing rate of the units. Then we applied a comb filter to label each unit as target speech or intrusion. Compared with previous correlogrambased method, the proposed algorithm saves computing time and also yields better performance. Index Terms—Sound separation, Computational auditory scene analysis, Correlogram Xueliang Zhang 0001 |
INTERSPEECH | 1 |
| 2011 | Monaural voiced speech segregation based on elaborate harmonic grouping strategies
Xueliang Zhang 0001, Wei Jiang 0030, Peng Li 0030, Bo Xu 0002 |
Sci. China Inf. Sci. | 2 |
| 2009 | Monaural voiced speech segregation based on elaborate harmonic grouping strategyabstractMonaural speech segregation is a very challenging problem which has been studied by many researchers. In this paper, we focus on voiced speech segregation. Different strategies are used to segregate resolved and unresolved harmonics respectively. For resolved harmonics, “harmonicity” principle and a novel mechanism based on “minimum amplitude” principle are employed. Amplitude modulation rate is extracted by “enhanced” autocorrelation function of envelope to segregate unresolved harmonics which is more robust than previous method. An elaborate rule is also introduced to determine the regions dominated by resolved and by unresolved harmonics. Proposed algorithm is evaluated on Cooke's 100 mixtures and compared with a state-of-the-art algorithm Hu and Wang model. Results show that proposed algorithm is more robust than the Hu and Wang model. Xueliang Zhang 0001, Peng Li 0030, Bo Xu 0002 |
ICASSP | 1 |