Shengkui Zhao

dblp:48/1104 · DBLP profile ↗
← Back
40ranked-venue papers
22as first author
18since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 21 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 8 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
abstract
The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent representations and poor speech quality, especially in out-of-domain scenarios. In this work, we propose HiFi-SR, a unified network that leverages end-to-end adversarial training to achieve high-fidelity speech super-resolution. Our model features a unified transformer-convolutional generator designed to seamlessly handle both the prediction of latent representations and their conversion into time-domain waveforms. The transformer network serves as a powerful encoder, converting low-resolution mel-spectrograms into latent space representations, while the convolutional network upscales these representations into high-resolution waveforms. To enhance high-frequency fidelity, we incorporate a multi-band, multi-scale time-frequency discriminator, along with a multi-scale mel-reconstruction loss in the adversarial training process. HiFi-SR is versatile, capable of upscaling any input speech signal between 4 kHz and 32 kHz to a 48 kHz sampling rate. Experimental results demonstrate that HiFi-SR significantly outperforms existing speech SR methods across both objective metrics and ABX preference tests, for both in-domain and out-of-domain scenarios.
Shengkui Zhao, Kun Zhou 0003, Zexu Pan, Chong Zhang 0003, Bin Ma 0001
ICASSP1
2025 Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning
abstract
Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or spectral domains, leading to increased generation complexity and slower inference speeds. Additionally, these methods have primarily modelled clean speech distributions, with limited exploration of noise distributions, thereby constraining the discriminative capability of diffusion models for speech enhancement. To address these issues, we propose a novel approach that integrates a conditional latent diffusion model (cLDM) with dual-context learning (DCL). Our method utilizes a variational autoencoder (VAE) to compress mel-spectrograms into a low-dimensional latent space. We then apply cLDM to transform the latent representations of both clean speech and background noise into Gaussian noise by the DCL process, and a parameterized model is trained to reverse this process, conditioned on noisy latent representations and text embeddings. By operating in a lower-dimensional space, the latent representations reduce the complexity of the generation process, while the DCL process enhances the model’s ability to handle diverse and unseen noise environments. Our experiments demonstrate the strong performance of the proposed approach compared to existing diffusion-based methods, even with fewer iterative steps, and highlight the superior generalization capability of our models to out-of-domain noise datasets.
Shengkui Zhao, Zexu Pan, Kun Zhou 0003, Chong Zhang 0003, Bin Ma 0001
ICASSP1
2025 Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang 0003, Kun Zhou 0003, Bin Ma 0001
INTERSPEECH3
2025 Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Zexu Pan, Shengkui Zhao, Kun Zhou 0003, Chong Zhang 0003, Bin Ma 0001
INTERSPEECH2
2025 ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
Shengkui Zhao, Zexu Pan, Bin Ma 0001
INTERSPEECH1
2024 Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?
abstract
Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not many people understand how and why this is so. In this study, we aim to deepen our understanding of this emerging method by investigating the role of soft prompts in automatic speech recognition (ASR). Our findings highlight their role as zero-shot learners in improving ASR performance while also exposing them to the risk of malicious modifications. Soft prompts aid generalization but are not obligatory for inference. We also identify two primary roles of soft prompts: content refinement and noise information enhancement, which enhances robustness against background noise. Additionally, we propose an effective modification on noise prompts to show that they are capable of zero-shot learning on adapting to out-of-distribution noise environments.
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Fabian Ritter Gutierrez, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001
ICASSP8
2024 SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance
abstract
Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model’s parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model’s total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm
Jia Qi Yip, Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Dianwen Ng, Chng Eng Siong, Bin Ma 0001
ICASSP2
2024 MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
abstract
Our previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale recurrent patterns. In this paper, we introduce a novel hybrid model that provides the capabilities to model both long-range, coarse-scale dependencies and fine-scale recurrent patterns by integrating a recurrent module into the MossFormer framework. Instead of applying the recurrent neural networks (RNNs) that use traditional recurrent connections, we present a recurrent module based on a feedforward sequential memory network (FSMN), which is considered "RNN-free" recurrent network due to the ability to capture recurrent patterns without using recurrent connections. Our recurrent module mainly comprises an enhanced dilated FSMN block by using gated convolutional units (GCU) and dense connections. In addition, a bottleneck layer and an output layer are also added for controlling information flow. The recurrent module relies on linear projections and convolutions for seamless, parallel processing of the entire sequence. The integrated MossFormer2 hybrid model demonstrates remarkable enhancements over MossFormer and surpasses other state-of-the-art methods in WSJ0-2/3mix, Libri2Mix, and WHAM!/WHAMR! benchmarks.
Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Jia Qi Yip, Dianwen Ng, Bin Ma 0001
ICASSP1
2024 Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
Kun Zhou 0003, Shengkui Zhao, Chong Zhang 0003, Hao Wang 0199, Dianwen Ng, Chongjia Ni, Trung Hieu Nguyen 0001, Jia Qi Yip, Bin Ma 0001
INTERSPEECH2
2024 Towards Audio Codec-based Speech Separation
Jia Qi Yip, Shengkui Zhao, Dianwen Ng, Chng Eng Siong, Bin Ma 0001
INTERSPEECH2
2023 D2Former: A Fully Complex Dual-Path Dual-Decoder Conformer Network Using Joint Complex Masking and Complex Spectral Mapping for Monaural Speech Enhancement
abstract
Monaural speech enhancement has been widely studied using real networks in the time-frequency (TF) domain. However, the input and the target are naturally complex-valued in the TF domain, a fully complex network is highly desirable for effectively learning the feature representation and modelling the sequence in the complex domain. Moreover, phase, an important factor for perceptual quality of speech, has been proved learnable together with magnitude from noisy speech using complex masking or complex spectral mapping. Many recent studies focus on either complex masking or complex spectral mapping, ignoring their performance boundaries. To address above issues, we propose a fully complex dual-path dual-decoder conformer network (D2Former) using joint complex masking and complex spectral mapping for monaural speech enhancement. In D2Former, we extend the conformer network into the complex domain and form a dual-path complex TF self-attention architecture for effectively modelling the complex-valued TF sequence. We further boost the TF feature representation in the encoder and the decoders using a dual-path learning structure by exploiting complex dilated convolutions on time dependency and complex feedforward sequential memory networks (CFSMN) for frequency recurrence. In addition, we improve the performance boundaries of complex masking and complex spectral mapping by combining the strengths of the two training targets into a joint-learning framework. As a consequence, D2Former takes fully advantages of the complex-valued operations, the dual-path processing, and the joint-training targets. Compared to the previous models, D2Former achieves state-of-the-art results on the VoiceBank+Demand benchmark with the smallest model size of 0.87M parameters.
Shengkui Zhao, Bin Ma 0001
ICASSP1
2023 MossFormer: Pushing the Performance Limit of Monaural Speech Separation Using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions
abstract
Transformer based models have provided significant performance improvements in monaural speech separation. However, there is still a performance gap compared to a recent proposed upper bound. The major limitation of the current dual-path Transformer models is the inefficient modelling of long-range elemental interactions and local feature patterns. In this work, we achieve the upper bound by proposing a gated single-head transformer architecture with convolution-augmented joint self-attentions, named MossFormer (Monaural speech separation TransFormer). To effectively solve the indirect elemental interactions across chunks in the dual-path architecture, MossFormer employs a joint local and global self-attention architecture that simultaneously performs a full-computation self-attention on local chunks and a linearised low-cost self-attention over the full sequence. The joint attention enables MossFormer model full-sequence elemental interaction directly. In addition, we employ a powerful attentive gating mechanism with simplified single-head self-attentions. Besides the attentive long-range modelling, we also augment MossFormer with convolutions for the position-wise local pattern modelling. As a consequence, MossFormer significantly outperforms the previous models and achieves the state-of-the-art results on WSJ0-2/3mix and WHAM!/WHAMR! benchmarks. Our model achieves the SI-SDRi upper bound of 21.2 dB on WSJ0-3mix and only 0.3 dB below the upper bound of 23.1 dB on WSJ0-2mix.
Shengkui Zhao, Bin Ma 0001
ICASSP1
2023 Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001
INTERSPEECH7
2023 ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
Jia Qi Yip, Duc-Tuan Truong, Dianwen Ng, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001
INTERSPEECH8
2022 End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression
abstract
Echo and noise suppression is an integral part of a full-duplex communication system. Many recent acoustic echo cancellation (AEC) systems rely on a separate adaptive filtering module for linear echo suppression and a neural module for residual echo suppression. However, in practice, adaptive filtering modules require time to converge and remain susceptible to changes in the acoustic environment. This introduces unnecessary delays to AEC systems using this two-stage framework, despite neural modules already having the capability to suppress both linear and nonlinear echo components. In this paper, we exploit the offset-compensating property of complex time-frequency masks and propose an end-to-end complex-valued neural network architecture. The building block of the proposed model is a pseudocomplex extension of the densely-connected multidilated DenseNet (D3Net), resulting in a very small network of only 354K parameters. The architecture utilized the multi-resolution nature of the D3Net to eliminate the need for pooling, allowing feature extraction using large receptive fields without any loss of output resolution. We also propose a dual-mask technique for joint echo and noise suppression with simultaneous speech enhancement. Evaluation on both synthetic and real test sets demonstrated promising results across multiple energy-based metrics and perceptual proxies.
Karn Watcharasupat, Thi Ngoc Tho Nguyen, Woon-Seng Gan, Shengkui Zhao, Bin Ma 0001
ICASSP4
2022 FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement
abstract
Convolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across frequency context is highly constrained due to limited receptive fields in the convolutions of CED. In this paper, we propose a convolutional recurrent encoder-decoder (CRED) structure to boost feature representation along the frequency axis. The CRED applies frequency recurrence on 3D convolutional feature maps along the frequency axis following each convolution, therefore, it is capable of catching long-range frequency correlations and enhancing feature representations of speech inputs. The proposed frequency recurrence is realized efficiently using a feedforward sequential memory network (FSMN). Besides the CRED, we insert two stacked FSMN layers between the encoder and the decoder to model further temporal dynamics. We name the proposed framework as Frequency Recurrent CRN (FRCRN). We design FRCRN to predict complex Ideal Ratio Mask (cIRM) in complex-valued domain and optimize FRCRN using both time-frequency-domain and time-domain losses. Our proposed approach achieved state-of-the-art performance on wideband bench-mark datasets and achieved 2nd place for the real-time fullband track in terms of Mean Opinion Score (MOS) and Word Accuracy (WAcc) in the ICASSP 2022 Deep Noise Suppression (DNS) challenge.
Shengkui Zhao, Bin Ma 0001, Karn Watcharasupat, Woon-Seng Gan
ICASSP1
2021 Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
abstract
Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections, which heavily rely on the representation power of the complex-valued convolutional layers. In this paper, we propose a complex convolutional block attention module (CCBAM) to boost the representation power of the complex-valued convolutional layers by constructing more informative features. The CCBAM is a lightweight and general module which can be easily integrated into any complex-valued convolutional layers. We integrate CCBAM with the deep complex U-Net and CRN to enhance their performance for speech enhancement. We further propose a mixed loss function to jointly optimize the complex models in both time-frequency (TF) domain and time domain. By integrating CCBAM and the mixed loss, we form a new end-to-end (E2E) complex speech enhancement framework. Ablation experiments and objective evaluations show the superior performance of the proposed approaches.
Shengkui Zhao, Trung Hieu Nguyen 0001, Bin Ma 0001
ICASSP1
2021 Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram
abstract
Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model, i.e., FastSpeech, and LPCNet neural vocoder to design a new cross-lingual VC framework named FastSpeech-VC. We address the mismatches of the phonetic set and the speech prosody by applying Phonetic PosteriorGrams (PPGs), which have been proved to bridge across speaker and language boundaries. Moreover, we add normalized logarithm-scale fundamental frequency (Log-F0) to further compensate for the prosodic mismatches and significantly improve naturalness. Our experiments on English and Mandarin languages demonstrate that with only mono-lingual corpus, the proposed FastSpeech-VC can achieve high quality converted speech with mean opinion score (MOS) close to the professional records while maintaining good speaker similarity. Compared to the baselines using Tacotron2 and Transformer TTS models, the FastSpeech-VC can achieve controllable converted speech rate and much faster inference speed. More importantly, the FastSpeech-VC can easily be adapted to a speaker with limited training utterances.
Shengkui Zhao, Hao Wang 0199, Trung Hieu Nguyen 0001, Bin Ma 0001
ICASSP1
2020 Towards Natural Bilingual and Code-Switched Speech Synthesis Based on Mix of Monolingual Recordings and Cross-Lingual Voice Conversion
abstract
Recent state-of-the-art neural text-to-speech (TTS) synthesis models have dramatically improved intelligibility and naturalness of generated speech from text. However, building a good bilingual or code-switched TTS for a particular voice is still a challenge. The main reason is that it is not easy to obtain a bilingual corpus from a speaker who achieves native-level fluency in both languages. In this paper, we explore the use of Mandarin speech recordings from a Mandarin speaker, and English speech recordings from another English speaker to build high-quality bilingual and code-switched TTS for both speakers. A Tacotron2-based cross-lingual voice conversion system is employed to generate the Mandarin speaker's English speech and the English speaker's Mandarin speech, which show good naturalness and speaker similarity. The obtained bilingual data are then augmented with code-switched utterances synthesized using a Transformer model. With these data, three neural TTS models -- Tacotron2, Transformer and FastSpeech are applied for building bilingual and code-switched TTS. Subjective evaluation results show that all the three systems can produce (near-)native-level speech in both languages for each of the speaker.
Shengkui Zhao, Trung Hieu Nguyen 0001, Hao Wang 0199, Bin Ma 0001
INTERSPEECH1
2019 Multi-Task Multi-Network Joint-Learning of Deep Residual Networks and Cycle-Consistency Generative Adversarial Networks for Robust Speech Recognition
Shengkui Zhao, Chongjia Ni, Rong Tong, Bin Ma 0001
INTERSPEECH1
2019 Fast Learning for Non-Parallel Many-to-Many Voice Conversion with Residual Star Generative Adversarial Networks
Shengkui Zhao, Trung Hieu Nguyen 0001, Hao Wang 0199, Bin Ma 0001
INTERSPEECH1
2017 A novel sparse model for multi-source localization using distributed microphone array
abstract
When distances between microphone pairs are larger than the half-wavelength of signals, source localization methods using cross-correlation such as time-difference-of-arrival (TDOA), steered response power (SRP) are commonly used in practice. We present here a novel model that expresses microphone pairwise cross-correlations as a sum of autocorrelations of source signals shifted by the relative delays of the signals arriving at the microphone pairs, and weighted by the source power and the distances between the sources and the microphone pairs. The model is formulated as a linear inverse problem and is sparse with respect to the source power map. The source power map, which directly shows the locations of all the sound sources, can be reconstructed using ℓ1-norm minimization algorithms. We demonstrate the effectiveness of our model in a wildlife monitoring application, where the goal is to locate multiple frogs in a dense chorus.
Thi Ngoc Tho Nguyen, Cagdas Tuna, Shengkui Zhao, Douglas L. Jones
ICASSP3
2017 On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition
abstract
Acoustic beamforming has played a key role in the robust automatic speech recognition (ASR) applications. Accurate estimates of the speech and noise spatial covariance matrices (SCM) are crucial for successfully applying the minimum variance distortionless response (MVDR) beamforming. Reliable estimation of time-frequency (TF) masks can improve the estimation of the SCMs and significantly improve the performance of the MVDR beamforming in ASR tasks. In this paper, we focus on the TF mask estimation using recurrent neural networks (RNN). Specifically, our methods include training the RNN to estimate the speech and noise masks independently, training the RNN to minimize the ASR cost function directly, and performing multiple passes to iteratively improve the mask estimation. The proposed methods are evaluated individually and overally on the CHiME-4 challenge. The results show that the proposed methods improve the ASR performance individually and also work complementarily. The overall performance achieves a word error rate of 8.9% with 6-microphone configuration, which is much better than 12.0% achieved with the state-of-the-art MVDR implementation.
Shengkui Zhao, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
ICASSP2
2016 An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources
abstract
This paper presents an eigenvector clustering approach for estimating the direction of arrival (DOA) of multiple speech signals using a microphone array. Existing clustering approaches usually only use low frequencies to avoid spatial aliasing. In this study, we propose a probabilistic eigenvector clustering approach to use all frequencies. In our work, time-frequency (TF) bins dominated by only one source are first detected using a combination of noise-floor tracking, onset detection and coherence test. For each selected TF bin, the largest eigenvector of its spatial covariance matrix is extracted for clustering. A mixture density model is introduced to model the distribution of the eigenvectors, where each component distribution corresponds to one source and is parameterized by the source DOA. To use eigenvectors of all frequencies, the steering vectors of all frequencies of the sources are used in the distribution function. The DOAs of the sources can be estimated by maximizing the likelihood of the eigenvectors using an expectation-maximization (EM) algorithm. Simulation and experimental results show that the proposed approach significantly improves the root-mean-square error (RMSE) for DOA estimation of multiple speech sources compared to the MUSIC algorithm implemented on the single-source dominated TF bins and our previous clustering approach.
Shengkui Zhao, Thi Ngoc Tho Nguyen, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
ICASSP2
2016 Large region acoustic source mapping: A generalized sparse constrained deconvolution approach
abstract
This paper presents a generalized multiple-point sparse constrained deconvolution approach for mapping acoustic noise sources in large regions using a movable array. Extended from our previous MPSC-DAMAS approach, we first derive a generalized inverse problem relating to the source powers and the array manifold using a generic beamformer and an explicit measurement noise model. We then propose a generalized MPSC-DAMAS (GMPSC-DAMAS) approach for resolving the inverse problem. A new parameter setting method based on a multiple-point minimum-variance-distortionless-response (MVDR) beamformer is also presented. The realizations of the GMPSC-DAMAS approach using the delay- and-sum (DAS) beamformer and the MVDR beamformer are evaluated. Simulation results show the proposed GMPSC-DAMAS approach achieves much lower absolute power estimation errors and processing time than the MPSC-DAMAS approach in terms of number of sources and robustness to measurement noise.
Shengkui Zhao, Cagdas Tuna, Thi Ngoc Tho Nguyen, Douglas L. Jones
ICASSP1
2015 Robust speech recognition using beamforming with adaptive microphone gains and multichannel noise reduction
abstract
This paper presents a robust speech recognition system using a microphone array for the 3rd CHiME Challenge. A minimum variance distortionless response (MVDR) beamformer with adaptive microphone gains is proposed for robust beamforming. Two microphone gain estimation methods are studied using the speech-dominant time-frequency bins. A multichannel noise reduction (MCNR) postprocessing is also proposed to further reduce the interference in the MVDR processed signal. Experimental results for the ChiME-3 challenge show that both the proposed MVDR beamformer with microphone gains and the MCNR postprocessing improve the speech recognition performance significantly. With the state-of-the-art deep neural network (DNN) based acoustic model, our system achieves a word error rate (WER) of 11.67% on the real test data of the evaluation set.
Shengkui Zhao, Thi Ngoc Tho Nguyen, Xionghu Zhong, Bo Ren 0006, Longbiao Wang, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
ASRU1
2015 A learning-based approach to direction of arrival estimation in noisy and reverberant environments
abstract
This paper presents a learning-based approach to the task of direction of arrival estimation (DOA) from microphone array input. Traditional signal processing methods such as the classic least square (LS) method rely on strong assumptions on signal models and accurate estimations of time delay of arrival (TDOA) . They only work well in relatively clean conditions, but suffer from noise and reverberation distortions. In this paper, we propose a learning-based approach that can learn from a large amount of simulated noisy and reverberant microphone array inputs for robust DOA estimation. Specifically, we extract features from the generalised cross correlation (GCC) vectors and use a multilayer perceptron neural network to learn the nonlinear mapping from such features to the DOA. One advantage of the learning based method is that as more and more training data becomes available, the DOA estimation will become more and more accurate. Experimental results on simulated data show that the proposed learning based method produces much better results than the state-of-the-art LS method. The testing results on real data recorded in meeting rooms show improved root-mean-square error (RMSE) compared to the LS method.
Shengkui Zhao, Xionghu Zhong, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
ICASSP2
2015 Large region acoustic source mapping using movable arrays
abstract
Mapping environmental noise with high resolution on a large scale (such as a city) is prohibitively expensive with current approaches, which use a large, dense array spanning the entire region of interest, or sequential noise measurements at thousands of locations on a dense grid. We propose instead a new acoustic measurement scheme using a small movable array (for example, mounted on a vehicle driving along the streets of a city) to rapidly acquire measurements at many different locations. A multiple-point sparse constrained deconvolution approach for the mapping of acoustic sources (MPSC-DAMAS) and a multiple-point covariance matrix fitting (MP-CMF) approach are developed to accurately estimate the locations and powers of stationary noise sources across the region of interest. Computer simulations of large region acoustic mapping demonstrate that superior resolution and much lower power estimation errors are achieved by the proposed approaches compared to the state-of-the-art SC-DAMAS approach and CMF approach.
Shengkui Zhao, Thi Ngoc Tho Nguyen, Douglas L. Jones
ICASSP1
2015 Learning to estimate reverberation time in noisy and reverberant rooms
Shengkui Zhao, Xionghu Zhong, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH2
2014 Robust DOA estimation of multiple speech sources
abstract
It is challenging to determine the directions of arrival of speech signals when there are fewer sensors than sources, particularly in noisy and reverberant environments. The coherence test by Mohan et al. exploits the time-frequency sparseness of non-stationary speech signals to select more relevant time-frequency bins to estimate directions of arrival. With no prior knowledge about the incoming sources, this work proposes a combination of noise-floor tracking, onset detection and a coherence test to robustly identify time-frequency bins where only one source is dominant. After that, the largest eigenvectors of covariance matrices corresponding to these bins are clustered and the directions of arrival of the sources are estimated based on the cluster centroids. Simulation and experimental results show that this method is able to localize 8 sources with small errors using only 3 omnidirectional microphones. The proposed method is robust to background noise and reverberation.
Thi Ngoc Tho Nguyen, Shengkui Zhao, Douglas L. Jones
ICASSP2
2014 A new auxiliary-vector algorithm with conjugate orthogonality for speech enhancement
abstract
In this paper, we propose a new auxiliary-vector (AV) algorithm using the conjugate orthogonality for speech enhancement. When only a limited data record is available, the AV algorithm is the state-of-the-art for obtaining the minimumvariance-distortionless (MVDR) filter. However, the current AV algorithms suffer from convergence problems when applied to the speech enhancement. Based on the conjugate GramSchmidt process, we develop new auxiliary vectors that are conjugate orthogonal and apply them to the AV algorithm. The proposed conjugate AV algorithm converges to the optimal MVDR solution within finite steps no greater than the filter dimension. Theoretical analysis establishes formal convergence of the proposed conjugate AV algorithm. Our experiments using the synthetic and real speech data show favorites of the new proposal over the state-of-the-art approaches. Index Terms: speech enhancement, microphone arrays, correlation, convergence, adaptive signal processing
Shengkui Zhao, Douglas L. Jones
INTERSPEECH1
2014 Underdetermined direction of arrival estimation using acoustic vector sensor
Shengkui Zhao, Tigran Saluev, Douglas L. Jones
Signal Process.1
2013 Spatialized audio multiparty teleconferencing with commodity miniature microphone array
abstract
This paper presents a Spatialized Audio Multiparty Teleconferencing (SAMT) system with a radically new communication experience for group teleconferencing. The system includes our recently developed 3D audio technologies: 3D sound source localization (SSL) and 3D audio capture and reproduction using a low-cost and compact design microphone array. In essence, the SAMT system offers 3D audio capture capability and spatial audio perception with multiple participants at a site, which still falls short in teleconferencing solutions. In addition to being able to identify and automatically track the active speaker, the system allows more compelling visual presentation for effective communication. Requiring only a low-cost microphone array and a consumer depth camera, the proposed system runs reliably and comfortably in real time on a commodity laptop or desktop PC. With such a minimal deployment requirement, we present a variety of user experiences created by SAMT.
Shengkui Zhao, Tien Dung Vu, Douglas L. Jones, Minh N. Do
ACM Multimedia2
2012 Real-time implementation and performance optimization of 3D sound localization on GPUs
abstract
Real-time 3D sound localization is an important technology for various applications such as camera steering systems, robotics audition, and gunshot direction. 3D sound localization adds a new dimension, but also significantly increases the computational requirements. Real-time 3D sound localization continuously processes large volumes of data for each possible 3D direction and acoustic frequency range. Such highly demanding compute requirements outpace current CPU compute abilities. This paper develops a real-time implementation of 3D sound localization on Graphical Processing Units (GPUs). Massively parallel GPU architectures are shown to be well suited for 3D sound localization. We optimize various aspects of GPU implementation, such as number of threads per thread block, register allocation per thread, and memory data layout for performance improvement. Experiments indicate that our GPU implementation achieves 501X and 130X speedup compared to a single-thread and a multi-thread CPU implementation respectively, thus enabling real-time operation of 3D sound localization.
Yun Liang 0001, Shengkui Zhao, Kyle Rupnow, Douglas L. Jones, Deming Chen
DATE3
2012 A Fast-Converging Adaptive Frequency-Domain MVDR Beamformer for Speech Enhancement
abstract
In this paper, we present a fast-converging adaptive frequency-domain minimum-variance-distortionlessresponse (MVDR) beamformer (FMV) for speech enhancement. The well-known FMV solution is optimum in the microphone array processing. However, the direct computation of the optimum FMV solution is often undesirable due to the the inversion of the spatio-spectral correlation matrix which is often unstable and is expensive for large arrays. To avoid the matrix inversion, we develop a fast-converging conjugate gradient (CG) algorithm for iteratively computing the FMV solution. Compared to the existing steepest descent (SD) algorithm, the CG algorithm can dramatically improve the convergence speed for the case of multiple interfering signals in speech enhancement. Therefore, the computational load and processing time can be significantly reduced. The speech enhancement experiments using a four-channel acousticvector-sensor (AVS) microphone array are demonstrated for the target speech signal corrupted by two and five interfering speech signals and superior performance are achieved.
Shengkui Zhao, Douglas L. Jones
INTERSPEECH1
2010 Nonlinear image restoration using recurrent radial basis function network
abstract
For nonlinear distorted images, the performance of the existing image restoration methods is limited in either visual quality or computational complexity. In this paper, we apply the recently developed technique called recurrent radial basis function network (RBFN) for nonlinear image restoration. We give the details of the construction of the recurrent RBFN network and the determination of the network parameters. Simulation results show that the proposed recurrent RBFN scheme outperforms the existing RBFN based methods in both visual quality and complexity when the degraded process is recursive.
Shengkui Zhao, Jianfei Cai 0001, Zhihong Man
ISCAS1
2009 A generalized data windowing scheme for adaptive conjugate gradient algorithms
Shengkui Zhao, Zhihong Man, Suiyang Khoo
Signal Process.1
2009 Variable step-size LMS algorithm with a quotient form
Shengkui Zhao, Zhihong Man, Suiyang Khoo, Hong Ren Wu
Signal Process.1
2006 Sliding Mode Control of Fuzzy Dynamic Systems
abstract
In this paper, a sliding mode control scheme is developed for a class of complex nonlinear systems with their T-S fuzzy models. It is shown that a set of extreme fuzzy subsystems are first derived, and a constructive sliding mode control law is then developed to guarantee the stability of the closed-loop fuzzy system. Simulation results are presented in support of the proposed scheme
Suiyang Khoo, Zhihong Man, Shengkui Zhao
ICARCV3
2006 Modified LMS and NLMS Algorithms with a New Variable Step Size
abstract
In this paper, the modified LMS and NLMS algorithms with variable step-size are presented. It is shown that the variable step size is computed using a ratio of the sums of weighted energy of the output error with two exponential factors alpha and beta, thus the fast error convergence of the modified LMS and NLMS algorithms can then be achieved. Also, by properly choosing the values of alpha and beta, the misadjustment can be further improved. A few simulation results are presented in support of the good performance of the proposed algorithms by comparing with other LMS-type algorithms
Shengkui Zhao, Zhihong Man, Suiyang Khoo
ICARCV1