Shixiong Zhang 0001

dblp:19/906-1 · also Shi-Xiong Zhang 0001 · DBLP profile ↗
← Back
64ranked-venue papers
15as first author
33since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 12 first-author · 29 since 2021Artificial intelligence and machine learning · 34 · 7 first-author · 19 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2024 SECap: Speech Emotion Captioning with Large Language Model
abstract
Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests.
Yaoxun Xu, Hangting Chen, Jianwei Yu 0001, Qiaochu Huang, Zhiyong Wu 0001, Shixiong Zhang 0001, Guangzhi Li, Yi Luo 0004, Rongzhi Gu
AAAI6
2024 UniX-Encoder: A Universal X-Channel Speech Encoder for AD-HOC Microphone Array Speech Processing
abstract
The speech field is evolving to solve more challenging scenarios, such as multi-channel recordings with multiple simultaneous talkers. In response to the diversity of microphone configurations in use, we introduce the UniX-Encoder, a universal encoder for multi-channel speech recordings. The UniX-Encoder is versatile, catering to a variety of speech tasks, and seamlessly integrates with any microphone array, whether in single-talker or multi-talker environments. Our research enhances previous multi-channel speech processing efforts in four aspects: 1) Adaptability: Contrasting traditional models constrained to certain microphone array configurations, our encoder is universally compatible. 2) Multi-Task Capability: Contrasting previous systems that were designed for single-task applications, the UniX-Encoder serves as a versatile upstream model, capable of extracting features for diverse speech tasks. 3) Self-Supervised Training: The UniX-Encoder is pretrained without the need for labeled multi-channel data. 4) End-to-End Integration: In contrast to models that first beamform then process single-channels, our encoder offers an end-to-end solution, bypassing explicit beamforming or separation. To validate its effectiveness, we tested the UniX-Encoder on a synthetic multi-channel dataset from the LibriSpeech corpus. Across various tasks, including ASR and speaker diarization, our encoder consistently outperformed combinations such as the WavLM model with the BeamformIt frontend.
Zili Huang, Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
ICASSP3
2024 LibriheavyMix: A 20, 000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
Zengrui Jin, Yifan Yang 0005, Mohan Shi, Wei Kang 0006, Xiaoyu Yang 0005, Zengwei Yao, Liyong Guo, Lingwei Meng, Long Lin, Yong Xu 0004, Shixiong Zhang 0001, Daniel Povey
INTERSPEECH12
2024 RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios
Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH2
2024 Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Daniel Povey, Sanjeev Khudanpur
INTERSPEECH2
2024 LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
Zheshu Song, Jianheng Zhuo, Yifan Yang 0005, Ziyang Ma 0001, Shixiong Zhang 0001, Xie Chen 0001
INTERSPEECH5
2024 Comparing Discrete and Continuous Space LLMs for Speech Recognition
Yaoxun Xu, Shixiong Zhang 0001, Jianwei Yu 0001, Zhiyong Wu 0001, Dong Yu 0001
INTERSPEECH2
2024 Advancing Multi-Talker ASR Performance With Large Language Models
abstract
Recognizing overlapping speech from multiple speakers in conversational scenarios is one of the most challenging problem for automatic speech recognition (ASR). Serialized output training (SOT) is a classic method to address multi-talker ASR, with the idea of concatenating transcriptions from multiple speakers according to the emission times of their speech for training. However, SOT-style transcriptions, derived from concatenating multiple related utterances in a conversation, depend significantly on modeling long contexts. Therefore, compared to traditional methods that primarily emphasize encoder performance in attention-based encoderdecoder (AED) architectures, a novel approach utilizing large language models (LLMs) that leverages the capabilities of pre-trained decoders may be better suited for such complex and challenging scenarios. In this paper, we propose an LLM-based SOT approach for multi-talker ASR, leveraging pre-trained speech encoder and LLM, fine-tuning them on multi-talker dataset using appropriate strategies. Experimental results demonstrate that our approach surpasses traditional AED-based methods on the simulated dataset LibriMix and achieves state-of-the-art performance on the evaluation set of the real-world dataset AMI, outperforming the AED model trained with 1000 times more supervised data in previous works.
Mohan Shi, Zengrui Jin, Yaoxun Xu, Yong Xu 0004, Shixiong Zhang 0001, Yiwen Shao, Dong Yu 0001
SLT5
2023 Neuralecho: Hybrid of Full-Band and Sub-Band Recurrent Neural Network For Acoustic Echo Cancellation and Speech Enhancement
abstract
This paper presents a hybrid of full-band and sub-band recurrent neural network (RNN) model, named NeuralEcho, to jointly solve echo and noise suppression. The full-band model part processes the signal’s entire frequency bands as a whole, while the sub-band model part divides the features into sub-bands and processes each sub-band separately. This approach allows the model to capture both the fine-grained local details of the sub-band processing and the global context of the full-band processing. The single-channel model is then generalized to accommodate a range of input channel numbers. Experimental results show that the hybrid model outperforms the conventional full-band models in terms of objective speech quality metrics and speech recognition accuracy. This suggests that the hybrid approach of full-band and sub-band processing can be a promising direction for future research in the field of speech enhancement.
Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001
ASRU4
2023 Deep Neural Mel-Subband Beamformer for in-Car Speech Separation
abstract
While current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies independently which results in higher computational costs and inference times, making them unsuitable for real-world use. In this paper, we propose DL-based mel-subband spatio-temporal beamformer to perform speech separation in a car environment with reduced computation cost and inference time. As opposed to conventional subband (SB) approaches, our framework uses a mel-scale based subband selection strategy which ensures a fine-grained processing for lower frequencies where most speech formant structure is present, and coarse-grained processing for higher frequencies. In a recursive way, robust frame-level beamforming weights are determined for each speaker location/zone in a car from the estimated subband speech and noise covariance matrices. Furthermore, proposed framework also estimates and suppresses any echoes from the loudspeaker(s) by using the echo reference signals. We compare the performance of our proposed framework to several NB, SB, and full-band (FB) processing techniques in terms of speech quality and recognition metrics. Based on experimental evaluations on simulated and real-world recordings, we find that our proposed framework achieves better separation performance over all SB and FB approaches and achieves performance closer to NB processing techniques while requiring lower computing cost.
Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
ICASSP4
2023 MMCosine: Multi-Modal Cosine Loss Towards Balanced Audio-Visual Fine-Grained Learning
abstract
Audio-visual learning helps to comprehensively under-stand the world by fusing practical information from multiple modalities. However, recent studies show that the imbalanced optimization of uni-modal encoders in a joint-learning model is a bottleneck to enhancing the model’s performance. We further find that the up-to-date imbalance-mitigating methods fail on some audio-visual fine-grained tasks, which have a higher demand for distinguishable feature distribution. Fueled by the success of cosine loss that builds hyperspherical feature spaces and achieves lower intra-class angular variability, this paper proposes Multi-Modal Cosine loss, MMCosine. It performs a modality-wise L2normalization to features and weights towards balanced and better multi-modal fine-grained learning. We demonstrate that our method can alleviate the imbalanced optimization from the perspective of weight norm and fully exploit the discriminability of the cosine metric. Extensive experiments prove the effectiveness of our method and the versatility with advanced multi-modal fusion strategies and up-to-date imbalance-mitigating methods. The project page is https://gewu-lab.github.io/MMCosine/.
Ruize Xu, Ruoxuan Feng, Shixiong Zhang 0001, Di Hu 0001
ICASSP3
2023 Zoneformer: On-device Neural Beamformer For In-car Multi-zone Speech Separation, Enhancement and Echo Cancellation
Yong Xu 0004, Vinay Kothapally, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH4
2023 Towards Unified All-Neural Beamforming for Time and Frequency Domain Speech Separation
abstract
Recently, frequency domain all-neural beamforming methods have achieved remarkable progress for multichannel speech separation. In parallel, the integration of time domain network structure and beamforming also gains significant attention. This study proposes a novel all-neural beamforming method in time domain and makes an attempt to unify the all-neural beamforming pipelines for time domain and frequency domain multichannel speech separation. The proposed model consists of two modules: separation and beamforming. Both modules perform temporal-spectral-spatial modeling and are trained from end-to-end using a joint loss function. The novelty of this study lies in two folds. Firstly, a time domain directional feature conditioned on the direction of the target speaker is proposed, which can be jointly optimized within the time domain architecture to enhance target signal estimation. Secondly, an all-neural beamforming network in time domain is designed to refine the pre-separated results. This module features with parametric time-variant beamforming coefficient estimation, without explicitly following the derivation of optimal filters that may lead to an upper bound. The proposed method is evaluated on simulated reverberant overlapped speech data derived from the AISHELL-1 corpus. Experimental results demonstrate significant performance improvements over frequency domain state-of-the-arts, ideal magnitude masks and existing time domain neural beamforming methods.
Rongzhi Gu, Shixiong Zhang 0001, Yuexian Zou, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Fast-Rir: Fast Neural Diffuse Room Impulse Response Generator
abstract
We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RIR takes rectangular room dimensions, listener and speaker positions, and reverberation time (T60) as inputs and generates specular and diffuse reflections for a given acoustic environment. Our FAST-RIR is capable of generating RIRs for a given input T60with an average error of 0.02s. We evaluate our generated RIRs in automatic speech recognition (ASR) applications using Google Speech API, Microsoft Speech API, and Kaldi tools. We show that our proposed FAST-RIR with batch size 1 is 400 times faster than a state-of-the-art diffuse acoustic simulator (DAS) on a CPU and gives similar performance to DAS in ASR experiments. Our FAST-RIR is 12 times faster than an existing GPU-based RIR generator (gpuRIR). We show that our FAST-RIR outperforms gpuRIR by 2.5% in an AMI far-field ASR benchmark.
Anton Ratnarajah, Shixiong Zhang 0001, Meng Yu 0003, Zhenyu Tang 0001, Dinesh Manocha, Dong Yu 0001
ICASSP2
2022 Multi-Channel Multi-Speaker ASR Using 3D Spatial Feature
abstract
Automatic speech recognition (ASR) of multi-channel multi-speaker overlapped speech remains one of the most challenging tasks to the speech community. In this paper, we look into this challenge by utilizing the location information of target speakers in the 3D space for the first time. To explore the strength of proposed the 3D spatial feature, two paradigms are investigated. 1) a pipelined system with a multi-channel speech separation module followed by the state-of-the-art single-channel ASR module; 2) a "All-In-One" model where the 3D spatial feature is directly used as an input to ASR system without explicit separation modules. Both of them are fully differentiable and can be back-propagated end-to-end. We test them on simulated overlapped speech and real recordings. Experimental results show that 1) the proposed ALL-In-One model achieved a comparable error rate to the pipelined system while reducing the inference time by half; 2) the proposed 3D spatial feature significantly outperformed (31% CERR) all previous works of using the 1D directional information in both paradigms.
Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
ICASSP2
2022 Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI
abstract
Recently, End-to-End (E2E) frameworks have achieved remarkable results on various Automatic Speech Recognition (ASR) tasks. However, Lattice-Free Maximum Mutual Information (LF-MMI), as one of the discriminative training criteria that show superior performance in hybrid ASR systems, is rarely adopted in E2E ASR frameworks. In this work, we propose a novel approach to introduce LF-MMI criterion into E2E ASR frameworks in both training and decoding stages. The proposed approach shows its effectiveness on two of the most widely used E2E frameworks including Attention-Based Encoder-Decoders (AEDs) and Neural Transducers (NTs). Experiments suggest that the introduction of the LF-MMI criterion consistently leads to significant performance improvements on various datasets and different E2E ASR frameworks. The best of our models achieves competitive CER of 4.1% / 4.4% on Aishell-1 dev/test set; significant error reduction is also achieved on Aishell-2 and Librispeech datasets over strong baselines. Code is released1.
Jinchuan Tian, Jianwei Yu 0001, Chao Weng, Shixiong Zhang 0001, Dan Su 0002, Dong Yu 0001, Yuexian Zou
ICASSP4
2022 Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
abstract
Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a general framework to jointly model the likelihoods of the monolingual and code-switch sub-tasks that comprise bilingual speech recognition. By defining the monolingual sub-tasks with label-to-frame synchronization, our joint modeling framework can be conditionally factorized such that the final bilingual output, which may or may not be code-switched, is obtained given only monolingual information. We show that this conditionally factorized joint framework can be modeled by an end-to-end differentiable neural network. We demonstrate the efficacy of our proposed model on bilingual Mandarin-English speech recognition across both monolingual and code-switched corpora.
Brian Yan, Meng Yu 0003, Shixiong Zhang 0001, Siddharth Dalmia, Dan Berrebbi, Chao Weng, Shinji Watanabe 0001, Dong Yu 0001
ICASSP4
2022 Joint Neural AEC and Beamforming with Double-Talk Detection
abstract
Acoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback.However, nonlinear distortions induced by audio devices, background noise, reverberation, and double-talk reduce the efficiency of conventional AEC systems.Several hybrid AEC models were proposed to address this, which use deep learning models to suppress residual echo from standard adaptive filtering.This paper proposes deep learning-based joint AEC and beamforming model (JAECBF) building on our previous self-attentive recurrent neural network (RNN) beamformer.The proposed network consists of two modules: (i) multi-channel neural-AEC, and (ii) joint AEC-RNN beamformer with a double-talk detection (DTD) that computes time-frequency (T-F) beamforming weights.We train the proposed model in an end-to-end approach to eliminate background noise and echoes from far-end audio devices, which include nonlinear distortions.From experimental evaluations, we find the proposed network outperforms other multi-channel AEC and denoising systems in terms of speech recognition rate and overall speech quality.
Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH4
2022 EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers
abstract
In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neural diarization (EEND) models, speaker counting with encoder-decoder based attractors (EDA), and speech separation using Conv-TasNet. In addition, we propose a multiple$1 \times 1$convolutional layer architecture for estimating the separation masks corresponding to a flexible number of speakers and a fusion technique for refining the separated speech signal with obtained speaker diarization information to improve the joint framework. Experiments using the LibriMix dataset show that our proposed method outperforms the single-task baselines in both diarization and separation metrics for fixed and flexible numbers of speakers and improves speaker counting performance for flexible numbers of speakers. All materials will be open-sourced and reproducible in ESPnet toolkit11https://github.com/espnet/espnet.
Soumi Maiti, Yushi Ueda, Shinji Watanabe 0001, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004
SLT6
2021 3D Spatial Features for Multi-Channel Target Speech Separation
abstract
The use of speaker's directional information for speech sepa-ration and speech recognition has demonstrated the state-of-the-art performances on multi-talker scenarios. One major limitation of previous approaches using speaker's directional information is the significant performance degradation when the coming directions of two sound sources are close. To address these challenges, this paper proposed a set of new three-dimensional (3D) spatial features for target speech sep-aration, by leveraging all the 3D location information of the target speaker, including azimuth, elevation, and the distance to the microphone array center. Previous works in this area are extended in two important directions. First, the traditional 1D directional features are generalized to 3D spatial features. Thus more discriminative spatial diversity between speakers is achieved. Second, to unleash the full power of these 3D spatial features, a microphone pair-wise attention model is also proposed. The proposed features and models were evaluated on both simulated reverberant datasets and real recordings under near and far-field conditions. Exper-imental results show that both proposed 3D spatial features and attention models can significantly improve the separation performance as well as reducing the recognition error rate.
Rongzhi Gu, Shixiong Zhang 0001, Meng Yu 0003, Dong Yu 0001
ASRU2
2021 Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
abstract
This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end (E2E) neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the azimuth angle of the sources with respect to the microphone array is defined as a latent variable. This angle controls the quality of separation, which in turn determines the ASR performance. All three functionalities of D-ASR: localization, separation, and recognition are connected as a single differentiable neural network and trained solely based on ASR error minimization objectives. The advantages of D-ASR over existing methods are threefold: (1) it provides explicit speaker locations, (2) it improves the explainability factor, and (3) it achieves better ASR performance as the process is more streamlined. In addition, D-ASR does not require explicit direction of arrival (DOA) supervision like existing data-driven localization models, which makes it more appropriate for realistic data. For the case of two source mixtures, D-ASR achieves an average DOA prediction error of less than three degrees. It also outperforms a strong far-field multi-speaker end-to-end system in both separation quality and ASR performance.
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001
ICASSP6
2021 ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech Separation
abstract
Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech recognition (ASR) systems. The conventional mask-based minimum variance distortionless response (MVDR) beamformer can be used to minimize the distortion, but comes with high level of residual noise. Furthermore, the matrix operations (e.g., matrix inversion) involved in the conventional MVDR solution are sometimes numerically unstable when jointly trained with neural networks. In this paper, we propose a novel all deep learning MVDR framework, where the matrix inversion and eigenvalue decomposition are replaced by two recurrent neural networks (RNNs), to resolve both issues at the same time. The proposed method can greatly reduce the residual noise while keeping the target speech undistorted by leveraging on the RNN-predicted frame-wise beamforming weights. The system is evaluated on a Mandarin audio-visual corpus and compared against several state-of-the-art (SOTA) speech separation systems. Experimental results demonstrate the superiority of the proposed method across several objective metrics and ASR accuracy.
Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001
ICASSP4
2021 Multi-Channel Speaker Verification for Single and Multi-Talker Speech
abstract
To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features.Specifically, we combine spectral, spatial, and directional features, which includes inter-channel phase difference, multichannel sinc convolutions, directional power ratio features, and angle features.To maximally leverage supervised learning, our framework is also equipped with multi-channel speech enhancement and voice activity detection.On all simulated, replayed, and real recordings, we observe large and consistent improvements at various degradation levels.On real recordings of multi-talker speech, we achieve a 36% relative reduction in equal error rate w.r.t.single-channel baseline.We find the improvements from speaker-dependent directional features more consistent in multi-talker conditions than clean.Lastly, we investigate if the learned multi-channel speaker embedding space can be made more discriminative through a contrastive loss-based fine-tuning.With a simple choice of Triplet loss, we observe a further 8.3% relative reduction in EER.
Saurabh Kataria 0001, Shixiong Zhang 0001, Dong Yu 0001
Interspeech2
2021 MIMO Self-Attentive RNN Beamformer for Multi-Speaker Speech Separation
abstract
Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion and eigenvalue decomposition with two RNNs.In this work, we present a self-attentive RNN beamformer to further improve our previous RNN-based beamformer by leveraging on the powerful modeling capability of self-attention.Temporal-spatial self-attention module is proposed to better learn the beamforming weights from the speech and noise spatial covariance matrices.The temporal self-attention module could help RNN to learn global statistics of covariance matrices.The spatial self-attention module is designed to attend on the cross-channel correlation in the covariance matrices.Furthermore, a multi-channel input with multi-speaker directional features and multi-speaker speech separation outputs (MIMO) model is developed to improve the inference efficiency.The evaluations demonstrate that our proposed MIMO self-attentive RNN beamformer improves both the automatic speech recognition (ASR) accuracy and the perceptual estimation of speech quality (PESQ) against prior arts.
Xiyun Li, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Jiaming Xu 0001, Bo Xu 0002, Dong Yu 0001
Interspeech4
2021 TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation
abstract
In this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach on the deep neural network (DNN) for environment-aware speech dereverberation, which can adaptively attend to the contextual information. More specifically, a FullBand based Temporal Attention approach (FTA) is proposed, which models the correlations between the fullband information of the context frames. In addition, considering the difference between the attenuation of high frequency bands and low frequency bands (high frequency bands attenuate faster than low frequency bands) in the room impulse response (RIR), we also propose a SubBand based Temporal Attention approach (STA). In order to guide the network to be more aware of the reverberant environments, we jointly optimize the dereverberation network and the reverberation time (RT60) estimator in a multi-task manner. Our experimental results indicate that the proposed method outperforms our previously proposed reverberation-time-aware DNN and the learned attention weights are fully physical consistent. We also report a preliminary yet promising dereverberation and recognition experiment on real test data.
Helin Wang, Bo Wu 0011, Lianwu Chen, Meng Yu 0003, Jianwei Yu 0001, Yong Xu 0004, Shixiong Zhang 0001, Chao Weng, Dan Su 0002, Dong Yu 0001
Interspeech7
2021 Generalized Spatio-Temporal RNN Beamformer for Target Speech Separation
abstract
Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal recurrent neural network based beamformer (RNN-BF) for target speech separation. This new beamforming framework directly learns the beamforming weights from the estimated speech and noise spatial covariance matrices. Leveraging on the temporal modeling capability of RNNs, the RNN-BF could automatically accumulate the statistics of the speech and noise covariance matrices to learn the frame-level beamforming weights in a recursive way. An RNN-based generalized eigenvalue (RNN-GEV) beamformer and a more generalized RNN beamformer (GRNN-BF) are proposed. We further improve the RNN-GEV and the GRNN-BF by using layer normalization to replace the commonly used mask normalization on the covariance matrices. The proposed GRNN-BF obtains better performance against prior arts in terms of speech quality (PESQ), speech-to-noise ratio (SNR) and word error rate (WER).
Yong Xu 0004, Zhuohuang Zhang, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
Interspeech4
2021 MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment
abstract
The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion score (MOS) tests. Non-intrusive speech quality assessment has attracted much attention recently due to the lack of access to clean reference signals for objective evaluations in real scenarios. In this paper, we propose a novel non-intrusive speech quality measurement model, MetricNet, which leverages label distribution learning and joint speech reconstruction learning to achieve significantly improved performance compared to the existing non-intrusive speech quality measurement models. We demonstrate that the proposed approach yields promisingly high correlation to the intrusive objective evaluation of speech quality on clean, noisy and processed speech data.
Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001
Interspeech4
2021 Neural Mask based Multi-channel Convolutional Beamforming for Joint Dereverberation, Echo Cancellation and Denoising
abstract
This paper proposes a new joint optimization framework for simultaneous dereverberation, acoustic echo cancellation, and denoising, which is motivated by the recently proposed con-volutional beamformer for simultaneous denoising and dereverberation. Using the echo aware mask based beamforming framework, the proposed algorithm could effectively deal with double-talk case and local inference, etc. The evaluations based on ERLE for echo only, and PESQ for double-talk demonstrate that the proposed algorithm could significantly improve the performance.
Meng Yu 0003, Yong Xu 0004, Chao Weng, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001
SLT5
2021 WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation
abstract
This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted Power minimization Distortionless response (WPD) beamformer can perform separation and dereverberation simultaneously, the noise cancellation component still has the potential to progress. We propose an improved neural WPD beamformer called "WPD++" by an enhanced beamforming module in the conventional WPD and a multi-objective loss function for the joint training. The beamforming module is improved by utilizing the spatio-temporal correlation. A multi-objective loss, including the complex spectra domain scale-invariant signal-to-noise ratio (C-Si-SNR) and the magnitude domain mean square error (Mag-MSE), is properly designed to make multiple constraints on the enhanced speech and the desired power of the dry clean signal. Joint training is conducted to optimize the complex-valued mask estimator and the WPD++ beamformer in an end-to-end way. The results show that the proposed WPD++ outperforms several state-of-the-art beamformers on the enhanced speech quality and word error rate (WER) of ASR.
Zhaoheng Ni, Yong Xu 0004, Meng Yu 0003, Bo Wu 0011, Shixiong Zhang 0001, Dong Yu 0001, Michael I. Mandel
SLT5
2021 Complex Neural Spatial Filter: Enhancing Multi-Channel Target Speech Separation in Complex Domain
abstract
To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of target speech in time-frequency domain under supervised deep learning framework. However, the existing methods are designed in the way that the real and imaginary parts of the cRM are separately modeled using real-valued training data pairs. The research motivation of this study is to design a deep model that fully exploits the temporal-spectral-spatial information of multi-channel signals for estimating cRM directly and efficiently in complex domain. As a result, a novel TSS network is designed consisting of two modules, a complex neural spatial filter (cNSF) and an MVDR. Essentially, cNSF is a cRM estimation model and an MVDR module is cascaded to the cNSF module to reduce the nonlinear speech distortions introduced by neural network. Specifically, to fit the cRM target, all input features of cNSF are reformulated into complex-valued representations. Then, to achieve good hierarchical feature abstraction, a complex deep neural network (cDNN) is delicately designed with U-Net structure. Experiments conducted on simulated multi-channel speech data demonstrate the proposed cNSF outperforms the baseline NSF by 12.1% scale-invariant signal-to-distortion ratio and 33.1% word error rate.
Rongzhi Gu, Shixiong Zhang 0001, Yuexian Zou, Dong Yu 0001
IEEE Signal Process. Lett.2
2021 An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
abstract
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance.
Daniel Michelsanti, Zheng-Hua Tan, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Audio-Visual Multi-Channel Integration and Recognition of Overlapped Speech
abstract
Automatic speech recognition (ASR) technologies have been significantly advanced in the past few decades. However, recognition of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in current ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption and the additional cues they provide to separate the target speaker from the interfering sound sources, this paper presents an audio-visual multi-channel based recognition system for overlapped speech. It benefits from a tight integration between a speech separation front-end and recognition back-end, both of which incorporate additional video input. A series of audio-visual multi-channel speech separation front-end components based on TF masking, Filter&Sum and mask-based MVDR neural channel integration approaches are developed. To reduce the error cost mismatch between the separation and the recognition components, the entire system is jointly fine-tuned using a multi-task criterion interpolation of the scale-invariant signal to noise ratio (Si-SNR) with either the connectionist temporal classification (CTC), or lattice-free maximum mutual information (LF-MMI) loss function. Experiments suggest that: the proposed audio-visual multi-channel recognition system outperforms the baseline audio-only multi-channel ASR system by up to 8.04% (31.68% relative) and 22.86% (58.51% relative) absolute WER reduction on overlapped speech constructed using either simulation or replaying of the LRS2 dataset respectively. Consistent performance improvements are also obtained using the proposed audio-visual multi-channel recognition system when using occluded video input with the lip region randomly covered up to 60%.
Jianwei Yu 0001, Shixiong Zhang 0001, Bo Wu 0011, Shansong Liu, Shoukang Hu, Mengzhe Geng, Xunying Liu, Helen M. Meng, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Multi-Channel Multi-Frame ADL-MVDR for Target Speech Separation
abstract
Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems. Minimum variance distortionless response (MVDR) filters are often adopted to remove nonlinear distortions, however, conventional neural mask-based MVDR systems still result in relatively high levels of residual noise. Moreover, the matrix inverse involved in the MVDR solution is sometimes numerically unstable during joint training with neural networks. In this study, we propose a multi-channel multi-frame (MCMF) all deep learning (ADL)-MVDR approach for target speech separation, which extends our preliminary multi-channel ADL-MVDR approach. The proposed MCMF ADL-MVDR system addresses linear and nonlinear distortions. Spatio-temporal cross correlations are also fully utilized in the proposed approach. The proposed systems are evaluated using a Mandarin audio-visual corpus and are compared with several state-of-the-art approaches. Experimental results demonstrate the superiority of our proposed systems under different scenarios and across several objective evaluation metrics, including ASR performance.
Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Donald S. Williamson, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Self-Supervised Learning for Audio-Visual Speaker Diarization
abstract
Speaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propose a self-supervised audio-video synchronization learning method to address the problem of speaker diarization without massive labeling effort. We improve the previous approaches by introducing two new loss functions: the dynamic triplet loss and the multinomial loss. We test them on a real-world human-computer interaction system and the results show our best model yields a remarkable gain of +8% F1-scores as well as diarization error rate reduction. Finally, we introduce a new large scale audio-video corpus designed to fill the vacancy of audio-video dataset in Chinese.
Yifan Ding 0002, Yong Xu 0004, Shixiong Zhang 0001, Yahuan Cong, Liqiang Wang 0001
ICASSP3
2020 Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning
abstract
Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In this work, we propose an integrated architecture for learning spatial features directly from the multi-channel speech waveforms within an end-to-end speech separation framework. In this architecture, time-domain filters spanning signal channels are trained to perform adaptive spatial filtering. These filters are implemented by a 2d convolution (conv2d) layer and their parameters are optimized using a speech separation objective function in a purely data-driven fashion. Furthermore, inspired by the IPD formulation, we design a conv2d kernel to compute the inter-channel convolution differences (ICDs), which are expected to provide the spatial cues that help to distinguish the directional sources. Evaluation results on simulated multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed ICD based MCSS model improves the overall signal-to-distortion ratio by 10.4% over the IPD based MCSS model.
Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001
ICASSP2
2020 Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition Objectives
abstract
Target speech extraction is a specific case of source separation where an auxiliary information like the location or some pre-saved anchor speech examples of the target speaker is used to resolve the permutation ambiguity. Traditionally such systems are optimized based on signal reconstruction objectives. Recently end-to-end automatic speech recognition (ASR) methods have enabled to optimize source separation systems with only the transcription based objective. This paper proposes a method to jointly optimize a location guided target speech extraction module along with a speech recognition module only with ASR error minimization criteria. Experimental comparisons with corresponding conventional pipeline systems verify that this task can be realized by end-to-end ASR training objectives without using parallel clean data. We show promising target speech recognition results in mixtures of two speakers and noise, and discuss interesting properties of the proposed system in terms of speech enhancement/separation objectives and word error rates. Finally, we design a system that can take both location and anchor speech as input at the same time and show that the performance can be further improved.
Aswin Shanmugam Subramanian, Chao Weng, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004, Shinji Watanabe 0001, Dong Yu 0001
ICASSP4
2020 Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset
abstract
Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of audio-visual technologies for overlapped speech recognition. Three issues associated with the construction of audio-visual speech recognition (AVSR) systems are addressed. First, the basic architecture designs i.e. end-to-end and hybrid of AVSR systems are investigated. Second, purposefully designed modality fusion gates are used to robustly integrate the audio and visual features. Third, in contrast to a traditional pipelined architecture containing explicit speech separation and recognition components, a streamlined and integrated AVSR system optimized consistently using the lattice-free MMI (LF-MMI) discriminative criterion is also proposed. The proposed LF-MMI time-delay neural network (TDNN) system establishes the state-of-the-art for the LRS2 dataset. Experiments on overlapped speech simulated from the LRS2 dataset suggest the proposed AVSR system outperformed the audio only baseline LF-MMI DNN system by up to 29.98% absolute in word error rate (WER) reduction, and produced recognition performance comparable to a more complex pipelined system. Consistent performance improvements of 4.89% absolute in WER reduction over the baseline AVSR system using feature fusion are also obtained.
Jianwei Yu 0001, Shixiong Zhang 0001, Jian Wu 0027, Shahram Ghorbani, Bo Wu 0011, Shiyin Kang, Shansong Liu, Xunying Liu, Helen M. Meng, Dong Yu 0001
ICASSP2
2020 Exploiting Cross-Domain Visual Feature Generation for Disordered Speech Recognition
Shansong Liu, Xurong Xie, Jianwei Yu 0001, Shoukang Hu, Mengzhe Geng, Rongfeng Su, Shixiong Zhang 0001, Xunying Liu, Helen M. Meng
INTERSPEECH7
2020 Neural Spatio-Temporal Beamformer for Target Speech Separation
abstract
Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR).On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability.In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement.Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the intermicrophone correlation that is already utilized in prior arts.Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN.The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts.
Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Chao Weng, Dong Yu 0001
INTERSPEECH3
2020 Audio-Visual Multi-Channel Recognition of Overlapped Speech
abstract
Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption, this paper presents an audio-visual multi-channel overlapped speech recognition system featuring tightly integrated separation front-end and recognition back-end. A series of audio-visual multi-channel speech separation front-end components based on \textit{TF masking}, \textit{filter\&sum} and \textit{mask-based MVDR} beamforming approaches were developed. To reduce the error cost mismatch between the separation and recognition components, they were jointly fine-tuned using the connectionist temporal classification (CTC) loss function, or a multi-task criterion interpolation with scale-invariant signal to noise ratio (Si-SNR) error cost. Experiments suggest that the proposed multi-channel AVSR system outperforms the baseline audio-only ASR system by up to 6.81\% (26.83\% relative) and 22.22\% (56.87\% relative) absolute word error rate (WER) reduction on overlapped speech constructed using either simulation or replaying of the lipreading sentence 2 (LRS2) dataset respectively.
Jianwei Yu 0001, Bo Wu 0011, Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Dong Yu 0001, Xunying Liu, Helen M. Meng
INTERSPEECH4
2019 Time Domain Audio Visual Speech Separation
abstract
Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker extraction from monaural mixtures. The architecture generalizes the previous TasNet (time-domain speech separation network) to enable multi-modal learning and at meanwhile it extends the classical audio-visual speech separation from frequency-domain to time-domain. The main components of proposed architecture include an audio encoder, a video encoder that extracts lip embedding from video streams, a multi-modal separation network and an audio decoder. Experiments on simulated mixtures based on recently released LRS2 dataset show that our method can bring 3dB+ and 4dB+ Si-SNR improvements on two- and three-speaker cases respectively, compared to audio-only TasNet and frequency-domain audio-visual networks.
Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001
ASRU3
2019 Encrypted Speech Recognition Using Deep Polynomial Networks
abstract
The cloud-based speech recognition/API provides developers or enterprises an easy way to create speech-enabled features in their applications. However, sending audios about personal or company internal information to the cloud, raises concerns about the privacy and security issues. The recognition results generated in cloud may also reveal some sensitive information. This paper proposes a deep polynomial network (DPN) that can be applied to the encrypted speech as an acoustic model. It allows clients to send their data in an encrypted form to the cloud to ensure that their data remains confidential, at mean while the DPN can still make frame-level predictions over the encrypted speech and return them in encrypted form. One good property of the DPN is that it can be trained on unencrypted speech features in the traditional way. To keep the cloud away from the raw audio and recognition results, a cloud-local joint decoding framework is also proposed. We demonstrate the effectiveness of model and framework on the Switchboard and Cortana voice assistant tasks with small performance degradation and latency increased comparing with the traditional cloud-based DNNs.
Shixiong Zhang 0001, Yifan Gong 0001, Dong Yu 0001
ICASSP1
2019 A Comprehensive Study of Speech Separation: Spectrogram vs Waveform Separation
abstract
Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain.Recently, a raw audio waveform separation network (TasNet) is introduced for single-channel data, with achieving high Si-SNR (scale-invariant source-to-noise ratio) and SDR (sourceto-distortion ratio) comparing against the state-of-the-art solution in frequency-domain.In this study, we incorporate effective components of the TasNet into a frequency-domain separation method.We compare both for alternative scenarios.We introduce a solution for directly optimizing the separation criterion in frequency-domain networks.In addition to speech separation objective and subjective measurements, we evaluate the separation performance on a speech recognition task as well.We study the speech separation problem for far-field data (more similar to naturalistic audio streams) and develop multi-channel solutions for both frequency and time-domain separators with utilizing spectral, spatial and speaker location information.For our experiments, we simulated multi-channel spatialized reverberate WSJ0-2mix dataset.Our experimental results show that spectrogram separation can achieve competitive performance with better network design.Multi-channel framework as well is shown to improve the single-channel performance relatively up to +35.5% and +46% in terms of WER and SDR, respectively.
Fahimeh Bahmaninezhad, Jian Wu 0027, Rongzhi Gu, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001
INTERSPEECH4
2019 Neural Spatial Filter: Target Speaker Speech Separation Assisted with Directional Information
Rongzhi Gu, Lianwu Chen, Shixiong Zhang 0001, Jimeng Zheng, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001
INTERSPEECH3
2019 Improved Speaker-Dependent Separation for CHiME-5 Challenge
abstract
This paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-channel, highly-overlapped conversational speech recognition in a dinner party scenario with reverberations and nonstationary noises.We adopt a speaker-aware training method by using i-vector as the target speaker information for multi-talker speech separation.With only one unified separation model for all speakers, we achieve a 10% absolute improvement in terms of word error rate (WER) over the previous baseline of 80.28% on the development set by leveraging our newly proposed data processing techniques and beamforming approach.With our improved back-end acoustic model, we further reduce WER to 60.15% which surpasses the result of our submitted CHiME-5 challenge system without applying any fusion techniques.
Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001
INTERSPEECH3
2018 Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker Verification
abstract
In this paper, given the speaker bottleneck feature vectors extracted with speaker discriminant neural networks, we focus on using the sequential speaker characteristics for text-dependent speaker verification. In each evaluation trial, speaker supervectors are used as the representations of the sequential speaker characteristics rendered in the compared speech utterances. To this end, dynamic time warping is used to warp the variable-length speaker feature vector sequences of the utterances to the same length. Thereafter for every utterance, a speaker supervector can be obtained as the concatenation of its speaker feature vectors. We use Euclidean distance and support vector machine (SVM) to compute the decision score on the speaker supervectors. Our experiments on a Microsoft internal keyword-spotting database showed the effectiveness of the proposed speaker supervector for text-dependent speaker verification. Moreover, when SVM backend was used in scoring, the speaker supervector achieved the best EER performance 1.627%, better than the combination of i-vector and probabilistic linear discriminant analysis.
Yong Zhao 0008, Shixiong Zhang 0001, Guoli Ye, Frank K. Soong
ICASSP3
2018 Domain and Speaker Adaptation for Cortana Speech Recognition
abstract
Voice assistant represents one of the most popular and important scenarios for speech recognition. In this paper, we propose two adaptation approaches to customize a multi-style well-trained acoustic model towards its subsidiary domain of Cortana assistant. First, we present anchor-based speaker adaptation by extracting the speaker information, i-vector or d-vector embeddings, from the anchor segments of ‘Hey Cortana’. The anchor embeddings are mapped to layer-wise parameters to control the transformations of both weight matrices and biases of multiple layers. Second, we directly update the existing model parameters for domain adaptation. We demonstrate that prior distribution should be updated along with the network adaptation to compensate the label bias from the development data. Updating the priors may have a significant impact when the target domain features high occurrence of anchor words. Experiments on Hey Cortana desktop test set show that both approaches improve the recognition accuracy significantly. The anchor-based adaptation using the anchor d-vector and the prior interpolation achieves 32% relative reduction in WER over the generic model.
Yong Zhao 0008, Jinyu Li 0001, Shixiong Zhang 0001, Yifan Gong 0001
ICASSP3
2016 Simplifying long short-term memory acoustic models for fast training and decoding
abstract
On acoustic modeling, recurrent neural networks (RNNs) using Long Short-Term Memory (LSTM) units have recently been shown to outperform deep neural networks (DNNs) models. This paper focuses on resolving two challenges faced by LSTM models: high model complexity and poor decoding efficiency. Motivated by our analysis of the gates activation and function, we present two LSTM simplifications: deriving input gates from forget gates, and removing recurrent inputs from output gates. To accelerate decoding of LSTMs, we propose to apply frame skipping during training, and frame skipping and posterior copying (FSPC) during decoding. In the experiments, model simplifications reduce the size of LSTM models by 26%, resulting in a simpler model structure. Meanwhile, the application of FSPC speeds up model computation by 2 times during LSTM decoding. All these improvements are achieved at the cost of 1% WER degradation.
Yajie Miao, Jinyu Li 0001, Shixiong Zhang 0001, Yifan Gong 0001
ICASSP4
2016 Recurrent support vector machines for speech recognition
abstract
Recurrent Neural Networks (RNNs) using Long-Short Term Memory (LSTM) architecture have demonstrated the state-of-the-art performances on speech recognition. Most of deep RNNs use the softmax activation function in the last layer for classification. This paper illustrates small but consistent advantages of replacing the softmax layer in RNN with Support Vector Machines (SVMs). The parameters of RNNs and SVMs are jointly learned using a sequence-level max-margin criteria, instead of cross-entropy. The resulting model is termed Recurrent SVM. The conventional SVMs need to predefine a feature space and do not have internal states to deal with arbitrary long-term dependencies in sequences. The proposed recurrent SVM uses LSTMs to learn the feature space and to capture temporal dependencies, while using the SVM (in the last layer) for sequence classification. The model is evaluated on the Windows phone task for large vocabulary continuous speech recognition.
Shixiong Zhang 0001, Rui Zhao 0017, Chaojun Liu, Jinyu Li 0001, Yifan Gong 0001
ICASSP1
2016 End-to-End attention based text-dependent speaker verification
abstract
A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetic discriminate/speaker discriminate DNN as a feature extractor for speaker verification has shown promising results. The extracted frame-level (bottleneck, posterior or d-vector) features are equally weighted and aggregated to compute an utterance-level speaker representation (d-vector or i-vector). In this work we use a speaker discriminate CNN to extract the noise-robust frame-level features. These features are smartly combined to form an utterance-level speaker vector through an attention mechanism. The proposed attention model takes the speaker discriminate information and the phonetic information to learn the weights. The whole system, including the CNN and attention model, is joint optimized using an end-to-end criterion. The training algorithm imitates exactly the evaluation process — directly mapping a test utterance and a few target speaker utterances into a single verification score. The algorithm can smartly select the most similar impostor for each target speaker to train the network. We demonstrated the effectiveness of the proposed end-to-end system on Windows 10 “Hey Cortana” speaker verification task.
Shixiong Zhang 0001, Zhuo Chen 0006, Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001
SLT1
2015 Deep neural support vector machines for speech recognition
abstract
A new type of deep neural networks (DNNs) is presented in this paper. Traditional DNNs use the multinomial logistic regression (softmax activation) at the top layer for classification. The new DNN instead uses a support vector machine (SVM) at the top layer. Two training algorithms are proposed at the frame and sequence-level to learn parameters of SVM and DNN in the maximum-margin criteria. In the frame-level training, the new model is shown to be related to the multiclass SVM with DNN features; In the sequence-level training, it is related to the structured SVM with DNN features and HMM state transition features. Its decoding process is similar to the DNN-HMM hybrid system but with frame-level posterior probabilities replaced by scores from the SVM. We term the new model deep neural support vector machine (DNSVM). We have verified its effectiveness on the TIMIT task for continuous speech recognition.
Shixiong Zhang 0001, Chaojun Liu, Kaisheng Yao, Yifan Gong 0001
ICASSP1
2014 Infinite structured support vector machines for speech recognition
abstract
Discriminative models, like support vector machines (SVMs), have been successfully applied to speech recognition and improved performance. A Bayesian non-parametric version of the SVM, the infinite SVM, improves on the SVM by allowing more flexible decision boundaries. However, like SVMs, infinite SVMs model each class separately, which restricts them to classifying one word at a time. A generalisation of the SVM is the structured SVM, whose classes can be sequences of words that share parameters. This paper studies a combination of Bayesian non-parametrics and structured models. One specific instance called infinite structured SVM is discussed in detail, which brings the advantages of the infinite SVM to continuous speech recognition.
Rogier C. van Dalen, Shixiong Zhang 0001, Mark J. F. Gales
ICASSP3
2013 Investigation of multilingual deep neural networks for spoken term detection
abstract
The development of high-performance speech processing systems for low-resource languages is a challenging area. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to use bottleneck features, or hybrid systems, trained on multilingual data for speech-to-text (STT) systems. This paper presents an investigation into the application of these multilingual approaches to spoken term detection. Experiments were run using the IARPA Babel limited language pack corpora (~10 hours/language) with 4 languages for initial multilingual system development and an additional held-out target language. STT gains achieved through using multilingual bottleneck features in a Tandem configuration are shown to also apply to keyword search (KWS). Further improvements in both STT and KWS were observed by incorporating language questions into the Tandem GMM-HMM decision trees for the training set languages. Adapted hybrid systems performed slightly worse on average than the adapted Tandem systems. A language independent acoustic model test on the target language showed that retraining or adapting of the acoustic models to the target language is currently minimally needed to achieve reasonable performance.
Kate M. Knill, Mark J. F. Gales, Shakti P. Rath, Philip C. Woodland, Chao Zhang 0031, Shixiong Zhang 0001
ASRU6
2013 Kernelized log linear models for continuous speech recognition
abstract
Large margin criteria and discriminative models are two effective improvements for HMM-based speech recognition. This paper proposed a large margin trained log linear model with kernels for CSR. To avoid explicitly computing in the high dimensional feature space and to achieve the nonlinear decision boundaries, a kernel based training and decoding framework is proposed in this work. To make the system robust to noise a kernel adaptation scheme is also presented. Previous work in this area is extended in two directions. First, most kernels for CSR focus on measuring the similarity between two observation sequences. The proposed joint kernels defined a similarity between two observation-label sequence pairs on the sentence level. Second, this paper addresses how to efficiently employ kernels in large margin training and decoding with lattices. To the best of our knowledge, this is the first attempt at using large margin kernel-based log linear models for CSR. The model is evaluated on a noise corrupted continuous digit task: AURORA 2.0.
Shixiong Zhang 0001, Mark J. F. Gales
ICASSP1
2013 Structured SVMs for Automatic Speech Recognition
abstract
Structured discriminative models are a flexible sequence classification approach that enable a wide variety of features to be used. This paper describes a particular model in this framework, structured support vector machines (SSVM), and how it can be applied to medium to large vocabulary speech recognition tasks. An important aspect of SSVMs is the form of the joint feature spaces. Here, context-dependent generative models, hidden Markov models, are used to obtain the features. To apply this form of combined generative and discriminative model to medium and larger vocabulary tasks, a number of issues need to be addressed. First, the features extracted are a function of the segmentation of the utterance. A Viterbi-like scheme for obtaining the “optimal” segmentation is described. Second, SSVMs can be viewed as large margin log linear models using a zero mean Gaussian prior of the discriminative parameter. However this form of prior is not appropriate for all features. A modified training algorithm is proposed that allows general Gaussian priors to be incorporated into the large margin criterion. Finally to speed up the training process, a 1-slack algorithm, caching competing hypotheses and parallelization strategies are also described. The performance of SSVMs is evaluated on small and medium to large speech recognition tasks: AURORA 2 and 4.
Shixiong Zhang 0001, Mark J. F. Gales
IEEE Trans. Speech Audio Process.1
2011 Extending noise robust structured support vector machines to larger vocabulary tasks
abstract
This paper describes a structured SVM framework suitable for noise-robust medium/large vocabulary speech recognition. Several theoretical and practical extensions to previous work on small vocabulary tasks are detailed. The joint feature space based on word models is extended to allow context-dependent triphone models to be used. By interpreting the structured SVM as a large margin log-linear model, illustrates that there is an implicit assumption that the prior of the discriminative parameter is a zero mean Gaussian. However, depending on the definition of likelihood feature space, a non-zero prior may be more appropriate. A general Gaussian prior is incorporated into the large margin training criterion in a form that allows the cutting plan algorithm to be directly applied. To further speed up the training process, 1-slack algorithm, caching competing hypothesis and parallelization strategies are also proposed. The performance of structured SVMs is evaluated on noise corrupted medium vocabulary speech recognition task: AURORA 4.
Shixiong Zhang 0001, Mark J. F. Gales
ASRU1
2011 Structured Support Vector Machines for Noise Robust Continuous Speech Recognition
abstract
The use of discriminative models is an interesting alternative to generative models for speech recognition. This paper examines one form of these models, structured support vector machines (SVMs), for noise robust speech recognition. One important aspect of structured SVMs is the form of the joint feature space. In this work features based on generative models are used, which allows model-based compensation schemes to be applied to yield robust joint features. However, these features require the segmentation of frames into words, or subwords, to be specified. In previous work this segmentation was obtained using generative models. Here the segmentations are refined using the parameters of the structured SVM. A Viterbilike scheme for obtaining "optimal" segmentations, and modifications to the training algorithm to allow them to be efficiently used, are described. The performance of the approach is evaluated on a noise corrupted continuous digit task: AURORA 2. Copyright © 2011 ISCA.
Shixiong Zhang 0001, Mark J. F. Gales
INTERSPEECH1
2011 Optimized Discriminative Kernel for SVM Scoring and Its Application to Speaker Verification
abstract
The decision-making process of many binary classification systems is based on the likelihood ratio (LR) scores of test patterns. This paper shows that LR scores can be expressed in terms of the similarity between the supervectors (SVs) formed by stacking the mean vectors of Gaussian mixture models corresponding to the test patterns, the target model, and the background model. By interpreting the support vector machine (SVM) kernels as a specific similarity (or discriminant) function between SVs, this paper shows that LR scoring is a special case of SVM scoring and that most sequence kernels can be obtained by assuming a specific form for the similarity function of SVs. This paper further shows that this assumption can be relaxed to derive a new general kernel. The kernel function is general in that it is a linear combination of any kernels belonging to the reproducing kernel Hilbert space. The combination weights are obtained by optimizing the ability of a discriminant function to separate the positive and negative classes using either regression analysis or SVM training. The idea was applied to both high-and low-level speaker verification. In both cases, results show that the proposed kernels achieve better performance than several state-of-the-art sequence kernels. Further performance enhancement was also observed when the high-level scores were combined with acoustic scores.
Shixiong Zhang 0001, Man-Wai Mak
IEEE Trans. Neural Networks1
2010 Structured Log Linear Models for Noise Robust Speech Recognition
abstract
The use of discriminative models for structured classification tasks, such as speech recognition is becoming increasingly popular. This letter examines the use of structured log-linear models for noise robust speech recognition. An important aspect of log-linear models is the form of the features. By using generative models to derive the features, state-of-the-art model-based compensation schemes can be used to make the system robust to noise. Previous work in this area is extended in two important directions. First, a large margin training of sentence-level log linear models is proposed for automatic speech recognition (ASR). This form of model is shown to be similar to the recently proposed structured Support Vector Machines (SVM). Second, based on the designed joint features, efficient lattice-based training and decoding are performed. This novel model combines generative kernels, discriminative models, efficient lattice-based large margin training and model-based noise compensation. It is evaluated on a noise corrupted continuous digit task: AURORA 2.0.
Shixiong Zhang 0001, Anton Ragni, Mark J. F. Gales
IEEE Signal Process. Lett.1
2009 Optimization of discriminative kernels in SVM speaker verification
abstract
An important aspect of SVM-based speaker verification systems is the design of sequence kernels. These kernels should be able to map variable-length observation sequences to fixed-size supervectors that capture the dynamic characteristics of speech utterances and allow speakers to be easily distinguished. Most existing kernels in SVM speaker verification are obtained by assuming a specific form for the similarity function of supervectors. This paper relaxes this assumption to derive a new general kernel. The kernel function is general in that it is a linear combination of any kernels belonging to the reproducing kernel Hilbert space. The combination weights are obtained by optimizing the ability of a discriminant function to separate a target speaker from impostors using either regression analysis or SVM training. The idea was applied to both low- and high-level speaker verification. In both cases, results show that the proposed kernels outperform the state-of-the-art sequence kernels. Further performance enhancement was also observed when the high-level scores were combined with acoustic scores. Index Terms — speaker verification; optimal kernels; sequence kernels; SVM; high-level features. 1.
Shixiong Zhang 0001, Man-Wai Mak
INTERSPEECH1
2009 A new adaptation approach to high-level speaker-model creation in speaker verification
Shixiong Zhang 0001, Man-Wai Mak
Speech Commun.1
2008 High-level speaker verification via articulatory-feature based sequence kernels and SVM
abstract
Interspeech 2008, Brisbane, Australia, 22-26 September 2008
Shixiong Zhang 0001, Man-Wai Mak
INTERSPEECH1
2007 High-level feature-based speaker verification via articulatory phonetic-class pronunciation modeling
abstract
Although articulatory feature-based conditional pronunciation models (AFCPMs) can capture the pronunciation characteristics of speakers, they requires one discrete density function for each phoneme, which may lead to inaccurate models when the amount of training data is limited. This paper proposes a phonetic-class based AFCPM in which the density functions in speaker models are conditioned on phonetic classes instead of phonemes. Phonemes are mapped to phonetic classes by (1) vector quantizing the phoneme-dependent universal background models, (2) grouping phonemes according to the classical phoneme tree, and (3) combination of (1) and (2). A new scoring method that uses an SVM to combine the scores of phonetic-class models is also proposed. Evaluations based on 2000 NIST SRE show that the proposed approach can effectively solve the data sparseness problem encountered in conventional AFCPM. 1.
Shixiong Zhang 0001, Man-Wai Mak, Helen M. Meng
INTERSPEECH1
2007 Speaker Verification via High-Level Feature Based Phonetic-Class Pronunciation Modeling
abstract
It has been shown recently that the pronunciation characteristics of speakers can be represented by articulatory feature-based conditional pronunciation models (AFCPMs). However, the pronunciation models are phoneme-dependent, which may lead to speaker models with low discriminative power when the amount of enrollment data is limited. This paper proposes to mitigate this problem by grouping similar phonemes into phonetic classes and representing background and speaker models as phonetic-class dependent density functions. Phonemes are grouped by (1) vector quantizing the discrete densities in the phoneme-dependent universal background models, (2) using the phone properties specified in the classical phoneme tree, or (3) combining vector quantization and phone properties. Evaluations based on 2000 NIST SRE show that this phonetic-class approach effectively alleviates the data spareness problem encountered in conventional AFCPM, which results in better performance when fused with acoustic features.
Shixiong Zhang 0001, Man-Wai Mak, Helen M. Meng
IEEE Trans. Computers1