Xiao-Lei Zhang 0001

dblp:03/5004-1 · also Xiaolei Zhang 0001 · DBLP profile ↗
← Back
56ranked-venue papers
16as first author
29since 2021 · last 2026
0000-0001-7694-193XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 32 · 11 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DualSpec: Text-to-Spatial-Audio Generation via Dual-Spectrogram Guided Diffusion Model
abstract
Text-to-audio (TTA), which generates audio signals from textual descriptions, has received huge attention in recent years. However, recent works focused on text to monaural audio only. As we know, spatial audio provides more immersive auditory experience than monaural audio, e.g. in virtual reality. To address this issue, we propose a text-to-spatial-audio (TTSA) generation framework named DualSpec. Specifically, it first trains variational autoencoders (VAEs) for extracting the latent acoustic representations from sound event audio. Then, given text that describes sound events and event directions, the proposed method uses the encoder of a pretrained large language model to transform the text into text features. Finally, it trains a diffusion model from the latent acoustic representations and text features for the spatial audio generation. In the inference stage, only the text description is needed to generate spatial audio. Particularly, to improve the synthesis quality and azimuth accuracy of the spatial sound events simultaneously, we propose to use two kinds of acoustic features. One is the Mel spectrograms which is good for improving the synthesis quality, and the other is the short-time Fourier transform spectrograms which is good at improving the azimuth accuracy. We provide a pipeline of constructing spatial audio dataset with text prompts, for the training of the VAEs and diffusion model. We also introduce new spatial-aware evaluation metrics to quantify the azimuth errors of the generated spatial audio recordings. Experimental results demonstrate that the proposed method can generate spatial audio with high directional and event consistency.
Lei Zhao 0031, Sizhou Chen, Linfeng Feng, Jichao Zhang, Xiao-Lei Zhang 0001, Xuelong Li 0001
IEEE Trans. Multim.5
2025 Co-Attention Based Multi-Channel TF-GridNet for Speech Separation with Ad-Hoc Microphone Arrays
abstract
Speech separation using ad-hoc microphone arrays has been explored, but there is still significant room for improvement, especially in complex scenarios with varying channel conditions. Co-attention, a feature fusion mechanism, is widely used in multimodal fusion to capture the cooperation between modalities and enhance the representation of extracted features. In this paper, we propose a co-attention-based multi-channel model for speech separation with ad-hoc microphone arrays. The co-attention mechanism is integrated into the model to enhance the interaction between different speakers across multiple channels, enabling efficient channel fusion. To the best of our knowledge, this is the first work to apply co-attention for speech separation. Experimental results demonstrate that the proposed method significantly outperforms existing approaches, underscoring the importance of co-attention in optimizing channel fusion for speech separation in challenging acoustic environments.
Hongmei Guo, Linfeng Feng, Xueqing Li 0003, Boyu Zhu, Xiao-Lei Zhang 0001, Xuelong Li 0001
ICASSP7
2025 Eliminating quantization errors in classification-based sound source localization
Linfeng Feng, Xiao-Lei Zhang 0001, Xuelong Li 0001
Neural Networks2
2025 Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization With Target Speaker ASR
abstract
Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spectrogram mask for the target speech. However, imperfections in these masks often result in over-/under-suppression of target/non-target speech, degrading perceptual quality. Generative methods, by contrast, re-synthesize target speech based on the mixture and target speaker cues, achieving superior perceptual quality. Nevertheless, these methods often overlook speech intelligibility, leading to alterations or loss of semantic content in the re-synthesized speech. Inspired by the Whisper model's success in target speaker ASR, we propose a generative TSE framework based on the pre-trained Whisper model to address the above issues. This framework integrates semantic modeling with flow-based acoustic modeling to achieve both high intelligibility and perceptual quality. Results from multiple benchmarks demonstrate that the proposed method outperforms existing generative and discriminative baselines. We present speech samples onhttps://aisaka0v0.github.io/GenerativeTSE_demo/.
Rujin Chen, Xiao-Lei Zhang 0001, Xuelong Li 0001
IEEE Signal Process. Lett.3
2024 Exploiting A Quantum Multiple Kernel Learning Approach For Low-Resource Spoken Command Recognition
abstract
We propose a theoretical analysis of quantum projection learning (QPL) that employs multiple kernels, highlighting its advantages through representation error analysis. Building upon previous studies that utilized a single quantum kernel-based method, we further investigate a quantum projection framework that incorporates multiple Gaussian kernels for low-resource spoken command recognition. Our empirical results align with our theoretical insights, suggesting that methods based on multiple kernels can further enhance the performance of QPL. By leveraging the quantum-to-classical projected output embeddings, we integrate this with a prototypical network for acoustic modeling. When evaluated using Arabic, Chuvash, Irish, and Lithuanian low-resource speech from CommonVoice, our proposed method surpasses the recurrent neural network and single kernel-based classifier baselines by an average of +5.28%.
Xianyan Fu, Xiao-Lei Zhang 0001, Chao-Han Huck Yang, Jun Qi 0002
ICASSP2
2024 Graph Attention Based Multi-Channel U-Net for Speech Dereverberation With Ad-Hoc Microphone Arrays
Hongmei Guo, Xiao-Lei Zhang 0001, Xuelong Li 0001
INTERSPEECH3
2024 Transformer-Based End-to-End Speech Translation With Rotary Position Embedding
abstract
Recently, many Transformer-based models have been applied to end-to-end speech translation because of their capability to model global dependencies. Position embedding is crucial in Transformer models as it facilitates the modeling of dependencies between elements at various positions within the input sequence. Most position embedding methods employed in speech translation such as the absolute and relative position embedding, often encounter challenges in leveraging relative positional information or adding computational burden to the model. In this letter, we introduce a novel approach by incorporating rotary position embedding into Transformer-based speech translation (RoPE-ST). RoPE-ST first adds absolute position information by multiplying the input vector with rotation matrices, and then implements relative position embedding through the dot-product of the self-attention mechanism. The main advantage of the proposed method over the original method is that rotary position embedding combines the benefits of absolute and relative position embedding, which is suited for position embedding in speech translation tasks. We conduct experiments on a multilingual speech translation corpus MuST-C. Results show that RoPE-ST achieves an average improvement of 2.91 BLEU over the method without rotary position embedding in eight translation directions.
Xueqing Li 0003, Shengqiang Li, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE Signal Process. Lett.3
2024 Learning Multi-Dimensional Speaker Localization: Axis Partitioning, Unbiased Label Distribution, and Data Augmentation
abstract
Multi-dimensional speaker localization (SL) aims to estimate the two- or three-dimensional locations of speakers. A recent advancement in multi-dimensional SL is the end-to-end deep neural networks (DNNs) with ad-hoc microphone arrays. This method transforms the SL problem into a classification problem, i.e. a problem of identifying the grids where speakers are located. However, the classification formulation has two closely connected weaknesses. Firstly, this approach introduces quantization error, which needs a large number of grids to mitigate the error. However, increasing the number of grids leads to the curse of dimensionality. To address the problems, we propose an efficient multi-dimensional SL algorithm, which has the following three novel contributions. First, we decouple the high-dimensional grid partitioning intoaxis partitioning, which substantially mitigates the curse-of-dimensionality. Particularly, for the multi-speaker localization problem, we employ a separator to circumvent the permutation ambiguity of the axis partitioning in the inference stage. Second, we introduce a comprehensiveunbiased label distributionscheme to further eliminate quantization errors. Finally, a set of data augmentation techniques are proposed, including coordinate transformation, stochastic node selection, and mixed training, to alleviate overfitting and sample imbalance problems. The proposed methods were evaluated on both simulated and real-world data, and the experimental results confirm the effectiveness.
Linfeng Feng, Yijun Gong, Xiao-Lei Zhang 0001, Xuelong Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Interpretable Spectrum Transformation Attacks to Speaker Recognition Systems
abstract
The success of adversarial attacks on speaker recognition is mainly in white-box scenarios. When applying the adversarial voices that are generated by attacking white-box surrogate models to black-box victim models, i.e. transfer-based black-box attacks, the transferability of the adversarial voices is not only far from satisfactory, but also lacks interpretable basis. To address these issues, in this paper, we propose a general framework, named spectral transformation attack based on modified discrete cosine transform (STA-MDCT), to improve the transferability of the adversarial voices to a black-box victim model. Specifically, we first apply MDCT to the input voice. Then, we slightly modify the energy of different frequency bands for capturing the salient regions of the adversarial noise in the time-frequency domain that are critical to a successful attack. Unlike existing approaches that operate voices in the time domain, the proposed framework operates voices in the time-frequency domain, which improves the interpretability, transferability, and imperceptibility of the attack. Moreover, it can be implemented with any gradient-based attackers. To utilize the advantage of model ensembling, we not only implement STA-MDCT with a single white-box surrogate model but also with an ensemble of surrogate models. Finally, we visualize the saliency maps of adversarial voices by the class activation maps (CAM), which offer an interpretable basis for transfer-based attacks in speaker recognition for the first time. Extensive comparison results with six representative attackers show that the CAM visualization clearly explains the effectiveness of STA-MDCT and the weaknesses of the comparison methods; the proposed method outperforms the comparison methods by a large margin. Our audio samples are available on the demo website.
Jiadi Yao, Jun Qi 0002, Xiao-Lei Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Multi-Resolution Convolutional Residual Neural Networks for Monaural Speech Dereverberation
abstract
It is known that the reverberant speech in different acoustic environments varies according to reverberation time. However, most deep learning based speech dereverberation methods rely on a single deep model to learn the context information. It may make the deep model biased to only part of the reverberant time durations. In this paper, we propose a multi-resolution framework to address this issue. The framework integrates the dereverberant ability of multiple deep subnetworks with different time resolutions into a unified model by transferring the dereverberant information from high-resolution subnetworks to low-resolution subnetworks. By doing so, the unified model can perform well in both long and short reverberant time. We further propose two implementations of the framework based on advanced convolutional residual neural networks. The first implementation, named multi-resolution UNet, uses our new implementation of UNet based on convolutional blocks as the dereverberation subnetwork. The second implementation, named multi-resolution stacked convolutional blocks, uses our new stacked convolutional blocks as the subnetwork. Experimental results in both simulated and real-world environments show that the proposed algorithms outperform the state-of-the-art dereverberation methods in terms of both the evaluation metrics for speech dereverberation and word error rate (WER) for speech recognition.
Lei Zhao 0031, Shengqiang Li, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
abstract
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The core idea of fast-U2++ is to output partial results of the bottom layers in its encoder with a small chunk, while using a large chunk in the top layers of its encoder to compensate the performance degradation caused by the small chunk. More-over, we use knowledge distillation method to reduce the token emission latency. We present extensive experiments on Aishell-1 dataset. Experiments and ablation studies show that compared to U2++, fast-U2++ reduces model latency from 320ms to 80ms, and achieves a character error rate (CER) of 5.06% with a streaming setup.
Chengdong Liang, Xiao-Lei Zhang 0001, Di Wu 0061, Shengqiang Li, Xingchen Song, Zhendong Peng, Fuping Pan
ICASSP2
2023 Optimizing Quantum Federated Learning Based on Federated Quantum Natural Gradient Descent
abstract
Quantum federated learning (QFL) is a quantum extension of the classical federated learning model across multiple local quantum devices. An efficient optimization algorithm is always expected to minimize the communication overhead among different quantum participants. In this work, we propose an efficient optimization algorithm, namely federated quantum natural gradient descent (FQNGD), and further, apply it to a QFL framework that is com-posed of a variational quantum circuit (VQC)-based quantum neural networks (QNN). Compared with stochastic gradient descent methods like Adam and Adagrad, the FQNGD algorithm admits much fewer training iterations for the QFL to get converged. Moreover, it can significantly reduce the total communication overhead among local quantum devices. Our experiments on a handwritten digit classification dataset justify the effectiveness of the FQNGD for the QFL framework in terms of a faster convergence rate on the training set and higher accuracy on the test set.
Jun Qi 0002, Xiao-Lei Zhang 0001, Javier Tejedor
ICASSP2
2023 Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit
abstract
Keyword spotting (KWS) enables speech-based user interaction and gradually becomes an indispensable component of smart devices. Recently, end-to-end (E2E) methods have be-come the most popular approach for on-device KWS tasks. However, there is still a gap between the research and deployment of E2E KWS methods. In this paper, we introduce WeKws, a production-quality, easy-to-build, and convenient-to-be-applied E2E KWS toolkit. WeKws contains the implementations of several state-of-the-art backbone networks, making it achieve highly competitive results on three publicly available datasets. To make WeKws a pure E2E toolkit, we utilize a refined max-pooling loss to make the model learn the ending position of the keyword by itself, which significantly simplifies the training pipeline and makes WeKws very efficient to be applied in real-world scenarios. The toolkit is publicly available at https://github.com/wenet-e2e/wekws.
Menglong Xu, Jingyong Hou, Xiao-Lei Zhang 0001, Lei Xie 0001, Fuping Pan
ICASSP5
2023 Branch-ECAPA-TDNN: A Parallel Branch Architecture to Capture Local and Global Features for Speaker Verification
Jiadi Yao, Chengdong Liang, Zhendong Peng, Xiao-Lei Zhang 0001
INTERSPEECH5
2023 Deep NMF topic modeling
Xiao-Lei Zhang 0001
Neurocomputing2
2023 Symmetric Saliency-Based Adversarial Attack to Speaker Identification
abstract
Adversarial attack approaches to speaker identification either need high computational cost or are not very effective, to our knowledge. To address this issue, in this letter, we propose a novel generation-network-based approach, called symmetric saliency-based encoder-decoder (SSED), to generate adversarial voice examples to speaker identification. It contains two novel components. First, it uses a novel saliency map decoder to learn the importance of speech samples to the decision of a targeted speaker identification system, so as to make the attacker focus on generating artificial noise to the important samples. It also proposes an angular loss function to push the speaker embedding far away from the source speaker. Our experimental results demonstrate that the proposed SSED yields the state-of-the-art performance, i.e. over 97% targeted attack success rate and a signal-to-noise level of over 39 dB on both the open-set and close-set speaker identification tasks, with a low computational cost.
Jiadi Yao, Xing Chen 0011, Xiao-Lei Zhang 0001, Weiqiang Zhang 0001, Kunde Yang
IEEE Signal Process. Lett.3
2023 LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification
abstract
Although the security of automatic speaker verification (ASV) is seriously threatened by recently emerged adversarial attacks, there have been some countermeasures to alleviate the threat. However, many defense approaches not only require the prior knowledge of the attackers but also possess weak interpretability. To address this issue, in this paper, we propose anattacker-independentandinterpretablemethod, namedlearnable mask detector(LMD), to separate adversarial examples from the genuine ones. It utilizes score variation as an indicator to detect adversarial examples, where the score variation is the absolute discrepancy between the ASV scores of an original audio recording and its transformed audio synthesized from its masked complex spectrogram. A core component of the score variation detector is to generate the masked spectrogram by a neural network. The neural network needs only genuine examples for training, which makes it an attacker-independent approach. Its interpretability lies that the neural network is trained to minimize the score variation of the targeted ASV, and maximize the number of the masked spectrogram bins of the genuine training examples. Its foundation is based on the observation that, masking out the vast majority of the spectrogram bins with little speaker information will inevitably introduce a large score variation to the adversarial example, and a small score variation to the genuine example. Experimental results with 12 attackers and two representative ASV systems show that our proposed method outperforms five state-of-the-art baselines. The extensive experimental results can also be a benchmark for the detection-based ASV defenses.
Xing Chen 0011, Xiao-Lei Zhang 0001, Weiqiang Zhang 0001, Kunde Yang
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 End-to-End Multi-Modal Speech Recognition on an Air and Bone Conducted Speech Corpus
abstract
Automatic speech recognition (ASR) has been significantly improved in the past years. However, most robust ASR systems are based on air-conducted (AC) speech, and their performances in low signal-to-noise-ratio (SNR) conditions are not satisfactory. Bone-conducted (BC) speech is intrinsically insensitive to environmental noise, and therefore can be used as an auxiliary source for improving the performance of an ASR at low SNR. In this paper, we first develop a multi-modal Mandarin corpus, which contains air- and bone-conducted synchronized speech (ABCS). The multi-modal speeches are recorded with a headset equipped with both AC and BC microphones. To our knowledge, it is by far the largest corpus for conducting bone conduction ASR research. Then, we propose a multi-modal conformer ASR system based on a novel multi-modal transducer (MMT). The proposed system extracts semantic embeddings from the AC and BC speech signals by a conformer-based encoder and a transformer-based truncated decoder. The semantic embeddings of the two speech sources are fused dynamically with adaptive weights by the MMT module. Experimental results demonstrate the proposed multi-modal system outperforms single-modal systems with either AC or BC modality and multi-modal baseline system by a large margin at various SNR levels. It also shows the two modalities complement with each other, and our method can effectively utilize the complementary information of different sources.
Mou Wang, Junqi Chen 0001, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 End-To-End Multi-Modal Speech Recognition with Air and Bone Conducted Speech
abstract
Improving the performance of automatic speech recognition (ASR) in adverse acoustic environments is a long-term tough task. Although many robust ASR systems based on conventional microphones have been developed, their performance with air-conducted (AC) speech is still far from satisfactory in low signal-to-noise-ratio (SNR) environments. Bone-conducted (BC) speech is relatively insensitive to ambient noise, and has a potential of promoting the ASR performance at such low SNR environments as an auxiliary source. In this paper, we propose a conformer-based multi-modal speech recognition system. It uses a conformer encoder and a transformer-based truncated decoder to extract the semantic information from AC and BC channels respectively. The semantic information of the two channels are re-weighted and integrated by a novel multi-modal transducer. Experimental results show the effectiveness of the proposed method. For example, given a 0 dB SNR environment, it yields a character error rate of over 59.0% lower than a noise-robust baseline conducted on AC channel only, and over 12.7% lower than a multi-modal baseline that takes the concatenated features of AC and BC speech as the input.
Junqi Chen 0001, Mou Wang, Xiao-Lei Zhang 0001, Zhiyong Huang 0001, Susanto Rahardja
ICASSP3
2022 Multi-Channel Far-Field Speaker Verification with Large-Scale Ad-hoc Microphone Arrays
abstract
Speaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments.However, existing approaches extract utterancelevel speaker embeddings from each channel of an ad-hoc microphone array, which does not consider fully the spatialtemporal information across the devices.In this paper, we propose to aggregate the multichannel signals of the ad-hoc microphone array at the frame-level by exploring the cross-channel information deeply with two attention mechanisms.The first one is a self-attention method.It consists of a cross-frame selfattention layer and a cross-channel self-attention layer successively, both working at the frame level.The second one learns the cross-frame and cross-channel information via two graph attention layers.Experimental results demonstrate that the proposed methods reach the state-of-the-art performance.Moreover, the graph-attention method is better than the self-attention method in most cases.
Chengdong Liang, Jiadi Yao, Xiao-Lei Zhang 0001
INTERSPEECH4
2022 Multi-class AUC Optimization for Robust Small-footprint Keyword Spotting with Limited Training Data
Menglong Xu, Shengqiang Li, Chengdong Liang, Xiao-Lei Zhang 0001
INTERSPEECH4
2022 Deep ad-hoc beamforming based on speaker extraction for target-dependent speech separation
Ziye Yang, Shanzheng Guan, Xiao-Lei Zhang 0001
Speech Commun.3
2022 End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-Entropy
abstract
End-to-end speaker verification achieves the verification through estimating directly the similarity score between a pair of utterances, which is formulated as a binary (i.e., target versus non-target) classification problem. Unlike the stage-wise method, an end-to-end verification approach optimizes the evaluation metrics directly and its output layer is parameter-free, which can save great computing and memory resources. However, there are two important issues that need to be meticulously handled in training an end-to-end speaker verification model. The first one is how to deal with severely imbalanced trials, i.e., the number of target trials is much smaller than that of nontarget trials, and the other is about how to handle easy trials that do not help improve the model in training. To circumvent these two issues, we propose in this paper a binary cross-entropy (BCE) type of loss function and present a method to train the deep neural network (DNN) models based on the proposed loss function for end-to-end speaker verification. The training process employs a bipartite ranking method to deal with the trial imbalance problem and a curriculum learning method to help improve both the training stability and performance of the model by selecting non-target trials from easy to hard ones gradually along the convergence process. Since the training process employs bipartite ranking and curriculum learning and the loss function is of the generalized BCE form, we name the new approach \textit{curriculum bipartite ranking weighted binary cross-entropy} (CBRW-BCE). Experimental results show that the model trained with CBRW-BCE not only achieves the state-of-the-art performance but is also well calibrated.
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Speech Enhancement Aided End-To-End Multi-Task Learning for Voice Activity Detection
abstract
Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address this issue, here we propose a speech enhancement aided end-to-end multi-task model for VAD. The model has two decoders, one for speech enhancement and the other for VAD. The two decoders share the same encoder and speech separation network. Unlike the direct thought that takes two separated objectives for VAD and speech enhancement respectively, here we propose a new joint optimization objective—VAD-masked scale-invariant source-to-distortion ratio (mSI-SDR). mSI-SDR uses VAD information to mask the output of the speech enhancement decoder in the training process. It makes the VAD and speech enhancement tasks jointly optimized not only at the shared encoder and separation network, but also at the objective level. It also satisfies real-time working requirement theoretically. Experimental results show that the multi-task method significantly outperforms its single-task VAD counterpart. Moreover, mSI-SDR outperforms SI-SDR in the same multi-task setting.
Xiao-Lei Zhang 0001
ICASSP2
2021 Transformer-Based End-to-End Speech Recognition with Local Dense Synthesizer Attention
abstract
Recently, several studies reported that dot-product self-attention (SA) may not be indispensable to the state-of-the-art Transformer models. Motivated by the fact that dense synthesizer attention (DSA), which dispenses with dot products and pairwise interactions, achieved competitive results in many language processing tasks, in this paper, we first propose a DSA-based speech recognition, as an alternative to SA. To reduce the computational complexity and improve the performance, we further propose local DSA (LDSA) to restrict the attention scope of DSA to a local range around the current central frame for speech recognition. Finally, we combine LDSA with SA to extract the local and global information simultaneously. Experimental results on the Ai-shell1 Mandarin speech recognition corpus show that the proposed LDSA-Transformer achieves a character error rate (CER) of 6.49%, which is slightly better than that of the SA-Transformer. Meanwhile, the LDSA-Transformer requires less computation than the SA-Transformer. The proposed combination method not only achieves a CER of 6.18%, which significantly outperforms the SA-Transformer, but also has roughly the same number of parameters and computational complexity as the latter. The implementation of the multi-head LDSA is available at https://github.com/mlxu995/multihead-LDSA
Menglong Xu, Shengqiang Li, Xiao-Lei Zhang 0001
ICASSP3
2021 Scaling Sparsemax Based Channel Selection for Speech Recognition with ad-hoc Microphone Arrays
abstract
Recently, speech recognition with ad-hoc microphone arrays has received much attention. It is known that channel selection is an important problem of ad-hoc microphone arrays, however, this topic seems far from explored in speech recognition yet, particularly with a large-scale ad-hoc microphone array. To address this problem, we propose a Scaling Sparsemax algorithm for the channel selection problem of the speech recognition with large-scale ad-hoc microphone arrays. Specifically, we first replace the conventional Softmax operator in the stream attention mechanism of a multichannel end-to-end speech recognition system with Sparsemax, which conducts channel selection by forcing the channel weights of noisy channels to zero. Because Sparsemax punishes the weights of many channels to zero harshly, we propose Scaling Sparsemax which punishes the channels mildly by setting the weights of very noisy channels to zero only. Experimental results with ad-hoc microphone arrays of over 30 channels under the conformer speech recognition architecture show that the proposed Scaling Sparsemax yields a word error rate of over 30% lower than Softmax on simulation data sets, and over 20% lower on semi-real data sets, in test scenarios with both matched and mismatched channel numbers.
Junqi Chen 0001, Xiao-Lei Zhang 0001
Interspeech2
2021 Deep ad-hoc beamforming
Xiao-Lei Zhang 0001
Comput. Speech Lang.1
2021 Speaker recognition based on deep learning: An overview
Zhongxin Bai, Xiao-Lei Zhang 0001
Neural Networks2
2021 Minimum-Volume Multichannel Nonnegative Matrix Factorization for Blind Audio Source Separation
abstract
Multichannel blind audio source separation aims to recover the latent sources from their multichannel mixtures without supervised information. One state-of-the-art blind audio source separation method, named independent low-rank matrix analysis (ILRMA), unifies independent vector analysis (IVA) and nonnegative matrix factorization (NMF). However, the spectra matrix produced from NMF may not find a compact spectral basis. It may not guarantee the identifiability of each source as well. To address this problem, here we propose to enhance the identifiability of the source model by a minimum-volume prior distribution. We further regularize a multichannel NMF (MNMF) and ILRMA respectively with the minimum-volume regularizer. The proposed methods maximize the posterior distribution of the separated sources, which ensures the stability of the convergence. Experimental results demonstrate the effectiveness of the proposed methods compared with auxiliary independent vector analysis, MNMF, ILRMA and its extensions.
Shanzheng Guan, Shupei Liu, Xiao-Lei Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Partial AUC Optimization Based Deep Speaker Embeddings with Class-Center Learning for Text-Independent Speaker Verification
abstract
Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and identification. The verification loss functions match the pipeline of speaker verification, but their implementations are difficult. Thus, most state-of-the-art deep embedding methods use the identification loss functions with softmax output units or their variants. In this paper, we propose a verification loss function, named the maximization of partial area under the Receiver-operating-characteristic (ROC) curve (pAUC), for deep embedding based text-independent speaker verification. We also propose a class-center based training trial construction method to improve the training efficiency, which is critical for the proposed loss function to be comparable to the identification loss in performance. Experiments on the Speaker in the Wild (SITW) and NIST SRE 2016 datasets show that the proposed pAUC loss function is highly competitive with the state-of-the-art identification loss functions.
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen
ICASSP2
2020 Deep Topic Modeling by Multilayer Bootstrap Network and Lasso
abstract
Topic modeling is widely studied for the dimension reduction and analysis of documents. However, it is formulated as a difficult optimization problem. Current approximate solutions also suffer from inaccurate model-or data-assumptions. To deal with the above problems, we propose a polynomial-time deep topic model with no model and data assumptions. Specifically, we first apply multilayer bootstrap network (MBN), which is an unsupervised deep model, to reduce the dimension of documents, and then use the low-dimensional data representations or their clustering results as the target of supervised Lasso for topic word discovery. To our knowledge, this is the first time that MBN and Lasso are applied to unsupervised topic modeling. Experimental comparison results with five representative topic models on the 20-newsgroups and TDT2 corpora illustrate the effectiveness of the proposed algorithm.
Xiao-Lei Zhang 0001
ICPR2
2020 Depthwise Separable Convolutional ResNet with Squeeze-and-Excitation Blocks for Small-Footprint Keyword Spotting
abstract
One difficult problem of keyword spotting is how to miniaturize its memory footprint while maintain a high precision. Although convolutional neural networks have shown to be effective to the small-footprint keyword spotting problem, they still need hundreds of thousands of parameters to achieve good performance. In this paper, we propose an efficient model based on depthwise separable convolution layers and squeeze-and-excitation blocks. Specifically, we replace the standard convolution by the depthwise separable convolution, which reduces the number of the parameters of the standard convolution without significant performance degradation. We further improve the performance of the depthwise separable convolution by reweighting the output feature maps of the first convolution layer with a so-called squeeze-and-excitation block. We compared the proposed method with five representative models on two experimental settings of the Google Speech Commands dataset. Experimental results show that the proposed method achieves the state-of-the-art performance. For example, it achieves a classification error rate of 3.29% with a number of parameters of 72K in the first experiment, which significantly outperforms the comparison methods given a similar model size. It achieves an error rate of 3.97% with a number of parameters of 10K, which is also slightly better than the state-of-the-art comparison method given a similar model size.
Menglong Xu, Xiao-Lei Zhang 0001
INTERSPEECH2
2020 Cosine metric learning based speaker verification
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen
Speech Commun.2
2020 Speaker Verification by Partial AUC Optimization With Mahalanobis Distance Metric Learning
abstract
Receiver operating characteristic (ROC) and detection error tradeoff (DET) curves are two widely used evaluation metrics for speaker verification. They are equivalent since the latter can be obtained by transforming the former's true positive y-axis to false negative y-axis and then re-scaling both axes by a probit operator. Real-world speaker verification systems, however, usually work on part of the ROC curve instead of the entire ROC curve given an application. Therefore, we propose in this article to use the area under part of the ROC curve (pAUC) as a more efficient evaluation metric for speaker verification. A Mahalanobis distance metric learning based back-end is applied to optimize pAUC, where the Mahalanobis distance metric learning guarantees that the optimization objective of the back-end is a convex one so that the global optimum solution is achievable. To improve the performance of the state-of-the-art speaker verification systems by the proposed back-end, we further propose two feature preprocessing techniques based on length-normalization and probabilistic linear discriminant analysis respectively. We evaluate the proposed systems on the major languages of NIST SRE16 and the core tasks of SITW. Experimental results show that the proposed back-end outperforms the state-of-the-art speaker verification back-ends in terms of seven evaluation metrics.
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 AUC Optimization for Deep Learning Based Voice Activity Detection
abstract
Voice activity detection (VAD) based on deep neural networks (DNN) has demonstrated good performance in adverse acoustic environments. Current DNN based VAD optimizes a surrogate function, e.g. minimum cross-entropy or minimum squared error, at a given decision threshold. However, VAD usually works on-the-fly with a dynamic decision threshold; and ROC curve is a global evaluation metric of VAD that reflects the performance of VAD at all possible decision thresholds. In this paper, we propose to optimize the area under ROC curve (AUC) by DNN, which can maximize the performance of VAD in terms of the ROC curve. Experimental results show that optimizing AUC by DNN results in higher performance than the common method of optimizing the minimum squared error by DNN.
Zi-Chen Fan, Zhongxin Bai, Xiao-Lei Zhang 0001, Susanto Rahardja, Jingdong Chen
ICASSP3
2019 Robust Sparse Multichannel Active Noise Control
abstract
Multichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ1-norm constraint to the standard MC-ANC helps to reduce the complexity of the system and accelerate the convergence rate. However, the convergence performance of ℓ1-norm constrained MC-ANC (cℓ1-MC-ANC) degrades significantly in reverberant environments. In this paper, we analyze the necessity of using sparsity-inducing algorithms with distinct zero-attracting strengths over loudspeakers, and then derive three algorithms of this kind in the complex domain. Simulation results show that, compared to cℓ1-MC-ANC, the proposed algorithms exhibit faster convergence or higher noise reduction at steady state in both free field and reverberant environments.
Jingli Xie, Danqi Jin, Wen Zhang 0002, Xiao-Lei Zhang 0001, Jie Chen 0022, DeLiang Wang
ICASSP4
2019 Phase-Aware Speech Enhancement Based on Deep Neural Networks
abstract
Short-time frequency transform (STFT) is fundamental in speech processing. Because of the difficulty of processing highly unstructured STFT phase, most speech-processing algorithms only operate with STFT magnitude, leaving the STFT phase far from explored. However, with the recent development of deep neural network (DNN) based speech processing, e.g., speech enhancement and recognition, phase processing is becoming more important than ever before as a new growing point of DNN-based methods. In this paper, we propose a phase-aware speech enhancement algorithm based on DNN. Specifically, in the training stage, when incorporating phase as a target, our core idea is to transform an unstructured phase spectrogram to its derivative along the time axis, i.e., instantaneous frequency deviation (IFD), which has a similar structure with its corresponding magnitude spectrogram. We further propose to optimize both IFD and magnitude jointly in a multiobjective learning framework. In the test stage, we propose a postprocessing method to recover the phase spectrogram from the estimated IFD. Experimental results demonstrate the effectiveness of the proposed method.
Naijun Zheng, Xiao-Lei Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Cosine Metric Learning for Speaker Verification in the I-vector Space
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen
INTERSPEECH2
2018 Multilayer bootstrap networks
abstract
Multilayer bootstrap network builds a gradually narrowed multilayer nonlinear network from bottom up for unsupervised nonlinear dimensionality reduction. Each layer of the network is a nonparametric density estimator. It consists of a group of k-centroids clusterings. Each clustering randomly selects data points with randomly selected features as its centroids, and learns a one-hot encoder by one-nearest-neighbor optimization. Geometrically, the nonparametric density estimator at each layer projects the input data space to a uniformly-distributed discrete feature space, where the similarity of two data points in the discrete feature space is measured by the number of the nearest centroids they share in common. The multilayer network gradually reduces the nonlinear variations of data from bottom up by building a vast number of hierarchical trees implicitly on the original data space. Theoretically, the estimation error caused by the nonparametric density estimator is proportional to the correlation between the clusterings, both of which are reduced by the randomization steps.
Xiao-Lei Zhang 0001
Neural Networks1
2016 Universal Background Sparse Coding and Multilayer Bootstrap Network for Speaker Clustering
Xiao-Lei Zhang 0001
INTERSPEECH1
2016 Boosting Contextual Information for Deep Neural Network Based Voice Activity Detection
abstract
Voice activity detection (VAD) is an important topic in audio signal processing. Contextual information is important for improving the performance of VAD at low signal-to-noise ratios. Here we explore contextual information by machine learning methods at three levels. At the top level, we employ an ensemble learning framework, named multi-resolution stacking (MRS), which is a stack of ensemble classifiers. Each classifier in a building block inputs the concatenation of the predictions of its lower building blocks and the expansion of the raw acoustic feature by a given window (called a resolution). At the middle level, we describe a base classifier in MRS, named boosted deep neural network (bDNN). bDNN first generates multiple base predictions from different contexts of a single frame by only one DNN and then aggregates the base predictions for a better prediction of the frame, and it is different from computationally-expensive boosting methods that train ensembles of classifiers for multiple base predictions. At the bottom level, we employ the multi-resolution cochleagram feature, which incorporates the contextual information by concatenating the cochleagram features at multiple spectrotemporal resolutions. Experimental results show that the MRS-based VAD outperforms other VADs by a considerable margin. Moreover, when trained on a large amount of noise types and a wide range of signal-to-noise ratios, the MRS-based VAD demonstrates surprisingly good generalization performance on unseen test scenarios, approaching the performance with noise-dependent training.
Xiao-Lei Zhang 0001, DeLiang Wang
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 A Deep Ensemble Learning Method for Monaural Speech Separation
abstract
Monaural speech separation is a fundamental problem in robust speech processing. Recently, deep neural network (DNN)-based speech separation methods, which predict either clean speech or an ideal time-frequency mask, have demonstrated remarkable performance improvement. However, a single DNN with a given window length does not leverage contextual information sufficiently, and the differences between the two optimization objectives are not well understood. In this paper, we propose a deep ensemble method, named multicontext networks, to address monaural speech separation. The first multicontext network averages the outputs of multiple DNNs whose inputs employ different window lengths. The second multicontext network is a stack of multiple DNNs. Each DNN in a module of the stack takes the concatenation of original acoustic features and expansion of the soft output of the lower module as its input, and predicts the ratio mask of the target speaker; the DNNs in the same module employ different contexts. We have conducted extensive experiments with three speech corpora. The results demonstrate the effectiveness of the proposed method. We have also compared the two optimization objectives systematically and found that predicting the ideal time-frequency mask is more efficient in utilizing clean training speech, while predicting clean speech is less sensitive to SNR variations.
Xiao-Lei Zhang 0001, DeLiang Wang
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Multi-resolution stacking for speech separation based on boosted DNN
abstract
Recent progress in speech separation shows that deep neural networks (DNN) based supervised methods can improve the performance in difficult noise conditions and exhibit good generalization to unseen noise scenarios. However, existing approaches do not explore contextual information sufficiently. In this paper, we focus on exploring contextual information using DNN. The proposed method has two parts—a multi-resolution stacking (MRS) framework and a boosted DNN (bDNN) classifier. The MRS framework trains a stack of classifier ensembles, where each classifier in an ensemble concatenates the raw acoustic feature and the outputs of its bottom ensemble as a new feature, and different classifiers in an ensemble work with different window lengths. The bDNN classifier first generates multiple base predictions for a frame from a given window that is centered on the frame and contains multiple neighboring frames, and then aggregates the base predictions for the final prediction. Our experimental comparison with DNN based speech separation in difficult noise scenarios demonstrates the effectiveness of the proposed method in terms of both prediction accuracy and objective speech intelligibility.
Xiao-Lei Zhang 0001, DeLiang Wang
INTERSPEECH1
2015 Convex Discriminative Multitask Clustering
abstract
Multitask clustering tries to improve the clustering performance of multiple tasks simultaneously by taking their relationship into account. Most existing multitask clustering algorithms fall into the type of generative clustering, and none are formulated as convex optimization problems. In this paper, we propose two convex Discriminative Multitask Clustering (DMTC) objectives to address the problems. The first one aims to learn a shared feature representation, which can be seen as a technical combination of the convex multitask feature learning and the convex Multiclass Maximum Margin Clustering (M3C). The second one aims to learn the task relationship, which can be seen as a combination of the convex multitask relationship learning and M3C. The objectives of the two algorithms are solved in a uniform procedure by the efficient cutting-plane algorithm and further unified in the Bayesian framework. Experimental results on a toy problem and two benchmark data sets demonstrate the effectiveness of the proposed algorithms.
Xiao-Lei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Heuristic Ternary Error-Correcting Output Codes Via Weight Optimization and Layered Clustering-Based Approach
abstract
One important classifier ensemble for multiclass classification problems is error-correcting output codes (ECOCs). It bridges multiclass problems and binary-class classifiers by decomposing multiclass problems to a serial binary-class problems. In this paper, we present a heuristic ternary code, named weight optimization and layered clustering-based ECOC (WOLC-ECOC). It starts with an arbitrary valid ECOC and iterates the following two steps until the training risk converges. The first step, named layered clustering-based ECOC (LC-ECOC), constructs multiple strong classifiers on the most confusing binary-class problem. The second step adds the new classifiers to ECOC by a novel optimized weighted (OW) decoding algorithm, where the optimization problem of the decoding is solved by the cutting plane algorithm. Technically, LC-ECOC makes the heuristic training process not blocked by some difficult binary-class problem. OW decoding guarantees the nonincrease of the training risk for ensuring a small code length. Results on 14 UCI datasets and a music genre classification problem demonstrate the effectiveness of WOLC-ECOC.
Xiao-Lei Zhang 0001
IEEE Trans. Cybern.1
2014 Nonlinear Dimensionality Reduction of Data by Deep Distributed Random Samplings
Xiao-Lei Zhang 0001
ACML1
2014 Unsupervised domain adaptation for deep neural network based voice activity detection
abstract
The mismatching problem between the training and test speech corpora hinders the practical use of the machine-learning-based voice activity detection (VAD). In this paper, we try to address this problem by the unsupervised domain adaptation techniques, which try to find a shared feature subspace between the mismatching corpora. The denoising deep neural network is used as the learning machine. Three domain adaptation techniques are used for analysis. Experimental results show that the unsupervised domain adaptation technique is promising to the mismatching problem of VAD.
Xiao-Lei Zhang 0001
ICASSP1
2014 Boosted deep neural networks and multi-resolution cochleagram features for voice activity detection
abstract
Voice activity detection (VAD) is an important frontend of many speech processing systems. In this paper, we describe a new VAD algorithm based on boosted deep neural networks (bDNNs). The proposed algorithm first generates multiple base predictions for a single frame from only one DNN and then aggregates the base predictions for a better prediction of the frame. Moreover, we employ a new acoustic feature, multi-resolution cochleagram (MRCG), that concatenates the cochleagram features at multiple spectrotemporal resolutions and shows superior speech separation results over many acoustic features. Experimental results show that bDNN-based VAD with the MRCG feature outperforms state-of-the-art VADs by a considerable margin.
Xiao-Lei Zhang 0001, DeLiang Wang
INTERSPEECH1
2013 Denoising deep neural networks based voice activity detection
abstract
Recently, the deep-belief-networks (DBN) based voice activity detection (VAD) has been proposed. It is powerful in fusing the advantages of multiple features, and achieves the state-of-the-art performance. However, the deep layers of the DBN-based VAD do not show an apparent superiority to the shallower layers. In this paper, we propose a denoising-deep-neural-network (DDNN) based VAD to address the aforementioned problem. Specifically, we pre-train a deep neural network in a special unsupervised denoising greedy layer-wise mode, and then fine-tune the whole network in a supervised way by the common back-propagation algorithm. In the pre-training phase, we take the noisy speech signals as the visible layer and try to extract a new feature that minimizes the reconstruction cross-entropy loss between the noisy speech signals and its corresponding clean speech signals. Experimental results show that the proposed DDNN-based VAD not only outperforms the DBN-based VAD but also shows an apparent performance improvement of the deep layers over shallower layers.
Xiao-Lei Zhang 0001, Ji Wu 0002
ICASSP1
2013 Weight optimization and layered clustering-based ECOC
abstract
Error correcting output code (ECOC) is a general framework of solving a multiclass classification problem via a binary-class classifier ensemble. In this paper, we propose a new heuristic coding method, named weight optimization and layered clustering-based ECOC (WOLC-ECOC). It iterates the following two steps until the training risk converges. The first step employs the layered clustering-based approach [1]. The approach can construct multiple different strong binary-class classifiers on a given binary-class problem, so that the heuristic training process will not be blocked by some difficult binary-class problems. The second step is the weight optimization technique [2]. It guarantees the non-increasing of the heuristic training process whenever we add new classifiers to the ECOC ensemble. Experimental results on several benchmark sets demonstrate that WOLC-ECOC is more effective than 15 referenced coding-decoding ECOC pairs.
Xiao-Lei Zhang 0001, Ji Wu 0002
ICASSP1
2013 Deep Belief Networks Based Voice Activity Detection
abstract
Fusing the advantages of multiple acoustic features is important for the robustness of voice activity detection (VAD). Recently, the machine-learning-based VADs have shown a superiority to traditional VADs on multiple feature fusion tasks. However, existing machine-learning-based VADs only utilize shallow models, which cannot explore the underlying manifold of the features. In this paper, we propose to fuse multiple features via a deep model, called deep belief network (DBN). DBN is a powerful hierarchical generative model for feature extraction. It can describe highly variant functions and discover the manifold of the features. We take the multiple serially-concatenated features as the input layer of DBN, and then extract a new feature by transferring these features through multiple nonlinear hidden layers. Finally, we predict the class of the new feature by a linear classifier. We further analyze that even a single-hidden-layer-based belief network is as powerful as the state-of-the-art models in the machine-learning-based VADs. In our empirical comparison, ten common features are used for performance analysis. Extensive experimental results on the AURORA2 corpus show that the DBN-based VAD not only outperforms eleven referenced VADs, but also can meet the real-time detection demand of VAD. The results also show that the DBN-based VAD can fuse the advantages of multiple features effectively.
Xiao-Lei Zhang 0001, Ji Wu 0002
IEEE Trans. Speech Audio Process.1
2012 Optimized weighted decoding for error-correcting output codes
abstract
A common method to solve a multiclass classification problem is to reduce the problem to a serial binary classification problems and combine them via Error-Correcting Output Codes (ECOC). The ECOC contains three parts: coding design, decoding algorithm, and base dichotomizer. Recently, the Loss-Weighted (LW) decoding algorithm (Escalera et al., PAMI2010), which introduces a weight matrix to the Loss-Based (LB) decoding (Allwein et al., JMLR2001), achieves improved performance over traditional decoding methods. However, the weight matrix is assigned empirically. In this paper, we present a theoretical global optimization method for the weight matrix, so as to achieve the minimal training risk. Although the experimental results on real-world image, audio and text classification tasks show that the proposed decoding method only leads to slightly better performances than others in the case of discrete outputs of the dichotomizers, the proposed method provides a new screen on the decoding methods of the ECOC.
Xiao-Lei Zhang 0001, Ji Wu 0002, Ping Lv
ICASSP1
2012 Linearithmic Time Sparse and Convex Maximum Margin Clustering
abstract
Recently, a new clustering method called maximum margin clustering (MMC) was proposed and has shown promising performances. It was originally formulated as a difficult nonconvex integer problem. To make the MMC problem practical, the researchers either relaxed the original MMC problem to inefficient convex optimization problems or reformulated it to nonconvex optimization problems, which sacrifice the convexity for efficiency. However, no approaches can both hold the convexity and be efficient. In this paper, a new linearithmic time sparse and convex MMC algorithm, called support-vector-regression-based MMC (SVR-MMC), is proposed. Generally, it first uses the SVR as the core of the MMC. Then, it is relaxed as a convex optimization problem, which is iteratively solved by the cutting-plane algorithm. Each cutting-plane subproblem is further decomposed to a serial supervised SVR problem by a new global extended-level method (GELM). Finally, each supervised SVR problem is solved in a linear time complexity by a new sparse-kernel SVR (SKSVR) algorithm. We further extend the SVR-MMC algorithm to the multiple-kernel clustering (MKC) problem and the multiclass MMC (M3C) problem, which are denoted as SVR-MKC and SVR-M3C, respectively. One key point of the algorithms is the utilization of the SVR. It can prevent the MMC and its extensions meeting an integer matrix programming problem. Another key point is the new SKSVR. It provides a linear time interface to the nonlinear kernel scenarios, so that the SVR-MMC and its extensions can keep a linearthmic time complexity in nonlinear kernel scenarios. Our experimental results on various real-world data sets demonstrate the effectiveness and the efficiency of the SVR-MMC and its two extensions. Moreover, the unsupervised application of the SVR-MKC to the voice activity detection (VAD) shows that the SVR-MKC can achieve good performances that are close to its supervised counterpart, meet the real-time demand of the VAD, and need no labeling for model training.
Xiao-Lei Zhang 0001, Ji Wu 0002
IEEE Trans. Syst. Man Cybern. Part B1
2011 Maximum Margin Clustering Based Statistical VAD With Multiple Observation Compound Feature
abstract
In this letter, we propose a new robust feature and an unsupervised learning approach for statistical voice activity detection (VAD). Maximum margin clustering (MMC), as an unsupervised classifier, can improve the robustness of support vector machine (SVM) based VAD while requiring no data labeling for model training. In the MMC framework, the multiple observation compound feature (MO-CF) is proposed to improve accuracy. MO-CF is composed of two subfeatures—multiple observation signal-to-noise ratio (MO-SNR) and multiple observation maximum probability (MO-MP). The contributions of the two subfeatures are balanced by a factor which is chosen to yield the largest area under the ROC curve (AUC) of the performance. The proposed approach obtains improved performance over seven commonly used VAD techniques in the experiments covering various noisy scenarios with low SNRs.
Ji Wu 0002, Xiao-Lei Zhang 0001
IEEE Signal Process. Lett.2
2011 Efficient Multiple Kernel Support Vector Machine Based Voice Activity Detection
abstract
In this letter, we propose a multiple kernel support vector machine (MK-SVM) method for multiple feature based VAD. To make the MK-SVM based VAD practical, we adapt the multiple kernel learning (MKL) thought to an efficient cutting-plane structural SVM solver. We further discuss the performances of the MK-SVM with two different optimization objectives, in terms of minimum classification errors (MCE) and improvement of receiver operating characteristic (ROC) curves. Our experimental results show that the proposed method not only leads to better global performances by taking the advantages of multiple features but also has a low computational complexity.
Ji Wu 0002, Xiao-Lei Zhang 0001
IEEE Signal Process. Lett.2
2010 A new VAD framework using statistical model and human knowledge based empirical rule
Ji Wu 0002, Xiao-Lei Zhang 0001
INTERSPEECH2