VLDB 2026 Research / reviewers in the wild / expert
Shilei Zhang
dblp:10/3523
· DBLP profile ↗
47ranked-venue papers
14as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 13 first-author · 19 since 2021Artificial intelligence and machine learning · 29 · 8 first-author · 17 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech CodecabstractTao Li, Wenshuo Ge, Zhichao Wang, Zihao Cui, Yong Ma, Yingying Gao, Chao Deng, Shilei Zhang, Junlan Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenshuo Ge, Zihao Cui, Yingying Gao, Chao Deng 0002, Shilei Zhang, Junlan Feng |
ACL (1) | 8 |
| 2026 | Modeling bidirectional modes of an event for temporal knowledge graph reasoning
Zepeng Li 0003, Rikui Huang, Shilei Zhang, Zhenwen Zhang, Jianghong Zhu |
Expert Syst. Appl. | 4 |
| 2025 | DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable StylesabstractHuman speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional diffusion module and an improved classifier-free guidance, which hierarchically models speech prosodic features, and controls different prosodic styles to guide prosody prediction. Experiments show that our method outperforms all baselines in naturalness and achieves superior synthesis speed compared to three diffusion-based baselines. Additionally, by adjusting the guiding scale, DiffStyleTTS effectively controls the guidance intensity of the synthetic prosody. Zhaoci Liu, Yajun Hu, Yingying Gao, Shilei Zhang, Zhen-Hua Ling |
COLING | 5 |
| 2025 | Energy-based Model Guided Self-Supervised Learning for Speaker VerificationabstractSelf-supervised learning (SSL) has significantly advanced speaker verification, especially in scenarios with limited labeled data. This paper introduces Energy-based Confidence-Aware Distillation (EBCA-DINO), an SSL enhancement for speaker verification that integrates Energy-Based Models (EBMs) into the DINO (Distillation with No Labels) framework. EBMs use energy scores to assess data complexity and uncertainty, guiding label-free self-distillation. The adaptive temperature scaling tailors the learning process to data characteristics, allowing the teacher model to dynamically adjust the student model’s focus based on sample difficulty. This energy-aware distillation optimizes speaker verification performance. Experimental results demonstrate that EBCA-DINO improves speaker verification with relative performance gains of 4.3%, 4.9%, and 8.7% on the Vox1-O, E, and H test trials, respectively. Yaqian Hao, Chenguang Hu, Chong Bian, Junlan Feng, Yingying Gao, Shilei Zhang |
ICASSP | 6 |
| 2025 | Codec-ASV: Exploring Neural Audio Codec For Speaker Representation LearningabstractDiscrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be largely preserved in the compression process since the reconstructed voice is almost the same in human listening. In this paper, we explore various training strategies and codec types for NAC-based speaker representation learning. Using ECAPA-TDNN as the model backbone, our approach achieves state-of-the-art performance with a 2.08% EER in NAC-based speaker verification scenarios. To better retain speaker information in early, more compressed layers, we introduce mask-layer augmentation and embedding fusion techniques during the training process. Experimental results show the effectiveness of our methods, particularly when inferring with limited codec layers. Yuke Lin, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026 |
ICASSP | 4 |
| 2025 | Efficient Extreme Large-Scale Speaker Verification: Dynamic Active Sub Fully-Connected Layers for Faster Training and Memory OptimizationabstractUsing larger scale datasets in the training stage of speaker verification model usually leads to better performance. However, when the speaker number of the training dataset becomes extreme large (e.g., more than 1 million), the training speed and GPU memory demand will become bottlenecks which are mainly brought by the extreme large dimension of last fully-connected(FC) layer’s weight matrix. We propose dynamic active sub FC layers (DAS-FC) to tackle this problem. Firstly, all speakers are dynamically divided into speaker groups by clustering rows of last FC layer’s weight matrix. Then, sub FC layers are generated according to speaker groups for model training. We also introduce Mini-Batch K-means and speaker based dataloader to further reduce time and resource costing. Experiments on an extreme large dataset with 1,068,237 speakers show that compared to traditional FC layer, DAS-FC can save up to 87% training time and save 56% GPU memory occupancy with only a 4.2% drop in model performance. Fulin Zhang, Chenguang Hu, Yingying Gao, Shilei Zhang, Junlan Feng |
ICASSP | 5 |
| 2025 | Privacy-Preserving Speaker Verification via End-to-End Secure Representation Learning
Chenguang Hu, Yaqian Hao, Fulin Zhang, Xiaoxue Luo, Yingying Gao, Chao Deng 0002, Shilei Zhang, Junlan Feng |
INTERSPEECH | 8 |
| 2024 | MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
Haiyang Sun 0004, Fulin Zhang, Yingying Gao, Shilei Zhang, Zheng Lian 0004, Junlan Feng |
INTERSPEECH | 4 |
| 2024 | GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
Yingying Gao, Shilei Zhang, Chao Deng 0002, Junlan Feng |
INTERSPEECH | 2 |
| 2024 | Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
Yaqian Hao, Chenguang Hu, Yingying Gao, Shilei Zhang, Junlan Feng |
INTERSPEECH | 4 |
| 2024 | On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
Yaqian Hao, Chenguang Hu, Yingying Gao, Shilei Zhang, Junlan Feng |
INTERSPEECH | 4 |
| 2024 | VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Yuke Lin, Ming Cheng 0005, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026 |
INTERSPEECH | 5 |
| 2024 | CEC: A Noisy Label Detection Method for Speaker Recognition
Yingying Gao, Yaqian Hao, Chenguang Hu, Fulin Zhang, Junlan Feng, Shilei Zhang |
INTERSPEECH | 7 |
| 2023 | Semi-Supervised Speech Enhancement Based On Speech PurityabstractWe tend to assume most available speech corpora we use are either completely clean or completely noised. However, the reality is most of them are a mix of both. In this paper, we propose a semi-supervised speech enhancement framework to enhance such typical speech datasets. This framework includes an estimator to measure the speech purity. Utterances with high speech purity are considered clean, otherwise noised. For clean speech utterances, we follow the supervised learning mechanism to train a deep learning speech enhancement model. For noised speech, we update the model in an unsupervised manner. Hence, we design our training loss as a combination of the supervised loss and unsupervised loss. We refer to this framework as SemiEnhance. Experimental results show that SemiEnhance substantially improves the speech quality, and achieves new state-of-the-art results on benchmark datasets: 2022 DNS Challenge and NoiseX-92. Zihao Cui, Shilei Zhang, Yingying Gao, Chao Deng 0002, Junlan Feng |
ICASSP | 2 |
| 2023 | VE-KWS: Visual Modality Enhanced End-to-End Keyword SpottingabstractThe performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting, which leverages complementary relationships over multiple modalities, has recently gained much attention. However, current studies mainly focus on combining the exclusively learned representations of different modalities, instead of exploring the modal relationships during each respective modeling. In this paper, we propose a novel visual modality enhanced end-to-end KWS framework (VE-KWS), which fuses audio and visual modalities from two aspects. The first one is utilizing the speaker location information obtained from the lip region in videos to assist the training of multi-channel audio beamformer. By involving the beamformer as an audio enhancement module, the acoustic distortions, caused by the far field or noisy environments, could be significantly suppressed. The other one is conducting cross-attention between different modalities to capture the inter-modal relationships and help the representation learning of each modality. Experiments on the MSIP challenge corpus show that our proposed model achieves a 2.79% false rejection rate and a 2.95% false alarm rate on the Eval set, resulting in a new SOTA performance compared with the top-ranking systems in the ICASSP2022 MISP challenge. He Wang 0022, Yihui Fu, Lei Xie 0001, Yingying Gao, Shilei Zhang, Junlan Feng |
ICASSP | 7 |
| 2023 | Cascaded Multi-task Adaptive Learning Based on Neural Architecture Search
Yingying Gao, Shilei Zhang, Zihao Cui, Chao Deng 0002, Junlan Feng |
INTERSPEECH | 2 |
| 2023 | Flow-VAE VC: End-to-End Flow Framework with Contrastive Loss for Zero-shot Voice Conversion
Rongxiu Zhong, Huibao Yang, Shilei Zhang |
INTERSPEECH | 5 |
| 2023 | Harmonic Attention for Monaural Speech EnhancementabstractTo further improve the quality of the enhanced speech, it is appealing that more profound articulatory and auditory knowledge should be introduced into the speech enhancement model. Among these, harmonics seriously affect speech timbre and play a crucial role in speech intelligibility. Especially in the frequency domain, harmonics appear as the local maximum peaks of energy, which could be expected to serve as anchors to recover the distorted speech. In this paper, an explicit modeling method, harmonic attention, is presented, patching the harmonics with the help of residual ones. In order to maintain the spectral structure of speech during the processing and to enable the network to support harmonic modeling, a harmonic attention-based progressive enhancement network (HAPNet) is applied, which gradually approaches clean speech with stacked modules of harmonic attention. In addition, to make enhanced speech more consistent with hearing, a loss function based on the loudness power compression (LC-SNR) is used, which measures both magnitude and phase values with appropriate auditory effects. The experimental visualization indicates that the harmonic attention can capture and recover the harmonics of speech. And the objective evaluations show that the presented HAPNet and LC-SNR outperform the referenced methods. Furthermore, the presented model trained on 100 hours of data achieves competitive results with the referenced models trained on 3000+ hours of data, and one trained on 500 hours of data yields the state-of-the-art performance. Tianrui Wang, Weibin Zhu, Yingying Gao, Shilei Zhang, Junlan Feng |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS ChallengeabstractThe harmonic structure of speech is resistant to noise, but the harmonics may still be partially masked by noise. Therefore, we previously proposed a harmonic gated compensation network (HGCN) to predict the full harmonic locations based on the unmasked harmonics and process the result of a coarse enhancement module to recover the masked harmonics. In addition, the auditory loudness loss function is used to train the network. For the DNS Challenge, we update HGCN with the following aspects, resulting in HGCN+. First, a high-band module is employed to help the model handle full-band signals. Second, cosine is used to model the harmonic structure more accurately. Then, the dual-path encoder and dual-path rnn (DPRNN) are introduced to take full advantage of the features. Finally, a gated residual linear structure replaces the gated convolution in the compensation module to increase the receptive field of frequency. The experimental results show that each updated module brings performance improvement to the model. HGCN+ also outperforms the referenced models on both wide-band and full-band test sets. Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng, Shilei Zhang |
ICASSP | 6 |
| 2022 | HGCN: Harmonic Gated Compensation Network for Speech EnhancementabstractMask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmonics are partially masked by noise. To tackle this challenge, we propose a harmonic gated compensation network (HGCN). We design a high-resolution harmonic integral spectrum to improve the accuracy of harmonic locations prediction. Then we add voice activity detection (VAD) and voiced region detection (VRD) to the convolutional recurrent network (CRN) to filter harmonic locations. Finally, the harmonic gating mechanism is used to guide the compensation model to adjust the coarse results from CRN to obtain the refinedly enhanced results. Our experiments show HGCN achieves substantial gain over a number of advanced approaches in the community. Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng, Shilei Zhang |
ICASSP | 5 |
| 2022 | Meta Auxiliary Learning for Low-resource Spoken Language UnderstandingabstractSpoken language understanding (SLU) treats automatic speech recognition (ASR) and natural language understanding (NLU) as a unified task and usually suffers from data scarcity.We exploit an ASR and NLU joint training method based on meta auxiliary learning to improve the performance of low-resource SLU task by only taking advantage of abundant manual transcriptions of speech data.One obvious advantage of such method is that it provides a flexible framework to implement a lowresource SLU training task without requiring access to any further semantic annotations.In particular, a NLU model is taken as label generation network to predict intent and slot tags from texts; a multi-task network trains ASR task and SLU task synchronously from speech; and the predictions of label generation network are delivered to the multi-task network as semantic targets.The efficiency of the proposed algorithm is demonstrated with experiments on the public CATSLU dataset, which produces more suitable ASR hypotheses for the downstream NLU task. Yingying Gao, Junlan Feng, Chao Deng 0002, Shilei Zhang |
INTERSPEECH | 4 |
| 2022 | Two-stage streaming keyword detection and localization with multi-scale depthwise temporal convolution
Jingyong Hou, Lei Xie 0001, Shilei Zhang |
Neural Networks | 3 |
| 2021 | Boundary and Context Aware Training for CIF-Based Non-Autoregressive End-to-End ASRabstractContinuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with competitive performance compared with other NAR methods. However, such an alignment learning strategy may suffer from an erroneous acoustic boundary estimation, severely hindering the convergence speed as well as the system performance. In this paper, we propose a boundary and context aware training approach for CIF based NAR models. Firstly, the connectionist temporal classification (CTC) spike information is utilized to guide the learning of acoustic boundaries in the CIF. Besides, an additional contextual decoder is introduced behind the CIF decoder, aiming to capture the linguistic dependencies within a sentence. Finally, we adopt a recently proposed Conformer architecture to improve the capacity of acoustic modeling. Experiments on the open-source Mandarin AISHELL-1 corpus show that the proposed method achieves a comparable character error rates (CERs) of 4.9% with only 1/24 latency compared with a state-of-the-art autoregressive (AR) Conformer model. Futhermore, when evaluating on an internal 7500 hours Mandarin corpus, our model still outperforms other NAR methods and even reaches the AR Conformer model on a challenging real-world noisy test set. Fan Yu 0002, Haoneng Luo, Yuhao Liang, Zhuoyuan Yao, Lei Xie 0001, Yingying Gao, Leijing Hou, Shilei Zhang |
ASRU | 9 |
| 2021 | Our Learned Lessons from Cross-Lingual Speaker Verification: The CRMI-DKU System Description for the Short-Duration Speaker Verification Challenge 2021
Xiaoyi Qin, Chao Wang 0111, Shilei Zhang, Ming Li 0026 |
Interspeech | 5 |
| 2019 | Few-Shot Audio Classification with Attentional Graph Neural Networks
Shilei Zhang, Yong Qin 0001, Kewei Sun, Yonghua Lin |
INTERSPEECH | 1 |
| 2018 | Identity-Enhanced Network for Facial Expression Recognition
Xingang Wang 0003, Shilei Zhang, Lingxi Xie, Hongyuan Yu |
ACCV (4) | 3 |
| 2017 | Service failure diagnosis in service function chainabstractNetwork function virtualization (NFV) is a powerful emerging technique with widespread applicability. It provides Network Functions (NFs) through software virtualization techniques that decouple software and hardware. Some connected network functions constitute a service function chain (SFC). Therefore, the deployment of SFCs is much agile and simple. However, this leads to more service failure. The service failure includes service availability failure and service quality degradation. Aiming at the problem that the existing service function chain detection methods have high detection cost and cannot locate the failure accurately. This paper presents a method based on minimum detection cost. The method consists of failure detection and failure localization. In failure detection, we calculate detection paths according to the topology of network functions to avoid duplicate probing of links between network functions. In failure localization, we locate service availability failure and service quality degradation respectively and add timestamp fields to network service header to analyze locations of service quality degradation. Experiments show that the method reduces active detection cost and improves the recall and false-positive of service failure localization. Shilei Zhang, Ying Wang 0002, Wenjing Li 0001, Xuesong Qiu 0001 |
APNOMS | 1 |
| 2017 | Load balancing for multiple controllers in SDN based on switches groupabstractSoftware-Defined Networking (SDN) develops a centralized control plane to manage the whole network, but the traditional centralized control plane suffers from the issues of reliability and scalability. Although some methods solves the issues, they do not balance load among controllers. In this paper, we propose a load balancing mechanism based on switches group for multiple controllers. The mechanism not only balances the load among controllers, but also solves the load oscillation and improves time efficiency. Experiments based on floodlight show that our mechanism can improve network balancing and the time efficiency, at same time, avoid load oscillation problem among controllers. Yaning Zhou, Jinke Yu, Junhua Ba, Shilei Zhang |
APNOMS | 5 |
| 2017 | Emotion recognition with multimodal features and temporal modelsabstractThis paper presents our methods to the Audio-Video Based Emotion Recognition subtask in the 2017 Emotion Recognition in the Wild (EmotiW) Challenge. The task aims to predict one of the seven basic emotions for short video segments. We extract different features from audio and facial expression modalities. We also explore the temporal LSTM model with the input of frame facial features, which improves the performance of the non-temporal model. The fusion of different modality features and the temporal model lead us to achieve a 58.5% accuracy on the testing set, which shows the effectiveness of our methods. Wenxuan Wang 0001, Jinming Zhao, Shizhe Chen, Qin Jin, Shilei Zhang, Yong Qin 0001 |
ICMI | 6 |
| 2017 | Autoencoder Regularized Network For Driving Style Representation LearningabstractIn this paper, we study learning generalized driving style representations from automobile GPS trip data. We propose a novel Autoencoder Regularized deep neural Network (ARNet) and a trip encoding framework trip2vec to learn drivers' driving styles directly from GPS records, by combining supervised and unsupervised feature learning in a unified architecture. Experiments on a challenging driver number estimation problem and the driver identification problem show that ARNet can learn a good generalized driving style representation: It significantly outperforms existing methods and alternative architectures by reaching the least estimation error on average (0.68, less than one driver) and the highest identification accuracy (by at least 3% improvement) compared with traditional supervised learning methods. Weishan Dong, Shilei Zhang |
IJCAI | 5 |
| 2016 | Video emotion recognition in the wild based on fusion of multimodal featuresabstractIn this paper, we present our methods to the Audio-Video Based Emotion Recognition subtask in the 2016 Emotion Recognition in the Wild (EmotiW) Challenge. The task is to predict one of the seven basic emotions for the characters in the video clips extracted from movies or TV shows. In our approach, we explore various multimodal features from audio, facial image and video motion modalities. The audio features contain statistical acoustic features, MFCC Bag-of-Audio-Words and MFCC Fisher Vectors. For image related features, we extract hand-crafted features (LBP-TOP and SPM Dense SIFT) and learned features (CNN features). The improved Dense Trajectory is used as the motion related features. We train SVM, Random Forest and Logistic Regression classifiers for each kind of feature. Among them, MFCC fisher vector is the best acoustic features and the facial CNN feature is the most discriminative feature for emotion recognition. We utilize late fusion to combine different modality features and achieve a 50.76% accuracy on the testing set, which significantly outperforms the baseline test accuracy of 40.47%. Shizhe Chen, Qin Jin, Shilei Zhang, Yong Qin 0001 |
ICMI | 4 |
| 2016 | Wake-up-word spotting using end-to-end deep neural network systemabstractDeep neural networks (DNNs) have tremendously improved the performance of automatic speech recognition (ASR). On the other hand, end-to-end speech recognition system can achieve state-of-the-art performance using Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) and Connectionist Temporal Classification (CTC) method for unsegmented sequence data. In this paper, we therefor propose a lightweight wake-up-word (WUW) spotting system based on end-to-end DNN architecture, which is intended to provide a great balance between decoding speed, accuracy and model size. The objective is to introduce CTC framework on spotting process, and to enhance the system by WUW-oriented model training and refinement steps. We test the performance of the proposed architecture on a conversational telephone dataset which illustrate that the computation time can be significantly reduced without a significant decrease in the spotting accuracy. Shilei Zhang, Yong Qin 0001 |
ICPR | 1 |
| 2016 | Rapid feature space MLLR speaker adaptation for deep neural network acoustic modelingabstractBilinear models based feature space Maximum Likelihood Linear Regression (FMLLR) speaker adaptation have showed good performance for GMM-HMMs especially when the amount of adaptation data is limited. In this paper, we propose using bilinear models feature as inputs to deep neural networks (DNNs) for rapid speaker adaptation of acoustic modeling to facilitate utterance-level normalization. The effectiveness of the proposed method is demonstrated with experiments on the Mandarin short message dictation and voice query dataset. Shilei Zhang, Yong Qin 0001 |
ICPR | 1 |
| 2016 | Text-independent voice conversion using deep neural network based phonetic level featuresabstractThis paper presents a phonetically-aware joint density Gaussian mixture model (JD-GMM) framework for voice conversion that no longer requires parallel data from source speaker at the training stage. Considering that the phonetic level features contain text information which should be preserved in the conversion task, we propose a method that only concatenates phonetic discriminant features and spectral features extracted from the same target speakers speech to train a JD-GMM. After the mapping relationship of these two features is trained, we can use phonetic discriminant features from source speaker to estimate target speaker's spectral features at conversion stage. The phonetic discriminant features are extracted using PCA from the output layer of a deep neural network (DNN) in an automatic speaker recognition (ASR) system. It can be seen as a low dimensional representation of the senone posteriors. We compare the proposed phonetically-aware method with conventional JD-GMM method on the Voice Conversion Challenge 2016 training database. The experimental results show that our proposed phonetically-aware feature method can obtain similar performance compared to the conventional JD-GMM in the case of using only target speech as training data. Huadi Zheng, Weicheng Cai, Tianyan Zhou, Shilei Zhang, Ming Li 0026 |
ICPR | 4 |
| 2013 | Semi-supervised accent detection and modelingabstractIn this paper, we propose an iterative refinement framework for semi-supervised accent detection, where the accent labels of training corpus were generated by the user's self-judgement with poor accuracy. Firstly, we get the initial accent detection models based on cross-validation (CV) method, and then select the pure accent samples iteratively based on cost criterion derived from neighbor function, which is sensitive to the accent class purity. SVM based accent recognition approach is applied as the basic accent detection method which assumes that certain phones are realized differently across accents. Finally, we update the accent specific acoustic models via adaptation based on the detected specific accent data. The efficiency of the proposed method is demonstrated with experiments on English dictation database. Shilei Zhang, Yong Qin 0001 |
ICASSP | 1 |
| 2012 | Model dimensionality selection in bilinear transformation for feature space MLLR rapid speaker adaptationabstractBilinear models based feature space Maximum Likelihood Linear Regression (FMLLR) speaker adaptation have showed good performance especially when the amount of adaptation data is limited. However, the model dimensionality selection is very critical to the performance of bilinear models and need more work to find the optimal selection method. In this paper, we present an empirical study on this issue and suggest using a piecewise log-linear function to describe the relationship between the relatively optimal dimensionality parameter and the variant amount of data. This relationship can be used to efficiently select the bilinear model dimensionality in FMLLR speaker adaptation with the variant amount of data for each test speaker to improve recognition performance on the English voice control dataset. Shilei Zhang, Yong Qin 0001 |
ICASSP | 1 |
| 2011 | Rapid feature space MLLR speaker adaptation with bilinear modelsabstractIn this paper, we propose a novel method for rapid feature space Maximum Likelihood Linear Regression (FMLLR) speaker adaptation based on bilinear models. When the amount of adaptation data is limited, the conventional FMLLR transforms can be easily over-trained and can even degrade the performance. In such cases, usually by introducing structural constraints on the FMLLR transformation, the original FMLLR adaptation method can be modified for rapid adaptation. The objective of our bilinear model is to introduce a prior knowledge analysis on the training speakers based on Singular Vector Decomposition (SVD), and to incorporate it in the decoding process. This can effectively reduce the number of free parameters of FMLLR transformation and achieve performance improvements even with limited adaptation data. The efficiency of the proposed algorithm is demonstrated with experiments on the Mandarin digital dataset and the Mandarin voice search dataset respectively. Shilei Zhang, Peder A. Olsen, Yong Qin 0001 |
ICASSP | 1 |
| 2010 | The 2009 IBM GALE Mandarin broadcast transcription systemabstractThis paper gives an up-to-date description of the IBM Mandarin broadcast transcription system developed under the DARPA GALE program. Technical advances over our previous system include a novel acoustic modeling approach using subspace Gaussian mixture models, a speaking rate adaptation method using frame rate normalization, and an effective recipe for lattice combination. We present results on three consortium-defined test sets. It is shown that with these advances, the new system attains a 9% relative reduction in character error rate compared to our previous GALE evaluation system. The reported 9.1% error rate on the phase three evaluation set represents the state of the art in Mandarin broadcast speech transcription. Stephen M. Chu, Daniel Povey, Hong-Kwang Jeff Kuo, Lidia Mangu, Shilei Zhang, Qin Shi 0001, Yong Qin 0001 |
ICASSP | 5 |
| 2010 | Modeling Syllable-Based Pronunciation Variation for Accented Mandarin Speech RecognitionabstractPronunciation variation is a natural and inevitable phenomenon in an accented Mandarin speech recognition application. In this paper, we integrate knowledge-based and data-driven approaches together for syllable-based pronunciation variation modeling to improve the performance of Mandarin speech recognition system for speakers with Southern accent. First, we generate the syllable-based pronunciation variation rules of Southern accent observed from the training corpus by Chinese linguistic expert. Second, dictionary augmentation with multiple pronunciation variants and pronunciation probability derived from forced alignment statistics of training data. The acoustic models will be retrained based on the new expansion dictionary. Finally, pronunciation variation adaptation will be performed to further fit the data on the decoding stage by taking distribution of variation rules clusters of testing set into account. The experimental results show that the proposed method provides a flexible framework to improve the recognition performance for accented speech effectively. Shilei Zhang, Qin Shi 0001, Yong Qin 0001 |
ICPR | 1 |
| 2010 | Automatic Pronunciation Transliteration for Chinese-English Mixed Language Keyword SpottingabstractThis paper presents automatic pronunciation transliteration method with acoustic and contextual analysis for Chinese-English mixed language keyword spotting (KWS) system. More often, we need to develop robust Chinese-English mixed language spoken language technology without Chinese accented English acoustic data. In this paper, we exploit pronunciation conversion method based on syllable-based characteristic analysis of pronunciation and data-driven phoneme pairs mappings to solve mixed language problem by only using well-trained Chinese models. One obvious advantage of such method is that it provides a flexible framework to implement the pronunciation conversion of English keywords to Chinese automatically. The efficiency of the proposed method was demonstrated under KWS task on mixed language database. Shilei Zhang, Zhiwei Shuang, Yong Qin 0001 |
ICPR | 1 |
| 2010 | Improved Mandarin Keyword Spotting Using Confusion Garbage ModelabstractThis paper presents an improved acoustic keyword spotting (KWS) algorithm using a novel confusion garbage model in Mandarin conversational speech. Observing the KWS corpus, we found there are many words with similar pronunciation with predefined keywords, although they have different Chinese characters and different meanings, which easily result in high false alarm rate. In this paper, an improved acoustic KWS method with confusion garbage models was developed that absorbs similar pronunciation words confused with specific keywords for a given task. One obvious advantage of such method is that it provides a flexible framework to implement the selection procedure and reduce false alarm rate effectively for a specific task. The efficiency of the proposed architecture was evaluated under HMM-based confidence measures (CM) methods and demonstrated on a conversational telephone dataset. Shilei Zhang, Zhiwei Shuang, Qin Shi 0001, Yong Qin 0001 |
ICPR | 1 |
| 2010 | Spoken English assessment system for non-native speakers using acoustic and prosodic featuresabstractThe absence of real-time and targeted feedback is often critical in spoken foreign language learning. Computer-assisted language assessment systems are playing an ever more important role in this domain. This work considers the idiosyncratic pronunciation patterns of Chinese English speakers and uses both acoustic and prosody features to capture pronunciation, word stress, and rhythm information. The proposed system uses a. automatic speech recognition and alignment for pronunciation assessment, b. a set of special features with appropriate normalization for word stress detection, and c. a prosody phrase prediction model for rhythm assessment; and is shown to give immediate and accurate analyses to speakers to improve learning efficiency. Qin Shi 0001, Shilei Zhang, Stephen M. Chu, Ji Xiao, Zhijian Ou |
INTERSPEECH | 3 |
| 2009 | Utterance verification using improved confidence measures based on alignment confusion rate in Chinese digits recognitionabstractIn this paper, we explore an approach to improved confidence measures based on a novel alignment confusion rate (ACR) which integrates alignment information from two different modeling unit sets in Chinese digits recognition system. Both initial-final (IF) phone set and head-body-tail (HBT) models have proven to obtain good recognition performance for connected digit strings. These two different modeling can produce similar results but with different time-marked word boundaries. The objective of our proposed method is combining posterior probability with alignment confusion rate score provided by word alignment of IF-based results to HBT-based reference results that minimizes word error rate to get an effective confidence measure for utterance verification. The efficiency of the proposed algorithm is demonstrated with various experiments on data collected from car-kit microphone. Shilei Zhang, Danning Jiang, Yong Qin 0001 |
ICASSP | 1 |
| 2009 | Main vowel domain tone modeling with lexical and prosodic analysis for Mandarin ASRabstractThe tone is a distinctive discriminative feature in Mandarin Chinese. Often functional, yet seldom thorough are most large-scale Mandarin speech recognition systems in treating tone modeling. In particular, many lack the necessary sophistication to deal with the myriad variations arising from the combination of acoustic and lexical contexts. This paper reports an attempt to account for these variabilities and to bring richer tone modeling into the IBM Mandarin broadcast transcription system. In particular, we describe a system that combines the embedded approach and a novel explicit tone modeling technique characterized by a. robust tone tracking in the main-vowel domain, and b. context-dependent models with lexical and prosodic contexts. The proposed method is validated on a connected-digits set and subsequently evaluated on a large-vocabulary broadcast transcription task. It is shown that 14.8% and 5.4% relative reductions in character error rate are achieved respectively. Shilei Zhang, Qin Shi 0001, Stephen M. Chu, Yong Qin 0001 |
ICASSP | 1 |
| 2008 | Recent advances in the IBM GALE Mandarin transcription systemabstractThis paper describes the system and algorithmic developments in the automatic transcription of Mandarin broadcast speech made at IBM in the second year of the DARPA GALE program. Technical advances over our previous system include improved acoustic models using embedded tone modeling, and a new topic-adaptive language model (LM) rescoring technique based on dynamically generated LMs. We present results on three community-defined test sets designed to cover both the broadcast news and the broadcast conversation domain. It is shown that our new baseline system attains a 15.4% relative reduction in character error rate compared with our previous GALE evaluation system. And a further 13.6% improvement over the baseline is achieved with the two described techniques. Selina M. Chu, Hong-Kwang Jeff Kuo, Lidia Mangu, Yi Y. Liu 0002, Yong Qin 0001, Qin Shi 0001, Shilei Zhang, Hagai Aronowitz |
ICASSP | 7 |
| 2006 | Fast SVM training based on the choice of effective samples for audio classification
Shilei Zhang, Hongchen Jiang, Shuwu Zhang, Bo Xu 0002 |
INTERSPEECH | 1 |
| 2005 | Optimal model order selection based on regression tree in speaker identification
Shilei Zhang, Junmei Bai, Shuwu Zhang, Bo Xu 0002 |
INTERSPEECH | 1 |