VLDB 2026 Research / reviewers in the wild / expert
Jia Liu 0001
dblp:49/1245-1
· DBLP profile ↗
76ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 38 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive PruningabstractThe goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability of labeled data. To alleviate these issues, in this paper, a data-efficient and low-complexity ASC system is built with a new model architecture and better training strategies. Specifically, we firstly design a new low-complexity architecture named Rep-Mobile by integrating multi-convolution branches which can be reparameterized at inference. Compared to other models, it achieves better performance and less computational complexity. Then we apply the knowledge distillation strategy and provide a comparison of the data efficiency of the teacher model with different architectures. Finally, we propose a progressive pruning strategy, which involves pruning the model multiple times in small amounts, resulting in better performance compared to a single step pruning. Experiments are conducted on the TAU dataset. With Rep-Mobile and these training strategies, our proposed ASC system achieves the state-of-the-art (SOTA) results so far, while also winning the first place with a significant advantage over others in the DCASE2024 Challenge. Bing Han 0008, Wen Huang 0004, Zhengyang Chen, Anbai Jiang, Pingyi Fan, Cheng Lu 0007, Zhiqiang Lv, Jia Liu 0001, Weiqiang Zhang 0001, Yanmin Qian |
ICASSP | 8 |
| 2025 | Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known AttributesabstractAdapting pre-trained models has become the dominant approach for anomalous sound detection (ASD), where classifying the attributes of machine working status is commonly chosen as the deputy task for fine-tuning. However, attributes might be intractable to collect for some machines, causing the label to bear mixed granularity and thus deprecating the ASD performance. Therefore, we propose an adaptive proto-type learning scheme for fine-tuning pre-trained models, which adaptively scales coarse-grained labels to sub-centers so as to keep consistency with fine-grained labels. To deal with domain shift, we employ SMOTE to over-sample the prototypes of the target domain. The experiment on the DCASE 2024 ASD dataset demonstrates the efficacy of the proposed scheme, setting up a new milestone of 65.01% on both sets and outperforming the best system of the challenge. A detailed ablation study is also conducted to validate the effectiveness. Anbai Jiang, Xinhu Zheng, Yihong Qiu, Pingyi Fan, Cheng Lu 0007, Jia Liu 0001 |
ICASSP | 8 |
| 2024 | Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound DetectionabstractMachine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages self-supervised pre-trained models on large-scale speech data. Specifically, we assign different weights to the features from different layers of the pre-trained model and then use the working condition as the label for self-supervised classification fine-tuning. Moreover, we introduce a data augmentation method that simulates different operating states of the machine to enrich the dataset. Furthermore, we devise a transformer pooling method that fuses the features of different segments. Experiments on the DCASE2023 dataset show that our proposed method outperforms the commonly used reconstruction-based autoencoder and classification-based convolutional network by a large margin, demonstrating the effectiveness of large-scale pre-training for enhancing the generalization and robustness of machine anomalous sound detection. In Task2 of DCASE2023, we achieve 2nd place with these methods. Bing Han 0008, Zhiqiang Lv, Anbai Jiang, Wen Huang 0004, Zhengyang Chen, Yufeng Deng, Cheng Lu 0007, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001, Yanmin Qian |
ICASSP | 11 |
| 2024 | AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
Anbai Jiang, Bing Han 0008, Zhiqiang Lv, Yufeng Deng, Weiqiang Zhang 0001, Xie Chen 0001, Yanmin Qian, Jia Liu 0001, Pingyi Fan |
INTERSPEECH | 8 |
| 2024 | Improving Anomalous Sound Detection Via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio ModelsabstractAnomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme. Xinhu Zheng, Anbai Jiang, Bing Han 0008, Yanmin Qian, Pingyi Fan, Jia Liu 0001, Weiqiang Zhang 0001 |
SLT | 6 |
| 2023 | Decoupling Detectors for Scalable Anomaly Detection in AIoT Systems with Multiple MachinesabstractThe fast-developing Artificial Internet of Things (AIoT) technologies enable the consistent monitoring of multiple machines, by which machine failures can be detected in the early phases, and production efficiency and system management can be greatly promoted, bringing huge significance for anomaly detection. However, in most cases, anomalies are not provided for training, and the lack of direct supervision deprecates the anomaly detection performance. For the application viewpoint, the detector is required to generalize well on multiple machines, except for being computationally efficient. The computational cost is strictly limited, which is a great challenge for mobile and embedded devices. In face of these issues, we propose MobileAnoNet, which decouples an end-to-end detector into a front-end feature extractor and a back-end anomaly detector. The front-end extractor, consuming most computation, is unified for all machine types, while the back-end detector is specialized for each machine type, improving the detection capacity. The model is trained by handy labels of machine types and working conditions, in which multiple classification heads are attached behind the feature extractor during training. The performance of the model is evaluated on two DCASE datasets focusing on machine audio anomaly detection. It's shown that MobileAnoNet achieves a general improvement of 6.9% and 8.8% on two datasets, respectively. The ablation study demonstrates that multi-task learning promotes the general representation capacity. The source code is available at: www.github.com/hqj-les30/MobileAnoNet. Qijun Hou, Anbai Jiang, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001 |
GLOBECOM | 5 |
| 2023 | Unsupervised Anomaly Detection and Localization of Machine Audio: A Gan-Based ApproachabstractAutomatic detection of machine anomaly remains challenging for machine learning. We believe the capability of generative adversarial network (GAN) suits the need of machine audio anomaly detection, yet rarely has this been investigated by previous work. In this paper, we propose AEGAN-AD, a totally unsupervised approach in which the generator (also an autoencoder) is trained to reconstruct input spectrograms. It is pointed out that the denoising nature of reconstruction deprecates its capacity. Thus, the discriminator is redesigned to aid the generator during both training stage and detection stage. The performance of AEGAN-AD on the dataset of DCASE 2022 Challenge TASK 2 demonstrates the state-of-the-art result on five machine types. A novel anomaly localization method is also investigated. Source code available at: www.github.com/jianganbai/AEGAN-AD Anbai Jiang, Weiqiang Zhang 0001, Yufeng Deng, Pingyi Fan, Jia Liu 0001 |
ICASSP | 5 |
| 2020 | Staged Training Strategy and Multi-Activation for Audio Tagging with Noisy and Sparse Multi-Label Data
Kexin He, Yuhan Shen, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 4 |
| 2019 | Multi-objective Optimization Training of PLDA for Speaker VerificationabstractMost current state-of-the-art text-independent speaker verifi-cation systems take probabilistic linear discriminant analysis (PLDA) as their backend classifiers. The parameters of PL-DA are often estimated by maximizing the objective function, which focuses on increasing the value of log-likelihood function, but ignoring the distinction between speakers. In order to better distinguish speakers, we propose a multi-objective optimization training for PLDA. Experiment results show that the proposed method has more than 10% relative performance improvement in both EER and MinDCF on the NIST SRE14 i-vector challenge dataset, and about 20% relative performance improvement in EER on the MCE18 dataset. Liang He 0003, Xianhong Chen, Can Xu 0003, Jia Liu 0001 |
ICASSP | 4 |
| 2019 | Large Margin Softmax Loss for Speaker VerificationabstractIn neural network based speaker verification, speaker embedding is expected to be discriminative between speakers while the intra-speaker distance should remain small.A variety of loss functions have been proposed to achieve this goal.In this paper, we investigate the large margin softmax loss with different configurations in speaker verification.Ring loss and minimum hyperspherical energy criterion are introduced to further improve the performance.Results on VoxCeleb show that our best system outperforms the baseline approach by 15% in EER, and by 13%, 33% in minDCF08 and minDCF10, respectively. Yi Liu 0049, Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2019 | Distance-Dependent Metric LearningabstractIn most existing metric learning methods, the data pairs are equally treated without considering their diversity. In fact, for pairs with different distances, the main purpose of the metric acted on them has some differences. If the pairs have smaller distance, metric should focus more on expanding the negative pairs, which are more easily misjudged. While if the pairs have larger distance, metric should focus more on shrinking the positive pairs. However, most metric learning methods neglect these differences. In this letter, we propose a distance-dependent metric learning (D2ML) method. It partitions data pairs into different clusters according to the ℓ2distance between them. Each cluster is associated with a Mahalanobis metric that learns the pairs' distance. This not only allows us to make each metric more targeted and adapt to the data diversity flexibly, but also avoids the problem of computing the distance between points assigned to different clusters, which happens in some local metric learning methods. D2ML is further extended to D3ML to embrace the nonlinear capacity of neural network. Experiments on UCI datasets and speaker recognition i-vector machine learning challenge show that the proposed methods are superior to other metric learning methods. Xianhong Chen, Liang He 0003, Can Xu 0003, Jia Liu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Speaker Embedding Extraction with Phonetic InformationabstractSpeaker embeddings achieve promising results on many speaker verification tasks.Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings.In this paper, we introduce phonetic information to the speaker embedding extraction based on the x-vector architecture.Two methods using phonetic vectors and multi-task learning are proposed.On the Fisher dataset, our best system outperforms the original x-vector approach by 20% in EER, and by 15%, 15% in minDCF08 and minDCF10, respectively.Experiments conducted on NIST SRE10 further demonstrate the effectiveness of the proposed methods. Yi Liu 0049, Liang He 0003, Jia Liu 0001, Michael T. Johnson |
INTERSPEECH | 3 |
| 2018 | Local Pairwise Linear Discriminant Analysis for Speaker VerificationabstractLinear discriminant analysis-probabilistic linear discriminant analysis (LDA-PLDA) is a standard and effective backend in the field of speaker verification. The object of LDA is to perform dimensionality reduction while minimizing within-class covariance and maximizing between-class covariance. For a target class (or speaker), our task is to make a binary decision about whether a test utterance is from a specific target speaker. Generally, the nontarget test utterances that are close to the target speaker are easily misjudged. Inspired by this idea, we propose a local pairwise linear discriminant analysis (LPLDA) algorithm. This new method focuses on maximizing the local pairwise covariance, which represents the local structure between the target class samples and neighboring nontarget class samples, instead of the between-class covariance, which represents the global structure of the data. Experiments on the NIST SRE 2010, 2014, and 2016 database show that, the proposed LPLDA-PLDA backend has significant performance improvements over the LDA-PLDA backend. Liang He 0003, Xianhong Chen, Can Xu 0003, Jia Liu 0001, Michael T. Johnson |
IEEE Signal Process. Lett. | 4 |
| 2017 | Gated convolutional networks based hybrid acoustic models for low resource speech recognitionabstractIn acoustic modeling for large vocabulary speech recognition, recurrent neural networks (RNN) have shown great abilities to model temporal dependencies. However, the performance of RNN is not prominent in resource limited tasks, even worse than the traditional feedforward neural networks (FNN). Furthermore, training time for RNN is much more than that for FNN. In recent years, some novel models are provided. They use non-recurrent architectures to model long term dependencies. In these architectures, they show that using gate mechanism is an effective method to construct acoustic models. On the other hand, it has been proved that using convolution operation is a good method to learn acoustic features. We hope to take advantages of both these two methods. In this paper we present a gated convolutional approach to low resource speech recognition tasks. The gated convolutional networks use convolutional architectures to learn input features and a gate to control information. Experiments are conducted on the OpenKWS, a series of low resource keyword search evaluations. From the results, the gated convolutional networks relatively decrease the WER about 6% over the baseline LSTM models, 5% over the DNN models and 3% over the BLSTM models. In addition, the new models accelerate the learning speed by more than 1.8 and 3.2 times compared to that of the baseline LSTM and BLSTM models. Jian Kang 0006, Weiqiang Zhang 0001, Jia Liu 0001 |
ASRU | 3 |
| 2017 | Comparison of multiple features and modeling methods for text-dependent speaker verificationabstractText-dependent speaker verification is becoming popular in the speaker recognition society. However, the conventional i-vector framework which has been successful for speaker identification and other similar tasks works relatively poorly in this task. Researchers have proposed several new methods to improve performance, but it is still unclear that which model is the best choice, especially when the pass-phrases are prompted during enrollment and test. In this paper, we introduce four modeling methods and compare their performance on the newly published RedDots dataset. To further explore the influence of different frame alignments, Viterbi and forward-backward algorithms are both used in the HMM-based models. Several bottleneck features are also investigated. Our experiments show that, by explicitly modeling the lexical content, the HMM-based modeling achieves good results in the fixed-phrase condition. In the prompted-phrase condition, GMM-HMM and i-vector/HMM are not as successful. In both conditions, the forward-backward algorithm brings more benefits to the i-vector/HMM system. Additionally, we also find that even though bottleneck features perform well for text-independent speaker verification, they do not outperform MFCCs on the most challenging Imposter-Correct trials on RedDots. Yi Liu 0049, Liang He 0003, Zhuzi Chen, Jia Liu 0001, Michael T. Johnson |
ASRU | 5 |
| 2017 | An LSTM-CTC based verification system for proxy-word based OOV keyword searchabstractProxy-word based out of vocabulary (OOV) keyword search has been proven to be quite effective in keyword search. In proxy-word based OOV keyword search, each OOV keyword is assigned several proxies and detections of the proxies are regarded as detections of the OOV keywords. However, the confidence scores of these detections are still those of the proxies from lattices. To obtain a better confidence measure, we employ an LSTM-CTC verification method in this work and the confidence scores are regenerated. OOV keyword search results on the evalpart1 dataset of the OpenKWS16 Evaluation have shown consistent improvement and the maximum relative improvement can reach 21.06% for the MWTW metric. Zhiqiang Lv, Jian Kang 0006, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 4 |
| 2017 | Deep neural networks based speaker modeling at different levels of phonetic granularityabstractRecently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impact of different phonetic precision to speaker verification tasks, three levels of phonetic granularity are evaluated when doing frame alignments, which are tied-triphone state, monophone state and monophone. And the distribution of the features associated to a given phonetic unit is further modeled with multiple Gaussians rather than a single Gaussian. We also propose a fast and efficient way to generate phonetic units of different granularity by tying DNN's outputs according to the clustering results based on DNN derived senone embeddings. Experiments are carried out on the NIST SRE 2008 female tasks. Results show that using DNNs with less precise phonetic units and more Gaussians per phonetic unit for speaker modeling generalize better to different speaker verification tasks. Liang He 0003, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 5 |
| 2016 | THU-EE System Description for NIST LRE 2015
Liang He 0003, Yi Liu 0049, Weiwei Liu 0001, Cai Meng, Jia Liu 0001 |
INTERSPEECH | 7 |
| 2016 | Investigating Various Diarization Algorithms for Speaker in the Wild (SITW) Speaker Recognition Challenge
Yi Liu 0049, Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 4 |
| 2016 | A Novel Discriminative Score Calibration Method for Keyword Search
Zhiqiang Lv, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 4 |
| 2016 | Improving Deep Neural Networks Based Speaker Verification Using Unlabeled Data
Liang He 0003, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 5 |
| 2016 | Maxout neurons for deep convolutional and LSTM neural networks in speech recognition
Jia Liu 0001 |
Speech Commun. | 2 |
| 2015 | High-performance Swahili keyword search with very limited language pack: The THUEE system for the OpenKWS15 evaluationabstractThis paper presents the Swahili keyword search system developed by the THUEE team for the OpenKWS15 evaluation, which is conducted by NIST under the IARPA Babel program. There are several highlights in the development of the system, including automatic generation of the pronunciation lexicon, aggressive data augmentation, the multilingual bottleneck feature extractor trained from 6 languages, text selection from web data for language model training, semi-supervised training for acoustic models and language models, out-of-vocabulary keyword detection using morphemes and a rich diversity of the systems for combination. A wide variety of acoustic modeling techniques are explored and compared. Up to 12 different individual systems are used for combination. The system achieves the state-of-the-art performance in the required condition of the evaluation. Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Jia Liu 0001 |
ASRU | 7 |
| 2015 | Improved system fusion for keyword searchabstractIt has been demonstrated that system fusion can significantly improve the performance of keyword search. In this paper, we compare the performance of several widely-used arithmetic-based fusion methods using different normalization pipeline and try to find the best pipeline. A novel arithmetic-based fusion method is proposed in this work. The method supplies a more effective way to incorporate the number of systems which have non-zero scores for a detection. When tested on the development test dataset of the OpenKWS15 Evaluation, the proposed method achieves the highest maximum term-weighted value (MTWV) and actual term-weighted value (ATWV) among all other arithmetic-based fusion methods. Usually, discriminative fusion methods employing classifiers can outperform arithmetic-based fusion methods. A DNN-based fusion method is explored in this work. After word-burst information is added, the DNN-based fusion method outperforms all other methods. In addition, it is notable that our arithmetic-based method achieves the same MTWV as the DNN-based method. Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Weiqiang Zhang 0001, Jia Liu 0001 |
ASRU | 7 |
| 2015 | The THUEE system for the openKWS14 keyword search evaluationabstractThe OpenKWS14 keyword search evaluation is one of the most challenging and influential evaluations in the field of speech recognition. Its goal is to build a high-performance keyword search system for a minority language with limited training data in a short period of time. We present the system of the Department of Electronic Engineering, Tsinghua University (THUEE team) for the OpenKWS14 keyword search evaluation. The highlights of the system include the use of convolutional maxout neural networks for acoustic modeling and the use of neural network language models for one-pass lattice generation. The final system is a fusion of 8 sub-systems. The system has achieved an actual term weighted value (ATWV) of 0.5107 for the full language pack (FullLP) condition in the evaluation, ranking third among the participating teams. Zhiqiang Lv, Beili Song, Yongzhe Shi, Wei-lan Wu, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 8 |
| 2015 | Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognitionabstractExploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold, and set the outputs of inactive nodes to zero. Thus, only a few nonzero outputs are fed to the next layer. Using this method, the output vector of each hidden layer becomes very sparse, so that the computational cost of the feed-forward algorithm can be reduced by adopting sparse matrix operations. The proposed method is evaluated in both small and large vocabulary speech recognition tasks, and results demonstrate that we can reduce the nonzero outputs to fewer than 20% of the total number of hidden nodes, without sacrificing speech recognition performance. Jian Kang 0006, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 5 |
| 2015 | Simultaneous utilization of spectral magnitude and phase information to extract supervectors for speaker verification anti-spoofingabstractProtection from spoofing attacks is an essential component of speaker verification systems. This paper proposes a novel approach to detect such attacks by utilizing supervectors derived from spectral magnitude and phase information. Three countermeasures are chosen to represent these important information. To combine different countermeasures, score fusion and an antispoofing supervector (ASSV) are used. Experiments conducted on ASVspoof 2015 show that the combination of magnitude and phase information obtains relative 90% improvement in terms of the equal error rate (EER) compared to the best subsystem in the development set. The two systems can also be fused to further improve the performance. In addition to accuracy improvements, the new supervector framework is extensible and allows for a more flexible interface to the back-end classifier design. Yi Liu 0049, Liang He 0003, Jia Liu 0001, Michael T. Johnson |
INTERSPEECH | 4 |
| 2015 | Investigation of bottleneck features and multilingual deep neural networks for speaker verification
Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 4 |
| 2015 | Using word confusion networks for slot filling in spoken language understanding
Xiaohao Yang, Jia Liu 0001 |
INTERSPEECH | 2 |
| 2015 | Dialog state tracking using long short-term memory neural networks
Xiaohao Yang, Jia Liu 0001 |
INTERSPEECH | 2 |
| 2015 | Multi-resolution time frequency feature and complementary combination for short utterance speaker recognition
Weiqiang Zhang 0001, Jia Liu 0001 |
Multim. Tools Appl. | 3 |
| 2014 | Stochastic pooling maxout networks for low-resource speech recognitionabstractMaxout network is a powerful alternate to traditional sigmoid neural networks and is showing success in speech recognition. However, maxout network is prone to overfitting thus regularization methods such as dropout are often needed. In this paper, a stochastic pooling regularization method for max-out networks is proposed to control overfitting. In stochastic pooling, a distribution is produced for each pooling region by the softmax normalization of the piece values. The active piece is selected based on the distribution during training, and an effective probability weighting is conducted during testing. We apply the stochastic pooling maxout (SPM) networks within the DNN-HMM framework and evaluate its effectiveness under a low-resource speech recognition condition. On benchmark test sets, the SPM network yields 4.7-8.6% relative improvements over the baseline maxout network. Further evaluations show the superiority of stochastic pooling over dropout for low-resource speech recognition. Yongzhe Shi, Jia Liu 0001 |
ICASSP | 3 |
| 2014 | Improved phonotactic language recognition based on RNN feature reconstructionabstractNowadays phone recognition followed by support vector machine (PR-SVM) has been proposed in language recognition tasks and shown encouraging results. However, it still suffers from the problems such as the curse of dimensionality led by the increasing order of the N-gram feature supervector, the fast increasing number of possible parameters because of fast exact match of the phoneme history, etc. These problems hamper the capability of N-gram vector space model (VSM) of handling long-term contexts. In this paper, a recurrent neural networks (RNN) based feature reconstruction (FR) method is presented to compensate for the deficiency of the N-grams feature for phonotactic language recognition in this paper. Experiments are implemented on 2009 National Institute of Standards and Technology language recognition evaluation (NIST LRE) database. The results show that the proposed method gives 8.76%, 3.82%, 11.93% relative error rate reduction for 30s, 10s, 3s respectively comparing with the baseline system. Weiwei Liu 0001, Weiqiang Zhang 0001, Yongzhe Shi, An Ji, Jia Liu 0001 |
ICASSP | 6 |
| 2014 | Variance regularization of RNNLM for speech recognitionabstractRecurrent neural network language models (RNNLMs) have been proved superior to many other competitive language modeling techniques in terms of perplexity and word error rate. The remaining problem is the great computational complexity of RNNLMs in the output layer, resulting in long time for evaluation. Typically, a class-based RNNLM with the output layer factorized was proposed for speedup, which was still not fast enough for real-time systems. In this paper, a novel variance regularization algorithm is proposed for RNNLMs to address this problem. All the softmax-normalizing factors in the output layers are penalized to make them converge to one during the training phase, so that the output probability can be estimated efficiently via one dot-product of vectors in the output layer. The computational complexity of the output layer is reduced significantly from O(|V|H) to O(H). We further use this model for rescoring in an advanced CD-HMM-DNN system. Experimental results show that our proposed variance regularization algorithm works quite well, and the word prediction of the model is about 300 times faster than that of RNNLM without any obvious deteriorations in word error rate. Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 4 |
| 2014 | Phonotactic language recognition based on time-gap-weighted lattice kernels
Weiwei Liu 0001, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2014 | Spoken language recognition based on gap-weighted subsequence kernels
Weiqiang Zhang 0001, Weiwei Liu 0001, Yongzhe Shi, Jia Liu 0001 |
Speech Commun. | 5 |
| 2014 | Efficient One-Pass Decoding with NNLM for Speech RecognitionabstractNeural network language model (NNLM) has achieved very good results in the field of speech recognition, machine translation, etc. Direct decoding with NNLM is challenging for the overwhelmingly heavy burden in complexity. Most of the previous work focused on rescoring the N-best list and lattice with NNLM in the second pass. In this work, several techniques are explored to directly incorporate the NNLM into the decoder of speech recognition. A novel training algorithm based on variance regularization is proposed to approximate the softmax-normalizing factor as a constant for fast evaluation. Also, the evaluation of NNLM is further speeded up via our advanced storage. Moreover, a simple cache-based strategy is explored to avoid redundant computations during the decoding process. To the authors' knowledge, it is the first time to directly incorporate NNLM into decoding. We evaluate our proposed methods on an English-Switchboard phone-call speech-to-text task. Experimental results show that incorporating the NNLM into the decoder significantly reduces the word error rate (WER) by 1.5% and 1.4% absolutely on the Hub5'00-SWB and RT03S-FSH sets, respectively. Also, the decoding with NNLM is twice as fast as the baseline at the same word error rate. Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2013 | Deep maxout neural networks for speech recognitionabstractA recently introduced type of neural network called maxout has worked well in many domains. In this paper, we propose to apply maxout for acoustic models in speech recognition. The maxout neuron picks the maximum value within a group of linear pieces as its activation. This nonlinearity is a generalization to the rectified nonlinearity and has the ability to approximate any form of activation functions. We apply maxout networks to the Switchboard phone-call transcription task and evaluate the performances under both a 24-hour low-resource condition and a 300-hour core condition. Experimental results demonstrate that maxout networks converge faster, generalize better and are easier to optimize than rectified linear networks and sigmoid networks. Furthermore, experiments show that maxout networks reduce underfitting and are able to achieve good results without dropout training. Under both conditions, maxout networks yield relative improvements of 1.1-5.1% over rectified linear networks and 2.6-14.5% over sigmoid networks on benchmark test sets. Yongzhe Shi, Jia Liu 0001 |
ASRU | 3 |
| 2013 | Combination of data borrowing strategies for low-resource LVCSRabstractLarge vocabulary continuous speech recognition (LVCSR) is particularly difficult for low-resource languages, where only very limited manually transcribed data are available. However, it is often feasible to obtain large amount of untranscribed data of the low-resource target language or sufficient transcribed data of some non-target languages. Borrowing data from these additional sources to help LVCSR for low-resource language becomes an important research direction. This paper presents an integrated data borrowing framework in this scenario. Three data borrowing approaches were first investigated in detail, including feature, model and data corpus. They borrow data at different levels from additional sources, and all get substantial performance improvements. As these strategies work independently, the obtained gains are likely additive. The three strategies are then combined to form an integrated data borrowing framework. Experiments showed that with the integrated data borrowing framework, significant improvement of more than 10% absolute WER reduction over a conventional baseline was obtained. In particular, the gain under the extreme limited low-resource scenario is 16%. Yanmin Qian, Kai Yu 0004, Jia Liu 0001 |
ASRU | 3 |
| 2013 | I-matrix for text-independent speaker recognitionabstractThis paper proposes an i-matrix for text-independent speaker recognition. The framework of the proposed i-matrix is similar to an i-vector. However, the presented method takes short-time cepstral feature matrices as inputs to explore both cepstral feature distribution and temporal information for the recognition task in the phase of statistical modeling. In the i-matrix, the variability of an utterance is constrained by two subspaces U and V, which are estimated by an iterative method on a large database. When U and V are well built, each utterance is represented by an i-matrix. Decision function is a cosine kernel. Experiments were carried out on the tel-tel-English condition of NIST SRE 2008 core task. Compared with an i-vector-LDA, the average EER and MDCF of an i-matrix-LDA showed a relative decrease of 4.82% and 5.12% respectively. Liang He 0003, Jia Liu 0001 |
ICASSP | 2 |
| 2013 | Temporal kernel neural network language modelabstractUsing neural networks to estimate the probabilities of word sequences has shown significant promise for statistical language modeling. Typical modeling methods include multi-layer neural networks, log-bilinear networks and recurrent neural networks, etc. In this paper, we propose the temporal kernel neural network language model, a variant of models mentioned above. This model explicitly captures long-term dependencies of words with exponential kernel, where the memory of history is decayed exponentially. Additionally, several sentences with variable lengths as a mini-batch are efficiently implemented for speeding up. Experimental results show that the proposed model is very competitive to the recurrent neural network language model and obtains the lower perplexity of 111.6 (more than 10% reduction) than the state-of-the-art results reported in the standard Penn Treebank Corpus. We further apply this model to Wall Street Journal speech recognition task, and observe significant improvements in word error rate. Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP | 4 |
| 2013 | Simplified domain transfer multiple kernel learning for language recognitionabstractDistribution mismatch between training and test data can greatly deteriorate the performance of language recognition. Some effective methods for compensation have been proposed, such as nuisance attribute projection (NAP). In real-world applications, there are often sufficient training samples from a different domain and only a limited number of labeled training samples from target domain, performance of a system will be degraded and needs to be further improved. In this paper, we introduce transfer learning to solve this problem. We propose a novel transfer learning algorithm referred to as simplified domain transfer multiple kernel learning (SDTMKL). Our aim is to discover a good representation of feature space that minimizes the distribution mismatch between samples from the source and target domains. Robust models can be learned in this suitable feature space. Results on a NIST language recognition task show that the SDTMKL method is quite effective and can further improve system performance when combined with NAP. Jia Liu 0001, Shanhong Xia |
ICASSP | 2 |
| 2013 | Audiovisual synthesis of exaggerated speech for corrective feedback in computer-assisted pronunciation trainingabstractIn second language learning, unawareness of the differences between correct and incorrect pronunciations is one of the largest obstacles for mispronunciation correction. In order to make the feedback more discriminatively perceptible, this paper presents a novel method for corrective feedback generation, namely, exaggerated feedback, for language learning. To produce exaggeration effect, the neutral audio and visual speech are both exaggerated and then re-synthesized synchronously based on the audiovisual synthesis technology. The audio speech exaggeration is realized by adjusting the acoustic features related to duration, pitch and energy of the speech according to different phonemes conditions. The visual speech exaggeration is realized by increasing the range of articulatory movement and slowing down the movement around the key actions. The results show that our methods can effectively generate bimodal exaggeration effect for feedback provision and make them more distinctive to be perceived. Junhong Zhao, Wai-Kim Leung, Helen M. Meng, Jia Liu 0001, Shanhong Xia |
ICASSP | 5 |
| 2013 | Parallel absolute-relative feature based phonotactic language recognition
Weiwei Liu 0001, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 4 |
| 2013 | MLP-HMM two-stage unsupervised training for low-resource languages on conversational telephone speech recognition
Yanmin Qian, Jia Liu 0001 |
INTERSPEECH | 2 |
| 2013 | THU-EE system fusion for the NIST 2012 speaker recognition evaluation
Weiqiang Zhang 0001, Weiwei Liu 0001, Jia Liu 0001 |
INTERSPEECH | 4 |
| 2013 | Exploiting articulatory features for pitch accent detectionabstractArticulatory features describe how articulators are involved in making sounds. Speakers often use a more exaggerated way to pronounce accented phonemes, so articulatory features can be helpful in pitch accent detection. Instead of using the actual articulatory features obtained by direct measurement of articulators, we use the posterior probabilities produced by multi-layer perceptrons (MLPs) as articulatory features. The inputs of MLPs are frame-level acoustic features pre-processed using the split temporal context-2 (STC-2) approach. The outputs are the posterior probabilities of a set of articulatory attributes. These posterior probabilities are averaged piecewise within the range of syllables and eventually act as syllable-level articulatory features. This work is the first to introduce articulatory features into pitch accent detection. Using the articulatory features extracted in this way, together with other traditional acoustic features, can improve the accuracy of pitch accent detection by about 2%. Junhong Zhao, Weiqiang Zhang 0001, Jia Liu 0001, Shanhong Xia |
J. Zhejiang Univ. Sci. C | 5 |
| 2012 | Cross-Lingual and Ensemble MLPs Strategies for Low-Resource Speech Recognition
Yanmin Qian, Jia Liu 0001 |
INTERSPEECH | 2 |
| 2012 | Articulatory Feature based Multilingual MLPs for Low-Resource Speech Recognition
Yanmin Qian, Jia Liu 0001 |
INTERSPEECH | 2 |
| 2011 | Strategies for using MLP based features with limited target-language training dataabstractRecently there has been some interest in the question of how to build LVCSR systems when there is only a limited amount of acoustic training data in the target language, but possibly more plentiful data in other languages. In this paper we investigate approaches using MLP based features. We experiment with two approaches: One is based on Automatic Speech Attribute Transcription (ASAT), in which we train classifiers to learn articulatory features. The other approach uses only the target-language data and relies on combination of multiple MLPs trained on different subsets. After system combination we get large improvements of more than 10% relative versus a conventional baseline. These feature-level approaches may also be combined with other, model-level methods for the multilingual or low-resource scenario. Yanmin Qian, Daniel Povey, Jia Liu 0001 |
ASRU | 4 |
| 2011 | State-Level Data Borrowing for Low-Resource Speech Recognition Based on Subspace GMMsabstractLarge vocabulary continuous speech recognition is always a difficult task, and it is particularly so for low-resource languages. The scenario we focus on here is having only 1 hour of acoustic training data in the “target” language. This paper presents work on a data borrowing strategy combined with the recently proposed Subspace Gaussian Mixture Model (SGMM). We developed data borrowing strategies based on two approaches: one based on minimizing K-L Divergence, and one that also takes into account state occupation counts. We demonstrate improvements versus the baseline SGMM setup, which itself is better than a conventional HMM-GMM system. The SGMMs are more robustly estimated by borrowing data from the non-target language at the acousticstate level. Although we tested the approach for SGMMs, we expect the general idea of borrowing data from a non-target language to be applicable for conventional GMMs as well. Index Terms: speech recognition, low-resource language, subspace gaussian mixture model Yanmin Qian, Daniel Povey, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2011 | Combining Lattice-Based Language Dependent and Independent Approaches for Out-of-Language Detection in LVCSR
Yuxiang Shan, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2011 | Robust Audio Fingerprinting Based on Local Spectral Luminance Maxima Scheme
Yongzhe Shi, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2011 | Robust speaker recognition in cross-channel condition based on Gaussian mixture model
Yuxiang Shan, Jia Liu 0001 |
Multim. Tools Appl. | 2 |
| 2011 | Time-Frequency Cepstral Features and Heteroscedastic Linear Discriminant Analysis for Language RecognitionabstractThe shifted delta cepstrum (SDC) is a widely used feature extraction for language recognition (LRE). With a high context width due to incorporation of multiple frames, SDC outperforms traditional delta and acceleration feature vectors. However, it also introduces correlation into the concatenated feature vector, which increases redundancy and may degrade the performance of backend classifiers. In this paper, we first propose a time-frequency cepstral (TFC) feature vector, which is obtained by performing a temporal discrete cosine transform (DCT) on the cepstrum matrix and selecting the transformed elements in a zigzag scan order. Beyond this, we increase discriminability through a heteroscedastic linear discriminant analysis (HLDA) on the full cepstrum matrix. By utilizing block diagonal matrix constraints, the large HLDA problem is then reduced to several smaller HLDA problems, creating a block diagonal HLDA (BDHLDA) algorithm which has much lower computational complexity. The BDHLDA method is finally extended to the GMM domain, using the simpler TFC features during re-estimation to provide significantly improved computation speed. Experiments on NIST 2003 and 2007 LRE evaluation corpora show that TFC is more effective than SDC, and that the GMM-based BDHLDA results in lower equal error rate (EER) and minimum average cost (Cavg) than either TFC or SDC approaches. Weiqiang Zhang 0001, Liang He 0003, Jia Liu 0001, Michael T. Johnson |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Automatic player labeling, tracking and field registration and trajectory mapping in broadcast soccer videoabstractIn this article, we present a method to perform automatic player trajectories mapping based on player detection, unsupervised labeling, efficient multi-object tracking, and playfield registration in broadcast soccer videos. Player detector determines the players' positions and scales by combining the ability of dominant color based background subtraction and a boosting detector with Haar features. We first learn the dominant color with accumulate color histogram at the beginning of processing, then use the player detector to collect hundreds of player samples, and learn player appearance codebook by unsupervised clustering. In a soccer game, a player can be labeled as one of four categories: two teams, referee or outlier. The learning capability enables the method to be generalized well to different videos without any manual initialization. With the dominant color and player appearance model, we can locate and label each player. After that, we perform multi-object tracking by using Markov Chain Monte Carlo (MCMC) data association to generate player trajectories. Some data driven dynamics are proposed to improve the Markov chain's efficiency, such as label consistency, motion consistency, and track length, etc. Finally, we extract key-points and find the mapping from an image plane to the standard field model, and then map players' position and trajectories to the field. A large quantity of experimental results on FIFA World Cup 2006 videos demonstrate that this method can reach high detection and labeling precision, reliably tracking in scenes of player occlusion, moderate camera motion and pose variation, and yield promising field registration results. Xiaofeng Tong, Jia Liu 0001, Tao Wang 0003, Yimin Zhang 0002 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2010 | Phone modeling and combining discriminative training for mandarinenglish bilingual speech recognitionabstractAutomatic multilingual speech recognition is always a difficult task. This paper presents recent work on the development of a Mandarin-English bilingual speech recognition system. A unified single set of bilingual acoustic models based on a novel State-Time-Alignment (STA) method is proposed to balance the performance and the complexity of the bilingual speech recognition system, and a comparison with the acoustic-likelihood method is presented. Discriminative training approaches such as MPE and fMPE have been shown to improve monolingual recognition performance, but have not yet been applied to bilingual speech recognition. This paper investigates the use of discriminative training methods on bilingual speech recognition, including MPE and fMPE. Experimental results show that the STA phone clustering method outperforms other existing phone clustering methods, and both forms of discriminative training reduce the word error rate of the multilingual system. Yanmin Qian, Jia Liu 0001 |
ICASSP | 2 |
| 2010 | A CMLLR supervector kernel for SVM language recognitionabstractThis paper explores the use of constrained maximum likelihood linear regression (CMLLR) transforms as features for language recognition. Modeling is carried out through support vector machine (SVM). This work proposes a novel CMLLR supervector kernel. Results on the NIST LRE09 task show that feature-domain CMLLR transforms contain more language dependent information than model-domain MLLRs, and the proposed CMLLR supervector kernel outperforms some other ones. We also compare our CMLLR-SVM system with some state-of-the-art systems, and combine them for a further improvement. Jia Liu 0001 |
ICASSP | 2 |
| 2010 | Combining Chinese spoken term detection systems via side-information conditioned linear logistic regression
Sha Meng, Weiqiang Zhang 0001, Jia Liu 0001 |
INTERSPEECH | 3 |
| 2010 | A fast query by humming system based on notesabstractQuery by humming (QBH), a content-based retrieval method, is an efficient way to search the song from a large database. The frame-based systems can achieve a good performance, but it is time-consuming. In this paper, we proposed an efficient note-based system, which is mainly comprised of noted-based linear scaling (NLS) and noted-based recursive align (NRA). The system after post-processing can achieve 96.1 % in Top5 and 0.211s in time. Index Terms: query by humming (QBH), musical information retrieval (MIR), note-based linear scaling (NLS), note-based recursive align (NRA) 1. Jia Liu 0001, Weiqiang Zhang 0001 |
INTERSPEECH | 2 |
| 2010 | Variant time-frequency cepstral features for speaker recognition
Weiqiang Zhang 0001, Liang He 0003, Jia Liu 0001 |
INTERSPEECH | 4 |
| 2009 | Automatic player detection, labeling and tracking in broadcast soccer video
Jia Liu 0001, Xiaofeng Tong, Wenlong Li 0003, Tao Wang 0003, Yimin Zhang 0002 |
Pattern Recognit. Lett. | 1 |
| 2008 | Fusing multiple systems into a compact lattice index for chinese spoken term detectionabstractWe examine the task of spoken term detection in Chinese spontaneous speech with a lattice-based approach. We first compare lattices generated with different units: word, character, tonal and toneless syllables, and also lattices converted from one unit to another unit. Then we combine lattices from multiple systems into a single lattice. By fully exploiting the redundant information in the combined lattice with a time-based node/arc merging, we achieve the result of a compact lattice index with the accuracy improved to 79.2% from 73.9% using the best subsystem. Sha Meng, Jia Liu 0001, Frank Seide |
ICASSP | 3 |
| 2008 | Addressing the out-of-vocabulary problem for large-scale Chinese spoken term detection
Sha Meng, Jian Shao 0001, Roger Peng Yu, Jia Liu 0001, Frank Seide |
INTERSPEECH | 4 |
| 2008 | An Equalized Heteroscedastic Linear Discriminant Analysis AlgorithmabstractHeteroscedastic linear discriminant analysis (HLDA) is a widely used feature extraction algorithm. This method, however, suffers from unbalanced training data in some cases. In this letter, we equalize the objective function and statistics of HLDA and present an equalized HLDA algorithm, which balances the training data according to the class prior probability. Simulations as well as experimental results for the task of language identification are used to demonstrate the effectiveness of the proposed method. Weiqiang Zhang 0001, Jia Liu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2007 | A study of lattice-based spoken term detection for Chinese spontaneous speechabstractWe examine the task of spoken term detection in Chinese spontaneous speech with a lattice-based approach. We compare lattices generated with different units: word, character, tonal syllable and toneless syllable, and also look into methods of converting lattices from one unit to another one. We find the best system is with toneless-syllable lattices converted from word lattices. Further improvement is achieved by lattice post-processing and system combination. Our best system has an accuracy of 80.2% on a keyword spotting task. Sha Meng, Frank Seide, Jia Liu 0001 |
ASRU | 4 |
| 2007 | Automatic Player Detection, Labeling and Tracking in Broadcast Soccer VideoabstractAutomatic player detection, labeling and tracking in broadcast soccer video are significant while quite challenging tasks. In this paper, we present a solution to perform automatic multiple player detection, unsupervised labeling and efficient tracking. Players ’ position and scale are determined by a boosting based detector. Players ’ appearance models are unsupervised learned from hundreds of samples automatically collected by detection. Thereafter, these models can be utilized for player labeling (Team A, Team B and Referee). Player tracking is achieved by Markov Chain Monte Carlo (MCMC) data association. Some data driven dynamics are proposed to improve the Markov chain’s efficiency. The testing results on FIFA World Cup 2006 video demonstrate that our method can reach high detection and labeling precision, and reliably tracking in cases of scenes such as multiple player occlusion, moderate camera motion and pose variation. 1 Jia Liu 0001, Xiaofeng Tong, Wenlong Li 0003, Tao Wang 0003, Yimin Zhang 0002, Bo Yang 0008, Lifeng Sun, Shiqiang Yang |
BMVC | 1 |
| 2007 | Two-Stage Method for Specific Audio RetrievalabstractSpecific audio retrieval, also referred as similarity-based audio retrieval, means to detect and locate a given query audio segment in a long stored audio signal. In this paper, we proposed a two-stage method for specific audio retrieval. In the first stage, the histogram pruning algorithm is used for coarse detection. In the second stage, the partial distance technique is used for fine verification and localization. Experimental results show that the two-stage coarse-to-fine method offers fast search speed and improves the robustness to additive noise and compression encoding. Weiqiang Zhang 0001, Jia Liu 0001 |
ICASSP (4) | 2 |
| 2004 | Embedded speech recognition system on 8-bit MCU coreabstractThis is a small vocabulary, speaker independent, discrete word speech recognition system based on the system on chip (SOC) philosophy. It is implemented on an 8-bit MCU (micro control unit). The system adopts linear predictive cepstral coefficient (LPCC) related features followed by a vector quantization (VQ) step as the front end, and a hidden Markov model (HMM) as the speech model. Confidence measures based on likelihood scores (LLS) are given for rejection of out-of-vocabulary (OOV) words. The recognition rate is improved with corrective training, and robustness is acquired by integrating the confidence measure into the system. The recognition accuracy is nearly 97% with a vocabulary up to 30 phrases under normal conditions. A simple speech codec is also implemented for all speech I/O purposes. Jia Liu 0001, Runsheng Liu |
ICASSP (5) | 3 |
| 2003 | A novel efficient decoding algorithm for CDHMM-based speech recognizer on chipabstractThe efficiency of decoding algorithm is the main limitation of realizing a CDHMM-based large-vocabulary name dialing application on consumer electronic products. To solve this problem, a novel efficient decoding algorithm, TPVD (two-pass Viterbi decoding), is represented. By coarse matching in the first pass, several high-confidence candidates are selected from hundreds. Then, fine matching in the second pass runs on base of those high-confidence candidates. Compared with traditional Viterbi beam search, the hardware resource requirement is significantly reduced. Meanwhile, the recognition accuracy and speed are satisfactory. This conclusion has been proven in an embedded Mandarin 600-name dialing system. Jia Liu 0001, Runsheng Liu |
ICASSP (2) | 3 |
| 2003 | Voice conversion with smoothed GMM and MAP adaptation
Min Chu, Eric Chang, Jia Liu 0001, Runsheng Liu |
INTERSPEECH | 4 |
| 2003 | Towards Robustness to Speech Rate in Mandarin All-Syllable Recognition
Jia Liu 0001, Runsheng Liu |
J. Comput. Sci. Technol. | 3 |
| 2002 | A Rejection Model Based on Multi-Layer Perceptrons for Mandarin Digit Recognition
Zhong Lin, Jia Liu 0001, Runsheng Liu |
J. Comput. Sci. Technol. | 2 |
| 2000 | Rejection based on a posteriori probability estimated by MLP with application for Mandarin voice dialer on ASICabstractHigh performance Mandarin voice dialer is much more difficult than its English counterpart to achieve, especially on inexpensive hardware as ASIC. One way to improve its performance is to incorporate rejecters into the system. In our study, an MLP based postprocessor, an a posteriori probability estimator, is applied after HMM Viterbi recognition. Poor utterances, which are recognized by HMMs but have low a posteriori probability, are then rejected. Rejecting 4.9% of all the testing utterances, the MLP rejector boosts the HMM-based system's single digit accuracy from 97.1% to 99.6% for the Mandarin voice dialer, a ten-syllable speaker independent task. The performance is better than those of rejection based on linear discrimination, anti-digit models or likelihood ratio. Lin Zhong 0001, Jia Liu 0001, Runsheng Liu |
ICASSP | 2 |
| 2000 | Confidence measure based unsupervised speaker adaptation
Husheng Li, Jia Liu 0001, Runsheng Liu |
INTERSPEECH | 2 |
| 1998 | A novel robust speech recognition algorithm based on multi-models and integrated decision method
Shengxi Pan, Jia Liu 0001, Jintao Jiang, Zuoying Wang, Dajin Lu |
ICSLP | 2 |