Yunxin Zhao

dblp:98/1521 · DBLP profile ↗
← Back
121ranked-venue papers
26as first author
8since 2021 · last 2026
0000-0001-5511-3692ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 89 · 18 first-author · 6 since 2021Artificial intelligence and machine learning · 61 · 12 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Improve NNLMs by text generation from pre-trained language models
Minguang Song, Yunxin Zhao
Comput. Speech Lang.2
2023 Joint Estimation of DOA and Distance in Noisy Reverberant Conditions
abstract
Sound Source localization (SSL) using microphone arrays is an active research topic with many applications, but noise and reverberation make the direction-of-arrival (DOA) and distance estimation a challenging problem. In this work, we propose a novel method to jointly estimate the DOA and distance in noisy and reverberant environments. Our method exploits the linear phase structure across frequencies in a steering vector (SV). We convert the joint estimation issue into an optimization problem, which can be solved by Newton’s method augmented by a gradient ascent method. Our method does not depend on certain microphone array geometry, and it can also be extended to estimate the elevation angle. We conducted experimental evaluations in simulated noisy and reverberant acoustic conditions, which verified the superiority of our proposed method to several established methods in estimation accuracy and computation efficiency.
Suliang Bu, Tuo Zhao, Yunxin Zhao
ICASSP3
2022 Enhance Rnnlms with Hierarchical Multi-Task Learning for ASR
abstract
It is known that neural language models (NLMs) can implicitly learn certain linguistic information from text. While generally NLMs only use word feature input, the success of factored NLMs has indicated a benefit of using additional linguistic feature inputs for language modeling. On the other hand, multi-task learning (MTL) has shown positive effects on the generalization performance of various natural language processing (NLP) tasks, including language modeling. However, how to best share information among related tasks in MTL remains to be addressed. In this current work, we propose a hierarchical multi-task learning (HMTL) approach to incorporate linguistic knowledge into recurrent neural network language models (RNNLM), instead of using linguistic features as word factors. Specifically, we consider the auxiliary tasks of chunking, part of speech tagging, and named entity recognition, and supervise the learning of these auxiliary tasks in a hierarchical way. Our proposed method has the potential of helping language models learn knowledge of linguistic hierarchy from the auxiliary tasks, and improve the performance of RNNLMs on automatic speech recognition (ASR). We have evaluated our proposed HMTL method on WSJ and AMI speech recognition tasks. Our experiment results demonstrate the effectiveness of the proposed approach.
Minguang Song, Yunxin Zhao
ICASSP2
2022 Steering vector correction in MVDR beamformer for speech enhancement
Suliang Bu, Yunxin Zhao, Tuo Zhao
INTERSPEECH2
2022 TDOA Estimation of Speech Source in Noisy Reverberant Environments
abstract
Sound source localization is important in many applications, but noise and reverberation make the time difference of arrival (TDOA) estimation a challenging problem. In this work, we propose two novel methods to effectively estimate TDOA in noisy and reverberant environments. Our methods exploit the linear phase structure across frequencies in a steering vector (SV). To reduce potential noise and mathematical issues, we utilize absolute phases of SVs. We convert TDOA estimation into an optimization problem that is solvable by Newton's method. We conducted experimental evaluations in simulated acoustic conditions. In conditions with moderate-to-high input SNR and low reverberation, our fast-search method is superior in TDOA accuracy and is very efficient in computation. In conditions with low input SNR and high reverberation, our detailed-search method shows strong robustness and maintains advantageous performance in TDOA estimation.
Suliang Bu, Tuo Zhao, Yunxin Zhao
SLT3
2022 Modeling Speech Structure to Improve T-F Masks for Speech Enhancement and Recognition
abstract
Time-frequency (TF) masks are widely used in speech enhancement (SE). However, accurately estimating TF masks from noisy speech remains a challenge to both statistical or neural network (NN) approaches. Statistical model based mask estimation usually depends on a good parameter initialization, while NN-based method relies on setting proper and stable learning targets. To address these issues, we propose to extract TF speech structure from clean speech and partition noisy speech spectrogram into mutually exclusive regions. We investigate modeling clean speech by utterance-specific narrowband complex Gaussian mixture models to derive the regions, and using the region targets to supervise the training of UNet++, a high-performance NN, for predicting regions from noisy speech. For multichannel SE, we consider two scenarios of using speech regions: 1) integrating the regions with TF masks by constraining the mask values or the model parameter updates, and 2) using the predicted regions in place of TF masks. For single-channel SE, we consider using the region targets to improve TF mask targets. Furthermore, we propose to use UNet++ for TF mask estimation. Our experiment results on speech recognition (CHiME-3) and SE (CHiME-3 and LibriSpeech) have demonstrated the effectiveness of our proposed approach of modeling speech region structure to improve TF masks for speech recognition and enhancement.
Suliang Bu, Yunxin Zhao, Tuo Zhao
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Learning Speech Structure to Improve Time-Frequency Masks
Suliang Bu, Yunxin Zhao
Interspeech2
2021 Word Similarity Based Label Smoothing in Rnnlm Training for ASR
abstract
Label smoothing has been shown as an effective regularization approach for deep neural networks. Recently, a context-sensitive label smoothing approach was proposed for training RNNLMs that improved word error rates on speech recognition tasks. Despite the performance gains, its plausible candidate words for label smoothing were confined to n-grams observed in training data. To investigate the potential of label smoothing in model training with insufficient data, in this current work, we propose to utilize the similarity between word embeddings to build a candidate word set for each target word, where by doing so, plausible words outside the n-grams in training data may be found and introduced into candidate word sets for label smoothing. Moreover, we propose to combine the smoothing labels from the n-gram based and the word similarity based methods to improve the generalization capability of RNNLMs. Our proposed approach to RNNLM training has been evaluated for n-best list rescoring on speech recognition tasks of WSJ and AMI, with improved experimental results on word error rates confirming its effectiveness.
Minguang Song, Yunxin Zhao
SLT2
2020 Learning Recurrent Neural Network Language Models With Context-Sensitive Label Smoothing for Automatic Speech Recognition
abstract
Recurrent neural network language models (RNNLMs) have become very successful in many natural language processing tasks. However, RNNLMs trained with a cross entropy loss function and hard output targets are prone to overfitting, which weakens the language models’ generalization power. In the current work, we investigate a new strategy of label smoothing in place of hard output targets to regularize RNNLM training. We propose an approach of context-sensitive candidate label smoothing that has two advantages. First, it not only helps prevent overfitted model but also distinguishes plausible words from implausible ones. Second, it helps alleviate the problems of data sparsity and unbalanced word occurrence in training data. We evaluate our proposed candidate label smoothing method on RNNLM training for two speech recognition tasks, and demonstrate its positive impacts on test set word error rate and perplexity.
Minguang Song, Yunxin Zhao
ICASSP2
2020 Voice Conversion for Persons with Amyotrophic Lateral Sclerosis
abstract
Amyotrophic lateral sclerosis (ALS) results in progressive paralysis of voluntary muscles throughout the body. As speech deteriorates, individuals rely on pre-programmed messages available on commercial speech generating devices to communicate using one of the generic electronic voices on the device. To replace these generic voices and restore vocal identity, our aim is to develop personalized voices for people with ALS via the approach of voice conversion. The task is challenging because very few people have large quantities of their premorbid healthy speech recorded. Therefore, we have to rely on small quantities of dysarthric speech concomitant with an individual's disease stage. Further, progressive fatigue prohibits acquisition of large speech datasets and individuals display a range of dysarthria severities resulting from breathing, voice, articulation, resonance, and prosody disturbances. As the first step to address these problems, we use healthy source speakers and propose the approach of combining a structured sparse spectral transform with multiple linear regression-based frequency warping prediction for spectral conversion, and interpolating the transformed spectral frames for speech rate modification. Our experimental data included four healthy source speakers from the ARCTIC dataset, and four target ALS speakers with mild to severe dysarthria, forming 16 speaker pairs. Subjective listening evaluations showed that on average, (i) the proposed approach improved speech intelligibility by about 80% over the target speakers' speech, (ii) the converted voice was 3 times more similar to the target speakers' speech than to the source speakers' speech, and (iii) the converted speech quality was close to the MOS scale "good" relative to the source speakers' speech being "excellent."
Yunxin Zhao, Mili Kuruvilla-Dugdale, Minguang Song
IEEE J. Biomed. Health Informatics1
2019 A Novel Method to Correct Steering Vectors in MVDR Beamformer for Noise Robust ASR
Suliang Bu, Yunxin Zhao, Mei-Yuh Hwang
INTERSPEECH2
2018 Slim Embedding Layers for Recurrent Neural Language Models
abstract
Recurrent neural language models are the state-of-the-art models for language modeling. When the vocabulary size is large, the space taken to store the model parameters becomes the bottleneck for the use of recurrent neural language models. In this paper, we introduce a simple space compression method that randomly shares the structured parameters at both the input and output embedding layers of the recurrent neural language models to significantly reduce the size of model parameters, but still compactly represent the original input and output embedding layers. The method is easy to implement and tune. Experiments on several data sets showthat the new method can get similar perplexity and BLEU score results whileonly using a very tiny fraction of parameters.
Raymond Kulhanek, Yunxin Zhao
AAAI4
2018 A Probability Weighted Beamformer for Noise Robust ASR
Suliang Bu, Yunxin Zhao, Mei-Yuh Hwang, Sining Sun
INTERSPEECH2
2018 Multi-Objective Multi-Task Learning on RNNLM for Speech Recognition
abstract
The cross entropy (CE) loss function is commonly adopted for neural network language model (NNLM) training. Although this criterion is largely successful, as evidenced by the quick advance of NNLM, minimizing CE only maximizes likelihood of training data. When training data is insufficient, the generalization power of the resulting LM is limited on test data. In this paper, we propose to integrate a pairwise ranking (PR) loss with the CE loss for multi-objective training on recurrent neural network language model (RNNLM). The PR loss emphasizes discrimination between target and non-target words and also reserves probabilities for low-frequency correct words, which complements the distribution learning role of the CE loss. Combining the two losses may therefore help improve the performance of RNNLM. In addition, we incorporate multi-task learning (MTL) into the proposed multi-objective learning to regularize the primary task of RNNLM by an auxiliary task of part-of-speech (POS) tagging. The proposed approach to RNNLM learning has been evaluated on two speech recognition tasks of WSJ and AMI with encouraging results achieved on word error rate reductions.
Minguang Song, Yunxin Zhao
SLT2
2018 Structured Sparse Spectral Transforms and Structural Measures for Voice Conversion
abstract
We investigate a structured sparse spectral transform method for voice conversion (VC) to perform frequency warping and spectral shaping simultaneously on high-dimensional (D) STRAIGHT spectra. Learning a large transform matrix for high-D data often results in an overfit matrix with low sparsity, which leads to muffled speech in VC. We address this problem by using the frequency-warping characteristic of a source-target speaker pair to define a region of support (ROS) in a transform matrix, and further optimize it by nonnegative matrix factorization (NMF) to obtain structured sparse transform. We also investigate structural measures of spectral and temporal covariance and variance at different scales for assessing VC speech quality. Our experiments on ARCTIC dataset of 12 speaker pairs show that embedding the ROS in spectral transforms offers flexibility in tradeoffs between spectral distortion and structure preservation, and the structural measures provide quantitatively reasonable results on converted speech. Our subjective listening tests show that the proposed VC method achieves a mean opinion score of "very good" relative to natural speech, and in comparison with three other VC methods, it is the most preferred one in naturalness and in voice similarity to target speakers.
Yunxin Zhao, Mili Kuruvilla-Dugdale, Minguang Song
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Exploiting different word clusterings for class-based RNN language modeling in speech recognition
abstract
We propose to exploit the potential of multiple word clusterings in class-based recurrent neural network (RNN) language models for ensemble RNN language modeling. By varying the clustering criteria and the space of word embedding, different word clusterings are obtained to define different word/class factorizations. For each such word/class factorization, several base RNNLMs are learned, and the word prediction probabilities of the base RNNLMs are then combined to form an ensemble prediction. We use a greedy backward model selection procedure to select a subset of models and combine these models for word prediction. The proposed ensemble language modeling method has been evaluated on Penn Treebank test set as well as Wall Street Journal (WSJ) Eval 92 and 93 test sets, where it improved test set perplexity and word error rate over the state-of-the-art single RNNLMs as well as multiple RNNLMs produced by varying RNN learning conditions.
Minguang Song, Yunxin Zhao
ICASSP2
2015 A novel static parameter calculation method for model compensation
abstract
Vector Taylor Series (VTS) based model compensation approach has been successfully applied to various robust speech recognition tasks. In this paper, we propose a novel method of variable transformation to calculate the static statistics. In addition, we provide a detailed explanation of VTS and random variable transformations adopted in some recent papers. Experiments on Aurora 4 showed that the proposed approach obtained 22.8% relative WER reduction over the traditional first-order VTS methods.
Suliang Bu, Yunxin Zhao, Yanmin Qian, Kai Yu 0004
ICASSP2
2015 Time-frequency kernel-based CNN for speech recognition
Tuo Zhao, Yunxin Zhao
INTERSPEECH2
2013 Real and imaginary modulation spectral subtraction for speech enhancement
Yunxin Zhao
Speech Commun.2
2013 Modulation domain blind speech separation in noisy environments
Yunxin Zhao
Speech Commun.2
2013 Building Acoustic Model Ensembles by Data Sampling With Enhanced Trainings and Features
abstract
We propose a novel approach of using Cross Validation (CV) and Speaker Clustering (SC) based data samplings to construct an ensemble of acoustic models for speech recognition. We also investigate the effects of the existing techniques of Cross Validation Expectation Maximization (CVEM), Discriminative Training (DT), and Multiple Layer Perceptron (MLP) features on the quality of the proposed ensemble acoustic models (EAMs). We have evaluated the proposed methods on TIMIT phoneme recognition task as well as on a telemedicine automatic captioning task. The proposed methods have led to significant improvements in recognition accuracy over conventional Hidden Markov Model (HMM) baseline systems, and the integration of EAMs with CVEM, DT, and MLP has also significantly improved the accuracy performances of the single model systems based on CVEM, DT, and MLP, where the increased inter-model diversity is shown to have played an important role in the performance gain.
Yunxin Zhao
IEEE Trans. Speech Audio Process.2
2012 Modulation domain blind source separation for noisy speech mixture
Yunxin Zhao
INTERSPEECH2
2012 The Latent Maximum Entropy Principle
abstract
We present an extension to Jaynes’ maximum entropy principle that incorporates latent variables. The principle oflatent maximum entropywe propose is different from both Jaynes’ maximum entropy principle and maximum likelihood estimation, but can yield better estimates in the presence of hidden variables and limited training data. We first show that solving for a latent maximum entropy model poses a hard nonlinear constrained optimization problem in general. However, we then show that feasible solutions to this problem can be obtained efficiently for the special case of log-linear models---which forms the basis for an efficient approximation to the latent maximum entropy principle. We derive an algorithm that combines expectation-maximization with iterative scaling to produce feasible log-linear solutions. This algorithm can be interpreted as an alternating minimization algorithm in the information divergence, and reveals an intimate connection between the latent maximum entropy and maximum likelihood principles. To select a final model, we generate a series of feasible candidates, calculate the entropy of each, and choose the model that attains the highest entropy. Our experimental results show that estimation based on the latent maximum entropy principle generally gives better results than maximum likelihood when estimating latent variable models on small observed data samples.
Dale Schuurmans, Yunxin Zhao
ACM Trans. Knowl. Discov. Data3
2011 Clustering of bootstrapped acoustic model with full covariance
abstract
HMM-based acoustic models built from bootstrap are generally very large, especially when full covariance matrices are used for Gaussians. Therefore, clustering is needed to compact the acoustic model to a reasonable size for practical applications. This paper discusses and investigates multiple distance measurements and algorithms for the clustering. The distance measurements include Entropy, KL, Bhattacharyya, Chernoff and their weighted versions. For clustering algorithms, besides conventional greedy bottom-up, algorithms such as N-Best distance Refinement (NBR), K-step Look-Ahead (KLA), Breadth-First Searched (BFS) best path are proposed. A two-pass optimization approach is also proposed to improve the model structure. Experiments in the Bootstrap and Restructuring (B SRS) frame work on Pashto show that the discussed clustering approach can lead to better quality of the restructured model. It also shows that final acoustic model that is diagonalized from the full covariance yields good improvement over BSRS model directly with diagonal model and yields significant improvement over the conventional diagonal model.
Peder A. Olsen, John R. Hershey, Bowen Zhou 0006, Yunxin Zhao
ICASSP7
2011 Spectral subtraction on real and imaginary modulation spectra
abstract
In this paper, we propose a novel phase preserving spectral subtraction method for enhancing speech in noise. Instead of the conventional approach of carrying out subtraction on the magnitude spectrum in the acoustic frequency domain, we propose to perform subtraction on the real and imaginary spectra separately in the modulation frequency domain. By doing so, we are able to enhance magnitude as well as phase through spectral subtraction. Our experimental results have demonstrated that the proposed method improved magnitude and phase estimation significantly at low SNR, and accordingly it significantly outperformed existing methods in speech enhancement evaluated on the criteria of segmental SNR and PESQ.
Yunxin Zhao
ICASSP2
2011 On the Effectiveness of Statistical Modeling Based Template Matching Approach for Continuous Speech Recognition
Xie Sun, Yunxin Zhao
INTERSPEECH3
2011 New Methods for Template Selection and Compression in Continuous Speech Recognition
Xie Sun, Yunxin Zhao
INTERSPEECH2
2010 Data sampling ensemble acoustic modelling in speaker independent speech recognition
abstract
In this paper, we extend our recent data-sampling based ensemble acoustic modeling technique for the speaker-independent task of TIMIT and propose new methods to further improve the effectiveness of the ensemble acoustic models. We propose applying overlapped speaker clustering in data sampling to construct an ensemble of acoustic models for speaker independent speech recognition. In addition, we evaluate the method of data sampling in recurrent neural network for constructing a RNN based frame classifier. We also investigate using CVEM in place of EM in our ensemble acoustic model training. By using these methods on the speaker independent TIMIT phone recognition task, we have obtained a 2.5% absolute gain on phone accuracy over a standard HMM baseline system.
Yunxin Zhao
ICASSP2
2010 Integrating MLP features and discriminative training in data sampling based ensemble acoustic modeling
Yunxin Zhao
INTERSPEECH2
2010 Integrate template matching and statistical modeling for speech recognition
Xie Sun, Yunxin Zhao
INTERSPEECH2
2009 Data sampling based ensemble acoustic modelling
abstract
In this paper, we propose a novel technique of using cross validation (CV) data sampling to construct an ensemble of acoustic models for conversational speech recognition. We further propose using hierarchical Gaussian mixture model (HGMM) and repartition training data to increase the ensemble size and diversity. The proposed methods are found to work well together for ensemble acoustic modeling. We also evaluated the quality of the ensemble acoustic models by using the measures of classification margin, average correct score and variance of correct score. We have found that the ensemble of acoustic models increases the margin and the average correct score, and reduces the variance. We compared the performance of our proposed method with a recently reported method of CV expectation maximization (CVEM) for single acoustic models. Our experimental results on a telemedicine automatic captioning task showed that the proposed ensemble acoustic modeling has led to significant improvements in word recognition accuracy.
Yunxin Zhao
ICASSP2
2009 Semi-tied covariance matrices for acoustic models based on random forests of phonetic decision trees
abstract
In this paper, we investigate combining semi-tied covariance matrices and random forests (RFs) based phonetic decision trees (PDTs) for acoustic modeling in conversational speech recognition. We first use the RF method to train multiple PDTs for each phone state unit, and generate multiple sets of acoustic models accordingly. We then apply semi-tied covariance matrices to each set of acoustic models to improve their fit to data. In decoding search we combine the likelihood scores from the multiple acoustic models for each speech frame. The viability of semi-tied covariance matrices with different tying classes are studied from their effects on the diversity of RF-based acoustic models as well as on the word accuracy of our task of telehealth automatic captioning. Experimental results indicate that semi-tied covariance matrices help enhance the diversity of the RFs-PDTs based acoustic models as well as increase word accuracy.
L. Che, Yunxin Zhao
ICASSP3
2008 Random-forests-based phonetic decision trees for conversational speech recognition
abstract
In this paper we present a novel technique of constructing phonetic decision trees (PDTs) for acoustic modeling in conversational speech recognition. We use random forests (RF) to train a set of PDTs for each phone-state unit and obtain multiple acoustic models accordingly, and we extend the PDT-based state tying to RF-based state-tying. We combine acoustic scores at the model level in decoding search. Several methods are investigated to estimate the weight parameters for model combination, including maximum likelihood estimation of the weights from training data, as well as using confidence scores of P-value or relative entropy to obtain the weights dynamically from online data. Experimental results on a telemedicine automatic captioning task demonstrate that the proposed RF-PDT technique leads to significant improvements in word recognition accuracy.
Yunxin Zhao
ICASSP2
2008 Random Forests of Phonetic Decision Trees for Acoustic Modeling in Conversational Speech Recognition
abstract
In this paper, we present a novel technique of constructing phonetic decision trees (PDTs) for acoustic modeling in conversational speech recognition. We use random forests (RFs) to train a set of PDTs for each phone state unit and obtain multiple acoustic models accordingly. We investigate several methods of combining acoustic scores from the multiple models, including maximum-likelihood estimation of the weights of different acoustic models from training data, as well as using confidence score of -value or relative entropy to obtain the weights dynamically from online data. Since computing acoustic scores from the multiple models slows down decoding search, we propose clustering methods to compact the RF-generated acoustic models. The conventional concept of PDT-based state tying is extended to RF-based state tying. On each RF tied state, we cluster the Gaussian density functions (GDFs) from multiple acoustic models into classes and compute a prototype for each class to represent the original GDFs. In this way, the number of GDFs in each RF tied state is decreased greatly, which significantly reduces the time for computing acoustic scores. Experimental results on a telemedicine automatic captioning task demonstrate that the proposed RF-PDT technique leads to significant improvements in word recognition accuracy.
Yunxin Zhao
IEEE Trans. Speech Audio Process.2
2007 A Bayesian Approach for Phonetic Decision Tree State Tying in Conversational Speech Recognition
abstract
This paper presents a new method of constructing phonetic decision trees (PDTs) for acoustic model state tying based on implicitly induced prior knowledge. Our hypothesis is that knowledge of pronunciation variation in spontaneous, conversational speech contained in a relatively large corpus can be used for building domain-specific or speaker-dependent PDTs. In the view of tree structure adaptation, this method leads to transformation of tree topology in contrast to keeping fixed tree structure as in traditional methods of speaker adaptation. A Bayesian learning framework is proposed to incorporate prior knowledge of decision rules in a greedy search of new decision trees, where the prior is generated by a decision tree growing process on a large data set. Experimental results on the telemedicine automatic captioning task demonstrate that the proposed approach results in consistent improvement in model quality and recognition accuracy.
Rusheng Hu, Yunxin Zhao
ICASSP (4)2
2007 Prior knowledge guided maximum expected likelihood based model selection and adaptation for nonnative speech recognition
Xiaodong He 0001, Yunxin Zhao
Comput. Speech Lang.2
2007 A fast and memory-efficient N-gram language model lookup method for large vocabulary continuous speech recognition
Yunxin Zhao
Comput. Speech Lang.2
2007 Knowledge-Based Adaptive Decision Tree State Tying for Conversational Speech Recognition
abstract
This paper presents a new method of constructing phonetic decision trees (PDTs) for acoustic model state tying based on implicitly induced prior knowledge. Our hypothesis is that knowledge on pronunciation variation in spontaneous, conversational speech contained in a relatively large corpus can be used for building domain-specific or speaker-dependent PDTs. In view of tree-structure adaptation, this method leads to transformation of tree topology in contrast to keeping fixed tree structure as in traditional methods of speaker adaptation. A Bayesian learning framework is proposed to incorporate prior knowledge on decision rules in a greedy search of new decision trees, where the prior is generated by a decision tree growing process on a large data set. Experimental results on the telemedicine automatic captioning task demonstrate that the proposed approach results in consistent improvement in model quality and recognition accuracy.
Rusheng Hu, Yunxin Zhao
IEEE Trans. Speech Audio Process.2
2007 A Novel Method of Language Modeling for Automatic Captioning in TC Video Teleconferencing
abstract
We are developing an automatic captioning system for teleconsultation video teleconferencing (TC-VTC) in telemedicine, based on large vocabulary conversational speech recognition. In TC-VTC, doctors' speech contains a large number of infrequently used medical terms in spontaneous styles. Due to insufficiency of data, we adopted mixture language modeling, with models trained from several datasets of medical and nonmedical domains. This paper proposes novel modeling and estimation methods for the mixture language model (LM). Component LMs are trained from individual datasets, with class n-gram LMs trained from in-domain datasets and word n-gram LMs trained from out-of-domain datasets, and they are interpolated into a mixture LM. For class LMs, semantic categories are used for class definition on medical terms, names, and digits. The interpolation weights of a mixture LM are estimated by a greedy algorithm of forward weight adjustment (FWA). The proposed mixing of in-domain class LMs and out-of-domain word LMs, the semantic definitions of word classes, as well as the weight-estimation algorithm of FWA are effective on the TC-VTC task. As compared with using mixtures of word LMs with weights estimated by the conventional expectation-maximization algorithm, the proposed methods led to a 21% reduction of perplexity on test sets of five doctors, which translated into improvements of captioning accuracy.
Xiaojia Zhang, Yunxin Zhao, Laura Schopp
IEEE Trans. Inf. Technol. Biomed.2
2006 Gradient Boosting Learning of Hidden Markov Models
abstract
In this paper, we present a new training algorithm, gradient boosting learning, for Gaussian mixture density (GMD) based acoustic models. This algorithm is based on a function approximation scheme from the perspective of optimization in function space rather than parameter space, i.e., stage-wise additive expansions of GMDs are used to search for optimal models instead of gradient descent optimization of model parameters. In the proposed approach, GMD starts from a single Gaussian and is built up by sequentially adding new components. Each new component is globally selected to produce optimal gain in the objective function. MLE and MMI are unified under the H-criterion, which is optimized by the extended BW (EBW) algorithm. A partial extended EM algorithm is developed for stage-wise optimization of new components. Experimental results on WSJ task demonstrate that the new algorithm leads to improved model quality and recognition performance
Rusheng Hu, Yunxin Zhao
ICASSP (1)3
2006 Fast Noise Compensation for Speech Separation in Diffuse Noise
abstract
In this paper, a fast noise compensation (FNC) algorithm is proposed for the adaptive decorrelation filtering (ADF) speech separation system in the presence of diffuse noise. The adaptation of ADF is a dynamic process, making noise effects at ADF outputs time-varying in nature. Such changing noise effects need to be tracked and adaptively removed. Under the assumption that acoustic paths are slow in change and utilizing the filtering structure of the separation model, noise compensation terms were adapted with an FFT-based fast algorithm. Experiments were based on both simulated and real recorded diffuse noises. Strong diffuse noise distracts part of the attention of ADF to do noise cancellation while separating speech, and FNC works by forcing ADF to stay focused on speech separation task. The proposed algorithm significantly improved the separation performance of ADF system in diffuse noise.
Yunxin Zhao
ICASSP (5)2
2006 Random Forests-Based Confidence Annotation Using Novel Features from Confusion Network
abstract
In this paper, we propose a set of new features for confidence annotation, including three features derived from confusion network and one from statistical significance test. We also propose using random forests as confidence classifier. The new features are combined with a set of eight previously proposed confidence features, and the random forests is compared with decision tree and support vector machine. Experiments were conducted on telehealth captioning task with a vocabulary size of 46,489. Average confidence annotation accuracy of 84.69% was achieved on 5 doctors' test set. In addition, random forests was shown useful for feature importance ranking. The proposed features are shown important in confidence annotation and random forests achieved best results among the three classifiers
Yunxin Zhao
ICASSP (1)2
2006 An Automatic Captioning System for Telemedicine
abstract
In this paper, we present a first exposition of an automatic closed captioning system designed to assist hearing impaired users in telemedicine. This system automatically separates telehealth conversation speech between a health care provider and a client into two streams and provides real-time captions of health care provider's speech to client. The captioning system is based on the state-of-the-art technology of large vocabulary conversational speech recognition, encompassing speech stream separation, acoustic modeling, language modeling, real-time decoding, confidence annotation, and human-computer interface, with innovations made in several components. The system currently handles a vocabulary size over 46 K. Real-time captioning performance at the average word accuracy of 77.95% is reported
Yunxin Zhao, Xiaojia Zhang, Rusheng Hu, Lili Che, Laura Schopp
ICASSP (1)1
2006 Bayesian decision tree state tying for conversational speech recognition
Rusheng Hu, Yunxin Zhao
INTERSPEECH2
2006 Adaptive speech enhancement for speech separation in diffuse noise
Yunxin Zhao
INTERSPEECH2
2006 New improvements in decoding speed and latency for automatic captioning
Rusheng Hu, Yunxin Zhao
INTERSPEECH3
2006 Speedup convergence and reduce noise for enhanced speech separation and recognition
abstract
Novel techniques are proposed to enhance time-domain adaptive decorrelation filtering (ADF) for separation and recognition of cochannel speech in reverberant room conditions. The enhancement techniques include whitening filtering on cochannel speech to improve condition of adaptive estimation, block-iterative formulation of ADF to speed up convergence, and integration of multiple ADF outputs through post filtering to reduce reverberation noise. Experimental data were generated by convolving TIMIT speech with acoustic path impulse responses measured in real room environment, with approximately 2 m microphone-source distance and initial target-to-interference ratio of about 0 dB. The proposed techniques significantly improved ADF convergence rate, target-to-interference ratio, and accuracy of phone recognition
Yunxin Zhao
IEEE Trans. Speech Audio Process.1
2005 Acoustic Model Training Using Greedy EM
abstract
We present a greedy EM (GEM) method for training Gaussian mixture density (GMD) based acoustic models. In the proposed approach, starting from a single Gaussian, GMD is built up by sequentially adding new components. Each new component is globally selected to avoid local optima. The sequential procedure offers more control over the model structure to achieve better coverage of data. GEM also provides a natural way of integrating information criteria for model complexity selection. Experimental results on a WSJ task show that the new method performs consistently better than the conventional method in speech recognition word error rate.
Rusheng Hu, Yunxin Zhao
ICASSP (1)3
2005 Adaptive Decorrelation Filtering Algorithm for Speech Source Separation in Uncorrelated Noises
abstract
A time-domain vector formulation is presented on analysis and solution of the adaptive decorrelation filtering (ADF) system for blind speech source separation in additive background noise. The formulation leads to a derivation of a gradient descent algorithm and offers insights into the impact of uncorrelated white noise on ADF systems. A new noise-adapted ADF algorithm is then derived by modifying the decorrelation criterion function to exclude noise impact on cross-correlation information of ADF outputs. Speech separation simulations were based on convolutive mixtures of TIMIT speech data generated with long impulse response measured in a real reverberant acoustic experiment. The proposed algorithm significantly improved convergence rate, gains in target-to-interference ratio at ADF outputs, and phone accuracy of separated target speech source.
Yunxin Zhao
ICASSP (1)2
2005 Improved Confusion Network Algorithm and Shortest Path Search from Word Lattice
abstract
We propose a novel confusion network (CN) generation algorithm with linear time complexity, O(T), which is capable of transforming a very large lattice into a confusion network with insignificant time. We further extend the confusion network concept to incorporate the case that a long word is split into short words. Finally, we develop a shortest path search algorithm that finds a sentence hypothesis from a word lattice to minimize the expected word error rate directly. The proposed algorithms are evaluated on the Switchboard task, where significant reduction of computation time was observed for the proposed confusion network algorithm as compared with a previously proposed confusion network algorithm, and improved word accuracy performance was observed for both the proposed CN algorithm and the shortest path algorithm as compared with one-best beam search decoding.
Yunxin Zhao
ICASSP (1)2
2005 Incremental largest margin linear regression and MAP adaptation for speech separation in telemedicine applications
Rusheng Hu, Yunxin Zhao
INTERSPEECH3
2005 Variable step size adaptive decorrelation filtering for competing speech separation
Yunxin Zhao
INTERSPEECH2
2005 Combining Statistical Language Models via the Latent Maximum Entropy Principle
Dale Schuurmans, Fuchun Peng, Yunxin Zhao
Mach. Learn.4
2004 Prior knowledge guided MEL based model selection and adaptation for nonnative speech recognition
abstract
An improved method of model complexity selection for nonnative speech recognition is proposed by using maximum a posteriori estimation of bias distributions. An algorithm is described for estimating the hyper-parameters of the prior distributions, and an automatic accent detection algorithm is also proposed for integration with dynamic model selection and adaptation. Experiments were performed on the WSJ1 task with American English speech, British accent speech, and Mandarin Chinese accent speech. Results show that the use of prior knowledge of accents enabled reliable estimation of bias distributions in the case of a very small amount of adaptation speech, or without adaptation speech. Recognition results show that the new approach is superior to the previous MEL (maximum expected likelihood) method, especially when the adaptation data are extremely limited.
Xiaodong He 0001, Yunxin Zhao
ICASSP (1)2
2004 Fast convergence speech source separation in reverberant acoustic environment
abstract
Three significant enhancements to time-domain adaptive decorrelation filtering (ADF) are proposed for effective separation and recognition of simultaneous speech sources in reverberant room conditions. The methods include whitening filtering on cochannel speech prior to ADF to improve condition of adaptive estimation, a novel block-iterative implementation of ADF to speed up convergence rate, and an integration of multiple ADF outputs through optimal post-filtering. Experimental data were generated by convolving TIMIT speech with acoustic path impulse responses measured in a real acoustic environment, with a 2m microphone-source distance and an initial target-to-interference ratio of about 0 dB. The proposed methods are shown to have speeded up the convergence rate of ADF to a level feasible for online applications, and they have significantly improved target-to-interference ratio and accuracy of phone recognition.
Yunxin Zhao
ICASSP (3)1
2004 Learning mixture models with the regularized latent maximum entropy principle
abstract
This paper presents a new approach to estimating mixture models based on a recent inference principle we have proposed: the latent maximum entropy principle (LME). LME is different from Jaynes' maximum entropy principle, standard maximum likelihood, and maximum aposteriori probability estimation. We demonstrate the LME principle by deriving new algorithms for mixture model estimation, and show how robust new variants of the expectation maximization (EM) algorithm can be developed. We show that a regularized version of LME (RLME), is effective at estimating mixture models. It generally yields better results than plain LME, which in turn is often better than maximum likelihood and maximum a posterior estimation, particularly when inferring latent variable models from small amounts of data.
Dale Schuurmans, Fuchun Peng, Yunxin Zhao
IEEE Trans. Neural Networks4
2003 Semantic n-gram language modeling with the latent maximum entropy principle
abstract
We describe a unified probabilistic framework for statistical language modeling-the latent maximum entropy principle-which can effectively incorporate various aspects of natural language, such as local word interaction, syntactic structure and semantic document information. Unlike previous work on maximum entropy methods for language modeling, which only allow explicit features to be modeled, our framework also allows relationships over hidden features to be captured, resulting in a more expressive language model. We describe efficient algorithms for marginalization, inference and normalization in our extended models. We then present experimental results for our approach on the Wall Street Journal corpus.
Dale Schuurmans, Fuchun Peng, Yunxin Zhao
ICASSP (1)4
2003 Learning Mixture Models with the Latent Maximum Entropy Principle
Dale Schuurmans, Fuchun Peng, Yunxin Zhao
ICML4
2003 Exploiting order-preserving perfect hashing to speedup n-gram language model lookahead
abstract
Minimum Perfect Hashing (MPH) has recently been shown successful in reducing Language Model (LM) lookahead time in LVCSR decoding. In this paper we propose to exploit the orderpreserving (OP) property of a string-key based MPH function to further reduce hashing operation and speed up LM lookahead. A subtree structure is proposed for LM lookahead and an orderpreserving MPH is integrated into the structure design. Subtrees are generated on demand and stored in caches. Experiments were performed on Switchboard data. By using the proposed method of OP MPH and subtree cache structure for both trigrams and backoff bigrams, the LM lookahead time was reduced by a factor of 2.9 in comparison with the baseline case of using MPH alone.
Yunxin Zhao
INTERSPEECH2
2003 Boltzmann Machine Learning with the Latent Maximum Entropy Principle
Dale Schuurmans, Fuchun Peng, Yunxin Zhao
UAI4
2003 Fast model selection based speaker adaptation for nonnative speech
abstract
The problem of adapting acoustic models of native English speech to nonnative speakers is addressed from a perspective of adaptive model complexity selection. The goal is to select model complexity dynamically for each nonnative talker so as to optimize the balance between model robustness to pronunciation variations and model detailedness for discrimination of speech sounds. A maximum expected likelihood (MEL) based technique is proposed to enable reliable complexity selection when adaptation data are sparse, where expectation of log-likelihood (EL) of adaptation data is computed based on distributions of mismatch biases between model and data, and model complexity is selected to maximize EL. The MEL based complexity selection is further combined with MLLR (maximum likelihood linear regression) to enable adaptation of both complexity and parameters of acoustic models. Experiments were performed on WSJ1 data of speakers with a wide range of foreign accents. Results show that the MEL based complexity selection is feasible when using as little as one adaptation utterance, and it is able to select dynamically the proper model complexity as the adaptation data increases. Compared with the standard MLLR, the MEL+MLLR method leads to consistent and significant improvement to recognition accuracy on nonnative speakers, without performance degradation on native speakers.
Xiaodong He 0001, Yunxin Zhao
IEEE Trans. Speech Audio Process.2
2003 Training DHMMs of mine and clutter to minimize landmine detection errors
abstract
Minimum classification error (MCE) training is proposed to improve performance of a discrete hidden Markov model (DHMM)-based landmine detection system. The system (baseline) was proposed previously for detection of both metal and nonmetal mines from ground-penetrating radar signatures collected by moving vehicles. An initial DHMM model is trained by conventional methods of vector quantization and the Baum-Welch algorithm. A sequential generalized probabilistic descent (GPD) algorithm that minimizes an empirical loss function is then used to estimate the landmine/background DHMM parameters, and an evolutionary algorithm (EA) based on fitness score of classification accuracy is used to generate and select codebooks. The landmine data of one geographical site was used for model training, and those of two different sites were used for evaluation of system performance. Three scenarios were studied: 1) apply MCE/GPD alone to DHMM estimation, 2) apply EA alone to codebook generation, and 3) first apply EA to codebook generation and then apply MCE/GPD to DHMM estimation. Overall, the combined EA and MCE/GPD training led to the best performance. At the same level of detection rate as the baseline DHMM system, the false-alarm rate was reduced by a factor of two, indicating significant performance improvement.
Yunxin Zhao, Paul D. Gader
IEEE Trans. Geosci. Remote. Sens.1
2002 Fast model adaptation and complexity selection for nonnative English speakers
abstract
In this paper, the problem of fast model adaptation and complexity selection for nonnative speaker is investigated. The key challenge lies in reliable complexity selection when only a small amount of adaptation data is available. A novel technique of combining a maximum likelihood (ML) based state-tying with a pseudo likelihood (PL) based state-tying is proposed to enable model complexity selection from using as little as three adaptation speech sentences. In MUPL, ML model complexity selection is performed on nodes with sufficient adaptation data, and PL based state tying is performed on nodes with insufficient adaptation data. Experiments were performed on WSJ data of six nonnative speakers. The combined model adaptation and complexity selection method led to consistent and significant improvement on recognition accuracy over MLLR, with an average error reduction of 13% when a varying number of adaptation speech sentences were taken from each speaker.
Xiaodong He 0001, Yunxin Zhao
ICASSP2
2002 Co-channel speech separation for assistive listening
abstract
A study is made on applying adaptive decorrelation filtering (ADF) to the design of an assistive listening system. The speech signal of a desired talker was corrupted by three simultaneous speech jammers and a speech-shaped diffusive noise. ADF was used to extract the desired speech from the jammers and noise. The effectiveness of the assistive listening system was measured by improvement in signal-to-noise ratio (SNR) and in word correct percentage, with the latter evaluated by a formal clinical test on subjects of both normal and impaired hearings. Significant gains in SNR and word correct percentage were obtained with the use of the assistive listening system. On eight subjects with normal hearing, the speech reception threshold was improved by 3 to 5 dBA, and on three subjects with hearing impairments, the threshold was improved by 4 to 8 dBA.
Yunxin Zhao, Kuan-Chieh Yen, Sigfrid D. Soli, Shawn X. Gao, Andy Vermiglio
ICASSP1
2002 Maximum expected likelihood based model selection and adaptation for nonnative English speakers
abstract
In this paper, the problem of fast model adaptation for nonnative speakers is addressed from a perspective of model complexity selection. The key challenge lies in reliable complexity selection when only a small amount of adaptation data is available. A novel maximum expected likelihood (MEL) based technique is proposed to enable model complexity selection from using as little as one adaptation sentence. In MEL, the expectation of loglikelihood is computed based on the mismatch bias between model and data which is measured by a small amount of adaptation data, and model complexity is selected to maximize EL. Experiments were performed on WSJ data of speakers with a wide range of foreign accents. The proposed method led to consistent and significant improvement on recognition accuracy over MLLR for nonnative speakers, without performance degradation on native speakers. The proposed method was able to dynamically select optimal model complexity as the available adaptation data increased. 1.
Xiaodong He 0001, Yunxin Zhao
INTERSPEECH2
2002 Minimum perfect hashing for fast n-gram language model lookup
Yunxin Zhao
INTERSPEECH2
2001 Lattice-ladder decorrelation filters developed for co-channel speech separation
abstract
The previously proposed lattice-ladder adaptive decorrelation filtering (LL-ADF) algorithm (Ken and Zhao 1999) is further studied and improved in this work, with the aim of developing a more efficient co-channel speech separation system. The effect of the joint linear predictions is first analyzed and the conversions between the lattice coefficients and the prediction and filter vectors are formulated. The implementation issues on the estimation of lattice coefficients are then discussed and the adaptation equations are further refined. Experimental results demonstrate the effectiveness of the algorithm in reducing cross-interference between co-channel speech sources as well as the significant performance improvement over the previous direct-form ADF algorithm. A simplified LL-ADF is also proposed as a compromise between computational cost and system performance.
Kuan-Chieh Yen, Yunxin Zhao
ICASSP2
2001 Recursive estimation of time-varying environments for robust speech recognition
abstract
An EM-type of recursive estimation algorithm is formulated in the DFT domain for joint estimation of time-varying parameters of distortion channel and additive noise from online degraded speech. Speech features are estimated from the posterior estimates of short-time speech power spectra in an on-the-fly fashion. Experiments were performed on speaker-independent continuous speech recognition using features of perceptually based linear prediction cepstral coefficients, log energy, and temporal regression coefficients. Speech data were taken from the TIMIT database and were degraded by simulated time-varying channel and noise. Experimental results showed significant improvement in recognition word accuracy due to the proposed recursive estimation as compared with the results from direct recognition using a baseline system and from performing speech feature estimation using a batch EM algorithm.
Yunxin Zhao, Kuan-Chieh Yen
ICASSP1
2001 Model complexity optimization for nonnative English speakers
abstract
In this paper, a study is made on selecting existing acoustic models that are trained from native English speech for improving recognition of nonnative English talkers ’ speech. The problem is addressed from the perspective that foreign accents prevent detailed triphone models that are commonly used in highperformance speech recognition systems to match well with these talkers ’ speech, and therefore an appropriate level of context-dependent acoustic modeling is needed for foreign accent speakers. In this work, model complexity selection is accomplished by empirically choosing a set of model tying thresholds and by using the principle of MDL. An experiment was performed on the Wall Street Journal task on three nonnative English talkers with Chinese accent (276 sentences). Compared to the result obtained from using the models optimized to native English speakers, the best model tying threshold and MDL yielded similar and significant reduction to recognition word errors by 23%. 1.
Xiaodong He 0001, Yunxin Zhao
INTERSPEECH2
2001 Online Bayesian tree-structured transformation of HMMs with optimal model selection for speaker adaptation
abstract
This paper presents a new recursive Bayesian learning approach for transformation parameter estimation in speaker adaptation. Our goal is to incrementally transform or adapt a set of hidden Markov model (HMM) parameters for a new speaker and gain large performance improvement from a small amount of adaptation data. By constructing a clustering tree of HMM Gaussian mixture components, the linear regression (LR) or affine transformation parameters for HMM Gaussian mixture components are dynamically searched. An online Bayesian learning technique is proposed for recursive maximum a posteriori (MAP) estimation of LR and affine transformation parameters. This technique has the advantages of being able to accommodate flexible forms of transformation functions as well as a priori probability density functions (PDFs). To balance between model complexity and goodness of fit to adaptation data, a dynamic programming algorithm is developed for selecting models using a Bayesian variant of the "minimum description length" (MDL) principle. Speaker adaptation experiments with a 26-letter English alphabet vocabulary were conducted, and the results confirmed effectiveness of the online learning framework.
Yunxin Zhao
IEEE Trans. Speech Audio Process.2
2001 Landmine detection with ground penetrating radar using hidden Markov models
abstract
Novel, general methods for detecting landmine signatures in ground penetrating radar (GPR) using hidden Markov models (HMMs) are proposed and evaluated. The methods are evaluated on real data collected by a GPR mounted on a moving vehicle at three different geographical locations. A large library of digital GPR signatures of both landmines and clutter/background was constructed and used for training. Simple, but effective, observation vector representations are constructed to naturally model the time-varying signatures produced by the interaction of the GPR and the landmines as the vehicle moves. The number and definition of the states of the HMMs are based on qualitative signature models. The model parameters are optimized using the Baum-Welch algorithm. The models were trained on landmine and background/clutter signatures from one geographical location and successfully tested at two different locations. The data used in the test were acquired from over 6000 m/sup 2/ of simulated dirt and gravel roads, and also off-road conditions. These data contained approximately 300 landmine signatures, over half of which were plastic-cased or completely nonmetal.
Paul D. Gader, Miroslaw Mystkowski, Yunxin Zhao
IEEE Trans. Geosci. Remote. Sens.3
2000 On-line Bayesian speaker adaptation using tree-structured transformation and robust priors
abstract
This paper presents new results by using our previously proposed on-line Bayesian learning approach for affine transformation parameter estimation in speaker adaptation. The on-line Bayesian learning technique allows updating parameter estimates after each utterance and it can accommodate flexible forms of transformation functions as well as prior probability density functions. We show through experimental results the robustness of heavy tailed priors to mismatch in prior density estimation. We also show that by properly choosing the transformation matrices and depths of hierarchical trees, recognition performance improved significantly.
Yunxin Zhao
ICASSP2
2000 Lattice-ladder structured adaptive decorrelation filtering for co-channel speech separation
abstract
A lattice-ladder adaptive decorrelation filtering (ADF) algorithm is proposed with the aim of developing a more efficient co-channel speech separation system. It is shown that based on the joint forward and backward linear predictions, a lattice-ladder structure can be derived for ADF. Experimental results show that the proposed algorithm is effective in reducing cross-interference between co-channel speech sources and has better tracking ability on the variations in the acoustic environment compared to the original direct-form ADF algorithm.
Kuan-Chieh Yen, Yunxin Zhao
ICASSP2
2000 Maximum likelihood joint estimation of channel and noise for robust speech recognition
abstract
An EM algorithm is formulated in the DFT domain for joint estimation of parameters of distortion channel and additive noise from online degraded speech, and the posterior estimates of short-time speech power spectra are obtained at the convergence of the EM algorithm. Any speech features derivable from power spectra can then be approximately estimated by minimum mean-squared error estimation. Experiments were performed on speaker-independent continuous speech recognition using as features the perceptually based linear prediction cepstral coefficients, energy, and temporal regression coefficients. Speech data were taken from the TIMIT database and were degraded by a distortion channel and colored noise at various SNR levels. Experimental results indicate that the proposed technique leads to convergent identification of channel and noise and significantly improved recognition accuracy.
Yunxin Zhao
ICASSP1
2000 Optimal on-line Bayesian model selection for speaker adaptation
abstract
In this paper, we show how to accommodate a Bayesian variant of Rissanen's MDL into on-line Bayesian adaptation to control both model structural complexity and parameterization complexity to best fit an available amount of adaptation data, the goal being minimization of resulting recognition error. An efficient bottom-up dynamic programming based pruning algorithm is developed for selecting models using the MDL principle. Speaker adaptation experiments using a 26-letter English alphabet vocabulary were conducted and the proposed Bayesian variant MDL method is shown to provide an optimal trade-off between recognition accuracy and complexity of model structure and parameterization over a full range of adaptation data size. It in general is capable of automatically selecting a set of model parameters that leads to best recognition performance for a given amount of adaptation data.
Yunxin Zhao
INTERSPEECH2
2000 A combined adaptive and decision tree based speech separation technique for telemedicine applications
abstract
We present a novel technique for separation of doctor and patient’s speech in conversations over a telemedicine network. The mixed speech signals acquired at doctor’s site is first broken into single talkers ’ speech segments and background by using thresholds of energy and duration. The speech segments are then identified as spoken by doctor or patient in two steps. In the first step, Gaussian mixture models (GMM) of doctor and patient are used, where the doctor’s model is obtained from his/her training speech, and the patient’s model is initialized by a general speaker model and then adapted by the patient’s speech. In the second step, a decision tree that uses contextual and confidence features is applied to refine the identification results. Preliminary experiments were performed on three data sets collected in telemedicine. Without adaptation and decision tree, error rates at the segment-level and frame-level were 25.44 % and 16.53%, respectively. With adaptation, segment and frame error rates were reduced to 13.11 % and 7.85%, and with decision tree, the error rates were further reduced to 10.48 % and 6.73%, respectively. 1.
Yunxin Zhao, Xiaodong He 0001, Laura Schopp
INTERSPEECH1
2000 Subband-based adaptive decorrelation filtering for co-channel speech separation
abstract
A subband-based adaptive decorrelation filtering algorithm (SBADF) is proposed for co-channel speech separation. The SBADF decomposes the input signals into several frequency subbands, and uses the adaptive decorrelation filtering algorithm (ADF) to process the signals in each subband independently. The processed subband signals are then combined for each channel to form the separated speech. Experimental results show that while the fullband ADF can achieve better separation performance after reaching convergence, the SBADF has the advantage of improved convergence rate and reduced computational complexity by a factor of approximately two.
Jonathan Huang, Kuan-Chieh Yen, Yunxin Zhao
IEEE Trans. Speech Audio Process.3
2000 A DCT-based fast signal subspace technique for robust speech recognition
abstract
In this correspondence, a fast computational method is proposed to approximate the Karhunen-Loeve transform (KLT) for the covariance matrix of the autoregressive process. A fast algorithm which reduces the computation of eigenvalues of an N/spl times/N symmetric Toeplitz matrix from O(N/sup 3/) in KLT to N/sup 2/ is further developed. Experimental results demonstrate that the performance of the fast algorithm is very close to the KLT in eigenvalue computation and in energy constrained signal subspace speech enhancement for speech recognition in a car environment.
Yunxin Zhao
IEEE Trans. Speech Audio Process.2
2000 Frequency-domain maximum likelihood estimation for automatic speech recognition in additive and convolutive noises
abstract
A feature estimation technique is proposed for speech signals that are degraded by both additive and convolutive noises. An EM algorithm is formulated in the frequency-domain for identification of the magnitude response of the distortion channel and power spectrum of additive noise, and posterior estimates of short-time power spectra of speech are obtained based on the identified channel and noise. The estimated posterior power spectra are used to calculate perceptually-based linear prediction cepstral coefficients, and the estimated cepstral features and their temporal regression coefficients are used for automatic speech recognition using acoustic models trained from clean speech. Experiments were performed on speaker independent continuous speech recognition, where the speech data were taken from the TIMIT database and were degraded by a distortion channel and simulated additive noises with white or colored spectral characteristics at various SNR levels. Experimental results indicate that the proposed technique leads to convergent identification of channel and noise and significantly improved recognition accuracy for speaker-independent continuous speech.
Yunxin Zhao
IEEE Trans. Speech Audio Process.1
1999 Adaptive decorrelation filtering for separation of co-channel speech signals from m>2 sources
abstract
The ADF algorithm for separating two signal sources by Weinstein, Feder, and Oppenheim (1993) is generalized for separation of co-channel speech signals from more than two sources. The system configuration, its accompanied ADF algorithm, and the choice of adaptation gain are derived. The applicability and limitation of the derived algorithm are also discussed. Experiments were conducted for separation of three speech sources with the acoustic paths measured from an office environment, and the algorithm was shown to improve the average target-to-interference ratio for the three sources by approximately 15 dB.
Kuan-Chieh Yen, Yunxin Zhao
ICASSP2
1999 A DCT-based fast enhancement technique for robust speech recognition in automobile usage
Yunxin Zhao, Stephen E. Levinson
EUROSPEECH2
1999 Co-channel speech separation in the presence of correlated and uncorrelated noises
Kuan-Chieh Yen, Yunxin Zhao
EUROSPEECH3
1999 Channel identification and spectrum estimation for robust automatic speech recognition
Yunxin Zhao
EUROSPEECH1
1999 Adaptive co-channel speech separation and recognition
abstract
An improved technique of co-channel speech separation, S-AADP/LMS, and its integration with automatic speech recognition is presented. The S-AADF/LMS technique is based on the algorithms of accelerated adaptive decorrelation filtering (AADP) and LMS noise cancellation, where a switching between the two algorithms is made depending upon the active/inactive status of the co-channel signal sources. The AADF improves the previous adaptive decorrelation algorithm in terms of system stability and estimation efficiency, and leads to better estimation of time-varying and reverberant channels. The S-AADF/LMS further improves the estimation accuracy when only one source signal remains active during certain periods of time. A coherence-function based source signal detection algorithm is also presented, which is successfully used in the switching between AADF and LMS and in extracting speech signals from leakage-corrupted background. Experiments were conducted under a simulated environment based on the measurements made of certain real room-acoustic conditions, and the results demonstrated the effectiveness of the proposed technique for co-channel speech separation and recognition.
Kuan-Chieh Yen, Yunxin Zhao
IEEE Trans. Speech Audio Process.2
1999 An EM algorithm for linear distortion channel estimation based on observations from a mixture of Gaussian sources
abstract
In this work, an expectation maximization (EM) algorithm is derived for maximum likelihood estimation of the autocorrelation function of a linear distortion channel as well as the level of additive noise, under the assumption that the source signal comes from a mixture of Gaussian sources. To facilitate parameter initialization in the EM algorithm, a correlation-matching based estimation algorithm is developed for the channel autocorrelation function. The proposed EM algorithm was evaluated on speech-derived simulated data of multiple autoregressive Gaussian sources and real speech of isolated digits under signal-to-noise ratios (SNRs) of 20 dB down to 0 dB. The algorithm is shown to produce convergent estimation results as well as estimates of signal statistics that lead to significantly improved classification accuracy under additive and convolutive noise conditions.
Yunxin Zhao
IEEE Trans. Speech Audio Process.1
1998 An energy-constrained signal subspace method for speech enhancement and recognition in colored noise
abstract
An energy-constrained signal subspace (ECSS) method is proposed for speech enhancement and recognition under an additive colored noise condition. The key idea is to match the short-time energy of the enhanced speech signal to the unbiased estimate of the short-time energy of the clean speech, which is proven very effective for improving the estimation of the noise-like, low-energy segments in the speech signal. The colored noise is modelled by an autoregressive (AR) process. A modified covariance method is used to estimate the AR parameters of the colored noise and a prewhitening filter is constructed based on the estimated parameters. The performance of the proposed algorithm was evaluated using the TI46 digit database and the TIMIT continuous speech database. It was found that the ECSS method can significantly improve the signal-to-noise ratio (SNR) and word recognition accuracy (WRA) for isolated digits and continuous speech under various SNR conditions.
Yunxin Zhao
ICASSP2
1998 Improvements on co-channel speech separation using ADF: low complexity, fast convergence, and generalization
abstract
Three modifications on the adaptive decorrelation filtering (ADF) algorithm are proposed to improve the performance of a co-channel speech separation system. Firstly, a simplified ADF (SADF) is suggested to reduce the computational complexity of ADF from O(N/sup 2/) to O(N) per sample, where N is the filter length used in the channel estimation. Secondly, a transform-domain ADF (TDADF) is developed to accelerate the convergence of the filter estimates while maintaining computational complexity at O(N). Thirdly, a generalized ADF (GADF) is derived to handle the noncausal filter estimation problem often encountered in co-channel speech separation. Experimental results showed that when the average signal-to-interference ratios (SIRs) in the co-channel signals were 6.15 and 5.38 dB, respectively, both the SADF and TDADF improved the SIRs to around 18 to 19 dB, and the GADF further improved the SIRs to around 19 to 24 dB.
Kuan-Chieh Yen, Yunxin Zhao
ICASSP2
1998 Robust speech recognition using discriminative stream weighting and parameter interpolation
Stephen M. Chu, Yunxin Zhao
ICSLP2
1998 Recognizing emotions in speech using short-term and long-term features
abstract
The acoustic characteristics of speech are influenced by speakers’ emotional status. In this study, we attempted to recognize the emotional status of individual speakers by using speech features that were extracted from short-time analysis frames as well as speech features that represented entire utterances. Principal component analysis was used to analyze the importance of individual features in representing emotional categories. Three classification methods including vector quantization, artificial neural networks and Gaussian mixture density model were used. Classifications using short-term features only, long-term features only and both short-term and long-term features were conducted. The best recognition performance of 62% accuracy was achieved by using the Gaussian mixture density method with both short-term and longterm features.
Yunxin Zhao
ICSLP2
1998 An energy-constrained signal subspace method for speech enhancement and recognition in white and colored noises
Yunxin Zhao
Speech Commun.2
1998 Channel identification and signal spectrum estimation for robust automatic speech recognition
abstract
A feature estimation technique is proposed for speech signals that are corrupted by both additive and convolutive noises via combining channel identification with power spectrum estimation. A correlation-matching algorithm is developed for channel identification, and a Gaussian mixture density model of speech DFT spectra is formulated for estimation of speech power spectra. Cepstral features of speech are calculated from the estimated power spectra. Using the proposed method, significantly improved accuracy was achieved on speaker-independent continuous speech recognition where the speech data were corrupted by a simulated linear distortion channel and additive white noise.
Yunxin Zhao
IEEE Signal Process. Lett.1
1998 A general model for bidirectional associative memories
abstract
This paper proposes a general model for bidirectional associative memories that associate patterns between the X-space and the Y-space. The general model does not require the usual assumption that the interconnection weight from a neuron in the X-space to a neuron in the Y-space is the same as the one from the Y-space to the X-space. We start by defining a supporting function to measure how well a state supports another state in a general bidirectional associative memory (GBAM). We then use the supporting function to formulate the associative recalling process as a dynamic system, explore its stability and asymptotic stability conditions, and develop an algorithm for learning the asymptotic stability conditions using the Rosenblatt perceptron rule. The effectiveness of the proposed model for recognition of noisy patterns and the performance of the model in terms of storage capacity, attraction, and spurious memories are demonstrated by some outstanding experimental results.
Hongchi Shi, Yunxin Zhao, Xinhua Zhuang
IEEE Trans. Syst. Man Cybern. Part B2
1997 A Visual Computing Environment for Very Large Scale Biomolecular Modeling
abstract
Knowledge of the complex molecular structures of living cells is being accumulated at a tremendous rate. Key technologies enabling this success have been, high performance computing and powerful molecular graphics applications, but the technology is beginning to seriously lag behind challenges posed by the size and number of new structures and by the emerging opportunities in drug design and genetic engineering. A visual computing environment is being developed which permits interactive modeling of biopolymers by linking a 3D molecular graphics program with an efficient molecular dynamics simulation program executed on remote high-performance parallel computers. The system will be ideally suited for distributed computing environments, by utilizing both local 3D graphics facilities and the peak capacity of high-performance computers for the purpose of interactive biomolecular modeling. To create an interactive 3D environment three input methods will be explored: (1) a six degree of freedom "mouse" for controlling the space shared by the model and the user; (2) voice commands monitored through a microphone and recognized by a speech recognition interface; (3) hand gestures, detected through cameras and interpreted using computer vision techniques. Controlling 3D graphics connected to real time simulations and the use of voice with suitable language semantics, as well as hand gestures, promise great benefits for many types of problem solving environments. Our focus on structural biology takes advantage of existing sophisticated software, provides concrete objectives, defines a well-posed domain of tasks and offers a well-developed vocabulary for spoken communication.
Michael Zeller, James C. Phillips, Andrew Dalke, William Humphrey, Klaus Schulten, Thomas S. Huang, Vladimir Pavlovic 0001, Yunxin Zhao, Zion Lo, Stephen M. Chu, Rajeev Sharma
ASAP8
1997 Co-channel speech separation for robust automatic speech recognition: stability and efficiency
abstract
A signal-separation front-end based on adaptive decorrelation filtering (ADF) was integrated with an HMM based speaker independent continuous speech recognition system for co-channel speech recognition. The ADF is improved by addressing the adaptation gain for system stability and efficiency: an upper bound of adaptation rate is derived for system stability, and an accelerated sequence of adaptation gain is introduced for system efficiency. The system was evaluated under simulated room acoustic conditions with both time-invariant and time-varying channels. It is shown that the system significantly improved the signal-to-interference ratio and the recognition word accuracy, and that the combination of the derived upper bound for adaptation rate with the accelerated adaptation gain sequence achieved the best performance for system stability and efficiency.
Kuan-Chieh Yen, Yunxin Zhao
ICASSP2
1997 Parallel, finite-convergence learning algorithms for relaxation labeling processes
abstract
This paper is theoretical. We present sufficient and "almost" necessary conditions for learning compatibility coefficients in relaxation labeling whose satisfaction will guarantee each desired sample labeling to become consistent and each ambiguous or erroneous input sample labeling to be attracted to the corresponding desired sample labeling. The derived learning conditions are parallel and local information based. In fact, they are organized as linear inequalities in unit wise and thus the perceptron like algorithms can be used to solve them efficiently with finite convergence.
Xinhua Zhuang, Yunxin Zhao
ICASSP2
1997 High performance CELP coder utilizing a novel adaptive forward-backward LPC quantization
abstract
A highly efficient algorithm termed adaptive forward-backward vector quantization (AFBVQ) is developed for variable bit rate quantization of linear predictive coding (LPC) coefficients and integrated with the FS1016 Federal Standard Code Excited Linear Predictive (CELP) coder. This results in a high performance low bit rate speech coder called as AFBVQ-CELP which brings in two-fold bit rate reduction by backward LPC indexing and by forward LPC VQ. In AFBVQ, a previously decoded and temporally close speech signal is re-segmented into overlapping blocks. As the LPC coefficients calculated from one of those synthetic blocks are spectrally close to the current unquantized LPC coefficients, the backward LPC indexing is used to encode the current speech block; otherwise, the forward linear prediction is practised with the split vector quantization supported by a very efficient codebook initialization termed Mixture Gaussian Clustering (MGC). When compared to FS1016 CELP coder, AFBVQ-CELP reduces the LPC bit rate by 18 bit-per-frame (bpf) at the same spectral distortion. It means the overall bit rate is reduced from 4.8 kbps (FS1016 CELP) to 4.2 kbps. Furthermore, the proposed AFBVQ consistently outperforms the traditional forward LPC VQ by 3 bpf with the same spectral distortion. Subjective listening tests show that with AFBVQ-CELP the LPC bit rate can be further reduced to 8.4 bpf, resulting in 3.94 kbps overall bit rate without compromising the decoded speech quality.
Zijun Yang, Jozsef Vass, Yunxin Zhao, Xinhua Zhuang
MMSP3
1997 Energy-constrained signal subspace method for speech enhancement and recognition
abstract
In this letter, an improved signal-subspace-based speech enhancement algorithm is proposed for automatic speech recognition under an additive noise environment. The key idea is to match the short-time energy of the enhanced speech signal to the unbiased estimate of the short-time energy of the clean speech, which is proven very effective for improving the estimation of the low-energy segments of continuous speech under low signal-to-noise ratio (SNR) conditions. Experimental results show significant improvement in both the segmental SNR and the word recognition accuracy of the enhanced speech under SNR conditions of 10-20 dB.
Yunxin Zhao
IEEE Signal Process. Lett.2
1997 Adaptive forward-backward quantizer for low bit rate high-quality speech coding
abstract
A novel variable-rate linear predictive coding (LPC) parameter quantization scheme is proposed, in which linear prediction is done by using either the current (forward LPC) or previously decoded (backward LPC) speech blocks. The proposed LPC quantization scheme was integrated into the FS1016 Federal Standard CELP coder. Significant LPC bit rate reduction is achieved without compromising the decoded speech quality.
Jozsef Vass, Yunxin Zhao, Xinhua Zhuang
IEEE Trans. Speech Audio Process.2
1996 A unification of relaxation labeling and associative memory
abstract
This paper attempts to consolidate the theoretical foundation of the relaxation labeling processes, explore the connections between the relaxation labeling model and the Hopfield associative memory model, and seek their unification. We start by defining a new labeling assignment space and then formulate the relaxation labeling process as a dynamic system of Lyapunov type, which is equipped with a well-defined energy function and described by a naturally fitted updating rule. We present a consistency condition and show that each /spl omega/-limit point of the dynamic system gives a consistent labeling. We finally make a peace between the multi-label and one-label relaxation labeling and reveal an interesting result that, for a one-label case, the newly formulated relaxation labeling model reduces to the Hopfield associative memory model.
Xinhua Zhuang, Yunxin Zhao
ICASSP2
1996 Binary linear decision tree with genetic algorithm
abstract
A linear decision binary tree structure is proposed in constructing piecewise linear classifiers with the genetic algorithm (GA) being shaped and employed at each nonterminal node to search for a linear decision function optimal in the sense of maximum impurity reduction. The methodology works for both the two-class and multiclass cases. In comparison to several other well known methods, the proposed binary tree-genetic algorithm (BTGA) is demonstrated to produce a much lower cross validation misclassification rate. Finally, a modified BTGA is applied to the important pap smear cell classification. This results in a spectrum for the combination of the highest desirable sensitivity along with the lowest possible false alarm rate. The multiple choices offered by the spectrum for the sensitivity-false alarm rate combination will provide the flexibility needed for the pap smear slide classification.
Bing-Bing Chai, Xinhua Zhuang, Yunxin Zhao, Jack Sklansky
ICPR3
1996 Speech/gesture interface to a visual computing environment for molecular biologists
abstract
Recent progress in 3-D, immersive display and virtual reality (VR) technologies has made possible many exciting applications, for example interactive visualization of complex scientific data. To fully exploit this potential there is a need for "natural" interfaces that allow the manipulation of such displays without cumbersome attachments. In this paper we describe the use of visual hand gesture analysis and speech recognition for developing a speech/gesture interface for controlling a 3-D display. The interface enhances an existing application, VMD, which is a VR visual computing environment far molecular biologists. The free hand gestures are used for manipulating the 3-D graphical display together with a set of speech commands. We describe the visual gesture analysis and the speech analysis techniques used in developing this interface. The dual modality of speech/gesture is found to greatly aid the interaction capability.
Rajeev Sharma, Thomas S. Huang, Vladimir Pavlovic 0001, Yunxin Zhao, Zion Lo, Stephen M. Chu, Klaus Schulten
ICPR4
1996 Robust automatic speech recognition using a multi-channel signal separation front-end
abstract
A multi-channel signal separation front-end for robust automatic speech recognition under time-varying interference conditions is developed.The speech signals acquired by a dual-channel system are restored by adaptive decorrelation filtering, and then examined by a time-domain or frequency-domain source signal detection technique to determine the active regions of each source signal.The front-end is integrated with an HMM-based speaker-independent continuous speech recognition system by providing the restored signals within the active regions for recognition.Under a simulated room acoustic condition, the overall system shows very promising performance.For the conditions with SNR above -10 dB, the achieved word recognition accuracies are very close to that of the interference-free condition.
Kuan-Chieh Yen, Yunxin Zhao
ICSLP2
1996 Piecewise linear classifiers using binary tree structure and genetic algorithm
Bing-Bing Chai, Xinhua Zhuang, Yunxin Zhao, Jack Sklansky
Pattern Recognit.4
1996 Self-learning speaker and channel adaptation based on spectral variation source decomposition
Yunxin Zhao
Speech Commun.1
1996 Gaussian mixture density modeling, decomposition, and applications
abstract
We present a new approach to the modeling and decomposition of Gaussian mixtures by using robust statistical methods. The mixture distribution is viewed as a contaminated Gaussian density. Using this model and the model-fitting (MF) estimator, we propose a recursive algorithm called the Gaussian mixture density decomposition (GMDD) algorithm for successively identifying each Gaussian component in the mixture. The proposed decomposition scheme has advantages that are desirable but lacking in most existing techniques. In the GMDD algorithm the number of components does not need to be specified a priori, the proportion of noisy data in the mixture can be large, the parameter estimation of each component is virtually initial independent, and the variability in the shape and size of the component densities in the mixture is taken into account. Gaussian mixture density modeling and decomposition has been widely applied in a variety of disciplines that require signal or waveform characterization for classification and recognition. We apply the proposed GMDD algorithm to the identification and extraction of clusters, and the estimation of unknown probability densities. Probability density estimation by identifying a decomposition using the GMDD algorithm, that is, a superposition of normal distributions, is successfully applied to automated cell classification. Computer experiments using both real data and simulated data demonstrate the validity and power of the GMDD algorithm for various models and different noise assumptions.
Xinhua Zhuang, Yan Huang 0010, Kannappan Palaniappan, Yunxin Zhao
IEEE Trans. Image Process.4
1995 Iterative self-learning speaker and channel adaptation under various initial conditions
abstract
A self-learning adaptation technique is presented which handles the speaker and channel induced spectral variations without enrolment speech. At the acoustic level, the distortion spectral bias is estimated in two steps using the unsupervised maximum likelihood estimation: in the first step, the probability distributions of the speech spectral features are assumed uniform for severely mismatched channels; in the second step, the spectral bias is reestimated assuming Gaussian distributions for the spectral features. At the phone unit level, unsupervised sequential adaptation is performed via Bayesian estimation from the online, bias-removed speech data, and iterative adaptation is further performed for dictation applications. Over four 198-sentence test sets, on a continuous speech recognition task with vocabulary size=853 and grammar perplexity=105, the largest increase of average word accuracy is 85.2% from the baseline accuracy of -0.3%, and the maximum average word accuracy is 89.4% from the baseline accuracy of 56.5%.
Yunxin Zhao
ICASSP1
1995 Hierarchical mixture models and phonological rules in open-vocabulary speech recognition
Yunxin Zhao
EUROSPEECH1
1994 An acoustic-phonetic-based speaker adaptation technique for improving speaker-independent continuous speech recognition
abstract
A new speaker adaptation technique is proposed for improving speaker-independent continuous speech recognition based on a decomposition of spectral variation sources. In this technique, the spectral variations are separated into two categories, one acoustic and the other phone-specific, where each variation source is modeled by a linear transformation system. The technique consists of two sequential steps: first, acoustic normalization is performed, and second, phone model parameters are adapted. Experiments of speaker adaptation on the TIMIT database using short calibration speech (5 s per speaker) have shown significant performance improvement over the baseline speaker-independent continuous speech recognition, where the recognition system uses Gaussian mixture density based hidden Markov models of phone units. For a vocabulary size of 853 and test set perplexity of 104, the recognition word accuracy has been improved from 86.9% for the baseline system to 90.5% after adaptation, corresponding to an error reduction of 27.5%. On a more difficult test set that contains an additional variation source due to recording channel mismatch, a more significant performance improvement has been obtained: for the same vocabulary and a test set perplexity of 101, the recognition word accuracy has been improved from 65.4% for the baseline to 86.0% after adaptation, corresponding to an error reduction of 59.5%.>
Yunxin Zhao
IEEE Trans. Speech Audio Process.1
1993 A new speaker adaptation technique using very short calibration speech
Yunxin Zhao
ICASSP (2)1
1993 Speaker normalization using constrained spectra shifts in auditory filter domain
Yoshio Ono, Hisashi Wakita, Yunxin Zhao
EUROSPEECH3
1993 Self-learning speaker adaptation based on spectral variation source decomposition
Yunxin Zhao
EUROSPEECH1
1993 A speaker-independent continuous speech recognition system using continuous mixture Gaussian density HMM of phoneme-sized units
abstract
The author describes a large vocabulary, speaker-independent, continuous speech recognition system which is based on hidden Markov modeling (HMM) of phoneme-sized acoustic units using continuous mixture Gaussian densities. A bottom-up merging algorithm is developed for estimating the parameters of the mixture Gaussian densities, where the resultant number of mixture components is proportional to both the sample size and dispersion of training data. A compression procedure is developed to construct a word transcription dictionary from the acoustic-phonetic labels of sentence utterances. A modified word-pair grammar using context-sensitive grammatical parts is incorporated to constrain task difficulty. The Viterbi beam search is used for decoding. The segmental K-means algorithm is implemented as a baseline for evaluating the bottom-up merging technique. The system has been evaluated on the TIMIT database (1990) for a vocabulary size of 853. For test set perplexities of 24, 104, and 853, the decoding word accuracies are 90.9%, 86.0%, and 62.9%, respectively. For the perplexity of 104, the decoding accuracy achieved by using the merging algorithm is 4.1% higher than that using the segmental K-means (22.8% error reduction), and the decoding accuracy using the compressed dictionary is 3.0% higher than that using a standard dictionary (18.1% error reduction).>
Yunxin Zhao
IEEE Trans. Speech Audio Process.1
1992 Parameter estimation and restoration of noisy images using Gibbs distributions in hidden Markov models
Yunxin Zhao, Xinhua Zhuang, Les E. Atlas, Lars Anderson
CVGIP Graph. Model. Image Process.1
1991 An HMM based speaker-independent continuous speech recognition system with experiments on the TIMIT database
abstract
The authors recently designed and implemented a large-vocabulary, speaker-independent, continuous speech recognition system. The system is based on hidden Markov modeling (HMM) of phoneme-sized acoustic units using continuous mixture Gaussian densities. The main structure of the system is outlined with a focus on a method of generating mixture Gaussian density models through a merging procedure whose efficiency was recently improved significantly. The system has been evaluated on the TIMIT database on a task of vocabulary size 853 and various grammar perplexities. The word accuracies are 92.2%, 84.9%, and 60.1% for the test set perplexities of 25, 106, and 853 (no grammar), respectively.>
Yunxin Zhao, Hisashi Wakita, Xinhua Zhuang
ICASSP1
1991 Morphological structuring image decomposition
abstract
Two major structuring element decomposition techniques are compared, and the superiority of the two pixel decomposition techniques over the cellular decomposition technique is shown in terms of the number of pipeline stages. As for the general structuring function decomposition, to the authors' knowledge, there is no efficient algorithm that has been found. The difficulty can be overcome by using an adequate representation of the general gray-scale structuring function. Representing a gray-scale image as a specific 3-D set, i.e. an umbra, makes it easier to shift all morphological theorems from the binary domain to the gray-scale domain; however, a direct umbra representation is not appropriate for the general gray-scale structuring function decomposition. The authors also provide a morphologically realizable representation and decomposition for the general gray-scale structuring function and show recursive algorithms which are pipelineable for efficiently performing gray-scale morphological operations.>
Xinhua Zhuang, Yunxin Zhao, Ming-Yee Chiu
ICASSP2
1991 Generate word transcription dictionary from sentence utterances and evaluate its effect on speaker-independent continuous speech recognition
Yunxin Zhao, Hisashi Wakita, Xinhua Zhuang
EUROSPEECH1
1991 A neural net algorithm for multidimensional maximum entropy spectrum estimation
Xinhua Zhuang, Yunxin Zhao, Thomas S. Huang
Neural Networks2
1990 An analog neural net performing multidimensional maximum entropy spectral estimation
abstract
A general algorithm is presented for computationally efficient multidimensional maximum entropy (ME) spectral estimation. The estimator is equivalent to an analog neural net that is governed by an energy function that measures the degree of constraint satisfaction, i.e., correlation-matching property. The multidimensional ME (MDME) spectral-estimation problem is defined. ME spectral estimation is formulated as an initial-value problem. The MDME spectral estimators, or algorithms, for solving the initial-value problem are developed. The neural net algorithm is derived, and simulated experiments with 1D or 2D signals are conducted. In each, the assumed true spectrum is given and autocorrelations at a number of horizontally and vertically equally spaced lags are calculated from the Fourier transform of the true spectrum. The overall results indicate very good performance in estimating ME spectra, even in cases where there are few autocorrelation measurements.>
Xinhua Zhuang, Hyonam Joo, Seho Oh, Yunxin Zhao, Thomas S. Huang
ICASSP4
1990 Experiments with a speaker-independent continuous speech recognition system on the timit database
Yunxin Zhao, Hisashi Wakita
ICSLP1
1988 From depth and optical flow to rigid body motion
abstract
The authors develop an algorithm to determine uniquely the rigid body motion from optical flow and depth, where the depth, however, does not involve any derivative information. Thus, the original assumptions made by D.H. Ballard and O.A. Kimball (1983) and by R.M. Haralick and X. Zhuang (1986) are relaxed. The proposed algorithm is appealing; in contrast to the existing linear optical flow-motion algorithms, it requires only three instead of eight optical flow image points.>
Xinhua Zhuang, Robert M. Haralick, Yunxin Zhao
CVPR3
1988 Application of the Gibbs distribution to hidden Markov modeling in isolated word recognition
abstract
A new method of formulating hidden Markov models (HMM) for isolated word recognition is presented. The authors model probabilities of hidden state sequences as Gibbs distributions (GDs) instead of the conventional products of transition probabilities. This formulation is based on the Hammersley-Clifford theorem which establishes the equivalence between Markov random fields (MRF) and GDs. The Markov chains in HMM are equivalent to one-dimensional, first order neighborhood MRFs. The observation sequences are modeled by the usual autoregressive Gaussian densities. The flexibility in the choice of energy functions in GDs makes it possible to use only a few parameters while maintaining a powerful model. The authors have developed a learning algorithm to estimate the parameters using maximum likelihood estimation and an algorithm to efficiently compute 1-D, first order neighborhood GDs using a lattice structure.>
Yunxin Zhao, Les E. Atlas, Xinhua Zhuang
ICASSP1