VLDB 2026 Research / reviewers in the wild / expert
Richard M. Stern
dblp:69/2347
· DBLP profile ↗
140ranked-venue papers
5as first author
12since 2021 · last 2027
0000-0003-0557-7282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 128 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 72 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Removing artifacts and residual noise of enhanced speech using GAN-based estimation
Simon Estrada, Rodrigo Mahú, Simon Repolt, Richard M. Stern, Néstor Becerra Yoma |
Comput. Speech Lang. | 4 |
| 2026 | The Value of Corrective Feedback in the Online Active Learning ParadigmabstractOnline Active Learning (OAL) is a powerful tool for classifying evolving data streams using limited annotations from a human operator who is a domain expert. The objective of the OAL learning paradigm is to minimize jointly the classification error rate and the annotation cost across the data stream by posing periodic Active Learning (AL) queries. In this paper, this objective is extended to include identification of classifier errors by the expert during the typical workflow. To this end, Corrective Feedback (CF) is introduced as a second channel of interaction between the expert and the learning algorithm, complementary to the AL channel, that allows the algorithm to obtain additional training labels without disrupting the expert's workflow. Online Active Learning with Corrective Feedback (OAL-CF) is formally defined as a paradigm, and its efficacy is proven through experimental application to two binary classification tasks, Spoken Language Verification and Voice-Type Discrimination. Finally, the effects of adding CF to the OAL paradigm are analyzed in terms of classification performance, annotation cost, trends over time, and class balance of the collected training data. Overall, the addition of CF results in a 53% relative reduction in cost compared to OAL without CF. Mark Lindsey, Francis Kubala, Richard M. Stern |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Iterative Feedback in the Online Active Learning ParadigmabstractOnline Active Learning with Corrective Feedback (OALCF) is a new machine learning paradigm that learns to detect events of interest in streaming audio by adapting to feedback from an operational user, who is a domain expert. The machine learns from each batch of the data stream by posing active learning queries to the expert and updating its model before making predictions. The expert reviews the predictions and indicates which of them are incorrect. This feedback is used to update the model for the next batch. In this paper, we introduce iterative feedback, where active learning and corrective feedback are performed multiple times per batch. We validate this approach on a large Spoken Language Verification task with 35 low-prevalence languages. The evaluation metric used (IMLM) accounts for the total cost of the method (i.e., error and feedback cost combined). The iterative algorithm achieves a 47.7% relative reduction in total cost compared to the original OAL-CF algorithm. Mark Lindsey, Francis Kubala, Richard M. Stern |
ASRU | 3 |
| 2025 | A Unified Metric for Simultaneous Evaluation of Error Rate and Annotation CostabstractPattern classification systems have traditionally been trained using a set of labeled training data and subsequently evaluated using different testing data. The cost of labeling the training data is typically substantial. Online Human-In-The-Loop (HITL) algorithms present an alternate approach that enables useful classification for many real-world applications using much less labeled data. These classifiers begin with a very small amount of training data and iteratively improve their performance by labeling a selected small number of utterances manually. Unfortunately, there is no unified evaluation metric that considers both classifier performance and annotation cost, which makes it difficult to evaluate these algorithms objectively. Furthermore, the lack of such a metric restricts the evaluation of online learning algorithms to prequential evaluation (before the classifier is adapted to the newly-labeled evaluation data), which does not realistically reflect the algorithm’s ability to adapt to the data stream in real time. This paper introduces the Interactive Machine Learning Metric (IMLM), a new unified evaluation metric that makes the combination of performance and annotation cost for binary classification tasks far less arbitrary. This metric is well suited for the evaluation of online HITL algorithms and also allows for fair comparison of different algorithms after adapting to the evaluation data. The value and appropriateness of IMLM is demonstrated by evaluating a series of Online Active Learning algorithms on a Spoken Language Verification task. Mark Lindsey, Francis Kubala, Richard M. Stern |
ICASSP | 3 |
| 2023 | Reducing the Cost of Spoof Detection Labeling using Mixed-Strategy Active Learning and Pretrained ModelsabstractActive learning is a powerful method for reducing the amount of labeled training data needed for a machine learning model to learn a task without degrading performance. This is accomplished by iteratively selecting the most informative samples from an unlabeled dataset to be labeled by an oracle (i.e., a human annotator) using an active learning sampling strategy. Pretrained models have been used in recent years as frontends for active learning neural networks to increase efficiency. This work applies active learning with pretrained models to the spoof detection task with the following two goals: 1) the identification of which pretrained speech models and active learning strategies are most effective for the spoof detection task, and 2) the development of an active learning method that selects the optimal sampling strategy from a list of available strategies at each step of the active learning process. This mixed strategy is shown to outperform all individual strategies for the task. Mark Lindsey, Nathaniel R. Robinson, Francis Kubala, Richard M. Stern |
ASRU | 4 |
| 2023 | Unsupervised Voice Type Discrimination Score Adaptation Using X-Vector ClustersabstractVoice type discrimination (VTD) is the task of automatically detecting speech produced in the same room as a recording device ("live speech") among other speech and non-speech noises, such as traffic noises or radio broadcasts ("distractor audio"). Existing work has described methods for performing the VTD task. This paper presents a method for adapting the output of these existing methods in an unsupervised manner via x-vector clustering and correlation. This adaptation method can be applied to the output of any VTD algorithm, requires no additional training data, and has been shown to yield a relative decrease in decision cost function (DCF) score of up to 47% on a standardized database collected for the task. Mark Lindsey, Tyler Vuong, Richard M. Stern |
ICASSP | 3 |
| 2023 | Respiratory distress estimation in human-robot interaction scenario
Eduardo Alvarado, Nicolás Grágeda, Alejandro Luzanto, Rodrigo Mahú, Jorge Wuth, Laura Mendoza, Richard M. Stern, Néstor Becerra Yoma |
INTERSPEECH | 7 |
| 2022 | Improved Modulation-Domain Loss for Neural-Network-based Speech Enhancement
Tyler Vuong, Richard M. Stern |
INTERSPEECH | 2 |
| 2022 | Investigating the Important Temporal Modulations for Deep-Learning-Based Speech Activity DetectionabstractWe describe a learnable modulation spectrogram feature for speech activity detection (SAD). Modulation features capture the temporal dynamics of each frequency subband. We compute learnable modulation spectrogram features by first calculating the log-mel spectrogram. Next, we filter each frequency subband with a bandpass filter that contains a learnable center frequency. The resulting SAD system was evaluated on the Fearless Steps Phase-04 SAD challenge. Experimental results showed that temporal modulations around the 4–6 Hz range are crucial for deep-learning-based SAD. These experimental results align with previous studies that found slow temporal modulation to be most important for speech-processing tasks and speech intelligibility. Additionally, we found that the learnable modulation spectrogram feature outperforms both the standard log-mel and fixed modulation spectrogram features on the Fearless Steps Phase-04 SAD test set. Tyler Vuong, Nikhil Madaan, Rohan Panda, Richard M. Stern |
SLT | 4 |
| 2021 | A Modulation-Domain Loss for Neural-Network-Based Real-Time Speech EnhancementabstractWe describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then used to calculate a weighted mean-squared error (MSE) in the modulation domain for training a speech enhancement system. Experiments showed that adding the modulation-domain MSE to the MSE in the spectro-temporal domain substantially improved the objective prediction of speech quality and intelligibility for real-time speech enhancement systems without incurring additional computation during inference. Tyler Vuong, Yangyang Xia, Richard M. Stern |
ICASSP | 3 |
| 2021 | The Application of Learnable STRF Kernels to the 2021 Fearless Steps Phase-03 SAD Challenge
Tyler Vuong, Yangyang Xia, Richard M. Stern |
Interspeech | 3 |
| 2021 | Temporal Context in Speech Emotion Recognition
Yangyang Xia, Alexander I. Rudnicky, Richard M. Stern |
Interspeech | 4 |
| 2020 | Learnable Spectro-Temporal Receptive Fields for Robust Voice Type DiscriminationabstractVoice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio that were played back such as traffic noise and television broadcasts ("Distractor Audio"). In this work, we propose a deep-learning-based VTD system that features an initial layer of learnable spectro-temporal receptive fields (STRFs). Our approach is also shown to provide very strong performance on a similar spoofing detection task in the ASVspoof 2019 challenge. We evaluate our approach on a new standardized VTD database that was collected to support research in this area. In particular, we study the effect of using learnable STRFs compared to static STRFs or unconstrained kernels. We also show that our system consistently improves a competitive baseline system across a wide range of signal-to-noise ratios on spoofing detection in the presence of VTD distractor noise. Tyler Vuong, Yangyang Xia, Richard M. Stern |
INTERSPEECH | 3 |
| 2019 | Robust Recognition of Reverberant and Noisy Speech Using Coherence-based ProcessingabstractThis paper describes a combination of techniques for improving speech recognition accuracy using two microphones in reverberant and noisy environments. These techniques include both monaural and binaural processing. The first stage is monaural precedence-based processing that enhances the onsets of the incoming speech signal, and hence suppresses later components that are more affected by reverberation. Onset enhancement has been shown to be useful to the human auditory system in separating the direct field from the reverberant field in reverberant environments. The second stage applies emphasis or suppression to signal components based on an estimation of the inter-microphone coherence of the incoming speech signal. Specifically, portions of the speech signal that are less coherent are suppressed, which is intended to reduce the contributions of components that are dominated by diffuse noise or high degrees of reverberation in the input signal. A combination of these techniques is shown to lead to significant improvements in speech recognition accuracy. A DNN-based automatic speech recognition system was used to evaluate the techniques described in this study over a range of reverberation times and signal-to-interferer ratios. Anjali Menon, Chanwoo Kim 0001, Richard M. Stern |
ICASSP | 3 |
| 2018 | Sound Source Separation Using Phase Difference and Reliable Mask Selection SelectionabstractIn this paper, we present an algorithm called Reliable Mask Selection-Phase Difference Channel Weighting (RMS-PDCW) which selects the target source masked by a noise source using the Angle of Arrival (AoA) information calculated using the phase difference information. The RMS-PDCW algorithm selects masks to apply using the information about the localized sound source and the onset detection of speech. We demonstrate that this algorithm shows relatively 5.3 percent improvement over the baseline acoustic model, which was multistyle-trained using 22 million utterances on the simulated test set consisting of real-world and interfering-speaker noise with reverberation time distribution between 0 ms and 900 ms and SNR distribution between 0 dB up to clean. Chanwoo Kim 0001, Anjali Menon, Michiel Bacchiani, Richard M. Stern |
ICASSP | 4 |
| 2018 | A Priori SNR Estimation Based on a Recurrent Neural Network for Robust Speech Enhancement
Yangyang Xia, Richard M. Stern |
INTERSPEECH | 2 |
| 2018 | A Comparative Study of Spatial Speech Separation Techniques to Improve Speech Recognition
Xinhui Zhou, Chiman Kwan, Bulent Ayhan, Chanwoo Kim 0001, Kshitiz Kumar, Richard M. Stern |
ISNN | 6 |
| 2017 | Binaural processing for robust recognition of degraded speechabstractThis paper discusses a new combination of techniques that help in improving the accuracy of speech recognition in adverse conditions using two microphones. Classic approaches toward binaural speech processing use some form of cross-correlation over time across the two sensors to effectively isolate target speech from interferers. Several additional techniques using temporal and spatial masking have been proposed in the past to improve recognition accuracy in the presence of reverberation and interfering talkers. In this paper, we consider the use of cross-correlation across frequency over some limited range of frequency channels in addition to the existing methods of monaural and binaural processing. This has the effect of locating and reinforcing coincident peaks across frequency over the representation of binaural interaction and provides local smoothing over the specified range of frequencies. Combined with the temporal and spatial masking techniques mentioned above, this leads to significant improvements in binaural speech recognition. Anjali Menon, Chanwoo Kim 0001, Umpei Kurokawa, Richard M. Stern |
ASRU | 4 |
| 2017 | Robust Speech Recognition Based on Binaural Auditory Processing
Anjali Menon, Chanwoo Kim 0001, Richard M. Stern |
INTERSPEECH | 3 |
| 2017 | Robustness Over Time-Varying Channels in DNN-HMM ASR Based Human-Robot Interaction
José Novoa, Jorge Wuth, Juan Pablo Escudero, Josué Fredes, Rodrigo Mahú, Richard M. Stern, Néstor Becerra Yoma |
INTERSPEECH | 6 |
| 2017 | Locally Normalized Filter Banks Applied to Deep Neural-Network-Based Robust Speech RecognitionabstractThis letter describes modifications to locally normalized filter banks (LNFB), which substantially improve their performance on the Aurora-4 robust speech recognition task using a Deep Neural Network-Hidden Markov Model (DNN-HMM)-based speech recognition system. The modified coefficients, referred to as LNFB features, are a filter-bank version of locally normalized cepstral coefficients (LNCC), which have been described previously. The ability of the LNFB features is enhanced through the use of newly proposed dynamic versions of them, which are developed using an approach that differs somewhat from the traditional development of delta and delta-delta features. Further enhancements are obtained through the use of mean normalization and mean-variance normalization, which is evaluated both on a per-speaker and a per-utterance basis. The best performing feature combination (typically LNFB combined with LNFB delta and delta-delta features and mean-variance normalization) provides an average relative reduction in word error rate of 11.4% and 9.4%, respectively, compared to comparable features derived from Mel filter banks when clean and multinoise training are used for the Aurora-4 evaluation. The results presented here suggest that the proposed technique is more robust to channel mismatches between training and testing data than MFCC-derived features and is more effective in dealing with channel diversity. Josué Fredes, José Novoa, Simon King 0001, Richard M. Stern, Néstor Becerra Yoma |
IEEE Signal Process. Lett. | 4 |
| 2017 | Synchrony-Based Feature Extraction for Robust Automatic Speech RecognitionabstractThis letter discusses the application of models of temporal patterns of auditory-nerve firings to enhance robustness of automatic speech recognition systems. Most conventional feature extraction schemes (such as mel-frequency cepstral coefficients and perceptual linear processing coefficients) are based on short-time energy in each frequency band, and the temporal patterns of auditory-nerve activity are discarded. We compare the impact on speech recognition accuracy of several types of feature extraction schemes based on the putative synchrony of auditory-nerve activity, including feature extraction based on a modified version of the generalized synchrony detector proposed by Seneff, and a modified version of the averaged localized synchrony response proposed by Young and Sachs. It was found that the use of features based on auditory-nerve synchrony can indeed improve speech recognition accuracy in the presence of additive noise based on experiments using multiple standard speech databases. Recognition accuracy obtained using the synchrony-based features is further increased if some form of noise removal is applied to the signal before the synchrony measure is estimated. Signal processing for noise removal based on the noise suppression that is a part of PNCC feature extraction is more effective toward this end than conventional spectral subtraction. Fernando de-la-Calle-Silos, Richard M. Stern |
IEEE Signal Process. Lett. | 2 |
| 2016 | Fusion Strategies for Robust Speech Recognition and Keyword Spotting for Channel- and Noise-Degraded Speech
Vikramjit Mitra, Julien van Hout, Wen Wang 0001, Chris Bartels, Horacio Franco, Dimitra Vergyri, Abeer Alwan, Adam Janin, John H. L. Hansen, Richard M. Stern, Abhijeet Sangwan, Nelson Morgan |
INTERSPEECH | 10 |
| 2016 | The Use of Locally Normalized Cepstral Coefficients (LNCC) to Improve Speaker Recognition Accuracy in Highly Reverberant Rooms
Víctor Poblete, Juan Pablo Escudero, Josué Fredes, José Novoa, Richard M. Stern, Simon King 0001, Néstor Becerra Yoma |
INTERSPEECH | 5 |
| 2016 | A Subband-Based Stationary-Component Suppression Method Using Harmonics and Power Ratio for Reverberant Speech RecognitionabstractThis letter describes a preprocessing method called subband-based stationary-component suppression method using harmonics and power ratio (SHARP) processing for reverberant speech recognition. SHARP processing extends a previous algorithm called Suppression of Slowly varying components and the Falling edge (SSF), which suppresses the steady-state portions of subband spectral envelopes. The SSF algorithm tends to over-subtract these envelopes in highly reverberant environments when there are high levels of power in previous analysis frames. The proposed SHARP method prevents excessive suppression both by boosting the floor value using the harmonics in voiced speech segments and by inhibiting the subtraction for unvoiced speech by detecting frames in which power is concentrated in high-frequency channels. These modifications enable the SHARP algorithm to improve recognition accuracy by further reducing the mismatch between power contours of clean and reverberated speech. Experimental results indicate that the SHARP method provides better recognition accuracy in highly reverberant environments compared to the SSF algorithm. It is also shown that the performance of the SHARP method can be further improved by combining it with feature-space maximum likelihood linear regression (fMLLR). Byung Joon Cho, Haeyong Kwon, Ji-Won Cho, Chanwoo Kim 0001, Richard M. Stern, Hyung-Min Park |
IEEE Signal Process. Lett. | 5 |
| 2016 | Power-Normalized Cepstral Coefficients (PNCC) for Robust Speech RecognitionabstractThis paper presents a new feature extraction algorithm called power normalized Cepstral coefficients (PNCC) that is motivated by auditory processing. Major new features of PNCC processing include the use of a power-law nonlinearity that replaces the traditional log nonlinearity used in MFCC coefficients, a noise-suppression algorithm based on asymmetric filtering that suppresses background excitation, and a module that accomplishes temporal masking. We also propose the use of medium-time power analysis in which environmental parameters are estimated over a longer duration than is commonly used for speech, as well as frequency smoothing. Experimental results demonstrate that PNCC processing provides substantial improvements in recognition accuracy compared to MFCC and PLP processing for speech in the presence of various types of additive noise and in reverberant environments, with only slightly greater computational cost than conventional MFCC processing, and without degrading the recognition accuracy that is observed while training and testing using clean speech. PNCC processing also provides better recognition accuracy in noisy environments than techniques such as vector Taylor series (VTS) and the ETSI advanced front end (AFE) while requiring much less computation. We describe an implementation of PNCC using “online processing” that does not require future knowledge of the input. Chanwoo Kim 0001, Richard M. Stern |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Efficient audio declipping using regularized least squaresabstractWhile many recently proposed audio declipping algorithms are highly effective in their ability to restore clipped speech, the algorithms' computational complexities inhibit their use in many practical situations. Real-time or nearly real-time performance is impossible using a typical laptop computer, with some algorithms taking as long as 400 times the actual duration of the input to complete restoration. This paper introduces a novel declipping algorithm, referred to as Regularized Blind Amplitude Reconstruction, which is capable of restoring clipped audio at rates much faster than real time and at restoration qualities comparable to existing algorithms. The quality of declipping is evaluated in terms of automatic speech recognition performance on declipped speech, as well as the degree to which each declipping algorithm improves the audio's signal-to-noise ratio. Mark Harvilla, Richard M. Stern |
ICASSP | 2 |
| 2015 | Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial WorkshopabstractA group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which the estimator was trained. The paper describes the problem and summarizes approaches that were taken by the group1. Hynek Hermansky, Lukás Burget, Jordan Cohen, Emmanuel Dupoux, Naomi Feldman, John Godfrey, Sanjeev Khudanpur, Matthew Maciejewski, Sri Harish Reddy Mallidi, Anjali Menon, Tetsuji Ogawa, Vijayaditya Peddinti, Richard C. Rose, Richard M. Stern, Matthew Wiesner, Karel Veselý |
ICASSP | 14 |
| 2015 | Robustness to additive noise of locally-normalized cepstral coefficients in speaker verification
Josué Fredes, José Novoa, Víctor Poblete, Simon King 0001, Richard M. Stern, Néstor Becerra Yoma |
INTERSPEECH | 5 |
| 2015 | Robust parameter estimation for audio declipping in noise
Mark Harvilla, Richard M. Stern |
INTERSPEECH | 2 |
| 2015 | A perceptually-motivated low-complexity instantaneous linear channel normalization technique applied to speaker verification
Víctor Poblete, Felipe Espic, Simon King 0001, Richard M. Stern, Fernando Huenupán, Josué Fredes, Néstor Becerra Yoma |
Comput. Speech Lang. | 4 |
| 2014 | An analysis of binaural spectro-temporal masking as nonlinear beamformingabstractArray-based time-frequency masking algorithms are an important type of nonlinear array processing. In this paper we develop a model that characterizes the directional sensitivity of these algorithms in a fashion similar to commonly-used the beam patterns used to characterize linear array processing. Two alternative formulations are described, and it is shown that one of these formulations predicts signal distortion and processing gain in time-frequency masking accurately, as well as speech recognition accuracy afforded by time-frequency masking in the presence of additive interfering sources. Amir R. Moghimi, Richard M. Stern |
ICASSP | 2 |
| 2014 | Least squares signal declipping for robust speech recognitionabstractThis paper introduces a novel declipping algorithm based on constrained least-squares minimization. Digital speech signals are often sampled at 16 kHz and classic declipping algorithms fail to accurately reconstruct the signal at this sampling rate due to the scarcity of reliable samples after clipping. The Constrained Blind Amplitude Reconstruction algorithm interpolates missing data points such that the resulting function is smooth while ensuring the inferred data fall in a legitimate range. The inclusion of explicit constraints helps to guide an accurate interpolation. Evaluation of declipping performance is based on automatic speech recognition word error rate and Constrained Blind Amplitude Reconstruction is shown to outperform the current state-of-the-art declipping technology under a variety of conditions. Declipping performance in additive noise is also considered. Mark Harvilla, Richard M. Stern |
INTERSPEECH | 2 |
| 2014 | Robust speech recognition using temporal masking and thresholding algorithmabstractIn this paper, we present a new dereverberation algorithm called Temporal Masking and Thresholding (TMT) to enhance the temporal spectra of spectral features for robust speech recognition in reverberant environments. This algorithm is motivated by the precedence effect and temporal masking of human auditory perception. This work is an improvement of our previous dereverberation work called Suppression of Slowlyvarying components and the falling edge of the power envelope (SSF). The TMT algorithm uses a different mathematical model to characterize temporal masking and thresholding compared to the model that had been used to characterize the SSF algorithm. Specifically, the nonlinear highpass filtering used in the SSF algorithm has been replaced by a masking mechanism based on a combination of peak detection and dynamic thresholding. Speech recognition results show that the TMT algorithm provides superior recognition accuracy compared to other algorithms such as LTLSS, VTS, or SSF in reverberant environments. Chanwoo Kim 0001, Kean K. Chin, Michiel Bacchiani, Richard M. Stern |
INTERSPEECH | 4 |
| 2014 | Post-masking: a hybrid approach to array processing for speech recognition
Amir R. Moghimi, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 3 |
| 2014 | Robust speech recognition in reverberant environments using subband-based steady-state monaural and binaural suppression
Hyung-Min Park, Matthew Maciejewski, Chanwoo Kim 0001, Richard M. Stern |
INTERSPEECH | 4 |
| 2014 | Optimization of the parameters characterizing sigmoidal rate-level functions based on acoustic features
Víctor Poblete, Néstor Becerra Yoma, Richard M. Stern |
Speech Commun. | 3 |
| 2013 | Optimization of sigmoidal rate-level function based on acoustic features
Víctor Poblete, Néstor Becerra Yoma, Richard M. Stern |
INTERSPEECH | 3 |
| 2013 | Perceptual Properties of Current Speech Recognition TechnologyabstractIn recent years, a number of feature extraction procedures for automatic speech recognition (ASR) systems have been based on models of human auditory processing, and one often hears arguments in favor of implementing knowledge of human auditory perception and cognition into machines for ASR. This paper takes a reverse route, and argues that the engineering techniques for automatic recognition of speech that are already in widespread use are often consistent with some well-known properties of the human auditory system. Hynek Hermansky, Jordan Cohen, Richard M. Stern |
Proc. IEEE | 3 |
| 2012 | Histogram-based subband powerwarping and spectral averaging for robust speech recognition under matched and multistyle trainingabstractThis paper describes a new algorithm that increases the robustness of speech recognition systems by matching the power histograms of the input in each frequency band to those obtained over clean training data, and then mixing together the processed and unprocessed spectra. Before calculating prototype histograms over the training data, the power signals in each channel are normalized by the local maximum and minimum of the channel. In contrast, histograms calculated over the testing data are normalized by the global maximum and minimum of the power spectrum. This mode of normalization leads to a significant reduction in noise. Following the histogram-based processing, it is shown that taking a weighted average between the processed and unprocessed power spectra contributes to further gains in recognition accuracy. Results are obtained for multiple speech recognition systems, noise types, and training conditions illustrating the broad utility of this approach. Mark Harvilla, Richard M. Stern |
ICASSP | 2 |
| 2012 | Two-microphone source separation algorithm based on statistical modeling of angle distributionsabstractIn this paper we present a novel two-microphone sound source separation algorithm, which selects speech from the target speaker while suppressing signals from interfering sources. In this algorithm, which is refered to as SMAD-CW, we first estimate the direction of sound sources for each time-frequency bin using phase differences in the spectral domain. For each frame we assume that the angle distribution is a mixture of two distributions, one from the target and the other from the dominant noise source. For each mixture component we use the von Mises distribution, which is a close approximation to the wrapped normal distribution. The expectation-maximization (EM) algorithm is employed to obtain parameters of this mixture distribution. Using this statistical model, we perform maximum a posteriori (MAP) hypothesis testing in order to obtain appropriate binary masks. We demonstrate that the algorithm described in this paper provides speech recognition accuracy that is significantly better than that obtained using conventional approaches. Chanwoo Kim 0001, Charbel El Khawand, Richard M. Stern |
ICASSP | 3 |
| 2012 | Power-Normalized Cepstral Coefficients (PNCC) for robust speech recognitionabstractThis paper presents a new feature extraction algorithm called Power Normalized Cepstral Coefficients (PNCC) that is based on auditory processing. Major new features of PNCC processing include the use of a power-law nonlinearity that replaces the traditional log nonlinearity used in MFCC coefficients, a noise-suppression algorithm based on asymmetric filtering that suppress background excitation, and a module that accomplishes temporal masking. We also propose the use of medium-time power analysis, in which environmental parameters are estimated over a longer duration than is commonly used for speech, as well as frequency smoothing. Experimental results demonstrate that PNCC processing provides substantial improvements in recognition accuracy compared to MFCC and PLP processing for speech in the presence of various types of additive noise and in reverberant environments, with only slightly greater computational cost than conventional MFCC processing, and without degrading the recognition accuracy that is observed while training and testing using clean speech. PNCC processing also provides better recognition accuracy in noisy environments than techniques such as Vector Taylor Series (VTS) and the ETSI Advanced Front End (AFE) while requiring much less computation. We describe an implementation of PNCC using “on-line processing” that does not require future knowledge of the input. Chanwoo Kim 0001, Richard M. Stern |
ICASSP | 2 |
| 2012 | Optimization of the DET curve in speaker verificationabstractSpeaker verification systems are, in essence, statistical pattern detectors which can trade off false rejections for false acceptances. Any operating point characterized by a specific tradeoff between false rejections and false acceptances may be chosen. Training paradigms in speaker verification systems however either learn the parameters of the classifier employed without actually considering this tradeoff, or optimize the parameters for a particular operating point exemplified by the ratio of positive and negative training instances supplied. In this paper we investigate the optimization of training paradigms to explicitly consider the tradeoff between false rejections and false acceptances, by minimizing the area under the curve of the detection error tradeoff curve. To optimize the parameters, we explicitly minimize a mathematical characterization of the area under the detection error tradeoff curve, through generalized probabilistic descent. Experiments on the NIST 2008 database show that for clean signals the proposed optimization approach is at least as effective as conventional learning. On noisy data, verification performance obtained with the proposed approach is considerably better than that obtained with conventional learning methods. L. Paola García-Perera, Juan A. Nolazco-Flores, Bhiksha Raj, Richard M. Stern |
SLT | 4 |
| 2012 | Learning-Based Auditory Encoding for Robust Speech RecognitionabstractThis paper describes an approach to the optimization of the nonlinear component of a physiologically motivated feature extraction system for automatic speech recognition. Most computational models of the peripheral auditory system include a sigmoidal nonlinear function that relates the log of signal intensity to output level, which we represent by a set of frequency dependent logistic functions. The parameters of these rate-level functions are estimated to maximize the a posteriori probability of the correct class in training data. The performance of this approach was verified by the results of a series of experiments conducted with the CMU S phinx-III speech recognition system on the DARPA Resource Management, Wall Street Journal databases, and on the AURORA 2 database. In general, it was shown that feature extraction that incorporates the learned rate-nonlinearity, combined with a complementary loudness compensation function, results in better recognition accuracy in the presence of background noise than traditional MFCC feature extraction without the optimized nonlinearity when the system is trained on clean speech and tested in noise. We also describe the use of lattice structure that constraints the training process, enabling training with much more complicated acoustic models. Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Binaural sound source separation motivated by auditory processingabstractIn this paper we present a new method of signal processing for robust speech recognition using two microphones. The method, loosely based on the human binaural hearing system, consists of passing the speech signals detected by two microphones through bandpass filtering. We develop a spatial masking function based on normalized cross-correlation, which provides rejection of off-axis interfering signals. To obtain improvements in reverberant environments, a temporal masking component, which is closely related to our previously-described de-reverberation technique known as SSF. We demonstrate that this approach provides substantially better recognition accuracy than conventional binaural sound-source separation algorithms. Chanwoo Kim 0001, Kshitiz Kumar, Richard M. Stern |
ICASSP | 3 |
| 2011 | Delta-spectral cepstral coefficients for robust speech recognitionabstractAlmost all current automatic speech recognition (ASR) systems conventionally append delta and double-delta cepstral features to static cepstral features. In this work we describe a modified feature-extraction procedure in which the time-difference operation is performed in the spectral domain, rather than the cepstral domain as is generally presently done. We argue that this approach based on "delta-spectral" features is needed because even though delta-cepstral features capture dynamic speech information and generally greatly improve ASR recognition accuracy, they are not robust to noise and reverberation. We support the validity of the delta-spectral approach both with observations about the modulation spectrum of speech and noise, and with objective experiments that document the benefit that the delta-spectral approach brings to a variety of currently popular feature extraction algorithms. We found that the use of delta-spectral features, rather than the more traditional delta-cepstral features, improves the effective SNR by between 5 and 8 dB for background music and white noise, and recognition accuracy in reverberant environments is improved as well. Kshitiz Kumar, Chanwoo Kim 0001, Richard M. Stern |
ICASSP | 3 |
| 2011 | An iterative least-squares technique for dereverberationabstractSome recent dereverberation approaches that have been effective for automatic speech recognition (ASR) applications, model reverberation as a linear convolution operation in the spectral domain, and derive a factorization to decompose spectra of reverberated speech in to those of clean speech and room-response filter. Typically, a general non-negative matrix factorization (NMF) framework is employed for this. In this work we present an alternative to NMF and propose an iterative least-squares deconvolution technique for spectral factorization. We propose an efficient algorithm for this and experimentally demonstrate it's effectiveness in improving ASR performance. The new method results in 40-50% relative reduction in word error rates over standard baselines on artificially reverberated speech. Kshitiz Kumar, Bhiksha Raj, Rita Singh, Richard M. Stern |
ICASSP | 4 |
| 2011 | Gammatone sub-band magnitude-domain dereverberation for ASRabstractWe present an algorithm for dereverberation of speech signals for automatic speech recognition (ASR) applications. Often ASR systems are presented with speech that has been recorded in environments that include noise and reverberation. The performance of ASR systems degrades with increasing levels of noise and reverberation. While many algorithms have been proposed for robust ASR in noisy environments, reverberation is still a challenging problem. In this paper, we present ' an approach for dereverberation that models reverberation as a convolution operation in the speech spectral domain. Using a least-squares error criterion we decompose reverberated spectra into clean spectra convolved with a filter. We incorporate non-negativity and sparsity of the speech spectra as constraints within a non-negative matrix factorization (NMF) frame work to achieve the decomposition. In ASR experiments where the system is trained with unreverberated and reverberated speech, we show that the proposed approach can provide upto 40% and 19% relative reduction respectively in performance. Kshitiz Kumar, Rita Singh, Bhiksha Raj, Richard M. Stern |
ICASSP | 4 |
| 2011 | Mask classification for missing-feature reconstruction for robust speech recognition in unknown background noise
Wooil Kim, Richard M. Stern |
Speech Commun. | 2 |
| 2010 | A hybrid physical and statistical dynamic articulatory framework incorporating analysis-by-synthesis for improved phone classificationabstractIn this paper, we present a dynamic articulatory model for phone classification. The model integrates real articulatory information derived from ElectroMagnetic Articulograph (EMA) data into its inner states. It maps from the articulatory space to the acoustic one using an adapted vocal tract model for each speaker and a physiologically-motivated articulatory synthesis approach. We apply the analysis-by-synthesis paradigm in a statistical fashion. We first present a fast approach for deriving analysis-by-synthesis distortion features. Next, the distortion between the speech synthesized from the articulatory states and the incoming speech signal is used to compute the output observation probabilities of the Hidden Markov Model (HMM) used for classification. Experiments with the novel framework show improvements over baseline in phone classification accuracy. Ziad Al Bawab, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2010 | Learning-based auditory encoding for robust speech recognitionabstractThis paper describes ways of speeding up the optimization process for learning physiologically-motivated components of a feature computation module directly from data. During training, word lattices generated by the speech decoder and conjugate gradient descent were included to train the parameters of logistic functions in a fashion that maximizes the a posteriori probability of the correct class in the training data. These functions represent the rate-level nonlinearities found in most mammalian auditory systems. Experiments conducted using the CMU SPHINX-III system on the DARPA Resource Management and Wall Street Journal tasks show that the use of discriminative training to estimate the shape of the rate-level nonlinearity provides better recognition accuracy in the presence of background noise than traditional procedures which do not employ learning. More importantly, the inclusion of conjugate gradient descent optimization and a word lattice to reduce the number of hypotheses considered greatly increases the training speed, which makes training with much more complicated models possible. Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2010 | Feature extraction for robust speech recognition based on maximizing the sharpness of the power distribution and on power flooringabstractThis paper presents a new robust feature extraction algorithm based on a modified approach to power bias subtraction combined with applying a threshold to the power spectral density. Power bias level is selected as a level above which the signal power distribution is sharpest. The sharpness is measured using the ratio of arithmetic mean to the geometric mean of medium-duration power. When subtracting this bias level, power flooring is applied to enhance robustness. These new ideas are employed to enhance our recently introduced feature extraction algorithm PNCC (Power Normalized Cepstral Coefficient). While simpler than our previous PNCC, experimental results show that this new PNCC is showing better performance than our previous implementation. Chanwoo Kim 0001, Richard M. Stern |
ICASSP | 2 |
| 2010 | Maximum-likelihood-based cepstral inverse filtering for blind speech dereverberationabstractCurrent state-of-the-art speech recognition systems work quite well in controlled environments but their performance degrades severely in realistic acoustical conditions in reverberant environments. In this paper we build on the recent developments that represent reverberation in the cepstral feature domain as a filtering operation and we formulate a maximum likelihood objective to obtain an inverse reverberation filter. We show analytically that the optimal inverse filter can be approximately obtained under certain assumptions about the corresponding clean speech signal. We demonstrate that our approach reduces the relative gap in word error rate by 30 percent in large as well as small reverberation times. Kshitiz Kumar, Richard M. Stern |
ICASSP | 2 |
| 2010 | Nonlinear enhancement of onset for robust speech recognitionabstractIn this paper we present a novel algorithm called Suppression of Slowly-varying components and the Falling edge of the power envelope (SSF) to enhance spectral features for robust speech recognition, especially in reverberant environments. This algorithm is motivated by the precedence effect and by the modulation frequency characteristics of the human auditory system. We describe two slightly different types of processing that differ in whether or not the falling edges of power trajectories are suppressed using a lowpassed power envelope signal. The SSF algorithms can be implemented for online processing. Speech recognition results show that this algorithm provides especially good robustness in reverberant environments. 1 Chanwoo Kim 0001, Richard M. Stern |
INTERSPEECH | 2 |
| 2010 | Automatic selection of thresholds for signal separation algorithms based on interaural delayabstractIn this paper we describe a system that separates signals by comparing the interaural time delays (ITDs) of their timefrequency components to a fixed threshold ITD. While in previous algorithms the fixed threshold ITD had been obtained empirically from training data in a specific environment, in real environments the characteristics that affect the optimal value of this threshold are unknown and possibly time varying. If these configurations are different from the environment under which the ITD threshold had been pre-computed, the performance of the source separation system is degraded. In this paper, we present an algorithm which chooses a threshold ITD that minimizes the cross-correlation of the target and interfering signals, after a compressive nonlinearity. We demonstrate that the algorithm described in this paper provides speech recognition accuracy that is much more robust to changes in environment than would be obtained using a fixed threshold ITD. Index Terms: Robust speech recognition, speech enhancement, signal separation, time delay analysis, phase difference analysis, cross correlation 1. Chanwoo Kim 0001, Richard M. Stern, Kiwan Eom |
INTERSPEECH | 2 |
| 2009 | Robust speech recognition using a Small Power Boosting algorithmabstractIn this paper, we present a noise robustness algorithm called small power boosting (SPB). We observe that in the spectral domain, time-frequency bins with smaller power are more affected by additive noise. The conventional way of handling this problem is estimating the noise from the test utterance and doing normalization or subtraction. In our work, in contrast, we intentionally boost the power of time-frequency bins with small energy for both the training and testing datasets. Since time-frequency bins with small power no longer exist after this power boosting, the spectral distortion between the clean and corrupt test sets becomes reduced. This type of small power boosting is also highly related to physiological nonlinearity. We observe that when small power boosting is done, suitable weighting smoothing becomes highly important. Our experimental results indicate that this simple idea is very helpful for very difficult noisy environments such as corruption by background music. Chanwoo Kim 0001, Kshitiz Kumar, Richard M. Stern |
ASRU | 3 |
| 2009 | Power function-based power distribution normalization algorithm for robust speech recognitionabstractA novel algorithm that normalizes the distribution of spectral power coefficients is described in this paper. The algorithm, called power-function-based power distribution (PPDN) is based on the observation that the ratio of arithmetic mean to geometric mean changes as speech is corrupted by noise, and a parametric power function is used to equalize this ratio. We also observe that a longer ¿medium-duration¿ observation window (of approximately 100 ms) is better suited for parameter estimation for noise compensation than the briefer window typically used for automatic speech recognition. We also describe the implementation of an online version of PPDN based on exponentially weighted temporal averaging. Experimental results shows that PPDN provides comparable or slightly better results than state of- the-art algorithms such as vector Taylor series for speech recognition while requiring much less computation. Hence, the algorithm is suitable for both real-time speech communication or as a real-time preprocessing stage for speech recognition systems. Chanwoo Kim 0001, Richard M. Stern |
ASRU | 2 |
| 2009 | Minimum variance modulation filter for robust speech recognitionabstractThis paper describes a way of designing modulation filter by data driven analysis which improves the performance of automatic speech recognition systems that operate in real environments. The filter for each nonlinear channel output is obtained by a constrained optimization process which jointly minimizes the environmental distortion as well as the distortion caused by the filter itself. Recognition accuracy is measured using the CMU SPHINX-III speech recognition system, and the DARPA resource management and Wall Street Journal speech corpus for training and testing. It is shown that feature extraction followed by modulation filtering provides better performance than traditional MFCC processing under different types of background noise and reverberation. Yu-Hsiang Bosco Chiu, Richard M. Stern |
ICASSP | 2 |
| 2009 | Deriving vocal tract shapes from electromagnetic articulograph data via geometric adaptation and matchingabstractIn this paper, we present our efforts towards deriving vocal tract shapes from ElectroMagnetic Articulograph data (EMA) via geometric adaptation and matching. We describe a novel approach for adapting Maeda’s geometric model of the vocal tract to one speaker in the MOCHA database. We show how we can rely solely on the EMA data for adaptation. We present our search technique for the vocal tract shapes that best fit the given EMA data. We then describe our approach of synthesizing speech from these shapes. Results on Mel-cepstral distortion reflect improvement in synthesis over the approach we used before without adaptation. Index Terms: MOCHA EMA data, Maeda Model, vocal tract adaptation, articulatory model fitting, articulatory synthesis Ziad Al Bawab, Lorenzo Turicchia, Richard M. Stern, Bhiksha Raj |
INTERSPEECH | 3 |
| 2009 | Unsupervised training scheme with non-stereo data for empirical feature vector compensation
Luis Buera, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Richard M. Stern |
INTERSPEECH | 5 |
| 2009 | Towards fusion of feature extraction and acoustic model training: a top down process for robust speech recognitionabstractThis paper presents a strategy to learn physiologicallymotivated components in a feature computation module discriminatively, directly from data, in a manner that is inspired by the presence of efferent processes in the human auditory system. In our model a set of logistic functions which represent the rate-level nonlinearities found in most mammal hearing system are put in as part of the feature extraction process. The parameters of these rate-level functions are estimated to maximize the a posteriori probability of the correct class in the training data. The estimated feature computation is observed to be robust against environmental noise. Experiments conducted with the CMU Sphinx-III on the DARPA Resource Management task show that the discriminatively estimated rate-nonlinearity results in better performance in the presence of background noise than traditional procedures which separate the feature extraction and model training into two distinct parts without feed back from the latter to the former. Index Terms: automatic speech recognition, discriminative training, auditory model, data analysis Yu-Hsiang Bosco Chiu, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 3 |
| 2009 | Speaker segmentation and clustering for simultaneously presented speechabstractThis paper proposes a new scheme used to segment and cluster speech segments on an unsupervised basis in cases where multiple speakers are presented simultaneously at different SNRs. The new elements in our work are in the development of new feature for segmenting and clustering simultaneously-presented speech, the procedure for identifying a candidate set of possible speaker-change points, and the use of pair-wise cross-segment distance distributions to cluster segments by speaker. The proposed system is evaluated in terms of the F measure that is obtained. The system is compared to a baseline system that uses MFCC for acoustic features, the Bayesian Information Criterion (BIC) for detecting speaker-change points, and the Kullback-Leibler distance for clustering the segments. Experimental indicate that the new system consistently provides better performance than the baseline system with very small computational cost. 1 Index Terms: speech segmentation, speaker clustering, feature extraction Lingyun Gu, Richard M. Stern |
INTERSPEECH | 2 |
| 2009 | Signal separation for robust speech recognition based on phase difference information obtained in the frequency domainabstractIn this paper, we present a new two-microphone approach that improves speech recognition accuracy when speech is masked by other speech. The algorithm improves on previous systems that have been successful in separating signals based on differences in arrival time of signal components from two microphones. The present algorithm differs from these efforts in that the signal selection takes place in the frequency domain. We observe that additional smoothing of the phase estimates over time and frequency is needed to support adequate speech recognition performance. We demonstrate that the algorithm described in this paper provides better recognition accuracy than timedomain-based signal separation algorithms, and at less than 10 percent of the computation cost. Index Terms: Robust speech recognition, signal separation, time delay analysis, phase difference analysis Chanwoo Kim 0001, Kshitiz Kumar, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 4 |
| 2009 | Feature extraction for robust speech recognition using a power-law nonlinearity and power-bias subtractionabstractThis paper presents a new robust feature extraction algorithm based on a modified approach to power bias subtraction combined with applying a threshold to the power spectral density. Power bias level is selected as a level above which the signal power distribution is sharpest. The sharpness is measured using the ratio of arithmetic mean to the geometric mean of medium-duration power. When subtracting this bias level, power flooring is applied to enhance robustness. These new ideas are employed to enhance our recently introduced feature extraction algorithm PNCC (Power Normalized Cepstral Coefficient). While simpler than our previous PNCC, experimental results show that this new PNCC is showing better performance than our previous implementation. Index Terms — Robust speech recognition, physiological modeling, sharpness of power distribution, power flooring, auditory threshold 1. Chanwoo Kim 0001, Richard M. Stern |
INTERSPEECH | 2 |
| 2009 | Spatial separation of speech signals using amplitude estimation based on interaural comparisons of zero-crossings
Hyung-Min Park, Richard M. Stern |
Speech Commun. | 2 |
| 2008 | Analysis-by-synthesis features for speech recognitionabstractWe present a framework for speech recognition that accounts for hidden articulatory information. We model the articulatory space using a codebook of articulatory configurations geometrically derived from EMA measurements available in the MOCHA database. The articulatory parameter set we derive is in the form of Maeda parameters. In turn, these parameters are used in a physiologically motivated articulatory speech synthesizer based on the model by Sondhi and Schroeter. We use the distortion between the speech synthesized from each of the articulatory configurations and the original speech as features for recognition. We setup a segmented phoneme recognition task on the MOCHA database using Gaussian mixture models (GMMs). Improvements are achieved when combining the probability scores generated using the distortion features with the scores using acoustic features. Ziad Al Bawab, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2008 | Single-channel speech separation based on modulation frequencyabstractThis paper describes an algorithm that performs a simple form of computational auditory scene analysis to separate multiple speech signals from one another on the basis of the modulation frequencies of the components. The most novel aspect of the algorithm is the use of the cross-correlation of the instantaneous frequencies of the components of a signal to identify and separate those components that are likely have been produced by a common sound source. The putative desired target speech signal is reconstructed by choosing those components that have the greatest mutual correlation, and then using extrinsic information such as fundamental frequency or speaker identification to determine which component clusters belong to which speaker. The system was evaluated by comparing speech recognition accuracy of a target speech signal that was extracted from a mixture of two speakers. It was found that recognition accuracy obtained when the separation was based on cross-correlation of changes in instantaneous frequency was better than the accuracy obtained when the separation was performed on the basis of fundamental frequency alone, for both the DARPA Resource Management Database and the Grid database used in the 2006 Speech Separation Challenge. Lingyun Gu, Richard M. Stern |
ICASSP | 2 |
| 2008 | Environment-invariant compensation for reverberation using linear post-filtering for minimum distortionabstractSpeaker identification systems work quite well in controlled environments but their performance degrades severely in the presence of the reverberation that is frequently encountered in realistic acoustical environments. In this paper we develop an algorithm to make speaker identification systems more robust to reverberation by passing sequences of cepstral features through a short FIR filter. The coefficients of the filter are chosen to minimize the mean square differences between compensated features in the training and testing environments. Surprisingly, the resulting filter coefficients are relatively invariant to the actual nature of the reverberation. The use of the post-filtering approach is shown to improve speaker identification accuracy, especially when reverberation times are relatively long. Kshitiz Kumar, Richard M. Stern |
ICASSP | 2 |
| 2008 | Analysis of physiologically-motivated signal processing for robust speech recognitionabstractThis paper discusses the relative impact that different stages of a popular auditory model have on improving the accuracy of automatic speech recognition in the presence of additive noise. Recognition accuracy is measured using the CMU SPHINX-III speech recognition system, and the DARPA Resource Management speech corpus for training and testing. It is shown that feature extraction based on auditory processing provides better performance in the presence of additive background noise than traditional MFCC processing and it is argued that an expansive nonlinearity in the auditory model contributes the most to noise robustness. Index Terms: auditory modeling, robust speech recognition, auditory analysis Yu-Hsiang Bosco Chiu, Richard M. Stern |
INTERSPEECH | 2 |
| 2008 | Robust signal-to-noise ratio estimation based on waveform amplitude distribution analysisabstractIn this paper, we introduce a new algorithm for estimating the signal-to-noise ratio (SNR) of speech signals, called WADA-SNR (Waveform Amplitude Distribution Analysis). In this algorithm we assume that the amplitude distribution of clean speech can be approximated by the Gamma distribution with a shaping parameter of 0.4, and that an additive noise signal is Gaussian. Based on this assumption, we can estimate the SNR by examining the amplitude distribution of the noisecorrupted speech. We evaluate the performance of the WADA-SNR algorithm on databases corrupted by white noise, background music, and interfering speech. The WADA-SNR algorithm shows significantly less bias and less variability with respect to the type of noise compared to the standard NIST STNR algorithm. In addition, the algorithm is quite computationally efficient. Chanwoo Kim 0001, Richard M. Stern |
INTERSPEECH | 2 |
| 2007 | Profile View Lip ReadingabstractIn this paper, we introduce profile view (PV) lip reading, a scheme for speaker-dependent isolated word speech recognition. We provide historic motivation for PV from the importance of profile images in facial animation for lip reading, and we present feature extraction schemes for PV as well as for the traditional frontal view (FV) approach. We compare lip reading results for PV and FV, which demonstrate a significant improvement for PV over FV. We show improvement in speech recognition with the integration of audio and visual features. We also found it advantageous to process the visual features over a longer duration than the duration marked by the endpoints of the speech utterance. Kshitiz Kumar, Tsuhan Chen, Richard M. Stern |
ICASSP (4) | 3 |
| 2007 | Missing Feature Speech Recognition using Dereverberation and Echo Suppression in Reverberant EnvironmentsabstractThis paper describes an algorithm that efficiently segregates desired speech features from spatially-separated interfering sources in reverberant environments. Although most binaural segregation techniques successfully remove interference components in the absence of reverberation, source segregation in reverberant environments remains a challenging problem. In order to reduce the effects of reverberation, we present a method that dereverberates input signals before they are segregated. The dereverberation filter is estimated from the autocorrelation of the observations and primarily deals with early reflections, while late reflections are effectively suppressed by an inhibitory mechanism that estimates their relative contribution in each time-frequency segment. Information about the salience of the target in a given time-frequency segment based on source separation is combined with the corresponding information based on reverberation suppression through the use of a continually-variable weighting function or mask. Use of the novel reverberation processing results in a relative decrease in WER of 11.5% to 20.9% and use of the combined approaches reduces relative WER by as much as 65.3%. Hyung-Min Park, Richard M. Stern |
ICASSP (4) | 2 |
| 2007 | "polyaural" array processing for automatic speech recognition in degraded environmentsabstractIn this paper we present a new method of signal processing for robust speech recognition using multiple microphones. The method, loosely based on the human binaural hearing system, consists of passing the speech signals detected by multiple microphones through bandpass filtering and nonlinear halfwave rectification operations, and then cross-correlating the outputs from each channel within each frequency band. These operations provide rejection of off-axis interfering signals. These operations are repeated (in a non-physiological fashion) for the negative of the signal, and an estimate of the desired signal is obtained by combining the positive and negative outputs. We demonstrate that the use of this approach provides substantially better recognition accuracy than delay-and-sum beamforming using the same sensors for target signals in the presence of additive broadband and speech maskers. Improvements in reverberant environments are tangible but more modest. Index Terms: robust speech recognition, binaural hearing, auditory processing, speech enhancement Richard M. Stern, Evandro B. Gouvêa, Govindarajan Thattai |
INTERSPEECH | 1 |
| 2006 | Band-Independent Mask Estimation for Missing-Feature Reconstruction in the Presence of Unknown Background NoiseabstractAn effective mask estimation scheme for missing-feature reconstruction is described that achieves robust speech recognition in the presence of unknown noise. In previous work on Bayesian classification for mask estimation, white noise and colored noise were used for training mask estimators. This paper, which is concerned with both the simulation of a more diverse set of background environments and with mitigating the "sparse training" problem, describes a new Bayesian mask-estimation procedure in which each frequency band is trained independently. The new method employs colored noise for training, which is obtained by partitioning each frequency subband. We also propose a reevaluation method of voiced/unvoiced decisions to alleviate performance degradation that is caused by errors in pitch detection. Experimental results indicate that the proposed procedure in conjunction with cluster-based missing-feature imputation improves speech recognition accuracy on the Aurora 2.0 database in the presence for all types of background noise considered. Wooil Kim, Richard M. Stern |
ICASSP (1) | 2 |
| 2006 | Spatial Separation of Speech Signals Using Continuously-Variable Masks Estimated From Comparisons of Zero CrossingsabstractThis paper describes an algorithm that achieves noise robustness in speech recognition by reconstructing the desired signal from a mixture of two signals using continuously-variable masks. In contrast to current methods which use binary masks, this approach estimates the relative contribution of the desired source in a mixture of sources and reconstructs the desired signal in proportion to its estimated contribution to each time-frequency segment. Estimation of the continuously-variable masks is based on the relationship between the relative intensity of each source and the interaural time difference (ITD). Estimation of the ITD is accomplished using zero-crossing-based methods. It is shown that the use of zero-crossing approaches to estimate ITDs and continuously-variable masks provide better speech recognition accuracy than cross-correlation-based approaches to ITD estimation and binary masks. Hyung-Min Park, Richard M. Stern |
ICASSP (4) | 2 |
| 2006 | An integrated approach to improve speech recognition rate for non-native speakers
Yunbin Deng, Xiaokun Li, Chiman Kwan, Roger Xu, Bhiksha Raj, Richard M. Stern, David Williamson |
INTERSPEECH | 6 |
| 2006 | Physiologically-motivated synchrony-based processing for robust automatic speech recognitionabstractThis paper describes the structure and performance of a new signal processing scheme, motivated by the physiology of the peripheral auditory system, that improves speech recognition accuracy in the presence of broadband noise. An important attribute of the peripheral processing is a novel mechanism to represent the cycle-by-cycle synchrony in the response of low-frequency auditory-nerve fibers, in addition to the more conventional processing based on mean rate of response. It is shown that the use of the physiologically-motivated peripheral processing improves recognition accuracy in the presence of both broadband and transient noise, and that the use of the synchrony mechanism provides further improvement beyond that which is provided by the mean rate mechanism. Index Terms: auditory modeling, robust speech recognition, auditory snchrony. Chanwoo Kim 0001, Yu-Hsiang Bosco Chiu, Richard M. Stern |
INTERSPEECH | 3 |
| 2006 | Voting for two speaker segmentationabstractThe process of locating the end points of each speakers voice in an audio file and then clustering segments based in speaker identity is called speaker segmentation. In this paper we present a method for two speaker segmentation, though it can be extended to more than two speakers. Most methods for speaker segmentation and clustering start with an initial computationally inexpensive speaker segmentation method, followed by a more accurate segment clustering. In this paper we describe a simple algorithm that improves the accuracy of the segment clustering while not increasing the computational complexity. Since the clustering is done iteratively, the improvement in each segment clustering step results in a significant overall increase in segmentation accuracy and cluster purity. We borrow ideas from speaker recognition to perform segment clustering by frame based voting. We look at each frame as an independent classifier deciding which speaker generated that segment. These ’classifiers’ are combined by voting to make a decision as to which segments should be clustered together. This simple change leads to 56.9% decrease in error rates on a segmentation task for the SWITCHBOARD corpus. Index Terms: Speaker segmentation, Voting based classifier combination, Speaker change detection, Speaker clustering. Rashmi Gangadharaiah, Richard M. Stern |
INTERSPEECH | 3 |
| 2006 | Subband Likelihood-Maximizing Beamforming for Speech Recognition in Reverberant EnvironmentsabstractIn this paper, we introduce subband likelihood-maximizing beamforming (S-LIMABEAM), a new microphone-array processing algorithm specifically designed for speech recognition applications. The proposed algorithm is an extension of the previously developed LIMABEAM array processing algorithm. Unlike most array processing algorithms which operate according to some waveform-level objective function, the goal of LIMABEAM is to find the set of array parameters that maximizes the likelihood of the correct recognition hypothesis. Optimizing the array parameters in this manner results in significant improvements in recognition accuracy over conventional array processing methods when speech is corrupted by additive noise and moderate levels of reverberation. Despite the success of the LIMABEAM algorithm in such environments, little improvement was achieved in highly reverberant environments. In such situations where the noise is highly correlated to the speech signal and the number of filter parameters to estimate is large, subband processing has been used to improve the performance of LMS-type adaptive filtering algorithms. We use subband processing principles to design a novel array processing architecture in which select groups of subbands are processed jointly to maximize the likelihood of the resulting speech recognition features, as measured by the recognizer itself. By creating a subband filtering architecture that explicitly accounts for the manner in which recognition features are computed, we can effectively apply the LIMABEAM framework to highly reverberant environments. By doing so, we are able to achieve improvements in word error rate of over 20% compared to conventional methods in highly reverberant environments Michael L. Seltzer, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Environment-independent mask estimation for missing-feature reconstructionabstractIn this paper, we propose an effective mask-estimation method for missing-feature reconstruction in order to achieve robust speech recognition in unknown noise environments. In previous work, it was found that training a model for mask estimation on speech corrupted by white noise did not provide environment-independent recognition accuracy. In this paper we describe a training method based on bands of colored noise that is more effective in reflecting spectral variations across neighboring frames and subbands. We also achieved further improvement in recognition accuracy by reconsidering frames that appeared to be unvoiced in the initial pitch analysis. Performance is evaluated using the Aurora 2.0 database in the presence of various types of noise maskers. Experimental results indicate that the proposed methods are effective in estimating masks for missing-feature reconstruction while remaining more independent of the noise conditions. 1. Wooil Kim, Richard M. Stern, Hanseok Ko |
INTERSPEECH | 2 |
| 2005 | Feature compensation based on switching linear dynamic modelabstractIn this letter, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. We employ the switching linear dynamic model (SLDM) as a parametric model for the clean speech distribution, which enables us to exploit temporal correlations inherent in speech signals. Both the background noise and clean speech components are simultaneously estimated by means of the interacting multiple model (IMM) algorithm. Nam Soo Kim, Woohyung Lim, Richard M. Stern |
IEEE Signal Process. Lett. | 3 |
| 2004 | Feature generation based on maximum normalized acoustic likelihood for improved speech recognitionabstractFeature representation is a very important factor that has a great effect on the performance of speech recognition systems. In this paper we focus on a feature generation process that is based on the linear transformation of an original log-spectral representation. While conventional linear feature generation methods generally use objective functions that are not closely related to recognition accuracy, our linear feature generation method attempts to find a transformation matrix that maximizes the normalized acoustic likelihood of the most likely state training data, a measure that is directly related to the classification error rate in speech recognition. The transformation matrix is generated using a gradient ascent optimization process, with the normalized acoustic likelihood of the most likely state sequence as the objective function. Experimental results using the DARPA RM corpus show that the proposed method consistently decreases word error rates compared to conventional linear feature generation methods. Xiang Li 0071, Richard M. Stern |
ICASSP (1) | 2 |
| 2004 | On tracking noise with linear dynamical system modelsabstractThis paper investigates the use of higher-order autoregressive vector predictors for tracking the noise in noisy speech signals. The autoregressive predictors form the state equation of a linear dynamical system that models the spectral dynamics of the noise process. Experiments show that the use of such models to track noise can lead to large gains in recognition performance on speech compensated for the estimated noise. However, predictors of order greater than 1 are not observed to improve the performance beyond that obtained with a first-order predictor. We analyze and explain why this is so. Bhiksha Raj, Rita Singh, Richard M. Stern |
ICASSP (1) | 3 |
| 2004 | Parameter sharing in subband likelihood-maximizing beamforming for speech recognition using microphone arraysabstractIn this paper, we present methods to improve the computational efficiency of our previously proposed algorithm for microphone array processing for speech recognition, called subband likelihood-maximizing beamforming (S-LIMABEAM). In S-LIMABEAM, the parameters of a subband filter-and-sum beamformer are optimized to maximize the likelihood of the correct transcription of the utterance, as measured by the speech recognizer itself. This approach has been shown to produce significant improvements in recognition accuracy over conventional array processing methods in a variety of noisy and reverberant environments. However, because of the manner in which recognition features are computed, the number of subband parameters that have to be jointly optimized may be large, which slows the convergence of the algorithm. To address this problem, we present two methods of sharing parameters among multiple subband filters in order to significantly reduce the number of parameters to be optimized. Both of these methods exploit the spectral smoothing that occurs in the feature extraction process, but do so in different ways. By sharing parameters in the proposed manner, we are able to obtain a significant reduction in the time to convergence of S-LIMABEAM with a minimal degradation in speech recognition accuracy. Michael L. Seltzer, Richard M. Stern |
ICASSP (1) | 2 |
| 2004 | Parallel feature generation based on maximizing normalized acoustic likelihoodabstractCombining information from parallel feature streams generally improves speech recognition accuracy.While many studies have attempted to determine the stage of the recognition system that provides best combination performance and the specific nature of how features are combined, relatively little attention has been paid to the design or selection of parallel feature sets when used in combination.In this paper we propose a new parallel feature generation algorithm based on the criterion of maximizing the normalized acoustic likelihood of the features after they are combined, which is closely related to the recognition accuracy obtained using the combination of these features.We use a gradient ascent procedure to manipulate the values of a set of transformation matrices through which individual features are passed before they are combined in a fashion that maximizes the normalized acoustic likelihood term after the features are combined.The function that combine the parallel features together is an intrinsic part of the optimization process.The use of the optimal linear transformation provides a relative decrease of 12.7 percent Word Error Rate on the DARPA Resource Management task. Xiang Li 0071, Richard M. Stern |
INTERSPEECH | 2 |
| 2004 | Reconstruction of missing features for robust speech recognition
Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
Speech Commun. | 3 |
| 2004 | A Bayesian classifier for spectrographic mask estimation for missing feature speech recognition
Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
Speech Commun. | 3 |
| 2004 | Likelihood-maximizing beamforming for robust hands-free speech recognitionabstractSpeech recognition performance degrades significantly in distant-talking environments, where the speech signals can be severely distorted by additive noise and reverberation. In such environments, the use of microphone arrays has been proposed as a means of improving the quality of captured speech signals. Currently, microphone-array-based speech recognition is performed in two independent stages: array processing and then recognition. Array processing algorithms, designed for signal enhancement, are applied in order to reduce the distortion in the speech waveform prior to feature extraction and recognition. This approach assumes that improving the quality of the speech waveform will necessarily result in improved recognition performance and ignores the manner in which speech recognition systems operate. In this paper a new approach to microphone-array processing is proposed in which the goal of the array processing is not to generate an enhanced output waveform but rather to generate a sequence of features which maximizes the likelihood of generating the correct hypothesis. In this approach, called likelihood-maximizing beamforming, information from the speech recognition system itself is used to optimize a filter-and-sum beamformer. Speech recognition experiments performed in a real distant-talking environment confirm the efficacy of the proposed approach. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Training of stream weights for the decoding of speech using parallel feature streamsabstractIn speech recognition systems, information from multiple sources such as different feature streams can be combined in many different ways to yield better recognition accuracy. In general, information may be combined at the level of the incoming feature vectors, at the level of the decoding process, or after hypothesis generation. We focus on the specific case where parallel streams of features are used simultaneously during search to generate a hypothesis, or a set of hypotheses. In this case the contributions of the individual features to the score associated with a frame of speech must be weighted appropriately during search. We present an offline data-driven algorithm for determining the weights to be associated with each feature stream for combining acoustic likelihoods for each frame. Experimental results show that the word error rates (WERs) obtained using the proposed algorithm are lower than those obtained using conventional schemes for parallel feature combination. Xiang Li 0071, Richard M. Stern |
ICASSP (1) | 2 |
| 2003 | Subband parameter optimization of microphone arrays for speech recognition in reverberant environmentsabstractWe present a new subband microphone array processing algorithm specifically designed for speech recognition applications. We previously proposed a speech recognizer-based array processing algorithm which resulted in significant improvements in recognition accuracy when the speech was corrupted by additive noise and moderate levels of reverberation. However, little improvement was achieved over conventional beamforming methods in highly reverberant environments. Subband processing has been used to improve the poor performance of LMS-type algorithms when the number of filter parameters to estimate is large and the noise is highly correlated to the speech signal, e.g. in highly reverberant environments. We apply a subband approach to a new array processing architecture in which select groups of subbands are processed jointly to maximize the likelihood of the resulting speech recognition features, as measured by the recognition system itself. By incorporating the recognizer into the filter optimization scheme we ensure that signal components important for recognition are emphasized without undue emphasis on less critical components. By utilizing a subband approach, we can effectively apply this framework to highly reverberant environments. In doing so, we are able to achieve improvements in word error rate of over 20% compared to conventional methods in highly reverberant environments. Michael L. Seltzer, Richard M. Stern |
ICASSP (1) | 2 |
| 2003 | Feature generation based on maximum classification probability for improved speech recognitionabstractFeature representation is a very important factor that has great effect on the performance of speech recognition systems. In this paper we focus on a feature generation process that is based on linear transformation of the original log-spectral representation. We first discuss several three popular linear transformation methods, Mel-Frequency Cepstral Coefficients (MFCC), Principal Component Analysis (PCA), and Linear Discriminant Analysis (LDA). We then propose a new method of linear transformation that maximizes the normalized acoustic likelihood of the most likely state sequences of training data, a measure that directly related to our ultimate objective of reducing Bayesian classification error rate in speech recognition. Experimental results show that the proposed method decreases the relative word error rate by more than 9.1 % compared to the best implementation of LDA, and by more than 25.9 % compared to MFCC features. 1. Xiang Li 0071, Richard M. Stern |
INTERSPEECH | 2 |
| 2003 | Duration normalization and hypothesis combination for improved spontaneous speech recognitionabstractWhen phone segmentations are known a priori, normalizing the duration of each phone has been shown to be effective in overcoming weaknesses in duration modeling of Hidden Markov Models (HMMs). While we have observed potential relative reductions in word error rate (WER) of up to 34.6% with oracle segmentation information, it has been difficult to achieve significant improvement in WER with segmentation boundaries that are estimated blindly. In this paper we present simple variants of our duration normalization algorithm, which make use of blindly-estimated segmentation boundaries to produce different recognition hypotheses for a given utterance. These hypotheses can then be combined for significant improvements in WER. With oracle segmentations, WER reductions of up to 38.5% are possible. With automatically-derived segmentations, this approach has achieved a reduction of WER of 3.9% for the Broadcast News corpus, 6.2% for the spontaneous register of the MULT_REG corpus, and 7.7% for a spontaneous corpus of connected Spanish digits collected by Telefonica de Investigacion y Desarrollo. Jon P. Nedel, Richard M. Stern |
INTERSPEECH | 2 |
| 2003 | Normalization of time-derivative parameters using histogram equalizationabstractIn this paper we describe a new framework of feature compensation for robust speech recognition. We introduce Delta-Cepstrum Normalization (DCN) that normalizes not only cepstral coefficients, but also their time-derivatives. In previous work, the mean and the variance of cepstral coefficients are normalized to reduce the irrelevant information, but such a normalization was not applied to time-derivative parameters because the reduction of the irrelevant information was not enough. However, Histogram Equalization provides better compensation and can be applied even to delta and delta-delta cepstra. We investigate various implementation of DCN, and show that we can achieve the best performance when the normalization of the cepstra and delta cepstra can be mutually interdependent. We evaluate the performance of DCN using speech data recorded by a PDA. DCN provides significant improvements compared to HEQ. We also examine the possibility of combining Vector Taylor Series (VTS) and DCN. Even though some combinations do not improve the performance of VTS, it is shown that the best combination gives better performance than VTS alone. Finally, the advantages of DCN in terms of the computation speed are also discussed. Yasunari Obuchi, Richard M. Stern |
INTERSPEECH | 2 |
| 2002 | Speech recognizer-based microphone array processing for robust hands-free speech recognitionabstractWe present a new array processing algorithm for microphone array speech recognition. Conventionally, the goal of array processing is to take distorted signals captured by the array and generate a cleaner output waveform. However, speech recognition systems operate on a set of features derived from the waveform, rather than the waveform itself. The goal of an array processor used in conjunction with a recognition system is to generate a waveform which produces a set of recognition features which maximize die likelihood for the words that are spoken, rather than to minimize the waveform distortion. We propose a new array processing algorithm which maximizes the likelihood of the recognition features. This is accomplished through the use of a new objective function which utilizes information from the recognition system itself, obtained in an unsupervised manner, to optimize the parameters of a filter-and-sum array processor. Using the proposed method, improvements in word error rate of up to 36% over conventional methods are achieved on real microphone array tasks in a wide range of environments. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2002 | Combining search spaces of heterogeneous recognizers for improved speech recogniton
Xiang Li 0071, Rita Singh, Richard M. Stern |
INTERSPEECH | 3 |
| 2002 | Automatic generation of subword units for speech recognition systemsabstractLarge vocabulary continuous speech recognition (LVCSR) systems traditionally represent words in terms of smaller subword units. Both during training and during recognition, they require a mapping table, called the dictionary, which maps words into sequences of these subword units. The performance of the LVCSR system depends critically on the definition of the subword units and the accuracy of the dictionary. In current LVCSR systems, both these components are manually designed. While manually designed subword units generalize well, they may not be the optimal units of classification for the specific task or environment for which an LVCSR system is trained. Moreover, when human expertise is not available, it may not be possible to design good subword units manually. There is clearly a need for data-driven design of these LVCSR components. In this paper, we present a complete probabilistic formulation for the automatic design of subword units and dictionary, given only the acoustic data and their transcriptions. The proposed framework permits easy incorporation of external sources of information, such as the spellings of words in terms of a nonideographic script. Rita Singh, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 3 |
| 2001 | Duration normalization for improved recognition of spontaneous and read speech via missing feature methodsabstractHidden Markov models (HMMs) are known to model the duration of sound units poorly. We present a technique to normalize the duration of each phone to overcome this weakness, with the conjecture that speech with normalized phone durations may be better modeled and discriminated using standard HMM acoustic models. Duration normalization is accomplished by dropping frames if a phone is longer than the desired duration and by adding "missing" frames and reconstructing them if a phone is shorter than the desired duration. If phone segmentations are known a priori, we achieve a 15.8% reduction in relative word error rate (WER) on spontaneous speech and a 10.3% reduction in relative WER on read speech. Preliminary work with automatic phone segmentations derived from the data is also presented. Jon P. Nedel, Richard M. Stern |
ICASSP | 2 |
| 2001 | Speech in Noisy Environments: robust automatic segmentation, feature extraction, and hypothesis combinationabstractThe first evaluation for Speech in Noisy Environments (SPINE1) was conducted by the Naval Research Labs (NRL) in August, 2000. The purpose of the evaluation was to test existing core speech recognition technologies for speech in the presence of varying types and levels of noise. In this case the noises were taken from military settings. Among the strategies used by Carnegie Mellon University's successful systems designed for this task were session-adaptive segmentation, robust mel-scale filtering for the computation of cepstra, the use of parallel front-end features and noise-compensation algorithms, and parallel hypotheses combination through word-graphs. This paper describes the motivations behind the design decisions taken for these components, supported by observations and experiments. Rita Singh, Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
ICASSP | 4 |
| 2001 | Distortion-class modeling for robust speech recognition under GSM RPE-LTP coding
Juan M. Huerta, Richard M. Stern |
Speech Commun. | 2 |
| 2000 | Inter-class MLLR for speaker adaptationabstractThis paper examines the use of interdependencies of parameter classes in transformation-based speaker adaptation algorithms such as maximum likelihood linear regression (MLLR). In transformation-based adaptation, increasing the number of transformation classes can provide more detailed information for adaptation, but at the expense of greater estimation error with small amounts of data. In this paper we introduce a new procedure, inter-class MLLR, which utilizes the relationship between different classes to achieve both detailed and reliable transformation-based adaptation using limited data, In this method, the inter-class relation is given by a linear regression which is estimated from training data. In experiments using non-native English speakers from the Spoke 3 data in the 1994 DARPA Wall Street Journal evaluation, inter-class MLLR provided a relative reduction in word error rates of 11.3% compared to conventional MLLR. Sam-Joo Doh, Richard M. Stern |
ICASSP | 2 |
| 2000 | Automatic generation of phone sets and lexical transcriptionsabstractLarge vocabulary automatic speech recognition systems model words as sequences of a small set of basic sub-word units (the phoneset), which the systems are trained to classify. All words in the system's vocabulary are transcribed in terms of this set in a dictionary. The phoneset and dictionary are specific to a language and are typically designed manually. The system's performance is critically dependent on the quality of the phoneset and the accuracy of the dictionary. The authors attempt to generate the phoneset and dictionary automatically, using only the training data and their transcriptions. We treat this as a joint optimization problem with a maximum a posteriori solution for the dictionary and a maximum likelihood solution for the phoneset and its acoustic models. Experiments with the DARPA Resource Management corpus show that the automatically generated phoneset and dictionary result in recognition accuracies close to those obtained using manually designed ones. Rita Singh, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 2000 | Using class weighting in inter-class MLLRabstractA new adaptation method called inter-class MLLR has recently been introduced. Inter-class MLLR utilizes relationships among different transformation functions to achieve more reliable estimates of MLLR parameters across multiple classes, and it produces lower word error rates (WER) than conventional MLLR in circumstances where very little speaker-specific adaptation data are available. This paper describes the application of weights to the neighboring classes to improve the effectiveness with which they are combined with the target class in inter-class MLLR. These weights are obtained from the variance of the estimation error considering the weighted least squares estimation in classical linear regression. In our experiments, the weights provided small improvements in WER for supervised adaptation but almost no improvement in unsupervised adaptation using only a small amount of adaptation data. We also discuss the effect of decreasing the number of neighboring classes as more adaptation data become available, the development of inter-class transformations from the test speaker, and the combination of inter-class MLLR with principal-component MLLR. None of the feasible variations of weighted inter-class MLLR provided significant improvements to recognition accuracy. 1. Sam-Joo Doh, Richard M. Stern |
INTERSPEECH | 2 |
| 2000 | Instantaneous-distortion based weighted acoustic modeling for robust recognition of coded speechabstractIn this paper we apply the Weighted Acoustic Modeling (WAM) technique to the recognition of speech coded by the full-rate GSM codec or the FS-1016 CELP codec employing various estimates of instantaneous distortion.In the WAM method, separate hidden Markov models are developed for regions of speech that exhibit low levels of codec-induced distortion and for regions with higher levels of such distortion.At recognition time, the contributions of these models are mixed together with a weighting that is determined by estimating the instantaneous distortion.In this paper instantaneous distortion was estimated from the instantaneous cepstral distortion, the long-term gain parameter of the codec, the long-term predictability of the reconstructed signal, and measurements of recoding sensitivity.We observe that the use of the long-term gain parameter produces results that are similar to those obtained by use of cepstral distortion (which can only be obtained if the original cepstra are transmitted along with the speech signal) for the GSM codec.Overall, the effect of the degradation in error rate introduced by coding can be reduced by up to 55% with these techniques for GSM coding, and by up to 38% for the CELP coding. Juan M. Huerta, Richard M. Stern |
INTERSPEECH | 2 |
| 2000 | Phone transition acoustic modeling: application to speaker independent and spontaneous speech systemsabstractHMM-based large vocabulary speech recognition systems usually have a very large number of statistical parameters. For better estimation, the number of parameters is reduced by sharing them across models. The parameter sharing is decided by regression trees which are built using phonetic classes designed either by a human expert or by data-driven methods. In situations where neither of these are reliable, it may be useful to have techniques for non-decision-tree based state tying which perform comparably to those based on traditional methods. In this paper we propose two methods for non-decision tree based parameter learning in HMM-based systems. In the first method (context-dependent state tying), we restructure acoustic models to explicitly capture the transitions between phones in continuous speech. In the second method (transition-based subword units), we redefine the basic sound units used to model speech to model transitions between sounds explicitly. Experiments show that context-dependent state tying is a viable option for large vocabulary systems. They also show that using transition-based subword units can improve performance on spontaneous speech. Jon P. Nedel, Rita Singh, Richard M. Stern |
INTERSPEECH | 3 |
| 2000 | Automatic subword unit refinement for spontaneous speech recognition via phone splittingabstractSpontaneous speech is highly variable and rarely conforms to conventional assumptions and linguistically defined pronunciation rules. Specifically, there may be many different continuous speech realizations for each expertly defined phonetic unit in the dictionary. The phones may be realized in a clean and complete fashion as in read speech, or they may be realized in a sloppy and incomplete fashion as in highly spontaneous speech. For spontaneous speech, therefore, it may be beneficial to model incompletely realized variants of any phonetic unit as separate units. In this paper we test this hypothesis by introducing two possible modeling classes for the phones AA and IY in the standard English CMU recognition dictionary. We propose three different automatic methods of segregating the training data properly in order to identify and label the appropriate variants. Each of these methods results in improved recognition performance over the baseline, leading to the conclusion that finer modeling frameworks can be helpful to parameterize properly and recognize spontaneous speech. Jon P. Nedel, Rita Singh, Richard M. Stern |
INTERSPEECH | 3 |
| 2000 | Reconstruction of damaged spectrographic features for robust speech recognitionabstractWe present two missing-feature based algorithms that recover noise-corrupted regions of spectrographic representations of speech for noise-robust speech recognition. These algorithms modify the incoming feature vector without any changes to the speech recognition system, in contrast to previously-described approaches. The first approach clusters the feature vectors representing clean speech. Missing data are recovered by estimating the spectral cluster in each analysis frame based on the uncorrupted feature values. The second approach uses MAP procedures to estimate the values of missing data elements based on their correlations with the features that are present. Both methods take into account bounds on the clean spectrogram implied by the noisy spectrogram. Large improvements in recognition accuracy are observed when these methods are used on speech corrupted by non-stationary noise when the locations of the corrupt regions of the spectrogram are known. We also present a new method of estimating the locations of corrupt regions in spectrograms that treats the problem of identifying these regions as one of Bayesian classification. This method, when used along with the best method to reconstruct them, results in recognition accuracies comparable with the best previous data compensation algorithm on speech corrupted by white noise. It also provides significant improvement on speech corrupted by music when the global SNR of the corrupted signal is known a priori. 1. Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
INTERSPEECH | 3 |
| 2000 | Classifier-based mask estimation for missing feature methods of robust speech recognitionabstractMissing feature methods of noise compensation for speech recognition operate by removing components of a spectrographic representation of speech that are considered to be corrupt, as indicated by a low signal-to-noise ratio. Recognition is either performed directly on the incomplete spectrograms or the missing components are reconstructed prior to recognition. These methods require a spectrographic mask which accurately labels the reliable and corrupt regions of the spectrogram. Current methods of mask estimation rely on assumptions about the corrupting noise such as stationarity. This is a significant drawback since the missing feature methods themselves have no such restrictions. We present a new mask estimation technique that uses a Bayesian classifier to determine the reliability of spectrographic elements. Features were designed that make no assumptions about the corrupting noise signal, but rather exploit characteristics of the speech signal itself. Missing feature compensation experiments were performed on speech corrupted by a variety of noises. In all cases, classifier-based mask estimation resulted in significantly better recognition accuracy than conventional mask estimation methods. 1. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 3 |
| 2000 | Structured redefinition of sound units by merging and splitting for improved speech recognitionabstractThe performance of speech recognition systems degrades when the basic sound units used are poorly defined or inconsistently used. Several attempts have been made to improve dictionaries automatically, either by redefining pronunciations of words in terms of existing sound units, or by redefining the sound units themselves completely. The problem with these approaches is that, while the former is limited by the sound units used, the latter discards all human information that has been incorporated into an expert-designed recognition dictionary. In this paper we propose a new merging-andsplitting algorithm that attempts to redefine the basic sound units used in the dictionary, while maintaining the expert knowledge built into a manually designed dictionary. Sound units from an existing dictionary are merged based on their inherent confusability, as measured by a Monte-Carlo based metric, and subsequently split to maximize the likelihood of the training data. Experiments with the Resource Management database indicate that this approach results in an improvement in recognition accuracy when context-independent models are used for recognition. When context-dependent models are used, the improvement observed is reduced. 1. Rita Singh, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 3 |
| 1999 | Automatic clustering and generation of contextual questions for tied states in hidden Markov modelsabstractMost current automatic speech recognition systems based on HMMs cluster or tie together subsets of the subword units with which speech is represented. This tying improves the recognition accuracy when systems are trained with limited data, and is performed by classifying the sub-phonetic units using a series of binary tests based on speech production, called "linguistic questions". This paper describes a new method for automatically determining the best combinations of subword units to form these questions. The hybrid algorithm proposed clusters state distributions of context-independent phones to obtain questions for triphonetic contexts. Experiments confirm that the questions thus generated can replace manually generated questions and can provide improved recognition accuracy. Automatic generation of questions has the additional important advantage of extensibility to languages for which the phonetic structure is not well understood by the system designer, and can be effectively used in situations where the subword units are not phonetically motivated. Rita Singh, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 1999 | Domain adduced state tying for cross-domain acoustic modellingabstractHomophone words is one of the specific problems of Automatic Speech Recognition (ASR) in French. Moreover, this phenomenon is particularly high for some inflections like the singular/plural inflection (72% of the 40.7K lemma of our 240K word dictionary have inflected forms which are homophonic). In order to take into account worddependencies spanning over a variable number of words, it is interesting to merge local language models, like 3-gram or 3-class models, with largespan models. We present in this paper two kinds of models : a phrase-based model, using phrases obtained from a training corpus by means of a finite-state parser; a homophone cache-based model, using derivation of constraints from word histories stored in a cache memory. Rita Singh, Bhiksha Raj, Richard M. Stern |
EUROSPEECH | 3 |
| 1998 | Speech recognition from GSM codec parametersabstractSpeech coding affects speech recognition performance, with recognition accuracy deteriorating as the coded bit rate decreases.Virtually all systems that recognize coded speech reconstruct the speech waveform from the coded parameters, and then perform recognition (after possible noise and/or channel compensation) using conventional techniques.In this paper we compare the recognition accuracy of coded speech obtained by reconstructing the speech waveform with the speech recognition accuracy obtained when using cepstral features derived from the coding parameters.We focus our efforts on speech that has been coded using the 13-kbps full-rate GSM codec, a Regular Pulse Excited Long Term Prediction (RPE-LTP) codec.The GSM codec develops separate representations for the linear prediction (LPC) filter and the residual signal components of the coded speech.We measure the effects of quantization and coding on the accuracy with which these parameters are represented, and present two different methods for recombining them for speech recognition purposes.We observe that by selectively combining the cepstral streams representing the LPC parameters and the residual signal it is possible to obtain recognition accuracy directly from the coded parameters that equals or exceeds the recognition accuracy obtained from the reconstructed waveforms. Juan M. Huerta, Richard M. Stern |
ICSLP | 2 |
| 1998 | Inference of missing spectrographic features for robust speech recognitionabstractTwo types of algorithms are introduced that recover missing time-frequency regions of log-spectral representations of speech. These compensation algorithms modify the incoming feature vector without any changes to the speech recognition system, in contrast to previously-described approaches. The first approach clusters the log-spectral vectors representing clean speech. Missing data are recovered by estimating the spectral cluster in each analysis frame on the basis of the feature values that are present. The second approach uses MAP procedures to estimate the values of missing data elements based on their correlation with the features that are present. Greatest recognition accuracy was obtained using the correlation-based approach, presumably because of its ability to exploit the temporal as well as spectral structure of speech. The recognition accuracy provided by these algorithms approaches but does not exceed that obtained by traditional marginalization. Nevertheless, it is believed that these algorithms provide greater computational efficiency and enable greater flexibility in recognition system structure. 1. Bhiksha Raj, Rita Singh, Richard M. Stern |
ICSLP | 3 |
| 1998 | Data-driven environmental compensation for speech recognition: A unified approach
Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
Speech Commun. | 3 |
| 1997 | The effects of background music on speech recognition accuracyabstractRecognition of broadcast data, such as TV and radio programs is a topic of great interest. One of the problems with such data is the frequent presence of background music that degrades the performance of speech recognition systems. In this paper we examine the effects of different kinds of music on automatic speech recognition systems by comparing the effects of music with the relatively well-known effects of white noise on these systems. We also examine the extent to which compensation algorithms that have been successfully applied to noisy speech are also helpful in improving recognition accuracy for speech that is corrupted by music. It is hoped that these experimental comparisons will lead to a better understanding of how to compensate for the effects of background music. 1. Bhiksha Raj, Vipul N. Parikh, Richard M. Stern |
ICASSP | 3 |
| 1997 | Speaker normalization through formant-based warping of the frequency scaleabstractSpeaker-dependent automatic speech recognition systems are known to outperform speaker-independent systems when enough training data are available to model acoustical variability among speakers. Speaker normalization techniques modify the spectral representation of incoming speech waveforms in an attempt to reduce variability between speakers. Recent successful speaker normalization algorithms have incorporated a speaker-specific frequency warping to the initial signal processing stages. These algorithms, however, do not make extensive use of acoustic features contained in the incoming speech. In this paper we study the possible benefits of the use of acoustic features in speaker normalization algorithms using frequency warping. We study the extent to which the use of such features, including specifically the use of formant frequencies, can improve recognition accuracy and reduce computational complexity for speaker normalization. We examine the characteristics and limitations of several types of feature sets and warping functions as we compare their performance relative to existing algorithms. 1. Evandro B. Gouvêa, Richard M. Stern |
EUROSPEECH | 2 |
| 1997 | Compensation for environmental and speaker variability by normalization of pole locations
Juan M. Huerta, Richard M. Stern |
EUROSPEECH | 2 |
| 1996 | A vector Taylor series approach for environment-independent speech recognitionabstractIn this paper we introduce a new analytical approach to environment compensation for speech recognition. Previous attempts at solving analytically the problem of noisy speech recognition have either used an overly-simplified mathematical description of the effects of noise on the statistics of speech or they have relied on the availability of large environment-specific adaptation sets. Some of the previous methods required the use of adaptation data that consists of simultaneously-recorded or "stereo" recordings of clean and degraded speech. In this work we introduce the use of a vector Taylor series (VTS) expansion to characterize efficiently and accurately the effects on speech statistics of unknown additive noise and unknown linear filtering in a transmission channel. The VTS approach is computationally efficient. It can be applied either to the incoming speech feature vectors, or to the statistics representing these vectors. In the first case the speech is compensated and then recognized; in the second case HMM statistics are modified using the VTS formulation. Both approaches use only the actual speech segment being recognized to compute the parameters required for environmental compensation. We evaluate the performance of two implementations of VTS algorithms using the CMU SPHINX-II system on the 100-word alphanumeric CENSUS database and on the 1993 5000-word ARPA Wall Street Journal database. Artificial white Gaussian noise is added to both databases. The VTS approaches provide significant improvements in recognition accuracy compared to previous algorithms. Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
ICASSP | 3 |
| 1996 | Cepstral compensation by polynomial approximation for environment-independent speech recognition
Bhiksha Raj, Evandro B. Gouvêa, Pedro J. Moreno 0001, Richard M. Stern |
ICSLP | 4 |
| 1995 | Multivariate-Gaussian-based cepstral normalization for robust speech recognitionabstractWe introduce a new family of environmental compensation algorithms called multivariate gaussian based cepstral normalization (RATZ). RATZ assumes that the effects of unknown noise and filtering on speech features can be compensated by corrections to the mean and variance of components of Gaussian mixtures, and an efficient procedure for estimating the correction factors is provided. The RATZ algorithm can be implemented to work with or without the use of "stereo" development data that had been simultaneously recorded in the training and testing environments. "Blind" RATZ partially overcomes the loss of information that would have been provided by stereo training through the use of a more accurate description of how noisy environments affect clean speech. We evaluate the performance of the two RATZ algorithms using the CMU SPHINX-II system on the alphanumeric census database and compare their performance with that of previous environmental-robustness developed at CMU. Pedro J. Moreno 0001, Bhiksha Raj, Evandro B. Gouvêa, Richard M. Stern |
ICASSP | 4 |
| 1995 | On the effects of speech rate in large vocabulary speech recognition systemsabstractIt is well known that a higher-than-normal speech rate will cause the rate of recognition errors in large vocabulary automatic speech recognition (ASR) systems to increase. In this paper we attempt to identify and correct for errors due to fast speech. We first suggest that phone rate is a more meaningful measure of speech rate than the more common word rate. We find that when data sets are clustered according to the phone rate metric, recognition errors increase when the phone rate is more than 1 standard deviation greater than the mean. We propose three methods to improve the recognition accuracy of fast speech, each addressing different aspects of performance degradation. The first method is an implementation of Baum-Welch codebook adaptation. The second method is based on the adaptation of HMM state-transition probabilities. In the third method, the pronunciation dictionaries are modified using rule-based techniques and compound words are added. We compare improvements in recognition accuracy for each method using data sets clustered according to the phone rate metric. Adaptation of the HMM state-transition probabilities to fast speech improves recognition of fast speech by a relative amount of 4 to 6 percent. Matthew A. Siegler, Richard M. Stern |
ICASSP | 2 |
| 1995 | A unified approach for robust speech recognitionabstractThere are two major structural approaches to robust speech recognition.In the first approach to the problem, compensation is performed by modifying the incoming cepstral stream using ML or MMSE methods to estimate parameters characterizing environmental degradation, from direct frame-by-frame comparisons between speech recorded in high-quality and degraded acoustical environments, or by signal processing techniques such as spectral subtraction.The second approach tackles the problem by modifying the statistics of the internal representation of speech cepstra in the classifier to make them more closely resemble the statistics of degraded speech.This paper attempts to unify these approaches to robust speech recognition by presenting three techniques that share the same basic assumptions and internal structure but differ in whether they modify the incoming speech cepstra or whether they modify the classifier statistics.We present SNR-dependent multi-vaRiate gAussian-based cepsTral normaliZation (SNR-RATZ) and SNR-based Blind RATZ (SNR-BRATZ), which modify incoming cepstra, along with STAR (STAtistical Re-estimation), which modifies the internal statistics of the classifier.The algorithms were tested using the SPHINX-II speech recognition system on the CENSUS database, a database of strings of letters and numbers to which unknown added and unknown linear filtering was introduced artificially.While all the algorithms showed good performance, STAR was observed to provide lower error rates as SNR decreases than any of the algorithms that modify incoming cepstra. Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
EUROSPEECH | 3 |
| 1994 | Environment normalization for robust speech recognition using direct cepstral comparisonabstractIn this paper we describe and evaluate a series of new algorithms that compensate for the effects of unknown acoustical environments or changes in environment. The algorithms use compensation vectors that are added to the cepstral representations of speech that is input to a speech recognition system. While these vectors are computed from direct frame-by-frame comparisons of cepstra of speech simultaneously recorded in the training environment and various prototype testing environments, the compensation algorithms do not assume that the acoustical characteristics of the actual testing environment are known. The specific compensation vector applied in a given frame depends on either physical attributes such as SNR or presumed phonetic identity. The compensation algorithms are evaluated using the 1992 ARPA 5000 word WSJ/CSR corpus. The best system combines phoneme-based and SNR-based cepstral compensation with cepstral mean normalization, and provides a 66.8% reduction in error rate over baseline processing when tested using a standard suite of unknown microphones.> Fu-Hua Liu, Richard M. Stern, Alex Acero, Pedro J. Moreno 0001 |
ICASSP (2) | 2 |
| 1994 | Sources of degradation of speech recognition in the telephone networkabstractWe compare speech recognition accuracy for high-quality speech recorded under controlled conditions with speech as it appears over long-distance telephone lines. In addition to comparing recognition accuracy we use telephone-channel simulation to identify the sources of degradation of speech over telephone lines that have the greatest impact on speech recognition accuracy. We first compare the performance of the CMU SPHINX-I system on the TIMIT and NTIMIT databases. We found that other factors beyond a mere decrease in bandwidth cause the observed degradation in recognition accuracy, and that the environmental compensation algorithms RASTA and CDCN fail to compensate completely for degradations introduced by the telephone network. We identify the most problematic telephone-channel impairments using a commercial telephone channel simulator and the SPHINX-II system. Of the various effects considered, additive noise and linear filtering appear to have the greatest impact on recognition accuracy. Finally, we examined the performance of three cepstral compensation algorithms in the presence of the most damaging conditions. We found the compensation algorithms to be effective except for the worst 1% of the telephone channels.> Pedro J. Moreno 0001, Richard M. Stern |
ICASSP (1) | 2 |
| 1994 | Robust speech recognition in the automobileabstractIn this paper we discuss a number of the ways in which the recognition accuracy of automatic speech recognition systems is affected by ambient noise in the automobile, along with the extent to which various techniques for robust speech recognition can provide for more robust recognition. We consider separately the effects of engine noise, interference by turbulent air outside the car, interference by sounds from the car’s radio, and interference by the sounds of the car’s windshield wipers. Recognition accuracy was compared using baseline processing, cepstral mean normalization (CMN), and codeword-dependent cepstral normalization (CDCN). The greatest degradation in recognition accuracy was produced by interference from AM-radio talk shows. The use of CMN and especially CDCN was found to be significantly improve recognition accuracy, except for the effects of interference from radio talk shows at low car speeds. This type of interference is effectively suppressed through the use of adaptive noise cancellation techniques. Nobutoshi Hanai, Richard M. Stern |
ICSLP | 2 |
| 1994 | Environmental robustness in automatic speech recognition using physiologic ally-motivated signal processing
Yoshiaki Ohshima, Richard M. Stern |
ICSLP | 2 |
| 1994 | Signal processing for robust speech recognitionabstractThis paper describes several new cepstral-based compensation procedures that render the SPHINX-II system more robust with respect to acoustical environment.The first algorithm, phonedependent cepstral compensation, is similar in concept to the previously-described MFCDCN method, except that cepstral compensation vectors are selected according to the current phonetic hypothesis, rather than on the basis of SNR or VQ codeword identity.We also describe two procedures to accomplish adaptation of the VQ codebook for new environments.Use of the various compensation algorithms in consort produces a reduction of error rates for SPHINX-II by as much as 40 percent relative to the rate achieved with cepstral mean normalization alone. Richard M. Stern, Fu-Hua Liu, Pedro J. Moreno 0001, Alex Acero |
ICSLP | 1 |
| 1993 | Multi-microphone correlation-based processing for robust speech recognition
Thomas M. Sullivan, Richard M. Stern |
ICASSP (2) | 2 |
| 1992 | Efficient joint compensation of speech for the effects of additive noise and linear filteringabstractThe authors describe two algorithms that provide robustness for automatic speech recognition systems in a fashion that is suitable for real-time environmental normalization for workstations of moderate size. The first algorithm is a modification of the SNR-dependent cepstral normalization (SDCN) and the fixed code-word dependent cepstral normalization (FCDCN) algorithms given by Acero and Stern (1990), except that unlike these algorithms it provides computationally-efficient environment normalization without prior knowledge of the acoustical characteristics of the environment in which the system will be operated. The second algorithm is a modification of the more complex CDCN algorithm that enables it to perform environmental compensation in better than real time. The authors compare the recognition accuracy, computational complexity, and amount of training data needed to adapt to new acoustical environments using these algorithms with several different types of headset-mounted and desktop microphones.> Fu-Hua Liu, Alex Acero, Richard M. Stern |
ICASSP | 3 |
| 1992 | Multiple approaches to robust speech recognitionabstractThis paper compares several different approaches to robust speechWe have found that two major factors degrading the performance of recognition.We review CMU's ongoing research in the use of speech recognition systems using desktop microphones in normal acoustical pre-processing to achieve robust speech recognition, inoffice environments are additive noise and unknown linear filtering.cluding the first evaluation of pre-processing in the context of the We showed in [2, 6] that simultaneous joint compensation for the DARPA standard ATIS domain for spoken language systems.We effects of additive noise and linear filtering is needed to achieve also describe and compare the effectiveness of three complementary maximal robustness with respect to acoustical differences between the methods of signal processing for robust speech recognition: acoustical training and testing environments of a speech recognition system.pre-processing, microphone array processing, and the use of We described in [2, 6] two algorithms that perform such joint comphysiologically-motivated models of peripheral signal processing.pensation, based on additive corrections to the cepstral coefficients of Recognition error rates are presented using these three approaches in the speech waveform.The more effective and adaptive of these isolation and in combination with each other for the speakeralgorithms, Codeword-Dependent Cepstral Normalization (CDCN) independent continuous alphanumeric census speech recognition task.[2], uses EM techniques to compute ML estimates of the additive noise and linear filtering that corrupt "clean" speech signals.The CDCN algorithm adapts automatically to new testing environments, Richard M. Stern, Fu-Hua Liu, Yoshiaki Ohshima, Thomas M. Sullivan, Alex Acero |
ICSLP | 1 |
| 1991 | Robust speech recognition by normalization of the acoustic spaceabstractSeveral algorithms are presented that increase the robustness of SPHINX, the CMU (Carnegie Mellon University) continuous-speech speaker-independent recognition systems, by normalizing the acoustic space via minimization of the overall VQ distortion. The authors propose an affine transformation of the cepstrum in which a matrix multiplication perform frequency normalization and a vector addition attempts environment normalization. The algorithms for environment normalization are efficient and improve the recognition accuracy when the system is tested on a microphone other than the one on which it was trained. The frequency normalization algorithm applies a different warping on the frequency axis to different speakers and it achieves a 10% decrease in error rate.> Alex Acero, Richard M. Stern |
ICASSP | 2 |
| 1991 | Speaker adaptation in continuous speech recognition via estimation of correlated mean vectorsabstractRecent attempts to improve the recognition performance of a semi-continuous version of the CMU SPHINX system (SPHINX-SC) through the use of speaker adaptation are described. The authors' approach to speaker adaptation is to use multivariate parameter estimation procedures to update the mean values of the component densities which comprise the system's codebook, given the speaker-specific observations. The authors have developed a least mean square (LMS) algorithm which produces a faster rate of convergence than the Bayesian estimator, at the expense of a finite misadjustment. This estimate is similar in form to an LMS transversal filter, and is computationally more efficient than the Bayesian estimate. Results show an overall reduction of 2.0 to 3.4% in word error rate due to adaptation for a set of 11 speakers from the DARPA resource management task.> William A. Rozzi, Richard M. Stern |
ICASSP | 2 |
| 1990 | Environmental robustness in automatic speech recognitionabstractInitial efforts to make Sphinx, a continuous-speech speaker-independent recognition system, robust to changes in the environment are reported. To deal with differences in noise level and spectral tilt between close-talking and desk-top microphones, two novel methods based on additive corrections in the cepstral domain are proposed. In the first algorithm, the additive correction depends on the instantaneous SNR of the signal. In the second technique, expectation-maximization techniques are used to best match the cepstral vectors of the input utterances to the ensemble of codebook entries representing a standard acoustical ambience. Use of the algorithms dramatically improves recognition accuracy when the system is tested on a microphone other than the one on which it was trained.> Alex Acero, Richard M. Stern |
ICASSP | 2 |
| 1990 | Acoustical pre-processing for robust spoken language systemsabstractIn this paper we report our initial efforts to make SPHINX, the CMU continuous-speech speaker-independent recognition system, robust to changes in the environment. To deal with differences in noise level and spectral tile between close-talking and desktop microphones, we propose two novel methods based on additive corrections in the cepstral domain. In the first algorithm, the additive correction depends on the instantaneous SNR of the signal. In the second technique, EM techniques are used to best match the cepstral vectors of the input utterances to the ensemble of codebook algorithms dramatically improves recognition accuracy when the system is tested on a microphone other than the one on which it was trained. Alex Acero, Richard M. Stern |
ICSLP | 2 |
| 1988 | Parsing spoken phrases despite missing wordsabstractThe authors compare the recognition accuracy obtained in forming sentence hypotheses using island-driven sentence parsers with parsers that hypothesize sentences in left-to-right fashion. Island-driven parsing algorithms are especially valuable in speech recognition systems because they can function more gracefully when not all of the correct words of an utterance were produced by the word hypothesizer. The inputs to both types of parsers consist of a lattice of candidate words, which are identified by their begin and end times, and the quality of the acoustic-phonetic match. Grammatical constraints are expressed by trigram models of sequences of lexical and semantic labels. The authors found that the island-driven parser produces parses with a higher percentage of correct words than the left-to-right parser is all cases considered. When the quality of the input lattices is extremely high, differences in parsing accuracy can be directly attributed to the superior ability of the island-driven parser to handle lattices with missing words. With lower-quality input, the accuracy of both types of parsers degrades, which is due to the creation of garden-path hypotheses and a lack of good words to serve as seeds for island formation.> Wayne H. Ward, Alex Hauptmann 0001, Richard M. Stern, Thomas Chanak |
ICASSP | 3 |
| 1987 | Sentence parsing with weak grammatical constraintsabstractThis paper compares the recognition accuracy obtained in forming sentence hypotheses using several parsers based on different types of weak statistical models of syntax and semantics. The inputs to the parsers were word hypotheses generated from simulated acoustic-phonetic labels. Grammatical constraints are expressed by trigram models of sequences of lexical or semantic labels, or by a finite-state network of the semantic labels. When the input to the parser is of high quality, the more restrictive trigram models were found to perform as well as or better than the finite-state language model. The more restrictive trigram and network models of language produce better recognition accuracy when all correct words are actually hypothesized, but strong constraints can degrade performance when many correct words are missing from the parser input. Richard M. Stern, Wayne H. Ward, Alex Hauptmann 0001, Juan Leon |
ICASSP | 1 |
| 1984 | Unsupervised adaptation to new speakers in feature-based letter recognitionabstractThis paper describes two new methods by which the CMU feature-based recognition system can learn the acoustical characteristics of individual speakers without feedback from the user. We have previously described how the system uses MAP techniques to update its estimates of the mean values of features used by the classifier in recognizing the letters of the English alphabet on the basis of a priori information and labelled observations. In the first of the new procedures described in this paper the system assumes a correct decision every time it classifies a new utterance with a sufficiently high confidence level. In the second new procedure the system adjusts its estimates of the means on the basis of their correlation with the average values of the features over all utterances. Experiments were conducted on two confusable sets of letters using both speaker adaptation procedures. In each case classification performance using the unsupervised estimation procedures could equal that obtained using speaker adaptation with feedback from the user, although which method provided the better performance depended on which set of letters was being classified. Moshé J. Lasry, Richard M. Stern |
ICASSP | 2 |
| 1984 | Fast Computation of the Difference of Low-Pass TransformabstractThis paper defines the difference of low-pass (DOLP) transform and describes a fast algorithm for its computation. The DOLP is a reversible transform which converts an image into a set of bandpass images. A DOLP transform is shown to require O(N2) multiplies and produce O(N log(N)) samples from an N sample image. When Gaussian low-pass filters are used, the result is a set of images which have been convolved with difference of Gaussian (DOG) filters from an exponential set of sizes. A fast computation technique based on ``resampling'' is described and shown to reduce the DOLP transform complexity to O(N log(N)) multiplies and O(N) storage locations. A second technique, ``cascaded convolution with expansion,'' is then defined and also shown to reduce the computational cost to O(N log(N)) multiplies. Combining these two techniques yields an algorithm for a DOLP transform that requires O(N) storage cells and requires O(N) multiplies. James L. Crowley, Richard M. Stern |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1984 | A Posteriori Estimation of Correlated Jointly Gaussian Mean VectorsabstractThis paper describes the use of maximum a posteriori probability (MAP) techniques to estimate the mean values of features used in statistical pattern classification problems, when these mean feature values from the various decision classes are jointly Gaussian random vectors that are correlated across the decision classes. A set of mathematical formalisms is proposed and used to derive closed-form expressions for the estimates of the class-conditional mean vectors, and for the covariance matrix of the errors of these estimates. Finally, the performance of these algorithms is described for the simple case of a two-class one-feature pattern recognition problem, and compared to the performance of classical estimators that do not exploit the class-to-class correlations of the features' mean values. Moshé J. Lasry, Richard M. Stern |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1983 | Feature-based speaker-independent recognition of isolated english lettersabstractFEATURE is a speaker-independent isolated letter recognition system. The system performs a series of feature measurements on an input utterance, then classifies the sound as one of 26 English letters using statistical pattern classification techniques. Performance was evaluated for 10 male and 10 female speakers. For each speaker tested, the system was trained on 4 tokens of each letter provided by the remaining 19 speakers. The average error rate was 10.5%. The system can be used in either a speaker-independent or dynamic adaptation mode. In this latter mode, the user provides feedback when an error is made, and the system changes the statistical parameters that are used during classification. In this way, the system dynamically adapts to the speech patterns of the current user. The use of tuning produced a decrease in the error rate from 10.5% to 6.2%, averaged across the 20 speakers. FEATURE is significant because it is able to perform fine phonetic distinctions (such as between the letters B-D-E, P-T-G, V-Z, M-N, J-K, I-R) in a speaker- independent mode. Ronald A. Cole, Richard M. Stern, Michael S. Phillips 0001, Scott M. Brill, Andrew P. Pilant, Philippe Specker |
ICASSP | 2 |
| 1983 | Dynamic speaker adaptation for isolated letter recognition using MAP estimationabstractA dynamic speaker-adaptation algorithm for the C-MU feature-based isolated letter recognition system, FEATURE, is described. The algorithm, based on maximum a posteriori probability estimation techniques, uses the labelled observations input thus far to the classifier, as well as the a priori correlations of the features within and across the various letters or sets of letters (classes). The probability density functions (pdf) of all the classes are updated simultaneously rather than on a class-by-class basis so that the pdf of a given class is updated before any observation from that class has been input. A significant improvement in the recognition performance was observed for different vocabularies as the system tuned to the the characteristics of a new speaker. Finally, the algorithm was compared to simpler forms of dynamic adaptation. It produced a faster decrease of the error rate than the other tuning procedures. After a small number of iterations, however, the various procedures yielded similar results. Richard M. Stern, Moshé J. Lasry |
ICASSP | 1 |