VLDB 2026 Research / reviewers in the wild / expert
Hynek Hermansky
dblp:70/158
· DBLP profile ↗
200ranked-venue papers
30as first author
6since 2021 · last 2025
0000-0001-8032-4811ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 181 · 27 first-author · 6 since 2021Artificial intelligence and machine learning · 114 · 10 first-author · 4 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
Man-Wai Mak, Johan Rohdin, Kong-Aik Lee, Hynek Hermansky |
INTERSPEECH | 5 |
| 2023 | Importance of Different Temporal Modulations of Speech: a Tale of two PerspectivesabstractHow important are different temporal speech modulations for speech recognition? We answer this question from two complementary perspectives. Firstly, we quantify the amount of phonetic information in the modulation spectrum of speech by computing the mutual information between temporal modulations with frame-wise phoneme labels. Looking from another perspective, we ask - which speech modulations an Automatic Speech Recognition (ASR) system prefers for its operation. Data-driven weights are learned over the modulation spectrum and optimized for an end-to-end ASR task. Both methods unanimously agree that speech information is mostly contained in slow modulation. Maximum mutual information occurs around 3-6 Hz which also happens to be the range of modulations most preferred by the ASR. In addition, we show that the incorporation of this knowledge into ASRs significantly reduces their dependency on the amount of training data. Samik Sadhu, Hynek Hermansky |
ICASSP | 2 |
| 2022 | Complex Frequency Domain Linear Prediction: A Tool to Compute Modulation Spectrum of Speech
Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 2 |
| 2022 | Dealing with Unknowns in Continual Learning for End-to-end Automatic Speech Recognition
Martin Sustek, Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 3 |
| 2021 | Radically Old Way of Computing Spectra: Applications in End-to-End ASRabstractWe propose a technique to compute spectrograms using Frequency Domain Linear Prediction (FDLP) that uses all-pole models to fit the squared Hilbert envelope of speech in different frequency sub-bands. The spectrogram of a complete speech utterance is computed by overlap-add of contiguous all-pole model responses. A long context window of 1.5 seconds allows us to capture the low frequency temporal modulations of speech in the spectrogram. For an end-to-end automatic speech recognition task, the FDLP spectrogram performs on par with the standard mel spectrogram features for clean read speech training and test data. For more realistic speech data with train-test domain mismatches or reverberations, FDLP spectrogram shows up to 25% and 22% relative WER improvements over mel spectrogram respectively. Samik Sadhu, Hynek Hermansky |
Interspeech | 2 |
| 2021 | Two-Stage Augmentation and Adaptive CTC Fusion for Improved Robustness of Multi-Stream end-to-end ASRabstractPerformance degradation of an Automatic Speech Recognition (ASR) system is commonly observed when the test acoustic condition is different from training. Hence, it is essential to make ASR systems robust against various environmental distortions, such as background noises and reverberations. In a multi-stream paradigm, improving robustness takes account of handling a variety of unseen single-stream conditions and inter-stream dynamics. Previously, a practical two-stage training strategy was proposed within multi-stream end-to-end ASR, where Stage-2 formulates the multi-stream model with features from Stage-1 Universal Feature Extractor (UFE). In this paper, as an extension, we introduce a two-stage augmentation scheme focusing on mismatch scenarios: Stage-1 Augmentation aims to address single-stream input varieties with data augmentation techniques; Stage-2 Time Masking applies temporal masks on UFE features of randomly selected streams to simulate diverse stream combinations. During inference, we also present adaptive Connectionist Temporal Classification (CTC) fusion with the help of hierarchical attention mechanisms. Experiments have been conducted on two datasets, DIRHA and AMI, as a multi-stream scenario. Compared with the previous training strategy, substantial improvements are reported with relative word error rate reductions of 29.7 - 59.3% across several unseen stream combinations. Gregory Sell, Hynek Hermansky |
SLT | 3 |
| 2020 | A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech RecognitionabstractThe multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic speech recognition, where parallel encoders aim to capture diverse information followed by a stream-level fusion based on attention mechanisms to combine the different views. However, with an increasing number of streams resulting in an increasing number of encoders, the previous approach could require substantial memory and massive amounts of parallel data for joint training. In this work, we propose a practical two-stage training scheme. Stage-1 is to train a Universal Feature Extractor (UFE), where encoder outputs are produced from a single-stream model trained with all data. Stage-2 formulates a multi-stream scheme intending to solely train the attention fusion module using the UFE features and pretrained components from Stage-1. Experiments have been conducted on two datasets, DIRHA and AMI, as a multi-stream scenario. Compared with our previous method, this strategy achieves relative word error rate reductions of 8.2-32.4%, while consistently outperforming several conventional combination methods. Gregory Sell, Xiaofei Wang 0007, Shinji Watanabe 0001, Hynek Hermansky |
ICASSP | 5 |
| 2020 | An Alternative to MFCCs for ASR
Pegah Ghahremani, Hossein Hadian, Daniel Povey, Hynek Hermansky, Sanjeev Khudanpur |
INTERSPEECH | 4 |
| 2020 | Continual Learning in Automatic Speech Recognition
Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 2 |
| 2020 | Multi-Stream End-to-End Speech RecognitionabstractAttention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved great success by utilizing both architectures during multi-task training and joint decoding. In this article, we present a multi-stream framework based on joint CTC/Attention E2E ASR with parallel streams represented by separate encoders aiming to capture diverse information. On top of the regular attention networks, the Hierarchical Attention Network (HAN) is introduced to steer the decoder toward the most informative encoders. A separate CTC network is assigned to each stream to force monotonic alignments. Two representative framework have been proposed and discussed, which are Multi-Encoder Multi-Resolution (MEM-Res) framework and Multi-Encoder Multi-Array (MEM-Array) framework, respectively. In MEM-Res framework, two heterogeneous encoders with different architectures, temporal resolutions and separate CTC networks work in parallel to extract complementary information from same acoustics. Experiments are conducted on Wall Street Journal (WSJ) and CHiME-4, resulting in relative Word Error Rate (WER) reduction of 18.0-32.1% and the best WER of 3.6% in the WSJ eval92 test set. The MEM-Array framework aims at improving the far-field ASR robustness using multiple microphone arrays which are activated by separate encoders. Compared with the best single-array results, the proposed framework has achieved relative WER reduction of 3.7% and 9.7% in AMI and DIRHA multi-array corpora, respectively, which also outperforms conventional fusion strategies. Xiaofei Wang 0007, Sri Harish Reddy Mallidi, Shinji Watanabe 0001, Takaaki Hori, Hynek Hermansky |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2019 | Deriving Spectro-temporal Properties of Hearing from Speech DataabstractHuman hearing and human speech are intrinsically tied together, as the properties of speech almost certainly developed in order to be heard by human ears. As a result of this connection, it has been shown that certain properties of human hearing are mimicked within data-driven systems that are trained to understand human speech. In this paper, we further explore this phenomenon by measuring the spectro-temporal responses of data-derived filters in a front-end convolutional layer of a deep network trained to classify the phonemes of clean speech. The analyses show that the filters do indeed exhibit spectro-temporal responses similar to those measured in mammals, and also that the filters exhibit an additional level of frequency selectivity, similar to the processing pipeline assumed within the Articulation Index. Lucas Ondel Yang, Gregory Sell, Hynek Hermansky |
ICASSP | 4 |
| 2019 | M-vectors: Sub-band Based Energy Modulation Features for Multi-stream Automatic Speech RecognitionabstractIn this paper, we propose a novel method to capture energy modulations from different frequency bands in speech into frame-level feature vectors, called Modulation-vectors or M-vectors, for use in Automatic Speech Recognition (ASR) systems. We show that in different multi-stream setups, with parallel streams for M-vectors and the popular Mel-frequency Cepstral Coefficient (MFCC) features, we can realize a boost in word recognition performance of end-to-end systems by ≈ 5%, and that of a monophone and triphone HMM-GMM ASR system by ≈ 18% and ≈ 16% respectively over using the traditional MFCC features. Samik Sadhu, Hynek Hermansky |
ICASSP | 3 |
| 2019 | Stream Attention-based Multi-array End-to-end Speech RecognitionabstractAutomatic Speech Recognition (ASR) using multiple microphone arrays has achieved great success in the far-field robustness. Taking advantage of all the information that each array shares and contributes is crucial in this task. Motivated by the advances of joint Connectionist Temporal Classification (CTC)/attention mechanism in the End-to-End (E2E) ASR, a stream attention-based multi-array framework is proposed in this work. Microphone arrays, acting as information streams, are activated by separate encoders and decoded under the instruction of both CTC and attention networks. In terms of attention, a hierarchical structure is adopted. On top of the regular attention networks, stream attention is introduced to steer the decoder toward the most informative encoders. Experiments have been conducted on AMI and DIRHA multi-array corpora using the encoder-decoder architecture. Compared with the best single-array results, the proposed framework has achieved relative Word Error Rates (WERs) reduction of 3.7% and 9.7% in the two datasets, respectively, which is better than conventional strategies as well. Xiaofei Wang 0007, Sri Harish Reddy Mallidi, Takaaki Hori, Shinji Watanabe 0001, Hynek Hermansky |
ICASSP | 6 |
| 2019 | Towards Automatic Methods to Detect Errors in Transcriptions of Speech RecordingsabstractThis work explores different methods to detect errors in transcriptions of speech recordings. We artificially corrupt well transcribed speech transcriptions with three types of errors: substitution, insertion and deletion on TIMIT phonemic transcriptions and WSJ word transcriptions. First, we use Bayesian model selection method by comparing the log-likelihoods from alignment and phone recognizer, a final score is computed to make decision. In this method, we consider two models, Bayesian Hidden Markov Model (HMM) and a Variational Auto-Encoder (VAE) combined with a HMM. Alternately, we build a biased ASR system with language models trained on individual transcriptions, detection decision is based on Levenshtein distance (LD) between transcription and oracle path from decoded lattice. We evaluate the methods of detecting errors in corrupted TIMIT transcription, the best result (either using model selection with VAE model or biased ASR) achieves 7% equal error rate on the Detection Error Tradeoff (DET) curve; we also evaluate the methods of detecting errors in corrupted WSJ transcriptions, and the best result (using biased ASR) achieves 3% equal error rate. Jinyi Yang, Lucas Ondel Yang, Vimal Manohar, Hynek Hermansky |
ICASSP | 4 |
| 2019 | Performance Monitoring for End-to-End Speech RecognitionabstractMeasuring performance of an automatic speech recognition (ASR) system without ground-truth could be beneficial in many scenarios, especially with data from unseen domains, where performance can be highly inconsistent.In conventional ASR systems, several performance monitoring (PM) techniques have been well-developed to monitor performance by looking at triphone posteriors or pre-softmax activations from neural network acoustic modeling.However, strategies for monitoring more recently developed end-to-end ASR systems have not yet been explored, and so that is the focus of this paper.We adapt previous PM measures (Entropy, M-measure and Autoencoder) and apply our proposed RNN predictor in the end-toend setting.These measures utilize the decoder output layer and attention probability vectors, and their predictive power is measured with simple linear models.Our findings suggest that decoder-level features are more feasible and informative than attention-level probabilities for PM measures, and that Mmeasure on the decoder posteriors achieves the best overall predictive performance with an average prediction error 8.8%.Entropy measures and RNN-based prediction also show competitive predictability, especially for unseen conditions. Gregory Sell, Hynek Hermansky |
INTERSPEECH | 3 |
| 2019 | Modulation Vectors as Robust Feature Representation for ASR in Domain Mismatched Conditions
Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 2 |
| 2019 | Exploring Methods for the Automatic Detection of Errors in Manual TranscriptionabstractQuality of data plays an important role in most deep learning tasks.In the speech community, transcription of speech recording is indispensable.Since the transcription is usually generated artificially, automatically finding errors in manual transcriptions not only saves time and labors but benefits the performance of tasks that need the training process.Inspired by the success of hybrid automatic speech recognition using both language model and acoustic model, two approaches of automatic error detection in the transcriptions have been explored in this work.Previous study using a biased language model approach, relying on a strong transcription-dependent language model, has been reviewed.In this work, we propose a novel acoustic model based approach, focusing on the phonetic sequence of speech.Both methods have been evaluated on a completely real dataset, which was originally transcribed with errors and strictly corrected manually afterwards. Xiaofei Wang 0007, Jinyi Yang, Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 5 |
| 2019 | Coding and decoding of messages in human speech communication: Implications for machine recognition of speech
Hynek Hermansky |
Speech Commun. | 1 |
| 2019 | DNN-based performance measures for predicting error rates in automatic speech recognition and optimizing hearing aid parameters
Angel Mario Castro Martinez, Lukas Gerlach 0001, Guillermo Payá-Vayá, Hynek Hermansky, Jasper Ooster, Bernd T. Meyer |
Speech Commun. | 4 |
| 2018 | Stream Attention for Distributed Multi-Microphone Speech Recognition
Xiaofei Wang 0007, Hynek Hermansky |
INTERSPEECH | 3 |
| 2017 | Predicting error rates for unknown data in automatic speech recognitionabstractIn this paper we investigate methods to predict word error rates in automatic speech recognition in the presence of unknown noise types, which have not been seen during training. The performance measures operate on phoneme posteriorgrams that are obtained from neural nets. We compare average frame-wise entropy as a baseline measure to the mean temporal distance (M-Measure) and to the number of phonetic events. The latter is obtained by learning typical phoneme activations from clean training data, which are later applied as phoneme-specific matched filters to posteriorgrams (MaP). When exceeding a threshold after filtering, we register this as phonetic event. For test sets using 10 unknown noise types and a wide range of signal-to-noise ratios, we find M-Measure and MaP to produce predictions twice as accurate as the baseline measure. When excluding noise types that contain speech segments, a prediction error of 3.1% is achieved, compared to 15.0% for the baseline measure. Bernd T. Meyer, Sri Harish Reddy Mallidi, Hendrik Kayser, Hynek Hermansky |
ICASSP | 4 |
| 2016 | Novel neural network based fusion for multistream ASRabstractRobustness of automatic speech recognition (ASR) to acoustic mismatches can be improved by multistream framework. Frequently used approach to combine decisions from individual streams involve training large number of neural networks, one for each possible stream combination. In this work, we propose to simplify the fusion by replacing the large number of fusion networks with a single fusion network. During training of the proposed fusion network, features from a stream are randomly dropped out. At test time, corrupted streams are identified and dropped out to improve robustness. Using the proposed approach, we were able to achieve significant reduction in number of parameters, while remaining in less than 2.5 % relative degradation of conventional fusion technique. Furthermore, proposed fusion network is also applied in a multistream ASR system to improve noise robustness of Aurora4 speech recognition task. Noticeable improvements were observed over baseline systems (relative improvement of 9.2 % in microphone mismatch and 3.2 % in additive noise conditions). Sri Harish Reddy Mallidi, Hynek Hermansky |
ICASSP | 2 |
| 2016 | A new efficient measure for accuracy prediction and its application to multistream-based unsupervised adaptationabstractA new efficient measure for predicting estimation accuracy is proposed and successfully applied to multistream-based unsupervised adaptation of ASR systems to address data uncertainty when the ground-truth is unknown. The proposed measure is an extension of the M-measure, which predicts confidence in the output of a probability estimator by measuring the divergences of probability estimates spaced at specific time intervals. In this study, the M-measure was extended by considering the latent phoneme information, resulting in an improved reliability. Experimental comparisons carried out in a multistream-based ASR paradigm demonstrated that the extended M-measure yields a significant improvement over the original M-measure, especially under narrow-band noise conditions. Tetsuji Ogawa, Sri Harish Reddy Mallidi, Emmanuel Dupoux, Jordan Cohen, Naomi Feldman, Hynek Hermansky |
ICPR | 6 |
| 2016 | A Framework for Practical Multistream ASR
Sri Harish Reddy Mallidi, Hynek Hermansky |
INTERSPEECH | 2 |
| 2016 | Assessing Speech Quality in Speech-Aware Hearing Aids Based on Phoneme Posteriorgrams
Constantin Spille, Hendrik Kayser, Hynek Hermansky, Bernd T. Meyer |
INTERSPEECH | 3 |
| 2016 | Performance monitoring for automatic speech recognition in noisy multi-channel environmentsabstractIn many applications of machine listening it is useful to know how well an automatic speech recognition system will do before the actual recognition is performed. In this study we investigate different performance measures with the aim of predicting word error rates (WERs) in spatial acoustic scenes in which the type of noise, the signal-to-noise ratio, parameters for spatial filtering, and the amount of reverberation are varied. All measures under consideration are based on phoneme posteriorgrams obtained from a deep neural net. While frame-wise entropy exhibits only medium predictive power for factors other than additive noise, we found the medium temporal distance between posterior vectors (M-Measure) as well as matched phoneme filters (MaP) to exhibit excellent correlations with WER across all conditions. Since our results were obtained with simulated behind-the-ear hearing aid signals, we discuss possible applications for speech-aware hearing devices. Bernd T. Meyer, Sri Harish Reddy Mallidi, Angel Mario Castro Martinez, Guillermo Payá-Vayá, Hendrik Kayser, Hynek Hermansky |
SLT | 6 |
| 2015 | Robust speech recognition in unknown reverberant and noisy conditionsabstractIn this paper, we describe our work on the ASpIRE (Automatic Speech recognition In Reverberant Environments) challenge, which aims to assess the robustness of automatic speech recognition (ASR) systems. The main characteristic of the challenge is developing a high-performance system without access to matched training and development data. While the evaluation data are recorded with far-field microphones in noisy and reverberant rooms, the training data are telephone speech and close talking. Our approach to this challenge includes speech enhancement, neural network methods and acoustic model adaptation, We show that these techniques can successfully alleviate the performance degradation due to noisy audio and data mismatch. Roger Hsiao, Jeff Z. Ma, William Hartmann, Martin Karafiát, Frantisek Grézl, Lukás Burget, Igor Szöke, Jan Cernocký, Shinji Watanabe 0001, Zhuo Chen 0006, Sri Harish Reddy Mallidi, Hynek Hermansky, Stavros Tsakalidis, Richard M. Schwartz |
ASRU | 12 |
| 2015 | Uncertainty estimation of DNN classifiersabstractNew efficient measures for estimating uncertainty of deep neural network (DNN) classifiers are proposed and successfully applied to multistream-based unsupervised adaptation of ASR systems to address uncertainty derived from noise. The proposed measure is the error from associative memory models trained on outputs of a DNN. In the present study, an attempt is made to use autoencoders for remembering the property of data. Another measure proposed is an extension of the M-measure, which computes the divergences of probability estimates spaced at specific time intervals. The extended measure results in an improved reliability by considering the latent information of phoneme duration. Experimental comparisons carried out in a multistream-based ASR paradigm demonstrates that the proposed measures yielded improvements over the multistyle trained system and system selected based on existing measures. Fusion of the proposed measures achieved almost the same performance as the oracle system selection. Sri Harish Reddy Mallidi, Tetsuji Ogawa, Hynek Hermansky |
ASRU | 3 |
| 2015 | Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial WorkshopabstractA group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which the estimator was trained. The paper describes the problem and summarizes approaches that were taken by the group1. Hynek Hermansky, Lukás Burget, Jordan Cohen, Emmanuel Dupoux, Naomi Feldman, John Godfrey, Sanjeev Khudanpur, Matthew Maciejewski, Sri Harish Reddy Mallidi, Anjali Menon, Tetsuji Ogawa, Vijayaditya Peddinti, Richard C. Rose, Richard M. Stern, Matthew Wiesner, Karel Veselý |
ICASSP | 1 |
| 2015 | Autoencoder based multi-stream combination for noise robust speech recognition
Sri Harish Reddy Mallidi, Tetsuji Ogawa, Karel Veselý, Phani S. Nidadavolu, Hynek Hermansky |
INTERSPEECH | 5 |
| 2015 | DNN derived filters for processing of modulation spectrum of speechabstractWe propose a novel approach to design modulation frequency filters for the first stage processing of critical band spectrum of speech using deep neural network (DNN). These filters replace conventional modulation frequency filters currently used in state-of-the-art BUT speech recognition system and yield about 10% relative improvement in phoneme recognition accuracy. The resulting filters are consistent with some known temporal properties of higher levels of mammalian auditory processing and suggest more efficient scheme for pre-processing of speech for ASR. Index Terms: deep neural network, convolutive layer, modulation filters, mammalian auditory processing Jan Pesán, Lukás Burget, Hynek Hermansky, Karel Veselý |
INTERSPEECH | 3 |
| 2014 | Featherweight phonetic keyword search for conversational speechabstractThe point process model (PPM) for keyword search is a phonetic event-driven approach that provides a whole-word focused alternative to fast lattice matching techniques. Recent efforts in PPMs have been focused on improved model estimation techniques and efficient search algorithms, but past evaluations have been limited to searching relatively easy scripted corpora for simple unigram queries, preventing comprehensive benchmarking against standard search methods. In this paper, we present techniques for score normalization and the processing of multi-word and out of training query terms as required by the 2006 NIST Spoken Term Detection (STD) evaluation, permitting the first comprehensive benchmark of PPM search technology against state-of-the-art word and phonetic-based search systems. We demonstrate PPM to be the fastest phonetic system while posting accuracies competitive with the best phonetic alternatives. Moreover, index construction time and size are better than any keyword search system entered in the NIST evaluation. Keith Kintzley, Aren Jansen, Hynek Hermansky |
ICASSP | 3 |
| 2014 | A long, deep and wide artificial neural net for robust speech recognition in unknown noiseabstractA long deep and wide artificial neural net (LDWNN) with multiple ensemble neural nets for individual frequency subbands is proposed for robust speech recognition in unknown noise. It is assumed that the effect of arbitrary additive noise on speech recognition can be approximated by white noise (or speech-shaped noise) of similar level across multiple frequency subbands. The ensemble neural nets are trained in clean and speech-shaped noise at 20, 10, and 5 dB SNR to accommodate noise of different levels, followed by a neural net trained to select the most suitable neural net for optimum information extraction within a frequency subband. The posteriors from multiple frequency subbands are fused by another neural net to give a more reliable estimation. Experimental results show that the subband ensemble net adapts well to unknow noise. Feipeng Li, Phani S. Nidadavolu, Hynek Hermansky |
INTERSPEECH | 3 |
| 2014 | Principal components of auditory spectro-temporal receptive fields
Nagaraj Mahajan, Nima Mesgarani, Hynek Hermansky |
INTERSPEECH | 3 |
| 2014 | Evaluating speech features with the minimal-pair ABX task (II): resistance to noiseabstractThe Minimal-Pair ABX (MP-ABX) paradigm has been pro-posed as a method for evaluating speech features for zero-resource/unsupervised speech technologies. We apply it in a phoneme discrimination task on the Articulation Index corpus to evaluate the resistance to noise of various speech features. In Experiment 1, we evaluate the robustness to additive noise at different signal-to-noise ratios, using car and babble noise from the Aurora-4 database and white noise. In Experiment 2, we ex-amine the robustness to different kinds of convolutional noise. In both experiments we consider two classes of techniques to induce noise resistance: smoothing of the time-frequency rep-resentation and short-term adaptation in the time-domain. We consider smoothing along the spectral axis (as in PLP) and along the time axis (as in FDLP). For short-term adaptation in the time-domain, we compare the use of a static compressive non-linearity followed by RASTA filtering to an adaptive com-pression scheme. Index Terms: noise resistance, zero-resource, speech features, evaluation framework, minimal-pair ABX task Thomas Schatz, Vijayaditya Peddinti, Xuan-Nga Cao, Francis R. Bach, Hynek Hermansky, Emmanuel Dupoux |
INTERSPEECH | 5 |
| 2014 | Robust Feature Extraction Using Modulation Filtering of Autoregressive ModelsabstractSpeaker and language recognition in noisy and degraded channel conditions continue to be a challenging problem mainly due to the mismatch between clean training and noisy test conditions. In the presence of noise, the most reliable portions of the signal are the high energy regions which can be used for robust feature extraction. In this paper, we propose a front end processing scheme based on autoregressive (AR) models that represent the high energy regions with good accuracy followed by a modulation filtering process. The AR model of the spectrogram is derived using two separable time and frequency AR transforms. The first AR model (temporal AR model) of the sub-band Hilbert envelopes is derived using frequency domain linear prediction (FDLP). This is followed by a spectral AR model applied on the FDLP envelopes. The output 2-D AR model represents a low-pass modulation filtered spectrogram of the speech signal. The band-pass modulation filtered spectrograms can further be derived by dividing two AR models with different model orders (cut-off frequencies). The modulation filtered spectrograms are converted to cepstral coefficients and are used for a speaker recognition task in noisy and reverberant conditions. Various speaker recognition experiments are performed with clean and noisy versions of the NIST-2010 speaker recognition evaluation (SRE) database using the state-of-the-art speaker recognition system. In these experiments, the proposed front-end analysis provides substantial improvements (relative improvements of up to 25%) compared to baseline techniques. Furthermore, we also illustrate the generalizability of the proposed methods using language identification (LID) experiments on highly degraded high-frequency (HF) radio channels and speech recognition experiments on noisy data. Sriram Ganapathy, Sri Harish Reddy Mallidi, Hynek Hermansky |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Frequency offset correction in speech without detecting pitchabstractRadio-transmitted speech sometimes contains a residual frequency shift or offset, resulting from incorrect demodulation in single-sideband channels. Frequency-shifted speech can mask speaker identity and reduce intelligibility. Therefore, frequency offset will degrade the performance of downstream speech technologies. Existing offset correction methods require a pitch estimate of the speech signal, which is difficult in noisy radio channels. We present a new, automatic algorithm for detecting and correcting frequency offset, based on third-order modulation spectral analysis. Our method is remarkably simple and does not require pitch estimation. We provide derivations, examples, and a pilot study demonstrating how offset correction improves speaker verification for radio-transmitted speech. Pascal Clark, Sri Harish Reddy Mallidi, Aren Jansen, Hynek Hermansky |
ICASSP | 4 |
| 2013 | Mean temporal distance: Predicting ASR error from temporal properties of speech signalabstractExtending previous work on prediction of phoneme recognition error from unlabeled data that were corrupted by unpredictable factors, the current work investigates a simple but effective method of estimating ASR performance by computing a function M(Δt), which represents the mean distance between speech feature vectors evaluated over certain finite time interval, determined as a function of temporal distance Δt between the vectors. It is shown that M(Δt) is a function of signal-to-noise ratio of speech signal. Comparing M(Δt) curves, derived on data used for training of the classifier, and on test utterances, allows for predicting error on the test data. Another interesting observation is that M(Δt) remains approximately constant, as temporal separation Δt exceeds certain critical interval (about 200 ms), indicating the extent of coarticulation in speech sounds. Hynek Hermansky, Ehsan Variani, Vijayaditya Peddinti |
ICASSP | 1 |
| 2013 | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisitionabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies. Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001 |
ICASSP | 8 |
| 2013 | Weak top-down constraints for unsupervised acoustic model trainingabstractTypical supervised acoustic model training relies on strong top-down constraints provided by dynamic programming alignment of the input observations to phonetic sequences derived from orthographic word transcripts and pronunciation dictionaries. This paper investigates a much weaker form of top-down supervision for use in place of transcripts and dictionaries in the zero resource setting. Our proposed constraints, which can be produced using recent spoken term discovery systems, come in the form of pairs of isolated word examples that share the same unknown type. For each pair, we perform a dynamic programming alignment of the acoustic observations of the two constituent examples, generating an inventory of cross-speaker frame pairs that each provide evidence that the same subword unit model should account for them. We find these weak top-down constraints are capable of improving model speaker independence by up to 57% relative over bottom-up training alone. Aren Jansen, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 3 |
| 2013 | Effect of filter bandwidth and spectral sampling rate of analysis filterbank on automatic phoneme recognitionabstractIn this study we investigate the effect of filter bandwidth and spectral sampling rate of analysis filterbank for speech recognition. Two experiments are conducted to evaluate the performance of an automatic phoneme recognition system on clean speech and speech in noise as the filter bandwidth increases from 0.5 to 3.5 ERB and the spectral resolution changes from 1, 1.5, 2, 3, 4, to 6 samples per Bark. Results indicate that the optimum filter bandwidth varies for different speech sounds at different frequency ranges. A spectral sampling of 4 filters per Bark with the filter bandwidth being ≈ 1 ERB produces the best performance on average. Feipeng Li, Hynek Hermansky |
ICASSP | 2 |
| 2013 | Filter-bank optimization for Frequency Domain Linear PredictionabstractThe sub-band Frequency Domain Linear Prediction (FDLP) technique estimates autoregressive models of Hilbert envelopes of subband signals, from segments of discrete cosine transform (DCT) of a speech signal, using windows. Shapes of the windows and their positions on the cosine transform of the signal determine implied filtering of the signal. Thus, the choices of shape, position and number of these windows can be critical for the performance of the FDLP technique. So far, we have used Gaussian or rectangular windows. In this paper asymmetric cochlear-like filters are being studied. Further, a frequency differentiation operation, that introduces an additional set of parameters describing local spectral slope in each frequency sub-band, is introduced to increase the robustness of sub-band envelopes in noise. The performance gains achieved by these changes are reported in a variety of additive noise conditions, with an average relative improvement of 8.04% in phoneme recognition accuracy. Vijayaditya Peddinti, Hynek Hermansky |
ICASSP | 2 |
| 2013 | Developing a speaker identification system for the DARPA RATS projectabstractThis paper describes the speaker identification (SID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present results using multiple SID systems differing mainly in the algorithm used for voice activity detection (VAD) and feature extraction. We show that (a) unsupervised VAD performs as well supervised methods in terms of downstream SID performance, (b) noise-robust feature extraction methods such as CFCCs out-perform MFCC front-ends on noisy audio, and (c) fusion of multiple systems provides 24% relative improvement in EER compared to the single best system when using a novel SVM-based fusion algorithm that uses side information such as gender, language, and channel id. Oldrich Plchot, Spyridon Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas 0001, Bing Zhang 0004, Xinhui Zhou |
ICASSP | 8 |
| 2013 | Deep neural network features and semi-supervised training for low resource speech recognitionabstractWe propose a new technique for training deep neural networks (DNNs) as data-driven feature front-ends for large vocabulary continuous speech recognition (LVCSR) in low resource settings. To circumvent the lack of sufficient training data for acoustic modeling in these scenarios, we use transcribed multilingual data and semi-supervised training to build the proposed feature front-ends. In our experiments, the proposed features provide an absolute improvement of 16% in a low-resource LVCSR setting with only one hour of in-domain training data. While close to three-fourths of these gains come from DNN-based features, the remaining are from semi-supervised training. Samuel Thomas 0001, Michael L. Seltzer, Kenneth Church 0001, Hynek Hermansky |
ICASSP | 4 |
| 2013 | Text-to-speech inspired duration modeling for improved whole-word acoustic modelsabstractIn the construction of whole-word acoustic models, we have previously demonstrated substantial gains by using MAP estimation to introduce a simple prior model of phonetic timing. Based solely on the word’s phonetic (dictionary) pronunciation, this simple model included no information about the individual durations of constituent phones. However, the problem of modeling segmental duration has long been studied in the textto-speech (TTS) community. We draw upon this work to develop a classification and regression tree (CART) approach for constructing prior models of phonetic timing which considers factors such as syllable stress, syllable position, adjacent phone class and voicing. This improved prior model closes 33% of the gap in keyword spotting performance between highly supervised whole-word models and those estimated without any examples. Keith Kintzley, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 3 |
| 2013 | Improvements in language identification on the RATS noisy speech corpus
Jeff Z. Ma, Bing Zhang 0004, Spyridon Matsoukas, Sri Harish Reddy Mallidi, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 6 |
| 2013 | Robust speaker recognition using spectro-temporal autoregressive modelsabstractSpeaker recognition in noisy environments is challenging when there is a mis-match in the data used for enrollment and veri-fication. In this paper, we propose a robust feature extraction scheme based on spectro-temporal modulation filtering using two-dimensional (2-D) autoregressive (AR) models. The first step is the AR modeling of the sub-band temporal envelopes by the application of the linear prediction on the sub-band dis-crete cosine transform (DCT) components. These sub-band en-velopes are stacked together and used for a second AR mod-eling step. The spectral envelope across the sub-bands is ap-proximated in this AR model and cepstral features are derived which are used for speaker recognition. The use of AR models emphasizes the focus on the high energy regions which are rel-atively well preserved in the presence of noise. The degree of modulation filtering is controlled using AR model order param-eter. Experiments are performed using noisy versions of NIST 2010 speaker recognition evaluation (SRE) data with a state-of-art speaker recognition system. In these experiments, the proposed features provide significant improvements compared to baseline features (relative improvements of 20 % in terms of equal error rate (EER) and 35 % in terms of miss rate at 10 % false alarm). Sri Harish Reddy Mallidi, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2013 | Stream selection and integration in multistream ASR using GMM-based performance monitoringabstractA moderately deep and rather wide artificial neural net is applied in phoneme recognition of noisy speech. The net is formed by first estimating posterior probabilities of phonemes in 21 band-limited streams covering the whole speech spectrum. These 21 band-limited streams are subdivided into three seven band-limited stream subsets, by differently sub-sampling the original 21 band-limited streams. In the second processing stage, all non-empty combinations of seven band-limited streams from each subset are formed as inputs to 127 artificial neural nets that are again trained to yield phoneme posteriors. In this way, 127 × 3 = 381 processing streams are formed. A novel technique for finding the best combination of the resulting 381 parallel processing streams, which uses the likelihood of a single-state Gaussian mixture model of the final classifier output is applied to selecting the most efficient streams. The technique is efficient in phoneme recognition of speech that is corrupted by realistic additive noise. Tetsuji Ogawa, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 3 |
| 2013 | Evaluating speech features with the minimal-pair ABX task: analysis of the classical MFC/PLP pipelineabstractWe present a new framework for the evaluation of speech representations in zero-resource settings, that extends and complements previous work by Carlin, Jansen and Hermansky [1].In particular, we replace their Same/Different discrimination task by several Minimal-Pair ABX (MP-ABX) tasks.We explain the analytical advantages of this new framework and apply it to decompose the standard signal processing pipelines for computing PLP and MFC coefficients.This method enables us to confirm and quantify a variety of well-known and not-so-well-known results in a single framework. Thomas Schatz, Vijayaditya Peddinti, Francis R. Bach, Aren Jansen, Hynek Hermansky, Emmanuel Dupoux |
INTERSPEECH | 5 |
| 2013 | Multi-stream recognition of noisy speech with performance monitoringabstractA prototype multi-stream system with a performance monitor for stream selection is proposed to recognize speech in un-known noise. The speech signal is decomposed into seven band-limited streams. Posterior probabilities of phonemes are estimated by a multi-layer perceptron (MLP) in each of these band-limited streams. Estimated posterior vectors of all 127 combinations (processing streams) of the seven band-limited streams form inputs to a second-stage MLP that esti-mates posterior probabilities of phonemes in each processing stream. A performance monitor is designed to predict the re-liability of individual processing streams based on the outputs from these streams. The top N streams that are least affected by noise are selected and their outputs are averaged to yield the final posterior probability vector used in Viterbi search for the best phoneme sequence. Experimental results show that the proposed technique is effective in dealing with noise. Index Terms: Multi-stream speech recognition, Performance monitoring Ehsan Variani, Feipeng Li, Hynek Hermansky |
INTERSPEECH | 3 |
| 2013 | Multistream Recognition of Speech: Dealing With Unknown UnknownsabstractThe paper discusses an approach for dealing with unexpected acoustic elements in speech. The approach is motivated by observations of human performance on such problems, which indicate the existence of multiple parallel processing streams in the human speech processing cognitive system, combined with the human ability to know when the correct information is being received. Some earlier relevant engineering approaches in multistream automatic recognition of speech (ASR) that aimed at processing of noisy speech and at dealing with unexpected out-of-vocabulary words are reviewed. The paper also reviews some currently active research in multistream ASR, focusing mainly on feedback-based techniques involving fusion of information between individual processing streams. The difference between the system behavior on its training data and during its operation is proposed as a substitute for the human ability of “knowing when knowing.” Most recent results indicate 9% relative improvement in error rates in phoneme recognition of high signal-to-noise ratio speech and as high as 30% relative improvements in moderate noise. Hynek Hermansky |
Proc. IEEE | 1 |
| 2013 | Perceptual Properties of Current Speech Recognition TechnologyabstractIn recent years, a number of feature extraction procedures for automatic speech recognition (ASR) systems have been based on models of human auditory processing, and one often hears arguments in favor of implementing knowledge of human auditory perception and cognition into machines for ASR. This paper takes a reverse route, and argues that the engineering techniques for automatic recognition of speech that are already in widespread use are often consistent with some well-known properties of the human auditory system. Hynek Hermansky, Jordan Cohen, Richard M. Stern |
Proc. IEEE | 1 |
| 2013 | Factor Analysis of Auto-Associative Neural Networks With Application in Speaker VerificationabstractAuto-associative neural network (AANN) is a fully connected feed-forward neural network, trained to reconstruct its input at its output through a hidden compression layer, which has fewer numbers of nodes than the dimensionality of input. AANNs are used to model speakers in speaker verification, where a speaker-specific AANN model is obtained by adapting (or retraining) the universal background model (UBM) AANN, an AANN trained on multiple held out speakers, using corresponding speaker data. When the amount of speaker data is limited, this adaptation procedure may lead to overfitting as all the parameters of UBM-AANN are adapted. In this paper, we introduce and develop the factor analysis theory of AANNs to alleviate this problem. We hypothesize that only the weight matrix connecting the last nonlinear hidden layer and the output layer is speaker-specific, and further restrict it to a common low-dimensional subspace during adaptation. The subspace is learned using large amounts of development data, and is held fixed during adaptation. Thus, only the coordinates in a subspace, also known as i-vector, need to be estimated using speaker-specific data. The update equations are derived for learning both the common low-dimensional subspace and the i-vectors corresponding to speakers in the subspace. The resultant i-vector representation is used as a feature for the probabilistic linear discriminant analysis model. The proposed system shows promising results on the NIST-08 speaker recognition evaluation (SRE), and yields a 23% relative improvement in equal error rate over the previously proposed weighted least squares-based subspace AANNs system. The experiments on NIST-10 SRE confirm that these improvements are consistent and generalize across datasets. Sri Garimella, Hynek Hermansky |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 16 |
| 2012 | Multilingual MLP features for low-resource LVCSR systemsabstractWe introduce a new approach to training multilayer perceptrons (MLPs) for large vocabulary continuous speech recognition (LVCSR) in new languages which have only few hours of annotated in-domain training data (for example, 1 hour of data). In our approach, large amounts of annotated out-of-domain data from multiple languages are used to train multilingual MLP systems without dealing with the different phoneme sets for these languages. Features extracted from these MLP systems are used to train LVCSR systems in the low-resource language similar to the Tandem approach. In our experiments, the proposed features provide a relative improvement of about 30% in an low-resource LVCSR setting with only one hour of training data. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
ICASSP | 3 |
| 2012 | Analysis of Temporal Resolution in Frequency Domain Linear Prediction
Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 2 |
| 2012 | Intrinsic Spectral Analysis for Zero and High Resource Speech RecognitionabstractThe constraints of the speech production apparatus imply that our vocalizations are approximately restricted to a lowdimensional manifold embedded in a high-dimensional space. Manifold learning algorithms provide a means to recover the approximate embedding from untranscribed data and enable use of the manifold’s intrinsic distance metric to characterize acoustic similarity for downstream automatic speech applications. In this paper, we consider a previously unevaluated nonlinear outof-sample extension for intrinsic spectral analysis (ISA), investigating its performance in both unsupervised and supervised tasks. In the zero resource regime, where the lack of transcribed resources forces us to rely solely on the phonetic salience of the acoustic features themselves, ISA provides substantial gains relative to canonical acoustic front-ends. When large amounts of transcribed speech for supervised acoustic model training are also available, we find that the data-driven intrinsic spectrogram matches the performance of and is complementary to these signal processing derived counterparts. Index Terms: intrinsic spectral analysis, manifold learning, speech recognition, zero resource Aren Jansen, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2012 | Inverting the Point Process Model for Fast Phonetic Keyword SearchabstractNormally, we represent speech as a long sequence of frames and model the keyword with a relatively small set of parameters, commonly with a hidden Markov model (HMM). However, since the input speech is much longer than the keyword, suppose instead that we represent the speech as a relatively sparse set of impulses (roughly one per phoneme) and model the keyword as a filter-bank where each filter’s impulse response relates to the likelihood of a phone at a given position within a word. Evaluating keyword detections can then be seen as a convolution of an impulse train with an array of filters. This view enables huge speedups; runtime no longer depends on the frame rate and is instead linear in the number of events (impulses). We apply this intuition to redesign the runtime engine behind the point process model for keyword spotting. We demonstrate impressive real-time speedups (500,000x faster than real-time) with minimal loss in search accuracy. Keith Kintzley, Aren Jansen, Kenneth Church 0001, Hynek Hermansky |
INTERSPEECH | 4 |
| 2012 | MAP Estimation of Whole-Word Acoustic Models with Dictionary PriorsabstractThe intrinsic advantages of whole-word acoustic modeling are offset by the problem of data sparsity. To address this, we present several parametric approaches to estimating intra-word phonetic timing models under the assumption that relative timing is independent of word duration. We show evidence that the timing of phonetic events is well described by the Gaussian distribution. We explore the construction of models in the absence Keith Kintzley, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 3 |
| 2012 | Phone recognition in critical bands using sub-band temporal modulations
Feipeng Li, Sri Harish Reddy Mallidi, Hynek Hermansky |
INTERSPEECH | 3 |
| 2012 | Data-driven Posterior Features for Low Resource Speech Recognition ApplicationsabstractIn low resource settings, with very few hours of training data, state-of-the-art speech recognition systems that require large amounts of task specific training data perform very poorly. We address this issue by building data-driven speech recognition front-ends on significant amounts of task independent data from different languages and genres collected in similar acoustic conditions as the data in the low resource scenario. We show that features derived from these trained front-ends perform significantly better and can alleviate the effect of reduced task specific training data in low resource settings. The proposed features provide a absolute improvement of about 12 % (18 % relative) in an low-resource LVCSR setting with only one hour of training data. We also demonstrate the usefulness of these features for zero-resource speech applications like spoken term discovery, which operate without any transcribed speech to train systems. The proposed features provide significant gains over conventional acoustic features on various information retrieval metrics for this task. Index Terms: Low-resource speech recognition, spoken term discovery, posterior features. Samuel Thomas 0001, Sriram Ganapathy, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 4 |
| 2012 | Acoustic and Data-driven Features for Robust Speech Activity DetectionabstractIn this paper we evaluate different features for speech activity detection (SAD). Several signal processing techniques are used to derive acoustic features that capture attributes of speech useful in differentiating speech segments in noise. The acoustic features include short-term spectral features, long-term modulation features both derived using Frequency Domain Linear Prediction (FDLP), and joint spectro-temporal features extracted using 2D filters on a cortical representation of speech. Posteriors of speech and non-speech from a trained multi-layer perceptron are also used as data-driven features for this task. These feature extraction techniques form part of an elaborate feature extraction front-end where information spanning several hundreds of milliseconds of the signal are used along with heteroscedastic linear discriminant analysis for dimensionality reduction. Processed feature outputs from the proposed front-end are used to train SAD systems based on Gaussian mixture models for processing of speech from multiple languages transmitted over noisy radio communication channels under the ongoing DARPA Robust Automatic Transcription of Speech (RATS) program. The proposed front-end performs significantly better than standard acoustic feature extraction techniques in these noisy conditions. Samuel Thomas 0001, Sri Harish Reddy Mallidi, Thomas Janu, Hynek Hermansky, Nima Mesgarani, Xinhui Zhou, Shihab A. Shamma, Tim Ng, Bing Zhang 0004, Long Nguyen 0001, Spyridon Matsoukas |
INTERSPEECH | 4 |
| 2012 | Estimating Classifier Performance in Unknown NoiseabstractWe propose and investigate a non-parametric method for identifying regions of speech that have unexpected distortions not seen in the training data. The method does not require knowledge of correct labels and relies only on divergence between statistics of the test and training data. Our experiments show that the proposed method re-quires a relatively small amount of test data of the order of several seconds to stabilize, and correlates well with recognition error observed on the test data. Index Terms: Unexpected distortions, confidence esti-mation, machine recognition of speech Ehsan Variani, Hynek Hermansky |
INTERSPEECH | 2 |
| 2012 | Beyond Novelty Detection: Incongruent Events, When General and Specific Classifiers DisagreeabstractUnexpected stimuli are a challenge to any machine learning algorithm. Here, we identify distinct types of unexpected events when general-level and specific-level classifiers give conflicting predictions. We define a formal framework for the representation and processing of incongruent events: Starting from the notion of label hierarchy, we show how partial order on labels can be deduced from such hierarchies. For each event, we compute its probability in different ways, based on adjacent levels in the label hierarchy. An incongruent event is an event where the probability computed based on some more specific level is much smaller than the probability computed based on some more general level, leading to conflicting predictions. Algorithms are derived to detect incongruent events from different types of hierarchies, different applications, and a variety of data types. We present promising results for the detection of novel visual and audio objects, and new patterns of motion in video. We also discuss the detection of Out-Of- Vocabulary words in speech recognition, and the detection of incongruent events in a multimodal audiovisual scenario. Daphna Weinshall, Alon Zweig, Hynek Hermansky, Stefan Kombrink, Frank W. Ohl, Jörn Anemüller, Jörg-Hendrik Bach, Luc Van Gool, Fabian Nater, Tomás Pajdla, Michal Havlena, Misha Pavel |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Phase AutoCorrelation (PAC) features for noise robust speech recognition
Shajith Ikbal, Hemant Misra, Hynek Hermansky, Mathew Magimai-Doss |
Speech Commun. | 3 |
| 2012 | Regularized Auto-Associative Neural Networks for Speaker VerificationabstractAuto-Associative Neural Network (AANN) is a fully connected feed-forward neural network, trained to reconstruct its input at its output through a hidden compression layer. AANNs are used to model speakers in speaker verification, where a speaker-specific AANN model is obtained by adapting (or retraining) the Universal Background Model (UBM) AANN, an AANN trained on multiple held out speakers, using corresponding speaker data. When the amount of speaker data is limited, this adaptation procedure leads to overfitting. Additionally, the resultant speaker-specific parameters become noisy due to outliers in data. Thus, we propose to regularize the parameters of an AANN during speaker adaptation. A closed-form expression for updating the parameters is derived. Further, these speaker-specific AANN parameters are directly used as features in linear discriminant analysis (LDA)/probabilistic discriminant (PLDA) analysis based speaker verification system. The proposed speaker verification system outperforms the previously proposed weighted least squares (WLS) based AANN speaker verification system on NIST-08 speaker recognition evaluation (SRE). Moreover, the proposed speaker verification system obviates the need for an intermediate dimensionality reduction (or i-vector extraction) step. Sri Garimella, Sri Harish Reddy Mallidi, Hynek Hermansky |
IEEE Signal Process. Lett. | 3 |
| 2012 | Sparse Multilayer Perceptron for Phoneme RecognitionabstractThis paper introduces the sparse multilayer perceptron (SMLP) which jointly learns a sparse feature representation and nonlinear classifier boundaries to optimally discriminate multiple output classes. SMLP learns the transformation from the inputs to the targets as in multilayer perceptron (MLP) while the outputs of one of the internal hidden layers is forced to be sparse. This is achieved by adding a sparse regularization term to the cross-entropy cost and updating the parameters of the network to minimize the joint cost. On the TIMIT phoneme recognition task, SMLP-based systems trained on individual speech recognition feature streams perform significantly better than the corresponding MLP-based systems. Phoneme error rate of 19.6% is achieved using the combination of SMLP-based systems, a relative improvement of 3.0% over the combination of MLP-based systems. Garimella S. V. S. Sivaram, Hynek Hermansky |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Multilayer perceptron with sparse hidden outputs for phoneme recognitionabstractThis paper introduces the sparse multilayer perceptron (SMLP) which learns the transformation from the inputs to the targets as in multilayer perceptron (MLP) while the outputs of one of the internal hidden layers is forced to be sparse. This is achieved by adding a sparse regularization term to the cross-entropy cost and learning the parameters of the network to minimize the joint cost. On the TIMIT phoneme recognition task, the SMLP based system trained using perceptual linear prediction (PLP) features performs better than the conventional MLP based system. Furthermore, their combination yields a phoneme error rate of 21.2%, a relative improvement of 6.2% over the baseline. Garimella S. V. S. Sivaram, Hynek Hermansky |
ICASSP | 2 |
| 2011 | MLP based phoneme detectors for Automatic Speech RecognitionabstractPhoneme posterior probabilities estimated using Multi-Layer Perceptrons (MLPs) are extensively used both as acoustic scores and features for speech recognition. In this paper we explore a different application of these posteriors as phonetic event detectors for speech recognition. We show how these detectors can be built to reliably capture phonetic events in the acoustic signal by integrating both acoustic and phonetic information about sound classes. These event detectors are used along with Segmental Conditional Random Fields (SCRFs) to improve the performance of speech recognition systems on the Broadcast News task. Samuel Thomas 0001, Patrick Nguyen, Geoffrey Zweig, Hynek Hermansky |
ICASSP | 4 |
| 2011 | Speech recognitionwith segmental conditional random fields: A summary of the JHU CLSP 2010 Summer WorkshopabstractThis paper summarizes the 2010 CLSP Summer Workshop on speech recognition at Johns Hopkins University. The key theme of the workshop was to improve on state-of-the-art speech recognition systems by using Segmental Conditional Random Fields (SCRFs) to integrate multiple types of information. This approach uses a state of-the-art baseline as a springboard from which to add a suite of novel features including ones derived from acoustic templates, deep neural net phoneme detections, duration models, modulation features, and whole word point-process models. The SCRF framework is able to appropriately weight these different information sources to produce significant gains on both die Broadcast News and Wall Street Journal tasks. Geoffrey Zweig, Patrick Nguyen, Dirk Van Compernolle, Kris Demuynck, Les E. Atlas, Pascal Clark, Gregory Sell, Meihong Wang, Fei Sha, Hynek Hermansky, Damianos Karakos, Aren Jansen, Samuel Thomas 0001, Sivaram G. S. V. S., Samuel R. Bowman, Justine T. Kao |
ICASSP | 10 |
| 2011 | Rapid Evaluation of Speech Representations for Spoken Term DiscoveryabstractAcoustic front-ends are typically developed for supervised learning tasks and are thus optimized to minimize word error rate, phone error rate, etc. However, in recent efforts to develop zero-resource speech technologies, the goal is not to use transcribed speech to train systems but instead to discover the acoustic structure of the spoken language automatically. For this new setting, we require a framework for evaluating the quality of speech representations without coupling to a particular recognition architecture. Motivated by the spoken term discovery task, we present a dynamic time warping-based framework for quantifying how well a representation can associate words of the same type spoken by different speakers. We benchmark the quality of a wide range of speech representations using multiple frame-level distance metrics and demonstrate that our performance metrics can also accurately predict phone recognition accuracies. Index Terms: evaluation methods, acoustic front-end, spoken term discovery, zero resource Michael A. Carlin, Samuel Thomas 0001, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 4 |
| 2011 | Event Selection from Phone Posteriorgrams Using Matched FiltersabstractIn this paper we address the issue of how to select a minimal set of phonetic events from a phone posteriorgram while minimizing the loss of information. We derive phone posteriorgrams from two sources, Gaussian mixture models and sparse multilayer perceptrons, and apply phone-specific matched filters to the posteriorgrams to yield a smaller set of phonetic events. We introduce a mutual information based performance measure to compare phonetic event selection techniques and demonstrate that events extracted using matched filters can reduce input data while significantly improving performance of an event-based Keith Kintzley, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 3 |
| 2011 | Modulation Spectrum Analysis for Recognition of Reverberant SpeechabstractRecognition of reverberant speech constitutes a challenging problem for typical speech recognition systems. This is mainly due to the conventional short-term analysis/compensation tech-niques. In this paper, we present a feature extraction technique based on modeling long segments of temporal envelopes of the speech signal in narrow sub-bands using frequency domain lin-ear prediction (FDLP). FDLP provides an all-pole approxima-tion of the Hilbert envelope of the signal by linear prediction on cosine transform of the signal. We show that the FDLP modulation spectrum plays an important role in the robustness of the proposed feature extraction. Automatic speech recogni-tion (ASR) experiments on speech data degraded with a number of room impulse responses (with varying degrees of distortion) show significant performance improvements for the proposed FDLP features when compared to other robust feature extrac-tion techniques (average relative reduction of 40 % in word er-ror rate). Similar improvements are also obtained for far-field data which contain natural reverberation in background noise. Sri Harish Reddy Mallidi, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2011 | Adaptive Stream Fusion in Multistream Recognition of SpeechabstractA new method to deal with variable distortions of speech during the operation of the system is proposed. First, multiple processing streams are formed by extracting different spectral and temporal modulation components from the speech signal. Information in each stream is used to estimate posterior probabilities of phonemes. Initial values for a weighted integration of these individual estimates are found by normalized cross-correlation of the estimates with the actual phoneme labels on the training data. A statistical model of the final estimated posterior probabilities is used to characterize the system performance. During the operation, the weights in the linear fusion are adapted using particle filtering to optimize the performance. Results on phoneme recognition from noisy speech indicate the effectiveness of the proposed method. Index Terms: multistream speech recognition, spectrotemporal modulations Nima Mesgarani, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2011 | Mixture of Auto-Associative Neural Networks for Speaker VerificationabstractThe paper introduces a mixture of auto-associative neural networks for speaker verification. A new objective function based on posterior probabilities of phoneme classes is used for training the mixture. This objective function allows each component of the mixture to model part of the acoustic space corresponding to a broad phonetic class. This paper also proposes how factor analysis can be applied in this setting. The proposed techniques show promising results on a subset of NIST-08 speaker recognition evaluation (SRE) and yield about 10% relative improvement when combined with the state-of-the-art Gaussian Mixture Model i-vector system. Garimella S. V. S. Sivaram, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2011 | Analysis of MLP-Based Hierarchical Phoneme Posterior Probability EstimatorabstractWe analyze a simple hierarchical architecture consisting of two multilayer perceptron (MLP) classifiers in tandem to estimate the phonetic class conditional probabilities. In this hierarchical setup, the first MLP classifier is trained using standard acoustic features. The second MLP is trained using the posterior probabilities of phonemes estimated by the first, but with a long temporal context of around 150-230 ms. Through extensive phoneme recognition experiments, and the analysis of the trained second MLP using Volterra series, we show that 1) the hierarchical system yields higher phoneme recognition accuracies-an absolute improvement of 3.5% and 9.3% on TIMIT and CTS respectively-over the conventional single MLP-based system, 2) there exists useful information in the temporal trajectories of the posterior feature space, spanning around 230 ms of context, 3) the second MLP learns the phonetic temporal patterns in the posterior features, which include the phonetic confusions at the output of the first MLP as well as the phonotactics of the language as observed in the training data, and 4) the second MLP classifier requires fewer number of parameters and can be trained using lesser amount of training data. Joel Pinto, Garimella S. V. S. Sivaram, Mathew Magimai-Doss, Hynek Hermansky, Hervé Bourlard |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Robust spectro-temporal features based on autoregressive models of Hilbert envelopesabstractIn this paper, we present a robust spectro-temporal feature extraction technique using autoregressive models (AR) of sub-band Hilbert envelopes. AR models of Hilbert envelopes are derived using frequency domain linear prediction (FDLP). From the sub-band Hilbert envelopes, spectral features are derived by integrating these envelopes in short-term frames and the temporal features are formed by converting these envelopes into modulation frequency components. The spectral and temporal feature streams are then combined at the phoneme posterior level and are used as the input features for a recognition system. For the proposed features, robustness is achieved by using novel techniques of noise compensation and gain normalization. Phoneme recognition experiments on telephone speech in the HTIMIT database show significant performance improvements for the proposed features when compared to other robust feature techniques (average relative reduction of 10.6 % in phoneme error rate). In addition to the overall phoneme recognition rates, the performance with broad phonetic classes is also reported. Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 3 |
| 2010 | Comparison of modulation features for phoneme recognitionabstractIn this paper, we compare several approaches for the extraction of modulation frequency features from speech signal using a phoneme recognition system. The general framework in these approaches is to decompose the speech signal into a set of sub-bands. Amplitude modulations (AM) in the sub-band signal are used to derive features for automatic speech recognition (ASR). Then, we propose a feature extraction technique which uses autoregressive models (AR) of sub-band Hilbert envelopes in relatively long segments of speech signal. AR models of Hilbert envelopes are derived using frequency domain linear prediction (FDLP). Features are formed by converting the FDLP envelopes into static and dynamic modulation frequency components. In the phoneme recognition experiments using the TIMIT database, the FDLP based modulation frequency features provide significant improvements compared to other techniques (average relative improvement of 7.5% over the base-line features). Furthermore, a detailed analysis is performed to determine the relative contribution of various processing stages in the proposed technique. Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 3 |
| 2010 | History of modulation spectrum in ASRabstractPersonal view of recent history of applications of modulation spectrum (spectral dynamics) in automatic recognition of speech (ASR) is reviewed and references to some past works are given. Hynek Hermansky |
ICASSP | 1 |
| 2010 | Sparse coding for speech recognitionabstractThis paper proposes a novel feature extraction technique for speech recognition based on the principles of sparse coding. The idea is to express a spectro-temporal pattern of speech as a linear combination of an overcomplete set of basis functions such that the weights of the linear combination are sparse. These weights (features) are subsequently used for acoustic modeling. We learn a set of overcomplete basis functions (dictionary) from the training set by adopting a previously proposed algorithm which iteratively minimizes the reconstruction error and maximizes the sparsity of weights. Furthermore, features are derived using the learned basis functions by applying the well established principles of compressive sensing. Phoneme recognition experiments show that the proposed features outperform the conventional features in both clean and noisy conditions. Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Trac D. Tran, Hynek Hermansky |
ICASSP | 5 |
| 2010 | Towards spoken term discovery at scale with zero resourcesabstractThe spoken term discovery task takes speech as input and identifies terms of possible interest. The challenge is to perform this task efficiently on large amounts of speech with zero resources (no training data and no dictionaries), where we must fall back to more basic properties of language. We find that long (∼ 1 s) repetitions tend to be contentful phrases (e.g. University of Pennsylvania) and propose an algorithm to search for these long repetitions without first recognizing the speech. To address efficiency concerns, we take advantage of (i) sparse feature representations and (ii) inherent low occurrence frequency of long content terms to achieve orders-of-magnitude speedup relative to the prior art. We frame our evaluation in the context of spoken document information retrieval, and demonstrate our method’s competence at identifying repeated terms in conversational telephone speech. Index Terms: spoken term discovery, zero resource speech recognition, dotplots Aren Jansen, Kenneth Church 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2010 | A multistream multiresolution framework for phoneme recognitionabstractSpectrotemporal representation of speech has already shown promising results in speech processing technologies, however, many inherent issues of such representation, such as high dimensionality have limited their use in speech and speaker recognition. Multistream framework fits very well to such representation where different regions can be separately mapped into posterior probabilities of classes before merging. In this study, we investigated the effective ways of forming streams out of this representation for robust phoneme recognition. We also investigated multiple ways of fusing the posteriors of different streams based on their individual confidence or interactions between them. We observed 8.6% relative improvement in clean and 4 % in noise. We developed a simple yet effective linear combination technique that provides intuitive understanding of stream combinations and how even systematic errors can be learned to reduce confusions. Index Terms: speech recognition, spectrotemporal modulations, multistream Nima Mesgarani, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2010 | Sparse auto-associative neural networks: theory and application to speech recognition
Garimella S. V. S. Sivaram, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2010 | Cross-lingual and multi-stream posterior features for low resource LVCSR systemsabstractWe investigate approaches for large vocabulary continuous speech recognition (LVCSR) system for new languages or new domains using limited amounts of transcribed training data. In these low resource conditions, the performance of conventional LVCSR systems degrade significantly. We propose to train low resource LVCSR system with additional sources of information like annotated data from other languages (German and Spanish) and various acoustic feature streams (short-term and modulation features). We train multilayer perceptrons (MLPs) on these sources of information and use Tandem features derived from the MLPs for low resource LVCSR. In our experiments, the proposed system trained using only one hour of English conversational telephone speech (CTS) provides a relative improvement of 11% over the baseline system. Index Terms: Cross-lingual posterior features, Multi-stream features, Low resource ASR, Tandem features Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2010 | A phoneme recognition framework based on auditory spectro-temporal receptive fieldsabstractWe propose to incorporate features derived using spectrotemporal receptive fields (STRFs) of neurons in the auditory cortex for phoneme recognition. Each of these STRFs is tuned to different auditory frequencies, scales and modulation rates. We select different sets of STRFs which are specific for phonemes in different broad phonetic classes (BPC) of sounds. These STRFs are then used as spectro-temporal filters on spectrograms of speech to extract features for phoneme recognition. For the phoneme recognition task on the TIMIT database, the proposed features show a relative improvement of about 5% over conventional feature extraction techniques. Samuel Thomas 0001, Kailash Patil, Sriram Ganapathy, Nima Mesgarani, Hynek Hermansky |
INTERSPEECH | 5 |
| 2010 | Fully integrated 500uW speech detection wake-up circuitabstractSpeech analysis requires substantial computation. It is desirable to run this analysis only when needed and at other times to go to a low power state. Here we propose a self-biased low power speech detection wake up circuit which interfaces directly to standard electret microphones. The speech detector includes a microphone preamplifier, a power extraction squaring circuit, a bandpass filter passing power of the modulation spectrum in the speech band from 2-12 Hz, a half rectifier which extracts this phoneme band power, and a PFM silicon neuron which emits spikes indicating phoneme-rate modulation of the audio spectrum. The output of the speech detector circuit is an asynchronous stream of digital spikes at a rate of 1Hz to 20Hz whose temporal structure indicates the presence of speech. A subsequent conventional processor will go to sleep between spikes and only wake up for full power speech analysis when the temporal structure indicates speech. The circuit is built in 1.6um 2P-2M CMOS and consumes 500uW with a 3V supply when attached to a standard electret microphone. Tobi Delbruck, Raphael Berner, Hynek Hermansky |
ISCAS | 4 |
| 2010 | The use of spike-based representations for hardware audition systemsabstractHumans are able to process speech and other sounds effectively in adverse environments, hearing through noise, reverberation, and interference from other speakers. To date, machines have been unable to match human performance. One profound difference between biological and engineering systems comes at the input stage. In machines, an acoustic signal is typically chopped into short equally spaced segments in time. In biological systems, the cochlea outputs asynchronous spikes that respond in real-time to acoustic inputs. In this paper we describe a spiking cochlea implementation and recent experiments in both speaker and speech recognition that use spikes as input. Shih-Chii Liu, Nima Mesgarani, John G. Harris, Hynek Hermansky |
ISCAS | 4 |
| 2010 | Data-Driven and Feedback Based Spectro-Temporal Features for Speech RecognitionabstractThis paper proposes novel data-driven and feedback based discriminative spectro-temporal filters for feature extraction in automatic speech recognition (ASR). Initially a first set of spectro-temporal filters are designed to separate each phoneme from the rest of the phonemes. A hybrid Hidden Markov Model/Multilayer Perceptron (HMM/MLP) phoneme recognition system is trained on the features derived using these filters. As a feedback to the feature extraction stage, top confusions of this system are identified, and a second set of filters are designed specifically to address these confusions. Phoneme recognition experiments on TIMIT show that the features derived from the combined set of discriminative filters outperform conventional speech recognition features, and also contain significant complementary information. Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Nima Mesgarani, Hynek Hermansky |
IEEE Signal Process. Lett. | 4 |
| 2010 | Autoregressive Models of Amplitude Modulations in Audio CompressionabstractWe present a scalable medium bit-rate wide-band audio coding technique based on frequency-domain linear prediction (FDLP). FDLP is an efficient method for representing the long-term amplitude modulations of speech/audio signals using autoregressive models. For the proposed audio codec, relatively long temporal segments (1000 ms) of the input audio signal are decomposed into a set of critically sampled sub-bands using a quadrature mirror filter (QMF) bank. The technique of FDLP is applied on each sub-band to model the sub-band temporal envelopes. The residual of the linear prediction, which represents the frequency modulations in the sub-band signal, are encoded and transmitted along with the envelope parameters. These steps are reversed at the decoder to reconstruct the signal. The proposed codec utilizes a simple signal independent nonadaptive compression mechanism for a wide class of speech and audio signals. The subjective and objective quality evaluations show that the reconstruction signal quality for the proposed FDLP codec compares well with the state-of-the-art audio codecs in the 32-64 kbps range. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Temporal envelope subtraction for robust speech recognition using modulation spectrumabstractIn this paper, we present a new noise compensation technique for modulation frequency features derived from syllable length segments of subband temporal envelopes. The subband temporal envelopes are estimated using frequency domain linear prediction (FDLP). We propose a technique for noise compensation in FDLP where an estimate of the noise envelope is subtracted from the noisy speech envelope. The noise compensated FDLP envelopes are compressed with static (logarithmic) and dynamic (adaptive loops) compression and are transformed into modulation spectral features. Experiments are performed on a phoneme recognition task as well as a connected digit recognition task where the test data is corrupted with variety of noise types at different signal to noise ratios. In these experiments with mismatched train and test conditions, the proposed features provide considerable improvements compared to other state of the art noise robust feature extraction techniques (average relative improvement of 25% and 35% over the baseline PLP features for phoneme and word recognition tasks respectively). Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ASRU | 3 |
| 2009 | Reconciliation of human and machine speech recognition performanceabstractThis paper focuses on resolving a number of issues that appear when the performance of human speech recognition is compared to that of automatic speech recognition. In particular human experimental data suggest that the resulting error is a product of the individual streams. On the other hand, Bayesian combination requires a multiplication of the estimates of prior probabilities and likelihoods. We show that, in principle, there is no discrepancy. The product of errors is a performance measure and human and machine performance may be consistent with this empirically established regularity. The product of probabilities is step in an algorithm to achieve the performance that may or may not be consistent with the product of errors. The main problem is that most of prior discussions failed to distinguish the performance measures from the estimates of the parameters used in the algorithm. Misha Pavel, Malcolm Slaney, Hynek Hermansky |
ICASSP | 3 |
| 2009 | Volterra series for analyzing MLP based phoneme posterior estimatorabstractWe present a framework to apply Volterra series to analyze multi-layered perceptrons trained to estimate the posterior probabilities of phonemes in automatic speech recognition. The identified Volterra kernels reveal the spectro-temporal patterns that are learned by the trained system for each phoneme. To demonstrate the applicability of Volterra series, we analyze a multilayered perceptron trained using Mel filter bank energy features and analyze its first order Volterra kernels. Joel Pinto, Garimella S. V. S. Sivaram, Hynek Hermansky, Mathew Magimai-Doss |
ICASSP | 3 |
| 2009 | Phoneme recognition using spectral envelope and modulation frequency featuresabstractWe present a new feature extraction technique for phoneme recognition that uses short-term spectral envelope and modulation frequency features. These features are derived from sub-band temporal envelopes of speech estimated using frequency domain linear prediction (FDLP). While spectral envelope features are obtained by the short-term integration of the sub-band envelopes, the modulation frequency components are derived from the long-term evolution of the sub-band envelopes. These features are combined at the phoneme posterior level and used as features for a hybrid HMM-ANN phoneme recognizer. For the phoneme recognition task on the TIMIT database, the proposed features show an improvement of 4.7% over the other feature extraction techniques. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
ICASSP | 3 |
| 2009 | Static and dynamic modulation spectrum for speech recognitionabstractWe present a feature extraction technique based on static and dynamic modulation spectrum derived from long-term envelopes in sub-bands. Estimation of the sub-band temporal envelopes is done using Frequency Domain Linear Prediction (FDLP). These sub-band envelopes are compressed with a static (logarithmic) and dynamic (adaptive loops) compression. The compressed sub-band envelopes are transformed into modulation spectral components which are used as features for speech recognition. Experiments are performed on a phoneme recognition task using a hybrid HMM-ANN phoneme recognition system and an ASR task using the TANDEM speech recognition system. The proposed features provide a relative improvements of 3.8 % and 11.5 % in phoneme recognition accuracies for TIMIT and conversation telephone speech (CTS) respectively. Further, these improvements are found to be consistent for ASR tasks on OGI-Digits database (relative improvement of 13.5 %). Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2009 | Posterior-based out of vocabulary word detection in telephone speechabstractIn this paper we present an out-of-vocabulary word detector suitable for English conversational and read speech. We use an approach based on phone posteriors created by a Large Vocab-ulary Continuous Speech Recognition system and an additional phone recognizer, that allows detection of OOV and misrecog-nized words. In addition, the recognized word output can be transcribed more detailed using several classes. Reported re-sults are on CallHome English and Wall Street Journal data. Index Terms: confidence measures, out-of-vocabulary word detection, phone posteriors, neural net, OOV Stefan Kombrink, Lukás Burget, Pavel Matejka, Martin Karafiát, Hynek Hermansky |
INTERSPEECH | 5 |
| 2009 | Discriminant spectrotemporal features for phoneme recognitionabstractWe propose discriminant methods for deriving twodimensional spectrotemporal features for phoneme recognition that are estimated to maximize the separation between the representations of phoneme classes. The linearity of the filters results in their intuitive interpretation enabling us to investigate the working principles of the system and to improve its performance by locating the sources of error. Two methods for the estimation of filters are proposed: Regularized Least Square (RLS) and Modified Linear Discriminant Analysis (MLDA). Both methods reach a comparable improvement over the baseline condition demonstrating the advantage of the discriminant spectrotemporal filters. Nima Mesgarani, Garimella S. V. S. Sivaram, Sridhar Krishna Nemala, Mounya Elhilali, Hynek Hermansky |
INTERSPEECH | 5 |
| 2009 | Arithmetic coding of sub-band residuals in FDLP speech/audio codecabstractA speech/audio codec based on Frequency Domain Linear Prediction (FDLP) exploits auto-regressive modeling to approximate instantaneous energy in critical frequency sub-bands of relatively long input segments. The current version of the FDLP codec operating at 66 kbps has been shown to provide comparable subjective listening quality results to state-of-the-art codecs on similar bit-rates even without employing standard blocks such as entropy coding or simultaneous masking. This paper describes an experimental work to increase compression efficiency of the FDLP codec by employing entropy coding. Unlike conventional Huffman coding employed in current speech/audio coding systems, we describe an efficient way to exploit arithmetic coding to entropy compress quantized spectral magnitudes of the sub-band FDLP residuals. Such an approach provides 11% (∼ 3 kbps) bit-rate reduction compared to the Huffman coding algorithm (∼ 1 kbps). Petr Motlícek, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2009 | Tandem representations of spectral envelope and modulation frequency features for ASRabstractWe present a feature extraction technique for automatic speech recognition that uses Tandem representation of short-term spectral envelope and modulation frequency features. These features, derived from sub-band temporal envelopes of speech estimated using frequency domain linear prediction, are combined at the phoneme posterior level. Tandem representations derived from these phoneme posteriors are used along with HMM based ASR systems for both small and large vocabulary continuous speech recognition (LVCSR) tasks. For a small vocabulary continuous digit task on the OGI Digits database, the proposed features reduce the word error rate (WER) by 13 % relative to other feature extraction techniques. We obtain a relative reduction of about 14 % in WER for an LVCSR task using the NIST RT05 evaluation data. For phoneme recognition tasks on the TIMIT database these features provide a relative improvement of 13% compared to other techniques. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2008 | Combination of strongly and weakly constrained recognizers for reliable detection of OOVSabstractThis paper addresses the detection of OOV segments in the output of a large vocabulary continuous speech recognition (LVCSR) system. First, standard confidence measures from frame-based wordand phone- posteriors are investigated. Substantial improvement is obtained when posteriors from two systems — strongly constrained (LVCSR) and weakly constrained (phone posterior estimator) are combined. We show that this approach is also suitable for detection of general recognition errors. All results are presented on WSJ task with reduced recognition vocabulary. Lukás Burget, Petr Schwarz, Pavel Matejka, Mirko Hannemann, Ariya Rastrow, Christopher M. White, Sanjeev Khudanpur, Hynek Hermansky, Jan Cernocký |
ICASSP | 8 |
| 2008 | Temporal masking for bit-rate reduction in audio codec based on Frequency Domain Linear PredictionabstractAudio coding based on frequency domain linear prediction (FDLP) uses auto-regressive model to approximate Hilbert envelopes in frequency sub-bands for relatively long temporal segments. Although the basic technique achieves good quality of the reconstructed signal, there is a need for improving the coding efficiency. In this paper, we present a novel method for the application of temporal masking to reduce the bit-rate in a FDLP based codec. Temporal masking refers to the hearing phenomenon, where the exposure to a sound reduces response to following sounds for a certain period of time (up to 200 ms). In the proposed version of the codec, a first order forward masking model of the human ear is implemented and informal listening experiments using additive white noise are performed to obtain the exact noise masking thresholds. Subsequently, this masking model is employed in encoding the sub- band FDLP carrier signal. Application of the temporal masking in the FDLP codec results in a bit-rate reduction of about 10% without degrading the quality. Performance evaluation is done with perceptual evaluation of audio quality (PEAQ) scores and with subjective listening tests. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky, Harinath Garudadri |
ICASSP | 3 |
| 2008 | Exploiting contextual information for improved phoneme recognitionabstractIn this paper, we investigate the significance of contextual information in a phoneme recognition system using the hidden Markov model - artificial neural network paradigm. Contextual information is probed at the feature level as well as at the output of the multilayered perceptron. At the feature level, we analyze and compare different methods to model sub-phonemic classes. To exploit the contextual information at the output of the multilayered perceptron, we propose the hierarchical estimation of phoneme posterior probabilities. The best phoneme (excluding silence) recognition accuracy of 73.4% on the TIMIT database is comparable to that of the state-of- the-art systems, but more emphasis is on analysis of the contextual information. Joel Pinto, Bayya Yegnanarayana, Hynek Hermansky, Mathew Magimai-Doss |
ICASSP | 3 |
| 2008 | Hierarchical and parallel processing of modulation spectrum for ASR applicationsabstractThe modulation spectrum is an efficient representation for describing dynamic information in signals. In this work we investigate how to exploit different elements of the modulation spectrum for extraction of information in automatic recognition of speech (ASR). Parallel and hierarchical (sequential) approaches are investigated. Parallel processing combines outputs of independent classifiers applied to different modulation frequency channels. Hierarchical processing uses different modulation frequency channels sequentially. Experiments are run on a LVCSR task for meetings transcription and results are reported on the RT05 evaluation data. Processing modulation frequencies channels with different classifiers provides a consistent reduction in WER (2% absolute w.r.t. PLP baseline). Hierarchical processing outperforms parallel processing. The largest WER reduction is obtained through sequential processing moving from high to low modulation frequencies. This model is consistent with several perceptual and physiological studies on auditory processing. Fabio Valente, Hynek Hermansky |
ICASSP | 2 |
| 2008 | Confidence estimation, OOV detection and language ID using phone-to-word transduction and phone-level alignmentsabstractAutomatic speech recognition (ASR) systems continue to make errors during search when handling various phenomena including noise, pronunciation variation, and out of vocabulary (OOV) words. Predicting the probability that a word is incorrect can prevent the error from propagating and perhaps allow the system to recover. This paper addresses the problem of detecting errors and OOVs for read Wall Street Journal speech when the word error rate (WER) is very low. It augments a traditional confidence estimate by introducing two novel methods: phone-level comparison using multi-string alignment (MSA) and word-level comparison using phone-to-word transduction. We show that features from phone and word string comparisons can be added to a standard maximum entropy framework thereby substantially improving performance in detecting both errors and OOVs. Additionally we show an extension to detecting English and accented English for the language identification (LID) task. Christopher M. White, Geoffrey Zweig, Lukás Burget, Petr Schwarz, Hynek Hermansky |
ICASSP | 5 |
| 2008 | The DIRAC AWEAR audio-visual platform for detection of unexpected and incongruent eventsabstractIt is of prime importance in everyday human life to cope with and respond appropriately to events that are not foreseen by prior experience. Machines to a large extent lack the ability to respond appropriately to such inputs. An important class of unexpected events is defined by incongruent combinations of inputs from different modalities and therefore multimodal information provides a crucial cue for the identification of such events, e.g., the sound of a voice is being heard while the person in the field-of-view does not move her lips. In the project DIRAC ("Detection and Identification of Rare Audio-visual Cues") we have been developing algorithmic approaches to the detection of such events, as well as an experimental hardware platform to test it. An audio-visual platform ("AWEAR" - audio-visual wearable device) has been constructed with the goal to help users with disabilities or a high cognitive load to deal with unexpected events. Key hardware components include stereo panoramic vision sensors and 6-channel worn-behind-the-ear (hearing aid) microphone arrays. Data have been recorded to study audio-visual tracking, a/v scene/object classification and a/v detection of incongruencies. Jörn Anemüller, Jörg-Hendrik Bach, Barbara Caputo, Michal Havlena, Jie Luo 0018, Hendrik Kayser, Bastian Leibe, Petr Motlícek, Tomás Pajdla, Misha Pavel, Akihiko Torii, Luc Van Gool, Alon Zweig, Hynek Hermansky |
ICMI | 14 |
| 2008 | Spectral noise shaping: improvements in speech/audio codec based on linear prediction in spectral domainabstractAudio coding based on Frequency Domain Linear Prediction (FDLP) uses auto-regressive models to approximate Hilbert envelopes in frequency sub-bands. Although the basic technique achieves good coding efficiency, there is a need to improve the reconstructed signal quality for tonal signals with impulsive spectral content. For such signals, the quantization noise in the FDLP codec appears as frequency components not present in the input signal. In this paper, we propose a technique of Spectral Noise Shaping (SNS) for improving the quality of tonal signals by applying a Time Domain Linear Prediction (TDLP) filter prior to the FDLP processing. The inverse TDLP filter at the decoder shapes the quantization noise to reduce the artifacts. Application of the SNS technique to the FDLP codec improves the quality of the tonal signals without affecting the bit-rate. Performance evaluation is done with Perceptual Evaluation of Audio Quality (PEAQ) scores and with subjective listening tests. Sriram Ganapathy, Petr Motlícek, Hynek Hermansky, Harinath Garudadri |
INTERSPEECH | 3 |
| 2008 | Front-end for far-field speech recognition based on frequency domain linear predictionabstractAbstract. Automatic Speech Recognition (ASR) systems usually fail when they encounter speech from far-field microphone in reverberant environments. This is due to the application of short-term feature extraction techniques which do not compensate for the artifacts introduced by long room impulse responses. In this paper, we propose a front-end, based on Frequency Domain Linear Prediction (FDLP), that tries to remove reverberation artifacts present in far-field speech. Long temporal segments of far-field speech are analyzed in narrow frequency sub-bands to extract FDLP envelopes and residual signals. Filtering the residual signals with gain normalized inverse FDLP filters result in a set of sub-band signals which are synthesized to reconstruct the signal back. ASR experiments on far-field speech data processed by the proposed front-end show significant improvements (relative reduction of 30 % in word error rate) compared to other robust feature extraction techniques. 2 IDIAP–RR 08-17 1 Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 3 |
| 2008 | Combining evidence from a generative and a discriminative model in phoneme recognitionabstractWe investigate the use of the log-likelihood of the features obtained from a generative Gaussian mixture model, and the posterior probability of phonemes from a discriminative multilayered perceptron in multi-stream combination for recognition of phonemes. Multi-stream combination techniques, namely early integration and late integration are used to combine the evidence from these models. By using multi-stream combination, we obtain a phoneme recognition accuracy of 74\\% on the standard TIMIT database, an absolute improvement of 2.5\\% over the single best stream. Joel Pinto, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Introducing temporal asymmetries in feature extraction for automatic speech recognitionabstractAbstract. We propose a new auditory inspired feature extraction technique for automatic speech recognition (ASR). Features are extracted by filtering the temporal trajectory of spectral energies in each critical band of speech by a bank of finite impulse response (FIR) filters. Impulse responses of these filters are derived from a modified Gabor envelope in order to emulate asymmetries of the temporal receptive field (TRF) profiles observed in higher level auditory neurons. We obtain 11.4% relative improvement in word error rate on OGI-Digits database and, 3.2 % relative improvement in phoneme error rate on TIMIT database over the MRASTA technique. 2 IDIAP–RR 08-25 1 Garimella S. V. S. Sivaram, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Hilbert envelope based spectro-temporal features for phoneme recognition in telephone speechabstractAbstract. In this paper, we present a spectro-temporal feature extraction technique using sub-band Hilbert envelopes of relatively long segments of speech signal. Hilbert envelopes of the sub-bands are estimated using Frequency Domain Linear Prediction (FDLP). Spectral features are derived by integrating the sub-band Hilbert envelopes in short-term frames and the temporal features are formed by converting the FDLP envelopes into modulation frequency components. These are then combined at the phoneme posterior level and are used as the input features for a phoneme recognition system. In order to improve the robustness of the proposed features to telephone speech, the sub-band temporal envelopes are gain normalized prior to feature extraction. Phoneme recognition experiments on telephone speech in the HTIMIT database show significant performance improvements for the proposed features when compared to other robust feature techniques (average relative reduction of 11 % in phoneme error rate).2 IDIAP–RR 08-18 n this paper, we present a spectro-temporal feature extraction technique using sub-band Hilbert envelopes of relatively long segments of speech signal. Hilbert envelopes of the sub-bands are estimated using Frequency Domain Linear Prediction (FDLP). Spectral features are derived by integrating the sub-band Hilbert envelopes in short-term frames and the temporal features are formed by converting Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 3 |
| 2008 | On the combination of auditory and modulation frequency channels for ASR applicationsabstractThis paper investigates the combination of evidence coming from different frequency channels obtained filtering the speech signal at different auditory and modulation frequencies. In our previous work \\cite{icassp2008}, we showed that combination of classifiers trained on different ranges of {\\it modulation} frequencies is more effective if performed in sequential (hierarchical) fashion. In this work we verify that combination of classifiers trained on different ranges of {\\it auditory} frequencies is more effective if performed in parallel fashion. Furthermore we propose an architecture based on neural networks for combining evidence coming from different auditory-modulation frequency sub-bands that takes advantages of previous findings. This reduces the final WER by 6.2\\% (from 45.8\\% to 39.6\\%) w.r.t the single classifier approach in a LVCSR task. Fabio Valente, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Beyond Novelty Detection: Incongruent Events, when General and Specific Classifiers DisagreeabstractUnexpected stimuli are a challenge to any machine learning algorithm. Here we identify distinct types of unexpected events, focusing on 'incongruent events' - when 'general level' and 'specific level' classifiers give conflicting predictions. We define a formal framework for the representation and processing of incongruent events: starting from the notion of label hierarchy, we show how partial order on labels can be deduced from such hierarchies. For each event, we compute its probability in different ways, based on adjacent levels (according to the partial order) in the label hierarchy . An incongruent event is an event where the probability computed based on some more specific level (in accordance with the partial order) is much smaller than the probability computed based on some more general level, leading to conflicting predictions. We derive algorithms to detect incongruent events from different types of hierarchies, corresponding to class membership or part membership. Respectively, we show promising results with real data on two specific problems: Out Of Vocabulary words in speech recognition, and the identification of a new sub-class (e.g., the face of a new individual) in audio-visual facial object recognition. Daphna Weinshall, Hynek Hermansky, Alon Zweig, Holly Brügge Jimison, Frank W. Ohl, Misha Pavel |
NIPS | 2 |
| 2008 | Recognition of Reverberant Speech Using Frequency Domain Linear PredictionabstractPerformance of a typical automatic speech recognition (ASR) system severely degrades when it encounters speech from reverberant environments. Part of the reason for this degradation is the feature extraction techniques that use analysis windows which are much shorter than typical room impulse responses. We present a feature extraction technique based on modeling temporal envelopes of the speech signal in narrow subbands using frequency domain linear prediction (FDLP). FDLP provides an all-pole approximation of the Hilbert envelope of the signal obtained by linear prediction on cosine transform of the signal. ASR experiments on speech data degraded with a number of room impulse responses (with varying degrees of distortion) show significant performance improvements for the proposed FDLP features when compared to other robust feature extraction techniques (average relative reduction of 24% in word error rate). Similar improvements are also obtained for far-field data which contain natural reverberation in background noise. These results are achieved without any noticeable degradation in performance for clean speech. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
IEEE Signal Process. Lett. | 3 |
| 2007 | Wide-Band Perceptual Audio Coding Based on Frequency-Domain Linear PredictionabstractIn this paper we propose an extension of the very low bit-rate speech coding technique, exploiting predictability of the temporal evolution of spectral envelopes, for wide-band audio coding applications. Temporal envelopes in critically band-sized sub-bands are estimated using frequency domain linear prediction applied on relatively long time segments. The sub-band residual signals, which play an important role in acquiring high quality reconstruction, are processed using a heterodyning-based signal analysis technique. For reconstruction, their optimal parameters are estimated using a closed-loop analysis-by-synthesis technique driven by a perceptual model emulating simultaneous masking properties of the human auditory system. We discuss the advantages of the approach and show some properties on challenging audio recordings. The proposed technique is capable of encoding high quality, variable rate audio signals on bit-rates below 1 bit/sample. Petr Motlícek, Vijay Ullal, Hynek Hermansky |
ICASSP (1) | 3 |
| 2007 | Combination of Acoustic Classifiers Based on Dempster-Shafer Theory of EvidenceabstractIn this paper we investigate combination of neural net based classifiers using Dempster-Shafer theory of evidence. Under some assumptions, combination rule resembles a product of errors rule observed in human speech perception. Different combination are tested in ASR experiments both in matched and mismatched conditions and compared with more conventional probability combination rules. Proposed techniques are particularly effective in mismatched conditions. Fabio Valente, Hynek Hermansky |
ICASSP (4) | 2 |
| 2007 | Detection of out-of-vocabulary words in posterior based ASRabstractOver the years, sophisticated techniques for utilizing the prior knowledge in the form of text-derived language model and in pronunciation lexicon evolved. However, their use has an undesirable effect: unexpected lexical items (words) in the phrase are replaced by acoustically acceptable in-vocabulary items [1]. This is the major source of error since the replacement often introduces additional errors [2, 3]. Improving the machine ability to handle these unexpected words would considerably increase the utility of speech recognition technology. Hamed Ketabdar, Mirko Hannemann, Hynek Hermansky |
INTERSPEECH | 3 |
| 2007 | Exploiting phoneme similarities in hybrid HMM-ANN keyword spottingabstractWe propose a technique for generating alternative models for keywords in a hybrid hidden Markov model - artificial neural network (HMM-ANN) keyword spotting paradigm. Given a base pronunciation for a keyword from the lookup dictionary, our algorithm generates a new model for a keyword which takes into account the systematic errors made by the neural network and avoiding those models that can be confused with other words in the language. The new keyword model improves the keyword detection rate while minimally increasing the number of false alarms. Joel Pinto, Andrew Lovitt, Hynek Hermansky |
INTERSPEECH | 3 |
| 2007 | MRASTA and PLP in automatic speech recognition
S. R. Mahadeva Prasanna, Hynek Hermansky |
INTERSPEECH | 2 |
| 2007 | Multi-stream features combination based on dempster-shafer rule for LVCSR systemabstractThis paper investigates the combination of two streams of acoustic features. Extending our previous work on small vocabulary task, we show that combination based on Dempster-Shafer rule outperforms several classical rules like sum, product and inverse entropy weighting even in LVCSR systems. We analyze results in terms of Frame Error Rate and Cross Entropy measures. Experimental framework uses meeting transcription task and results are provided on RT05 evaluation data. Results are consistent with what has been previously observed on smaller databases. Fabio Valente, Jithendra Vepa, Hynek Hermansky |
INTERSPEECH | 3 |
| 2007 | Hierarchical neural networks feature extraction for LVCSR systemabstractThis paper investigates the use of a hierarchy of Neural Networks for performing data driven feature extraction.Two different hierarchical structures based on long and short temporal context are considered.Features are tested on two different LVCSR systems for Meetings data (RT05 evaluation data) and for Arabic Broadcast News (BNAT05 evaluation data).The hierarchical NNs outperforms the single NN features consistently on different type of data and tasks and provides significant improvements w.r.t.respective baselines systems.Best results are obtained when different time resolutions are used at different level of the hierarchy. Fabio Valente, Jithendra Vepa, Christian Plahl, Christian Gollan, Hynek Hermansky, Ralf Schlüter |
INTERSPEECH | 5 |
| 2006 | Towards ASR Based on Hierarchical Posterior-Based Keyword RecognitionabstractThe paper presents an alternative approach to automatic recognition of speech in which each targeted word is classified by a separate binary classifier against all other sounds. No time alignment is done. To build a recognizer for N words, N parallel binary classifiers are applied. The system first estimates uniformly sampled posterior probabilities of phoneme classes, followed by a second step in which a rather long sliding time window is applied to the phoneme posterior estimates and its content is classified by an artificial neural network to yield posterior probability of the keyword. On a small vocabulary ASR task, the system still does not reach the performance of the state-of-the-art system but its conceptual simplicity, the ease of adding new target words, and its inherent resistance to out-of-vocabulary sounds may prove significant advantage in many applications Petr Fousek, Hynek Hermansky |
ICASSP (1) | 2 |
| 2006 | Discriminant linear processing of time-frequency planeabstractExtending previous works done on considerably smaller data sets, the paper studies linear discriminant analysis of about 30 hours of phoneme-labeled speech data in the time-frequency domain. Analysis is carried both independently in time and frequency and jointly. Data driven spectral basis show similar frequency sensitivity as human hearing. LDA-derived temporal FIR filters are consistent with temporal lateral inhibition. Considerable improvement is obtained using first temporal discriminant. Fabio Valente, Hynek Hermansky |
INTERSPEECH | 2 |
| 2005 | Multi-resolution RASTA filtering for TANDEM-based ASRabstractNew speech representation based on multiple filtering of temporal trajectories of speech energies in frequency sub-bands is proposed and tested. The technique extends earlier works on delta features and RASTA filtering by processing temporal trajectories by a bank of band-pass filters with varying resolutions. In initial tests on OGI Digits database the technique yields about 30\\% relative improvement in word error rate over the conventional PLP features. Since the applied filters have zero-mean impulse responses, the technique is inherently robust to linear distortions. Hynek Hermansky, Petr Fousek |
INTERSPEECH | 1 |
| 2004 | Phase autocorrelation (PAC) features in entropy based multi-stream for robust speech recognitionabstractMethods to improve noise robustness of speech recognition systems often result in degradation of recognition performance for clean speech. Recently proposed phase autocorrelation (PAC) based features (S. Ikbal et al., Proc. ICASSP-03, p.II-133-6, 2003; Proc. IEEE ASRU 2003 Workshop, 2003), while showing noticeable improvement in noise robustness, also suffer from this drawback. We try to alleviate this problem by using the PAC based features along with regular speech features in a multi-stream framework. The multi-stream system uses the entropy of the posterior probability distribution, computed during recognition, as a confidence measure to combine evidence from different feature streams adaptively (Misra, H. et al., Proc. ICASSP-03, p.II-741-4, 2003). Experimental results obtained on OGI Numbers95 database and Noisex92 noise database show that such a system yields the best possible recognition performance in all conditions. Actually, the combination always performs better than the best performing stream for all the conditions. Shajith Ikbal, Hemant Misra, Hervé Bourlard, Hynek Hermansky |
ICASSP (1) | 4 |
| 2004 | Spectral entropy based feature for robust ASRabstractIn general, entropy gives us a measure of the number of bits required to represent some information. When applied to the probability mass function (PMF), entropy can also be used to measure the "peakiness" of a distribution. We propose using the entropy of a short time Fourier transform spectrum, normalised as PMF, as an additional feature for automatic speech recognition (ASR). It is indeed expected that a peaky spectrum, representation of clear formant structure in the case of voiced sounds, will have low entropy, while a flatter spectrum, corresponding to nonspeech or noisy regions, will have higher entropy. Extending this reasoning further, we introduce the idea of a multiband/multiresolution entropy feature where we divide the spectrum into equal size subbands and compute entropy in each subband. The results show that multiband entropy features used in conjunction with normal cepstral features improve the performance of an ASR system. Hemant Misra, Shajith Ikbal, Hervé Bourlard, Hynek Hermansky |
ICASSP (1) | 4 |
| 2004 | On use of task independent training data in tandem feature extractionabstractThe problem we address in this paper is, whether the feature extraction module trained on large amounts of task independent data, can improve the performance of stochastic models? We show that when there is only a small amount of task specific training data available, tandem features trained on task independent data give considerable improvement over perceptual linear prediction (PLP) cepstral features in hidden Markov model (HMM) based speech recognition systems. Sunil Sivadas, Hynek Hermansky |
ICASSP (1) | 2 |
| 2004 | LP-TRAP: linear predictive temporal patternsabstractAutoregressive modeling is applied for approximating the temporal evolution of spectral density in critical-band-sized subbands of a segment of speech signal. The generalized autocorrelation linear predictive technique allows for a compromise between fitting the peaks and the troughs of the Hilbert envelope of the signal in the sub-band. The cosine transform coefficients of the approximated sub-band envelopes, computed recursively from the all-pole polynomials, are used as inputs to a TRAP-based speech recognition system and are shown to improve recognition accuracy. Marios Athineos, Hynek Hermansky, Daniel P. W. Ellis |
INTERSPEECH | 2 |
| 2004 | New nonsense syllables database - analyses and preliminary ASR experimentsabstractThe paper presents analyses, modifications, and first ex-periments with a new nonsense syllables database. Results of preliminary experiments with phoneme recognition are given and discussed. 1. Petr Fousek, Frantisek Grézl, Hynek Hermansky, Petr Svojanovsky |
INTERSPEECH | 3 |
| 2004 | Entropy based combination of tandem representations for noise robust ASRabstractIn this paper, we present an entropy based method to combine tandem representations of the recently proposed Phase AutoCorrelation (PAC) based features and Mel-Frequency Cepstral Coefficients (MFCC) features. PAC based features, derived from a nonlinear transformation of autocorrelation coefficients and shown to be noise robust, improve their robustness to additive noise in their tandem representation. On the other hand, MFCC features in their tandem representation show a significant improvement in recognition performance on clean speech. An entropy based combination method investigated in this paper adaptively gives a higher weighting to the representation of MFCC features in clean speech and to the representation of PAC based features in noisy speech, thus yielding a robust recognition performance in all conditions. Shajith Ikbal, Hemant Misra, Sunil Sivadas, Hynek Hermansky, Hervé Bourlard |
INTERSPEECH | 4 |
| 2003 | Generalized tandem feature extractionabstractWe study the use of a generalized multilayer perceptron (MLP) architecture to tandem feature extraction. In the tandem feature extraction scheme an MLP with a softmax output layer is discriminatively trained to estimate phoneme posterior probabilities on a labeled database. The outputs of the MLP after nonlinear transformation and whitening are used as features in a Gaussian mixture model (GMM) based speech recognizer. We consider three layer MLPs with a linear output layer. They nonlinearly transform the input data to a higher dimensional space defined by the output of hidden units and perform linear discriminant analysis (LDA) on the hidden unit outputs. We compare the performances of these features with the direct application of LDA on input data, which is equivalent to MLP with linear hidden and output layers. The tandem features outperform those obtained from LDA and linear output MLPs on a connected digit recognition task. Sunil Sivadas, Hynek Hermansky |
ICASSP (1) | 2 |
| 2003 | Segmentation of speech for speaker and language recognitionabstractCurrent Automatic Speech Recognition systems convert the speech signal into a sequence of discrete units, such as phonemes, and then apply statistical methods on the units to produce the linguistic message. Similar methodology has also been applied to recognize speaker and language, except that the output of the system can be the speaker or language information. Therefore, we propose the use of temporal trajectories of fundamental frequency and short-term energy to segment and label the speech signal into a small set of discrete units that can be used to characterize speaker and/or language. The proposed approach is evaluated using the NIST Extended Data Speaker Detection task and the NIST Language Identification task. André Adami, Hynek Hermansky |
INTERSPEECH | 2 |
| 2003 | Local averaging and differentiating of spectral plane for TRAP-based ASR
Frantisek Grézl, Hynek Hermansky |
INTERSPEECH | 2 |
| 2003 | Band-independent speech-event categories for TRAP based ASR
Hynek Hermansky, Pratibha Jain |
INTERSPEECH | 1 |
| 2003 | Beyond a single critical-band in TRAP based ASR
Pratibha Jain, Hynek Hermansky |
INTERSPEECH | 2 |
| 2003 | Novel approaches for one- and two-speaker detection
Sachin S. Kajarekar, André Adami, Hynek Hermansky |
INTERSPEECH | 3 |
| 2003 | In search of target class definition in tandem feature extraction
Sunil Sivadas, Hynek Hermansky |
INTERSPEECH | 2 |
| 2003 | Data-driven spectral basis functions for automatic speech recognition
Naren Malayath, Hynek Hermansky |
Speech Commun. | 2 |
| 2002 | A new speaker change detection method for two-speaker segmentationabstractIn absence of prior information about speakers, an important step in speaker segmentation is to obtain initial estimates for training speaker models. In this paper, we present a new method for obtaining these estimates. The method assumes that a conversation must be initiated by one of the speakers. Thus one speaker model is estimated from the small segment at the beginning of the conversation and the segment that has the largest distance from the initial segment is used to train second speaker model. We describe a system based on this method and evaluate it on two different tasks: a controlled task with variations in the duration of the initial speaker segment and amount of overlapped speech and 2001 NIST Speaker Recognition Evaluation task that contains natural conversations. This system shows significant improvements over the conventional system in absence of overlapped speech on the controlled task. André Adami, Sachin S. Kajarekar, Hynek Hermansky |
ICASSP | 3 |
| 2002 | Hierarchical tandem feature extractionabstractWe present a hierarchical architecture for tandem acoustic modeling. In the tandem acoustic modeling paradigm a Multi Layer Perceptron (MLP) is discriminatively trained to estimate phoneme posterior probabilities on a labeled database. The outputs of the MLP after nonlinear transformation and whitening are used as features in a Gaussian Mixture Model (GMM) based recognizer. In this paper we replace the large monolithic MLP with hierarchies of MLP experts. We apply this approach on Speech in Noisy Environments (SPINE 1) evaluation conducted by the Naval Research Laboratory (NRL). We observe a reduction in word error rate of 30% with context-independent models and 5% WER with context-dependent models relative to PLP features. Sunil Sivadas, Hynek Hermansky |
ICASSP | 2 |
| 2002 | Qualcomm-ICSI-OGI features for ASRabstractOur feature extraction module for the Aurora task is based on a combination of a conventional noise supression technique (Wiener filtering) with our temporal processing technigues (linear discriminant RASTA filtering and nonlinear TempoRAl Pattern (TRAP) classifier). We observe better than 58% relative error improvement on the prescribed Aurora Digit Task, a performance level that is somewhat better than the new ETSI Advanced Feature standard. Further- more, to test generalization of our approach to an independent test set not available during development, we evaluate performance on American English SpeechDatCar digits and show 10.54% relative improvement over the new ETSI stan- dard. André Adami, Lukás Burget, Stéphane Dupont, Harinath Garudadri, Frantisek Grézl, Hynek Hermansky, Pratibha Jain, Sachin S. Kajarekar, Nelson Morgan, Sunil Sivadas |
INTERSPEECH | 6 |
| 2002 | Distributed speech recognition using noise-robust MFCC and traps-estimated manner features
Pratibha Jain, Hynek Hermansky, Brian Kingsbury |
INTERSPEECH | 2 |
| 2002 | Bark resolution from speech data
Naren Malayath, Hynek Hermansky |
INTERSPEECH | 2 |
| 2002 | Analysis of Information in Speech Based on MANOVAabstractWe propose analysis of information in speech using three sources - language (phone), speaker and channeL Information in speech is measured as mutual information between the source and the set of features extracted from speech signaL We assume that distribu(cid:173) tion of features can be modeled using Gaussian distribution. The mutual information is computed using the results of analysis of variability in speech. We observe similarity in the results of phone variability and phone information, and show that the results of the proposed analysis have more meaningful interpretations than the analysis of variability. Sachin S. Kajarekar, Hynek Hermansky |
NIPS | 2 |
| 2001 | A study of two dimensional linear discriminants for ASRabstractWe study the information in the joint time-frequency domain using 1515 dimensional-15 spectral energies and temporal span of 1s-block of spectrogram as features. In this feature space, we first derive 20 joint linear discriminants (JLDs) using linear discriminant analysis (LDA). Using principal component analysis (PCA), we conclude that information in this block of the spectrogram can be analyzed independently across the time and frequency domains. Under this assumption, we propose a sequential design of two dimensional discriminants (CLDs), i.e., spectral discriminants followed by temporal discriminants. We show that these CLDs are similar to first few JLDs and the discriminant features derived from the CLDs outperform those obtained from JLDs in the continuous-digit recognition task. Sachin S. Kajarekar, Bayya Yegnanarayana, Hynek Hermansky |
ICASSP | 3 |
| 2001 | Robust ASR front-end using spectral-based and discriminant features: experiments on the Aurora tasksabstractThis paper describes an automatic speech recognition frontend that combines low-level robust ASR feature extraction techniques, and higher-level linear and non-linear feature transformations. The low-level algorithms use data-derived filters, mean and variance normalization of the feature vectors, and dropping of noise frames. The feature vectors are then linearly transformed using Principal Components Analysis (PCA). An Artificial Neural Network (ANN) is also used to compute features that are useful for classification of speech sounds. It is trained for phoneme probability estimation on a large corpus of noisy speech. These transformations lead to two feature streams whose vectors are concatenated and then used for speech recognition. This method was tested on the set of speech corpora used for the “Aurora” evaluation. Using the feature stream generated without the ANN yields an overall 41% reduction of the error rate over Mel-Frequency Cepstral Coefficients (MFCC) reference features. Adding the ANN stream further reduces the error rate yielding a 46% reduction over the reference features. M. Carmen Benítez, Lukás Burget, Barry Y. Chen, Stéphane Dupont, Harinath Garudadri, Hynek Hermansky, Pratibha Jain, Sachin S. Kajarekar, Nelson Morgan, Sunil Sivadas |
INTERSPEECH | 6 |
| 2000 | Tandem connectionist feature extraction for conventional HMM systemsabstractHidden Markov model speech recognition systems typically use Gaussian mixture models to estimate the distributions of decorrelated acoustic feature vectors that correspond to individual subword units. By contrast, hybrid connectionist-HMM systems use discriminatively-trained neural networks to estimate the probability distribution among subword units given the acoustic observations. In this work we show a large improvement in word recognition performance by combining neural-net discriminative feature processing with Gaussian-mixture distribution modeling. By training the network to generate the subword probability posteriors, then using transformations of these estimates as the base features for a conventionally-trained Gaussian-mixture based system, we achieve relative error rate reductions of 35% or more on the multicondition Aurora noisy continuous digits task. Hynek Hermansky, Daniel P. W. Ellis, Sangita Sharma |
ICASSP | 1 |
| 2000 | Feature extraction using non-linear transformation for robust speech recognition on the Aurora databaseabstractWe evaluate the performance of several feature sets on the Aurora task as defined by ETSI. We show that after a non-linear transformation, a number of features can be effectively used in a HMM-based recognition system. The non-linear transformation is computed using a neural network which is discriminatively trained on the phonetically labeled (forcibly aligned) training data. A combination of the non-linearly transformed PLP (perceptive linear predictive coefficients), MSG (modulation filtered spectrogram) and TRAP (temporal pattern) features yields a 63% improvement in error rate as compared to baseline me frequency cepstral coefficients features. The use of the non-linearly transformed RASTA-like features, with system parameters scaled down to take into account the ETSI imposed memory and latency constraints, still yields a 40% improvement in error rate. Sangita Sharma, Daniel P. W. Ellis, Sachin S. Kajarekar, Pratibha Jain, Hynek Hermansky |
ICASSP | 5 |
| 2000 | Temporal patterns of critical-band spectrum for text-to-speech
Pratibha Jain, Hynek Hermansky |
INTERSPEECH | 2 |
| 2000 | Optimization of units for continuous-digit recognition task
Sachin S. Kajarekar, Hynek Hermansky |
INTERSPEECH | 2 |
| 2000 | Discriminative MLPs in HMM-based recognition of speech in cellular telephony
Sunil Sivadas, Pratibha Jain, Hynek Hermansky |
INTERSPEECH | 3 |
| 2000 | Relevance of time-frequency features for phonetic and speaker-channel classification
Howard Hua Yang, Sarel van Vuuren, Sangita Sharma, Hynek Hermansky |
Speech Commun. | 4 |
| 1999 | Temporal patterns (TRAPs) in ASR of noisy speechabstractWe study a new approach to processing temporal information for automatic speech recognition (ASR). Specifically, we study the use of rather long-time temporal patterns (TRAPs) of spectral energies in place of the conventional spectral patterns for ASR. The proposed neural TRAPs are found to yield significant amount of complementary information to that of the conventional spectral feature based ASR system. A combination of these two ASR systems is shown to result in improved robustness to several types of additive and convolutive environmental degradations. Hynek Hermansky, Sangita Sharma |
ICASSP | 1 |
| 1999 | Relevancy of time-frequency features for phonetic classification measured by mutual informationabstractIn this paper we use mutual information to study the distribution in time and frequency of information relevant for phonetic classification. A large database of hand-labeled fluent speech is used to (a) compute the mutual information between phoneme labels and a point of logarithmic energy in the time-frequency plane and (b) compute the joint mutual information between phoneme labels and two points of logarithmic energy in the time-frequency plane. Howard Hua Yang, Sarel van Vuuren, Hynek Hermansky |
ICASSP | 3 |
| 1999 | Down-sampling speech representation in ASR
Hynek Hermansky, Pratibha Jain |
EUROSPEECH | 1 |
| 1999 | Analysis of sources of variability in speechabstractThe variability in the speech signal can be attributed to the following sources: (a) Phonetic content, (b) Speaker and Channel, and (c) Coarticulation or context. In this paper, the variability in speech is decomposed using Two Factor Analysis of Variance (ANOVA) with the above mentioned sources as factors. The speech variability is decomposed in temporal and spectral domain separately and structure of these sources of variability in time-frequency plane is described. Although these factors are not indepdendent, it is shown that they can be studied independently after modeling the interaction between the factors. 1 Introduction Speech signal contains various sources of information, e.g., phoneme, context, speaker, environment, etc. Hence, the information in speech can be decomposed into the information in these sources. Usually, not all the sources are relevant to the given task and we believe that decomposing the information in speech can help in suppressing the influence of unwanted... Sachin S. Kajarekar, Naren Malayath, Hynek Hermansky |
EUROSPEECH | 3 |
| 1999 | Speech variability in the modulation spectral domain - SANOVA technique -
Sarel van Vuuren, Hynek Hermansky |
EUROSPEECH | 2 |
| 1999 | Search for Information Bearing Components in Speech
Howard Hua Yang, Hynek Hermansky |
NIPS | 2 |
| 1999 | On the relative importance of various components of the modulation spectrum for automatic speech recognition
Noboru Kanedera, Takayuki Arai, Hynek Hermansky, Misha Pavel |
Speech Commun. | 3 |
| 1999 | Speech enhancement using linear prediction residual
Bayya Yegnanarayana, Carlos Avendaño, Hynek Hermansky, P. Satyanarayana Murthy |
Speech Commun. | 3 |
| 1998 | On properties of modulation spectrum for robust automatic speech recognitionabstractWe report on the effect of band-pass filtering of the time trajectories of spectral envelopes on speech recognition. Several types of filter (linear-phase FIR, DCT, and DFT) are studied. Results indicate the relative importance of different components of the modulation spectrum of speech for ASR. General conclusions are: (1) most of the useful linguistic information is in modulation frequency components from the range between 1 and 16 Hz, with the dominant component at around 4 Hz, (2) it is important to preserve the phase information in the modulation frequency domain, (3) the features which include components at around 4 Hz in the modulation spectrum outperform the conventional delta features, (4) the features which represent the several modulation frequency bands with appropriate center frequency and bandwidth increase recognition performance. Noboru Kanedera, Hynek Hermansky, Takayuki Arai |
ICASSP | 2 |
| 1998 | Enhancement of reverberant speech using LP residualabstractIn this paper we propose a new method of processing speech degraded by reverberation. The method is based on analysis of short (2 ms) segments of data to enhance the regions in the speech signal having high signal to reverberant component ratio (SRR). The short segment analysis shows that SRR is different in different segments of speech. The processing method involves identifying and manipulating the linear prediction residual in three different regions of the speech signal, namely, the high SRR region, the low SRR region and only the reverberation component region. A weighting function is derived to modify the LP residual. The weighted residual samples are used to excite the time-varying LP all-pole filter to obtain perceptually enhanced speech. Bayya Yegnanarayana, P. Satyanarayana Murthy, Carlos Avendaño, Hynek Hermansky |
ICASSP | 4 |
| 1998 | Spectral basis functions from discriminant analysisabstractThe work examines Karhunen-Loeve Transform and Linear Discriminant Analysis as means for designing optimized spectral bases for the projection of the critical-band auditory-like spectrum. 1. INTRODUCTION 1.1. The state-of-art Typical large vocabulary automatic recognition of speech (ASR) consists of three main components: feature extraction, pattern classification, and language modeling. The feature extraction attempts to reduce the information rate of raw speech data by alleviating irrelevant variability such as speaker characteristics or environmental noise, the pattern classification further reduces information rate by classifying each time instant into one of (phoneme-like) subword-unit classes, and language modeling compensates for possible errors of classification by emphasizing more likely word combinations. Over the past two decades we witnessed the introduction of stochastic approaches in both the pattern classification and the language modeling modules. Stochastic technique... Hynek Hermansky, Naren Malayath |
ICSLP | 1 |
| 1998 | TRAPS - classifiers of temporal patternsabstractThe work proposes a radically different set of features for ASR where TempoRAl Patterns of spectral energies are used in place of the conventional spectral patterns. The approach has several inherent advantages, among them robustness to stationary or slowly varying disturbances. 1. INTRODUCTION 1.1. Spectral features In 1665 Isaac Newton made the following observation: 'The filling of a very deepe flaggon with a constant streame of beere or water sounds yer vowells in this order w, u, !, o, a, e, i, y' [8]. What young Newton observed was the spectral resonance peak which enhanced the spectrum of the beer pouring sound and moved up in frequency as the "deepe flaggon" was filling up. Since then, attempts to find acoustic correlates of phonetic categories mostly followed Newton's lead and studied the spectrum of speech. Spectrum-based techniques form the basis of most feature extraction methods in current ASR. A problem with the spectrum of sound is that it can easily be modified by v... Hynek Hermansky, Sangita Sharma |
ICSLP | 1 |
| 1998 | On the importance of components of the modulation spectrum for speaker verificationabstractWe provide an analysis of the relative importance of components of the modulation spectrum for speaker verification. The aim is to remove less relevant components and reduce system sensitivity to acoustic disturbances while improving verification accuracy. Spectral components between 0.1 Hz and 10 Hz are found to contain the most useful speaker information. We discuss this result in the context of RASTA processing and cepstral mean subtraction. When compared to cepstral mean subtraction that retains components up to 50 Hz, lowpass filtering to 10 Hz with downsampling by 75% is found to significantly improve robustness in mismatched conditions. The downsampling results in a large computational savings. 1. INTRODUCTION Many speaker verification systems attempt to characterize a speaker using acoustic features based on filtered logarithmic spectral energies derived from a short-time analysis [1, 5, 2]. Spectral components of the time sequences of logarithmic spectral energies, aka. the ... Sarel van Vuuren, Hynek Hermansky |
ICSLP | 2 |
| 1998 | Should recognizers have ears?
Hynek Hermansky |
Speech Commun. | 1 |
| 1997 | Sub-band based recognition of noisy speechabstractA new approach to automatic speech recognition based on independent class-conditional probability estimates in several frequency sub-bands is presented. The approach is shown to be especially applicable to environments which cause partial corruption of the frequency spectrum of the signal. Some of the issues involved in the implementation of the approach are also addressed. Sangita Tibrewala, Hynek Hermansky |
ICASSP | 2 |
| 1997 | Multiresolution channel normalization for ASR in reverberant environmentsabstractTo overcome the problems related with the long impulse responses produced by reverberation, we use a long time window (high frequency resolution) analysis during the channel normalization steps of the feature extraction process in automatic speech recognition (ASR). After normalization, a trade between frequency and time resolution is used to increase the rate at which the time information is sampled (short-time domain), yielding an appropriate domain to derive ASR features. Experiments on data with reverberation times of about 0:5 s show that the new technique achieves significant performance improvement of a speech recognizer under reverberation, with only some performance degradation on clean speech. 1. INTRODUCTION In spite of the efforts by many researches during the last 50 years, there has been very little success in reducing the effects of reverberation in speech communications. The effects of reverberation on the details of the speech signal are complicated and difficult to ... Carlos Avendaño, Sangita Tibrewala, Hynek Hermansky |
EUROSPEECH | 3 |
| 1997 | On the importance of various modulation frequencies for speech recognition
Noboru Kanedera, Takayuki Arai, Hynek Hermansky, Misha Pavel |
EUROSPEECH | 3 |
| 1997 | Towards decomposing the sources of variability in speech
Naren Malayath, Hynek Hermansky, Alexander Kain |
EUROSPEECH | 2 |
| 1997 | Multi-band and adaptation approaches to robust speech recognition
Sangita Tibrewala, Hynek Hermansky |
EUROSPEECH | 2 |
| 1997 | Data-driven design of RASTA-like filtersabstractWe describe use of Linear Discriminant Analysis LDA for data-driven automatic design of RASTA-like lters.The LDA applied to rather long segments of time trajectories of critical-band energies yields FIR lters to be applied to these time trajectories in the feature extraction module.Frequency responses of the rst three discriminant v ectors are in principle consistent with the ad hoc designed RASTA, delta and double-delta lters.On a connected digit task the new features outperform the original RASTA processing. Sarel van Vuuren, Hynek Hermansky |
EUROSPEECH | 2 |
| 1997 | Processing linear prediction residual for speech enhancementabstractIn this paper we propose a method for enhancement of speech in the presence of additive noise. The objective is to selectively enhance the high SNR regions in the noisy speech in the temporal and spectral domains, without causing significant distortion in the resulting enhanced speech. This is proposed to be done at three different levels: (a) At the gross level, by identifying the regions of speech and noise in the temporal domain, (b) At the finer level, by identifying the regions of high and low SNR portions in the noisy speech, and (c) At the short--time spectrum level, by enhancing the spectral peaks over spectral valleys. Processing of noisy speech for enhancement involves mostly weighting the LP residual samples. The weighted residual samples are used to excite the time-- varying LP filter to produce enhanced speech. 1. INTRODUCTION Speech signal collected under normal environmental conditions is usually degraded due to noise and distortions. Performance of speech systems depe... Bayya Yegnanarayana, Carlos Avendaño, Hynek Hermansky, P. Satyanarayana Murthy |
EUROSPEECH | 3 |
| 1997 | On the effects of short-term spectrum smoothing in channel normalizationabstractWe present a simple analysis showing that channel normalization techniques are less effective when applied to spectral energies obtained by (weighted) summation of components of the short-time Fourier power spectrum of speech. We show that applying channel normalization processing prior to critical band integration or linear predictive all-pole modeling improves the effectiveness of the techniques. Carlos Avendaño, Hynek Hermansky |
IEEE Trans. Speech Audio Process. | 2 |
| 1996 | Intelligibility of speech with filtered time trajectories of spectral envelopes
Takayuki Arai, Misha Pavel, Hynek Hermansky, Carlos Avendaño |
ICSLP | 3 |
| 1996 | Study on the dereverberation of speech based on temporal envelope filteringabstractIn this paper we explore speech dereverberation techniques whose principle is the recovery of the envelope modulations of the original (anechoic) speech [ 3 ], [4].Based on our previous experience with such kind of processing for additive noise reduction applications, we apply a data designed lterbank technique [2] to the reverberant speech.Comparing our results with other works we discuss the eectiveness and limitations of this type of approaches. Carlos Avendaño, Hynek Hermansky |
ICSLP | 2 |
| 1996 | Data based filter design for RASTA-like channel normalization in ASR
Carlos Avendaño, Sarel van Vuuren, Hynek Hermansky |
ICSLP | 3 |
| 1996 | Towards ASR on partially corrupted speechabstractA new highly parallel approach to automatic recognition of speech, inspired by early Fletcher's research on Articulation Index, and based on independent probability estimates in several sub-bands of the available speech spectrum, is presented. The approach is especially suitable for situations when part of the spectrum of speech is corrupted. In such cases, it can yield an order-of-magnitude improvement in the error rate over a conventional full-band recognizer. 1. Hynek Hermansky, Sangita Tibrewala, Misha Pavel |
ICSLP | 1 |
| 1996 | Towards increasing speech recognition error rates
Hervé Bourlard, Hynek Hermansky, Nelson Morgan |
Speech Commun. | 2 |
| 1995 | Speech enhancement based on temporal processingabstractFinite impulse response (FIR) Wiener-like filters are applied to time trajectories of the cubic-root compressed short-term power spectrum of noisy speech recorded over cellular telephone communications. Informal listenings indicate that the technique brings a noticeable improvement to the quality of processed noisy speech while not causing any significant degradation to clean speech. Alternative filter structures are being investigated as well as other potential applications in cellular channel compensation and narrowband to wideband speech mapping. Hynek Hermansky, Eric A. Wan, Carlos Avendaño |
ICASSP | 1 |
| 1995 | Stochastic perceptual models of speechabstractWe have developed a statistical model of speech (based on auditory perceptual criteria) that avoids a number of current constraining assumptions for statistical speech recognition systems, particularly the model of speech as a sequence of stationary segments consisting of uncorrelated acoustic vectors. We further wish to focus statistical modeling power on perceptually-dominant and information-rich portions of the speech signal, which may also be the parts of the speech signal with a better chance to withstand adverse acoustical conditions. We describe some of the theory, along with some preliminary experiments. These experiments suggest that the regions of acoustic signal containing significant spectral change are critical to the recognition of continuous speech. Nelson Morgan, Hervé Bourlard, Steven Greenberg, Hynek Hermansky, Su-Lin Wu |
ICASSP | 4 |
| 1995 | Beyond NYQUIST: towards the recovery of broad-bandwidth speech from narrow-bandwidth speechabstractA new technique is presented which improves the subjective quality of band-limited speech. The approach is based on a linear model of speech production, in which we independently estimate the spectral envelope and excitation function for a broad-bandwidth speech signal to reconstruct missing frequency components in narrow-bandwidth speech. Carlos Avendaño, Hynek Hermansky, Eric A. Wan |
EUROSPEECH | 2 |
| 1995 | The challenge of spoken language systems: research directions for the ninetiesabstractA spoken language system combines speech recognition, natural language processing and human interface technology. It functions by recognizing the person's words, interpreting the sequence of words to obtain a meaning in terms of the application, and providing an appropriate response back to the user. Potential applications of spoken language systems range from simple tasks, such as retrieving information from an existing database (traffic reports, airline schedules), to interactive problem solving tasks involving complex planning and reasoning (travel planning, traffic routing), to support for multilingual interactions. We examine eight key areas in which basic research is needed to produce spoken language systems: (1) robust speech recognition; (2) automatic training and adaptation; (3) spontaneous speech; (4) dialogue models; (5) natural language response generation; (6) speech synthesis and speech generation; (7) multilingual systems; and (8) interactive multimodal systems. In each area, we identify key research challenges, the infrastructure needed to support research, and the expected benefits. We conclude by reviewing the need for multidisciplinary research, for development of shared corpora and related resources, for computational support and far rapid communication among researchers. The successful development of this technology will increase accessibility of computers to a wide range of users, will facilitate multinational communication and trade, and will create new research specialties and jobs in this rapidly expanding area.> Ronald A. Cole, Lynette Hirschman, Les E. Atlas, Mary E. Beckman, Alan Biermann, Marcia A. Bush, Mark A. Clements, Jordan Cohen, Oscar Garcia, Brian A. Hanson, Hynek Hermansky, Steve Levinson, Kathy McKeown, Nelson Morgan, David G. Novick, Mari Ostendorf, Sharon L. Oviatt, Patti Price, Harvey F. Silverman, Judy Spitz, Alex Waibel, Clifford J. Weinstein, Stephen A. Zahorian, Victor Zue |
IEEE Trans. Speech Audio Process. | 11 |
| 1994 | Integrating RASTA-PLP into speech recognitionabstractIn previous work, we and others have shown that bandpass filtering of temporal trajectories of simple functions of the critical band spectrum can lead to more robust speech recognizers in the presence of additive and convolutional error. In this study we report results on several mechanisms for incorporating this analysis technique into training, in a way that is consistent with on-line approaches to speech recognition. In particular, we show improved robustness to these forms of degradation for a system that maps the filtered spectral points using a linear regression computed from results of the different transformations.> Joachim Koehler, Nelson Morgan, Hynek Hermansky, Hans-Günter Hirsch, Grace Tong |
ICASSP (1) | 3 |
| 1994 | Stochastic perceptual auditory-event-based models for speech recognition
Nelson Morgan, Hervé Bourlard, Steven Greenberg, Hynek Hermansky |
ICSLP | 4 |
| 1994 | RASTA processing of speechabstractPerformance of even the best current stochastic recognizers severely degrades in an unexpected communications environment. In some cases, the environmental effect can be modeled by a set of simple transformations and, in particular, by convolution with an environmental impulse response and the addition of some environmental noise. Often, the temporal properties of these environmental effects are quite different from the temporal properties of speech. We have been experimenting with filtering approaches that attempt to exploit these differences to produce robust representations for speech recognition and enhancement and have called this class of representations relative spectra (RASTA). In this paper, we review the theoretical and experimental foundations of the method, discuss the relationship with human auditory perception, and extend the original method to combinations of additive noise and convolutional noise. We discuss the relationship between RASTA features and the nature of the recognition models that are required and the relationship of these features to delta features and to cepstral mean subtraction. Finally, we show an application of the RASTA technique to speech enhancement.> Hynek Hermansky, Nelson Morgan |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | Recognition of speech in additive and convolutional noise based on RASTA spectral processing
Hynek Hermansky, Nelson Morgan, Hans-Günter Hirsch |
ICASSP (2) | 1 |
| 1993 | Evaluation and optimization of perceptually-based ASR front-endabstractSeveral recently proposed automatic speech recognition (ASR) front-ends are experimentally compared in speaker-dependent, speaker-independent (or cross-speaker) recognition. The perceptually based linear predictive (PLP) front-end, with the root-power sums (RPS) distance measure, yields generally the highest accuracies, especially in cross-speaker recognition., It is experimentally shown that one can optimize the system and further improve recognition accuracy for speaker-independent recognition by controlling the distance measure's sensitivity to spectral peaks and the spectral tilt and by utilizing the speech dynamic features. For a digit vocabulary and five reference templates obtained with a clustering algorithm, the optimization improves recognition accuracy from 97% to 98.1%, with respect to the PL-PRPS front-end.> Jean-Claude Junqua, Hisashi Wakita, Hynek Hermansky |
IEEE Trans. Speech Audio Process. | 3 |
| 1992 | RASTA-PLP speech analysis techniqueabstractMost speech parameter estimation techniques are easily influenced by the frequency response of the communication channel. The authors have developed a technique that is more robust to such steady-state spectral factors in speech. The approach is conceptually simple and computationally efficient. The new method is described, and experimental results are proposed that show significant advantages for the proposed method.> Hynek Hermansky, Nelson Morgan, Aruna Bayya, Phil Kohn |
ICASSP | 1 |
| 1992 | Towards handling the acoustic environment in spoken language processing
Hynek Hermansky, Nelson Morgan |
ICSLP | 1 |
| 1991 | Continuous speech recognition using PLP analysis with multilayer perceptronsabstractThe authors investigate the use of continuous features derived by perceptual linear predictive (PLP) analysis, examine the effect of adding temporal features, and compare it to the previously studied use of multiframe input. Comparisons of the MLP (multilayer perceptron) and conventional Gaussian classifiers are also reported. The speaker-dependent portion of the Resource Management database was used for this test. Additionally, some experiments were performed with a perplexity-2200 speaker-independent recognition task on a subset of the TIMIT database. In each case, the PLP features were used as input to the networks. The experiments show the advantage of continuous PLP features and their first and second temporal derivatives.> Nelson Morgan, Hynek Hermansky, Hervé Bourlard, Phil Kohn, Chuck Wooters |
ICASSP | 2 |
| 1991 | Perceptual linear predictive (PLP) analysis-resynthesis technique
Hynek Hermansky, Louis Anthony Cox Jr. |
EUROSPEECH | 1 |
| 1991 | Compensation for the effect of the communication channel in auditory-like analysis of speech (RASTA-PLP)
Hynek Hermansky, Nelson Morgan, Aruna Bayya, Phil Kohn |
EUROSPEECH | 1 |
| 1990 | Towards feature-based speech metricabstractA speech metric which directly uses spectral features such as spectral peak frequencies and bandwidths is proposed and evaluated. The spectral features either are derived directly by solving the all-pole model polynomial to get spectral peak frequencies and bandwidths and fitting the linear regression line to the logarithmic spectrum of the model or are estimated as a linear combination of the several lower cepstral coefficients of the all-pole model spectrum. The performance of the studied metric in speaker-independent speech recognition of telephone-quality speech approaches the performance of the best weighted cepstral metrics.> Aruna Bayya, Hynek Hermansky |
ICASSP | 2 |
| 1989 | The effective second formant F2' and the vocal tract front-cavityabstractThe authors advance the hypothesis that the equivalent perceptual second formant F2' carries information about the front cavity of the vocal tract. They note a previous result that peaks found by perceptually based linear predictive (PLP) analysis closely track the F2'. The authors recall P. Mermelstein's (1967) observation of the relative invariance of the front cavity shape and size in vowels synthesized to maintain a constant F-pattern under variations of the vocal tract length. They show by tracing X-rays that humans with radically different vocal tract lengths tend to preserve the front cavity shape and size when producing vowels with identical phonetic values. Articulatory synthesis with variable vocal tract lengths is used to demonstrate that the F2' from the PLP model follows the resonance frequency of the front cavity.> Hynek Hermansky, David J. Broad |
ICASSP | 1 |
| 1988 | Optimization of perceptually-based ASR front-end [automatic speech recognition]abstractSeveral recently proposed automatic speech recognition (ASR) front-ends are experimentally compared for speaker-dependent and cross-speaker ASR. The perceptually based linear predictive front-end yields the highest accuracies. By modifying its sensitivity to spectral peaks and to spectral tilt and by utilizing the speech dynamics the authors further improve, by about 10%, its error rate in speaker-independent ASR.> Hynek Hermansky, Jean-Claude Junqua |
ICASSP | 1 |
| 1987 | An efficient speaker-independent automatic speech recognition by simulation of some properties of human auditory perceptionabstractAn auditory model of speech perception, the Perceptually based linear predictive analysis with Root power sum metric (PLP-RPS), is applied as the front-end of an automatic speech recognizer (ASR). The PLP-RPS front-end is compared with standard linear predictive-cepstral metric (LP-CEP) front-end, and with LP-RPS and PLP-CEP front-ends. The two-spectral-peak models are the most efficient in modeling of linguistic information in speech. Consequently, in speaker-independent ASR, high analysis order front-ends are less effective than low-order front-ends. Synthetic speech is used for front-end evaluation. Some of perceptual inconsistencies of standard LP front-ends are alleviated in PLP front-ends. The PLP-RPS front-end is most sensitive to harmonic structure of speech spectrum. Perceptual experiments indicate similar tendencies in human auditory perception. Hynek Hermansky |
ICASSP | 1 |
| 1986 | Perceptually based processing in automatic speech recognitionabstractThe perceptually based linear predictive (PLP) speech analysis method is applied to isolated word automatic speech recognition (ASR). Low dimensionality of the PLP analysis vector, which is otherwise identical in form to the standard linear predictive (LP) analysis vector, allows for computational and storage savings in ASR. We show that in speaker-dependent recognition of the alpha-numeric vocabulary, the PLP method in VQ-based ASR yields similar recognition scores as does the standard ASR system. The main focus of the paper is on cross-speaker ASR. We demonstrate in experiments with vowel centroids of two male and one female speakers that PLP speech representation is more consistent with the underlying phonetic information than the standard LP method. Conclusions from the experiments are confirmed by superior performance of the PLP method in cross-speaker isolated word recognition. Hynek Hermansky, Kazuhiro Tsuga, Shozo Makino, Hisashi Wakita |
ICASSP | 1 |
| 1985 | Perceptually based linear predictive analysis of speechabstractA novel speech analysis method which uses several established psychoacoustic concepts, the perceptually based linear predictive analysis (PLP), models the auditory spectrum by the spectrum of the low-order all-pole model. The auditory spectrum is derived from the speech waveform by critical-band filtering, equal-loudness curve pre-emphasis, and intensity-loudness root compression. We demonstrate through analysis of both synthetic and natural speech that psychoacoustic concepts of spectral auditory integration in vowel perception, namely the F1, F2' concept of Carlson and Fant and the 3.5 Bark auditory integration concept of Chistovich, are well modeled by the PLP method. A complete speech analysis-synthesis system based on the PLP method is also described in the paper. Hynek Hermansky, Brian A. Hanson, Hisashi Wakita |
ICASSP | 1 |
| 1985 | Low-dimensional representation of vowels based on all-pole modeling in the psychophysical domain
Hynek Hermansky, Brian A. Hanson, Hisashi Wakita |
Speech Commun. | 1 |
| 1984 | Spectral envelope sampling and interpolation in linear predictive analysis of speechabstractIn spite of its extensive use, speech analysis based on linear prediction (LP) is liable to various causes of inaccuracy. This paper presents a novel approach to improve the accuracy in the estimation of the voiced speech production model based on the LP method. The presented method uses interpolation between spectral points which are least influenced by artifacts in the spectral analysis and by noise in the signal. We show, on analyses of both synthetic and natural speech, that the averaged parabolic approximation between harmonic peaks of voiced speech spectrum reduces the sensitivity of the LP analysis to changes in the fundamental frequency Fo and to noise. The method is well suited for combination with the Spectral Transform LP method, previously proposed by the authors [1]. Hynek Hermansky, Hiroya Fujisaki, Yasuo Sato |
ICASSP | 1 |
| 1983 | Analysis and synthesis of speech based on spectral transform linear predictive methodabstractThe Linear Predictive (LP) method has been widely used in speech analysis, mainly because of the simple mathematical formulation of the model and the straightforward computation of its parameters. However, there still remain certain difficulties that cause errors in the result of the analysis. Often encountered are the errors due to harmonic structure of the excitation source. The Spectral Transform LP (STLP) method proposed in the present paper aims at reducing these errors. Amplitude transforms on the input spectrum and on the spectrum of the model are introduced to modify the error criterion and the model adopted in the standard LP analysis. We show by analyses of both synthetic and natural speech that the STLP method offers significant improvement over the standard LP method. A method of STLP speech synthesis using the standard LP model is proposed. A perceptual experiment confirms the superiority of the STLP method in analysis of speech. Hynek Hermansky, Hiroya Fujisaki, Yasuo Sato |
ICASSP | 1 |