EDBT 2026 Demo / reviewers in the wild / expert
Lalit R. Bahl
dblp:03/6454
· DBLP profile ↗
64ranked-venue papers
44as first author
0since 2021 · last 1999
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 31 first-authorArtificial intelligence and machine learning · 15 · 9 first-authorTheory of computation · 11 · 8 first-authorSystems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Speech recognition and synthesis · 64% Language models and text generation · 17% Probabilistic and Bayesian machine learning · 10% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 67% Image and video processing · 33% | |
| Theoretical computer science
10 papers |
Coding theory · 94% Algorithms and data structures · 3% Information theory · 2% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.0 | 3 | 1999 | Partitioning the feature space of a classifier with linear hyperplanes · IEEE Trans. Speech Audio Process. 1999 A method for the construction of acoustic Markov models for words · IEEE Trans. Speech Audio Process. 1993 Multonic Markov word models for large vocabulary continuous speech recognition · IEEE Trans. Speech Audio Process. 1993 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 5 | 1993 | A fast approximate acoustic match for large vocabulary speech recognition · IEEE Trans. Speech Audio Process. 1993 A method for the construction of acoustic Markov models for words · IEEE Trans. Speech Audio Process. 1993 Estimating hidden Markov model parameters so as to maximize speech recognition accuracy · IEEE Trans. Speech Audio Process. 1993 |
Natural language and speech › Language models and text generation
decoding |
0.0 | 1 | 1999 | Partitioning the feature space of a classifier with linear hyperplanes · IEEE Trans. Speech Audio Process. 1999 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › feature selection
feature grouping |
0.0 | 1 | 1999 | Partitioning the feature space of a classifier with linear hyperplanes · IEEE Trans. Speech Audio Process. 1999 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.0 | 1 | 1999 | Partitioning the feature space of a classifier with linear hyperplanes · IEEE Trans. Speech Audio Process. 1999 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
speaker adaptation |
0.0 | 1 | 1998 | Speaker clustering and transformation for speaker adaptation in speech recognition systems · IEEE Trans. Speech Audio Process. 1998 |
Natural language and speech › Speech recognition and synthesis › speaker diarization
speaker clustering |
0.0 | 1 | 1998 | Speaker clustering and transformation for speaker adaptation in speech recognition systems · IEEE Trans. Speech Audio Process. 1998 |
Natural language and speech › Language models and text generation › tokenization
subword units |
0.0 | 2 | 1993 | A method for the construction of acoustic Markov models for words · IEEE Trans. Speech Audio Process. 1993 Multonic Markov word models for large vocabulary continuous speech recognition · IEEE Trans. Speech Audio Process. 1993 |
Image and video processing
data augmentation |
0.0 | 1 | 1994 | The metamorphic algorithm: a speaker mapping approach to data augmentation · IEEE Trans. Speech Audio Process. 1994 |
Audio and music processing › speech recognition
speaker adaptation |
0.0 | 1 | 1994 | The metamorphic algorithm: a speaker mapping approach to data augmentation · IEEE Trans. Speech Audio Process. 1994 |
Audio and music processing
speech recognition |
0.0 | 1 | 1994 | The metamorphic algorithm: a speaker mapping approach to data augmentation · IEEE Trans. Speech Audio Process. 1994 |
Natural language and speech › Speech recognition and synthesis
acoustic model training |
0.0 | 1 | 1993 | Estimating hidden Markov model parameters so as to maximize speech recognition accuracy · IEEE Trans. Speech Audio Process. 1993 |
Natural language and speech › Speech recognition and synthesis › acoustic model training › discriminative training
maximum mutual information estimation |
0.0 | 1 | 1993 | Estimating hidden Markov model parameters so as to maximize speech recognition accuracy · IEEE Trans. Speech Audio Process. 1993 |
Natural language and speech › Speech recognition and synthesis
search and decoding |
0.0 | 1 | 1993 | A fast approximate acoustic match for large vocabulary speech recognition · IEEE Trans. Speech Audio Process. 1993 |
Coding theory › error-correcting codes › decoding
decoding algorithms |
0.0 | 3 | 1975 | Decoding for channels with insertions, deletions, and substitutions with applications to speech recognition · IEEE Trans. Inf. Theory 1975 Optimal decoding of linear codes for minimizing symbol error rate (Corresp.) · IEEE Trans. Inf. Theory 1974 Multiple-Burst-Error Correction by Threshold Decoding · Inf. Control. 1969 |
Coding theory › error-correcting codes
convolutional codes |
0.0 | 4 | 1974 | On the structure of rate 1/n convolutional codes · IEEE Trans. Inf. Theory 1972 An efficient algorithm for computing free distance (Corresp.) · IEEE Trans. Inf. Theory 1972 Rate 1/2 convolutional codes with complementary generators · IEEE Trans. Inf. Theory 1971 |
Coding theory
error-correcting codes |
0.0 | 4 | 1971 | Single- and multiple-burst-correcting properties of a class of cyclic product codes · IEEE Trans. Inf. Theory 1971 Correction of two erasure bursts (Corresp.) · IEEE Trans. Inf. Theory 1969 On Gilbert burst-error-correcting codes (Corresp.) · IEEE Trans. Inf. Theory 1969 |
Network optimization and economics › network design
concentrator location |
0.0 | 1 | 1978 | Optimization of Teleprocessing Networks with Concentrators and Multiconnected Terminals · IEEE Trans. Computers 1978 |
Network optimization and economics
network design |
0.0 | 1 | 1978 | Optimization of Teleprocessing Networks with Concentrators and Multiconnected Terminals · IEEE Trans. Computers 1978 |
Network optimization and economics
resource allocation |
0.0 | 1 | 1978 | Optimization of Teleprocessing Networks with Concentrators and Multiconnected Terminals · IEEE Trans. Computers 1978 |
Coding theory › error-correcting codes
burst error correction |
0.0 | 3 | 1971 | Single- and multiple-burst-correcting properties of a class of cyclic product codes · IEEE Trans. Inf. Theory 1971 On Gilbert burst-error-correcting codes (Corresp.) · IEEE Trans. Inf. Theory 1969 Multiple-Burst-Error Correction by Threshold Decoding · Inf. Control. 1969 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
acoustic-phonetic modeling |
0.0 | 1 | 1975 | Design of a linguistic statistical decoder for the recognition of continuous speech · IEEE Trans. Inf. Theory 1975 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
continuous speech recognition |
0.0 | 1 | 1975 | Design of a linguistic statistical decoder for the recognition of continuous speech · IEEE Trans. Inf. Theory 1975 |
Coding theory › error-correcting codes › insertion and deletion
insertion-deletion channel |
0.0 | 1 | 1975 | Decoding for channels with insertions, deletions, and substitutions with applications to speech recognition · IEEE Trans. Inf. Theory 1975 |
Coding theory › error-correcting codes › decoding › sequential decoding
stack decoding |
0.0 | 1 | 1975 | Decoding for channels with insertions, deletions, and substitutions with applications to speech recognition · IEEE Trans. Inf. Theory 1975 |
Coding theory › error-correcting codes › decoding › iterative decoding › soft-input soft-output decoding
BCJR algorithm |
0.0 | 1 | 1974 | Optimal decoding of linear codes for minimizing symbol error rate (Corresp.) · IEEE Trans. Inf. Theory 1974 |
Coding theory › source coding › source modeling
markov sources |
0.0 | 1 | 1974 | Optimal decoding of linear codes for minimizing symbol error rate (Corresp.) · IEEE Trans. Inf. Theory 1974 |
Algorithms and data structures › search algorithms
bidirectional search |
0.0 | 1 | 1972 | An efficient algorithm for computing free distance (Corresp.) · IEEE Trans. Inf. Theory 1972 |
Coding theory › error-correcting codes › convolutional codes › free distance
free distance bounds |
0.0 | 1 | 1972 | On the structure of rate 1/n convolutional codes · IEEE Trans. Inf. Theory 1972 |
Coding theory › error-correcting codes › convolutional codes › free distance
free distance computation |
0.0 | 1 | 1972 | An efficient algorithm for computing free distance (Corresp.) · IEEE Trans. Inf. Theory 1972 |
Methods — techniques the papers use, named apart from their topics
mutual information · 0.0linear discriminant analysis · 0.0decision tree · 0.0linear transformation · 0.0gaussian mean reestimation · 0.0forward-backward algorithm · 0.0hidden markov model · 0.0speaker normalization · 0.0piecewise linear mapping · 0.0parameter reestimation · 0.0error-correcting training · 0.0dynamic programming matching · 0.0upper bound derivation · 0.0stack decoding · 0.0sequential assignment algorithm · 0.0markov chain inference · 0.0likelihood function derivation · 0.0dynamic programming · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1999 | Partitioning the feature space of a classifier with linear hyperplanesabstractWe describe the design and use of linear hyperplanes to partition the feature space of a classifier. The objective of the partitioning is to minimize the average entropy of the class distribution in the final partitions. The hyperplanes are characterized by a vector /spl nu//sub n/ and scalar h/sub n/, which are computed with the objective of maximizing the mutual information associated with the partitioning. We show that the problem of designing the /spl nu//sub n/ can be simplified to an approximately equivalent problem, the solution of which, interestingly enough, turns out to be the linear discriminant of the data. We also describe a decision-tree based technique that partitions the feature space in a hierarchical fashion where each node in the decision tree represents a linear hyperplane that further partitions the feature space into two regions. The end result of the tree-growing process is that the entire feature space is partitioned into nonoverlapping regions where each region is bounded by a number of hyperplanes. As the criterion for the design of the hyperplanes is the minimization of the average class entropy of the regions, each region is characterized by the occurrence of only a small number of classes. The partitioning information provided by the decision-tree can then be used to eliminate a number of classes from being considered, and hence simplifies the job of the classifier. We show the application of the decision-tree as a preprocessor to a classifier in the speech recognition problem. Here the classes are modeled with mixtures of tens to hundreds of thousands of Gaussians, and the use of the partitioning information reduces the computation associated with the classification process by factors larger than 20 with negligible degradation in the word error rate (note, however, that the overall decoding process comprises other steps in addition to the classification-the computational speedup mentioned above does not affect the computation of these other steps). We also compare the performance of this quantization scheme with standard Gaussian clustering schemes and show that the decision-tree based quantization provides better performance in most cases. Mukund Padmanabhan, Lalit R. Bahl, David Nahamoo |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | A discriminant measure for model complexity adaptationabstractWe present a discriminant measure that can be used to determine the model complexity in a speech recognition system. In the speech recognition process, given a test feature vector the conditional probability of the feature vector has to be obtained for several allophone (sub-phonetic units) classes using a Gaussian-mixture density model for each class. The Gaussian-mixture models are constructed from the training data belonging to the allophone classes, and the number of mixture components that are required to adequately model the PDF of each class is determined by using some simple rule of thumb-for instance the number of components has to be sufficient to model the data reasonably well but not so many as to overmodel the data. A typical example of the choice of the number is to make it proportional to the number of data samples. However, such methods may result in models that are sub-optimal as far as classification accuracy is concerned. We present a new discriminant measure that can be used to determine in an objective fashion, the number of Gaussians required to best model the PDF of an allophone class. We also present the results of experiments showing the improvement in recognition performance when the number of mixture components is chosen based on the discriminant measure as opposed to the rule of thumb. These results are presented both for the speaker-independent and speaker-adapted case. Lalit R. Bahl, Mukund Padmanabhan |
ICASSP | 1 |
| 1998 | Speech recognition performance on a voicemail transcription taskabstractWe describe a new testbed for developing speech recognition algorithms-the ARRPA-sponsored voicemail transcription task, analogous to other tasks such as the Switchboard, CallHome and the Hub 4 tasks. The task involves the transcription of voicemail conversations. Voicemail represents a very large volume of real-world speech data, which is however not particularly well represented in existing databases. For instance, the Switchboard and CallHome databases contain telephone conversations between two humans, representing telephone-bandwidth spontaneous speech; the Hub 4 database contains radio broadcasts which represents different kinds of speech data such as spontaneous speech from a well-trained speaker, conversations between two humans possibly over the telephone, etc. The voicemail database on the other hand also represents telephone bandwidth spontaneous speech, however the difference with respect to the Switchboard and CallHome tasks is that the interaction is not between two humans, but rather between a human and a machine-consequently, the speech is expected to be a little more formal in its nature, without the problems of crosstalk, barge-in etc. This eliminates some of the variables and provides more controlled conditions enabling one to concentrate on the aspects of spontaneous speech and effects of the telephone channel. We describe the modality of collection of the speech data, and some algorithmic techniques that were devised based on this data. We also describe the initial results of the transcription performance on this task. Mukund Padmanabhan, Ellen Eide, Bhuvana Ramabhadran, Ganesh N. Ramaswamy, Lalit R. Bahl |
ICASSP | 5 |
| 1998 | Acoustics-only based automatic phonetic baseform generationabstractPhonetic baseforms are the basic recognition units in most speech recognition systems. These baseforms are usually determined by linguists once a vocabulary is chosen and not modified thereafter. However, several applications, such as name dialing, require the user be able to add new words to the vocabulary. These new words are often names, or task-specific jargon, that have user-specific pronunciations. This paper describes a novel method for generating phonetic transcriptions (baseforms) of words based on acoustic evidence alone. It does not require either the spelling or any prior acoustic representation of the new word, is vocabulary independent, and does not have any linguistic constraints (pronunciation rules). Our experiments demonstrate the high decoding accuracies obtained when baseforms deduced using this approach are incorporated into our speech recognizer. Also, the error rates on the added words were found to be comparable to or better than when the baseforms were derived by hand. Bhuvana Ramabhadran, Lalit R. Bahl, Peter DeSouza, Mukund Padmanabhan |
ICASSP | 2 |
| 1998 | A method for modeling liaison in a speech recognition system for FrenchabstractIn French the pronunciations of many words change dramatically depending on the word immediately preceding it. The result of this phenomenon, known as “liaison”, in an ASR system that does not model “liaison” is the requirement of unnatural pronunciation and much user dissatisfaction. We present, in this paper, the development of an acoustic model which takes into account the wide variability of word pronunciations caused by the liaison, the integration of this model into a French continuous speech recognition system and decoding results. Lalit R. Bahl, Steven V. De Gennaro, Pieter de Souza, Edward A. Epstein, J. M. Le Roux, Burn L. Lewis, Claire Waast-Richard |
ICSLP | 1 |
| 1998 | A time-synchronous, tree-based search strategy in the acoustic fast match of an asynchronous speech recognition system
Ellen Eide, Lalit R. Bahl |
ICSLP | 2 |
| 1998 | Speaker clustering and transformation for speaker adaptation in speech recognition systemsabstractA speaker adaptation strategy is described that is based on finding a subset of speakers, from the training set, who are acoustically close to the test speaker, and using only the data from these speakers (rather than the complete training corpus) to reestimate the system parameters. Further, a linear transformation is computed for every one of the selected training speakers to better map the training speaker's data to the test speaker's acoustic space. Finally, the system parameters (Gaussian means) are reestimated specifically for the test speaker using the transformed data from the selected training speakers. Experiments showed that this scheme is capable of providing an 18% relative improvement in the error rate on a large-vocabulary task with the use of as little as three sentences of adaptation data. Mukund Padmanabhan, Lalit R. Bahl, David Nahamoo, Michael Picheny |
IEEE Trans. Speech Audio Process. | 2 |
| 1997 | Decision-tree based quantization of the feature space of a speech recognizerabstractWe present a decision-tree based procedure to quantize the feature-space of a speech recognizer, with the motivation of reducing the computation time required for evaluating gaussians in a speech recognition system. The entire feature space is quantized into non overlapping regions where each region is bounded by a number of hyperplanes. Further, each region is characterized by the occurence of only a small number of the total alphabet of allophones (sub-phonetic speech units); by identifying the region in which a test feature vector lies, only the gaussians that model the density of allophones that exist in that region need be evaluated. The quantization of the feature space is done in a heirarchical manner using a binary decision tree. Each node of the decision tree represents a region of the feature space, and is further characterized by a hyperplane (a vector v n and a scalar threshold value hn ), that subdivides the region corresponding to the current node into two non-overlapping... Mukund Padmanabhan, Lalit R. Bahl, David Nahamoo, Pieter de Souza |
EUROSPEECH | 2 |
| 1996 | Discriminative training of Gaussian mixture models for large vocabulary speech recognition systemsabstractTwo discriminative techniques are described (and evaluated) for estimating the parameters of the Gaussians in a large vocabulary speech-recognition system. The first technique is based on using a modification of the maximum mutual information (MMI) objective function, and appears to provide no improvement over standard ML estimation. The second technique is based on a heuristic correction of the Gaussian parameters, and is seen to give a 2-5% improvement over ML estimation. Lalit R. Bahl, Mukund Padmanabhan, David Nahamoo, Ponani S. Gopalakrishnan |
ICASSP | 1 |
| 1996 | Speaker clustering and transformation for speaker adaptation in large-vocabulary speech recognition systemsabstractA speaker adaptation strategy is described that is based on finding a subset of speakers, from the training set, who are acoustically close to the test speaker, and using only the data from these speakers (rather than the complete training corpus) to re-estimate the system parameters. Further, a linear transformation is computed for every one of the selected training speakers to better map the training speaker's data to the test speaker's acoustic space. Finally, the system parameters (Gaussian means) are re-estimated specifically for the test speaker using the transformed data from the selected training speakers. Experiments showed that this scheme is capable of reducing the error rate by 10-15% with the use of as little as 3 sentences of adaptation data. Mukund Padmanabhan, Lalit R. Bahl, David Nahamoo, Michael Picheny |
ICASSP | 2 |
| 1995 | Performance of the IBM large vocabulary continuous speech recognition system on the ARPA Wall Street Journal taskabstractIn this paper we discuss various experimental results using our continuous speech recognition system on the Wall Street Journal task. Experiments with different feature extraction methods, varying amounts and type of training data, and different vocabulary sizes are reported. Lalit R. Bahl, S. Balakrishnan-Aiyer, Jerome R. Bellegarda, Martin Franz, Ponani S. Gopalakrishnan, David Nahamoo, Miroslav Novak, Mukund Padmanabhan, Michael Picheny, Salim Roukos |
ICASSP | 1 |
| 1995 | Experiments using data augmentation for speaker adaptationabstractSpeaker adaptation typically involves customizing some existing (reference) models in order to account for the characteristics of a new speaker. This work considers the slightly different paradigm of customizing some reference data for the purpose of populating the new speaker's space, and then using the resulting (augmented) data to derive the customized models. The data augmentation technique is based on the metamorphic algorithm first proposed in Bellegarda et al. [1992], assuming that a relatively modest amount of data (100 sentences) is available from each new speaker. This contraint requires that reference speakers be selected with some care. The performance of this method is illustrated on a portion of the Wall Street Journal task. Jerome R. Bellegarda, Peter V. de Souza, David Nahamoo, Mukund Padmanabhan, Michael Picheny, Lalit R. Bahl |
ICASSP | 6 |
| 1995 | A tree search strategy for large-vocabulary continuous speech recognitionabstractWe describe a tree search strategy, called the Envelope Search, which is a time-asynchronous search scheme that combines aspects of the A* heuristic search algorithm with those of the time-synchronous Viterbi search algorithm. This search technique is used in the large-vocabulary continuous speech recognition system developed at the IBM Research Center. Ponani S. Gopalakrishnan, Lalit R. Bahl, Robert L. Mercer |
ICASSP | 2 |
| 1995 | Fast match based on decision tree
Claire Waast-Richard, Lalit R. Bahl, Marc El-Bèze |
EUROSPEECH | 2 |
| 1994 | Robust methods for using context-dependent features and models in a continuous speech recognizerabstractIn this paper we describe the method we use to derive acoustic features that reflect some of the dynamics of frame-based parameter vectors. Models for such observations must be context dependent. Such models were outlined in an earlier paper. Here we describe a method for using these models in a recognition system. The method is more robust than using continuous parameter models in recognition. At the same time it does not suffer from the possible information loss in vector quantization based systems.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP (1) | 1 |
| 1994 | The metamorphic algorithm: a speaker mapping approach to data augmentationabstractLarge vocabulary speaker-dependent speech recognition systems adjust to the acoustic peculiarities of each new speaker based on some enrolment data provided by this speaker. As the amount of data required increases with the sophistication of the underlying acoustic models, the enrolment may get lengthy. To streamline it, it is therefore desirable to make use of previously acquired speech data. The authors describe a data augmentation strategy based on a piecewise linear mapping between the feature space of a new speaker and that of a reference speaker. This speaker-normalizing mapping is used to transform the previously acquired data of the reference speaker onto the space of the new speaker. The performance of the resulting procedure, dubbed the metamorphic algorithm, is illustrated on an isolated utterance speech recognition task with a vocabulary of 20000 words. Results show that the metamorphic algorithm can substantially reduce the word error rate when only a limited amount of enrolment data is available. Alternatively, it leads to a level of performance comparable to that obtained when a much greater amount of enrolment data is required from the new speaker. In addition, it can also be used for tracking spectral evolution over time, thus providing a possible means for robust speaker self-adaptation.> Jerome R. Bellegarda, Peter V. de Souza, Arthur Nádas, David Nahamoo, Michael Picheny, Lalit R. Bahl |
IEEE Trans. Speech Audio Process. | 6 |
| 1993 | Context dependent vector quantization for continuous speech recognition
Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, Michael Picheny |
ICASSP (2) | 1 |
| 1993 | A supervised approach to the construction of context-sensitive acoustic prototypes
Jerome R. Bellegarda, Peter V. de Souza, David Nahamoo, Michael Picheny, Lalit R. Bahl |
ICASSP (2) | 5 |
| 1993 | Word lookahead scheme for cross-word right context models in a stack decoder
Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
EUROSPEECH | 1 |
| 1993 | Multonic Markov word models for large vocabulary continuous speech recognitionabstractA new class of hidden Markov models is proposed for the acoustic representation of words in an automatic speech recognition system. The models, built from combinations of acoustically based sub-word units called fenones, are derived automatically from one or more sample utterances of a word. Because they are more flexible than previously reported fenone-based word models, they lead to an improved capability of modeling variations in pronunciation. They are therefore particularly useful in the recognition of continuous speech. In addition, their construction is relatively simple, because it can be done using the well-known forward-backward algorithm for parameter estimation of hidden Markov models. Appropriate reestimation formulas are derived for this purpose. Experimental results obtained on a 5000-word vocabulary natural language continuous speech recognition task are presented to illustrate the enhanced power of discrimination of the new models.> Lalit R. Bahl, Jerome R. Bellegarda, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | Estimating hidden Markov model parameters so as to maximize speech recognition accuracyabstractThe problem of estimating the parameter values of hidden Markov word models for speech recognition is addressed. It is argued that maximum-likelihood estimation of the parameters via the forward-backward algorithm may not lead to values which maximize recognition accuracy. An alternative estimation procedure called corrective training, which is aimed at minimizing the number of recognition errors, is described. Corrective training is similar to a well-known error-correcting training procedure for linear classifiers and works by iteratively adjusting the parameter values so as to make correct words more probable and incorrect words less probable. There are strong parallels between corrective training and maximum mutual information estimation; the relationship of these two techniques is discussed and a comparison is made of their performance. Although it has not been proved that the corrective training algorithm converges, experimental evidence suggests that it does, and that it leads to fewer recognition errors that can be obtained with conventional training methods.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | A method for the construction of acoustic Markov models for wordsabstractA technique for constructing Markov models for the acoustic representation of words is described. Word models are constructed from models of subword units called fenones. Fenones represent very short speech events and are obtained automatically through the use of a vector quantizer. The fenonic baseform for a word-i.e., the sequence of fenones used to represent the word-is derived automatically from one or more utterances of that word. Since the word models are all composed from a small inventory of subword models, training for large-vocabulary speech recognition systems can be accomplished with a small training script. A method for combining phonetic and fenonic models is presented. Results of experiments with speaker-dependent and speaker-independent models on several isolated-word recognition tasks are reported. The results are compared with those for phonetics-based Markov models and template-based dynamic programming (DP) matching.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer, Michael Picheny |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | A fast approximate acoustic match for large vocabulary speech recognitionabstractIn a large vocabulary speech recognition system using hidden Markov models, calculating the likelihood of an acoustic signal segment for all the words in the vocabulary involves a large amount of computation. In order to run in real time on a modest amount of hardware, it is important that these detailed acoustic likelihood computations be performed only on words which have a reasonable probability of being the word that was spoken. The authors describe a scheme for rapidly obtaining an approximate acoustic match for all the words in the vocabulary in such a way as to ensure that the correct word is, with high probability, one of a small number of words examined in detail. Using fast search methods, they obtain a matching algorithm that is about a hundred times faster than doing a detailed acoustic likelihood computation on all the words in the IBM Office Correspondence isolated word dictation task, which has a vocabulary of 20000 words. Experimental results showing the effectiveness of such a fast match for a number of talkers are given.> Lalit R. Bahl, Steven V. De Gennaro, Ponani S. Gopalakrishnan, Robert L. Mercer |
IEEE Trans. Speech Audio Process. | 1 |
| 1992 | A fast match for continuous speech recognition using allophonic modelsabstractIn a large vocabulary real-time speech recognition system, there is a need for a fast method for selecting a list of candidate words from the vocabulary that match well with a given acoustic input. The authors describe a highly accurate fast acoustic match for continuous speech recognition. The algorithm uses allophonic models and efficient search techniques to select a set of candidate words. The allophonic models are derived by constructing decision trees that query the context in which each phone occurs to arrive at an allophone in a given context. The models for all the words in the vocabulary are arranged in a tree structure and efficient tree search algorithms are used to select a list of candidate words using these models. Using this method, the authors are able to obtain over 99% accuracy in the fast match for a continuous speech recognition task which has a vocabulary of 5000 words.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP | 1 |
| 1992 | Adaptation of large vocabulary recognition system parametersabstractThe authors report on a series of experiments in which the hidden Markov model baseforms and the language model probabilities were updated from spontaneously dictated speech captured during recognition sessions with the IBM Tangora system. The basic technique for baseform modification consisted of constructing new fenonic baseforms for all recognized words. To modify the language model probabilities, a simplified version of a cache language model was implemented. The word error rate across six talkers was 3.7%. Baseform adaptation reduced the average error rate to 3.5%, and using the cache language model reduced the error rate to 3.2%. Combining both techniques further reduced the error rate to 3.1%-a respectable improvement over the original error rate, especially given that the system was speaker-trained prior to adaptation.> Lalit R. Bahl, Peter V. de Souza, David Nahamoo, Michael Picheny, Salim Roukos |
ICASSP | 1 |
| 1992 | Robust speaker adaptation using a piecewise linear acoustic mappingabstractIn a large vocabulary speech recognition system, it is desirable to make use of previously acquired speech data when encountering new speakers. The authors describe an adaptation strategy based on a piecewise linear mapping between the feature space of a new speaker and that of a reference speaker. This speaker-normalizing mapping is used to transform the previously acquired parameters of the reference speaker onto the space of the new speaker. This results in a robust speaker adaptation procedure which allows for a drastic reduction in the amount of training data required from the new speaker. The performance of this method is illustrated on an isolated utterance speech recognition task with a vocabulary of 20000 words.> Jerome R. Bellegarda, Peter V. de Souza, Arthur Nádas, David Nahamoo, Michael Picheny, Lalit R. Bahl |
ICASSP | 6 |
| 1991 | A new class of fenonic Markov word models for large vocabulary continuous speech recognitionabstractA technique for constructing hidden Markov models for the acoustic representation of words is described. The models, built from combinations of acoustically based subword units called fenones, are derived automatically from one or more sample utterances of words. They are more flexible than previously reported fenone-based word models and lead to an improved capability of modeling variations in pronunciation. In addition, their construction is simplified, because it can be done using the well-known forward-backward algorithm for the parameter estimation of hidden Markov models. Experimental results obtained on a 5000-word vocabulary continuous speech recognition task are presented to illustrate some of the benefits associated with the new models. Multonic baseforms resulted in a reduction of 16% in the average error rate obtained for ten speakers.> Lalit R. Bahl, Jerome R. Bellegarda, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP | 1 |
| 1991 | Automatic phonetic baseform determinationabstractThe authors describe a series of experiments in which the phonetic baseform is deduced automatically for new words by utilizing actual utterances of the new word in conjunction with a set of automatically derived spelling-to-sound rules. Recognition performance was evaluated on new words spoken by two different speakers when the phonetic baseforms were extracted via the above approach. The error rates on these new words were found to be comparable to or better than when the phonetic baseforms were derived by hand, thus validating the basic approach.> Lalit R. Bahl, Subrata K. Das, Peter DeSouza, M. Epstein, Robert L. Mercer, Bernard Mérialdo, David Nahamoo, Michael Picheny, J. Powell |
ICASSP | 1 |
| 1991 | Decision trees for phonological rules in continuous speechabstractThe authors present an automatic method for modeling phonological variation using decision trees. For each phone they construct a decision tree that specifies the acoustic realization of the phone as a function of the context in which it appears. Several-thousand sentences from a natural language corpus spoken by several speakers are used to construct these decision trees. Experimental results on a 5000-word vocabulary natural language speech recognition task are presented.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP | 1 |
| 1991 | A fast algorithm for deleted interpolation
Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer, David Nahamoo |
EUROSPEECH | 1 |
| 1990 | Constructing groups of acoustically confusable wordsabstractMethod for constructing groups of acoustically similar words in a large vocabulary speech recognition system are studied. The study is based on the idea that given a dictionary of words one should be able to take each word and construct a reasonably short list of words from the vocabulary that sound similar to the chosen word. Groups of acoustically similar words are constructed using the hidden Markov models for the words. They are then used in a fast match scheme in an isolated utterance recognition system. The experiments show that a short list of words can be constructed during recognition in a very quick fashion using this scheme and that the list will contain the correct word with a high probability.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, Dimitri Kanevsky, David Nahamoo |
ICASSP | 1 |
| 1989 | Large vocabulary natural language continuous speech recognitionabstractA description is presented of the authors' current research on automatic speech recognition of continuously read sentences from a naturally-occurring corpus: office correspondence. The recognition system combines features from their current isolated-word recognition system and from their previously developed continuous-speech recognition system. It consists of an acoustic processor, an acoustic channel model, a language model, and a linguistic decoder. Some new features in the recognizer relative to the isolated-word speech recognition system include the use of a fast match to prune rapidly to a manageable number the candidates considered by the detailed match, multiple pronunciations of all function words, and modeling of interphone coarticulatory behavior. The authors recorded training and test data from a set of ten male talkers. The perplexity of the test sentences was found to be 93; none of sentences was part of the data used to generate the language model. Preliminary (speaker-dependent) recognition results on these talkers yielded an average word error rate of 11.0%.> Lalit R. Bahl, Raimo Bakis, Jerome R. Bellegarda, Peter F. Brown, David Burshtein, Subrata K. Das, Peter V. de Souza, Ponani S. Gopalakrishnan, Frederick Jelinek, Dimitri Kanevsky, Robert L. Mercer, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 1 |
| 1989 | Matrix fast match: a fast method for identifying a short list of candidate words for decodingabstractA rapid method is presented for identifying a short list of candidate words that match well with some acoustic input to serve as a fast matching stage in a large-vocabulary speech recognition system that uses hidden Markov models and maximum a posteriori decoding. Given hidden Markov models for all the words in the vocabulary the authors derive a class of algorithms that are faster than a detailed likelihood computation using these models by constructing an estimator of the likelihood. Using such an estimator they produce a list of candidate words that match well with the given acoustic input which has the property that it is guaranteed to contain the correct word in all the cases where a detailed likelihood computation would assign the maximum likelihood to that word.> Lalit R. Bahl, Ponani S. Gopalakrishnan, Dimitri Kanevsky, David Nahamoo |
ICASSP | 1 |
| 1989 | A fast approximate acoustic match for large vocabulary speech recognition
Lalit R. Bahl, Steven V. De Gennaro, Ponani S. Gopalakrishnan, Robert L. Mercer |
EUROSPEECH | 1 |
| 1988 | Speech recognition with continuous-parameter hidden Markov modelsabstractThe acoustic-modelling problem in automatic speech recognition is examined from an information theoretic point of view. This problem is to design a speech-recognition system which can extract from the speech waveform as much information as possible about the corresponding word sequence. The information extraction process is factored into two steps: a signal-processing step which converts a speech waveform into a sequence of informative acoustic feature vectors, and a step which models such a sequence. The authors are primarily concerned with the use of hidden Markov models to model sequences of feature vectors which lie in a continuous space. They explore the trade-off between packing information into such sequences and being able to model them accurately. The difficulty of developing accurate models of continuous-parameter sequences is addressed by investigating a method of parameter estimation which is designed to cope with inaccurate modeling assumptions.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer |
ICASSP | 1 |
| 1988 | Obtaining candidate words by polling in a large vocabulary speech recognition systemabstractConsiders the problem of rapidly obtaining a short list of candidate words for more detailed inspection in a large vocabulary, vector-quantizing speech recognition system. An approach called polling is advocated, in which each label produced by the vector quantizer casts a varying, real-valed vote for each word in the vocabulary. The words receiving the highest votes are placed on a short list to be matched in detail at a later stage of processing. Expressions are derived for these votes under the assumption that for any given word, the observed label frequencies have Poisson distributions. Although the method is more general, particular attention is paid to the implementation of polling in speech recognition systems which use hidden Markov models during the acoustic match computation. Results are presented of experiments with speaker-dependent and speaker-independent Markov models on two different isolated word recognition tasks.> Lalit R. Bahl, Raimo Bakis, Peter V. de Souza, Robert L. Mercer |
ICASSP | 1 |
| 1988 | A new algorithm for the estimation of hidden Markov model parametersabstractDiscusses the problem of estimating the parameter values of hidden Markov word models for speech recognition. The authors argue that maximum-likelihood estimation of the parameters does not lead to values which maximize recognition accuracy and describe an alternative estimation procedure called corrective training which is aimed at minimizing the number of recognition errors. Corrective training is similar to a well-known error-correcting training procedure for linear classifiers and works by iteratively adjusting the parameter values so as to make correct words more probable and incorrect words less probable. There are also strong parallels between corrective training and maximum mutual information estimation. They do not prove that the corrective training algorithm converges, but experimental evidence suggests that it does, and that it leads to significantly fewer recognition errors than maximum likelihood estimation.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer |
ICASSP | 1 |
| 1988 | Acoustic Markov models used in the Tangora speech recognition systemabstractThe Speech Recognition Group at IBM Research has developed a real-time, isolated-word speech recognizer called Tangora, which accepts natural English sentences drawn from a vocabulary of 20000 words. Despite its large vocabulary, the Tangora recognizer requires only about 20 minutes of speech from each new user for training purposes. The accuracy of the system and its ease of training are largely attributable to the use of hidden Markov models in its acoustic match component. An automatic technique for constructing Markov word models is described and results are included of experiments with speaker-dependent and speaker-independent models on several isolated-word recognition tasks.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer, Michael Picheny |
ICASSP | 1 |
| 1987 | Experiments with the Tangora 20, 000 word speech recognizerabstractThe Speech Recognition Group at IBM Research in Yorktown Heights has developed a real-time, isolated-utterance speech recognizer for natural language based on the IBM Personal Computer AT and IBM Signal Processors. The system has recently been enhanced by expanding the vocabulary from 5,000 words to 20,000 words and by the addition of a speech workstation to support usability studies on document creation by voice. The system supports spelling and interactive personalization to augment the vocabularies. This paper describes the implementation, user interface, and comparative performance of the recognizer. Amir Averbuch, Lalit R. Bahl, Raimo Bakis, Peter F. Brown, Gregg Daggett, Ken Davies, Steven V. De Gennaro, Peter V. de Souza, E. Epstein, D. Fraleigh, Frederick Jelinek, B. Lewis, Robert L. Mercer, J. Moorhead, Arthur Nádas, David Nahamoo, Michael Picheny, G. Shichman, P. Spinelli, Dirk Van Compernolle, H. Wilkens |
ICASSP | 2 |
| 1986 | An IBM PC based large-vocabulary isolated-utterance speech recognizerabstractThe Speech Recognition Group at IBM Research in Yorktown Heights has designed a real-time, isolated-utterance speech recognizer for natural language with a 5,000-word vocabulary based on the IBM Personal Computer (PC) AT model and two IBM Signal Processors realized in VLSI technology. The enrollment period for a new user is approximately 20 minutes. The basic vocabulary is chosen from the most common words in several collections of documents such as office memoranda and business letters. The system supports spelling and interactive personalization to augment this vocabulary. Signal processing, vector quantization, and acoustic matching algorithms are programmed on the IBM Signal Processors which fit into the PC AT chassis. The PC AT controls the Processors and implements the decoder stack search and the language model, as well as the application-specific interface. The modular architecture of the design is expandable to a 20,000-word vocabulary system by the addition of two more IBM Signal Processors housed in a PC Expansion Unit. Amir Averbuch, Lalit R. Bahl, Raimo Bakis, Peter F. Brown, A. G. Cole, Gregg Daggett, Subrata K. Das, Ken Davies, S. DeGennaro, Peter V. de Souza, E. Epstein, D. Fraleigh, Frederick Jelinek, Slava M. Katz, B. Lewis, Robert L. Mercer, Arthur Nádas, David Nahamoo, Michael Picheny, G. Shichman, P. Spinelli |
ICASSP | 2 |
| 1986 | Maximum mutual information estimation of hidden Markov model parameters for speech recognitionabstractA method for estimating the parameters of hidden Markov models of speech is described. Parameter values are chosen to maximize the mutual information between an acoustic observation sequence and the corresponding word sequence. Recognition results are presented comparing this method with maximum likelihood estimation. Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer |
ICASSP | 1 |
| 1984 | Some experiments with large-vocabulary isolated-word sentence recognitionabstractThis paper deals with two experiments with a large vocabulary isolated word recognizer. The first compares word error rates for 1) meaningful sentences belonging to actual documents and 2) random word lists from the same vocabulary. The error rate is considerably lower for random word lists. The second experiment investigates the performance of the recognition system on sentences containing words outside the vocabulary of the recognizer. Sentences from a 5000 word vocabulary task are recognized with a recognizer limited to a 2000 word subvocabulary. The error rate is only slightly higher than it would be if recognition of the full 5000 word vocabulary was allowed. Lalit R. Bahl, Subrata K. Das, Peter V. de Souza, Frederick Jelinek, Slava M. Katz, Robert L. Mercer, Michael Picheny |
ICASSP | 1 |
| 1983 | Recognition of isolated-word sentences from a 5000-word vocabulary office correspondence taskabstractRecognition results on sentences from a 5000-word vocabulary drawn from office correspondence are presented. The sentences were read with pauses between the words. The vocabulary comprises the 5000 most frequently occurring words in a data-base of 14,000 office memoranda and letters, and has a perplexity of 90, measured from a trigram language model. Experiments were carried out with 6 speakers (4 male, 2 female) in an office environment using a close-talking microphone. The recognition system was automatically trained to each speaker by having the speaker read 100 typical sentences from the office correspondence data-base. Recognition was carried out for each speaker on 20 test sentences, consisting of 299 words. The recognition rate (% words correct) averaged across the 6 speakers was 94.5%. Lalit R. Bahl, A. G. Cole, Frederick Jelinek, Robert L. Mercer, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 1 |
| 1983 | A Maximum Likelihood Approach to Continuous Speech RecognitionabstractSpeech recognition is formulated as a problem of maximum likelihood decoding. This formulation requires statistical models of the speech production process. In this paper, we describe a number of statistical models for use in speech recognition. We give special attention to determining the parameters for such models from sparse data. We also describe two decoding methods, one appropriate for constrained artificial languages and one appropriate for more realistic decoding tasks. To illustrate the usefulness of the methods described, we review a number of decoding results that have been obtained with them. Lalit R. Bahl, Frederick Jelinek, Robert L. Mercer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1981 | Continuous parameter acoustic processing for recognition of a natural speech corpusabstractThe stack linguistic decoder described in references (1,3,6) has been modified to operate directly on continuous parameter vectors produced by an acoustic processor, thereby bypassing entirely labeling and segmentation of the speech signal. A recognition experiment has been carried out on a natural speech corpus using five-dimensional vectors composed of the four Dutch vowel factors of Klein, Plomp, and Pols, plus an indicator of short-term changes in overall power. The results are compared to those from an earlier experiment in which a centisecond labeling acoustic processor was used. Lalit R. Bahl, Raimo Bakis, Paul S. Cohen, A. G. Cole, Frederick Jelinek, Burn L. Lewis, Robert L. Mercer |
ICASSP | 1 |
| 1981 | Speech recognition of a natural text read as isolated wordsabstractThis paper describes the results of an experiment in which we have applied our speech recognition system to the problem of recognizing sentences from our restricted laser-patent corpus when the sentences are read with pauses between the words. Except for changes to the phonology and the training data, nothing has been done to adapt the system to isolated word recognition. On 20 sentences a word error rate of 3.1% was obtained. This compares with 8.7% for the same sentences when spoken continuously by the same talker. Lalit R. Bahl, Raimo Bakis, Paul S. Cohen, A. G. Cole, Frederick Jelinek, Burn L. Lewis, Robert L. Mercer |
ICASSP | 1 |
| 1981 | Continuous speech recognition with automatically selected acoustic prototypes obtained by either bootstrapping or clusteringabstractAutomatic selection of acoustic prototypes is an important step towards making speech recognition systems automatically adaptable to new speakers. Two methods for automatically obtaining a set of acoustic prototypes for use by a centisecond labeling acoustic processor are described. One method is based on bootstrapping, the other on clustering. Recognition results using these automatically obtained prototypes on the 1000-word vocabulary natural language Laser Patent task are presented. These results are compared to those from an experiment in which the acoustic prototypes were manually selected. Arthur Nádas, Robert L. Mercer, Lalit R. Bahl, Raimo Bakis, Paul S. Cohen, A. G. Cole, Frederick Jelinek, Burn L. Lewis |
ICASSP | 3 |
| 1980 | Further results on the recognition of a continuously read natural corpusabstractFurther results have been obtained on the recognition of continuously read sentences from a natural language corpus of laser patents. The vocabulary is limited to the 1000 most frequently occurring words in the corpus. Our model of the task language has a perplexity of 24.1 words (corresponding to an entropy of 4.6 bits/word). This paper describes modifications and improvements to the system which have resulted in the lowering of the word error rate from the previously reported 33.1% to 8.9%. Lalit R. Bahl, Raimo Bakis, Paul S. Cohen, A. G. Cole, Frederick Jelinek, Burn L. Lewis, Robert L. Mercer |
ICASSP | 1 |
| 1979 | Recognition results for several experimental acoustic processorsabstractThe statistical training and decoding procedures developed at IBM Research can be used with a wide variety of acoustic processors. We have recently (July and August 1978) achieved error-free or nearly error-free decoding results with several different acoustic processors on sentences from the New Raleigh Language (vocabulary 250 words, perplexity 7.27 words). One of these processors, which achieved a 0% error rate on New Raleigh sentences, has been used to decode sentences from the New Raleigh Language without the benefit of syntactic guidance during the decoding process. On this much more difficult task, it has achieved an error rate of 8.8% at the word level, corresponding to a sentence error rate of 53%. All of these processors are non-segmenting processors which produce output once every 10ms. Lalit R. Bahl, Raimo Bakis, Paul S. Cohen, A. G. Cole, Frederick Jelinek, Burn L. Lewis, Robert L. Mercer |
ICASSP | 1 |
| 1978 | Automatic recognition of continuously spoken sentences from a finite state grammerabstractWe report performance results on the recognition of continuously spoken sentences from the finite state grammar for the "New Raleigh Language" (vocabulary-250 words; average sentence length-8 words; entropy-2.86 bits/word; perplexity-7.27 words). Sentence and word error rates of 5% and 0.6% , respectively, are achieved, using a new centisecond-level model for the acoustic processor. We also report results for the "CMU-AIX05 Language" (vocabulary-1011 words; average sentence length-about 7 words; entropy-2.18 bits/word; perplexity-4.53 words), using both our earlier phone-level model and the centisecond-level model. With the phone-level acoustic-processor model, sentence and word error rates of 2% and 0.8%, respectively, are achieved. With the centisecond-level model, sentence and word error rates are 1% and 0.1%, respectively. Lalit R. Bahl, James K. Baker, Paul S. Cohen, A. G. Cole, Frederick Jelinek, Burn L. Lewis, Robert L. Mercer |
ICASSP | 1 |
| 1978 | Recognition of continuously read natural corpusabstractPreliminary results have been obtained with a system for recognizing continuously read sentences from a naturally-occurring corpus (Laser Patents), restricted to a 1000-word vocabulary. Our model of the task language has an entropy of about 4.8 bits/word and a perplexity of 21.11 words. Many new problems arise in recognition of a substantial natural corpus (compared to recognition of an artificially constrained language). Some techniques are described for treating these problems. On a test set consisting of 20 sentences having a total of 486 words, there was a word error rate of 33.1%. Lalit R. Bahl, James K. Baker, Paul S. Cohen, Frederick Jelinek, Burn L. Lewis, Robert L. Mercer |
ICASSP | 1 |
| 1978 | Optimization of Teleprocessing Networks with Concentrators and Multiconnected TerminalsabstractIn this paper we consider the optimization of teleprocessing (TP) networks with concentrators. Each terminal may in general have multiple connections to several concentrators. The first part of the paper is concerned with the optimal assignment of terminals to a given set of concentrators. A sequential assignment algorithm is presented with its efficiency and optimality established. The second part of the paper is concerned with the optimal number and location of the concentrators to be installed. A generalized "ADD" algorithm is suggested. It is shown that the total cost function is convex with respect to the nwnber of concentrators installed in such an algorithm. Donald T. Tang, Lin S. Woo, Lalit R. Bahl |
IEEE Trans. Computers | 3 |
| 1976 | Preliminary results on the performance of a system for the automatic recognition of continuous speechabstractThis report presents results obtained in some experiments on the computer recognition of continuous speech. The experiments deal with two simple languages having vocabularies of 11 and 250 words. Lalit R. Bahl, James K. Baker, Paul S. Cohen, N. Rex Dixon, Frederick Jelinek, Robert L. Mercer, Harvey F. Silverman |
ICASSP | 1 |
| 1975 | Decoding for channels with insertions, deletions, and substitutions with applications to speech recognitionabstractA model for channels in which an input sequence can produce output sequences of varying length is described. An efficient computational procedure for calculating Pr\{Y\mid X\}is devised, whereX = x_1,x_2,\cdots,x_MandY = y_1,y_2,\cdots,y_Nare the input and output of the channel. A stack decoding algorithm for decoding on such channels is presented. The appropriate likelihood function is derived. Channels with memory are considered. Some applications to speech and character recognition are discussed. Lalit R. Bahl, Frederick Jelinek |
IEEE Trans. Inf. Theory | 1 |
| 1975 | Design of a linguistic statistical decoder for the recognition of continuous speechabstractMost current attempts at automatic speech recognition are formulated in an artificial intelligence framework. In this paper we approach the problem from an information-theoretic point of view. We describe the overall structure of a linguistic statistical decoder (LSD) for the recognition of continuous speech. The input to the decoder is a string of phonetic symbols estimated by an acoustic processor (AP). For each phonetic string, the decoder finds the most likely input sentence. The decoder consists of four major subparts: 1) a statistical model of the language being recognized; 2) a phonemic dictionary and statistical phonological rules characterizing the speaker; 3) a phonetic matching algorithm that computes the similarity between phonetic strings, using the performance characteristics of the AP; 4) a word level search control. The details of each of the subparts and their interaction during the decoding process are discussed. Frederick Jelinek, Lalit R. Bahl, Robert L. Mercer |
IEEE Trans. Inf. Theory | 2 |
| 1974 | Optimal decoding of linear codes for minimizing symbol error rate (Corresp.)abstractThe general problem of estimating the a posteriori probabilities of the states and transitions of a Markov source observed through a discrete memoryless channel is considered. The decoding of linear block and convolutional codes to minimize symbol error probability is shown to be a special case of this problem. An optimal decoding algorithm is derived. Lalit R. Bahl, John Cocke, Frederick Jelinek, Josef Raviv |
IEEE Trans. Inf. Theory | 1 |
| 1972 | An efficient algorithm for computing free distance (Corresp.)abstractAn efficient bidirectional search algorithm for computing the free distance of convolutional codes is described. Lalit R. Bahl, C. Cullum, W. Donald Frazer, Frederick Jelinek |
IEEE Trans. Inf. Theory | 1 |
| 1972 | On the structure of rate 1/n convolutional codesabstractWe show what choice there is in assigning output digits to transitions of a binary rate1/ncode trellis so that the latter will correspond to a convolutional code. We then prove that in any rate\frac{1}{2}noncatastrophic code of constraint length\upsiloneach binary sequence of length2j(1 \leq j \leq \upsilon - 1)is associated with exactly2^{\upsilon -j -1}distinct pathsjbranches long. As a consequence of the above properties nondegenerate codes with branch complementarity are fully determined by the topological relationship of the trellis transitions associated with output pairs 00. Finally, we derive a new upper bound on free distance of rate1/nconvolutional codes and use our results to determine the length of the largest input sequence that can conceivably result in an output whose weight is Lalit R. Bahl, Frederick Jelinek |
IEEE Trans. Inf. Theory | 1 |
| 1971 | Single- and multiple-burst-correcting properties of a class of cyclic product codesabstractThe direct product ofpsingle parity-check codes of block lengthsn_1,n_2, \cdots ,n_pis a cyclic code of block lengthn_1 \times n_2 \times \cdots \times n_pwith(n_1 - 1) \times (n_2 - 1) \times \cdots \times (n_p - 1)information symbols per block, if the integersn_1,n_2 \cdots ,n_pare relatively prime in pairs. A lower bound for the single-burst-correction (SBC) capability of these codes is obtained. Then, a detailed analysis is made forp = 3, and it is shown that the codes can correct one long burst or two short bursts of errors. A lower bound for the double-burst-correction (DBC) capability is derived, and a simple decoding algorithm is obtained. The generalization to correcting an arbitrary number of bursts is discussed. Lalit R. Bahl, Robert T. Chien |
IEEE Trans. Inf. Theory | 1 |
| 1971 | Rate 1/2 convolutional codes with complementary generatorsabstractThe structure of a class of rate\frac{1}{2}convolutional codes called complementary codes is investigated. Of special interest are properties that permit a simplified evaluation of free distance. Methods for finding the codes with largest free distance in this class are obtained. A synthesis procedure and a search procedure that result in good codes up to constraint lengths of 24 are described. The free and minimum distances of the best complementary codes are compared with the best known bounds and with the distances of other known codes. Lalit R. Bahl, Frederick Jelinek |
IEEE Trans. Inf. Theory | 1 |
| 1970 | Block Codes for a Class of Constrained Noiseless Channels
Donald T. Tang, Lalit R. Bahl |
Inf. Control. | 2 |
| 1969 | Multiple-Burst-Error Correction by Threshold Decoding
Lalit R. Bahl, Robert T. Chien |
Inf. Control. | 1 |
| 1969 | On Gilbert burst-error-correcting codes (Corresp.)
Lalit R. Bahl, Robert T. Chien |
IEEE Trans. Inf. Theory | 1 |
| 1969 | Correction of two erasure bursts (Corresp.)
Robert T. Chien, Lalit R. Bahl, Donald T. Tang |
IEEE Trans. Inf. Theory | 2 |