Stephen E. Levinson

dblp:l/SELevinson · DBLP profile ↗
← Back
52ranked-venue papers
13as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 11 first-authorArtificial intelligence and machine learning · 18 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Face, body and person analysis · 26% Representation and self-supervised learning · 22% Speech recognition and synthesis · 12%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%
Computer graphics and multimedia
2 papers
Audio and music processing · 100%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
affect recognition
0.112007
Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007
Computer vision › Face, body and person analysis › affect recognition
audiovisual emotion recognition
0.112007
Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007
Human-AI interaction › affective computing
affective human-computer interaction
0.112007
Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.112006
Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
nonlinear manifold learning
0.112006
Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006
Machine learning › Time series and sequential data › time series analysis
time series learning
0.112006
Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006
Human-AI interaction
affective computing
0.112005
Audio-Visual Affect Recognition through Multi-Stream Fused HMM for HCI · CVPR (2) 2005
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition
0.012002
A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002
Audio and music processing › sound source localization
binaural localization
0.012002
A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002
Audio and music processing
sound source localization
0.012002
A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002
Computer vision › 3D vision
camera pose estimation
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001
Computer vision › 3D vision › multi-view geometry
epipolar geometry estimation
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001
Computer vision › Video understanding and tracking
feature tracking
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001
Computer vision › Video understanding and tracking
object tracking
0.012001
Tracking of Object with SVM Regression · CVPR (2) 2001
Natural language and speech › Speech recognition and synthesis › speech analysis
prosody analysis
0.012007
Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007
Natural language and speech › Language models and text generation › language acquisition
spoken language acquisition
0.011994
An experiment in spoken language acquisition · IEEE Trans. Speech Audio Process. 1994
Wireless sensing and localization › ranging
time difference of arrival
0.012002
A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002
High-performance computing
scientific computing systems
0.011993
Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial Intelligence · IEEE Trans. Knowl. Data Eng. 1993
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.021985
Structural methods in automatic speech recognition · Proc. IEEE 1985
Isolated and Connected Word Recognition-Theory and Selected Applications · IEEE Trans. Commun. 1981
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.021994
An experiment in spoken language acquisition · IEEE Trans. Speech Audio Process. 1994
The Vocal Speech Understanding System · IJCAI 1975
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
0.011986
Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.011986
Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986
Audio and music processing › speech processing
speech modeling
0.011986
Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986
Audio and music processing
speech processing
0.011986
Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.011985
Structural methods in automatic speech recognition · Proc. IEEE 1985
Computer vision › Image recognition and object detection
template matching
0.011985
Structural methods in automatic speech recognition · Proc. IEEE 1985
Audio and music processing
speech signal
0.011986
Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986

Methods — techniques the papers use, named apart from their topics

voting · 0.1feature selection · 0.1phase unwrapping · 0.1cross power spectrum · 0.1TDOA · 0.1multistream HMM · 0.1multi-stream HMM · 0.1maximum mutual information · 0.1maximum entropy principle · 0.1hidden markov model · 0.1outlier detection · 0.0affine transform estimation · 0.0SVM regression · 0.0connectionist network · 0.0heuristic search · 0.0connectionist systems · 0.0gaussian mixture · 0.0expectation-maximization · 0.0
YearPublicationVenuePosition
2010 Nonlinear Dynamical Multi-Scale Model of Associative Memory
abstract
How can we get such reliable behavior from the mind when the brain is made up of such unreliable elements as neurons? We propose that the answer is related to the emergence of stable brain states and we offer a model that illustrates how such states could arise. We discuss a new ab initio nonlinear dynamical multi-scale model that will serve as the foundation for an associative memory. Scale 0 consists of spiking Hodgkin-Huxley (HH) neurons. Scale 1 consists of components that are made up of large populations of HH neurons whose topological structure evolves according to a Hebbian-plasticity rule based on synchronous firing. The component's state is captured by the variance of phase synchrony for the population. Many such components are sparsely connected to form a large network, whose state can be captured by the n-tuple consisting of the individual states of each member component. Scale 2 takes the state of the overall network and upon examining the particular interrelationships of each component (determining how the state of one component affects the state of others) is able to generate a class of trajectories that is multistationary and stable periodic. Such a class we consider a memory, the encoding of many such memories leads to the creation of a robust associative memory. The details of the different scales are examined.
Alexander M. Duda, Stephen E. Levinson
ICMLA2
2007 HMM-Based Concept Learning for a Mobile Robot
abstract
We are developing an intelligent robot and attempting to teach it language. While there are many aspects of this research, for the purposes here the most important are the following ideas. Language is primarily based on semantics, not syntax, which is still the focus in speech recognition research these days. To truly learn meaning, a language engine cannot simply be a computer program running on a desktop computer analyzing speech. It must be part of a more general, embodied intelligent system, one capable of using associative learning to form concepts from the perception of experiences in the world, and further capable of manipulating those concepts symbolically. In this paper, we present a general cascade model for learning concepts, and explore the use of hidden Markov models (HMMs) as part of the cascade model. HMMs are capable of automatically learning and extracting the underlying structure of continuous-valued inputs and representing that structure in the states of the model. These states can then be treated as symbolic representations of the inputs. We show how a cascade of HMMs can be embedded in a small mobile robot and used to find correlations among sensory inputs to learn a set of symbolic concepts, which are used for decision making and could eventually be manipulated linguistically
Kevin M. Squire, Stephen E. Levinson
IEEE Trans. Evol. Comput.2
2007 Audio-Visual Affect Recognition
abstract
The ability of a computer to detect and appropriately respond to changes in a user's affective state has significant implications to human-computer interaction (HCI). In this paper, we present our efforts toward audio-visual affect recognition on 11 affective states customized for HCI application (four cognitive/motivational and seven basic affective states) of 20 nonactor subjects. A smoothing method is proposed to reduce the detrimental influence of speech on facial expression recognition. The feature selection analysis shows that subjects are prone to use brow movement in face, pitch and energy in prosody to express their affects while speaking. For person-dependent recognition, we apply the voting method to combine the frame-based classification results from both audio and visual channels. The result shows 7.5% improvement over the best unimodal performance. For person-independent test, we apply multistream HMM to combine the information from multiple component streams. This test shows 6.1% improvement over the best component performance
Zhihong Zeng, Jilin Tu, Ming Liu 0009, Thomas S. Huang, Brian Pianfetti, Dan Roth 0001, Stephen E. Levinson
IEEE Trans. Multim.7
2006 Learning Nonlinear Manifolds from Time Series
Ruei-Sung Lin, Che-Bin Liu, Ming-Hsuan Yang 0001, Narendra Ahuja, Stephen E. Levinson
ECCV (2)5
2006 Extraction of pragmatic and semantic salience from spontaneous spoken English
Tong Zhang 0005, Mark Hasegawa-Johnson, Stephen E. Levinson
Speech Commun.3
2006 Cognitive state classification in a spoken tutorial dialogue system
Tong Zhang 0005, Mark Hasegawa-Johnson, Stephen E. Levinson
Speech Commun.3
2005 Audio-Visual Affect Recognition through Multi-Stream Fused HMM for HCI
abstract
Advances in computer processing power and emerging algorithms are allowing new ways of envisioning human computer interaction. This paper focuses on the development of a computing algorithm that uses audio and visual sensors to detect and track a user's affective state to aid computer decision making. Using our multi-stream fused hidden Markov model (MFHMM), we analyzed coupled audio and visual streams to detect 11 cognitive/emotive states. The MFHMM allows the building of an optimal connection among multiple streams according to the maximum entropy principle and the maximum mutual information criterion. Person-independent experimental results from 20 subjects in 660 sequences show that the MFHMM approach performs with an accuracy of 80.61% which outperforms face-only HMM, pitch-only HMM, energy-only HMM, and independent HMM fusion.
Zhihong Zeng, Jilin Tu, Brian Pianfetti, Ming Liu 0009, Tong Zhang 0005, ZhenQiu Zhang, Thomas S. Huang, Stephen E. Levinson
CVPR (2)8
2004 Adaptive Discriminative Generative Model for Object Tracking
Ruei-Sung Lin, Ming-Hsuan Yang 0001, Stephen E. Levinson
ECAI3
2004 Bimodal HCI-related affect recognition
abstract
Perhaps the most fundamental application of affective computing will be Human-Computer Interaction (HCI) in which the computer should have the ability to detect and track the user's affective states, and make corresponding feedback. The human multi-sensor affect system defines the expectation of multimodal affect analyzer. In this paper, we present our efforts toward audio-visual HCI-related affect recognition. With HCI applications in mind, we take into account some special affective states which indicate users' cognitive/motivational states. Facing the fact that a facial expression is influenced by both an affective state and speech content, we apply a smoothing method to extract the information of the affective state from facial features. In our fusion stage, a voting method is applied to combine audio and visual modalities so that the final affect recognition accuracy is greatly improved. We test our bimodal affect recognition approach on 38 subjects with 11 HCI-related affect states. The extensive experimental results show that the average person-dependent affect recognition accuracy is almost 90% for our bimodal fusion.
Zhihong Zeng, Jilin Tu, Ming Liu 0009, Tong Zhang 0005, Nick Rizzolo, ZhenQiu Zhang, Thomas S. Huang, Dan Roth 0001, Stephen E. Levinson
ICMI9
2004 Automatic detection of contrast for speech understanding
abstract
Contrast is a very popular phenomenon in spoken language, and carries very important information to help understanding contents and structures of spoken language. In this paper, we propose an idea of automatic contrast detection as an effort for better speech understanding. We study the automatic tagging of three specific types of contrast: symmetric contrast, contrastive focus, and contrastive topic. We label the three types of contrasted words as contrast (C), and other words as noncontrast (¬C). The classification of contrast events is based on prosodic, spectral, and part-of-speech (POS) information sources. The integration of different knowledge sources is realized by a time-delay recursive neural network (TDRNN). The approach we proposed was testified on 235 spontaneous utterances consisting of 3500 words (samples). The contrast detection was speaker independent. The tests yielded an average of 87.9% classification rate.
Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang 0005
INTERSPEECH2
2004 Children's emotion recognition in an intelligent tutoring scenario
abstract
This paper presents an approach to automatically recognize emotion which children exhibit in an intelligent tutoring system. Emotion recognition can assist the computer agent to adapt its tutorial strategies to improve the efficiency of knowledge transmission. In this study, we detect three emotional classes: confidence, puzzle, and hesitation. Emotion is detected by means of lexical, prosodic, spectral, and syntactic analyses of users’ speech. An automatic speech recognition system serves as the fundamental constituent of the system. A robust classification and regression tree (CART) integrates the various information sources together for final decision. The effectiveness of the proposed approach has been tested on data collected by Wizard-of-Oz (WoZ) experiments. Our emotion recognition was speaker-independent, and yielded 91.3% accuracy. The test results showed that the spectral and duration-related prosodic features played very important roles in emotion recognition.
Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang 0005
INTERSPEECH2
2004 Semantic analysis for a speech user interface in an intelligent tutoring system
abstract
In this paper, we describe the strategy of semantic analysis for a speech user interface that is designed for a multimodal intelligent tutoring system. The semantic analysis involves three phases: semantic parsing, salient words/phrases spotting, and accented word detection. Semantic parsing attempts to represent the recognized sentence with a well-formed semantic frame. The recognized sentence consists of the a posterior most probably hypothesized words given the acoustic evidence, and is compliant with the grammatical knowledge that is represented by a semantic language model. The salient words/phrases are useful when semantic parsing fails. The accented words are useful when the user response is out of our expectations, and assist to make the computer agent smarter and smarter.
Yuexi Ren, Mark Hasegawa-Johnson, Stephen E. Levinson
IUI3
2003 A Bayes-rule based hierarchical system for binaural sound source localization
abstract
A Bayes-rule based hierarchical binaural sound source localization system is proposed. By combining three localization cues: interaural time differences (ITDs), interaural intensity differences (IIDs), and spectral cues, and a hierarchical decision making structure, this system enables a sound source to be located in a 3D space by using only the binaural inputs. Preliminary simulations have shown the effectiveness of this system. It can be used in studying binaural localization mechanism and applications such as in hearing aids and robotics.
Danfeng Li, Stephen E. Levinson
ICASSP (5)2
2003 Automatic language acquisition by an autonomous robot
abstract
There is no such thing as a disembodied mind. We posit that cognitive development can only occur through interaction with the physical world. To this end, we are developing a robotic platform for the purpose of studying cognition. We suggest that the central component of cognition is a memory which is primarily associative, one where learning occurs as the correlation of events from diverse inputs. We also posit that human-like cognition requires a well-integrated sensory-motor system, to provide these diverse inputs. As implemented in our robot, this system includes binaural hearing, stereo vision, tactile sense, and basic proprioceptive control. On top of these abilities, we are implementing and studying various models of processing, learning and decision making. Our goal is to produce a robot that will learn to carry out simple tasks in response to natural language requests. The robot's understanding of language will be learned concurrently with its other cognitive abilities. We have already developed a robust system and conducted a number or experiments on the way to this goal, some details of which appear in this paper. This is a first progress report of what we believe will be a long term project with significant implications.
Stephen E. Levinson, Weiyu Zhu, Danfeng Li, Kevin Squire, Ruei-Sung Lin, Matthew Kleffner, Matthew McClain, Johnny Lee
IJCNN1
2002 Articulatory speech synthesis based upon fluid dynamic principles
abstract
In this paper, an articulatory speech synthesizer based on fluid dynamic principles is presented. The key idea is to devise a refined speech production model based on the most fundamental physics of the human vocal apparatus. Our articulatory synthesizer essentially contains two parts: a vocal fold model which represents the excitation source and a vocal tract model, which describes the positions of articulators. First, we propose a combined minimum error and minimum jerk criterion to estimated the moving vocal tract shapes during speech production. Second, we propose a nonlinear mechanical model to generate the vocal fold excitation signals. Finally, a computational fluid dynamics (CFD) approach is used to solve the Reynolds-averaged Navier-Stokes (RANS) equations, which are the governing equations of speech production inside vocal apparatus. Experimental results show that our system can synthesize intelligible continuous speech sentences while naturally handling the co-articulation effects during speech production.
Stephen E. Levinson, Donald Davis, Scott Slimon
ICASSP2
2002 A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot
abstract
A robust linear phase unwrapping method is proposed to solve the 2/spl pi/ discontinuities in the phase of the cross power spectrum from the binaural inputs using two omnidirectional microphones. The relative incident angle of the interested sound is then estimated according to the time difference of arrival (TDOA) which is obtained from the unwrapped phase of the cross power spectrum. The frequency components associated with the high power are clustered into groups by the phase and frequency distance, and the dominant group is then used to obtain the initial slope estimation. The phase is unwrapped by checking the difference between the actual and the predicted phase by the estimated slope. The re-estimation is then performed by the unwrapped phase. The algorithm is tested under different incident angles and signal to noise ratio (SNR) using real speech signal and white Gaussian noise. The simulation results show the high accuracy and the robustness. This method is also Implemented to control a robot to adaptively adjust itself to the position facing the sound source directly. The satisfactory result was achieved in an open house demonstration.
Danfeng Li, Stephen E. Levinson
ICRA2
2001 Tracking of Object with SVM Regression
abstract
This paper presents a novel feature-matching based approach for rigid object tracking. The proposed method models the tracking problem as discovering the affine transforms of object images between frames according to the extracted feature correspondences. False feature matches (outliers) are automatically detected and removed with a new SVM regression technique, where outliers are iteratively identified as support vectors with the gradually decreased insensitive margin /spl epsi/. This method, in addition to object tracking, can also be used for general feature-based epipolar constraint estimation, in which it can quickly detect outliers even if they make up, in theory, over 50% of the whole data. We have applied the proposed method to track real objects under cluttering backgrounds with very encouraging results.
Weiyu Zhu, Ruei-Sung Lin, Stephen E. Levinson
CVPR (2)4
2001 Spoken Language Acquisition Via Human-Robot Interaction
abstract
This paper presents a subproject of a challenging project that explores teaching a computer human-intelligence. In the subproject, a multisensory mobile robot is used as the interface for human-computer interaction, and spoken language is taught to the computer through natural human-robot interaction. Different from state-of-the-art speech recognizers, our approach associates speech patterns directly with sensory inputs of the robot. This approach allows our system to learn multilingual speech patterns online. Further investigation of this project will include human-computer interaction that involves more modalities, and applications that use the proposed idea to train home appliances. 1.
Qiong Liu 0003, Thomas S. Huang, Ying Wu 0001, Stephen E. Levinson
ICME4
2000 Edge Orientation-Based Multi-View Object Recognition
abstract
An edge orientation-based algorithm for multi-view object recognition is presented. The distribution of edge point orientations, combined with the normalized second moments, is taken as a feature vector to describe and index each object instance. For each unknown test object, a set of likelihood weights for all the possible candidate objects is obtained by computing the Euclidean distances between the unknown feature set and all the available template feature vectors. A convincing coefficient is introduced to evaluate the confidence of the best match. New views (photo shots) will be automatically taken if the best match is thought to be insufficiently convincing. In experiments, our algorithm has achieved an average of 91.5% correct recognition rate under the 5-view scheme for 320 testing images taken from eight natural objects.
Weiyu Zhu, Stephen E. Levinson
ICPR2
2000 Signal approximation in Hilbert space and its application on articulatory speech synthesis
Stephen E. Levinson, Mark Hasegawa-Johnson
INTERSPEECH2
2000 Word concept model: a knowledge representation for dialogue agents
abstract
Information extraction is a key component in dialogue systems. Knowledge about the world as well as knowledge specific to each word should be used for robust semantic processing. An intelligent agent is necessary for a dialogue system when meanings are strictly defined by using a world state model. A layered concept structure is proposed to represent knowledge associated with each word in a "speech-friendly" way. By considering knowledge stored in the word concept model as well as knowledge base of the world model, meaning of a given sentence can be correctly identified. This paper describes the layered concept structure and how knowledge about words can be stored in this concept model.
Tong Zhang 0005, Stephen E. Levinson
INTERSPEECH3
1999 Video Sequence Learning and Recognition Via Dynamic Som
abstract
Information contained in video sequences is crucial for an autonomous robot or a computer to learn and respond to its surrounding environment. In the past, robot vision mainly concentrated on still image processing and small "image cube" processing. Continuous video sequence learning and recognition is rarely addressed in the literature due to its high requirement of dynamic processing. In this paper, we propose a novel neural network structure called dynamic self-organizing map (DSOM) for video sequence processing. The proposed technique has been tested on simulation data sets, and the results validate its learning/recognition ability.
Qiong Liu 0003, Yong Rui, Thomas S. Huang, Stephen E. Levinson
ICIP (4)4
1999 Temporal sequence learning and recognition with dynamic SOM
abstract
The purpose of the paper is to propose a map-like artificial neural network for temporal sequence pattern clustering. The map construction in our presentation is related to the self-organizing map (SOM) idea. The SOM idea was originally designed for static pattern learning and recognition. It has been found efficient for organizing high dimensional data sets. One of the biggest limitations of the traditional SOM technique is caused by its static characteristics. We propose a new neural network construction model and its corresponding training algorithm based on traditional SOM training technology and backpropagation training technology. It overcomes the static limitation of traditional SOM and tries to reach a new stage for dynamic pattern clustering, and recognition. At the end of the paper, we give some experimental results for testing this proposed method on real speech data.
Qiong Liu 0003, Sylvian R. Ray, Stephen E. Levinson, Thomas S. Huang
IJCNN3
1999 A DCT-based fast enhancement technique for robust speech recognition in automobile usage
Yunxin Zhao, Stephen E. Levinson
EUROSPEECH3
1995 Numerical simulations of fluid flow in the vocal tract
Gaël Richard, D. Snider, H. Duncan, Qiguang Lin, James L. Flanagan, Stephen E. Levinson, Donald Davis, Scott Slimon
EUROSPEECH7
1994 An experiment in spoken language acquisition
abstract
The paper continues the authors' investigation of machines that adaptively acquire language through interaction with a complex environment. In particular, the present work focuses on the problem of spoken word acquisition, using the authors' proposed principles to motivate a method to govern the emergence of word symbols from the speech signal. The mechanism involves a connectionist network embedded in a feedback control system. The resulting system has two unique characteristics. First, no text is utilized by the device, in contrast to all other speech understanding systems. Second, the vocabulary and grammar is unconstrained, being acquired by the device during the course of performing its task. This is also in contrast to all other systems, in which the salient vocabulary and grammar are preprogrammed. A rudimentary baseline experiment is described, involving 1105 natural language utterances in an automated call routing application scenario.
Allen L. Gorin, Stephen E. Levinson, Ananth Sankar
IEEE Trans. Speech Audio Process.2
1993 Some experiments in spoken language acquisition
Allen L. Gorin, Laura G. Miller, Stephen E. Levinson
ICASSP (1)3
1993 Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial Intelligence
abstract
The findings of a workshop, the goals of which were to identify applications, research problems, and designs of high performance computing and communications (HPCC) systems for supporting applications are discussed. In computer vision, the main scientific issues are machine learning, surface reconstruction, inverse optics and integration, model acquisition, and perception and action. In speech and natural language processing (SNLP), issues were identified statistical analysis in corpus-based speech and language understanding, search strategies for language analysis, auditory and vocal-tract modeling, integration of multiple levels of speech and language analyses, and connectionist systems. In AI, important issues that need immediate attention include the development of efficient machine learning and heuristic search methods that can adapt to different architectural configurations, and the design and construction of scalable and verifiable knowledge bases, active memories, and artificial neural networks.>
Benjamin W. Wah, Thomas S. Huang, Aravind K. Joshi, Dan I. Moldovan, Yiannis Aloimonos, Ruzena Bajcsy, Dana H. Ballard, Doug DeGroot, Kenneth A. De Jong, Charles R. Dyer, Scott E. Fahlman, Ralph Grishman, Lynette Hirschman, Richard E. Korf, Stephen E. Levinson, Daniel P. Miranker, N. H. Morgan, Sergei Nirenburg, Tomaso A. Poggio, Edward M. Riseman, Craig Stanfil, Salvatore J. Stolfo, Steven L. Tanimoto, Charles C. Weems
IEEE Trans. Knowl. Data Eng.15
1991 Adaptive acquisition of spoken language
abstract
The problem of building a device that acquires language during the course of performing its task, called learning by doing, is considered. Some basic principles and mechanisms upon which such a device might be constructed are described. In particular, a language acquisition mechanism that is based upon the intuition of building associations between messages and appropriate responses to them and a mechanism for human-machine interaction based on control theory methods are investigated. A conversational-mode system that demonstrates and evaluates the proposed principles and mechanisms is described. Experimental results that validate the approach are reported.>
Allen L. Gorin, Stephen E. Levinson, A. N. Gertner
ICASSP2
1990 On adaptive acquisition of language
abstract
A system that automatically acquires a language model for a particular task from semantic-level information is described. This is in contrast to systems with predefined vocabulary and syntax. The purpose of the system is to map spoken or typed input into a machine action. To accomplish this task a medium-grain neural network is used. An adaptive training procedure is introduced for estimating the connection weights. It has the advantages of rapid, single-pass and order-invariant learning. The resulting weights have information-theoretic significance and do not require gradient search techniques for their estimation. The system was experimentally evaluated on three text-based tasks; a three-class inward-call manager with an acquired vocabulary of over 1600 words, a 15-action subset of the DARPA Resource Manager with an acquired vocabulary of over 700 words, and discrimination between idiomatic phrases meaning yes or no.>
Allen L. Gorin, Stephen E. Levinson, Laura G. Miller, A. N. Gertner, Andrej Ljolje, E. R. Goldman
ICASSP2
1990 Continuous speech recognition from a phonetic transcription
abstract
A widely accepted linguistic theory holds that speech recognition in humans proceeds from an intermediate representation of the acoustic signal in terms of a small number of phonetic symbols. A novel speech recognition system based on this theory in which the acoustic-to-phonetic mapping is accomplished by means of a particular form of hidden Markov model and is independent of lexical and syntactic constraint is described. Word recognition is then treated as a classical string-to-string editing problem which is solved with a two-level dynamic programming algorithm that accounts for lexical and syntactic structure. The system was tested on speaker-independent recognition of fluent speech from the 991-word DARPA resource management task, on which 76.6% word accuracy was achieved. In informal tests it was observed that the phonetic transcription can be resynthesized to provide a 100-bit/s vocoder with word intelligibility rates of approximately 75%.>
Stephen E. Levinson, Andrej Ljolje, Laura G. Miller
ICASSP1
1989 Speaker independent phonetic transcription of fluent speech for large vocabulary speech recognition
abstract
Results are presented of experiments on speaker independent phonetic transcription of fluent speech. The acoustic-phonetic model is a 38505-parameter continuously variable duration hidden Markov model which allows real-time phonetic transcription to be performed by means of a modified Viterbi algorithm. The model was trained on 3020 sentences from the TIMIT database. Testing was performed on the remaining 180 sentences. In a test without lexical or syntactic constraints, the authors obtained 52% correct phonetic transcription with 12% insertions. The design of a system for recognition of fluent speech based on the technique for phonetic transcription is described.>
Stephen E. Levinson, M. Y. Liberman, Andrej Ljolje, Laura G. Miller
ICASSP1
1988 Large vocabulary speech recognition using a hidden Markov model for acoustic/phonetic classification
abstract
Experiments with a speech recognition system are reported. The system comprises an acoustic/phonetic decoder, a lexical access mechanism and a syntax analyzer. The acoustic, phonetic and lexical processing are based on a continuously variable duration hidden Markov model (CVDHMM). The syntactic component is based on the Cocke-Kasami-Young (CKY) parser and a content-free covering grammar of English. Lexical items are represented in terms of the 43 phonetic units. In recognition tests conducted on a separate data set, a 70% correct recognition rate on phonetic units in fluent speech was observed. In two additional tests on isolated words, a 40% word recognition was observed with the complete 52000 word lexicon. When the vocabulary size was reduced to 1040 words, the recognition rate improved to 80%. After syntax analysis the word recognition rate rose to 90%.>
Stephen E. Levinson, Andrej Ljolje, Laura G. Miller
ICASSP1
1988 Syntactic analysis for large vocabulary speech recognition using a context-free covering grammar
abstract
The authors describe a syntactive component for large vocabulary speech recognition that incorporates a context-free covering grammar as the language model. The component comprises two modules: a preprocessing module and a syntactic analysis module. The preprocessing module consists of a lexical and a grammar preprocessor. The preprocessors facilitate grammar development. The syntactic analysis module consists of an error correcting maximum likelihood Cocke-Younger-Kasami parser and a context-free covering grammar. The parsar determines the sentence of maximum likelihood accepted by the grammar with respect to a word log likelihood matrix supplied b the acoustic/phonetic component of the speech recognition system. The effectiveness of this type of syntactic analysis was tested in a simulation using a 1040 word vocabulary. Results and error correcting examples are given.>
Laura G. Miller, Stephen E. Levinson
ICASSP2
1987 Continuous speech recognition by means of acoustic/ Phonetic classification obtained from a hidden Markov model
abstract
This paper describes an experimental continuous speech recognition system comprising procedures for acoustic/phonetic classification, lexical access and sentence retrieval. Speech is assumed to be composed of a small number of phonetic units which may be identified with the states of a hidden Markov model. The acoustic correlates of the phonetic units are then characterized by the observable Gaussian process associated with the corresponding state of the underlying Markov chain. Once the parameters of such a model are determined, a phonetic transcription of an utterance can be obtained by means of a Viterbi-like algorithm. Given a lexicon in which each entry is orthographically represented in terms of the chosen phonetic units, a word lattice is produced by a lexical access procedure. Lexical items whose orthography matches subsequences of the phonetic transcription are sought by means of a hash coding technique and their likelihoods are computed directly from the corresponding interval of acoustic measurements. The recognition process is completed by recovering from the word lattice, the string of words of maximum likelihood conditioned on the measurements. The desired string is derived by a best-first search algorithm. In an experimental evaluation of the system, the parameters of an acoustic/phonetic model were estimated from fluent utterances of 37 seven-digit numbers. A digit recognition rate of 96% was then observed on an independent test set of 59 utterances of the same form from the same speaker. Half of the observed errors resulted from insertions while deletions and substitutions accounted equally for the other half.
Stephen E. Levinson
ICASSP1
1986 Continuously variable duration hidden Markov models for speech analysis
abstract
During the past decade, the applicability of hidden Markov models (HMM) to various facets of speech analysis had been demonstrated in several different experiments. These investigations all rest on the assumption that speech is a quasi-stationary process whose stationary intervals can be identified with the occupancy of a single state of an appropriate HMM. In the traditional form of the HMM, the probability of duration of a state decreases exponentially with time. This behavior does not provide an adequate representation of the temporal structure of speech. The solution proposed here is to replace the probability distributions of duration with continuous probability density functions to form a continuously variable duration hidden Markov model (CVDHMM). The gamma distribution is ideally suited to specification of the durational density since it is one-sided and has only two parameters which, together, define both mean and variance. The main result is a derivation and proof of convergence of reestimation formulae for all the parameters of the CVDHMM. It is interesting to note that if the state durations are gamma distributed, one of the formulae is nonalgebraic but, fortuitously, has properties such that it is easily and rapidly solved numerically to any desired degree of accuracy. Other results are presented including the performance of the formulae on simulated data.
Stephen E. Levinson
ICASSP1
1986 Report of the 1985 IEEE ASSPS workshop on speech recognition
abstract
The IEEE Acoustics, Speech, and Signal Processing (ASSPS) Workshoonp Speech Recognition was held at the Arden House conference center of Columbia University in Harriman, NY, USA on 3-6, December 1985. This is the first time in five years that the ASSPS has sponsored a workshop devoted to speech recognition and it ends a 10 year absence of the society from Arden House. One hundred thirty seven researchers representing twelve countries and nearly as many disciplines gathered to discuss the frontiers of researcohn speech recognition. Each of the meeting's six sessions comprised a keynote address on a particular topic of interest, four brief presentations on specific aspects of the main theme and spontaneous and unrehearsed discussion by all attendees moderated by the session chairman. Approximately half of each session was allocated to open discussion. There were no parallel sessions so that all attendees could participate in all technical discussions. Taken as a whole, the workshop demonstrated that speech recognition is an active but severely factionated area of research.
Stephen E. Levinson
ICASSP1
1986 Maximum likelihood estimation for multivariate mixture observations of markov chains
abstract
To use probabilistic functions of a Markov chain to model certain parameterizations of the speech signal, we extend an estimation technique of Liporace to the eases of multivariate mixtures, such as Gaussian sums, and products of mixtures. We also show how these problems relate to Liporace's original framework.
Biing-Hwang Juang, Stephen E. Levinson, Man Mohan Sondhi
IEEE Trans. Inf. Theory2
1985 Recent developments in the application of hidden Markov models to speaker-independent isolated word recognition
abstract
In this paper we extend previous work on isolated word recognition based on hidden Markov models by replacing the discrete symbol representation of the speech signal by a continuous Gaussian mixture density. In this manner the inherent quantization error introduced by the discrete representation is essentially eliminated. The resulting recognizer was tested on a vocabulary of the 10 digits across a wide range of talkers and test conditions, and shown to have an error rate at least comparable to that of the best template recognizers and significantly lower than that of the discrete symbol hidden Markov model system. Several issues involved in the training of the continuous density models and in the implementation of the recognizer are discussed.
Biing-Hwang Juang, Lawrence R. Rabiner, Stephen E. Levinson, Man Mohan Sondhi
ICASSP3
1985 Structural methods in automatic speech recognition
abstract
The past decade has witnessed substantial progress toward the goal of constructing a machine capable of understanding colloquial discourse. Central to this progress has been the development and application of mathematical methods that permit modeling the speech signal as a complex code with several coexisting levels of structure. The most successful of these are "template matching," stochastic modeling, and probabilistic parsing. The manifestation of common themes such as dynamic programming and finite-state descriptions accentuates a superficial likeness amongst the methods which is often mistaken for the deeper similarity arising from their shared Bayesian foundation. In this paper, we outline the mathematical bases of these methods, invariant metrics, hidden Markov chains, and formal grammars, respectively. We then recount and briefly interpret the results of experiments in speech recognition to which the various methods were applied. Since these mathematical principles seem to bear little resemblance to traditional linguistic characterizations of speech, the success of the experiments is occasionally attributed, even by their authors, merely to excellent engineering. We conclude by speculating that, quite to the contrary, these methods actually constitute a powerful theory of speech that can be reconciled with and elucidate conventional linguistic theories while being used to build truly competent mechanical speech recognizers.
Stephen E. Levinson
Proc. IEEE1
1984 A vector quantizer incorporating both LPC shape and energy
abstract
The theory of vector quantization (VQ) of linear predictive coding (LPC) coefficients has established a wide variety of techniques for quantizing LPC spectral shape to minimize overall spectral distortion. Such vector quantizers have been widely used in the areas of speech coding and speech recognition. The conventional vector quantizer utilizes only spectral shape information and essentially disregards the energy or gain term associated with the optimal LPC fit to the signal being modelled. In this paper we present a method of incorporating LPC spectral shape and energy into the codebook entries of the vector quantizer. To do this we postulate a distortion measure for comparing two LPC vectors which uses a weighted sum of an LPC shape distortion and a log energy distortion. Based on this combined distortion measure we have designed and studied vector quantizers of several sizes for use in isolated word speech recognition experiments. We have found that a fairly significant correlation exists between LPC shape and signal energy; hence a combined LPC shape plus energy vector quantizer with a given distortion requires far fewer codebook entries than one in which LPC shape and energy are quantized separately. Based on isolated word recognition tests on both a 10-digit and a 129 word airlines vocabulary, we have found improvements in recognition accuracy by using the VQ with both LPC shape and energy over that obtained using a VQ with LPC shape alone.
Lawrence R. Rabiner, Man Mohan Sondhi, Stephen E. Levinson
ICASSP3
1983 Speaker independent isolated digit recognition using hidden Markov models
abstract
A method for speaker independent isolated digit recognition based on modeling entire words as discrete probabilistic functions of a Markov chain is described. Training is a three part process comprising conventional methods of linear prediction coding (LPC) and vector quantization of the LPCs followed by an algorithm for estimating the parameters of a hidden Markov process. Recognition utilizes linear prediction and vector quantization steps prior to maximum likelihood classification based on the Viterbi algorithm. Vector quantization is performed by a K-means algorithm which finds a codebook of 64 prototypical vectors that minimize the distortion measure (Itakura distance) over the training set. After training based on a 1,000 token set, recognition experiments were conducted on a separate 1,000 token test set obtained from the same talkers. In this test a 3.5% error rate was observed which is comparable to that measured in an identical test of an LPC/DTW (dynamic time warping) system. The computational demand for recognition under the new system is reduced by a factor of approximately 10 in both time and memory compared to that of the LPC/DTW system. It is also of interest that the classification errors made by the two systems are virtually disjoint; thus the possibility exists to obtain error rates near 1% by a combination of the methods. In describing our experiments we discuss several issues of theoretical importance, namely: 1) Alternatives to the Baum-Welch algorithm for model parameter estimation, e.g., Lagrangian techniques; 2) Model combining techniques by means of a bipartite graph matching algorithm providing improved model stability; 3) Methods for treating the finite training data problem by modifications to both the Baum-Welch algorithm and Lagrangian techniques; and 4) Use of non-ergodic Markov chains for isolated word recognition. We note that the experiments reported here are the first in which a direct comparison is made between two conceptually different (i.e. parametric and non-parametric) methods of treating the non-stationarity problem in speech recognition by implicitly dividing the speech signal into quasi-stationary intervals.
Stephen E. Levinson, Lawrence R. Rabiner, Man Mohan Sondhi
ICASSP1
1981 Connected word recognition using a syntax-directed dynamic programming temporal alignment procedure
abstract
In this paper we describe a system for connected word recognition in which a sentence in a formal language, uttered without pauses between words, is recognized by finding the grammatically well formed sequence of isolated word templates to which its distance is least. This is accomplished by means of a single monolithic algorithm in which temporal registration, segmentation and grammatical analysis are performed simultaneously. The algorithm is a syntax-directed version of the level building dynamic time warping algorithm of Myers and Rabiner. A test was conducted on a total of 208 sentences comprising 1781 words and spoken by two male and two female speakers. The sentences were composed from a 127 word vocabulary according to a moderately complex grammar and semantic structure appropriate to an airline information and reservation task. Test results revealed a 13% sentence error rate and a 6% word error rate.
Cory S. Myers, Stephen E. Levinson
ICASSP2
1981 A preliminary study on the use of demisyllables in automatic speech recognition
abstract
A speech recognition system is described for recognizing isolated words from reference templates created by concatenating demisyllables from a corpus of about 1000 demisyllables. The composition (in terms of demisyllables) of each reference word is specified in a lexicon with one or more entries for each word of the vocabulary. Experiments were carried out, using a 100-word vocabulary, to investigate the usefulness of such a representation and the effect on performance of some simple modifications in demisyllable specification and durations of reference patterns. Recognition accuracy of 97.6% was obtained using 132 reference templates for the 100-word vocabulary.
Aaron E. Rosenberg, Lawrence R. Rabiner, Stephen E. Levinson, Jay G. Wilpon
ICASSP3
1981 Isolated and Connected Word Recognition-Theory and Selected Applications
abstract
The art and science of speech recognition have been advanced to the state where it is now possible to communicate reliably with a computer by speaking to it in a disciplined manner using a vocabulary of moderate size. It is the purpose of this paper to outline two aspects of speech-recognition research. First, we discuss word recognition as a classical pattern-recognition problem and show how some fundamental concepts of signal processing, information theory, and computer science can be combined to give us the capability of robust recognition of isolated words and simple connected word sequences. We then describe methods whereby these principles, augmented by modern theories of formal language and semantic analysis, can be used to study some of the more general problems in speech recognition. It is anticipated that these methods will ultimately lead to accurate mechanical recognition of fluent speech under certain controlled conditions.
Lawrence R. Rabiner, Stephen E. Levinson
IEEE Trans. Commun.2
1980 A conversational mode airline information and reservation system using speech input and output
abstract
We describe a conversational mode speech understanding system which enables its user to make airline reservations and obtain timetable information through a spoken dialog. The system is structured as a three level hierarchy consisting of an acoustic word recognizer, a syntax analyzer and a semantic processor. The semantic level controls an audio response system making two way speech communication possible. The system is highly robust and operates on-line in a few times real time on a laboratory minicomputer. The speech communication channel is a standard telephone set connected to the computer by an ordinary dialed-up line.
Stephen E. Levinson, Kathleen L. Shipley
ICASSP1
1979 A new system for continuous speech recognition - preliminary results
abstract
A speaker dependent system for recognizing carefully articulated continuous speech is described. The system accepts English sentences composed from a 127 word vocabulary appropriate to an airline information reservation task. The system is controlled by a finite state parser which generates word candidates and established their temporal locations in hypothetical sentences. The word candidates are evaluated by an LPC distance measure and a dynamic programming algorithm which nonlinearly time aligns isolated word reference templates with the input speech stream. The input is recognized as the hypothetical sentence having the lowest distance according to a well-defined criterion. In a preliminary test based on 100 sentences spoken over dialed up telephone lines by two male talkers, 90% word accuracy, resulting in 75% sentence recognition, was achieved.
Stephen E. Levinson, Aaron E. Rosenberg
ICASSP1
1979 Speaker independent recognition of isolated words using clustering techniques
abstract
A speaker independent, isolated word recognition system is proposed which is based on the use of multiple templates for each word in the vocabulary. The word templates are obtained from a statistical clustering analysis of a large data base consisting of 100 replications of each word (i.e. once by each of 100 talkers). The recognition system, which uses telephone recordings, is based on an LPC analysis of the unknown word, dynamic time warping of each reference template to the unknown word (using the Itakura LPC distance measure), and the application of a K-nearest neighbor (KNN) decision rule to lower the probability of error. Results are presented on two test sets of data which show error rates that are comparable to, or better than, those obtained with speaker trained, isolated word recognition systems.
Lawrence R. Rabiner, Stephen E. Levinson, Aaron E. Rosenberg, Jay G. Wilpon
ICASSP2
1978 Some experiments with a syntax directed speech recognition system
abstract
In syntax directed speech recognition communication between the acoustic and syntactic processors of the system causes the acoustic analyzer to evaluate only those grammatically correct word hypotheses generated by the syntax analysis algorithm. In this paper we discuss two such systems which are much faster, but no less accurate than, systems in which the acoustic and syntactic analyses are sequential and isolated. Theoretical details of the systems are given and experimental results of tests are presented. Finally we discuss methods for using syntax directed recognition to improve overall system accuracy. We also show the relevance of one of our methods to the recognition of connected speech.
Stephen E. Levinson, Aaron E. Rosenberg
ICASSP1
1978 Computing relative redundancy to measure grammatical constraint in speech recognition tasks
abstract
In this paper we present new and computationally efficient algorithms for computing some statistical properties of finite languages. In particular, the relative redundancy which measures grammatical constraint, is computed for several speech recognition task languages which have appeared in the literature.
Man Mohan Sondhi, Stephen E. Levinson
ICASSP2
1976 Measuring pitch and formant frequencies for a speech understanding system
abstract
We present new techniques for the measurement of pitch period and formant frequencies. Individual pitch periods are estimated by finding local maxima in the power spectrum averaged over a frequency band that includes a formant. A spectral analysis is then performed for each pitch period over a data interval that is of shorter duration than the estimated pitch period. The pitch synchronous spectral analysis followed by post detection smoothing - preferably carried out over both time and frequency - leads to a spectrogram from which formant trajectories can be estimated by simple algorithms.
Donald W. Tufts, Stephen E. Levinson, R. Rao
ICASSP2
1975 The Vocal Speech Understanding System
Stephen E. Levinson
IJCAI1