EDBT 2026 Demo / reviewers in the wild / expert
Stephen E. Levinson
dblp:l/SELevinson
· DBLP profile ↗
52ranked-venue papers
13as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 11 first-authorArtificial intelligence and machine learning · 18 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Face, body and person analysis · 26% Representation and self-supervised learning · 22% Speech recognition and synthesis · 12% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 100% | |
| Computer graphics and multimedia
2 papers |
Audio and music processing · 100% |
Topics — the 27 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
affect recognition |
0.1 | 1 | 2007 | Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007 |
Computer vision › Face, body and person analysis › affect recognition
audiovisual emotion recognition |
0.1 | 1 | 2007 | Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007 |
Human-AI interaction › affective computing
affective human-computer interaction |
0.1 | 1 | 2007 | Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
0.1 | 1 | 2006 | Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
nonlinear manifold learning |
0.1 | 1 | 2006 | Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006 |
Machine learning › Time series and sequential data › time series analysis
time series learning |
0.1 | 1 | 2006 | Learning Nonlinear Manifolds from Time Series · ECCV (2) 2006 |
Human-AI interaction
affective computing |
0.1 | 1 | 2005 | Audio-Visual Affect Recognition through Multi-Stream Fused HMM for HCI · CVPR (2) 2005 |
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition |
0.0 | 1 | 2002 | A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002 |
Audio and music processing › sound source localization
binaural localization |
0.0 | 1 | 2002 | A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002 |
Audio and music processing
sound source localization |
0.0 | 1 | 2002 | A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002 |
Computer vision › 3D vision
camera pose estimation |
0.0 | 1 | 2001 | Tracking of Object with SVM Regression · CVPR (2) 2001 |
Computer vision › 3D vision › multi-view geometry
epipolar geometry estimation |
0.0 | 1 | 2001 | Tracking of Object with SVM Regression · CVPR (2) 2001 |
Computer vision › Video understanding and tracking
feature tracking |
0.0 | 1 | 2001 | Tracking of Object with SVM Regression · CVPR (2) 2001 |
Computer vision › Video understanding and tracking
object tracking |
0.0 | 1 | 2001 | Tracking of Object with SVM Regression · CVPR (2) 2001 |
Natural language and speech › Speech recognition and synthesis › speech analysis
prosody analysis |
0.0 | 1 | 2007 | Audio-Visual Affect Recognition · IEEE Trans. Multim. 2007 |
Natural language and speech › Language models and text generation › language acquisition
spoken language acquisition |
0.0 | 1 | 1994 | An experiment in spoken language acquisition · IEEE Trans. Speech Audio Process. 1994 |
Wireless sensing and localization › ranging
time difference of arrival |
0.0 | 1 | 2002 | A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a Robot · ICRA 2002 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1993 | Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial Intelligence · IEEE Trans. Knowl. Data Eng. 1993 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.0 | 2 | 1985 | Structural methods in automatic speech recognition · Proc. IEEE 1985 Isolated and Connected Word Recognition-Theory and Selected Applications · IEEE Trans. Commun. 1981 |
Natural language and speech › Speech recognition and synthesis
spoken language understanding |
0.0 | 2 | 1994 | An experiment in spoken language acquisition · IEEE Trans. Speech Audio Process. 1994 The Vocal Speech Understanding System · IJCAI 1975 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
0.0 | 1 | 1986 | Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation |
0.0 | 1 | 1986 | Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986 |
Audio and music processing › speech processing
speech modeling |
0.0 | 1 | 1986 | Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986 |
Audio and music processing
speech processing |
0.0 | 1 | 1986 | Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model |
0.0 | 1 | 1985 | Structural methods in automatic speech recognition · Proc. IEEE 1985 |
Computer vision › Image recognition and object detection
template matching |
0.0 | 1 | 1985 | Structural methods in automatic speech recognition · Proc. IEEE 1985 |
Audio and music processing
speech signal |
0.0 | 1 | 1986 | Maximum likelihood estimation for multivariate mixture observations of markov chains · IEEE Trans. Inf. Theory 1986 |
Methods — techniques the papers use, named apart from their topics
voting · 0.1feature selection · 0.1phase unwrapping · 0.1cross power spectrum · 0.1TDOA · 0.1multistream HMM · 0.1multi-stream HMM · 0.1maximum mutual information · 0.1maximum entropy principle · 0.1hidden markov model · 0.1outlier detection · 0.0affine transform estimation · 0.0SVM regression · 0.0connectionist network · 0.0heuristic search · 0.0connectionist systems · 0.0gaussian mixture · 0.0expectation-maximization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Nonlinear Dynamical Multi-Scale Model of Associative MemoryabstractHow can we get such reliable behavior from the mind when the brain is made up of such unreliable elements as neurons? We propose that the answer is related to the emergence of stable brain states and we offer a model that illustrates how such states could arise. We discuss a new ab initio nonlinear dynamical multi-scale model that will serve as the foundation for an associative memory. Scale 0 consists of spiking Hodgkin-Huxley (HH) neurons. Scale 1 consists of components that are made up of large populations of HH neurons whose topological structure evolves according to a Hebbian-plasticity rule based on synchronous firing. The component's state is captured by the variance of phase synchrony for the population. Many such components are sparsely connected to form a large network, whose state can be captured by the n-tuple consisting of the individual states of each member component. Scale 2 takes the state of the overall network and upon examining the particular interrelationships of each component (determining how the state of one component affects the state of others) is able to generate a class of trajectories that is multistationary and stable periodic. Such a class we consider a memory, the encoding of many such memories leads to the creation of a robust associative memory. The details of the different scales are examined. Alexander M. Duda, Stephen E. Levinson |
ICMLA | 2 |
| 2007 | HMM-Based Concept Learning for a Mobile RobotabstractWe are developing an intelligent robot and attempting to teach it language. While there are many aspects of this research, for the purposes here the most important are the following ideas. Language is primarily based on semantics, not syntax, which is still the focus in speech recognition research these days. To truly learn meaning, a language engine cannot simply be a computer program running on a desktop computer analyzing speech. It must be part of a more general, embodied intelligent system, one capable of using associative learning to form concepts from the perception of experiences in the world, and further capable of manipulating those concepts symbolically. In this paper, we present a general cascade model for learning concepts, and explore the use of hidden Markov models (HMMs) as part of the cascade model. HMMs are capable of automatically learning and extracting the underlying structure of continuous-valued inputs and representing that structure in the states of the model. These states can then be treated as symbolic representations of the inputs. We show how a cascade of HMMs can be embedded in a small mobile robot and used to find correlations among sensory inputs to learn a set of symbolic concepts, which are used for decision making and could eventually be manipulated linguistically Kevin M. Squire, Stephen E. Levinson |
IEEE Trans. Evol. Comput. | 2 |
| 2007 | Audio-Visual Affect RecognitionabstractThe ability of a computer to detect and appropriately respond to changes in a user's affective state has significant implications to human-computer interaction (HCI). In this paper, we present our efforts toward audio-visual affect recognition on 11 affective states customized for HCI application (four cognitive/motivational and seven basic affective states) of 20 nonactor subjects. A smoothing method is proposed to reduce the detrimental influence of speech on facial expression recognition. The feature selection analysis shows that subjects are prone to use brow movement in face, pitch and energy in prosody to express their affects while speaking. For person-dependent recognition, we apply the voting method to combine the frame-based classification results from both audio and visual channels. The result shows 7.5% improvement over the best unimodal performance. For person-independent test, we apply multistream HMM to combine the information from multiple component streams. This test shows 6.1% improvement over the best component performance Zhihong Zeng, Jilin Tu, Ming Liu 0009, Thomas S. Huang, Brian Pianfetti, Dan Roth 0001, Stephen E. Levinson |
IEEE Trans. Multim. | 7 |
| 2006 | Learning Nonlinear Manifolds from Time Series
Ruei-Sung Lin, Che-Bin Liu, Ming-Hsuan Yang 0001, Narendra Ahuja, Stephen E. Levinson |
ECCV (2) | 5 |
| 2006 | Extraction of pragmatic and semantic salience from spontaneous spoken English
Tong Zhang 0005, Mark Hasegawa-Johnson, Stephen E. Levinson |
Speech Commun. | 3 |
| 2006 | Cognitive state classification in a spoken tutorial dialogue system
Tong Zhang 0005, Mark Hasegawa-Johnson, Stephen E. Levinson |
Speech Commun. | 3 |
| 2005 | Audio-Visual Affect Recognition through Multi-Stream Fused HMM for HCIabstractAdvances in computer processing power and emerging algorithms are allowing new ways of envisioning human computer interaction. This paper focuses on the development of a computing algorithm that uses audio and visual sensors to detect and track a user's affective state to aid computer decision making. Using our multi-stream fused hidden Markov model (MFHMM), we analyzed coupled audio and visual streams to detect 11 cognitive/emotive states. The MFHMM allows the building of an optimal connection among multiple streams according to the maximum entropy principle and the maximum mutual information criterion. Person-independent experimental results from 20 subjects in 660 sequences show that the MFHMM approach performs with an accuracy of 80.61% which outperforms face-only HMM, pitch-only HMM, energy-only HMM, and independent HMM fusion. Zhihong Zeng, Jilin Tu, Brian Pianfetti, Ming Liu 0009, Tong Zhang 0005, ZhenQiu Zhang, Thomas S. Huang, Stephen E. Levinson |
CVPR (2) | 8 |
| 2004 | Adaptive Discriminative Generative Model for Object Tracking
Ruei-Sung Lin, Ming-Hsuan Yang 0001, Stephen E. Levinson |
ECAI | 3 |
| 2004 | Bimodal HCI-related affect recognitionabstractPerhaps the most fundamental application of affective computing will be Human-Computer Interaction (HCI) in which the computer should have the ability to detect and track the user's affective states, and make corresponding feedback. The human multi-sensor affect system defines the expectation of multimodal affect analyzer. In this paper, we present our efforts toward audio-visual HCI-related affect recognition. With HCI applications in mind, we take into account some special affective states which indicate users' cognitive/motivational states. Facing the fact that a facial expression is influenced by both an affective state and speech content, we apply a smoothing method to extract the information of the affective state from facial features. In our fusion stage, a voting method is applied to combine audio and visual modalities so that the final affect recognition accuracy is greatly improved. We test our bimodal affect recognition approach on 38 subjects with 11 HCI-related affect states. The extensive experimental results show that the average person-dependent affect recognition accuracy is almost 90% for our bimodal fusion. Zhihong Zeng, Jilin Tu, Ming Liu 0009, Tong Zhang 0005, Nick Rizzolo, ZhenQiu Zhang, Thomas S. Huang, Dan Roth 0001, Stephen E. Levinson |
ICMI | 9 |
| 2004 | Automatic detection of contrast for speech understandingabstractContrast is a very popular phenomenon in spoken language, and carries very important information to help understanding contents and structures of spoken language. In this paper, we propose an idea of automatic contrast detection as an effort for better speech understanding. We study the automatic tagging of three specific types of contrast: symmetric contrast, contrastive focus, and contrastive topic. We label the three types of contrasted words as contrast (C), and other words as noncontrast (¬C). The classification of contrast events is based on prosodic, spectral, and part-of-speech (POS) information sources. The integration of different knowledge sources is realized by a time-delay recursive neural network (TDRNN). The approach we proposed was testified on 235 spontaneous utterances consisting of 3500 words (samples). The contrast detection was speaker independent. The tests yielded an average of 87.9% classification rate. Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang 0005 |
INTERSPEECH | 2 |
| 2004 | Children's emotion recognition in an intelligent tutoring scenarioabstractThis paper presents an approach to automatically recognize emotion which children exhibit in an intelligent tutoring system. Emotion recognition can assist the computer agent to adapt its tutorial strategies to improve the efficiency of knowledge transmission. In this study, we detect three emotional classes: confidence, puzzle, and hesitation. Emotion is detected by means of lexical, prosodic, spectral, and syntactic analyses of users’ speech. An automatic speech recognition system serves as the fundamental constituent of the system. A robust classification and regression tree (CART) integrates the various information sources together for final decision. The effectiveness of the proposed approach has been tested on data collected by Wizard-of-Oz (WoZ) experiments. Our emotion recognition was speaker-independent, and yielded 91.3% accuracy. The test results showed that the spectral and duration-related prosodic features played very important roles in emotion recognition. Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang 0005 |
INTERSPEECH | 2 |
| 2004 | Semantic analysis for a speech user interface in an intelligent tutoring systemabstractIn this paper, we describe the strategy of semantic analysis for a speech user interface that is designed for a multimodal intelligent tutoring system. The semantic analysis involves three phases: semantic parsing, salient words/phrases spotting, and accented word detection. Semantic parsing attempts to represent the recognized sentence with a well-formed semantic frame. The recognized sentence consists of the a posterior most probably hypothesized words given the acoustic evidence, and is compliant with the grammatical knowledge that is represented by a semantic language model. The salient words/phrases are useful when semantic parsing fails. The accented words are useful when the user response is out of our expectations, and assist to make the computer agent smarter and smarter. Yuexi Ren, Mark Hasegawa-Johnson, Stephen E. Levinson |
IUI | 3 |
| 2003 | A Bayes-rule based hierarchical system for binaural sound source localizationabstractA Bayes-rule based hierarchical binaural sound source localization system is proposed. By combining three localization cues: interaural time differences (ITDs), interaural intensity differences (IIDs), and spectral cues, and a hierarchical decision making structure, this system enables a sound source to be located in a 3D space by using only the binaural inputs. Preliminary simulations have shown the effectiveness of this system. It can be used in studying binaural localization mechanism and applications such as in hearing aids and robotics. Danfeng Li, Stephen E. Levinson |
ICASSP (5) | 2 |
| 2003 | Automatic language acquisition by an autonomous robotabstractThere is no such thing as a disembodied mind. We posit that cognitive development can only occur through interaction with the physical world. To this end, we are developing a robotic platform for the purpose of studying cognition. We suggest that the central component of cognition is a memory which is primarily associative, one where learning occurs as the correlation of events from diverse inputs. We also posit that human-like cognition requires a well-integrated sensory-motor system, to provide these diverse inputs. As implemented in our robot, this system includes binaural hearing, stereo vision, tactile sense, and basic proprioceptive control. On top of these abilities, we are implementing and studying various models of processing, learning and decision making. Our goal is to produce a robot that will learn to carry out simple tasks in response to natural language requests. The robot's understanding of language will be learned concurrently with its other cognitive abilities. We have already developed a robust system and conducted a number or experiments on the way to this goal, some details of which appear in this paper. This is a first progress report of what we believe will be a long term project with significant implications. Stephen E. Levinson, Weiyu Zhu, Danfeng Li, Kevin Squire, Ruei-Sung Lin, Matthew Kleffner, Matthew McClain, Johnny Lee |
IJCNN | 1 |
| 2002 | Articulatory speech synthesis based upon fluid dynamic principlesabstractIn this paper, an articulatory speech synthesizer based on fluid dynamic principles is presented. The key idea is to devise a refined speech production model based on the most fundamental physics of the human vocal apparatus. Our articulatory synthesizer essentially contains two parts: a vocal fold model which represents the excitation source and a vocal tract model, which describes the positions of articulators. First, we propose a combined minimum error and minimum jerk criterion to estimated the moving vocal tract shapes during speech production. Second, we propose a nonlinear mechanical model to generate the vocal fold excitation signals. Finally, a computational fluid dynamics (CFD) approach is used to solve the Reynolds-averaged Navier-Stokes (RANS) equations, which are the governing equations of speech production inside vocal apparatus. Experimental results show that our system can synthesize intelligible continuous speech sentences while naturally handling the co-articulation effects during speech production. Stephen E. Levinson, Donald Davis, Scott Slimon |
ICASSP | 2 |
| 2002 | A Linear Phase Unwrapping Method for Binaural Sound Source Localization on a RobotabstractA robust linear phase unwrapping method is proposed to solve the 2/spl pi/ discontinuities in the phase of the cross power spectrum from the binaural inputs using two omnidirectional microphones. The relative incident angle of the interested sound is then estimated according to the time difference of arrival (TDOA) which is obtained from the unwrapped phase of the cross power spectrum. The frequency components associated with the high power are clustered into groups by the phase and frequency distance, and the dominant group is then used to obtain the initial slope estimation. The phase is unwrapped by checking the difference between the actual and the predicted phase by the estimated slope. The re-estimation is then performed by the unwrapped phase. The algorithm is tested under different incident angles and signal to noise ratio (SNR) using real speech signal and white Gaussian noise. The simulation results show the high accuracy and the robustness. This method is also Implemented to control a robot to adaptively adjust itself to the position facing the sound source directly. The satisfactory result was achieved in an open house demonstration. Danfeng Li, Stephen E. Levinson |
ICRA | 2 |
| 2001 | Tracking of Object with SVM RegressionabstractThis paper presents a novel feature-matching based approach for rigid object tracking. The proposed method models the tracking problem as discovering the affine transforms of object images between frames according to the extracted feature correspondences. False feature matches (outliers) are automatically detected and removed with a new SVM regression technique, where outliers are iteratively identified as support vectors with the gradually decreased insensitive margin /spl epsi/. This method, in addition to object tracking, can also be used for general feature-based epipolar constraint estimation, in which it can quickly detect outliers even if they make up, in theory, over 50% of the whole data. We have applied the proposed method to track real objects under cluttering backgrounds with very encouraging results. Weiyu Zhu, Ruei-Sung Lin, Stephen E. Levinson |
CVPR (2) | 4 |
| 2001 | Spoken Language Acquisition Via Human-Robot InteractionabstractThis paper presents a subproject of a challenging project that explores teaching a computer human-intelligence. In the subproject, a multisensory mobile robot is used as the interface for human-computer interaction, and spoken language is taught to the computer through natural human-robot interaction. Different from state-of-the-art speech recognizers, our approach associates speech patterns directly with sensory inputs of the robot. This approach allows our system to learn multilingual speech patterns online. Further investigation of this project will include human-computer interaction that involves more modalities, and applications that use the proposed idea to train home appliances. 1. Qiong Liu 0003, Thomas S. Huang, Ying Wu 0001, Stephen E. Levinson |
ICME | 4 |
| 2000 | Edge Orientation-Based Multi-View Object RecognitionabstractAn edge orientation-based algorithm for multi-view object recognition is presented. The distribution of edge point orientations, combined with the normalized second moments, is taken as a feature vector to describe and index each object instance. For each unknown test object, a set of likelihood weights for all the possible candidate objects is obtained by computing the Euclidean distances between the unknown feature set and all the available template feature vectors. A convincing coefficient is introduced to evaluate the confidence of the best match. New views (photo shots) will be automatically taken if the best match is thought to be insufficiently convincing. In experiments, our algorithm has achieved an average of 91.5% correct recognition rate under the 5-view scheme for 320 testing images taken from eight natural objects. Weiyu Zhu, Stephen E. Levinson |
ICPR | 2 |
| 2000 | Signal approximation in Hilbert space and its application on articulatory speech synthesis
Stephen E. Levinson, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2000 | Word concept model: a knowledge representation for dialogue agentsabstractInformation extraction is a key component in dialogue systems. Knowledge about the world as well as knowledge specific to each word should be used for robust semantic processing. An intelligent agent is necessary for a dialogue system when meanings are strictly defined by using a world state model. A layered concept structure is proposed to represent knowledge associated with each word in a "speech-friendly" way. By considering knowledge stored in the word concept model as well as knowledge base of the world model, meaning of a given sentence can be correctly identified. This paper describes the layered concept structure and how knowledge about words can be stored in this concept model. Tong Zhang 0005, Stephen E. Levinson |
INTERSPEECH | 3 |
| 1999 | Video Sequence Learning and Recognition Via Dynamic SomabstractInformation contained in video sequences is crucial for an autonomous robot or a computer to learn and respond to its surrounding environment. In the past, robot vision mainly concentrated on still image processing and small "image cube" processing. Continuous video sequence learning and recognition is rarely addressed in the literature due to its high requirement of dynamic processing. In this paper, we propose a novel neural network structure called dynamic self-organizing map (DSOM) for video sequence processing. The proposed technique has been tested on simulation data sets, and the results validate its learning/recognition ability. Qiong Liu 0003, Yong Rui, Thomas S. Huang, Stephen E. Levinson |
ICIP (4) | 4 |
| 1999 | Temporal sequence learning and recognition with dynamic SOMabstractThe purpose of the paper is to propose a map-like artificial neural network for temporal sequence pattern clustering. The map construction in our presentation is related to the self-organizing map (SOM) idea. The SOM idea was originally designed for static pattern learning and recognition. It has been found efficient for organizing high dimensional data sets. One of the biggest limitations of the traditional SOM technique is caused by its static characteristics. We propose a new neural network construction model and its corresponding training algorithm based on traditional SOM training technology and backpropagation training technology. It overcomes the static limitation of traditional SOM and tries to reach a new stage for dynamic pattern clustering, and recognition. At the end of the paper, we give some experimental results for testing this proposed method on real speech data. Qiong Liu 0003, Sylvian R. Ray, Stephen E. Levinson, Thomas S. Huang |
IJCNN | 3 |
| 1999 | A DCT-based fast enhancement technique for robust speech recognition in automobile usage
Yunxin Zhao, Stephen E. Levinson |
EUROSPEECH | 3 |
| 1995 | Numerical simulations of fluid flow in the vocal tract
Gaël Richard, D. Snider, H. Duncan, Qiguang Lin, James L. Flanagan, Stephen E. Levinson, Donald Davis, Scott Slimon |
EUROSPEECH | 7 |
| 1994 | An experiment in spoken language acquisitionabstractThe paper continues the authors' investigation of machines that adaptively acquire language through interaction with a complex environment. In particular, the present work focuses on the problem of spoken word acquisition, using the authors' proposed principles to motivate a method to govern the emergence of word symbols from the speech signal. The mechanism involves a connectionist network embedded in a feedback control system. The resulting system has two unique characteristics. First, no text is utilized by the device, in contrast to all other speech understanding systems. Second, the vocabulary and grammar is unconstrained, being acquired by the device during the course of performing its task. This is also in contrast to all other systems, in which the salient vocabulary and grammar are preprogrammed. A rudimentary baseline experiment is described, involving 1105 natural language utterances in an automated call routing application scenario. Allen L. Gorin, Stephen E. Levinson, Ananth Sankar |
IEEE Trans. Speech Audio Process. | 2 |
| 1993 | Some experiments in spoken language acquisition
Allen L. Gorin, Laura G. Miller, Stephen E. Levinson |
ICASSP (1) | 3 |
| 1993 | Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial IntelligenceabstractThe findings of a workshop, the goals of which were to identify applications, research problems, and designs of high performance computing and communications (HPCC) systems for supporting applications are discussed. In computer vision, the main scientific issues are machine learning, surface reconstruction, inverse optics and integration, model acquisition, and perception and action. In speech and natural language processing (SNLP), issues were identified statistical analysis in corpus-based speech and language understanding, search strategies for language analysis, auditory and vocal-tract modeling, integration of multiple levels of speech and language analyses, and connectionist systems. In AI, important issues that need immediate attention include the development of efficient machine learning and heuristic search methods that can adapt to different architectural configurations, and the design and construction of scalable and verifiable knowledge bases, active memories, and artificial neural networks.> Benjamin W. Wah, Thomas S. Huang, Aravind K. Joshi, Dan I. Moldovan, Yiannis Aloimonos, Ruzena Bajcsy, Dana H. Ballard, Doug DeGroot, Kenneth A. De Jong, Charles R. Dyer, Scott E. Fahlman, Ralph Grishman, Lynette Hirschman, Richard E. Korf, Stephen E. Levinson, Daniel P. Miranker, N. H. Morgan, Sergei Nirenburg, Tomaso A. Poggio, Edward M. Riseman, Craig Stanfil, Salvatore J. Stolfo, Steven L. Tanimoto, Charles C. Weems |
IEEE Trans. Knowl. Data Eng. | 15 |
| 1991 | Adaptive acquisition of spoken languageabstractThe problem of building a device that acquires language during the course of performing its task, called learning by doing, is considered. Some basic principles and mechanisms upon which such a device might be constructed are described. In particular, a language acquisition mechanism that is based upon the intuition of building associations between messages and appropriate responses to them and a mechanism for human-machine interaction based on control theory methods are investigated. A conversational-mode system that demonstrates and evaluates the proposed principles and mechanisms is described. Experimental results that validate the approach are reported.> Allen L. Gorin, Stephen E. Levinson, A. N. Gertner |
ICASSP | 2 |
| 1990 | On adaptive acquisition of languageabstractA system that automatically acquires a language model for a particular task from semantic-level information is described. This is in contrast to systems with predefined vocabulary and syntax. The purpose of the system is to map spoken or typed input into a machine action. To accomplish this task a medium-grain neural network is used. An adaptive training procedure is introduced for estimating the connection weights. It has the advantages of rapid, single-pass and order-invariant learning. The resulting weights have information-theoretic significance and do not require gradient search techniques for their estimation. The system was experimentally evaluated on three text-based tasks; a three-class inward-call manager with an acquired vocabulary of over 1600 words, a 15-action subset of the DARPA Resource Manager with an acquired vocabulary of over 700 words, and discrimination between idiomatic phrases meaning yes or no.> Allen L. Gorin, Stephen E. Levinson, Laura G. Miller, A. N. Gertner, Andrej Ljolje, E. R. Goldman |
ICASSP | 2 |
| 1990 | Continuous speech recognition from a phonetic transcriptionabstractA widely accepted linguistic theory holds that speech recognition in humans proceeds from an intermediate representation of the acoustic signal in terms of a small number of phonetic symbols. A novel speech recognition system based on this theory in which the acoustic-to-phonetic mapping is accomplished by means of a particular form of hidden Markov model and is independent of lexical and syntactic constraint is described. Word recognition is then treated as a classical string-to-string editing problem which is solved with a two-level dynamic programming algorithm that accounts for lexical and syntactic structure. The system was tested on speaker-independent recognition of fluent speech from the 991-word DARPA resource management task, on which 76.6% word accuracy was achieved. In informal tests it was observed that the phonetic transcription can be resynthesized to provide a 100-bit/s vocoder with word intelligibility rates of approximately 75%.> Stephen E. Levinson, Andrej Ljolje, Laura G. Miller |
ICASSP | 1 |
| 1989 | Speaker independent phonetic transcription of fluent speech for large vocabulary speech recognitionabstractResults are presented of experiments on speaker independent phonetic transcription of fluent speech. The acoustic-phonetic model is a 38505-parameter continuously variable duration hidden Markov model which allows real-time phonetic transcription to be performed by means of a modified Viterbi algorithm. The model was trained on 3020 sentences from the TIMIT database. Testing was performed on the remaining 180 sentences. In a test without lexical or syntactic constraints, the authors obtained 52% correct phonetic transcription with 12% insertions. The design of a system for recognition of fluent speech based on the technique for phonetic transcription is described.> Stephen E. Levinson, M. Y. Liberman, Andrej Ljolje, Laura G. Miller |
ICASSP | 1 |
| 1988 | Large vocabulary speech recognition using a hidden Markov model for acoustic/phonetic classificationabstractExperiments with a speech recognition system are reported. The system comprises an acoustic/phonetic decoder, a lexical access mechanism and a syntax analyzer. The acoustic, phonetic and lexical processing are based on a continuously variable duration hidden Markov model (CVDHMM). The syntactic component is based on the Cocke-Kasami-Young (CKY) parser and a content-free covering grammar of English. Lexical items are represented in terms of the 43 phonetic units. In recognition tests conducted on a separate data set, a 70% correct recognition rate on phonetic units in fluent speech was observed. In two additional tests on isolated words, a 40% word recognition was observed with the complete 52000 word lexicon. When the vocabulary size was reduced to 1040 words, the recognition rate improved to 80%. After syntax analysis the word recognition rate rose to 90%.> Stephen E. Levinson, Andrej Ljolje, Laura G. Miller |
ICASSP | 1 |
| 1988 | Syntactic analysis for large vocabulary speech recognition using a context-free covering grammarabstractThe authors describe a syntactive component for large vocabulary speech recognition that incorporates a context-free covering grammar as the language model. The component comprises two modules: a preprocessing module and a syntactic analysis module. The preprocessing module consists of a lexical and a grammar preprocessor. The preprocessors facilitate grammar development. The syntactic analysis module consists of an error correcting maximum likelihood Cocke-Younger-Kasami parser and a context-free covering grammar. The parsar determines the sentence of maximum likelihood accepted by the grammar with respect to a word log likelihood matrix supplied b the acoustic/phonetic component of the speech recognition system. The effectiveness of this type of syntactic analysis was tested in a simulation using a 1040 word vocabulary. Results and error correcting examples are given.> Laura G. Miller, Stephen E. Levinson |
ICASSP | 2 |
| 1987 | Continuous speech recognition by means of acoustic/ Phonetic classification obtained from a hidden Markov modelabstractThis paper describes an experimental continuous speech recognition system comprising procedures for acoustic/phonetic classification, lexical access and sentence retrieval. Speech is assumed to be composed of a small number of phonetic units which may be identified with the states of a hidden Markov model. The acoustic correlates of the phonetic units are then characterized by the observable Gaussian process associated with the corresponding state of the underlying Markov chain. Once the parameters of such a model are determined, a phonetic transcription of an utterance can be obtained by means of a Viterbi-like algorithm. Given a lexicon in which each entry is orthographically represented in terms of the chosen phonetic units, a word lattice is produced by a lexical access procedure. Lexical items whose orthography matches subsequences of the phonetic transcription are sought by means of a hash coding technique and their likelihoods are computed directly from the corresponding interval of acoustic measurements. The recognition process is completed by recovering from the word lattice, the string of words of maximum likelihood conditioned on the measurements. The desired string is derived by a best-first search algorithm. In an experimental evaluation of the system, the parameters of an acoustic/phonetic model were estimated from fluent utterances of 37 seven-digit numbers. A digit recognition rate of 96% was then observed on an independent test set of 59 utterances of the same form from the same speaker. Half of the observed errors resulted from insertions while deletions and substitutions accounted equally for the other half. Stephen E. Levinson |
ICASSP | 1 |
| 1986 | Continuously variable duration hidden Markov models for speech analysisabstractDuring the past decade, the applicability of hidden Markov models (HMM) to various facets of speech analysis had been demonstrated in several different experiments. These investigations all rest on the assumption that speech is a quasi-stationary process whose stationary intervals can be identified with the occupancy of a single state of an appropriate HMM. In the traditional form of the HMM, the probability of duration of a state decreases exponentially with time. This behavior does not provide an adequate representation of the temporal structure of speech. The solution proposed here is to replace the probability distributions of duration with continuous probability density functions to form a continuously variable duration hidden Markov model (CVDHMM). The gamma distribution is ideally suited to specification of the durational density since it is one-sided and has only two parameters which, together, define both mean and variance. The main result is a derivation and proof of convergence of reestimation formulae for all the parameters of the CVDHMM. It is interesting to note that if the state durations are gamma distributed, one of the formulae is nonalgebraic but, fortuitously, has properties such that it is easily and rapidly solved numerically to any desired degree of accuracy. Other results are presented including the performance of the formulae on simulated data. Stephen E. Levinson |
ICASSP | 1 |
| 1986 | Report of the 1985 IEEE ASSPS workshop on speech recognitionabstractThe IEEE Acoustics, Speech, and Signal Processing (ASSPS) Workshoonp Speech Recognition was held at the Arden House conference center of Columbia University in Harriman, NY, USA on 3-6, December 1985. This is the first time in five years that the ASSPS has sponsored a workshop devoted to speech recognition and it ends a 10 year absence of the society from Arden House. One hundred thirty seven researchers representing twelve countries and nearly as many disciplines gathered to discuss the frontiers of researcohn speech recognition. Each of the meeting's six sessions comprised a keynote address on a particular topic of interest, four brief presentations on specific aspects of the main theme and spontaneous and unrehearsed discussion by all attendees moderated by the session chairman. Approximately half of each session was allocated to open discussion. There were no parallel sessions so that all attendees could participate in all technical discussions. Taken as a whole, the workshop demonstrated that speech recognition is an active but severely factionated area of research. Stephen E. Levinson |
ICASSP | 1 |
| 1986 | Maximum likelihood estimation for multivariate mixture observations of markov chainsabstractTo use probabilistic functions of a Markov chain to model certain parameterizations of the speech signal, we extend an estimation technique of Liporace to the eases of multivariate mixtures, such as Gaussian sums, and products of mixtures. We also show how these problems relate to Liporace's original framework. Biing-Hwang Juang, Stephen E. Levinson, Man Mohan Sondhi |
IEEE Trans. Inf. Theory | 2 |
| 1985 | Recent developments in the application of hidden Markov models to speaker-independent isolated word recognitionabstractIn this paper we extend previous work on isolated word recognition based on hidden Markov models by replacing the discrete symbol representation of the speech signal by a continuous Gaussian mixture density. In this manner the inherent quantization error introduced by the discrete representation is essentially eliminated. The resulting recognizer was tested on a vocabulary of the 10 digits across a wide range of talkers and test conditions, and shown to have an error rate at least comparable to that of the best template recognizers and significantly lower than that of the discrete symbol hidden Markov model system. Several issues involved in the training of the continuous density models and in the implementation of the recognizer are discussed. Biing-Hwang Juang, Lawrence R. Rabiner, Stephen E. Levinson, Man Mohan Sondhi |
ICASSP | 3 |
| 1985 | Structural methods in automatic speech recognitionabstractThe past decade has witnessed substantial progress toward the goal of constructing a machine capable of understanding colloquial discourse. Central to this progress has been the development and application of mathematical methods that permit modeling the speech signal as a complex code with several coexisting levels of structure. The most successful of these are "template matching," stochastic modeling, and probabilistic parsing. The manifestation of common themes such as dynamic programming and finite-state descriptions accentuates a superficial likeness amongst the methods which is often mistaken for the deeper similarity arising from their shared Bayesian foundation. In this paper, we outline the mathematical bases of these methods, invariant metrics, hidden Markov chains, and formal grammars, respectively. We then recount and briefly interpret the results of experiments in speech recognition to which the various methods were applied. Since these mathematical principles seem to bear little resemblance to traditional linguistic characterizations of speech, the success of the experiments is occasionally attributed, even by their authors, merely to excellent engineering. We conclude by speculating that, quite to the contrary, these methods actually constitute a powerful theory of speech that can be reconciled with and elucidate conventional linguistic theories while being used to build truly competent mechanical speech recognizers. Stephen E. Levinson |
Proc. IEEE | 1 |
| 1984 | A vector quantizer incorporating both LPC shape and energyabstractThe theory of vector quantization (VQ) of linear predictive coding (LPC) coefficients has established a wide variety of techniques for quantizing LPC spectral shape to minimize overall spectral distortion. Such vector quantizers have been widely used in the areas of speech coding and speech recognition. The conventional vector quantizer utilizes only spectral shape information and essentially disregards the energy or gain term associated with the optimal LPC fit to the signal being modelled. In this paper we present a method of incorporating LPC spectral shape and energy into the codebook entries of the vector quantizer. To do this we postulate a distortion measure for comparing two LPC vectors which uses a weighted sum of an LPC shape distortion and a log energy distortion. Based on this combined distortion measure we have designed and studied vector quantizers of several sizes for use in isolated word speech recognition experiments. We have found that a fairly significant correlation exists between LPC shape and signal energy; hence a combined LPC shape plus energy vector quantizer with a given distortion requires far fewer codebook entries than one in which LPC shape and energy are quantized separately. Based on isolated word recognition tests on both a 10-digit and a 129 word airlines vocabulary, we have found improvements in recognition accuracy by using the VQ with both LPC shape and energy over that obtained using a VQ with LPC shape alone. Lawrence R. Rabiner, Man Mohan Sondhi, Stephen E. Levinson |
ICASSP | 3 |
| 1983 | Speaker independent isolated digit recognition using hidden Markov modelsabstractA method for speaker independent isolated digit recognition based on modeling entire words as discrete probabilistic functions of a Markov chain is described. Training is a three part process comprising conventional methods of linear prediction coding (LPC) and vector quantization of the LPCs followed by an algorithm for estimating the parameters of a hidden Markov process. Recognition utilizes linear prediction and vector quantization steps prior to maximum likelihood classification based on the Viterbi algorithm. Vector quantization is performed by a K-means algorithm which finds a codebook of 64 prototypical vectors that minimize the distortion measure (Itakura distance) over the training set. After training based on a 1,000 token set, recognition experiments were conducted on a separate 1,000 token test set obtained from the same talkers. In this test a 3.5% error rate was observed which is comparable to that measured in an identical test of an LPC/DTW (dynamic time warping) system. The computational demand for recognition under the new system is reduced by a factor of approximately 10 in both time and memory compared to that of the LPC/DTW system. It is also of interest that the classification errors made by the two systems are virtually disjoint; thus the possibility exists to obtain error rates near 1% by a combination of the methods. In describing our experiments we discuss several issues of theoretical importance, namely: 1) Alternatives to the Baum-Welch algorithm for model parameter estimation, e.g., Lagrangian techniques; 2) Model combining techniques by means of a bipartite graph matching algorithm providing improved model stability; 3) Methods for treating the finite training data problem by modifications to both the Baum-Welch algorithm and Lagrangian techniques; and 4) Use of non-ergodic Markov chains for isolated word recognition. We note that the experiments reported here are the first in which a direct comparison is made between two conceptually different (i.e. parametric and non-parametric) methods of treating the non-stationarity problem in speech recognition by implicitly dividing the speech signal into quasi-stationary intervals. Stephen E. Levinson, Lawrence R. Rabiner, Man Mohan Sondhi |
ICASSP | 1 |
| 1981 | Connected word recognition using a syntax-directed dynamic programming temporal alignment procedureabstractIn this paper we describe a system for connected word recognition in which a sentence in a formal language, uttered without pauses between words, is recognized by finding the grammatically well formed sequence of isolated word templates to which its distance is least. This is accomplished by means of a single monolithic algorithm in which temporal registration, segmentation and grammatical analysis are performed simultaneously. The algorithm is a syntax-directed version of the level building dynamic time warping algorithm of Myers and Rabiner. A test was conducted on a total of 208 sentences comprising 1781 words and spoken by two male and two female speakers. The sentences were composed from a 127 word vocabulary according to a moderately complex grammar and semantic structure appropriate to an airline information and reservation task. Test results revealed a 13% sentence error rate and a 6% word error rate. Cory S. Myers, Stephen E. Levinson |
ICASSP | 2 |
| 1981 | A preliminary study on the use of demisyllables in automatic speech recognitionabstractA speech recognition system is described for recognizing isolated words from reference templates created by concatenating demisyllables from a corpus of about 1000 demisyllables. The composition (in terms of demisyllables) of each reference word is specified in a lexicon with one or more entries for each word of the vocabulary. Experiments were carried out, using a 100-word vocabulary, to investigate the usefulness of such a representation and the effect on performance of some simple modifications in demisyllable specification and durations of reference patterns. Recognition accuracy of 97.6% was obtained using 132 reference templates for the 100-word vocabulary. Aaron E. Rosenberg, Lawrence R. Rabiner, Stephen E. Levinson, Jay G. Wilpon |
ICASSP | 3 |
| 1981 | Isolated and Connected Word Recognition-Theory and Selected ApplicationsabstractThe art and science of speech recognition have been advanced to the state where it is now possible to communicate reliably with a computer by speaking to it in a disciplined manner using a vocabulary of moderate size. It is the purpose of this paper to outline two aspects of speech-recognition research. First, we discuss word recognition as a classical pattern-recognition problem and show how some fundamental concepts of signal processing, information theory, and computer science can be combined to give us the capability of robust recognition of isolated words and simple connected word sequences. We then describe methods whereby these principles, augmented by modern theories of formal language and semantic analysis, can be used to study some of the more general problems in speech recognition. It is anticipated that these methods will ultimately lead to accurate mechanical recognition of fluent speech under certain controlled conditions. Lawrence R. Rabiner, Stephen E. Levinson |
IEEE Trans. Commun. | 2 |
| 1980 | A conversational mode airline information and reservation system using speech input and outputabstractWe describe a conversational mode speech understanding system which enables its user to make airline reservations and obtain timetable information through a spoken dialog. The system is structured as a three level hierarchy consisting of an acoustic word recognizer, a syntax analyzer and a semantic processor. The semantic level controls an audio response system making two way speech communication possible. The system is highly robust and operates on-line in a few times real time on a laboratory minicomputer. The speech communication channel is a standard telephone set connected to the computer by an ordinary dialed-up line. Stephen E. Levinson, Kathleen L. Shipley |
ICASSP | 1 |
| 1979 | A new system for continuous speech recognition - preliminary resultsabstractA speaker dependent system for recognizing carefully articulated continuous speech is described. The system accepts English sentences composed from a 127 word vocabulary appropriate to an airline information reservation task. The system is controlled by a finite state parser which generates word candidates and established their temporal locations in hypothetical sentences. The word candidates are evaluated by an LPC distance measure and a dynamic programming algorithm which nonlinearly time aligns isolated word reference templates with the input speech stream. The input is recognized as the hypothetical sentence having the lowest distance according to a well-defined criterion. In a preliminary test based on 100 sentences spoken over dialed up telephone lines by two male talkers, 90% word accuracy, resulting in 75% sentence recognition, was achieved. Stephen E. Levinson, Aaron E. Rosenberg |
ICASSP | 1 |
| 1979 | Speaker independent recognition of isolated words using clustering techniquesabstractA speaker independent, isolated word recognition system is proposed which is based on the use of multiple templates for each word in the vocabulary. The word templates are obtained from a statistical clustering analysis of a large data base consisting of 100 replications of each word (i.e. once by each of 100 talkers). The recognition system, which uses telephone recordings, is based on an LPC analysis of the unknown word, dynamic time warping of each reference template to the unknown word (using the Itakura LPC distance measure), and the application of a K-nearest neighbor (KNN) decision rule to lower the probability of error. Results are presented on two test sets of data which show error rates that are comparable to, or better than, those obtained with speaker trained, isolated word recognition systems. Lawrence R. Rabiner, Stephen E. Levinson, Aaron E. Rosenberg, Jay G. Wilpon |
ICASSP | 2 |
| 1978 | Some experiments with a syntax directed speech recognition systemabstractIn syntax directed speech recognition communication between the acoustic and syntactic processors of the system causes the acoustic analyzer to evaluate only those grammatically correct word hypotheses generated by the syntax analysis algorithm. In this paper we discuss two such systems which are much faster, but no less accurate than, systems in which the acoustic and syntactic analyses are sequential and isolated. Theoretical details of the systems are given and experimental results of tests are presented. Finally we discuss methods for using syntax directed recognition to improve overall system accuracy. We also show the relevance of one of our methods to the recognition of connected speech. Stephen E. Levinson, Aaron E. Rosenberg |
ICASSP | 1 |
| 1978 | Computing relative redundancy to measure grammatical constraint in speech recognition tasksabstractIn this paper we present new and computationally efficient algorithms for computing some statistical properties of finite languages. In particular, the relative redundancy which measures grammatical constraint, is computed for several speech recognition task languages which have appeared in the literature. Man Mohan Sondhi, Stephen E. Levinson |
ICASSP | 2 |
| 1976 | Measuring pitch and formant frequencies for a speech understanding systemabstractWe present new techniques for the measurement of pitch period and formant frequencies. Individual pitch periods are estimated by finding local maxima in the power spectrum averaged over a frequency band that includes a formant. A spectral analysis is then performed for each pitch period over a data interval that is of shorter duration than the estimated pitch period. The pitch synchronous spectral analysis followed by post detection smoothing - preferably carried out over both time and frequency - leads to a spectrogram from which formant trajectories can be estimated by simple algorithms. Donald W. Tufts, Stephen E. Levinson, R. Rao |
ICASSP | 2 |
| 1975 | The Vocal Speech Understanding System
Stephen E. Levinson |
IJCAI | 1 |