Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Joseph Picone

dblp:37/1933 · also Joe Picone · DBLP profile ↗
← Back
53ranked-venue papers
13as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 10 first-authorArtificial intelligence and machine learning · 21 · 1 first-authorComputer networks · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 70% Speech recognition and synthesis · 29% Language models and text generation · 1%
Computer graphics and multimedia
2 papers
Audio and music processing · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
acoustic modeling
0.322016
A Doubly Hierarchical Dirichlet Process Hidden Markov Model with a Non-Ergodic Structure · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Syllable-based large vocabulary continuous speech recognition · IEEE Trans. Speech Audio Process. 2001
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.212016
A Doubly Hierarchical Dirichlet Process Hidden Markov Model with a Non-Ergodic Structure · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.212016
A Doubly Hierarchical Dirichlet Process Hidden Markov Model with a Non-Ergodic Structure · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
hierarchical dirichlet process
0.212016
A Doubly Hierarchical Dirichlet Process Hidden Markov Model with a Non-Ergodic Structure · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › continuous speech recognition
large vocabulary continuous speech recognition
0.012001
Syllable-based large vocabulary continuous speech recognition · IEEE Trans. Speech Audio Process. 2001
Audio and music processing
speech processing
0.021999
Kanji-to-Hiragana conversion based on a length-constrained n-gram analysis · IEEE Trans. Speech Audio Process. 1999
Design and implementation of a robust pitch detector based on a parallel processing technique · IEEE J. Sel. Areas Commun. 1988
Audio and music processing
grapheme-to-phoneme conversion
0.011999
Kanji-to-Hiragana conversion based on a length-constrained n-gram analysis · IEEE Trans. Speech Audio Process. 1999
Natural language and speech › Language models and text generation
japanese text processing
0.011999
Kanji-to-Hiragana conversion based on a length-constrained n-gram analysis · IEEE Trans. Speech Audio Process. 1999
Audio and music processing › speech processing
pitch estimation
0.011988
Design and implementation of a robust pitch detector based on a parallel processing technique · IEEE J. Sel. Areas Commun. 1988
Physical-layer communications
signal processing for communications
0.011982
An Application of the Barrett-Lampard Expansion to Obtain the Slope Characteristics of a Speech Signal · IEEE Trans. Commun. 1982
Audio and music processing
speech coding
0.011988
Design and implementation of a robust pitch detector based on a parallel processing technique · IEEE J. Sel. Areas Commun. 1988
Audio and music processing › speech coding
vocoder
0.011988
Design and implementation of a robust pitch detector based on a parallel processing technique · IEEE J. Sel. Areas Commun. 1988
Environmental and earth informatics
oceanography
0.011985
Acoustic waveguides applications to oceanic science · Proc. IEEE 1985
Coding theory
source coding
0.011982
An Application of the Barrett-Lampard Expansion to Obtain the Slope Characteristics of a Speech Signal · IEEE Trans. Commun. 1982
Coding theory › source coding › lossy source coding
speech coding
0.011982
An Application of the Barrett-Lampard Expansion to Obtain the Slope Characteristics of a Speech Signal · IEEE Trans. Commun. 1982

Methods — techniques the papers use, named apart from their topics

non-ergodic structure · 0.2hierarchical dirichlet process · 0.2n-gram analysis · 0.0length-constrained matching · 0.0triphone comparison · 0.0syllable-level acoustic unit · 0.0parallel processing · 0.0fixed-point implementation · 0.0digital filtering · 0.0barrett-lampard expansion · 0.0
YearPublicationVenuePosition
2018 Deep Architectures for Spatio-Temporal Modeling: Automated Seizure Detection in Scalp EEGs
abstract
Automated seizure detection using clinical electroencephalograms is a challenging machine learning problem because the multichannel signal often has an extremely low signal to noise ratio. Events of interest such as seizures are easily confused with signal artifacts (e.g., eye movements) or benign variants (e.g., slowing). Commercially available systems suffer from unacceptably high false alarm rates. Deep learning algorithms that employ high dimensional models have not previously been effective due to the lack of big data resources. In this paper, we use the TUH EEG Seizure Corpus to evaluate a variety of hybrid deep structures including Convolutional Neural Networks and Long Short-Term Memory Networks. We introduce a novel recurrent convolutional architecture that delivers 30% sensitivity at 7 false alarms per 24 hours. We have also evaluated our system on a held-out evaluation set based on the Duke University Seizure Corpus and demonstrate that performance trends are similar to the TUH EEG Seizure Corpus. This is a significant finding because the Duke corpus was collected with different instrumentation and at different hospitals. Our work shows that deep learning architectures that integrate spatial and temporal contexts are critical to achieving state of the art performance and will enable a new generation of clinically-acceptable technology.
Meysam Golmohammadi, Saeedeh Ziyabari, Vinit Shah, Iyad Obeid, Joseph Picone
ICMLA5
2016 A Nonparametric Bayesian Approach for Spoken Term Detection by Example Query
abstract
State of the art speech recognition systems use data-intensive context-dependent phonemes as acoustic units. However, these approaches do not translate well to low resourced languages where large amounts of training data is not available. For such languages, automatic discovery of acoustic units is critical. In this paper, we demonstrate the application of nonparametric Bayesian models to acoustic unit discovery. We show that the discovered units are correlated with phonemes and therefore are linguistically meaningful. We also present a spoken term detection (STD) by example query algorithm based on these automatically learned units. We show that our proposed system produces a P@N of 61.2% and an EER of 13.95% on the TIMIT dataset. The improvement in the EER is 5% while P@N is only slightly lower than the best reported system in the literature.
Amir Hossein Harati Nejad Torbati, Joseph Picone
INTERSPEECH2
2016 A nonparametric Bayesian approach for automatic discovery of a lexicon and acoustic units
abstract
State of the art speech recognition systems use context-dependent phonemes as acoustic units. However, these approaches do not work well for low resourced languages where large amounts of training data or resources such as a lexicon are not available. For such languages, automatic discovery of acoustic units can be important. In this paper, we demonstrate the application of nonparametric Bayesian models to acoustic unit discovery. We show that the discovered units are linguistically meaningful. We also present a semi-supervised learning algorithm that uses a nonparametric Bayesian model to learn a mapping between words and acoustic units. We demonstrate that a speech recognition system using these discovered resources can approach the performance of a speech recognizer trained using resources developed by experts. We show that unsupervised discovery of acoustic units combined with semi-supervised discovery of the lexicon achieved performance (9.8% WER) comparable to other published high complexity systems. This nonparametric approach enables the rapid development of speech recognition systems in low resourced languages.
Amir Hossein Harati Nejad Torbati, Joseph Picone
SLT2
2016 A Doubly Hierarchical Dirichlet Process Hidden Markov Model with a Non-Ergodic Structure
abstract
Nonparametric Bayesian models use a Bayesian framework to learn model complexity automatically from the data, eliminating the need for a complex model selection process. A Hierarchical Dirichlet Process Hidden Markov Model (HDPHMM) is the nonparametric Bayesian equivalent of a hidden Markov model (HMM), but is restricted to an ergodic topology that uses a Dirichlet Process Model to achieve a mixture distribution-like model. For applications involving ordered sequences (e.g., speech recognition), it is desirable to impose a left-to-right structure on the model. In this paper, we introduce a model based on HDPHMM that: 1) shares data points between states, 2) models non-ergodic structures, and 3) models non-emitting states. The first point is particularly important because Gaussian mixture models, which support such sharing, have been very effective at modeling modalities in a signal (e.g., speaker variability). Further, sharing data points allows models to be estimated more accurately, an important consideration for applications such as speech recognition in which some mixture components occur infrequently. We demonstrate that this new model produces a 20% relative reduction in error rate for phoneme classification and an 18% relative reduction on a speech recognition task on the TIMIT Corpus compared to a baseline system consisting of a parametric HMM.
Amir Hossein Harati Nejad Torbati, Joseph Picone
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Speech acoustic unit segmentation using hierarchical dirichlet processes
Amir Hossein Harati Nejad Torbati, Joseph Picone, Marc Sobel
INTERSPEECH2
2012 Applications of Dirichlet Process Mixtures to speaker adaptation
abstract
Balancing unique acoustic characteristics of a speaker such as identity and accent, with general acoustic behavior that describes phoneme identity, is one of the great challenges in applying nonparametric Bayesian approaches to speaker adaptation. The Dirichlet Process Mixture (DPM) is a relatively new model that provides an elegant framework in which individual characteristics can be balanced with aggregate behavior without diluting the quality of the individual models. Unlike Gaussian Mixture models (GMMs), which tend to smear multimodal behavior through averaging, the DPM model attempts to preserve unique behaviors through use of an infinite mixture model. In this paper, we present some exploratory research on applying these models to the acoustic modeling component of the speaker adaptation problem. DPM based models are shown to provide up to 10% reduction in WER over maximum likelihood linear regression (MLLR) on a speaker adaptation task based on the Resource Management database.
Amir Hossein Harati Nejad Torbati, Joseph Picone, Marc Sobel
ICASSP2
2008 Nonlinear mixture autoregressive hidden Markov models for speech recognition
abstract
Gaussian mixture models are a very successful method for modeling the output distribution of a state in a hidden Markov model (HMM). However, this approach is limited by the assumption that the dynamics of speech features are linear and can be modeled with static features and their derivatives. In this paper, a nonlinear mixture autoregressive model is used to model state output distributions (MAR-HMM). Estimation of model parameters is extended to handle vector features. MAR-HMMs are shown to provide superior performance to comparable Gaussian mixture model-based HMMs (GMM-HMM) with lower complexity on two pilot classification tasks.
Sundararajan Srinivasan, Daniel May, Georgios Y. Lazarou, Joseph Picone
INTERSPEECH5
2007 Report on the NSF-sponsored Human Language Technology Workshop on Industrial Centers
Mary P. Harper, Alex Acero, Srinivas Bangalore, Jordan Cohen, Barbara Cuthill, Carol Y. Espy-Wilson, Christiane Fellbaum, John Garofolo, Chin-Hui Lee 0001, Jim Lester, Andrew McCallum, Nelson Morgan, Michael Picheney, Joseph Picone, Lance Ramshaw, Jeffrey C. Reynar, Hadar Shemtov, Clare Voss
MTSummit15
2007 A cluster-based power-efficient MAC scheme for event-driven sensing applications
Georgios Y. Lazarou, Joseph Picone
Ad Hoc Networks3
2006 Nonlinear dynamical invariants for speech recognition
Sundararajan Srinivasan, M. Pannuri, Georgios Y. Lazarou, Joseph Picone
INTERSPEECH5
2005 Effects of displayless navigational interfaces on user prosodics
Julie Baca, Joseph Picone
Speech Commun.2
2004 Effects on transcription errors on supervised learning in speech recognition
abstract
Hidden Markov model-based speech recognition systems use supervised learning to train acoustic models. On difficult tasks such as conversational speech, there has been concern over the impact erroneous transcriptions have on the parameter estimation process. This work analyzes the effects of mislabeled data on recognition accuracy. Training is performed using manually corrupted transcriptions, and results are presented on three tasks: TIdigits, alphadigits and switchboard. For alphadigits, with 16% of the training data mislabeled, the performance of the system degrades by 12% relative to the baseline. On switchboard, at 16% mislabeled training data, the performance of the system degrades by 8.5% relative to the baseline. An analysis of these results revealed that the Gaussian mixture model contributes significantly to the robustness of the supervised learning training process.
Ram Sundaram, Joseph Picone
ICASSP (1)2
2003 Dialog systems for automotive environments
Julie Baca, Hualin Gao, Joseph Picone
INTERSPEECH4
2003 Analysis of the Aurora large vocabulary evaluations
Naveen Parihar, Joseph Picone
INTERSPEECH2
2002 Open-Domain Voice-Activated Question Answering
Sanda M. Harabagiu, Dan I. Moldovan, Joseph Picone
COLING3
2002 A sparse modeling approach to speech recognition based on relevance vector machines
Jonathan Hamaker, Joseph Picone, Aravind Ganapathiraju
INTERSPEECH2
2001 Syllable-based large vocabulary continuous speech recognition
abstract
Most large vocabulary continuous speech recognition (LVCSR) systems in the past decade have used a context-dependent (CD) phone as the fundamental acoustic unit. We present one of the first robust LVCSR systems that uses a syllable-level acoustic unit for LVCSR on telephone-bandwidth speech. This effort is motivated by the inherent limitations in phone-based approaches-namely the lack of an easy and efficient way for modeling long-term temporal dependencies. A syllable unit spans a longer time frame, typically three phones, thereby offering a more parsimonious framework for modeling pronunciation variation in spontaneous speech. We present encouraging results which show that a syllable-based system exceeds the performance of a comparable triphone system both in terms of word error rate (WER) and complexity. The WER of the best syllabic system reported here is 49.1% on a standard Switchboard evaluation, a small improvement over the triphone system. We also report results on a much smaller recognition task, OGI Alphadigits, which was used to validate some of the benefits syllables offer over triphones. The syllable-based system exceeds the performance of the triphone system by nearly 20%, an impressive accomplishment since the alphadigits application consists mostly of phone-level minimal pair distinctions.
Aravind Ganapathiraju, Jonathan Hamaker, Joseph Picone, Mark Ordowski, George R. Doddington
IEEE Trans. Speech Audio Process.3
2000 Towards language independent acoustic modeling
abstract
We describe procedures and experimental results using speech from diverse source languages to build an ASR system for a single target language. This work is intended to improve ASR in languages for which large amounts of training data are not available. We have developed both knowledge-based and automatic methods to map phonetic units from the source languages to the target language. We employed HMM adaptation techniques and discriminative model combination to combine acoustic models from the individual source languages for recognition of speech in the target language. Experiments are described in which Czech Broadcast News is transcribed using acoustic models trained from small amounts of Czech read speech augmented by English, Spanish, Russian, and Mandarin acoustic models.
William J. Byrne, Peter Beyerlein, Juan M. Huerta, Sanjeev Khudanpur, B. Marthi, John Morgan, Nino Peterek, Joseph Picone, Dimitra Vergyri
ICASSP8
2000 Hybrid SVM/HMM architectures for speech recognition
abstract
In this paper, we describe the use of a powerful machine learning scheme, Support Vector Machines (SVM), within the framework of hidden Markov model (HMM) based speech recognition. The hybrid SVM/HMM system has been developed based on our public domain toolkit. The hybrid system has been evaluated on the OGI Alphadigits corpus and performs at 11.6% WER, as compared to 12.7% with a triphone mixture-Gaussian HMM system, while using only a fifth of the training data used by triphone system. Several important issues that arise out of the nature of SVM classifiers have been addressed. We are in the process of migrating this technology to large vocabulary recognition tasks like SWITCHBOARD. 1. INTRODUCTION Speech recogn i t i on can be v i ewed as a pa t t ern recognition problem where we desire each unique sound t o be d i s t i ngu i shab l e f r om a l l o t he r sounds . Traditionally statistical models, such as Gaussian mixture models, have been used to "represent" th...
Aravind Ganapathiraju, Jonathan Hamaker, Joseph Picone
INTERSPEECH3
2000 Support vector machines for automatic data cleanup
Aravind Ganapathiraju, Joseph Picone
INTERSPEECH2
1999 Initial evaluation of hidden dynamic models on conversational speech
abstract
Conversational speech recognition is a challenging problem primarily because speakers rarely fully articulate sounds. A successful speech recognition approach must infer intended spectral targets from the speech data, or develop a method of dealing with large variances in the data. Hidden dynamic models (HDMs) attempt to automatically learn such targets in a hidden feature space using models that integrate linguistic information with constrained temporal trajectory models. HDMs are a radical departure from conventional hidden Markov models (HMMs), which simply account for variation in the observed data. We present an initial evaluation of such models on a conversational speech recognition task involving a subset of the SWITCHBOARD corpus. We show that in an N-best rescoring paradigm, HDMs are capable of delivering performance competitive with HMMs.
Joseph Picone, Sandi Pike, Roland Reagan, Terri Kamm, John S. Bridle, Z. Ma, Hywel B. Richards, Mike Schuster
ICASSP1
1999 A public domain speech-to-text system
abstract
The lack of freely available state-of-the-art Speech-to-Text (STT) software has been a major hindrance to the development of new audio information processing technology. The high cost of the infrastructure required to conduct state-of-the-art speech recognition research prevents many small research groups from evaluating new ideas on large-scale tasks. In this paper, we present the core components of an available state-of-the-art STT system: an acoustic processor which converts the speech signal into a sequence of feature vectors; a training module which estimates the parameters for a Hidden Markov Model; a linguistic processor which predicts the next word given a sequence of previously recognized words; and a search engine which finds the most probable word sequence given a set of feature vectors. 1.
Mark Ordowski, Neeraj Deshmukh, Aravind Ganapathiraju, Jonathan Hamaker, Joseph Picone
EUROSPEECH5
1999 Kanji-to-Hiragana conversion based on a length-constrained n-gram analysis
abstract
A common problem in speech processing is the conversion of the written form of a language to a set of phonetic symbols representing the pronunciation. In this paper, we focus on an aspect of this problem specific to the Japanese language. Written Japanese consists of a mixture of three types of symbols: Kanji, Hiragana, and Katakana. We describe an algorithm for converting conventional Japanese orthography to a Hiragana-like symbol set that closely approximates the most common pronunciation of the text. The algorithm is based on two hypotheses: (1) the correct reading of a Kanji character can be determined by examining a small number of adjacent characters and (2) the number of such combinations required in a dictionary is manageable. The algorithm described here converts the input test by selecting the most probable sequence of orthographic units (n-grams) that can be concatenated to form the input text. In closed-set testing, the n-gram algorithm was shown to provide better performance than several public domain algorithms, achieving a sentence error rate of 3% on a wide range of text material. Though the focus of this paper is written Japanese, the pattern matching algorithm described here has applications to similar problems in other languages.
Joseph Picone, Tom Staples, Kazuhiro Kondo, Nozomi Arai
IEEE Trans. Speech Audio Process.1
1998 Advances in alphadigit recognition using syllables
abstract
We present a set of experiments which explore the use of syllables for recognition of continuous alphadigit utterances. In this system, syllables are used as the primary unit of recognition. This work was motivated by our need to verify and isolate phenomena seen when performing syllable-based experiments on the Switchboard corpus. The performance of our base syllable system is better than a crossword triphone system while requiring a small portion of the resources necessary for triphone systems. All experiments were performed on the OGI Alphadigits corpus, which consists of telephone-bandwidth alphadigit strings. The word error rate (WER) of the best syllable system (context-independent syllables) reported here is 11.1% compared to 12.2% for a crossword triphone system.
Jonathan Hamaker, Aravind Ganapathiraju, Joseph Picone, John J. Godfrey
ICASSP3
1998 Visualization of signal processing concepts
abstract
One of the key difficulties in a Signals and Systems course is visualization of the mathematically complex concepts presented. Thus, there is a need for graphical tools which enhance the students' comprehension of these difficult concepts by allowing interactive learning. We present a software package to assist in the explanation and visualization of signal processing concepts for an educational environment. We provide a set of Java-based tools for understanding the concepts of convolution, spectral analysis, and pole/zero system response. The wide availability and platform-independence provided by Java make this tool highly portable and easily accessible to a broader audience of students than comparable systems based on Matlab or other commercial software. The software described in this paper is available in the public domain at our website: http://isip.msstate.edu/.
Janna Shaffer, Jonathan Hamaker, Joseph Picone
ICASSP3
1998 Resegmentation of SWITCHBOARD
abstract
The SWITCHBOARD (SWB) corpus is one of the most important benchmarks for recognition tasks involving large vocabulary conversational speech (LVCSR). The high error rates on SWB are largely attributable to an acoustic model mismatch, the high frequency of poorly articulated monosyllabic words, and large variations in pronunciations. It is imperative to improve the quality of segmentations and transcriptions of the training data to achieve better acoustic modeling. By adapting existing acoustic models to only a small subset of such improved transcriptions, we have achieved a 2% absolute improvement in performance.
Neeraj Deshmukh, Aravind Ganapathiraju, Andi Gleeson, Jonathan Hamaker, Joseph Picone
ICSLP5
1998 Support vector machines for speech recognition
abstract
Hidden Markov models (HMM) with Gaussian mixture observation densities are the dominant approach in speech recognition. These systems typically use a representational model for acoustic modeling which can often be prone to overfitting and does not translate to improved discrimination. We propose a new paradigm centered on principles of structural risk minimization using a discriminative framework for speech recognition based on support vector machines (SVMs). SVMs have the ability to simultaneously optimize the representational and discriminative ability of the acoustic classifiers. We have developed the first SVM-based large vocabulary speech recognition system that improves performance over traditional HMM-based systems. This hybrid system achieves a state-of-the-art word error rate of 10.6% on a continuous alphadigit task—a 10% improvement relative to an HMM system. On SWITCHBOARD, a large vocabulary task, the system improves performance over a traditional HMM system from 41.6% word error rate to 40.6%. This dissertation discusses several practical issues that arise when SVMs are incorporated into the hybrid system.
Aravind Ganapathiraju, Jonathan Hamaker, Joseph Picone
ICSLP3
1998 Information theoretic approaches to model selection
abstract
The p r imary p rob l em in l a rge vocabu l a ry conversational speech recognition (LVCSR) is poor acoustic-level matching due to large variability in pronunciations. There is much to explore about the “quality” of states in an HMM and the interrelationships between inter-state and intra-state Gaussians used to model speech. Of particular interest is the variable discriminating power of the individual states. The fundamental concept addressed in this paper is to investigate means of exploiting such dependencies through model topology optimization based on the Bayesian Information Criterion (BIC) and the Minimum Description Length (MDL) principle.
Jonathan Hamaker, Aravind Ganapathiraju, Joseph Picone
ICSLP3
1998 Improved surname pronunciations using decision trees
Julie Ngan, Aravind Ganapathiraju, Joseph Picone
ICSLP3
1997 An advanced system to generate pronunciations of proper nouns
abstract
Accurate recognition of proper nouns is a critical component of automatic speech recognition (ASR). Since there are no obvious letter-to-sound conversion rules that govern the pronunciation of any large set of proper nouns, this is an open-ended problem that evolves constantly under various sociolinguistic influences. A Boltzmann machine neural network is well-suited for the task of generating the most likely pronunciations of a proper noun. This pronunciation output can be used to build better acoustic models for the noun that result in improved recognition performance. We present an advanced version of this N-best pronunciations system; and a multiple pronunciations dictionary of 18000 surnames and 25000 pronunciations used as a training database. The database and software are available in the public domain.
Neeraj Deshmukh, Julie Ngan, Jonathan Hamaker, Joseph Picone
ICASSP4
1997 Microsegment-based connected digit recognition
abstract
By building acoustic phonetic models which explicitly represent as much knowledge of pronunciation in a small domain (the digits) as possible, we can create a recognition system which not only performs well but allows for meaningful error analysis and improvement. An HMM-based recognizer for the digits and a few associated words was constructed in accord with these principles. About 65 phonetic models were trained on 140 carefully labeled utterances, then iteratively trained on unlabeled data under orthographic supervision. The basic system achieved less than 3% word error rate on digit strings of unknown length from unseen test speakers, and 1.4% on 7-digit strings of known length. This is competitive with word-based models using the same HMM engine and similar parameter settings. As an R&D system, it allows meaningful analysis of errors and relatively straightforward means of improvement.
John J. Godfrey, Aravind Ganapathiraju, Coimbatore S. Ramalingam, Joseph Picone
ICASSP4
1996 Automated generation of N-best pronunciations of proper nouns
abstract
The problem of proper noun recognition is key to developing pervasive voice interfaces in applications such as directory assistance and data entry for telecommunications. Recognition of such words requires an ability to generate reasonably accurate pronunciation networks. This is a very challenging problem because a large percentage of proper nouns, such as personal names, appear to have no obvious (or simple) letter to mapping rules that can be used to generate the pronunciations. It appears to be an open-ended problem that is constantly evolving as a function of numerous sociological factors. Yet humans do amazingly well at generating and recognizing the pronunciation of a name never encountered before. We present an algorithm based on a Boltzmann machine type of neural network that generates the most likely pronunciations of a proper noun from the text-only spellings of the name. This method does not require voice data containing the spelling or nominal pronunciation.
Neeraj Deshmukh, Mary Weber, Joseph Picone
ICASSP3
1996 Benchmarking human performance for continuous speech recognition
Neeraj Deshmukh, Richard Duncan, Aravind Ganapathiraju, Joseph Picone
ICSLP4
1995 Voice across Hispanic America: a telephone speech corpus of American Spanish
abstract
As part of the Polyphone project, Texas Instruments is in the process of collecting and developing a corpus of telephone speech in American Spanish. The corpus, called Voice Across Hispanic America (VAHA), will attempt to provide balanced phonetic coverage of the language, in addition to containing widely used vocabulary items such as digits, letter strings, yes/no responses, proper names, and selected command words and phrases used in automated telephone service applications. The speakers are native speakers of Spanish living in the United States. The collection and development of the corpus is expected to be completed by June 1995. So far, the authors have collected about 500 speakers from various parts of the U.S. They describe the design issues in various aspects of the project, such as subject recruitment, corpus and prompt sheet design, the data acquisition system, and validation and transcription. They conclude with a brief statistical profile of the data collected.
Yeshwant K. Muthusamy, Edward Holliman, Barbara Wheatley, Joseph Picone, John J. Godfrey
ICASSP4
1994 A comparative analysis of Japanese and English digit recognition
abstract
This paper presents initial results of comparisons between fluently spoken Japanese and English on a common task: speaker independent digit recognition with applications in voice dialing. The complexity of this task across these languages is comparable in terms of lexicon size and perplexity of the language model. The English lexicon contained 11 words, and the Japanese lexicon contained 13 words. The durations of the words, as well as phones proved to be longer and have greater variation in English than in Japanese. An analysis of several key recognition parameters, namely the frame duration, LPC order, and feature vector dimensionality are also included. None of the above parameters seems to show language dependency in our test.>
Kazuhiro Kondo, Joseph Picone, Barbara Wheatley
ICASSP (1)2
1994 The voice across Japan database-the Japanese language contribution to Polyphone
abstract
Texas Instruments' Voice Across Japan (VAJ) database, modeled after the highly successful Voice Across America project, consists of a wide range of diverse speech material including digit strings, yes/no questions, and phonetically-rich read sentences. The data is being collected using long distance telephone lines and an analog telephone interface. The target size is 14 items per speaker by 10,000 speakers. Greater emphasis is being placed on the collection of phonetically-rich read sentence data. Four randomly selected sentences are included in each session: one from the 512 sentence ATR PB set, and three from a 10,000 sentence set developed specifically for this project. This latter sentence set, designed to maximize the triphone coverage of the database, is described. The VAJ database is planned to be included in the Linguistic Data Consortium's (LDC) Polyphone (multi-language) database.>
Thomas Staples, Joseph Picone, Nozomi Arai
ICASSP (1)2
1993 Managing software complexity in signal processing research
Joseph Picone
ICASSP (3)1
1990 Robust pitch determination via SVD based cepstral methods
abstract
A greatly enhanced cepstral-based pitch estimator which uses the MUSIC algorithm for estimating background noise characteristics is presented. This approach couples the signal enhancement capabilities of MUSIC, which is based on singular value decomposition (SVD) orthogonalization, with the harmonic spectrum estimation capabilities of the cepstrum. Marked improvements are demonstrated over standard fast Fourier transform (FFT)-based cepstral processing. Objective performance evaluations show increased frequency determination capability at low signal-to-noise ratios.>
M. Scott Andrews, Joseph Picone, Ronald D. DeGroat
ICASSP2
1990 The demographics of speaker independent digit recognition
abstract
A database designed to provide a statistically significant model of the demographics of the continental US population is presented. Recognition performance on a simple digit recognition task is analyzed and shown to be most highly correlated with signal-to-noise ratio, dialect, and age. Several other demographic features, including sex, income level, education level, and market size, are found to have little correlation with recognition error rate.>
Joseph Picone
ICASSP1
1990 Duration in context clustering for speech recognition
Joseph Picone
Speech Commun.1
1989 Speech recognition in a unification grammar framework
abstract
The authors describe a stochastic unification grammar system that is a generalization of the conventional hidden Markov model (HMM) approach. Unification grammars concisely model context, providing a more powerful characterization of the acoustic data than the first-order Markov process. It is shown that this approach generalizes traditional FSA (finite-state automaton)-based HMM systems and that a stochastic chart parsing algorithm produces the exact same solutions as an existing FSA-based system. The shift from automata to grammars allows efficient processing of complex language models by hypothesizing symbols once per frame, no matter how many times they are needed. As an added benefit, the chart parsing algorithm allows parallel processing of lower level hypotheses autonomously with no fundamental algorithm changes.>
Charles T. Hemphill, Joseph Picone
ICASSP2
1989 On modeling duration in context in speech recognition
abstract
A clustering algorithm is introduced that allows clustering of HMM (hidden Markov models) models directly. This clustering algorithm determines the appropriate duration profile for a recognition unit. High-performance speaker-independent digit recognition on a studio-quality connected-digit database is demonstrated using this algorithm.>
Joseph Picone
ICASSP1
1989 A phonetic vocoder
abstract
The authors study the problem of coding spectral information in speech at bit rates in the range of 100-400 b/s using speaker-independent phone-based recognition. Spectral information is coded as a sequence of phonetic events and a sequence of transitions through the corresponding hidden Markov model (HMM)-based phone models. This simple phonetic speech-coding system has been shown to be a promising approach. A simple inventory of phonemes is sufficient for capturing the bulk of the acoustic information.>
Joseph Picone, George R. Doddington
ICASSP1
1988 Enhancing the performance of speech recognition with echo cancellation
abstract
The use of echo cancellation to improve speech recognition performance over telephone channels is described. Echo cancellation is shown to provide an increase of 25 dB in signal-to-noise ratio, thereby increasing recognition performance to a level which can be attained over telephone channels with no echo. A prototype system which includes the echo canceller and an isolated-word speaker-independent speech recognizer has been implemented within a single AT&T WE DSP-32.>
Joseph Picone, M. A. Johnson, Walter T. Hartwell
ICASSP1
1988 Design and implementation of a robust pitch detector based on a parallel processing technique
abstract
The design and implementation of a parallel-processing-based pitch detector is presented. Pitch information is extracted by performing pitch detection on four different waveforms derived from the speech signal. Pitch information from the four pitch-detection processes is then combined to determine a final pitch estimate. The performance of this pitch detector is evaluated on a large database and compared to other well-known pitch detection algorithms. It has been implemented in real time on a TMS32020 fixed-point digital signal processor as part of a 2.4 kb/s vocoder. A performance comparison of the real-time fixed-point implementation and a computer simulation are also given. The results show that the pitch detector performance is maintained in the real-time implementation mainly because the majority of the algorithm computations are integer arithmetic and logic-type operations.>
Rafid A. Sukkar, Joseph L. LoCicero, Joseph Picone
IEEE J. Sel. Areas Commun.3
1987 Harmonic coding of speech at 4.8 kb/s
abstract
This paper describes a new speech coding technique which yields improved speech quality over existing 2.4 kb/s LPC vocoders. The method is computationally efficient and operates at a data rate of 4.8 kb/s. Each speech frame is initially classified as voiced or unvoiced. Unvoiced frames are synthesized using a linear predictive coding filter with noise or multipulse excitation. Voiced frames are synthesized using a sum of sinusoids. The frequency of each sinusoid is defined by peaks in the frequency spectrum. A new interpolation technique provides a computationally efficient method of locating the spectral peaks. A real-time, fully quantized version has been implemented in hardware.
Edward C. Bronson, Douglas A. Carlone, W. Bastiaan Kleijn, Kevin M. O'Dell, Joseph Picone, David L. Thomson
ICASSP5
1987 Low rate speech coding using contour quantization
abstract
Vector quantization-based approaches to speech coding have generated new interest in very low bit rate speech coding, that is, speech coded to bit rates below 1200 bits/sec. To achieve such low bit rates, it is necessary to quantize the pitch and energy parameters at rates below 100 bits/sec. Contour quantization is introduced as a technique in which the contour of a given parameter is normalized by a nominal value and vector quantized. Contour quantization is shown to be extremely robust and efficient in encoding the pitch and energy parameters of the LPC vocoder. In this paper, a low rate speech coding system which uses contour quantization to encode the LPC excitation is presented. The system is a fixed bit rate system which is intended to operate at bit rates ranging from 400 bits/s to 800 bits/s. The overall system delay varies from 300 ms at 800 bits/s to 400 ms at 400 bits/s. At 800 bits/s, the system achieved a score of 89 on a three male speaker DRT, and a score of 81 on a three female speaker DRT.
Joseph Picone, George R. Doddington
ICASSP1
1987 Robust pitch detection in a noisy telephone environment
abstract
While many readily available pitch tracking algorithms are capable of accurately tracking pitch on studio quality speech data, robust performance in real operational environments is still an elusive goal. In this paper, three pitch detection algorithms are evaluated over a database consisting of speech data collected over a wide range of telephone lines including long distance exchanges. The speech material contained in the database consists of excerpts from typical telephone conversations, collected at the receiving end of a two party exchange. Subjective and objective evaluations were conducted on three pitch tracking algorithms: an improved version of the Integrated Correlation[1,2] pitch tracker, the Gold-Rabiner [3,4] parallel processing algorithm, and the NSA LPC-10 DYPTRACK Version 43 [5-8] algorithm. A comparative analysis of these algorithms indicates that the integrated correlation pitch tracker provides significantly better performance, mainly due to its ability to make accurate voicing decisions in noisy environments. In addition, intelligibility tests demonstrate that synthetic speech intelligibility is correlated with the degree of accuracy of pitch estimation. This result reinforces our belief that accurate pitch tracking is crucial to the operational acceptance of the speech quality produced by the LPC pitch-excited vocoder.
Joseph Picone, George R. Doddington, Bruce G. Secrest
ICASSP1
1986 Fast and accurate pitch detection using pattern recognition and adaptive time-domain analysis
abstract
A method of determining pitch and voicing information from speech signals is presented. The algorithm, which employs time-domain analysis and pattern recognition techniques, is fast and yields accurate pitch and voicing estimates. A search routine is employed to find periodicity in each of four signals derived from the speech waveform and the results are combined to form a pitch estimate. The voicing decision uses linear discriminant analysis, and declares speech frames voiced or unvoiced based on a weighted sum of 13 parameters. Performance comparisons with other pitch detectors are reported.
Dimitrios P. Prezas, Joseph Picone, David L. Thomson
ICASSP2
1986 Recognition of speech under stress and in noise
abstract
Speech recognizers trained in one condition but operating in a different condition degrade in performance. Typical of this situation is when the recognizer is trained under normal conditions but operated in a stressful and noisy environment as in military applications. This paper reports on recognition experiments conducted with a "simulated stress" data base using a baseline algorithm and its modifications. These algorithms perform acceptably well (1 % substitution rate) for a vocabulary of 105 words under normal conditions, but degrade by an order of magnitude under the "stress" conditions. The experiments also show that the speech production variation caused by noise exposure at the ear is far more deleterious than ambient acoustic noise with a noise cancelling microphone.
Periagaram K. Rajasekaran, George R. Doddington, Joseph Picone
ICASSP3
1986 Joint estimation of the LPC parameters and the multi-pulse excitation
Joseph Picone, Dimitrios P. Prezas, Walter T. Hartwell, Joseph L. LoCicero
Speech Commun.1
1985 Acoustic waveguides applications to oceanic science
Joseph Picone
Proc. IEEE1
1982 An Application of the Barrett-Lampard Expansion to Obtain the Slope Characteristics of a Speech Signal
abstract
Many methods of speech waveform encoding utilize differential techniques for bandwidth compression. The channel bit rate of these systems is directly related to the slope characteristics of the speech signal. Employing well-known statistical characteristics of speech, this paper presents a derivation of the probability density function (pdf) for the differential slope of a speech signal and evaluates two families of pdfs. The technique used here, based on the Barrett-Lampard expansion, can also be applied to the general problem of obtaining the pdf of a random signal that has been digitally filtered. In this way, the frequency reliance upon a Gaussian assumption can be avoided.
Joseph Picone, Joseph L. LoCicero
IEEE Trans. Commun.1