EDBT 2026 Demo / reviewers in the wild / expert
Joseph P. Campbell
dblp:39/2844
· DBLP profile ↗
40ranked-venue papers
6as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-authorArtificial intelligence and machine learning · 23 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Speech recognition and synthesis · 95% Kernel, tree and ensemble methods · 5% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.2 | 1 | 2014 | Characterizing Phonetic Transformations and Acoustic Differences Across English Dialects · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Natural language and speech › Speech recognition and synthesis
speech analysis |
0.2 | 1 | 2014 | Characterizing Phonetic Transformations and Acoustic Differences Across English Dialects · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Natural language and speech › Speech recognition and synthesis
speaker recognition |
0.1 | 3 | 2007 | Speaker Verification Using Support Vector Machines and High-Level Features · IEEE Trans. Speech Audio Process. 2007 Phonetic Speaker Recognition with Support Vector Machines · NIPS 2003 Speaker recognition: a tutorial · Proc. IEEE 1997 |
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker verification |
0.1 | 2 | 2007 | Speaker Verification Using Support Vector Machines and High-Level Features · IEEE Trans. Speech Audio Process. 2007 Speaker recognition: a tutorial · Proc. IEEE 1997 |
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding |
0.1 | 1 | 2006 | Exploiting nonacoustic sensors for speech encoding · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.1 | 1 | 2006 | Exploiting nonacoustic sensors for speech encoding · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Speech recognition and synthesis
speech enhancement |
0.1 | 1 | 2006 | Exploiting nonacoustic sensors for speech encoding · IEEE Trans. Speech Audio Process. 2006 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.0 | 1 | 2003 | Phonetic Speaker Recognition with Support Vector Machines · NIPS 2003 |
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker identification |
0.0 | 1 | 1997 | Speaker recognition: a tutorial · Proc. IEEE 1997 |
Biometric security
biometric recognition |
0.0 | 1 | 1997 | Speaker recognition: a tutorial · Proc. IEEE 1997 |
Biometric security
speaker recognition |
0.0 | 1 | 1997 | Speaker recognition: a tutorial · Proc. IEEE 1997 |
Methods — techniques the papers use, named apart from their topics
insertion and deletion transformations · 0.2hidden markov model · 0.2skin vibration sensor · 0.1microwave radar · 0.1bone conduction sensor · 0.1MELPe coder · 0.1support vector machine · 0.1n-gram modeling · 0.1log likelihood ratio kernel · 0.1term frequency analysis · 0.0speech processing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Speaker Comparison for Forensic and Investigative Applications II
Jean-François Bonastre, Joseph P. Campbell, Anders P. Eriksson, Hirotaka Nakasone, Reva Schwartz |
INTERSPEECH | 2 |
| 2016 | Corpora for the Evaluation of Robust Speaker Recognition Systems
Douglas E. Sturim, Pedro A. Torres-Carrasquillo, Joseph P. Campbell |
INTERSPEECH | 3 |
| 2014 | Characterizing Phonetic Transformations and Acoustic Differences Across English DialectsabstractIn this work, we propose a framework that automatically discovers dialect-specific phonetic rules. These rules characterize when certain phonetic or acoustic transformations occur across dialects. To explicitly characterize these dialect-specific rules, we adapt the conventional hidden Markov model to handle insertion and deletion transformations. The proposed framework is able to convert pronunciation of one dialect to another using learned rules, recognize dialects using learned rules, retrieve dialect-specific regions, and refine linguistic rules. Potential applications of our proposed framework include computer-assisted language learning, sociolinguistics, and diagnosis tools for phonological disorders. Nancy F. Chen, Sharon W. Tam, Wade Shen, Joseph P. Campbell |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2012 | Analyzing and Interpreting Automatically Learned Rules Across DialectsabstractIn this paper, we demonstrate how informative dialect recogni-tion systems such as acoustic pronunciation model (APM) help speech scientists locate and analyze phonetic rules efficiently. In particular, we analyze dialect-specific characteristics auto-matically learned from APM across two American English di-alects. We show that unsupervised rule retrieval performs sim-ilarly to supervised retrieval, indicating that APM is useful in practical applications, where word transcripts are often unavail-able. We also demonstrate that the top-ranking rules learned from APM generally correspond to the linguistic literature, and can even pinpoint potential research directions to refine existing knowledge. Thus, the APM system can help phoneticians ana-lyze rules efficiently by characterizing large amounts of data to postulate rule candidates, so they can reserve time to conduct more targeted investigations. Potential applications of informa-tive dialect recognition systems include forensic phonetics and diagnosis of spoken language disorders. Index Terms: informative dialect recognition, rule retrieval, phonological rules, forensic phonetics Nancy F. Chen, Wade Shen, Joseph P. Campbell |
INTERSPEECH | 3 |
| 2011 | Informative dialect recognition using context-dependent pronunciation modelingabstractWe propose an informative dialect recognition system that learns phonetic transformation rules, and uses them to identify dialects. A hidden Markov model is used to align reference phones with dialect specific pronunciations to characterize when and how often substitutions, insertions, and deletions occur. Decision tree clustering is used to find context-dependent phonetic rules. We ran recognition tasks on 4 Arabic dialects. Not only do the proposed systems perform well on their own, but when fused with baselines they improve performance by 21-36% relative. In addition, our proposed decision-tree system beats the baseline monophone system in recovering phonetic rules by 21% relative. Pronunciation rules learned by our proposed system quantify the occurrence frequency of known rules, and suggest rule candidates for further linguistic studies. Nancy F. Chen, Wade Shen, Joseph P. Campbell, Pedro A. Torres-Carrasquillo |
ICASSP | 3 |
| 2011 | USSS-MITLL 2010 human assisted speaker recognitionabstractThe United States Secret Service (USSS) teamed with MIT Lincoln Laboratory (MIT/LL) in the US National Institute of Standards and Technology's 2010 Speaker Recognition Evaluation of Human Assisted Speaker Recognition (HASR). We describe our qualitative and automatic speaker comparison processes and our fusion of these processes, which are adapted from USSS casework. The USSS-MIT/LL 2010 HASR results are presented. We also present post-evaluation results. The results are encouraging within the resolving power of the evaluation, which was limited to enable reasonable levels of human effort. Future ideas and efforts are discussed, including new features and capitalizing on naive listeners. Reva Schwartz, Joseph P. Campbell, Wade Shen, Douglas E. Sturim, William M. Campbell, Fred Richardson, Robert B. Dunn, Robert Granville |
ICASSP | 2 |
| 2011 | Assessing the speaker recognition performance of naive listeners using mechanical turkabstractIn this paper we attempt to quantify the ability of naive listeners to perform speaker recognition in the context of the NIST evaluation task. We describe our protocol: a series of listening experiments using large numbers of naive listeners (432) on Amazon's Mechanical Turk that attempts to measure the ability of the average human listener to perform speaker recognition. Our goal was to compare the performance of the average human listener to both forensic experts and state-of-the-art automatic systems. We show that naive listeners vary substantially in their performance, but that an aggregation of listener responses can achieve performance similar to that of expert forensic examiners. Wade Shen, Joseph P. Campbell, Derek Straub, Reva Schwartz |
ICASSP | 2 |
| 2011 | Characterizing Deletion Transformations Across Dialects Using a Sophisticated Tying MechanismabstractIn this work, we propose extensions of our Phone-based Pronunciation Model (PPM) for analyzing dialect differences. We compared these systems using 3 metrics and 2 datasets of English. Empirical results suggest that (1) sophisticated tying is suitable in modeling deletion transformations across dialects, beating standard tying by 33% relative, and (2) APM (Acousticbased Pronunciation Model) improves performance in generating dialect-specific pronunciations, dialect identification and rule retrieval, achieving relative gains beyond 34%. Index Terms: acoustic model, pronunciation model, phonetic rule, phonetic transformation, dialect recognition Nancy F. Chen, Wade Shen, Joseph P. Campbell |
INTERSPEECH | 3 |
| 2010 | A linguistically-informative approach to dialect recognition using dialect-discriminating context-dependent phonetic modelsabstractWe propose supervised and unsupervised learning algorithms to extract dialect discriminating phonetic rules and use these rules to adapt biphones to identify dialects. Despite many challenges (e.g., sub-dialect issues and no word transcriptions), we discovered dialect discriminating biphones compatible with the linguistic literature, while outperforming a baseline monophone system by 7.5% (relative). Our proposed dialect discriminating biphone system achieves similar performance to a baseline all-biphone system despite using 25% fewer biphone models. In addition, our system complements PRLM (Phone Recognition followed by Language Modeling), verified by obtaining relative gains of 15-29% when fused with PRLM. Our work is an encouraging first step towards a linguistically-informative dialect recognition system, with potential applications in forensic phonetics, accent training, and language learning. Nancy F. Chen, Wade Shen, Joseph P. Campbell |
ICASSP | 3 |
| 2010 | Transcript-dependent speaker recognition using mixer 1 and 2abstractTranscript-dependent speaker-recognition experiments are performed with the Mixer 1 and 2 read-transcription corpus using the Lincoln Laboratory speaker recognition system. Our analysis shows how widely speaker-recognition performance can vary on transcript-dependent data compared to conversational data of the same durations, given enrollment data from the same spontaneous conversational speech. A description of the techniques used to deal with the unaudited data in order to create 171 male and 198 female text-dependent experiments from the Mixer 1 and 2 read transcription corpus is given. Fred Richardson, Joseph P. Campbell |
INTERSPEECH | 2 |
| 2009 | Large-scale analysis of formant frequency estimation variability in conversational telephone speechabstractWe quantify how the telephone channel and regional dialect influence formant estimates extracted from Wavesurfer [1, 2] in spontaneous conversational speech from over 3,600 native American English speakers. To the best of our knowledge, this is the largest scale study on this topic. We found that F1 estimates are higher in cellular channels than those in landline, while F2 in general shows an opposite trend. We also characterized vowel shift trends in northern states in U.S.A. and compared them with the Northern city chain shift (NCCS) [3]. Our analysis is useful in forensic applications where it is important to distinguish between speaker, dialect, and channel characterisitcs. Index Terms: formant frequency, Northern city chain shift, spontaneous conversational speech, telephone channel, American Nancy F. Chen, Wade Shen, Joseph P. Campbell, Reva Schwartz |
INTERSPEECH | 3 |
| 2008 | Bridging the Gap between Linguists and Technology Developers: Large-Scale, Sociolinguistic Annotation for Dialect and Speaker Recognition
Christopher Cieri, Stephanie M. Strassel, Meghan Lammie Glenn, Reva Schwartz, Wade Shen, Joseph P. Campbell |
LREC | 6 |
| 2007 | Construction of a phonotactic dialect corpus using semiautomatic annotationabstractAbstract : In this paper, we discuss rapid, semiautomatic annotation techniques of detailed phonological phenomena for large corpora. We describe the use of these techniques for the development of a corpus of American English dialects. The resulting annotations and corpora will support both large scale linguistic dialect analysis and automatic dialect identification. We delineate the semiautomatic annotation process that we are currently employing and, a set of experiments we ran to validate this process. From these experiments, we learned that the use of ASR techniques could significantly increase the throughput and consistency of human annotators. Reva Schwartz, Wade Shen, Joseph P. Campbell, Shelley Paget, Julie Vonwiller, Dominique Estival, Christopher Cieri |
INTERSPEECH | 3 |
| 2007 | Speaker Verification Using Support Vector Machines and High-Level FeaturesabstractHigh-level characteristics such as word usage, pronunciation, phonotactics, prosody, etc., have seen a resurgence for automatic speaker recognition over the last several years. With the availability of many conversation sides per speaker in current corpora, high-level systems now have the amount of data needed to sufficiently characterize a speaker. Although a significant amount of work has been done in finding novel high-level features, less work has been done on modeling these features. We describe a method of speaker modeling based upon support vector machines. Current high-level feature extraction produces sequences or lattices of tokens for a given conversation side. These sequences can be converted to counts and then frequencies ofn-gram for a given conversation side. We use support vector machine modeling of these n-gram frequencies for speaker verification. We derive a new kernel based upon linearizing a log likelihood ratio scoring system. Generalizations of this method are shown to produce excellent results on a variety of high-level features. We demonstrate that our methods produce results significantly better than standard log-likelihood ratio modeling. We also demonstrate that our system can perform well in conjunction with standard cesptral speaker recognition systems. William M. Campbell, Joseph P. Campbell, Terry P. Gleason, Douglas A. Reynolds, Wade Shen |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | The Mixer and Transcript Reading Corpora: Resources for Multilingual, Crosschannel Speaker Recognition Research
Christopher Cieri, Walter D. Andrews, Joseph P. Campbell, George R. Doddington, John J. Godfrey, Shudong Huang, Mark Y. Liberman, Alvin F. Martin, Hirotaka Nakasone, Mark A. Przybocki, Kevin Walker |
LREC | 3 |
| 2006 | Support vector machines for speaker and language recognition
William M. Campbell, Joseph P. Campbell, Douglas A. Reynolds, Elliot Singer, Pedro A. Torres-Carrasquillo |
Comput. Speech Lang. | 2 |
| 2006 | Editorial
Joseph P. Campbell, John S. D. Mason, Javier Ortega-Garcia |
Comput. Speech Lang. | 1 |
| 2006 | Exploiting nonacoustic sensors for speech encodingabstractThe intelligibility of speech transmitted through low-rate coders is severely degraded when high levels of acoustic noise are present in the acoustic environment. Recent advances in nonacoustic sensors, including microwave radar, skin vibration, and bone conduction sensors, provide the exciting possibility of both glottal excitation and, more generally, vocal tract measurements that are relatively immune to acoustic disturbances and can supplement the acoustic speech waveform. We are currently investigating methods of combining the output of these sensors for use in low-rate encoding according to their capability in representing specific speech characteristics in different frequency bands. Nonacoustic sensors have the ability to reveal certain speech attributes lost in the noisy acoustic signal; for example, low-energy consonant voice bars, nasality, and glottalized excitation. By fusing nonacoustic low-frequency and pitch content with acoustic-microphone content, we have achieved significant intelligibility performance gains using the DRT across a variety of environments over the government standard 2400-bps MELPe coder. By fusing quantized high-band 4-to-8-kHz speech, requiring only an additional 116 bps, we obtain further DRT performance gains by exploiting the ear's insensitivity to fine spectral detail in this frequency region. Thomas F. Quatieri, Kevin Brady 0001, D. Messing, Joseph P. Campbell, William M. Campbell, Michael S. Brandstein, Clifford J. Weinstein, John D. Tardelli, Paul D. Gatewood |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | Estimating and Evaluating Confidence for Forensic Speaker RecognitionabstractEstimating and evaluating confidence has become a key aspect of the speaker recognition problem because of the increased use of this technology in forensic applications. We discuss evaluation measures for speaker recognition and some of their properties. We then propose a framework for confidence estimation based upon scores and meta-information, such as utterance duration, channel type, and SNR. The framework uses regression techniques with multilayer perceptrons to estimate confidence with a data-driven methodology. As an application, we show the use of the framework in a speaker comparison task drawn from the NIST 2000 evaluation. A relative comparison of different types of meta-information is given. We demonstrate that the new framework can give substantial improvements over standard distribution methods of estimating confidence. William M. Campbell, Douglas A. Reynolds, Joseph P. Campbell, Kevin Brady 0001 |
ICASSP (1) | 3 |
| 2004 | Multisensor MELPe using parameter substitutionabstractThe estimation of speech parameters and the intelligibility of speech transmitted through low-rate coders, such as MELP (mixed excitation linear prediction), are severely degraded when there are high levels of acoustic noise in the speaking environment. The application of nonacoustic and nontraditional sensors, which are less sensitive to acoustic noise than the standard microphone, is being investigated as a means to address this problem. Sensors being investigated include the general electromagnetic motion sensor (GEMS) and the physiological microphone (P-mic). As an initial effort in this direction, a multisensor MELPe coder (MELP coder with the addition of a noise preprocessor) using parameter substitution has been developed, where pitch and voicing parameters are obtained from GEMS and P-Mic sensors, respectively, and the remaining parameters are obtained as usual from a standard acoustic microphone. This parameter substitution technique is shown to produce significant and promising DRT (diagnostic rhyme test) intelligibility improvements over the standard 2400 bps MELPe coder in several high-noise military environments. Further work is in progress aimed at utilizing the nontraditional sensors for additional intelligibility improvements and for more effective lower-rate coding in noise. Kevin Brady 0001, Thomas F. Quatieri, Joseph P. Campbell, William M. Campbell, Michael S. Brandstein, Clifford J. Weinstein |
ICASSP (1) | 3 |
| 2004 | High-level speaker verification with support vector machinesabstractRecently, high-level features such as word idiolect, pronunciation, phone usage, prosody, etc., have been successfully used in speaker verification. The benefit of these features was demonstrated in the NIST extended data task for speaker verification; with enough conversational data, a recognition system can become "familiar" with a speaker and achieve excellent accuracy. Typically, high-level-feature recognition systems produce a sequence of symbols from the acoustic signal and then perform recognition using the frequency and co-occurrence of symbols. We propose the use of support vector machines for performing the speaker verification task from these symbol frequencies. Support vector machines have been applied to text classification problems with much success. A potential difficulty in applying these methods is that standard text classification methods tend to "smooth" frequencies which could potentially degrade speaker verification. We derive a new kernel based upon standard log likelihood ratio scoring to address limitations of text classification methods. We show that our methods achieve significant gains over standard methods for processing high-level features. William M. Campbell, Joseph P. Campbell, Douglas A. Reynolds, Douglas A. Jones, Timothy R. Leek |
ICASSP (1) | 2 |
| 2004 | The Mixer Corpus of Multilingual, Multichannel Speaker Recognition Data
Christopher Cieri, Joseph P. Campbell, Hirotaka Nakasone, David Miller 0002, Kevin Walker |
LREC | 2 |
| 2004 | Conversational Telephone Speech Corpus Collection for the NIST Speaker Recognition Evaluation 2004
Alvin F. Martin, David Miller 0002, Mark A. Przybocki, Joseph P. Campbell, Hirotaka Nakasone |
LREC | 4 |
| 2003 | Combining cross-stream and time dimensions in phonetic speaker recognitionabstractRecent studies show that phonetic sequences from multiple languages can provide effective features for speaker recognition. So far, only pronunciation dynamics in the time dimension, i.e., n-gram modeling on each of the phone sequences, have been examined. In the JHU 2002 Summer Workshop, we explored modeling the statistical pronunciation dynamics across streams in multiple languages (cross-stream dimension) as an additional component to the time dimension. We found that bigram modeling in the cross-stream dimension achieves improved performance over that in the time dimension on the NIST 2001 Speaker Recognition Evaluation Extended Data Task. Moreover, a linear combination of information from both dimensions at the score level further improves the performance, showing that the two dimensions contain complementary information. Qin Jin, Jirí Navrátil 0001, Douglas A. Reynolds, Joseph P. Campbell, Walter D. Andrews, Joy S. Abramson |
ICASSP (4) | 4 |
| 2003 | Conditional pronunciation modeling in speaker detectionabstractWe present a conditional pronunciation modeling method for the speaker detection task that does not rely on acoustic vectors. Aiming at exploiting higher-level information carried by the speech signal, it uses time-aligned streams of phones and phonemes to model a speaker's specific pronunciation. Our system uses phonemes drawn from a lexicon of pronunciations of words recognized by an automatic speech recognition system to generate the phoneme stream and an open-loop phone recognizer to generate a phone stream. The phoneme and phone streams are aligned at the frame level and conditional probabilities of a phone, given a phoneme, are estimated using cooccurrence counts. A likelihood detector is then applied to these probabilities. Performance is measured using the NIST Extended Data paradigm and the Switchboard-I corpus. Using 8 training conversations for enrollment, a 2.1% equal error rate was achieved. Extensions and alternatives, as well as fusion experiments, are presented and discussed. David Klusácek, Jirí Navrátil 0001, Douglas A. Reynolds, Joseph P. Campbell |
ICASSP (4) | 4 |
| 2003 | Phonetic speaker recognition using maximum-likelihood binary-decision tree modelsabstractRecent work in phonetic speaker recognition has shown that modeling phone sequences using n-grams is a viable and effective approach to speaker recognition, primarily aiming at capturing speaker-dependent pronunciation and also word usage. The paper describes a method involving binary-tree-structured statistical models for extending the phonetic context beyond that of standard n-grams (particularly bigrams) by exploiting statistical dependencies within a longer sequence window without exponentially increasing the model complexity, as is the case with n-grams. Two ways of dealing with data sparsity are also studied; namely, model adaptation and a recursive bottom-up smoothing of symbol distributions. Results obtained under a variety of experimental conditions using the NIST 2001 Speaker Recognition Extended Data Task indicate consistent improvements in equal-error rate performance as compared to standard bigram models. The described approach confirms the relevance of long phonetic context in phonetic speaker recognition and represents an intermediate stage between short phone context and word-level modeling without the need for any lexical knowledge, which suggests its language independence. Jirí Navrátil 0001, Qin Jin, Walter D. Andrews, Joseph P. Campbell |
ICASSP (4) | 4 |
| 2003 | The SuperSID project: exploiting high-level information for high-accuracy speaker recognitionabstractThe area of automatic speaker recognition has been dominated by systems using only short-term, low-level acoustic information, such as cepstral features. While these systems have indeed produced very low error rates, they ignore other levels of information beyond low-level acoustics that convey speaker information. Recently published work has shown examples that such high-level information can be used successfully in automatic speaker recognition systems and has the potential to improve accuracy and add robustness. For the 2002 JHU CLSP summer workshop, the SuperSID project (http://www.clsp.jhu.edu/ws2002/groups/supersid/) was undertaken to exploit these high-level information sources and dramatically increase speaker recognition accuracy on a defined NIST evaluation corpus and task. The paper provides an overview of the structure, data, task, tools, and accomplishments of this project. Wide ranging approaches using pronunciation models, prosodic dynamics, pitch and duration features, phone streams, and conversational interactions were explored and developed. We show how these novel features and classifiers indeed provide complementary information and can be fused together to drive down the equal error rate on the 2001 NIST extended data task to 0.2% - a 71% relative reduction in error over the previous state of the art. Douglas A. Reynolds, Walter D. Andrews, Joseph P. Campbell, Jirí Navrátil 0001, Barbara Peskin, André Adami, Qin Jin, David Klusácek, Joy S. Abramson, Radu Mihaescu, John J. Godfrey, Douglas A. Jones, Bing Xiang |
ICASSP (4) | 3 |
| 2003 | Person authentication by voice: a need for cautionabstractBecause of recent events and as members of the scientific community working in the field of speech processing, we feel compelled to publicize our views concerning the possibility of identifying or authenticating a person from his or her voice. The need for a clear and common message was indeed shown by the diversity of information that has been circulating on this matter in the media and general public over the past year. In a press release initiated by the AFCP and further elaborated in collaboration with the SpLC ISCA-SIG, the two groups herein discuss and present a summary of the current state of scientific knowledge and technological development in the field of speaker recognition, in accessible wording for nonspecialists. Our main conclusion is that, despite the existence of technological solutions to some constrained applications, at the present time, there is no scientific process that enables one to uniquely characterize a person’s voice or to identify with absolute certainty an individual from his or her voice. 1. Jean-François Bonastre, Frédéric Bimbot, Louis-Jean Boë, Joseph P. Campbell, Douglas A. Reynolds, Ivan Magrin-Chagnolleau |
INTERSPEECH | 4 |
| 2003 | Fusing high- and low-level features for speaker recognitionabstractThe area of automatic speaker recognition has been dominated by systems using only short-term, low-level acoustic information, such as cepstral features. While these systems have produced low error rates, they ignore higher levels of information beyond low-level acoustics that convey speaker information. Recently published works have demonstrated that such high-level information can be used successfully in automatic speaker recognition systems by improving accuracy and potentially increasing robustness. Wide ranging high-levelfeature-based approaches using pronunciation models, prosodic dynamics, pitch gestures, phone streams, and conversational interactions were explored and developed under the SuperSID Joseph P. Campbell, Douglas A. Reynolds, Robert B. Dunn |
INTERSPEECH | 1 |
| 2003 | Phonetic Speaker Recognition with Support Vector MachinesabstractA recent area of significant progress in speaker recognition is the use of high level features—idiolect, phonetic relations, prosody, discourse structure, etc. A speaker not only has a distinctive acoustic sound but uses language in a characteristic manner. Large corpora of speech data available in recent years allow experimentation with long term statistics of phone patterns, word patterns, etc. of an individual. We propose the use of support vector machines and term frequency analysis of phone se- quences to model a given speaker. To this end, we explore techniques for text categorization applied to the problem. We derive a new kernel based upon a linearization of likelihood ratio scoring. We introduce a new phone-based SVM speaker recognition approach that halves the er- ror rate of conventional phone-based approaches. William M. Campbell, Joseph P. Campbell, Douglas A. Reynolds, Douglas A. Jones, Timothy R. Leek |
NIPS | 2 |
| 2002 | Gender-dependent phonetic refraction for speaker recognitionabstractThis paper describes improvements to an innovative high-performance speaker recognition system. Recent experiments showed that with sufficient training data phone strings from multiple languages are exceptional features for speaker recognition. The prototype phonetic speaker recognition system used phone sequences from six languages to produce an equal error rate of 11.5% on Switchboard-I audio files. The improved system described in this paper reduces the equal error rate to less then 4%. This is accomplished by incorporating gender-dependent phone models, pre-processing the speech files to remove cross-talk, and developing more sophisticated fusion techniques for the multi-language likelihood scores. Walter D. Andrews, Mary A. Kohler, Joseph P. Campbell, John J. Godfrey, Jaime Hernandez-Cordero |
ICASSP | 3 |
| 2001 | Speaker indexing in large audio databases using anchor modelsabstractIntroduces the technique of anchor modeling in the applications of speaker detection and speaker indexing. The anchor modeling algorithm is refined by pruning the number of models needed. The system is applied to the speaker detection problem where its performance is shown to fall short of the state-of-the-art Gaussian mixture model with universal background model (GMM-UBM) system. However, it is further shown that its computational efficiency lends itself to speaker indexing for searching large audio databases for desired speakers. Here, excessive computation may prohibit the use of the GMM-UBM recognition system. Finally, the paper presents a method for cascading anchor model and GMM-UBM detectors for speaker indexing. This approach benefits from the efficiency of anchor modeling and high accuracy of GMM-UBM recognition. Douglas E. Sturim, Douglas A. Reynolds, Elliot Singer, Joseph P. Campbell |
ICASSP | 4 |
| 2001 | Phonetic speaker recognition
Walter D. Andrews, Mary A. Kohler, Joseph P. Campbell |
INTERSPEECH | 3 |
| 2000 | Speaker recognition using G.729 speech codec parametersabstractExperiments in Gaussian-mixture-model speaker recognition from mel-cepstra, derived from mel-filter bank energies (MFBs) of the G.729 codec all-pole spectral envelope, showed significant performance loss relative to the standard mel-cepstral coefficients of G.729 synthesized (coded) speech (Quatieri et al. 1999). In this paper, we investigate two approaches to recover speaker recognition performance from G.729 parameters. The first is a parametric approach that makes explicit use of G.729 parameters, rather than deriving cepstra from MFBs of an all-pole spectrum. Specifically, the G.729 LSFs are converted to "direct" cepstral coefficients for which there exists a one-to-one correspondence with the LSFs. The G.729 residual is also considered; in particular, appending G.729 pitch as a single parameter to the direct cepstral coefficients gives further performance gain. The second nonparametric approach uses the original MFB paradigm, but adds harmonic striations to the G.729 all-pole spectral envelope. Although obtaining considerable performance gains with these methods, we have yet to match the performance of G.729 synthesized speech, motivating the need for representing additional fine structure of the G.729 residual. Thomas F. Quatieri, Robert B. Dunn, Douglas A. Reynolds, Joseph P. Campbell, Elliot Singer |
ICASSP | 4 |
| 2000 | Bootstrapping for speaker recognition
Walter D. Andrews, Joseph P. Campbell, Douglas A. Reynolds |
INTERSPEECH | 2 |
| 1999 | Corpora for the evaluation of speaker recognition systemsabstractUsing standard speech corpora for development and evaluation has proven to be very valuable in promoting progress in speech and speaker recognition research. In this paper, we present an overview of current publicly available corpora intended for speaker recognition research and evaluation. We outline the corpora's salient features with respect to their suitability for conducting speaker recognition experiments and evaluations. We hope to increase the awareness and use of these standard corpora and corresponding evaluation procedures throughout the speaker recognition community. Joseph P. Campbell, Douglas A. Reynolds |
ICASSP | 1 |
| 1999 | Speaker and language recognition using speech codec parametersabstractThis paper proposes our new text-to-speech (TTS) system that concatenates large numbers of speech segments to produce very natural and intelligible synthetic speech. One novel point of our system is its new synthesis unit, which is has three remarkable characteristics as follows; The synthesis units contain all Japanese syllables together with all possible vowel sequences, so very smooth synthetic speech is produced. Both previous and succeeding phoneme environments are considered when speech segments are concatenated, so natural sounding transients from a vowel to a consonant, which is the only concatenation point with the proposed unit, are present in the synthetic speech. Each unit has various fundamental frequency (F0) contours. Therefore, F0 modification rates are very small in any synthesis event, and the F0 modification process causes only minor distortion. To develop a unit database efficiently and effectively, we analyzed 4,850,000 Japanese phrases (breath-group) containing 87,810,000 phonemes and ranked them in order of appearance frequency. Listening tests confirm the high intelligibility and naturalness of speech produced by our new TTS system. It uses the 50,000 highest frequency units that cover over 77% of Japanese texts. Thomas F. Quatieri, Elliot Singer, Robert B. Dunn, Douglas A. Reynolds, Joseph P. Campbell |
EUROSPEECH | 5 |
| 1997 | Speaker recognition: a tutorialabstractA tutorial on the design and development of automatic speaker-recognition systems is presented. Automatic speaker recognition is the use of a machine to recognize a person from a spoken phrase. These systems can operate in two modes: to identify a particular person or to verify a person's claimed identity. Speech processing and the basic components of automatic speaker-recognition systems are shown and design tradeoffs are discussed. Then, a new automatic speaker-recognition system is given. This recognizer performs with 98.9% correct decalcification. Last, the performances of various systems are compared. Joseph P. Campbell |
Proc. IEEE | 1 |
| 1996 | In Memory of Thomas E. Tremain 1934-1995abstractA brief biography of Thomas E. Tremain (1934-1995), a pioneer in the field of digital speech coding, is given highlighting his professional achievements. Joseph P. Campbell |
IEEE Trans. Speech Audio Process. | 1 |
| 1995 | Testing with the YOHO CD-ROM voice verification corpusabstractA standard database for testing voice verification systems, called YOHO, is now available from the Linguistic Data Consortium (LDC). The purpose of this database is to enable research, spark competition, and provide a means for comparative performance assessments between various voice verification systems. A test plan is presented for the suggested use of the LDC's YOHO CD-ROM for testing voice verification systems. This plan is based upon ITT's voice verification test methodology as described by Higgins, et al. (1992), but differs slightly in order to match the LDC's CD-ROM version of YOHO and to accommodate different systems. Test results of several algorithms using YOHO are also presented. Joseph P. Campbell |
ICASSP | 1 |