Beth A. Carlson

dblp:78/5922 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 1998
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-authorArtificial intelligence and machine learning · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Speech recognition and synthesis · 75% Probabilistic and Bayesian machine learning · 25%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.011994
A projection-based likelihood measure for speech recognition in noise · IEEE Trans. Speech Audio Process. 1994
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
noise-robust speech recognition
0.011994
A projection-based likelihood measure for speech recognition in noise · IEEE Trans. Speech Audio Process. 1994
Machine learning › Probabilistic and Bayesian machine learning
divergence measure
0.011991
A Computationally Compact Divergence Measure for Speech Processing · IEEE Trans. Pattern Anal. Mach. Intell. 1991
Information theory › information measures › divergence measures
directed divergence
0.011991
A Computationally Compact Divergence Measure for Speech Processing · IEEE Trans. Pattern Anal. Mach. Intell. 1991
Information theory › information measures › divergence measures
kullback-leibler divergence
0.011991
A Computationally Compact Divergence Measure for Speech Processing · IEEE Trans. Pattern Anal. Mach. Intell. 1991

Methods — techniques the papers use, named apart from their topics

itakura-saito distance · 0.0gaussian autoregressive modeling · 0.0weighted projection measure · 0.0hidden markov model · 0.0gaussian density · 0.0
YearPublicationVenuePosition
1998 Blind clustering of speech utterances based on speaker and language characteristics
Douglas A. Reynolds, Elliot Singer, Beth A. Carlson, Gerald C. O'Leary, Jack McLaughlin, Marc A. Zissman
ICSLP3
1997 Using missing feature theory to actively select features for robust speech recognition with interruptions, filtering and noise KN-37
Richard Lippmann, Beth A. Carlson
EUROSPEECH2
1996 Unsupervised topic clustering of switchboard speech messages
abstract
This paper presents a statistical technique which can be used to automatically group speech data records based on the similarity of their content. A tree-based clustering algorithm is used to generate a hierarchical structure for the corpus. This structure can then be used to guide the search for similar material in data from other corpora. The SWITCHBOARD Speech Corpus was used to demonstrate these techniques, since it provides sets of speech files which are nominally on the same topic. Excellent automatic clustering was achieved on the truth text transcripts provided with the SWITCHBOARD corpus, with an average cluster purity of 97.3%. Degraded clustering was achieved using the output transcriptions of a speech recognizer, with a clustering purity of 61.4%.
Beth A. Carlson
ICASSP1
1995 The effects of telephone transmission degradations on speaker recognition performance
abstract
The two largest factors affecting automatic speaker identification performance are the size of the population and the degradations introduced by noisy communication channels (e.g., telephone transmission). To examine experimentally these two factors, this paper presents text-independent speaker identification results for varying speaker population sizes up to 630 speakers for both clean, wideband speech and telephone speech. A system based on Gaussian mixture speaker models is used for speaker identification and experiments are conducted on the TIMIT and NTIMIT databases. This is believed to be the first speaker identification experiments on the complete 630 speaker TIMIT and NTIMIT databases and the largest text-independent speaker identification task reported to date. Identification accuracies of 99.5% and 60.7% are achieved on the TIMIT and NTIMIT databases, respectively. This paper also presents experiments which examine and attempt to quantify the performance loss associated with various telephone degradations by systematically degrading the TIMIT speech in a manner consistent with measured NTIMIT degradations and measuring the performance loss at each step. It is found that the standard degradations of filtering and additive noise do not account for all of the performance gap between the TIMIT and NTIMIT data. Measurements of nonlinear microphone distortions are also described which may explain the additional performance loss.
Douglas A. Reynolds, Marc A. Zissman, Thomas F. Quatieri, Gerald C. O'Leary, Beth A. Carlson
ICASSP5
1995 Text-dependent speaker verification using decoupled and integrated speaker and speech recognizers
Douglas A. Reynolds, Beth A. Carlson
EUROSPEECH2
1994 A projection-based likelihood measure for speech recognition in noise
abstract
Investigates a projection-based likelihood measure that significantly improves automatic speech recognition performance in the presence of additive broadband noise. The measure was developed by modifying likelihood scores in continuous Gaussian density hidden Markov models (HMMs), resulting in the weighted projection measure (WPM). Experimental results using the proposed measure are reported for several performance factors: different cepstral-based parameters, normal and multistyle speech, and various noise signals, including white, jittering white, and broadband colored noise. In all cases, significant improvements in speaker-dependent, isolated word recognition were achieved using the WPM instead of the standard Gaussian likelihood measure (weighted Euclidean distance (WED)). As an example, at a SNR of 5 dB, the WPM resulted in improvement in recognition accuracy from 19.4 to 80.6% compared with the standard WED for the DFT mel-cepstral representation.
Beth A. Carlson, Mark A. Clements
IEEE Trans. Speech Audio Process.1
1992 Speech recognition in noise using a projection-based likelihood measure for mixture density HMM's
abstract
In this study, a cepstral likelihood measure based on the projection operation is incorporated into a mixture density hidden Markov model (HMM) scheme to improve recognition in the presence of additive noise. The case in which the models are determined only under noise-free conditions is addressed. A background discussion and a derivation of the measure are provided. Recognition experiments are presented showing the usefulness of the proposed measure over the standard Gaussian measure (weighted Euclidean distance) for speaker independent, isolated word recognition in noise. It was found that the proposed mixture weighted projection measure significantly improved performance in several noise types, including white, jittering white, and colored noise. As an example, at an SNR of 10-dB white noise, recognition improved from only 38.4% correct using the Gaussian measure to 83.6% using the developed measure.>
Beth A. Carlson, Mark A. Clements
ICASSP1
1991 Application of a weighted projection measure for robust hidden Markov model based speech recognition
abstract
The use of a projection-based cepstral measure for speech recognition in noise is investigated. Interpretations of the measure's spectral and perceptual significance are given, along with its application to cepstral, melcepstral, and delta parameters. It is shown how the projection measure can be incorporated into a continuous density hidden Markov model system in the form of a weighted measure. Both the case where these densities are unimodal Gaussians and mixtures of Gaussians are addressed. In recognition experiments, the weighted projection measure significantly outperformed the standard weighted Euclidean or Gaussian distance measure.>
Beth A. Carlson, Mark A. Clements
ICASSP1
1991 A Computationally Compact Divergence Measure for Speech Processing
abstract
The directed divergence, which is a measure based on the discrimination information between two signal classes, is investigated. A simplified expression for computing the directed divergence is derived for comparing two Gaussian autoregressive processes such as those found in speech. This expression alleviates both the computational cost (reduced by two thirds) and the numerical problems encountered in computing the directed divergence. In addition, the simplified expression is compared with the Itakura-Saito distance (which asymptotically approaches the directed divergence). Although the expressions for these two distances closely resemble each other, only moderate correlations between the two were found on a set of actual speech data.>
Beth A. Carlson, Mark A. Clements
IEEE Trans. Pattern Anal. Mach. Intell.1