Jennifer Cole 0001

dblp:66/4279 · also Jennifer S. Cole · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-0465-4920ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Crowdsourced and Automatic Speech Prominence Estimation
abstract
The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the prominence of each word in an utterance. These prominence labels are useful for linguistic analysis, as well as training automated systems to perform emphasis-controlled text-to-speech or emotion recognition. Manually annotating prominence is time-consuming and expensive, which motivates the development of automated methods for speech prominence estimation. However, developing such an automated system using machine-learning methods requires human-annotated training data. Using our system for acquiring such human annotations, we collect and open-source crowd-sourced annotations of a portion of the LibriTTS dataset. We use these annotations as ground truth to train a neural speech prominence estimator that generalizes to unseen speakers, datasets, and speaking styles. We investigate design decisions for neural prominence estimation as well as how neural prominence estimation improves as a function of two key factors of annotation cost: dataset size and the number of annotations per utterance.
Max Morrison, Pranav Pawar, Nathan Pruyne, Jennifer Cole 0001, Bryan Pardo
ICASSP4
2024 K-means and hierarchical clustering of f0 contours
Constantijn Kaland, Jeremy Steffman, Jennifer Cole 0001
INTERSPEECH3
2024 Phonological Symmetry Does Not Predict Generalization of Perceptual Adaptation to Vowels
Zuheyra Tokac, Jennifer Cole 0001
INTERSPEECH2
2023 Pitch Accent Variation and the Interpretation of Rising and Falling Intonation in American English
Thomas Sostarics, Jennifer Cole 0001
INTERSPEECH2
2022 DeepFry: Identifying Vocal Fry Using Deep Neural Networks
abstract
Vocal fry or creaky voice refers to a voice quality characterized by irregular glottal opening and low pitch. It occurs in diverse languages and is prevalent in American English, where it is used not only to mark phrase finality, but also sociolinguistic factors and affect. Due to its irregular periodicity, creaky voice challenges automatic speech processing and recognition systems, particularly for languages where creak is frequently used. This paper proposes a deep learning model to detect creaky voice in fluent speech. The model is composed of an encoder and a classifier trained together. The encoder takes the raw waveform and learns a representation using a convolutional neural network. The classifier is implemented as a multi-headed fully-connected network trained to detect creaky voice, voicing, and pitch, where the last two are used to refine creak prediction. The model is trained and tested on speech of American English speakers, annotated for creak by trained phoneticians. We evaluated the performance of our system using two encoders: one is tailored for the task, and the other is based on a state-of-the-art unsupervised representation. Results suggest our best-performing system has improved recall and F1 scores compared to previous methods on unseen data.
Bronya Roni Chernyak, Talia Ben Simon, Yael Segal-Feldman, Jeremy Steffman, Eleanor Chodroff, Jennifer Cole 0001, Joseph Keshet
INTERSPEECH6
2019 Testing the Distinctiveness of Intonational Tunes: Evidence from Imitative Productions in American English
abstract
Understanding the structure of intonational variation is a longstanding issue in prosodic research.A given utterance can be realized with countless intonational contours, and while variation in prosodic meaning is also large, listeners nevertheless converge on relatively consistent form-function mappings.While this suggests the existence of abstract intonational representations, it has been unclear how exactly these are defined.The present study examines the validity of a well-defined set of phonological representations for the generation of intonation in the nuclear region of an intonational phrase in American English: namely, the combination of binary pitch accents (H*/L*), phrase accents (H-/L-), and boundary tones (H%/L%) proposed in Pierrehumbert (1980).In an exploratory study, we examined whether speakers maintained the eight-way distinction among intonational contours posited to exist in this representational system.We created eight synthesized contours according to Pierrehumbert (1980) and examined whether listeners generalized these contours to novel productions.Speakers largely distinguished rising from nonrising contours in production, but few other distinctions were maintained.While this does not rule out the existence of additional contours in production, these findings do suggest that the representation of rising and non-rising contours may be privileged and more readily accessible in the intonational grammar.
Eleanor Chodroff, Jennifer Cole 0001
INTERSPEECH2
2019 Perception of Pitch Contours in Speech and Nonspeech
Daniel R. Turner, Ann R. Bradlow, Jennifer Cole 0001
INTERSPEECH3
2018 Information Structure, Affect and Prenuclear Prominence in American English
abstract
The influence of information structure (IS: givenness, accessibility, newness and focus) on pitch accent assignment and acoustic prominence measures of prenuclear words was investigated for American English speech elicited through read production of mini-stories. Results showed a consistent pattern of accenting the initial content word in the sentence, supporting an analysis of prenuclear accent as structural, or ‘rhythmic’. While no association was observed between IS and accent type (e.g., H*, L*, L+H*, L*+H), the acoustic-phonetic realization of prominence was modulated by information structure. In particular, words that carry contrastive focus generally showed more extreme f0 excursions relative to the average. In addition, there was a strong influence of speaking style or ‘affect’ on both pitch accent type and the acoustic-phonetic realization of prominence. Speakers were more likely to produce L+H* accents in a lively than a neutral speaking style. Differences in affect were also strongly reflected in f0 excursion, duration, and amplitude within the target word. Overall, this study indicates both linguistic (information structure) and paralinguistic (affect) influences on the phonetic implementation of prenuclear prominence, with varying influence of these two factors on the phonological assignment of prenuclear pitch accents.
Eleanor Chodroff, Jennifer Cole 0001
INTERSPEECH2
2017 Crowd-sourcing prosodic annotation
Jennifer Cole 0001, Tim Mahrt, Joseph Roy
Comput. Speech Lang.1
2015 Analysis and classification of cooperative and competitive dialogs
abstract
Cooperative and competitive game dialogs are comparatively examined with respect to temporal, basic text-based, and dialog act characteristics.The condition-specific speaker strategies are amongst others well reflected in distinct dialog act probability distributions, which are discussed in the context of the Gricean Cooperative Principle and of Relevance Theory.Based on the extracted features, we trained Bayes classifiers and support vector machines to predict the dialog condition, that yielded accuracies from 90 to 100%.Taken together the simplicity of the condition classification task and its probabilistic expressiveness for dialog acts suggests a two-stage classification of condition and dialog acts.
Uwe D. Reichel, Nina Pörner, Dianne Nowack, Jennifer Cole 0001
INTERSPEECH4
2014 Detecting articulatory compensation in acoustic data through linear regression modeling
abstract
Examining articulatory compensation has been important in understanding how the speech production system is organized, and how it relates to the acoustic and ultimately phonological levels. This paper offers a method that detects articulatory compensation in the acoustic signal, which is based on linear regression modeling of co-variation patterns between acoustic cues. We demonstrate the method on selected acoustic cues for spontaneously produced American English stop consonants. Compensatory patterns of cue variation were observed for voiced stops in some cue pairs, while uniform patterns of cue variation were found for stops as a function of place of articulation or position in the word. Overall, the results suggest that this method can be useful for observing articulatory strategies indirectly from acoustic data and testing hypotheses about the conditions under which articulatory compensation is most likely.
Alina Khasanova, Jennifer Cole 0001, Mark Hasegawa-Johnson
INTERSPEECH2
2012 F0 and the Perception of Prominence
abstract
This study investigates the role F0 plays in the perception of prominence in American English. Raw, log and locally normalized measures of F0 were extracted from words in a 35K word corpus of spontaneous speech. Linear regression analyses were conducted to test the strength of these measures as cues to prominence, with prominence based on judgments made by ordinary listeners in real-time auditory perception. The Bayesian Information Criterion was used to further investigate whether these F0 measures cue prominence in a linear or piecewise linear function, corresponding to a linguistic model of prominence as a gradient or discrete feature. The results of this study show that F0 measures are similar to intensity measures in both their strength as cues to perceived prominence, and in signaling a discrete prominence distinction that distinguishes nonor weaklyprominent words from words with greater prominence. Our finding that F0 and intensity cue a discrete prominence distinction is compared with our prior finding that duration and word frequency signal gradient prominence distinctions. This apparent discrepancy is discussed in terms of the dual nature of prominence in English, as an expression of layered metrical (stress) structure in phonology, and as an expression of pragmatic focus.
Tim Mahrt, Jennifer Cole 0001, Margaret M. Fleck, Mark Hasegawa-Johnson
INTERSPEECH2
2011 The Phonology and Phonetics of Perceived Prosody: What do Listeners Imitate?
abstract
An imitation experiment tests the hypothesis that when asked to reproduce a spontaneously-spoken utterance that they hear, speakers imitate the prosody of the stimulus in its phonological structure more accurately than the phonetic details. Results suggest that speakers rarely distort the presence of a pitch accent or an intonational phrase boundary, but more often change the nature of the phonetic cues, e.g. the duration of a pause or the occurrence of irregular pitch periods associated with boundaries and accents in American English. These findings argue for an encoding of phonological prosodic structure that is separate from the phonetic cues that signal that structure.
Jennifer Cole 0001, Stefanie Shattuck-Hufnagel
INTERSPEECH1
2011 Optimal Models of Prosodic Prominence Using the Bayesian Information Criterion
abstract
This study investigated the relation between various acoustic features and prominence. Past research has suggested that duration, pitch, and intensity all play a role in the perception of prominence. In our past work, we found a correlation between these acoustic features and speaker agreement over the placement of prominence. The current study was motivated by a need to enrich our understanding of this correlation. Using the Bayesian information criterion, we show that the best model for a feature that cues prosody is not necessarily a single Gaussian. Rather, the best model depends on the feature. This finding has consequences for our understanding of the role of these features in the perception of prosody and for prosody recognition systems.
Tim Mahrt, Jui-Ting Huang, Yoonsook Mo, Margaret M. Fleck, Mark Hasegawa-Johnson, Jennifer Cole 0001
INTERSPEECH6
2009 Prosodic effects on vowel production: evidence from formant structure
abstract
Speakers communicate pragmatic and discourse meaning through the prosodic form assigned to an utterance, and listeners must attend to the acoustic cues to prosodic form to fully recover the speaker’s intended meaning. While much of the research on prosody examines supra-segmental cues such as F0 and temporal patterns, prosody is also known to affect the phonetic properties of segments as well. This paper reports on the effect of prosodic prominence on the formant patterns of vowels using speech data from the Buckeye corpus of spontaneous American English. A prosody annotation was obtained for a subset of this corpus based on the auditory perception of 97 ordinary, untrained listeners. To understand the relationship between prominence perception and formant structure, as a measure of the ‘strength ’ of the vowel
Yoonsook Mo, Jennifer Cole 0001, Mark Hasegawa-Johnson
INTERSPEECH2
2006 Prosody dependent speech recognition on radio news corpus of American English
abstract
Does prosody help word recognition? This paper proposes a novel probabilistic framework in which word and phoneme are dependent on prosody in a way that reduces word error rates (WER) relative to a prosody-independent recognizer with comparable parameter count. In the proposed prosody-dependent speech recognizer, word and phoneme models are conditioned on two important prosodic variables: the intonational phrase boundary and the pitch accent. An information-theoretic analysis is provided to show that prosody dependent acoustic and language modeling can increase the mutual information between the true word hypothesis and the acoustic observation by exciting the interaction between prosody dependent acoustic model and prosody dependent language model. Empirically, results indicate that the influence of these prosodic variables on allophonic models are mainly restricted to a small subset of distributions: the duration PDFs (modeled using an explicit duration hidden Markov model or EDHMM) and the acoustic-prosodic observation PDFs (normalized pitch frequency). Influence of prosody on cepstral features is limited to a subset of phonemes: for example, vowels may be influenced by both accent and phrase position, but phrase-initial and phrase-final consonants are independent of accent. Leveraging these results, effective prosody dependent allophonic models are built with minimal increase in parameter count. These prosody dependent speech recognizers are able to reduce word error rates by up to 11% relative to prosody independent recognizers with comparable parameter count, in experiments based on the prosodically-transcribed Boston Radio News corpus.
Ken Chen 0001, Mark Hasegawa-Johnson, Aaron Cohen, Sarah Borys, Sung-Suk Kim, Jennifer Cole 0001, Jeung-Yoon Choi
IEEE Trans. Speech Audio Process.6
2005 The stress foot as a unit of planned timing: evidence from shortening in the prosodic phrase
Jennifer Cole 0001
INTERSPEECH2
2005 Simultaneous recognition of words and prosody in the Boston University Radio Speech Corpus
Mark Hasegawa-Johnson, Ken Chen 0001, Jennifer Cole 0001, Sarah Borys, Sung-Suk Kim, Aaron Cohen, Tong Zhang 0005, Jeung-Yoon Choi, Taejin Yoon
Speech Commun.3
2004 Modeling and recognition of phonetic and prosodic factors for improvements to acoustic speech recognition models
abstract
This paper examines the usefulness of including prosodic and phonetic context information in the phoneme model of a speech recognizer. This is done by creating a series of prosodic and phonetic models and then comparing the mutual information between the observations and each possible context variable. Prosodic variables show improvement less often than phone context variables, however, prosodic variables generally show a larger increase in mutual information. A recognizer with allophones defined using the maximum mutual information prosodic and phonetic variables outperforms a recognizer with allophones defined exclusively using phonetic variables. 1.
Sarah Borys, Aaron Cohen, Mark Hasegawa-Johnson, Jennifer Cole 0001
INTERSPEECH4
2004 Intertranscriber reliability of prosodic labeling on telephone conversation using toBI
abstract
Two transcribers have labeled prosodic events indepen-dently on a subset of Switchboard corpus using adapted ToBI (TOnes and Break Indices) system. Transcriptions of two types of pitch accents (H * and L*), phrasal accents (H- and L-) and boundary tones (H % and L%) encoded independently by two transcribers are compared for intertranscriber reliabil-ity. Two commonly used methods of reliability measurement, ‘transcriber-pair-word ’ comparison and kappa statistic, are used for comparison with previous reports on the intertranscriber consistency. The results obtained from transcriber-pair-word comparison are: The overall agreement on the presence or ab-sence and choice of pitch accent is 86.57%. The agreement on the presence or absence and the choice of phrasal accent is 85.63%. The presence and choice of boundary tone is 89.33%. When both transcribers agreed that there is at least a phrasal tone, the agreement on the choice of the type of either phrasal accent or boundary tone is 73.86%. The kappa coefficient of agreement (K) of 0.7 to 1 indicates the degree of reliability to be from good to perfect. A kappa coefficient of 0.75 is obtained for agreement on the presence or absence of pitch accents, 0.67 for the presence of phrasal accents, and 0.61 for the strength of disjuncture between phrasal accent and boundary tone. Com-parison of the present results with those of previous reliability studies [1][2][3][4] suggests that some higher agreement rates for this study may result from our adoption of fewer labeling distinctions in the transcription of pitch accent events. The results for phrase boundary labeling suggest that spontaneous speech of the type found in the Switchboard corpus is harder to code for the degree of disjuncture between prosodic domains than is read speech. 1.
Taejin Yoon, Sandra Chavarria, Jennifer Cole 0001, Mark Hasegawa-Johnson
INTERSPEECH3
2003 Prosody dependent speech recognition with explicit duration modelling at intonational phrase boundaries
abstract
Does prosody help word recognition? In this paper, we propose a novel probabilistic framework in which word and phoneme are dependent on prosody in a way that improves word recognition. The prosody attribute that we investigate in this study is the duration lengthening effects of the speech segments in the vicinity of intonational phrase boundaries. Explicit Duration Hidden Markov Model (EDHMM) is implemented to provide an accurate phoneme duration model. This study is conducted on Boston University Radio New Corpus with prosodic boundaries marked using ToBI labelling system. We found that lengthening of the phrase final rhymes can be reliably modelled by EDHMM, which significantly improves the prosody dependent acoustic modelling. Conversely, no systematic duration variation is found at phrase initial position. With prosody dependence implemented in acoustic model, pronunciation model and language model, both word recognition accuracy and boundary recognition accuracy are improved by 1% over systems without prosody dependence.
Ken Chen 0001, Sarah Borys, Mark Hasegawa-Johnson, Jennifer Cole 0001
INTERSPEECH4