Jiahong Yuan

dblp:78/3632 · DBLP profile ↗
← Back
48ranked-venue papers
24as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 24 first-author · 9 since 2021Artificial intelligence and machine learning · 30 · 17 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond manual transcripts: Exploring the potential of automatic speech recognition errors in improving Alzheimer's disease detection
Yin-Long Liu, Yuanchao Li, Jiahong Yuan, Zhen-Hua Ling
J. Biomed. Informatics8
2025 Transformer-based Speech Model Learns Well as Infants and Encodes Abstractions through Exemplars in the Poverty of the Stimulus Environment
abstract
Infants are capable of learning language, predominantly through speech and associations, in impoverished environments—a phenomenon known as the Poverty of the Stimulus (POS). Is this ability uniquely human, as an innate linguistic predisposition, or can it be empirically learned through potential linguistic structures from sparse and noisy exemplars? As an early exploratory work, we systematically designed a series of tasks, scenarios, and metrics to simulate the POS. We found that the emerging speech model wav2vec2.0 with pretrained weights from an English corpus can learn well in noisy and sparse Mandarin environments. We then tested various hypotheses and observed three pieces of evidence for abstraction: label correction, categorical patterns, and clustering effects. We concluded that models can encode hierarchical linguistic abstractions through exemplars in POS environments. We hope this work offers new insights into language acquisition from a speech perspective and inspires further research.
Jiahong Yuan
COLING3
2025 The USTC System for EEG-Music Emotion Recognition Challenge
abstract
This paper presents the Neural Harmony team’s submission to Task 1 (Person Identification) of the ICASSP 2025 EEG-Music Emotion Recognition Challenge, which aims to identify the subject from a given EEG segment. To enhance performance, we propose a novel architecture incorporating the Multiscale ConvBlock and integrating attention mechanisms with convolutional networks. We also reprocessed the data and trained multiple models with different train-validation splits, which were ensembled during testing to further improve robustness. Our final results on the test data exceed the challenge baseline, achieving 100% accuracy in Person Identification. Additionally, unseen subjects were introduced to evaluate the model’s generalization ability, and the results confirm the model’s strong adaptability to new subjects.
Yin-Long Liu, Jiahong Yuan, Zhen-Hua Ling
ICASSP5
2025 Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
abstract
Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an exploration of parameter-efficient fine-tuning strategies in monolingual, cross-lingual, and transfer learning contexts. We further compare the SER ability of models and humans at both utterance- and segment-levels. Additionally, we investigate the impact of dialect on cross-lingual SER through human evaluation. Our findings reveal that models, with appropriate knowledge transfer, can adapt to the target language and achieve performance comparable to native speakers. We also demonstrate the significant effect of dialect on SER for individuals without prior linguistic and paralinguistic background. Moreover, both humans and models exhibit distinct behaviors across different emotions. These results offer new insights into the cross-lingual SER capabilities of SSL models, underscoring both their similarities to and differences from human emotion perception.
Zhichen Han, Tianqi Geng, Jiahong Yuan, Korin Richmond, Yuanchao Li
ICASSP4
2023 The Ustc System for Adress-m Challenge
abstract
This paper describes our submission to the ICASSP 2023 Signal Processing Grand Challenge (SPGC), which focuses on multilingual Alzheimer’s disease (AD) recognition through spontaneous speech. Our approaches include using a variety of acoustic features and silence-related information for AD detection and mini-mental state examination (MMSE) score prediction, and fine-tuning wav2vec2.0 models on speech in various frequency bands for AD detection. Our overall results on the test data outperform the baseline provided by the organizers, achieving 73.9% accuracy in AD detection by fine-tuning our bilingual wav2vec2.0 pre-trained model on the 0-1000Hz frequency band speech, and 4.610 RMSE (r = 0.565) in MMSE prediction through the fusion of eGeMAPS and silence features.
Kangdi Mei, Xinyun Ding, Yinlong Liu, Zhiqiang Guo, Feiyang Xu, Xin Li 0064, Tuya Naren, Jiahong Yuan, Zhen-Hua Ling
ICASSP8
2023 Improved Contextualized Speech Representations for Tonal Analysis
Jiahong Yuan, Xingyu Cai, Kenneth Church 0001
INTERSPEECH1
2022 Text2video: Text-Driven Talking-Head Video Synthesis with Personalized Phoneme - Pose Dictionary
abstract
With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds a phoneme-pose dictionary and trains a generative adversarial network (GAN) to generate video from interpolated phoneme poses. Compared to audio-driven video generation algorithms, our approach has a number of advantages: 1) It only needs about 1 min of the training data, which is significantly less than audio-driven approaches; 2) It is more flexible and not subject to vulnerability due to speaker variation; 3) It significantly reduces the preprocessing and training time from several days for audio-based methods to 4 hours, which is 10 times faster. We perform extensive experiments to compare the proposed method with state-of-the-art talking face generation methods on a benchmark dataset and datasets of our own. The results demonstrate the effectiveness and superiority of our approach.
Jiahong Yuan, Miao Liao, Liangjun Zhang
ICASSP2
2022 W-CTC: a Connectionist Temporal Classification Loss with Wild Cards
Xingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun, Jiaji Huang, Kenneth Church 0001
ICLR2
2021 Decoupling Recognition and Transcription in Mandarin ASR
abstract
Much of the recent literature on automatic speech recognition (ASR) is taking an end-to-end approach. Unlike English where the writing system is closely related to sound, Chinese characters (Hanzi) represent meaning, not sound. We propose factoring audio → Hanzi into two sub-tasks: (1) audio → Pinyin and (2) Pinyin → Hanzi, where Pinyin is a system of phonetic transcription of standard Chinese. Factoring the audio → Hanzi task in this way achieves 3.9% CER (character error rate) on the Aishell-1 corpus, the best result reported on this dataset so far.
Jiahong Yuan, Xingyu Cai, Dongji Gao, Renjie Zheng, Liang Huang 0001, Kenneth Church 0001
ASRU1
2021 Speaking Rate and Tonal Realization in Mandarin Chinese: What Can We Learn From Large Speech Corpora?
abstract
Two Mandarin speech corpora were used to investigate tonal realization in terms of duration and pitch. The data consist of nearly 1000 hours of speech from more than 1600 speakers. The two corpora, both developed for ASR, differ in speaking rate by approximately 25%. This provides an opportunity to examine the influence of speaking rate on the realization of tones in natural speech. Our analysis found two differences for slower speaking rates: (1) lower "static" tones and (2) more change for "dynamic" tones. Tone 1 was higher and Tone 3 was lower on the first syllable of disyllabic words, suggesting a metrical structure of left-prominence. On the other hand, however, the second syllable was longer, and the slope of Tone 2 and Tone 4 was higher on the second syllable in one of the corpora, both of which suggest right-prominence. We also found a shift from right-prominence to left-prominence, with respect to the realization of the "dynamic" tones, when the speaking rate became slower. Our study demonstrated that both phrasing and metrical structure play an important role in tonal realization.
Jiahong Yuan, Kenneth Church 0001
ICASSP1
2021 Pause-Encoded Language Models for Recognition of Alzheimer's Disease and Emotion
abstract
We propose enhancing Transformer language models (BERT, RoBERTa) to take advantage of pauses. Pauses play an important role in speech. In previous work we developed a method to encode pauses in transcripts for recognition of Alzheimer's disease. In this study, we extend this idea to language models. We re-train BERT and RoBERTa using a large collection of pause-encoded transcripts, and conduct fine- tuning for two downstream tasks, recognition of Alzheimer's disease and emotion. Pause-encoded language models outperform text-only language models on these tasks. Pause augmentation by duration perturbation for training is shown to improve pause-encoded language models.
Jiahong Yuan, Xingyu Cai, Kenneth Church 0001
ICASSP1
2021 Speech Emotion Recognition with Multi-Task Learning
Xingyu Cai, Jiahong Yuan, Renjie Zheng, Liang Huang 0001, Kenneth Church 0001
Interspeech2
2021 On Attention Redundancy: A Comprehensive Study
abstract
Yuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, Kenneth Church. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yuchen Bian, Jiaji Huang, Xingyu Cai, Jiahong Yuan, Kenneth Church 0001
NAACL-HLT4
2020 Detection and Analysis of T/D Deletion in Librispeech
abstract
In this study we developed a new method for automatic identification of t/d deletion. Our method achieved 94% accuracy on TIMIT and 87% on human-annotated data from Librispeech. We then conducted an analysis of t/d deletion on more than 500k tokens in Librispeech. The following results were found: (1) /d/ is more likely to be deleted than /t/; (2) t/d is more likely to be deleted when preceded by a nasal or a coronal obstruent; (3) In terms of the following phone, the rate of t/d deletion from low to high was: vowels and pausesemi-weak past tense > regular past tense; (5) t/d is less likely to be deleted when the phonological neighborhood density (PND) is higher.
Jiahong Yuan
ICASSP1
2020 LAIX Corpus of Chinese Learner English: Towards a Benchmark for L2 English ASR
Huan Luan, Jiahong Yuan
INTERSPEECH3
2020 Disfluencies and Fine-Tuning Pre-Trained Language Models for Detection of Alzheimer's Disease
Jiahong Yuan, Yuchen Bian, Xingyu Cai, Jiaji Huang, Kenneth Church 0001
INTERSPEECH1
2020 Fuzzy Correlation Measurement Algorithms for Big Data and Application to Exchange Rates and Stock Prices
abstract
In the era of Internet of people and things, big data are merging. Conventional computation algorithms including correlation measures become inefficient to deal with big data problems. Motivated by this observation, we present three fuzzy correlation measurement algorithms, that is, the centroid-based measure, the integral-based measure, and the α-cut-based measure using fuzzy techniques. Data of Shanghai stock price index (SSI) and exchange rates of main foreign currencies over China Yuan from 22 January 2013 to 17 May 2018 are used to check the effectiveness of our algorithms, and, more importantly, to observe the causality relationship between SSI and these main exchange rates. We have observed some findings as follows. First, the usage of the highest, lowest, or closing values in daily exchange rates and stock prices has impact on the significant Granger causes of exchange rates over SSI, but does not produce any opposite cause from SSI to exchange rates. Second, no matter which of our fuzzy measurement algorithms is used, Hongkong Dollar over China Yuan and U.S. Dollar over China Yuan are positively related with SSI, and Euro over China Yuan negatively correlated with SSI is always recognized as a Granger cause to SSI with the significance level being 1%. Finally, both the optimism level and the uncertainty level are observed having impact on the correlation coefficients, but the later brings more significant changes to results of the Granger causality tests.
Junhu Ruan, Jiahong Yuan, Yan Shi 0008, Yuchun Zhu, Felix T. S. Chan, Weizhen Rao
IEEE Trans. Ind. Informatics3
2019 Classification of Chinese Dialect Regions from L2 English Speech
abstract
This paper presents an effort to classify Chinese speakers' L1 dialect regions from their L2 English speech. By applying LightGBM (a gradient boosting classifier based on decision trees) to softmax-based features from deep neural networks, our system achieved 68% accuracy on five dialect regions using one sentence and 82% accuracy using 23 words. The results represent a nearly 50% error reduction over a baseline system based on HMM/GMM and forced alignment. We demonstrated that modeling phone boundaries and vowel stress yielded a relative error reduction of 18%, with phone boundaries being more useful than vowels and consonants. Furthermore, in terms of classification models, LightGBM was extremely robust on this task, which we believe deserves further investigation.
Jiahong Yuan, Zhengqiang Rao
ICASSP1
2019 On the Role of Style in Parsing Speech with Neural Models
abstract
The differences in written text and conversational speech are substantial; previous parsers trained on treebanked text have given very poor results on spontaneous speech. For spoken language, the mismatch in style also extends to prosodic cues, though it is less well understood. This paper re-examines the use of written text in parsing speech in the context of recent advances in neural language processing. We show that neural approaches facilitate using written text to improve parsing of spontaneous speech, and that prosody further improves over this state-of-the-art result. Further, we find an asymmetric degradation from read vs. spontaneous mismatch, with spontaneous speech more generally useful for training parsers.
Trang Tran 0001, Jiahong Yuan, Yang Liu 0004, Mari Ostendorf
INTERSPEECH2
2018 GlobalTIMIT: Acoustic-Phonetic Datasets for the World's Languages
Nattanun Chanchaochai, Christopher Cieri, Japhet Debrah, Sishi Liao, Mark Y. Liberman, Jonathan Wright, Jiahong Yuan, Juhong Zhan, Yuqing Zhan
INTERSPEECH9
2018 Pitch Characteristics of L2 English Speech by Chinese Speakers: A Large-scale Study
Jiahong Yuan, Qiusi Dong, Huan Luan
INTERSPEECH1
2016 The Rhythmic Constraint on Prosodic Boundaries in Mandarin Chinese Based on Corpora of Silent Reading and Speech Perception
Jiahong Yuan, Xiaoying Xu, Mark Y. Liberman
INTERSPEECH2
2016 Phoneme, Phone Boundary, and Tone in Automatic Scoring of Mandarin Proficiency
Jiahong Yuan, Mark Y. Liberman
INTERSPEECH1
2015 Investigating consonant reduction in Mandarin Chinese with improved forced alignment
abstract
Phonetic reduction has been an important topic in linguistics research. It also presents a great challenge for forced alignment, a technique widely used for automatic phonetic segmentation. In this study, we employed skip-state HMMs to improve forced alignment quality and to make forced alignment applicable to the investigation of phonetic reduction and deletion. With skip-state HMMs, forced alignment accuracy at 10 ms agreement was improved from 73.3% to 75.6% on a corpus of Mandarin Chinese broadcast news speech. Our analysis based on the improved forced alignment of Mandarin broadcast news speech – verified by hand segmentation of a random sample of cases – shows that: 1. The durations of frication and aspiration are additive in the production of plosives and affricates; 2. Plosives are more likely to be deleted than affricates; and 3. Plosives and affricates in higher-frequency words and at word-medial position are more likely to be reduced.
Jiahong Yuan, Mark Y. Liberman
INTERSPEECH1
2014 Mandarin tone classification without pitch tracking
abstract
A deep neural network (DNN) based classifier achieved 27.38% frame error rate (FER) and 15.62% segment error rate (SER) in recognizing five tonal categories in Mandarin Chinese broadcast news, based on 40 mel-frequency cepstral coefficients (MFCCs). The same architecture scored substantially lower when trained and tested with F0and amplitude parameters alone: 40.05% FER and 22.66% SER. These results are substantially better than the best previously-reported results on broadcast-news tone classification [1] and are also better than a human listener achieved in categorizing test stimuli created by amplitude- and frequency-modulating complex tones to match the extracted F0and amplitude parameters.
Neville Ryant, Jiahong Yuan, Mark Y. Liberman
ICASSP2
2014 Highly accurate phonetic segmentation using boundary correction models and system fusion
abstract
Accurate phone-level segmentation of speech remains an important task for many subfields of speech research. We investigate techniques for boosting the accuracy of automatic phonetic segmentation based on HMM acoustic-phonetic models. In prior work [25] we were able to improve on state-of-the-art alignment accuracy by employing special phone boundary HMM models, trained on phonetically segmented training data, in conjunction with a simple boundary-time correction model. Here we present further improved results by using more powerful statistical models for boundary correction that are conditioned on phonetic context and duration features. Furthermore, we find that combining multiple acoustic front-ends gives additional gains in accuracy, and that conditioning the combiner on phonetic context and side information helps. Overall, we reduce segmentation errors on the TIMIT corpus by almost one half, from 93.9% to 96.8% boundary accuracy with a 20-ms tolerance.
Andreas Stolcke, Neville Ryant, Vikramjit Mitra, Jiahong Yuan, Wen Wang 0001, Mark Y. Liberman
ICASSP4
2014 Automatic phonetic segmentation in Mandarin Chinese: Boundary models, glottal features and tone
abstract
We conducted experiments on forced alignment in Mandarin Chinese. A corpus of 7,849 utterances was created for the purpose of the study. Systems differing in their use of explicit phone boundary models, glottal features, and tone information were trained and evaluated on the corpus. Results showed that employing special one-state phone boundary HMM models significantly improved forced alignment accuracy, even when no manual phonetic segmentation was available for training. Spectral features extracted from glottal waveforms (by performing glottal inverse filtering from the speech waveforms) also improved forced alignment accuracy. Tone dependent models only slightly outperformed tone independent models. The best system achieved 93.1% agreement (of phone boundaries) within 20 ms compared to manual segmentation without boundary correction.
Jiahong Yuan, Neville Ryant, Mark Y. Liberman
ICASSP1
2014 F0 declination in English and Mandarin Broadcast News Speech
Jiahong Yuan, Mark Y. Liberman
Speech Commun.1
2013 Using multiple versions of speech input in phone recognition
abstract
This study investigates the use of multiple versions of the same speech unit in automatic phone recognition. Two methods were applied to combine multiple utterance versions in decoding: cross forced-alignment and n-best ROVER. The phone error rate was reduced from 15% to 2% on isolated words and from 33% to 19% on TIMIT sentences. The error rate was reduced the most when the second version was added, and less so as each additional version was added. Depending on the language model weight, it might be better to use the language model only in n-best generation, but omit it in scoring the hypotheses applied to the combination methods. N-best ROVER effectiveness may be enhanced by lowering the language model weight.
Mark Y. Liberman, Jiahong Yuan, Andreas Stolcke, Wen Wang 0001, Vikramjit Mitra
ICASSP2
2013 Articulatory trajectories for large-vocabulary speech recognition
abstract
Studies have demonstrated that articulatory information can model speech variability effectively and can potentially help to improve speech recognition performance. Most of the studies involving articulatory information have focused on effectively estimating them from speech, and few studies have actually used such features for speech recognition. Speech recognition studies using articulatory information have been mostly confined to digit or medium vocabulary speech recognition, and efforts to incorporate them into large vocabulary systems have been limited. We present a neural network model to estimate articulatory trajectories from speech signals where the model was trained using synthetic speech signals generated by Haskins Laboratories' task-dynamic model of speech production. The trained model was applied to natural speech, and the estimated articulatory trajectories obtained from the models were used in conjunction with standard cepstral features to train acoustic models for large-vocabulary recognition systems. Two different large-vocabulary English datasets were used in the experiments reported here. Results indicate that employing articulatory information improves speech recognition performance not only under clean conditions but also under noisy background conditions. Perceptually motivated robust features were also explored in this study and the best performance was obtained when systems based on articulatory, standard cepstral and perceptually motivated feature were all combined.
Vikramjit Mitra, Wen Wang 0001, Andreas Stolcke, Hosung Nam, Colleen Richey, Jiahong Yuan, Mark Y. Liberman
ICASSP6
2013 Scale-space expansion of acoustic features improves speech event detection
abstract
In a system for detecting and measuring phonetic events (here bursts, voice onsets, and voice-onset times), we show that the addition of features smoothed at multiple scales can improve both recall (the proportion of events correctly identified) and measurement accuracy (the timing of events and the difference between event times, relative to expert human judgments). Multi-scale (or “scale space”) features had an especially strong positive effect on robustness across datasets with different materials and recording conditions. Standard machine-learning classifiers were able to integrate information across scales, without any special treatment of the multi-scale features.
Neville Ryant, Jiahong Yuan, Mark Y. Liberman
ICASSP2
2013 Speech activity detection on youtube using deep neural networks
abstract
Speech activity detection (SAD) is an important first step in speech processing. Commonly used methods (e.g., frame-level classification using gaussian mixture models (GMMs)) work well under stationary noise conditions, but do not generalize well to domains such as YouTube, where videos may exhibit a diverse range of environmental conditions. One solution is to augment the conventional cepstral features with additional, hand-engineered features (e.g., spectral flux, spectral centroid, multiband spectral entropies) which are robust to changes in environment and recording condition. An alternative approach, explored here, is to learn robust features during the course of training using an appropriate architecture such as deep neural networks (DNNs). In this paper we demonstrate that a DNN with input consisting of multiple frames of mel frequency cepstral coefficients (MFCCs) yields drastically lower frame-wise error rates (19.6%) on YouTube videos compared to a conventional GMM based system (40%).
Neville Ryant, Mark Y. Liberman, Jiahong Yuan
INTERSPEECH3
2013 The spectral dynamics of vowels in Mandarin Chinese
Jiahong Yuan
INTERSPEECH1
2013 Automatic phonetic segmentation using boundary models
abstract
This study attempts to improve automatic phonetic segmentation within the HMM framework. Experiments were conducted to investigate the use of phone boundary models, the use of precise phonetic segmentation for training HMMs, and the difference between context-dependent and contextindependent phone models in terms of forced alignment performance. Results show that the combination of special one-state phone boundary models and monophone HMMs can significantly improve forced alignment accuracy. HMM-based forced alignment systems can also benefit from using precise phonetic segmentation for training HMMs. Context-dependent phone models are not better than context-independent models when combined with phone boundary models. The proposed system achieves 93.92% agreement (of phone boundaries) within 20 ms compared to manual segmentation on the TIMIT corpus. This is the best reported result on TIMIT to our knowledge.
Jiahong Yuan, Neville Ryant, Mark Y. Liberman, Andreas Stolcke, Vikramjit Mitra, Wen Wang 0001
INTERSPEECH1
2013 A Cross-language Study on Automatic Speech Disfluency Detection
Wen Wang 0001, Andreas Stolcke, Jiahong Yuan, Mark Y. Liberman
HLT-NAACL3
2011 Automatic detection of "g-dropping" in American English using forced alignment
abstract
This study investigated the use of forced alignment for automatic detection of “g-dropping” in American English (e.g., walkin'). Two acoustic models were trained, one for -in' and the other for -ing. The models were added to the Penn Phonetics Lab Forced Aligner, and forced alignment will choose the more probable pronunciation from the two alternatives. The agreement rates between the forced alignment method and native English speakers ranged from 79% to 90%, which were comparable to the agreement rates among the native speakers (79% - 96%). The two variations of pronunciation not only differed in their nasal codas, but also - and even more so - in their vowel quality. This is shown by both the KL-divergence between the two models, and that native Mandarin speakers performed poorly on classification of “g-dropping”.
Jiahong Yuan, Mark Y. Liberman
ASRU1
2010 Robust speaking rate estimation using broad phonetic class recognition
abstract
Robust speaking rate estimation can be useful in automatic speech recognition and speaker identification, and accurate, automatic measures of speaking rate are also relevant for research in linguistics, psychology, and social sciences. In this study we built a broad phonetic class recognizer for speaking rate estimation. We tested the recognizer on a variety of data sets, including laboratory speech, telephone conversations, foreign accented speech, and speech in different languages, and we found that the recognizer's estimates are robust under these sources of variation. We also found that the acoustic models of the broad phonetic classes are more robust than those of the monophones for syllable detection.
Jiahong Yuan, Mark Y. Liberman
ICASSP1
2010 Linguistic rhythm in foreign accent
Jiahong Yuan
INTERSPEECH1
2010 F0 declination in English and Mandarin broadcast news speech
Jiahong Yuan, Mark Y. Liberman
INTERSPEECH1
2009 Comparison of vowel structures of Japanese and English in articulatory and auditory spaces
abstract
In previous work [1] we investigated the vowel structures of Japanese in both articulatory space and auditory perceptual space using Laplacian eigenmaps, and examined relations between speech production and perception. The results showed that the inherent structures of Japanese vowels were consistent in the two spaces. To verify whether such a property generalizes to other languages, we use the same approach to investigate the more crowded English vowel space. Results show that the vowel structure reflects the articulatory features for both languages. The degree of tongue-palate approximation is the most important feature for vowels, followed by the open ratio of the mouth to oral cavity. The topological relations of the vowel structures are consistent with both the articulatory and auditory perceptual spaces; in particular the lip-protruded vowel /UW / of English was distinct from the unrounded Japanese /�/. The rhotic vowel /ER / was located apart from the surface constructed by the other vowels, where the same phenomena appeared in both spaces. Index Terms: vowels, speech production, speech perception 1.
Mark K. Tiede, Jiahong Yuan
INTERSPEECH3
2009 Investigating /l/ variation in English through forced alignment
abstract
We present a new method for measuring the "darkness " of /l/, and use it to investigate the variation of English /l / in a large speech corpus that is automatically aligned with phones predicted from an orthographic transcript. We found a correlation between the rime duration and /l/-darkness for syllable-final /l/, but no correlation between /l / duration and darkness for syllable-initial /l/. The data showed a clear difference between clear and dark /l / in English, and also showed that syllable-final /l / was less dark preceding an unstressed vowel than preceding a consonant or a word boundary. Index Terms: gestural phonology, forced alignment, variation 1.
Jiahong Yuan, Mark Y. Liberman
INTERSPEECH1
2008 Covariations of English segmental durations across speakers
Jiahong Yuan
INTERSPEECH1
2008 Different roles of pitch and duration in distinguishing word stress in English
Jiahong Yuan, Stephen Isard, Mark Y. Liberman
INTERSPEECH1
2007 A corpus study of the 3rd tone sandhi in standard Chinese
abstract
In Standard Chinese, a Low tone (Tone3) is often realized with a rising F0 contour before another Low tone, known as the 3 rd tone Sandhi. This study investigates the acoustic characteristics of the 3 rd tone Sandhi in Standard Chinese using a large telephone conversation speech corpus. Sandhi Rising was found to be different from the underlying Rising tone (Tone2) in bi-syllabic words in two measures: the magnitude of the F0 rising and the time span of the F0 rising. We also found different effects of word frequency on Sandhi Rising and the underlying Rising tones. Finally, for trisyllabic constituents with Low tone only, constituent boundary showed interesting but puzzling effects on the 3 rd tone Sandhi.
Yiya Chen, Jiahong Yuan
INTERSPEECH2
2007 Perception of disfluency: language differences and listener bias
abstract
This paper describes a crosslinguistic disfluency perception experiment. We tested the recognizability of pause fillers and partial words in English, German and Mandarin. Subjects were speakers of English with no knowledge of Mandarin or German. We found that subjects could identify disfluent from fluent utterances at a level above chance. Pause fillers were easier to identify than partial words. Accuracy rates were highest for English, followed by German and then Mandarin. Although German accuracy rates were higher than those for Mandarin, discriminability analysis suggests that this is due to conservative bias towards false negatives rather than non-recognition of the acoustic material. The fact that subjects could identify disfluent speech in languages they did not know shows that there are real phonetic crosslinguistic cues to disfluency. Index Terms: crosslinguistic perception, disfluency, pause filler, partial words.
Catherine Lai, Kyle Gorman, Jiahong Yuan, Mark Y. Liberman
INTERSPEECH3
2006 Towards an integrated understanding of speaking rate in conversation
abstract
We investigate factors that affect speaking rate in conversation, using large corpora of conversational telephone speech in English and Chinese. We find that speaking rate as a function of "turn" length rises rapidly for turns from one to seven words; remains level (when final words are included) or falls gradually (if final words are excluded) for turns of medium length; and rises slowly for longer turns. When talking with strangers or discussing certain topics, people tend to use longer turns but slower speech rates. In general older people have a slower speech, and males tend to speak slightly faster than females. Finally, we find that the effect of L1 (native language) on L2 (second language) speaking rate is L1 dependent.
Jiahong Yuan, Mark Y. Liberman, Christopher Cieri
INTERSPEECH1
2005 Pitch accent prediction: effects of genre and speaker
abstract
To build a robust pitch accent prediction system, we need to understand the effects of speech genre and speaker variation. This paper reports our studies on genre and speaker variation in pitch accent placement and their effects on automatic pitch accent prediction. We find some interesting accentuation pattern differences that can be attributed to speech genre, and a set of textual features that are robust to genre in accent prediction. We also find that although there is significant variation among speakers in pitch accent placement, speaker dependent models are not needed in accent prediction. Finally, we show that after taking speaker variation into account, there is little room to improve for state-of-the-art classifiers on read news speech. 1.
Jiahong Yuan, Jason M. Brenier, Daniel Jurafsky
INTERSPEECH1
2002 The acoustic realization of anger, fear, joy and sadness in Chinese
abstract
This paper studies the acoustic realization of anger, fear, joy and sadness in Chinese. An emotion database of total 288 sentences was collected from nine speakers. Four listeners were asked to judge the emotion type of each sentence, choosing from anger, fear, joy, sadness and neutral. The results suggest that there are two dimensions in the acoustic realization of anger, fear, joy and sadness in Chinese. We then conducted acoustic analyses on the database. The acoustic attributes of each emotion in the domains of phonation, articulation and prosody are reported. We concluded that anger and fear are mainly realized on phonation; joy is mainly realized on prosody of F0; and sadness is realized on both of the two dimensions. Articulation and prosody of duration play a secondary role in the acoustic realization of the emotions. 1.
Jiahong Yuan, Li Qin Shen, Fangxin Chen
INTERSPEECH1