EDBT 2026 Demo / reviewers in the wild / expert
Frank K. Soong
dblp:84/6490
· DBLP profile ↗
243ranked-venue papers
19as first author
18since 2021 · last 2024
0000-0002-9088-3577ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 211 · 18 first-author · 12 since 2021Artificial intelligence and machine learning · 122 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level QualityabstractText-to-speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality, how to define/judge that quality, and how to achieve it. In this paper, we answer these questions by first defining the human-level quality based on the statistical significance of subjective measure and introducing appropriate guidelines to judge it, and then developing a TTS system called NaturalSpeech that achieves human-level quality on benchmark datasets. Specifically, we leverage a variational auto-encoder (VAE) for end-to-end text-to-waveform generation, with several key modules to enhance the capacity of the prior from text and reduce the complexity of the posterior from speech, including phoneme pre-training, differentiable duration modeling, bidirectional prior/posterior modeling, and a memory mechanism in VAE. Experimental evaluations on the popular LJSpeech dataset show that our proposed NaturalSpeech achieves -0.01 CMOS (comparative mean opinion score) to human recordings at the sentence level, with Wilcoxon signed rank test at p-level p >> 0.05, which demonstrates no statistically significant difference from human recordings for the first time. Xu Tan 0003, Jiawei Chen 0008, Haohe Liu, Jian Cong, Chen Zhang 0020, Xi Wang 0016, Yichong Leng, Yuanhao Yi, Lei He 0005, Sheng Zhao 0002, Tao Qin 0001, Frank K. Soong, Tie-Yan Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 13 |
| 2023 | ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph ReadingabstractWhile state-of-the-art Text-to-Speech systems can generate natural speech of very high quality at sentence level, they still meet great challenges in speech generation for paragraph / long-form reading.Such deficiencies are due to i) ignorance of cross-sentence contextual information, and ii) high computation and memory cost for long-form synthesis.To address these issues, this work develops a lightweight yet effective TTS system, ContextSpeech.Specifically, we first design a memory-cached recurrence mechanism to incorporate global text and speech context into sentence encoding.Then we construct hierarchically-structured textual semantics to broaden the scope for global context enhancement.Additionally, we integrate linearized self-attention to improve model efficiency.Experiments show that ContextSpeech significantly improves the voice quality and prosody expressiveness in paragraph reading with competitive model efficiency. Yujia Xiao, Shaofei Zhang, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Sheng Zhao 0002, Frank K. Soong, Tan Lee |
INTERSPEECH | 7 |
| 2023 | MSMC-TTS: Multi-Stage Multi-Codebook VQ-VAE Based Neural TTSabstractThis paper aims to improve neural TTS with vector-quantized, compact speech representations. We propose a Vector-Quantized Variational AutoEncoder (VQ-VAE) based feature analyzer to encode acoustic features into sequences with different time resolutions, and quantize them with multiple VQ codebooks to form the Multi-Stage Multi-Codebook Representation (MSMCR). The TTS system, MSMC-TTS, is proposed to predict better speech via this representation. In prediction, the multi-stage predictor is trained to map the input text sequence to MSMCRs in stages, by minimizing Euclidean distance and “triplet loss”. In synthesis, the neural vocoder converts ground-truth or predicted MSMCRs into speech waveforms. The proposed system is trained with single-speaker TTS datasets and tested in various scenarios for comprehensive evaluation. In TTS evaluation, MSMC-TTS obtains MOS of 4.34 and 4.10 on English and Chinese datasets, which significantly outperforms VITS with scores of 3.78 and 3.90. Meanwhile, compared with Mel-Spectrograms, the domain discrepancy between prediction and ground truth is lower in MSMCRs with the higher Domain-classification Error Rate (DER). Furthermore, this system shows lower modeling complexity and data size requirements, preserving excellent performance even with fewer model parameters or training data. The noticeable improvement in analysis-synthesis and TTS from multiple codebooks and stages also validate them as vital components in seeking a more profitable speech representation and building high-performance neural TTS. Haohan Guo, Fenglong Xie, Xixin Wu, Frank K. Soong, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | A Universal Ordinal Regression for Assessing Phoneme-Level PronunciationabstractThe efficacy and robustness of Ordinal Regression (OR) in assessing speech pronunciation for language learning at phrase level has been shown before. However, for assessing phoneme pronunciation, we need to: 1. collect human scoring annotations for phoneme tokens of a short duration (60-70 ms); 2. train an ordinal regression model for each phoneme with the corresponding training and inference costs. In this paper, we propose to train a Universal Ordinal Regression (UOR) model instead of multiple, separate models for different phonemes, and evaluate its performance accordingly. A single universal binary classifier in UOR is trained to make a binary preference decision (better or worse) between a pair of two tokens with the same phoneme ID. In inference, labeled anchored tokens of specific phoneme ID in the training data are paired with test phoneme token to make binary preference decisions. By evaluating the new UOR on Speechocean762, a public speech database for pronunciation evaluation, we show the advantages of the proposed new approach. Improvements of Pearson Correlation Coefficient by 16.7% and Mean Square Error by 25.0%, all relatively, are obtained against the state-of-the-art systems. Shaoguang Mao, Frank K. Soong, Yan Xia 0005, Jonathan Tien |
ICASSP | 2 |
| 2022 | Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward NetworkabstractFastSpeech, as a feed-forward transformer based TTS, can avoid the slow serial, autoregressive inference to generate the target mel-spectrogram in a parallel way. As a non-autoregressive TTS, the latency and computation load in inference is shifted from vocoder to transformer where the efficiency is limited by the quadratic time and memory complexity in the self-attention mechanism, particularly for a long text sequence. To tackle this challenges, We propose two models, ProbSparseFS and LinearizedFS, which have efficient self-attention arrangements to improve the inference speed and memory complexity. LinearizedFS has achieved 3.4x memory savings and 2.1x inference speedup, compared with the those of the baseline FastSpeech. A further optimized LinearizedFS with a lightweight FFN can accelerate the inference speed by 3.6x more. We do subjective voice quality evaluations in MOS and CMOS of News report and Audiobook applications, for multi-speaker and multi-style scenarios. Test results verified that the proposed models yield a TTS quality which is on-par with that of the baseline system but with much better memory efficiency and inference speed. Yujia Xiao, Xi Wang 0016, Lei He 0005, Frank K. Soong |
ICASSP | 4 |
| 2022 | An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) EmbeddingsabstractMany mispronunciation detection and diagnosis (MD&D) research approaches try to exploit both the acoustic and linguistic features as input. Yet the improvement of the performance is limited, partially due to the shortage of large amount annotated training data at the phoneme level. Phonetic embeddings, extracted from ASR models trained with huge amount of word level annotations, can serve as a good representation of the content of input speech, in a noise-robust and speaker-independent manner. These embeddings, when used as implicit phonetic supplementary information, can alleviate the data shortage of explicit phoneme annotations. We propose to utilize Acoustic, Phonetic and Linguistic (APL) embedding features jointly for building a more powerful MD&D system. Experimental results obtained on the L2-ARCTIC database show the proposed approach outperforms the baseline by 9.93%, 10.13% and 6.17% on the detection accuracy, diagnosis error rate and the F-measure, respectively. Wenxuan Ye, Shaoguang Mao, Frank K. Soong, Wenshan Wu, Yan Xia 0005, Jonathan Tien, Zhiyong Wu 0001 |
ICASSP | 3 |
| 2022 | Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual KnowledgeabstractEnd-to-end TTS requires a large amount of speech/text paired data to cover all necessary knowledge, particularly how to pronounce different words in diverse contexts, so that a neural model may learn such knowledge accordingly. But in real applications, such high demand of training data is hard to be satisfied and additional knowledge often needs to be injected manually. For example, to capture pronunciation knowledge on languages without regular orthography, a complicated grapheme-to-phoneme pipeline needs to be built based on a large structured pronunciation lexicon, leading to extra, sometimes high, costs to extend neural TTS to such languages. In this paper, we propose a framework to learn to automatically extract knowledge from unstructured external resources using a novel Token2Knowledge attention module. The framework is applied to build a TTS model named Neural Lexicon Reader that extracts pronunciations from raw lexicon texts in an end-to-end manner. Experiments show the proposed model significantly reduces pronunciation errors in low-resource, end-to-end Chinese TTS, and the lexicon-reading capability can be transferred to other languages with a smaller amount of data. Mutian He 0001, Lei He 0005, Frank K. Soong |
INTERSPEECH | 4 |
| 2022 | A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTSabstractWe propose a Multi-Stage, Multi-Codebook (MSMC) approach to high performance neural TTS synthesis.A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple stages into MSMC Representations (MSMCRs) with different time resolutions, and quantizing them with multiple VQ codebooks, respectively.Multi-stage predictors are trained to map the input text sequence to MSMCRs progressively by minimizing a combined loss of the reconstruction Mean Square Error (MSE) and "triplet loss".In synthesis, the neural vocoder converts the predicted MSM-CRs into final speech waveforms.The proposed approach is trained and tested with an English TTS database of 16 hours by a female speaker.The proposed TTS achieves an MOS score of 4.41, which outperforms the baseline with an MOS of 3.62.Compact versions of the proposed TTS with much less parameters can still preserve high MOS scores.Ablation studies show that both multiple stages and multiple codebooks are effective for achieving high TTS performance. Haohan Guo, Fenglong Xie, Frank K. Soong, Xixin Wu, Helen M. Meng |
INTERSPEECH | 3 |
| 2022 | Disentangling Style and Speaker Attributes for TTS Style TransferabstractEnd-to-end neural TTS has shown improved performance in speech style transfer. However, the improvement is still limited by the available training data in both target styles and speakers. Additionally, degenerated performance is observed when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style. In this paper, we propose a new approach to seen and unseen style transfer training on disjoint, multi-style datasets, i. e., datasets of different styles are recorded, one individual style by one speaker in multiple utterances. An inverse autoregressive flow (IAF) technique is first introduced to improve the variational inference for learning an expressive style representation. A speaker encoder network is then developed for learning a discriminative speaker embedding, which is jointly trained with the rest neural TTS modules. The proposed approach of seen and unseen style transfer is effectively trained with six specifically-designed objectives: reconstruction loss, adversarial loss, style distortion loss, cycle consistency loss, style classification loss, and speaker classification loss. Experiments demonstrate, both objectively and subjectively, the effectiveness of the proposed approach for seen and unseen style transfer tasks. The performance of our approach is superior to and more robust than those of four other reference systems of prior art. Xiaochun An, Frank K. Soong, Lei Xie 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | ParaTTS: Learning Linguistic and Prosodic Cross-Sentence Information in Paragraph-Based TTSabstractRecent advancements in neural end-to-end text-to-speech (TTS) models have shown high-quality, natural synthesized speech in a conventional sentence-based TTS. However, it is still challenging to reproduce similar high quality when a whole paragraph is considered in TTS, where a large amount of contextual information needs to be considered in building a paragraph-based TTS model. To alleviate the difficulty in training, we propose to model linguistic and prosodic information by considering cross-sentence, embedded structure in training. Three sub-modules, including linguistics-aware, prosody-aware and sentence-position networks, are trained together with a modified Tacotron2. Specifically, to learn the information embedded in a paragraph and the relations among the corresponding component sentences, we utilize linguistics-aware and prosody-aware networks. The information in a paragraph is captured by encoders and the inter-sentence information in a paragraph is learned with multi-head attention mechanisms. The relative sentence position in a paragraph is explicitly exploited by a sentence-position network. Trained on a storytelling audio-book corpus (4.08 hours), recorded by a female Mandarin Chinese speaker, the proposed TTS model demonstrates that it can produce rather natural and good-quality speech paragraph-wise. The cross-sentence contextual information, such as break and prosodic variations between consecutive sentences, can be better predicted and rendered than the sentence-based model. Tested on paragraph texts, of which the lengths are similar to, longer than, or much longer than the typical paragraph length of the training data, the TTS speech produced by the new model is consistently preferred over the sentence-based model in subjective tests and confirmed in objective measures. Liumeng Xue, Frank K. Soong, Shaofei Zhang, Lei Xie 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Speech Bert Embedding for Improving Prosody in Neural TTSabstractThis paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). As a pre-trained model, it can learn prosody attributes from a large amount of speech data, which can utilize more data than the original training data used by the target TTS. The embedding is extracted from the previous segment of a fixed length in the proposed BERT. The extracted embedding is then used together with the mel-spectrogram to predict the following segment in the TTS decoder. Experimental results obtained by the Transformer TTS show that the proposed BERT can extract fine-grained, segment-level prosody, which is complementary to utterance-level prosody to improve the final prosody of the TTS speech. The objective distortions measured on a single speaker TTS are reduced between the generated speech and original recordings. Subjective listening tests also show that the proposed approach is favorably preferred over the TTS without the BERT prosody embedding module, for both in-domain and out-of-domain applications. For Microsoft professional, single/multiple speakers and the LJ Speaker in the public database, subjective preference is similarly confirmed with the new BERT prosody embedding. TTS demo audio samples are in https://judy44chen.github.io/TTSSpeechBERT/. Xi Wang 0016, Frank K. Soong, Lei He 0005 |
ICASSP | 4 |
| 2021 | MBNET: MOS Prediction for Synthesized Speech with Mean-Bias NetworkabstractMean opinion score (MOS) is a popular subjective metric to assess the quality of synthesized speech, and usually involves multiple human judges to evaluate each speech utterance. To reduce the labor cost in MOS test, multiple methods have been proposed to automatically predict MOS scores. To our knowledge, for a speech utterance, all previous works only used the average of multiple scores from different judges as the training target and discarded the score of each individual judge, which did not well exploit the precious MOS training data. In this paper, we propose MBNet, a MOS predictor with a mean subnet and a bias subnet to better utilize every judge score in MOS datasets, where the mean subnet is used to predict the mean score of each utterance similar to that in previous works, and the bias subnet to predict the bias score (the difference between the mean score and each individual judge score) and capture the personal preference of individual judges. Experiments show that compared with MOSNet baseline that only leverages mean score for training, MBNet improves the system-level spearmans rank correlation coefficient (SRCC) by 2.9% on VCC 2018 dataset and 6.7% on VCC 2016 dataset. Yichong Leng, Xu Tan 0003, Sheng Zhao 0002, Frank K. Soong, Xiang-Yang Li 0001, Tao Qin 0001 |
ICASSP | 4 |
| 2021 | Improving Pronunciation Assessment Via Ordinal Regression with Anchored Reference SamplesabstractSentence level pronunciation assessment is important for Computer Assisted Language Learning (CALL). Traditional speech pronunciation assessment, based on the Goodness of Pronunciation (GOP) algorithm, has some weakness in assessing a speech utterance: 1) Phoneme GOP scores cannot be easily translated into a sentence score with a simple average for effective assessment; 2) The rank ordering information has not been well exploited in GOP scoring for delivering a robust assessment and correlate well with a human rater’s evaluations. In this paper, we propose two new statistical features, average GOP (aGOP) and confusion GOP (cGOP) and use them to train a binary classifier in Ordinal Regression with Anchored Reference Samples (ORARS). When the proposed approach is tested on Microsoft mTutor ESL Dataset, a relative improvement of Pearson correlation coefficient of 26.9% is obtained over the conventional GOP-based one. The performance is at a human-parity level or better than human raters. Shaoguang Mao, Frank K. Soong, Yan Xia 0005, Jonathan Tien, Zhiyong Wu 0001 |
ICASSP | 3 |
| 2021 | A New High Quality Trajectory Tiling Based Hybrid TTS In Real TimeabstractA trajectory tiling based, hybrid TTS is revisited in this study for improving its synthesis performance. A combination of Transformer encoder and RNN based decoder architecture where two-level, at both word and Chinese phonetic alphabet letter levels, linguistic representation is exploited to generate a cogent and smooth speech parameter trajectory. And then a segment candidate lattice is constructed by minimizing the log spectral distortion of mel-spectrograms and RMSE of F0 between the generated trajectory and candidates. Normalized cross-correlation is used to find the best sequence of "wave-form tiles" in the lattice for synthesizing the final speech waveforms. Subjective A/B preference tests show that the new hybrid system outperforms our earlier trajectory-tiling hybrid baseline TTS (67% vs 11%) and the state-of-the-art, real-time TTS system constructed with Tacotron 2 and LPC-Net (56% vs 27%). Fenglong Xie, Wen-Chao Su, Frank K. Soong |
ICASSP | 5 |
| 2021 | Improving Performance of Seen and Unseen Speech Style Transfer in End-to-End Neural TTSabstractEnd-to-end neural TTS training has shown improved performance in speech style transfer.However, the improvement is still limited by the training data in both target styles and speakers.Inadequate style transfer performance occurs when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style.In this paper, we propose a new approach to style transfer for both seen and unseen styles, with disjoint, multi-style datasets, i.e., datasets of different styles are recorded, each individual style is by one speaker with multiple utterances.To encode the style information, we adopt an inverse autoregressive flow (IAF) structure to improve the variational inference.The whole system is optimized to minimize a weighed sum of four different loss functions: 1) a reconstruction loss to measure the distortions in both source and target reconstructions; 2) an adversarial loss to "fool" a well-trained discriminator; 3) a style distortion loss to measure the expected style loss after the transfer; 4) a cycle consistency loss to preserve the speaker identity of the source after the transfer.Experiments demonstrate, both objectively and subjectively, the effectiveness of the proposed approach for seen and unseen style transfer tasks.The performance of the new approach is better and more robust than those of four baseline systems of the prior art. Xiaochun An, Frank K. Soong, Lei Xie 0001 |
Interspeech | 2 |
| 2021 | Conversational End-to-End TTS for Voice AgentsabstractEnd-to-end neural TTS has achieved excellent performance on reading style speech synthesis. However, it is still a challenge to build a high-quality conversational TTS due to the limitations of corpus and modeling capability. This study aims at building a conversational TTS for a voice agent under sequence to sequence modeling framework. We firstly construct a spontaneous conversational speech corpus well designed for the voice agent with a new recording scheme ensuring both recording quality and conversational speaking style. Secondly, we propose a conversation context-aware end-to-end TTS approach that employs an auxiliary encoder and a conversational context encoder to specifically reinforce the information about the current utterance and its context in a conversation as well. Experimental results show that the proposed approach produces more natural prosody in accordance with the conversational context, with significant preference gains at both utterance-level and conversation-level. Moreover, we find that the model has the ability to express some spontaneous behaviors like fillers and repeated words, which makes the conversational speaking style more realistic. Haohan Guo, Shaofei Zhang, Frank K. Soong, Lei He 0005, Lei Xie 0001 |
SLT | 3 |
| 2021 | Effective and direct control of neural TTS prosody by removing interactions between different attributes
Xiaochun An, Frank K. Soong, Shan Yang 0001, Lei Xie 0001 |
Neural Networks | 2 |
| 2021 | Cycle consistent network for end-to-end style transfer TTS training
Liumeng Xue, Shifeng Pan, Lei He 0005, Lei Xie 0001, Frank K. Soong |
Neural Networks | 5 |
| 2020 | Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density NetworkabstractIn this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully achieved high-quality and lightweight speech synthesis systems by combining a vocal tract LP filter with a WaveRNN-based vocal source (i.e., excitation) generator. However, the quality of synthesized speech is often unstable because the vocal source component is insufficiently represented by the μ-law quantization method, and the model is trained without considering the entire speech production mechanism. To address this problem, we first introduce LP-MDN, which enables the autoregressive neural vocoder to structurally represent the interactions between the vocal tract and vocal source components. Then, we propose to incorporate the LP-MDN to the LPCNet vocoder by replacing the conventional discretized output with continuous density distribution. The experimental results verify that the proposed system provides high quality synthetic speech by achieving a mean opinion score of 4.41 within a text-to-speech framework. Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto, Frank K. Soong, Hong-Goo Kang |
ICASSP | 4 |
| 2020 | Improving Prosody with Linguistic and Bert Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTSabstractRecent advances of neural TTS have made "human parity" synthesized speech possible when a large amount of studio-quality training data from a voice talent is available. However, with only limited, casual recordings from an ordinary speaker, human-like TTS is still a big challenge, in addition to other artifacts like incomplete sentences, repetition of words, etc. Chinese, a language, of which the text is different from that of other roman-letter based languages like English, has no blank space between adjacent words, hence word segmentation errors can cause serious semantic confusions and unnatural prosody. In this study, with a multi-speaker TTS to accommodate the insufficient training data of a target speaker, we investigate linguistic features and Bert-derived information to improve the prosody of our Mandarin Chinese TTS. Three factors are studied: phone-related and prosody-related linguistic features; better predicted breaks with a refined Bert-CRF model; augmented phoneme sequence with character embedding derived from a Bert model. Subjective tests on in- and out-domain tasks of News, Chat and Audiobook, have shown that all factors are effective for improving prosody of our Mandarin TTS. The model with additional character embeddings from Bert is the best one, which outperforms the baseline by 0.17 MOS gain. Yujia Xiao, Lei He 0005, Huaiping Ming, Frank K. Soong |
ICASSP | 4 |
| 2020 | An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training DataabstractA frame-unit-selection based voice conversion system proposed earlier by us is revisited here to enhance its performance in both speech naturalness and speaker similarity. Speaker independent, bilingual (Mandarin Chinese and American English) deep neural net (DNN) acoustic model’s output, frame-level phone posterior probability (PPP), is used to represent the phonetic information. The corresponding frame-level F0 is used as the prosodic information. Kullback-Leibler divergence (KLD) between source and target PPPs (phonetic distortion) and the absolute difference between normalized source and target F0 (prosodic distortion) are used for selecting target frame candidates to construct a search lattice. The optimal target unit trajectory is obtained by Viterbi algorithm which tries to minimize the dynamic acoustic difference between the acoustic trajectory of the source speech and target candidates. The obtained spectral trajectory together with the enhanced pitch period and pitch correlation trajectory are sent to LPCNet vocoder to synthesize the converted waveforms. Compared with the top-rank system in Voice Conversion Challenge 2018, our new system can achieve on-par performance on studio to studio American English VC test, and better performance on non-studio to studio Mandarin Chinese VC test, in both speech naturalness MOS and speaker similarity DMOS. Fenglong Xie, Yibin Zheng, Frank K. Soong |
ICASSP | 7 |
| 2020 | An Efficient Subband Linear Prediction for LPCNet-Based Neural Synthesis
Xi Wang 0016, Lei He 0005, Frank K. Soong |
INTERSPEECH | 4 |
| 2020 | Transfer Learning for Improving Singing-Voice Detection in Polyphonic Instrumental MusicabstractDetecting singing-voice in polyphonic instrumental music is critical to music information retrieval.To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential.However, frame-level labeling is time-consuming and labor expensive, resulting there is little well-labeled dataset available for singing-voice detection (S-VD).Hence, we propose a data augmentation method for S-VD by transfer learning.In this study, clean speech clips with voice activity endpoints and separate instrumental music clips are artificially added together to simulate polyphonic vocals to train a vocal /non-vocal detector.Due to the different articulation and phonation between speaking and singing, the vocal detector trained with the artificial dataset does not match well with the polyphonic music which is singing vocals together with the instrumental accompaniments.To reduce this mismatch, transfer learning is used to transfer the knowledge learned from the artificial speech-plus-music training set to a small but matched polyphonic dataset, i.e., singing vocals with accompaniments.By transferring the related knowledge to make up for the lack of well-labeled training data in S-VD, the proposed data augmentation method by transfer learning can improve S-VD performance with an F-score improvement from 89.5% to 93.2%. Yuanbo Hou, Frank K. Soong, Jian Luan 0001, Shengchen Li |
INTERSPEECH | 2 |
| 2019 | Domain Adversarial Training for Improving Keyword Spotting Performance of ESL SpeechabstractA second language (L2) learner usually cannot speak L2 well in both pronunciations and forming-of-words. Hence his/her L2 speech cannot be well recognized by a recognizer trained with native data. Domain adversarial training (DAT), capable of reducing the acoustic mismatch between training and testing, can be useful for improving speech recognition of L2 learners. To get around the ungrammatical L2 speech in scenario-based conversation training, keyword spotting (KWS) is an effective solution by relaxing the language model constraint in decoding. On the acoustic pronunciation side, DAT is investigated in this study for training a neural net-based acoustic model. DAT model is trained with both native and English as second language (ESL) learners' speech to extract more invariant features from native to ESL speech by equalizing their intrinsic difference. The model is jointly optimized for improved senone classification in training. Testing on ESL learners' speech and native English, the DAT model improves recognition performance which is comparable to jointly trained multi-condition model but significantly improves the performance of native speech recognition. In KWS, DAT shows a consistent better performance than the multi-condition training. The improved performance of proposed model is also obtained without increasing its computation complexity or the model size. Jingyong Hou, Sining Sun, Frank K. Soong, Wenping Hu, Lei Xie 0001 |
ICASSP | 4 |
| 2019 | NN-based Ordinal Regression for Assessing Fluency of ESL SpeechabstractAutomatic assessment of a language learner's speech fluency is highly desirable for language education, e.g. for English as a Second Language (ESL) learning. In this paper, we formulate the fluency assessment as a problem of Ordinal Regression with Anchored Reference Samples (ORARS), where the fluency of a speech utterance is predicted by an ordinal regression neural network (NN) trained with anchored reference samples. The ORARS is trained and tested by: picking human expert labeled samples in each mean opinion score (MOS) bucket as the anchored reference samples and pairing them with input speech samples as training couplets; training an NN-based binary classifier to determine which sample in a pair is better in fluency; predicting the rank (MOS) of a test sample based upon the posteriors of all binary comparisons between the test sample and all anchored reference samples. Experimentally, our proposed approach outperforms the traditional NN-based methods and reaches a performance of "human parity", i.e. as comparable as human experts, in its fluency assessment of collected ESL speech. To the best of our knowledge, this is the first attempt to assess speech fluency with an ordinal regression framework where a test input is paired with bucketed and anchored reference samples. Shaoguang Mao, Zhiyong Wu 0001, Jingshuai Jiang, Peiyun Liu, Frank K. Soong |
ICASSP | 5 |
| 2019 | A Pitch-aware Approach to Single-channel Speech SeparationabstractDespite significant advancements of deep learning on separating speech sources mixed in a single channel, same gender speaker mix, i.e., male-male or female-female, is still more difficult to separate than the case of opposite gender mix. In this study, we propose a pitch-aware speech separation approach to improve the speech separation performance. The proposed approach performs speech separation in the following steps: 1) training a pre-separation model to separate the mixed sources; 2) training a pitch-tracking network to perform polyphonic pitch tracking; 3) incorporating the estimated pitch for the final pitch-aware speech separation. Experimental results of the new approach, tested on the WSJ0-2mix public dataset, show that the new approach improves speech separation performance for both same and opposite gender mixture. The improved performance in signal-to-distortion (SDR) of 12.0 dB is the best reported result without using any phase enhancement. Frank K. Soong, Lei Xie 0001 |
ICASSP | 2 |
| 2019 | A New GAN-Based End-to-End TTS Training AlgorithmabstractEnd-to-end, autoregressive model-based TTS has shown significant performance improvements over the conventional one.However, the autoregressive module training is affected by the exposure bias, or the mismatch between the different distributions of real and predicted data.While real data is available in training, but in testing, only predicted data is available to feed the autoregressive module.By introducing both real and generated data sequences in training, we can alleviate the effects of the exposure bias.We propose to use Generative Adversarial Network (GAN) along with the key idea of Professor Forcing in training.A discriminator in GAN is jointly trained to equalize the difference between real and predicted data.In AB subjective listening test, the results show that the new approach is preferred over the standard transfer learning with a CMOS improvement of 0.1.Sentence level intelligibility tests show significant improvement in a pathological test set.The GAN-trained new model is also more stable than the baseline to produce better alignments for the Tacotron output. Haohan Guo, Frank K. Soong, Lei He 0005, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2019 | Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTSabstractThe end-to-end TTS, which can predict speech directly from a given sequence of graphemes or phonemes, has shown improved performance over the conventional TTS.However, its predicting capability is still limited by the acoustic/phonetic coverage of the training data, usually constrained by the training set size.To further improve the TTS quality in pronunciation, prosody and perceived naturalness, we propose to exploit the information embedded in a syntactically parsed tree where the inter-phrase/word information of a sentence is organized in a multilevel tree structure.Specifically, two key features: phrase structure and relations between adjacent words are investigated.Experimental results in subjective listening, measured on three test sets, show that the proposed approach is effective to improve the pronunciation clarity, prosody and naturalness of the synthesized speech of the baseline system. Haohan Guo, Frank K. Soong, Lei He 0005, Lei Xie 0001 |
INTERSPEECH | 2 |
| 2019 | Forward-Backward Decoding for Regularizing End-to-End TTSabstractNeural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scaling it to challenging test sets. One concern is that the encoder-decoder with attention-based network adopts autoregressive generative sequence model with the limitation of exposure bias To address this issue, we propose two novel methods, which learn to predict future by improving agreement between forward and backward decoding sequence. The first one is achieved by introducing divergence regularization terms into model training objective to reduce the mismatch between two directional models, namely L2R and R2L (which generates targets from left-to-right and right-to-left, respectively). While the second one operates on decoder-level and exploits the future information during decoding. In addition, we employ a joint training strategy to allow forward and backward decoding to improve each other in an interactive process. Experimental results show our proposed methods especially the second one (bidirectional decoder regularization), leads a significantly improvement on both robustness and overall naturalness, as outperforming baseline (the revised version of Tacotron2) with a MOS gap of 0.14 in a challenging test, and achieving close to human quality (4.42 vs. 4.49 in MOS) on general test. Yibin Zheng, Xi Wang 0016, Lei He 0005, Shifeng Pan, Frank K. Soong, Zhengqi Wen, Jianhua Tao 0001 |
INTERSPEECH | 5 |
| 2019 | Voice conversion with SI-DNN and KL divergence based mapping without parallel training data
Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
Speech Commun. | 2 |
| 2018 | Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker VerificationabstractIn this paper, given the speaker bottleneck feature vectors extracted with speaker discriminant neural networks, we focus on using the sequential speaker characteristics for text-dependent speaker verification. In each evaluation trial, speaker supervectors are used as the representations of the sequential speaker characteristics rendered in the compared speech utterances. To this end, dynamic time warping is used to warp the variable-length speaker feature vector sequences of the utterances to the same length. Thereafter for every utterance, a speaker supervector can be obtained as the concatenation of its speaker feature vectors. We use Euclidean distance and support vector machine (SVM) to compute the decision score on the speaker supervectors. Our experiments on a Microsoft internal keyword-spotting database showed the effectiveness of the proposed speaker supervector for text-dependent speaker verification. Moreover, when SVM backend was used in scoring, the speaker supervector achieved the best EER performance 1.627%, better than the combination of i-vector and probabilistic linear discriminant analysis. Yong Zhao 0008, Shixiong Zhang 0001, Guoli Ye, Frank K. Soong |
ICASSP | 6 |
| 2018 | A New Glottal Neural Vocoder for Speech Synthesis
Xi Wang 0016, Lei He 0005, Frank K. Soong |
INTERSPEECH | 4 |
| 2018 | Paired Phone-Posteriors Approach to ESL Pronunciation Quality Assessment
Yujia Xiao, Frank K. Soong, Wenping Hu |
INTERSPEECH | 2 |
| 2017 | Improving native language (L1) identifation with better VAD and TDNN trained separately on native and non-native English corporaabstractIdentifying a speaker's native language (L1), i.e., mother tongue, based upon non-native English (L2) speech input, is both challenging and useful for many human-machine voice interface applications, e.g., computer assisted language learning (CALL). In this paper, we improve our sub-phone TDNN based i-vector approach to L1 recognition with a more accurate TDNN-derived VAD and a highly discriminative classifier. Two TDNNs are separately trained on native and non-native English, LVCSR corpora, for contrasting their corresponding sub-phone posteriors and resultant supervectors. The derived i-vectors are then exploited for improving the performance further. Experimental results on a database of 25 L1s show a 3.1% identification rate improvement, from 78.7% to 81.8%, compared with a high performance baseline system which has already achieved the best published results on the 2016 ComParE corpus of only 11 L1s. The statistical analysis of the features used in our system provides useful findings, e.g. pronunciation similarity among the non-native English speakers with different L1s, for research on second-language (L2) learning and assessment. Yao Qian, Keelan Evanini, Patrick L. Lange, Robert A. Pugh, Rutuja Ubale, Frank K. Soong |
ASRU | 6 |
| 2017 | Perceptual quality and modeling accuracy of excitation parameters in DLSTM-based speech synthesis systemsabstractThis paper investigates how the perceptual quality of the synthesized speech is affected by reconstruction errors in excitation signals generated by a deep learning-based statistical model. In this framework, the excitation signal obtained by an LPC inverse filter is first decomposed into harmonic and noise components using an improved time-frequency trajectory excitation (ITFTE) scheme, then they are trained and generated by a deep long short-term memory (DLSTM)-based speech synthesis system. By controlling the parametric dimension of the ITFTE vocoder, we analyze the impact of the harmonic and noise components to the perceptual quality of the synthesized speech. Both objective and subjective experimental results confirm that the maximum perceptually allowable spectral distortion for the harmonic spectrum of the generated excitation is ~0.08 dB. On the other hand, the absolute spectral distortion in the noise components is meaningless, and only the spectral envelope is relevant to the perceptual quality. Eunwoo Song, Frank K. Soong, Hong-Goo Kang |
ASRU | 2 |
| 2017 | Improving Sub-Phone Modeling for Better Native Language Identification with Non-Native English Speech
Yao Qian, Keelan Evanini, David Suendermann-Oeft, Robert A. Pugh, Patrick L. Lange, Hillary Molloy, Frank K. Soong |
INTERSPEECH | 8 |
| 2017 | Proficiency Assessment of ESL Learner's Sentence Prosody with TTS Synthesized Voice as Reference
Yujia Xiao, Frank K. Soong |
INTERSPEECH | 2 |
| 2017 | DNN i-Vector Speaker Verification with Short, Text-Constrained Test Utterances
Jinghua Zhong, Wenping Hu, Frank K. Soong, Helen M. Meng |
INTERSPEECH | 3 |
| 2017 | Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis SystemsabstractIn this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a finite window length is inadequate to capture the time evolution of the ITFTE parameters. We propose to use the LSTM to exploit the time-varying nature of both trajectories of the excitation and filter parameters, where the LSTM is implemented to use the linguistic text input and to predict both ITFTE and LPC parameters holistically. In the case of LPC parameters, we further enhance the generated spectrum by applying LP bandwidth expansion and line spectral frequency-sharpening filters. These filters are not only beneficial for reducing unstable synthesis filter conditions but also advantageous toward minimizing the muffling problem in the generated spectrum. Experimental results have shown that the proposed LSTM-RNN system with the ITFTE vocoder significantly outperforms both similarly configured band aperiodicity-based systems and our best prior DNN-trainecounterpart, both objectively and subjectively. Eunwoo Song, Frank K. Soong, Hong-Goo Kang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Unsupervised speaker adaptation for DNN-based TTS synthesisabstractMulti-speaker TTS trained with a general DNN has outperformed individually modelled baseline [1]. Multi-speaker DNN takes advantages of larger amount of training data from multiple speakers to find robust transformations in the hidden layers and covers more speaker variability in the output regression layer. In this paper, we propose a new approach to unsupervised speaker adaptation with multi-speaker DNN. It takes advantage of shared hidden transformation to search for the labels of unlabelled acoustic frames and the found labels are used for speaker adaption. Experimental results show that the new approach of unsupervised adaptation can achieve comparable performance with supervised adaptation both objectively and subjectively. We further extend it to cross-lingual adaptation. It can remove non-native accent and improve the naturalness while keep the same speaker's characteristics. Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 3 |
| 2016 | Speaker and language factorization in DNN-based TTS synthesisabstractWe have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, where speaker-specific layers are used for multi-speaker modelling, and shared layers and language-specific layers are employed for multi-language, linguistic feature transformation. Experimental results on a speech corpus of multiple speakers in both Mandarin and English show that the proposed factorized DNN can not only achieve a similar voice quality as that of a multi-speaker DNN, but also perform polyglot synthesis with a monolingual speaker's voice. Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 3 |
| 2016 | A KL divergence and DNN approach to cross-lingual TTSabstractWe propose a Kullback-Leibler divergence (KLD) and deep neural net (DNN) based approach to cross-lingual TTS (CL-TTS) training. A speaker independent DNN (SI-DNN) ASR is used to equalize the speaker difference between a source speaker in L1 and a reference speaker in L2. Two speaker dependent GMM-HMM parametric TTS systems are first trained in the respective languages. The senones sets of the two TTS are matched in the SI-DNN ASR in terms of their output posteriors distributions in KLD. The minimum KLD criterion is used to transform the senones in the source speaker's TTS (L1) to the corresponding "closest" senones in the target language (L2). The new CL-TTS thus trained has been shown to achieve high speaker similarity to the source speaker in L1 while high intelligibility and naturalness are preserved. For untranscribed source speaker's recordings, say, conversational speech, a frame mapping, instead of "senone mapping" is also proposed to achieve a high but slightly inferior CL-TTS. Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
ICASSP | 2 |
| 2016 | Improved Time-Frequency Trajectory Excitation Vocoder for DNN-Based Speech Synthesis
Eunwoo Song, Frank K. Soong, Hong-Goo Kang |
INTERSPEECH | 2 |
| 2016 | A KL Divergence and DNN-Based Approach to Voice Conversion without Parallel Training Sentences
Fenglong Xie, Frank K. Soong, Haifeng Li 0001 |
INTERSPEECH | 2 |
| 2016 | Learning Distributed Word Representations For Bidirectional LSTM Recurrent Neural NetworkabstractPeilu Wang, Yao Qian, Frank K. Soong, Lei He, Hai Zhao. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Peilu Wang, Yao Qian, Frank K. Soong, Lei He 0005 |
HLT-NAACL | 3 |
| 2016 | A deep bidirectional LSTM approach for video-realistic talking head
Lei Xie 0001, Shan Yang 0001, Frank K. Soong |
Multim. Tools Appl. | 5 |
| 2016 | Improving speaker verification performance against long-term speaker variability
Jun Wang 0073, Lantian Li, Thomas Fang Zheng, Frank K. Soong |
Speech Commun. | 5 |
| 2016 | Modeling F0 trajectories in hierarchically structured deep neural networks
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001 |
Speech Commun. | 4 |
| 2016 | A Two-Pass Framework of Mispronunciation Detection and Diagnosis for Computer-Aided Pronunciation TrainingabstractThis paper presents a two-pass framework with discriminative acoustic modeling for mispronunciation detection and diagnoses (MD&D). The first pass of mispronunciation detection does not require explicit phonetic error pattern modeling. The framework instantiates a set of antiphones and a filler model to augment the original phone model for each canonical phone. This guarantees full coverage of all possible error patterns while maximally exploiting the phonetic information derived from the text prompt. The antiphones can be used to detect substitutions. The filler model can detect insertions, and phone skips are allowed to detect deletions. As such, there is no prior assumption on the possible error patterns that can occur. The second pass of mispronunciation diagnosis expands the detected insertions and substitutions into phone networks, and another recognition pass attempts to reveal the phonetic identities of the detected mispronunciation errors. Discriminative training (DT) is applied respectively to the acoustic models of the mispronunciation detection pass and the mispronunciation diagnosis pass. DT effectively separates the acoustic models of the canonical phones and the antiphones. Overall, with DT in both passes of MD&D, the error rate is reduced by 40.4% relative, compared with the maximum likelihood baseline. After DT, the error rates of the respective passes are also lower than those of a strong single-pass baseline with DT by 1.3% and 5.1% relative which are statistically significant. Xiaojun Qian, Helen M. Meng, Frank K. Soong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesisabstractIn DNN-based TTS synthesis, DNNs hidden layers can be viewed as deep transformation for linguistic features and the output layers as representation of acoustic space to regress the transformed linguistic features to acoustic parameters. The deep-layered architectures of DNN can not only represent highly-complex transformation compactly, but also take advantage of huge amount of training data. In this paper, we propose an approach to model multiple speakers TTS with a general DNN, where the same hidden layers are shared among different speakers while the output layers are composed of speaker-dependent nodes explaining the target of each speaker. The experimental results show that our approach can significantly improve the quality of synthesized speech objectively and subjectively, comparing with speech synthesized from the individual, speaker-dependent DNN-based TTS. We further transfer the hidden layers for a new speaker with limited training data and the resultant synthesized speech of the new speaker can also achieve a good quality in term of naturalness and speaker similarity. Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 3 |
| 2015 | Photo-real talking head with deep bidirectional LSTMabstractLong short-term memory (LSTM) is a specific recurrent neural network (RNN) architecture that is designed to model temporal sequences and their long-range dependencies more accurately than conventional RNNs. In this paper, we propose to use deep bidirectional LSTM (BLSTM) for audio/visual modeling in our photo-real talking head system. An audio/visual database of a subject's talking is firstly recorded as our training data. The audio/visual stereo data are converted into two parallel temporal sequences, i.e., contextual label sequences obtained by forced aligning audio against text, and visual feature sequences by applying active-appearance-model (AAM) on the lower face region among all the training image samples. The deep BLSTM is then trained to learn the regression model by minimizing the sum of square error (SSE) of predicting visual sequence from label sequence. After testing different network topologies, we interestingly found the best network is two BLSTM layers sitting on top of one feed-forward layer on our datasets. Compared with our previous HMM-based system, the newly proposed deep BLSTM-based one is better on both objective measurement and subjective A/B test. Frank K. Soong, Lei Xie 0001 |
ICASSP | 3 |
| 2015 | Word embedding for recurrent neural network based TTS synthesisabstractThe current state of the art TTS synthesis can produce synthesized speech with highly decent quality if rich segmental and suprasegmental information are given. However, some suprasegmental features, e.g., Tone and Break (TOBI), are time consuming due to being manually labeled with a high inconsistency among different annotators. In this paper, we investigate the use of word embedding, which represents word with low dimensional continuous-valued vector and being assumed to carry a certain syntactic and semantic information, for bidirectional long short term memory (BLSTM), recurrent neural network (RNN) based TTS synthesis. Experimental results show that word embedding can significantly improve the performance of BLSTM-RNN based TTS synthesis without using features of TOBI and Part of Speech (POS). Peilu Wang, Yao Qian, Frank K. Soong, Lei He 0005 |
ICASSP | 3 |
| 2015 | AA spectral space warping approach to cross-lingual voice transformation in HMM-based TTSabstractThis paper presents a new approach to cross-lingual voice transformation in HMM-based TTS with only the recordings from two monolingual speakers in different languages (e.g. Mandarin and English). We aim to synthesize one speaker's speech in the other language. We regard the spectral space of any speaker to be composed of universal elementary units (i.e. tied-states) of speech in different languages. Our approach first forces the spectral spaces of the two speakers to have the same number of tied-states. Then we find an optimal one-to-one tied-state mapping between the two spectral spaces. Hence, the mapped speech trajectory in the spectral space of the target speaker can be found according to that generated in the spectral space of the reference speaker. Consequently, we can synthesize high-quality speech for the target monolingual speaker's voice in the other language. This can also be used as training data for a new TTS system. Hao Wang 0077, Frank K. Soong, Helen M. Meng |
ICASSP | 2 |
| 2015 | Sequence generation error (SGE) minimization based deep neural networks training for text-to-speech synthesis
Yuchen Fan 0001, Yao Qian, Frank K. Soong, Lei He 0005 |
INTERSPEECH | 3 |
| 2015 | HMM trajectory-guided sample selection for photo-realistic talking head
Frank K. Soong |
Multim. Tools Appl. | 2 |
| 2015 | Improved mispronunciation detection with deep neural network trained acoustic models and transfer learning based logistic regression classifiers
Wenping Hu, Yao Qian, Frank K. Soong |
Speech Commun. | 3 |
| 2014 | A DNN-based acoustic modeling of tonal language and its application to Mandarin pronunciation trainingabstractIn this paper we investigate a Deep Neural Network (DNN) based approach to acoustic modeling of tonal language and assess its speech recognition performance with different features and modeling techniques. Mandarin Chinese, the most widely spoken tonal language, is chosen for testing the tone related ASR performance. Furthermore, the DNN-trained, tone-sensitive model is evaluated in automatic detection of mispronunciation among L2 Mandarin learners. The best DNN-HMM acoustic model of tonal syllable (initial and tonal final), trained with embedded F0 features, has shown improved ASR performance, when compared with the baseline DNN system of 39 MFCC features. The proposed system achieves better ASR performance than the baseline system, i.e., by 32% and 35% in relative tone error rate reduction and 20% and 23% in relative tonal syllable error rate reduction, for female and male speakers, respectively. In a speech database of L2 Mandarin learners (native speakers of European languages), 2% equal error rate reduction, from 27.5% to 25.5%, has been obtained with our DNN-HMM system in detecting mispronunciations, compared with the baseline system. Wenping Hu, Yao Qian, Frank K. Soong |
ICASSP | 3 |
| 2014 | On the training aspects of Deep Neural Network (DNN) for parametric TTS synthesisabstractDeep Neural Network (DNN), which can model a long-span, intricate transform compactly with a deep-layered structure, has recently been investigated for parametric TTS synthesis with a fairly large corpus (33,000 utterances) [6]. In this paper, we examine DNN TTS synthesis with a moderate size corpus of 5 hours, which is more commonly used for parametric TTS training. DNN is used to map input text features into output acoustic features (LSP, F0 and V/U). Experimental results show that DNN can outperform the conventional HMM, which is trained in ML first and then refined by MGE. Both objective and subjective measures indicate that DNN can synthesize speech better than HMM-based baseline. The improvement is mainly on the prosody, i.e., the RMSE of natural and generated F0 trajectories by DNN is improved by 2 Hz. This benefit is likely from the key characteristics of DNN, which can exploit feature correlations, e.g., between F0 and spectrum, without using a more restricted, e.g. diagonal Gaussian probability family. Our experimental results also show: the layer-wise BP pre-training can drive weights to a better starting point than random initialization and result in a more effective DNN; state boundary info is important for training DNN to yield better synthesized speech; and a hyperbolic tangent activation function in DNN hidden layers yields faster convergence than a sigmoidal one. Yao Qian, Yuchen Fan 0001, Wenping Hu, Frank K. Soong |
ICASSP | 4 |
| 2014 | A maximum a Posterior-based reconstruction approach to speech bandwidth expansion in noiseabstractWe propose a novel bandwidth expansion algorithm for extending narrowband speech signal to wideband by exploiting segment examples pre-stored in a speaker independent database. Both narrowband and wideband representation of speech signals are pre-stored in the corpus and they are dynamically chopped into variable length segments. Narrowband segments are used dynamically to explain a given narrowband input sentence while the wideband expanded version of the input sentence is constructed correspondingly. The matching process in the narrowband favors a longer segment patch by the chosen Maximum A Posterior (MAP) criterion. As a result, the multiple choices in matching process are significantly reduced with the MAP criterion in decoding. The approach is further generalized to deal with noise corrupted narrowband input signals and the well-known Vector Taylor Series (VTS) noise adaptation algorithm is incorporated into the matching and bandwidth expansion process. A series of experiments is performed to validate the approach on both clean and noise corrupted narrowband speech where both car noise and babble noise corrupted samples are tested. Hyunson Seo, Hong-Goo Kang, Frank K. Soong |
ICASSP | 3 |
| 2014 | TTS synthesis with bidirectional LSTM based recurrent neural networksabstractFeed-forward, Deep neural networks (DNN)-based text-tospeech (TTS) systems have been recently shown to outperform decision-tree clustered context-dependent HMM TTS systems [1, 4]. However, the long time span contextual effect in a speech utterance is still not easy to accommodate, due to the intrinsic, feed-forward nature in DNN-based modeling. Also, to synthesize a smooth speech trajectory, the dynamic features are commonly used to constrain speech parameter trajectory generation in HMM-based TTS [2]. In this paper, Recurrent Neural Networks (RNNs) with Bidirectional Long Short Term Memory (BLSTM) cells are adopted to capture the correlation or co-occurrence information between any two instants in a speech utterance for parametric TTS synthesis. Experimental results show that a hybrid system of DNN and BLSTM-RNN, i.e., lower hidden layers with a feed-forward structure which is cascaded with upper hidden layers with a bidirectional RNN structure of LSTM, can outperform either the conventional, decision tree-based HMM, or a DNN TTS system, both objectively and subjectively. The speech trajectory generated by the BLSTM-RNN TTS is fairly smooth and no dynamic constraints are needed. Yuchen Fan 0001, Yao Qian, Fenglong Xie, Frank K. Soong |
INTERSPEECH | 4 |
| 2014 | Sequence error (SE) minimization training of neural network for voice conversionabstractNeural network (NN) based voice conversion, which employs a nonlinear function to map the features from a source to a target speaker, has been shown to outperform GMM-based voice conversion approach [4-7]. However, there are still limitations to be overcome in NN-based voice conversion, e.g. NN is trained on a Frame Error (FE) minimization criterion and the corresponding weights are adjusted to minimize the error squares over the whole source-target, stereo training data set. In this paper, we use the idea of sentence optimization based, minimum generation error (MGE) training in HMM-based TTS synthesis, and modify the FE minimization to Sequence Error (SE) minimization in NN training for voice conversion. The conversion error over a training sentence from a source speaker to a target speaker is minimized via a gradient descent-based, back propagation (BP) procedure. Experimental results show that the speech converted by the NN, which is first trained with frame error minimization and then refined with sequence error minimization, sounds subjectively better than the converted speech by NN trained with frame error minimization only. Scores on both naturalness and similarity to the target speaker are improved. Index Terms: voice conversion, neural network, pre-training, sequence error minimization Fenglong Xie, Yao Qian, Yuchen Fan 0001, Frank K. Soong, Haifeng Li 0001 |
INTERSPEECH | 4 |
| 2014 | Modeling DCT parameterized F0 trajectory at intonation phrase level with DNN or decision tree
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001 |
INTERSPEECH | 4 |
| 2013 | A fast table lookup based, statistical model driven non-uniform unit selection TTSabstractFor multi-channel TTS applications, e.g. in a cloud service, it is highly desirable that high quality speech can be synthesized in low complexity. In this paper, we propose a fast table lookup based, statistical model driven approach to non-uniform unit selection TTS for that purpose. In TTS training, the voice font of all waveform segments is organized as a Gaussian kernel coded hash table and a table for storing quantized costs of all possible concatenation segment pairs. In synthesis, waveform segments with non-uniform lengths are first selected to construct a candidate lattice by looking up the Gaussian kernel coded hash table, and the best path is searched in the lattice by minimizing the accumulated concatenation scores, which are retrieved from the quantization table for possible concatenations. Experimental results show that the new approach can significantly reduce the search complexity while keep a high TTS voice quality. Yao Qian, Frank K. Soong, Yundi Qian |
ICASSP | 2 |
| 2013 | A new DNN-based high quality pronunciation evaluation for computer-aided language learning (CALL)
Wenping Hu, Yao Qian, Frank K. Soong |
INTERSPEECH | 3 |
| 2013 | A source-filter based adaptive harmonic model and its application to speech prosody modification
JeeSok Lee, Frank K. Soong, Hong-Goo Kang |
INTERSPEECH | 2 |
| 2013 | Binocular photometric stereo acquisition and reconstruction for 3d talking head applicationsabstractIn order to render a high quality, versatile 3D talking head, a stable, high frame rate AV data acquisition system is con-structed. It can capture 3D position, surface orientation and albedo texture of the talking head video images along with the corresponding speech signals. The system consists of a com-puter controlled LED lighting subsystem; high speed stereo cameras; a microphone; and a computer for synchronous re-cording of multi-stream AV data. The visual image data col-lected is processed through a binocular photometric stereo 3D reconstruction pipeline. The pipeline automatically segments out the face; computes the depth map with binocular stereo; computes the normal map with photometric stereo; generates albedo texture; and finally constructs a high-detailed 3d model with depth and normal cues as constraints. By using the data collected with the built system, we can capture high quality dynamic facial performance, synchronized with the subject’s uttered speech. Index Terms: talking head, binocular photometric stereo, fa-cial performance capture Chaoyang Wang 0001, Yasuyuki Matsushita, Bojun Huang, Magnetro Chen, Frank K. Soong |
INTERSPEECH | 6 |
| 2013 | A new language independent, photo-realistic talking head driven by voice onlyabstractWe propose a new photo-realistic, voice driven only (i.e. no linguistic info of the voice input is needed) talking head. The core of the new talking head is a context-dependent, multilayer, Deep Neural Network (DNN), which is discriminatively trained over hundreds of hours, speaker independent speech data. The trained DNN is then used to map acoustic speech input to 9,000 tied “senone” states probabilistically. For each photo-realistic talking head, an HMM-based lips motion synthesizer is trained over the speaker’s audio/visual training data where states are statistically mapped to the corresponding lips images. In test, for given speech input, DNN predicts the likely states in their posterior probabilities and photo-realistic lips animation is then rendered through the DNN predicted state lattice. The DNN trained on English, speaker independent data has also been tested with other language input, e.g. Mandarin, Spanish, etc. to mimic the lips movements cross-lingually. Subjective experiments show that lip motions thus rendered for 15 non-English languages are highly synchronized with the audio input and photo-realistic to human eyes perceptually. Xinjian Zhang, Gang Li 0012, Frank Seide, Frank K. Soong |
INTERSPEECH | 5 |
| 2013 | A Unified Trajectory Tiling Approach to High Quality Speech RenderingabstractIt is technically challenging to make a machine talk as naturally as a human so as to facilitate “frictionless” interactions between machine and human. We propose a trajectory tiling-based approach to high-quality speech rendering, where speech parameter trajectories, extracted from natural, processed, or synthesized speech, are used to guide the search for the best sequence of waveform “tiles” stored in a pre-recorded speech database. We test the proposed unified algorithm in both Text-To-Speech (TTS) synthesis and cross-lingual voice transformation applications. Experimental results show that the proposed trajectory tiling approach can render speech which is both natural and highly intelligible. The perceived high quality of rendered speech is also confirmed in both objective and subjective evaluations. Yao Qian, Frank K. Soong, Zhijie Yan |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Improved minimum converted trajectory error training for real-time speech-to-lips conversionabstractGaussian mixture model (GMM) based speech-to-lips conversion often operates in two alternative ways: batch conversion and sliding window-based conversion for real-time processing. Previously, Minimum Converted Trajectory Error (MCTE) training has been proposed to improve the performance of batch conversion. In this paper, we extend previous work and propose a new training criteria, MCTE for Real-time conversion (R-MCTE), to explicitly optimize the quality of sliding window-based conversion. In R-MCTE, we use the probabilistic descent method to refine model parameters by minimizing the error on real-time converted visual trajectories over training data. Objective evaluations on the LIPS 2008 Visual Speech Synthesis Challenge data set shows that the proposed method achieves both good lip animation performance and low delay in real-time conversion. Frank K. Soong |
ICASSP | 3 |
| 2012 | High quality lip-sync animation for 3D photo-realistic talking headabstractWe propose a new 3D photo-realistic talking head with high quality, lip-sync animation. It extends our prior high-quality 2D photo-realistic talking head to 3D. An a/v recording of a person speaking a set of prompted sentences with good phonetic coverage for ~20-minutes is first made. We then use a 2D-to-3D reconstruction algorithm to automatically adapt a general 3D head mesh model to the person. In training, super feature vectors consisting of 3D geometry, texture and speech are augmented together to train a statistical, multi-streamed, Hidden Markov Model (HMM). The HMM is then used to synthesize both the trajectories of head motion animation and the corresponding dynamics of texture. The resultant 3D talking head animation can be controlled by the model predicted geometric trajectory while the articulator movements, e.g., lips, are rendered with dynamic 2D texture image sequences. Head motions and facial expression can also be separately controlled by manipulating corresponding parameters. In a real-time demonstration, the life-like 3D talking head can take any input text, convert it into speech and render lip-synced speech animation photo-realistically. Frank K. Soong |
ICASSP | 3 |
| 2012 | Modeling pitch trajectory by hierarchical HMM with minimum generation error trainingabstractA hierarchical pitch model (HPM) was recently proposed to HMM-based speech synthesis. In HPM, pitch trajectory is modeled as an additive combination of hierarchical layers (including state, phone, syllable, etc), and a minimum generation error (MGE) criterion is used to re-estimate model parameters. In this paper, we extend the MGE criterion to a tree-based model clustering process to simultaneously cluster the context-dependent models at all layers, and construct a full MGE training process for HPM training. Experiments were conducted to investigate the effects of HPM with different training criteria and different hierarchical layer combinations. Experimental results show that the full MGE training can significantly improve HPM's ability to predict F0 trajectory in TTS over the ML-based approach on test data. The new HPM also outperforms the conventional state-level HMM in F0 prediction. Yi-Jian Wu, Frank K. Soong |
ICASSP | 2 |
| 2012 | Noise estimation using a constrained sequential HMM IN log-spectral domainabstractHow to utilize the time correlation of speech/nonspeech presence is a crucial problem faced by noise estimators. The popular technique of exploiting such correlation is to smooth noisy spectra by using a temporal recursive filter with a time-varying smoothing factor. But this technique cannot warrant the statistical optimality. In theory, hidden Markov model (HMM) is more desirable than this technique. It can give an elaborate description of speech/nonspeech transition. Moreover, some theoretical frameworks, such as maximum likelihood (ML), are available for optimal estimation. This paper presents a constrained sequential HMM to model the time correlation of speech/nonspeech presence of an individual log-power sequence. Its parameter set is on-line adapted to varying signals based on a ML framework. We compared its performance with that of well-established algorithms by speech enhancement experiments. The results confirmed its promising performance. Dongwen Ying, Xugang Lu, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
ICASSP | 6 |
| 2012 | Turning a Monolingual Speaker into Multilingual for a Mixed-language TTS
Yao Qian, Frank K. Soong, Sheng Zhao 0002 |
INTERSPEECH | 3 |
| 2012 | The Use of DBN-HMMs for Mispronunciation Detection and Diagnosis in L2 English to Support Computer-Aided Pronunciation TrainingabstractThis paper investigates acoustic modeling using the hybrid DBN-HMM framework in mispronunciation detection and diagnosis of L2 English. This is one of the first efforts that compare the performance of DBN-HMM with that of the best-tuned GMM-HMM trained in ML and MWE on the same set of features. Previous work in ASR has also shown the necessity of unsupervised pre-training for DBNs to work well. We explore further the effect of training our ASR engine in an unsupervised manner with additional unannotated L2 data from the test speakers. This is compared with the original ASR that has been trained with annotated data in a supervised manner. Experiments show that DBN-HMM can give significant improvement (between 13-18 % relative in word pronunciation error rate) but is computationally more expensive. Index Terms: mispronunciation detection and diagnosis, restricted boltzmann machine, deep belief network 1. Xiaojun Qian, Helen M. Meng, Frank K. Soong |
INTERSPEECH | 3 |
| 2012 | Objective Intelligibility Assessment of Text-to-Speech System using Template Constrained Generalized Posterior Probability
Linfang Wang, Zhe Geng, Frank K. Soong |
INTERSPEECH | 5 |
| 2012 | Constrained Multichannel Speech Dereverberation
Meng Yu 0003, Frank K. Soong |
INTERSPEECH | 2 |
| 2012 | Tip tap tones: mobile microtraining of mandarin soundsabstractLearning a second language is hard, especially when the learner's brain must be retrained to identify sounds not present in his or her native language. It also requires regular practice, but many learners struggle to find the time and motivation. Our solution is to break down the challenge of mastering a foreign sound system into minute-long episodes of "microtraining" delivered through mobile gaming. We present the example of Tip Tap Tones - a mobile game with the purpose of helping learners acquire the tonal sound system of Mandarin Chinese. In a 3-week, 12-user study of this system, we found that an average of 71 minutes' gameplay significantly improved tone identification by around 25%, regardless of whether the underlying sounds had been used to train tone perception. Overall, results suggest that mobile microtraining is an efficient, effective, and enjoyable way to master the sounds of Mandarin Chinese, with applications to other languages and domains. Darren Edge, Kai-Yin Cheng, Michael Whitney, Yao Qian, Zhijie Yan, Frank K. Soong |
Mobile HCI | 6 |
| 2011 | Improved F0 modeling and generation in voice conversionabstractF0 is an acoustic feature that varies largely from one speaker to an other. F0 is characterized by a discontinuity in the transition between voiced and unvoiced sounds that presents an obstacle to GMM modeling for use in voice conversion. A Multi-Space Distribution (MSD) [5] can be used to model unvoiced and voiced F0 regions in a linearly weighted mixture. However, the use of two incompatible probabilistic spaces, for example a continuous probability density for voiced observations, and a discrete probability for unvoiced observations, may result in an imprecise voiced/unvoiced (v/u) conversion in a maximum likelihood (ML) sense. In this paper we propose to use voicing strength, characterized by the normalized correlation coefficient magnitude, as calculated from F0 feature extraction, as an additional feature for improving F0 modeling and the v/u decision in the context of voice conversion. The proposed method was evaluated on male-to-female voice conversion tasks in both Mandarin and English. Objective tests showed that the approach is effective in reducing the Root Mean Square Error, while the results for subjective metrics including AB preference and ABX speaker similarity tests also showed gains. Aki Kunikoshi, Yao Qian, Frank K. Soong, Nobuaki Minematsu |
ICASSP | 3 |
| 2011 | Speaker characterization using spectral subband energy ratio based on Harmonic plus Noise ModelabstractThis paper proposes a feature extraction for speaker characterization by exploring the relationship between the two distinct components of the speech signal, one is harmonics accounting for the periodicity of the signal and the other is modulated noise accounting for the turbulences of the glottal airflow. The harmonic and noise parts of the speech signal are decomposed based on the Harmonic plus Noise Model approach. We estimate the spectral subband energy ratios (SSERs) as the speaker characteristic features, which are expected to reflect the interaction property of the vocal tract and glottal airflow of individual speakers for speaker verification. The speaker verification experiments based on a GMM-UBM system have shown the efficiency of the SSER features, reducing the error equal rate by 27.2% by combining with the conventional MFCC features. Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo |
ICASSP | 3 |
| 2011 | A frame mapping based HMM approach to cross-lingual voice transformationabstractCross-lingual voice transformation is challenging when source language (L1) and target language (L2) are very different in corresponding phonetics and prosodies. We propose a frame mapping based HMM approach to this problem. The source speaker's speech data is first warped in frequency toward the target speaker by mapping corresponding formants of selected vowels. The parameter trajectories of the warped data are then "tiled" with the frames in target speaker's L2 data. The tiled new trajectories then form a simulated training set of target speaker in L1 and it is used to train an HMM TTS. With a bilingual (Mandarin and English) source speaker and a monolingual (English) target speaker, the frame mapping-based approach is capable of generating highly intelligible, good quality speech data in L1 (Mandarin), which sounds rather close to the target speaker. The good performance of the cross-lingual voice transformation is confirmed with speaker similarity, naturalness and intelligibility evaluations subjectively. Yao Qian, Frank K. Soong |
ICASSP | 3 |
| 2011 | Synthesizing visual speech trajectory with minimum generation errorabstractIn this paper, we propose a minimum generation error (MGE) training method to refine the audio-visual HMM to improve visual speech trajectory synthesis. Compared with the traditional maximum likelihood (ML) estimation, the proposed MGE training explicitly optimizes the quality of generated visual speech trajectory, where the audio-visual HMM modeling is jointly refined by using a heuristic method to find the optimal state alignment and a probabilistic descent algorithm to optimize the model parameters under the MGE criterion. In objective evaluation, compared with the ML-based method, the proposed MGE-based method achieves consistent improvement in the mean square error reduction, correlation increase, and recovery of global variance. It also improves the naturalness and audio-visual consistency perceptually in the subjective test. Yi-Jian Wu, Xiaodan Zhuang, Frank K. Soong |
ICASSP | 4 |
| 2011 | A Sparse and Low-rank approach to efficient face alignment for photo-real talking head synthesisabstractIn this paper, we propose a framework for practical large-scale face alignment, based on the recent development of Robust Alignment by Sparse and Low-rank Decomposition for linearly correlated images (RASL). Unfortunately, the original implementation is not applicable in large image dataset. We extend this technique to deal with the situation with millions of images, with the aid of l1-regularized least squares. Our proposal is applied onto the photo-real talking head, a challenging application which requires highly precise alignments of faces from video sequences. We verify the efficacy of our algorithm with experiments using real talking head data. Our method attains comparable quality to RASL in the experiments. King Keung Wu, Frank K. Soong, Yeung Yam |
ICASSP | 3 |
| 2011 | Improvements in Speaker Characterization Using Spectral Subband Energy Based on Harmonic plus Noise Model
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo |
INTERSPEECH | 3 |
| 2011 | A New Phonetic Candidate Generator for Improving Search Query Efficiency
Yao Qian, Frank K. Soong |
INTERSPEECH | 3 |
| 2011 | On Mispronunciation Lexicon Generation Using Joint-Sequence Multigrams in Computer-Aided Pronunciation Training (CAPT)abstractWe investigate the use of joint-sequence multigrams to generate L2 mispronunciation lexicons for mispronunciation detection and diagnosis. In the joint-sequence framework, a pair of parallel strings (namely, the input string of either graphemes or phonemes of the canonical pronunciation and the phonetic string of the mispronunciation) are aligned to form joint units for probabilistic estimation. We compare results on lexicons produced by phoneme-to-mispronunciation conversion and those by grapheme-to-mispronunciation conversion. Results reflect the hypothesized advantage (1.1 % reduction in expected miss rate) in unifying phonetic confusion due to L1 negative transfer with those due to grapheme-to-phoneme errors. The impact of mispronunciation by mis-use of analogy is also studied. Recognition results show the benefit of a lexicon with proper priors. Index Terms: mispronunciation detection and diagnosis, lexicon extension, joint-sequence multigrams Xiaojun Qian, Helen M. Meng, Frank K. Soong |
INTERSPEECH | 3 |
| 2011 | Text Driven 3D Photo-Realistic Talking Head
Frank K. Soong, Qiang Huo |
INTERSPEECH | 3 |
| 2011 | Improved Prosody Generation by Maximizing Joint Probability of State and Longer UnitsabstractThe current state-of-the-art hidden Markov model (HMM)-based text-to-speech (TTS) can produce highly intelligible, synthesized speech with decent segmental quality. However, its prosody, especially at phrase or sentence level, still tends to be bland. This blandness is partially due to the fact that the state-based HMM is inadequate in capturing global, hierarchical suprasegmental information in speech signals. In this paper, to improve the TTS prosody, longer units are first explicitly modeled with appropriate parametric distributions. The resultant models are then integrated with the state-based baseline models in generating better prosody by maximizing the joint probability. Experimental results in both Mandarin and English show consistent improvements over our baseline system with only state-based prosody model. The improvements are both objectively measurable and subjectively perceivable. Yao Qian, Zhizheng Wu 0001, Boyang Gao, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Voice Activity Detection Based on an Unsupervised Learning FrameworkabstractHow to construct models for speech/nonspeech discrimination is a crucial point for voice activity detectors (VADs). Semi-supervised learning is the most popular way for model construction in conventional VADs. In this correspondence, we propose an unsupervised learning framework to construct statistical models for VAD. This framework is realized by a sequential Gaussian mixture model. It comprises an initialization process and an updating process. At each subband, the GMM is firstly initialized using EM algorithm, and then sequentially updated frame by frame. From the GMM, a self-regulatory threshold for discrimination is derived at each subband. Some constraints are introduced to this GMM for the sake of reliability. For the reason of unsupervised learning, the proposed VAD does not rely on an assumption that the first several frames of an utterance are nonspeech, which is widely used in most VADs. Moreover, the speech presence probability in the time-frequency domain is a byproduct of this VAD. We tested it on speech from TIMIT database and noise from NOISEX-92 database. The evaluations effectively showed its promising performance in comparison with VADs such as ITU G.729B, GSM AMR, and a typical semi-supervised VAD. Dongwen Ying, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2010 | RIch-context Unit Selection (RUS) approach to high quality TTSabstractThis paper presents a Rich-context Unit Selection (RUS) approach to high quality speech synthesis. Based upon our previous work on rich context modeling, we use the corresponding parametric HMMs to represent waveform units and form a “sausage-like” lattice. A prune-and-search procedure is proposed, in which Kullback-Leibler divergence is adopted to select potential candidate units, and normalized cross-correlation is used as the final objective measure to search for the optimal unit path. The maximum cross-correlation criterion provides the optimal concatenation between successive units, in terms of spectral similarity, phase continuity and best connecting timing instants. Subjectively, both preference and MOS tests were conducted to compare RUS with our current Weight-table based Unit Selection (WUS) synthesis. Experimental results show that the voice quality of synthesized speech is significantly improved by RUS over the conventional WUS. Zhijie Yan, Yao Qian, Frank K. Soong |
ICASSP | 3 |
| 2010 | Improved modeling for F0 generation and V/U decision in HMM-based TTSabstractThe HMM-based TTS can produce a highly intelligible and decent quality voice. However, sometimes the synthesized speech exhibits perceptibly annoying glitches due to F0 extraction errors in the training data and voiced/unvoiced swapping errors in F0 generation. In the conventional MSD based F0 modeling [10], the dual but incompatible two probabilistic spaces, the continuous probability density for voiced observations or the discrete probability for unvoiced observations, prevent us from using likelihood based frame occupancy to alleviate the deteriorating effect of F0 extraction errors in training a more robust model for synthesis. In this paper, we propose a new approach to improved modeling the piece-wise continuous F0 trajectory and v/u decision for HMM-based TTS. Voicing strength, characterized by the normalized correlation coefficient magnitude calculated in F0 feature extraction, is used as an additional feature in F0 modeling and for v/u decision. Experimental results show the new approach to F0 modeling and generation outperforms MSD-HMM method and a newly proposed GTD-HMM method [9] significantly. The improvements are both objectively measurable and subjectively perceivable. Frank K. Soong, Yao Qian, Zhijie Yan, Jielin Pan, Yonghong Yan 0002 |
ICASSP | 2 |
| 2010 | Cross-validation based decision tree clustering for HMM-based TTSabstractIn HMM-based speech synthesis, we usually use complex, context dependent models to characterize prosodically and linguistically rich speech units. It is therefore difficult to prepare training data which can cover all combinatorial possibilities of contexts. A common approach to cope with this insufficient training data problem is to build a clustered tree via the MDL criterion. However, an MDL-based tree still tends to be inadequate in its power to predict unseen data. In this paper, we adopt the cross-validation principle to build such a decision tree to minimize the generation error of unseen contexts. An efficient training algorithm is implemented by exploiting the sufficient statistics. Experimental results show that the proposed method can achieve better speech synthesis results, both objectively and subjectively, than the baseline results of the MDL-based decision tree. Yu Zhang 0007, Zhijie Yan, Frank K. Soong |
ICASSP | 3 |
| 2010 | A perceptual study of acceleration parameters in HMM-based TTS
Zhijie Yan, Frank K. Soong |
INTERSPEECH | 3 |
| 2010 | A hierarchical F0 modeling method for HMM-based speech synthesis
Yi-Jian Wu, Frank K. Soong, Zhen-Hua Ling, Li-Rong Dai 0001 |
INTERSPEECH | 3 |
| 2010 | Discriminative acoustic model for improving mispronunciation detection and diagnosis in computer-aided pronunciation training (CAPT)abstractIn this study, we propose a discriminative training algorithm to jointly minimize mispronunciation detection errors (i.e., false rejections and false acceptances) and diagnosis errors (i.e., correctly pinpointing mispronunciations but incorrectly stating how they are wrong). An optimization procedure, similar to Minimum Word Error (MWE) discriminative training, is developed to refine the ML-trained HMMs. The errors to be minimized are obtained by comparing transcribed training utterances (including mispronunciations) with Extended Recognition Networks [3] which contain both canonical pronunciations and explicitly modeled mispronunciations. The ERN is compiled by handcrafted rules, or data-driven rules. Several conclusions can be drawn from the experiments: (1) data-driven rules are more effective than hand-crafted ones in capturing mispronunciations; (2) compared with the ML training baseline, discriminative training can reduce false rejections and diagnostic errors, though false acceptances increase slightly due to a small number of false-acceptance samples in the training set. Xiaojun Qian, Frank K. Soong, Helen M. Meng |
INTERSPEECH | 2 |
| 2010 | An HMM trajectory tiling (HTT) approach to high quality TTS
Yao Qian, Zhijie Yan, Yi-Jian Wu, Frank K. Soong, Xin Zhuang, Shengyi Kong |
INTERSPEECH | 4 |
| 2010 | Synthesizing photo-real talking head via trajectory-guided sample selectionabstractIn this paper, we propose an HMM trajectory-guided, real image sample concatenation approach to photo-real talking head synthesis. It renders a smooth and natural video of articulators in sync with given speech signals. An audio-visual database is used to train a statistical Hidden Markov Model (HMM) of lips movement first and the trained model is then used to generate a visual parameter trajectory of lips movement for given speech signals, all in the maximum likelihood sense. The HMM generated trajectory is then used as a guide to select, in the original training database, an optimal sequence of mouth images which are then stitched back to a background head video. The whole procedure is fully automatic and data driven. With an audio/video footage as short as 20 minutes from a speaker, the proposed system can synthesize a highly photo-real video in sync with the given speech signals. This system won the FIRST place in the Audio-Visual match contest in LIPS2009 Challenge, which was perceptually evaluated by recruited human subjects. Index Terms: visual speech synthesis, photo-real, talking head, trajectory-guided Xiaojun Qian, Frank K. Soong |
INTERSPEECH | 4 |
| 2010 | Formant-based frequency warping for improving speaker adaptation in HMM TTS
Xin Zhuang, Yao Qian, Frank K. Soong, Yi-Jian Wu |
INTERSPEECH | 3 |
| 2010 | A minimum converted trajectory error (MCTE) approach to high quality speech-to-lips conversionabstractHigh quality speech-to-lips conversion, investigated in this work, ren-ders realistic lips movement (video) consistent with input speech (audio) without knowing its linguistic content. Instead of memoryless frame-based conversion, we adopt maximum likelihood estimation of the vi-sual parameter trajectories using an audio-visual joint Gaussian Mixture Model (GMM). We propose a minimum converted trajectory error ap-proach (MCTE) to further refine the converted visual parameters. First, we reduce the conversion error by training the joint audio-visual GMM with weighted audio and visual likelihood. Then MCTE uses the gen-eralized probabilistic descent algorithm to minimize a conversion error of the visual parameter trajectories defined on the optimal Gaussian ker-nel sequence according to the input speech. We demonstrate the effec-tiveness of the proposed methods using the LIPS 2009 Visual Speech Synthesis Challenge dataset, without knowing the linguistic (phonetic) content of the input speech. Index Terms: visual speech synthesis, speech-to-lips conversion, mini-mum conversion error, minimum generation error Xiaodan Zhuang, Frank K. Soong, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2009 | Improving mispronunciation detection using machine learningabstractIn this paper, we investigate the problem of mispronunciation detection by considering the influence of speaker and syllables. Machine learning techniques are used to make our method more convenient and flexible for new features, such as syllables normalization. The experimental results on our database, consisting of 9898 syllables pronounced by 100 speakers, show the effectiveness of our method by reducing the average false acceptance rate (FAR) by 42.5% using data set generated by model without adaptation to observation set and reducing average FAR by 32.5% using data set generated by model with adaptation to observation set. Yuqiang Chen, Chao Huang 0011, Frank K. Soong |
ICASSP | 3 |
| 2009 | State mapping for cross-language speaker adaptation in TTSabstractCross-language speaker adaptation has many interesting applications, e.g. speech-to-speech translation. However, in cross-language speaker adaptation, a common phoneme set, assumed to be used by different speakers of the same language, does not exist any longer. Instead, a nearest neighbor based phoneme mapping from one language to the other has been adopted. In this study, we used our recently proposed sub-phonemic HMM state mapping for cross-language adaptations. The sub-phonemic HMM states, due to their phonetic segment nature, tend to be more sharable across different languages than phonemes. Kullback-Leibler divergence, an information-theoretic measure, is chosen here to measure the similarity between given states in different languages. Experimental results show that new state mapping outperforms the phoneme mapping baseline system in terms of three objective measures: log spectral distance, F0 adaptation error and F0 correlations. In comparing with intra-language adaptation, the cross-language result of the new algorithm is also fairly decent. Yao Qian, Frank K. Soong |
ICASSP | 4 |
| 2009 | Improved prosody generation by maximizing joint likelihood of state and longer unitsabstractThe current state-of-art HMM-bsed TTS can produce highly intelligible output speech and deliver a decent segmental quality. However, its prosody, especially at the phrase or sentence level, tends to be bland. The blandness of synthesized prosody is partially due to the fact that a state-based HMM is rather inadequate in modeling a global, hierarchical prosodic structure at a sentence or phrase level. In this study, the prosody of longer units are first modeled explicitly by appropriate parametric distributions. The resultant models are then integrated with the state-level baseline models to generate an optimal prosody by maximizing the joint likelihood of all, from state to longer, units. Experimental results in both Mandarin and English show consistent improvements over the state-based baseline system. The improvements are both objectively measurable and subjectively perceivable. Yao Qian, Zhizheng Wu 0001, Frank K. Soong |
ICASSP | 3 |
| 2009 | An evidence framework for Bayesian learning of continuous-density hidden Markov modelsabstractWe present an evidence Bayesian framework, which can learn both the prior distributions and posterior distributions from data, for continuous-density hidden Markov models (CDHMM). The goal of this study is to build the regularized CDHMMs to improve model generalization, and achieve desirable recognition performance for unknown test speech. Under this framework, we develop an EM iterative procedure to estimate the marginal distribution or the evidence function for exponential family distributions. By adopting the variational Bayesian inference, we derive an empirical Bayesian solution to CDHMM parameters and their hyperparameters. Such a regularized CDHMM compensates the model uncertainty and the ill-posed conditions. Compared with maximum likelihood (ML) or other Bayesian approaches with heuristic hyperparameters, the proposed approach can utilize available data more effectively. The experiments on noisy speech recognition using Aurora2 show that the proposed Bayesian approach performs better than the baseline ML CDHMMs especially with mismatched test data or limited training data. Yu Zhang 0007, Peng Liu 0001, Jen-Tzung Chien, Frank K. Soong |
ICASSP | 4 |
| 2009 | Model-based speech separation: identifying transcription using orthogonalityabstractSpectral envelopes and harmonics are the building elements of a speech signal. By estimating these elements, individual speech sources in a mixture observation can be reconstructed and hence separated. Transcription gives the spoken content. More important, it describes the expected sequence of spectral envelopes, if modeling of different speech sounds is acquired. Our recently proposed single-microphone speech separation algorithm exploits this to derive the spectral envelope trajectories of individual sources and remove interference accordingly. The correctness of such transcription becomes critical to the separation performance. This paper investigates the relationship between the correctness of transcription hypotheses and the orthogonality of associated source estimates. An orthogonality measure is introduced to quantify the correlation between spectrograms. Experiments verify that underlying true transcriptions lead to a salient orthogonality distribution, which is distinguishable from the counterfeit transcription one. Accordingly a transcription identification technique is developed, which succeeds in identifying true transcriptions in 99.74% of the experimental trials. 1 Siu Wa Lee, Frank K. Soong, Tan Lee |
INTERSPEECH | 2 |
| 2009 | A minimum v/u error approach to F0 generation in HMM-based TTS
Yao Qian, Frank K. Soong, Zhizheng Wu 0001 |
INTERSPEECH | 2 |
| 2009 | Auto-checking speech transcriptions by multiple template constrained posteriorabstractChecking transcription errors in speech database is an important but tedious task that traditionally requires intensive manual labor. In [9], Template Constrained Posterior (TCP) was proposed to automate the checking process by screening potential erroneous sentences with a single context template. However, single template-based method is not robust and requires parameter optimization that still involves some manual work. In this work, we propose to use multiple templates which is more robust and requires no development data for parameter optimization. By using its multiple hypothesis sifting capabilities -from well-defined, full context to loosely defined context like wild card, the confidence for a focus unit can be measured at different expected accuracy. The joint verification by multiple TCP improves measured confidence of each unit in the transcription and is robust across different speech databases. Experimental results show that the checking process automatically separates erroneous sentences from correct ones: the sentence error hit rate decrease rapidly in the sorted TCP values, from 59% to 7% for the Mexican Spanish database and from 63% to 11% for the American English database, among the top 10% sentences in the rank lists. Shenghao Qin, Frank K. Soong |
INTERSPEECH | 3 |
| 2009 | Rich context modeling for high quality HMM-based TTSabstractThis paper presents a rich context modeling approach to high quality HMM-based speech synthesis. We first analyze the over-smoothing problem in conventional decision tree tyingbased HMM, and then propose to model the training speech tokens with rich context models. Special training procedure is adopted for reliable estimation of the rich context model parameters. In synthesis, a search algorithm following a contextbased pre-selection is performed to determine the optimal rich context model sequence which generates natural and crisp output speech. Experimental results show that spectral envelopes synthesized by the rich context models are with crisper formant structures and evolve with richer details than those obtained by the conventional models. The speech quality improvement is also perceived by listeners in a subjective preference test, in which 76% of the sentences synthesized using rich context modeling are preferred. Index Terms: HMM-based TTS, rich context modeling Zhijie Yan, Yao Qian, Frank K. Soong |
INTERSPEECH | 3 |
| 2009 | A Multi-Space Distribution (MSD) and two-stream tone modeling approach to Mandarin speech recognition
Yao Qian, Frank K. Soong |
Speech Commun. | 2 |
| 2009 | A Quadratic Optimization Approach to Discriminative Training of CDHMMsabstractIn this letter, we reformulate the discriminative training (DT) of continuous density hidden Markov models (CDHMMs) as an ellipsoid constrained quadratic programming (ECQP) problem, and we solve it by a line search algorithm. The ellipsoid constraint intrinsically arises from the step-size control in each iteration of DT optimization, which leads to an efficient solution without relaxing the objective function to be convex, as in a general quadratic programming problem. Moreover, the problem can be equivalently converted to a lower-dimensional one under some conditions, which helps further to simplify the solution. We show that under a Kullback–Leibler divergence (KLD) constraint, DT of CDHMM parameters such as Gaussian means and variances can be efficiently solved by the proposed algorithm, with only mild assumptions adopted. Experimental results on two tasks show that the ECQP approach considerably outperforms other popular algorithms in terms of both final recognition accuracy and convergence speed. Peng Liu 0001, Frank K. Soong |
IEEE Signal Process. Lett. | 2 |
| 2009 | Graph-Based Partial Hypothesis Fusion for Pen-Aided Speech InputabstractWe study a specificpartialhypothesisfusionproblem in sequential data labeling. The problem arises in the multimodal applications where a decision is made by merging complete hypothesis from one input and partial hypothesis from the other. For example, in a pen-aided speech interface, appropriate pen input can provide partial but crucial information. We address the problem in a Bayesian framework, and reformulate the solution as a revised search in a representation. A dynamic programming algorithm is proposed to efficiently solve the partial hypothesis fusion via the graph. It is shown that the computational cost of the graph based partial hypothesis fusion is proportional to the size of the graph, which is highly feasible for a given compact graph. We apply the proposed algorithm to two real applications: an intelligent pen-based dictation error correction system and an automatic handwritten character completion with a speech ldquoshortcutrdquo. Experimental results show that the algorithm is effective in utilizing the partial information from one modality to enhance the bimodal interface performance. Peng Liu 0001, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | A Cross-Language State Sharing and Mapping Approach to Bilingual (Mandarin-English) TTSabstractWe propose a hidden Markov model (HMM)-based bilingual (Mandarin and English) text-to-speech (TTS) system to synthesize natural speech for given bilingual text. A simple baseline system consisting of two independent monolingual HMM synthesizers is built first from corresponding Mandarin and English data recorded by a bilingual speaker. A new, mixed language TTS is then constructed by asking language-independent and language-specific questions for sharing HMM states across the two languages in decision-tree based clustering. By sharing states, the new system has a smaller footprint than the baseline system. Speech synthesized by the new system sounds very similar to the baseline for non-mixed, Mandarin or English, monolingual sentences but much better for mixed-language sentences. This higher quality of mixed-language output is confirmed by a preference score, 60.2% to 39.8%, in a subjective listening test. A cross-language state mapping algorithm is further proposed for cross-language synthesis when only monolingual (English) recorded data from a source language speaker is available. Mandarin speech is then synthesized with the HMM model parameters in the nearest neighbor leaf nodes of the English decision tree. The nearest neighbor is measured with the Kullback-Leibler divergence (KLD) and mappings between leaf nodes in the decision trees of the source and target languages are established via the speech data recorded by a different, bilingual speaker. High voice (speaker) similarity is preserved in the synthesized target language sentences by using the recording of a source language from a monolingual speaker. Perceptual test results conducted on synthesized Mandarin speech show 1) high intelligibility which is confirmed by a Chinese character transcription accuracy of 92.1% and 2) decent speech quality with an average MOS score of 3.1. Yao Qian, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Discriminative training for improving letter-to-sound conversion performanceabstractIn this paper, we propose to use discriminative training (DT) for improving letter-to-sound (LTS) conversion performance. LTS is a critical component in both ASR and TTS for predicting the correct pronunciation of a word not included in the lexicon. For TTS applications, predicting the proper pronunciation of an out-of-vocabulary person/place name, especially a name with foreign origin can be challenging. We utilize discriminative training, which has been successfully used in speech recognition, to sharpen the baseline N-grams of grapheme-phoneme pairs. We address the problem in a unified framework of discriminative training. Two criteria, maximum mutual information (MMI) and minimum phoneme error (MPE), are investigated. Experimental results show that DT yields a small (3.8-4.6% relative) but consistent error reduction across all databases tested. In addition, we observe that by pinpointing the local errors in a finer resolution, we can obtain a better discriminative model. Peng Liu 0001, Jiali You 0001, Frank K. Soong |
ICASSP | 4 |
| 2008 | A cross-language state mapping approach to bilingual (Mandarin-English) TTSabstractWe propose a cross-language state mapping approach to HMM-based bilingual TTS. Two language-dependent decision trees are built first with a bilingual speech database recorded by a single speaker. A state mapping for every leaf node in the decision tree of a target language is created by finding the nearest leaf node in the tree of a source language. Kullback-Leibler divergence between two distributions is used to find the nearest leaf node. To synthesize target language speech by a monolingual, (source language) speaker's voice, we find HMM parameters trained by the monolingual (source language) speaker in the mapped leaf nodes. Similar mappings can be constructed by reversing the source and target languages. With these bi-directional cross-lingual mappings, we can synthesize bilingual or mixed-code speech by HMMs trained by any monolingual speaker. High voice (speaker) similarity is preserved in synthesized speech of the target language. Two perceptual tests on synthesized Mandarin speech confirms high intelligibility with a Chinese character transcription accuracy of 92.1% and an MOS score of 3.08. Yao Qian, Frank K. Soong, Gongshen Liu |
ICASSP | 3 |
| 2008 | Prefix tree based auto-completion for convenient bi-modal chinese character inputabstractWe address the problem of character auto-completion (CAC) for predicting Chinese characters with partial, cursive handwriting input. A prefix tree decoder based CAC algorithm is proposed. The approach is based upon HMM pen trajectory modeling and radical structure of Chinese characters. Without finishing the strokes, high quality character candidates can be efficiently predicted. As a result, significant improvement of the recognition throughput can be obtained. We combine further the handwriting CAC with speech recognition candidates in the posterior sense, and come up with a flexible, rapid bi- modal Chinese character input system. The system was tested on a large Chinese corpus and shown that: more than 90% of input attempts can be correctly finished with only 50% of the whole character trajectory written. Peng Liu 0001, Frank K. Soong |
ICASSP | 3 |
| 2008 | Symbol graph based discriminative training and rescoring for improved math symbol recognitionabstractIn the symbol recognition stage of online handwritten math expression recognition, the one-pass dynamic programming algorithm can produce high-quality symbol graphs in addition of the best recognized hypotheses [1]. In this paper, we exploit the rich hypotheses embedded in a symbol graph to discriminatively train the exponential weights of different model likelihoods and the insertion penalty. The training is investigated in two different criteria: Maximum Mutual Information (MMI) and Minimum Symbol Error (MSE). After discriminative training, trigram-based graph rescoring is performed in a post-processing stage. Experimental results finally show a 97% symbol accuracy on a test set of 2,574 written expressions with 43,300 symbols, a signi..cant improvement of symbol accuracy obtained. Zhen Xuan Luo, Yu Shi 0001, Frank K. Soong |
ICASSP | 3 |
| 2008 | Template constrained posterior for verifying phone transcriptionsabstractA new statistical confidence measure, template constrained posterior (TCP), is proposed for verifying phone transcriptions of speech databases. Different from generalized posterior probability (GPP), TCP is computed by considering string hypotheses that bear a focused unit, e.g., phone with partially matched left and right contexts. Parameters used for TCP include context window length, partial matching ratio, KLD threshold for selecting confusable phones, and verification threshold. They are determined by minimizing verification errors in a development set. Evaluated on a test set which contains 52.1% sentence errors and 0.62% phone errors, TCP achieves 92% and 88% error hit rate in rejected sentences, when the corresponding acceptance ratios are set at 90% and 80%, respectively. Frank K. Soong |
ICASSP | 3 |
| 2008 | Improving letter-to-sound conversion performance with automatically generated new wordsabstractWe propose a novel way to alleviate the data sparseness problem in training letter-to-sound (LTS) N-gram models by adding automatically generated new words to the training set. The proposed method consists of two procedures: (1) generating a large pool of new words automatically; (2) selecting good new word candidates from the new word pool via semi-supervised learning. The new words are created by replacing stressed syllables of an existing word with other stressed syllables under specified contextual constraints. The new word selection by semi-supervised learning is based upon consistent pronunciation predictions by different LTS models. After adding new words to the training set, the performance of LTS conversion is significantly improved. For the NetTalk dictionary, compared with the performance from the N-gram baseline model, 21.6% relative word error rate reduction is obtained. For the CMU dictionary, 9.1% and 5.6% relative word error rate reductions are obtained, respectively, with/without considering the stress. Jiali You 0001, Frank K. Soong, Jinlin Wang 0001 |
ICASSP | 3 |
| 2008 | Automatic mispronunciation detection for MandarinabstractThis paper presents the methods to improve the performance of mispronunciation detection at syllable level for Mandarin from two aspects: proposing scaled log-posterior probability (SLPP) and weighted phone SLPP to get the better measure of pronunciation quality; introducing speaker normalization of speaker adaptive training (SAT) and speaker adaptation of selective maximum likelihood linear regression (SMLLR) to get a better statistical model. Experiments based on a database, consisting of 8000 syllables pronounced by 40 speakers with varied pronunciation proficiency, confirm the promising effectiveness of these strategies by reducing FAR from 41.1% to 31.4% at 90% FRR and 36.0% to 16.3%at 95%FRR. Chao Huang 0011, Frank K. Soong, Min Chu, Renhua Wang |
ICASSP | 3 |
| 2008 | Radical based fine trajectory HMMs of online handwritten charactersabstractWe study models that characterize pen trajectories of online handwritten characters in a fine manner. We propose radical based fine trajectory hidden Markov models (HMMs), which adopt radicals as basic units, and a multi-path HMM topology that emits observations with multi-space distributions (MSD) is built for each radical. Meanwhile, various stroke orders, writing styles and realness of sub-strokes are reasonably modeled. The radical based fine trajectory HMMs lead to handwriting recognition with effective prediction, and their generative nature can be utilized for a novel handwriting synthesis framework. Experimental show that along with the model precision increasing, about 50% recognition error can be reduced, and the fine models can generate decent character samples. Peng Liu 0001, Frank K. Soong |
ICPR | 3 |
| 2008 | A symbol graph based handwritten math expression recognitionabstractIn online handwritten math expression recognition, one-pass dynamic programming can produce high-quality symbol graphs in addition to best symbol sequence hypotheses, especially after discriminative training and trigram graph rescoring. Impact of symbol graphs on whole expression recognition, however, has not been referred to yet, since the interface of structure analysis module does not work well with symbol graphs on the basis of typical tree search. In this paper, we propose a method to convert symbol graph to segment graph to make the tree search efficient and effective, i.e., search of best segmentations in symbol graph without pruning becomes possible. With trigram rescoring, the overall expression recognition accuracy has been improved by 10% relative in comparison with the baseline. Yu Shi 0001, Frank K. Soong |
ICPR | 2 |
| 2008 | Duration refinement by jointly optimizing state and longer unit likelihood
Boyang Gao, Yao Qian, Zhizheng Wu 0001, Frank K. Soong |
INTERSPEECH | 4 |
| 2008 | Mispronunciation detection for Mandarin Chinese
Chao Huang 0011, Frank K. Soong, Min Chu |
INTERSPEECH | 3 |
| 2008 | An ellipsoid constrained quadratic programming perspective to discriminative training of HMMs
Peng Liu 0001, Frank K. Soong |
INTERSPEECH | 2 |
| 2008 | Generating natural F0 trajectory with additive trees
Yao Qian, Frank K. Soong |
INTERSPEECH | 3 |
| 2008 | GPU-accelerated Gaussian clustering for fMPE discriminative trainingabstractThe Graphics Processing Unit (GPU) has extended its applications from its original graphic rendering to more general scientific computation. Through massive parallelization, state-ofthe-art GPUs can deliver 200 billion floating-point operations per second (0.2 TFLOPS) on a single consumer-priced graphics card. This paper describes our attempt in leveraging GPUs for efficient HMM model training. We show that using GPUs for a specific example of Gaussian clustering, as required in fMPE, or feature-domain Minimum Phone Error discriminative training, can be highly desirable. The clustering of huge number of Gaussians is very time consuming due to the enormous model size in current LVCSR systems. Comparing an NVidia Geforce 8800 Ultra GPU against an Intel Pentium 4 implementation, we find that our brute-force GPU implementation is 14 times faster overall than a CPU implementation that uses approximate speed-up heuristics. GPU accelerated fMPE reduces the WER 6% relatively, compared to the maximumlikelihood trained baseline on two conversational-speech recognition tasks. Yu Shi 0001, Frank Seide, Frank K. Soong |
INTERSPEECH | 3 |
| 2008 | Efficient handwriting correction of speech recognition errors with template constrained posterior (TCP)abstractMore mobile devices are starting to use automatic speech recognition for command or text input. However, correcting recognition errors in a small compact mobile device is usually inconvenient and it may take several finger operations on a small keypad to correct errors. In this paper, we propose a new multimodal input method and a novel confidence measure ― template constrained posterior (TCP) to simplify the correction process. The method works by interactively integrating a handwriting recognizer with a speech recognizer. Information obtained in pen-based error marking, like error location, error type, etc., is fed back to the speech recognizer, and speech recognition errors are automatically corrected using the TCP confidence measure. Experimental results on Aurora2, Wall Street Journal, Switchboard, and two Chinese databases show that compared with speech recognition baseline, the proposed method achieves relative error reduction of 64.9%, 43.9%, 26.1%, 39.0%, 31.4%, respectively, after the auto correction. Peng Liu 0001, Frank K. Soong |
INTERSPEECH | 4 |
| 2008 | A real-time text to audio-visual speech synthesis system
Xiaojun Qian, Yao Qian, Frank K. Soong |
INTERSPEECH | 6 |
| 2008 | Prosody for Mandarin speech recognition: a comparative study of read and spontaneous speech
Yu Ting Yeung, Yao Qian, Tan Lee, Frank K. Soong |
INTERSPEECH | 4 |
| 2008 | Tone-enhanced generalized character posterior probability (GCPP) for Cantonese LVCSR
Yao Qian, Frank K. Soong, Tan Lee |
Comput. Speech Lang. | 2 |
| 2008 | A Constrained Line Search Optimization Method for Discriminative Training of HMMsabstractIn this paper, we propose a novel optimization algorithm called constrained line search (CLS) for discriminative training (DT) of Gaussian mixture continuous density hidden Markov model (CDHMM) in speech recognition. The CLS method is formulated under a general framework for optimizing any discriminative objective functions including maximum mutual information (MMI), minimum classification error (MCE), minimum phone error (MPE)/minimum word error (MWE), etc. In this method, discriminative training of HMM is first cast as a constrained optimization problem, where Kullback-Leibler divergence (KLD) between models is explicitly imposed as a constraint during optimization. Based upon the idea of line search, we show that a simple formula of HMM parameters can be found by constraining the KLD between HMM of two successive iterations in an quadratic form. The proposed CLS method can be applied to optimize all model parameters in Gaussian mixture CDHMMs, including means, covariances, and mixture weights. We have investigated the proposed CLS approach on several benchmark speech recognition databases, including TIDIGITS, Resource Management (RM), and Switchboard. Experimental results show that the new CLS optimization method consistently outperforms the conventional EBW method in both recognition performance and convergence behavior. Peng Liu 0001, Cong Liu 0006, Hui Jiang 0001, Frank K. Soong, Renhua Wang |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Identifying Language Origin of Named Entity With Multiple Information SourcesabstractTo identify the language origin of a named entity, morphological information associated with its letter spelling, such as letter N-grams, is commonly employed. However, with this information only, named entities with similar spellings but from different language origins are difficult to differentiate. In this paper, a measure of "popularity," in terms of frequency or page count of the named entity in language-specific Web search, is proposed for identifying its language origin. Morphological information, including letter or letter-chunk N-grams, is used to enhance the performance of language identification in conjunction with Web-based page counts. Six languages, including English, German, French, Portuguese, Chinese, and Japanese (Chinese and Japanese named entities are shown in their corresponding phonetic alphabets, i.e., Pinyin and Romaji), are tested. Experiments show that when classifying four Latin languages, including English, German, French, and Portuguese, which are written in Latin alphabets, features from different information sources yield substantial performance improvements in the classification accuracy over a letter 4-gram-based baseline system. The accuracy increases from 75.0% to 86.3%, or a 45.2% relative error reduction. Jiali You 0001, Min Chu, Frank K. Soong, Jinlin Wang 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | A constrained line search approach to general discriminative HMM trainingabstractRecently, we proposed a novel optimization algorithm called constrained line search (CLS) to train Gaussian mean vectors of HMMs in the MMI sense. In this paper, we extend and re-formulate it in a more general framework. The new CLS can optimize any discriminative objective functions including MMI, MCE, MPE/MWE etc. Also, closed-form solutions to update all Gaussian mixture parameters, including means, covariances and mixture weights, are obtained. We investigate the new CLS on several benchmark speech recognition databases, including TIDIGITS, Switchboard mini-train and Switchboard full h5train00 sets. Experimental results show that the new CLS optimization method outperforms the conventional EBW method in both performance and convergence behavior. Peng Liu 0001, Cong Liu 0006, Hui Jiang 0001, Frank K. Soong, Renhua Wang |
ASRU | 4 |
| 2007 | Divergence-Based Similarity Measure for Spoken Document RetrievalabstractWe propose a novel, divergence-based similarity measure for spoken document retrieval (SDR). We derive a dynamic programming algorithm that measures Kullback-Leibler divergence between two HMMs first. The measure is further generalized to a graph matching algorithm, which is efficient for SDR application. The proposed approach compares the underlying acoustic models of keywords and a target database to alleviate the impact of mismatched vocabulary and language model, e.g. different domains. Experimental results on the Wall Street Journal (WSJ) database show that the proposed approach achieves a comparable performance, compared with the word posterior based approach. It outperforms the latter when there is a mismatch in language model. The approach is promising for building an open-vocabulary, domain independent SDR application. Peng Liu 0001, Frank K. Soong, Jian-Lai Zhou |
ICASSP (4) | 2 |
| 2007 | A New Minimum Divergence Approach to Discriminative TrainingabstractWe propose to use minimum divergence, where acoustic similarity between HMMs is characterized by Kullback-Leibler divergence, for discriminative training. The MD objective function is defined as a posterior weighted divergence measured over the whole training set. Different from our earlier work, where KLD-based acoustic similarity is pre-computed for all initial models and stays invariant in the optimization procedure, here we propose to jointly optimize the whole variable MD by adjusting HMM parameters since MD is a function of the adjusted HMM parameters. An EBW optimization method is derived to minimize the whole MD objective function. The new MD formulation is evaluated on the TIDIGITS and Switchboard databases. Experimental results show that the new MD yields relative word error rate reductions of 62.1% on TIDIGITS and 8.8% on Switchboard databases when compared with the best ML-trained systems. It is also shown the new MD consistently outperforms other discriminative training criteria, such as MPE. Jun Du 0002, Peng Liu 0001, Hui Jiang 0001, Frank K. Soong, Renhua Wang |
ICASSP (4) | 4 |
| 2007 | A Constrained Line Search Optimization for Discriminative Training in Speech RecognitionabstractIn this paper, we propose a novel constrained line search to optimize the MMEE objective function for training discriminative HMMs. In our method, the MMI estimation is cast as a constrained maximization problem, where Kullback-Leibler divergence between models before and after parameters adjustment is introduced as a constraint during optimization. Then, based on the idea of line search, we show that a simple, closed-form solution can be derived under some approximation assumptions. The proposed optimization method have been investigated in two speech recognition tasks: TIDIGITS and Switchboard (mini-train). Experimental results show that the new training method achieves significant word error rate reduction when comparing with our best MLE models, i.e., relatively 63.8% on TIDIGITS and 6.1% on the Switchboard mini-train set, respectively. Our results also show that the constrained line search method consistently outperforms the popular EBW method in both tasks. Cong Liu 0006, Peng Liu 0001, Hui Jiang 0001, Frank K. Soong, Renhua Wang |
ICASSP (4) | 4 |
| 2007 | Agreement Learning for Automatic Accent AnnotationabstractAutomatic accent annotation is important in both speech synthesis and speech recognition. Existing statistical learning algorithms rely heavily on a sufficiently large set of labeled training samples that are expensive and time consuming to collect. For unlabeled data, unsupervised learning can be initiated with a small set of manually labeled data. This paper shows that the accuracy of automatic accent annotation can be improved by augmenting a small amount of manually labeled data with a large pool of unlabeled data. We introduce an agreement-learning algorithm for this propose. Experimental results show that it is possible to reduce human-labeling effort significantly while reducing up to 50% errors. Xinqiang Ni, Min Chu, Frank K. Soong, Yong Zhao 0008 |
ICASSP (4) | 4 |
| 2007 | Full HMM Training for Minimizing Generation Error in SynthesisabstractIn maximum-likelihood (ML) based HMM synthesis, the generated trajectory of a sentence in the training set is in general does not reproduce the trajectory of the original one. To overcome this shortcoming, a minimum generation error (MGE) criterion has been previously proposed. In this paper, a complete MGE-based HMM training is introduced, where the MGE criterion is applied to the entire training process, including context-dependent HMM training, context-dependent HMM clustering and clustered HMM training. In this procedure, the HMMs are trained to minimize the generation error of training data, which is in line with the HMM-based synthesis. From the experiments, the quality of synthesized speech is improved after applying the MGE criterion to the whole training process. Yi-Jian Wu, Renhua Wang, Frank K. Soong |
ICASSP (4) | 3 |
| 2007 | A Segmentation Posterior Based Endpointing AlgorithmabstractA segmentation posterior probability based endpointing algorithm for robust ASR is proposed. First, each speech signal is partitioned into homogeneous segments via auto-segmentation. Then posterior probabilities of all possible endpoints are computed, based on the segmentation likelihoods of all levels in a selected range. Endpoints with the highest posterior probabilities are finally selected. The new method differs from the previous auto-segmentation and clustering based algorithm on that the former considers hypotheses from several levels, while the latter depends only on one appropriate level. Another potential benefit of the proposed method is that any endpointing or VAD results can be integrated, as hypotheses, into the posterior probability framework. Experiments based on the AURORA2 digit database show the robustness of the proposed method. Yanlu Xie, Yu Shi 0001, Frank K. Soong, Beiqian Dai |
ICASSP (4) | 3 |
| 2007 | Word Graph Based Feature Enhancement for Noisy Speech RecognitionabstractThis paper presents a word graph based feature enhancement method for robust speech recognition in noise. The approach uses signal processing based speech enhancement as a starting point, and then performs Wiener filtering to remove residual noise. During the process, a decoded word graph is used to directly guide the feature enhancement with respect to the HMM for recognition, so that the enhanced feature can match the clean speech model better in the acoustic space. The proposed word graph based feature enhancement method was tested on the Aurora 2 database. Experimental results show that an improved recognition performance can be obtained comparing with conventional signal processing based and GMM based feature enhancement methods. With signal processing based weighted noise estimation and GMM based method, the relative error rate reductions are 35.44% and 42.58%, respectively. The proposed word graph based method improves the performance further, and a relative error rate reduction of 57.89% is obtained. Zhijie Yan, Frank K. Soong, Renhua Wang |
ICASSP (4) | 2 |
| 2007 | Generalized Segment Posterior Probability for Automatic Mandarin Pronunciation EvaluationabstractIn this paper, we investigate the automatic pronunciation evaluation method for native Mandarin. Multi-space distribution (MSD) hidden Markov model (HMM) is adopted to train the gold standard model. Machine scores derived from the generalized segment posterior probability on both syllables and phone level are proposed and investigated to measure the goodness of pronunciation (GOP). They are evaluated on the database collected internally and shown better performance than other well-known methods. In addition, detailed analyses of human scoring such as inter/intra-rater on utterance/speaker level are also given. Chao Huang 0011, Min Chu, Frank K. Soong, Weiping Ye |
ICASSP (4) | 4 |
| 2007 | A MSD-HMM Approach to Pen Trajectory Modeling for Online Handwriting RecognitionabstractIn modeling online handwritten characters, imaginary strokes have been conveniently generated by connecting adjacent real strokes together to form a continuous trajectory. However, this approach causes confusions among characters with similar but actually different trajectories. In this paper, we propose to use multi-space probability distribution (MSD) to model imaginary strokes jointly with real strokes. With the proposed MSD, real and imaginary strokes become observations from different probability spaces and they are modeled stochastically. Also, the flexibility in MSD to assign different feature dimensions to each individual space enables us to ignore certain features that can cause singularity problem in modeling. Experimental results obtained in handwritten Chinese character recognition indicate MSD provides 1.3%-2.8% character recognition accuracy improvement across different recognition systems where MSD significantly improves discrimination among confusable characters with similar trajectories. Frank K. Soong, Peng Liu 0001, Yi-Jian Wu |
ICDAR | 2 |
| 2007 | A Unified Framework for Symbol Segmentation and Recognition of Handwritten Mathematical ExpressionsabstractA symbol decoding and graph generation algorithm for online handwritten mathematical expression recognition is formulated. It differs from our previous system and most other systems in two aspects: (1) it embeds stroke grouping into symbol identification to form a unified probabilistic framework for symbol recognition; and (2) a symbol graph rather than a list of symbol sequence hypotheses is generated, which makes post-processing with new information possible. Experimental results show that high quality symbol graph can be generated by the proposed algorithm. Symbol sequence corresponding to the best path in the graph demonstrates much higher symbol recognition accuracy than before, especially after rescoring with trigram. Math formula recognition performance is significantly improved. Yu Shi 0001, Frank K. Soong |
ICDAR | 3 |
| 2007 | Minimum Error Discriminative Training for Radical-Based Online Chinese Handwriting RecognitionabstractFree style Chinese handwriting recognition continues to pose a challenge to researchers due to the variety of writing styles. To recognize handwritten characters in an online mode, Hidden Markov Model (HMM) has been naturally adopted to model the pen trajectory of a character and a decent recognition performance is achieved. In this study, we start from a maximum likelihood trained HMM model and focus on minimizing recognition errors at the radical (sub- character) level to optimize the recognition performance. A novel Minimum Radical Error discriminative training criterion is proposed, and compared with the discrimination at the character level, our new approach further reduces the character errors by 15.6% relatively (29.0% overall reduction from the maximum likelihood baseline model) on a Chinese database. Yu Zhang 0007, Peng Liu 0001, Frank K. Soong |
ICDAR | 3 |
| 2007 | Model-based speech separation with single-microphone input
Siu Wa Lee, Frank K. Soong, Pak-Chung Ching |
INTERSPEECH | 2 |
| 2007 | Iterative unit selection with unnatural prosody detectionabstractCorpus-driven speech synthesis is hampered by the occurrence of occasional glitches which ruin the impression of the whole utterance. We propose an iterative unit selection integrated with an unnatural prosody detection model to identify any unnatural prosody. The system searches an optimal path in the lattice, verifies its naturalness by the unnatural prosody model and replaces the bad section with a better candidate, until it passes the verification test. In light of hypothesis testing, we show this trial-and-error approach takes effective advantage of abundant candidate samples in the database. Also, in contrast to conventional prosody prediction, an unnatural prosody detection model still leaves enough room for the prosody variations. Unnaturalness confidence measures are studied. The combined model can reduce the objective distortion by 16.3%. Perceptual experiments also confirm the proposed approach improves the synthetic speech quality appreciably. Index Terms: speech synthesis, unit selection, confidence measure, unnatural prosody detection, iterative synthesis Dacheng Lin, Yong Zhao 0008, Frank K. Soong, Min Chu |
INTERSPEECH | 3 |
| 2007 | An unsupervised approach to automatic prosodic annotation
Xinqiang Ni, Frank K. Soong, Min Chu |
INTERSPEECH | 3 |
| 2007 | Robust F0 modeling for Mandarin speech recognition in noise
Sheng Qiang, Yao Qian, Frank K. Soong, Congfu Xu |
INTERSPEECH | 3 |
| 2007 | Context constrained-generalized posterior probability for verifying phone transcriptionsabstractA new statistical confidence measure, Context ConstrainedGeneralized Posterior probability (CC-GPP), is proposed for verifying phone transcriptions in speech databases. Different from generalized posterior probability (GPP), CC-GPP is computed by considering string hypotheses that bear a focused phone with partially matched left and right contexts. Parameters used for CC-GPP include context window length, a minimal number of matched context phones, and verification thresholds. They are determined by minimizing verification errors in a development set. Evaluated on a test set of 500 sentences that consist of 2.1% phone errors, CCGPP achieves 99.6% accuracy and 78.7% recall when 90% of the phones are accepted. Hua Zhang 0009, Frank K. Soong |
INTERSPEECH | 3 |
| 2007 | A Syllable Lattice Approach to Speaker VerificationabstractThis paper proposes a syllable-lattice-based speaker verification algorithm for Mandarin Chinese input. For each speech utterance, a syllable lattice is generated with a speaker-independent large-vocabulary continuous speech recognition system in free syllable decoding. The verification decision is made based upon the likelihood ratio between a target-speaker model and a speaker-independent background model, computed on the decoded syllable lattice. The likelihood function is calculated efficiently in a forward algorithm by considering all paths in the lattice. The proposed algorithm was evaluated using a Mandarin Chinese database, where 1832 true and 26 250 impostor trials were recorded by 19 target speakers and 180 impostors. The average duration of each trial is 2 s long without silence. The target-speaker model was adapted from the speaker-independent background model using enrollment data of two minutes with silence. The proposed algorithm achieved an equal-error rate of 0.857% which is better than 1.21% of the hidden Markov model-based speaker verification algorithm without using syllable lattices. The equal-error rate was further reduced to 0.617% by incorporating the Goussian mixture model-universal background model algorithm with 2048 Gaussian kernels whose equal error rate is 0.990%. Minho Jin, Frank K. Soong, Chang Dong Yoo |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | A Cohort-Based Speaker Model Synthesis for Mismatched Channels in Speaker VerificationabstractMismatch between enrollment and test data is one of the top performance degrading factors in speaker recognition applications. This mismatch is particularly true over public telephone networks, where input speech data is collected over different handsets and transmitted over different channels from one trial to the next. In this paper, a cohort-based speaker model synthesis (SMS) algorithm, designed for synthesizing robust speaker models without requiring channel-specific enrollment data, is proposed. This algorithm utilizes a priori knowledge of channels extracted from speaker-specific cohort sets to synthesize such speaker models. The cohort selection in the proposed new SMS can be either speaker-specific or Gaussian component based. Results on the China Criminal Police College (CCPC) speaker recognition corpus, which contains utterances from both landline and mobile channel, show the new algorithms yield significant speaker verification performance improvement over Htnorm and universal background model (UBM)-based speaker model synthesis. Thomas Fang Zheng, Mingxing Xu, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Static and Dynamic Spectral Features: Their Noise Robustness and Optimal Weights for ASRabstractIn this paper, we investigate the relative noise robustness of dynamic and static spectral features in speech recognition. It is found that the dynamic cepstrum is more robust to additive noise than its static counterpart. The results are consistent across different types of noise and over a wide range of noise levels. To exploit this unequal robustness, we propose a simple yet effective strategy of exponentially weighting the likelihoods that are contributed by the static and dynamic features during the decoding process. The optimal weights are discriminatively trained with a small amount of development data. This method is evaluated on two speaker-independent, connected digit databases, one in English (Aurora 2) and the other in Cantonese (CUDIGIT). For various types of noise at different signal-to-noise ratios (SNRs), the average relative word error rate reductions attained with the discriminatively trained weights are 36.6% and 41.9% for Aurora 2 and CUDIGIT, respectively. Noticeable performance improvement can be observed even when there is channel distortion. The proposed approach is appealing to practical applications because 1) noise estimation is not required, 2) model adaptation is not required, 3)only a minor modification of the decoding process is needed, and 4) only a few feature weights need to be trained Frank K. Soong, Tan Lee |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Weighted Likelihood Ratio (WLR) Hidden Markov Model for Noisy Speech RecognitionabstractIn this paper we present a weighted likelihood ratio (WLR) based hidden Markov model and apply it to speech recognition in noise. The WLR measure emphasizes spectral peaks than valleys in comparing two given speech spectra. The measure is more consistent with human perception of speech formants where natural resonances of vocal track are and tends to be more robust to broad-band noise interferences than other measures. A complete HMM framework of this measure is derived and a mixture of exponential kernels is used to model the output probability density function. The new WLR-HMM is tested on the Aurora2 connected digits database in noise. It shows more robust performance than the MFCC trained GMM baseline system. When combined with the dynamic cepstral features, the multiple-stream WLR-HMM shows a 39% relative improvement over the baseline system Chao Huang 0011, Yingchun Huang, Frank K. Soong, Jian-Lai Zhou |
ICASSP (1) | 3 |
| 2006 | Syllable Lattice Based Re-Scoring For Speaker VerificationabstractThe Gaussian mixture based GMM-UBM approaches have shown good performance in speaker verification without using contextual information. In this paper, we exploit the information provided in the arcs of a decoded syllable lattice for speaker verification. The forward algorithm is used to summarize this information in the syllable lattice instead of the best decoded string. The performance is evaluated on a Mandarin Chinese database. With two minutes of target speaker's enrollment data, the proposed algorithm shows 1.03% of equal-error rate for short input utterances with an average duration of two seconds. By combining with the GMM-UBM, the system shows a 0.74% of equal-error rate. Minho Jin, Frank K. Soong, Chang Dong Yoo |
ICASSP (1) | 2 |
| 2006 | An Iterative Trajectory Regeneration Algorithm for Separating Mixed Speech SourcesabstractHarmonicity and continuity are two important perceptual cues for separating mixed speech sources. This paper focuses on the separation of two speech sources with a single-microphone input. An iterative, least-squares (LS) based trajectory regeneration algorithm is proposed to estimate the magnitude spectrum of each source. Time-derivatives of the spectrum, or the dynamic spectral information, is used as a constraint in solving the resultant weighted normal equations. Each estimated spectral trajectory, as a result, exhibits similar temporal variations as the original source. Asymptotically, we also prove that the regenerated trajectory yields the same time variations as the given dynamic information. When cascaded with our previously proposed harmonic filtering algorithm to separate mixed voiced signals, the new trajectory regeneration is shown to be very effective to reduce mean squared errors by 82.2% and 69.5%, relatively, with ideal and approximated dynamic information, respectively Siu Wa Lee, Frank K. Soong, Pak-Chung Ching |
ICASSP (1) | 2 |
| 2006 | Tone-Enhanced Generalized Character Posterior Probability (GCPP) for Cantonese LVCSRabstractTone-enhanced, generalized character posterior probability (GCPP), a generalized form of posterior probability at subword (Chinese character) level, is proposed as a rescoring metric for improving Cantonese LVCSR performance. The search network is constructed first by converting the original word graph to a restructured word graph, then a character graph and finally, a character confusion network (CCN). Based upon GCPP enhanced with tone information, the character error rate (CER) is minimized or the GCPP product is maximized over a chosen graph. Experimental results show that the tone enhanced GCPP can improve character error rate by up to 15.1%, relatively Yao Qian, Frank K. Soong, Tan Lee |
ICASSP (1) | 2 |
| 2006 | Auto-Segmentation Based Partitioning and Clustering Approach to Robust EndpointingabstractAn auto segmentation based partitioning and clustering approach to robust Voice Activity Detection (VAD) is proposed. It is done in two successive steps: homogeneous frame partitioning and segment clustering. The first step, due to its auto segmentation nature, does not need a noise model, and is applicable to different noise types and SNR's. The algorithm is a dynamic programming based procedure and provides a graceful performance in finding segmentation thresholds. Multiple parameters like energy, pitch and voicing information can be easily incorporated into the procedure. The algorithm is evaluated on the test sets in the Aurora2 database. The algorithm shows its robustness at low SNR operating environments; the endpoint estimate errors are shown to have small variance. Yu Shi 0001, Frank K. Soong, Jian-Lai Zhou |
ICASSP (1) | 2 |
| 2006 | A Comparative Study of Discriminative Methods for Reranking LVCSR N-Best Hypotheses in Domain Adaptation and GeneralizationabstractThis paper is an empirical study on the performance of different discriminative approaches to reranking the N-best hypotheses output from a large vocabulary continuous speech recognizer (LVCSR). Four algorithms, namely perceptron, boosting, ranking support vector machine (SVM) and minimum sample risk (MSR), are compared in terms of domain adaptation, generalization and time efficiency. In our experiments on Mandarin dictation speech, we found that for domain adaptation, perceptron performs the best; for generalization, boosting performs the best. The best result on a domain-specific test set is achieved by the perceptron algorithm. A relative character error rate (CER) reduction of 11% over the baseline was obtained. The best result on a general test set is 3.4% CER reduction over the baseline, achieved by the boosting algorithm. Zhengyu Zhou, Jianfeng Gao 0001, Frank K. Soong, Helen M. Meng |
ICASSP (1) | 3 |
| 2006 | Improved Chinese Character Input by Merging Speech and Handwriting Recognition HypothesesabstractIn this paper we propose to merge speech and handwriting recognition hypotheses together for improving the performance of Chinese character input. The recognition result of handwriting character input can be reliable when the character is written rather squarely. However, more legible of square handwriting tends to slow down the input (stroke writing) speed. On the other hand, speech input is fairly efficient but a large number of homonyms and its vulnerability to adverse environment prevent speech from being used as a robust Chinese character input method. The handwriting stroke information and acoustic speech information, in many cases, are complementary to each other. In this study we use independent, statistically trained HMMs for recognizing each input mode individually but merge recognition hypotheses from the two recognizers. Generalized posterior probabilities are used to synchronize, compare and merge hypotheses appropriately. Experimental results have shown that significant input speedup can be obtained while maintaining the same recognition performance. Jian-Lai Zhou, Frank K. Soong, Beiqian Dai |
ICASSP (1) | 4 |
| 2006 | Word graph based speech rcognition error correction by handwriting inputabstractWe propose a convenient handwriting user interface for correcting speech recognition errors efficiently. Via the proposed hand-marked correction on the displayed recognition result, substitution, deletion and insertion errors can be corrected efficiently by rescoring the word graph generated in the recognition pass. A new path in the graph that matches the user's feedback in the maximum likelihood sense is found.With the aid of language model and hand corrections part in the best decoded path, rescoring the word graph can correct more errors than user provides. All recognition errors can be corrected after finite number of corrections. Experimental results show that by indicating one word error in user feedback, 33.8% of the erroneous sentences can be corrected; while by indicating one character error, 12.9% of the erroneous sentences can be corrected. Peng Liu 0001, Frank K. Soong |
ICMI | 2 |
| 2006 | Minimum divergence based discriminative trainingabstractWe propose to use Minimum Divergence(MD) as a new measure of errors in discriminative training. To focus on improving discrimination between any two given acoustic models, we refine the error definition in terms of Kullback-Leibler Divergence (KLD) between them. The new measure can be regarded as a modified version of Minimum Phone Error (MPE) but with a higher resolution than just a symbol matching based criterion. Experimental recognition results show the new MD based training yields relative word error rate reductions of 57.8% and 6.1% on TIDigits and Switchboard databases, respectively, in comparing with the ML trained baseline systems. The recognition performance of MD is also shown to be consistently better than that of MPE. Jun Du 0002, Peng Liu 0001, Frank K. Soong, Jian-Lai Zhou, Renhua Wang |
INTERSPEECH | 3 |
| 2006 | Generalization of the minimum classification error (MCE) training based on maximizing generalized posterior probability (GPP)
Qiang Fu 0008, Antonio Moreno-Daniel, Biing-Hwang Juang, Jian-Lai Zhou, Frank K. Soong |
INTERSPEECH | 5 |
| 2006 | Auto-segmentation based VAD for robust ASRabstractAn auto-segmentation based endpointing algorithm for robust ASR is proposed. The algorithm consists of two successive steps: (1) homogeneous segment partitioning and (2) segment clustering. The first step, due to its self-segmentation nature, does not need a noise model, and is applicable to different noises at various SNR’s. The dynamic programming based segment partitioning, which can generate more homogeneous segments than individual frames for clustering, yields a more robust VAD mechanism. Experiments are performed on the AURORA2 digit database by comparing the new algorithm with the ETSI standard for DSR. Quantitative assessment of the new algorithm is performed via different evaluation criteria, including: ROC curves, speech/non-speech discrimination, and speech recognition performance. Yu Shi 0001, Frank K. Soong, Jian-Lai Zhou |
INTERSPEECH | 2 |
| 2006 | A multi-space distribution (MSD) approach to speech recognition of tonal languages
Huanliang Wang, Yao Qian, Frank K. Soong, Jian-Lai Zhou, Jiqing Han 0001 |
INTERSPEECH | 3 |
| 2006 | A tree-based kernel selection approach to efficient Gaussian mixture model-universal background model based speaker identification
Thomas Fang Zheng, Zhanjiang Song, Frank K. Soong, Wenhu Wu |
Speech Commun. | 4 |
| 2005 | Optimal Clustering and Non-Uniform Allocation of Gaussian Kernels in Scalar Dimension for HMM CompressionabstractWe propose an algorithm for optimal clustering and nonuniform allocation of Gaussian kernels in scalar (feature) dimension to compress complex, Gaussian mixture-based, continuous density HMMs into computationally efficient, small footprint models. The symmetric Kullback-Leibler divergence (KLD) is used as the universal distortion measure and it is minimized in both kernel clustering and allocation procedures. The algorithm was tested on the resource management (RM) database. The original context-dependent HMMs can be compressed to any resolution, measured by the total number of clustered scalar kernel components. Good trade-offs between the recognition performance and model complexities have been obtained; the HMM can be compressed to 15-20% of the original model size, which needs 1-5% of multiplication/division operations, and results in almost negligible recognition performance degradation. Xiao-Bing Li, Frank K. Soong, Tor André Myrvoll, Renhua Wang |
ICASSP (1) | 2 |
| 2005 | Generalized Posterior Probability for Minimum Error Verification of Recognized SentencesabstractGeneralized posterior probability (GPP) is investigated in this paper as a statistical confidence measure for verifying recognized sentences of a large vocabulary continuous speech recognition system (LVCSR). We optimize the GPP by training the exponential weights of the acoustic and language models and decision threshold to minimize total verification errors. Two utterance level confidence measures: generalized utterance posterior probability (GUPP) and product of generalized word posterior probabilities (GWPP) of component words in a string hypothesis are tested. When evaluated on the Chinese Basic Travel Expression Corpus (BTEC), 47.9% and 53.9% relative improvement of utterance confidence error rate (CER) have been obtained for the GUPP and product of GWPP confidence measures, respectively. Wai Kit Lo, Frank K. Soong |
ICASSP (1) | 2 |
| 2005 | Static and Dynamic Spectral Features: Their Noise Robustness and Optimal Weights for ASRabstractIn this paper, we investigate the relative noise robustness between dynamic and static spectral features, by using two speaker independent continuous digit databases in English (Aurora2) and Cantonese (CUDigit). It is found that the dynamic cepstrum is more robust to additive noise than its static counterpart. The results are consistent across different types of noise and under various SNRs. Optimal exponential weights for exploiting unequal noise robustness of the two features are discriminatively trained in a development set. When tested under various noise conditions, the optimal weights yielded relative word error rate reductions of 36.6% and 41.9% for Aurora2 and CUDigit, respectively. The proposed weighting is attractive for many ASR applications in noise because: (1) no noise estimation for feature compensation; (2) no adaptation of clean HMMs to a noisy environment; and (3) only a trivial change in the decoding process by weighting log likelihoods of static and dynamic components separately. Frank K. Soong, Tan Lee |
ICASSP (1) | 2 |
| 2005 | Harmonic filtering for joint estimation of pitch and voiced source with single-microphone input
Siu Wa Lee, Frank K. Soong, Pak-Chung Ching |
INTERSPEECH | 2 |
| 2005 | Background model based posterior probability for measuring confidenceabstractWord posterior probability (WPP) computed over LVCSR word graphs has been used successfully in measuring confidence of speech recognition output. However, for certain applications the word graph is too sparse to warrant reliable WPP estimation. In this paper, we incorporate subword units as background models to generate a subword graph for estimating posterior probability. Experiments on both English and Chinese databases show that syllable background models can repopulate the dynamic hypothesis space for effective computation of confidence measure. The resultant posterior probability confidence measure achieves 94.3% and 95.2% Out-Of-Vocabulary (OOV) word detection / rejection in English and Chinese, respectively. Correspondingly, confidence error rates are at 6.0% and 6.4%, respectively. Peng Liu 0001, Jian-Lai Zhou, Frank K. Soong |
INTERSPEECH | 4 |
| 2005 | Phonetic transcription verification with generalized posterior probabilityabstractAccurate phonetic transcription is critical to high quality concatenation based text-to-speech synthesis. In this paper, we propose to use generalized syllable posterior probability (GSPP) as a statistical confidence measure to verify errors in phonetic transcriptions, such as reading errors, inadequate alternatives of pronunciations in the lexicon, letter-to-sound errors in transcribing out-of-vocabulary words, idiosyncratic pronunciations, etc. in a TTS speech database. GSPP is computed based upon a syllable graph generated by a recognition decoder. Testing on two data sets, the proposed GSPP is shown to be effective in locating phonetic transcription errors. Equal error rates (EERs) of 8.2% and 8.4%, are obtained on two testing sets, respectively. It is also found that the GSPP verification performance is fairly stable over a wide range around the optimal value of acoustic model exponential weight used in computing GSPP. Yong Zhao 0008, Min Chu, Frank K. Soong |
INTERSPEECH | 4 |
| 2005 | Refining phoneme segmentations using speaker-adaptive context dependent boundary modelsabstractConsistent phoneme segmentation is essential in building high quality Text-to-Speech (TTS) voice fonts. In this paper we propose to adapt an existing well-trained Context Dependent Boundary Model (CDBM) for refining segment boundaries to a new speaker with limited, manually segmented data. Three adaptation approaches: MLLR, MAP, and a combination of the two, are studied. The combined one, MLLR+MAP, delivers the best boundary refinement performance. In comparison with other boundary segmentation methods, the adapted CDBM yields better results, especially with a limited amount of adaptation data. Given 400 manually segmented boundary tokens in about 20 sentences as a development set, the segmentation precision can reach 90% of human labeled boundaries within a tolerance of 20 ms. Yong Zhao 0008, Min Chu, Frank K. Soong |
INTERSPEECH | 4 |
| 2005 | A Dynamic In-Search Data Selection Method With Its Applications to Acoustic Modeling and Utterance VerificationabstractIn this paper, we propose a dynamic in-search data selection method to diagnose competing information automatically from speech data. In our method, the Viterbi beam search is used to decode all training data. During decoding, all partial paths within the beam are examined to identify the so-called competing-token and true-token sets for each individual hidden Markov model (HMM). In this work, the collected data tokens are used for acoustic modeling and utterance verification as two specific examples. In acoustic modeling, the true-token sets are used to adapt HMMs with a sequential maximum a posteriori adaptation method, while a generalized probabilistic descent-based discriminative training method is proposed to improve HMMs based on competing-token sets. In utterance verification, under the framework of likelihood ratio testing, the true-token sets are employed to train positive models for the null hypothesis and the competing-token sets are used to estimate negative models for the alternative hypothesis. All the proposed methods are evaluated in Bell Laboratories communicator system. Experimental results show that the new acoustic modeling method can consistently improve recognition performance over our best maximum likelihood estimation models, roughly 1% absolute reduction in word error rate. The results also show the new verification models can significantly improve the performance of utterance verification over the conventional anti models, almost relatively 30% reduction of equal error rate when identifying misrecognized words from the recognition results. Hui Jiang 0001, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | A Unified Approach in Speech-to-Speech Translation: Integrating Features of Speech recognition and Machine Translation
Ruiqiang Zhang, Gen-ichiro Kikui, Hirofumi Yamamoto, Frank K. Soong, Taro Watanabe, Wai Kit Lo |
COLING | 4 |
| 2004 | Robust verification of recognized words in noiseabstractIn this paper we investigate robust word verification in noise using the generalized word posterior probability (GWPP). In computing GWPP, reduced search space, relaxed time registrations of hypothesized words in the word graph, and optimal acoustic and language model weights are employed. The sensitivity of word verification errors with respect to the parameters of GWPP was tested under different SNR conditions. We found that around the optimal parameter settings, there exists a relatively stable region where the total number of word verification errors is fairly insensitive (robust) to the exact choice of the optimal values. Cross-SNR condition tests using a large vocabulary, speaker independent, continuous Japanese speech database (Basic Travel Expression Corpus) confirms the robustness of the GWPP based word verification in different SNR’s. 1. Wai Kit Lo, Frank K. Soong, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2004 | Tone information as a confidence measure for improving Cantonese LVCSR
Yao Qian, Tan Lee, Frank K. Soong |
INTERSPEECH | 3 |
| 2004 | Optimal acoustic and language model weights for minimizing word verification errorsabstractGeneralized word posterior probability (GWPP), a confidence measure for verifying recognized words, needs to equalize and weight acoustic and language model likelihood contributions to minimize verification errors. In this study, we investigate the word verification error surface and use it to optimize these weights and the corresponding verification threshold in a development set. We test three different search algorithms for finding the optimal parameters, including: a full grid search, a gradient-based steepest descent search, and a downhill simplex search. The three search methods yield very similar solutions. Proper acoustic and language model weights, especially the ratio between them, changes with the relative importance (reliability) between the two knowledge sources. For a narrow beam width, the role of the acoustic model is less critical than language model in GWPP-based word verification, which is due to the noisy acoustic information maintained in a narrow beam. Using a large vocabulary continuous Japanese speech database (Basic Travel Expression Corpus), the largest relative improvement obtained is 33.2 % for confidence error rate and 38.7 % for a modified word accuracy. 1. Frank K. Soong, Wai Kit Lo, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2004 | Improved spoken language translation using n-best speech recognition hypothesesabstractWe intended to demonstrate the effect of using N-best speech recognition hypotheses for improving speech translation performance. A log-linear model, which integrated features from speech recognition and statistical machine translation, was used to rescore the translation candidates. Model parameters were estimated by optimizing an objectively measurable but subjectively relevant translation quality metric. Experimental results have shown that the proposed N-best approach improved translation quality over the conventional single-best approach. The improvements were confirmed consistently by several automatic translation evaluation metrics. 1. Ruiqiang Zhang, Gen-ichiro Kikui, Hirofumi Yamamoto, Frank K. Soong, Taro Watanabe, Eiichiro Sumita, Wai Kit Lo |
INTERSPEECH | 4 |
| 2003 | Combining neighboring filter channels to improve quantile based histogram equalizationabstractA mismatch between the training data and the test condition of an automatic speech recognition system usually deteriorates the recognition performance. Quantile based histogram equalization can increase the system's robustness by approximating the cumulative density function of the current signal and then reducing an eventual mismatch based on this estimate. In a first step each output of the mel scaled filter bank can be transformed independent from the others. This paper describes an improved version of the algorithm that combines neighboring filter channels. On several databases recorded in real car environment the recognition error rates could be significantly reduced with this new approach. Florian Hilger, Hermann Ney, Olivier Siohan, Frank K. Soong |
ICASSP (1) | 4 |
| 2003 | Optimal clustering of multivariate normal distributions using divergence and its application to HMM adaptationabstractWe present an optimal clustering algorithm for grouping multivariate normal distributions into clusters using the divergence, a symmetric, information-theoretic distortion measure based on the Kullback-Liebler distance. Optimal solutions for normal distributions are shown to be obtained by solving a set of Riccati matrix equations and the optimal centroids are found by alternating the mean and covariance matrix intermediate solutions. The clustering performance of the new algorithm compared favorably against the conventional, non-optimal clustering solutions of sample mean and sample covariance in its overall rate-distortion and even distributions of samples across clusters. The resultant clusters were further tested on unsupervised adaptation of HMM parameters in a framework of structured maximum a posterior linear regression (SMAPLR). The Wall Street Journal database was used for the adaptation experiment. The recognition performance with respect to the word error rate, was significantly improved from a nonoptimal centroid (sample mean and covariance) of 32.6% to 27.6% and 27.5% for the diagonal and full covariance matrix cases, respectively. Tor André Myrvoll, Frank K. Soong |
ICASSP (1) | 2 |
| 2003 | Modeling Cantonese pronunciation variation by acoustic model refinementabstractPronunciation variations can be roughly classified into two types: a phone change or a sound change [1][2]. A phone change happens when a canonical phone is produced as a different phone. Such a change can be modeled by converting the baseform (standard) phone to a surfaceform (actual) phone. A sound change happens at a lower, phonetic or subphonetic level within a phone and it cannot be modeled well by either the baseform or the surfaceform phone alone. We propose here to refine the acoustic models to cope with sound changes by (1) sharing the Gaussian mixture components of HMM states in the baseform and the surfaceform models; (2) adapting the mixture components of the baseform models towards those of the surfaceform models; (3) selectively reconstructing new acoustic models through sharing or adapting. The proposed pronunciation modeling algorithms are generic and can, in principle, be applied to different languages. Specifically, they were tested in a Cantonese speech recognition database. Relative word error rate reductions of 5.45%, 2.53%, and 3.04 % have been achieved using the three approaches, respectively. 1. Patgi Kam, Tan Lee, Frank K. Soong |
INTERSPEECH | 3 |
| 2003 | On divergence based clustering of normal distributions and its application to HMM adaptation
Tor André Myrvoll, Frank K. Soong |
INTERSPEECH | 2 |
| 2002 | A dynamic in-search discriminative training approach for large vocabulary speech recognitionabstractIn this paper, we propose a dynamic in-search discriminative training approach of a large-scale HMM model for large vocabulary speech recognition. A previously proposed data selection method is used to choose competing hypotheses dynamically during Viterbi beam search procedure. Particularly, all active word-ending paths are examined during search with reference transcription to identify competing tokens for different HMM's. Then HMMs are re-estimated based on an GPD-based discriminative training to minimize total number of possible error tokens among all collected competing tokens. In this way, recognition errors, e.g., word error rate, in training data can be reduced indirectly. The proposed approach is flexible enough to run in a batch or incremental mode. Also, the method can efficiently be implemented to process large amount of training data and update a large-scale state-tied HMM: set for large vocabulary recognition tasks. Some preliminary results on DARPA communicator task show the new discriminative training method can improve recognition performance over our best ML-trained system. Hui Jiang 0001, Olivier Siohan, Frank K. Soong |
ICASSP | 3 |
| 2002 | Classifier design for verification of multi-class recognition decisionabstractThis paper investigates a 2-class classifier approach with the aim of improving the word verification performance. The classifier operates on a discriminant function which is a linear combination of the smoothed likelihood ratios for the N-best candidates and the background (BG) and out-of-vocabulary (OOV) filler models, and is optimized using discriminative training to minimize the classification error. This paper discusses several strategies involving the likelihood ratio based formulation and the use of N-best candidates and the BG and OOV models in the classifier. In word verification experiments using a connected-digit database containing utterances recorded in a moving car with a hands-free microphone, the likelihood ratio based formulation achieved a relative error reduction of 35% in comparison with a likelihood based formulation. In addition, we observed that the use of N-best candidates and the BG and OOV models improved the performance with a relative error reduction of roughly 10%. Tomoko Matsui, Frank K. Soong, Biing-Hwang Juang |
ICASSP | 2 |
| 2002 | Bell labs approach to Aurora evaluation on connected digit recognitionabstractABSTRACTIn this paper we study various front-endfeatures, mod-eling and adaptation algorithms on the Aurora 3 databases,including auditory, moment, and AM-FMmodulation fea-tures, context-dependentdigit models, segmental K-meanstraining, discriminative training, and model adaptations.The evaluation results on Aurora 3 are presented with abrief summary of our Aurora 2 results.1. INTRODUCTIONThe Aurora evaluation is for researchers to test their algo-rithms on noise robustness and compare results measuredon the same databases. So far, there are two tasks on theAurora evaluation, Aurora 2 and 3, both are for connecteddigit recognition. While the Aurora 2 databases use thecontrolled experiments by adding noise digitally to cleanEnglish digit strings [1], the Aurora 3 databases are col-lected in a real-worldcar environment in 4 languages. Inthis paper, we report our evaluation results on two of thelanguages, Spanish and German.2. BELL LABS APPROACHESIn this section, we present our baseline system then describethe different feature sets that have been used for this eval-uation. Alternative training strategies and acoustic modeladaptation techniques are also reviewed.A. Context-DependentModel: Similar to last year ap-proach [1], we have decided to use context-dependent(CD)digit models, together with Bell Labs recognition engine asbackend. This contrasts with the officialAurora backendthat is based on whole-worddigit models and the HTK en-gine. The officialbackend setup typically leads to poorerresults, especially in larger databases, and we believe that abetter baseline is beneficialto properly study the effect ofdifferent front-endson the finalrecognition performance.Last year, we investigated several approaches to buildCD digit models. Given the limited amount of trainingdata, especially in the Aurora3 databases, it is required torely on some tying techniques to build CD digit models.The Head-Body-Tail digit model structure (HBT) assumesthat CD digit models are built by concatenating a left-context-dependentunit (head) with a context-independentunit(body)followedbyaright-context-dependentunit(tail).For example, assuming that the lexicon contains 10 digitsplus a silence model, each digit model consists of a set of 1body, 11 heads and 11 tails (representing all left/right con-texts) [2]. We typically model each head and tail with a3-stateHMM, while a 4-stateHMM is used for each body.Most of the experiments done this year have been based onthe HBT structure. CD digit models can also be built as tri-phone modelsusing a decision tree. This is the approach weintroduced last year [1], and some of this year experimentshave been carried out using this model topology.B. Auditory Feature: The new auditory front-endinour recognition system was developed to mimic the robusthuman hearing in adverse acoustic environments [3, 4]. Inthe front-end,efficientsignal processing functions were im-plemented to satisfy both real-timeand computation costrequirements. Based on the analysis of the outer and mid-dle ear, a transfer function was constructed to replace thecommonly used preemphasis filter, and then a new set ofdigital auditory filters,which simulate auditory filteringinthe cochlea, replaces those used in the MFCC and PLP.The auditory feature extraction procedure consists of: anouter-middle-eartransfer function, FFT, frequency conver-sion from linear to the Bark scale, auditory filtering,non-linearity,and discrete cosine transform (DCT). In our previ-ous study[3], the feature has been evaluated in two tasks:connected-digit and large vocabulary, continuous speechrecognitionundervariousnoiseconditions,usingbothhand-setand hands-freedatainlandlineand wirelesstransmissionwith additive car and babble noise. Compared with theLPCC, MFCC, MEL-LPCC,and PLP features, the audi-tory feature achieved significantperformance improvement Jingdong Chen, Dimitris Dimitriadis, Hui Jiang 0001, Tor André Myrvoll, Olivier Siohan, Frank K. Soong |
INTERSPEECH | 7 |
| 2002 | Recognition of noisy speech using normalized moments
Jingdong Chen, Yiteng Huang, Frank K. Soong |
INTERSPEECH | 4 |
| 2001 | Hierarchical stochastic feature matching for robust speech recognitionabstractIn this paper we investigate how to improve the robustness of a speech recognizer in a noisy, mismatched environment when only a single or a few test utterances are available for compensating the mismatch. A new hierarchical tree-based transformation is proposed to enhance the conventional stochastic matching algorithm in the cepstral feature space. The tree-based hierarchical transformation is estimated in two criteria: i) maximum likelihood (ML) using the current test utterance; ii) sequential maximum a posterior (MAP) using the current and previous utterances. Recognition results obtained using a hands-free database show the proposed feature compensation is robust. Significant performance improvement has been observed over the conventional stochastic matching. Hui Jiang 0001, Frank K. Soong |
ICASSP | 2 |
| 2001 | Evaluating the Aurora connected digit recognition task - a bell labs approach
Mohamed Afify, Hui Jiang 0001, Filipp Korkmazskiy, Olivier Siohan, Frank K. Soong, Arun C. Surendran |
INTERSPEECH | 7 |
| 2001 | A data selection strategy for utterance verification in continuous speech recognitionabstractABSTRACTIn this paper, we propose the concept of rival for verifying hy-pothesis in speech recognition. A likelihood ratio test, based onthe rivals model, are investigated for utterance verification in con-tinuous speech recognition. We present a data selection strategyto identity useful subsets of training data to train rival model auto-matically from training data. And a single pass strategy for utter-ance verification, namely verification-in-search, is also proposed.Some preliminary experiments on DARPA Communicator traveltask have shown the rival models give better verification perfor-mance in terms of identifying mis-recognized words from the out-put of our baseline recognizer.1. INTRODUCTIONRecent advances in automatic speech recognition (ASR) technol-ogy have enabled ASRsystems tomigrate from laboratory to manyservices and products. However, in many practical applications, itbecomes more desirable and urgent to equip a speech recognizerwith utterance verification(UV).[4] Utterance verification isa pro-cedure used to verify how reliable are the results from a speechrecognizer. Usually, a quantitative score, also called confidencemeasure, is used to indicate the reliability of every recognition de-cision. Based on the confidence measure, a series of further ac-tions can be taken after recognition, e.g., to reject or remedy therecognition results. Utterance verification is a crucial technique tomake today’s speech recognizers more “intelligent” than ever be-fore. For instance, a speech recognizer with a powerful UV capa-bility will be able to smartly reject non-speech noises, detect/rejectout-of-vocabulary words, even correct some potential recognitionmistakes, guide the system to perform unsupervised learning, andprovide side information to assist high level speech understanding,etc.Extensivestudies on utterance verification havebeen performedrecently in the literature. One of the most important progressesis to cast utterance verification scenario as a statistical hypothesistesting problem.[4, 7] According to the Neyman-Pearson Lemma,an optimal test is to evaluate a likelihood ratio between two hy-potheses, Hui Jiang 0001, Frank K. Soong |
INTERSPEECH | 2 |
| 2001 | An auditory system-based feature for robust speech recognitionabstractAn auditory feature extraction algorithm for robust speech recognition in adverse acoustic environments is presented. The feature computation is comprised of an outer-middle-ear transfer function, FFT, frequency conversion from linear to the Bark scale, auditory filtering, nonlinearity, and discrete cosine transform. The feature is evaluated in two tasks: connected-digit recognition and large vocabulary continuous speech recognition. The tested data were under various noise conditions, including handset and hands-free speech data in landline and wireless communications with additive car and babble noise. Compared with the LPCC, MFCC, MEL-LPCC, and PLPfeatures, the proposed feature has an average 20 % to 30 % string error rate reduction on the connected-digit task, and 8 % to 14 % word error rate reduction on the Wall Street Journal task in various additive noise conditions. 1. Frank K. Soong, Olivier Siohan |
INTERSPEECH | 2 |
| 2001 | A real-time Japanese broadcast news closed-captioning systemabstractThis paper describes a collaboration between Bell Labs and NHK (Japan Broadcasting Corp.) STRL to develop a real-time large vocabulary speech recognition system for live closed-captioning of NHK news programs. Bell Labs broadcast news recognition engine consists of a two-pass decoder using bigram language models (LM) and right biphone models during the first pass, and trigram LM with within-word triphone models in the second pass. Various pruning strategies are used to achieve real time decoding, together with a noise compensation procedure aimed at improving recognition on noisy segments of the program. The system operates in a real-time mode and delivers less than 2% of word error rate (WER) on studio news conditions and about 5% of WER on noisy news and reporter speech when evaluated on a real broadcast news program. Olivier Siohan, Akio Ando, Mohamed Afify, Hui Jiang 0001, Kazuo Onoe, Frank K. Soong, Qiru Zhou |
INTERSPEECH | 9 |
| 2000 | A high-performance auditory feature for robust speech recognitionabstractAn auditory feature extraction algorithm for robust speech recognition in adverse acoustic environments is proposed. Based on the analysis of human auditory system, the feature extraction algorithm consists of several modules: FFT, outer-middle-ear transfer function, frequency conversion from linear to Bark scales, auditory filtering, nonlinearity, and discrete cosine transform. Three recognition experiments have been conducted on connected digit recognition in wireless and land-line communications using handsets and handsfree microphones. Compared to LPCC and MFCC features, the proposed feature has shown 11% to 23% error-rate reductions on average in handset and hands-free acoustic environments in the experiments. Frank K. Soong, Olivier Siohan |
INTERSPEECH | 2 |
| 2000 | Hands-free human-machine dialogue - corpora, technology and evaluation
Frank K. Soong, Eric A. Woudenberg |
INTERSPEECH | 1 |
| 1999 | Hidden Markov models with divergence based vector quantized variancesabstractThis paper describes a method to significantly reduce the complexity of continuous density HMM with only a small degradation in performance. The proposed method is noise-robust and may perform even better than the standard algorithm if training and testing noise conditions are not matched. The method is based on approximating the variance vectors of the Gaussian kernels by a vector quantization (VQ) codebook of a small size. The quantization of the variance vectors is done using an information theoretic distortion measure. Closed form expressions are given for the computation of the VQ codebook and the superiority of the proposed distortion measure over the Euclidean distance is demonstrated. The effectiveness of the proposed method is shown using the connected TI digits database and a noisy version of it. For the connected TI digit database, the proposed method shows that by quantizing the variance to 16 levels we can maintain recognition performance within 1% degradation of the original VR system. In comparison, with Euclidean distortion, a size 256 codebook is needed for a similar error rate. Jae H. Kim, Raziel Haimi-Cohen, Frank K. Soong |
ICASSP | 3 |
| 1999 | A block least squares approach to acoustic echo cancellationabstractWe propose an efficient block least squares (BLS) algorithm for acoustic echo cancellation. The high computation and memory requirements associated with a long room echo make the simple, gradient-based LMS filter a more acceptable commercial solution than a full-fledged LS canceler. However, the LMS echo canceler has slower convergence and worse steady-state performance than its LS counterpart. In the proposed BLS approach, the autocorrelation and cross-correlation of the source and echo, required in solving the LS normal equations, are performed once per block using FFTs. With appropriate data windowing the autocorrelation matrix is constrained to be Toeplitz, allowing the corresponding normal equations to be solved efficiently. The positive definiteness of the autocorrelation function eliminates the stability problems of other fast LS algorithms. BLS can reduce the echo residual to the level of background noise, allowing a residual power based, statistical near-end speech detector to be devised. Performance in real environments under various settings of filter length, SNR, near-end speech presence, etc., is investigated. Eric A. Woudenberg, Frank K. Soong, Biing-Hwang Juang |
ICASSP | 2 |
| 1998 | Improved utterance rejection using length dependent thresholds
Sunil K. Gupta, Frank K. Soong |
ICSLP | 2 |
| 1997 | Generalized mixture of HMMs for continuous speech recognitionabstractThis paper presents a new technique for modeling heterogeneous data sources such as speech signals received via distinctly different channels. Such a scenario arises when an automatic speech recognition system is deployed in wireless telephony in which highly heterogeneous channels coexist and interoperate. The problem is that a simple model may become inadequate to describe accurately the diversity of the signal, resulting in an unsatisfactory recognition performance. To deal with such a problem, we propose a generalized mixture model (GMM) approach. For speech signals, in particular, we use mixtures of hidden Markov models (i.e., GMHMM, generalized mixture of HMMs). By applying discriminative training for GMHMM we obtained 1.0% word error rate for the recognition of the digits strings from the wireless database, comparing to 1.4% word error rate for the conventional HMM based discriminative technique. Filipp Korkmazskiy, Biing-Hwang Juang, Frank K. Soong |
ICASSP | 3 |
| 1996 | High-accuracy connected digit recognition for mobile applicationsabstractWe present a connected digit recognition system with low storage and computational complexity which achieves good performance in car noise. Our system uses the TI-DIGITS database with additive car noise for training whole-word digit and background models. A digit accuracy of 96.1% is obtained on a 15-speaker database collected in a car using an open microphone with an average SNR of approximately 2 dB. There is a further error reduction of almost 35% if the top two candidate strings are considered using a traceback based N-best algorithm. The system can be implemented on a currently available fixed-point DSP chip. We show that significant performance improvements are obtained by using two-level cepstral mean subtraction (CMS), gender-dependent models and a decoding grammar constraining the possible lengths of digit strings. Sunil K. Gupta, Frank K. Soong, Raziel Haimi-Cohen |
ICASSP | 2 |
| 1996 | Quantizing mixture-weights in a tied-mixture HMMabstractIn this paper, we describe new techniques to signicantly reduce computational, storage and memory access requirements of a tied-mixture HMM based speech recognition system.Although continuous mixture HMMs oer improved recognition performance, we show that tied-mixture HMMs may oer signicant advantage in complexity reduction for low-cost implementations.In particular, we consider two tasks: (a) connected digit recognition in car noise; and (b) sub-word modeling for command word recognition in a noisy oce environment.We show that quantization of mixture weights can provide an almost three fold reduction in mixture-weight storage requirements without any signicant loss in recognition performance.Furthermore, we show that by combining mixture-weight quantization with techniques such a s V Q-Assist the computational and memory access requirements can be reduced by almost 60-80% without any degradation in recognition performance. Sunil K. Gupta, Frank K. Soong, Raziel Haimi-Cohen |
ICSLP | 2 |
| 1995 | An orthogonal polynomial representation of speech signals and its probabilistic model for text independent speaker verificationabstractA segmental probabilistic model based on an orthogonal polynomial representation of speech signals is proposed. Unlike the conventional frame based probabilistic model, this segment based model concatenates the similar acoustic characteristics of consecutive frames into an acoustic segment and represents the segment by an orthogonal polynomial function. An algorithm which iteratively performs recognition and segmentation processes is proposed for estimating the parameters of the segment model. This segment model is applied in the text independent speaker verification. For a 20-speaker database, the experimental results show that the performance by using segment models is better than that by using the conventional frame based probabilistic model. The equal error rate can be reduced by 3.6% when the models are represented by 64-mixture density functions. Chi-Shi Liu, Hsiao-Chuan Wang, Frank K. Soong, Chao-Shih Huang |
ICASSP | 3 |
| 1995 | Large vocabulary, word-based Mandarin dictation system
Jung-Kuei Chen, Lin-Shan Lee, Frank K. Soong |
EUROSPEECH | 3 |
| 1995 | Optimizing baseforms for HMM-based speech recognition
Torbjørn Svendsen, Frank K. Soong, Heiko Purnhagen |
EUROSPEECH | 2 |
| 1994 | Discriminative training of high performance speech recognizer using N best candidatesabstractProposes an N-best candidates based, discriminative training procedure for constructing high performance HMM speech recognizers. The algorithm has two features: (1) a new frame-level loss function; (2) N best candidates are used for training. The new frame-level loss function, defined as a rectified log likelihood difference between the correct and other competing hypotheses, is minimized over all training utterances. Two speech recognition applications have been tested: speaker independent, small vocabulary (10 Mandarin Chinese digits), continuous speech recognition; and a speaker-trained, large vocabulary (5,000 commonly used Chinese words), isolated word recognition. Significant performance improvement over the traditional maximum likelihood trained HMMs has been obtained. In the connected Chinese digit recognition experiment, the string error rate is reduced from 17% to 10.8% for unknown length decoding and from 8.2% to 5.2% for known length decoding. In the large vocabulary, isolated word recognition experiment, the recognition error rate is improved from 6.8% to 3.8%.> Jung-Kuei Chen, Frank K. Soong |
ICASSP (1) | 2 |
| 1994 | Large vocabulary word recognition based on tree-trellis searchabstractIn this paper we propose a large vocabulary (90000 words), Chinese (Mandarin) word recognizer based on the tree-trellis fast search algorithm. The recognizer is divided into 3 modules: local likelihood computation, a forward trellis search and a backward tree search. In the forward trellis search, a free syllable decoding is performed without a language model and a partial path map is created. The best-first tree search is then applied backward along a lexicon, which is arranged as a syllabic tree, to find the N-best word candidates. In the experiment, context-dependent subsyllabic HMMs were trained with a new discriminative training method. When it is evaluated on a speaker-trained database, the recognizer achieved a word error rate of 5% for the full size (90000 words) vocabulary and 1.7% for a smaller subset (5000 words) vocabulary. A real-time demo system has also been implemented on an SGI R-4000 workstation.> Jung-Kuei Chen, Frank K. Soong, Lin-Shan Lee |
ICASSP (2) | 2 |
| 1994 | Cepstral channel normalization techniques for HMM-based speaker verification
Aaron E. Rosenberg, Frank K. Soong |
ICSLP | 3 |
| 1994 | The use of tree-trellis search for large-vocabulary Mandarin polysyllabic word speech recognition
Eng-Fong Huang, Frank K. Soong, Hsiao-Chuan Wang |
Comput. Speech Lang. | 2 |
| 1994 | A Minimum Error Rate Pattern Recognition Approach to Speech RecognitionabstractIn this paper, a minimum error rate pattern recognition approach to speech recognition is studied with particular emphasis on the speech recognizer designs based on hidden Markov models (HMMs) and Viterbi decoding. This approach differs from the traditional maximum likelihood based approach in that the objective of the recognition error rate minimization is established through a specially designed loss function, and is not based on the assumptions made about the speech generation process. Various theoretical and practical issues concerning this minimum error rate pattern recognition approach in speech recognition are investigated. The formulation and the algorithmic structures of several minimum error rate training algorithms for an HMM-based speech recognizer are discussed. The tree-trellis based N-best decoding method and a robust speech recognition scheme based on the combined string models are described. This approach can be applied to large vocabulary, continuous speech recognition tasks and to speech recognizers using word or subword based speech recognition units. Various experimental results have shown that significant error rate reduction can be achieved through the proposed approach. Wu Chou, Biing-Hwang Juang, Frank K. Soong |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 1994 | An N-best candidates-based discriminative training for speech recognition applicationsabstractThe authors propose an N-best candidates-based discriminative training procedure for constructing high-performance HMM speech recognizers. The algorithm has two distinct features: N-best hypotheses are used for training discriminative models; and a new frame-level loss function is minimized to improve the separation between the correct and incorrect hypotheses. The N-best candidates are decoded based on their recently proposed tree-trellis fast search algorithm. The new frame-level loss function, which is defined as a halfwave rectified log-likelihood difference between the correct and competing hypotheses, is minimized over all training tokens. The minimization is carried out by adjusting the HMM parameters along a gradient descent direction. Two speech recognition applications have been tested, including a speaker independent, small vocabulary (ten Mandarin Chinese digits), continuous speech recognition, and a speaker-trained, large vocabulary (5000 commonly used Chinese words), isolated word recognition. Significant performance improvement over the traditional maximum likelihood trained HMMs has been obtained. In the connected Chinese digit recognition experiment, the string error rate is reduced from 17.0 to 10.8% for unknown length decoding and from 8.2 to 5.2% for known length decoding. In the large vocabulary, isolated word recognition experiment, the recognition error rate is reduced from 7.2 to 3.8%. Additionally, they have found that using more relaxed decoding constraints in preparing N-best hypotheses yields better recognition results.> Jung-Kuei Chen, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 2 |
| 1994 | A fast algorithm for large vocabulary keyword spotting applicationabstractPresents a fast algorithm for spotting a large number of keywords in unconstrained, continuous speech using an HMM-based continuous speech recognizer. This fast algorithm is based on a two stage scheme. In the first stage, the forward backward search is performed for detecting N most likely common subwords. In the second stage, the tree-trellis search is carried out to determine the optimum keyword by traversing the tree-structural vocabulary in an effective way. Compared with the conventional whole-word based keyword spotting algorithm, the proposed fast algorithm can drastically reduce the computational cost.> Eng-Fong Huang, Hsiao-Chuan Wang, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 3 |
| 1993 | Optimal quantization of LSP parametersabstractTwo nonuniform aspects of the line spectrum pair (LSP) linear predictive coding (LPC) parameters are investigated, including nonuniform statistical distributions and spectral sensitivities of adjacent LSP frequency differences. Based upon these two nonuniform properties, a globally optimal scalar quantizer is designed for each differential LSP frequency. The design algorithm is dynamic programming based and minimization of a nontrivial data dependent spectral distortion is adopted as the optimality criterion. At 32 bits/frame, the new LSP quantizer achieves a 1-dB average log spectral distortion, a commonly accepted level for reproducing perceptually transparent spectral information. The quantization performance has also been shown to be robust across different speakers and databases.> Frank K. Soong, Biing-Hwang Juang |
IEEE Trans. Speech Audio Process. | 1 |
| 1992 | Continuous probabilistic acoustic map for speaker recognitionabstractA continuous probabilistic acoustic map (CPAM) approach to speaker recognition is investigated. In the CPAM formulation, the speech input of a speaker is parameterized as a mixture of tied, universal probability density functions (PDFs) with either a CPAM model alone for text-independent operation or a CPAM-based hidden Markov model (HMM) for text-dependent operation. A continuously spoken digit database of 20 speakers (10 M, 10 F) is used to evaluate the CPAM approach in both identification and verification performance. The CPAM approach is shown to perform better than a vector quantization based approach in text-independent speaker recognition, and as well as the text-dependent, conventional, continuous mixture HMM approach with significant representation efficiency. In particular, the CPAM-based HMM achieves an identification error rate of 1.7% and a verification equal-error rate of 4.0% with a CPAM of 128 PDFs while a conventional, continuous mixtures HMM needs 400 PDFs to achieve corresponding error rates of 1.9% and 4.0% using the same combined cepstral features and three-digit test utterances.> Belle L. Tseng, Frank K. Soong, Aaron E. Rosenberg |
ICASSP | 2 |
| 1992 | The use of cohort normalized scores for speaker verification
Aaron E. Rosenberg, Joel DeLong, Biing-Hwang Juang, Frank K. Soong |
ICSLP | 5 |
| 1992 | Continuous mixture HMM-LR using the a* algorithm for continuous speech recognition
Kouichi Yamaguchi, Shigeki Sagayama, Kenji Kita, Frank K. Soong |
ICSLP | 4 |
| 1991 | A tree-trellis based fast search for finding the N-best sentence hypotheses in continuous speech recognitionabstractA novel tree-trellis based fast search for finding the N-best sentence hypotheses in continuous speech recognition is presented. The search consists of a forward time-synchronous trellis search and a backward time-asynchronous tree search. The Viterbi algorithm is used for recording the scores of all partial paths in a trellis time synchronously. Then a backward A* algorithm based tree search is used to extend partial paths time asynchronously. Extended partial paths in the backward tree search are rank ordered in a stack by their corresponding best possible scores of the remaining paths which are prerecorded in the forward trellis path map. In each path growing cycle, the current best partial path, which is at the top of the stack, is extended by the best possible one arc (word) extension. The tree-trellis search is different from the traditional time synchronous Viterbi search in its ability to find not just the best but the N best paths of different word content.> Frank K. Soong, Eng-Fong Huang |
ICASSP | 1 |
| 1990 | Statistical segmentation and word modeling techniques in isolated word recognitionabstractA speech recognition system is described using a combination of statistical segment and word modeling. Segment models are constructed by first segmenting training data automatically and then grouping the resultant segments into clusters. Mixtures of Gaussian densities are used to model each segment cluster. In order to integrate the segment models into word models, a generalization of the hidden Markov model approach is proposed. Experimental results on a multispeaker recognition system for alpha-digits demonstrate that the new approach improved the performance of conventional whole-word-based models. In particular, the word models show good discrimination abilities for differentiating phonetically similar words such as the E-set alphabet.> S. A. Euler, Biing-Hwang Juang, Frank K. Soong |
ICASSP | 4 |
| 1990 | A probabilistic acoustic map based discriminative HMM trainingabstractA hidden Markov model (HMM) training procedure is proposed for improving the discriminative power of a maximum-likelihood (ML)-based HMM. The discriminative HMM consists of two component models, a master model and a slave model. The master model is the conventional ML model. The slave model is obtained by aligning training tokens of a word with all but the correct word master models. All models are trained by projecting acoustic observations onto a set of common probabilistic basis functions, which is called a probabilistic acoustic map, and the output probability of a model is represented as a weighted sum of the basis functions. The proposed algorithm was tested on a 39-word, alpha-digit database of 100 speakers (50 male and 50 female). Specifically, the highly confusable E-set words were separately tested. Experimental results indicate that the new training procedure improved the recognition accuracy of the E-set words by 4.3% and of all 39 words by 3.6%.> Eng-Fong Huang, Frank K. Soong |
ICASSP | 2 |
| 1990 | Speaker recognition based on source coding approachesabstractThe use of nonmemoryless source coders in speaker recognition problems is studied, and the effects of source variations, including speaking inconsistency and channel mismatch, in source coder designs for the intended application are discussed. It is found that incorporation of memory in source coders in general enhances the speaker recognition accuracy but that more remarkable improvements can be accomplished by properly including potential source variations in the coder design/training. An experiment with a 100-speaker database shows a 99.5% recognition accuracy.> Biing-Hwang Juang, Frank K. Soong |
ICASSP | 2 |
| 1990 | Sub-word unit talker verification using hidden Markov modelsabstractA talker verification system based on characterizing talker utterances as sequences of subword units represented by hidden Markov models (HMMs) was implemented and tested. Two types of subword units were studied: phonelike units (PLUs) and acoustic segment units (ASUs). PLUs are based on phonetic transcriptions of spoken utterances and ASUs are extracted directly from the acoustic signal without use of any linguistic knowledge. The ASU representation has the advantage of not requiring transcriptions of training utterances. Verification performance was evaluated on a 100-talker database of 20000 isolated digit utterances. The experiments show only small differences in performance between PLU- and ASU-based representations. Overall, the verification equal-error rate is approximately 7 to 8% for one-digit test utterances (approximately 0.5 s in duration) and 1% or less for seven-digit test utterances (approximately 3.5 s in duration). In addition, a technique for updating models, using data from current test utterances, was devised and implemented. Using this adaptation technique, the error rate falls to 6% for one-digit utterances and less than 0.5% for seven-digit utterances. The experiment confirms that excellent verification performance can be obtained using HMMs of subword units.> Aaron E. Rosenberg, Frank K. Soong |
ICASSP | 3 |
| 1990 | Optimal quantization of LSP parameters using delayed decisionsabstractA previously published study by the authors (Proc. ICASSP, p.394-7, 1988) of optimal quantization of line spectral pair (LSP) parameters is extended by incorporating delayed decisions coding (in frequency). The A* algorithm is proposed for finding the best quantization bit pattern of LSP frequency differences. The best coding pattern is obtained efficiently without an exhaustive, hence prohibitive, search. The proposed search achieves a better rate-distortion performance than the best results obtained in the previous study. At 30 bits/frame, a net gain of 2 bits/frame over the previous results, the novel method achieves 1-dB average spectral distortion. Most importantly, the number of frames with large spectral distortions (>2 dB), which can be subjectively disturbing and degrade the perceived quality of a speech coder, is significantly reduced. The search complexity of the A* algorithm is moderate. While the peak load is comparable to a nonoptimal M-algorithm, the average load is about an order of magnitude lower.> Frank K. Soong, Biing-Hwang Juang |
ICASSP | 1 |
| 1990 | Experiments in automatic talker verification using sub-word unit hidden Markov models
Aaron E. Rosenberg, Frank K. Soong, Maureen A. McGee |
ICSLP | 3 |
| 1990 | A tree-trellis based fast search for finding the n best sentence hypotheses in continuous speech recognitionabstractA novel tree-trellis based fast search for finding the N-best sentence hypotheses in continuous speech recognition is presented. The search consists of a forward time-synchronous trellis search and a backward time-asynchronous tree search. The Viterbi algorithm is used for recording the scores of all partial paths in a trellis time synchronously. Then a backward A* algorithm based tree search is used to extend partial paths time asynchronously. Extended partial paths in the backward tree search are rank ordered in a stack by their corresponding best possible scores of the remaining paths which are prerecorded in the forward trellis path map. In each path growing cycle, the current best partial path, which is at the top of the stack, is extended by the best possible one arc (word) extension. The tree-trellis search is different from the traditional time synchronous Viterbi search in its ability to find not just the best but the N best paths of different word content. > Frank K. Soong, Eng-Fong Huang |
ICSLP | 1 |
| 1989 | Word recognition using whole word and subword modelsabstractThe problem of how to select and construct a set of fundamental unit statistical models suitable for speech recognition is addressed. A unified framework is discussed which can be used to accomplish the goal of creating effective basic models of speech. The performances of three types of fundamental units, namely whole word, phoneme-like, and acoustic segment units, in a 1109-word vocabulary speech recognition task are compared. The authors point out the relative advantages of each type of speech unit based on the results of a series of recognition experiments.> Biing-Hwang Juang, Frank K. Soong, Lawrence R. Rabiner |
ICASSP | 3 |
| 1989 | A phonetically labeled acoustic segment (PLAS) approach to speech analysis-synthesisabstractA phonetically labeled acoustic segment (PLAS) approach is proposed for speech analysis-synthesis. The goal is to develop a unified framework for general speech processing by means of a bidirectional context-constrained mapping between a phonetic space and an acoustic space. The PLAS analysis module is a continuous phone (phoneme) recognizer, while the PLAS synthesis module is a phonetically organized acoustic database. To regulate the proposed mapping in a phonetically structured manner, phone context-dependency was imposed in phone modeling, recognition, and synthesis. The PLAS approach was tested successfully on a database of continuously spoken Japanese utterances recorded by a single male talker. The automatic segmentation boundaries derived from modeling PLAS units agreed well with corresponding manual segmentation points, i.e. they were within a +or-20-ms interval 95% of the time. A 4% phoneme recognition error rate was obtained in a continuous recognition test. Natural-sounding speech was synthesized at an average bit rate of 55 b/s allocated to segmental information.> Frank K. Soong |
ICASSP | 1 |
| 1988 | A segment model based approach to speech recognitionabstractProposes a global acoustic segment model for characterizing fundamental speech sound units and their interactions based upon a general framework of hidden Markov models (HMM). Each segment model represents a class of acoustically similar sounds. The intra-segment variability of each sound class is modeled by an HMM, and the sound-to-sound transition rules are characterized by a probabilistic intersegment transition matrix. An acoustically-derived lexicon is used to construct word models based upon subword segment models. The proposed segment model was tested on a speaker-trained, isolated word, speech recognition task with a vocabulary of 1109 basic English words. In the current study, only 128 segment models were used, and recognition was performed by optimally aligning the test utterance with all acoustic lexicon entries using a maximum likelihood Viterbi decoding algorithm. Based upon a database of three male speakers, the average word recognition accuracy for the top candidate was 85% and increased to 96% and 98% for the top 3 and top 5 candidates, respectively.> Frank K. Soong, Biing-Hwang Juang |
ICASSP | 2 |
| 1988 | High performance connected digit recognition, using hidden Markov modelsabstractAlgorithms for connected-word recognition based on whole-word reference patterns have become increasingly sophisticated and have been shown capable of achieving high recognition performance for small or syntax-constrained moderate-size vocabularies in a speaker-trained mode. An enhanced analysis feature set consisting of both instantaneous and transitional spectral information is used and the hidden-Markov-model-based connected digit recognizer is tested in speaker-trained, multispeaker, and speaker-independent modes. The performance achieved was 0.35, 1.65 and 1.75% string error rates, respectively, for known length strings and 0.78, 2.85 and 2.94% string error rates, respectively, for unknown length strings.> Lawrence R. Rabiner, Jay G. Wilpon, Frank K. Soong |
ICASSP | 3 |
| 1988 | Optimal quantization of LSP parameters [speech coding]abstractTwo nonuniform aspects of the line spectrum pair (LSP) linear predictive coding parameters are investigated, including nonuniform statistical distributions and spectral sensitivities of adjacent LSP frequency differences. Based on these two nonuniform properties, a novel, globally optimal, scalar quantizer is designed for each differential LSP frequency. The design algorithm is based on dynamic programming, and minimization of a nontrivial data dependent spectral distortion is adopted as the optimality criterion. The LSP quantizer can achieve a 1-dB average log spectral distortion, which is the experimental difference limen (DL) for producing perceptually transparent spectral information at 32 bits/frame. The quantization performance is shown to be fairly robust across different speakers and databases.> Frank K. Soong, Bling-Hwang Juang |
ICASSP | 1 |
| 1987 | A training procedure for a segment-based-network approach to isolated word recognitionabstractIn this paper, we propose a complete training procedure for creating a subword-based network and test it in an isolated word recognition experiment. We first hand segment one training token per word into contiguous subword segments with the aid of an interactive program that can display and playback various acoustic features of an utterance. The subword segmental units adopted in this paper consist of four different sound classes including: stationary sounds, fast transitional sounds, slow transitional sounds plus consonant clusters and others. The hand segmented token is used to initialize a subword-based word network which is then refined by using more training tokens. The refinement is carried out with a two-level dynamic programming (DP) procedure. At the first level, or the word level, an endpoint-relaxed DP algorithm is used to remove any possible endpointing errors and to mark tentative segment boundaries. Between the marked segment boundaries, another endpoint-relaxed DP algorithm is employed at the segment level to refine the segments extracted at the word level. A segment-based word network, which consists of serial and parallel branches, is generated from this training procedure. While serial branches are generated by using acoustically similar segments aligned at the segment level parallel branches are created for accomodating different acoustic manifestations of the same sound class in different phonetic contexts or different pronunciations. A speaker-dependent, isolated word, recognition experiment was carried out. For a four-speaker(2 male and 2 female), English alphabet data base, the segment-based network, when compared with a conventional word-template-based approach, gives improved performance. The word error rate is reduced from 11.2% for the word-based recognizer down to 7.7% for the network-based recognizer; or correspondingly, the number of misrecognized words is reduced from 116 to 80 out of 1040 recognition trials. Frank K. Soong |
ICASSP | 1 |
| 1987 | A frequency-weighted Itakura spectral distortion measure and its application to speech recognition in noiseabstractThe performance of a recognizer based on the Itakura spectral distortion measure deteriorates when speech signals are corrupted by noise, specially if it is not feasible to train and to test the recognizer under similar noise conditions. To alleviate this problem, we consider a more noise-resistant, weighted spectral distortion measure which weights the high SNR regions in frequency more than the low SNR regions. For the weighting function we choose a "bandwidth broadened" test spectrum; it weights spectral distortion more at the peaks than at the valleys of the spectrum. The amount of weighting is adapted according to an estimate of SNR, and becomes essentially constant in the noise-free case. The new measure has the dot product form and computaional efficiency of the Itakura distortion measure in the autocorrelation domain. It has been tested on a 10 speaker, isolated digit data base in a series of speaker independent speech recognition experiments. Additive white Gaussian noise was used to simulate different SNR conditions (from 5 dB to ∞ dB). The new measure performs as well as the original unweighted Itakura distortion measure at high SNR's, and significantly better at medium to low SNRs. At an SNR of 5 dB, the new measure achieves a digit error rate of 12.49% while the original Itakura distortion gives an error rate of 27.6%. The equivalent SNR improvement at low SNR's, is about 5 - 7 dB. Frank K. Soong, Man Mohan Sondhi |
ICASSP | 1 |
| 1987 | On the automatic segmentation of speech signalsabstractFor large vocabulary and continuous speech recognition, the sub-word-unit-based approach is a viable alternative to the whole-word-unit-based approach. For preparing a large inventory of subword units, an automatic segmentation is preferrable to manual segmentation as it substantially reduces the work associated with the generation of templates and gives more consistent results. In this paper we discuss some methods for automatically segmenting speech into phonetic units. Three different approaches are described, one based on template matching, one based on detecting the spectral changes that occur at the boundaries between phonetic units and one based on a constrained-clustering vector quantization approach. An evaluation of the performance of the automatic segmentation methods is given. Torbjørn Svendsen, Frank K. Soong |
ICASSP | 2 |
| 1986 | Evaluation of a vector quantization talker recognition system in text independent and text dependent modesabstractA vector quantization based talker recognition system is described and evaluated. The system is based on constructing highly efficient short-term spectral representations of individual talkers using vector quantization codebook construction techniques. Although the approach is intrinsically text-independent, the system can be easily extended to text-dependent operation for improved performance and security by encoding specified training word utterances to form word prototypes. The system has been evaluated using a 100-talker database of 20,000 spoken digits. In a talker verification mode, average equal-error rate performance of 2.2% for text-independent operation and 0.3% for text-dependent operation is obtained for 7-digit long test utterances. Aaron E. Rosenberg, Frank K. Soong |
ICASSP | 2 |
| 1986 | A high quality subband speech coder with backward adaptive predictor and optimal time-frequency bit assignmentabstractIn this paper we propose a hybrid speech coder which utilizes properties of APC, ATC, and subband coding. Subband splitting is used to reduce the dynamic range of a full-band APC. As a result, the prediction or power gain is reduced and the instability problem associated with the APC coder is alleviated to a large extent. The optimal bit allocations used in ATC is extended in the new coder where the available coding bits are optimally allocated, both in the time and frequency domains. Frame boundary artifacts (such as those found in ATC) are not present due to the time-domain processing nature of SBC. In order to improve the efficiency in quantizing the side information (time and frequency domain signal power and the corresponding bit allocations), vector quantization is used. Adaptive backward predictors are used to further reduce the number of bits allocated to side information, leaving more bits to encode the prediction residual signals. At 16 kbps the coder achieves better than 20 dB segmental signal-to-noise ratio and sounds transparent for most speakers. Frank K. Soong, Richard V. Cox, Nikil Jayant |
ICASSP | 1 |
| 1986 | On the use of instantaneous and transitional spectral information in speaker recognitionabstractThe use of instantaneous and transitional spectral representations of spoken utterances for speaker recognition is investigated. LPC derived-cepstral coefficients are used to represent instantaneous spectral information and best linear fits of each cepstral coefficient over a specified time window are used to represent transitional information. An evaluation has been carried out using a data base of isolated digit utterances over dialed-up telephone lines by 10 talkers. Two vector quantization (VQ) codebooks, instantaneous and transitional, are constructed from training utterances for each speaker. The experimental results show that the instantaneous and transitional representations are relatively uncorrelated thus providing complementary information for speaker recognition. A rectangular window of approximately 100-150 ms duration provides an effective estimate of spectral transitions for speaker recognition. Also, simple transmission channel variations are shown to affect the instantaneous spectral representations and the corresponding recognition performance significantly, while the transitional representations and performance are relatively resistant. Frank K. Soong, Aaron E. Rosenberg |
ICASSP | 1 |
| 1985 | Comparative study of several distortion measures for speech recognitionabstractIn this study we compared several different spectral distortion measures including the Itakura-Saito (IS), the log likelihood ratio (LLR), the likelihood ratio (LR), the cepstral (CEP), and two perceptually based distortion measures, the weighted likelihood ratio (WLR) and the weighted slope metric (WSM) distortion measures, in terms of their effects on the performance of a standard dynamic time warping (DTW) based, isolated word, speech recognizer. Two modifications of the basic forms of each measure were also investigated, namely a Bark-scale frequency warping and the incorporation of suprasegmental energy information. All distortion measures and their modifications were tested on an alpha-digit vocabulary, 4-talker, telephone recording data base. The results can be summarized as: (1) All LPC-based distortion measures performed reasonably well. The LLR and WSM distortion measures gave the highest recognition accuracy, while the IS distortion measure gave the lowest score; (2) Whereas the addition of suprasegmental energy information helped the recognition performance, the use of gain and absolute loudness degraded the performance; (3) Bark-scale frequency warping did not perform as well as its unwarped counterpart; (4) The WLR distortion measure did not perform as well as its unweighted counterpart. N. Nocerino, Frank K. Soong, Lawrence R. Rabiner, Dennis H. Klatt |
ICASSP | 2 |
| 1985 | An efficient vector-quantization preprocessor for speaker independent isolated word recognitionabstractRecently a new structure for isolated word recognition was proposed based on the ideas of vector quantization (VQ). In this scheme a separate VQ codebook, for each word in the vocabulary, was designed, based on a training sequence of several tokens of each word by one or more talkers. In the original implementation, the recognizer chose the word in the vocabulary whose average quantization distortion (according to its particular codebook) was minimum. In the proposed implementation, the word-based VQ's are used as a front end preprocessor to eliminate word candidates whose distortion scores are large; a DTW processor then resolves the choice among the remaining word candidates (i.e. those which are passed on by the preprocessor). Both of the above schemes work very well for small vocabularies; however the major flaw is the lack of temporal information in the word-based VQ processor. As such, as the vocabulary for recognition grows in size and complexity, the ability of the VQ processor to resolve among similar sounding words decreases dramatically, and the effectiveness of the proposed recognition structure similarly decreases. To alleviate this difficulty a technique for incorporating temporal structure into the preprocessor is also proposed. In particular, the probability density function of the time of occurrence for each vector in the codebook is estimated from the same training sequence used to derive the codebook vectors. In the recognizer, the spectral distance score of the VQ is combined with a (scaled) temporal distance score, for each frame in the word. An evaluation of the proposed recognizer showed good performance on both the digits vocabulary, and on a vocabulary of 129 airlines terms. Kuk-Chin Pan, Frank K. Soong, Lawrence R. Rabiner, A. F. Bergh |
ICASSP | 2 |
| 1985 | Subband coding of speech using backward adaptive prediction and bit allocationabstractIn this study it is our goal to improve the performance of ADPCM and subband speech coders at medium bit rates (9.6∼16 kb/s) without increasing the coder complexity substantially. Various major building blocks including the predictors (both fixed and adaptive), subband quadrature mirror filter (QMF) and bit assignment strategy (both static and dynamic) are investigated in detail. We have found that (1) the Least-Squares (LS) adaptive lattice predictor outperforms both the pole-zero adaptive predictor recommended by CCITT and a first-order fixed predictor. (2) more subbands can improve the coder performance (3) longer QMF can reduce the interband aliasing and improve the subjective performance of a subband coder (4) an optimal dynamic bit allocation scheme with an improvement of SNR as high as 5 dB is much more favorable than a fixed bit allocation. With all of the above finding we propose a 4-band hybrid subband coder with an LS adaptive lattice predictor and an optimal dynamic bit allocation strategy. Frank K. Soong, Richard V. Cox, Nikil Jayant |
ICASSP | 1 |
| 1985 | A vector quantization approach to speaker recognitionabstractIn this study a vector quantization (VQ) codebook was used as an efficient means of characterizing the short-time spectral features of a speaker. A set of such codebooks were then used to recognize the identity of an unknown speaker from his/her unlabelled spoken utterances based on a minimum distance (distortion) classification rule. A series of speaker recognition experiments was performed using a 100-talker (50 male and 50 female) telephone recording database consisting of isolated digit utterances. For ten random but different isolated digits, over 98% speaker identification accuracy was achieved. The effects, on performance, of different system parameters such as codebook sizes, the number of test digits, phonetic richness of the text, and difference in recording sessions were also studied in detail. Frank K. Soong, Aaron E. Rosenberg, Lawrence R. Rabiner, Biing-Hwang Juang |
ICASSP | 1 |
| 1985 | Comparative study of several distortion measures for speech recognition
N. Nocerino, Frank K. Soong, Lawrence R. Rabiner, Dennis H. Klatt |
Speech Commun. | 2 |
| 1984 | On the use of transient information in speech recognitionabstractIn this paper we investigate the effects of signal processing on the performance of isolated-word recognition by changing various time-resolution related parameters. The vocabulary used,{"P", "B", "T", "D", "V", "Z"}, is a highly confusable subset of the 39-word alpha-digit database. We showed that the recognition performance is significantly improved by trace segmentation which compresses the steady-state parts of speech signals and refines the endpoints. By changing the cutoff frequency of the low-pass filter in the filterbank analysis, we observed the existence of an optimal region of cutoff frequencies ranging from 50 to 100 Hz (at -6 dB). Outside this region, the performance does not deteriorate completely even at a very low cutoff frequency where the transients are severely distorted. This phenomenon was explained by the fact of spectral modification of the steady-state vowels following the initial transients. Jean-Sylvain Liénard, Frank K. Soong |
ICASSP | 2 |
| 1984 | Line spectrum pair (LSP) and speech data compressionabstractLine Spectrum Pair (LSP) was first introduced by Itakura [1,2] as an alternative LPC spectral representations. It was found that this new representation has such interesting properties as (1) all zeros of LSP polynomials are on the unit circle, (2) the corresponding zeros of the symmetric and anti-symmetric LSP polynomials are interlaced, and (3) the reconstructed LPC all-pole filter preserves its minimum phase property if (1) and (2) are kept intact through a quantization procedure. In this paper we prove all these properties via a "phase function." The statistical characteristics of LSP frequencies are investigated by analyzing a speech data base. In addition, we derive an expression for spectral sensitivity with respect to single LSP frequency deviation such that some insight on their quantization effects can be obtained. Results on multi-pulse LPC using LSP for spectral information compression are finally presented. Frank K. Soong, Biing-Hwang Juang |
ICASSP | 1 |
| 1982 | On the high resolution and unbiased frequency estimates of sinusoids in white noise-A new adaptive approachabstractA new adaptive algorithm is proposed to give an unbiased and high resolution frequency estimate of sinusoids in white noise. Base on the normalized Least-Squares (LS) lattice algorithm and the inverse power iteration method, the eigenvector associated with the minimum eigenvalue of the signal covariance matrix is estimated in the algorithm. The zeros of the "eigenvector polynomial" thus obtained are all on the unit circle and at the angles of the sinusoid frequencies. It is an adaptive realization of the Pisarneko's "harmonic retrieval" method. But differing from the adaptive method proposed by Thompson where the eigenvector is obtained through a gradient search in a constrained optimization formulation, in the new method the eigenvector is computed by an inverse power iteration. It enjoys all the advantages of the normalized LS lattice such as fast computations, low round-off noise and an easy stability check, etc. as well as fast convergence rate of the inverse power iteration method. Computer simulation results are shown. Frank K. Soong, Allen M. Peterson |
ICASSP | 1 |
| 1982 | Fast least-squares (LS) in the voice echo cancellation applicationabstractThe existing echo cancellation methods are primarily based on the LMS adaptive algorithm. Despite the fact that the LMS echo canceller works better than its predecessor-the echo suppressor, its performance can be substantially improved if the Recursive LS (RLS) algorithm is used instead. However the αp2operations (p: filter order) per sample required prevents the RLS algorithm from being used in this and many other applications where the filter order is relatively high. The computational complexity of the RLS has recently been brought down to αp by exploiting the shifting structure of the signal covariance matrix. Two fast algorithms, namely the LS lattice and the "fast Kalman", are used here. Comparisons between the two fast LS algorithms and the LMS gradient algorithm are made and the performance difference is demonstrated. Two important problems in voice echo cancellation: the flat delay estimation and the near-end speech detection, are approached novelly through a minimum-mean-squared-error flat delay estimator and a likelihood near-end speech detector. Simulation results are very satsifactory. Frank K. Soong, Allen M. Peterson |
ICASSP | 1 |
| 1981 | On the asymptotic behavior of a complex adaptive line enchancer (CALE)abstractThe Adaptive Line Enhancer (ALE) was first introduced by Widrow and has since been used in applications such as noise cancellation, coherent interference rejection, detection and estimation of sinusoid(s) in white noise, etc. In this paper, a Complex Adaptive Line Enhancer (CALE) is adopted to process complex-valued signals in their analytic form. The performance and asymptotic behavior of the CALE in cases of single and multiple sinusoid(s) in white noise are investigated in detail. It is shown that the frequency estimation obtained by CALE in the case of a single sinusoid is unbiased and independent of the bulk delay, Δ, the predictor length, L, and the Signal-to-Noise Ratio (SNR). If a conventional real ALE is used, this result, in general, cannot be obtained. A precise analysis is made for the case of two sinusoids. It is shown that an optimal choice of Δ, exists. For more than two sinusoids, an approximation is used to derive more insight into the analytical structure of a CALE. An intuitive explanation for the optimal choice of Δ, is given. Simulation results are also shown. Frank K. Soong, S. Shankar Narayan, Allen M. Peterson |
ICASSP | 1 |
| 1980 | Fast spectral estimation of speech signal in analytic formabstractThe analytic form of speech signals has been generated by combining the real-valued speech with its imaginary-valued discrete-Hilbert-transformed counter-part. Instead of using an 8th order real-valued linear prediction to analyze its spectral components, a 4th order complex-valued linear prediction is used. Furthermore, a fast and closed-form rooting procedure is adopted to extract the roots (i.e., bandwidth and frequency information of the speech formants) of the obtained 4th order polynomial of the corresponding inverse filter. The numerical complexity and convergence problems are thus avoided and the computational load is greatly reduced. Promising results in both spectral and pitch estimation are shown. Its advantages in efficient speech coding are discussed. Frank K. Soong, Allen M. Peterson |
ICASSP | 1 |
| 1978 | Observations on linear estimationabstractHeisey and Griffiths have proposed a generalization of linear prediction, called "linear estimation", in which both past and future data samples are used to predict (estimate) the present sample. They report that although the mean-square error from this formulation is usually smaller than from standard linear prediction, the corresponding spectral estimate is a poorer fit to the true spectrum. We give a general explanation for this apparent paradox in terms of the zeros of the estimated inverse filter and examine specifically the case of frequency estimation for a single complex sinusoid in noise. The intuitively appealing idea that future as well as past data should be included in the estimates is best implemented by a combined forward-backward prediction method. Leland B. Jackson, Frank K. Soong |
ICASSP | 2 |
| 1978 | Frequency estimation by linear predictionabstractThe application of linear prediction to frequency estimation for sinusoidal signals in noise is investigated. It is shown that improved performance is obtained by processing a complex-valued version of the real-valued input signal, with the corresponsing sampling rate reduced by one-half. The case of a single sinusoid in white noise is studied in detail, including the eigenvalues of the covariance matrix, zeros of the inverse filter polynomial, frequency bias, and frequency variance as a function of input SNR and prediction order. Leland B. Jackson, Donald W. Tufts, Frank K. Soong, Rahul M. Rao |
ICASSP | 3 |