Yu Shi 0001

dblp:55/4736-1 · DBLP profile ↗
← Back
30ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0003-1872-3429ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2025 Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
abstract
End-to-End Automatic Speech Recognition (ASR) has advanced significantly yet still struggles with rare and domain-specific entities. This paper introduces a simple yet efficient prompt-based biasing technique for contextualized ASR, enhancing recognition accuracy by leverage a unified multi-task learning framework. The approach comprises two key components: a prompt biasing model which is trained to determine when to focus on entities in prompt, and a entity filtering mechanism which efficiently filters out irrelevant entities. Our method significantly enhances ASR accuracy on entities, achieving a relative $30.7 \%$ and $18.0 \%$ reduction in Entity Word Error Rate compared to the baseline model with shallow fusion on in-house domain dataset with small and large entity lists, respectively. The primary advantage of this method lies in its efficiency and simplicity without any structure change, making it lightweight and highly efficient.
Yu Shi 0001, Jinyu Li 0001
ASRU2
2023 i-Code: An Integrative and Composable Multimodal Learning Framework
abstract
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.
Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001
AAAI6
2023 Z-Code++: A Pre-trained Language Model Optimized for Abstractive Summarization
abstract
Pengcheng He, Baolin Peng, Song Wang, Yang Liu, Ruochen Xu, Hany Hassan, Yu Shi, Chenguang Zhu, Wayne Xiong, Michael Zeng, Jianfeng Gao, Xuedong Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Baolin Peng, Song Wang 0012, Yang Liu 0124, Ruochen Xu, Hany Hassan, Yu Shi 0001, Chenguang Zhu 0001, Wayne Xiong, Michael Zeng 0001, Jianfeng Gao 0001, Xuedong Huang 0001
ACL (1)7
2023 Code-Switching Text Generation and Injection in Mandarin-English ASR
abstract
Code-switching speech refers to a means of expression by mixing two or more languages within a single utterance. Automatic Speech Recognition (ASR) with End-to-End (E2E) modeling for such speech can be a challenging task due to the lack of data. In this study, we investigate text generation and injection for improving the performance of an industry commonly-used streaming model, Transformer-Transducer (T-T), in Mandarin-English code-switching speech recognition. We first propose a strategy to generate codeswitching text data and then investigate injecting generated text into T-T model explicitly by Text-To-Speech (TTS) conversion or implicitly by tying speech and text latent spaces. Experimental results on the T-T model trained with a dataset containing 1,800 hours of real Mandarin-English code-switched speech show that our approaches to inject generated code-switching text significantly boost the performance of T-T models, i.e., 16% relative Token-based Error Rate (TER) reduction averaged on three evaluation sets, and the approach of tying speech and text latent spaces is superior to that of TTS conversion on the evaluation set which contains more homogeneous data with the training set.
Yuxuan Hu 0003, Yao Qian, Ma Jin, Linquan Liu, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Edward Lin, Michael Zeng 0001
ICASSP7
2023 Improving Readability for Automatic Speech Recognition Transcription
abstract
Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to grammatical errors, disfluency, and other noises common in spoken communication. These readable issues introduced by speakers and ASR systems will impair the performance of downstream tasks and the understanding of human readers. In this work, we present a task called ASR post-processing for readability (APR) and formulate it as a sequence-to-sequence text generation problem. The APR task aims to transform the noisy ASR output into a readable text for humans and downstream tasks while maintaining the semantic meaning of speakers. We further study the APR task from the benchmark dataset, evaluation metrics, and baseline models: First, to address the lack of task-specific data, we propose a method to construct a dataset for the APR task by using the data collected for grammatical error correction. Second, we utilize metrics adapted or borrowed from similar tasks to evaluate model performance on the APR task. Lastly, we use several typical or adapted pre-trained models as the baseline models for the APR task. Furthermore, we fine-tune the baseline models on the constructed dataset and compare their performance with a traditional pipeline method in terms of proposed evaluation metrics. Experimental results show that all the fine-tuned baseline models perform better than the traditional pipeline method, and our adapted RoBERTa model outperforms the pipeline method by 4.95 and 6.63 BLEU points on two test sets, respectively. The human evaluation and case study further reveal the ability of the proposed model to improve the readability of ASR transcripts.
Junwei Liao, Sefik Emre Eskimez, Liyang Lu, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Hong Qu 0002, Michael Zeng 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2022 Optimizing Alignment of Speech and Language Latent Spaces for End-To-End Speech Recognition and Understanding
abstract
The advances in attention-based encoder-decoder (AED) networks have brought great progress to end-to-end (E2E) automatic speech recognition (ASR). One way to further improve the performance of AED-based E2E ASR is to introduce an extra text encoder for leveraging extensive text data and thus capture more context-aware linguistic information. However, this approach brings a mismatch problem between the speech encoder and the text encoder due to the different units used for modeling. In this paper, we propose an embedding aligner and modality switch training to better align the speech and text latent spaces. The embedding aligner is a shared linear projection between text encoder and speech encoder trained by masked language modeling (MLM) loss and connectionist temporal classification (CTC), respectively. The modality switch training randomly swaps speech and text embeddings based on the forced alignment result to learn a joint representation space. Experimental results show that our proposed approach achieves a relative 14% to 19% word error rate (WER) reduction on Librispeech ASR task. We further verify its effectiveness on spoken language understanding (SLU), i.e., an absolute 2.5% to 2.8% F1 score improvement on SNIPS slot filling task.
Wei Wang 0010, Shuo Ren 0002, Yao Qian, Shujie Liu 0001, Yu Shi 0001, Yanmin Qian, Michael Zeng 0001
ICASSP5
2021 Generating Human Readable Transcript for Automatic Speech Recognition with Pre-Trained Language Model
abstract
Modern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to disfluency, filter words, and other errata common in spoken communication. Many downstream tasks and human readers rely on the output of the ASR system; therefore, errors introduced by the speaker and ASR system alike will be propagated to the next task in the pipeline. In this work, we propose an ASR post-processing model that aims to transform the incorrect and noisy ASR output into a readable text for humans and downstream tasks. We leverage the Metadata Extraction (MDE) corpus to construct a task-specific dataset for our study. Since the dataset is small, we propose a novel data augmentation method and use a two-stage training strategy to fine-tune the RoBERTa pre-trained model. On the constructed test set, our model outperforms a production two-step pipeline-based post-processing method by a large margin of 13.26 on readability-aware WER (RA-WER) and 17.53 on BLEU metrics. Human evaluation also demonstrates that our method can generate more human-readable transcripts than the baseline method.
Junwei Liao, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Sefik Emre Eskimez, Liyang Lu, Hong Qu 0002, Michael Zeng 0001
ICASSP2
2021 Speech-Language Pre-Training for End-to-End Spoken Language Understanding
abstract
End-to-end (E2E) spoken language understanding (SLU) can infer semantics directly from speech signal without cascading an automatic speech recognizer (ASR) with a natural language understanding (NLU) module. However, paired utterance recordings and corresponding semantics may not always be available or sufficient to train an E2E SLU model in a real production environment. In this paper, we propose to unify a well-optimized E2E ASR encoder (speech) and a pre-trained language model encoder (language) into a transformer decoder. The unified speech-language pre-trained model (SLP) is continually enhanced on limited labeled data from a target domain by using a conditional masked language model (MLM) objective, and thus can effectively generate a sequence of intent, slot type, and slot value for given input speech in the inference. The experimental results on two public corpora show that our approach to E2E SLU is superior to the conventional cascaded method. It also outperforms the present state-of-the-art approaches to E2E SLU with much less paired data.
Yao Qian, Ximo Bian, Yu Shi 0001, Naoyuki Kanda, Leo Shen, Michael Zeng 0001
ICASSP3
2021 Improving Zero-shot Neural Machine Translation on Language-specific Encoders- Decoders
abstract
Recently, universal neural machine translation (NMT) with shared encoder-decoder gained good performance on zero-shot translation. Unlike universal NMT, jointly trained language-specific encoders-decoders aim to achieve universal representation across non-shared modules, each of which is for a language or language family. The non-shared architecture has the advantage of mitigating internal language competition, especially when the shared vocabulary and model parameters are restricted in their size. However, the performance of using multiple encoders and decoders on zero-shot translation still lags behind universal NMT. In this work, we study zero-shot translation using language-specific encoders-decoders. We propose to generalize the non-shared architecture and universal NMT by differentiating the Transformer layers between language-specific and interlingua. By selectively sharing parameters and applying cross-attentions, we explore maximizing the representation universality and realizing the best alignment of language-agnostic information. We also introduce a denoising auto-encoding (DAE) objective to jointly train the model with the translation task in a multi-task manner. Experiments on two public multilingual parallel datasets show that our proposed model achieves competitive or better results than universal NMT and the strong pivot baseline. Moreover, we experiment incrementally adding new language to the trained model by only updating the new model parameters. With this little effort, the zero-shot translation between this newly added language and existing languages achieves a comparable result with the model trained jointly from scratch on all languages.
Junwei Liao, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Hong Qu 0002, Michael Zeng 0001
IJCNN2
2021 Listen, Look and Deliberate: Visual Context-Aware Speech Recognition Using Pre-Trained Text-Video Representations
abstract
In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore novel VC-ASR approaches to leverage video and text representations extracted by a self-supervised pre-trained text-video embedding model. Firstly, we propose a multi-stream attention architecture to leverage signals from both audio and video modalities. This architecture consists of separate encoders for the two modalities and a single decoder that at-tends over them. We show that this architecture is better than fusing modalities at the signal level. Additionally, we also explore lever-aging the visual information in a second pass model, which has also been referred to as a `deliberation model'. The deliberation model accepts audio representations and text hypotheses from the first pass ASR and combines them with a visual stream for an improved visual context-aware recognition. The proposed deliberation scheme can work on top of any well trained ASR and also enabled us to leverage the pre-trained text model to ground the hypotheses with the visual features. Our experiments on HOW2 dataset show that multi-stream and deliberation architectures are very effective at the VC-ASR task. We evaluate the proposed models for two scenarios; clean audio stream and distorted audio in which we mask out some specific words in the audio. The deliberation model outperforms the multi-stream model and achieves a relative WER improvement of 6% and 8.7% for the clean and masked data, respectively, compared to an audio-only model. The deliberation model also improves re-covering the masked words by 59% relative.
Shahram Ghorbani, Yashesh Gaur, Yu Shi 0001, Jinyu Li 0001
SLT3
2020 Discriminative Transfer Learning for Optimizing ASR and Semantic Labeling in Task-Oriented Spoken Dialog
Yao Qian, Yu Shi 0001, Michael Zeng 0001
INTERSPEECH2
2010 A Study of Discriminative Training for HMM-Based Online Handwritten Chinese/Japanese Character Recognition
abstract
We present a study of discriminative training of classifiers using both maximum mutual information (MMI) and minimum classification error (MCE) criteria for online handwritten Chinese/Japanese character recognition based on continuous-density hidden Markov models. It is observed that MCE-trained classifiers can achieve a much higher recognition accuracy than that of MMI-trained ones. Benchmark results of MCE-trained classifiers for simplified Chinese, traditional Chinese and Japanese characters are reported on three recognition tasks with a vocabulary of 9119, 20924, and 12333 characters respectively.
Yongqiang Wang 0008, Qiang Huo, Yu Shi 0001
ICFHR3
2010 A study of irrelevant variability normalization based training and unsupervised online adaptation for LVCSR
abstract
This paper presents an experimental study of a maximum likelihood (ML) approach to irrelevant variability normalization (IVN) based training and unsupervised online adaptation for large vocabulary continuous speech recognition. A moving window based frame labeling method is used for acoustic sniffing. The IVN-based approach achieves a 10% relative word error rate reduction over an ML-trained baseline system on a Switchboard-1 conversational telephone speech transcription task.
Guangchuan Shi, Yu Shi 0001, Qiang Huo
INTERSPEECH2
2009 A Study of Feature Design for Online Handwritten Chinese Character Recognition Based on Continuous-Density Hidden Markov Models
abstract
We present a new feature extraction approach to online Chinese handwriting recognition based on continuous-density hidden Markov models (CDHMM). Given an online handwriting sample, a sequence of time-ordered dominant points are extracted first, which include stroke-endings, points corresponding to local extrema of curvature, and points with a large distance to the chords formed by pairs of previously identified neighboring dominant points. Then, at each dominant point, a 6-dimensional feature vector is extracted, which consists of two coordinate features, two delta features, and two double-delta features. Its effectiveness has been confirmed by experiments for a recognition task with a vocabulary of 9119 Chinese characters and CDHMMs trained from about 10 million samples using both maximum likelihood and discriminative training criteria.
Qiang Huo, Yu Shi 0001
ICDAR3
2008 Symbol graph based discriminative training and rescoring for improved math symbol recognition
abstract
In the symbol recognition stage of online handwritten math expression recognition, the one-pass dynamic programming algorithm can produce high-quality symbol graphs in addition of the best recognized hypotheses [1]. In this paper, we exploit the rich hypotheses embedded in a symbol graph to discriminatively train the exponential weights of different model likelihoods and the insertion penalty. The training is investigated in two different criteria: Maximum Mutual Information (MMI) and Minimum Symbol Error (MSE). After discriminative training, trigram-based graph rescoring is performed in a post-processing stage. Experimental results finally show a 97% symbol accuracy on a test set of 2,574 written expressions with 43,300 symbols, a signi..cant improvement of symbol accuracy obtained.
Zhen Xuan Luo, Yu Shi 0001, Frank K. Soong
ICASSP2
2008 Approximateword-lattice indexing with text indexers: Time-Anchored Lattice Expansion
abstract
We address the problem of how to represent or approximate speech lattices to be indexed with existing text indexers. We present a method named Time-Anchored Lattice Expansion (TALE), which can be implemented by a Standard Text Indexer (STI). On a 170-hour lecture set, we compare TALE with other lattice indexing methods: confusion networks, Position-Specific Posterior Lattices (PSPL), and Time-based Merging for Index (TMI). All methods achieve accuracies comparable to searching raw lattices when the corresponding index structures and phrase matching algorithms are used. However, when implemented with an STI, TALE significantly outperforms all other methods. Compared to indexing linear text, TALE improves accuracy by 30–60% for multi-word phrase searches and by 130% for two-term AND queries.
Yu Shi 0001, Frank Seide
ICASSP2
2008 A symbol graph based handwritten math expression recognition
abstract
In online handwritten math expression recognition, one-pass dynamic programming can produce high-quality symbol graphs in addition to best symbol sequence hypotheses, especially after discriminative training and trigram graph rescoring. Impact of symbol graphs on whole expression recognition, however, has not been referred to yet, since the interface of structure analysis module does not work well with symbol graphs on the basis of typical tree search. In this paper, we propose a method to convert symbol graph to segment graph to make the tree search efficient and effective, i.e., search of best segmentations in symbol graph without pruning becomes possible. With trigram rescoring, the overall expression recognition accuracy has been improved by 10% relative in comparison with the baseline.
Yu Shi 0001, Frank K. Soong
ICPR1
2008 GPU-accelerated Gaussian clustering for fMPE discriminative training
abstract
The Graphics Processing Unit (GPU) has extended its applications from its original graphic rendering to more general scientific computation. Through massive parallelization, state-ofthe-art GPUs can deliver 200 billion floating-point operations per second (0.2 TFLOPS) on a single consumer-priced graphics card. This paper describes our attempt in leveraging GPUs for efficient HMM model training. We show that using GPUs for a specific example of Gaussian clustering, as required in fMPE, or feature-domain Minimum Phone Error discriminative training, can be highly desirable. The clustering of huge number of Gaussians is very time consuming due to the enormous model size in current LVCSR systems. Comparing an NVidia Geforce 8800 Ultra GPU against an Intel Pentium 4 implementation, we find that our brute-force GPU implementation is 14 times faster overall than a CPU implementation that uses approximate speed-up heuristics. GPU accelerated fMPE reduces the WER 6% relatively, compared to the maximumlikelihood trained baseline on two conversational-speech recognition tasks.
Yu Shi 0001, Frank Seide, Frank K. Soong
INTERSPEECH1
2007 Towards spoken-document retrieval for the enterprise: Approximate word-lattice indexing with text indexers
abstract
Enterprise-scale search engines are generally designed for linear text. Linear text is suboptimal for audio search, where accuracy can be significantly improved if the search includes alternate recognition candidates, commonly represented as word lattices. We propose two methods to enable text indexers to approximately index lattices with little or no code change: “TMI” (Time-based Merging for Indexing) aims at lattice-index size reduction, and the “sausage”-like “TALE” (Time-Anchored Lattice Expansion) approximation requires no indexer-code or data-format changes at all. On four enterprise-type data sets (meetings, phone calls, lectures, and voicemail), TMI and TALE improve accuracy by 30–60% for multi-word phrase searches and by 130% for two-term AND queries, compared to indexing linear text.
Frank Seide, Yu Shi 0001
ASRU3
2007 A Segmentation Posterior Based Endpointing Algorithm
abstract
A segmentation posterior probability based endpointing algorithm for robust ASR is proposed. First, each speech signal is partitioned into homogeneous segments via auto-segmentation. Then posterior probabilities of all possible endpoints are computed, based on the segmentation likelihoods of all levels in a selected range. Endpoints with the highest posterior probabilities are finally selected. The new method differs from the previous auto-segmentation and clustering based algorithm on that the former considers hypotheses from several levels, while the latter depends only on one appropriate level. Another potential benefit of the proposed method is that any endpointing or VAD results can be integrated, as hypotheses, into the posterior probability framework. Experiments based on the AURORA2 digit database show the robustness of the proposed method.
Yanlu Xie, Yu Shi 0001, Frank K. Soong, Beiqian Dai
ICASSP (4)2
2007 A Unified Framework for Symbol Segmentation and Recognition of Handwritten Mathematical Expressions
abstract
A symbol decoding and graph generation algorithm for online handwritten mathematical expression recognition is formulated. It differs from our previous system and most other systems in two aspects: (1) it embeds stroke grouping into symbol identification to form a unified probabilistic framework for symbol recognition; and (2) a symbol graph rather than a list of symbol sequence hypotheses is generated, which makes post-processing with new information possible. Experimental results show that high quality symbol graph can be generated by the proposed algorithm. Symbol sequence corresponding to the best path in the graph demonstrates much higher symbol recognition accuracy than before, especially after rescoring with trigram. Math formula recognition performance is significantly improved.
Yu Shi 0001, Frank K. Soong
ICDAR1
2006 Auto-Segmentation Based Partitioning and Clustering Approach to Robust Endpointing
abstract
An auto segmentation based partitioning and clustering approach to robust Voice Activity Detection (VAD) is proposed. It is done in two successive steps: homogeneous frame partitioning and segment clustering. The first step, due to its auto segmentation nature, does not need a noise model, and is applicable to different noise types and SNR's. The algorithm is a dynamic programming based procedure and provides a graceful performance in finding segmentation thresholds. Multiple parameters like energy, pitch and voicing information can be easily incorporated into the procedure. The algorithm is evaluated on the test sets in the Aurora2 database. The algorithm shows its robustness at low SNR operating environments; the endpoint estimate errors are shown to have small variance.
Yu Shi 0001, Frank K. Soong, Jian-Lai Zhou
ICASSP (1)1
2006 Auto-segmentation based VAD for robust ASR
abstract
An auto-segmentation based endpointing algorithm for robust ASR is proposed. The algorithm consists of two successive steps: (1) homogeneous segment partitioning and (2) segment clustering. The first step, due to its self-segmentation nature, does not need a noise model, and is applicable to different noises at various SNR’s. The dynamic programming based segment partitioning, which can generate more homogeneous segments than individual frames for clustering, yields a more robust VAD mechanism. Experiments are performed on the AURORA2 digit database by comparing the new algorithm with the ETSI standard for DSR. Quantitative assessment of the new algorithm is performed via different evaluation criteria, including: ROC curves, speech/non-speech discrimination, and speech recognition performance.
Yu Shi 0001, Frank K. Soong, Jian-Lai Zhou
INTERSPEECH1
2004 Segmental tonal modeling for phone set design in Mandarin LVCSR
abstract
Modeling units play a very important role in state-of-art speech recognition systems. The design and selection of them will directly impact the performance of the final speech recognition engine. As a tonal language, Mandarin's modeling units are more special for the tonal processing. In this paper, after fully investigating several dominant modeling strategies, we propose a new phone set design strategy for Mandarin, called segmental tonal modeling. Instead of modeling tone types directly, we realize them implicitly and jointly by two segments, which both carry tonal information. Both HTK and SAPI based experiments confirmed that such a method is very efficient. In addition to improving the accuracy by 9-23%, it greatly reduces the decoding time by 30-45%. Given the similar decoding speed, the new phone set configuration can reduce the error rate by relatively 35%.
Chao Huang 0011, Yu Shi 0001, Jian-Lai Zhou, Min Chu, Terry Wang, Eric Chang
ICASSP (1)2
2004 Studies in massively speaker-specific speech recognition
abstract
Over the past several years, the primary focus for the speech-recognition research community has been speaker-independent speech recognition, with the emphasis of working on databases with larger and larger numbers of speakers. For example, the most recent EARS program, which is sponsored by DARPA, calls for recordings of thousands of speakers. However, we are interested in making a speech interface work well for one particular individual, and we propose using massive amounts of speaker-specific training data recorded in daily life. We call this massively speaker-specific recognition (MSSR). As a pre-research, we leverage the large corpus we have available from speech-synthesis work to study the benefit of MSSR only from the acoustic-modeling aspect. Initial results show that, by changing the focus to MSSR, word error rates can drop very significantly. In comparison with speaker-adaptive speech recognition systems, MSSR also performs better since model parameters can be tuned to be suitable to one particular individual.
Yu Shi 0001, Eric Chang
ICASSP (1)1
2004 Tone articulation modeling for Mandarin spontaneous speech recognition
abstract
Tone modeling is an unavoidable problem in Mandarin speech recognition. In continuous speech, the pitch contour exhibits variable patterns, and it is strongly influenced by its tone context. Although several effective methods have been proposed to improve the accuracy for tonal syllables in Mandarin continuous speech recognition, many recognition errors are caused by poor tone discrimination capability of the acoustic model. Furthermore, the case becomes worse for the recognition of spontaneous speech. In this paper, we report our work on tone articulation modeling. Tone context dependent models are used to model unstable pitch patterns caused by co-articulation in continuous speech. Corresponding acoustic features are investigated as well. Our methods are evaluated on two test sets: one is reading-style speech data, the other is spontaneous. The experimental results show that for the test set of casual speech, the proposed method turns out to be more effective than tone context independent model, while they are comparable for the test set of reading-style speech. Several factors which have potential to improve the proposed method are discussed in the final part in this paper.
Jian-Lai Zhou, Yu Shi 0001, Chao Huang 0011, Eric Chang
ICASSP (1)3
2003 Spectrogram-based formant tracking via particle filters
abstract
The paper presents a particle-filtering method for estimating formant frequencies of speech signals from spectrograms. First, frequency bands corresponding to the analyzed formants are extracted via a two-step dynamic programming based algorithm. A particle-filtering method is then used to locate accurately formants in every formant area based on the posterior PDF described by a set of support points with associated weights. Formant trajectories of voiced frames of a group of 81 utterances were manually tracked and labeled, partly for model training and partly for algorithm evaluation. In the experiments, the proposed method obtains average estimation errors of 72, 115, and 113 Hz for the first three formants, respectively, whereas the LPC based method induces 118, 172, and 250 Hz deviations. The experimental results show that the formants estimated by the proposed method are quite reliable and the trajectories are more accurate than LPC.
Yu Shi 0001, Eric Chang
ICASSP (1)1
2002 Power spectral density based channel equalization of large speech database for concatenative TTS system
abstract
This paper proposes a channel equalization algorithm for a large speech database with application in concatenative TTS systems. The convolutional channel distortion is equalized by comparing the power spectral densities (PSDs) of utterances of different recording sessions. Autoregressive linear filters are designed on a corpus level and are used offline to filter the corresponding sentences to compensate for the relative distortions caused by the channel effects. Two experiments are carried out to evaluate the benefit of the channel equalization approach. First, this method is used to reduce the distance of their PSDs between two recording sessions to verify the effectiveness of the method. Secondly, it is applied practically in the TTS system. The whole TTS speech database is processed to reduce the PSDs variance over all sessions. Moreover, a subjective listening test is carried out to obtain human evaluation of the new TTS system. Almost all listeners prefer the synthetic speech generated by the new TTS system. Furthermore, an analysis of variance (ANOVA) on this subjective listening test demonstrates that the channel equalization process has significant effect on increasing the perceived voice-quality consistency of the TTS system.
Yu Shi 0001, Eric Chang, Hu Peng, Min Chu
INTERSPEECH1
2002 A system for spoken query information retrieval on mobile devices
abstract
With the proliferation of handheld devices, information access on mobile devices is a topic of growing relevance. This paper presents a system that allows the user to search for information on mobile devices using spoken natural-language queries. We explore several issues related to the creation of this system, which combines state-of-the-art speech-recognition and information-retrieval technologies. This is the first work that we are aware of which evaluates spoken query based information retrieval on a commonly available and well researched text database, the Chinese news corpus used in the National Institute of Standards and Technology (NIST)s TREC-5 and TREC-6 benchmarks. To compare spoken-query retrieval performance for different relevant scenarios and recognition accuracies, the benchmark queries-read verbatim by 20 speakers-were recorded simultaneously through three channels: headset microphone, PDA microphone, and cellular phone. Our results show that for mobile devices with high-quality microphones, spoken-query retrieval based on existing technologies yields retrieval precisions that come close to that for perfect text input (mean average precision 0.459 and 0.489, respectively, on TREC-6).
Eric Chang, Frank Seide, Helen M. Meng, Zhuoran Chen, Yu Shi 0001, Yuk-Chi Li
IEEE Trans. Speech Audio Process.5
2001 Speech lab in a box: a Mandarin speech toolbox to jumpstart speech related research
abstract
The necessity of gathering data has been an impediment for researchers and students who are interested in getting started in the fields related to speech recognition. We are proposing a new approach of distributing data that is designed to quickly help researchers and students achieve a set of baseline results to build upon. Furthermore, by leveraging publicly available programs, all researchers will be able to exactly reproduce results that are described in this paper. We also aim to facilitate comparison of recognition results in the field of Mandarin speech recognition by including a testing set in the toolbox. We describe a toolbox that includes Mandarin speech data from 125 speakers, suitable language model, scripts and data files required for recreating a set of baseline experiments, and a copy of Microsoft SAPI 5.0 SDK that can help professors and students who wish to jumpstart research programs in speech technologies. By lowering the barrier of entry to the field, we hope to encourage more participation in the study of Mandarin speech recognition.
Eric Chang, Yu Shi 0001, Jian-Lai Zhou, Chao Huang 0011
INTERSPEECH2