Lin-Shan Lee

dblp:40/176 · DBLP profile ↗
← Back
283ranked-venue papers
26as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 214 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 168 · 12 first-author · 3 since 2021Computer networks · 20 · 10 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 SpeechDPR: End-To-End Spoken Passage Retrieval For Open-Domain Spoken Question Answering
abstract
Spoken Question Answering (SQA) is essential for machines to reply to user’s question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domain SQA (openSQA), in which the machine needs to first retrieve passages that possibly contain the answer from a spoken archive in addition, was never considered. This paper proposes the first known end-to-end frame-work, Speech Dense Passage Retriever (SpeechDPR), for the retrieval component of the openSQA problem. SpeechDPR learns a sentence-level semantic representation by distilling knowledge from the cascading model of unsupervised ASR (UASR) and text dense retriever (TDR). No manually transcribed speech data is needed. Initial experiments showed performance comparable to the cascading model of UASR and TDR, and significantly better when UASR was poor, verifying this approach is more robust to speech recognition errors.
Chyi-Jiunn Lin, Guan-Ting Lin, Yung-Sung Chuang, Wei-Lun Wu, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Lin-Shan Lee
ICASSP8
2024 REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR
abstract
Unsupervised automatic speech recognition (ASR) aims to learn the mapping between the speech signal and its corresponding textual transcription without the supervision of paired speech-text data. A word/phoneme in the speech signal is represented by a segment of speech signal with variable length and unknown boundary, and this segmental structure makes learning the mapping between speech and text challenging, especially without paired data. In this paper, we propose REBORN, Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR. REBORN alternates between (1) training a segmentation model that predicts the boundaries of the segmental structures in speech signals and (2) training the phoneme prediction model, whose input is a segmental structure segmented by the segmentation model, to predict a phoneme transcription. Since supervised data for training the segmentation model is not available, we use reinforcement learning to train the segmentation model to favor segmentations that yield phoneme sequence predictions with a lower perplexity. We conduct extensive experiments and find that under the same setting, REBORN outperforms all prior unsupervised ASR models on LibriSpeech, TIMIT, and five non-English languages in Multilingual LibriSpeech. We comprehensively analyze why the boundaries learned by REBORN improve the unsupervised ASR performance.
Liang-Hsuan Tseng, En-Pei Hu, Cheng-Han Chiang, Yuan Tseng, Hung-yi Lee, Lin-Shan Lee, Shao-Hua Sun
NeurIPS6
2022 DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering
abstract
Spoken Question Answering (SQA) is to find the answer from a spoken document given a question, which is crucial for personal assistants when replying to the queries from the users.Existing SQA methods all rely on Automatic Speech Recognition (ASR) transcripts.Not only does ASR need to be trained with massive annotated data that are time and cost-prohibitive to collect for low-resourced languages, but more importantly, very often the answers to the questions include name entities or out-of-vocabulary words that cannot be recognized correctly.Also, ASR aims to minimize recognition errors equally over all words, including many function words irrelevant to the SQA task.Therefore, SQA without ASR transcripts (textless) is always highly desired, although known to be very difficult.This work proposes Discrete Spoken Unit Adaptive Learning (DUAL), leveraging unlabeled data for pre-training and finetuned by the SQA downstream task.The time intervals of spoken answers can be directly predicted from spoken documents.We also release a new SQA benchmark corpus, NMSQA, for data with more realistic scenarios.We empirically showed that DUAL yields results comparable to those obtained by cascading ASR and text QA model and robust to real-world data.Our code and model will be open-sourced 1 .
Guan-Ting Lin, Yung-Sung Chuang, Ho-Lam Chung, Shu-Wen Yang, Hsuan-Jui Chen, Shuyan Dong, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH10
2021 Fragmentvc: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with Attention
abstract
Any-to-any voice conversion aims to convert the voice from and to any speakers even unseen during training, which is much more challenging compared to one-to-one or many-to-many tasks, but much more attractive in real-world scenarios. In this paper we proposed FragmentVC, in which the latent phonetic structure of the utterance from the source speaker is obtained from Wav2Vec 2.0, while the spectral features of the utterance(s) from the target speaker are obtained from log mel-spectrograms. By aligning the hidden structures of the two different feature spaces with a two-stage training process, FragmentVC is able to extract fine-grained voice fragments from the target speaker utterance(s) and fuse them into the desired utterance, all based on the attention mechanism of Transformer as verified with analysis on attention maps, and is accomplished end-to-end. This approach is trained with reconstruction loss only without any disentanglement considerations between content and speaker information and doesn't require parallel data. Objective evaluation based on speaker verification and subjective evaluation with MOS both showed that this approach outperformed SOTA approaches, such as AdaIN-VC and AutoVC.
Yist Y. Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, Lin-Shan Lee
ICASSP5
2021 Towards Lifelong Learning of End-to-End ASR
abstract
Automatic speech recognition (ASR) technologies today are primarily optimized for given datasets; thus, any changes in the application environment (e.g., acoustic conditions or topic domains) may inevitably degrade the performance. We can collect new data describing the new environment and fine-tune the system, but this naturally leads to higher error rates for the earlier datasets, referred to as catastrophic forgetting. The concept of lifelong learning (LLL) aiming to enable a machine to sequentially learn new tasks from new datasets describing the changing real world without forgetting the previously learned knowledge is thus brought to attention. This paper reports, to our knowledge, the first effort to extensively consider and analyze the use of various approaches of LLL in end-to-end (E2E) ASR, including proposing novel methods in saving data for past domains to mitigate the catastrophic forgetting problem. An overall relative reduction of 28.7% in WER was achieved compared to the fine-tuning baseline when sequentially learning on three very different benchmark corpora. This can be the first step toward the highly desired ASR technologies capable of synchronizing with the continuously changing real world.
Heng-Jui Chang, Hung-yi Lee, Lin-Shan Lee
Interspeech3
2021 End-to-End Whispered Speech Recognition with Frequency-Weighted Approaches and Pseudo Whisper Pre-training
abstract
Whispering is an important mode of human speech, but no end-to-end recognition results for it were reported yet, probably due to the scarcity of available whispered speech data. In this paper, we present several approaches for end-to-end (E2E) recognition of whispered speech considering the special characteristics of whispered speech and the scarcity of data. This includes a frequency-weighted SpecAugment policy and a frequency-divided CNN feature extractor for better capturing the high-frequency structures of whispered speech, and a layer-wise transfer learning approach to pre-train a model with normal or normal-to-whispered converted speech then fine-tune it with whispered speech to bridge the gap between whispered and normal speech. We achieve an overall relative reduction of 19.8% in PER and 44.4% in CER on a relatively small whispered TIMIT corpus. The results indicate as long as we have a good E2E model pre-trained on normal or pseudo-whispered speech, a relatively small set of whispered speech may suffice to obtain a reasonably good E2E whispered speech recognizer.
Heng-Jui Chang, Alexander H. Liu, Hung-yi Lee, Lin-Shan Lee
SLT4
2021 Defending Your Voice: Adversarial Attack on Voice Conversion
abstract
Substantial improvements have been achieved in recent years in voice conversion, which converts the speaker characteristics of an utterance into those of another speaker without changing the linguistic content of the utterance. Nonetheless, the improved conversion technologies also led to concerns about privacy and authentication. It thus becomes highly desired to be able to prevent one's voice from being improperly utilized with such voice conversion technologies. This is why we report in this paper the first known attempt to perform adversarial attack on voice conversion. We introduce human imperceptible noise into the utterances of a speaker whose voice is to be defended. Given these adversarial examples, voice conversion models cannot convert other utterances so as to sound like being produced by the defended speaker. Preliminary experiments were conducted on two currently state-of-the-art zero-shot voice conversion models. Objective and subjective evaluation results in both white-box and black-box scenarios are reported. It was shown that the speaker characteristics of the converted utterances were made obviously different from those of the defended speaker, while the adversarial examples of the defended speaker are not distinguishable from the authentic utterances.
Chien-Yu Huang, Yist Y. Lin, Hung-yi Lee, Lin-Shan Lee
SLT4
2020 Sequence-to-Sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding
abstract
In this paper, we investigate the benefit that off-the-shelf word embedding can bring to the sequence-to-sequence (seq-to-seq) automatic speech recognition (ASR). We first introduced the word embedding regularization by maximizing the cosine similarity between a transformed decoder feature and the target word embedding. Based on the regularized decoder, we further proposed the fused decoding mechanism. This allows the decoder to consider the semantic consistency during decoding by absorbing the information carried by the transformed decoder feature, which is learned to be close to the target word embedding. Initial results on LibriSpeech demonstrated that pre-trained word embedding can significantly lower ASR recognition error with a negligible cost, and the choice of word embedding algorithms among Skip-gram, CBOW and BERT is important.
Alexander H. Liu, Tzu-Wei Sung, Shun-Po Chuang, Hung-yi Lee, Lin-Shan Lee
ICASSP5
2020 Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
abstract
In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is achieved by proper temporal segmentation to make the representations phoneme-synchronized, and proper phonetic clustering to have total number of distinct representations close to the number of phonemes. Mapping between the distinct representations and phonemes is learned from a small amount of annotated paired data. Preliminary experiments on LJSpeech demonstrated the learned representations for vowels have relative locations in latent space in good parallel to that shown in the IPA vowel chart defined by linguistics experts. With less than 20 minutes of annotated speech, our method outperformed existing methods on phoneme recognition and is able to synthesize intelligible speech that beats our baseline model.
Alexander H. Liu, Tao Tu 0002, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2020 Interrupted and Cascaded Permutation Invariant Training for Speech Separation
abstract
Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the minimum cost label assignments dynamically, very few studies considered the separation problem to be optimizing both the model parameters and the label assignments, but focused on searching for good model architecture and parameters. In this paper, we investigate instead for a given model architecture the various flexible label assignment strategies for training the model, rather than directly using PIT. Surprisingly, we discover a significant performance boost compared to PIT is possible if the model is trained with fixed label assignments and a good set of labels is chosen. With fixed label training cascaded between two sections of PIT, we achieved the state-of-the-art performance on WSJ0-2mix without changing the model architecture at all.
Gene-Ping Yang, Szu-Lin Wu, Yao-Wen Mao, Hung-yi Lee, Lin-Shan Lee
ICASSP5
2020 SpeechBERT: An Audio-and-Text Jointly Learned Language Model for End-to-End Spoken Question Answering
abstract
While various end-to-end models for spoken language understanding tasks have been explored recently, this paper is probably the first known attempt to challenge the very difficult task of end-to-end spoken question answering (SQA).Learning from the very successful BERT model for various text processing tasks, here we proposed an audio-and-text jointly learned SpeechBERT model.This model outperformed the conventional approach of cascading ASR with the following text question answering (TQA) model on datasets including ASR errors in answer spans, because the end-to-end model was shown to be able to extract information out of audio data before ASR produced errors.When ensembling the proposed end-to-end model with the cascade architecture, even better performance was achieved.In addition to the potential of end-to-end SQA, the SpeechBERT can also be considered for many other spoken language understanding tasks just as BERT for many text processing tasks.
Yung-Sung Chuang, Chi-Liang Liu, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2020 Doing Something we Never could with Spoken Language Technologies-from early days to the era of deep learning
Lin-Shan Lee
INTERSPEECH1
2019 Adversarial Training of End-to-end Speech Recognition Using a Criticizing Language Model
abstract
In this paper we proposed a novel Adversarial Training (AT) approach for end-to-end speech recognition using a Criticizing Language Model (CLM). In this way the CLM and the automatic speech recognition (ASR) model can challenge and learn from each other iteratively to improve the performance. Since the CLM only takes the text as input, huge quantities of unpaired text data can be utilized in this approach within end-to-end training. Moreover, AT can be applied to any end-to-end ASR model using any deep-learning-based language modeling frameworks, and compatible with any existing end-to-end decoding method. Initial results with an example experimental setup demonstrated the proposed approach is able to gain consistent improvements efficiently from auxiliary text data under different scenarios.
Alexander H. Liu, Hung-yi Lee, Lin-Shan Lee
ICASSP3
2019 Towards End-to-end Speech-to-text Translation with Two-pass Decoding
abstract
Speech-to-text translation (ST) refers to transforming the audio in source language to the text in target language. Mainstream solutions for such tasks are to cascade automatic speech recognition with machine translation, for which the transcriptions of the source language are needed in training. End-to-end approaches for ST tasks have been investigated because of not only technical interests such as to achieve globally optimized solution, but the need for ST tasks for the many source languages worldwide which do not have written form. In this paper, we propose a new end-to-end ST framework with two decoders to handle the relatively deeper relationships between the source language audio and target language text. The first-pass decoder generates some useful latent representations, and the second-pass decoder then integrates the output of both the encoder and the first-pass decoder to generate the text translation in target language. Only paired source language audio and target language text are used in training. Preliminary experiments on several language pairs showed improved performance, and offered some initial analysis.
Tzu-Wei Sung, Jun-You Liu, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2019 Completely Unsupervised Phoneme Recognition by a Generative Adversarial Network Harmonized with Iteratively Refined Hidden Markov Models
Kuan-Yu Chen 0005, Che-Ping Tsai, Da-Rong Liu, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH5
2019 Improved Speech Separation with Time-and-Frequency Cross-Domain Joint Embedding and Clustering
abstract
Speech separation has been very successful with deep learning techniques.Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for speech signals.It is highly correlated to the phonetic structure of speech, or "how the speech sounds" when perceived by human, but primarily frequency domain features carrying temporal behaviour.Very impressive work achieving speech separation over time domain was reported recently, probably because waveforms in time domain may describe the different realizations of speech in a more precise way than spectrogram.In this paper, we propose a framework properly integrating the above two directions, hoping to achieve both purposes.We construct a time-and-frequency feature map by concatenating the 1-dim convolution encoded feature map (for time domain) and the spectrogram (for frequency domain), which was then processed by an embedding network and clustering approaches very similar to those used in time and frequency domain prior works.In this way, the information in the time and frequency domains, as well as the interactions between them, can be jointly considered during embedding and clustering.Very encouraging results (state-of-the-art to our knowledge) were obtained with WSJ0-2mix dataset in preliminary experiments.
Gene-Ping Yang, Chao-I Tuan, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2018 Scalable Sentiment for Sequence-to-Sequence Chatbot Response with Performance Analysis
abstract
Conventional seq2seq chatbot models only try to find the sentences with the highest probabilities conditioned on the input sequences, without considering the sentiment of the output sentences. Some research works trying to modify the sentiment of the output sequences were reported. In this paper, we propose five models to scale or adjust the sentiment of the chatbot response: persona-based model, reinforcement learning, plug and play model, sentiment transformation network and cycleGAN, all based on the conventional seq2seq model. We also develop two evaluation metrics to estimate if the responses are reasonable given the input. These metrics together with other two popularly used metrics were used to analyze the performance of the five proposed models on different aspects, and reinforcement learning and cycleGAN were shown to be very attractive. The evaluation metrics were also found to be well correlated with human evaluation.
Chih-Wei Lee, Yau-Shian Wang, Tsung-Yuan Hsu, Kuan-Yu Chen 0005, Hung-yi Lee, Lin-Shan Lee
ICASSP6
2018 Domain Independent Key Term Extraction from Spoken Content Based on Context and Term Location Information in the Utterances
abstract
This paper proposes a domain independent approach for extracting key terms from spoken content based on context and term location information, or the sentence structures. Once it is trained with data of enough different domains, it is able to extract key terms in other unseen domains. This is obviously very attractive because of the unlimited number of domains over the Internet. Its performance degrades only very slightly with recognition errors, so very useful for spoken content. The basic idea here is that the sentence structures or context and term location information are in general domain independent, and remain essentially unchanged with recognition errors. For example, the fact that the key term for the sentence “The subject of this article is primarily about neural networks” is “neural networks” can be extended to any other unseen term other than “neural networks” in any other unseen domain, and this is more or less preserved under recognition errors. In the experiments a model trained with data for five different domains can extract key terms from data in the sixth unseen domain with very good performance.
Hsien-Chin Lin, Chi-Yu Yang, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2018 Transcribing Lyrics from Commercial Song Audio: the First Step Towards Singing Content Processing
abstract
Spoken content processing (such as retrieval and browsing) is maturing, but the singing content is still almost completely left out. Songs are human voice carrying plenty of semantic information just as speech, and may be considered as a special type of speech with highly flexible prosody. The various problems in song audio, for example the significantly changing phone duration over highly flexible pitch contours, make the recognition of lyrics from song audio much more difficult. This paper reports an initial attempt towards this goal. We collected music-removed version of English songs directly from commercial singing content. The best results were obtained by TDNN-BLSTM with data augmentation with 3-fold speed perturbation plus some special approaches. The WER achieved (73.90%) was significantly lower than the baseline (96.21 %), but still relatively high.
Che-Ping Tsai, Yi-Lin Tuan, Lin-Shan Lee
ICASSP3
2018 Segmental Audio Word2Vec: Representing Utterances as Sequences of Vectors with Applications in Spoken Term Detection
abstract
While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be trained in an unsupervised way from an unlabeled corpus, except the word boundaries are needed. In this paper, we extend audio Word2Vec from word-level to utterance-level by proposing a new segmental audio Word2Vec, in which unsupervised spoken word boundary segmentation and audio Word2Vec are jointly learned and mutually enhanced, so an utterance can be directly represented as a sequence of vectors carrying phonetic structure information. This is achieved by a segmental sequence-to-sequence autoencoder (SSAE), in which a segmentation gate trained with reinforcement learning is inserted in the encoder. Experiments on English, Czech, French and German show very good performance in both unsupervised spoken word segmentation and spoken term detection applications (significantly better than frame-based DTW).
Yu-Hsuan Wang, Hung-yi Lee, Lin-Shan Lee
ICASSP3
2018 Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations
abstract
Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for each target speaker.In this paper, we propose an adversarial learning framework for voice conversion, with which a single model can be trained to convert the voice to many different speakers, all without parallel data, by separating the speaker characteristics from the linguistic content in speech signals.An autoencoder is first trained to extract speaker-independent latent representations and speaker embedding separately using another auxiliary speaker classifier to regularize the latent representation.The decoder then takes the speaker-independent latent representation and the target speaker embedding as the input to generate the voice of the target speaker with the linguistic content of the source utterance.The quality of decoder output is further improved by patching with the residual signal produced by another pair of generator and discriminator.A target speaker set size of 20 was tested in the preliminary experiments, and very good voice quality was obtained.Conventional voice conversion metrics are reported.We also show that the speaker information has been properly reduced from the latent representations.
Ju-Chieh Chou, Cheng-chieh Yeh, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2018 Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings
abstract
Unsupervised discovery of acoustic tokens from audio corpora without annotation and learning vector representations for these tokens have been widely studied.Although these techniques have been shown successful in some applications such as query-byexample Spoken Term Detection (STD), the lack of mapping relationships between these discovered tokens and real phonemes have limited the down-stream applications.This paper represents probably the first attempt towards the goal of completely unsupervised phoneme recognition, or mapping audio signals to phoneme sequences without phoneme-labeled audio data.The basic idea is to cluster the embedded acoustic tokens and learn the mapping between the cluster sequences and the unknown phoneme sequences with a Generative Adversarial Network (GAN).An unsupervised phoneme recognition accuracy of 36% was achieved in the preliminary experiments.
Da-Rong Liu, Kuan-Yu Chen 0005, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2018 Phonetic-and-Semantic Embedding of Spoken words with Applications in Spoken Content Retrieval
abstract
Word embedding or Word2Vec has been successful in offering semantics for text words learned from the context of words. Audio Word2Vec was shown to offer phonetic structures for spoken words (signal segments for words) learned from signals within spoken words. This paper proposes a two-stage framework to perform phonetic-and-semantic embedding on spoken words considering the context of the spoken words. Stage 1 performs phonetic embedding with speaker characteristics disentangled. Stage 2 then performs semantic embedding in addition. We further propose to evaluate the phonetic-and-semantic nature of the audio embeddings obtained in Stage 2 by parallelizing with text embeddings. In general, phonetic structure and semantics inevitably disturb each other. For example the words “brother” and “sister” are close in semantics but very different in phonetic structure, while the words “brother” and “bother” are in the other way around. But phonetic-and-semantic embedding is attractive, as shown in the initial experiments on spoken document retrieval. Not only spoken documents including the spoken query can be retrieved based on the phonetic structures, but spoken documents semantically related to the query but not including the query can also be retrieved based on the semantics.
Sung-Feng Huang, Chia-Hao Shen, Hung-yi Lee, Lin-Shan Lee
SLT5
2018 Rhythm-Flexible Voice Conversion Without Parallel Data Using Cycle-GAN Over Phoneme Posteriorgram Sequences
abstract
Speaking rate refers to the average number of phonemes within some unit time, while the rhythmic patterns refer to duration distributions for realizations of different phonemes within different phonetic structures. Both are key components of prosody in speech, which is different for different speakers. Models like cycle-consistent adversarial network (Cycle-GAN) and variational auto-encoder (VAE) have been successfully applied to voice conversion tasks without parallel data. However, due to the neural network architectures and feature vectors chosen for these approaches, the length of the predicted utterance has to be fixed to that of the input utterance, which limits the flexibility in mimicking the speaking rates and rhythmic patterns for the target speaker. On the other hand, sequence-to-sequence learning model was used to remove the above length constraint, but parallel training data are needed. In this paper, we propose an approach utilizing sequence-to-sequence model trained with unsupervised Cycle-GAN to perform the transformation between the phoneme posteriorgram sequences for different speakers. In this way, the length constraint mentioned above is removed to offer rhythm-flexible voice conversion without requiring parallel data. Preliminary evaluation on two datasets showed very encouraging results.
Cheng-chieh Yeh, Po-Chun Hsu, Ju-Chieh Chou, Hung-yi Lee, Lin-Shan Lee
SLT5
2018 Unsupervised Discovery of Structured Acoustic Tokens With Applications to Spoken Term Detection
abstract
In this paper, we compare two paradigms for unsupervised discovery of structured acoustic tokens directly from speech corpora without any human annotation. The multigranular paradigm seeks to capture all available information in the corpora with multiple sets of tokens for different model granularities. The hierarchical paradigm attempts to jointly learn several levels of signal representations in a hierarchical structure. The two paradigms are unified within a theoretical framework in this paper. Query-by-example spoken term detection (QbE-STD) experiments on the query by example search on speech task dataset of MediaEval 2015 verifies the competitiveness of the acoustic tokens. The enhanced relevance score proposed in this work improves both paradigms for the task of QbE-STD. We also list results on the ABX evaluation task of the Zero Resource Challenge 2015 for comparison of the paradigms.
Cheng-Tao Chung, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Personalized word representations carrying personalized semantics learned from social network posts
abstract
Distributed word representations have been shown to be very useful in various natural language processing (NLP) application tasks. These word vectors learned from huge corpora very often carry both semantic and syntactic information of words. However, it is well known that each individual user has his own language patterns because of different factors such as interested topics, friend groups, social activities, wording habits, etc., which may imply some kind of personalized semantics. With such personalized semantics, the same word may imply slightly differently for different users. For example, the word “Cappuccino” may imply “Leisure”, “Joy”, “Excellent” for a user enjoying coffee, by only a kind of drink for someone else. Such personalized semantics of course cannot be carried by the standard universal word vectors trained with huge corpora produced by many people. In this paper, we propose a framework to train different personalized word vectors for different users based on the very successful continuous skip-gram model using the social network data posted by many individual users. In this framework, universal background word vectors are first learned from the background corpora, and then adapted by the personalized corpus for each individual user to learn the personalized word vectors. We use two application tasks to evaluate the quality of the personalized word vectors obtained in this way, the user prediction task and the sentence completion task. These personalized word vectors were shown to carry some personalized semantics and offer improved performance on these two evaluation tasks.
Zih-Wei Lin, Tzu-Wei Sung, Hung-yi Lee, Lin-Shan Lee
ASRU4
2017 Personalized acoustic modeling by weakly supervised multi-task deep learning using acoustic tokens discovered from unlabeled data
abstract
It is well known that recognizers personalized to each user are much more effective than user-independent recognizers. With the popularity of smartphones today, although it is not difficult to collect a large set of audio data for each user, it is difficult to transcribe it. However, it is now possible to automatically discover acoustic tokens from unlabeled personal data in an unsupervised way. We therefore propose a multi-task deep learning framework called a phoneme-token deep neural network (PTDNN), jointly trained from unsupervised acoustic tokens discovered from unlabeled data and very limited transcribed data for personalized acoustic modeling. We term this scenario “weakly supervised”. The underlying intuition is that the high degree of similarity between the HMM states of acoustic token models and phoneme models may help them learn from each other in this multi-task learning framework. Initial experiments performed over a personalized audio data set recorded from Facebook posts demonstrated that very good improvements can be achieved in both frame accuracy and word accuracy over popularly-considered baselines such as fDLR, speaker code and lightly supervised adaptation. This approach complements existing speaker adaptation approaches and can be used jointly with such techniques to yield improved results.
Cheng-Kuang Wei, Cheng-Tao Chung, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2017 Order-Preserving Abstractive Summarization for Spoken Content Based on Connectionist Temporal Classification
abstract
Connectionist temporal classification (CTC) is a powerful approach for sequence-to-sequence learning, and has been popularly used in speech recognition.The central ideas of CTC include adding a label "blank" during training.With this mechanism, CTC eliminates the need of segment alignment, and hence has been applied to various sequence-to-sequence learning problems.In this work, we applied CTC to abstractive summarization for spoken content.The "blank" in this case implies the corresponding input data are less important or noisy; thus it can be ignored.This approach was shown to outperform the existing methods in term of ROUGE scores over Chinese Gigaword and MATBN corpora.This approach also has the nice property that the ordering of words or characters in the input documents can be better preserved in the generated summaries.
Bo-Ru Lu, Frank Shyu, Yun-Nung Chen, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH5
2017 Unsupervised Iterative Deep Learning of Speech Features and Acoustic Tokens with Applications to Spoken Term Detection
abstract
In this paper, we aim to automatically discover high-quality frame-level speech features and acoustic tokens directly from unlabeled speech data. A multigranular acoustic tokenizer (MAT) was proposed for automatic discovery of multiple sets of acoustic tokens from the given corpus. Each acoustic token set is specified by a set of hyperparameters describing the model configuration. These different sets of acoustic tokens carry different characteristics for the given corpus and the language behind and, thus, can be mutually reinforced. The multiple sets of token labels are then used as the targets of a multitarget deep neural network (MDNN) trained on frame-level acoustic features. Bottleneck features extracted from the MDNN are then used as the feedback input to the MAT and the MDNN itself in the next iteration. The multigranular acoustic token sets and the frame-level speech features can be iteratively optimized in the iterative deep learning framework. We call this framework the MAT deep neural network. The results were evaluated using the metrics and corpora defined in the Zero Resource Speech Challenge organized at Interspeech 2015, and improved performance was obtained with a set of experiments of query-by-example spoken term detection on the same corpora. Visualization for the discovered tokens against the English phonemes was also shown.
Cheng-Tao Chung, Cheng-Yu Tsai, Chia-Hsiang Liu, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Audio Word2Vec: Unsupervised Learning of Audio Segment Representations Using Sequence-to-Sequence Autoencoder
abstract
The vector representations of fixed dimensionality for words (in text) offered by Word2Vec have been shown to be very useful in many application scenarios, in particular due to the semantic information they carry.This paper proposes a parallel version, the Audio Word2Vec.It offers the vector representations of fixed dimensionality for variable-length audio segments.These vector representations are shown to describe the sequential phonetic structures of the audio segments to a good degree, with very attractive real world applications such as query-by-example Spoken Term Detection (STD).In this STD application, the proposed approach significantly outperformed the conventional Dynamic Time Warping (DTW) based approaches at significantly lower computation requirements.We propose unsupervised learning of Audio Word2Vec from audio data without human annotation using Sequence-to-sequence Autoencoder (SA).SA consists of two RNNs equipped with Long Short-Term Memory (LSTM) units: the first RNN (encoder) maps the input audio sequence into a vector representation of fixed dimensionality, and the second RNN (decoder) maps the representation back to the input audio sequence.The two RNNs are jointly trained by minimizing the reconstruction error.Denoising Sequence-to-sequence Autoencoder (DSA) is further proposed offering more robust learning.
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH5
2016 Towards Machine Comprehension of Spoken Content: Initial TOEFL Listening Comprehension Test by Machine
abstract
Multimedia or spoken content presents more attractive information than plain text content, but it's more difficult to display on a screen and be selected by a user. As a result, accessing large collections of the former is much more difficult and time-consuming than the latter for humans. It's highly attractive to develop a machine which can automatically understand spoken content and summarize the key information for humans to browse over. In this endeavor, we propose a new task of machine comprehension of spoken content. We define the initial goal as the listening comprehension test of TOEFL, a challenging academic English examination for English learners whose native language is not English. We further propose an Attention-based Multi-hop Recurrent Neural Network (AMRNN) architecture for this task, achieving encouraging results in the initial tests. Initial results also have shown that word-level attention is probably more robust than sentence-level attention for this task with ASR errors.
Bo-Hsiang Tseng, Sheng-syun Shen, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2016 Interactive Spoken Content Retrieval by Deep Reinforcement Learning
abstract
User-machine interaction is important for spoken content retrieval.For text content retrieval, the user can easily scan through and select on a list of retrieved item.This is impossible for spoken content retrieval, because the retrieved items are difficult to show on screen.Besides, due to the high degree of uncertainty for speech recognition, the retrieval results can be very noisy.One way to counter such difficulties is through user-machine interaction.The machine can take different actions to interact with the user to obtain better retrieval results before showing to the user.The suitable actions depend on the retrieval status, for example requesting for extra information from the user, returning a list of topics for user to select, etc.In our previous work, some hand-crafted states estimated from the present retrieval results are used to determine the proper actions.In this paper, we propose to use Deep-Q-Learning techniques instead to determine the machine actions for interactive spoken content retrieval.Deep-Q-Learning bypasses the need for estimation of the hand-crafted states, and directly determine the best action base on the present retrieval status even without any human knowledge.It is shown to achieve significantly better performance compared with the previous hand-crafted states.
Yen-Chen Wu, Tzu-Hsiang Lin, Yang-De Chen, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH5
2016 Hierarchical attention model for improved machine comprehension of spoken content
abstract
Multimedia or spoken content presents more attractive information than plain text content, but the former is more difficult to display on a screen and be selected by a user. As a result, accessing large collections of the former is much more difficult and time-consuming than the latter for humans. It's therefore highly attractive to develop machines which can automatically understand spoken content and summarize the key information for humans to browse over. In this endeavor, a new task of machine comprehension of spoken content was proposed recently. The initial goal was defined as the listening comprehension test of TOEFL, a challenging academic English examination for English learners whose native languages are not English. An Attention-based Multi-hop Recurrent Neural Network (AMRNN) architecture was also proposed for this task, which considered only the sequential relationship within the speech utterances. In this paper, we propose a new Hierarchical Attention Model (HAM), which constructs multi-hopped attention mechanism over tree-structured rather than sequential representations for the utterances. Improved comprehension performance robust with respect to ASR errors were obtained.
Juei-Yang Hsu, Hung-yi Lee, Lin-Shan Lee
SLT4
2016 Abstractive headline generation for spoken content by attentive recurrent neural networks with ASR error modeling
abstract
Headline generation for spoken content is important since spoken content is difficult to be shown on the screen and browsed by the user. It is a special type of abstractive summarization, for which the summaries are generated word by word from scratch without using any part of the original content. Many deep learning approaches for headline generation from text document have been proposed recently, all requiring huge quantities of training data, which is difficult for spoken document summarization. In this paper, we propose an ASR error modeling approach to learn the underlying structure of ASR error patterns and incorporate this model in an Attentive Recurrent Neural Network (ARNN) architecture. In this way, the model for abstractive headline generation for spoken content can be learned from abundant text data and the ASR data for some recognizers. Experiments showed very encouraging results and verified that the proposed ASR error model works well even when the input spoken content is recognized by a recognizer very different from the one the model learned from.
Lang-Chi Yu, Hung-yi Lee, Lin-Shan Lee
SLT3
2015 An iterative deep learning framework for unsupervised discovery of speech features and linguistic units with applications on spoken term detection
abstract
In this work we aim to discover high quality speech features and Linguistic units directly from unlabeled speech data in a zero resource scenario. The results are evaluated using the metrics and corpora proposed in the Zero Resource Speech Challenge organized at Interspeech 2015. A Multi-layered Acoustic Tokenizer (MAT) was proposed for automatic discovery of multiple sets of acoustic tokens from the given corpus. Each acoustic token set is specified by a set of hyperparameters that describe the model configuration. These sets of acoustic tokens carry different characteristics fof the given corpus and the language behind, thus can be mutually reinforced. The multiple sets of token labels are then used as the targets of a Multi-target Deep Neural Network (MDNN) trained on low-level acoustic features. Bottleneck features extracted from the MDNN are then used as the feedback input to the MAT and the MDNN itself in the next iteration. We call this iterative deep learning framework the Multi-layered Acoustic Tokenizing Deep Neural Network (MAT-DNN), which generates both high quality speech features for the Track 1 of the Challenge and acoustic tokens for the Track 2 of the Challenge. In addition, we performed extra experiments on the same corpora on the application of query-by-example spoken term detection. The experimental results showed the iterative deep learning framework of MAT-DNN improved the detection performance due to better underlying speech features and acoustic tokens.
Cheng-Tao Chung, Cheng-Yu Tsai, Hsiang-Hung Lu, Chia-Hsiang Liu, Hung-yi Lee, Lin-Shan Lee
ASRU6
2015 Towards structured deep neural network for automatic speech recognition
abstract
In this paper we propose the Structured Deep Neural Network (structured DNN) as a structured and deep learning framework. This approach can learn to find the best structured object (such as a label sequence) given a structured input (such as a vector sequence) by globally considering the mapping relationships between the structures rather than item by item. When automatic speech recognition is viewed as a special case of such a structured learning problem, where we have the acoustic vector sequence as the input and the phoneme label sequence as the output, it becomes possible to comprehensively learn utterance by utterance as a whole, rather than frame by frame. Structured Support Vector Machine (structured SVM) was proposed to perform ASR with structured learning previously, but limited by the linear nature of SVM. Here we propose structured DNN to use nonlinear transformations in multi-layers as a structured and deep learning approach. This approach was shown to beat structured SVM in preliminary experiments on TIMIT.
Yi-Hsiu Liao, Hung-yi Lee, Lin-Shan Lee
ASRU3
2015 Personalizing universal recurrent neural network language model with user characteristic features by social network crowdsourcing
abstract
With the popularity of mobile devices, personalized speech recognizer becomes more realizable today and highly attractive. Each mobile device is primarily used by a single user, so it's possible to have a personalized recognizer well matching to the characteristics of individual user. Although acoustic model personalization has been investigated for decades, much less work have been reported on personalizing language model, probably because of the difficulties in collecting enough personalized corpora. Previous work used the corpora collected from social networks to solve the problem, but constructing a personalized model for each user is troublesome. In this paper, we propose a universal recurrent neural network language model with user characteristic features, so all users share the same model, except each with different user characteristic features. These user characteristic features can be obtained by crowdsouring over social networks, which include huge quantity of texts posted by users with known friend relationships, who may share some subject topics and wording patterns. The preliminary experiments on Facebook corpus showed that this proposed approach not only drastically reduced the model perplexity, but offered very good improvement in recognition accuracy in n-best rescoring tests. This approach also mitigated the data sparseness problem for personalized language models.
Bo-Hsiang Tseng, Hung-yi Lee, Lin-Shan Lee
ASRU3
2015 Enhancing automatically discovered multi-level acoustic patterns considering context consistency with applications in spoken term detection
abstract
This paper presents a novel approach for enhancing the multiple sets of acoustic patterns automatically discovered from a given corpus. In a previous work it was proposed that different HMM configurations (number of states per model, number of distinct models) for the acoustic patterns form a two-dimensional space. Multiple sets of acoustic patterns automatically discovered with the HMM configurations properly located on different points over this two-dimensional space were shown to be complementary to one another, jointly capturing the characteristics of the given corpus. By representing the given corpus as sequences of acoustic patterns on different HMM sets, the pattern indices in these sequences can be relabeled considering the context consistency across the different sequences. Good improvements were observed in preliminary experiments of pattern spoken term detection (STD) performed on both TIMIT and Mandarin Broadcast News with such enhanced patterns.
Cheng-Tao Chung, Wei-Ning Hsu, Lin-Shan Lee
ICASSP4
2015 Enhancing sparse voice annotation for semantic retrieval of personal photos by continuous space word representations
abstract
It is very attractive for the user to retrieve photos from a huge collection using high-level personal queries (e.q. uncle Bill's house), but technically very challenging. The previous work proposed a set of approaches to achieve the goal assuming only 30% of the photos are annotated by sparse spoken descriptions when the photos are taken. This includes fusing the sparse spontaneously spoken features with visual features of the photos by non-negative matrix factorization (NMF), and enhancing the results with two-layer mutually reinforced random walk. However, because the speech annotation is very sparse, the retrieval is very often dominated by the very complete visual features. In this paper, we propose to use continuous space word representations to extend the sparse speech information and expand the photo representation to enhance the retrieval model. Very encouraging improvements were observed in the preliminary experiments.
Yuan-ming Liou, Hung-tsung Lu, Yi-Sheng Fu, Winston H. Hsu, Lin-Shan Lee
ICASSP5
2015 Semantic retrieval of personal photos using a deep autoencoder fusing visual features with speech annotations represented as word/paragraph vectors
Hung-tsung Lu, Yuan-ming Liou, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2015 Structuring lectures in massive open online courses (MOOCs) for efficient learning by linking similar sections and predicting prerequisites
abstract
The increasing popularity of Massive Open Online Courses (MOOCs) has resulted in huge number of courses available over the Internet. Typically, a learner can type a search query into the look-up window of a MOOC platform and receive a set of course suggestions. But it is difficult for the learner to select lectures out of those suggested courses and learn the desired information efficiently. In this paper, we propose to structure the lectures of the various suggested courses into a map (graph) for each query entered by the learner, indicating the lectures with very similar content and reasonable sequence order of learning. In this way the learner can define his own learning path on the map based on his interests and backgrounds, and learn the desired information from lectures in different courses without too much difficulties in minimum time. We propose a series of approaches for linking lectures of very similar content and predicting the prerequisites for this purpose. Preliminary results show that the proposed approaches have the potential to achieve the above goal.
Sheng-syun Shen, Hung-yi Lee, Shang-Wen Li 0001, Victor Zue, Lin-Shan Lee
INTERSPEECH5
2015 Personalized speech recognizer with keyword-based personalized lexicon and language model using word vector representations
Ching-Feng Yeh, Yuan-ming Liou, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2015 Finding Complex Features for Guest Language Fragment Recovery in Resource-Limited Code-Mixed Speech Recognition
abstract
The rise of mobile devices and online learning brings into sharp focus the importance of speech recognition not only for the many languages of the world but also for code-mixed speech, especially where English is the second language. The recognition of code-mixed speech, where the speaker mixes languages within a single utterance, is a challenge for both computers and humans, not least because of the limited training data. We conduct research on a Mandarin-English code-mixed lecture corpus, where Mandarin is the host language and English the guest language, and attempt to find complex features for the recovery of English segments that were misrecognized in the initial recognition pass. We propose a multi-level framework wherein both low-level and high-level cues are jointly considered; we use phonotactic, prosodic, and linguistic cues in addition to acoustic-phonetic cues to discriminate at the frame level between English- and Chinese-language segments. We develop a simple and exact method for CRF feature induction, and improved methods for using cascaded features derived from the training corpus. By additionally tuning the data imbalance ratio between English and Chinese, we demonstrate highly significant improvements over previous work in the recovery of English-language segments, and demonstrate performance superior to DNN-based methods. We demonstrate considerable performance improvements not only with the traditional GMM-HMM recognition paradigm but also with a state-of-the-art hybrid CD-HMM-DNN recognition framework.
Aaron Heidel, Hsiang-Hung Lu, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Spoken Content Retrieval - Beyond Cascading Speech Recognition with Text Retrieval
abstract
Spoken content retrieval refers to directly indexing and retrieving spoken content based on the audio rather than text descriptions. This potentially eliminates the requirement of producing text descriptions for multimedia content for indexing and retrieval purposes, and is able to precisely locate the exact time the desired information appears in the multimedia. Spoken content retrieval has been very successfully achieved with the basic approach of cascading automatic speech recognition (ASR) with text information retrieval: after the spoken content is transcribed into text or lattice format, a text retrieval engine searches over the ASR output to find desired information. This framework works well when the ASR accuracy is relatively high, but becomes less adequate when more challenging real-world scenarios are considered, since retrieval performance depends heavily on ASR accuracy. This challenge leads to the emergence of another approach to spoken content retrieval: to go beyond the basic framework of cascading ASR with text retrieval in order to have retrieval performances that are less dependent on ASR accuracy. This overview article is intended to provide a thorough overview of the concepts, principles, approaches, and achievements of major technical contributions along this line of investigation. This includes five major directions: 1) Modified ASR for Retrieval Purposes: cascading ASR with text retrieval, but the ASR is modified or optimized for spoken content retrieval purposes; 2) Exploiting the Information not present in ASR outputs: to try to utilize the information in speech signals inevitably lost when transcribed into phonemes and words; 3) Directly Matching at the Acoustic Level without ASR: for spoken queries, the signals can be directly matched at the acoustic level, rather than at the phoneme or word levels, bypassing all ASR issues; 4) Semantic Retrieval of Spoken Content: trying to retrieve spoken content that is semantically related to the query, but not necessarily including the query terms themselves; 5) Interactive Retrieval and Efficient Presentation of the Retrieved Objects: with efficient presentation of the retrieved objects, an interactive retrieval process incorporating user actions may produce better retrieval results and user experiences.
Lin-Shan Lee, James R. Glass, Hung-yi Lee, Chun-an Chan
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 A Recursive Dialogue Game for Personalized Computer-Aided Pronunciation Training
abstract
Learning languages in addition to the native language is very important for all people in the globalized world today, and computer-aided pronunciation training (CAPT) is attractive since the software can be used anywhere at any time, and repeated as many times as desired. In this paper, we introduce the immersive interaction scenario offered by spoken dialogues to CAPT by proposing a recursive dialogue game to make CAPT personalized. A number of tree-structured sub-dialogues are linked sequentially and recursively as the script for the game. The system policy at each dialogue turn is to select in real-time along the dialogue the best training sentence for each specific individual learner within the dialogue script, considering the learner's learning status and the future possible dialogue paths in the script, such that the learner can have the scores for all pronunciation units considered reaching a predefined standard in a minimum number of turns. The purpose here is that those pronunciation units poorly produced by the specific learner can be offered with more practice opportunities in the future sentences along the dialogue, which enables the learner to improve the pronunciation without having to repeat the same training sentences many times. This makes the learning process for each learner completely personalized. The dialogue policy is modeled by Markov decision process (MDP) with high-dimensional continuous state space, and trained with fitted value iteration using a huge number of simulated learners. These simulated leaners have the behavior similar to real learners, and were generated from a corpus of real learner data. The experiments demonstrated very promising results and a real cloud-based system is also successfully implemented.
Pei-hao Su, Chuan-Hsun Wu, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Supervised Detection and Unsupervised Discovery of Pronunciation Error Patterns for Computer-Assisted Language Learning
abstract
Pronunciation error patterns (EPs) are patterns of mispronunciation frequently produced by language learners, and are usually different for different pairs of target and native languages. Accurate information of EPs can offer helpful feedbacks to the learners to improve their language skills. However, the major difficulty of EP detection comes from the fact that EPs are intrinsically similar to their corresponding canonical pronunciation, and different EPs corresponding to same canonical pronunciation are also intrinsically similar to each other. As a result, distinguishing EPs from their corresponding canonical pronunciation and between different EPs of the same phoneme is a difficult task-perhaps even more difficult than distinguishing between different phonemes in one language. On the other hand, the cost of deriving all EPs for each pair of target and native languages is high, usually requiring extensive expert knowledge or high-quality annotated data. Unsupervised EP discovery from a corpus of learner recordings would thus be an attractive addition to the field. In this paper, we propose new frameworks for both supervised EP detection and unsupervised EP discovery. For supervised EP detection, we use hierarchical multi-layer perceptrons (MLPs) as the EP classifiers to be integrated with the baseline using HMM/GMM in a two-pass Viterbi decoding architecture. Experimental results show that the new framework enhances the power of EP diagnosis. For unsupervised EP discovery we propose the first known framework, using the hierarchical agglomerative clustering (HAC) algorithm to explore sub-segmental variation within phoneme segments and produce fixed-length segment-level feature vectors in order to distinguish different EPs. We tested K-means (assuming a known number of EPs) and the Gaussian mixture model with the minimum description length principle (estimating an unknown number of EPs) for EP discovery. Preliminary experiments offered very encouraging results, although there is still a long way to go to approach the performance of human experts. We also propose to use the universal phoneme posteriorgram (UPP), derived from an MLP trained on corpora of mixed languages, as frame-level features in both supervised detection and unsupervised discovery of EPs. Experimental results show that using UPP not only achieves the best performance, but also is useful in analyzing the mispronunciation produced by language learners.
Yow-Bang Wang, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 An Improved Framework for Recognizing Highly Imbalanced Bilingual Code-Switched Lectures with Cross-Language Acoustic Modeling and Frame-Level Language Identification
abstract
This paper considers the recognition of a widely observed type of bilingual code-switched speech: the speaker speaks primarily the host language (usually his native language), but with a few words or phrases in the guest language (usually his second language) inserted in many utterances of the host language. In this case, not only the languages are switched back and forth within an utterance so the language identification is difficult, but much less data are available for the guest language, which results in poor recognition accuracy for the guest language part. Unit merging approaches on three levels of acoustic modeling (triphone models, HMM states and Gaussians) have been proposed for cross-lingual data sharing for such highly imbalanced bilingual code-switched speech. In this paper, we present an improved overall framework on top of the previously proposed unit merging approaches for recognizing such code-switched speech. This includes unit recovery for reconstructing the identity for units of the two languages after being merged, unit occupancy ranking to offer much more flexible data sharing between units both across languages and within the language based on the accumulated occupancy of the HMM states, and estimation of frame-level language posteriors using blurred posteriorgram features (BPFs) to be used in decoding. We also present a complete set of experimental results comparing all approaches involved for a real-world application scenario under unified conditions, and show very good improvement achieved with the proposed approaches.
Ching-Feng Yeh, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Unsupervised spoken term detection with spoken queries by multi-level acoustic patterns with varying model granularity
abstract
This paper presents a new approach for unsupervised Spoken Term Detection with spoken queries using multiple sets of acoustic patterns automatically discovered from the target corpus. The different pattern HMM configurations(number of states per model, number of distinct models, number of Gaussians per state)form a three-dimensional model granularity space. Different sets of acoustic patterns automatically discovered on different points properly distributed over this three-dimensional space are complementary to one another, thus can jointly capture the characteristics of the spoken terms. By representing the spoken content and spoken query as sequences of acoustic patterns, a series of approaches for matching the pattern index sequences while considering the signal variations are developed. In this way, not only the on-line computation load can be reduced, but the signal distributions caused by different speakers and acoustic conditions can be reasonably taken care of. The results indicate that this approach significantly outperformed the unsupervised feature-based DTW baseline by 16.16% in mean average precision on the TIMIT corpus.
Cheng-Tao Chung, Chun-an Chan, Lin-Shan Lee
ICASSP3
2014 Transcribing code-switched bilingual lectures using deep neural networks with unit merging in acoustic modeling
abstract
This paper considers the transcription of the widely observed yet less investigated bilingual code-switched speech: the words or phrases of the guest language are inserted within the utterances of the host language, so the languages are switched back and forth within an utterance, and much less data are available for the guest language. Two approaches utilizing the deep neural network (DNN) were tested and analyzed, including using DNN bottleneck features in HMM/GMM (BF-HMM/GMM) and modeling context-dependent HMM senones by DNN (CD-DNN-HMM). In both cases the unit merging (and recovery) techniques in acoustic modeling were used to handle the data imbalance problem. Improved recognition accuracies were observed with unit merging (and recovery) for the two approaches under different conditions.
Ching-Feng Yeh, Lin-Shan Lee
ICASSP2
2014 Semantic retrieval of personal photos using matrix factorization and two-layer random walk fusing sparse speech annotations with visual features
Yuan-ming Liou, Yi-Sheng Fu, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH4
2014 Alignment of spoken utterances with slide content for easier learning with recorded lectures using structured support vector machine (SVM)
Sheng-syun Shen, Sz-Rung Shiang, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH5
2014 Spoken question answering using tree-structured conditional random fields and two-layer random walk
Sz-Rung Shiang, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH3
2014 Improved open-vocabulary spoken content retrieval with word and subword lattices using acoustic feature similarity
Hung-yi Lee, Po-wei Chou, Lin-Shan Lee
Comput. Speech Lang.3
2014 Improved Semantic Retrieval of Spoken Content by Document/Query Expansion with Random Walk Over Acoustic Similarity Graphs
abstract
In a text context, document/query expansion has proven very useful in retrieving objects semantically related to the query. However, when applying text-based techniques on spoken content, the inevitable recognition errors seriously degrade performance even when the retrieval process is performed over lattices. We propose the estimation of more accurate term distributions (or unigram language models) for the spoken documents by acoustic similarity graphs. In this approach, a graph is constructed for each term describing the acoustic similarity among all signal regions hypothesized to be the considered term. Score propagation based on a random walk over the graph offers more reliable scores of the term hypotheses, which in turn yield more accurate term distributions (or unigram language models). This approach was applied with the language modeling retrieval approach, including using document expansion based on latent topic analysis and query expansion with a query-regularized mixture model. We extend these approaches from words to subword n-grams, and the query expansion from document-level to utterance-level and from term-based to topic-based. Experiments performed on Mandarin broadcast news showed improved performance under almost all tested conditions.
Hung-yi Lee, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Spoken Knowledge Organization by Semantic Structuring and a Prototype Course Lecture System for Personalized Learning
abstract
It takes very long time to go through a complete online course. Without proper background, it is also difficult to understand retrieved spoken paragraphs. This paper therefore presents a new approach of spoken knowledge organization for course lectures for efficient personalized learning. Automatically extracted key terms are taken as the fundamental elements of the semantics of the course. Key term graph constructed by connecting related key terms forms the backbone of the global semantic structure. Audio/video signals are divided into multi-layer temporal structure including paragraphs, sections and chapters, each of which includes a summary as the local semantic structure. The interconnection between semantic structure and temporal structure together with spoken term detection jointly offer to the learners efficient ways to navigate across the course knowledge with personalized learning paths considering their personal interests, available time and background knowledge. A preliminary prototype system has also been successfully developed.
Hung-yi Lee, Sz-Rung Shiang, Ching-Feng Yeh, Yun-Nung Chen, Sheng-yi Kong, Lin-Shan Lee
IEEE ACM Trans. Audio Speech Lang. Process.7
2013 Towards unsupervised semantic retrieval of spoken content with query expansion based on automatically discovered acoustic patterns
abstract
This paper presents an initial effort to retrieve semantically related spoken content in a completely unsupervised way. Unsupervised approaches of spoken content retrieval is attractive because the need for annotated data reasonably matched to the spoken content for training acoustic and language models can be bypassed. However, almost all such unsupervised approaches focus on spoken term detection, or returning the spoken segments containing the query, using either template matching techniques such as dynamic time warping (DTW) or model-based approaches. However, users usually prefer to retrieve all objects semantically related to the query, but not necessarily including the query terms. This paper proposes a different approach. We transcribe the spoken segments in the archive to be retrieved through into sequences of acoustic patterns automatically discovered in an unsupervised method. For an input query in spoken form, the top-N spoken segments from the archive obtained with the first-pass retrieval with DTW are taken as pseudo-relevant. The acoustic patterns frequently occurring in these segments are therefore considered as query-related and used for query expansion. Preliminary experiments performed on Mandarin broadcast news offered very encouraging results.
Yun-Chiao Li, Hung-yi Lee, Cheng-Tao Chung, Chun-an Chan, Lin-Shan Lee
ASRU5
2013 Toward unsupervised model-based spoken term detection with spoken queries without annotated data
abstract
We present a two-stage model-based approach for unsupervised query-by-example spoken term detection (STD) without any annotated data. Compared to the prevailing DTW approaches for the unsupervised STD task, HMMs used by model-based approaches can better capture the signal distributions and time trajectories of speech with a more global view of the spoken archive; matching with model states also significantly reduces the computational load. The utterances in the spoken archive are first offline decoded into acoustic patterns automatically discovered in an unsupervised way from the spoken archive. In the first stage, we propose a document state matching (DSM) approach, where query frames are matched to the HMM state sequences for the spoken documents. In this process, a novel duration-constrained Viterbi (DC-Vite) algorithm is proposed to avoid unrealistic speaking rate distortion. In the second stage, pseudo relevant/irrelevant examples retrieved from the first stage are respectively used to construct query/anti-query HMMs. Each spoken term hypothesis is then rescored with the likelihood ratio to these two HMMs. Experimental results show an absolute 11.8% of mean average precision improvement with a more than 50% reduction in computation time compared to the segmental DTW approach on a Mandarin broadcast news corpus.
Chun-an Chan, Cheng-Tao Chung, Yu-Hsin Kuo, Lin-Shan Lee
ICASSP4
2013 Unsupervised discovery of linguistic structure including two-level acoustic patterns using three cascaded stages of iterative optimization
abstract
Techniques for unsupervised discovery of acoustic patterns are getting increasingly attractive, because huge quantities of speech data are becoming available but manual annotations remain hard to acquire. In this paper, we propose an approach for unsupervised discovery of linguistic structure for the target spoken language given raw speech data. This linguistic structure includes two-level (subword-like and word-like) acoustic patterns, the lexicon of word-like patterns in terms of subword-like patterns and the N-gram language model based on word-like patterns. All patterns, models, and parameters can be automatically learned from the unlabelled speech corpus. This is achieved by an initialization step followed by three cascaded stages for acoustic, linguistic, and lexical iterative optimization. The lexicon of word-like patterns defines allowed consecutive sequence of HMMs for subword-like patterns. In each iteration, model training and decoding produces updated labels from which the lexicon and HMMs can be further updated. In this way, model parameters and decoded labels are respectively optimized in each iteration, and the knowledge about the linguistic structure is learned gradually layer after layer. The proposed approach was tested in preliminary experiments on a corpus of Mandarin broadcast news, including a task of spoken term detection with performance compared to a parallel test using models trained in a supervised way. Results show that the proposed system not only yields reasonable performance on its own, but is also complimentary to existing large vocabulary ASR systems.
Cheng-Tao Chung, Chun-an Chan, Lin-Shan Lee
ICASSP3
2013 Unsupervised domain adaptation for spoken document summarization with structured support vector machine
abstract
Supervised approaches can learn a spoken document summarizer generating high-quality summaries using a set of training examples matched to the domain of target documents. However, preparing a sufficient number of in-domain training examples is expensive. In this paper we propose an approach for unsupervised domain adaptation for spoken document summarization, so no in-domain training examples are needed. A summarizer is first learned from a set of out-of-domain training examples by a supervised summarization approach based on structured support vector machine, and this summarizer is used to generate a set of initial summaries for the target spoken documents. The target documents and their initial machine-generated summaries then serve as extra training examples for learning a new summarizer, which further updates the summaries of the target spoken documents. This process is continued iteratively to incrementally improve the summarizer for the target spoken documents. Moreover, extra approaches transforming the feature representations based on the data distribution in the target domain and augmenting the representations with an extra set of domain-specific features are also proposed. Encouraging results were obtained in summarizing Mandarin-English code-switching course lectures using training examples from Mandarin broadcast news.
Hung-yi Lee, Yu-Yu Chou, Yow-Bang Wang, Lin-Shan Lee
ICASSP4
2013 Enhancing query expansion for semantic retrieval of spoken content with automatically discovered acoustic patterns
abstract
Query expansion techniques were originally developed for text information retrieval in order to retrieve the documents not containing the query terms but semantically related to the query. This is achieved by assuming the terms frequently occurring in the top-ranked documents in the first-pass retrieval results to be query-related and using them to expand the query to do the second-pass retrieval. However, when this approach was used for spoken content retrieval, the inevitable recognition errors and the OOV problems in ASR make it difficult for many query-related terms to be included in the expanded query, and much of the information carried by the speech signal is lost during recognition and not recoverable. In this paper, we propose to use a second ASR engine based on acoustic patterns automatically discovered from the spoken archive used for retrieval. These acoustic patterns are discovered directly based on the signal characteristics, and therefore can compensate for the information lost during recognition to a good extent. When a text query is entered, the system generates the first-pass retrieval results based on the transcriptions of the spoken segments obtained via the conventional ASR. The acoustic patterns frequently occurring in the spoken segments ranked on top of the first-pass results are considered as query-related, and the spoken segments containing these query-related acoustic patterns are retrieved. In this way, even though some query-related terms are OOV or incorrectly recognized, the segments including these terms can still be retrieved by acoustic patterns corresponding to these terms. Preliminary experiments performed on Mandarin broadcast news offered very encouraging results.
Hung-yi Lee, Yun-Chiao Li, Cheng-Tao Chung, Lin-Shan Lee
ICASSP4
2013 A dialogue game framework with personalized training using reinforcement learning for computer-assisted language learning
abstract
We propose a framework for computer-assisted language learning as a pedagogical dialogue game. The goal is to offer personalized learning sentences on-line for each individual learner considering the learner's learning status, in order to strike a balance between more practice on poorly-pronounced units and complete practice on the whole set of pronunciation units. This objective is achieved using a Markov decision process (MDP) trained with reinforcement learning using simulated learners generated from real learner data. Preliminary experimental results on a subset of the example dialogue script show the effectiveness of the framework.
Pei-hao Su, Yow-Bang Wang, Tien-han Yu, Lin-Shan Lee
ICASSP4
2013 Toward unsupervised discovery of pronunciation error patterns using universal phoneme posteriorgram for computer-assisted language learning
abstract
In Computer-Aided Pronunciation Training, we hope to specify the type of mispronunciation, or Error Pattern (EP), the language learner has made as a more effective feedback. But derivation of EPs usually requires expert knowledge and pedagogical experiences, which is not easy to obtain for each pair of target and native languages. In this paper we propose a preliminary framework toward unsupervised discovery of EPs from a corpus of learners' recordings. We use Universal Phoneme Posteriorgram, derived from Multi-Layer Perceptron trained with a corpus of mixed languages, as features to bring supervised knowledge into the unsupervised task. We also use Hierarchical Agglomerative Clustering algorithm to explore sub-segmental variation of phoneme segments for distinguishing EPs. We tested K-means (assuming known number of EPs) and Gaussian Mixture Model with minimum description length principle (estimating unknown number of EPs) for EP discovery. Preliminary experimental results illustrated the effectiveness of the proposed framework, although there is still a long way to go compared to human annotators.
Yow-Bang Wang, Lin-Shan Lee
ICASSP2
2013 Interactive spoken content retrieval by extended query model and continuous state space Markov Decision Process
abstract
Interactive retrieval is important for spoken content because the retrieved spoken items are not only difficult to be shown on the screen but also scanned and selected by the user, in addition to the speech recognition uncertainty. The user cannot playback and go through all the retrieved items to find out what he is looking for. Markov Decision Process (MDP) was used in a previous work to help the system take different actions to interact with the user based on an estimated retrieval performance, but the MDP state was represented by the less precise quantized retrieval performance metric. In this paper, we consider the retrieval performance metric as a continuous state variable in MDP and optimize the MDP by fitted value iteration (FVI).We also use query expansion with the language modeling retrieval framework to produce the next set of retrieval results. Improved performance was found in the preliminary experiments.
Tsung-Hsien Wen, Hung-yi Lee, Pei-hao Su, Lin-Shan Lee
ICASSP4
2013 Supervised spoken document summarization based on structured support vector machine with utterance clusters as hidden variables
Sz-Rung Shiang, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH3
2013 A recursive dialogue game framework with optimal Policy offering personalized computer-assisted language learning
abstract
This paper introduces a new recursive dialogue game frame-work for personalized computer-assisted language learning. A series of sub-dialogue trees are cascaded into a loop as the script for the game. At each dialogue turn there are a number of train-ing sentences to be selected. The dialogue policy is optimized to offer the most appropriate training sentence for an individ-ual learner at each dialogue turn considering the learning status, such that the learner can have the scores for all pronunciation units exceeding a pre-defined threshold in minimum number of turns. The policy is modeled as a Markov Decision Process (MDP) with high dimensional continuous state space. Experi-ments demonstrate promising results for the approach.
Pei-hao Su, Yow-Bang Wang, Tsung-Hsien Wen, Tien-han Yu, Lin-Shan Lee
INTERSPEECH5
2013 Recurrent neural network based language model personalization by social network crowdsourcing
abstract
Speech recognition has become an important feature in smartphones in recent years. Different from traditional au-tomatic speech recognition, the speech recognition on smart-phones can take advantage of personalized language models to model the linguistic patterns and wording habits of a particu-lar smartphone owner better. Owing to the popularity of social networks in recent years, personal texts and messages are no longer inaccessible. However, data sparseness is still an un-solved problem. In this paper, we propose a three-step adapta-tion approach to personalize recurrent neural network language models (RNNLMs). We believe that its capability to model word histories as distributed representations of arbitrary length can help mitigate the data sparseness problem. Furthermore, we also propose additional user-oriented features to empower the RNNLMs with stronger capabilities for personalization. The experiments on a Facebook dataset showed that the proposed method not only drastically reduced the model perplexity in preliminary experiments, but also moderately reduced the word error rate in n-best rescoring tests.
Tsung-Hsien Wen, Aaron Heidel, Hung-yi Lee, Yu Tsao 0001, Lin-Shan Lee
INTERSPEECH5
2013 Speaking rate normalization with lattice-based context-dependent phoneme duration modeling for personalized speech recognizers on mobile devices
Ching-Feng Yeh, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH3
2013 Model-Based Unsupervised Spoken Term Detection with Spoken Queries
abstract
We present a set of model-based approaches for unsupervised spoken term detection (STD) with spoken queries that requires neither speech recognition nor annotated data. This work shows the possibilities in migrating from DTW-based to model-based approaches for unsupervised STD. The proposed approach consists of three components: self-organizing models, query matching, and query modeling. To construct the self-organizing models, repeated patterns are captured and modeled using acoustic segment models (ASMs). In the query matching phase, a document state matching (DSM) approach is proposed to represent documents as ASM sequences, which are matched to the query frames. In this way, not only do the ASMs better model the signal distributions and time trajectories of speech, but the much-smaller number of states than frames for the documents leads to a much lower computational load. A novel duration-constrained Viterbi (DC-Vite) algorithm is further proposed for the above matching process to handle the speaking rate distortion problem. In the query modeling phase, a pseudo likelihood ratio (PLR) approach is proposed in the pseudo relevance feedback (PRF) framework. A likelihood ratio evaluated with query/anti-query HMMs trained with pseudo relevant/irrelevant examples is used to verify the detected spoken term hypotheses. The proposed framework demonstrates the usefulness of ASMs for STD in zero-resource settings and the potential of an instantly responding STD system using ASM indexing. The best performance is achieved by integrating DTW-based approaches into the rescoring steps in the proposed framework. Experimental results show an absolute 14.2% of mean average precision improvement with 77% CPU time reduction compared with the segmental DTW approach on a Mandarin broadcast news corpus. Consistent improvements were found on TIMIT and MediaEval 2011 Spoken Web Search corpus.
Chun-an Chan, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2013 Enhanced Spoken Term Detection Using Support Vector Machines and Weighted Pseudo Examples
abstract
Spoken term detection (STD) is a key technology for retrieval of spoken content, which will be very important to retrieve and browse multimedia content over the Internet. The discriminative capability of machine learning methods has recently been used to facilitate STD. This paper presents a new approach to improve STD using support vector machines (SVM) based on acoustic information. The concept of pseudo-relevance feedback (PRF) well used in the retrieval of text, image and video is used here. The basic idea of using PRF here is to assume some spoken segments in the first-pass retrieved results are relevant (or pseudo-relevant) and some others irrelevant (or pseudo-irrelevant), and take these segments as positive and negative examples to train a query-specific SVM. This SVM is then used for re-ranking the first-pass retrieved results, and only the re-ranked results are shown to the user. In this paper, feature vectors representing the spoken segments based on acoustic information to be used in SVM are considered and analyzed. Furthermore, conventionally in PRF the items with the highest and lowest scores in the first-pass retrieved results are respectively taken as pseudo-relevant and -irrelevant, but in this way some incorrect examples are inevitably included in the training data especially when the recognition accuracy is poor. Here we further propose an enhanced SVM which not only better selects positive/negative examples considering the reliability of the spoken segments, but emphasizes more on more reliable training examples by modifying the SVM formulation. Experiments on two different sets of spoken archives with different speaking styles and different levels of recognition accuracies demonstrated significant improvements offered by the proposed approaches.
Hung-yi Lee, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2013 An Experimental Analysis on Integrating Multi-Stream Spectro-Temporal, Cepstral and Pitch Information for Mandarin Speech Recognition
abstract
Gabor features have been proposed for extracting spectro-temporal modulation information from speech signals, and have been shown to yield large improvements in recognition accuracy. We use a flexible Tandem system framework that integrates multi-stream information including Gabor, MFCC, and pitch features in various ways, by modeling either or both of the tone and phoneme variations in Mandarin speech recognition. We use either phonemes or tonal phonemes (tonemes) as either the target classes of MLP posterior estimation and/or the acoustic units of HMM recognition. The experiments yield a comprehensive analysis on the contributions to recognition accuracy made by either of the feature sets. We discuss their complementarities in tone, phoneme, and toneme classification. We show that Gabor features are better for recognition of vowels and unvoiced consonants, while MFCCs are better for voiced consonants. Also, Gabor features are capable of capturing changes in signals across time and frequency bands caused by Mandarin tone patterns, while pitch features further offer extra tonal information. This explains why the integration of Gabor, MFCC, and pitch features offers such significant improvements.
Yow-Bang Wang, Shang-Wen Li 0001, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2012 Two-dimensional frame-and-feature weighted Viterbi decoding for robust speech recognition
abstract
In this paper we propose a new approach of two-dimensional frame-and-feature weighted Viterbi decoding performed at the recognizer back-end for robust speech recognition. A new SVM-based frame weighting approach is proposed considering the energy distribution and harmonicity of the frame. The feature weighting is based on a previously proposed approach using an entropy measure considering confusion between phoneme classes. These two different weighting schemes on the two different dimensions are then properly integrated in Viterbi decoding in this paper. Extensive experiments performed with the Aurora 4 testing environment showed significant improvements.
Yang Chang, Lin-Shan Lee
ICASSP2
2012 Unsupervised two-stage keyword extraction from spoken documents by topic coherence and support vector machine
abstract
This paper proposes an unsupervised two-stage approach to automatically extract keywords from spoken documents. In the first stage, for each candidate term we compute a topic coherence and term significance measure (TCS) based on probabilistic latent semantic analysis (PLSA) models. In the second stage, we take the candidate terms with highest and lowest TCS scores as positive and negative examples to train an SVM classifier in an unsupervised way using prosodic, lexical, and semantic features, and then classify the candidate keyword using this SVM classifier. The experiments with course lectures showed that the first-stage offered very good precision, so the second-stage effectively extracted the keywords.
Yun-Nung Chen, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2012 Utterance-level latent topic transition modeling for spoken documents and its application in automatic summarization
abstract
In this paper, we propose to use an utterance-level latent topic transition model to estimate the latent topics behind the utterances, and test the performance of such model in extractive speech summarization. In this model, the latent topic weights behind an utterance are estimated, and these topic weights evolve from an utterance to the next in a spoken document based on a topic transition function represented by a matrix. We explore different ways of obtaining such topic transition matrices used in the model, and find using a set of matrices estimated with utterances clustered from a training spoken document set is very useful. This model was shown to be able to offer extra performance improvement when used with the popularly used Probability Latent Semantic Analysis (PLSA) in preliminary experiments on speech summarization.
Hung-yi Lee, Yun-Nung Chen, Lin-Shan Lee
ICASSP3
2012 Semantic query expansion and context-based discriminative term modeling for spoken document retrieval
abstract
In this paper, we propose a semantic query expansion approach by extending the query-regularized mixture model to include latent topics and apply it to spoken documents. We also propose to use context feature vectors for spoken segments to train SVM models to enhance the posterior-weighted normalized term frequencies in lattices. Experiments on Mandarin broadcast news showed that this approach offered good improvements when applied on spoken documents including relatively high recognition errors.
Tsung-wei Tu, Hung-yi Lee, Yu-Yu Chou, Lin-Shan Lee
ICASSP4
2012 Improved approaches of modeling and detecting Error Patterns with empirical analysis for Computer-Aided Pronunciation Training
abstract
Error pattern detection is very helpful in Computer-Aided Pronunciation Training (CAPT). This paper reports the work of modeling and detecting Error Patterns defined by language teachers based on their linguist knowledge and pedagogical experiences. We develop a model generation framework to create the Error Pattern models from existing phoneme models. We also propose a serial structure for integrating Goodness-of-Pronunciation with the Error Pattern detectors. Experimental results and analysis over different approaches for modeling and detecting Error Patterns are presented, and it is found that both the binary classification error rates and the capability of Error Pattern diagnosis can be improved effectively with the proposed approaches.
Yow-Bang Wang, Lin-Shan Lee
ICASSP2
2012 Recognition of highly imbalanced code-mixed bilingual speech with frame-level language detection based on blurred posteriorgram
abstract
In this work, we proposed a new framework for recognition of highly imbalanced code-mixed bilingual speech using an additional frame-level language detector in the conventional recognition system. Blurred posteriorgram features (BPFs) are also proposed to be used in the language detector. The approach was evaluated with real spontaneous lectures offered at National Taiwan University. The highly imbalanced language distribution in code-mixed speech makes the task difficult. Preliminary experimental results showed not only very good performance improvement, but the improvement is complementary to that brought by better acoustic models, whether due to better adaptation approach or increased training data. The code-mixed bilingual speech is frequently used in the daily lives of many people in the globalized world today.
Ching-Feng Yeh, Aaron Heidel, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2012 Discriminative Fuzzy Clustering Maximum a Posterior Linear Regression for Speaker Adaptation
abstract
We propose a discriminative fuzzy clustering maximum a posterior linear regression (DFCMAPLR) model adaptation approach to compensate the acoustic mismatch due to speaker variability. The DFCMAPLR approach adopts the MAP criterion and a discriminative objective function to estimate shared affine transform and fuzzy weight sets, respectively. Then, through a linear combination of the calculated fuzzy weights and shared affine transforms, more specific affine transforms are formed for model adaptation. By incorporating the MAP criterion and the discriminative information, DFCMAPLR can calculate shared affine transforms reliably and enhance the discriminative power of the adapted acoustic model. Based on the experimental results on the ASTTEL200 Mandarin corpus, we verified that DFCMAPLR outperforms not only the conventional maximum likelihood linear regression (MLLR) but also the fuzzy clustering MLLR(FCMLLR), which estimates the shared affine transform and fuzzy weight sets both based on the maximum likelihood criterion. Moreover, when compared to the baseline result, DFCMAPLR provides a clear improvement of 9.86% (24.04% to 21.67%) relative average phone error rate (PER) reduction.
Ting-Yao Hu, Yu Tsao 0001, Lin-Shan Lee
INTERSPEECH3
2012 Open-Vocabulary Retrieval of Spoken Content with Shorter/Longer Queries Considering Word/Subword-based Acoustic Feature Similarity
Hung-yi Lee, Po-wei Chou, Lin-Shan Lee
INTERSPEECH3
2012 Supervised Spoken Document Summarization jointly Considering Utterance Importance and Redundancy by Structured Support Vector Machine
Hung-yi Lee, Yu-Yu Chou, Yow-Bang Wang, Lin-Shan Lee
INTERSPEECH4
2012 Error Pattern Detection Integrating Generative and Discriminative Learning for Computer-Aided Pronunciation Training
Yow-Bang Wang, Lin-Shan Lee
INTERSPEECH2
2012 Interactive Spoken Content Retrieval with Different Types of Actions Optimized By a Markov Decision Process
abstract
Interaction with user is specially important for spoken con-tent retrieval, not only because of the recognition uncertainty, but because the retrieved spoken content items are difficult to be shown on the screen and difficult to be scanned and selected by the user. The user cannot playback and go through all the retrieved items and then find out they are not what he is look-ing for. In this paper, we propose a new approach for inter-active spoken content retrieval, in which the system can esti-mate the quality of the retrieved results, and take different types of actions to clarify the user’s intention based on an intrinsic policy. The policy is optimized by a Markov Decision Process (MDP) trained with Reinforcement Learning based on a set of pre-defined rewards considering the extra burden given to the user.
Tsung-Hsien Wen, Hung-yi Lee, Lin-Shan Lee
INTERSPEECH3
2012 Improved semantic retrieval of spoken content by language models enhanced with acoustic similarity graph
abstract
Retrieving objects semantically related to the query has been widely studied in text information retrieval. However, when applying the text-based techniques on spoken content, the inevitable recognition errors may seriously degrade the performance. In this paper, we propose to enhance the expected term frequencies estimated from spoken content by acoustic similarity graphs. For each word in the lexicon, a graph is constructed describing acoustic similarity among spoken segments in the archive. Score propagation over the graph helps in estimating the expected term frequencies. The enhanced expected term frequencies can be used in the language modeling retrieval approach, as well as semantic retrieval techniques such as the document expansion based on latent semantic analysis, and query expansion considering both words and latent topic information. Preliminary experiments performed on Mandarin broadcast news indicated that improved performance were achievable under different conditions.
Hung-yi Lee, Tsung-Hsien Wen, Lin-Shan Lee
SLT3
2012 Personalized language modeling by crowd sourcing with social network data for voice access of cloud applications
abstract
Voice access of cloud applications via smartphones is very attractive today, specifically because a smartphones is used by a single user, so personalized acoustic/language models become feasible. In particular, huge quantities of texts are available within the social networks over the Internet with known authors and given relationships, it is possible to train personalized language models because it is reasonable to assume users with those relationships may share some common subject topics, wording habits and linguistic patterns. In this paper, we propose an adaptation framework for building a robust personalized language model by incorporating the texts the target user and other users had posted on the social networks over the Internet to take care of the linguistic mismatch across different users. Experiments on Facebook dataset showed encouraging improvements in terms of both model perplexity and recognition accuracy with proposed approaches considering relationships among users, similarity based on latent topics, and random walk over a user graph.
Tsung-Hsien Wen, Hung-yi Lee, Tai-Yuan Chen, Lin-Shan Lee
SLT4
2012 Integrating Recognition and Retrieval With Relevance Feedback for Spoken Term Detection
abstract
Recognition and retrieval are typically viewed as two cascaded independent modules for spoken term detection (STD). Retrieval techniques are assumed to be applied on top of automatic speech recognition (ASR) output, with performance depending on ASR accuracy. We propose a framework that integrates recognition and retrieval and consider them jointly in order to yield better STD performance. This can be achieved either by adjusting the acoustic model parameters (model-based) or by considering detected examples (example-based) using relevance information provided by the user (user relevance feedback) or inferred by the system (pseudo-relevance feedback), either for a given query (short-term context) or by taking into account many previous queries (long-term context). Such relevance feedback approaches have long been used in text information retrieval, but are rarely considered and cannot be directly applied to the retrieval of spoken content. The proposed relevance feedback approaches are specific to spoken content retrieval and are hence very different from those developed for text retrieval, which are applied only to text symbols. We present not only these relevance feedback scenarios and approaches for STD, but also propose a framework to integrate them all together. Preliminary experiments showed significant improvements in each case.
Hung-yi Lee, Chia-Ping Chen, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2012 Interactive Spoken Document Retrieval With Suggested Key Terms Ranked by a Markov Decision Process
abstract
Interaction with users is a powerful strategy that potentially yields better information retrieval for all types of media, including text, images, and videos. While spoken document retrieval (SDR) is a crucial technology for multimedia access in the network era, it is also more challenging than text information retrieval because of the inevitable recognition errors. It is therefore reasonable to consider interactive functionalities for SDR systems. We propose an interactive SDR approach in which given the user's query, the system returns not only the retrieval results but also a short list of key terms describing distinct topics. The user selects these key terms to expand the query if the retrieval results are not satisfactory. The entire retrieval process is organized around a hierarchy of key terms that define the allowable state transitions; this is modeled by a Markov decision process, which is popularly used in spoken dialogue systems. By reinforcement learning with simulated users, the key terms on the short list are properly ranked such that the retrieval success rate is maximized while the number of interactive steps is minimized. Significant improvements over existing approaches were observed in preliminary experiments performed on information needs provided by real users. A prototype system was also implemented.
Yi-Cheng Pan, Hung-yi Lee, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2012 Modulation Spectrum Equalization for Improved Robust Speech Recognition
abstract
We propose novel approaches for equalizing the modulation spectrum for robust feature extraction in speech recognition. Common to all approaches in that the temporal trajectories of the feature parameters are first transformed into the magnitude modulation spectrum. In spectral histogram equalization (SHE) and two-band spectral histogram equalization (2B-SHE), we equalize the histogram of the modulation spectrum for each utterance to a reference histogram obtained from clean training data, or perform the equalization with two sub-bands on the modulation spectrum. In magnitude ratio equalization (MRE), we define the magnitude ratio of lower to higher modulation frequency components for each utterance, and equalize this to a reference value obtained from clean training data. These approaches can be viewed as temporal filters that are adapted to each testing utterance. Experiments performed on the Aurora 2 and 4 corpora for small and large vocabulary tasks indicate that significant performance improvements are achievable for all noise conditions. We also show that additional improvements can be obtained when these approaches are integrated with cepstral mean and variance normalization (CMVN), histogram equalization (HEQ), higher order cepstral moment normalization (HOCMN), or the advanced front-end (AFE). We analyze and discuss the reasons for these improvements from different viewpoints with different sets of data, including adaptive temporal filtering, noise behavior on the modulation spectrum, phoneme types, and modulation spectrum distance measures.
Liang-Che Sun, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2011 Improved spoken term detection using support vector machines with acoustic and context features from pseudo-relevance feedback
abstract
This paper reports a new approach to improving spoken term detection that uses support vector machine (SVM) with acoustic and linguistic features. As SVM is a good technique for discriminating different features in vector space, we recently proposed to use pseudo-relevance feedback to automatically generate training data for SVM training and use SVM to re-rank the first-pass results considering the context consistency in the lattices. In this paper, we further extend this concept by considering acoustic features at word, phone and HMM state levels and linguistic features of different order. Extensive experiments under various recognition environments demonstrate significant improvements in all cases. In particular, the acoustic features at the HMM state level offered the most significant improvements, and the improvements achieved by acoustic and linguistic features are shown to be additive.
Tsung-wei Tu, Hung-yi Lee, Lin-Shan Lee
ASRU3
2011 Integrating frame-based and segment-based dynamic time warping for unsupervised spoken term detection with spoken queries
abstract
ABSTRACT Rapidly increasing quantities of multimedia and spoken con tent today demand fast and accurate retrieval approaches for con venient browsing. The spoken documents with wide variety of different acoustic and linguistic conditions make supervised training of well-matched acoustic/language models very difficult. Unsuper vised methods using frame-based dynamic time warping (DTW) re quire no acoustic/language models but with high computation load. Therefore, segment-based DTW was proposed to relieve the computation load at the cost of degraded detection performance. In this pa per, we refine the segment-based DTW by allowing deletion of end segments of query to improve detection performance. The search space is also reduced by segment similarity constraints. We also pro posed a two-pass framework. The segment-baed DTW is performed in the first pass to locate hypothesized spoken term region and the frame-based DTW for precise rescoring in the second pass. Then the pseudo relevance feedback is used to expand acoustic variations of the query. We obtain significantly higher detection performance at significantly lower computation load as compared to frame-based DTW.
Chun-an Chan, Lin-Shan Lee
ICASSP2
2011 Improved spoken term detection with graph-based re-ranking in feature space
abstract
This paper presents a graph-based approach for spoken term detection. Each first-pass retrieved utterance is a node on a graph and the edge between two nodes is weighted by the similarity between the two utterances evaluated in feature space. The score of each node is then modified by the contributions from its neighbors by random walk or its modified version, because utterances similar to more utterances with higher scores should be given higher relevance scores. In this way the global similarity structure of all first-pass retrieved utterances can be jointly considered. Experimental results show that this new approach offers significantly better performance than the previously proposed pseudo-relevance feedback approach, which considers primarily the local similarity relationship between first-pass retrieved utterances, and these two different approaches can be cascaded to provide even better results.
Yun-Nung Chen, Chia-Ping Chen, Hung-yi Lee, Chun-an Chan, Lin-Shan Lee
ICASSP5
2011 Improved spoken term detection using support vector machines based on lattice context consistency
abstract
We propose an improved spoken term detection approach that uses support vector machines trained with lattice context consistency. The basic idea is that the same term usually have similar context, while quite different context usually implies the terms are different. Support vector machine can be trained using query context feature vectors obtained from the lattice to estimate better scores for ranking, and significant improvements can be obtained. This process can be performed iteratively and integrated with the pseudo relevance feedback in acoustic feature space proposed previously, both offering further improvements.
Hung-yi Lee, Tsung-wei Tu, Chia-Ping Chen, Chao-Yu Huang, Lin-Shan Lee
ICASSP5
2011 Multi-stream spectro-temporal and cepstral features based on data-driven hierarchical phoneme clusters
abstract
We propose a method to enhance multi-stream Gabor and MFCC features using data-driven hierarchical phoneme clusters to yield more discriminating posteriors. We take into account different hierarchy structures, and in addition perform mean and variance normalization. A relative improvement of 11.5% over the conventional MFCC Tandem system was achieved in experiments conducted on Mandarin broadcast news. We analyze the complementarity between Gabor and MFCC features for different types of phonemes, and investigate the benefits that come from using hierarchical phoneme clusters.
Shang-Wen Li 0001, Liang-Che Sun, Lin-Shan Lee
ICASSP3
2011 Bilingual acoustic modeling with state mapping and three-stage adaptation for transcribing unbalanced code-mixed lectures
abstract
This paper presents a bilingual acoustic modeling approach for transcribing Mandarin-English code-mixed lectures with highly unbalanced language distribution. Special terminologies for the content were produced in the guest language of English (about 15%) and embedded in the utterances produced in the host language of Mandarin (about 85%). The code-mixing nature of the target corpus and the very small percentage of the English data made the task difficult. State mapping and merging approaches plus three stages of model adaptation handles the above problem. Significant improvements in recognition accuracy were obtained in the experiment with a real bilingual code-mixed lecture corpus recorded at National Taiwan University. The code-mixing situation considered is actually very natural in the spoken language of the daily lives of many people in the globalized world today.
Ching-Feng Yeh, Liang-Che Sun, Chao-Yu Huang, Lin-Shan Lee
ICASSP4
2011 Unsupervised Hidden Markov Modeling of Spoken Queries for Spoken Term Detection without Speech Recognition
Chun-an Chan, Lin-Shan Lee
INTERSPEECH2
2011 Spoken Lecture Summarization by Random Walk over a Graph Constructed with Automatically Extracted Key Terms
abstract
This paper proposes an improved approach for spoken lecture summarization, in which random walk is performed on a graph constructed with automatically extracted key terms and proba-bilistic latent semantic analysis (PLSA). Each sentence of the document is represented as a node of the graph and the edge be-tween two nodes is weighted by the topical similarity between the two sentences. The basic idea is that sentences topically similar to more important sentences should be more important. In this way all sentences in the document can be jointly consid-ered more globally rather than individually. Experimental re-sults showed significant improvement in terms of ROUGE eval-uation. Index Terms: summarization, course lecture, probabilistic la-tent semantic analysis (PLSA), random walk, key term
Yun-Nung Chen, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH4
2011 Improved Tonal Language Speech Recognition by Integrating Spectro-Temporal Evidence and Pitch Information with Properly Chosen Tonal Acoustic Units
Shang-Wen Li 0001, Yow-Bang Wang, Liang-Che Sun, Lin-Shan Lee
INTERSPEECH4
2011 Bilingual Acoustic Model Adaptation by Unit Merging on Different Levels and Cross-Level Integration
Ching-Feng Yeh, Chao-Yu Huang, Lin-Shan Lee
INTERSPEECH3
2011 Semantic Analysis and Organization of Spoken Documents Based on Parameters Derived From Latent Topics
abstract
Spoken documents are audio signals and are thus not easily displayed on-screen and not easily scanned and browsed by the user. It is therefore highly desirable to automatically construct summaries, titles, latent topic trees and key term-based topic labels for these spoken documents to aid the user in browsing. We refer to this as semantic analysis and organization. Also, as network content is both copious and dynamic, with topics and domains changing everyday, the approaches here must be primarily unsupervised. We propose a framework for unsupervised semantic analysis and organization of spoken documents and for this purpose propose two measures derived from latent topic analysis: latent topic significance and latent topic entropy. We show that these can be integrated into an application system, with which the user can more easily navigate archives of spoken documents. Probabilistic latent semantic analysis is used as a typical example approach for unsupervised topic analysis in most experiments, although latent Dirichlet allocation is also used in some experiments to show that the proposed measures are equally applicable for different analysis approaches. All of the experiments were performed on Mandarin Chinese broadcast news.
Sheng-yi Kong, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2010 An initial attempt to improve spoken term detection by learning optimal weights for different indexing features
abstract
Because different indexing features actually have different discriminative capabilities for spoken term detection and different levels of reliability in recognition, it is reasonable to weight the indexing features in the transcribed lattices differently during spoken term detection. In this paper, we present an initial attempt of using two weighting schemes, one context independent (fixed weight for each feature) and one context dependent(different weights for the same feature in different context). These weights can be learned by optimizing a desired spoken term detection performance measure over a training document set and a training query set. Encouraging initial results based on unigrams of Chinese characters and syllables for the corpus of Mandarin broadcast news were obtained from the preliminary experiments.
Yu-Hui Chen, Chia-Chen Chou, Hung-yi Lee, Lin-Shan Lee
ICASSP4
2010 Integrating recognition and retrieval with user feedback: A new framework for spoken term detection
abstract
People usually consider recognition and retrieval as two cascaded independent modules for spoken term detection. Retrieval techniques were assumed to be applied on top of some ASR output, with performance depending on ASR accuracy. In this paper, we propose a new framework: to integrate the two parts into a single task. This can be achieved by adjusting the acoustic model parameters, borrowing the principle of Minimum Classification Error (MCE), based on user feedback. The modified acoustic models then give updated posterior probabilities for the lattice-based structures used in spoken term detection. Encouraging results were obtained on a bilingual course lecture corpus in preliminary experiments.
Hung-yi Lee, Lin-Shan Lee
ICASSP2
2010 An initial attempt for phoneme recognition using Structured Support Vector Machine (SVM)
abstract
Structured Support Vector Machine (SVM) is a recently developed extension of the very successful SVM approach, which can efficiently classify structured pattern with maximized margin. This paper presents an initial attempt for phoneme recognition using structured SVM. We simply learn the basic framework of HMMs in configuring the structured SVM. In the preliminary experiments with TIMIT corpus, the proposed approach was able to offer an absolute performance improvement of 1.33% over HMMs even with a highly simplified initial approach, probably because of the concept of maximized margin of SVM. We see the potential of this approach because of the high generality, high flexibility, and high power of structured SVM.
Hao Tang 0002, Chao-Hong Meng, Lin-Shan Lee
ICASSP3
2010 Unsupervised spoken-term detection with spoken queries using segment-based dynamic time warping
Chun-an Chan, Lin-Shan Lee
INTERSPEECH2
2010 Improved spoken term detection by feature space pseudo-relevance feedback
abstract
Abstract In this paper, we propose an improved approach for spokenterm detection using pseudo-relevance feedback. To remedy theproblem of unmatched acoustic models with respect to spokenutterances produced under different acoustic conditions, whichmay give relatively poor recognition output, we integrate therelevance scores derived from the lattices with the DTW dis-tances derived from the feature space of MFCC parametersor phonetic posteriorgrams. These DTW distances are evalu-ated for a carefully selected set of pseudo-relevant utterances,which obtained from the first-pass returned list given by thesearch engine. The utterances on the first-pass returned list arethen reranked accordingly and finally shown to the user. Veryencouraging, performance improvements were obtained in thepreliminary experiments, especially when the acoustic modelsare poorly matched to the spoken utterances.Index Terms: spoken term detection, pseudo-relevance feed-back 1. Introduction Spoken term detection is to return a list of spoken utterancescontaining the term requested by the user. In many approachesof spoken term detection, the spoken utterances are first recog-nized and transformed into transcriptions or lattices by speechrecognition technologies, and then the search engine looksthrough all the transcriptions or lattices very similar to the text-based information retrieval. In this process much of the in-formation in the acoustic signals may be lost in the stage ofspeech recognition, especially when the acoustic models usedare not well matched to the characteristics of the acoustic sig-nals, which naturally results in degraded recognition accuracyand poor detection performance. This is very common in thescenario of spoken term detection, because the huge quantitiesof spoken utterances available over the Internet are naturallyproduced by many different people under many different acous-tic conditions, it is thus very difficult to train a set of acousticmodels well matched to so many different acoustic conditions.As a result, when the relevance scores such as the posteriorprobabilities of the query term derived from transcriptions orlattices are used to rank the retrieved utterances, it is hard tojudge whether a word hypothesis of the query in the transcrip-tions or lattices is a positive target or a false alarm when therecognition output is unreliable. Although many efficient ap-proaches [1, 2, 3] have been proposed to enhance the detectionperformance due to the relatively poor recognition output, thecompensative information straightly from the feature space isnecessary.In text-based information retrieval, even if the texts to beretrieved include all precise words, it is still difficult to retrieveall documents relevant to the query term because many of themdo not include the very short query term entered by the user.However, because many related terms may co-occur in manyrelated documents, a document containing some words appear-ing in some documents identified to be relevant by the searchengine may have high probability to be relevant, even if it doesnot include the query term. For example, a document includingthe words ”George Bush”, ”US”, ”Middle East” may be relevantto a query term of ”White House”, even if it does not includethe query term of ”White House”. In other words, it is possi-ble to retrieve the relevant documents without the query termsince they are ”similar” to some retrieved relevant documentsin some way. Pseudo-relevance feedback, also known as blindrelevance feedback, is one way to realize the above idea. In thisapproach, it is assumed that the set of documents appearing onthe top of the retrieved document list are relevant (or ”pseudo-relevant”), so documents somehow similar to those ”pseudo-relevant” documents can be retrieved, for example, by expand-ing the query with keywords from those ”pseudo-relevant” doc-uments [4]. Similar idea of pseudo-relevance feedback has beenapplied on spoken term detection [5].In this paper, we try to perform similar pseudo-relevancefeedback for spoken term detection as shown in Figure 1. Theupper half of Figure 1 is the conventional spoken term detec-tion. MFCC features were obtained from all spoken utterancesin the archive, speech recognition produces lattices for the ut-terances, and the retrieved engine selects the utterances basedon the relevance scores evaluated from the lattices with respectto the query Qentered by the user. The approach proposedhere in this paper is shown in the lower half of Figure 1. Thefirst-pass returned list is not shown to the user, but instead a”pseudo-relevant utterance set X
Chia-Ping Chen, Hung-yi Lee, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH4
2010 Improved spoken term detection by discriminative training of acoustic models based on user relevance feedback
Hung-yi Lee, Chia-Ping Chen, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH4
2010 Improved phoneme recognition by integrating evidence from spectro-temporal and cepstral features
Shang-Wen Li 0001, Liang-Che Sun, Lin-Shan Lee
INTERSPEECH3
2010 Mandarin tone recognition using affine-invariant prosodic features and tone posteriorgram
Yow-Bang Wang, Lin-Shan Lee
INTERSPEECH2
2010 Automatic key term extraction from spoken course lectures using branching entropy and prosodic/semantic features
abstract
This paper proposes a set of approaches to automatically extract key terms from spoken course lectures including audio signals, ASR transcriptions and slides. We divide the key terms into two types: key phrases and keywords and develop different approaches to extract them in order. We extract key phrases using right/left branching entropy and extract keywords by learning from three sets of features: prosodic features, lexical features and semantic features from Probabilistic Latent Semantic Analysis (PLSA). The learning approaches include an unsupervised method (K-means exemplar) and two supervised ones (AdaBoost and neural network). Very encouraging preliminary results were obtained with a corpus of course lectures, and it is found that all approaches and all sets of features proposed here are useful.
Yun-Nung Chen, Sheng-yi Kong, Lin-Shan Lee
SLT4
2010 A framework integrating different relevance feedback scenarios and approaches for spoken term detection
abstract
This paper presents a new framework integrating different relevance feedback scenarios (pseudo relevance feedback and user relevance feedback in short- and long-term context) and different approaches (model- and example-based) in a spoken term detection system, and shows the retrieval performance can be improved step by step. It is found that short-term context user relevance feedback can further improve the retrieval performance after pseudo relevance feedback, regardless of whether the acoustic models have been adapted by matched data or long-term context user relevance feedback or not. Moreover, model-based and example-based methods are shown to be additive when integrated in short-term context user relevance feedback scenario.
Hung-yi Lee, Chia-Ping Chen, Ching-Feng Yeh, Lin-Shan Lee
SLT4
2010 Performance Analysis for Lattice-Based Speech Indexing Approaches Using Words and Subword Units
abstract
Lattice-based speech indexing approaches are attractive for the combination of short spoken segments, short queries, and low automatic speech recognition (ASR) accuracies, as lattices provide recognition alternatives and therefore tend to compensate for recognition errors. Position-specific posterior lattices (PSPLs) and confusion networks (CNs), two of the most popular lattice-based approaches, both reduce disk space requirements and are more efficient than raw lattices. When PSPLs and CNs are used in a word-based fashion, they cannot handle OOV or rare word queries. In this paper, we propose an efficient approach for the construction of subword-based PSPLs (S-PSPLs) and CNs (S-CNs) and present a comprehensive performance analysis of PSPL and CN structures using both words and subword units, taking into account basic principles and structures, and supported by experimental results on Mandarin Chinese. S-PSPLs and S-CNs are shown to yield significant mean average precision (MAP) improvements over word-based PSPLs and CNs for both out-of-vocabulary (OOV) and in-vocabulary queries while requiring much less disk space for indexing.
Yi-Cheng Pan, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2009 Discriminative Lexicon Adaptation for Improved Character Accuracy - A New Direction in Chinese Language Modeling
Yi-Cheng Pan, Lin-Shan Lee, Sadaoki Furui
ACL/IJCNLP2
2009 Voice-based information retrieval - how far are we from the text-based information retrieval ?
abstract
Although network content access is primarily text-based today, almost all roles of text can be accomplished by voice. Voice-based information retrieval refers to the situation that the user query and/or the content to be retried are in form of voice. This paper tries to compare the voice-based information retrieval with the currently very successful text-based information retrieval, and identifies two major issues in which voice-based information retrieval is far behind: retrieval accuracy and user-system interaction. These two issues are reviewed, analyzed and discussed in detail. It is found that very good approaches have been proposed and very good improvements have been achieved, although there is still a very long way to go. A few successful prototype systems, among many others are presented at the end.
Lin-Shan Lee, Yi-Cheng Pan
ASRU1
2009 Spoken term detection from bilingual spontaneous speech using code-switched lattice-based structures for words and subword units
abstract
This paper presents the first work known publicly on spoken term detection from bilingual spontaneous speech using code-switched lattice-based structures for word and subword units. The corpus used is the lectures with Chinese as the host language and English as the guest language recorded for a real course offered in National Taiwan University. The techniques reported here have been successfully implemented and tested in a real lecture system now available on-line over the Internet. We also present the approaches of using word fragment as the subword unit for English, and analyse the difficult issues when code-switched lattice-based structures for subword units are used for tasks involving languages of quite different natures.
Hung-yi Lee, Yueh-Lien Tang, Hao Tang 0002, Lin-Shan Lee
ASRU4
2009 Improved clustered hierarchical tandem system with bottom-up processing
abstract
The outputs of multi-layer perceptron (MLP) classifiers have been successfully used in tandem systems as features for HMM-based automatic speech recognition. In a previous paper, we proposed data-driven clustered hierarchical MLP (CHMLP) tandem system yielding improved performance by dividing the complicated global phone classification problem into simpler hierarchical tasks, in which specialized MLPs are trained to classify small clusters of confusing phones in a hierarchical structure. In this paper a bottom-up processing is further proposed to enhance the classification in the above CHMLP and offer even better performance. MLP rescoring for the tandem system is also investigated. The best result achieved 19.1% relative error reduction over the MFCC baseline.
Shuo-Yiin Chang, Lin-Shan Lee
ICASSP2
2009 Latent semantic retrieval of personal photos with sparse user annotation by fused image/speech/text features
abstract
While users prefer high-level semantic photo descriptions (e.g., who, what, when, where), we wish to minimize the need to annotate photos using such descriptions by the user. We propose a latent semantic personal photo retrieval approach using fused image/speech/text features. We use low-level image features to derive relatoionships among sparsely annotated photos, and probabilistic latent semantic analysis (PLSA) models based on fused image/speech/text features to analyze photo “topics”. We then retrieve the photos using text or speech queries of simple high-level semantic words only. In preliminary experiments, while only 10% of the photos were manually annotated, the photos could be well retrieved with very encouraging results.
Yi-Sheng Fu, Chia-Yu Wan, Lin-Shan Lee
ICASSP3
2009 Learning on demand - course lecture distillation by information extraction and semantic structuring for spoken documents
abstract
This paper presents a new approach of organizing the course lectures (as spoken documents) for efficient learning on demand by the users. By the properly matching the course lectures with the slides used, we divide the course lectures into hierarchical ldquomajor segmentsrdquo with variable length based on the topics discussed. Key term extraction, hierarchical summarization and semantic structuring are then performed over these ldquomajor segmentsrdquo. A key term graph is also constructed, based on which the various major segments of the course can be linked. In this way, the user can ask questions to the system, and develop his own road map of learning the knowledge he needs considering his available time and his background knowledge, based on the semantic structure provided by the system. A preliminary prototype system has been successfully developed with encouraging initial test results.
Sheng-yi Kong, Miao-ru Wu, Che-Kuang Lin, Yi-Sheng Fu, Lin-Shan Lee
ICASSP5
2009 Improved lattice-based spoken document retrieval by directly learning from the evaluation measures
abstract
Lattice-based approaches have been widely used in spoken document retrieval to handle the speech recognition uncertainty and errors. Position Specific Posterior Lattices (PSPL) and Confusion Network (CN) are good examples. It is therefore interesting to derive improved model for spoken document retrieval by properly integrating different versions of lattice-based approaches in order to achieve better performance. In this paper we borrow the framework of dasialearning to rankpsila from text document retrieval and try to integrate it into the scenario of lattice-based spoken document retrieval. Two approaches are considered here, AdaRank and SVM-map. With these approaches, we are able to learn and derived improved models using different versions of PSPL/CN. Preliminary experiments with broadcast news in Mandarin Chinese showed significant improvements.
Chao-Hong Meng, Hung-yi Lee, Lin-Shan Lee
ICASSP3
2009 Mandarin spontaneous narrative planning - prosodic evidence from national taiwan university lecture corpus
abstract
This paper discusses discourse planning of pre-organized spontaneous narratives (SpnNS) in comparison with read speech (RS). F0 and tempo modulations are compared by speech paragraph size and discourse boundaries. The speaking rate of SpnNS from university classroom lecture is 2 to 3 times to that of RS by professionals; paragraph phrasing of SpnNS is 6 times that of RS. Patterns of paragraph association are distinct for SpnNS and RS. Sub-paragraph and paragraph units in RS are marked by distinct relative F0 resets and boundary pause duration, but by patterns of intensity contrasts in SpnNS instead. Consistent to both data sets is the finding that combined relative supra-segmental cues reflecting global prosodic properties are more discriminative to distinguish discourse boundaries than any fragments of singular cue, supporting higher-level discourse planning in the acoustic signals. We believe these findings can be directly applied to speech technology development.
Chiu-yu Tseng, Zhao-yu Su, Lin-Shan Lee
INTERSPEECH3
2009 Higher Order Cepstral Moment Normalization for Improved Robust Speech Recognition
abstract
Cepstral normalization has widely been used as a powerful approach to produce robust features for speech recognition. Good examples of this approach include cepstral mean subtraction, and cepstral mean and variance normalization, in which either the first or both the first and the second moments of the Mel-frequency cepstral coefficients (MFCCs) are normalized. In this paper, we propose the family of higher order cepstral moment normalization, in which the MFCC parameters are normalized with respect to a few moments of orders higher than 1 or 2. The basic idea is that the higher order moments are more dominated by samples with larger values, which are very likely the primary sources of the asymmetry and abnormal flatness or tail size of the parameter distributions. Normalization with respect to these moments therefore puts more emphasis on these signal components and constrains the distributions to be more symmetric with more reasonable flatness and tail size. The fundamental principles behind this approach are also analyzed and discussed based on the statistical properties of the distributions of the MFCC parameters. Experimental results based on the AURORA 2, AURORA 3, AURORA 4, and Resource Management (RM) testing environments show that with the proposed approach, recognition accuracy can be significantly and consistently improved for all types of noise and all SNR conditions.
Chang-Wen Hsu, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2009 Improved Features and Models for Detecting Edit Disfluencies in Transcribing Spontaneous Mandarin Speech
abstract
Detection of edit disfluencies is key to transcribing spontaneous utterances. In this paper, we present improved features and models to detect edit disfluencies and enhance transcription of spontaneous Mandarin speech using hypothesized disfluency interruption points (IPs) and edit word detection. A comprehensive set of prosodic features that takes into account the special characteristics of edit disfluencies in Mandarin is developed, and an improved model combining decision trees and maximum entropy is proposed to detect IPs. This model is further adapted to desired prosodic conditions by latent prosodic modeling, a probabilistic framework for analyzing speech prosody in terms of a set of latent prosodic states. These techniques contribute to higher recognition accuracy (by rescoring with the hypothesized IPs) and better edit word detection (using conditional random fields defined on Chinese characters) in the final transcription, as verified by experiments on a spontaneous Mandarin speech corpus.
Che-Kuang Lin, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2008 Context dependent quantization for distributed and/or robust speech recognition
abstract
It is well-known that the high correlation existing in speech signals is very helpful in various speech processing applications. In this paper, we propose a new concept of context-dependent quantization, in which the representative parameter (whether a scalar or a vector) for a quantization partition cell is not fixed, but depends on the signal context on both sides, and the signal context dependencies can be trained with a clean speech corpus or estimated from a noisy speech corpus. This results in a much finer quantization based on local signal characteristics, without using any extra bit rate. This approach is equally applicable to all (scalar or vector) quantization approaches, and can be used either for signal compression in distributed speech recognition (DSR) or for feature transformation in robust speech recognition. In the latter case, each feature parameter is simply transformed into its representative parameter after quantization. In preliminary experiments with AURORA 2 and simulated GPRS channels, this concept is integrated with a recently proposed Histogram-based Quantization (HQ), the partition cells of which are also dynamic depending on local signal statistics. Significant performance improvements were obtained with the presence of both environmental noise and transmission errors.
Chia-Yu Wan, Lin-Shan Lee
ICASSP3
2008 Data-driven clustered hierarchical tandem system for LVCSR
abstract
This paper investigates the use of a hierarchy of Neural Networks for performing data driven feature extraction.Two different hierarchical structures based on long and short temporal context are considered.Features are tested on two different LVCSR systems for Meetings data (RT05 evaluation data) and for Arabic Broadcast News (BNAT05 evaluation data).The hierarchical NNs outperforms the single NN features consistently on different type of data and tasks and provides significant improvements w.r.t.respective baselines systems.Best results are obtained when different time resolutions are used at different level of the hierarchy.
Shuo-Yiin Chang, Lin-Shan Lee
INTERSPEECH2
2008 Confusion-based entropy-weighted decoding for robust speech recognition
Chia-Yu Wan, Lin-Shan Lee
INTERSPEECH3
2008 Improved large vocabulary Mandarin speech recognition by selectively using tone information with a two-stage prosodic model
Lin-Shan Lee
INTERSPEECH2
2008 Evaluation of modulation spectrum equalization techniques for large vocabulary robust speech recognition
Liang-Che Sun, Chang-Wen Hsu, Lin-Shan Lee
INTERSPEECH3
2008 Latent semantic retrieval of spoken documents over position specific posterior lattices
abstract
This paper presents a new approach of latent semantic retrieval of spoken documents over Position Specific Posterior Lattices (PSPL). This approach performs concept matching instead of literal term matching during retrieval based on the Probabilistic Latent Semantic Analysis (PLSA), so as to solve the problem of term mismatch between the query and the desired spoken documents. This approach is performed over PSPL to consider the multiple hypotheses generated by ASR process, as well as the position information for these hypotheses, so as to alleviate the problem of relatively poor ASR accuracy. We establish a framework to evaluate semantic relevance between terms and the relevance score between a query and a PSPL, both based on the latent topic information from PLSA. Preliminary experiments on Chinese broadcast news segments showed significant improvements can be obtained with the proposed approach.
Hung-lin Chang, Yi-Cheng Pan, Lin-Shan Lee
SLT3
2008 Automatic title generation for Chinese spoken documents with a delicate scored Viterbi algorithm
abstract
Automatic title generation for spoken documents is believed to be an important key for browsing and navigation over huge quantities of multimedia content. A new framework of automatic title generation for Chinese spoken documents is proposed in this paper using a delicate scored Viterbi algorithm performed over automatically generated text summaries of the testing spoken documents. The Viterbi beam search is guided by a delicate score evaluated from three sets of models: term selection model tells the most suitable terms to be included in the title, term ordering model gives the best ordering of the terms to make the title readable, and title length model tells the reasonable length of the title. The models are trained from a training corpus which is not required to be matched with the testing spoken documents. Both objective evaluation based on F1 measure and subjective human evaluation for relevance and readability indicated the approach is very attractive.
Sheng-yi Kong, Chien-Chih Wang, Ko-chien Kuo, Lin-Shan Lee
SLT4
2008 Robustness analysis on lattice-based speech indexing approaches with respect to varying recognition accuracies by refined simulations
abstract
We analyze the robustness of different lattice-based speech indexing approaches. While we believe such analysis is important, to our knowledge it has been neglected in prior works. In order to make up for the lack of corpora with various noise characteristics, we use refined approaches to simulate feature vector sequences directly from HMMs, including those with a wide range of recognition accuracies, as opposed to simply adding noise and channel distortion to the existing noisy corpora. We compare, analyze, and discuss the robustness of several state-of-the-art speech indexing approaches.
Yi-Cheng Pan, Hung-lin Chang, Lin-Shan Lee
SLT3
2008 Histogram-Based Quantization for Robust and/or Distributed Speech Recognition
abstract
In a distributed speech recognition (DSR) framework, the speech features are quantized and compressed at the client and recognized at the server. However, recognition accuracy is degraded by environmental noise at the input, quantization distortion, and transmission errors. In this paper, histogram-based quantization (HQ) is proposed, in which the partition cells for quantization are dynamically defined by the histogram or order statistics of a segment of the most recent past values of the parameter to be quantized. This scheme is shown to be able to solve to a good degree many problems related to DSR. A joint uncertainty decoding (JUD) approach is further developed to consider the uncertainty caused by both environmental noise and quantization errors. A three-stage error concealment (EC) framework is also developed to handle transmission errors. The proposed HQ is shown to be an attractive feature transformation approach for robust speech recognition outside of a DSR environment as well. All the claims have been verified by experiments using the Aurora 2 testing environment, and significant performance improvements for both robust and/or distributed speech recognition over conventional approaches have been achieved.
Chia-Yu Wan, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2007 Robust speech recognition by properly utilizing reliable frames and segments in corrupted signals
abstract
In this paper, we propose a new approach to detecting and utilizing reliable frames and segments in corrupted signals for robust speech recognition. Novel approaches to estimating an energy-based measure and a harmonicity measure for each frame are developed. SNR-dependent GMM Classifiers are then trained, together with a Reliable Frame Selection and Clustering module and a Reliable Segment Identification module, to detect the most reliable frames in an utterance. These reliable frames and segments thus obtained can be properly used in both front-end feature enhancement and back-end Viterbi decoding. In the extensive experiments reported here, very significant improvements in recognition accuracies were obtained with the proposed approaches for all types of noise and all SNR values defined in the Aurora 2 database.
Chia-Yu Wan, Lin-Shan Lee
ASRU3
2007 Robust topic inference for latent semantic language model adaptation
abstract
We perform topic-based, unsupervised language model adaptation under an N-best rescoring framework by using previous-pass system hypotheses to infer a topic mixture which is used to select topic-dependent LMs for interpolation with a topici-ndependent LM. Our primary focus is on techniques for improving the robustness of topic inference for a given utterance with respect to recognition errors, including the use of ASR confidence and contextual information from surrounding utterances. We describe a novel application of metadata-based pseudo-story segmentation to language model adaptation, and present good improvements to character error rate on multigenre GALE Project data in Mandarin Chinese.
Aaron Heidel, Lin-Shan Lee
ASRU2
2007 Analytical comparison between position specific posterior lattices and confusion networks based on words and subword units for spoken document indexing
abstract
In this paper we analytically compare the two widely accepted approaches of spoken document indexing, Position Specific Posterior Lattices (PSPL) and Confusion Network (CN), in terms of retrieval accuracy and index size. The fundamental distinctions between these two approaches in terms of construction units, posterior probabilities, number of clusters, indexing coverage and space requirements are discussed in detail. A new approach to approximate subword posterior probability in a word lattice is also incorporated in PSPL/CN to handle OOV/rare word problems, which were unaddressed in original PSPL and CN approaches. Extensive experimental results on Chinese broadcast news segments indicate that PSPL offers higher accuracy than CN but requiring much larger disk space, while subword-based PSPL turns out to be very attractive because it lowers the storage cost while offers even higher accuracies.
Yi-Cheng Pan, Hung-lin Chang, Lin-Shan Lee
ASRU3
2007 Type-II dialogue systems for information access from unstructured knowledge sources
abstract
In this paper, we present a new formulation and a new framework for a new type of dialogue system, referred to as the type-II dialogue systems in this paper. The distinct feature of such dialogue systems is their tasks of information access from unstructured knowledge sources, or the lack of a well-organized back-end database offering the information for the user. Typical example tasks of this type of dialogue systems include retrieval, browsing and question answering. The mainstream dialogue systems with a well-organized back-end database are then referred to as type-I dialogue systems here in the paper. The functionalities of each module in such type-II dialogue systems are analyzed, presented, and compared with the respective modules in type-I dialogue systems. A preliminary type-II dialogue system recently developed in National Taiwan University is also presented at the end as a typical example.
Yi-Cheng Pan, Lin-Shan Lee
ASRU2
2007 Modulation spectrum equalization for robust speech recognition
abstract
Two approaches for modulation spectrum equalization are proposed for robust feature extraction in speech recognition. In both cases the temporal trajectories of the feature parameters are first transformed into the modulation spectrum. In the spectral histogram equalization (SHE) approach, we equalize the histogram of the modulation spectrum for each utterance to a reference histogram obtained from clean training data. In the magnitude ratio equalization (MRE) approach, we equalize the magnitude ratio of lower to higher frequency components on the modulation spectrum to a reference value also obtained from clean training data. Preliminary experimental results performed on the AURORA 2 testing environment indicate that significant performance improvements are achievable with these approaches, when integrated with cepstral mean and variance normalization (CMVN), for all testing sets A, B, and C, all types of noise, for all SNR values. We also show that the approach of magnitude ratio equalization (MRE) offers additional performance improvements when integrated with other more advanced feature normalization approaches such as histogram equalization (HEQ) and higher-order cepstral moment normalization (HOCMN).
Liang-Che Sun, Chang-Wen Hsu, Lin-Shan Lee
ASRU3
2007 Pronunciation Modeling for Spontaneous Speech Recognition using Latent Pronunciation Analysis (LPA) and Prior Knowledge
abstract
In this paper, we propose a new framework for pronunciation modeling, in which the search algorithm tries to focus primarily on the clearly-pronounced portion of speech, while deemphasizing the observations of the slurred portion. This is based on the prior analysis that the pronunciation variation has to do with the predictability and the importance of the words in the spoken utterances, which may be estimated to some extent. We define a set of pronunciation-related features and develop a latent pronunciation analysis (LPA) to estimate the "latent pronunciation states" in the speech. The LPA probabilities, pronunciation-related features and another set of prior knowledge obtained from two distance measures between phonemes are integrated in a SVM classifier to produce a "pronunciation variation indicator" for each frame, based on which the Viterbi decoding was performed. Very encouraging initial results on Mandarin spontaneous speech were obtained in preliminary experiments.
Che-Kuang Lin, Lin-Shan Lee
ICASSP (4)2
2007 Three-Stage Error Concealment for Distributed Speech Recognition (DSR) with Histogram-Based Quantization (HQ) Under Noisy Environment
abstract
In this paper, a three-stage error concealment (EC) framework based on the recently proposed histogram-based quantization (HQ) for distributed speech recognition (DSR) is proposed, in which noisy input speech is assumed and both the transmission errors and environmental noise are considered jointly. The first stage detects the erroneous feature parameters at both the frame and subvector levels. The second stage then reconstructs the detected erroneous subvectors by MAP estimation, considering the prior speech source statistics, the channel transition probability, and the reliability of the received subvectors. The third stage then considers the uncertainty of the estimated vectors during Viterbi decoding. At each stage, the error concealment (EC) techniques properly exploit the inherent robust nature of histogram-based quantization (HQ). Extensive experiments with AURORA 2.0 testing environment and GPRS simulation indicated the proposed framework is able to offer significantly improved performance against a wide variety of environmental noise and transmission error conditions.
Chia-Yu Wan, Lin-Shan Lee
ICASSP (4)3
2007 Virtual Conduction System with Multi-Resolution Wall Display
abstract
The virtual conduction system (VCS) allows a user to conduct a photo-realistic pseudo orchestra. The VCS includes four modules: (1) gesture recognition module, (2) audio rendering module, (3) video rendering module, and (4) multi-resolution display module. With gesture recognition and tempo adjustment, the user not only can change the playback rate of an audio and video recording, but also can control the volume from different portion of an orchestra in real time. With the video rendering module and multi-resolution display module, the VCS provides a new visual experience with an interactive multi-resolution wall-size display.
Wei-Ting Peng, En-Wei Huang, Wei-Lun Chang, Po-Chung Huang, Jun-Ying Bai, Han-Ru Chen, Shao-Yi Chien, Shyh-Kang Jeng, Yi-Ping Hung, Li-Chen Fu, Lin-Shan Lee
ICME11
2007 Language model adaptation using latent dirichlet allocation and an efficient topic inference algorithm
abstract
We present an effort to perform topic mixture-based language model adaptation using latent Dirichlet allocation (LDA).We use probabilistic latent semantic analysis (PLSA) to automatically cluster a heterogeneous training corpus, and train an LDA model using the resultant topicdocument assignments.Using this LDA model, we then construct topic-specific corpora at the utterance level for interpolation with a background language model during language model adaptation.We also present a novel iterative algorithm for LDA topic inference.Very encouraging results were obtained in preliminary experiments with broadcast news in Mandarin Chinese.
Aaron Heidel, Hung-An Chang, Lin-Shan Lee
INTERSPEECH3
2007 Extended powered cepstral normalization (p-CN) with range equalization for robust features in speech recognition
Chang-Wen Hsu, Lin-Shan Lee
INTERSPEECH2
2007 Subword-based position specific posterior lattices (s-PSPL) for indexing speech information
Yi-Cheng Pan, Hung-lin Chang, Berlin Chen, Lin-Shan Lee
INTERSPEECH4
2007 Lexicon adaptation with reduced character error (LARCE) - a new direction in Chinese language modeling
abstract
Good language modeling relies on good predefined lexicons. For Chinese, since there are no text word boundaries and the concept of “word ” is not very well defined, constructing good lexicons is difficult. In this paper, we propose lexicon adapta-tion with reduced character error (LARCE), which learns new word tokens based on the criterion of reduced adaptation cor-pus error rate. In this approach, a multi-character string is taken as a new “word ” as long as it is helpful in reducing the er-ror rate, and minimum number of new, high-quality words can be obtained. This algorithm is based on character-based con-sensus networks. In initial experiments on Chinese broadcast news, it is shown that LARCE not only significantly outper-forms PAT-tree-based word extraction algorithms, but even out-performs manually augmented lexicons. It is believed the con-cept is equally useful for other character-based languages.
Yi-Cheng Pan, Lin-Shan Lee
INTERSPEECH2
2007 A Perceptually Constrained GSVD-Based Approach for Enhancing Speech Corrupted by Colored Noise
abstract
The singular value decomposition (SVD)-based method for single-channel speech enhancement has been shown to be very useful when the additive noise is white. For colored noise, with this approach, one needs to whiten the noise spectrum prior to SVD-based approach and perform the inverse whitening processing afterwards. A truncated quotient SVD (QSVD)-based approach has been proposed to handle this problem and found very useful. In this paper, a generalized SVD (GSVD)-based subspace approach for speech enhancement is first extended from the concept of the truncated QSVD-based approach, in which the dimension of the signal subspace can be precisely and automatically determined for each frame of the noisy signal. But with this new approach some residual noise is still perceivable under lower signal-to-noise ratio conditions. Therefore a perceptually constrained GSVD (PCGSVD)-based approach is further proposed to incorporate the masking properties of human auditory system to make sure the undesired residual noise to be nearly un-perceivable. Closed-form solutions are obtained for both the GSVD- and PCGSVD-based enhancement approaches. Very carefully performed objective evaluations and subjective listening tests show that the PCGSVD-based approach proposed here can offer improved speech quality, intelligibility and recognition accuracy, whether the noise is stationary or nonstationary, especially when the additive noise is nonwhite
Gwo-hwa Ju, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2006 Entropy-Based Feature Parameter Weighting for Robust Speech Recognition
abstract
In this work, we propose an entropy-based measure to determine the discriminating ability of a feature parameter in identifying the correct acoustic models, and a feature parameter weighting scheme using this measure during Viterbi decoding. The purpose is to emphasize the scores obtained with more discriminating parameters, and to de-emphasize the scores with less discriminating parameters. Extensive experiments verified that this approach is equally useful for different types of features, and can be easily integrated with typical existing robust speech recognition approaches.
Chia-Yu Wan, Lin-Shan Lee
ICASSP (1)3
2006 Improved Spoken Document Retrieval With Dynamic Key Term Lexicon and Probabilistic Latent Semantic Analysis (PLSA)
abstract
Spoken document retrieval will be very important in the future network era. In this paper, we propose using a "dynamic key term lexicon" automatically extracted from the ever-changing document archives as an extra feature set in the retrieval task. This lexicon is much more compact but semantically rich, thus it can retrieve relevant documents more efficiently. The key terms include named entities and others selected by a new metric referred to as the term entropy here derived from probabilistic latent semantic analysis (PLSA). Various configurations of retrieval models were tested with a broadcast news archive in Mandarin Chinese and significant performance improvements were obtained, especially with the new version of PLSA models based on a key term lexicon rather than the full lexicon.
Ya-chao Hsieh, Yu-tsun Huang, Chien-Chih Wang, Lin-Shan Lee
ICASSP (1)4
2006 Improved Spoken Document Summarization Using Probabilistic Latent Semantic Analysis (PLSA)
abstract
In this paper we propose a set of new methods exploring the topical information embedded in the spoken documents and using such information in automatic summarization of spoken documents. By introducing a set of latent topic variables, Probabilistic Latent Semantic Analysis (PLSA) is useful to find the underlying probabilistic relationships between documents and terms. Two useful measures, referred to as topic significance and term entropy in this paper, are proposed based on the PLSA modeling to determine the terms and thus sentences important for the document which can then be used to construct the summary. Experiment results for preliminary tests performed on broadcast news stories in Mandarin Chinese indicated improved performance as compared to some existing approaches.
Sheng-yi Kong, Lin-Shan Lee
ICASSP (1)2
2006 Joint Uncertainty Decoding (JUD) with Histogram-Based Quantization (HQ) for Robust and/or Distributed Speech Recognition
abstract
Histogram-based Quantization (HQ) has been recently proposed as a robust and scalable quantization approach for Distributed Speech Recognition (DSR). In this paper, Histogram-based Quantization (HQ) is further verified as an attractive feature transformation approach for robust speech recognition, Joint Uncertainty Decoding (JUD) is developed to be applied with HQ for improved recognition accuracy, and the approach was evaluated for both cases of robust speech recognition and DSR. In Joint Uncertainty Decoding (JUD), we jointly consider and estimate the uncertainty caused by both the environmental noise and the quantization errors in Viterbi decoding under the framework of HQ. For robust speech recognition, HQ was used as the front-end feature transformation and JUD as the enhancement approach at the back-end recognizer. For DSR, HQ was applied at the client end as a data compression process and JUD at the server. The evaluation with Aurora 2.0 testing environment showed very significant improvements for both cases of robust and/or distributed speech recognition.
Chia-Yu Wan, Lin-Shan Lee
ICASSP (1)2
2006 A new framework for system combination based on integrated hypothesis space
abstract
In this paper, a new concept of integrated hypothesis space for large vocabulary continuous speech recognition (LVCSR) system combination is proposed. Unlike the conventional systems combination approaches such as ROVER, the hypothesis spaces are directly integrated here without string alignment. In this way the timing information for all word hypotheses is well preserved and the new framework is more flexible on rescoring approaches used. Four rescoring criteria on the integrated hypothesis space were further explored and experiments on Chinese broadcast news corpus indicated improved performance.
I-Fan Chen, Lin-Shan Lee
INTERSPEECH2
2006 Extension and further analysis of higher order cepstral moment normalization (HOCMN) for robust features in speech recognition
Chang-Wen Hsu, Lin-Shan Lee
INTERSPEECH2
2006 Powered cepstral normalization (p-CN) for robust features in speech recognition
Chang-Wen Hsu, Lin-Shan Lee
INTERSPEECH2
2006 Prosodic modeling in large vocabulary Mandarin speech recognition
Jui-Ting Huang, Lin-Shan Lee
INTERSPEECH2
2006 Feature analysis for emotion recognition from Mandarin speech considering the special characteristics of Chinese language
abstract
Emotion recognition from speech signals is regarded as a critical step toward intelligent human-machine interface.However, feature parameters useful for this purpose may have to do with the special structures of the language.In this paper we present a detailed analysis of the feature parameters for emotion recognition considering the characteristics of the Chinese language, primarily the monosyllable structure and the tone behavior.The analysis is based on the feature parameters on three levels: frame-level, syllable-level, and word-level.The results show that the frame-level and syllable-level ones are good indicators, while taking the ensemble features on all three levels can yield a recognition accuracy of 90.0%.We also found that the pitch and power related features are the most important, and the fourth tone in Mandarin serves as the strongest indicator to emotions.All these findings are consistent with the characteristics of Mandarin Chinese.
Yi-Hao Kao, Lin-Shan Lee
INTERSPEECH2
2006 Multi-layered summarization of spoken document archives by information extraction and semantic structuring
Lin-Shan Lee, Sheng-yi Kong, Yi-Cheng Pan, Yi-Sheng Fu, Yu-tsun Huang
INTERSPEECH1
2006 Latent prosodic modeling (LPM) for speech with applications in recognizing spontaneous Mandarin speech with disfluencies
Che-Kuang Lin, Lin-Shan Lee
INTERSPEECH2
2006 Efficient interactive retrieval of spoken documents with key terms ranked by reinforcement learning
abstract
Unlike written documents, spoken documents are difficult to display on the screen; it is also difficult for users to browse these documents during retrieval.It has been proposed recently to use interactive multi-modal dialogues to help the user navigate through a spoken document archive to retrieve the desired documents.This interaction is based on a topic hierarchy constructed by the key terms extracted from the retrieved spoken documents.In this paper, the efficiency of the user interaction in such a system is further improved by a key term ranking algorithm using Reinforcement Learning with simulated users.Significant improvements in retrieval efficiency, which are relatively robust to the speech recognition errors, are observed in preliminary evaluations.
Yi-Cheng Pan, Yen-shin Lee, Yi-Sheng Fu, Lin-Shan Lee
INTERSPEECH5
2006 Improved Summarization of Chinese spoken Documents by Probabilistic Latent Semantic Analysis (PLSA) with Further Analysis and Integrated Scoring
abstract
In a previous paper [1] two new scoring measures, topic significance (TS) and topic entropy (TE), obtained from probabilistic latent semantic analysis (PLSA) were shown to outperform very successful baseline significance score (SS) in selecting the important sentences for summarization of spoken documents. In this paper extensive experiments using the ROUGE scores with respect to different parameters at different summarization ratios were carefully analyzed in great detail. It was also found that integration of these two scoring measures offered further improvements, and special considerations of the structure of Chinese language was also helpful when summarizing Chinese spoken documents.
Sheng-yi Kong, Lin-Shan Lee
SLT2
2006 Simulation Analysis for Interactive Retrieval of spoken Documents with Key Terms Ranked by Reinforcement Learning
abstract
Unlike written documents, spoken documents are difficult to display on the screen; it is also difficult for users to browse these documents during retrieval. It has been proposed recently to use interactive multi-modal dialogues to help the user navigate through a spoken document archive to retrieve the desired documents. This interaction is based on a topic hierarchy constructed by the key terms extracted from the retrieved spoken documents. In this paper, the efficiency of the user interaction in such a system is further improved by a key term ranking algorithm using reinforcement learning with simulated users. Extensive simulation analysis was performed, and significant improvements in retrieval efficiency were observed. These improvements show the relative robustness to speech recognition errors.
Yi-Cheng Pan, Lin-Shan Lee
SLT2
2006 Optimization of temporal filters for constructing robust features in speech recognition
abstract
Linear discriminant analysis (LDA) has long been used to derive data-driven temporal filters in order to improve the robustness of speech features used in speech recognition. In this paper, we proposed the use of new optimization criteria of principal component analysis (PCA) and the minimum classification error (MCE) for constructing the temporal filters. Detailed comparative performance analysis for the features obtained using the three optimization criteria, LDA, PCA, and MCE, with various types of noise and a wide range of SNR values is presented. It was found that the new criteria lead to superior performance over the original MFCC features, just as LDA-derived filters can. In addition, the newly proposed MCE-derived filters can often do better than the LDA-derived filters. Also, it is shown that further performance improvements are achievable if any of these LDA/PCA/MCE-derived filters are integrated with the conventional approach of cepstral mean and variance normalization (CMVN). The performance improvements obtained in recognition experiments are further supported by analyses conducted using two different distance measures.
Jeih-Weih Hung, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2005 Energy-based frame selection for reliable feature normalization and transformation in robust speech recognition
Lin-Shan Lee
INTERSPEECH2
2005 Hierarchical topic organization and visual presentation of spoken documents using probabilistic latent semantic analysis (PLSA) for efficient retrieval/browsing applications
abstract
The most attractive form of future network content will be multi-media including speech information, and such speech information usually carries the core concepts for the content. As a result, the spoken documents associated with the multi-media content very possibly can serve as the key for retrieval and browsing. This paper presents a new approach of hierarchical topic organization and visual presentation of spoken documents for such a purpose based on the Probabilistic Latent Semantic Analysis (PLSA). With this approach the spoken documents can be organized into a two-dimensional tree (or multi-layered map) of topic clusters, and the user can very efficiently retrieve or browse the network content or associated spoken documents. Different from the conventional document clustering approaches, with PLSA the relationships among the topic clusters and the appropriate terms as the topic labels can be very well derived. An initial prototype system with Chinese broadcast news as the example spoken documents including automatic generation of titles and summaries and retrieval/browsing functionalities is also presented. Choice of different units other than words to be used as the terms in the processing is also considered in the system based on the special structure of the Chinese language. 1.
Te-Hsuan Li, Ming-Han Lee, Berlin Chen, Lin-Shan Lee
INTERSPEECH4
2005 Improved spontaneous Mandarin speech recognition by disfluency interruption point (IP) detection using prosodic features
Che-Kuang Lin, Lin-Shan Lee
INTERSPEECH2
2005 Histogram-based quantization (HQ) for robust and scalable distributed speech recognition
Chia-Yu Wan, Lin-Shan Lee
INTERSPEECH2
2005 Segmental eigenvoice with delicate eigenspace for improved speaker adaptation
abstract
Eigenvoice techniques have been proposed to provide rapid speaker adaptation with very limited adaptation data, but the performance may be saturated when more adaptation data become available. This is because in these techniques an eigenspace with reduced dimensionality is established by properly utilizing the a priori knowledge from the large quantity of training data. The reduced dimensionality of the eigenspace requires less adaptation data to estimate the model parameters for the new speaker, but also makes it less easy to obtain more precise models with more adaptation data. In this paper, a new segmental eigenvoice approach is proposed, in which the eigenspace can be further segmented into N subeigenspaces by properly classifying the model parameters into N clusters. These N subeigenspaces can help to construct a more delicate eigenspace and more precise models when more adaptation data are available. It will be shown that there can be at least mixture-based, model-based and feature-based segmental eigenvoice approaches. Not only improved performance can be obtained, but these different approaches can be properly integrated to offer better performance. Two further approaches leading to improved segmental eigenvoice techniques with even better performance are also proposed. The experiments were performed with both a large vocabulary and a small vocabulary recognition tasks.
Yu Tsao 0001, Shang-Ming Lee, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2004 Efficient and robust distributed speech recognition (DSR) over wireless fading channels: 2D-DCT compression, iterative bit allocation, short BCH code and interleaving
abstract
In this paper, a new framework for distributed speech recognition (DSR) over wireless fading channels is proposed. A 2D-DCT based compression method and an iterative bit allocation algorithm are developed for source coding, while short BCH code integrated with interleaving is proposed for channel coding. The high correlation among speech feature parameters in the temporal domain is well exploited in the 2D-DCT based compression, and a carefully designed iterative bit allocation algorithm can very efficiently make use of every bit transmitted. The very low bit rate provides planning space for strong error control. Short BCH code integrated with interleaving is initially proposed to efficiently handle the severe bursty errors usually encountered in wireless fading channels. Overall system simulation based on a GPRS wireless channel simulator indicated significant recognition performance improvement is achievable at a bit rate of 3.4 kbit/s as compared to the conventional SVQ compression approach (with bit rate 4.4 or 4.8 kbit/s). The low computation requirements for all processes involved make it easy to implement them in the mobile telephone clients with existing technology.
Wei-Hao Hsu, Lin-Shan Lee
ICASSP (1)2
2004 Higher order cepstral moment normalization (HOCMN) for robust speech recognition
abstract
Cepstral mean subtraction (CMS) and cepstral normalization (CN) have been popularly used to normalize the first and the second moments of cepstral coefficients, and proved to be very helpful for robust speech recognition (Furui, S. 1981; Viikki, O. and Laurila, K., 1998). A unified formulation for higher order cepstral moment normalization (HOCMN) is developed by extending the concept of CMS and CN to orders much higher than three. A whole family of normalization techniques for different orders is thus proposed. Preliminary experimental results based on Aurora 2.0 showed that the recognition accuracy can be significantly improved with this approach under all noisy conditions. For example, HOCMN/sub (1,5,100)/ (normalization of the first, fifth and 100th order cepstral moments) is shown to offer an error rate reduction of 32.83% as compared to the conventional CN with a full-utterance processing interval, or an error rate reduction of 20.78% as compared to CN with a segmental processing interval.
Chang-Wen Hsu, Lin-Shan Lee
ICASSP (1)2
2004 Improved speech enhancement by applying time-shift property of DFT on hankel matrices for signal subspace decomposition
abstract
In previous studies, the signal subspace technique for speech enhancement was extended and a perceptually constrained generalized singular value decomposition (PCGSVD)-based algorithm [1] was developed which properly integrated the auditory masking effect and the GSVD algorithm. Both objective measures and subjective tests verified that this approach can offer better performance than the GSVD-based approach and the conventional spectral subtraction (SS) algorithm. But very high computational complexity is required in the PCGSVD-based method when performing the matrices decomposition via the GSVD algoruthm. In this paper, we properly utilize the time-shift property of DFT and the special structure of Hankel matrices to perform similar functions previously offered by GSVD, and a perceptually constrained minimum variance estimation algorithm is developed. By replacing GSVD algorithm with DFT, the computation complexity is significantly reduced, almost the same as the conventional SS algorithm. Experiments showed that comparable performance to that of the PCGSVD-based approach can be achieved, regardless of whether the additive noise is stationary or not, specially when it is non-white.
Gwo-hwa Ju, Lin-Shan Lee
INTERSPEECH2
2004 A new feature extraction front-end for robust speech recognition using progressive histogram equalization and multi-eigenvector temporal filtering
Shang-nien Tsai, Lin-Shan Lee
INTERSPEECH2
2004 A discriminative HMM/N-gram-based retrieval approach for mandarin spoken documents
abstract
In recent years, statistical modeling approaches have steadily gained in popularity in the field of information retrieval. This article presents an HMM/N-gram-based retrieval approach for Mandarin spoken documents. The underlying characteristics and the various structures of this approach were extensively investigated and analyzed. The retrieval capabilities were verified by tests with word- and syllable-level indexing features and comparisons to the conventional vector-space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. Fusion of information via indexing word- and syllable-level features was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained.
Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
ACM Trans. Asian Lang. Inf. Process.3
2003 Data-driven temporal filters based on multi-eigenvectors for robust features in speech recognition
abstract
It was previously proposed to use the principal component analysis (PCA) to derive the data-driven temporal filters for obtaining robust features in speech recognition, in which the first principal components are taken as the filter coefficients. In this paper, a multi-eigenvector approach is proposed instead, in which the first M eigenvectors obtained in PCA are weighted by their corresponding eigenvalues and summed to be used as the filter coefficients. Experimental results showed that the multi-eigenvector filters offer significant recognition performance as compared to the previously proposed PCA-derived filters under all different conditions tested with the AURORA2 database, especially when the training and testing environments are highly mismatched.
Ni-Chun Wang, Jeih-Weih Hung, Lin-Shan Lee
ICASSP (1)3
2003 Improved Chinese broadcast news transcription by language modeling with temporally consistent training corpora and iterative phrase extraction
abstract
In this paper an iterative Chinese new phrase extraction method based on the intra-phrase association and context variation statistics is proposed.A Chinese language model enhancement framework including lexicon expansion is then developed.Extensive experiments for Chinese broadcast news transcription were then performed to explore the achievable improvements with respect to the degree of temporal consistency for the adaptation corpora.Very encouraging results were obtained and detailed analysis discussed.
Pi-Chuan Chang, Shuo-Peng Liao, Lin-Shan Lee
INTERSPEECH3
2003 Automatic title generation for Chinese spoken documents using an adaptive k nearest-neighbor approach
abstract
The purpose of automatic title generation is to understand a document and to summarize it with only several but readable words or phrases. It is important for browsing and retrieving spoken documents, which may be automatically transcribed, but it will be much more helpful if given the titles indicating the content subjects of the documents. For title generation for Chinese language, additional problems such as word segmentation and key phrase extraction also have to be solved. In this paper, we developed a new approach of title generation for Chinese spoken documents. It includes key phrase extraction, topic classification, and a new title generation model based on an adaptive K nearest-neighbor concept. The tests were performed with a training corpus including 151,537 news stories in text form with human-generated titles and a testing corpus of 210 broadcast news stories. The evaluation included both objective F1 measures and 5-level subjective human evaluation. Very positive results were obtained. Keyword: title generation, Chinese spoken documents
Shun-Chuan Chen, Lin-Shan Lee
INTERSPEECH2
2003 Perceptually-constrained generalized singular value decomposition-based approach for enhancing speech corrupted by colored noise
abstract
In a previous work, we have successfully integrated the transformation-based signal subspace technique with the generalized singular value decomposition (GSVD) algorithm to develop an improved speech enhancement framework [1]. In this paper, we further incorporate the perceptual masking effect of the psychoacoustics model as extra constraints of the previously proposed GSVD-based algorithm to obtain improved sound feature, and furthermore make sure the undesired residual noise to be nearly unperceivable. Both subjective listening tests and spectrogram-plot comparison showed that the closed-form solution developed here can offer significantly better speech quality than either the conventional spectral subtraction algorithm or the previously proposed GSVD-based technique, regardless of whether the additive noise is white or not. 1.
Gwo-hwa Ju, Lin-Shan Lee
INTERSPEECH2
2003 Speech enhancement and improved recognition accuracy by integrating wavelet transform and spectral subtraction algorithm
Gwo-hwa Ju, Lin-Shan Lee
INTERSPEECH2
2003 Automatic title generation for Chinese spoken documents considering the special structure of the language
Lin-Shan Lee, Shun-Chuan Chen
INTERSPEECH1
2003 Cross domain Chinese speech understanding and answering based on named-entity extraction
Yun-Tien Lee, Shun-Chuan Chen, Lin-Shan Lee
INTERSPEECH3
2003 Why is the special structure of the language important for Chinese spoken language processing? - examples on spoken document retrieval, segmentation and summarization
abstract
The Chinese language is not only spoken by the largest population in the world, but quite different from many western languages with a very special structure. It is not alphabetic: large number of Chinese characters are ideographic symbols and pronounced as monosyllables. The open vocabulary nature, the flexible wording structure and the tone behavior are also good examples within the special structure. It is believed that better results and performance will be obtainable in developing Chinese spoken language processing technologies, if this special structure can be taken into account. In this paper, a set of “feature units” for Chinese spoken language processing is identified, and the retrieval, segmentation and summarization of Chinese spoken documents are taken as examples in analyzing the use of such “feature units”. Experimental results indicate that by careful considerations of the special structure and proper choice of the “feature units”, significantly better performance can be achieved.
Lin-Shan Lee, Yuan Ho, Jia-fu Chen, Shun-Chuan Chen
INTERSPEECH1
2002 Data-driven temporal filters for robust features in speech recognition obtained via Minimum Classification Error (MCE)
abstract
In deriving the data-driven temporal filters for speech features, the Linear Discriminant Analysis (LDA) and the Principal Component Analysis (PCA) have been shown to be successful in improving the feature robustness. In this paper, it's proposed that the criterion of Minimum Classification Error (MCE) can also be used to obtain the data-driven temporal filters. Two versions of MCE-derived temporal filters, Feature-based and Model-based, are proposed and it is shown that both of them can significantly improve the recognition performance of the original MFCC features as the LDA/PCA-derived filters do. Detailed comparative analysis among the different temporal filtering approaches is presented. It is also shown that the proposed MCE filters can be integrated with the conventional temporal filters, RASTA or CMS, to obtain improved recognition performance regardless of whether the training and testing environments are matched or mismatched, compressed or noise corrupted.
Jeih-Weih Hung, Lin-Shan Lee
ICASSP2
2002 Data-driven temporal filters obtained via different optimization criteria evaluated on Aurora2 database
Jeih-Weih Hung, Lin-Shan Lee
INTERSPEECH2
2002 Speech enhancement based on generalized singular value decomposition approach
Gwo-hwa Ju, Lin-Shan Lee
INTERSPEECH2
2002 Distributed Chinese keyword spotting and verification for spoken dialogues under wireless environment
Yun-Tien Lee, Cheng-Huang Wu, Yumin Lee, Lin-Shan Lee
INTERSPEECH4
2002 Improved Chinese spoken document retrieval with hybrid modeling and data-driven indexing features
abstract
Different models retrieve the documents based on different approaches of extracting the underlying content. Different levels of indexing features also offer different functionalities and discriminabilities when retrieving the documents. In this paper, we present results for Chinese spoken document retrieval with hybrid models to integrate the knowledge obtainable from three basic retrieval models, namely, the standard vector space model (VSM), the hidden Markov model (HMM), and the latent semantic indexing (LSI) model. The characteristics of retrieval performance using both word-level and syllable-level indexing features were extensively explored. In addition, a data-driven approach to derive variable-length indexing features is also presented. Very satisfactory performance can be achieved with these data-driven features while retaining very compact feature set size. Experiments showed that this approach has the potential to identify domain-specific terminologies or newly-generated phrases. It is therefore very useful not only in Chinese document retrieval, but also in detecting out of vocabulary (OOV) words in Chinese. Very encouraging results were obtained when the hybrid models were used with the data-driven indexing features as well. 1.
Chun-Jen Wang, Berlin Chen, Lin-Shan Lee
INTERSPEECH3
2002 A hierarchical tag-graph search scheme with layered grammar rules for spontaneous speech understanding
Bor-Shen Lin, Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
Pattern Recognit. Lett.4
2002 Discriminating capabilities of syllable-based features and approaches of utilizing them for voice retrieval of speech information in Mandarin Chinese
abstract
With the rapidly growing use of the audio and multimedia information over the Internet, the technology for retrieving speech information using voice queries is becoming more and more important. In this paper, considering the monosyllabic structure of the Chinese language, a whole class of syllable-based indexing features, including overlapping segments of syllables and syllable pairs separated by a few syllables, is extensively investigated based on a Mandarin broadcast news database. The strong discriminating capabilities of such syllable-based features were verified by comparing with the word- or character-based features. Good approaches for better utilizing such capabilities, including fusion with the word- and character-level information and improved approaches to obtain better syllable-based features and query expressions, were extensively investigated. Very encouraging experimental results were obtained.
Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2002 A set of corpus-based text-to-speech synthesis technologies for Mandarin Chinese
abstract
This paper presents a set of corpus-based text-to-speech synthesis technologies for Mandarin Chinese. A large speech corpus produced by a single speaker is used, and the speech output is. synthesized from waveform units of variable lengths, with desired linguistic properties, retrieved from this corpus. Detailed methodologies were developed for designing "phonetically rich" and "prosodically rich" corpora by automatically selecting sentences from a large text corpus to include as many desired phonetic combinations and prosodic features as possible. Automatic phonetic labeling with iterative correction rules and automatic prosodic labeling with a multi-pass top-down procedure were also developed such that the labeling process for the corpora can be completely automatic. A hierarchical prosodic structure for an arbitrary desired text sentence is then generated based on the identification of different levels of break indices, and the prosodic feature sets and appropriate waveform units are finally selected and retrieved from the corpus, modified if necessary, and concatenated to produce the output speech. The special structure of Mandarin Chinese has been carefully considered in all these technologies, and preliminary assessments indicated very encouraging synthesized speech quality.
Fu-Chiang Chou, Chiu-yu Tseng, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2001 Rapid speaker adaptation using a priori knowledge by eigenspace analysis of MLLR parameters
abstract
This paper considers the problem of rapid speaker adaptation in speech recognition. In particular, we exploit an approach based on combination of transformations, which utilizes the concepts of both maximum likelihood linear regression (MLLR) and eigenvoice adaptation. We analyze three different possible methods to realize the concept, and formulate a fast algorithm of maximum likelihood coefficient estimation for test speakers. It is found that the best approach can properly utilize the a priori knowledge of speaker-independent models in constructing the eigenspace for speaker characteristics, while using MLLR matrices in representing the specific speakers so as to reduce the on-line memory and computation requirement of the adaptation phase. This best approach leads to identical models relative to eigenvoice adaptation that is based on MLLR-adapted speaker models. The experimental results and discussions also provide a good analysis towards integration of the MLLR and eigenvoice approaches.
Nick Jui-Chang Wang, Sammy S.-M. Lee, Frank Seide, Lin-Shan Lee
ICASSP4
2001 Improved spoken document retrieval by exploring extra acoustic and linguistic cues
abstract
In this paper, we explored the use of various extra information to improve the performance of spoken document retrieval (SDR). From the speech recognition perspective, we incorporated the acoustic stress and word confusion information into the audio indexing. From the linguistic perspective, we applied the part-of-speech information in both the audio indexing and the query representation. From the information retrieval perspective, we integrated techniques such as the query expansion by word associations and the blind relevance feedback into the retrieval process. The SDR experiments were based on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). We used the Chinese newswire text stories as query exemplars and the Mandarin Chinese audio news stories as the spoken documents. With all the above acoustic and linguistic cues applied, the average precision was improved from 0.5122 to 0.6312 for the TDT-2 collection and from 0.6216 to 0.7172 for the TDT-3 collection. 1.
Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
INTERSPEECH3
2001 An HMM/n-gram-based linguistic processing approach for Mandarin spoken document retrieval
abstract
In this paper an HMM/N-gram-based linguistic processing approach for Mandarin spoken document retrieval is presented. The underlying characteristics and different structures of this approach were extensively investigated. The retrieval capabilities were verified by tests with indexing features of word- and syllable(subword)-levels and comparison with the conventional vector space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. The information fusion of indexing features of word- and syllable-levels was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained.
Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
INTERSPEECH3
2001 Credibility proof for speech content and speaker verification by fragile watermarking with consecutive frame-based processing
abstract
In this paper, we proposed an easy scheme to protect the integrity of a period of speech content. We first break the speech signal into a series of consecutive frame and do DCT transform on them, then we sequentially encode the watermark into those frames separately. Every watermark is modulated by the statistical characteristic of previous frame of speech content; therefore we can detect any signal discontinuity caused by telltale signal distortion on the watermarked speech. Experiment shows that even after cropping most of the signal and leaves only 0.1 second of watermarked signal, we can still detect the watermark.
Yiou-Wen Cheng, Lin-Shan Lee
INTERSPEECH2
2001 Comparative analysis for data-driven temporal filters obtained via principal component analysis (PCA) and linear discriminant analysis (LDA) in speech recognition
abstract
The Linear Discriminant Analysis (LDA) has been widely used to derive the data-driven temporal filtering of speech feature vectors. In this paper, we proposed that the Principal Component Analysis (PCA) can also be used in the optimization process just as LDA to obtain the temporal filters, and detailed comparative analysis between these two approaches are presented and discussed. It's found that the PCA-derived temporal filters significantly improve the recognition performance of the original MFCC features as LDA-derived filters do. Also, while PCA/LDA filters are combined with the conventional temporal filters, RASTA or CMS, the recognition performance will be further improved regardless the training and testing environments are matched or mismatched, compressed or noise corrupted. 1.
Jeih-Weih Hung, Hsin-Min Wang, Lin-Shan Lee
INTERSPEECH3
2001 Pronunciation variation analysis with respect to various linguistic levels and contextual conditions for Mandarin Chinese
abstract
Chinese language has quite different characteristic structures from those of English. There are at least word, character, syllable, Initial-Final levels in Chinese, each carrying different levels of information with complicated correlations among them. In this paper, we investigate the dependency of pronunciation variation in conversational Mandarin speech on these different levels under various contextual conditions considering the structural features of the language. The influence of speaking rate and word frequency on such pronunciation variation is also analyzed. Different pruning methods, for including pronunciation variation in speech recognition were also evaluated, and the experimental results showed that improved accuracy is obtainable if the characteristics of the pronunciation variation found in the analysis can be properly taken into account. All discussions here are based on tests with the LDC Mandarin Call Home corpus.
Ming-Yi Tsai, Fu-Chiang Chou, Lin-Shan Lee
INTERSPEECH3
2001 Segmental eigenvoice for rapid speaker adaptation
abstract
This paper presents a new approach to improve the conventional eigenvoice technique. In the conventional eigenvoice, an eigenspace is established by introducing a priori training speakers via PCA. The adaptation data is then used to determine a group of coefficients with respect to the eigenspace and build the SD model for the testing speaker. In the proposed approach, the eigenspace in the conventional eigenvoice is segmented into N sub-eigenspaces. Each subeigenspace is established by those components in the training speaker SD models with similar properties to each other. With the adaptation data, N groups of coefficients corresponding to the N sub-eigenspaces can be determined to build SD model for the new testing speaker. Here, both mixture-based and feature-based segmentation of eigenspace were tested, and improved results compared to the conventional eigenvoice were obtained in both cases. Even better results were obtained when these approaches were properly combined.
Yu Tsao 0001, Shang-Ming Lee, Fu-Chiang Chou, Lin-Shan Lee
INTERSPEECH4
2001 Eigen-MLLR coefficients as new feature parameters for speaker identification
Nick Jui-Chang Wang, Wei-Ho Tsai, Lin-Shan Lee
INTERSPEECH3
2001 Voice access of global information for broad-band wireless: technologies of today and challenges of tomorrow
abstract
The rapid development of the Internet and the World Wide Web has created a global network that will soon become a physical embodiment of the entire human knowledge and a complete integration of the global information activities. The traditional approach to access the network is through a computer physically tied to the network. As broad-band wireless takes off, the traditional tethered approach will gradually become obsolete. It is believed that one of the most natural and user-friendly approached for accessing the network will be via human voice, and the integration of spoken language processing technologies with broad-band wireless technologies will be a key to the evolution of a broad-band wireless information community. This paper offers a vision of the above concept. Technical considerations and some typical example applications of accessing the information and services using voice in a broad-band wireless environment are discussed. Fundamentals of spoken language processing technologies that are crucial in such a broad-band wireless environment are briefly reviewed. Technical challenges caused by the unique nature of wireless mobile communications are also presented along with some possible solutions.
Lin-Shan Lee, Yumin Lee
Proc. IEEE1
2001 New approaches for domain transformation and parameter combination for improved accuracy in parallel model combination (PMC) techniques
abstract
Parallel model combination (PMC) techniques have been very successful and popularly used in many applications to improve the performance of speech recognition systems under noisy environments. However, it is believed that some assumptions and approximations made in this approach, primarily in the domain transformation and parameter combination processes, are not necessarily accurate enough in certain practical situations, which may degrade the achievable performance of PMC. In this paper, the possible sources that cause the performance degradation in these processes are carefully analyzed and discussed. Three new approaches, including the truncated Gaussian approach and the split mixture approach for the domain transformation process and the estimated cross-term approach for parameter combination process, are proposed in this paper in order to handle these problems, minimize such degradation, and improve the accuracy of the PMC techniques. These proposed approaches were analyzed and discussed with two recognition tasks, one relatively simple, and the other more complicated and realistic. Both sets of experiments showed that these proposed approaches are able to provide significant improvements over the original PMC method, especially when the SNR condition is worse.
Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
2001 Computer-aided analysis and design for spoken dialogue systems based on quantitative simulations
abstract
In this paper, a complete development of computer-aided analysis and design approaches for spoken dialogue systems based on quantitative simulations is presented. With this approach the various performance metrics of a dialogue system can be flexibly defined and numerically evaluated, such that the behavior and performance of the dialogue system can be well predicted and efficiently analyzed before the implementation of the real spoken dialogue system is completed. How the different dialogue performance measures vary with respect to each of the many very complicated factors, regardless of whether it is caused by an individual component, by the overall system design, or by users' response patterns, can be separately identified, because all such factors can he precisely controlled in the simulation. Several analysis examples are presented to show how the approach can be used, including selection and tuning of the speech understanding front end, system strategy design considering query factors and confirmation factors, and objective estimates of user's degree of satisfaction. This approach is therefore very useful for the analysis and design of spoken dialogue systems, although the online test, corpus-based analysis and user survey can always follow after the system is online.
Bor-Shen Lin, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2000 Retrieval of broadcast news speech in Mandarin Chinese collected in Taiwan using syllable-level statistical characteristics
abstract
Spoken document retrieval has been extensively studied over the years because of its high potential in various applications in the near future. Considering the monosyllabic structure of the Chinese language, a whole class of indexing features for retrieval of spoken documents in Mandarin Chinese using syllable-level statistical characteristics has been studied, and very encouraging experimental results on retrieval of broadcast news speech collected in Taiwan were obtained. This paper reports some interesting initial results and findings obtained in this research.
Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
ICASSP3
2000 Fundamental performance analysis for spoken dialogue systems based on a quantitative simulation approach
abstract
The performance of dialogue systems is mostly measured based on the analysis of a large dialogue corpus. In this way, the dialogue performance can not be obtained before the system is on line, and the dialogue corpus should be recollected if the system is modified. Also, the effect of different factors, including system dialogue strategy, recognition and understanding accuracy or user response pattern, etc., on the dialogue performance can not be quantitatively identified because they can not be precisely controlled in different corpora. In this paper, a fundamental performance analysis scheme for dialogue systems based on a quantitative simulation approach is proposed. With this scheme the fundamental performance of a dialogue system can be predicted and analyzed efficiently without having any real spoken dialogue system implemented or having any dialogue corpus actually collected. How the dialogue performance varies with respect to each factor, from recognition accuracy to dialogue strategy, can be individually identified, because all such factors can be precisely controlled in the simulation. The quality of service for the spoken dialogue system can also be flexibly defined and the design parameters easily determined. This approach is therefore very useful for the design, development and improvement of spoken dialogue although the on-line system and real corpus will eventually be needed in the final evaluation and analysis of the system performance in any case.
Bor-Shen Lin, Lin-Shan Lee
ICASSP2
2000 Fast speaker adaptation using eigenspace-based maximum likelihood linear regression
abstract
This paper presents an eigenspace-based fast speaker adaptation approach which can improve the modeling accuracy of the conventional maximum likelihood linear regression (MLLR) techniques when only very limited adaptation data is available. The proposed eigenspace-based MLLR approach was developed by introducing a priori knowledge analysis on the training speakers via PCA, so as to construct an eigenspace for MLLR full regression matrices as well as to derive a set of bases called eigen-matrices. The full regression matrices for each outside speaker are then constrained to be located in the space spanned by the first K eigen-matrices. The proposed eigenspace-based regression matrices, serving as an initial estimate of the speaker-specific MLLR transformation, effectively reduces the number of free parameters, while precise modeling for the inter-dimensional correlation among the model parameters by full matrices was maintained. Experimental results showed that for supervised adaptation...
Kuan-Ting Chen, Wen-Wei Liau, Hsin-Min Wang, Lin-Shan Lee
INTERSPEECH4
2000 Retrieval of mandarin broadcast news using spoken queries
abstract
Considering the monosyllabic structure of the Chinese language, a whole class of indexing features for retrieval of Mandarin broadcast news using syllable-level statistical characteristics has been previously investigated. This paper presents the improvements achieved over the previous results. The major differences are: (1) Multi-scale character- and word-level indexing terms have been integrated with the syllable-level information. (2) Information cues from the contemporary newswire text corpus have been used to create more accurate syllable indexing terms. (3) Automatic document expansion, blind relevance feedback, and query expansion via the term association matrix have been applied in retrieval. With all these schemes, the average precision can be improved from 55.46 % to 71.29%. 1.
Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
INTERSPEECH3
2000 Automatic metric-based speech segmentation for broadcast news via principal component analysis
abstract
In this paper, we proposed an algorithm used to improve the performance of the metric-based segmentation techniques, by which the segmentation points are found at maxima of a distance measured between two contiguous windows shifted along the stream of speech features. In our proposed method, the PCA processes are first performed on the speech features to obtain more robust features, and then the above metric-based segmentation was applied on the PCA-derived features to decide the segmentation points. Experiment results show that our proposed method can efficiently improve the detection rates of the segmentation points up to 7% while the false alarm rates remain unchanged.
Jeih-Weih Hung, Hsin-Min Wang, Lin-Shan Lee
INTERSPEECH3
2000 MAT-2000 - design, collection, and validation of a Mandarin 2000-speaker telephone speech database
Hsiao-Chuan Wang, Frank Seide, Chiu-yu Tseng, Lin-Shan Lee
INTERSPEECH4
2000 Live Lexicons and Dynamic Corpora Adapted to the Network Resources for Chinese Spoken Language Processing Applications in an Internet Era
Lin-Shan Lee, Lee-Feng Chien
LREC1
1999 Improved parallel model combination techniques with split Gaussian mixtures for speech recognition under noisy conditions
abstract
The parallel model combination (PMC) technique has been very successful and frequently used to improve the performance of a speech recognition system under noisy environments. In this approach it is assumed that the log spectrum of speech signals is Gaussian-distributed, which is not always valid especially when the number of mixtures in the HMMs is few. In this paper, a simple approach is proposed to improve the PMC method by splitting the mixtures before the domain transformation process in the PMC is performed, and merging the mixtures back to the original number after the PMC processes are completed. Preliminary experimental results show that the increased number of mixtures during the PMC processes can in fact provide significant improvements over the original PMC method in terms of the recognition accuracies, especially when the SNR is low.
Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee
ICASSP3
1999 A framework of performance evaluation and error analysis methodology for speech understanding systems
abstract
With improved speech understanding technology, many successful working systems have been developed. However, the high degree of complexity and wide variety of design methodology make the performance evaluation and error analysis for such systems very difficult. The different metrics for individual modules such as the word accuracy, spotting rate, language model coverage and slot accuracy are very often helpful, but it is always difficult to select or tune each of the individual modules or determine which module contributed to how much percentage of understanding errors based on such metrics. A new framework for performance evaluation and error analysis for speech understanding systems is proposed based on the comparison with the 'best-matched' references obtained from the word graphs with the target words and tags given. In this framework, all test utterances can be classified based on the error types, and various understanding metrics can be obtained accordingly. Error analysis approaches based on an error plane are then proposed, with which the sources for understanding errors (e.g., poor acoustic recognition, language model, search error, etc.) can be identified for each utterance. Such a framework will be very helpful for design and analysis of speech understanding systems.
Bor-Shen Lin, Lin-Shan Lee
ICASSP2
1999 Selection of waveform units for corpus-based Mandarin speech synthesis based on decision trees and prosodic modification costs
abstract
The removal of noise from speech signals has applications ranging from speech enhancement for cellular communications to front ends for speech recognition systems. In this paper, we present a new nonlinear time-domain method called Noise-Regularized Adaptive Filtering. The approach is based on minimum mean-squared estimation using a modified cost function and allows designing both linear and nonlinear filters using only the observed noisy speech. 1. A GENERAL FRAMEWORK FOR MMSE ESTIMATION Given a noisy speech signal,
Fu-Chiang Chou, Chiu-yu Tseng, Lin-Shan Lee
EUROSPEECH3
1999 Phonetic state tied-mixture tone modeling for large vocabulary continuous Mandarin speech recognition
abstract
This paper describes a practical dictation system that is able to compensate for the lack of a large text database and reports on the results of eld tests in which the system was used to make medical rehabilitation diagnosis reports in a hospital. Our dictation system uses two recognition engines: continuous speech recognition and isolated word (including connected words) recognition engines. When a user makes a report, the two recognition engines can be selectively operated in a single window through voice commands or mouse click. The system operates on common personal computers under the Windows O.S. A eld test conducted in noisy therapeutic work rooms that compared the performance of speech input system comparing to conventional keyboard input system for the medical rehabilitation eld, demonstrated the e ectiveness of the dictation system.
Tai-Hsuan Ho, Chin-Jung Liu 0002, Herman Sun, Ming-Yi Tsai, Lin-Shan Lee
EUROSPEECH5
1999 Consistent dialogue across concurrent topics based on an expert system model
abstract
This paper describes the development and evaluation of objective methods for testing synthetic intonation. While subjective methods are available for assessing the quality of synthetic intonation, such tests consume time and resources, and are not useful for day-to-day model development. Therefore, objective measures of F0 modelling are necessary. Currently, objective evaluation of synthetic intonation involves the use of Root Mean Squared Error and Correlation. However, it is unclear how large an improvement in either score must be before it is reflected perceptually. It is also unclear how detailed an analysis these metrics provide. Therefore, two other metrics are to be tested, both of which are similar to a basic RMSE measurement. All of the evaluation results are compared to a perceptual study in order to determine how the objective measures relate to perceived differences in the contours. 1. INTRODUCTION One difficulty in building models for synthesizing intonation is determinin...
Bor-Shen Lin, Hsin-Min Wang, Lin-Shan Lee
EUROSPEECH3
1999 Automatic selection of phonetically distributed sentence sets for speaker adaptation with application to large vocabulary Mandarin speech recognition
Jia-Lin Shen, Hsin-Min Wang, Ren-Yuan Lyu, Lin-Shan Lee
Comput. Speech Lang.4
1998 Improved search strategy for large vocabulary continuous Mandarin speech recognition
abstract
This paper presents a new search strategy for large vocabulary continuous Mandarin speech recognition considering the special structure of the Chinese language. This strategy is composed of forward and backward passes, between which a high-quality syllable lattice is generated to bridge the syllable-level and word-level decoding processes. In the forward pass, considering the small number of syllables in the Chinese language, a frame-synchronous stack decoder is used to integrate the high-order syllable N-Gram language model, so as to generate a very accurate and compact syllable lattice. In the backward pass, considering the special monosyllabic wording structure in the Chinese language, the search space for the word-level decoding is expanded dynamically from the syllable lattice, and the best word sequence is extracted based on the knowledge provided by the word pronunciation lexicon and the word N-Gram language model. In the preliminary experiments, it was found that, with this strategy, the character error rate can be reduced by more than 20% as compared with a previous system using syllable-aligned lattice approach on a speaker-adaptive continuous speech recognition task.
Tai-Hsuan Ho, Kae-Cherng Yang, Kuo-Hsun Huang, Lin-Shan Lee
ICASSP4
1998 Improved robustness for speech recognition under noisy conditions using correlated parallel model combination
abstract
The parallel model combination (PMC) technique has been shown to achieve very good performance for speech recognition under noisy conditions. In this approach, the speech signal and the noise are assumed uncorrelated during modeling. A new correlated PMC is proposed by properly estimating and modeling the nonzero correlation between the speech signal and the noise. Preliminary experimental results show that this correlated PMC can provide significant improvements over the original PMC in terms of both the model differences and the recognition accuracies. Error rate reduction on the order of 14% can be achieved.
Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee
ICASSP3
1998 Statistics-based segment pattern lexicon-a new direction for Chinese language modeling
abstract
This paper presents a new direction for Chinese language modeling based on a different concept of the lexicon. Because every Chinese character has its own meaning and there are no "blanks" in Chinese sentences serving as word boundaries, also because the wording structure in the Chinese language is extremely flexible, the "words" in Chinese are actually not well defined, and there does not exist a commonly accepted lexicon. This makes language modeling very sophisticated in the Chinese language, and the "out of vocabulary (OOV)" problem specially serious. A new concept for the lexicon is thus proposed. The elements of this lexicon can be words or any other "segment patterns". They should be extracted from the training corpus by statistical approaches with a goal to minimize the overall perplexity. The language models can then be developed based on this new lexicon. Very encouraging experimental results have been obtained.
Kae-Cherng Yang, Tai-Hsuan Ho, Lee-Feng Chien, Lin-Shan Lee
ICASSP4
1998 A modified blind equalization technique based on a constant modulus algorithm
abstract
Improved blind equalisation techniques based on a modified constant modulus algorithm (CMA) are proposed. This algorithm can not only accomplish blind equalization and carrier phase recovery simultaneously, but also choose a nonlinear gain with a least-squares algorithm to avoid the gradient noise amplification problem and achieve improved stability and robustness without increasing the computation complexity. It can be switched over to a decision-directed equalization scheme once the error level is reasonably low to obtain higher convergence speed and lower residual intersymbol interference as well. Extensive computer simulation results have verified the analysis and indicated the very attractive behavior of the proposed algorithm.
Jia-Chin Lin 0001, Lin-Shan Lee
ICC2
1998 A*-admissible key-phrase spotting with sub-syllable level utterance verification
abstract
In this paper, we propose an A*-admissible key-phrase spotting framework, which needs little domain knowledge and is capable of extracting salient key-phrase fragments from an input utterance in real-time. There are two key features in our approach. Firstly, the acoustic models and the search framework are specially designed such that very high degree vocabulary flexibility can be achieved for any desired application tasks. Secondly, the search framework uses an efficient two-pass A* search to generate N-best key-phrase candidates and then several sub-syllable level verification functions are properly weighted and used to further improve the recognition accuracy. Experimental results show that the A*-admissible key-phrase spotting with sub-word level utterance method outperforms the baseline methods used in common approaches. 1. INTRODUCTION In recent years, various spoken dialog systems have been widely investigated for the fast growing demand for real-world applications. It is diff...
Berlin Chen, Hsin-Min Wang, Lee-Feng Chien, Lin-Shan Lee
ICSLP4
1998 Automatic segmental and prosodic labeling of Mandarin speech database
Fu-Chiang Chou, Chiu-yu Tseng, Lin-Shan Lee
ICSLP3
1998 Improved parallel model combination based on better domain transformation for speech recognition under noisy environments
Jeih-Weih Hung, Jia-Lin Shen, Lin-Shan Lee
ICSLP3
1998 Hierarchical tag-graph search for spontaneous speech understanding in spoken dialog systems
abstract
It has been relatively difficult to develop natural language parsers for spoken dialog systems, not only because of the possible recognition errors, pauses, hesitations, out-ofvocabulary words, and the grammatically incorrect sentence structures, but because of the great efforts required to develop a general enough grammar with satisfactory coverage and flexibility to handle different applications. In this paper, a new hierarchical graph-based search scheme with layered structure is presented, which is shown to provide more robust and flexible spontaneous speech understanding for spoken dialog systems. 1.INTRODUCTION Traditionally, natural language understanding is integrated with the speech recognizer with a N-best interface in spoken dialog systems [1][2], that is, the recognizer sequentially generates its best N sentence hypotheses until any one is accepted by the natural language understanding part. However, for spontaneous speech with fragments, disfluencies, OOV words, and ill-...
Bor-Shen Lin, Berlin Chen, Hsin-Min Wang, Lin-Shan Lee
ICSLP4
1998 Improved robust speech recognition considering signal correlation approximated by taylor series
Jia-Lin Shen, Jeih-Weih Hung, Lin-Shan Lee
ICSLP3
1998 Robust entropy-based endpoint detection for speech recognition in noisy environments
abstract
This paper presents an entropy-based algorithm for accurate and robust endpoint detection for speech recognition under noisy environments. Instead of using the conventional energy-based features, the spectral entropy is developed to identify the speech segments accurately. Experimental results show that this algorithm outperforms the energy-based algorithms in both detection accuracy and recognition performance under noisy environments, with an average error rate reduction of more than 16%.
Jia-Lin Shen, Jeih-Weih Hung, Lin-Shan Lee
ICSLP3
1998 A syllable-based Chinese spoken dialogue system for telephone directory services primarily trained with a corpus
Yen-Ju Yang, Lin-Shan Lee
ICSLP2
1998 Isolated Mandarin base-syllable recognition based upon the segmental probability model
abstract
A segmental probability model (SPM) is proposed for fast and accurate recognition of the highly confusing isolated Mandarin base-syllables by deleting the state transition probabilities of continuous density hidden Markov models (CHMM), abandoning the dynamic programming process, letting the states equally segment the base-syllables deterministically, and using several special approaches to improve the accuracy and speed. This is achieved by considering the special characteristics of the target vocabulary.
Ren-Yuan Lyu, I-Chung Hong, Jia-Lin Shen, Ming-Yu Lee, Lin-Shan Lee
IEEE Trans. Speech Audio Process.5
1997 A multi-phase approach for fast spotting of large vocabulary Chinese keywords from Mandarin speech using prosodic information
abstract
This paper presents a multi-phase approach for fast spotting of large vocabulary Chinese keywords from a spontaneous Mandarin speech utterance using prosodic knowledge. Without searching through the whole utterance using large number of keyword models, the multi-phase framework proposed including some special scoring schemes provides very good efficiency by considering the monosyllable-based structure of Mandarin Chinese. This approach is therefore very fast due to very good boundary estimations and the deletion of most impossible syllable and keyword candidates using context independent models, and is also very accurate due to the carefully designed scoring processes. A task with 2611 keywords was tested. An inclusion rate of 85.79% for the top 10 candidates is attained, at a speed requiring only 1.2 times that of the utterance length on a Sparc 20 workstation.
Bo-Ren Bai, Chiu-yu Tseng, Lin-Shan Lee
ICASSP3
1997 Internet Chinese information retrieval using unconstrained Mandarin speech queries based on a client-server architecture and a PAT-tree-based language model
abstract
In order to pursue high performance of Chinese information access on the Internet, this paper presents an attractive approach with a successful integration of efficient speech recognition and information retrieval techniques. A working system based on the proposed approach for speech retrieval of real-time Chinese netnews services has been implemented and tested. Very exciting performance has been achieved.
Lee-Feng Chien, Sung-Chien Lin, Jenn-Chau Hong, Ming-Chiuan Chen, Hsin-Min Wang, Jia-Lin Shen, Keh-Jiann Chen, Lin-Shan Lee
ICASSP8
1997 A Chinese text-to-speech system based on part-of-speech analysis, prosodic modeling and non-uniform units
abstract
This paper presents a new Chinese text-to-speech system that produces very natural and intelligible synthetic Mandarin speech based on part-of-speech analysis, prosodic modeling and non-uniform units. The distinguishing features and key technology for the system can be summarized as follows. (1) A text analysis module for word identification and tagging was developed based on part-of-speech modeling and using heuristic rules to achieve very high accuracy. (2) The required prosodic parameters for the synthetic speech are derived from a two-stage procedure. The prosodic structures of the input texts are first derived from a statistical model trained by a large speech database, and the prosodic parameters are then determined according to the structures. (3) A specially designed speech segments inventory constructed with non-uniform and pitch dependent units is used to improve the fluency and intelligibility of the system.
Fu-Chiang Chou, Chiu-yu Tseng, Keh-Jiann Chen, Lin-Shan Lee
ICASSP4
1997 Syllable-based relevance feedback techniques for Mandarin voice record retrieval using speech queries
abstract
In order to solve the problem with the new environment of fast growth of audio resources on the Internet, we have presented a syllable-based approach which is capable of retrieving Mandarin voice records using queries of unconstrained speech. However, the performance achieved by this previously proposed approach is still not satisfactory, and one of the reason is that very often the information provided by the speech query for the request subject may not be sufficient. We present approaches based the relevance feedback technique to improving the performance of the previous research. The proposed approaches include a relevance measure adjustment scheme using a relevance table for the voice database, a query expansion scheme to generate a new query including the feedback information, and a combination of these two schemes. Extensive preliminary experiments were performed and demonstrated.
Lin-Shan Lee, Bo-Ren Bai, Lee-Feng Chien
ICASSP1
1997 A Modified Code Tracking Loop for Direct-Sequence Spread-Spectrum Systems on Frequency-Selective Fading Channels
abstract
A modified fully-digital code tracking loop is proposed in this paper for direct-sequence spread-spectrum signaling on a frequency-selective fading channel. A data-modulated channel estimator is used to cope with the time-varying Rayleigh fading effect and the data modulation effect, and extract the desired error signal from each path independently in the multipath environments. By taking advantage of the inherent diversity with the maximal ratio combining (MRC) or a proposed Even/odd maximal ratio combining (EMRC) technique, this modified code tracking loop can avoid the problem due to the drift or flutter effects of the error characteristics, and provide better performance on frequency selective fading channels. Extensive computer simulation has verified the analysis and indicated very attractive behavior of the proposed digital tracking loop.
Jia-Chin Lin 0001, Lin-Shan Lee
ICC (1)2
1997 Intelligent retrieval of very large Chinese dictionaries with speech queries
abstract
To retrieve a Chinese word from a Chinese dictionary, it needs the user to know exactly the first character of the desired word. Because there is more than 10,000 Chinese characters, this makes the Chinese dictionary relatively difficult to be used. To reduce the problem, this paper presents intelligent retrieval techniques for very large Chinese dictionaries with speech queries. The proposed techniques properly integrate the technologies of Mandarin speech recognition and Chinese information retrieval with a syllable-based approach utilizing the mono-syllabic structure of the language. Moreover, it is very nice to provide the function of retrieving all relevant word entries from the dictionaries using speech queries describing “general concepts” of the desired words. To achieve the challenging function, the techniques of relevance feedback are also included. Based on these techniques, a retrieval system was implemented successfully on a Pentium PC for a very large Chinese dictionary which includes 160,000 word entries and the total length of the lexical information under the word entries exceeds 20,000,000 words.
Sung-Chien Lin, Lee-Feng Chien, Ming-Chiuan Chen, Lin-Shan Lee, Keh-Jiann Chen
EUROSPEECH4
1997 Chinese language model adaptation based on document classification and multiple domain-specific language models
Sung-Chien Lin, Chi-Lung Tsai, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH5
1997 Isolated Mandarin syllable recognition with limited training data specially considering the effect of tones
abstract
A set of new approaches is proposed to model the Mandarin syllables for accurate recognition with limited training data while specially considering the effect of tones, including improved initial values and state transition topologies, and making use of the durational cue. The results show that these approaches are very useful practically.
Yumin Lee, Lin-Shan Lee, Chiu-yu Tseng
IEEE Trans. Speech Audio Process.2
1997 Complete recognition of continuous Mandarin speech for Chinese language with very large vocabulary using limited training data
abstract
This correspondence presents the first known results of complete recognition of continuous Mandarin speech for the Chinese language with very large vocabulary but very limited training data. Various acoustic and linguistic processing techniques were developed, and a prototype system of a continuous speech Mandarin dictation machine has been successfully implemented. The best recognition accuracy achieved is 92.2% for finally decoded Chinese characters.
Hsin-Min Wang, Tai-Hsuan Ho, Rung-Chiung Yang, Jia-Lin Shen, Bo-Ren Bai, Jenn-Chau Hong, Wei-Peng Chen, Tong-Lo Yu, Lin-Shan Lee
IEEE Trans. Speech Audio Process.9
1996 An efficient voice retrieval system for very-large-vocabulary Chinese textual databases with a clustered language model
abstract
This paper presents an accurate and efficient voice retrieval system for very-large-vocabulary Chinese textual databases with a specially-designed clustered language model. To reduce the problems resulted from the complexity of unconstrained speech-input queries for retrieval, the system is completely syllable-based in both speech recognition and database retrieval by properly utilizing the mono-syllabic structure of Chinese language. In addition, it partitions the records in the database into clusters and trains the clustered language model using the clustering results. The proposed clustered language model with its augmented search algorithm are very useful to improve accuracy and speed of the speech retrieval system. In the preliminary tests using an experimental database with about 30,000 bibliographical records, it was found that the present system can accept unconstrained speech-input queries and achieve very good performance.
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ICASSP4
1996 Fast and accurate recognition of very-large-vocabulary continuous Mandarin speech for Chinese language with improved segmental probability modeling
abstract
This paper presents a fast and accurate recognition of continuous Mandarin speech with very large vocabulary using an improved segmental probability model (SPM) approach. In order to extensively utilize the acoustic and linguistic knowledge to further improve the recognition performance, a few special techniques are thus developed. Preliminary simulation results show that the final achievable rate for the base syllable recognition with the improved segmental probability modeling is as high as 91.62%, which indicates a 18.48% error rate reduction and more than 3 times faster than the well-studied sub-syllable-based CHMM. Also, a tone recognizer and a word-based Chinese language model are included and the achieved recognition accuracy for the final decoded Chinese characters is 92.10%.
Jia-Lin Shen, Lin-Shan Lee
ICASSP2
1996 Very-large-vocabulary Mandarin voice message file retrieval using speech queries
abstract
In order to solve the problem with the new environment of fast growth of audio resources on the Intemet, this paper presents a new approach which is capable of retrieving Mandarin voice message files using queries of unconstrained speech.By properly utilizing the monosyllabic structure of the Chinese language, the proposed approach perfoms the statistical similarity estimation between the speech queries and the voice message files, and executes the complete matching process directly at the phonetic level using syllable-based statistical information.Based on this approach, some experiments are tested and encouraging results are demonstrated.
Bo-Ren Bai, Lee-Feng Chien, Lin-Shan Lee
ICSLP3
1996 Automatic generation of prosodic structure for high quality Mandarin speech synthesis
abstract
A key problem for today's speech synthesis technology is to automatically generate an appropriate hierarchical prosodic structure for text input and incorporate it into synthesized sptech[l][2].This paper presents a method for such a problem in Mandarin Chinese.This method uses a speech database for the training of a statistical model to generate the prosodic structure and determine prosodic parameten such as syllable duration, pause, energy and intonation.The experimental results show that an accuraq of 83.1% in the prediction of prosodic structure can be achieved.Furthermore, a Chinese text-to-pech system can be developed based on the proposed prosodic st"Te.
Fu-Chiang Chou, Chiu-yu Tseng, Lin-Shan Lee
ICSLP3
1996 Use of prosodic information to integrate acoustic and linguistic knowledge in continuous Mandarin speech recognition with very large vocabulary
abstract
This paper presents a new approach to use prosodic information for the integration of acoustic and linguistic knowledge in continuous Mandarin speech with very large vocabulary.Since the overhead computation incurred from unification of search space is confined to the syllable boundaries, the use of prosodic information to reduce the syllable boundary hypotheses as well as the syllable matching length is shown to be effective.The inherent complexity with the very large vocabulary is also reduced by the use of phrase boundary hypotheses conjectured via the phrase-final lengthening.Experimental results show a 47.2% recognition time save with only 5.67% error rate increase using the syllable and phrase boundary hypotheses conjectured from prosodic information.
Hung-Yun Hsieh, Ren-Yuan Lyu, Lin-Shan Lee
ICSLP3
1996 Robust speech recognition features based on temporal trajectory filtering of frequency band spectrum
abstract
This paper presents the use of a variety of filters in the temporal trajectories of frequency band spectrum to extract speech recognition features for environmental robustness.Three kind of filters for em- phasizing the statistically important parts of speech are proposed.First, a bank of RASTA-like band-pass filters to fit the statistical peaks of modulation frequency band spectrum of speech are used.Secondly, a three-channel octave band-iilter band with a smoothed rectangular window spline is applied.Thirdly, a datadriven filter is developed.Experimental results show that significant improvements for speech recognition using the proposed feature extraction approach under noisy environments can be achieved.
Jia-Lin Shen, Wen-Liang Hwang, Lin-Shan Lee
ICSLP3
1996 Speaker intention modeling for large vocabulary Mandarin spoken dialogues
abstract
This paper presents a statistical speaker intention modeling approach of speech act types (SAT's)[l] prediction for large vocabulary Mandarin spoken dialogues.A SAT is an abstraction of speaker's intention in terms of the type of action thax the speaker intends by the utterance.With this approach, spoken dialogue systems can be constructed to predict speaker's intention and make a proper action in advance.
Yen-Ju Yang, Lee-Feng Chien, Lin-Shan Lee
ICSLP3
1995 Golden Mandarin (III)-a user-adaptive prosodic-segment-based Mandarin dictation machine for Chinese language with very large vocabulary
abstract
This paper presents a prototype prosodic-segment-based Mandarin dictation machine for the Chinese language with very large vocabulary. It accepts utterances continuous within a prosodic segment which is composed of one or a few word(s). It also possesses various on-line learning capabilities for fast adaptation to a new user in acoustic, lexical and linguistic levels. The overall system is implemented on an IBM/PC with an additional DSP card including a Motorola DSP 96002 chip. The word accuracy can achieve nearly 90% for a new user after he produces about 10 minutes of speech to train the system, and the accuracy can be further improved with the on-line learning functions.
Ren-Yuan Lyu, Lee-Feng Chien, Shiao-Hong Hwang, Hung-Yun Hsieh, Rung-Chiuan Yang, Bo-Ren Bai, Jia-Chi Weng, Yen-Ju Yang, Shi-Wei Lin, Keh-Jiann Chen, Chiu-yu Tseng, Lin-Shan Lee
ICASSP12
1995 Complete recognition of continuous Mandarin speech for Chinese language with very large vocabulary but limited training data
abstract
This paper presents the first known results for complete recognition of continuous Mandarin speech for Chinese language with very large vocabulary but very limited training data. Although some isolated-syllable-based or isolated-word-based large-vocabulary Mandarin speech recognition systems have been successfully developed, a continuous-speech-based system of this kind has never been reported before. For successful development of this system, several important techniques have been used, including acoustic modeling of a set of sub-syllabic models for base syllable recognition and another set of context-dependent models for tone recognition, a multiple candidate searching technique based on a concatenated syllable matching algorithm to synchronize base syllable and tone recognition, and a word-class-based Chinese language model for linguistic decoding. The best recognition accuracy achieved is 88.69% for finally decoded Chinese characters, with 88.69%, 91.57%, and 81.37% accuracy for base syllables, tones, and tonal syllables respectively.
Hsin-Min Wang, Jia-Lin Shen, Yen-Ju Yang, Chiu-yu Tseng, Lin-Shan Lee
ICASSP5
1995 Large vocabulary, word-based Mandarin dictation system
Jung-Kuei Chen, Lin-Shan Lee, Frank K. Soong
EUROSPEECH2
1995 Fast and accurate continuous speech recognition for Chinese language with very large vocabulary
Tai-Hsuan Ho, Hsin-Min Wang, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH5
1995 A syllable-based very-large-vocabulary voice retrieval system for Chinese databases with textual attributes
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH4
1995 Unconstrained speech retrieval for Chinese document databases with very large vocabulary and unlimited domains
Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
EUROSPEECH4
1995 A chernoff distance based segmental probability model (CD-SPM) approach for Mandarin syllable recognition
Jia-Lin Shen, Lin-Shan Lee
EUROSPEECH2
1995 Practically realizable digital transmission significantly below the Nyquist bandwidth
abstract
Proposes a practically realizable digital signaling scheme that requires a bandwidth significantly below the Nyquist limit, and at the same time achieves the same asymptotic performance as memoryless PAM. A five-step iterative procedure for constructing such schemes is presented, and examples that occupy down to about 60 percent of the Nyquist bandwidth are demonstrated using this simple procedure. Practical receiver structures are also briefly discussed.>
Cheng-Kun Wang, Lin-Shan Lee
IEEE Trans. Commun.2
1994 Large vocabulary word recognition based on tree-trellis search
abstract
In this paper we propose a large vocabulary (90000 words), Chinese (Mandarin) word recognizer based on the tree-trellis fast search algorithm. The recognizer is divided into 3 modules: local likelihood computation, a forward trellis search and a backward tree search. In the forward trellis search, a free syllable decoding is performed without a language model and a partial path map is created. The best-first tree search is then applied backward along a lexicon, which is arranged as a syllabic tree, to find the N-best word candidates. In the experiment, context-dependent subsyllabic HMMs were trained with a new discriminative training method. When it is evaluated on a speaker-trained database, the recognizer achieved a word error rate of 5% for the full size (90000 words) vocabulary and 1.7% for a smaller subset (5000 words) vocabulary. A real-time demo system has also been implemented on an SGI R-4000 workstation.>
Jung-Kuei Chen, Frank K. Soong, Lin-Shan Lee
ICASSP (2)3
1994 An initial study on a segmental probability model approach to large-vocabulary continuous Mandarin speech recognition
abstract
This paper presents an initial study to perform large-vocabulary continuous Mandarin speech recognition based on a segmental probability model (SPM) approach. SPM was first proposed for recognition of isolated Mandarin syllables, in which every syllable must be equally segmented before recognition. A concatenated syllable matching algorithm is therefore introduced in place of the conventional Viterbi search algorithm to perform the recognition process based on SPM. In addition, a training procedure is also proposed to reestimate the SPM parameters for continuous speech. Preliminary simulation results indicate that significant improvements in both recognition rates and speed can be achieved as compared to the conventional HMM-based Viterbi search approaches.>
Jia-Lin Shen, Hsin-Min Wang, Bo-Ren Bai, Lin-Shan Lee
ICASSP (2)4
1994 Incremental speaker adaptation using phonetically balanced training sentences for Mandarin syllable recognition based on segmental probability models
Jia-Lin Shen, Hsin-Min Wang, Ren-Yuan Lyu, Lin-Shan Lee
ICSLP4
1994 An intelligent and efficient word-class-based Chinese language model for Mandarin speech recognition with very large vocabulary
Yen-Ju Yang, Sung-Chien Lin, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ICSLP5
1994 An exact performance analysis of the clipped diversity combining receiver for FH/MFSK systems against a band multitone jammer
abstract
This paper applies the clipper receiver with diversity combining techniques to the frequency hopping (FH) M-ary frequency shift keying (MFSK) systems against a destructive band multitone jammer, and presents an exact error probability performance evaluation with further discussions. A noise free channel is first assumed for analysis, and this assumption is then verified by computer simulation. Both optimal jamming tone power from the jammer's view and the choice of the clipping level from the communicator's view are investigated. The worst case performance of the FH/MFSK system against an optimum jammer is discussed in detail with extensive numerical results. It is also found that the clipper receiver provides the most attractive performance as compared to several other diversity combining receivers previously proposed.>
Jinn-Ja Chang, Lin-Shan Lee
IEEE Trans. Commun.2
1993 Golden Mandarin (II)-an improved single-chip real-time Mandarin dictation machine for Chinese language with very large vocabulary
Lin-Shan Lee, Chiu-yu Tseng, Keh-Jiann Chen, I-Jung Hung, Ming-Yu Lee, Lee-Feng Chien, Yumin Lee, Ren-Yuan Lyu, Hsin-Min Wang, Yung-Chuan Wu, Tung-Sheng Lin, Hung-Yan Gu, Chi-ping Nee, Chun-Yi Liao, Yeng-Ju Yang, Yuan-Cheng Chang, Rung-Chiung Yang
ICASSP (2)1
1993 A new framework for recognition of Mandarin syllables with tones using sub-syllabic units
Chih-Heng Lin, Lin-Shan Lee, Pei-Yih Ting
ICASSP (2)2
1993 Continuous hidden Markov models integrating transitional and instantaneous features for Mandarin syllable recognition
Yumin Lee, Lin-Shan Lee
Comput. Speech Lang.2
1993 A novel scaling scheme for fast Hartley transform
Guey-Shya Chen, Ja-Ling Wu, Wei-Jou Duh, Lin-Shan Lee
Signal Process.4
1993 A best-first language processing model integrating the unification grammar and Markov language model for speech recognition applications
abstract
A language processing model is proposed in which the grammatical approach of unification grammar and the statistical approach of Markov language models are properly integrated in a word lattice chart parsing algorithm with different best-first parsing strategies. This model has been successfully implemented in experiments on Mandarin speech recognition although it is language-independent. Test results show that significant improvements in both correct rate of recognition and computation speed can be achieved. A correct rate of 93.8% and 5 s per sentence on an IBM PC/AT, as compared with 73.8% and 25 s using unification grammar alone and 82.2% and 3 s using a Markov language model alone, was achieved. This high performance is due to the effective rejection of noisy word hypothesis interferences; that is, the unification-based grammatical analysis eliminates all illegal combinations, while the Markovian probabilities of constituents combined with the considerations on constituent length indicate the correct direction of processing.>
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
1993 Golden Mandarin (I)-A real-time Mandarin speech dictation machine for Chinese language with very large vocabulary
abstract
The first successfully implemented real-time Mandarin dictation machine, which recognizes Mandarin speech with very large vocabulary and almost unlimited texts for the input of Chinese characters into computers, is described. The machine is speaker-dependent, and the input speech is in the form of sequences of isolated syllables. The machine can be decomposed into two subsystems. The first subsystem recognizes the syllables using hidden Markov models. Because every syllable can represent many different homonym characters and form different multisyllabic words with syllables on its right or left, the second subsystem is needed to identify the exact characters from the syllables and correct the errors in syllable recognition. The real-time implementation is on an IBM PC/AT, connected to three sets of specially designed hardware boards on which seven TMS 320C25 chips operate in parallel. The preliminary test results indicate that it takes only about 0.45 s to dictate a syllable (or character) with an accuracy on the order of 90%.>
Lin-Shan Lee, Chiu-yu Tseng, Hung-Yan Gu, Fu-hua Liu, Robert Chen-Hao Chang, Yueh-hong Lin, Yumin Lee, Shih-Lung Tu, Shew-Heng Hsieh, Chian-hung Chen
IEEE Trans. Speech Audio Process.1
1993 Improved tone concatenation rules in a formant-based Chinese text-to-speech system
abstract
A set of improved tone concatenation rules to be used in a formant-based Mandarin Chinese text-to-speech system is presented. This system concatenates prestored syllables superimposed by additional tone patterns to obtain speech sentences for unlimited text, with the acoustic properties of each syllable modified by a set of synthesis rules. The tone concatenation rules are the most important among these synthesis rules, because they tell how the tone patterns for the syllables should be modified in an arbitrary sentence under various conditions of concatenating syllables of different tones on both sides. The improved tone concatenation rules are obtained empirically by carefully analyzing the tone pattern behavior under various tone concatenation conditions for many sentences in a database. A total of 14 representative tone patterns are defined for the five tones, and different rules about which pattern should be used under what kind of tone concatenation conditions are organized in detail. Preliminary subjective tests indicate that these rules actually give better synthesized speech for a formant-based Chinese text-to-speech system.>
Lin-Shan Lee, Chiu-yu Tseng, Ching Jiang Hsieh
IEEE Trans. Speech Audio Process.1
1993 A direct-concatenation approach to train hidden Markov models to recognize the highly confusing Mandarin syllables with very limited training data
abstract
The recognition of a total of 408 very confusing Mandarin syllables is very difficult because this vocabulary consists of 38 confusing sets, each of which can have as many as 19 syllables. The recognition of these 408 syllables becomes even more difficult when only very limited training data are available. A special direct-concatenation approach for training hidden Markov models (HMMs) to recognize these syllables with very limited training data is developed in which each syllable is divided into INITIAL and FINAL parts and 408 right-context-dependent INITIAL HMMs and 38 left-context-independent FINAL HMMs are separately trained and the transition region carefully taken account of, and then these INITIAL and FINAL HMMs are directly concatenated to form syllable recognition. Experimental results show that this approach can utilize the very limited training data most efficiently and provide significant improvements in recognition performance. Although the results are obtained for Mandarin syllables, the approach is believed to be equally helpful for the recognition of other confusing vocabularies.>
Fu-hua Liu, Yumin Lee, Lin-Shan Lee
IEEE Trans. Speech Audio Process.3
1991 A Preference-first Language Processor Integrating the Unification Grammar and Markov Language Model for Speech Recognition Applications
abstract
A language processor is to find out a most promising sentence hypothesis for a given word lattice obtained from acoustic signal recognition. In this paper a new language processor is proposed, in which unification grammar and Markov language model are integrated in a word lattice parsing algorithm based on an augmented chart, and the island-driven parsing concept is combined with various preference-first parsing strategies defined by different construction principles and decision rules. Test results show that significant improvements in both correct rate of recognition and computation speed can be achieved.
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ACL3
1991 An Efficient Natural Language Processing System Specially Designed for the Chinese Language
Lin-Shan Lee, Lee-Feng Chien, Long Ji Lin, Keh-Jiann Chen
Comput. Linguistics1
1991 An augmented chart data structure with efficient word lattice parsing scheme in speech recognition applications
Lee-Feng Chien, Lin-Shan Lee, Keh-Jiann Chen
Speech Commun.2
1990 An Augmented Chart Data Structure with Efficient Word Lattice Parsing Scheme In Speech Recognition Applications
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
COLING3
1990 An augmented chart parsing algorithm integrating unification grammar and Markov language model for continuous speech recognition
abstract
An efficient algorithm is developed to handle the difficulties in parsing noise word lattices (sets of word hypotheses obtained in continuous-speech recognition) which include problems such as word boundary overlapping, homonyms, lexical ambiguities, recognition uncertainty and errors, etc. An augmented chart is proposed, and the algorithms is then derived on this chart. This algorithm properly integrates the global structural synthesis capabilities of the unification grammar and the local relation estimation capabilities of the Markov language model. The parsing algorithm is island driven and best first. In this way, the features of the grammatical and statistical approaches can be combined, and the effects of the two different approaches are reflected in a single algorithm such that the overall selectivity can be appropriately optimized.>
Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee
ICASSP3
1990 A real-time Mandarin dictation machine for Chinese language with unlimited texts and very large vocabulary
abstract
A successfully implemented real-time Mandarin dictation machine which recognizes Mandarin speech with unlimited texts and very large vocabulary for the input of Chinese characters to computers is described. Isolated syllables including the tones are first recognized using specially trained hidden Markov models with special feature parameters. The exact characters are then identified from the syllables using a Markov Chinese language model. The real-time implementation is on an IBM PC/AT, connected to a set of special hardware boards on which ten TMS 320C25 chips operate in parallel. It takes only 0.45 s to dictate a character.>
Lin-Shan Lee, Chiu-yu Tseng, Hung-Yan Gu, Fu-hua Liu, Robert Chen-Hao Chang, Shew-Heng Hsieh, Chian-hung Chen
ICASSP1
1990 An Efficient speech Recognition System for the initials of Mandarin syllables
abstract
In a long-term research project, the recognition of Mandarin speech for very large vocabulary and unlimited text is considered. Its first stage goal is to recognize the Mandarin syllables. In a previous paper, an initial/final two-phase recognition approach to recognize these very confusing syllables was proposed, in which each syllable is divided into initial and final parts and recognized separately, and efficient recognition techniques for the finals were proposed and discussed. This paper serves as a continuation and proposes an efficient system to recognize the Mandarin initials. In this system, a classification procedure is first used to categorize the unknown initials into two groups C1 and C2; different approaches are then separately applied and independently optimized to recognize C1 and C2. It is found that Finite State Vector Quantization (FSVQ) is very useful, whose two modified versions, Modified FSVQ (MFSVQ) and the Second Order FSVQ (SOFSVQ), can provide the best recognition performance for C1 and C2 by carefully adjusting a design parameter called characteristic interval. Experimental results show that a recognition rate of 94.1% to 94.7% can be achieved using this system. Such a design is accomplished by carefully considering the special characteristics of Mandarin syllables and initials.
Pei-Yih Ting, Chiu-yu Tseng, Lin-Shan Lee
Int. J. Pattern Recognit. Artif. Intell.3
1990 A Mandarin Dictation Machine Based Upon a Hierarchical Recognition Approach and Chinese Natural Language Analysis
abstract
An experimental Mandarin dictation machine for inputting Mandarin speech (spoken Chinese language) into computers is described. Because of the special characteristics of the Chinese language, syllables are chosen as the basic units for dictation. The machine is designed based on a hierarchical language recognition approach in which acoustic signals are first recognized as a sequence of syllables, possible word hypotheses are then formed from the syllables, and the complete sentences are finally obtained. This approach is implemented by two subsystems. The first recognizes the syllables using speech signal processing techniques, the second subsystem then identifies the exact characters from the syllable and corrects the errors in syllable recognition. The detailed syllable recognition algorithms, word formation rules, parser, grammar, and the syntactic checking algorithms are described. With newspaper text in the form of isolated syllables as input, the preliminary test results indicate that such a dictation machine is not only practically attractive, but technically feasible.>
Lin-Shan Lee, Chiu-yu Tseng, Keh-Jiann Chen, Chia-Hwa Hwang, Pei-Yih Ting, Long Ji Lin
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Multi-H phase-coded modulations with asymmetric modulation indexes
abstract
Multi-H phase-coded modulation (MHPM) is a bandwidth-efficient modulation scheme which offers substantial coding gain over conventional digital modulation schemes. MHPM with asymmetric modulation indices corresponding to the bipolar data +1 and -1 is considered, and numerical results for the minimum Euclidean distances are provided. It is shown that performance improvements on the error probability over conventional MHPM are gained with essentially the same bandwidth and a very slight modification in implementation. The upper bounds on the error probabilities as functions of observation intervals and received E/sub b//N/sub 0/ are also investigated in detail. It is concluded that the concept of asymmetric modulation indices for MHPM is attractive for bandwidth and power-efficient modulation.>
Hong-Kuang Hwang, Lin-Shan Lee, Sin-Horng Chen
IEEE J. Sel. Areas Commun.2
1988 New speech recognition approaches based upon finite state vector quantization with structural constraints
abstract
The label-transition finite-state vector-quantization (FSVQ) algorithm is extensively explored to exhibit the power of finite-state machines for speech recognition. It is found that the FSVQ algorithm combined with special structural constraints can discriminate a finite set of candidates very successfully. All the consonant initials of isolated Mandarin monosyllables from designated speakers are used as the example vocabulary in the simulation. In addition to utilizing the first order memory provided by FSVQ on speech recognition, an experiment is conducted that expands the FSVQ to use the second-order memory and the dynamic relationship among the components of this three-vector group are used for recognition. The simulation results show that a slightly higher recognition rate (94.4%) is obtained with a consistent prediction interval.>
Pei-Yih Ting, Chiu-yu Tseng, Lin-Shan Lee
ICASSP3
1988 Efficient speech Recognition Techniques for the finals of Mandarin syllables
abstract
A long-term research project toward Mandarin speech recognition techniques for very large vocabulary and unlimited text is considered. By carefully examining the special structures of Chinese language, the first-stage goal is set to be the design of efficient techniques to recognize the finals of Mandarin syllables. In this paper, three special approaches to do this are proposed. The Segmental Model Approach defines the final models by dividing the finals into several segments according to the acoustic structures of the speech signals. The Three-pass Approach uses three consecutive passes to classify the finals into small sets and improve the recognition efficiency. The Multi-section Vector Quantization (MSVQ) Approach, on the other hand, significantly reduces the necessary computation time by incorporating the branch-and-bound algorithm and common codebook concept with the MSVQ techniques. Extensive computer simulations are performed first to optimize each approach by choosing the best set of parameters then to compare the performance of the three approaches. It was found that all the three approaches are very efficient in terms of relatively high recognition rate and short computation time, and the MSVQ Approach provides the highest recognition rate at the shortest computation time, thus it is most attractive.
Chia-Hwa Hwang, Yen-Ming Hsu, Biing-Chin Wang, Chiu-yu Tseng, Lin-Shan Lee
Int. J. Pattern Recognit. Artif. Intell.5
1988 An improved sequential estimation scheme for PN acquisition
abstract
An improved sequential estimation (ISE) pseudonoise (PN) acquisition scheme based on an extended characteristic polynomial is proposed in which the PN despreader can work or both the m sequence and the inverted m sequence. The scheme can be easily implemented by a front-end bit detector followed by full digital circuitry. The mean acquisition time of ISE is derived.>
Lin-Shan Lee, Jung-Hui Chiu
IEEE Trans. Commun.1
1987 The Preliminary Results of a Mandarin Dictation Machine Based Upon Chinese Natural Language Analysis
Lin-Shan Lee, Chiu-yu Tseng, Keh-Jiann Chen
IJCAI1
1987 The Minimum Likelihood-A New Concept for Bit Synchronization
abstract
A new bit synchronization concept based on the "minimum likelihood" criterion instead of the conventional "maximum likelihood" concept is developed. The minimum likelihood situation is even easier to reach than the maximum likelihood because the derivative of the log likelihood function becomes identically zero there. Minimum likelihood implies "least likely" for synchronization (the worst case synchronization error) or an "orthogonal" timing condition which simply means that the locally generated clock is synchronized correctly, but with a delay of a half bit period. The structure and performance of the minimum likelihood bit synchronizer are discussed in detail in this paper. The results indicate that the minimum likelihood bit synchronizer has a much simpler structure, but with performance very close to the optimal maximum likelihood synchronizer.
Jung-Hui Chiu, Lin-Shan Lee
IEEE Trans. Commun.2
1986 A Chinese Natural Language Processing System Based Upon the Theory of Empty Categories
Long Ji Lin, Lin-Shan Lee, Keh-Jiann Chen
AAAI2
1986 A Chinese text-to-speech system based upon a syllable concatenation model
abstract
An unrestricted Chinese text-to-speech system has been successfully constructed. The text input which is written in Mandarin Phonetic Symbols II (MPS II) instead of traditional Chinese orthography is analyzed by a simple parser. Then each syllable in the text is assigned proper duration and pitch, while its gain is modified, and its reflection coefficients interpolated or decimated. The stress, intonation and inter-syllable pause are also considered. In other words, a model of concatenating syllables of Mandarin Chinese based on simplified phonetic rules and some preliminary analysis of natural speech data is proposed. We believe that our model is sufficient compared with some more complicated approaches [1]. After editing the syllables according to our proposed synthesis rules, a continuous speech signal is synthesized, and then transmitted to speech interface card for speech output. Samples of synthesized speech by our text-to-speech system have been tested on subjects with satisfactory results. Evaluation of the system in intelligibility, comprehensibility and fluency has been conducted and an average of 96%, 95.9% and 75% respectively was reached.
Ouhyoung Ming, Chin-jiang Shie, Chiu-yu Tseng, Lin-Shan Lee
ICASSP4
1986 Closed-Form Statistical Analysis for Square Law PN Acquisition. Detector Performance in Spread Spectrum Systems
Jung-Hui Chiu, Lin-Shan Lee
ICC2
1986 A General Theory for Asynchronous Speech Encryption Techniques
abstract
Speech encryption techniques have always been very important for military communications, but most useful techniques require perfect synchronization between the transmitter and the receiver. This not only complicates the implementation, but makes the transmission very sensitive to channel conditions because slight synchronization error might completely break the transmission. Two special techniques were proposed recently in which the synchronization becomes completely unnecessary. This improves the feasibility and reliability tremendously. In this paper, a general theory for such "asynchronous speech encryption techniques" is developed in detail, starting by defining the asynchronous approach and model and ending with a general solution. It will be found that the two techniques proposed earlier become two special cases of the general solution here.
Lin-Shan Lee, Ger-Chih Chou
IEEE J. Sel. Areas Commun.1
1984 A New Time Domain Speech Scrambling System Which Does Not Require Frame Synchronization
abstract
Communications security is becoming more and more Important today. There is thus an increasing interest in analog scramblers due to the desire for secure speech communications over existing telephone channels with standard telephone bandwidth at acceptable speech quality and reasonable cost. The concept of scrambling the sample values of the speech waveforms is attractive due to its higher degree of security compared to traditional scramblers. But all these sample value scramblers require frame synchronization, i.e., the signal segments used in scrambling and descrambling processes have to be exactly the same for signal recovery. This not only complicates the implementation but makes the transmission very sensitive to channel conditions. A new frequency domain scrambling system was proposed recently, which eliminates the requirement for frame synchronization. However, FFT algorithms are used in this system to transform speech samples back and forth between time and frequency domains. This not only requires higher complexity and cost, but introduces significant roundoff errors and enhances the channel noise and distortion. In this paper, a new time domain scrambling system was proposed which eliminates all the FFT's in the previous system but still does not require frame synchronization. This simplifies the implementation and improves the quality and reliability. The theoretical derivation, synchronization analysis, software simulation, hardware implementation, and residual intelligibility tests of this new system are discussed in great detail in this paper.
Lin-Shan Lee, Ger-Chih Chou
IEEE J. Sel. Areas Commun.1
1984 A New Formulation of Spectrum-Orbit Utilization Efficiency for Satellite Communications in Interference-Limited Situations
abstract
The spectrum-orbit utilization efficiency of satellite communications is becoming increasingly important. A first attempt is made in this paper to try to formulate the spectrum-orbit utilization efficiency in interference-limited situations by defining the self-efficiency for a single satellite, and the cross-efficiency and coordination efficiency for each pair of adjacent satellites. The various technical aspects and complicated interference considerations relevant to spectrum-orbit utilization can all be reasonably reflected by these simple efficiency parameters. They tell the behavior of each satellite by itself and with its neighbors. Such a formulation is therefore very useful for improving the spectrum-orbit utilization efficiency in future satellite communications.
Lin-Shan Lee
IEEE Trans. Commun.1
1984 A New Frequency Domain Speech Scrambling System Which Does Not Require Frame Synchronization
abstract
Communication security is becoming more and more important today. There is thus an increasing interest in analog scramblers due to the desire for secure speech communications over existing telephone channels with standard telephone bandwidth at acceptable speech quality and reasonable cost. The concept of scrambling the sample values of the speech waveforms becomes attractive due to its higher degree of security compared to the traditional scramblers. But all these sample value scramblers require frame synchronization, i.e., the signal segments used in scrambling and descrambling processes have to be exactly the same for signal recovery. This complicates the implementation and makes the transmission very sensitive to channel conditions. In this paper, a new frequency domain scrambling algorithm is presented, which is an extension of the DFT scrambler previously proposed. The use of short-time Fourier analyis and the filter bank concept leads to the special feature that frame synchronization is completely unnecessary. This simplifies the implementation and improves the reliability and feasibility. The theoretical developments, simulation results, hardware implementation, and test results are all discussed in detail.
Lin-Shan Lee, Ger-Chih Chou, Ching-Sung Chang
IEEE Trans. Commun.1
1984 An A, B-Polar Approach for Multimode Polarization Analysis in Satellite Communications
abstract
This paper describes a convenient method of polarization analysis for a circular multimode aperture antenna by decomposing the field distribution into two components calledA,B-polar. The new idea in this technique is the separation of theA,B-polar for field expression into a rotationally symmetric magnitude function and an orientation function independent of the antenna off-axis. This leads to easily sketched figures of the multimode polarization patterns which give a clear physical picture on their effects and applications. Several examples are done in the context of the frequency reuse capabilities of earth-station antennas.
Fan-Tung Tseng, Lin-Shan Lee
IEEE Trans. Commun.2
1984 A Differential Technique for Detections of Circularly Polarized Satellite Signal Parameters
abstract
This paper presents fundamental concepts, operational principles, and typical experimental results of a differential technique for detections of circularly polarized satellite signal parameters at low axial ratios. The approach is to equip a rotating π-polarizer with an orthomode transducer in the antenna front end. A two-channel phase-locked receiver is used to detect the differential phase (DP) and the differential level (DL) between the two output channels. The required parameters are then readily available from the displays of the DP and DL values on a strip chart recorder.
Fan-Tung Tseng, Lin-Shan Lee
IEEE Trans. Commun.2
1982 Accurate Detection of Satellite Beacon Polarization States Using Cascaded Heterodyne Phase-Locked Loops
abstract
The use of orthogonal polarizations sharing the same frequency is attractive in satellite communications due to its efficiency in utilizing the spectrum. The main problem remaining to be solved is that of the various cross polarization effects which cause interference between the orthogonal channels. The technique using cascaded heterodyne phase-locked loops (PLL's) to accurately detect the phase as well as amplitude of both copolarized and cross-polarized satellite beacon carriers under poor carrier-to-noise ratio conditions is developed and presented in this paper. Such a technique is very useful not only in satellite beacon measurements and cross polarization experiments, but also in the implementation of cross polarization restoration networks currently developed for satellite communication systems.
Fan-Tung Tseng, Lin-Shan Lee
IEEE Trans. Commun.2
1981 The Feasibility of Two One-Parameter Polarization Control Methods in Satellite Communications
abstract
The orthogonal polarization techniques will be widely used in satellite communications due to its efficient utilization of the spectrum, but the feasibility of different methods to compensate for the cross-polarization at ground stations will be essential for the applicability of the techniques. Although complicated four-variable systems have been designed by many groups, two simple one-parameter methods have been proposed recently. The first, rotational compensation, simply rotates the linear polarization directions of the receiving antenna to maximize XPD; the second, quadrature cancellation, simply injects quadrature cancelling signals. This paper studies the feasibility of these two methods in terms of system designer's viewpoints. A correlation approach is used to estimate the achievable XPD values for different situations including rain, ice particles, and Faraday rotation. The system availability and cost considerations are analyzed using a possible tradeoff between CNR and XPD values. The control errors and stability problems are also discussed. The results indicate that there exist many situations in which these one-parameter methods are feasible and attractive and provide possible solutions for designers desiring effective, low-cost receiving stations with satisfactory performance.
Lin-Shan Lee
IEEE Trans. Commun.1
1979 The Vector Space Formulation of the Rain Crosspolarization Problem and Its Compensations
abstract
The rain crosspolarization problem is formulated in a vector space. In this formulation signals are vectors, and the crosspolarization effect is an operator. The compensators currently designed for satellite communications become the inverse operator approach in this formulation. A new approach using eigenvectors is developed. The result indicates that use of two orthogonal linear polarizations and a rotation of their directions can eliminate the crosspolatization. A feedback loop can be used to control the rotation angle, and only one control variable is sufficient. Even if no satellite applications have been found, good potential in terrestrial systems is expected.
Lin-Shan Lee
IEEE Trans. Commun.1
1979 Polarization Control Schemes For Satellite Communications with Multiple Uplinks
abstract
Orthogonal polarizations will be more and more widely used in satellite communications in the near future because of their efficiency in utilizing the spectrum, but the currently designed compensation control systems can compensate for signals coming from only one uplink station. This paper therefore discusses generally the control schemes feasible to solve the multiple Uplink problem, in which there are many stations transmitting uplink signals simultaneously. The general approach proposed is that for each uplink station two Compensators can be used to compensate for the uplink and, downlink separately. This will require the station to transmit its own pilot signals and receive them back from the satellite as the control reference. Two basic techniques can be used to realize this approach. The first, correlation technique, is based on the assumption that an obtainable correlation between uplink and downlink cross-polarizations exists, and tries to compensate for uplink and downlink separately using this correlation. The second, separate link technique, uses an additional pair of pilotsignals to obtain seiJarate control references and decouples the compensations of uplink and downlink. The discussions indicate that the separate downlink technique, whose design considerations and system performance are examined carefully is most attractive.
Lin-Shan Lee
IEEE Trans. Commun.1
1979 New Radiation Limits For a Digital Radio World
abstract
Regulations on radiation limits have been Widely adopted to keep radio transmissions from interfering with each other. Most of them set an upper limit on the total power flux density radiated Within a certain frequency bandwidth,B. But new techniques have given rise to the possibility that some digital transmission systems could cause significant interference to other transmissions without violating the current regulations. This is due to the fact that bursty signals can concentrate most of their energy in a very short pulse with very high power but maintain a low enough average power spectral density, thus meeting the current regulations. New radiation limits are therefore suggested in this paper which set an upper bound on the total energy instead of average power radiated within a certain time period,H, as well as a certain frequency bandwidth,B. It is shown that the current regulations can be easily extended to establish the new limits, and these new limits are essentially identical to the current limits for the traditional continuous-wave signals but can limit the bursty digital signals efficiently as well. Practical consideration for the appropriate length of the time period,H, to be used are discussed, and a few example signals are used to demonstrate the use of the new limits.
Lin-Shan Lee, James M. Janky
IEEE Trans. Commun.1
1978 A Polarization Control System for Satellite Communications with Multiple Uplinks
abstract
Several groups have designed polarization control systems to solve the rain-crosspolarization problem in satellite communications. However, the multiple-uplink problem still exists, i.e., one system cannot receive signals from different uplink stations with different rain conditions simultaneously. In order to solve this problem, a new scheme is developed in this paper based on the assumption that an obtainable correlation exists between the uplink and downlink crosspolarizations and the concept that uplink and downlink can be compensated separately. The currently designed systems can be used directly in this scheme; only the addition of similar devices is necessary. The feasibility is checked by a practical example, and satisfactory performance is observed.
Lin-Shan Lee
IEEE Trans. Commun.1