Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sin-Horng Chen

dblp:00/3107 · DBLP profile ↗
← Back
68ranked-venue papers
15as first author
2since 2021 · last 2023
0000-0002-9820-2318ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 37 · 10 first-authorComputer networks · 5 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Speech recognition and synthesis · 95% Language models and text generation · 2% Information extraction and text analysis · 2%
Computer graphics and multimedia
2 papers
Audio and music processing · 78% Image and video coding · 14% Image and video processing · 8%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.542016
Speaker Adaptation of SR-HPM for Speaking Rate-Controlled Mandarin TTS · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Modeling of Speaking Rate Influences on Mandarin Speech Prosody and Its Application to Speaking Rate-controlled TTS · IEEE ACM Trans. Audio Speech Lang. Process. 2014
A new duration modeling approach for Mandarin speech · IEEE Trans. Speech Audio Process. 2003
Natural language and speech › Speech recognition and synthesis › speech synthesis
speaking rate control
0.422016
Speaker Adaptation of SR-HPM for Speaking Rate-Controlled Mandarin TTS · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Modeling of Speaking Rate Influences on Mandarin Speech Prosody and Its Application to Speaking Rate-controlled TTS · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
prosody modeling
0.432014
Modeling of Speaking Rate Influences on Mandarin Speech Prosody and Its Application to Speaking Rate-controlled TTS · IEEE ACM Trans. Audio Speech Lang. Process. 2014
A New Prosody-Assisted Mandarin ASR System · IEEE Trans. Speech Audio Process. 2012
A new duration modeling approach for Mandarin speech · IEEE Trans. Speech Audio Process. 2003
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
speaker adaptation
0.212016
Speaker Adaptation of SR-HPM for Speaking Rate-Controlled Mandarin TTS · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.252012
A New Prosody-Assisted Mandarin ASR System · IEEE Trans. Speech Audio Process. 2012
A modular RNN-based method for continuous Mandarin speech recognition · IEEE Trans. Speech Audio Process. 2001
An RNN-based preclassification method for fast continuous Mandarin speech recognition · IEEE Trans. Speech Audio Process. 1998
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
mandarin speech recognition
0.132001
A modular RNN-based method for continuous Mandarin speech recognition · IEEE Trans. Speech Audio Process. 2001
Tone recognition of continuous Mandarin speech based on neural networks · IEEE Trans. Speech Audio Process. 1995
An RNN-based preclassification method for fast continuous Mandarin speech recognition · IEEE Trans. Speech Audio Process. 1998
Natural language and speech › Language models and text generation
language modeling
0.012012
A New Prosody-Assisted Mandarin ASR System · IEEE Trans. Speech Audio Process. 2012
Natural language and speech › Information extraction and text analysis › temporal information extraction
duration modeling
0.012003
A new duration modeling approach for Mandarin speech · IEEE Trans. Speech Audio Process. 2003
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
continuous speech recognition
0.011998
An RNN-based preclassification method for fast continuous Mandarin speech recognition · IEEE Trans. Speech Audio Process. 1998
Machine learning › Deep learning architectures and training
recurrent neural network
0.011998
An RNN-based prosodic information synthesizer for Mandarin text-to-speech · IEEE Trans. Speech Audio Process. 1998
Natural language and speech › Speech recognition and synthesis
search space reduction
0.011998
An RNN-based preclassification method for fast continuous Mandarin speech recognition · IEEE Trans. Speech Audio Process. 1998
Audio and music processing
speech coding
0.021994
Variable frame rate speech coding using optimal interpolation · IEEE Trans. Commun. 1994
Vector quantization of pitch information in Mandarin speech · IEEE Trans. Commun. 1990
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
tone recognition
0.011995
Tone recognition of continuous Mandarin speech based on neural networks · IEEE Trans. Speech Audio Process. 1995
Audio and music processing › speech coding
linear predictive coding
0.011994
Variable frame rate speech coding using optimal interpolation · IEEE Trans. Commun. 1994
Audio and music processing
speech processing
0.011990
Vector quantization of pitch information in Mandarin speech · IEEE Trans. Commun. 1990
Image and video coding › quantization
vector quantization
0.011990
Vector quantization of pitch information in Mandarin speech · IEEE Trans. Commun. 1990
Physical-layer communications › modulation › continuous phase modulation
multi-h modulation
0.011989
Multi-H phase-coded modulations with asymmetric modulation indexes · IEEE J. Sel. Areas Commun. 1989
Integrated circuit design › digital signal processing circuits
digital filter
0.011988
Efficient bit-level systolic array implementation of FIR and IIR digital filters · IEEE J. Sel. Areas Commun. 1988
Cloud and datacenter computing › resource provisioning
dynamic resource provisioning
0.011988
Efficient bit-level systolic array implementation of FIR and IIR digital filters · IEEE J. Sel. Areas Commun. 1988
Hardware accelerators and domain-specific architectures
systolic array
0.011988
Efficient bit-level systolic array implementation of FIR and IIR digital filters · IEEE J. Sel. Areas Commun. 1988
Image and video processing › video frame interpolation
interpolation
0.011994
Variable frame rate speech coding using optimal interpolation · IEEE Trans. Commun. 1994
Physical-layer communications › modulation
bandwidth-efficient modulation
0.011989
Multi-H phase-coded modulations with asymmetric modulation indexes · IEEE J. Sel. Areas Commun. 1989
Integrated circuit design › digital arithmetic circuits
bit-level systolic array
0.011988
Efficient bit-level systolic array implementation of FIR and IIR digital filters · IEEE J. Sel. Areas Commun. 1988

Methods — techniques the papers use, named apart from their topics

hierarchical prosodic model · 0.4structural maximum a posteriori · 0.2decision tree parameter organization · 0.2feature normalization · 0.2data-driven modeling · 0.2hidden markov model · 0.2lattice rescoring · 0.1joint prosody labeling and modeling · 0.1recurrent neural network · 0.1EM algorithm · 0.0optimal interpolation · 0.0contour quantization · 0.0orthogonal polynomial representation · 0.0minimum euclidean distance analysis · 0.0error probability bounds · 0.0systolic array · 0.0inner-product computation · 0.0
YearPublicationVenuePosition
2023 Toward enriched decoding of mandarin spontaneous speech
Yu-Chih Deng, Yuan-Fu Liao, Yih-Ru Wang, Sin-Horng Chen
Speech Commun.4
2023 VoiceTalk: Multimedia-IoT Applications for Mixing Mandarin, Taiwanese, and English
abstract
The voice-based Internet of Multimedia Things (IoMT) is the combination of IoT interfaces and protocols with associated voice-related information, which enables advanced applications based on human-to-device interactions. An example is Automatic Speech Recognition (ASR) for live captioning and voice translation. Three major issues of ASR for IoMT are IoT development cost, speech recognition accuracy, and execution time complexity. For the first issue, most non-voice IoT applications are upgraded with the ASR feature through hard coding, which are error prone. For the second issue, recognition accuracy must be improved for ASR. For the third issue, many multimedia IoT services are real-time applications and, therefore, the ASR delay must be short. This article elaborates on the above issues based on an IoT platform called VoiceTalk. We built the largest Taiwanese spoken corpus to train VoiceTalk ASR (VT-ASR) and show how the VT-ASR mechanism can be transparently integrated with existing IoT applications. We consider two performance measures for VoiceTalk: speech recognition accuracy and VT-ASR delay. For the acoustic tests of PAL-Labs, VT-ASR's accuracy is 96.47%, while Google's accuracy is 94.28%. We are the first to develop an analytic model to investigate the probability that the VT-ASR delay for the first speaker is complete before the second speaker starts talking. From the measurements and analytic modeling, we show that the VT-ASR delay is short enough to result in a very good user experience. Our solution has won several important government and commercial TV contracts in Taiwan. VT-ASR has demonstrated better Taiwanese Mandarin speech recognition accuracy than famous commercial products (including Google and Iflytek) in Formosa Speech Recognition Challenge 2018 (FSR-2018) and was the best among all participating ASR systems for Taiwanese recognition accuracy in FSR-2020.
Yi-Bing Lin, Yuan-Fu Liao, Sin-Horng Chen, Shaw-Hwa Hwang, Yih-Ru Wang
ACM Trans. Internet Techn.3
2018 An Exploration of Local Speaking Rate Variations in Mandarin Read Speech
Guan-Ting Liou, Chen-Yu Chiang, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH4
2017 Fast and accurate image recognition using Deeply-Fused Branchy Networks
abstract
In order to achieve higher accuracy of image recognition, deeper and wider networks have been used. However, when the network size gets bigger, its forward inference time also takes longer. To address this problem, we propose Deeply-Fused Branchy Network (DFB-Net) by adding small but complete side branches to the target baseline main branch. DFB-Net allows easy-to-discriminate samples to be classified faster. For hard-to-discriminate samples, DFB-Net makes probability fusion by averaging softmax probabilities to make collaborative predictions. Extensive experiments on the two CIFAR datasets show that DFB-Net achieves state-of-the-art results to obtain an error rate of 3.07% on CIFAR-10 and 16.01% on CIFAR-100. Meanwhile, the forward inference time (with a batch size of 1 and averaged among all test samples) only takes 10.4 ms on CIFAR-10, 18.8 ms on CIFAR-100, using GTX 1080 GPU with cuDNN 5.1.
Mou-Yue Huang, Ching-Hao Lai, Sin-Horng Chen
ICIP3
2016 Structural maximum a posteriori speaker adaptation of speaking rate-dependent hierarchical prosodic model for Mandarin TTS
abstract
In this paper, a structural maximum a posterior speaker adaptation method to adjust the existing speaking rate (SR) dependent hierarchical prosodic model (SR-HPM) to a new speaker's data for realizing a new voice of any given SR is discussed. The adaptive SR-HPM is formulated based on MAP estimation with a reference SR-HPM serving as an informative prior. The prior information provided by the reference SR-HPM is hierarchically organized by decision trees. The results of objective and subjective evaluations showed that the proposed method not only performed slightly better than the maximum likelihood-based model in the observed SR range of the target speaker's data, but also was much better in the unseen SR range.
I-Bin Liao, Chen-Yu Chiang, Sin-Horng Chen
ICASSP3
2016 Speaker Adaptation of SR-HPM for Speaking Rate-Controlled Mandarin TTS
abstract
In this paper, a structural maximum a posteriori (SMAP) speaker adaptation approach to adjusting the speaking rate (SR)-dependent hierarchical prosodic model (SR-HPM) of an existing SR-controlled Mandarin text-to-speech system to a new speaker's data for producing a new voice is discussed. Two main issues are addressed. One is the small SR coverage of the adaptation data and is solved by using the existing SR-HPM that was trained from a speech corpus of wide SR coverage as an informative prior. Another is the data sparseness problem resulting from the large number of parameters of the SR-HPM to be adjusted. It is solved by hierarchically organizing the SR-HPM parameters into decision trees so as to be efficiently adjusted by the SMAP method. The effectiveness of the proposed approach is evaluated on speech databases of five new speakers. Both objective and subjective evaluations show that the proposed method not only performs better than the maximum likelihood-based method in the observed SR range of the target speaker's data, but also is much better in the unseen SR ranges.
I-Bin Liao, Chen-Yu Chiang, Yih-Ru Wang, Sin-Horng Chen
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Modeling of Speaking Rate Influences on Mandarin Speech Prosody and Its Application to Speaking Rate-controlled TTS
abstract
A new data-driven approach to building a speaking rate-dependent hierarchical prosodic model (SR-HPM), directly from a large prosody-unlabeled speech database containing utterances of various speaking rates, to describe the influences of speaking rate on Mandarin speech prosody is proposed. It is an extended version of the existing HPM model which contains 12 sub-models to describe various relationships of prosodic-acoustic features of speech signal, linguistic features of the associated text, and prosodic tags representing the prosodic structure of speech. Two main modifications are suggested. One is designing proper normalization functions from the statistics of the whole database to compensate the influences of speaking rate on all prosodic-acoustic features. Another is modifying the HPM training to let its parameters be speaking-rate dependent. Experimental results on a large Mandarin read speech corpus showed that the parameters of the SR-HPM together with these feature normalization functions interpreted the effects of speaking rate on Mandarin speech prosody very well. An application of the SR-HPM to design and implement a speaking rate-controlled Mandarin TTS system is demonstrated. The system can generate natural synthetic speech for any given speaking rate in a wide range of 3.4-6.8 syllables/sec. Two subjective tests, MOS and preference test, were conducted to compare the proposed system with the popular HTS system. The MOS scores of the proposed system were in the range of 3.58-3.83 for eight different speaking rates, while they were in 3.09-3.43 for HTS. Besides, the proposed system had higher preference scores (49.8%-79.6%) than those (9.8%-30.7%) of HTS. This confirmed the effectiveness of the speaking rate control method of the proposed TTS system.
Sin-Horng Chen, Chiao-Hua Hsieh, Chen-Yu Chiang, Hsi-Chun Hsiao, Yih-Ru Wang, Yuan-Fu Liao, Hsiu-Min Yu
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 A speaking rate-controlled Mandarin TTS system
abstract
In this paper, a new speaking rate-controlled Mandarin TTS system based on a speaking rate-dependent hierarchical prosodic model (SR-HPM) [6] is proposed. In the training phase, a data-driven approach is employed to automatically build the SR-HPM directly from a large prosody-unlabeled speech database containing utterances of various speaking rates. The SR-HPM comprises 15 sub-models designed to describe various relationships among 3 types of prosodic-acoustic features of speech utterances, two types of prosodic tags specifying a 4-layer prosody hierarchy, linguistic features of various levels of the associated texts, and the speaking rates. In the test phase, the SR-HPM is employed to generate 4 prosodic-acoustic features, including syllable pitch contours, syllable durations, syllable energy levels, and syllable juncture pause durations. Combining these prosodic features with the spectral features generated by the HTS synthesizer, the system can generate natural speech for any speaking rate in a wide range of 0.15-0.3 seconds/syllable. A distinct feature of the system to control the occurrence frequencies of breaks of various types as well as their pause durations according to the given speaking rate was demonstrated. A subjective test showed that MOS scores of 3.35, 3.44 and 3.28 were achieved respectively for fast (SR=0.17 sec/syllable), medium (SR=0.2 sec/syllable) and slow (SR=0.25 sec/syllable) synthetic speeches.
Chiao-Hua Hsieh, Yih-Ru Wang, Chen-Yu Chiang, Sin-Horng Chen
ICASSP4
2013 Knowledge integration for improving performance in LVCSR
abstract
This paper presents a knowledge integration framework to improve performance in large vocabulary continuous speech recognition. Two types of knowledge sources, manner attribute and prosodic structure, are incorporated. For manner of articulation, six attribute detectors trained with an American English corpus (WSJ0) are utilized to rescore hypothesized phones in word lattices obtained by a baseline ASR system. For the prosodic structure, models trained with an unsupervised joint prosody labeling and modeling (PLM) technique using WSJ0 are used in lattice rescoring. Experimental results on the American English WSJ word recognition task of the Nov92 test set show that the proposed approach significantly outperforms the baseline system that does not use articulatory and prosodic information. The results also demonstrate the effectiveness and usefulness of the PLM technique in constructing prosodic models for American English ASR.
Chen-Yu Chiang, Sabato Marco Siniscalchi, Sin-Horng Chen, Chin-Hui Lee 0001
INTERSPEECH3
2013 Alleviating the over-smoothing problem in GMM-based voice conversion with discriminative training
abstract
In this paper, we propose a discriminative training (DT) method to alleviate the muffled sound effect caused by over smoothing in the Gaussian mixture model (GMM)-based voice conversion (VC). For the conventional GMM-based VC, we often observed a large degree of ambiguities among acoustic classes (generative classes), determined by the source feature vectors for generating the converted feature vectors, causing the “muffled sound” effect on the converted voice. The proposed DT method is applied to refine the parameters in the maximum likelihood (ML)-trained joint density GMM (JDGMM) in the training stage to reduce the ambiguities among acoustic classes (generative classes) to alleviate the muffled sound effect. Experimental results demonstrate that the DT method significantly enhances the discriminative power between acoustic classes (generative classes) in the objective evaluation and effectively alleviates the muffled sound effect in the subjective evaluation.
Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH5
2013 A Lattice-Reduction Aided List Demapper for Coded MIMO Receiver
abstract
The max-log list demapper has been widely employed in the implementations of a coded multiple-input multiple-output (MIMO) receiver, where only a candidate-list of signal vectors is examined in the likelihood-ratio calculation to reduce complexity. Traditionally, the candidate-list is generated in the original-lattice domain which, unfortunately, results in a severe degradation in the performance of demapper if the channel is in ill-condition. In this paper, a new lattice-reduction aided max-log list demapper is proposed in which the candidate-list is generated after a successive cancellation of multi-layer interference in the lattice-reduced domain. Thanks to the newly designed metrics and search algorithms for the generation of the candidate-list, the proposed demapper provides significant gains over the existing methods, especially for the cases with a small list size and/or under a spatially-correlated channel
Tung-Jung Hsieh, Wern-Ho Sheen, Sin-Horng Chen, Jen-Yuan Hsu
VTC Fall3
2012 Punctuation generation inspired linguistic features for mandarin prosodic boundary prediction
abstract
A novel statistical linguistic feature, called punctuation confidence, is proposed in this paper for assisting in prosodic break prediction in Mandarin text-to-speech. The punctuation confidence calculated from the input text is a measure of the likelihood of inserting a major PM at a word boundary. Since a punctuation in text tends to be pronounced as a break, the punctuation confidence associated with a punctuation estimate should provide useful information for break prediction from text. The idea is realized in this study by first employing a conditional random field (CRF)-based model to generate a predicted punctuation and its associated punctuation confidence for each word boundary. Then, the predicted punctuation and its punctuation confidence are combined with contextual linguistic features to predict the break type of the word boundary by an MLP (multi-layer perceptrons). Experiment on the Treebank speech corpus confirmed the effectiveness of the proposed approach.
Chen-Yu Chiang, Yih-Ru Wang, Sin-Horng Chen
ICASSP3
2012 A New Approach of Speaking Rate Modeling for Mandarin Speech Prosody
Chiao-Hua Hsieh, Chen-Yu Chiang, Yih-Ru Wang, Hsiu-Min Yu, Sin-Horng Chen
INTERSPEECH5
2012 A Study of Mutual Information for GMM-Based Spectral Conversion
abstract
The Gaussian mixture model (GMM)-based method has dominated the field of voice conversion (VC) for last decade. However, the converted spectra are excessively smoothed and thus produce muffled converted sound. In this study, we improve the speech quality by enhancing the dependency between the source (natural sound) and converted feature vectors (converted sound). It is believed that enhancing this dependency can make the converted sound closer to the natural sound. To this end, we propose an integrated maximum a posteriori and mutual information (MAPMI) criterion for parameter generation on spectral conversion. Experimental results demonstrate that the quality of converted speech by the proposed MAPMI method outperforms that by the conventional method in terms of formal listening test.
Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH5
2012 A New Prosody-Assisted Mandarin ASR System
abstract
This paper presents a new prosody-assisted automatic speech recognition (ASR) system for Mandarin speech. It differs from the conventional approach of using simple prosodic cues on employing a sophisticated prosody modeling approach based on a four-layer prosody-hierarchy structure to automatically generate 12 prosodic models from a large unlabeled speech database by the joint prosody labeling and modeling (PLM) algorithm proposed previously. By incorporating these 12 prosodic models into a two-stage ASR system to rescore the word lattice generated in the first stage by the conventional hidden Markov model (HMM) recognizer, we can obtain a better recognized word string. Besides, some other information can also be decoded, including part of speech (POS), punctuation mark (PM), and two types of prosodic tags which can be used to construct the prosody-hierarchy structure of the testing speech. Experimental results on the TCC300 database, which consists of long paragraphic utterances, showed that the proposed system significantly outperformed the baseline scheme using an HMM recognizer with a factored language model which models word, POS, and PM. Performances of 20.7%, 14.4%, and 9.6% in word, character, and base-syllable error rates were obtained. They corresponded to 3.7%, 3.7%, and 2.4% absolute (or 15.2%, 20.4%, and 20% relative) error reductions. By an error analysis, we found that many word segmentation errors and tone recognition errors were corrected.
Sin-Horng Chen, Jyh-Her Yang, Chen-Yu Chiang, Ming-Chieh Liu, Yih-Ru Wang
IEEE Trans. Speech Audio Process.1
2011 Enriching Mandarin speech recognition by incorporating a hierarchical prosody model
abstract
This paper presents a new probabilistic framework of Mandarin speech recognition by incorporating a sophisticated hierarchical prosody model into the conventional HMM-based system. The prosody model describes the relations of linguistic cues of various levels, break types and prosodic states which represent the prosody hierarchical structure, and prosody-related acoustic features. Aside from producing the recognized word sequences, the system also decodes other information including word's part-of-speech, punctuation marks, inter-syllable break types, and prosodic states of syllables. Experimental results on the TCC300 corpus, which consists of paragraphic utterances, showed that the proposed system significantly outperformed the baseline system. The word and character error rates decreased from 24.4% and 18.1% to 20.7% and 14.4% (or 15.2% and 20.4% relative improvements), respectively.
Jyh-Her Yang, Ming-Chieh Liu, Hao-Hsiang Chang, Chen-Yu Chiang, Yih-Ru Wang, Sin-Horng Chen
ICASSP6
2011 A New Model-Based Mandarin-Speech Coding System
Chen-Yu Chiang, Jyh-Her Yang, Ming-Chieh Liu, Yih-Ru Wang, Yuan-Fu Liao, Sin-Horng Chen
INTERSPEECH6
2009 Advanced unsupervised joint prosody labeling and modeling for Mandarin speech and its application to prosody generation for TTS
Chen-Yu Chiang, Sin-Horng Chen, Yih-Ru Wang
INTERSPEECH2
2009 A novel model-based pitch conversion method for Mandarin speech
Hsin-Te Hwang, Chen-Yu Chiang, Po-Yi Sung, Sin-Horng Chen
INTERSPEECH4
2008 Exploration of high-level prosodic patterns for continuous mandarin speech
abstract
In this paper, the high-level prosodic patterns of prosodic word (PW), prosodic phrase (PPh) and breath group/prosodic phrase group (BG/PG) for syllable pitch-level and duration are explored using an automatic joint prosody labeling and modeling method. Experimental results on a treebank speech corpus showed that the explored high-level prosodic patterns not only matched well with our a priori knowledge about Mandarin prosody, but also conformed well to other previous studies. They can therefore be integrated to form a meaningful Mandarin prosody hierarchy.
Chen-Yu Chiang, Hsiu-Min Yu, Yih-Ru Wang, Sin-Horng Chen
ICASSP4
2008 Robust estimation for sparse data
abstract
Robust parameters estimation of sparse data is generally applied to the test cases of time-consuming or high cost data collection. This study concerns with the problem in small sample size which is often encountered in the client data processing for speaker verification. We found that there always exists a coverage mismatch problem between the samples and its population in terms of probability density function (pdf) when the sample size is less than 20. We call this special problem the distribution mismatch (DM) problem. The paper proposes to solve the DM problem through addressing a new coverage-based estimator.
Wen-Hui Lo, Sin-Horng Chen
ICPR2
2007 Latent Prosody Model of Continuous Mandarin Speech
abstract
The major difficulty of prosody modeling and automatic tone recognition of continuous Mandarin speech is the complex interaction of tones and prosody/intonation on FO contours. In this study, we propose a latent prosody model (LPM) aiming to jointly model the affections of tone and prosody state on FO. The main purposes are twofold including (1) automatic prosody state labeling and (2) improving tone recognition accuracy. The basic idea is to introduce latent prosody state variables into an additive statistic model of FO which already considers the affecting factors of tone and speaker. Experiments on the Tree-Bank corpus showed that LPM not only gave meaningful prosody state labeling results but also improved the average tone recognition rate from 80.86% of a multi-layer perceptron (MLP) baseline to 82.55%.
Chen-Yu Chiang, Yuan-Fu Liao, Yih-Ru Wang, Sin-Horng Chen, Keikichi Hirose
ICASSP (4)5
2007 An automatic prosody labeling method for Mandarin speech
Chen-Yu Chiang, Hsiu-Min Yu, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH4
2005 On the inter-syllable coarticulation effect of pitch modeling for Mandarin speech
abstract
In this paper, a new statistics-based pitch model for Mandarin speech is proposed. The model considers three major affecting factors on the syllable pitch contour, including lexical tone, prosodic state and inter-syllable coarticulation effect. The study emphasizes on the modeling of inter-syllable coarticulation effect. Interactive affections of neighboring tones and different inter-syllable coarticulation states are considered. Experimental results show that the model performed well even for connected pitch contours of several syllables. Automatic labeling of linguistically meaningful coarticulation states is an extra benefit. So it is a promising F0 model.
Chen-Yu Chiang, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH3
2004 A model-based tone labeling method for Min-Nan/Taiwanese speech
abstract
A model-based tone labeling method for Min-Nan/Taiwanese speech is proposed. It takes the mean and shape of syllable pitch contours as two modeling units and considers some major affecting factors that control their variations. By using the EM algorithm to estimate all parameters of the pitch mean and shape models from a speech database, we can decide the best tone sequences pronounced in all utterances of the database. Experimental results show that it outperforms the VQ classification method which suffers from the interference resulting from neighboring syllables and from the global prosodic phrase patterns.
Wei-Chih Kuo, Yih-Ru Wang, Sin-Horng Chen
ICASSP (1)3
2003 A high-performance Min-Nan/Taiwanese TTS system
abstract
The implementation of a high-performance Min-Nan/Taiwanese TTS system is presented. The system can convert both Min-Nan/Taiwanese texts, represented in a hybrid Han-Lo written form, and Chinese texts into natural Taiwanese speech. It is an improved version of the system developed previously (Huang, J.Y., "Implementation of Tone Sandhi Rules and Tagger for Taiwanese TTS", Master Thesis, Commun. Eng. Dept., National Chiao Tung Univ., 2001). Improvements include: the addition of a "Chinese-to-Min-Nan/Taiwanese" lexicon to solve the OOV problem and to increase the ability of processing Chinese text; the use of explicit tone sandhi rules to ease the learning of prosody generation; a further processing of the training database to detect all breaks not associated with PMs; and the use of four RNNs (recurrent neural nets) to generate four types of prosodic parameters separately. The system is implemented by software and runs in real-time on a PC. An informal subjective listening test confirmed that the system performed well. All synthetic speech sounded natural for well-tokenized Min-Nan/Taiwanese texts and for automatically tokenized Chinese texts.
Wei-Chih Kuo, Xiang-Rui Zhong, Yih-Ru Wang, Sin-Horng Chen
ICASSP (1)4
2003 A mismatch-aware stochastic matching algorithm for robust speech recognition
abstract
We present a mismatch-aware stochastic matching (MASM) algorithm to alleviate the performance degradation under mismatched training and testing conditions. MASM first computes a reliability measure of applying a set of pre-trained speech models to a mismatch test utterance along the time axis or among different feature vector components. It then estimates and compensates the mismatch using the reliability measure to guide the speech segmentation. Experiments on a serious mismatched condition with training on PSTN-speech database and testing on mobile GSM-speech database showed that MASM outperformed the stochastic match (SM) method, especially, for short utterances.
Yuan-Fu Liao, Jeng-Shien Lin, Sin-Horng Chen
ICASSP (2)3
2003 An NN-based approach to prosodic information generation for synthesizing English words embedded in Chinese text
Wei-Chih Kuo, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH4
2003 A new pitch modeling approach for Mandarin speech
Wen-Hsing Lai, Yih-Ru Wang, Sin-Horng Chen
INTERSPEECH3
2003 A new duration modeling approach for Mandarin speech
abstract
A new duration modeling approach for Mandarin speech is proposed. It explicitly takes several major affecting factors, such as multiplicative companding factors (CFs), and estimates all model parameters by an EM algorithm. The three basic Tone 3 patterns (i.e., full tone, half tone and sandhi tone) are also properly considered using three different CFs to separate how they affect syllable duration. Experimental results show that the variance of the syllable duration is greatly reduced from 180.17 to 2.52 frame/sup 2/ (1 frame = 5 ms) by the syllable duration modeling to eliminate effects from those affecting factors. Moreover, the estimated CFs of those affecting factors agree well with our prior linguistic knowledge. Two extensions of the duration modeling method are also performed. One is the use of the same technique to model initial and final durations. The other is to replace the multiplicative model with an additive one. Lastly, a preliminary study of applying the proposed model to predict syllable duration for TTS (text-to-speech) is also performed. Experimental results show that it outperforms the conventional regressive prediction method.
Sin-Horng Chen, Wen-Hsing Lai, Yih-Ru Wang
IEEE Trans. Speech Audio Process.1
2002 RNN-based prosodic modeling for mandarin speech and its application to speech-to-text conversion
Wern-Jun Wang, Yuan-Fu Liao, Sin-Horng Chen
Speech Commun.3
2001 A novel syllable duration modeling approach for Mandarin speech
abstract
In this paper, a novel syllable duration modeling approach for Mandarin speech is proposed. It explicitly takes several main affecting factors as multiplicative companding parameters and estimates all model parameters by an EM algorithm. Experimental results show that the variance of the observed syllable duration is greatly reduced from 183.4 frame/sup 2/ (1 frame=5 ms) to 18.5 frame/sup 2/ by eliminating effects from these affecting factors. Besides, the estimated companding values of these affecting factors agree well with our prior linguistic knowledge. A preliminary study of applying the proposed model to predict syllable duration for TTS is also performed. Experimental results show that it outperforms the conventional regressive prediction method. Lastly, an extension of the approach to incorporate initial and final duration modeling is presented. This leads to a better understanding of the relation between the companding factors of initial and final duration models and those of syllable duration model.
Wen-Hsing Lai, Sin-Horng Chen
ICASSP2
2001 Multi-keyword spotting of telephone speech using orthogonal transform-based SBR and RNN prosodic model
abstract
In this paper, orthogonal transform-based signal bias removal (OTSBR) approach and RNN prosodic model are proposed for multi-keyword spotting of telephone speech. OTSBR is employed in the pre-processing stage of acoustic decoding and aimed at channel bias estimation to eliminate the acoustic mismatch between training and testing environments. The RNN prosodic model is adopted in the post-processing stage of the acoustic decoding to detect word boundaries for reordering the keyword candidates from the keyword spotter. Simulations on the real speech database collected from the Phone Directory Assistant Service developed in Chunghwa Telecommunication Laboratories (CTL-PDAS) were performed to evaluate the proposed methods. Experimental results showed that 71.0% of keyword detection rate and 81.8% of top 5 keywords inclusion rate can be attained by incorporating OTSBR and RNN prosodic model into the system.
Wern-Jun Wang, Chun-Jen Lee, Eng-Fong Huang, Sin-Horng Chen
INTERSPEECH4
2001 A modular RNN-based method for continuous Mandarin speech recognition
abstract
A new modular recurrent neural network (MRNN)-based method for continuous Mandarin speech recognition (CMSR) is proposed. The MRNN recognizer is composed of four main modules. The first is a sub-MRNN module whose function is to generate discriminant functions for all 412 base-syllables. It accomplishes the task by using four recurrent neural network (RNN) submodules. The second is an RNN module which is designed to detect syllable boundaries for providing timing cues in order to help solve the time-alignment problem. The third is also an RNN module whose function is to generate discriminant functions for 143 intersyllable diphone-like units to compensate the intersyllable coarticulation effect. The fourth is a dynamic programming (DP)-based recognition search module. Its function is to integrate the other three modules and solve the time-alignment problem for generating the recognized base-syllable sequence. A new multilevel pruning scheme designed to speed up the recognition process is also proposed. The whole MRNN can be trained by a sophisticated three-stage minimum classification error/generalized probabilistic descent (MCE/GPD) algorithm. Experimental results showed that the proposed method performed better than the maximum likelihood (ML)-trained hidden Markov model (HMM) method and is comparable to the MCE/GPD-trained HMM method. The multilevel pruning scheme was also found to be very efficient.
Yuan-Fu Liao, Sin-Horng Chen
IEEE Trans. Speech Audio Process.2
2000 A hybrid statistical/RNN approach to prosody synthesis for taiwanese TTS
Sin-Horng Chen, Chen-Chung Ho
INTERSPEECH1
2000 A robust training algorithm for adverse speech recognition
Wei-Tyng Hong, Sin-Horng Chen
Speech Commun.2
1999 A segment-based C0 adaptation scheme for PMC-based noisy Mandarin speech recognition
abstract
A segment-based C/sub 0/ (the zero-th order of cepstral coefficient) adaptation scheme for PMC-based Mandarin speech recognition is proposed in this paper. It incorporates a new C/sub 0/ model of speech signal into the PMC method to improve the gain matching between the clean-speech HMM models and the current noise model. The C/sub 0/ model is constructed in the training phase by jointly modeling the normalized C/sub 0/ with other MFCC recognition features to form C/sub 0/-normalized HMM models. In the testing phase, it pre-segments the input utterance into syllable-like segments, performs C/sub 0/-denormalization operations to expand the C/sub 0/-normalized HMM models, and uses them in the PMC method. Compared with the conventional PMC method, the proposed method can achieve a much better noise compensation effect due to the use of more precise gain matching in the PMC model combination. Experimental results showed that the base-syllable accuracy rate was significantly upgraded for continuous noisy Mandarin speech recognition.
Wei-Tyng Hong, Sin-Horng Chen
ICASSP2
1999 A robust environment-effects suppression training algorithm for adverse Mandarin speech recognition
abstract
. This paper addresses the problem of speech recognition in the presence of additive noise. To deal with this problem, it is possible to estimate the noise characteristics using methods which have previously been developed for speech enhancement techniques. Spectral subtraction can then be used to reduce the eect of additive noise on speech in the spectral domain. Some techniques have also recently been proposed for recognition with missing data. These approaches require an estimation of the local SNR to detect the speech spectral features which are relatively free from noise so as to perform recognition on these parts only. In this article, we compare these two dierent strategies, spectral subtraction and "missing data", on continuous speech additively disturbed with real noise. It is shown that missing data methods can improve recognition performance under certain noise conditions but still need to be improved in order to to reach the performance of the spectral subtracti...
Wei-Tyng Hong, Sin-Horng Chen
EUROSPEECH2
1999 A prototype of Mandarin speech telephone number inquiry system
Peng-Ren Lu, Wei-Tyng Hong, Sheng-Lun Chiang, Yih-Ru Wang, Sin-Horng Chen
EUROSPEECH5
1999 Prosodic modeling of Mandarin speech and its application to lexical decoding
Wern-Jun Wang, Yuan-Fu Liao, Sin-Horng Chen
EUROSPEECH3
1998 An MRNN-based method for continuous Mandarin speech recognition
abstract
A new modular recurrent neural network (MRNN)-based method for continuous Mandarin speech recognition is proposed. The system uses five RNNs to accomplish many subtasks separately and then combine them to integrally solve the problem. They include two RNNs for the discrimination of the two sub-syllable groups of 100 right-final-dependent (RFD) initials and 39 context independent (CI) finals, two RNNs for the generation of dynamic weighting functions for sub-syllable's integration, and one RNN for syllable boundary detection. All RNN modules are combined using a delay-decision Viterbi search. The method differs from the ANN/HMM hybrid approach of using ANNs to perform not only sub-syllables discrimination but also temporal structure modeling of the speech signal. The system is trained using a three-stage training method embedding with the MCE/GPD algorithms. Besides, a fast recognition method using multi-level pruning is also proposed. Experimental results showed that it outperforms the HMM method on both the recognition accuracy and the computational complexity.
Yuan-Fu Liao, Sin-Horng Chen
ICASSP2
1998 Mandarin telephone speech recognition for automatic telephone number directory service
abstract
This paper discusses an HMM-based Mandarin telephone speech recognition method for implementing a prototype system of automatic telephone number directory service. It adopted the GPD/MCE training algorithm to train the HMM models for 100 final-dependent syllable initials and 40 syllable finals. The SBR method was used to compensate the speaker and channel effects. Besides, a recurrent neural network (RNN) based pre-classification scheme was employed to speed up the recognition search. A syllable recognition rate of 53.7% was achieved. This method was then used to implement an isolated-word recognizer for the prototype system to discriminate 1922 names of bank and insurance companies. Word recognition rates of 94.8% for top-1 and 97.9% for top-3 were achieved.
Yih-Ru Wang, Sin-Horng Chen
ICASSP2
1998 An RNN-based prosodic information synthesizer for Mandarin text-to-speech
abstract
A new RNN-based prosodic information synthesizer for Mandarin Chinese text-to-speech (TTS) is proposed in this paper. Its four-layer recurrent neural network (RNN) generates prosodic information such as syllable pitch contours, syllable energy levels, syllable initial and final durations, as well as intersyllable pause durations. The input layer and first hidden layer operate with a word-synchronized clock to represent current-word phonologic states within the prosodic structure of text to be synthesized. The second hidden layer and output layer operate on a syllable-synchronized clock and use outputs from the preceding layers, along with additional syllable-level inputs fed directly to the second hidden layer, to generate desired prosodic parameters. The RNN was trained on a large set of actual utterances accompanied by associated texts, and can automatically learn many human-prosody phonologic rules, including the well-known Sandhi Tone 3 F0-change rule. Experimental results show that all synthesized prosodic parameter sequences matched quite well with their original counterparts, and a pitch-synchronous-overlap-add-based (PSOLA-based) Mandarin TTS system was also used for testing of our approach. While subjective tests are difficult to perform and remain to be done in the future, we have carried out informal listening tests by a significant number of native Chinese speakers and the results confirmed that all synthesized speech sounded quite natural.
Sin-Horng Chen, Shaw-Hwa Hwang, Yih-Ru Wang
IEEE Trans. Speech Audio Process.1
1998 An RNN-based preclassification method for fast continuous Mandarin speech recognition
abstract
A novel recurrent neural network-based (RNN-based) front-end preclassification scheme for fast continuous Mandarin speech recognition is proposed. First, an RNN is employed to discriminate each input frame for the three broad classes of initial, final, and silence. A finite state machine (FSM) is then used to classify the input frame into four states including three stable states of initial (I), final (F), and silence (S), and a transient (T) state. The decision is made based on examining whether the RNN discriminates well between classes. We then restrict the search space for the three stable states in the following DP search to speed up the recognition process. The efficiency of the proposed scheme was examined by simulations in which we incorporate it with a hidden Markov model-based (HMM-based) continuous 411 Mandarin based-syllables recognizer. The experimental results showed that it can be used in conjunction with the beam search to greatly reduce the computational complexity of the HMM recognizer while keeping the recognition rate almost undegraded.
Sin-Horng Chen, Yuan-Fu Liao, Song-Mao Chiang, Saga Chang
IEEE Trans. Speech Audio Process.1
1998 Modular recurrent neural networks for Mandarin syllable recognition
abstract
A new modular recurrent neural network (MRNN)- based speech-recognition method that can recognize the entire vocabulary of 1280 highly confusable Mandarin syllables is proposed in this paper. The basic idea is to first split the complicated task, in both feature and temporal domains, into several much simpler subtasks involving subsyllable and tone discrimination, and then to use two weighting RNN's to generate several dynamic weighting functions to integrate the subsolutions into a complete solution. The novelty of the proposed method lies mainly in the use of appropriate a priori linguistic knowledge of simple initial-final structures of Mandarin syllables in the architecture design of the MRNN. The resulting MRNN is therefore effective and efficient in discriminating among highly confusable Mandarin syllables. Thus both the time-alignment and scaling problems of the ANN-based approach for large-vocabulary speech-recognition can be addressed. Experimental results show that the proposed method and its extensions, the reverse-time MRNN (Rev-MRNN) and bidirection MRNN (Bi-MRNN), all outperform an advanced HMM method trained with the MCE/GPD algorithm in both recognition-rate and system complexity.
Sin-Horng Chen, Yuan-Fu Liao
IEEE Trans. Neural Networks1
1997 A robust RNN-based pre-classification for noisy Mandarin speech recognition
Wei-Tyng Hong, Sin-Horng Chen
EUROSPEECH2
1997 An RNN-based spectral information generation for Mandarin text-to-speech
Shaw-Hwa Hwang, Sin-Horng Chen, Saga Chang
EUROSPEECH2
1996 Continuous Mandarin speech recognition using hierarchical recurrent neural networks
abstract
An ANN-based continuous Mandarin base-syllable recognition system is proposed. It adopts a hybrid approach to combine an HRNN with a Viterbi search. The HRNN is taken at a front-end processor and responsible for calculating discrimination scores for all 411 base-syllables. The Viterbi search is then followed to find out the best base-syllable sequence with highest score as the recognized output. Experimental results showed that the proposed system outperforms the conventional HMM method on both the recognition accuracy and the computational complexity. The system can also be further modified to reduce the computational complexity while retaining the recognition accuracy almost be ungraded.
Yuan-Fu Liao, Wen-Yuan Chen, Sin-Horng Chen
ICASSP3
1996 A Mandarin text-to-speech system
Shaw-Hwa Hwang, Sin-Horng Chen, Yih-Ru Wang
ICSLP2
1996 The broad study of homograph disambiguity for Mandarin speech synthesis
Wern-Jun Wang, Shaw-Hwa Hwang, Sin-Horng Chen
ICSLP3
1996 A speech recognition method based on the sequential multi-layer perceptrons
Wen-Yuan Chen, Sin-Horng Chen, Cheng-Jung Lin
Neural Networks2
1995 A prosodic model of Mandarin speech and its application to pitch level generation for text-to-speech
abstract
A prosodic model of Mandarin speech is proposed to simulate human's pronunciation mechanism for exploring the hidden pronunciation states embedded in the input text. Parameters representing these pronunciation states are then used to assist prosody information generation. A multirate recurrent neural network (MRNN) is employed to realize the prosodic model. Two learning methods were proposed to train the MRNN. One is an indirect method which firstly uses an additional SRNN to track the dynamics of the prosody information of the utterance; and then takes the outputs of its hidden layer as desired targets to train the MRNN. The other is a direct training method which integrates the MRNN and the following MLP prosody synthesizers to directly learn the relation between the input linguistic features and the output prosody information. Simulation results confirmed the effectiveness of the approach. Most synthesized prosodic parameter sequences match quite well with their original counterparts.
Shaw-Hwa Hwang, Sin-Horng Chen
ICASSP2
1995 An improvement on syllable-based continuous Mandarin speech recognition via using inter-syllable boundary models
Saga Chang, Sin-Horng Chen
EUROSPEECH2
1995 Speech recognition with hierarchical recurrent neural networks
Wen-Yuan Chen, Yuan-Fu Liao, Sin-Horng Chen
Pattern Recognit.3
1995 Generalized minimal distortion segmentation for ANN-based speech recognition
abstract
A generalized minimal distortion segmentation algorithm is proposed to solve the time alignment problem for ANN-based speech recognition. By modeling dynamics of spectral information of an acoustic segment with smooth curves obtained by orthonormal polynomial expansion, a speech signal is optimally divided into segments and then recognized by an MLP recognizer. Experimental results showed that the proposed method outperforms the standard CDHMM method.>
Sin-Horng Chen, Wen-Yuan Chen
IEEE Trans. Speech Audio Process.1
1995 Tone recognition of continuous Mandarin speech based on neural networks
abstract
Several neural network-based tone recognition schemes for continuous Mandarin speech are discussed. A basic MLP tone recognizer using recognition features extracted from the processing syllable is first introduced. Then, some additional features extracted from neighboring syllables are added to compensate for the coarticulation effect. It is then further improved to compensate For the effect of sandhi rules of tone pronunciation by including tone information of neighboring syllables. The recognition criterion is now changed to find the best tone sequence that minimizes the total risk that simultaneously considers tone recognition of all syllables in the input utterance. Last, two approaches using HCNN and HSMLP, respectively, to model the intonation pattern as a hidden Markov chain for assisting tone recognition are proposed. The effectiveness of these schemes was confirmed by simulations on a speaker-independent tone recognition task. A recognition rate of 86.72% was achieved.>
Sin-Horng Chen, Yih-Ru Wang
IEEE Trans. Speech Audio Process.1
1994 Application of a generalized probabilistic descent method to recurrent neural network based speech recognition
abstract
A new method is proposed to train recurrent neural networks (RNNs) for speech recognition such that the difficulty of selecting appropriate target functions can be avoided. A novel architecture of the RNN-based speech recognition system is also introduced for solving the problem related to large vocabulary speech recognition. Additionally, the proposed RNN-based recognizer is found to have the advantages of being capable of absorbing the temporal variation of speech patterns as well as possessing effective discrimination capabilities. Performance of the proposed system was examined using two speech recognition tasks of recognizing 10 Mandarin digits and 54 confusable Mandarin syllables. Experimental results show that the proposed method outperforms both the continuous observation densities hidden Markov models method and a RNN recognizer using the extended back propagation training algorithm.>
Sin-Horng Chen, Yuan-Fu Liao, Wen-Yuan Chen
ICASSP (2)1
1994 Tone Recognition of Continuous Mandarin Speech Based on Hidden Markov Model
abstract
In this paper, several tone recognition schemes for continuous Mandarin speech are discussed. First, an SCHMM is used to model the acoustic features of a syllable for tone discrimination. Parameters extracted from the F0 and energy contours of the syllable by discrete Legendre orthonormal transform are used as the recognition features. Then, a scheme using two-layer network is proposed to cope with the difficulty resulting from the declination effect on the F0 contour of the declarative sentential utterance. The declination effect is modeled by a sentence-level HMM on the upper layer and the acoustic features of each tone are modeled by a state-dependent SCHMM on the lower layer. Lastly, the coarticulation effect coming from neighboring syllables is considered in the scheme using context-dependent model. Performance of these recognition schemes was examined by simulations. A recognition rate of 86.34% was achieved.
Yih-Ru Wang, Jyh-Ming Shieh, Sin-Horng Chen
Int. J. Pattern Recognit. Artif. Intell.3
1994 Variable frame rate speech coding using optimal interpolation
abstract
A VFR LPC vocoder using optimal interpolation is presented in this paper. In the encoder, some representative frames of an utterance are selected for transmission. In the decoder, LPC parameters of all untransmitted frames are restored by optimal interpolation. Simulation results show that this coding scheme outperforms the conventional VFR vocoder using linear interpolation. By incorporating contour quantization of gain and pitch information, a low variable-rate LPC vocoder is realized. An informal listening test shows that very high intelligible reconstructed speech was obtained at an average data rate of 300 bps.
Chii-Jen Chung, Sin-Horng Chen
IEEE Trans. Commun.2
1992 A first study on neural net based generation of prosodic and spectral information for Mandarin text-to-speech
abstract
A neural-network-based approach to generating prosodic and spectral information of syllables for Mandarin text-to-speech synthesis is studied. Some contextual features are first extracted from a given input text by text analysis and taken as input signals for synthesis. Then, six multilayer perceptrons are employed to generate pause duration, syllable duration, and pitch mean and shape of one- and two-syllable synthesis units, several reproduction templates of proper size are first generated for each synthesis unit of syllable approach. The objective is to generate spectral patterns of the syllable that can be directly concatenated to synthesize natural speech without further modification. The validity of this novel approach was examined by simulation using a database of sentential utterances recorded from TV news, reported by a single female announcer. Experimental results confirmed that this is a promising approach for Mandarin text-to-speech synthesis.>
Sin-Horng Chen, Shaw-Hwa Hwang, Chun-Yu Tsai
ICASSP1
1991 Discriminative analysis of distortion sequences in speech recognition
abstract
The authors suggest a linear discriminant function to complete the distance score instead of a conventional average distance. Several discriminative algorithms are proposed to learn the discriminant function. These include one heuristic method, two methods based on the error propagation algorithm, and one method based on the generalized probabilistic descent (GPD) algorithm. The authors study these methods in a speaker-independent speech recognition task involving utterances of the highly confusable English E-set. The results show that the best performance is obtained by using the GPD method, which achieved a 78.1% accuracy, compared to 67.6% with the traditional average method.>
Pao-Chung Chang, Sin-Horng Chen, Biing-Hwang Juang
ICASSP2
1990 Mandarin tone recognition by multi-layer perceptron
abstract
Tone recognition of isolated Mandarin monosyllables using the multilayer perceptron (MLP) model is reported. Ten features extracted from the fundamental frequency and energy contours of a monosyllable are used as the recognition features. The backpropagation algorithm is used to train the internal representation of the MLP. Several variations of the MLP with regard to the number of layers and the number of neurons in each layer are considered. In a speaker-untrained test, a recognition rate of 93.8% was achieved. It outperforms a Gaussian classifier which has a recognition rate of 90.6%.>
Pao-Chung Chang, San-Wei Sun, Sin-Horng Chen
ICASSP3
1990 A Chinese fundamental frequency synthesizer based on a statistical model
Sin-Horng Chen, Su-Min Lee, Saga Chang
ICSLP1
1990 Vector quantization of pitch information in Mandarin speech
abstract
A method of quantizing the shape of pitch contour segments of Mandarin speech by using orthogonal polynomial representation and vector quantization techniques is proposed. Only a very limited number of representative pitch contour patterns of words can be found in Mandarin conversation; therefore, pitch information can be represented by the shape and the length of the pitch contour segment word by word instead of frame by frame. An average bit rate of 0.78 b/frame (34.67 b/s) for voiced sounds was achieved. The method is a variable-rate coding scheme with an average delay of 317 ms.>
Sin-Horng Chen, Yih-Ru Wang
IEEE Trans. Commun.1
1989 Multi-H phase-coded modulations with asymmetric modulation indexes
abstract
Multi-H phase-coded modulation (MHPM) is a bandwidth-efficient modulation scheme which offers substantial coding gain over conventional digital modulation schemes. MHPM with asymmetric modulation indices corresponding to the bipolar data +1 and -1 is considered, and numerical results for the minimum Euclidean distances are provided. It is shown that performance improvements on the error probability over conventional MHPM are gained with essentially the same bandwidth and a very slight modification in implementation. The upper bounds on the error probabilities as functions of observation intervals and received E/sub b//N/sub 0/ are also investigated in detail. It is concluded that the concept of asymmetric modulation indices for MHPM is attractive for bandwidth and power-efficient modulation.>
Hong-Kuang Hwang, Lin-Shan Lee, Sin-Horng Chen
IEEE J. Sel. Areas Commun.3
1988 Efficient bit-level systolic array implementation of FIR and IIR digital filters
abstract
Bit-level systolic architectures based on an inner-product computation scheme for finite-impulse response (FIR) and infinite-impulse-response (IIR) digital are presented. The FIR filter structure is optimized in the sense that for a given clock rate, both the utilization efficiency and average throughput are maximized. The IIR filter structure has approximately the same utilization efficiency and throughput rate as previous related techniques for processing a single data stream (channel), but it allows two data streams to be processed concurrently to double the performance. This feature makes the IIR system attractive for use in applications where multiple filtering and particularly bandpass analysis are required.>
Chin-Liang Wang, Che-Ho Wei, Sin-Horng Chen
IEEE J. Sel. Areas Commun.3
1987 On the use of pitch contour of Mandarin speech in text-independent speaker identification
abstract
By taking advantage of the four-tone structure in the pitch contour of Mandarin speech, text-independent speaker identification of using orthogonal pitch parameter is described. Slopes, mean, and duration of the pitch contour of each word in an utterance are taken as recognition features. An 85% identification rate is achieved by using parameters of pitch contour only. When incorporating parameters of pitch contour with parameters of vocal tract, this system outperforms that of using parameters of pitch contour or vocal tract only. A recognition rate of 99.2% is reached in such a system.
Sin-Horng Chen, Min-Tau Lin
ICASSP1
1984 Robust image estimation in signal-dependent noise
abstract
Several estimators for the digital restoration of a noisy image are presented. These estimators differ from those previously found in the literature in that they are robust for deviations from the assumed signal or noise statistics, and the image is assumed to have been corrupted by signal-dependent or nonlinear noise, rather than simple additive noise. The assumption of signal-dependence complicates the problem considerably for non-robust estimators; for robust estimators, the problems are such that analytic solutions become impossible, and numerical methods must be used to derive the estimators. Both point-estimators and multi-parameter estimators are considered. In addition to the description of the various robust estimators, a comparison of their performance on real images corrupted by simulated signal-dependent noise with various (Gaussian and non-Gaussian) distributions is also presented.
Sin-Horng Chen, John F. Walkup
ICASSP1