EDBT 2026 Demo / reviewers in the wild / expert
Chung-Hsien Wu 0001
dblp:20/924
· DBLP profile ↗
178ranked-venue papers
49as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 109 · 33 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 87 · 21 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorSystems, architecture and hardware · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Applying Emotion-Cause Entailment for Help-Seeker Guidance in Emotional Support ConversationsabstractDeveloping a dialogue system for emotional support conversations (ESCs) is challenging, as it requires addressing both general dialogue aspects and dynamic emotional interactions with help-seekers. Existing systems often fail to adapt to help-seekers' emotional changes, which are crucial for effective support. A supporter must understand when to initiate guidance; otherwise, help-seekers may reject support due to negative emotions. This study introduces the dynamic plug-and-play language model (DPPLM), which integrates a transformer-based language model with two attribute models to generate strategic and empathetic responses. Unlike the original plug-and-play framework, DPPLM tracks help-seekers' emotional states using emotion-cause entailment, enabling better timing and balance between strategy and empathy. Evaluated on the emotional support conversation (ESConv) dataset, DPPLM consistently delivered stage-appropriate responses, providing greater comfort and support. It achieved a BERTScore of 0.8451, a ROUGE-L of 12.10, a Distinct-1 of 4.63, a Distinct-2 of 32.24, an emotion accuracy of 65.42%, and a strategy accuracy of 58.03%. In human evaluations, DPPLM outperformed baseline systems in overall performance, demonstrating its effectiveness in ESCs. Jeremy Chang, Chung-Hsien Wu 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Dysarthric Speech Recognition Using Curriculum Learning and Multi-stream Architecture
I-Ting Hsieh, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2024 | Applying Reinforcement Learning and Multi-Generators for Stage Transition in an Emotional Support Dialogue System
Jeremy Chang, Chung-Hsien Wu 0001 |
INTERSPEECH | 3 |
| 2024 | Dysarthric Speech Recognition Using Curriculum Learning and Articulatory Feature Embedding
I-Ting Hsieh, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2024 | USD-AC: Unsupervised Speech Disentanglement for Accent Conversion
Jen-Hung Huang, Wei-Tsung Lee, Chung-Hsien Wu 0001 |
INTERSPEECH | 3 |
| 2024 | Digital Phenotyping-Based Bipolar Disorder Assessment Using Multiple Correlation Data Imputation and Lasso-MLPabstractClinical rating scales can be used to assess the severity of bipolar disorder; however, their use involves clinician–patient interactions, which is labor-intensive. Therefore, this study proposes a digital-phenotyping-based system that provides clinical ratings of bipolar disorder severity using global positioning system, self-scale, daily mood, user emotion, sleep time, and multimedia data; these ratings are given on Hamilton Depression Rating Scale (HAM-D) and Young Mania Rating Scale (YMRS). A K-nearest-neighbor-based imputation method was used to handle missing data. In this method, missing data points are filled in with the multiple correlations between different features. Furthermore, the Least Absolute Shrinkage and Selection Operator (Lasso)-regression-based multilayer perceptron (Lasso-MLP) method was adopted to predict the total and factor scores on the HAM-D and YMRS. Five-fold cross-validation were used in evaluation experiments. When the designed data imputation method was used with Lasso-MLP, the mean square errors of the total score and average factor score on HAM-D (the YMRS) were 0.56 (0.38) and 1.88 (0.98), respectively, which were smaller than the corresponding values obtained through Lasso regression (by 0.12 and 0.05, respectively, for HAM-D and by 0.12 and 0.10, respectively, for the YMRS). The experimental results also indicated that the models trained with the imputed data outperformed those trained without imputed data. Thus, the developed approaches can eliminate the missing data problem and provide accurate clinical ratings. Jia-Hao Hsu, Chung-Hsien Wu 0001, Wei-Kai Wang, Hung-Yi Su, Esther Ching-Lan Lin, Po See Chen |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Speech Emotion Recognition using Decomposed Speech via Multi-task Learning
Jia-Hao Hsu, Chung-Hsien Wu 0001, Yu-Hung Wei |
INTERSPEECH | 2 |
| 2023 | Applying Segment-Level Attention on Bi-Modal Transformer Encoder for Audio-Visual Emotion RecognitionabstractEmotions can be expressed through multiple complementary modalities. This study selected speech and facial expressions as modalities by which to recognize emotions. Current audiovisual emotion recognition models perform supervised learning using signal-level inputs. Such models are presumed to characterize the temporal relationships in signals. In this study, supervised learning was performed on segment-level signals, which are more granular than signal-level signals, to precisely train an emotion recognition model. Effectively fusing multimodal signals is challenging. In this study, sequential segments of audiovisual signals were obtained, and features were extracted and applied to estimate segment-level attention weights according to the emotional consistency of the two modalities using a neural tensor network. A proposed bimodal Transformer Encoder was trained using signal-level and segment-level emotion labels in which temporal context was incorporated into the signals to improve upon existing emotion recognition models. In bimodal emotion recognition, the experimental results demonstrated that the proposed method achieved 74.31% accuracy (3.05% higher than the method of fusing correlation features) on the audio-visual emotion dataset BAUM-1, which is based on fivefold cross-validation, and 76.81% accuracy (2.57% higher than the Multimodal Transformer Encoder) on the multimodal emotion data set CMU-MOSEI, which is composed of training, validation, and testing sets. Jia-Hao Hsu, Chung-Hsien Wu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Generalization Ability Improvement of Speaker Representation and Anti-Interference for Speaker VerificationabstractThe ability to generalize to mismatches between training and testing conditions and resist interference from other speakers is crucial for the performance of speaker verification. In this paper, we propose two novel approaches to improve the generalization ability to deal with the mismatched recorded scenarios and languages in test conditions and to reduce the influence of interference from other speakers on the similarity measurement of two speaker embeddings. First, parent embedding learning (PEL) is used for model training, which exploits the generalization ability of the shared structure to improve the representation of speaker embeddings. Second, partial adaptive score normalization (PAS-Norm) is used to reduce the influence of interference from other speakers on embedding-based similarity measures. In the experiments, the speaker embedding models are trained using the VoxCeleb2 dataset, and the performance is evaluated on four other datasets under different conditions, including VoxCeleb1, Librispeech, SITW, and CN-Celeb datasets. In the experiments on VoxCeleb1, evaluation results considering a large number of verification speakers and identity restrictions show that the proposed PEL-based system reduces the EER by 6.0% and 4.9% in these two cases, respectively, compared to the state-of-the-art (SOTA) system. Furthermore, in the experiments evaluating speaker verification in mismatch conditions on SITW and CN-Celeb, the proposed PEL-based system also outperforms the SOTA system. In the language mismatched conditions, the EER is reduced by 8.3%. For the evaluation of the influence of interference from other speakers, the EER is significantly reduced by 24.4% when PAS-Norm is used instead of the baseline AS-Norm score normalization method. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Decomposition and Reorganization of Phonetic Information for Speaker Embedding LearningabstractSpeech content is closely related to the stability of speaker embeddings in speaker verification tasks. In this paper, we propose a novel architecture based on self-constraint learning (SCL) and reconstruction task (RT) to remove the influence of phonetic information on speaker embedding generation. First, SCL is used to reduce the divergence of frame-level features, which can avoid ambiguity between the resulting embeddings of the two utterances being compared. Second, RT is used to further remove phonetic information in frame-level layers, focusing on speaker-discriminative feature transformation. In our experiments, the speaker embedding models were trained on the VoxCeleb2 dataset and evaluated on the VoxCeleb1, Librispeech, SITW and VoxMovies datasets. Experimental results on VoxCeleb1 show that the proposed DROP-TDNN system reduced the EER by 7.5%, compared to the state-of-the-art ECAPA-TDNN system. Furthermore, the proposed DROP-TDNN system also outperformed the ECAPA-TDNN system in the experiments on SITW, Librispeech and VoxMovies under cross-dataset conditions. In the experiments on SITW, the proposed system reduced the EER by 3.4% compared to the ECAPA-TDNN system. In the experiments on Librispeech, the proposed system demonstrated the advantage of removing phonetic information under the clean speech condition, with a significant reduction of 25.5% in EER compared to the ECAPA-TDNN system. In the experiments on VoxMovies, the proposed system reduced the EER by up to 7.9% compared to the ECAPA-TDNN system under different pronunciation and background conditions. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Empathetic Response Generation Based on Plug-and-Play Mechanism With Empathy PerturbationabstractSpoken dialogue systems have rapidly developed but are often viewed as inhumane because they lack empathetic communication skills. In this study, a transformer-based language model (DialoGPT fine-tuned on the EmpatheticDialogues dataset) was combined with two proposed attribute models for affective and cognitive empathy to improve its performance. The affective empathy model ensures that the user sentence and system response have similar emotional valence, and the cognitive empathy model ensures that the system response is relevant to the user's input by using a DialoGPT-based reverse generation model to calculate the cross-entropy loss. A plug-and-play structure with these empathy attribute models was used to perturb the language generation model to increase response empathy without fine-tuning or retraining the generation model. Experiments indicated that the proposed model responses had substantially higher affective empathy, cognitive empathy, and BLEU scores than did the baseline model. Subjective evaluations also indicated that the responses of the proposed model had greater empathy, relevance, and fluency than did the baseline model. Moreover, the proposed model outperformed other similar models in A/B tests. Jia-Hao Hsu, Jeremy Chang, Min-Hsueh Kuo, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Memory-Efficient Multi-Step Speech Enhancement with Neural ODE
Jen-Hung Huang, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2022 | Linear-time Mixed-Cell-Height Legalization for Minimizing Maximum DisplacementabstractDue to the aggressive scaling of advanced technology nodes, multiple-row-height cells have become more and more common in VLSI design. Consequently, the placement of cells is no longer independent among different rows, which makes the traditional row-based legalization techniques obsolete. In this work, we present a highly efficient linear-time mixed-cell-height legalization approach that optimizes both the total cell displacement and the maximum cell displacement. First, a fast window-based cell insertion technique introduced in [4] is applied to obtain a feasible initial row assignment and cell ordering which is known to be good for total displacement consideration. In the second stage, we use an iterative cell swapping algorithm to change the row assignment and the cell order of the critical cells for maximum displacement reduction. Then we develop an optimal linear time DAG-based fixed row and fixed order legalization algorithm to minimize the maximum cell displacement. Finally, we propose a cell shifting heuristic to reduce the total cell displacement without increasing the maximum cell displacement. Using the proposed approach, the quality provided by the global placement can be preserved as much as possible. Compared with the state-of-the-art work [4], experimental results show that our proposed algorithm can reduce the maximum cell displacement by more than 11% on average with similar average cell displacement. Chung-Hsien Wu 0001, Wai-Kei Mak, Chris C. N. Chu |
ISPD | 1 |
| 2021 | Assessment of Bipolar Disorder Using Heterogeneous Data of Smartphone-Based Digital PhenotypingabstractIn mental health disorder, Bipolar Disorder (BD) is one of the most common mental illness. Using rating scales for assessment is one of the approaches for diagnosing and tracking BD patients. However, the requirement for manpower and time is heavy in the process of evaluation. In order to reduce the cost of social and medical resources, this study collects the user’s data by the App on smartphones, consisting of location data (GPS), self-report scales, daily mood, sleeping time and records of multi-media (text, speech, video) which are heterogeneous digital phenotyping data, to build a database. The features of each heterogeneous digital phenotyping data are extracted independently. Lasso Regression and ElasticNet Regression methods are employed to predict the score of Hamilton Depression Rating Scale (HAM-D) and Young Mania Rating Scale (YMRS), as a reference for the evaluation of BD. As incomplete and missing data are very common in medical research, the ensemble method is adopted to combine the results from different models trained with different combinations of missing data. The collected heterogeneous digital phenotyping data from 84 BD patients were used for training and evaluation of the proposed approach based on five-fold cross validation method. Experimental results show that the performance of the assessment system using the proposed method are encouraging. Hung-Yi Su, Chung-Hsien Wu 0001, Cheng-Ray Liou, Esther Ching-Lan Lin, Po See Chen |
ICASSP | 2 |
| 2021 | Exploring Macroscopic and Microscopic Fluctuations of Elicited Facial Expressions for Mood Disorder ClassificationabstractIn the clinical diagnosis of mood disorder, a large proportion of patients with bipolar disorder (BD) are misdiagnosed as having unipolar depression (UD). Generally, long-term tracking is required for patients with BD to conduct an appropriate diagnosis by using traditional diagnosis tools. A one-time diagnosis system for facilitating diagnosis procedures is thus highly desirable. Accordingly, in this study, the facial expressions of patients with BD, patients with UD, and healthy controls elicited by emotional video clips were used for conducting mood disorder classification; the classification was performed by exploring the temporal fluctuation characteristics among the three groups. First, macroscopic facial expressions characterized by action units (AUs) were applied for describing the temporal transformation of muscles. Modulation spectrum analysis was applied to extract short-term intensity variations in the AUs. An interval-based multilayer perceptron (MLP) neural network was then used to classify mood disorder on the basis of the detected AU intensities. Moreover, motion vectors (MVs) were employed to describe subtle changes in facial expressions in the microscopic view. Eight basic orientations of MV change were considered for representing microfluctuation. Wavelet decomposition was then applied to extract entropy and energy features in different frequency bands. A long short-term memory model was finally used to model long-term variations for conducting mood disorder classification. A decision-level fusion approach was conducted on the combined results of macroscopic and microscopic facial expressions. For evaluating the described methods, the facial expressions elicited from the 36 subjects (12 from each of the BD, UD, and control groups) were used in 12-fold cross-validation experiments. Approaches for macroscopic and microscopic expressions achieved classification accuracies of 63.9 and 66.7 percent, respectively, and the accuracy of the fusion approach reached 72.2 percent. The results indicate that macroscopic and microscopic view descriptors are complementary to each other and helpful for conducting mood disorder classification. Qian-Bei Hong, Chung-Hsien Wu 0001, Ming-Hsiang Su, Chia-Cheng Chang |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Speech Emotion Recognition Considering Nonverbal Vocalization in Affective ConversationsabstractIn real-life communication, nonverbal vocalization such as laughter, cries or other emotion interjections, within an utterance play an important role for emotion expression. In previous studies, only few emotion recognition systems consider nonverbal vocalization, which naturally exists in our daily conversation. In this work, both verbal and nonverbal sounds within an utterance are considered for emotion recognition of real-life affective conversations. Firstly, a support vector machine (SVM)-based verbal and nonverbal sound detector is developed. A prosodic phrase auto-tagger is further employed to extract the verbal/nonverbal sound segments. For each segment, the emotion and sound feature embeddings are respectively extracted using the deep residual networks (ResNets). Finally, a sequence of the extracted feature embeddings for the entire dialog turn are fed to an attentive long short-term memory (LSTM)-based sequence-to-sequence model to output an emotional sequence as recognition result. The NNIME corpus (The NTHU-NTUA Chinese interactive multimodal emotion corpus), which consists of verbal and nonverbal sounds, was adopted for system training and testing. 4766 single speaker dialogue turns in the audio data of the NNIME corpus were selected for evaluation. The experimental results showed that nonverbal vocalization was helpful for speech emotion recognition. For comparison, the proposed method based on decision-level fusion achieved an accuracy of 61.92% for speech emotion recognition outperforming the traditional methods as well as the feature-level and model-level fusion approaches. Jia-Hao Hsu, Ming-Hsiang Su, Chung-Hsien Wu 0001, Yi-Hsuan Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Statistics Pooling Time Delay Neural Network Based on X-Vector for Speaker VerificationabstractThis paper aims to improve speaker embedding representation based on x-vector for extracting more detailed information for speaker verification. We propose a statistics pooling time delay neural network (TDNN), in which the TDNN structure integrates statistics pooling for each layer, to consider the variation of temporal context in frame-level transformation. The proposed feature vector, named as statsvector, are compared with the baseline x-vector features on the VoxCeleb dataset and the Speakers in the Wild (SITW) dataset for speaker verification. The experimental results showed that the proposed stats-vector with score fusion achieved the best performance on VoxCeleb1 dataset. Furthermore, considering the interference from other speakers in the recordings, we found that the proposed statsvector efficiently reduced the interference and improved the speaker verification performance on the SITW dataset. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang, Chien-Lin Huang |
ICASSP | 2 |
| 2020 | Combining Deep Embeddings of Acoustic and Articulatory Features for Speaker IdentificationabstractIn this study, deep embedding of acoustic and articulatory features are combined for speaker identification. First, a convolutional neural network (CNN)-based universal background model (UBM) is constructed to generate acoustic feature (AC) embedding. In addition, as the articulatory features (AFs) represent some important phonological properties during speech production, a multilayer perceptron (MLP)-based AF embedding extraction model is also constructed for AF embedding extraction. The extracted AC and AF embeddings are concatenated as a combined feature vector for speaker identification using a fully-connected neural network. This proposed system was evaluated by three corpora consisting of King-ASR, LibriSpeech and SITW, and the experiments were conducted according to the properties of the datasets. We adopted all three corpora to evaluate the effect of AF embedding, and the results showed that combining AF embedding into the input feature vector improved the performance of speaker identification. The LibriSpeech corpus was used to evaluate the effect of the number of enrolled speakers. The proposed system achieved an EER of 7.80% outperforming the method based on x-vector with PLDA (8.25%). And we further evaluated the effect of signal mismatch using the SITW corpus. The proposed system achieved an EER of 25.19%, which outperformed the other baseline methods. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang, Chien-Lin Huang |
ICASSP | 2 |
| 2020 | Detecting Unipolar and Bipolar Depressive Disorders from Elicited Speech Responses Using Latent Affective Structure ModelabstractMood disorders, including unipolar depression (UD) and bipolar disorder (BD) [1] , are reported to be one of the most common mental illnesses in recent years. In diagnostic evaluation on the outpatients with mood disorder, a large portion of BD patients are initially misdiagnosed as having UD [2] . As most previous research focused on long-term monitoring of mood disorders, short-term detection which could be used in early detection and intervention is thus desirable. This work proposes an approach to short-term detection of mood disorder based on the patterns in emotion of elicited speech responses. To the best of our knowledge, there is no database for short-term detection on the discrimination between BD and UD currently. This work collected two databases containing an emotional database (MHMC-EM) collected by the Multimedia Human Machine Communication (MHMC) lab and a mood disorder database (CHI-MEI) collected by the CHI-MEI Medical Center, Taiwan. As the collected CHI-MEI mood disorder database is quite small and emotion annotation is difficult, the MHMC-EM emotional database is selected as a reference database for data adaptation. For the CHI-MEI mood disorder data collection, six eliciting emotional videos are selected and used to elicit the participants' emotions. After watching each of the six eliciting emotional video clips, the participants answer the questions raised by the clinician. The speech responses are then used to construct the CHI-MEI mood disorder database. Hierarchical spectral clustering is used to adapt the collected MHMC-EM emotional database to fit the CHI-MEI mood disorder database for dealing with the data bias problem. The adapted MHMC-EM emotional data are then fed to a denoising autoencoder for bottleneck feature extraction. The bottleneck features are used to construct a long short term memory (LSTM)-based emotion detector for generation of emotion profiles from each speech response. The emotion profiles are then clustered into emotion codewords using the K-means algorithm. Finally, a class-specific latent affective structure model (LASM) is proposed to model the structural relationships among the emotion codewords with respect to six emotional videos for mood disorder detection. Leave-one-group-out cross validation scheme was employed for the evaluation of the proposed class-specific LASM-based approaches. Experimental results show that the proposed class-specific LASM-based method achieved an accuracy of 73.33 percent for mood disorder detection, outperforming the classifiers based on SVM and LSTM. Kun-Yi Huang, Chung-Hsien Wu 0001, Ming-Hsiang Su, Yu-Ting Kuo |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Attention-Based Response Generation Using Parallel Double Q-Learning for Dialog Policy Decision in a Conversational SystemabstractThis article proposes an approach to response generation using a Parallel Double Q-learning algorithm for dialog policy decision in a conversational system. First, a new semantic representation of the user's input sentence is presented by using the CKIP parser to derive the semantic dependency sequence of the input sentence. Then, a Gated Recurrent Unit-based Autoencoder is used to obtain the user's turn representation as well as context representation. A Parallel Double Q-learning algorithm with a Deep Neural Network (PD-DQN), combining two Double DQNs in parallel for the contextual and semantic information in the user's message, respectively, are proposed to determine the dialog act. Finally, the user's input and the determined dialog act are fed to an attention-based Transformer model to generate the response template. With the generated response template, the semantic slots are filled with their corresponding values to obtain the final sentence response. This article collects a multi-turn conversation database consisting of 4186 turns in the travel domain and 447 chitchat question-answer pairs as the evaluation corpus. Five-fold cross validation is employed for performance evaluation. Experimental results show that the proposed approach based on semantic dependency for intent detection increases the accuracy by 4.3%. For dialog policy decision, the PD-DQN achieves 87.57% task success rate, which is 13.9% higher than the baseline Double DQN (73.67%). Finally, using the attention-based Transformer for response template generation obtains a Bleu score of 13.6, improved by 1.5 compared to the Sequence-to-Sequence model. In subjective evaluation, both the dialog policy and sentence generation model achieve a higher appropriateness and grammatical correctness scores than the baseline system. Ming-Hsiang Su, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | A Two-Stage Transformer-Based Approach for Variable-Length Abstractive SummarizationabstractThis study proposes a two-stage method for variable-length abstractive summarization. This is an improvement over previous models, in that the proposed approach can simultaneously achieve fluent and variable-length abstractive summarization. The proposed abstractive summarization model consists of a text segmentation module and a two-stage Transformer-based summarization module. First, the text segmentation module utilizes a pre-trained Bidirectional Encoder Representations from Transformers (BERT) and a bidirectional long short-term memory (LSTM) to divide the input text into segments. An extractive model based on the BERT-based summarization model (BERTSUM) is then constructed to extract the most important sentence from each segment. For training the two-stage summarization model, first, the extracted sentences are used to train the document summarization module in the second stage. Next, the segments are used to train the segment summarization module in the first stage by simultaneously considering the outputs of the segment summarization module and the pre-trained second-stage document summarization module. The parameters of the segment summarization module are updated by considering the loss scores of the document summarization module as well as the segment summarization module. Finally, collaborative training is applied to alternately train the segment summarization module and the document summarization module until convergence. For testing, the outputs of the segment summarization module are concatenated to provide the variable-length abstractive summarization result. For evaluation, the BERT-biLSTM-based text segmentation model is evaluated using ChWiki_181k database and obtains a good effect in capturing the relationship between sentences. Finally, the proposed variable-length abstractive summarization system achieved a maximum of 70.0% accuracy in human subjective evaluation on the LCSTS dataset. Ming-Hsiang Su, Chung-Hsien Wu 0001, Hao-Tse Cheng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Sound Events Recognition and Retrieval Using Multi-Convolutional-Channel Sparse Coding Convolutional Neural NetworksabstractThis article proposes two novel deep convolutional neural networks (CNN), which are called the sparse coding convolutional neural network (SC-CNN) and the multi-convolutional-channel SC-CNN (MSC-CNN), to address the sound event recognition and retrieval problem. Unlike the general framework of a CNN, in which the feature learning process is performed hierarchically, the proposed framework models the whole memorization process in the human brain, including encoding, storage, and recollection. In particular, the MSC-CNN is designed to recognize multiple sound events that occur simultaneously. The experimental results indicate that the proposed SC-CNN and MSC-CNN outperforms the state-of-the-art systems in sound event recognition and retrieval. Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang, Andri Santoso, Seksan Mathulaprangsan, Chin-Chin Chiang, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2020 | Cell-Coupled Long Short-Term Memory With L-Skip Fusion Mechanism for Mood Disorder Detection Through Elicited Audiovisual FeaturesabstractIn early stages, patients with bipolar disorder are often diagnosed as having unipolar depression in mood disorder diagnosis. Because the long-term monitoring is limited by the delayed detection of mood disorder, an accurate and one-time diagnosis is desirable to avoid delay in appropriate treatment due to misdiagnosis. In this paper, an elicitation-based approach is proposed for realizing a one-time diagnosis by using responses elicited from patients by having them watch six emotion-eliciting videos. After watching each video clip, the conversations, including patient facial expressions and speech responses, between the participant and the clinician conducting the interview were recorded. Next, the hierarchical spectral clustering algorithm was employed to adapt the facial expression and speech response features by using the extended Cohn-Kanade and eNTERFACE databases. A denoizing autoencoder was further applied to extract the bottleneck features of the adapted data. Then, the facial and speech bottleneck features were input into support vector machines to obtain speech emotion profiles (EPs) and the modulation spectrum (MS) of the facial action unit sequence for each elicited response. Finally, a cell-coupled long short-term memory (LSTM) network with an L -skip fusion mechanism was proposed to model the temporal information of all elicited responses and to loosely fuse the EPs and the MS for conducting mood disorder detection. The experimental results revealed that the cell-coupled LSTM with the L -skip fusion mechanism has promising advantages and efficacy for mood disorder detection. Ming-Hsiang Su, Chung-Hsien Wu 0001, Kun-Yi Huang, Tsung-Hsien Yang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Speech Emotion Recognition Using Deep Neural Network Considering Verbal and Nonverbal Speech SoundsabstractSpeech emotion recognition is becoming increasingly important for many applications. In real-life communication, non-verbal sounds within an utterance also play an important role for people to recognize emotion. In current studies, only few emotion recognition systems considered nonverbal sounds, such as laughter, cries or other emotion interjection, which naturally exists in our daily conversation. In this work, both verbal and nonverbal sounds within an utterance were thus considered for emotion recognition of real-life conversations. Firstly, an SVM-based verbal/nonverbal sound detector was developed. A Prosodic Phrase (PPh) auto-tagger was further employed to extract the verbal/nonverbal segments. For each segment, the emotion and sound features were respectively extracted based on convolutional neural networks (CNNs) and then concatenated to form a CNN-based generic feature vector. Finally, a sequence of CNN-based feature vectors for an entire dialog turn was fed to an attentive long short-term memory (LSTM)-based sequence-to-sequence model to output an emotional sequence as recognition result. Experimental results on the recognition of seven emotional states in the NNIME (The NTHU-NTUA Chinese interactive multimodal emotion corpus) showed that the proposed method achieved a detection accuracy of 52.00% outperforming the traditional methods. Kun-Yi Huang, Chung-Hsien Wu 0001, Qian-Bei Hong, Ming-Hsiang Su, Yi-Hsuan Chen |
ICASSP | 2 |
| 2019 | Follow-Up Question Generation Using Neural Tensor Network-Based Domain Ontology Population in an Interview Coaching System
Ming-Hsiang Su, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2019 | Attention-based convolutional neural network and long short-term memory for short-term detection of mood disorders based on elicited speech responses
Kun-Yi Huang, Chung-Hsien Wu 0001, Ming-Hsiang Su |
Pattern Recognit. | 2 |
| 2019 | Response Selection and Automatic Message-Response Expansion in Retrieval-Based QA Systems using Semantic Dependency Pair ModelabstractThis article presents an approach to response selection and message-response (MR) database expansion from the unstructured data on the psychological consultation websites for a retrieval-based question answering (QA) system in a constrained domain for emotional support and comforting. First, we manually construct an initial MR database based on the articles collected from the psychological consultation websites. The Chinese Knowledge and Information Processing probabilistic context-free grammar is adopted to obtain the semantic dependency graphs (SDGs) of all the messages and responses in the initial MR database. For each sentence in the MR database, all the semantic dependencies, each composed of two words and their semantic relation, are extracted from the SDG of the sentence to form a semantic dependency set. Finally, a matrix with the element representing the correlation between the semantic dependencies of the messages and their corresponding responses is constructed as a semantic dependency pair model (SDPM) for response selection. Moreover, as the number of MR pairs in the psychological consultation websites is increasing day by day, the MR database in the QA system should be expanded to meet the needs of the users. For MR database expansion, the unstructured data from the message board are automatically collected. For the collected data, the supervised latent Dirichlet allocation is adopted for event detection and then the event-based delta Bayesian Information Criterion is used for message and response article segmentation. Each extracted message segment is then fed to the constructed retrieval-based QA system to find the best matched response segment and the matching score is also estimated to verify if the new MR pair is suitable to be included in the expanded MR database. Fivefold cross validation was employed to evaluate the performance of the proposed retrieval-based QA system over the expanded MR database based on SDPM. Compared to the vector space model-based method, the Okapi BM25 model, and the deep learning-based sequence-to-sequence with attention model, the proposed approach achieved a more favorable performance according to a statistical significance test. The retrieval accuracy based on MR expansion was also evaluated and a satisfactory result was obtained confirming the effectiveness of the expanded MR database. In addition, the user's satisfaction score of the proposed system was evaluated using the Cronbach's alpha value and the satisfaction score of the proposed SDPM was higher than those of the methods for comparison. Ming-Hsiang Su, Chung-Hsien Wu 0001, Kun-Yi Huang, Wu-Hsuan Lin |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2018 | Locality-Preserving Complex-Valued Gaussian Process Latent Variable Model for Robust Face RecognitionabstractLearning a low-dimensional image representation yields effective and efficient face recognition. The use of such a representation helps to weaken the curse of dimensionality. However, the traditional facial representation method is not robust against partial occlusions or variations of expression. To solve this problem, this paper proposes a more reliable, complex-valued representation of facial image. The robust representation is based on the proposed locality-preserving complex-valued Gaussian process latent variable model (LP-CGPLVM). In the LP-CGPLVM, the Euler formula is utilized to transform original facial images into the complex domain. A proper complex GP is employed to model the mapping between the complex-valued high-dimensional data and the corresponding low-dimensional representation. Moreover, the locality-preserving constraint is taken into consideration to preserve the neighborhood data structure. The experimental results indicate that our proposed method is robust against partial occlusions and various facial expressions. Sih-Huei Chen, Yuan-Shan Lee, Yu-Sheng Hsu, Chung-Hsien Wu 0001, Jia-Ching Wang |
ICASSP | 4 |
| 2018 | Attention-Based Dialog State Tracking for Conversational Interview CoachingabstractThis study proposes an approach to dialog state tracking (DST) in a conversational interview coaching system. For the interview coaching task, the semantic slots, used mostly in traditional dialog systems, are difficult to define manually. This study adopts the topic profile of the response from the interviewee as the dialog state representation. In addition, as the response generally consists of several sentences, the summary vector obtained from a long short-term memory neural network (LSTM) is likely to contain noisy information from many irrelevant sentences. This study proposes a sentence attention mechanism combining the sentence attention weights from a convolutional neural tensor network (CNTN) and the topic profile by selectively focusing on significant sentences for attention-based dialog state tracking. This study collected 260 interview dialogs consisting of 3,016 dialog turns for performance evaluation. A five-fold cross validation scheme was employed and the results show that the proposed method outperformed the semantic slot-based baseline method. Ming-Hsiang Su, Chung-Hsien Wu 0001, Kun-Yi Huang, Chu-Kwang Chen |
ICASSP | 2 |
| 2018 | Follow-up Question Generation Using Pattern-based Seq2seq with a Small Corpus for Interview Coaching
Ming-Hsiang Su, Chung-Hsien Wu 0001, Kun-Yi Huang, Qian-Bei Hong, Huai-Hung Huang |
INTERSPEECH | 2 |
| 2018 | Sound Event Recognition Using Auditory-Receptive-Field Binary Pattern and Hierarchical-Diving Deep Belief NetworkabstractAutomatic sound event recognition (SER) has recently attracted renewed interest. Although practical SER system has many useful applications in everyday life, SER is challenging owing to the variations among sounds and noises in the real-world environment. This paper presents a novel feature extraction and classification method to solve the problem of SER. An audio-visual descriptor, called the auditory-receptive-field binary pattern, is designed based on the spectrogram image feature, the cepstral features, and the human auditory receptive field model. The extracted features are then fed into a classifier to perform event classification. The proposed classifier, called the hierarchical-diving deep belief network, is a deep neural network system that hierarchically learns the discriminative characteristics from physical feature representation to the abstract concept. The performance of our proposed system was verified using several experiments under various conditions. Using the RWCP dataset, the proposed system achieved a recognition rate of 99.27% for real-world sound data in 105 categories. Under noisy conditions, the developed system is very robust, with which it achieved 95.06% recognition rate with 0 dB signal-to-noise ratio. Using the TUT sound event dataset, the proposed system achieves error rates of 0.81 and 0.73 in sound event detection in home and residential area scenes. The experimental results reveal that the proposed system outperformed the other systems in this field. Chien-Yao Wang, Jia-Ching Wang, Andri Santoso, Chin-Chin Chiang, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Mood detection from daily conversational speech using denoising autoencoder and LSTMabstractIn current studies, an extended subjective self-report method is generally used for measuring emotions. Even though it is commonly accepted that speech emotion perceived by the listener is close to the intended emotion conveyed by the speaker, research has indicated that there still remains a mismatch between them. In addition, the individuals with different personalities generally have different emotion expressions. Based on the investigation, in this study, a support vector machine (SVM)-based emotion model is first developed to detect perceived emotion from daily conversational speech. Then, a denoising autoencoder (DAE) is used to construct an emotion conversion model to characterize the relationship between the perceived emotion and the expressed emotion of the subject for a specific personality. Finally, a long short-term memory (LSTM)-based mood model is constructed to model the temporal fluctuation of speech emotions for mood detection. Experimental results show that the proposed method achieved a detection accuracy of 64.5%, improving by 5.0% compared to the HMM-based method. Kun-Yi Huang, Chung-Hsien Wu 0001, Ming-Hsiang Su, Hsiang-Chi Fu |
ICASSP | 2 |
| 2017 | Fully complex deep neural network for phase-incorporating monaural source separationabstractDeep neural network (DNN) have become a popular means of separating a target source from a mixed signal. Most of DNN-based methods modify only the magnitude spectrum of the mixture. The phase spectrum is left unchanged, which is inherent in the short-time Fourier transform (STFT) coefficients of the input signal. However, recent studies have revealed that incorporating phase information can improve the quality of separated sources. To estimate simultaneously the magnitude and the phase of STFT coefficients, this work paper developed a fully complex-valued deep neural network (FCDNN) that learns the nonlinear mapping from complex-valued STFT coefficients of a mixture to sources. In addition, to reinforce the sparsity of the estimated spectra, a sparse penalty term is incorporated into the objective function of the FCDNN. Finally, the proposed method is applied to singing source separation. Experimental results indicate that the proposed method outperforms the state-of-the-art DNN-based methods. Yuan-Shan Lee, Chien-Yao Wang, Shu-Fan Wang, Jia-Ching Wang, Chung-Hsien Wu 0001 |
ICASSP | 5 |
| 2017 | Speech emotion recognition with ensemble learning methodsabstractIn this paper, we propose to apply ensemble learning methods on neural networks to improve the performance of speech emotion recognition tasks. The basic idea is to first divide unbalanced data set into balanced subsets and then combine the predictions of the models trained on these subsets. Several methods regarding the decomposition of data and the exploitation of model predictions are investigated in this study. On the public-domain FAU-Aibo database, which is used in Interspeech Emotion Challenge evaluation, the best performance we achieve is an unweighted average (UA) recall rate of 45.5% for the 5-class classification task. Furthermore, such performance is achieved with a feature space of 40-dimension. Compared to the baseline system with 384-dimension feature vector per example and an UA of 38.9%, such a performance is very impressive. Indeed, this is one of the best performances on FAU-Aibo within the static modeling framework. Po-Yuan Shih, Chia-Ping Chen, Chung-Hsien Wu 0001 |
ICASSP | 3 |
| 2017 | Recognition and retrieval of sound events using sparse coding convolutional neural networkabstractThis paper proposes a novel deep convolutional neural network (CNN), called sparse coding convolutional neural network (SC-CNN), to address the problem of sound event recognition and retrieval task. Unlike the general framework of a CNN, in which feature learning process is performed hierarchically, the proposed framework models the whole memorizing procedures in the human brain, including encoding, storage, and recollection. Sound data from the RWCP sound scene dataset with added noise from NOISEX-92 noise dataset are used to compare the performance of the proposed system with the state-of-the-art baselines. The experimental results indicated that the proposed SC-CNN outperformed the state-of-the-art systems in sound event recognition and retrieval. In the sound event recognition task, the proposed system achieved an accuracy of 94.6%, 100% and 100% under 0db, 10db and clean noise conditions, respectively. In the retrieval task, the proposed system improves the mAP rate of the general CNN by approximately 6%. Chien-Yao Wang, Andri Santoso, Seksan Mathulaprangsan, Chin-Chin Chiang, Chung-Hsien Wu 0001, Jia-Ching Wang |
ICME | 5 |
| 2017 | Interaction Style Recognition Based on Multi-Layer Multi-View Profile RepresentationabstractInteraction Style (IS) refers to patterns of interaction containing highly contextual and innate information. Awareness of our IS can help us discover interpersonal conflicts and guide us how to interact with others. Recently, automatic IS recognition is becoming increasingly important in the design of a dialogue system for harmonious interaction. With the goal to select appropriate responses, four IS types proposed by Berens are selected as the basis for our study. In this study, multiple views (multi-views) of the utterances during interaction, including emotions and dialogue topics, are recognized first. Inspired by the emotion profile theory, the IS profiles are then extracted using the multi-view features to better characterize the IS of the interactional utterances. Similar to the multilayer architectures in deep neural networks, a multi-layer multi-view IS profile representation method, structured layer by layer through embedding the multi-views, is proposed to better interpret intermediate representations in the feature space of the interactional utterances based on a probabilistic fusion model. The IS is finally recognized by using the Support Vector Machine (SVM) based on the obtained IS profiles. Experimental results demonstrate that the proposed method achieved an encouraging IS recognition accuracy and outperformed the previous method. Wen-Li Wei, Jen-Chun Lin, Chung-Hsien Wu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Personalized Spontaneous Speech Synthesis Using a Small-Sized Unsegmented Semispontaneous SpeechabstractA systematic approach is proposed to synthesizing personalized spontaneous speech using a small-sized unsegmented speech corpus of the target speaker. First, an automatic segmentation algorithm is employed to segment and label a collected semispontaneous speech corpus of the target speaker. Then, a pretrained average voice model is adapted to the voice model of the target speaker by using the segmented data. A postfilter based on modulation spectrum is adopted to further improve the speaker similarity of the synthesized speech as well as alleviate the over-smoothing problem of the synthesized speech. For generating spontaneous speech, a smoothing method applied at the prosodic word level is proposed to improve speech fluency. For objective evaluation on spontaneous speech segmentation, the segmentation accuracy of the proposed method is superior to that of Viterbi-based forced alignment. The results of subjective listening test also show that the proposed method can improve the spontaneity and speaker similarity of the synthesized speech compared to the maximum likelihood linear regression based speaker adaptation method. Yi-Chin Huang, Chung-Hsien Wu 0001, Yan-You Chen, Ming-Ge Shie, Jhing-Fa Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Speaker Identification Using Discriminative Features and Sparse RepresentationabstractSpeaker identification is an important topic with relevance to various disciplines. This paper proposes a novel speaker identification system, which consists of two major components-feature extraction and sparse representation classifier (SRC). Although SRC has been utilized for many classification purposes, few studies have provided insight into the link between the commonly used speaker identification feature, i-vector, and SRC. To combine i-vector and SRC sufficiently, we use probabilistic principal component analysis and Bartlett test to extract high-quality i-vector to construct a discriminative dictionary in SRC, supporting effective speaker identification. Besides improving dictionary from the i-vector aspect, we also utilize dictionary learning to further enhance the content of the dictionary. Two learning methods are proposed-robust principal component analysis dictionary and SVD-dictionary. Furthermore, we propose constructing a noise dictionary and combine it with the original dictionary to absorb and suppress noise when implementing the sparse coding. Various coding methods are utilized and analyzed. A comparison to the methods for speaker identification reveals that the proposed method outperforms the baselines and confirms its feasibility. Yu-Hao Chin, Jia-Ching Wang, Chien-Lin Huang, Kuang-Yao Wang, Chung-Hsien Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2016 | Generation of Emotion Control Vector Using MDS-Based Space Transformation for Expressive Speech Synthesis
Yan-You Chen, Chung-Hsien Wu 0001, Yu-Fong Huang |
INTERSPEECH | 2 |
| 2016 | Unipolar Depression vs. Bipolar Disorder: An Elicitation-Based Approach to Short-Term Detection of Mood Disorder
Kun-Yi Huang, Chung-Hsien Wu 0001, Yu-Ting Kuo, Fong-Lin Jang |
INTERSPEECH | 2 |
| 2016 | Candidate Expansion and Prosody Adjustment for Natural Speech Synthesis Using a Small CorpusabstractThis study proposes a hybrid approach to natural-sounding speech synthesis based on candidate expansion, unit selection, and prosody adjustment using a small corpus. The proposed method is more specific to tonal language, in particular Mandarin. In conventional speech synthesis studies, the quality of synthesized speech depends heavily on the size of the speech corpus. However, it is highly time-consuming and labor-intensive to prepare a large labeled corpus. In this work, candidate expansion is proposed to retrieve potential candidates that are unlikely to be retrieved using only linguistic features. The optimal unit sequence is then obtained from the expanded candidates by using the proposed unit selection mechanism at the phoneme and prosodic word levels. Finally, a prosodic word-level prosody adjustment is proposed to improve the continuity and smoothness of the prosody of the synthesized speech. To evaluate the proposed method, the Tsing-Hua corpus of speech synthesis was adopted. The results of an objective evaluation demonstrate the effectiveness of candidate expansion and the improvement of the continuity and smoothness of the prosody of the synthesized speech. The results of a subjective evaluation also show the proposed system could synthesize the speech with improved quality and naturalness, in particular for a small-sized or resource-limited corpus. Yan-You Chen, Chung-Hsien Wu 0001, Yi-Chin Huang, Shih-Lun Lin, Jhing-Fa Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Improving Mandarin Prosody Generation Using Alternative Smoothing TechniquesabstractProsody plays a vital role for conveying both communicative meanings and specific speaking styles in speech communication. In recent years, Hidden Markov Model (HMM)-based synthesis system (HTS) has been developed in triumph, which can synthesize stable and smooth speech. However, the prosody of the synthesized speech suffers from the over-smoothing problem. Thus, a better prosodic model is required to improve the natural variability of the synthesized speech. This study exploits a hybrid method to alleviate this problem by combining the statistical and the template-based unit selection methods. First, a two-level clustering approach is proposed to obtain representative prosodic patterns (denoted by codewords) of the hierarchical prosodic structure modeled by a modified Fujisaki model. The prosodic codewords are then used to represent the prosody of each sentence in the parallel corpus consisting of the real speech corpus and the synthesized counterpart obtained from the HTS. The synthesized speech utterance is then used as the query for retrieving the prosodic codewords of the utterances in the synthesized corpus. The retrieved synthesized prosodic codewords are mapped to the prosodic codewords of the real speech based on linear mapping rules obtained from the parallel corpus. The prosodic codeword language models for prosodic word and prosodic phrase are employed respectively to choose the optimal codeword sequence of the real speech. Finally, the most likely sequence of prosodic codewords can be obtained based on the NURBS-based continuity measure for synthesizing speech with natural prosody. The experimental results of subjective and objective tests demonstrate that the proposed prosodic model substantially improves naturalness of the intonation of the synthesized speech compared to that of the HMM-based method. Yi-Chin Huang, Chung-Hsien Wu 0001, Si-Ting Weng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Exploiting Turn-Taking Temporal Evolution for Personality Trait Perception in Dyadic ConversationsabstractIn dyadic conversations, turn-taking is a dynamically evolving behavior strongly linked to paralinguistic communication. Turn-taking temporal evolution in a dyadic conversation is inevitable and can be incorporated into a modeling framework for characterizing and recognizing the personality traits (PTs) of two speakers. This study presents an approach to automatically predicting PTs in a dyadic conversation. First, a recurrent neural network (RNN) was used to model the relationship between Big Five Inventory 10 (BFI-10) items and linguistic features of spoken text in each turn of a speaker (speaker turn) to output a BFI-10 profile. The RNN applies a recurrent property to characterize the short-term temporal evolution of a dialog. Second, the coupled hidden Markov model (C-HMM) was employed to model the long-term turn-taking temporal evolution and cross-speaker contextual information for detecting the PTs of two individuals for the entire dialog represented by the BFI-10 profile sequence. The Mandarin Conversational Dialogue Corpus was used for evaluation. The evaluation result shows that an average perception accuracy of 79.66% for the big five traits was achieved using five-fold cross validation. Compared with conventional HMM and support vector machine-based methods, the proposed approach achieved a more favorable performance according to a statistical significance test. The encouraging results confirm the usability of this system for future applications. Ming-Hsiang Su, Chung-Hsien Wu 0001, Yu-Ting Zheng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Compressive Sensing-Based Speech EnhancementabstractThis study proposes a speech enhancement method based on compressive sensing. The main procedures involved in the proposed method are performed in the frequency domain. First, an overcomplete dictionary is constructed from the trained speech frames. The atoms of this redundant dictionary are spectrum vectors that are trained by the K-SVD algorithm to ensure the sparsity of the dictionary. For a noisy speech spectrum, formant detection and a quasi-SNR criterion are first utilized to determine whether a frequency bin in the spectrogram is reliable, and a corresponding mask is designed. The mask-extracted reliable components in a speech spectrum are regarded as partial observations and a measurement matrix is constructed. The problem can therefore be treated as a compressive sensing problem. The K atoms of a K-sparsity speech spectrum are found using an orthogonal matching pursuit algorithm. Because the K atoms form the speech signal subspace, the removal of the noise projected onto these K atoms is achieved by multiplying the noisy spectrum with the optimized gain that corresponds to each selected atom. The proposed method is experimentally compared with the baseline methods and demonstrates its superiority. Jia-Ching Wang, Yuan-Shan Lee, Chang-Hong Lin, Shu-Fan Wang, Chih-Hao Shih, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Emotion recognition of affective speech based on multiple classifiers using acoustic-prosodic information and semantic labels (Extended abstract)abstractThis work presents an approach to emotion recognition of affective speech based on multiple classifiers using acoustic-prosodic information (AP) and semantic labels (SLs). For AP-based recognition, acoustic and prosodic features are extracted from the detected emotional salient segments of the input speech. Three types of models GMMs, SVMs, and MLPs are adopted as the base-level classifiers. A Meta Decision Tree (MDT) is then employed for classifier fusion to obtain the AP-based emotion recognition confidence. For SL-based recognition, semantic labels are used to automatically extract Emotion Association Rules (EARs) from the recognized word sequence of the affective speech. The maximum entropy model (MaxEnt) is thereafter utilized to characterize the relationship between emotional states and EARs for emotion recognition. Finally, a weighted product fusion method is used to integrate the AP-based and SL-based recognition results for final emotion decision. For evaluation, 2,033 utterances for four emotional states were collected. The experimental results reveal that the emotion recognition performance for AP-based recognition using MDT achieved 80.00%. On the other hand, an average recognition accuracy of 80.92% was obtained for SL-based recognition. Finally, combining AP information and SLs achieved 83.55% accuracy for emotion recognition. Chung-Hsien Wu 0001, Wei-Bin Liang |
ACII | 1 |
| 2015 | Hierarchical modeling of temporal course in emotional expression for speech emotion recognitionabstractThis paper presents an approach to hierarchical modeling of temporal course in emotional expression for speech emotion recognition. In the proposed approach, a segmentation algorithm is employed to hierarchically chunk an input utterance into three-level temporal units, including low-level descriptors (LLDs)-based sub-utterance level, emotion profile (EP)-based sub-utterance level and utterance level. An emotion-oriented hierarchical structure is constructed based on the three-level units to describe the temporal emotion expression in an utterance. A hierarchical correlation model is also proposed to fuse the three-level outputs from the corresponding emotion recognizers and further model the correlation among them to determine the emotional state of the utterance. The EMO-DB corpus was used to evaluate the performance on speech emotion recognition. Experimental results show that the proposed method considering the temporal course in emotional expression provides the potential to improve the speech emotion recognition performance. Chung-Hsien Wu 0001, Wei-Bin Liang, Kuan-Chun Cheng, Jen-Chun Lin |
ACII | 1 |
| 2015 | Affective structure modeling of speech using probabilistic context free grammar for emotion recognitionabstractA complete emotional expression typically contains a complex temporal course in a natural conversation. Related research on utterance-level and segment-level processing lacks understanding of the underlying structure of emotional speech. In this study, a hierarchical affective structure of an emotional utterance characterized by the probabilistic context free grammars (PCFGs) is proposed for emotion modeling. SVM-based emotion profiles are obtained and employed to segment the utterance into emotionally consistent segments. Vector quantization is applied to convert the emotion profile of each segment into codewords. A binary tree in which each node represents a codeword is constructed to characterize the affective structure of the utterance modeled by PCFG. Given an input utterance, the output emotion is determined according to the PCFG-based emotion model with the highest likelihood of the speech segments along with the score of the affective structure. For evaluation, the EMO-DB database and its expansion in utterance length were conducted. Experimental results show that the proposed method achieved emotion recognition accuracy of 87.22% for long utterances and outperformed the SVM-based method. Kun-Yi Huang, Jia-Kuan Lin, Yu-Hsien Chiu, Chung-Hsien Wu 0001 |
ICASSP | 4 |
| 2015 | Fluent personalized speech synthesis with prosodic word-level spontaneous speech generation
Yi-Chin Huang, Chung-Hsien Wu 0001, Ming-Ge Shie |
INTERSPEECH | 2 |
| 2015 | Sentence extraction with topic modeling for question-answer pair generation
Chung-Hsien Wu 0001, Chao-Hong Liu, Po-Hsun Su |
Soft Comput. | 1 |
| 2015 | Model Generation of Accented Speech using Model Transformation and Verification for Bilingual Speech RecognitionabstractNowadays, bilingual or multilingual speech recognition is confronted with the accent-related problem caused by non-native speech in a variety of real-world applications. Accent modeling of non-native speech is definitely challenging, because the acoustic properties in highly-accented speech pronounced by non-native speakers are quite divergent. The aim of this study is to generate highly Mandarin-accented English models for speakers whose mother tongue is Mandarin. First, a two-stage, state-based verification method is proposed to extract the state-level, highly-accented speech segments automatically. Acoustic features and articulatory features are successively used for robust verification of the extracted speech segments. Second, Gaussian components of the highly-accented speech models are generated from the corresponding Gaussian components of the native speech models using a linear transformation function. A decision tree is constructed to categorize the transformation functions and used for transformation function retrieval to deal with the data sparseness problem. Third, a discrimination function is further applied to verify the generated accented acoustic models. Finally, the successfully verified accented English models are integrated into the native bilingual phone model set for Mandarin-English bilingual speech recognition. Experimental results show that the proposed approach can effectively alleviate recognition performance degradation due to accents and can obtain absolute improvements of 4.1%, 1.8%, and 2.7% in word accuracy for bilingual speech recognition compared to that using traditional ASR approaches, MAP-adapted, and MLLR-adapted ASR methods, respectively. Han-Ping Shen, Chung-Hsien Wu 0001, Pei-Shan Tsai |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Speech Emotion Verification Using Emotion Variance Modeling and Discriminant Scale-Frequency MapsabstractThis paper develops an approach to speech-based emotion verification based on emotion variance modeling and discriminant scale-frequency maps. The proposed system consists of two parts-feature extraction and emotion verification. In the first part, for each sound frame, important atoms from the Gabor dictionary are selected by using the matching pursuit algorithm. The scale, frequency, and magnitude of the atoms are extracted to construct a nonuniform scale-frequency map, which supports auditory discriminability by the analysis of critical bands. Next, sparse representation is used to transform scale-frequency maps into sparse coefficients to enhance the robustness against emotion variance and achieve error-tolerance improvement. In the second part, emotion verification, two scores are calculated. A novel sparse representation verification approach based on Gaussian-modeled residual errors is proposed to generate the first score from the sparse coefficients. Such a classifier can minimize emotion variance and improve recognition accuracy. The second score is calculated by using the emotional agreement index (EAI) from the same coefficients. These two scores are combined to obtain the final detection result. Experiments on an emotional database of spoken speech were conducted and indicate that the proposed approach can achieve an average equal error rate (EER) of as low as 6.61%. A comparison among different approaches reveals that the proposed method is superior to the others and confirms its feasibility. Jia-Ching Wang, Yu-Hao Chin, Bo-Wei Chen, Chang-Hong Lin, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2015 | Code-Switching Event Detection by Using a Latent Language Space Model and the Delta-Bayesian Information CriterionabstractThis paper proposes a new paradigm for code-switching event detection based on latent language space models (LLSMs) and the delta-Bayesian information criterion (ΔBIC). A phone-based Mandarin-English speech recognizer was first employed for obtaining the senone sequence of a speech utterance. For each senone, acoustic features and the posterior probability of the articulatory features (AFs) were extracted and applied to an eigenspace transformation, based on principal component analysis (PCA). Latent semantic analysis (LSA) was then adopted for constructing a matrix to model the importance of each principal component in the eigenspace for the senones and AFs in each language. The spatial relationships among the senones (or AFs) represented by the PCA-transformed eigenvalues in the LSA-based matrix were employed to construct an LLSM for characterizing a language. In code-switching event detection, the language likelihood between the input speech LLSM and each of the language-dependent LLSMs was estimated. The Euclidian-distance-based similarities and cosine-angle-distance-based similarities were adopted for estimating the language likelihood for senones and AFs. The ΔBIC was then used for estimating the language transition score for each hypothesized code-switching event. Finally, the dynamic programming algorithm was employed for obtaining the most likely code-switching language sequence. The proposed approach was evaluated using a Mandarin-English code-switching speech database and outperformed other conventional methods. A duration accuracy of 72.45% can be obtained from the proposed system with optimized parameters. Chung-Hsien Wu 0001, Han-Ping Shen, Chun-Shan Hsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Natural speech synthesis based on hybrid approach with candidate expansion and verificationabstractA hybrid Mandarin speech synthesis system combining concatenation-based and model-based methodology is investigated in this research. To effectively exploit a small-size corpus, the candidate sets for unit selection are expanded via clusters based on articulatory features (AF), which are estimated as the outputs of an artificial neural network. This is followed by a filtering operation incorporating residual compensation, to remove unsuitable units. Given an input text, an optimal unit sequence is decided by the minimization of a total cost, which depends on the spectral features, contextual articulatory features, formants, and pitch values. Furthermore, prosodic word verification is integrated to check the smoothness of the output speech. The units failing to pass the prosodic word verification are replaced by model-based synthesized units for better speech quality. Objective and subjective evaluations have been conducted. Comparisons among the proposed method, the HMM-based method, and the conventional hybrid method clearly show that candidate set expansion based on articulatory features lead to more units suitable for selection, and the verification process is effective in improving the naturalness of the output speech. Chung-Hsien Wu 0001, Yi-Chin Huang, Shih-Lun Lin, Chia-Ping Chen |
ICASSP | 1 |
| 2014 | Polyglot Speech Synthesis Based on Cross-Lingual Frame Selection Using Auditory and Articulatory FeaturesabstractIn this paper, an approach for polyglot speech synthesis based on cross-lingual frame selection is proposed. This method requires only mono-lingual speech data of different speakers in different languages for building a polyglot synthesis system, thus reducing the burden of data collection. Essentially, a set of artificial utterances in the second language for a target speaker is constructed based on the proposed cross-lingual frame-selection process, and this data set is used to adapt a synthesis model in the second language to the speaker. In the cross-lingual frame-selection process, we propose to use auditory and articulatory features to improve the quality of the synthesized polyglot speech. For evaluation, a Mandarin-English polyglot system is implemented where the target speaker only speaks Mandarin. The results show that decent performance regarding voice identity and speech quality can be achieved with the proposed method. Chia-Ping Chen, Yi-Chin Huang, Chung-Hsien Wu 0001, Kuan-De Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Exploiting Psychological Factors for Interaction Style Recognition in Spoken ConversationabstractDetermining how a speaker is engaged in a conversation is crucial for achieving harmonious interaction between computers and humans. In this study, a fusion approach was developed based on psychological factors to recognize Interaction Style ($IS$) in spoken conversation, which plays a key role in creating natural dialogue agents. The proposed Fused Cross-Correlation Model (FCCM) provides a unified probabilistic framework to model the relationships among the psychological factors of emotion, personality trait ($PT$), transient$IS$, and$IS$history, for recognizing$IS$. An emotional arousal-dependent speech recognizer was used to obtain the recognized spoken text for extracting linguistic features to estimate transient$IS$likelihood and recognize$PT$. A temporal course modeling approach and an emotional sub-state language model, based on the temporal phases of an emotional expression, were employed to obtain a better emotion recognition result. The experimental results indicate that the proposed FCCM yields satisfactory results in$IS$recognition and also demonstrate that combining psychological factors effectively improves$IS$recognition accuracy. Wen-Li Wei, Chung-Hsien Wu 0001, Jen-Chun Lin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Synthesis of Spontaneous Speech With Syllable Contraction Using State-Based Context-Dependent Voice TransformationabstractPronunciation normally varies in spontaneous speech, and is an integral aspect of spontaneous expression. This study describes a voice transformation-based approach to generating spontaneous speech with syllable contractions for Hidden Markov Model (HMM)-based speech synthesis. A multi-dimensional linear regression model is adopted as the context-dependent, state-based transformation function to convert the feature sequence of read speech to that of spontaneous speech with syllable contraction. With insufficient number of training data, the obtained transformation functions are categorized using a decision tree based on linguistic and articulatory features for better and efficient selection of suitable transformation functions. Furthermore, to cope with the problem of small parallel corpus, cross-validation of trained transformation function is performed to ensure correct transformation functions are obtained and prevent over-fitting. Consequently, pronunciation variations of syllable contraction for the trained and the unseen syllable-contracted words are generated from the transformation function retrieved from the decision tree using linguistic and articulatory features. Objective and subjective tests were used to evaluate the performance of the proposed approach. Evaluation results demonstrate that the proposed transformation function substantially improves apparent spontaneity of the synthesized speech compared to the conventional methods. Chung-Hsien Wu 0001, Yi-Chin Huang, Chung-Han Lee, Jun-Cheng Guo |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Chinese-English Phone Set Construction for Code-Switching ASR Using Acoustic and DNN-Extracted Articulatory FeaturesabstractThis study proposes a data-driven approach to phone set construction for code-switching automatic speech recognition (ASR). Acoustic and context-dependent cross-lingual articulatory features (AFs) are incorporated into the estimation of the distance between triphone units for constructing a Chinese-English phone set. The acoustic features of each triphone in the training corpus are extracted for constructing an acoustic triphone HMM. Furthermore, the articulatory features of the “last/first” state of the corresponding preceding/succeeding triphone in the training corpus are used to construct an AF-based GMM. The AFs, extracted using a deep neural network (DNN), are used for code-switching articulation modeling to alleviate the data sparseness problem due to the diverse context-dependent phone combinations in intra-sentential code-switching. The triphones are then clustered to obtain a Chinese-English phone set based on the acoustic HMMs and the AF-based GMMs using a hierarchical triphone clustering algorithm. Experimental results on code-switching ASR show that the proposed method for phone set construction outperformed other traditional methods. Chung-Hsien Wu 0001, Han-Ping Shen, Yan-Ting Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Automatic pronunciation clustering using a World English archive and pronunciation structure analysisabstractEnglish is the only language available for global communication. Due to the influence of speakers' mother tongue, however, those from different regions inevitably have different accents in their pronunciation of English. The ultimate goal of our project is creating a global pronunciation map of World Englishes on an individual basis, for speakers to use to locate similar English pronunciations. If the speaker is a learner, he can also know how his pronunciation compares to other varieties. Creating the map mathematically requires a matrix of pronunciation distances among all the speakers considered. This paper investigates invariant pronunciation structure analysis and Support Vector Regression (SVR) to predict the inter-speaker pronunciation distances. In experiments, the Speech Accent Archive (SAA), which contains speech data of worldwide accented English, is used as training and testing samples. IPA narrow transcriptions in the archive are used to prepare reference pronunciation distances, which are then predicted based on structural analysis and SVR, not with IPA transcriptions. Correlation between the reference distances and the predicted distances is calculated. Experimental results show very promising results and our proposed method outperforms by far a baseline system developed using an HMM-based phoneme recognizer. Han-Ping Shen, Nobuaki Minematsu, Takehiko Makino, Steven H. Weinberger, Teeraphon Pongkittiphan, Chung-Hsien Wu 0001 |
ASRU | 6 |
| 2013 | Personalized natural speech synthesis based on retrieval of pitch patterns using hierarchical Fujisaki modelabstractIn recent years, speech synthesis based on Hidden Markov Model (HMM) has been developed, which can synthesize stable and intelligible speech with flexibility and small footprint. However, synthesized prosodic features are still incapable to convey personalization and natural property. Previous prosody models, mainly constructed from the clustered prosodic features, are unable to characterize personalized prosodic information as the linguistic cues of the input sentence are indistinguishable for all speakers. An approach to retrieval of personalized pitch patterns from the real speech corpus of the target speaker, is proposed, incorporating with the HMM-based speech synthesizer, to generate a personalized natural pitch contour. The modified Fujisaki model is adopted to depict the hierarchical pitch patterns, aiming to model local pitch contour variation and global intonation of utterances in the corpus. The codeword sequences of utterances in the training and the synthesized corpora are constructed and used to obtain the relationship of pitch patterns between the real and synthesized speech. Finally, a language model of pitch pattern is constructed to obtain an optimal pitch pattern sequence of the input sentence. The experimental results using subjective and objective evaluations demonstrated the proposed approach can substantially outperform the conventional statistical synthesis methods, in terms of naturalness and speaker similarity. Yi-Chin Huang, Chung-Hsien Wu 0001, Shih-Lun Lin |
ICASSP | 2 |
| 2013 | Facial action unit prediction under partial occlusion based on Error Weighted Cross-Correlation ModelabstractOcclusive effect is a crucial issue that may dramatically degrade performance on facial expression recognition. As emotion recognition from facial expression is based on the entire facial feature, occlusive effect remains a challenging problem to be solved. To manage this problem, an Error Weighted Cross-Correlation Model (EWCCM) is proposed to effectively predict the facial Action Unit (AU) under partial facial occlusion from non-occluded facial regions for providing the correct AU information for emotion recognition. The Gaussian Mixture Model (GMM)-based Cross-Correlation Model (CCM) in EWCCM is first proposed not only modeling the extracted facial features but also constructing the statistical dependency among features from paired facial regions for AU prediction. The Bayesian classifier weighting scheme is then adopted to explore the contributions of the GMM-based CCMs to enhance the prediction accuracy. Experiments show that a promising result of the proposed approach can be obtained. Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei |
ICASSP | 2 |
| 2013 | Interaction style detection based on Fused Cross-Correlation Model in spoken conversationabstractIn recent years, much attention has been given to dialogue strategy design to achieve intelligent speech-based human-computer interaction. Since speakers generally express their intents in different Interaction Styles (ISs), the responses of a spoken dialogue system should be versatile instead of invariable and planned. This paper presents an approach to automatic detection of a user's IS using a Fused Cross-Correlation Model (FCCM). As IS generally involves high level psychological meaning, cross-correlation among various psychological factors including emotion, personality trait, and IS is thus considered for IS detection modeling. The Bayes' theorem is then used to integrate the cross-correlation into the IS detector for enhancing the IS detection accuracy. Experiments show a promising result of the proposed approach. Wen-Li Wei, Chung-Hsien Wu 0001, Jen-Chun Lin |
ICASSP | 2 |
| 2013 | Code-Switching event detection based on delta-BIC using phonetic eigenvoice models
Wei-Bin Liang, Chung-Hsien Wu 0001, Chun-Shan Hsu |
INTERSPEECH | 2 |
| 2013 | Emotion recognition of conversational affective speech using temporal course modeling
Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei |
INTERSPEECH | 2 |
| 2013 | Multiple visual concept discovery using concept-based visual word clustering
Jun-Bin Yeh, Chung-Hsien Wu 0001, Shi-Xin Mai |
Multim. Syst. | 2 |
| 2013 | Personalized Spectral and Prosody Conversion Using Frame-Based Codeword Distribution and Adaptive CRFabstractThis study proposes a voice conversion-based approach to personalized text-to-speech (TTS) synthesis. The conversion functions, trained using a small parallel corpus with source and target speech data, can impose the voice characteristics of a target speaker on an existing synthesizer. Frame alignment between a pair of sentences in the parallel corpus is generally used for training voice conversion functions. However, with incorrect alignment, the resultant conversion functions may generate unacceptable conversion results. Traditional frame alignment using minimal spectral distance between the frame-based feature vectors of the source and the target phone sequences can be imprecise because the voice properties of the source and target phones inherently differ. In the proposed method, feature vectors of the parallel corpus are transformed into codewords in an eigenspace. A more precise frame alignment can be obtained by integrating the codeword occurrence distributions into distance estimation. In addition to the spectral property, a prosodic word/phrase boundary prediction model was constructed using an adaptive conditional random field (CRF) to generate personalized prosodic information. Objective and subjective tests were conducted to evaluate the performance of the proposed approach. The experimental results showed that the proposed voice conversion method, based on distribution-based alignment and prosodic word boundary detection, can improve the speech quality and speaker similarity of the converted speech. Compared to other methods, the evaluation results verified the improved performance of the proposed method. Yi-Chin Huang, Chung-Hsien Wu 0001, Yu-Ting Chao |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Two-Level Hierarchical Alignment for Semi-Coupled HMM-Based Audiovisual Emotion Recognition With Temporal CourseabstractA complete emotional expression typically contains a complex temporal course in face-to-face natural conversation. To address this problem, a bimodal hidden Markov model (HMM)-based emotion recognition scheme, constructed in terms of sub-emotional states, which are defined to represent temporal phases of onset, apex, and offset, is adopted to model the temporal course of an emotional expression for audio and visual signal streams. A two-level hierarchical alignment mechanism is proposed to align the relationship within and between the temporal phases in the audio and visual HMM sequences at the model and state levels in a proposed semi-coupled hidden Markov model (SC-HMM). Furthermore, by integrating a sub-emotion language model, which considers the temporal transition between sub-emotional states, the proposed two-level hierarchical alignment-based SC-HMM (2H-SC-HMM) can provide a constraint on allowable temporal structures to determine an optimal emotional state. Experimental results show that the proposed approach can yield satisfactory results in both the posed MHMC and the naturalistic SEMAINE databases, and shows that modeling the complex temporal structure is useful to improve the emotion recognition performance, especially for the naturalistic database (i.e., natural conversation). The experimental results also confirm that the proposed 2H-SC-HMM can achieve an acceptable performance for the systems with sparse training data or noisy conditions. Chung-Hsien Wu 0001, Jen-Chun Lin, Wen-Li Wei |
IEEE Trans. Multim. | 1 |
| 2013 | Speaking Effect Removal on Emotion Recognition From Facial Expressions Based on Eigenface ConversionabstractSpeaking effect is a crucial issue that may dramatically degrade performance in emotion recognition from facial expressions. To manage this problem, an eigenface conversion-based approach is proposed to remove speaking effect on facial expressions for improving accuracy of emotion recognition. In the proposed approach, a context-dependent linear conversion function modeled by a statistical Gaussian Mixture Model (GMM) is constructed with parallel data from speaking and non-speaking facial expressions with emotions. To model the speaking effect in more detail, the conversion functions are categorized using a decision tree considering the visual temporal context of the Articulatory Attribute (AA) classes of the corresponding input speech segments. For verification of the identified quadrant of emotional expression on the Arousal-Valence (A-V) emotion plane, which is commonly used to dimensionally define the emotion classes, from the reconstructed facial feature points, an expression template is constructed to represent the feature points of the non-speaking facial expressions for each quadrant. With the verified quadrant, a regression scheme is further employed to estimate the A-V values of the facial expression as a precise point in the A-V emotion plane. Experimental results show that the proposed method outperforms current approaches and demonstrates that removing the speaking effect on facial expression is useful for improving the performance of emotion recognition. Chung-Hsien Wu 0001, Wen-Li Wei, Jen-Chun Lin, Wei-Yu Lee |
IEEE Trans. Multim. | 1 |
| 2012 | Cross-lingual frame selection method for polyglot speech synthesisabstractA novel approach is proposed to creating a polyglot speech synthesis system without the need of collecting speech data from a bilingual (or multilingual) speaker, which is often expensive or even infeasible. Given a target speaker with data in the first language (Mandarin in this study), the basic idea is to construct artificial utterances in the second language (English) via selection of speech sample frames of the given speaker in the first language. As the speaker needs not be polyglot, this method is generally applicable to any speaker and any languages. In the search for optimal frame sequence selection, the candidate set is constrained by a decision tree for phone segments in the speech data of both languages, and the cost function depends on the context-dependent articulatory and auditory features. Evaluation results show that good performance regarding similarity (speaker identity) and naturalness (speech quality) can be achieved with the proposed method. Chia-Ping Chen, Yi-Chin Huang, Chung-Hsien Wu 0001, Kuan-De Lee |
ICASSP | 3 |
| 2012 | Phone set construction based on context-sensitive articulatory attributes for code-switching speech recognitionabstractBilingual speakers are known for their ability to code-switch or mix their languages during communication. This phenomenon occurs when bilinguals substitute a word or phrase from one language with a phrase or word from another language. For code-switching speech recognition, it is essential to collect a large-scale code-switching speech database for model training. In order to ease the negative effect caused by the data sparseness problem in training code-switching speech recognizers, this study proposes a data-driven approach to phone set construction by integrating acoustic features and cross-lingual context-sensitive articulatory features into distance measure between phone units. KL-divergence and a hierarchical phone unit clustering algorithm are used in this study to cluster similar phone units to reduce the need of the training data for model construction. The experimental results show that the proposed method outperforms other traditional phone set construction methods. Chung-Hsien Wu 0001, Han-Ping Shen, Yan-Ting Yang |
ICASSP | 1 |
| 2012 | Error Diagnosis of Chinese Sentences Using Inductive Learning Algorithm and Decomposition-Based Testing MechanismabstractThis study presents a novel approach to error diagnosis of Chinese sentences for Chinese as second language (CSL) learners. A penalized probabilistic First-Order Inductive Learning (pFOIL) algorithm is presented for error diagnosis of Chinese sentences. The pFOIL algorithm integrates inductive logic programming (ILP), First-Order Inductive Learning (FOIL), and a penalized log-likelihood function for error diagnosis. This algorithm considers the uncertain, imperfect, and conflicting characteristics of Chinese sentences to infer error types and produce human-interpretable rules for further error correction. In a pFOIL algorithm, relation pattern background knowledge and quantized t -score background knowledge are proposed to characterize a sentence and then used for likelihood estimation. The relation pattern background knowledge captures the morphological, syntactic and semantic relations among the words in a sentence. One or two kinds of the extracted relations are then integrated into a pattern to characterize a sentence. The quantized t -score values are used to characterize various relations of a sentence for quantized t -score background knowledge representation. Afterwards, a decomposition-based testing mechanism which decomposes a sentence into background knowledge set needed for each error type is proposed to infer all potential error types and causes of the sentence. With the pFOIL method, not only the error types but also the error causes and positions can be provided for CSL learners. Experimental results reveal that the pFOIL method outperforms the C4.5, maximum entropy, and Naive Bayes classifiers in error classification. Ru-Yng Chang, Chung-Hsien Wu 0001, Philips Kokoh Prasetyo |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2012 | Error Weighted Semi-Coupled Hidden Markov Model for Audio-Visual Emotion RecognitionabstractThis paper presents an approach to the automatic recognition of human emotions from audio-visual bimodal signals using an error weighted semi-coupled hidden Markov model (EWSC-HMM). The proposed approach combines an SC-HMM with a state-based bimodal alignment strategy and a Bayesian classifier weighting scheme to obtain the optimal emotion recognition result based on audio-visual bimodal fusion. The state-based bimodal alignment strategy in SC-HMM is proposed to align the temporal relation between audio and visual streams. The Bayesian classifier weighting scheme is then adopted to explore the contributions of the SC-HMM-based classifiers for different audio-visual feature pairs in order to obtain the emotion recognition output. For performance evaluation, two databases are considered: the MHMC posed database and the SEMAINE naturalistic database. Experimental results show that the proposed approach not only outperforms other fusion-based bimodal emotion recognition methods for posed expressions but also provides satisfactory results for naturalistic expressions. Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei |
IEEE Trans. Multim. | 2 |
| 2012 | MFTL: A Design and Implementation for MLC Flash Memory Storage SystemsabstractNAND flash memory has gained its popularity in a variety of applications as a storage medium due to its low power consumption, nonvolatility, high performance, physical stability, and portability. In particular, Multi-Level Cell (MLC) flash memory, which provides a lower cost and higher density solution, has occupied the largest part of NAND flash-memory market share. However, MLC flash memory also introduces new challenges: (1) Pages in a block must be written sequentially. (2) Information to indicate a page being obsoleted cannot be recorded in its spare area due to the limitation on the number of partial programming. Since most of applications access NAND flash memory under FAT file system, this article designs an MLC Flash Translation Layer (MFTL) for flash-memory storage systems which takes constraints of MLC flash memory and access behaviors of FAT file system into consideration. A series of trace-driven simulations was conducted to evaluate the performance of the proposed scheme. Although MFTL is designed for MLC flash memory and FAT file system, it is applicable to SLC flash memory and other file systems as well. Our experiment results show that the proposed MFTL could achieve a good performance for various access patterns even on SLC flash memory. Jen-Wei Hsieh, Chung-Hsien Wu 0001, Ge-Ming Chiu |
ACM Trans. Storage | 2 |
| 2011 | Semi-Coupled Hidden Markov Model with State-Based Alignment Strategy for Audio-Visual Emotion Recognition
Jen-Chun Lin, Chung-Hsien Wu 0001, Wen-Li Wei |
ACII (1) | 2 |
| 2011 | A Regression Approach to Affective Rating of Chinese Words from ANEW
Wen-Li Wei, Chung-Hsien Wu 0001, Jen-Chun Lin |
ACII (2) | 2 |
| 2011 | Emotion Detection Based on Concept Inference and Spoken Sentence Analysis for Customer Service
Ren-Ying Fang, Bo-Wei Chen, Jhing-Fa Wang, Chung-Hsien Wu 0001 |
INTERSPEECH | 4 |
| 2011 | Speech Indexing Using Semantic Context InferenceabstractAbstract This study presents a novel approach to spoken document retrieval based on semantic context inference for speech indexing. Each recognized term in a spoken document is mapped onto a semantic inference vector containing a bag of semantic terms through a semantic relation matrix. The semantic context inference vector is then constructed by summing up all the semantic inference vectors. Such a semantic term expansion and re-weighting make the semantic context inference vector a suitable representation for speech indexing. The experiments were conducted on 1550 anchor news stories collected from Mandarin Chinese broadcast news of 198 hours. The experimental results indicate that the proposed speech indexing using the semantic context inference contributes to a substantial performance improvement of spoken document retrieval. Index Terms : speech indexing, semantic context inference, spoken document retrieval 1. Introduction Speech is the most convenient way for the interaction of human-to-human and human-to-machine. The applications of spoken document retrieval in education, business and entertainment are rapidly growing. The recent attempts include multilingual oral history archives access [1], MIT lecture browsing [2], and the management of National Gallery consisting of speeches, news broadcasts and recordings [3], voice search about spoken dialog, call-routing systems [4], etc. All of them focus on retrieving the information to meet users' requirements. We know that it is not straightforward to directly compare the speech query with the spoken documents in the database. In order to construct an efficient and effective retrieval system, the state-of-the-art spoken document retrieval (SDR) technologies adopt the transcription obtained from automatic speech recognition for indexing. Vector space model [5] and probabilistic models (HMM [6], GMM [7], KL-divergence [8]), rely on certain similarity functions that assume a document is more likely to be relevant to a query if it contains more occurrences of query terms. The indexing techniques of text-based information retrieval have been widely adopted in spoken document retrieval. However, due to imperfect speech recognition results, out-of-vocabulary, and the ambiguity in homophone and word tokenization, conventional text-based indexing techniques are not always appropriate for spoken document retrieval. The transcription errors may cause undesired semantic and syntactic expression, thus result in an inadequate indexing. Several approaches have been proposed to address these problems with various indexing units such as word, sub-word, phone, and so on. The multi-level knowledge indexing approach considers three information sources including the speech transcription, keywords extracted from spoken documents, and hypernyms of the extracted keywords [9]. Hui et al. applied the Chien-Lin Huang, Bin Ma 0001, Haizhou Li 0001, Chung-Hsien Wu 0001 |
INTERSPEECH | 4 |
| 2011 | An Efficient Pre-Processing Scheme to Improve the Sound Source Localization System in Noisy Environment
Sheng-Chieh Lee, K. Bharanitharan, Bo-Wei Chen, Jhing-Fa Wang, Chung-Hsien Wu 0001, Min-Jian Liao |
INTERSPEECH | 5 |
| 2011 | Interactional Style Detection for Versatile Dialogue Response Using Prosodic and Semantic Features
Wei-Bin Liang, Chung-Hsien Wu 0001, Chih-Hung Wang, Jhing-Fa Wang |
INTERSPEECH | 2 |
| 2011 | Candidate Generation for ASR Output Error Correction Using a Context-Dependent Syllable Cluster-Based Confusion Matrix
Chao-Hong Liu, Chung-Hsien Wu 0001, David Sarwono, Jhing-Fa Wang |
INTERSPEECH | 2 |
| 2011 | Emotion Recognition of Affective Speech Based on Multiple Classifiers Using Acoustic-Prosodic Information and Semantic LabelsabstractThis work presents an approach to emotion recognition of affective speech based on multiple classifiers using acoustic-prosodic information (AP) and semantic labels (SLs). For AP-based recognition, acoustic and prosodic features including spectrum, formant, and pitch-related features are extracted from the detected emotional salient segments of the input speech. Three types of models, GMMs, SVMs, and MLPs, are adopted as the base-level classifiers. A Meta Decision Tree (MDT) is then employed for classifier fusion to obtain the AP-based emotion recognition confidence. For SL-based recognition, semantic labels derived from an existing Chinese knowledge base called HowNet are used to automatically extract Emotion Association Rules (EARs) from the recognized word sequence of the affective speech. The maximum entropy model (MaxEnt) is thereafter utilized to characterize the relationship between emotional states and EARs for emotion recognition. Finally, a weighted product fusion method is used to integrate the AP-based and SL-based recognition results for the final emotion decision. For evaluation, 2,033 utterances for four emotional states (Neutral, Happy, Angry, and Sad) are collected. The speaker-independent experimental results reveal that the emotion recognition performance based on MDT can achieve 80.00 percent, which is better than each individual classifier. On the other hand, an average recognition accuracy of 80.92 percent can be obtained for SL-based recognition. Finally, combining acoustic-prosodic information and semantic labels can achieve 83.55 percent, which is superior to either AP-based or SL-Based approaches. Moreover, considering the individual personality trait for personalized application, the recognition accuracy of the proposed approach can be further improved to 85.79 percent. Chung-Hsien Wu 0001, Wei-Bin Liang |
IEEE Trans. Affect. Comput. | 1 |
| 2011 | Interruption Point Detection of Spontaneous Speech Using Inter-Syllable Boundary-Based Prosodic FeaturesabstractThis article presents a probabilistic scheme for detecting the interruption point (IP) in spontaneous speech based on inter-syllable boundary-based prosodic features. Because of the high error rate in spontaneous speech recognition, a combined acoustic model considering both syllable and subsyllable recognition units, is firstly used to determine the inter-syllable boundaries and output the recognition confidence of the input speech. Based on the finding that IPs always occur at inter-syllable boundaries, a probability distribution of the prosodic features at the current potential IP is estimated. The Conditional Random Field (CRF) model, which employs the clustered prosodic features of the current potential IP and its preceding and succeeding inter-syllable boundaries, is employed to output the IP likelihood measure. Finally, the confidence of the recognized speech, the probability distribution of the prosodic features and the CRF-based IP likelihood measure are integrated to determine the optimal IP sequence of the input spontaneous speech. In addition, pitch reset and lengthening are also applied to improve the IP detection performance. The Mandarin Conversional Dialogue Corpus is adopted for evaluation. Experimental results show that the proposed IP detection approach obtains 10.56% and 6.5% more effective results than the hidden Markov model and the Maximum Entropy model respectively under the same experimental conditions. Besides, the IP detection error rate can be further reduced by 9.15% using pitch reset and lengthening information. The experimental results confirm that the proposed model based on inter-syllable boundary-based prosodic features can effectively detect the interruption point in spontaneous Mandarin speech. Chung-Hsien Wu 0001, Wei-Bin Liang, Jui-Feng Yeh |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2011 | Articulation-Disordered Speech Recognition Using Speaker-Adaptive Acoustic Models and Personalized Articulation PatternsabstractThis article presents a novel approach to speaker-adaptive recognition of speech from articulation-disordered speakers without a large amount of adaptation data. An unsupervised, incremental adaptation method is adopted for personalized model adaptation based on the recognized syllables with high recognition confidence from an automatic speech recognition (ASR) system. For articulation pattern discovery, the manually transcribed syllables and the corresponding recognized syllables are associated with each other using articulatory features. The Apriori algorithm is applied to discover the articulation patterns in the corpus, which are then used to construct a personalized pronunciation dictionary to improve the recognition accuracy of the ASR. The experimental results indicate that the proposed adaptation method achieves a syllable error rate reduction of 6.1%, outperforming the conventional adaptation methods that have a syllable error rate reduction of 3.8%. In addition, an average syllable error rate reduction of 5.04% is obtained for the ASR using the expanded pronunciation dictionary. Chung-Hsien Wu 0001, Hung-Yu Su, Han-Ping Shen |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2011 | Speaker Clustering Using Decision Tree-Based Phone Cluster Models With Multi-Space Probability DistributionsabstractThis paper presents an approach to speaker clustering using decision tree-based phone cluster models (DT-PCMs). In this approach, phone clustering is first applied to construct the universal phone cluster models to accommodate acoustic characteristics from different speakers. Since pitch feature is highly speaker-related and beneficial for speaker identification, the decision trees based on multi-space probability distributions (MSDs), useful to model both pitch and cepstral features for voiced and unvoiced speech simultaneously, are constructed. In speaker clustering based on DT-PCMs, contextual, phonetic, and prosodic features of each input speech segment is used to select the speaker-related MSDs from the MSD decision trees to construct the initial phone cluster models. The maximum-likelihood linear regression (MLLR) method is then employed to adapt the initial models to the speaker-adapted phone cluster models according to the input speech segment. Finally, the agglomerative clustering algorithm is applied on all speaker-adapted phone cluster models, each representing one input speech segment, for speaker clustering. In addition, an efficient estimation method for phone model merging is proposed for model parameter combination. Experimental results show that the MSD-based DT-PCMs outperform the conventional GMM- and HMM-based approaches for speaker clustering on the RT09 tasks. Han-Ping Shen, Jui-Feng Yeh, Chung-Hsien Wu 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Unsupervised Alignment of News Video and Text Using Visual Patterns and Textual ConceptsabstractA brief preview of a news video can be generated by semantically aligning the textual sentences of the anchor report, summarized by the anchor, with the visual field shots. Since accurately detecting the object in a visual shot is difficult and a textual term may generally correspond to several synonyms, the alignment of an anchor sentence with a video shot remains challenging. In this study, the temporal relation among the frames in a visual shot is characterized by a visual language model. The language model-based temporal relation is then applied to sentence-based alignment. The bag-of-word representations for the main objects in the key frames of a visual shot are firstly mapped to the visual patterns trained from the news video database. Furthermore, the textual terms in the report sentence are mapped to the textual concepts that are obtained from the HowNet knowledge base. Finally, unsupervised alignment between the textual concepts and the visual patterns in the news videos is performed using the IBM model-1. For evaluation, the visual pattern language model yields an alignment score of 0.77, exceeding that, 0.66, from the DTW method. Considering the performance for different news categories, visual pattern discovery and textual concept discovery can indeed improve the alignment performance in most news categories. Jun-Bin Yeh, Chung-Hsien Wu 0001, Sheng-Xiong Chang |
IEEE Trans. Multim. | 2 |
| 2010 | Discriminative Training for Near-Synonym Substitution
Liang-Chih Yu, Hsiu-Min Shih, Yu-Ling Lai, Jui-Feng Yeh, Chung-Hsien Wu 0001 |
COLING | 5 |
| 2010 | Pronunciation variation generation for spontaneous speech synthesis using state-based voice transformationabstractThis study presents an approach to Hidden Markov Models (HMM)-based spontaneous speech synthesis with pronunciation variation for better spontaneity. Pronunciation variation generally occurs in spontaneous speech and plays an important role in expressing the spontaneity. In this study, a state-based transformation function is adopted to model the relation between read speech and the corresponding spontaneous speech with pronunciation variations. The transformation function is then used to generate the state-based pronunciation variations. Due to the lack of training data, the articulatory features are used to cluster the transformation functions using Classification and Regression Trees (CARTs) such that the unseen pronunciation variation with the same articulatory features can be generated from the transformation function in the same cluster. Objective and subjective tests are conducted to evaluate the performance of the proposed approach. The experimental results show that the proposed transformation function achieves a significant improvement on spontaneity in synthesized speech. Chung-Han Lee, Chung-Hsien Wu 0001, Jun-Cheng Guo |
ICASSP | 2 |
| 2010 | Dialogue act detection in error-prone spoken dialogue systems using partial sentence tree and latent dialogue act matrix
Wei-Bin Liang, Chung-Hsien Wu 0001, Yu-Cheng Hsiao |
INTERSPEECH | 2 |
| 2010 | Prosodic word-based error correction in speech recognition using prosodic word expansion and contextual informationabstractIn this study, considering the effect of phrase grouping in spontaneous speech, prosodic words, instead of lexical words, are adopted as the units for error correction of speech recognition results. The prosodic words and the corresponding mis-recognized word fragments are obtained from a speech database to construct a mis-recognized word fragment table for the extracted prosodic words. For each word fragment in a recognized word sequence, the potential prosodic words which are likely to be misrecognized as input word fragments are retrieved from the table for prosodic word candidate expansion. The prosodic word-based contextual information, considering substitution and concatenation scores, is then employed into a probabilistic model to find the best word fragment sequence as the corrected output. Experimental results show that the proposed method achieved a 0.32 F1 score, with improvements of 0.18 and 0.10 compared to the SMT-based and lexical word-based approaches, respectively. Chao-Hong Liu, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2010 | Extraction of robust visual phrases using graph mining for image retrievalabstractFor the images of an objects with multi-viewpoints, the visual words in a visual phrase may be covered by the object and thus degrades the visual phrase extraction performance. This paper presents an approach to robust visual phrase extraction using graph mining for content-based image retrieval. In this study, the concurrent appearance of two visual words can be estimated over all of the category-related images in a database. The appearance frequencies of the visual words at each image are then used to construct a relation graph of visual words. Graph mining is utilized to mine the frequent dense subgraphs from the visual word relation graphs to extract the visual phrases. Experiments were conducted on the Caltech101 database and the experimental results show that the extracted visual phrases are robust to achieve a better retrieval performance than the pair-wise visual phrase approach. Jun-Bin Yeh, Chung-Hsien Wu 0001 |
ISCAS | 2 |
| 2010 | Design and Implementation for Multi-level Cell Flash Memory Storage SystemsabstractNAND flash memory has gained its popularity in a variety of applications as a storage medium due to its low power consumption, non-volatility, high performance, physical stability, and portability. In particular, Multi-Level Cell (MLC) flash memory, which provides a lower cost and higher density solution, has occupied the largest part of NAND flash-memory market share. However, MLC flash memory also introduces new challenges: (1) Pages in a block must be written sequentially. (2) Information to indicate a page being obsoleted cannot be recorded in its spare area. This paper designs an MLC Flash Translation Layer (MFTL) for flash-memory storage systems which takes new constraints of MLC flash memory and access behaviors of file system into consideration. A series of trace-driven simulations is conducted to evaluate the performance of the proposed scheme. Our experiment results show that the proposed MFTL outperforms other related works in terms of the number of extra page writes, the number of total block erasures, and the memory requirement for the management. Jen-Wei Hsieh, Chung-Hsien Wu 0001, Ge-Ming Chiu |
RTCSA | 2 |
| 2010 | Annotation and verification of sense pools in OntoNotes
Liang-Chih Yu, Chung-Hsien Wu 0001, Ru-Yng Chang, Chao-Hong Liu, Eduard H. Hovy |
Inf. Process. Manag. | 2 |
| 2010 | Exploiting Prosody Hierarchy and Dynamic Features for Pitch Modeling and Generation in HMM-Based Speech SynthesisabstractThis paper proposes a method for modeling and generating pitch in hidden Markov model (HMM)-based Mandarin speech synthesis by exploiting prosody hierarchy and dynamic pitch features. The prosodic structure of a sentence is represented by a prosody hierarchy, which is constructed from the predicted prosodic breaks using a supervised classification and regression tree (S-CART). The S-CART is trained by maximizing the proportional reduction of entropy to minimize the errors in the prediction of the prosodic breaks. The pitch contour of a speech sentence is estimated using the STRAIGHT algorithm and decomposed into the prosodic features (static features) at prosodic word, syllable, and frame layers, based on the predicted prosodic structure. Dynamic features at each layer are estimated to preserve the temporal correlation between adjacent units. A hierarchical prosody model is constructed using an unsupervised CART (U-CART) for generating pitch contour. Minimum description length (MDL) is adopted in U-CART training. Objective and subjective evaluations with statistical hypothesis testing were conducted, and the results compared to corresponding results for HMM-based pitch modeling. The comparison confirms the improved performance of the proposed method. Chi-Chun Hsia, Chung-Hsien Wu 0001, Jung-Yun Wu |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Introduction to the Special Section on Voice TransformationabstractThe 11 papers in this special section focus on voice transformation. Yannis Stylianou, Tomoki Toda, Chung-Hsien Wu 0001, Alexander Kain, Olivier Rosec |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Hierarchical Prosody Conversion Using Regression-Based Clustering for Emotional Speech SynthesisabstractThis paper presents an approach to hierarchical prosody conversion for emotional speech synthesis. The pitch contour of the source speech is decomposed into a hierarchical prosodic structure consisting of sentence, prosodic word, and subsyllable levels. The pitch contour in the higher level is encoded by the discrete Legendre polynomial coefficients. The residual, the difference between the source pitch contour and the pitch contour decoded from the discrete Legendre polynomial coefficients, is then used for pitch modeling at the lower level. For prosody conversion, Gaussian mixture models (GMMs) are used for sentence- and prosodic word-level conversion. At subsyllable level, the pitch feature vectors are clustered via a proposed regression-based clustering method to generate the prosody conversion functions for selection. Linguistic and symbolic prosody features of the source speech are adopted to select the most suitable function using the classification and regression tree for prosody conversion. Three small-sized emotional parallel speech databases with happy, angry, and sad emotions, respectively, were designed and collected for training and evaluation. Objective and subjective evaluations were conducted and the comparison results to the GMM-based method for prosody conversion achieved an improved performance using the hierarchical prosodic structure and the proposed regression-based clustering method. Chung-Hsien Wu 0001, Chi-Chun Hsia, Chung-Han Lee, Mai-Chun Lin |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Sentence Correction Incorporating Relative Position and Parse Template Language ModelsabstractSentence correction has been an important emerging issue in computer-assisted language learning. However, existing techniques based on grammar rules or statistical machine translation are still not robust enough to tackle the common errors in sentences produced by second language learners. In this paper, a relative position language model and a parse template language model are proposed to complement traditional language modeling techniques in addressing this problem. A corpus of erroneous English-Chinese language transfer sentences along with their corrected counterparts is created and manually judged by human annotators. Experimental results show that compared to a state-of-the-art phrase-based statistical machine translation system, the error correction performance of the proposed approach achieves a significant improvement using human evaluation. Chung-Hsien Wu 0001, Chao-Hong Liu, Matthew Harris, Liang-Chih Yu |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Regression-based clustering for hierarchical pitch conversionabstractThis study presents a hierarchical pitch conversion method using regression-based clustering for conversion function modeling. The pitch contour of a speech utterance is first extracted and decomposed into sentence-, word and sub-syllable-level features in a top-down mechanism. The pair-wise source and target pitch feature vectors at each level are then clustered to generate the pitch conversion function. Regression-based clustering, which clusters the feature vectors to achieve a minimum conversion error between the predicted and the real feature vectors is proposed for conversion function generation. A classification and regression tree (CART), incorporating linguistic, phonetic and source prosodic features, is adopted to select the most suitable function for pitch conversion. Several objective and subjective evaluations were conducted and the comparison results to the GMM-based methods for pitch conversion confirm the performance of the proposed regression-based clustering approach. Chung-Han Lee, Chi-Chun Hsia, Chung-Hsien Wu 0001, Mai-Chun Lin |
ICASSP | 3 |
| 2009 | Semantic role labeling with discriminative feature selection for spoken language understanding
Chao-Hong Liu, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2009 | Extraction of Query Term-related Visual Phrases for News Video Retrieval using Mutual InformationabstractThis paper presents an approach to query term-related visual phrases extraction using mutual information for object-based news video retrieval. As visual words are useful for object representation, unstable visual words generally appear in the frame sequence of a shot. Using the appearance frequency of the visual words in a sliding window over the query term-related stories, the appearance pattern of a visual word is adopted to characterize the visual word. Based on the appearance pattern of a visual word, the mutual information between two visual words can be estimated over all of the extracted stories. The mutual information is then used to construct a visual word relation graph. Visual phrases are then extracted by discovering the complete sub-graphs from the visual word relation graph for news video retrieval. Experiments were conducted on the MATBN news video database and the experimental results show that a good precision rate for video news retrieval can be achieved. Jun-Bin Yeh, Chung-Hsien Wu 0001 |
ISCAS | 2 |
| 2009 | Psychiatric document retrieval using a discourse-aware model
Liang-Chih Yu, Chung-Hsien Wu 0001, Fong-Lin Jang |
Artif. Intell. | 2 |
| 2009 | Speech-Annotated Photo Retrieval Using Syllable-Transformed PatternsabstractThis study presents a novel indexing and retrieval scheme for digital photos with speech annotations based on syllable-transformed image-like patterns. Speech recognition error and out-of-vocabulary (OOV) problems generally result in incorrect indexing and degrade the retrieval performance. In this study, the recognizedn-best candidates used to deal with recognition error problems are transformed into an image-like pattern using multidimensional scaling. A hybrid mechanism integrating syllables, characters, words, and image-like patterns is exploited for speech indexing and retrieval. Experiments show the hybrid indexing method integrating the syllable-transformed image-like patterns can achieve a better result compared to previous indexing methods. Chung-Hsien Wu 0001, Chien-Lin Huang, Wei-Chuan Lee, Yu-Sheng Lai |
IEEE Signal Process. Lett. | 1 |
| 2009 | Introduction to the Special Issue on Recent Advances in Asian Language Spoken Document Retrievalabstractintroduction Introduction to the Special Issue on Recent Advances in Asian Language Spoken Document Retrieval Share on Authors: Chung-Hsien Wu National Cheng Kung University National Cheng Kung UniversityView Profile , Haizhou Li Institute for Infocomm Research Institute for Infocomm ResearchView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 8Issue 1March 2009 Article No.: 1pp 1–3https://doi.org/10.1145/1482343.1482344Online:01 March 2009Publication History 0citation192DownloadsMetricsTotal Citations0Total Downloads192Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Chung-Hsien Wu 0001, Haizhou Li 0001 |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2009 | Improving Structural Statistical Machine Translation for Sign Language With Small Corpus Using Thematic Role Templates as Translation MemoryabstractThis paper presents a structural statistical machine translation (SSMT) model to deal with the data sparseness problem that occurs as a result of the necessarily small corpus to translate Chinese into Taiwanese Sign Language (TSL). A parallel bilingual corpus was developed, and linguistic information from the Sinica Treebank is adopted for Chinese sentence analysis. The synchronous context free grammar (SCFG) was adopted to convert a Chinese structure to the corresponding TSL structure and then extract a translation memory which comprises the thematic relations between the grammar rules of both structures. In structural translation, the statistical MT (SMT) approach was used to align the thematic roles in the grammar rules and the translation memory provides the reference templates for TSL structure translation. Finally, the agreement information for TSL verbs was labeled for enriching the expressiveness of the translated TSL sequence. Several experiments were conducted to evaluate the translation performance and the communication effectiveness for the deaf. The evaluation results demonstrate that the proposed approach outperforms a baseline statistical MT system using the same small corpus, especially for the translation of long sentences. Hung-Yu Su, Chung-Hsien Wu 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Story Segmentation and Topic Classification of Broadcast News via a Topic-Based Segmental Model and a Genetic AlgorithmabstractThis paper presents a two-stage approach to story segmentation and topic classification of broadcast news. The two-stage paradigm adopts a decision tree and a maximum entropy model to identify the potential story boundaries in the broadcast news within a sliding window. The problem for story segmentation is thus transformed to the determination of a boundary position sequence from the potential boundary regions. A genetic algorithm is then applied to determine the chromosome, which corresponds to the final boundary position sequence. A topic-based segmental model is proposed to define the fitness function applied in the genetic algorithm. The syllable- and word-based story segmentation schemes are adopted to evaluate the proposed approach. Experimental results indicate that a miss probability of 0.1587 and a false alarm probability of 0.0859 are achieved for story segmentation on the collected broadcast news corpus. On the TDT-3 Mandarin audio corpus, a miss probability of 0.1232 and a false alarm probability of 0.1298 are achieved. Moreover, an outside classification accuracy of 74.55% is obtained for topic classification on the collected broadcast news, while an inside classification accuracy of 88.82% is achieved on the TDT-2 Mandarin audio corpus. Chung-Hsien Wu 0001, Chia-Hsin Hsieh |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Idiolect Extraction and Generation for Personalized Speaking Style ModelingabstractA person's speaking style, consisting of such attributes as voice, choice of vocabulary, and the physical motions employed, not only expresses the speaker's identity but also emphasizes the content of an utterance. Speech combining these aspects of speaking style becomes more vivid and expressive to listeners. Recent research on speaking style modeling has paid more attention to speech signal processing. This approach focuses on text processing for idiolect extraction and generation to model a specific person's speaking style for the application of text-to-speech (TTS) conversion. The first stage of this study adopts a statistical method to automatically detect the candidate idiolects from a personalized, transcribed speech corpus. Based on the categorization of the detected candidate idiolects, superfluous idiolects are extracted using the fluency measure while the remaining candidates are regarded as the nonsuperfluous idiolects. In idiolect generation, the input text is converted into a target text with a particular speaker's speaking style via the insertion of superfluous idiolect or synonym substitution of nonsuperfluous idiolect. To evaluate the performance of the proposed methods, experiments were conducted on a Chinese corpus collected and transcribed from the speech files of three Taiwanese politicians. The results show that the proposed method can effectively convert a source text into a target text with a personalized speaking style. Chung-Hsien Wu 0001, Chung-Han Lee, Chung-Hau Liang |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | OntoNotes: Corpus Cleanup of Mistaken Agreement Using Word Sense Disambiguation
Liang-Chih Yu, Chung-Hsien Wu 0001, Eduard H. Hovy |
COLING | 2 |
| 2008 | Automatic assessment of articulation disorders using confident unit-based model adaptationabstractThis paper presents an approach to automatic assessment on articulation disorders using unsupervised acoustic model adaptation. Prior knowledge is obtained via the phonological analysis of the speech data from 453 articulation disordered children. A confusion matrix of the recognition units for a specific subject is re-estimated based on the prior knowledge and the recognition results to choose the confident units for adaptation. The adapted acoustic models can effectively improve the recognition performance of the disordered speech and thus used for articulation assessment. In the experiments, the proposed unsupervised adaptation method achieved a significant performance improvement of 9.1% for disordered speech on syllable recognition rate. Automatic assessment also shows encouraging consistency to the assessment from the therapist. Hung-Yu Su, Chung-Hsien Wu 0001, Pei-Jen Tsai |
ICASSP | 2 |
| 2008 | Unsupervised pronunciation grammar growing using knowledge-based and data-driven approachesabstractThis study presents a novel approach to unsupervised pronunciation grammar growing for non-native speech recognition. Unsupervised pronunciation grammar growing includes pronunciation variation graph construction and non-native grammar generation. Knowledge-based and data-driven approaches are considered for variation graph construction. The measurement of confidence and support is used for grammar selection. Experiments show that unsupervised pronunciation grammar growing is suitable for the improvement of non-native speech recognition. Chien-Lin Huang, Chung-Hsien Wu 0001, Haizhou Li 0001, Chia-Hsin Hsieh, Bin Ma 0001 |
ICME | 2 |
| 2008 | Interruption point detection of spontaneous speech using prior knowledge and multiple featuresabstractThis paper presents an approach to interruption point (IP) detection of spontaneous speech based on conditional random fields using prior knowledge and multiple features. The features adopted in this study consist of subsyllable boundaries and prosodic features. Conditional random fields (CRFs) and variable-length contextual features are employed for IP modeling. In order to apply the features with continuous values to the CRF models, the K-means clustering algorithm is adopted for the quantization of the prosodic features. In the experimental results, Mandarin Conversional Dialogue Corpus (MCDC) was used to evaluate the proposed method. The IP detection error rate achieved almost 20% reduction in Rt04 measure. The experimental results show that the proposed model can effectively detect the interruption point in spontaneous speech. Wei-Bin Liang, Jui-Feng Yeh, Chung-Hsien Wu 0001, Chi-Chiuan Liou |
ICME | 3 |
| 2008 | Adaptive decision tree-based phone cluster models for speaker clustering
Chia-Hsin Hsieh, Chung-Hsien Wu 0001, Han-Ping Shen |
INTERSPEECH | 2 |
| 2008 | Robust speaker verification using short-time frequency with long-time window and fusion of multi-resolutionsabstractThis study presents a novel approach of feature analysis to speaker verification. There are two main contributions in this paper. First, the feature analysis of short-time frequency with long-time window (SFLW) is a compact feature for the efficiency of speaker verification. The purpose of SFLW is to take account of short-time frequency characteristics and longtime resolution at the same time. Secondly, the fusion of multi-resolutions is used for the effectiveness of robust speaker verification. The speaker verification system can be further improved using multi-resolution features. The experimental results indicate that the proposed approaches not only speed up the processing time but also improve the performance of speaker verification. Chien-Lin Huang, Bin Ma 0001, Chung-Hsien Wu 0001, Brian Kan-Wing Mak, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2008 | Ontology-based speech act identification in a bilingual dialog system using partial pattern treesabstractAbstract This article presents a bilingual ontology‐based dialog system with multiple services. An ontology‐alignment algorithm is proposed to integrate ontologies of different languages for cross‐language applications. A domain‐specific ontology is further extracted from the bilingual ontology using an island‐driven algorithm and a domain corpus. This study extracts the semantic words/concepts using latent semantic analysis (LSA). Based on the extracted semantic words and the domain ontology, a partial pattern tree is constructed to model the speech act of a spoken utterance. The partial pattern tree is used to deal with the ill‐formed sentence problem in a spoken‐dialog system. Concept expansion based on domain ontology is also adopted to improve system performance. For performance evaluation, a medical dialog system with multiple services, including registration information, clinic information, and FAQ information, is implemented. Four performance measures were used separately for evaluation. The speech act identification rate was 86.2%. A task success rate of 77% was obtained. The contextual appropriateness of the system response was 78.5%. Finally, the rate for correct FAQ retrieval was 82%, an improvement of 15% over the keyword‐based vector‐space model. The results show the proposed ontology‐based speech‐act identification is effective for dialog management. Jui-Feng Yeh, Chung-Hsien Wu 0001, Ming-Jun Chen |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Stochastic vector mapping-based feature enhancement using prior-models and model adaptation for noisy speech recognition
Chia-Hsin Hsieh, Chung-Hsien Wu 0001 |
Speech Commun. | 2 |
| 2008 | HAL-Based Evolutionary Inference for Pattern Induction From Psychiatry Web ResourcesabstractNegative and stressful life events play a significant role in triggering depressive episodes. Psychiatric services that can identify such events efficiently are vital for mental health care and prevention. Meaningful patterns, e.g.,, must be extracted from psychiatric texts before these services can be provided. This study presents an evolutionary text-mining framework capable of inducing variable-length patterns from unannotated psychiatry Web resources. The proposed framework can be divided into two parts: 1) a cognitive motivated model such as hyperspace analog to language (HAL) and 2) an evolutionary inference algorithm (EIA). The HAL model constructs a high-dimensional context space to represent words as well as combinations of words. Based on the HAL model, the EIA bootstraps with a small set of seed patterns, and then iteratively induces additional relevant patterns. To avoid moving in the wrong direction, the EIA further incorporates relevance feedback to guide the induction process. Experimental results indicate that combining the HAL model and relevance feedback enables the EIA to not only induce patterns from the unannotated Web corpora, but also achieve useful results in a reasonable amount of time. The proposed framework thus significantly reduces reliance on annotated corpora. Liang-Chih Yu, Chung-Hsien Wu 0001, Jui-Feng Yeh, Fong-Lin Jang |
IEEE Trans. Evol. Comput. | 2 |
| 2008 | Extended probabilistic HAL with close temporal association for psychiatric query document retrievalabstractPsychiatric query document retrieval can assist individuals to locate query documents relevant to their depression-related problems efficiently and effectively. By referring to relevant documents, individuals can understand how to alleviate their depression-related symptoms according to recommendations from health professionals. This work presents an extended probabilistic Hyperspace Analog to Language ( epHAL ) model to achieve this aim. The epHAL incorporates the close temporal associations between words in query documents to represent word cooccurrence relationships in a high-dimensional context space. The information flow mechanism further combines the query words in the epHAL space to infer related words for effective information retrieval. The language model perplexity is considered as the criterion for model optimization. Finally, the epHAL is adopted for psychiatric query document retrieval, and indicates its superiority in information retrieval over traditional approaches. Jui-Feng Yeh, Chung-Hsien Wu 0001, Liang-Chih Yu, Yu-Sheng Lai |
ACM Trans. Inf. Syst. | 2 |
| 2007 | Topic Analysis for Psychiatric Document Retrieval
Liang-Chih Yu, Chung-Hsien Wu 0001, Chin-Yew Lin, Eduard H. Hovy, Chia-Ling Lin |
ACL | 2 |
| 2007 | Conversion Function Clustering and Selection for Expressive Voice ConversionabstractIn this study, a conversion function clustering and selection approach to conversion-based expressive speech synthesis is proposed. First, a set of small-sized emotional parallel speech databases is designed and collected to train the conversion functions. Gaussian mixture bi-gram model (GMBM) is adopted as the conversion function to model the temporal and spectral evolution of speech. Conversion functions initially constructed from the parallel sub-syllable pairs in the speech database are clustered based on linguistic and spectral information. Subjective and objective evaluations with statistical hypothesis testing were conducted to evaluate the quality of the converted speech. The results show that the proposed method exhibits encouraging potential in conversion-based expressive speech synthesis. Chi-Chun Hsia, Chung-Hsien Wu 0001, Jian-Qi Wu |
ICASSP (4) | 2 |
| 2007 | Phone Set Generation Based on Acoustic and Contextual Analysis for Multilingual Speech RecognitionabstractThis study presents a novel approach to generating phone units generation for the recognition of multilingual speech. Acoustic and contextual analysis is performed to characterize multilingual phonetic units for phone set generation. A confusion matrix combining acoustic and contextual similarities between every two phonetic units is constructed for phonetic unit clustering. Acoustic likelihood and hyperspace analog to language (HAL) model are adopted for acoustic similarity and contextual similarity estimation of phone models, respectively. Experiments show that the generated phone set provides a compact and robust set that considers acoustic and contextual information for multilingual speech recognition. Chien-Lin Huang, Chung-Hsien Wu 0001 |
ICASSP (4) | 2 |
| 2007 | Disfluency correction of spontaneous speech using conditional random fields with variable-length features
Jui-Feng Yeh, Chung-Hsien Wu 0001, Wei-Yen Wu |
INTERSPEECH | 2 |
| 2007 | Magic MirrorabstractThis investigation describes a novel design and implementation of an interactive multimedia mirror system, called "Magic Mirror." The Magic Mirror can be easily implemented in existing personal computers or hand-held device with normal peripherals and regular reflective glass by integrating image/speech processing, Internet connectivity, and 3D and multimedia software. The integrated Magic Mirror, which includes speech recognition, speech synthesis, face detection/modified/recognition, 3D virtual genius, hidden LCD mirror, and camera, performs simple syndication to capture information about peripherals and network connections. The user can easily activate personal multimedia services using verbal commands. The Magic Mirror can function like a good friend who listens to the user's questions and automatically responds to these requests, providing relaxation and consolation. Moreover, the Magic Mirror can detect a user's feeling based on speech and image recognition features to select the appropriate music and speech to alter the user's mood. Jun-Ren Ding, Chien-Lin Huang, Ji-Kun Lin, Jar-Ferr Yang, Chung-Hsien Wu 0001 |
ISM | 5 |
| 2007 | Joint Optimization of Word Alignment and Epenthesis Generation for Chinese to Taiwanese Sign SynthesisabstractThis work proposes a novel approach to translate Chinese to Taiwanese sign language and to synthesize sign videos. An aligned bilingual corpus of Chinese and Taiwanese Sign Language (TSL) with linguistic and signing information is also presented for sign language translation. A two-pass alignment in syntax level and phrase level is developed to obtain the optimal alignment between Chinese sentences and Taiwanese sign sequences. For sign video synthesis, a scoring function is presented to develop motion transition-balanced sign videos with rich combinations of intersign transitions. Finally, the maximum a posteriori (MAP) algorithm is employed for sign video synthesis based on joint optimization of two-pass word alignment and intersign epenthesis generation. Several experiments are conducted in an educational environment to evaluate the performance on the comprehension of sign expression. The proposed approach outperforms the IBM Model 2 in sign language translation. Moreover, deaf students perceived sign videos generated by the proposed method to be satisfactory. Yu-Hsien Chiu, Chung-Hsien Wu 0001, Hung-Yu Su, Chih-Jen Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Transfer-based statistical translation of Taiwanese sign language using PCFGabstractThis article presents a transfer-based statistical model for Chinese to Taiwanese sign-language (TSL) translation. Two sets of probabilistic context-free grammars (PCFGs) are derived from a Chinese Treebank and a bilingual parallel corpus. In this approach, a three-stage translation model is proposed. First, the input Chinese sentence is parsed into possible phrase structure trees (PSTs) based on the Chinese PCFGs. Second, the Chinese PSTs are then transferred into TSL PSTs according to the transfer probabilities between the context-free grammar (CFG) rules of Chinese and TSL derived from the bilingual parallel corpus. Finally, the TSL PSTs are used to generate the possible translation results. The Viterbi algorithm is adopted to obtain the best translation result via the three-stage translation. For evaluation, three objective evaluation metrics including AER, Top-N, and BLUE and one subjective evaluation metric using MOS were used. Experimental results show that the proposed approach outperforms the IBM Model 3 in the task of Chinese to sign-language translation. Chung-Hsien Wu 0001, Hung-Yu Su, Yu-Hsien Chiu |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2007 | Spoken Document Retrieval Using Multilevel Knowledge and Semantic VerificationabstractThis study presents a novel approach to spoken document retrieval based on multilevel knowledge indexing and semantic verification. Multilevel knowledge indexing considers three information sources, namely transcription data, keywords extracted from spoken documents, and hypernyms of the extracted keywords. A semantic network with forward-backward propagation is presented for semantic verification of the retrieved documents. In the forward step for semantic verification, a bag of keywords is chosen based on word significance measures. Semantic relations are estimated and adopted for verification in the backward procedure. The verification score is then utilized to weight and rerank the retrieved documents to obtain the final results. Experiments are performed on 40 h of anchor speech extracted from 198 h of collected broadcast news. Experimental results indicate that multilevel knowledge indexing and semantic verification achieve better retrieval results than other indexing schemes. Chien-Lin Huang, Chung-Hsien Wu 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Variable-Length Unit Selection in TTS Using Structural Syntactic CostabstractThis paper presents a variable-length unit selection scheme based on syntactic cost to select text-to-speech (TTS) synthesis units. The syntactic structure of a sentence is derived from a probabilistic context-free grammar (PCFG), and represented as a syntactic vector. The syntactic difference between target and candidate units (words or phrases) is estimated by the cosine measure with the inside probability of PCFG acting as a weight. Latent semantic analysis (LSA) is applied to reduce the dimensionality of the syntactic vectors. The dynamic programming algorithm is adopted to obtain a concatenated unit sequence with minimum cost. A syntactic property-rich speech database is designed and collected as the unit inventory. Several experiments with statistical testing are conducted to assess the quality of the synthetic speech as perceived by human subjects. The proposed method outperforms the synthesizer without considering syntactic property. The structural syntax estimates the substitution cost better than the acoustic features alone Chung-Hsien Wu 0001, Chi-Chun Hsia, Jiun-Fu Chen, Jhing-Fa Wang |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Conversion Function Clustering and Selection Using Linguistic and Spectral Information for Emotional Voice ConversionabstractIn emotional speech synthesis, a large speech database is required for high-quality speech output. Voice conversion needs only a compact-sized speech database for each emotion. This study designs and accumulates a set of phonetically balanced small- sized emotional parallel speech databases to construct conversion functions. The Gaussian mixture bigram model (GMBM) is adopted as the conversion function to characterize the temporal and spectral evolution of the speech signal. The conversion function is initially constructed for each instance of parallel subsyllable pairs in the collected speech database. To reduce the total number of conversion functions and select an appropriate conversion function, this study presents a framework by incorporating linguistic and spectral information for conversion function clustering and selection. Subjective and objective evaluations with statistical hypothesis testing are conducted to evaluate the quality of the converted speech. The proposed method compares favorably with previous methods in conversion-based emotional speech synthesis. Chi-Chun Hsia, Chung-Hsien Wu 0001, Jian-Qi Wu |
IEEE Trans. Computers | 2 |
| 2007 | Generation of Phonetic Units for Mixed-Language Speech Recognition Based on Acoustic and Contextual AnalysisabstractThis work presents a novel approach to generating phonetic units in order to recognize mixed-language or multilingual speech. Acoustic and contextual analysis is performed to characterize multilingual phonetic units for phone set creation. Acoustic likelihood is utilized for similarity estimation of phone models. The hyperspace analog to language (HAL) model is adopted for contextual modeling and contextual similarity estimation. A confusion matrix combining acoustic and contextual similarities between every two phonetic units is built for phonetic unit clustering. Multidimensional scaling (MDS) method is applied to the confusion matrix for reducing dimensionality. Experimental results indicate that the created phonetic set provides a compact and robust set that considers acoustic and contextual information for mixed-language or multilingual speech recognition. Chien-Lin Huang, Chung-Hsien Wu 0001 |
IEEE Trans. Computers | 2 |
| 2007 | Psychiatric Consultation Record Retrieval Using Scenario-Based Representation and Multilevel Mixture ModelabstractPsychiatric consultation record retrieval attempts to help people to efficiently and effectively locate the consultation records relevant to their depressive problems. Consultation records can also make people aware that they are not alone, because many individuals have suffered from the same or similar problems. Additionally, people can understand how to alleviate their depressive symptoms according to recommendations from health professionals. To achieve this goal, this paper proposes the use of a scenario-based representation, i.e., a symptom-based structural representation, to capture the depressive symptoms and their semantic relations, such as cause-effect and temporal relations, for understanding users' queries clearly. The symptoms and relations are identified from semantic mining and analysis of consultation records. The multilevel mixture model is adopted to estimate the relevance of queries and consultation records based on the structural information. Experimental results show that the proposed approach achieves higher precision than does a term-based flat representation. An experiment is also conducted to examine the effect of error propagation resulting from incorrect identification of symptoms and relations. Experimental results demonstrate that combining different approaches can improve the retrieval robustness. Liang-Chih Yu, Chung-Hsien Wu 0001, Fong-Lin Jang |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2007 | Speech Sentence Compression Based on Speech Segment Extraction and ConcatenationabstractThis correspondence presents a speech sentence compression scheme. A compressed word sequence is first extracted. Speech segments, in the spoken document, corresponding to the extracted words are selected for concatenation. Evaluation of the proposed approach shows the compressed speech sentence retains important and meaningful information and naturalness Chung-Hsien Wu 0001, Chia-Hsin Hsieh, Chien-Lin Huang |
IEEE Trans. Multim. | 1 |
| 2006 | Stochastic Discourse Modeling in Spoken Dialogue Systems Using Semantic Dependency Graphs
Jui-Feng Yeh, Chung-Hsien Wu 0001, Mao-Zhu Yang |
ACL | 2 |
| 2006 | HAL-Based Cascaded Model for Variable-Length Semantic Pattern Induction from Psychiatry Web Resources
Liang-Chih Yu, Chung-Hsien Wu 0001, Fong-Lin Jang |
ACL | 2 |
| 2006 | Stochastic vector mapping-based feature enhancement using prior model and environment adaptation for noisy speech recognition
Chia-Hsin Hsieh, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2006 | Map-based adaptation for speech conversion using adaptation data selection and non-parallel training
Chung-Han Lee, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2006 | Emotion recognition from text using semantic labels and separable mixture modelsabstractThis study presents a novel approach to automatic emotion recognition from text. First, emotion generation rules (EGRs) are manually deduced from psychology to represent the conditions for generating emotion. Based on the EGRs, the emotional state of each sentence can be represented as a sequence of semantic labels (SLs) and attributes (ATTs); SLs are defined as the domain-independent features, while ATTs are domain-dependent. The emotion association rules (EARs) represented by SLs and ATTs for each emotion are automatically derived from the sentences in an emotional text corpus using the a priori algorithm. Finally, a separable mixture model (SMM) is adopted to estimate the similarity between an input sentence and the EARs of each emotional state. Since some features defined in this approach are domain-dependent, a dialog system focusing on the students' daily expressions is constructed, and only three emotional states, happy, unhappy, and neutral , are considered for performance evaluation. According to the results of the experiments, given the domain corpus, the proposed approach is promising, and easily ported into other domains. Chung-Hsien Wu 0001, Ze-Jing Chuang, Yu-Chung Lin |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2006 | Automatic segmentation and identification of mixed-language speech using delta-BIC and LSA-based GMMsabstractThis paper proposes an approach to segmenting and identifying mixed-language speech. A delta Bayesian information criterion (delta-BIC) is firstly applied to segment the input speech utterance into a sequence of language-dependent segments using acoustic features. A VQ-based bi-gram model is used to characterize the acoustic-phonetic dynamics of two consecutive codewords in a language. Accordingly the language-specific acoustic-phonetic property of sequence of phones was integrated in the identification process. A Gaussian mixture model (GMM) is used to model codeword occurrence vectors orthonormally transformed using latent semantic analysis (LSA) for each language-dependent segment. A filtering method is used to smooth the hypothesized language sequence and thus eliminate noise-like components of the detected language sequence generated by the maximum likelihood estimation. Finally, a dynamic programming method is used to determine globally the language boundaries. Experimental results show that for Mandarin, English, and Taiwanese, a recall rate of 0.87 for language boundary segmentation was obtained. Based on this recall rate, the proposed approach achieved language identification accuracies of 92.1% and 74.9% for single-language and mixed-language speech, respectively. Chung-Hsien Wu 0001, Yu-Hsien Chiu, Chi-Jiun Shia |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Multiple change-point audio segmentation and classification using an MDL-based Gaussian modelabstractThis study presents an approach for segmenting and classifying an audio stream based on audio type. First, a silence deletion procedure is employed to remove silence segments in the audio stream. A minimum description length (MDL)-based Gaussian model is then proposed to statistically characterize the audio features. Audio segmentation segments the audio stream into a sequence of homogeneous subsegments using the MDL-based Gaussian model. A hierarchical threshold-based classifier is then used to classify each subsegment into different audio types. Finally, a heuristic method is adopted to smooth the subsegment sequence and provide the final segmentation and classification results. Experimental results indicate that for TDT-3 news broadcast, a missed detection rate (MDR) of 0.1 and a false alarm rate (FAR) of 0.14 were achieved for audio segmentation. Given the same MDR and FAR values, segment-based audio classification achieved a better classification accuracy of 88% compared to a clip-based approach. Chung-Hsien Wu 0001, Chia-Hsin Hsieh |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Voice conversion using duration-embedded bi-HMMs for expressive speech synthesisabstractThis paper presents an expressive voice conversion model (DeBi-HMM) as the post processing of a text-to-speech (TTS) system for expressive speech synthesis. DeBi-HMM is named for its duration-embedded characteristic of the two HMMs for modeling the source and target speech signals, respectively. Joint estimation of source and target HMMs is exploited for spectrum conversion from neutral to expressive speech. Gamma distribution is embedded as the duration model for each state in source and target HMMs. The expressive style-dependent decision trees achieve prosodic conversion. The STRAIGHT algorithm is adopted for the analysis and synthesis process. A set of small-sized speech databases for each expressive style is designed and collected to train the DeBi-HMM voice conversion models. Several experiments with statistical hypothesis testing are conducted to evaluate the quality of synthetic speech as perceived by human subjects. Compared with previous voice conversion methods, the proposed method exhibits encouraging potential in expressive speech synthesis. Chung-Hsien Wu 0001, Chi-Chun Hsia, Te-Hsien Liu, Jhing-Fa Wang |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Edit disfluency detection and correction using a cleanup language model and an alignment modelabstractThis investigation presents a novel approach to detecting and correcting the edit disfluency in spontaneous speech. Hypothesis testing using acoustic features is first adopted to detect potential interruption points (IPs) in the input speech. The word order of the cleanup utterance is then cleaned up based on the potential IPs using a class-based cleanup language model, the deletable region and the correction are aligned using an alignment model. Finally, log linear weighting is applied to optimize the performance. Using the acoustic features, the IP detection rate is significantly improved especially in recall rate. Based on the positions of the potential IPs, the cleanup language model and the alignment model are able to detect and correct the edit disfluency efficiently. Experimental results demonstrate that the proposed approach has achieved error rates of 0.33 and 0.21 for IP detection and edit word deletion, respectively. Jui-Feng Yeh, Chung-Hsien Wu 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Movement Epenthesis Generation Using NURBS-Based Spatial InterpolationabstractThis study proposes a novel approach to the generation of video-based movement epenthesis for sign language and concentrates on a spatial interpolation approach using a nonuniform rational B-spline function to produce a smooth interpolation curve. To generate movement epenthesis, the beginning and end "cut points" are determined based on the concatenation cost, which is a linear combination of the distance, smoothness, and image distortion costs. The distance cost is defined as the normalized Euclidian distance between the palm locations at two cut points. The smoothness cost is determined by accumulating the second derivative of the curve. The image distortion cost is then estimated from the normalized Euclidian distance between the real and generated hand images. To evaluate the proposed approach, a set of sign video databases is collected and preprocessed with image calibration, content annotation, and principle component analysis. An image component overlapping procedure is also employed to yield a smooth sign video output. Evaluation results demonstrate that the synthesized trajectory is very close to the original trajectory. Moreover, the system was evaluated subjectively based on the results of ten hearing and five deaf people. The evaluation results also demonstrate the stability and feasibility of the proposed approach Ze-Jing Chuang, Chung-Hsien Wu 0001, Wei-Sheng Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Semantic Segment Extraction and Matching for Internet FAQ RetrievalabstractThis investigation presents a novel approach to semantic segment extraction and matching for retrieving information from Internet FAQs with natural language queries. Two semantic segments, the question category segment (QS) and the keyword segment (KS), are extracted from the input queries and the FAQ questions with a semiautomatically derived question-semantic grammar. A semantic matching method is presented to estimate the similarity between the semantic segments of the query and the questions in the FAQ collection. Additionally, the vector space model (VSM) is adopted to measure the similarity between the query and the answers of the QA pairs. Finally, a multistage ranking strategy is adopted to determine the optimally performing combination of similarity metrics. The experimental results illustrate that the proposed method achieves an average rank of 4.52 and a top-10 recall rate of 90.89 percent. Compared with the query-expansion method, this method improves the performance by 4.82 places in the average rank of correct answers, 25.34 percent in the top-5 recall rate, and 5.21 percent in the top-10 recall rate. Chung-Hsien Wu 0001, Jui-Feng Yeh, Yu-Sheng Lai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | Hand Motion Recognition for the Vision-based Taiwanese Sign Language Interpretation
Chia-Shiuan Cheng, Pi-Fuei Hsieh, Chung-Hsien Wu 0001 |
ACII | 3 |
| 2005 | IG-Based Feature Extraction and Compensation for Emotion Recognition from Speech
Ze-Jing Chuang, Chung-Hsien Wu 0001 |
ACII | 2 |
| 2005 | Vision-Based Recognition of Hand Shapes in Taiwanese Sign Language
Jung-Ning Huang, Pi-Fuei Hsieh, Chung-Hsien Wu 0001 |
ACII | 3 |
| 2005 | Facial Phoneme Extraction for Taiwanese Sign Language Recognition
Shi-Hou Lin, Pi-Fuei Hsieh, Chung-Hsien Wu 0001 |
ACII | 3 |
| 2005 | Spoken document summarization using acoustic, prosodic and semantic informationabstractThis paper presents a spoken document summarization scheme using acoustic, prosodic, and semantic information. First, speech recognition confidence is estimated to choose reliable words from the speech transcription. Prosodic information, including pitch and energy, is used for stressed word selection. Latent semantic indexing (LSI) is adopted to identify significant words. Finally, word trigram and semantic dependency is measured to include the syntactic and semantic information for speech summarization. The dynamic programming (DP) algorithm is used to find the best summarization result according to the summarization score estimated from the above five measures. Finally, the summarized result is presented by the concatenation of the summarized speech words. Experimental results indicate that the proposed approach effectively extracts important words and gives a promising speech summary. Chien-Lin Huang, Chia-Hsin Hsieh, Chung-Hsien Wu 0001 |
ICME | 3 |
| 2005 | Duration-embedded bi-HMM for expressive voice conversion
Chi-Chun Hsia, Chung-Hsien Wu 0001, Te-Hsien Liu |
INTERSPEECH | 2 |
| 2005 | Audio-video summarization of TV news using speech recognition and shot change detection
Chien-Lin Huang, Chia-Hsin Hsieh, Chung-Hsien Wu 0001 |
INTERSPEECH | 3 |
| 2005 | Domain-specific FAQ retrieval using independent aspectsabstractThis investigation presents an approach to domain-specific FAQ (frequently-asked question) retrieval using independent aspects. The data analysis classifies the questions in the collected QA (question-answer) pairs into ten question types in accordance with question stems. The answers in the QA pairs are then paragraphed and clustered using latent semantic analysis and the K-means algorithm. For semantic representation of the aspects, a domain-specific ontology is constructed based on WordNet and HowNet. A probabilistic mixture model is then used to interpret the query and QA pairs based on independent aspects; hence the retrieval process can be viewed as the maximum likelihood estimation problem. The expectation-maximization (EM) algorithm is employed to estimate the optimal mixing weights in the probabilistic mixture model. Experimental results indicate that the proposed approach outperformed the FAQ-Finder system in medical FAQ retrieval. Chung-Hsien Wu 0001, Jui-Feng Yeh, Ming-Jun Chen |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2005 | Speech act modeling and verification of spontaneous speech with disfluency in a spoken dialogue systemabstractThis work presents an approach to modeling speech acts and verifying spontaneous speech with disfluency in a spoken dialogue system. According to this approach, semantic information, syntactic structure and fragment class of an input utterance are statistically encapsulated in a proposed speech act hidden Markov model (SAHMM) to characterize the speech act. An interpolation mechanism is exploited to re-estimate the state transition probability in SAHMM, to deal with the problem of disfluency in a sparse training corpus. Finally, a Bayesian belief model (BBM), based on latent semantic analysis (LSA), is adopted to verify the potential speech acts and output the final speech act. Experiments were conducted to evaluate the proposed approach using a spoken dialogue system for providing air travel information. A testing database from 25 speakers, with 480 dialogues that include 3038 sentences, was established and used for evaluation. Experimental results show that the proposed approach identifies 95.3% of speech act at a rejection rate of 5%, and the semantic accuracy is 4.2% better than that obtained using a keyword-based system. The proposed strategy also effectively alleviates the disfluency problem in spontaneous speech. Chung-Hsien Wu 0001, Gwo-Lang Yan |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Automated Alignment and Extraction of Bilingual Domain Ontology for Cross-Language Domain-Specific Applications
Jui-Feng Yeh, Chung-Hsien Wu 0001, Ming-Jun Chen, Liang-Chih Yu |
COLING | 2 |
| 2004 | Language boundary detection and identification of mixed-language speech based on MAP estimationabstractThe paper proposes a maximum a posteriori (MAP) based approach to segment and identify jointly an utterance with mixed languages. A statistical framework for language boundary detection and language identification is proposed. First, the MAP estimation is used to determine the boundary number and positions. Further, an LSA-based GMM and a VQ-based bigram language model are proposed to characterize a language and used for language identification. Finally, a likelihood ratio test approach is used to determine the optimal number of language boundaries. Experimental results show that the proposed approach exhibits encouraging potential in mixed-language segmentation and identification. Chi-Jiun Shia, Yu-Hsien Chiu, Jia-Hsin Hsieh, Chung-Hsien Wu 0001 |
ICASSP (1) | 4 |
| 2004 | Emotion recognition using acoustic features and textual contentabstractThe paper presents an approach to emotion recognition from speech signals and textual content. In the analysis of speech signals, thirty-three acoustic features are extracted from the speech input. After principle component analysis (PCA), 14 principle components are selected for discriminative representation. In this representation, each principle component is the combination of the 33 original acoustic features and forms a feature subspace. Support vector machines (SVMs) are adopted to classify the emotional states. In text analysis, all emotional keywords and emotion modification words are manually defined. The emotion intensity levels of emotional keywords and emotion modification words are estimated from a collected emotion corpus. The final emotional state is determined based on the emotion outputs from the acoustic and textual approaches. The experimental result shows that the emotion recognition accuracy of the integrated system is better than each of the two individual approaches. Ze-Jing Chuang, Chung-Hsien Wu 0001 |
ICME | 2 |
| 2004 | Speech act identification using an ontology-based partial pattern treeabstractThis paper presents an ontology-based partial pattern tree to identify the speech act in a spoken dialogue system. This study first extracts the key words/concepts in an application domain using latent semantic analysis (LSA). A partial pattern tree is used to deal with the ill-formed sentence problem in a spoken dialogue system. Concept expansion based on domain ontology is adopted to improve system performance. For performance evaluation, a medical dialogue system with multiple services, including registration information, clinic information and FAQ information, is implemented. Four performance measures were separately used for evaluation. The speech act identification rate achieves 86.2%. A Task Success Rate of 77% is obtained. The contextual appropriateness of the system response is 78.5%. Finally, the correct rate for FAQ retrieval is 82% with an improvement of 15% in comparison with the keyword-based vector space model. The results show the proposed ontology-based partial pattern tree is effective for dialogue management. Chung-Hsien Wu 0001, Jui-Feng Yeh, Ming-Jun Chen |
INTERSPEECH | 1 |
| 2004 | Error-Tolerant Sign Retrieval Using Visual Features and Maximum A Posteriori EstimationabstractThis paper proposes an efficient error-tolerant approach to retrieving sign words from a Taiwanese Sign Language (TSL) database. This database is tagged with visual gesture features and organized as a multilist code tree. These features are defined in terms of the visual characteristics of sign gestures by which they are indexed for sign retrieval and displayed using an anthropomorphic interface. The maximum a posteriori estimation is exploited to retrieve the most likely sign word given the input feature sequence. An error-tolerant mechanism based on mutual information criterion is proposed to retrieve a sign word of interest efficiently and robustly. A user-friendly anthropomorphic interface is also developed to assist learning TSL. Several experiments were performed in an educational environment to investigate the system's retrieval accuracy. Our proposed approach outperformed a dynamic programming algorithm in its task and shows tolerance to user input errors. Chung-Hsien Wu 0001, Yu-Hsien Chiu, Kung-Wei Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Recovery from false rejection using statistical partial pattern trees for sentence verification
Chung-Hsien Wu 0001, Yeou-Jiunn Chen |
Speech Commun. | 1 |
| 2003 | Flexible speech act identification of spontaneous speech with disfluencyabstractThis paper describes an approach for flexible speech act identification of spontaneous speech with disfluency. In this approach, semantic information, syntactic structure, and fragment features of an input utterance are statistically encapsulated into a proposed sp eech act hidden Markov model (SAHMM) to characterize the speech act. To deal with the disfluency problem in a sparse training corpus, an interpolation mechanism is exploited to re-estimate the state transition probability in SAHMM. Finally, the dialog system accepts the speech act with best score and returns the corresponding response. Experiments were conducted to evaluate the proposed approach using a spoken dialogue system for the air travel information service. A testing database from 25 speakers containing 480 dialogues including 3038 sentences was collected and used for evaluation. Using the proposed approach, the experimental results show that the performance can achieve 90.3% in speech act correct rate (SACR) and 85.5% in fragment correct rate (FCR) for fluent speech and gains a significant improvement of 5.7% in SACR and 6.9% in FCR compared to the baseline system without considering filled pauses for disfluent speech. Chung-Hsien Wu 0001, Gwo-Lang Yan |
INTERSPEECH | 1 |
| 2002 | Perceptual speech modeling for noisy speech recognitionabstractThis paper proposes a perceptual modeling approach with a two-stage recognition to deal with the issues of recognition degradation in noisy environment. The auditory masking effect is used for speech enhancement and acoustic modeling in order to overcome the model inconsistencies between training speech and noisy input. In the two-stage recognition, the maximum a posteriori (MAP) based adaptation algorithm is used to incrementally adapt the noise model. In order to evaluate our proposed approach, a Mandarin keyword spotting system was constructed. The experimental results show our proposed method achieves a better recognition rate compared to the audible noise suppression (ANS) and parallel model combination (PMC) methods for both in 70km/hr (10.3dB) and 90km/hr (6.4dB) car environments. Chung-Hsien Wu 0001, Yu-Hsien Chiu, Huigan Lim |
ICASSP | 1 |
| 2002 | Parameter-based lip modeling for facial animation of general objectsabstractFor human-computer interaction, virtual characters become more and more important in many applications. However, the facial animation developed in the past years generally focused on the simulation of human face. For animation and interaction in virtual world the emphasis should be extended to general objects with human-like faces. We propose four parameter-based lip models that can apply the lip motion to general objects with lip-like meshes. We analyze the characteristics of lip motions and define four different lip models: inflexible-cylindrical model, flexible-cylindrical model, inflexible-square model, and flexible-square model. The parameters of lip motion in different lip models are extracted from the original real facial image sequence with lip motion. These parameters are then used to generate a lip-motion simulation system for general objects. Using the system, the user can apply the lip motion to any meshes by simply defining the lip model. Ze-Jing Chuang, Chung-Hsien Wu 0001 |
ICME (1) | 2 |
| 2002 | Emotion recognition from textual input using an emotional semantic network
Ze-Jing Chuang, Chung-Hsien Wu 0001 |
INTERSPEECH | 2 |
| 2002 | Sign language translation using an error tolerant retrieval algorithm
Chung-Hsien Wu 0001, Yu-Hsien Chiu, Kung-Wei Cheng |
INTERSPEECH | 1 |
| 2002 | Generation of robust phonetic set and decision tree for Mandarin using chi-square testing
Yeou-Jiunn Chen, Chung-Hsien Wu 0001, Yu-Hsien Chiu, Hsiang-Chuan Liao |
Speech Commun. | 2 |
| 2002 | Speech act modeling in a spoken dialog system using a fuzzy fragment-class Markov model
Chung-Hsien Wu 0001, Gwo-Lang Yan |
Speech Commun. | 1 |
| 2002 | Meaningful term extraction and discriminative term selection in text categorization via unknown-word methodologyabstractIn this article, an approach based on unknown words is proposed for meaningful term extraction and discriminative term selection in text categorization. For meaningful term extraction, a phrase-like unit (PLU)-based likelihood ratio is proposed to estimate the likelihood that a word sequence is an unknown word. On the other hand, a discriminative measure is proposed for term selection and is combined with the PLU-based likelihood ratio to determine the text category. We conducted several experiments on a news corpus, called MSDN. The MSDN corpus is collected from an online news Website maintained by the Min-Sheng Daily News, Taiwan. The corpus contains 44,675 articles with over 35 million words. The experimental results show that the system using a simple classifier achieved 95.31% accuracy. When using a state-of-the-art classifier, kNN, the average accuracy is 96.40%, outperforming all the other systems evaluated on the same collection, including the traditional term-word by kNN (88.52%); sleeping-experts (82.22%); sparse phrase by four-word sleeping-experts (86.34%); and Boolean combinations of words by RIPPER (87.54%). A proposed purification process can effectively reduce the dimensionality of the feature space from 50,576 terms in the word-based approach to 19,865 terms in the unknown word-based approach. In addition, more than 80% of automatically extracted terms are meaningful. Experiments also show that the proportion of meaningful terms extracted from training data is relative to the classification accuracy in outside testing. Yu-Sheng Lai, Chung-Hsien Wu 0001 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2001 | Discriminative disfluency modeling for spontaneous speech recognition
Chung-Hsien Wu 0001, Gwo-Lang Yan |
INTERSPEECH | 1 |
| 2001 | Multi-keyword spotting of telephone speech using a fuzzy search algorithm and keyword-driven two-level CBSM
Chung-Hsien Wu 0001, Yeou-Jiunn Chen |
Speech Commun. | 1 |
| 2001 | Automatic generation of synthesis units and prosodic information for Chinese concatenative synthesis
Chung-Hsien Wu 0001, Jau-Hung Chen |
Speech Commun. | 1 |
| 2000 | Intention extraction and semantic matching for internet FAQ retrieval using spoken language query
Yu-Sheng Lai, Kuen-Lin Lee, Chung-Hsien Wu 0001 |
INTERSPEECH | 3 |
| 2000 | Natural language processing for Taiwanese sign language to speech conversion
Chung-Hsien Wu 0001, Yu-Hsien Chiu, Chi-Shiang Guo |
INTERSPEECH | 1 |
| 2000 | Error recovery and sentence verification using statistical partial pattern tree for conversational speech
Chung-Hsien Wu 0001, Yeou-Jiunn Chen, Cher-Yao Yang |
INTERSPEECH | 1 |
| 1999 | Utterance verification using prosodic information for Mandarin telephone speech keyword spottingabstractIn this paper, the prosodic information, a very special and important feature in Mandarin speech, is used for Mandarin telephone speech utterance verification. A two-stage strategy, with recognition followed by verification, is adopted. For keyword recognition, 59 context-independent subsyllables, i.e., 22 INITIALs and 37 FINALs in Mandarin speech, and one background/silence model, are used as the basic recognition units. For utterance verification, 12 anti-subsyllable HMMs, 175 context-dependent prosodic HMMs, and five anti-prosodic HMMs, are constructed. A keyword verification function combining phonetic-phase and prosodic-phase verification is investigated. Using a test set of 2400 conversational speech utterances from 20 speakers (12 males and 8 females), at 8.5% false rejection, the proposed verification method resulted in 17.8% false alarm rate. Furthermore, this method was able to correctly reject 90.4% of nonkeywords. Comparison with a baseline system without prosodic-phase verification shows that the prosodic information can benefit the verification performance. Yeou-Jiunn Chen, Chung-Hsien Wu 0001, Gwo-Lang Yan |
ICASSP | 2 |
| 1999 | Speech act modeling in a spoken dialogue system using fuzzy hidden Markov model and bayes' decision criterion
Chung-Hsien Wu 0001, Gwo-Lang Yan |
EUROSPEECH | 1 |
| 1999 | Automatic Selection of Synthesis Units from a Large Speech Database
Jau-Hung Chen, Chung-Hsien Wu 0001 |
PACLIC | 2 |
| 1998 | Telephone speech multi-keyword spotting using fuzzy search algorithm and prosodic verificationabstractIn this paper a fuzzy search algorithm is proposed to deal with the recognition error for telephone speech. Since the prosodic information is a very special and important feature for Mandarin speech, we integrate the prosodic information into keyword verification. For multi-keyword detection, we define a keyword relation and a weighting function for reasonable keyword combinations. In the keyword recognizer, 94 INITIAL and 38 FINAL context-dependent Hidden Markov Models (HMM‘s) are used to construct the phonetic recognizer. For prosodic verification, a total of 175 context-dependent HMM’s and five anti-prosodic HMM‘s are used. In this system, 1275 faculty names and department names are selected as the keywords. Using a test set of 3595 conversional speech utterance from 37 speakers (21 male, 16 female), the proposed fuzzy search algorithm and prosodic verification can reduce the error rate from 17.64% to 11.29% for multiple keywords embedded in non-keyword speech. Chung-Hsien Wu 0001, Yeou-Jiunn Chen, Yu-Chun Hung |
ICSLP | 1 |
| 1998 | Spoken dialogue system using corpus-based hidden Markov modelabstractIn a spoken dialogue system, the intention is the most important component for speech understanding. In this paper, we propose a corpus-based hidden Markov model (HMM) to model the intention of a sentence. Each intention is represented by a sequence of word segment categories determined by a task-specific lexicon and a corpus. In the training procedure, five intention HMM’s are defined, each representing one intention in our approach. In the intention identification process, the phrase sequence is fed to each intention HMM. Given a speech utterance, the Viterbi algorithm is used to find the most likely intention sequences. The intention HMM considers not only the phrase frequency but also the syntactic and semantic structure in a phrase sequence. In order to evaluate the proposed method, a spoken dialogue model for air travel information service is investigated. The experiments were carried out using a test database from 25 speakers (15 male and 10 female). There are 120 dialogues, which contain 725 sentences in the test database. The experimental results show that the correct response rate can achieve about 80.3% using intention HMM. Chung-Hsien Wu 0001, Gwo-Lang Yan |
ICSLP | 1 |
| 1997 | A novel two-level method for the computation of the LSP frequencies using a decimation-in-degree algorithmabstractA novel two-level method is proposed for rapidly and accurately computing the line spectrum pair (LSP) frequencies. An efficient decimation-in-degree (DID) algorithm is also proposed in the first level, which can transform any symmetric or antisymmetric polynomial with real coefficients into the other polynomials with lower degrees and without any transcendental functions. The DID algorithm not only can avoid prior storage or large calculation of transcendental functions but can also be easily applied toward those fast root-finding methods. In the second level, if the transformed polynomial is of degree 4 or less, employing closed-form formulas is the fastest procedure of quite high accuracy. If it is of a higher degree, a modified Newton-Raphson method with cubic convergence is applied. Additionally, the process of the modified Newton-Raphson method can be accelerated by adopting a deflation scheme along with Descartes rule of signs and the interlacing property of LSP frequencies for selecting the better initial values. Besides this, Horner's method is extended to efficiently calculate the values of a polynomial and its first and second derivatives. A few conventional numerical methods are also implemented to make a comparison with the two-level method. Experimental results indicate that the two-level method is the fastest one. Furthermore, this method is more advantageous under the requirement of a high level of accuracy. Chung-Hsien Wu 0001, Jau-Hung Chen |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | A Mandarin Voice Organizer Based on a Template-Matching Speech Recognizer
Jhing-Fa Wang, Jyh-Shing Shyuu, Chung-Hsien Wu 0001 |
PACLIC | 3 |
| 1995 | Computer-aided analysis and classification of heart sounds based on neural networks and time analysisabstractThis paper describes a computer-aided heart sound analysis and classification system (CHACS) based on neural networks and time analysis. In this system, two subsystems in both time and frequency domains are proposed. In the first subsystem, a multilayer perceptron neural network is adopted to classify heart sound patterns. In the second subsystem, a set of heuristic rules is used to characterize heart sounds. The individual classification results of these two subsystems are combined to give the final suggestion. Using this system, heart sounds can be selectively stored, retrieved, enhanced, and replayed. Besides, the CHACS provides an online display of the heart beat rate and allows an objective and reliable classification of heart sounds. Experimental results show that a classification rate of 95.6% is obtained. Chung-Hsien Wu 0001, Ching-Wen Lo, Jhing-Fa Wang |
ICASSP | 1 |
| 1991 | Integrating neural nets and one-stage dynamic programming for speaker independent continuous Mandarin digit recognitionabstractA Bayesian neural network; a one-stage dynamic programming algorithm, and a Hopfield time-alignment network are integrated to form a speaker-independent continuous Mandarin digit recognizer. In this system, a Bayesian network trained with a splitting LVQ (learning vector quantisation) and the LVQ2 algorithms gives the a posteriori probability. The one-stage algorithm is then employed for coarse recognition. Finally, the Hopfield time-alignment network is used to eliminate unreasonable candidates. Experimental evaluation of this system, using 53 speakers (28 male, 25 female), each speaking 20 digit strings of varying length (1-7 digits/string) and at varying speaking rate (150-240 digits/min), gave an average recognition accuracy of 94.3%, with 1.4% insertion, 1.1% deletion, and 3.2% substitution errors.> Jhing-Fa Wang, Chung-Hsien Wu 0001, Chaug-Ching Haung, Jau-Yien Lee |
ICASSP | 2 |
| 1991 | Speaker-Independent Recognition of isolated Words using concatenated Neural NetworksabstractA speaker-independent isolated word recognizer is proposed. It is obtained by concatenating a Bayesian neural network and a Hopfield time-alignment network. In this system, the Bayesian network outputs the a posteriori probability for each speech frame, and the Hopfield network is then concatenated for time warping. A proposed splitting Learning Vector Quantization (LVQ) algorithm derived from the LBG clustering algorithm and the Kohonen LVQ algorithm is first used to train the Bayesian network. The LVQ2 algorithm is subsequently adopted as a final refinement step. A continuous mixture of Gaussian densities for each frame and multi-templates for each word are employed to characterize each word pattern. Experimental evaluation of this system with four templates/word and five mixtures/frame, using 53 speakers (28 males, 25 females) and isolated words (10 digits and 30 city names) databases, gave average recognition accuracies of 97.3%, for the speaker-trained mode and 95.7% for the speaker-independent mode, respectively. Comparisons with K-means and DTW algorithms show that the integration of the splitting LVQ and LVQ2 algorithms makes this system well suited to speaker-independent isolated word recognition. A cookbook approach for the determination of parameters in the Hopfield time-alignment network is also described. Chung-Hsien Wu 0001, Jhing-Fa Wang, Chaug-Ching Huang, Jau-Yien Lee |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1991 | A shunting multilayer perceptron network for confusing/composite pattern recognition
Chung-Hsien Wu 0001, Jhing-Fa Wang, Wen-Horng Wu |
Pattern Recognit. | 1 |