EDBT 2026 Demo / reviewers in the wild / expert
Tan Lee
dblp:57/5024
· DBLP profile ↗
156ranked-venue papers
14as first author
40since 2021 · last 2024
0000-0002-7089-3436ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 136 · 10 first-author · 36 since 2021Artificial intelligence and machine learning · 107 · 10 first-author · 30 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Efficient Black-Box Speaker Verification Model Adaptation With Reprogramming And Backend LearningabstractThe development of deep neural networks (DNN) has significantly enhanced the performance of speaker verification (SV) systems in recent years. However, a critical issue that persists when applying DNN-based SV systems in practical applications is domain mismatch. To mitigate the performance degradation caused by the mismatch, domain adaptation becomes necessary. This paper introduces an approach to adapt DNN-based SV models by manipulating the learnable model inputs, inspired by the concept of adversarial reprogramming. The pre-trained SV model remains fixed and functions solely in the forward process, resembling a black-box model. A lightweight network is utilized to estimate the gradients for the learnable parameters at the input, which bypasses the gradient back-propagation through the black-box model. The reprogrammed output is processed by a two-layer backend learning module as the final adapted speaker embedding. The number of parameters involved in the gradient calculation is small in our design. With few additional parameters, the proposed method achieves both memory and parameter efficiency. The experiments are conducted in language mismatch scenarios. Using much less computation cost, the proposed method obtains close or superior performance to the fully finetuned models in our experiments, which demonstrates its effectiveness. Tan Lee |
ICASSP | 2 |
| 2024 | Sparsely Shared Lora on Whisper for Child Speech RecognitionabstractWhisper is a powerful automatic speech recognition (ASR) model. Nevertheless, its zero-shot performance on low-resource speech requires further improvement. Child speech, as a representative type of low-resource speech, is leveraged for adaptation. Recently, parameter-efficient fine-tuning (PEFT) in NLP was shown to be comparable and even better than full fine-tuning, while only needing to tune a small set of trainable parameters. However, current PEFT methods have not been well examined for their effectiveness on Whisper. In this paper, only parameter composition types of PEFT approaches such as LoRA and Bitfit are investigated as they do not bring extra inference costs. Different popular PEFT methods are examined. Particularly, we compare LoRA and AdaLoRA and figure out the learnable rank coefficient is a good design. Inspired by the sparse rank distribution allocated by AdaLoRA, a novel PEFT approach Sparsely Shared LoRA (S2-LoRA) is proposed. The two low-rank decomposed matrices are globally shared. Each weight matrix only has to maintain its specific rank coefficients that are constrained to be sparse. Experiments on low-resource Chinese child speech show that with much fewer trainable parameters, S2-LoRA can achieve comparable in-domain adaptation performance to AdaLoRA and exhibit better generalization ability on out-of-domain data. In addition, the rank distribution automatically learned by S2-LoRA is found to have similar patterns to AdaLoRA’s allocation. Wei Liu 0147, Tan Lee |
ICASSP | 4 |
| 2024 | Modeling Intrapersonal and Interpersonal Influences for Automatic Estimation of Therapist Empathy in Counseling ConversationabstractCounseling is usually conducted through spoken conversation between a therapist and a client. The empathy level of therapist is a key indicator of outcomes. Presuming that therapist’s empathy expression is shaped by their past behavior and their perception of the client’s behavior, we propose a model to estimate the therapist empathy by considering both intrapersonal and interpersonal influences. These dynamic influences are captured by applying an attention mechanism to the therapist turn and the historical turns of both therapist and client. Our findings suggest that the integration of dynamic influences enhances empathy level estimation. The influence-derived embedding should constitute a minor portion of the target turn representation for optimal empathy estimation. The client’s turns (interpersonal influence) slightly surpass the therapist’s own turns (intrapersonal influence) in empathy estimation effectiveness. It is noted that concentrating exclusively on recent historical turns can significantly impact the estimation of therapist empathy. Dehua Tao, Tan Lee, Harold Chui, Sarah Luk |
ICASSP | 2 |
| 2024 | Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction LossabstractThis research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximally preserving the target speaker’s individuality in terms of both the voice timbre and speaking style. This is formulated as a task of learning from noisy labels. We propose to augment the commonly used speech reconstruction loss with two additional terms. The first term constitutes a regularization loss that mitigates the impact of distorted articulation in the training speech. The second term is a consistency loss that encourages correct articulation in the generated speech. These additional loss terms are obtained from frame-level articulation scores of original and generated speech, which are derived using a separately trained phone classifier. Experimental results on a real case of tongue cancer patient confirm that the synthetic voice achieves comparable articulation quality to unimpaired natural speech, while effectively maintaining the target speaker’s individuality. Audio samples are available at https://myspeechproject.github.io/ArticulationRepair/. Yusheng Tian, Tan Lee |
ICASSP | 3 |
| 2024 | A Parameter-efficient Language Extension Framework for Multilingual ASR
Wei Liu 0147, Jingyong Hou, Muyong Cao, Tan Lee |
INTERSPEECH | 5 |
| 2024 | LUPET: Incorporating Hierarchical Information Path into Multilingual ASRabstractToward high-performance multilingual automatic speech recognition (ASR), various types of linguistic information and model design have demonstrated their effectiveness independently.They include language identity (LID), phoneme information, language-specific processing modules, and crosslingual self-supervised speech representation.It is expected that leveraging their benefits synergistically in a unified solution would further improve the overall system performance.This paper presents a novel design of a hierarchical information path, named LUPET, which sequentially encodes, from the shallow layers to deep layers, multiple aspects of linguistic and acoustic information at diverse granularity scales.The path starts from LID prediction, followed by acoustic unit discovery, phoneme sharing, and finally token recognition routed by a mixture-ofexpert.ASR experiments are carried out on 10 languages in the Common Voice corpus.The results demonstrate the superior performance of LUPET as compared to the baseline systems.Most importantly, LUPET effectively mitigates the issue of performance compromise of high-resource languages with low-resource ones in the multilingual setting. Wei Liu 0147, Jingyong Hou, Muyong Cao, Tan Lee |
INTERSPEECH | 5 |
| 2024 | Learning Representation of Therapist Empathy in Counseling Conversation Using Siamese Hierarchical Attention NetworkabstractCounseling is an activity of conversational speaking between a therapist and a client.Therapist empathy is an essential indicator of counseling quality and assessed subjectively by considering the entire conversation.This paper proposes to encode long counseling conversation using a hierarchical attention network.Conversations with extreme values of empathy rating are used to train a Siamese network based encoder with contrastive loss.Two-level attention mechanisms are applied to learn the importance weights of individual speaker turns and groups of turns in the conversation.Experimental results show that the use of contrastive loss is effective in encouraging the conversation encoder to learn discriminative embeddings that are related to therapist empathy.The distances between conversation embeddings positively correlate with the differences in the respective empathy scores.The learned conversation embeddings can be used to predict the subjective rating of therapist empathy. Dehua Tao, Tan Lee, Harold Chui, Sarah Luk |
INTERSPEECH | 2 |
| 2024 | Contrastive Context-Speech Pretraining for Expressive Text-to-Speech SynthesisabstractThe latest Text-to-Speech (TTS) systems can produce speech with voice quality and naturalness comparable to human speech. Yet the demand for large amount of high-quality data from target speakers remains a significant challenge. Particularly for long-form expressive reading, target speaker's training speech that covers rich contextual information are needed. In this paper a novel design of context-aware speech pre-trained model is developed for expressive TTS based on contrastive learning. The model can be trained with abundant speech data without explicitly labelled speaker identities. It captures the intricate relationship between the speech expression of a spoken sentence and the contextual text information. By incorporating cross-modal text and speech features into the TTS model, it enables the generation of coherent and expressive speech, which is especially beneficial when there is a scarcity of target speaker data. The pre-trained model is evaluated first in the task of Context-Speech retrieval and then as the integral part of a zero-shot TTS system. Experimental results demonstrate that the pretraining framework effectively learns Context-Speech representations and significantly enhances the expressiveness of synthesized speech. Audio demos are available at: https://ccsp2024.github.io/demo/. Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Xinfa Zhu, Sheng Zhao 0002, Tan Lee |
ACM Multimedia | 7 |
| 2024 | Automatic Detection of Speech Sound Disorder in Cantonese-Speaking Pre-School ChildrenabstractSpeech sound disorder (SSD) is a type of developmental disorder in which children encounter persistent difficulties in correctly producing certain speech sounds. Conventionally, assessment of SSD relies largely on speech and language pathologists (SLPs) with appropriate language background. With the unsatisfied demand for qualified SLPs, automatic detection of SSD is highly desirable for assisting clinical work and improving the efficiency and quality of services. In this paper, methods and systems for fully automatic detection of SSD in young children are investigated. A microscopic approach and a macroscopic approach are developed. The microscopic system is based on detection of phonological errors in impaired child speech. A deep neural network (DNN) model is trained to learn the similarity and contrast between consonant segments. Phonological error is identified by contrasting a test speech segment to reference segments. The phone-level similarity scores are aggregated for speaker-level SSD detection. The macroscopic approach leverages holistic changes of speech characteristics related to disorders. Various types of speaker-level embeddings are investigated and compared. Experimental results show that the proposed microscopic system achieves unweighted average recall (UAR) from 84.0% to 91.9% on phone-level error detection. The proposed macroscopic approach can achieve a UAR of 89.0% on speaker-level SSD detection. The speaker embeddings adopted for macroscopic SSD detection can effectively discard the information related to speaker's personal identity. Si Ioi Ng, Cymie Wing-Yee Ng, Jiarui Wang 0003, Tan Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found DataabstractCreating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model and then use the enhanced data for text-to-speech (TTS) model training. This paper investigates the use of conditional diffusion models for generalized speech enhancement, which aims at addressing multiple types of audio degradation simultaneously. The enhancement is performed on the log Mel-spectrogram domain to align with the TTS training objective. Text information is introduced as an additional condition to improve the model robustness. Experiments on real-world recordings demonstrate that the synthetic voice built on data enhanced by the proposed model produces higher-quality synthetic speech, compared to those trained on data enhanced by strong baselines. Code and check-points of the proposed enhancement model are available at https://github.com/dmse4tts/DMSE4TTS. Yusheng Tian, Wei Liu 0147, Tan Lee |
ASRU | 3 |
| 2023 | Convolution-Based Channel-Frequency Attention for Text-Independent Speaker VerificationabstractDeep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolution-based attention module, namely C2D-Att. The interaction between the convolution channel and frequency is involved in the attention calculation by lightweight convolution layers. This requires only a small number of parameters. Fine-grained attention weights are produced to represent channel and frequency-specific information. The weights are imposed on the input features to improve the representation ability for speaker modeling. The C2D-Att is integrated into a modified version of ResNet for speaker embedding extraction. Experiments are conducted on VoxCeleb datasets. The results show that C2D-Att is effective in generating discriminative attention maps and outperforms other attention methods. The proposed model shows robust performance with different scales of model size and achieves state-of-the-art results. Yusheng Tian, Tan Lee |
ICASSP | 3 |
| 2023 | Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation ScoringabstractRecent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation quality representation. In this paper, we propose to use linguistic-acoustic similarity to explicitly measure the deviation of non-native production from its native reference for pronunciation assessment. Specifically, the deviation is first estimated by the cosine similarity between reference phone embedding and corresponding acoustic embedding. Next, a phone-level Goodness of pronunciation (GOP) pre-training stage is introduced to guide this similarity-based learning for better initialization of the aforementioned two embeddings. Finally, a transformer-based hierarchical pronunciation scorer is used to map a sequence of phone embeddings, acoustic embeddings along with their similarity measures to predict the final utterance-level score. Experimental results on the non-native databases suggest that the proposed system significantly outperforms the baselines, where the acoustic and phone embeddings are simply added or concatenated. A further examination shows that the phone embeddings learned in the proposed approach are able to capture linguistic-acoustic attributes of native pronunciation as references. Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee |
ICASSP | 7 |
| 2023 | An ASR-Free Fluency Scoring Approach with Self-Supervised LearningabstractA typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free approach for automatic fluency assessment using self-supervised learning (SSL). Specifically, wav2vec2.0 is used to extract frame-level speech features, followed by K-means clustering to assign a pseudo label (cluster index) to each frame. A BLSTM-based model is trained to predict an utterance-level fluency score from frame-level SSL features and the corresponding cluster indexes. Neither speech transcription nor time stamp information is required in the proposed system. It is ASR-free and can potentially avoid the ASR errors effect in practice. Experimental results carried out on non-native English databases show that the proposed approach significantly improves the performance in the "open response" scenario as compared to previous methods and matches the recently reported performance in the "read aloud" scenario. Wei Liu 0147, Kaiqi Fu, Xiaohai Tian, Shuju Shi, Wei Li 0119, Zejun Ma 0001, Tan Lee |
ICASSP | 7 |
| 2023 | Covariance Regularization for Probabilistic Linear Discriminant AnalysisabstractProbabilistic linear discriminant analysis (PLDA) is commonly used in speaker verification systems to score the similarity of speaker embeddings. Recent studies improved the performance of PLDA in domain-matched conditions by diagonalizing its covariance. We suspect such a brutal pruning approach could eliminate its capacity in modeling dimension correlation of speaker embeddings, leading to inadequate performance with domain adaptation. This paper explores two alternative covariance regularization approaches, namely, interpolated PLDA and sparse PLDA, to tackle the problem. The interpolated PLDA incorporates the prior knowledge from cosine scoring to interpolate the covariance of PLDA. The sparse PLDA introduces a sparsity penalty to update the covariance. Experimental results demonstrate that both approaches outperform diagonal regularization noticeably with domain adaptation. In addition, in-domain data can be significantly reduced when training sparse PLDA for domain adaptation. Mingjie Shao, Xuanji He, Xu Li 0015, Tan Lee, Guanglu Wan |
ICASSP | 5 |
| 2023 | Model Compression for DNN-based Speaker Verification Using Weight Quantization
Wei Liu 0147, Zhaoyang Zhang 0001, Tan Lee |
INTERSPEECH | 5 |
| 2023 | CoMFLP: Correlation Measure Based Fast Search on ASR Layer PruningabstractTransformer-based speech recognition (ASR) model with deep layers exhibited significant performance improvement.However, the model is inefficient for deployment on resourceconstrained devices.Layer pruning (LP) is a commonly used compression method to remove redundant layers.Previous studies on LP usually identify the redundant layers according to a task-specific evaluation metric.They are time-consuming for models with a large number of layers, even in a greedy search manner.To address this problem, we propose CoM-FLP, a fast search LP algorithm based on correlation measure.The correlation between layers is computed to generate a correlation matrix, which identifies the redundancy among layers.The search process is carried out in two steps: (1) coarse search: to determine top K candidates by pruning the most redundant layers based on the correlation matrix; (2) fine search: to select the best pruning proposal among K candidates using a task-specific evaluation metric.Experiments on an ASR task show that the pruning proposal determined by CoMFLP outperforms existing LP methods while only requiring constant time complexity.The code is publicly available at https://github.com/louislau1129/CoMFLP. Wei Liu 0147, Tan Lee |
INTERSPEECH | 3 |
| 2023 | A Study on Using Duration and Formant Features in Automatic Detection of Speech Sound Disorder in Childrenabstract24th Annual Conference of the International Speech Communication Association, INTERSPEECH 2023, Dublin, Ireland, August 20-24, 2023 Si Ioi Ng, Cymie Wing-Yee Ng, Tan Lee |
INTERSPEECH | 3 |
| 2023 | A Study on Prosodic Entrainment in Relation to Therapist Empathy in Counseling Conversation
Dehua Tao, Tan Lee, Harold Chui, Sarah Luk |
INTERSPEECH | 2 |
| 2023 | Creating Personalized Synthetic Voices from Post-Glossectomy Speech with Guided Diffusion Models
Yusheng Tian, Guangyan Zhang, Tan Lee |
INTERSPEECH | 3 |
| 2023 | ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph ReadingabstractWhile state-of-the-art Text-to-Speech systems can generate natural speech of very high quality at sentence level, they still meet great challenges in speech generation for paragraph / long-form reading.Such deficiencies are due to i) ignorance of cross-sentence contextual information, and ii) high computation and memory cost for long-form synthesis.To address these issues, this work develops a lightweight yet effective TTS system, ContextSpeech.Specifically, we first design a memory-cached recurrence mechanism to incorporate global text and speech context into sentence encoding.Then we construct hierarchically-structured textual semantics to broaden the scope for global context enhancement.Additionally, we integrate linearized self-attention to improve model efficiency.Experiments show that ContextSpeech significantly improves the voice quality and prosody expressiveness in paragraph reading with competitive model efficiency. Yujia Xiao, Shaofei Zhang, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Sheng Zhao 0002, Frank K. Soong, Tan Lee |
INTERSPEECH | 8 |
| 2023 | iEmoTTS: Toward Robust Cross-Speaker Emotion Transfer and Control for Speech Synthesis Based on Disentanglement Between Prosody and TimbreabstractThe capability of generating speech with a specific type of emotion is desired for many human-computer interaction applications. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech data with emotion labels from target speakers is not available for model training. This paper presents a novel cross-speaker emotion transfer system named iEmoTTS. The system is composed of an emotion encoder, a prosody predictor, and a timbre encoder. The emotion encoder extracts the identity of emotion type and the respective emotion intensity from the mel-spectrogram of input speech. The emotion intensity is measured by the posterior probability that the input utterance carries that emotion. The prosody predictor is used to provide prosodic features for emotion transfer. The timbre encoder provides timbre-related information for the system. Unlike many other studies which focus on disentangling speaker and style factors of speech, the iEmoTTS is designed to achieve cross-speaker emotion transfer via disentanglement between prosody and timbre. Prosody is considered the primary carrier of emotion-related speech characteristics, and timbre accounts for the essential characteristics for speaker identification. Zero-shot emotion transfer, meaning that the speech of target speakers is not seen in model training, is also realized with iEmoTTS. Extensive experiments of subjective evaluation have been carried out. The results demonstrate the effectiveness of iEmoTTS compared with other recently proposed systems of cross-speaker emotion transfer. It is shown that iEmoTTS can produce speech with designated emotion types and controllable emotion intensity. With appropriate information bottleneck capacity, iEmoTTS is able to transfer emotional information to a new speaker effectively. Audio samples are publicly available. Guangyan Zhang, Jialun Wu, Yutao Gai, Feijun Jiang, Tan Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2022 | A Study on the Efficacy of Model Pre-Training In Developing Neural Text-to-Speech SystemabstractIn the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers’ data is a common approach. However, in terms of ultimately achieved system performance for target speaker(s), the actual benefits of model pre-training are uncertain and unstable, depending very much on the quantity and text content of training data. This study aims to understand better why and how model pre-training can positively contribute to TTS system performance. It is postulated that the pre-training process plays a critical role in learning text-related variation in speech, while further training with the target speaker’s data aims to capture the speaker-related variation. Different test sets are created with varying degrees of similarity to target speaker data in terms of text content. Experiments show that leveraging a speaker-independent TTS trained on speech data with diverse text content can improve the target speaker TTS on domain-mismatched text. We also attempt to reduce the amount of pre-training data for a new text domain and improve the data and computational efficiency. It is found that the TTS system could achieve comparable performance when the pre-training data is reduced to 1/8 of its original size. Guangyan Zhang, Yichong Leng, Daxin Tan, Kaitao Song, Xu Tan 0003, Sheng Zhao 0002, Tan Lee |
ICASSP | 8 |
| 2022 | Durational Patterning at Discourse Boundaries in Relation to Therapist Empathy in Psychotherapy
Jonathan Him Nok Lee, Dehua Tao, Harold Chui, Tan Lee, Sarah Luk, Nicolette Wing Tung Lee, Koonkan Fung |
INTERSPEECH | 4 |
| 2022 | EDITnet: A Lightweight Network for Unsupervised Domain Adaptation in Speaker VerificationabstractPerformance degradation caused by language mismatch is a common problem when applying a speaker verification system on speech data in different languages.This paper proposes a domain transfer network, named EDITnet, to alleviate the language-mismatch problem on speaker embeddings without requiring speaker labels.The network leverages a conditional variational auto-encoder to transfer embeddings from the target domain into the source domain.A self-supervised learning strategy is imposed on the transferred embeddings so as to increase the cosine distance between embeddings from different speakers.In the training process of the EDITnet, the embedding extraction model is fixed without fine-tuning, which renders the training efficient and low-cost.Experiments on Voxceleb and CN-Celeb show that the embeddings transferred by ED-ITnet outperform the un-transferred ones by around 30% with the ECAPA-TDNN512.Performance improvement can also be achieved with other embedding extraction models, e.g., TDNN, SE-ResNet34. Wei Liu 0147, Tan Lee |
INTERSPEECH | 3 |
| 2022 | Automatic Detection of Speech Sound Disorder in Child Speech Using Posterior-based Speaker Representationsabstract23rd Annual Conference of the International Speech Communication Association, INTERSPEECH 2022, Incheon, Korea, September 18-22, 2022 Si Ioi Ng, Cymie Wing-Yee Ng, Jiarui Wang 0003, Tan Lee |
INTERSPEECH | 4 |
| 2022 | Unifying Cosine and PLDA Back-ends for Speaker VerificationabstractState-of-art speaker verification (SV) systems use a backend model to score the similarity of speaker embeddings extracted from a neural network model.The commonly used back-end models are the cosine scoring and the probabilistic linear discriminant analysis (PLDA) scoring.With the recently developed neural embeddings, the theoretically more appealing PLDA approach is found to have no advantage against or even be inferior the simple cosine scoring in terms of SV system performance.This paper presents an investigation on the relation between the two scoring approaches, aiming to explain the above counter-intuitive observation.It is shown that the cosine scoring is essentially a special case of PLDA scoring.In other words, by properly setting the parameters of PLDA, the two back-ends become equivalent.As a consequence, the cosine scoring not only inherits the basic assumptions for the PLDA but also introduces additional assumptions on the properties of input embeddings.Experiments show that the dimensional independence assumption required by the cosine scoring contributes most to the performance gap between the two methods under the domain-matched condition.When there is severe domain mismatch and the dimensional independence assumption does not hold, the PLDA would perform better than the cosine for domain adaptation. Xuanji He, Tan Lee, Guanglu Wan |
INTERSPEECH | 4 |
| 2022 | Environment Aware Text-to-Speech SynthesisabstractThis study aims at designing an environment-aware text-tospeech (TTS) system that can generate speech to suit specific acoustic environments.It is also motivated by the desire to leverage massive data of speech audio from heterogeneous sources in TTS system development.The key idea is to model the acoustic environment in speech audio as a factor of data variability and incorporate it as a condition in the process of neural network based speech synthesis.Two embedding extractors are trained with two purposely constructed datasets for characterization and disentanglement of speaker and environment factors in speech.A neural network model is trained to generate speech from extracted speaker and environment embeddings.Objective and subjective evaluation results demonstrate that the proposed TTS system is able to effectively disentangle speaker and environment factors and synthesize speech audio that carries designated speaker characteristics and environment attribute.Audio samples are available online for demonstration 1 . Daxin Tan, Guangyan Zhang, Tan Lee |
INTERSPEECH | 3 |
| 2022 | Characterizing Therapist's Speaking Style in Relation to Empathy in PsychotherapyabstractIn conversation-based psychotherapy, therapists use verbal techniques to help clients express thoughts and feelings, and change behavior.In particular, how well therapists convey empathy is an essential quality index of psychotherapy sessions and is associated with psychotherapy outcome.In this paper, we analyze the prosody of therapist speech and attempt to associate the therapist's speaking style with subjectively perceived empathy.An automatic speech and text processing system is developed to segment long recordings of psychotherapy sessions into pause-delimited utterances with text transcriptions.Data-driven clustering is applied to the utterances from different therapists in multiple sessions.For each cluster, a typological representation of utterance genre is derived based on quantized prosodic feature parameters.Prominent speaking styles of the therapist can be observed and interpreted from salient utterance genres that are correlated with empathy.Using the salient utterance genres, an accuracy of 71% is achieved in classifying psychotherapy sessions into "high" and "low" empathy level.Analysis of results suggests that empathy level tends to be (1) low if therapists speak long utterances slowly or speak short utterances quickly; and (2) high if therapists talk to clients with a steady tone and volume. Dehua Tao, Tan Lee, Harold Chui, Sarah Luk |
INTERSPEECH | 2 |
| 2022 | Hierarchical Attention Network for Evaluating Therapist Empathy in Counseling SessionabstractCounseling typically takes the form of spoken conversation between a therapist and a client.The empathy level expressed by the therapist is considered to be an essential quality factor of counseling outcome.This paper proposes a hierarchical recurrent network combined with two-level attention mechanisms to determine the therapist's empathy level solely from the acoustic features of conversational speech in a counseling session.The experimental results show that the proposed model can achieve an accuracy of 72.1% in classifying the therapist's empathy level as being "high" or "low".It is found that the speech from both the therapist and the client are contributing to predicting the empathy level that is subjectively rated by an expert observer.By analyzing speaker turns assigned with high attention weights, it is observed that 2 to 6 consecutive turns should be considered together to provide useful clues for detecting empathy, and the observer tends to take the whole session into consideration when rating the therapist empathy, instead of relying on a few specific speaker turns. Dehua Tao, Tan Lee, Harold Chui, Sarah Luk |
INTERSPEECH | 2 |
| 2022 | Transport-Oriented Feature Aggregation for Speaker Embedding LearningabstractPooling is needed to aggregate frame-level features into utterance-level representations for speaker modeling.Given the success of statistics-based pooling methods, we hypothesize that speaker characteristics are well represented in the statistical distribution over the pre-aggregation layer's output, and propose to use transport-oriented feature aggregation for deriving speaker embeddings.The aggregated representation encodes the geometric structure of the underlying feature distribution, which is expected to contain valuable speaker-specific information that may not be represented by the commonly used statistical measures like mean and variance.The original transportoriented feature aggregation is also extended to a weightedframe version to incorporate the attention mechanism.Experiments on speaker verification with the Voxceleb dataset show improvement over statistics pooling and its attentive variant. Yusheng Tian, Tan Lee |
INTERSPEECH | 3 |
| 2022 | Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
Guangyan Zhang, Kaitao Song, Xu Tan 0003, Daxin Tan, Yuzi Yan, Gang Wang 0001, Tao Qin 0001, Tan Lee, Sheng Zhao 0002 |
INTERSPEECH | 10 |
| 2022 | Enhancing Segment-Based Speech Emotion Recognition by Iterative Self-LearningabstractDespite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training. Recently, segment-based approaches for SER have been evolving, which train backbone networks on shorter segments instead of whole utterances, and thus naturally augments training examples without additional resources. However, one core challenge remains for segment-based approaches: most emotional corpora do not provide ground-truth labels at the segment level. To supervisely train a segment-based emotion model on such datasets, the most common way assigns each segment the corresponding utterance’s emotion label. However, this practice typically introduces noisy (incorrect) labels as emotional information is not uniformly distributed across the whole utterance. On the other hand, DNNs have been shown to easily over-fit a dataset when being trained with noisy labels. To this end, this work proposes a simple and effective iterative self-learning (ISL) framework, which comprises a procedure to progressively correct segment-level labels in an iterative learning manner. The ISL method produces dynamically-generated and soft emotion labels, leading to significant performance improvements. Experiments on three well-known emotional corpora demonstrate noticeable gains using the proposed method. Shuiyang Mao, Pak-Chung Ching, Tan Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Improving Text-Independent Speaker Verification with Auxiliary Speakers Using GraphabstractThe paper presents a novel approach to refining similarity scores between input utterances for robust speaker verification. Given the embeddings from a pair of input utterances, a graph model is designed to incorporate additional information from a group of embeddings representing the so-called auxiliary speakers. The relations between the input utterances and the auxiliary speakers are represented by the edges and vertices in the graph. The similarity scores are refined by iteratively updating the values of the graph's vertices using an algorithm similar to the random walk algorithm on graphs. Through this updating process, the information of auxiliary speakers is involved in determining the relation between input utterances and hence contributing to the verification process. We propose to create a set of artificial embeddings through the model training process. Utilizing the generated embeddings as auxiliary speakers, no extra data are required for the graph model in the verification stage. The proposed model is trained in an end-to-end manner within the whole system. Experiments are carried out with the Voxceleb datasets. The results indicate that involving auxiliary speakers with graph is effective to improve speaker verification performance. Si Ioi Ng, Tan Lee |
ASRU | 3 |
| 2021 | Utterance-Level Neural Confidence Measure for End-to-End Children Speech RecognitionabstractConfidence measure is a performance index of particular importance for automatic speech recognition (ASR) systems deployed in real-world scenarios. In the present study, utterance-level neural confidence measure (NCM) in end-to-end automatic speech recognition (E2E ASR) is investigated. The E2E system adopts the joint CTC-attention Transformer architecture. The prediction of NCM is formulated as a task of binary classification, i.e., accept/reject the input utterance, based on a set of predictor features acquired during the ASR decoding process. The investigation is focused on evaluating and comparing the efficacies of predictor features that are derived from different internal and external modules of the E2E system. Experiments are carried out on children speech, for which state-of-the-art ASR systems show less than satisfactory performance and robust confidence measure is particularly useful. It is noted that predictor features related to acoustic information of speech play a more important role in estimating confidence measure than those related to linguistic information. N-best score features show significantly better performance than single-best ones. It has also been shown that the metrics of EER and AUC are not appropriate to evaluate the NCM of a mismatched ASR with significant performance gap. Wei Liu 0147, Tan Lee |
ASRU | 2 |
| 2021 | EditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional FusionabstractThis paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without causing audible degradation in speech quality and naturalness. The EditSpeech system is developed upon a neural text-to-speech (NTTS) synthesis framework. Partial inference and bidirectional fusion are proposed to effectively incorporate the contextual information related to the edited region and achieve smooth transition at both left and right boundaries. Distortion introduced to the unmodified parts of the utterance is alleviated. The EditSpeech system is developed and evaluated on English and Chinese in multi-speaker scenarios. Objective and subjective evaluation demonstrate that EditSpeech outperforms a few baseline systems in terms of low spectral distortion and preferred speech quality. Audio samples are available online for demonstration11https://daxintan-cuhk.github.io/EditSpeech/. Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang 0002, Xiao Chen 0012, Tan Lee |
ASRU | 6 |
| 2021 | Detection of Consonant Errors in Disordered Speech Based on Consonant-Vowel Segment EmbeddingabstractSpeech sound disorder (SSD) refers to a type of developmental disorder in young children who encounter persistent difficulties in producing certain speech sounds at the expected age.Consonant errors are the major indicator of SSD in clinical assessment.Previous studies on automatic assessment of SSD revealed that detection of speech errors concerning short and transitory consonants is less satisfactory.This paper investigates a neural network based approach to detecting consonant errors in disordered speech using consonant-vowel (CV) diphone segment in comparison to using consonant monophone segment.The underlying assumption is that the vowel part of a CV segment carries important information of co-articulation from the consonant.Speech embeddings are extracted from CV segments by a recurrent neural network model.The similarity scores between the embeddings of the test segment and the reference segments are computed to determine if the test segment is the expected consonant or not.Experimental results show that using CV segments achieves improved performance on detecting speech errors concerning those "difficult" consonants reported in the previous studies. Si Ioi Ng, Cymie Wing-Yee Ng, Tan Lee |
Interspeech | 4 |
| 2021 | Pairing Weak with Strong: Twin Models for Defending Against Adversarial Attack on Speaker Verification
Xu Li 0015, Tan Lee |
Interspeech | 3 |
| 2021 | Fine-Grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style DisentanglementabstractThis paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style embeddings from the mel-spectrograms of phone-level speech segments. Collaborative learning and adversarial learning strategies are applied in order to achieve effective disentanglement of content and style factors in speech and alleviate the "content leakage" problem in style modeling. The proposed system can be used for varying-content speech style transfer in the single-speaker scenario. The results of objective and subjective evaluation show that our system performs better than other fine-grained speech style transfer models, especially in the aspect of content preservation. By incorporating a style predictor, the proposed system can also be used for text-to-speech synthesis. Audio samples are provided for system demonstration https://daxintan-cuhk.github.io/pl-csd-speech . Daxin Tan, Tan Lee |
Interspeech | 2 |
| 2021 | Applying the Information Bottleneck Principle to Prosodic Representation LearningabstractThis paper describes a novel design of a neural network-based speech generation model for learning prosodic representation.The problem of representation learning is formulated according to the information bottleneck (IB) principle.A modified VQ-VAE quantized layer is incorporated in the speech generation model to control the IB capacity and adjust the balance between reconstruction power and disentangle capability of the learned representation.The proposed model is able to learn word-level prosodic representations from speech data.With an optimized IB capacity, the learned representations not only are adequate to reconstruct the original speech but also can be used to transfer the prosody onto different textual content.Extensive results of the objective and subjective evaluation are presented to demonstrate the effect of IB capacity control, the effectiveness, and potential usage of the learned prosodic representation in controllable neural speech generation. Guangyan Zhang, Daxin Tan, Tan Lee |
Interspeech | 4 |
| 2021 | Bayesian Learning for Deep Neural Network AdaptationabstractA key task for speech recognition systems is to reduce the mismatch between training and evaluation data that is often attributable to speaker differences. Speaker adaptation techniques play a vital role to reduce the mismatch. Model-based speaker adaptation approaches often require sufficient amounts of target speaker data to ensure robustness. When the amount of speaker level data is limited, speaker adaptation is prone to overfitting and poor generalization. To address the issue, this paper proposes a full Bayesian learning based DNN speaker adaptation framework to model speaker-dependent (SD) parameter uncertainty given limited speaker specific adaptation data. This framework is investigated in three forms of model based DNN adaptation techniques: Bayesian learning of hidden unit contributions (BLHUC), Bayesian parameterized activation functions (BPAct), and Bayesian hidden unit bias vectors (BHUB). In the three methods, deterministic SD parameters are replaced by latent variable posterior distributions for each speaker, whose parameters are efficiently estimated using a variational inference based approach. Experiments conducted on 300-hour speed perturbed Switchboard corpus trained LF-MMI TDNN/CNN-TDNN systems suggest the proposed Bayesian adaptation approaches consistently outperform the deterministic adaptation on the NIST Hub5'00 and RT03 evaluation sets. When using only the first five utterances from each speaker as adaptation data, significant word error rate reductions up to 1.4% absolute (7.2% relative) were obtained on the CallHome subset. The efficacy of the proposed Bayesian adaptation techniques is further demonstrated in a comparison against the state-of-the-art performance obtained on the same task using the most recent systems reported in the literature. Xurong Xie, Xunying Liu, Tan Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Resting-State EEG-Based Biometrics with Signals Features Extracted by Multivariate Empirical Mode DecompositionabstractEEG-based biometrics has gained great attention in recent years due to its superiority over traditional biometrics in terms of its resistance to circumvention. While there are numerous choices of data acquisition protocol, the present study is carried out with the least demanding resting-state condition. Motivated by neurophysiological knowledge, a type of novel feature, namely the intrinsic mode correlation (IMCOR), is proposed. It is designed by combining the nonstationary multivariate empirical mode decomposition (NA-MEMD) and the concept of brain connectivity. With machine learning classifiers, our system yields promising performance in a 81-class classification (F1 score: 0.99) within a single session. For 32-class cross-session classification, an F1 score of 0.55 is attained. The results suggest that the proposed method might be vulnerable to temporal effects and between-session variability. This study highlights the uniqueness of the proposed non-stationary and connectivity-based feature and demonstrated its success as a biometrics. Further investigation is needed to make the method practically useful. Matthew K. H. Ma, Tan Lee, Manson Cheuk-Man Fong, William S.-Y. Wang |
ICASSP | 2 |
| 2020 | Mixture Factorized Auto-Encoder for Unsupervised Hierarchical Deep Factorization of Speech SignalabstractSpeech signal is constituted and contributed by various informative factors, such as linguistic content and speaker characteristic. There have been notable recent studies attempting to factorize speech signal into these individual factors without requiring any annotation. These studies typically assume continuous representation for linguistic content, which is not in accordance with general linguistic knowledge and may make the extraction of speaker information less successful. This paper proposes the mixture factorized auto-encoder (mFAE) for unsupervised deep factorization. The encoder part of mFAE comprises a frame tokenizer and an utterance embedder. The frame tokenizer models linguistic content of input speech with a discrete categorical distribution. It performs frame clustering by assigning each frame a soft mixture label. The utterance embedder generates an utterance-level vector representation. A frame decoder serves to reconstruct speech features from the encoders' outputs. The mFAE is evaluated on speaker verification (SV) task and unsupervised subword modeling (USM) task. The SV experiments on VoxCeleb 1 show that the utterance embedder is capable of extracting speaker-discriminative embeddings with performance comparable to a x-vector baseline. The USM experiments on ZeroSpeech 2017 dataset verify that the frame tokenizer is able to capture linguistic content and the utterance embedder can acquire speaker-related information. Siyuan Feng 0001, Tan Lee |
ICASSP | 3 |
| 2020 | Time-Frequency Feature Decomposition Based on Sound Duration for Acoustic Scene ClassificationabstractAcoustic scene classification is the task of identifying the type of acoustic environment in which a given audio signal is recorded. The signal is a mixture of sound events with various characteristics. In-depth and focused analysis is needed to find out the most representative sound patterns for recognizing and differentiating the scenes. In this paper, we propose a feature decomposition method based on temporal median filtering, and use convolutional neural network to model long-duration background sounds and transient sounds separately. Experiments on log-mel and wavelet based time-frequency features show that using the proposed method leads to better classification accuracy. Analysis of detailed experimental results reveals that (1) long-duration sounds are generally most informative for acoustic scene classification; and (2) the focus of sound duration may be different for classifying different types of acoustic scenes. Yuzhong Wu, Tan Lee |
ICASSP | 2 |
| 2020 | Text-Independent Speaker Verification with Dual Attention NetworkabstractThis paper presents a novel design of attention model for textindependent speaker verification.The model takes a pair of input utterances and generates an utterance-level embedding to represent speaker-specific characteristics in each utterance.The input utterances are expected to have highly similar embeddings if they are from the same speaker.The proposed attention model consists of a self-attention module and a mutual attention module, which jointly contributes to the generation of the utterance-level embedding.The self-attention weights are computed from the utterance itself while the mutual-attention weights are computed with the involvement of the other utterance in the input pairs.As a result, each utterance is represented by a self-attention weighted embedding and a mutual-attention weighted embedding.The similarity between the embeddings is measured by a cosine distance score and a binary classifier output score.The whole model, named Dual Attention Network, is trained end-to-end on Voxceleb database.The evaluation results on Voxceleb 1 test set show that the Dual Attention Network significantly outperforms the baseline systems.The best result yields an equal error rate of 1.6%. Tan Lee |
INTERSPEECH | 2 |
| 2020 | Advancing Multiple Instance Learning with Attention Modeling for Categorical Speech Emotion RecognitionabstractCategorical speech emotion recognition is typically performed as a sequence-to-label problem, i.e., to determine the discrete emotion label of the input utterance as a whole. One of the main challenges in practice is that most of the existing emotion corpora do not give ground truth labels for each segment; instead, we only have labels for whole utterances. To extract segment-level emotional information from such weakly labeled emotion corpora, we propose using multiple instance learning (MIL) to learn segment embeddings in a weakly supervised manner. Also, for a sufficiently long utterance, not all of the segments contain relevant emotional information. In this regard, three attention-based neural network models are then applied to the learned segment embeddings to attend the most salient part of a speech utterance. Experiments on the CASIA corpus and the IEMOCAP database show better or highly competitive results than other state-of-the-art approaches. Shuiyang Mao, Pak-Chung Ching, C.-C. Jay Kuo, Tan Lee |
INTERSPEECH | 4 |
| 2020 | Emotion Profile Refinery for Speech Emotion ClassificationabstractHuman emotions are inherently ambiguous and impure.When designing systems to anticipate human emotions based on speech, the lack of emotional purity must be considered.However, most of the current methods for speech emotion classification rest on the consensus, e. g., one single hard label for an utterance.This labeling principle imposes challenges for system performance considering emotional impurity.In this paper, we recommend the use of emotional profiles (EPs), which provides a time series of segment-level soft labels to capture the subtle blends of emotional cues present across a specific speech utterance.We further propose the emotion profile refinery (EPR), an iterative procedure to update EPs.The EPR method produces soft, dynamically-generated, multiple probabilistic class labels during successive stages of refinement, which results in significant improvements in the model accuracy.Experiments on three well-known emotion corpora show noticeable gain using the proposed method. Shuiyang Mao, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2020 | EigenEmo: Spectral Utterance Representation Using Dynamic Mode Decomposition for Speech Emotion ClassificationabstractHuman emotional speech is, by its very nature, a variant signal.This results in dynamics intrinsic to automatic emotion classification based on speech.In this work, we explore a spectral decomposition method stemming from fluid-dynamics, known as Dynamic Mode Decomposition (DMD), to computationally represent and analyze the global utterance-level dynamics of emotional speech.Specifically, segment-level emotion-specific representations are first learned through an Emotion Distillation process.This forms a multi-dimensional signal of emotion flow for each utterance, called Emotion Profiles (EPs).The DMD algorithm is then applied to the resultant EPs to capture the eigenfrequencies, and hence the fundamental transition dynamics of the emotion flow.Evaluation experiments using the proposed approach, which we call EigenEmo, show promising results.Moreover, due to the positive combination of their complementary properties, concatenating the utterance representations generated by EigenEmo with simple EPs averaging yields noticeable gains. Shuiyang Mao, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2020 | Automatic Detection of Phonological Errors in Child Speech Using Siamese Recurrent AutoencoderabstractSpeech sound disorder (SSD) refers to the developmental disorder in which children encounter persistent difficulties in correctly pronouncing words.Assessment of SSD has been relying largely on trained speech and language pathologists (SLPs).With the increasing demand for and long-lasting shortage of SLPs, automated assessment of speech disorder becomes a highly desirable approach to assisting clinical work.This paper describes a study on automatic detection of phonological errors in Cantonese speech of kindergarten children, based on a newly collected large speech corpus.The proposed approach to speech error detection involves the use of a Siamese recurrent autoencoder, which is trained to learn the similarity and discrepancy between phone segments in the embedding space.Training of the model requires only speech data from typically developing (TD) children.To distinguish disordered speech from typical one, cosine distance between the embeddings of the test segment and the reference segment is computed.Different model architectures and training strategies are experimented.Results on detecting the 6 most common consonant errors demonstrate satisfactory performance of the proposed model, with the average precision value from 0.82 to 0.93. Si Ioi Ng, Tan Lee |
INTERSPEECH | 2 |
| 2020 | CUCHILD: A Large-Scale Cantonese Corpus of Child Speech for Phonology and Articulation AssessmentabstractThis paper describes the design and development of CUCHILD, a large-scale Cantonese corpus of child speech. The corpus contains spoken words collected from 1,986 child speakers aged from 3 to 6 years old. The speech materials include 130 words of 1 to 4 syllables in length. The speakers cover both typically developing (TD) children and children with speech disorder. The intended use of the corpus is to support scientific and clinical research, as well as technology development related to child speech assessment. The design of the corpus, including selection of words, participants recruitment, data acquisition process, and data pre-processing are described in detail. The results of acoustical analysis are presented to illustrate the properties of child speech. Potential applications of the corpus in automatic speech recognition, phonological error detection and speaker diarization are also discussed. Si Ioi Ng, Cymie Wing-Yee Ng, Jiarui Wang 0003, Tan Lee, Kathy Yuet-Sheung Lee, Michael C. F. Tong |
INTERSPEECH | 4 |
| 2020 | Learning Syllable-Level Discrete Prosodic Representation for Expressive Speech Generation
Guangyan Zhang, Tan Lee |
INTERSPEECH | 3 |
| 2019 | Revisiting Hidden Markov Models for Speech Emotion RecognitionabstractHidden Markov models (HMMs) have a long tradition in automatic speech recognition (ASR) due to their capability of capturing temporal dynamic characteristics of speech. For emotion recognition from speech, three HMM based architectures are investigated and compared throughout the current paper, namely, the Gaussian mixture model based HMMs (GMM-HMMs), the subspace based Gaussian mixture model based HMMs (SGMM-HMMs) and the hybrid deep neural network HMMs (DNN-HMMs). Extensive emotion recognition experiments are carried out on these three architectures on the CASIA corpus, the Emo-DB corpus and the IEMOCAP database, respectively, and results are compared with those of state-of-the-art approaches. These HMM based architectures prove capable of constituting an effective model for speech emotion recognition. Also, the modeling accuracy is further enhanced by incorporating various advanced techniques from the ASR area. In particular, among all of the architectures, the SGMM-HMMs achieve the best performance in most of the experiments. Shuiyang Mao, Dehua Tao, Guangyan Zhang, Pak-Chung Ching, Tan Lee |
ICASSP | 5 |
| 2019 | Adversarial Multi-task Deep Features and Unsupervised Back-end Adaptation for Language RecognitionabstractThis paper presents an investigation into speaker-invariant feature learning and domain adaptation for language recognition (LR) with short utterances. While following the conventional design of i-vector front-end and probabilistic linear discriminant analysis (PLDA) back-end, we propose to apply speaker adversarial multi-task learning (AMTL) to aim explicitly at learning speaker-invariant multilingual bottleneck features and perform unsupervised PLDA adaptation to alleviate performance degradation caused by domain mismatch between training and test data. Through a demo experiment, we show the adverse effect of domain mismatch and motivate the necessity of domain adaptation. LR experiments are carried out with the AP17-OLR challenge dataset to evaluate the effectiveness of the proposed methods in comparison with the state of the art. The results show that both speaker AMTL and unsupervised PLDA adaptation contribute significantly to performance improvement on the short-duration LR task. The effectiveness of PLDA adaptation is found to be insensitive to the number of clusters assumed in unsupervised data labeling. Our best system outperforms the state-of-the-art system of AP17-OLR and shows relative improvements of 6.98% in terms of Cavgand 4.80% in terms of EER on 1-second test set. Siyuan Feng 0001, Tan Lee |
ICASSP | 3 |
| 2019 | Combining Phone Posteriorgrams from Strong and Weak Recognizers for Automatic Speech Assessment of People with AphasiaabstractThis paper presents an investigation on applying automatic speech recognition (ASR) to speech assessment of people with aphasia (PWA). A distinctive characteristic of PWA speech is paraphasia, which refers to frequent occurrence of phonemic errors, unintended words and non-verbal sounds. In view of the wide variety of paraphasias, we propose to view the ASR errors so caused as out-of-vocabulary (OOV) words. Inspired by previous research on OOV detection, paraphasias in PWA speech are captured by comparing the phone posteriorgrams of a strongly constrained speech recognizer and a weakly constrained one. The posteriorgrams also reveal other characteristics of impaired speech, e.g., change of speaking rate, voice abnormality. Siamese and 2-channel convolutional neural network (CNN) models are used for classifying the posteriorgram pairs and predicting the severity of aphasia. Experimental results on a Cantonese database of PWA speech confirm the effectiveness of the proposed methods. The best F1 score attained on binary classification (severe versus mild aphasia) is 0.891. Tan Lee, Anthony Pak-Hin Kong |
ICASSP | 2 |
| 2019 | Enhancing Sound Texture in CNN-based Acoustic Scene ClassificationabstractAcoustic scene classification is the task of identifying the scene from which the audio signal is recorded. Convolutional neural network (CNN) models are widely adopted with proven successes in acoustic scene classification. However, there is little insight on how an audio scene is perceived in CNN, as what have been demonstrated in image recognition research. In the present study, the Class Activation Mapping (CAM) is utilized to analyze how the log-magnitude Mel-scale filter-bank (log-Mel) features of different acoustic scenes are learned in a CNN classifier. It is noted that distinct high-energy time-frequency components of audio signals generally do not correspond to strong activation on CAM, while the background sound texture are well learned in CNN. In order to make the sound texture more salient, we propose to apply the Difference of Gaussian (DoG) and Sobel operator to process the log-Mel features and enhance edge information of the time-frequency image. Experimental results on the DCASE 2017 ASC challenge show that using edge enhanced log-Mel images as input feature of CNN significantly improves the performance of audio scene classification. Yuzhong Wu, Tan Lee |
ICASSP | 2 |
| 2019 | BLHUC: Bayesian Learning of Hidden Unit Contributions for Deep Neural Network Speaker AdaptationabstractSpeaker adaptation techniques play a key role in reducing the mismatch between speech recognition systems and target users. In order to robustly learn speaker-dependent adaptation parameters, model based DNN adaptation techniques often require a significant amount of data. For example, in the commonly used learning hidden unit contributions (LHUC) based DNN adaptation, speaker-dependent high-dimensional hidden layer output scaling vectors are used. When limited adaptation data are available, the standard L-HUC is prone to over-fitting and poor generalization. To address the issue, Bayesian learning of hidden unit contributions (BLHUC) is proposed in this paper. A posterior distribution over the LHUC scaling vectors is used to explicitly model the uncertainty associated with the adaptation parameters. An efficient variational inference based approach is adopted to estimate the LHUC parameter posterior distribution. Experiments conducted on a 300-hour Switchboard setup showed that the proposed BLHUC method outperformed the baseline speaker-independent DNN systems and LHUC adapted DNN systems by up to 1.4% and 1.1% absolute reductions of word error rate respectively, when only using 1 utterance of adaptation data from each speaker. Consistent performance improvements were also obtained over the baseline, LHUC adapted and LHUC SAT systems when increasing the amount of adaptation data. Xurong Xie, Xunying Liu, Tan Lee, Shoukang Hu |
ICASSP | 3 |
| 2019 | Improving Unsupervised Subword Modeling via Disentangled Speech Representation Learning and TransformationabstractThis study tackles unsupervised subword modeling in the zero-resource scenario, learning frame-level speech representation that is phonetically discriminative and speaker-invariant, using only untranscribed speech for target languages. Frame label acquisition is an essential step in solving this problem. High quality frame labels should be in good consistency with golden transcriptions and robust to speaker variation. We propose to improve frame label acquisition in our previously adopted deep neural network-bottleneck feature (DNN-BNF) architecture by applying the factorized hierarchical variational autoencoder (FHVAE). FHVAEs learn to disentangle linguistic content and speaker identity information encoded in speech. By discarding or unifying speaker information, speaker-invariant features are learned and fed as inputs to DPGMM frame clustering and DNN-BNF training. Experiments conducted on ZeroSpeech 2017 show that our proposed approaches achieve $2.4\%$ and $0.6\%$ absolute ABX error rate reductions in across- and within-speaker conditions, comparing to the baseline DNN-BNF system without applying FHVAEs. Our proposed approaches significantly outperform vocal tract length normalization in improving frame labeling and subword modeling. Siyuan Feng 0001, Tan Lee |
INTERSPEECH | 2 |
| 2019 | Combining Adversarial Training and Disentangled Speech Representation for Robust Zero-Resource Subword ModelingabstractThis study addresses the problem of unsupervised subword unit discovery from untranscribed speech. It forms the basis of the ultimate goal of ZeroSpeech 2019, building text-to-speech systems without text labels. In this work, unit discovery is formulated as a pipeline of phonetically discriminative feature learning and unit inference. One major difficulty in robust unsupervised feature learning is dealing with speaker variation. Here the robustness towards speaker variation is achieved by applying adversarial training and FHVAE based disentangled speech representation learning. A comparison of the two approaches as well as their combination is studied in a DNN-bottleneck feature (DNN-BNF) architecture. Experiments are conducted on ZeroSpeech 2019 and 2017. Experimental results on ZeroSpeech 2017 show that both approaches are effective while the latter is more prominent, and that their combination brings further marginal improvement in across-speaker condition. Results on ZeroSpeech 2019 show that in the ABX discriminability task, our approaches significantly outperform the official baseline, and are competitive to or even outperform the official topline. The proposed unit sequence smoothing algorithm improves synthesis quality, at a cost of slight decrease in ABX discriminability. Siyuan Feng 0001, Tan Lee |
INTERSPEECH | 2 |
| 2019 | Deep Learning of Segment-Level Feature Representation with Multiple Instance Learning for Utterance-Level Speech Emotion Recognition
Shuiyang Mao, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2019 | Automatic Assessment of Language Impairment Based on Raw ASR Output
Tan Lee, Anthony Pak-Hin Kong |
INTERSPEECH | 2 |
| 2019 | Child Speech Disorder Detection with Siamese Recurrent Network Using Speech Attribute Features
Jiarui Wang 0003, Tan Lee |
INTERSPEECH | 4 |
| 2019 | Fast DNN Acoustic Model Speaker Adaptation by Learning Hidden Unit Contribution Features
Xurong Xie, Xunying Liu, Tan Lee |
INTERSPEECH | 3 |
| 2019 | Exploiting Cross-Lingual Speaker and Phonetic Diversity for Unsupervised Subword ModelingabstractThis research addresses the problem of acoustic modeling of low-resource languages for which transcribed training data is absent. The goal is to learn robust frame-level feature representations that can be used to identify and distinguish subword-level speech units. The proposed feature representations comprise various types of multilingual bottleneck features (BNFs) that are obtained via multi-task learning of deep neural networks (MTL-DNN). One of the key problems is how to acquire high-quality frame labels for untranscribed training data to facilitate supervised DNN training. It is shown that learning of robust BNF representations can be achieved by effectively leveraging transcribed speech data and well-trained automatic speech recognition (ASR) systems from one or more out-of-domain (resource-rich) languages. Out-of-domain ASR systems can be applied to perform speaker adaptation with untranscribed training data of the target language, and to decode the training speech into frame-level labels for DNN training. It is also found that better frame labels can be generated by considering temporal dependency in speech when performing frame clustering. The proposed methods of feature learning are evaluated on the standard task of unsupervised subword modeling in Track 1 of the ZeroSpeech 2017 Challenge. The best performance achieved by our system is 9.7% in terms of across-speaker triphone minimal-pair ABX error rate, which is comparable to the best systems reported recently. Lastly, our investigation reveals that the closeness between target languages and out-of-domain languages and the amount of available training data for individual target languages could have significant impact on the goodness of learned features. Siyuan Feng 0001, Tan Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Acoustical Assessment of Voice Disorder With Continuous Speech Using ASR Posterior FeaturesabstractTraditionally acoustical assessment of voice disorder relies on simple and homogeneous speech samples like sustained vowels. Continuous speech is believed to be more representative of the daily function of voice and more preferable in clinical practice. This paper describes an attempt on automating voice assessment with continuous speech utterances. The proposed system makes use of a novel type of features that are derived from phone posterior probabilities outputted by a deep neural network based automatic speech recognition (ASR) system. These ASR-based voice features are designed to effectively quantify the mismatch between disordered voice and normal voice. Prediction of voice disorder severity is carried out first at utterance-level and subsequently the prediction scores for individual utterances from a subject are combined to give an overall-assessment on the subject. With a low-dimension ASR-based feature vector, the utterance-level prediction accuracy is comparable to that with conventional features with a much higher dimension. By jointly using the ASR features and conventional voice features, a subject-level prediction accuracy of over 80% on three severity classes can be achieved. Subjects with mild disorder and those with severe disorder could be perfectly distinguished by the proposed method. Yuanyuan Liu 0002, Tan Lee, Thomas K. T. Law, Kathy Yuet-Sheung Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Automatic Speech Assessment for Aphasic Patients Based on Syllable-Level Embedding and Supra-Segmental Duration FeaturesabstractAphasia is a type of acquired language impairment resulting from brain injury. Speech assessment is an important part of the comprehensive assessment process for aphasic patients. It is based on the acoustical and linguistic analysis of patients' speech elicited through pre-defined story-telling tasks. This type of narrative spontaneous speech embodies multi-fold atypical characteristics related to the underlying language impairment. This paper presents an investigation on automatic speech assessment for Cantonese-speaking aphasic patients using an automatic speech recognition (ASR) system. A novel approach to extracting robust text features from erroneous ASR output is developed based on word embedding methods. The text features can effectively distinguish the stories told by an impaired speaker from those by unimpaired ones. On the other hand, a set of supra-segmental duration features are derived from syllable-level time alignments produced by the ASR system, to characterize the atypical prosody of impaired speech. The proposed text features, duration features and their combination are evaluated in a binary classification experiment as well as in automatic prediction of subjective assessment score. The results clearly show that the text features are very effective in the intended task of aphasia assessment, while using duration features could provide additional benefit. Tan Lee, Anthony Pak-Hin Kong |
ICASSP | 2 |
| 2018 | Reducing Model Complexity for DNN Based Large-Scale Audio ClassificationabstractAudio classification is the task of identifying the sound categories that are associated with a given audio signal. This paper presents an investigation on large-scale audio classification based on the recently released AudioSet database. AudioSet comprises 2 millions of audio samples from YouTube, which are human-annotated with 527 sound category labels. Audio classification experiments with the balanced training set and the evaluation set of AudioSet are carried out by applying different types of neural network models. The classification performance and the model complexity of these models are compared and analyzed. While the CNN models show better performance than MLP and RNN, its model complexity is relatively high and undesirable for practical use. We propose two different strategies that aim at constructing low-dimensional embedding feature extractors and hence reducing the number of model parameters. It is shown that the simplified CNN model has only 1/22 model parameters of the original model, with only a slight degradation of performance. Yuzhong Wu, Tan Lee |
ICASSP | 2 |
| 2018 | Improving Cross-Lingual Knowledge Transferability Using Multilingual TDNN-BLSTM with Language-Dependent Pre-Final Layer
Siyuan Feng 0001, Tan Lee |
INTERSPEECH | 2 |
| 2018 | Exploiting Speaker and Phonetic Diversity of Mismatched Language Resources for Unsupervised Subword Modeling
Siyuan Feng 0001, Tan Lee |
INTERSPEECH | 2 |
| 2018 | Cross-cultural (A)symmetries in Audio-visual Attitude PerceptionabstractInternational audience Hansjörg Mixdorff, Albert Rilliard, Tan Lee, Matthew K. H. Ma, Angelika Hönemann |
INTERSPEECH | 3 |
| 2018 | Automatic Speech Assessment for People with Aphasia Using TDNN-BLSTM with Multi-Task LearningabstractThis paper describes an investigation on automatic speech assessment for people with aphasia (PWA) using a DNN based automatic speech recognition (ASR) system. The main problems being addressed are the lack of training speech in the intended application domain and the relevant degradation of ASR performance for impaired speech of PWA. We adopt the TDNN-BLSTM structure for acoustic modeling and apply the technique of multi-task learning with large amount of domain-mismatched data. This leads to a significant improvement on the recognition accuracy, as compared with a conventional single-task learning DNN system. To facilitate the extraction of robust text features for quantifying language impairment in PWA speech, we propose to incorporate N-best hypotheses and confusion network representation of the ASR output. The severity of impairment is predicted from text features and supra-segmental duration features using different regression models. Experimental results show a high correlation of 0.842 between the predicted severity level and the subjective Aphasia Quotient score. Tan Lee, Siyuan Feng 0001, Anthony Pak-Hin Kong |
INTERSPEECH | 2 |
| 2017 | Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogramabstractAutomatic music transcription is usually approached by using a time-frequency (TF) representation such as the short-time Fourier transform (STFT) spectrogram or the constant-Q transform. In this paper, we propose a novel yet simple TF representation that capitalizes the effectiveness of spectral flux features in highlighting note onset times. We refer to this representation as the differential spectrogram and investigate its usefulness for note-level piano transcription using two different non-negative matrix factorization (NMF) algorithms. Experiments on the MAPS ENSTDkCl dataset validate the advantages of the differential spectrogram over the STFT spectrogram for this task. Moreover, by adapting a state-of-the-art convolutional NMF algorithm with the differential spectrogram, we can achieve even better accuracy than the state-of-the-art on this dataset. Our analysis shows that the new representation suppresses unwanted TF patterns and performs particularly well in improving the recall rate. Lufei Gao, Li Su 0004, Yi-Hsuan Yang, Tan Lee |
ICASSP | 4 |
| 2017 | Shefce: A Cantonese-English bilingual speech corpus for pronunciation assessmentabstractThis paper introduces the development of ShefCE: a Cantonese-English bilingual speech corpus from L2 English speakers in Hong Kong. Bilingual parallel recording materials were chosen from TED online lectures. Script selection were carried out according to bilingual consistency (evaluated using a machine translation system) and the distribution balance of phonemes. 31 undergraduate to postgraduate students in Hong Kong aged 20-30 were recruited and recorded a 25-hour speech corpus (12 hours in Cantonese and 13 hours in English). Baseline phoneme/syllable recognition systems were trained on background data with and without the ShefCE training data. The final syllable error rate (SER) for Cantonese is 17.3% and final phoneme error rate (PER) for English is 34.5%. The automatic speech recognition performance on English showed a significant mismatch when applying L1 models on L2 data, suggesting the need for explicit accent adaptation. ShefCE and the corresponding baseline models will be made openly available for academic research. Raymond W. M. Ng, Alvin C. M. Kwan, Tan Lee, Thomas Hain |
ICASSP | 3 |
| 2017 | On the Linguistic Relevance of Speech Units Learned by Unsupervised Acoustic Modeling
Siyuan Feng 0001, Tan Lee |
INTERSPEECH | 2 |
| 2017 | Acoustic Assessment of Disordered Voice with Continuous Speech Based on Utterance-Level ASR Posterior Features
Yuanyuan Liu 0002, Tan Lee, Pak-Chung Ching, Thomas K. T. Law, Kathy Yuet-Sheung Lee |
INTERSPEECH | 2 |
| 2017 | RNN-LDA Clustering for Feature Based DNN Adaptation
Xurong Xie, Xunying Liu, Tan Lee |
INTERSPEECH | 3 |
| 2017 | Audio-visual expressions of attitude: How many different attitudes can perceivers decode?
Hansjörg Mixdorff, Angelika Hönemann, Albert Rilliard, Tan Lee, Matthew K. H. Ma |
Speech Commun. | 4 |
| 2016 | Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speechabstractThis paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the types and amount of speech materials used for acoustical assessment were very limited. With the ASR technology, we are able to perform acoustical and linguistic analyses with a large amount of natural speech from impaired speakers. The present study is focused on Cantonese, which is a major Chinese dialect. Two representative disorders of speech production are investigated: dysphonia and aphasia. ASR experiments are carried out with continuous and spontaneous speech utterances from Cantonese-speaking patients. The results confirm the feasibility and potential of using natural speech for acoustical assessment of voice and speech disorders, and reveal the challenging issues in acoustic modeling and language modeling of pathological speech. Tan Lee, Yuanyuan Liu 0002, Pei-Wen Huang, Jen-Tzung Chien, Wang-Kong Lam, Yu Ting Yeung, Thomas K. T. Law, Kathy Yuet-Sheung Lee, Anthony Pak-Hin Kong, Sam-Po Law |
ICASSP | 1 |
| 2016 | Hybrid Accelerated Optimization for Speech Recognition
Jen-Tzung Chien, Pei-Wen Huang, Tan Lee |
INTERSPEECH | 3 |
| 2016 | Predicting Severity of Voice Disorder from DNN-HMM Acoustic Posteriors
Tan Lee, Yuanyuan Liu 0002, Yu Ting Yeung, Thomas K. T. Law, Kathy Yuet-Sheung Lee |
INTERSPEECH | 1 |
| 2015 | Modeling temporal dependency for robust estimation of LP model parameters in speech enhancement
Chun Hoy Wong, Tan Lee, Yu Ting Yeung, Pak-Chung Ching |
INTERSPEECH | 2 |
| 2015 | Multi-pitch estimation based on sparse representation with pre-screened dictionaryabstractThis paper presents a study on frame-level multi-pitch estimation (MPE) for polyphonic piano music based on sparse representation approach. In this approach, a multi-pitch input spectrum is represented by a sparse linear combination of a large number of spectrum exemplars in a given dictionary. By estimating the sparse weight vector and identifying its non-zero elements, a set of possible pitch candidates can be found. This study is focused mainly on the construction and optimization of the exemplar dictionary. A complete dictionary is first built from single-note piano music. We propose to perform prescreening on the dictionary by which the exemplars of the notes belonging to certain octaves and chromas are excluded during the subsequent estimation. Experimental results show that the pre-screening process not only helps in reducing the computational complexity, but also leads to more accurate pitch estimation. On the formulation of sparse estimation problem, we introduce a probabilistic assumption on the estimation error, such that the estimation is converted into a constrained convex quadratic programming problem. We also propose to use spectral combination as a new scheme of pitch determination from the estimated sparse weights. Lufei Gao, Tan Lee |
MMSP | 2 |
| 2015 | Objective measures for quality assessment of noise-suppressed speech
Huijun Ding, Tan Lee, Ing Yann Soon, Chai Kiat Yeo, Peng Dai 0002, Guo Dan |
Speech Commun. | 2 |
| 2015 | A method of speech periodicity enhancement using transform-domain signal decomposition
Feng Huang 0002, Tan Lee, W. Bastiaan Kleijn, Ying-Yee Kong |
Speech Commun. | 2 |
| 2015 | Acoustic Segment Modeling with Spectral Clustering MethodsabstractThis paper presents a study of spectral clustering-based approaches to acoustic segment modeling (ASM). ASM aims at finding the underlying phoneme-like speech units and building the corresponding acoustic models in the unsupervised setting, where no prior linguistic knowledge and manual transcriptions are available. A typical ASM process involves three stages, namely initial segmentation, segment labeling, and iterative modeling. This work focuses on the improvement of segment labeling. Specifically, we use posterior features as the segment representations, and apply spectral clustering algorithms on the posterior representations. We propose a Gaussian component clustering (GCC) approach and a segment clustering (SC) approach. GCC applies spectral clustering on a set of Gaussian components, and SC applies spectral clustering on a large number of speech segments. Moreover, to exploit the complementary information of different posterior representations, a multiview segment clustering (MSC) approach is proposed. MSC simultaneously utilizes multiple posterior representations to cluster speech segments. To address the computational problem of spectral clustering in dealing with large numbers of speech segments, we use inner product similarity graph and make reformulations to avoid the explicit computation of the affinity matrix and Laplacian matrix. We carried out two sets of experiments for evaluation. First, we evaluated the ASM accuracy on the OGI-MTS dataset, and it was shown that our approach could yield 18.7% relative purity improvement and 15.1% relative NMI improvement compared with the baseline approach. Second, we examined the performances of our approaches in the real application of zero-resource query-by-example spoken term detection on SWS2012 dataset, and it was shown that our approaches could provide consistent improvement on four different testing scenarios with three evaluation metrics. Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Supervised Single-Microphone Multi-Talker Speech Separation with Conditional Random FieldsabstractWe apply conditional random field (CRF) for single-microphone speech separation in a supervised learning scenario. We train the parameters with mixture data in which the sources are competing with the same average signal power. Compared with factorial hidden Markov model (HMM) baselines, the CRF settings require fewer training mixture data to improve objective speech quality measures and speech recognition accuracy of the reconstructed sources, when mixing ratios of training and testing mixture data are matched. The CRF settings also handle minor mixing ratio mismatch after adjusting the gain factors of the sources with non-linear mappings inspired from the mixture-maximization model. When the mixing ratio mismatch further increases such that the speech mixture is dominated by only one source, factorial HMM finally catches up with and performs better than the CRF settings due to improved model accuracy. We also develop a convex statistical inference simplification based on linear-chain CRFs. The simplification achieves the same performance level as the original CRF settings after integrating additional observations. Yu Ting Yeung, Tan Lee, Cheung-Chi Leung |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | A graph-based Gaussian component clustering approach to unsupervised acoustic modeling
Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2014 | Large-margin conditional random fields for single-microphone speech separation
Yu Ting Yeung, Tan Lee, Cheung-Chi Leung |
INTERSPEECH | 2 |
| 2014 | Correcting Chord Classification Errors Based on Tonal Organization Information of Classical MusicabstractChord progression in Classical music is not random. It follows some specific structure and rules. In this paper, we present a generative account of chord progression described by a set of phrase-structure grammar rules from Martin Rohr Meier. With some modifications and simplifications, these rules can be used to correct chord errors. A chord classification system for classical music is developed. With the aid of beat and melodic information, processed pitch class vector is obtained from the extracted constant-Q transform spectra. Exploiting tonal grammar rules, the most probable and musically sensible chord sequence is derived from that sequence of pitch class vectors. Some examples of classical piano excerpts are evaluated. The experimental result shows that our system is useful to improve chord classification accuracy. Wang-Kong Lam, Tan Lee |
ISM | 2 |
| 2013 | Evaluation of pitch estimation algorithms on separated speechabstractTo post-process outputs of speech separation systems with harmonic enhancement, it is normally required to estimate the fundamental frequency. This paper evaluates the performance of a few representative robust pitch estimation algorithms on speech reconstructed from two-speaker mixture signals. The separation outputs obtained by two state-of-the-art single-channel separation algorithms are used for the evaluation. A recently proposed sparsity-based pitch estimation method is applied to the separated speech and a new pitch tracking algorithm is proposed. Experimental results show that on the separated speech the proposed method consistently surpasses the others with significantly low gross error rate, which is similar to the gross error rates of the other methods on clean speech. Feng Huang 0002, Yu Ting Yeung, Tan Lee |
ICASSP | 3 |
| 2013 | Using parallel tokenizers with DTW matrix combination for low-resource spoken term detectionabstractRecently the posteriorgram-based template matching framework has been successfully applied to query-by-example spoken term detection tasks for low-resource languages. This framework employs a tokenizer to derive posteriorgrams, and applies dynamic time warping (DTW) to the posteriorgrams to locate the possible occurrences of a query term. Based on this framework, we propose to improve the detection performance by using multiple tokenizers with DTW distance matrix combination. The proposed approach uses multiple tokenizers in parallel as the front-end to generate different posteriorgram representations, and combines the distance matrices of the different posteriorgrams into a single matrix. DTW detection is then applied to the combined distance matrix. Lastly score post-processing techniques including pseudo-relevance feedback and score normalization are used for further improvement. Experiments were conducted on the spoken web search datasets of MediaEval 2011 and MediaEval 2012. Experimental results show that combining multiple tokenizers significantly outperforms the best single tokenizer, and that the DTW matrix combination method consistently outperforms the score combination method when more than three tokenizers are involved. Score post-processing techniques show further gains on top of using multiple tokenizers. Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2013 | Using dynamic conditional random field on single-microphone speech separationabstractThe use of dynamic conditional random field (DCRF) for model-based single-microphone speech separation is investigated. The speech sources are represented by acoustic state sequences from speaker-dependent acoustic models. The posterior probabilities of the source acoustic states given a speech mixture are inferred with a maximum entropy probability distribution which is represented by DCRF. The posterior probabilities are needed for minimum mean-square error estimation of the speech sources. Loopy belief propagation is applied for the inference. Averaged stochastic gradient descent and limited-memory BFGS are compared for parameter estimation. With the log-magnitude spectrum of the speech mixture as input observation, the proposed method achieves better separation performance in terms of Blind Source Separation Metrics (SDR, SAR, SIR) and PESQ than a factorial hidden Markov model baseline system in our experiments. Yu Ting Yeung, Tan Lee, Cheung-Chi Leung |
ICASSP | 2 |
| 2013 | Unsupervised mining of acoustic subword units with segment-level Gaussian posteriorgrams
Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2013 | Shifted-Delta MLP Features for Spoken Language RecognitionabstractThis letter presents our study of applying phoneme posterior features for spoken language recognition (SLR). In our work, phoneme posterior features are estimated from a multilayer perceptron (MLP) based phoneme recognizer, and are further processed through transformations including taking logarithm, PCA transformation, and appending shifted delta coefficients. The resulting shifted-delta MLP (SDMLP) features show similar distribution as conventional shifted-delta cepstral (SDC) features, and are more robust compared to the SDC features. Experiments on the NIST LRE2005 dataset show that the SDMLP features fit well with the state-of-the-art GMM-based SLR systems, and SDMLP features outperform SDC features significantly. Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2013 | Pitch Estimation in Noisy Speech Using Accumulated Peak Spectrum and Sparse Estimation TechniqueabstractPitch estimation from acoustic signals is a fundamental problem in many areas of speech research. For noise-corrupted speech, reliable pitch estimation is difficult. This paper presents a study of pitch estimation in noisy speech based on robust temporal-spectral representation and sparse reconstruction. We propose to accumulate spectral peaks over consecutive time frames. Since harmonic structure of speech changes much more slowly than noise spectrum, spectral peaks related to pitch harmonics would stand out over the noise through the accumulation. Experimental results show that the accumulated peak spectrum is indeed a robust representation of pitch harmonics. Subsequently, the accumulated peak spectrum is expressed as a sparse linear combination of a large set of clean peak spectrum exemplars. Gaussian mixture density is used to model noise spectrum peaks. The weights of the linear combination are estimated so as to maximize the likelihood of the accumulated peak spectrum under sparsity constraint. Robust pitch estimation is done based on the sparse weights and the corresponding peak spectrum exemplars. The use of Gaussian mixture model leads to non-convexity of the objective function for sparse weight estimation. By approximation and reformulation, two convex optimization approaches are developed to estimate the weights. Extensive experimental studies are carried out to evaluate performance of the proposed pitch estimation algorithms on a wide variety of noise conditions. It is clearly shown that the proposed methods significantly and consistently outperform the conventional methods, particularly at very low signal-to-noise ratios (e.g., SNR <; -5 dB). Feng Huang 0002, Tan Lee |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Spoken Language Recognition With Prosodic FeaturesabstractSpeech prosody is believed to carry much language-specific information that can be used for spoken language recognition (SLR). In the past, the use of prosodic features for SLR has been studied sporadically and the reported performances were considered unsatisfactory. In this paper, we exploit a wide range of prosodic attributes for large-scale SLR tasks. These attributes describe the multifaceted variations of F0, intensity and duration in different spoken languages. Prosodic attributes are modeled by the bag of n-grams approach with support vector machine (SVM) as in the conventional phonotactic SLR systems. Experimental results on OGI and NIST-LRE tasks showed that the use of proposed attributes gives significantly better SLR performance than those previously reported. The full feature set includes 87 prosodic attributes and redundancy among attributes may exist. Attributes are broken down into particular bigrams called bins. Four entropy-based feature selection metrics with different selection criteria are derived. Attributes can be selected by individual bins, or by attributes as batches of bins. It can also be done in a language-dependent or language-independent manner. By comparing different selection sizes and criteria, an optimal attribute subset comprising 5,000 bins is found by using a bin-level language-independent criterion. Feature selection reduces model size by 2.5 times and shortens the runtime by 6 times. The optimal subset of bins gives the lowest EER of 20.18% on NIST-LRE 2007 SLR task in a prosodic attribute model (PAM) system which exclusively modeled prosodic attributes. In a phonotactic-prosodic fusion SLR system, the detection cost, Cavgis 2.09%. The relative detection cost reduction is 23%. Raymond W. M. Ng, Tan Lee, Cheung-Chi Leung, Bin Ma 0001, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Sparsity-based confidence measure for pitch estimation in noisy speechabstractIn this paper, we propose a confidence measure for the pitch estimation method presented in a parallel paper [1]. The confidence measure is derived based on a sparse representation of the speech harmonic structure. The measurement quantitatively reflects the harmonicity associated with an estimated pitch. It is used to indicate the reliability of the results. Histogram is employed to illustrate the distribution of the measurement obtained from speech signals at low signal-to-noise ratios. It is shown that with the confidence measure, correct results can be effectively identified. By using the confidence measure, an adaptive algorithm for pitch estimation is proposed. Reliable pitch values obtained from preceding estimations are identified and used to predict the local pitch range for subsequent estimations. Parameter of the estimation algorithm is dynamically adjusted according to the predicted pitch range. Experimental results show that with the adaptive algorithm, pitch estimation accuracy is noticeably improved. Feng Huang 0002, Tan Lee |
ICASSP | 2 |
| 2012 | Transform-domain Wiener filter for speech periodicity enhancementabstractIn this paper, we present a transform-domain Wiener filtering approach for enhancing speech periodicity. The enhancement is performed on the linear prediction residual signal. Two sequential lapped frequency transforms are applied to the residual in a pitch-synchronous manner. The residual signal is effectively represented by two separate sets of transform coefficients that correspond to the periodic and aperiodic components, respectively. A Wiener filter operating on the transform coefficients is developed to restore periodicity and reduce noise. Different filter parameters are designed for the transform coefficients of the periodic and aperiodic components. A template-driven method is used to estimate the filter parameters for the periodic component. For the aperiodic components, the filter parameters are computed based on a local SNR for effective noise reduction. Experimental results confirm that the harmonic structure of the signal can be effectively restored with the proposed approach. Feng Huang 0002, Tan Lee, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2012 | An acoustic segment modeling approach to query-by-example spoken term detectionabstractThe framework of posteriorgram-based template matching has been shown to be successful for query-by-example spoken term detection (STD). This framework employs a tokenizer to convert query examples and test utterances into frame-level posteriorgrams, and applies dynamic time warping to match the query posteriorgrams with test posteriorgrams to locate possible occurrences of the query term. It is not trivial to design a reliable tokenizer due to heterogeneous test conditions and the limitation of training resources. This paper presents a study of using acoustic segment models (ASMs) as the tokenizer. ASMs can be obtained following an unsupervised iterative procedure without any training transcriptions. The STD performance of the ASM tokenizer is evaluated on Fisher Corpus with comparison to three alternative tokenizers. Experimental results show that the ASM tokenizer outperforms a conventional GMM tokenizer and a language-mismatched phoneme recognizer. In addition, the performance is significantly improved by applying unsupervised speaker normalization techniques. Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 3 |
| 2012 | Integrating multiple observations for model-based single-microphone speech separation with conditional random fieldsabstractA single-microphone speech separation framework based on conditional random fields (CRFs) is proposed in this paper. Unlike factorial HMM, CRF does not have the conditional independence assumption on observations, thus different types of observations from the speech mixture can be integrated into the models through feature functions. Similar to factorial HMM, there is the statistical independence assumption on sources. Under this assumption, the two-source single-microphone speech separation problem can be expressed by two independent linear-chain CRFs. The separation problem becomes two pattern recognition problems, with respect to CRF models of the two sources. Experimental results show that by integrating initial separation outputs from factorial HMM with log power spectrum, fundamental frequency and speaker likelihoods of the mixture, CRF separation framework consistently improves the results from factorial HMM in terms of SNR, segmental SNR and PESQ. Yu Ting Yeung, Tan Lee, Cheung-Chi Leung |
ICASSP | 2 |
| 2012 | Robust Pitch Estimation Using l1-regularized Maximum Likelihood Estimation
Feng Huang 0002, Tan Lee |
INTERSPEECH | 2 |
| 2011 | Score fusion and calibration in multiple language detectors with large performance variationabstractIn a large-scale language detection task, performance variation found between different component systems and different target languages has an adverse effect to the pooled error statistics. Special care has to be taken in score fusion and calibration. In this paper, we use a prosodic LID system to fuse with a phonotactic LID system using NIST Language Recognition Evaluation 2009 experimental data. Among four logistic regression models, the one which gives the lowest Cavg is chosen. We further explore our previously proposed calibration algorithm based on the minimum erroneous deviation criterion. The algorithm is made more robust by removing the predetermined list of target languages to be calibrated, as well as by adding an optimization constraint which enforces calibration in the data portion with a large performance variation. The fusion and calibration operations together bring a 33.9% relative Cavg reduction compared with the original result from a phonotactic LID system. Raymond W. M. Ng, Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 3 |
| 2011 | Robust Speaker Recognition Using Denoised Vocal Source and Vocal Tract FeaturesabstractTo alleviate the problem of severe degradation of speaker recognition performance under noisy environments because of inadequate and inaccurate speaker-discriminative information, a method of robust feature estimation that can capture both vocal source- and vocal tract-related characteristics from noisy speech utterances is proposed. Spectral subtraction, a simple yet useful speech enhancement technique, is employed to remove the noise-specific components prior to the feature extraction process. It has been shown through analytical derivation, as well as by simulation results, that the proposed feature estimation method leads to robust recognition performance, especially at low signal-to-noise ratios. In the context of Gaussian mixture model-based speaker recognition with the presence of additive white Gaussian noise, the new approach produces consistent reduction of both identification error rate and equal error rate at signal-to-noise ratios ranging from 0 to 15 dB. Ning Wang 0052, Pak-Chung Ching, Nengheng Zheng, Tan Lee |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Prosodic attribute model for spoken language identificationabstractProsodic information is believed to carry language-specific information useful to spoken language recognition. Modeling prosodic features is a challenging problem, on which a wide diversity of approaches have been investigated. In this paper, a novel prosodic attribute model (PAM) is proposed to capture prosodic features with compact models. It models the language-specific co-occurrence statistics of a comprehensive set of prosodic features. When the prosodic LID system with PAM is evaluated in NIST Language Recognition Evaluations (LRE) 2007 and 2009, it demonstrates respectively 21% and 11% relative EER reduction compared to a phonotactic LID system. The contributions of prosodic features in detecting some of the target languages, including tonal languages, are even more substantial. It is also noted that most prosodic attributes in the comprehensive set are making positive contributions. Raymond W. M. Ng, Cheung-Chi Leung, Tan Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 3 |
| 2010 | Cross-lingual speaker adaptation via Gaussian component mapping
Houwei Cao, Tan Lee, Pak-Chung Ching |
INTERSPEECH | 2 |
| 2010 | Pitch estimation in noisy speech based on temporal accumulation of spectrum peaks
Feng Huang 0002, Tan Lee |
INTERSPEECH | 2 |
| 2010 | Perception-based automatic approximation of F0 contours in Cantonese speech
Tan Lee |
INTERSPEECH | 2 |
| 2010 | Towards long-range prosodic attribute modeling for language recognitionabstractAs a high-level feature, prosody may be an effective feature when it is modeled over longer ranges than the typical range of a syllable. This paper is about language recognition with the high-level prosodic attributes. It studies two important issues of long-range modeling, namely the data scarcity handling method, and the model which properly describes prosodic boundary events. Illustrated by NIST language recognition evaluation (LRE) 2009, long-range modeling is shown to bring a 7.2% relative improvement to a prosodic language detector. Score fusion between the long-range prosodic system and a phonotactic system gives an EER of 3.07%. Exploiting boundary N -grams is the main contributing factor to global EER reduction, while different long-range prosodic modeling factors benefit the detection of different languages. Analysis reveals the evidence of language-specific long-range prosodic attributes, which sheds light on robust long-range modeling methods for language recognition. Index Terms: language recognition, prosody, long-range modeling Raymond W. M. Ng, Cheung-Chi Leung, Ville Hautamäki, Tan Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2010 | Exploitation of phase information for speaker recognition
Ning Wang 0052, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2009 | Effects of language mixing for automatic recognition of Cantonese-English code-mixing utterances
Houwei Cao, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2009 | Model-based speech separation: identifying transcription using orthogonalityabstractSpectral envelopes and harmonics are the building elements of a speech signal. By estimating these elements, individual speech sources in a mixture observation can be reconstructed and hence separated. Transcription gives the spoken content. More important, it describes the expected sequence of spectral envelopes, if modeling of different speech sounds is acquired. Our recently proposed single-microphone speech separation algorithm exploits this to derive the spectral envelope trajectories of individual sources and remove interference accordingly. The correctness of such transcription becomes critical to the separation performance. This paper investigates the relationship between the correctness of transcription hypotheses and the orthogonality of associated source estimates. An orthogonality measure is introduced to quantify the correlation between spectrograms. Experiments verify that underlying true transcriptions lead to a salient orthogonality distribution, which is distinguishable from the counterfeit transcription one. Accordingly a transcription identification technique is developed, which succeeds in identifying true transcriptions in 99.74% of the experimental trials. 1 Siu Wa Lee, Frank K. Soong, Tan Lee |
INTERSPEECH | 3 |
| 2009 | Exploration of vocal excitation modulation features for speaker recognition
Ning Wang 0052, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2008 | Language modeling for speech recognition of spoken Cantonese
Yu Ting Yeung, Houwei Cao, Nengheng Zheng, Tan Lee, Pak-Chung Ching |
INTERSPEECH | 4 |
| 2008 | Prosody for Mandarin speech recognition: a comparative study of read and spontaneous speech
Yu Ting Yeung, Yao Qian, Tan Lee, Frank K. Soong |
INTERSPEECH | 3 |
| 2008 | Tone-enhanced generalized character posterior probability (GCPP) for Cantonese LVCSR
Yao Qian, Frank K. Soong, Tan Lee |
Comput. Speech Lang. | 3 |
| 2007 | Modeling tones in hakka on the basis of the command-response model
Wentao Gu, Rerrario Shui-Ching Ho, Tan Lee |
INTERSPEECH | 3 |
| 2007 | Perceptual equivalence of approximated Cantonese tone contoursabstractThis paper describes a perceptual study on approximated Cantonese tone contours. We believe that the perception of tone contours relies mainly on the major trend of pitch movement, and is not sensitive to the exact F0 values at particular time instants. The tone contours of individual syllables and the transition between them are approximated as a small number of linear movements. The effect of such approximation is assessed by perceptual experiments. It is found that the six Cantonese tones can be represented by one or two linear movements, and the transition between tones can be represented by a single linear movement, without creating noticeable perceptual difference. Such simple approximations are desirable for perception-driven F0 modeling for text-to-speech applications. Index Terms: perceptual equivalence, Cantonese tones, prosody modeling. Tan Lee |
INTERSPEECH | 2 |
| 2007 | Integration of Complementary Acoustic Features for Speaker RecognitionabstractThis letter describes a speaker verification system that uses complementary acoustic features derived from the vocal source excitation and the vocal tract system. A new feature set, named the wavelet octave coefficients of residues (WOCOR), is proposed to capture the spectro-temporal source excitation characteristics embedded in the linear predictive residual signal. WOCOR is used to supplement the conventional vocal tract-related features, in this case, the Mel-frequency cepstral coefficients (MFCC), for speaker verification. A novel confidence measure-based score fusion technique is applied to integrate WOCOR and MFCC. Speaker verification experiments are carried out on the NIST 2001 database. The equal error rate (EER) attained with the proposed method is 7.67%, in comparison to 9.30% of the conventional MFCC-based system Nengheng Zheng, Tan Lee, Pak-Chung Ching |
IEEE Signal Process. Lett. | 2 |
| 2007 | Discrimination Power of Vocal Source and Vocal Tract Related Features for Speaker SegmentationabstractThis paper presents an analysis of the speaker discrimination power of vocal source related features, in comparison to the conventional vocal tract related features. The vocal source features, named wavelet octave coefficients of residues (WOCOR), are extracted by pitch-synchronous wavelet transform of the linear predictive (LP) residual signals. Using a series of controlled experiments, it is shown that WOCOR is less sensitive to spoken content than the conventional MFCC features and thus more discriminative when the amount of training data is limited. These advantages of WOCOR are exploited in the task of speaker segmentation for telephone conversation, in which statistical speaker models need to be built upon short speech segments. Experimental results show that the proposed use of WOCOR leads to noticeable reduction of segmentation errors. Wai Nang Chan, Nengheng Zheng, Tan Lee |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Static and Dynamic Spectral Features: Their Noise Robustness and Optimal Weights for ASRabstractIn this paper, we investigate the relative noise robustness of dynamic and static spectral features in speech recognition. It is found that the dynamic cepstrum is more robust to additive noise than its static counterpart. The results are consistent across different types of noise and over a wide range of noise levels. To exploit this unequal robustness, we propose a simple yet effective strategy of exponentially weighting the likelihoods that are contributed by the static and dynamic features during the decoding process. The optimal weights are discriminatively trained with a small amount of development data. This method is evaluated on two speaker-independent, connected digit databases, one in English (Aurora 2) and the other in Cantonese (CUDIGIT). For various types of noise at different signal-to-noise ratios (SNRs), the average relative word error rate reductions attained with the discriminatively trained weights are 36.6% and 41.9% for Aurora 2 and CUDIGIT, respectively. Noticeable performance improvement can be observed even when there is channel distortion. The proposed approach is appealing to practical applications because 1) noise estimation is not required, 2) model adaptation is not required, 3)only a minor modification of the decoding process is needed, and 4) only a few feature weights need to be trained Frank K. Soong, Tan Lee |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Use of Vocal Source Features in Speaker SegmentationabstractThis paper addresses the problem of speaker segmentation in telephone conversation. The segmentation is done in three steps: 1) preliminary segmentation to hypothesize speaker turning points; 2) clustering of segments; and 3) re-segmentation to determine speaker identity of each segment. It is found that vocal source related features are more speaker-discriminative than the conventional vocal tract related features for small amount of data. This motivates us to thoughtfully incorporate vocal source features into early stages of the speaker segmentation process, where decisions have to be made with limited data. Speaker segmentation experiments are carried out on 36 summed channel conversations in the NIST 2004 Speaker Recognition Evaluation. The proposed use of vocal source features leads to noticeable performance improvement Wai Nang Chan, Tan Lee, Nengheng Zheng, Hua Ouyang |
ICASSP (1) | 2 |
| 2006 | Feature Extraction From Talking Mouths for Video-Based Bi-Modal Speaker VerificationabstractAs the low-cost video transmission becomes popular, video-based bi-modal (audio and visual) authentication has great potential in various applications that require access control over handheld terminals. In this paper, we propose to use the averaged mouth image (AMI) for speaker verification. The AMI is computed by averaging properly aligned mouth images over the whole video sequence. Despite its simplicity, the AMI not only contains appearance information but also describes stylistic articulation gestures of individual speakers. The AMI is found to be fairly invariant against the spoken content. The experimental results show that the AMI based features are very effective in discriminating speaking persons. Explicit and precise extraction of lip contours or other feature points are not required. For bi-modal verification, the proposed video features are found to be highly complementary to the audio features Hua Ouyang, Tan Lee, Wai Nang Chan |
ICASSP (5) | 2 |
| 2006 | Tone-Enhanced Generalized Character Posterior Probability (GCPP) for Cantonese LVCSRabstractTone-enhanced, generalized character posterior probability (GCPP), a generalized form of posterior probability at subword (Chinese character) level, is proposed as a rescoring metric for improving Cantonese LVCSR performance. The search network is constructed first by converting the original word graph to a restructured word graph, then a character graph and finally, a character confusion network (CCN). Based upon GCPP enhanced with tone information, the character error rate (CER) is minimized or the GCPP product is maximized over a chosen graph. Experimental results show that the tone enhanced GCPP can improve character error rate by up to 15.1%, relatively Yao Qian, Frank K. Soong, Tan Lee |
ICASSP (1) | 3 |
| 2006 | Automatic speech recognition of Cantonese-English code-mixing utterances
Joyce Y. C. Chan, Pak-Chung Ching, Tan Lee, Houwei Cao |
INTERSPEECH | 3 |
| 2006 | Improved tone modeling for Mandarin broadcast news speech recognitionabstractTone has a crucial role in Mandarin speech in distinguishing ambiguous words. Most state-of-the-art Mandarin automatic speech recognition systems adopt embedded tone modeling, where tonal acoustic units are used and F0 features are appended to the spectral feature vector. In this paper, we combine the embedded aproach (using improved F0 smoothing) with explicit tone modeling in rescoring the output lattices. Oracle experiments indicate 32% relative improvement can be achieved by rescoring with perfect tone information. Recognition experiments on Mandarin broad-cast news show that, even with an accuracy of only 70%, the explicit tone classifier offers complementary knowledge and improves performance significantly. Through the combination of tone modeling techniques, the character error rate on the CTV test set can be improved from 13.0% to 11.5%. Man-Hung Siu, Mei-Yuh Hwang, Mari Ostendorf, Tan Lee |
INTERSPEECH | 5 |
| 2006 | Towards automatic parameter extraction of command-response model for Cantonese
Raymond W. M. Ng, Tan Lee, Wentao Gu |
INTERSPEECH | 2 |
| 2005 | Static and Dynamic Spectral Features: Their Noise Robustness and Optimal Weights for ASRabstractIn this paper, we investigate the relative noise robustness between dynamic and static spectral features, by using two speaker independent continuous digit databases in English (Aurora2) and Cantonese (CUDigit). It is found that the dynamic cepstrum is more robust to additive noise than its static counterpart. The results are consistent across different types of noise and under various SNRs. Optimal exponential weights for exploiting unequal noise robustness of the two features are discriminatively trained in a development set. When tested under various noise conditions, the optimal weights yielded relative word error rate reductions of 36.6% and 41.9% for Aurora2 and CUDigit, respectively. The proposed weighting is attractive for many ASR applications in noise because: (1) no noise estimation for feature compensation; (2) no adaptation of clean HMMs to a noisy environment; and (3) only a trivial change in the decoding process by weighting log likelihoods of static and dynamic components separately. Frank K. Soong, Tan Lee |
ICASSP (1) | 3 |
| 2005 | Development of a Cantonese-English code-mixing speech corpus
Joyce Y. C. Chan, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2004 | Tone information as a confidence measure for improving Cantonese LVCSR
Yao Qian, Tan Lee, Frank K. Soong |
INTERSPEECH | 2 |
| 2004 | Time -frequency analysis of vocal source signal for speaker recognitionabstractThis paper investigates the importance of spectro-temporal characteristics of the source excitation signal for speaker recognition. We propose an effective feature extraction technique for obtaining essential time-frequency information from the linear prediction (LP) residual signal, which are closely related to the glottal excitation of individual speaker. With pitch synchro-nous analysis, wavelet transform is applied to every two pitch cycles of the LP residual signal to generate a new feature vector, called Wavelet Octave Coefficients of Residues (WOCOR), which provides additional speaker discriminative power to the commonly used linear predictive Cepstral coefficients (LPCC). Experimental evaluation over a Cantonese speaker recognition corpus demonstrates the effectiveness of WOCOR for speaker recognition. Recognition tests with WOCOR and LPCC outperforms the conventional methods of using Mel Frequency Cepstral Coefficients (MFCC). 1. Nengheng Zheng, Pak-Chung Ching, Tan Lee |
INTERSPEECH | 3 |
| 2004 | Explicit duration modeling for Cantonese connected-digit recognition
Tan Lee |
INTERSPEECH | 2 |
| 2004 | Analysis and modeling of F0 contours for cantonese text-to-speechabstractFor the generation of highly natural synthetic speech, the control of prosody is of primary importance. The fundamental frequency (F0) is one of the most important components of speech prosody. This research investigates the variation of F0 in continuous Cantonese speech, with the goal of establishing an effective mechanism of prosody control in Cantonese text-to-speech (TTS) applications. Cantonese is a commonly used Chinese dialect that is well known for being rich in tones. This article describes a simple yet effective approach to the analysis and modeling of F0. The surface F0 contour of a continuous Cantonese utterance is considered to be the combination of a global component--phrase-level intonation curve, and local components--syllable-level tone contoursA novel method of F0 normalization is proposed to separate the local components from the global one. As a result, the variation in tone contours is greatly reduced. Statistical analysis is performed for the phrase curves and context-dependent tone contours that are extracted from a large corpus of 1,200 utterances. Specifically, the analysis is focused on co-articulated tone contours for disyllabic words, cross-word contours, and phrase-initial tone contours. Based on the results of the analysis, a template-based model for F0 generation is established and integrated with a Cantonese TTS system. Subjective listening tests show that the proposed model significantly improves the naturalness of the output speech. Tan Lee, Yao Qian |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2003 | Modeling Cantonese pronunciation variation by acoustic model refinementabstractPronunciation variations can be roughly classified into two types: a phone change or a sound change [1][2]. A phone change happens when a canonical phone is produced as a different phone. Such a change can be modeled by converting the baseform (standard) phone to a surfaceform (actual) phone. A sound change happens at a lower, phonetic or subphonetic level within a phone and it cannot be modeled well by either the baseform or the surfaceform phone alone. We propose here to refine the acoustic models to cope with sound changes by (1) sharing the Gaussian mixture components of HMM states in the baseform and the surfaceform models; (2) adapting the mixture components of the baseform models towards those of the surfaceform models; (3) selectively reconstructing new acoustic models through sharing or adapting. The proposed pronunciation modeling algorithms are generic and can, in principle, be applied to different languages. Specifically, they were tested in a Cantonese speech recognition database. Relative word error rate reductions of 5.45%, 2.53%, and 3.04 % have been achieved using the three approaches, respectively. 1. Patgi Kam, Tan Lee, Frank K. Soong |
INTERSPEECH | 2 |
| 2003 | Overlapped di-tone modeling for tone recognition in continuous Cantonese speechabstractThis paper presents a novel approach to tone recognition in continuous Cantonese speech based on overlapped di-tone Gaussian mixture models (ODGMM). The ODGMM is designed with special consideration on the fact that Cantonese tone identification relies more on the relative pitch level than on the pitch contour. A di-tone unit covers a group of two consecutive tone occurrences. The tone sequence carried by a Cantonese utterance can be considered as the connection of such di-tone units. Adjacent di-tone units overlap with each other by exactly one tone. For each di-tone unit, a GMM is trained with a 10-dimensional feature vector that characterizes the F0 movement within the unit. In particular, the di-tone models capture the relative deviation between the F0 levels of the two tones. Viterbi decoding algorithm is adopted to search for the optimal tone sequence, under the phonological constraints on syllable-tone combination. Experimental results show the ODGMM approach significantly outperforms the previously proposed methods for tone recognition in continuous Cantonese speech. Yao Qian, Tan Lee |
INTERSPEECH | 2 |
| 2002 | Unsupervised n-best based model adaptation using model-level confidence measuresabstractThis paper presents a study on using confidence measures for unsupervised N-Best based adaptation of hidden Markov model (HMM) parameters. Confidence measures have been widely used for the detection of speech recognition errors. They are also useful in selecting and/or screening data for unsupervised adaptation of HMM. In this paper, a model-level confidence measure is proposed for model adaptation with the Maximum Likelihood Linear Regression (MLLR) technique. The modellevel confidence measure provides a finer selection of adaptation data than the word or utterance level measures. The proposed confidence measure is derived from the N-best hypotheses. The computation involves not only the recognized models but also other models that are easily confused with them. Experimental results show the proposed confidence measure improves the effectiveness of unsupervised model adaptation. The relative improvement in word error rate is up to 9.75%. Ka-Yan Kwan, Tan Lee |
INTERSPEECH | 2 |
| 2002 | Modeling tones in continuous Cantonese speechabstractCantonese is a major Chinese dialect with a complicated tone system. This research focuses on quantitative modeling of Cantonese tones. It uses Stem-ML, a language-independent framework for quantitative intonation modeling and generation. A set of F 0 prediction models are built, and trained on acoustic data. The prediction error is about 11 Hz or 1 semitone. The resulting optimal model parameters are analyzed in accordance with linguistic knowledge. Key observations include: (1) There is no obvious advantage to model the entering tones separately. They can be considered as simply truncated versions of the nonentering tones; (2) Cantonese appears to have a declining phrase intonation; (3) Tones at initial positions of a phrase or a sentence tend to have a greater prosodic strength than those at the final positions; (4) Content words are stronger than function words; (5) Long words are stronger than short words. Tan Lee, Greg Kochanski, Chilin Shih |
INTERSPEECH | 1 |
| 2002 | Spoken language resources for Cantonese speech processing
Tan Lee, Wai Kit Lo, Pak-Chung Ching, Helen M. Meng |
Speech Commun. | 1 |
| 2002 | Using tone information in Cantonese continuous speech recognitionabstractIn Chinese languages, tones carry important information at various linguistic levels. This research is based on the belief that tone information, if acquired accurately and utilized effectively, contributes to the automatic speech recognition of Chinese. In particular, we focus on the Cantonese dialect, which is spoken by tens of millions of people in Southern China and Hong Kong. Cantonese is well known for its complicated tone system, which makes automatic tone recognition very difficult. This article describes an effective approach to explicit tone recognition of Cantonese in continuously spoken utterances. Tone feature vectors are derived, on a short-time basis, to characterize the syllable-wide patterns of F0 (fundamental frequency) and energy movements. A moving-window normalization technique is proposed to reduce the tone-irrelevant fluctuation of F0 and energy features. Hidden Markov models are employed for context-dependent acoustic modeling of different tones. A tone recognition accuracy of 66.4% has been achieved in the speaker-independent case. The recognized tone patterns are then utilized to assist Cantonese large-vocabulary continuous speech recognition (LVCSR) via a lattice expansion approach. Experimental results show that reliable tone information helps to improve the overall performance of LVCSR. Tan Lee, Wai H. Lau, Yiu Wing Wong, Pak-Chung Ching |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2001 | Cantonese text-to-speech synthesis using sub-syllable unitsabstractThis paper describes our recent investigation on the use of both intra-syllable and cross-syllable acoustic units for Cantonese text-to-speech synthesis. In our previous work, isolated monosyllable units were used for concatenative speech synthesis of Cantonese. The synthetic speech was considered to be unnatural in such a way that there was an obvious lack of perceptual continuity. The proposed system adopts an acoustic inventory that covers all legitimate intrasyllable and cross-syllable acoustic units. Synthetic speech produced via concatenation of such sub-syllable units better captures the pertinent transitory effects that are crucial to perceived naturalness. Different strategies are used to concatenate speech segments with different acoustic-phonetic properties. Subjective listening test shows a noticeable performance improvement that is accounted for mainly by smoother transition between sonorant segments. Ka Man Law, Tan Lee, Wai H. Lau |
INTERSPEECH | 2 |
| 2001 | ISIS: a learning system with combined interaction and delegation dialogsabstractThis paper presents a progress update of our ISIS 1 trilingual spoken dialog system. As described in [8], this is a conversational system for the stocks domain, and supports interactions in the languages of our region – English and two dialects of Chinese (Mandarin and Cantonese). ISIS provides a system test-bed for our initial explorations with the CORBA architecture, and delegation to KQML (Knowledge Query and Manipulation Language) agents. CORBA offers the advantages of interoperability, scalability and location transparency in client/server systems development. Users can delegate tasks to software agents to help monitor information (e.g. a drop in the price of a pre-specified stock), and generate user alert messages. Our current work presents new research directions in the context of ISIS: (i) automatic incorporation of newly listed stocks into our system’s knowledge base; (ii) switching between on-line interaction and off-line delegation in a single dialog thread. We will also report on enhancements in the system’s architecture and features (e.g. automatic end-point detection). delegation in a single dialog thread. We will also report on enhancements in the system’s architecture and features (e.g. automatic end-point detection). 2. System Architecture Previous work in the development of software infrastructures for dialog systems include [11] 2 [2]. Over the past year, we have continued to explore the development of a spoken dialog system based on the CORBA architecture. This middleware resides in between the operating system and the application layer 1. Helen M. Meng, Shuk Fong Chan, Yee Fong Wong, Cheong Chat Chan, Yiu Wing Wong, Tien Ying Fung, Wai Ching Tsui, Ke Chen 0001, Ting-Yao Wu, Tan Lee, Wing Nin Choi, Pak-Chung Ching, Huisheng Chi |
INTERSPEECH | 12 |
| 2000 | Acoustic modeling for Chinese speech recognition: a comparative study of Mandarin and CantoneseabstractThis paper presents a comparative study on automatic speech recognition for two different Chinese dialects, namely Mandarin and Cantonese. It focuses on decision-tree based context-dependent acoustic modeling for large-vocabulary continuous speech recognition. Extensive phonological and phonetic knowledge are incorporated to design questions concerning the left and right context of sub-syllable units, namely INITIALs and FINALs. This results in a set of class-triphone models for each dialect. Syllable recognition accuracy of 81.7% and 75.5% are attained for Mandarin and Cantonese respectively. Such a performance gap is accountable by various linguistic and practical reasons, including: 1) phonological and phonetic discrepancies between the two dialects; 2) design of training databases; and 3) design of phonetic questions in decision-tree clustering. Tan Lee, Yiu Wing Wong, Bo Xu 0002, Pak-Chung Ching, Taiyi Huang 0001 |
ICASSP | 2 |
| 2000 | Lexical tree decoding with a class-based language model for Chinese speech recognition
Wing Nin Choi, Yiu Wing Wong, Tan Lee, Pak-Chung Ching |
INTERSPEECH | 3 |
| 2000 | Incorporating tone information into Cantonese large-vocabulary continuous speech recognition
Wai H. Lau, Tan Lee, Yiu Wing Wong, Pak-Chung Ching |
INTERSPEECH | 2 |
| 2000 | Using cross-syllable units for Cantonese speech synthesis
Ka Man Law, Tan Lee |
INTERSPEECH | 2 |
| 2000 | ISIS: A multilingual spoken dialog system developed with CORBA and KQML agentsabstractISIS, which abbreviates Intelligent Speech for Information Systems, is a trilingual spoken dialog system (SDS) for the financial domain. It handles two dialects of Chinese (Cantonese and Putonghua), as well as English the predominant languages in our region. The system supports spoken language queries regarding stock market information and simulated personal portfolios. Real-time information is retrieved directly from a dedicated Reuters satellite feed. ISIS provides a system test-bed for our work in multilingual speech recognition and generation, speaker authentication, language understanding and dialog modeling. Furthermore, ISIS supports our initial explorations in: (i) CORBA's interoperability and scalability for SDS development; in conjunction with (ii) asynchronous human-computer interaction by delegation to KQML software agents... Helen M. Meng, Shuk Fong Chan, Yee Fong Wong, Tien Ying Fung, Wai Ching Tsui, Tin Hang Lo, Cheong Chat Chan, Ke Chen 0001, Ting-Yao Wu, Tan Lee, Wing Nin Choi, Yiu Wing Wong, Pak-Chung Ching, Huisheng Chi |
INTERSPEECH | 12 |
| 1999 | Two-dimensional multi-resolution analysis of speech signals and its application to speech recognitionabstractThis paper describes a novel approach of using multi-resolution analysis (MRA) for automatic speech recognition. Two-dimensional MRA is applied to the short-time log spectrum of speech signal to extract the slowly varying spectral envelope that contains the most important articulatory and phonetic information. After passing through a standard cepstral analysis process, the MRA features are used for speech recognition in the same way as conventional short-time features like MFCCs, PLPs, etc. Preliminary experiments on both clean connected speech and noisy telephone conversation speech show that the use of MRA cepstra results in a significant reduction in insertion error when compared with MFCCs. Chun-Ping Chan, Yiu Wing Wong, Tan Lee, Pak-Chung Ching |
ICASSP | 3 |
| 1999 | Micro-prosodic control in cantonese text-to-speech synthesis
Tan Lee, Helen M. Meng, Wai H. Lau, Wai Kit Lo, Pak-Chung Ching |
EUROSPEECH | 1 |
| 1999 | Acoustic modeling and language modeling for cantonese LVCSRabstractThis paper describes our recent work on the development of a large-vocabulary, speaker-independent continuous speech recognition system for Cantonese (a major Chinese dialect). Both acoustic modeling and language modeling are being addressed. For acoustic modeling, we focus on right-context-dependent sub-syllable units. Tying of HMM at model as well as state level is applied based on phonetic knowledge and the decision-tree approach. Statistical language model is built from large amount of newspaper text. The overall recognition accuracy for syllable and Chinese character are 81.83% and 68.94% respectively. Keywords: LVCSR, Cantonese speech recognition, acoustic modeling, language modeling 1 INTRODUCTION Cantonese is one of the major Chinese dialects spoken by tens of millions of people in Hong Kong, Southern China as well as many overseas Chinese communities. With the great advancement of computer and information technology, there is an ever-increasing demand of largevocabulary con... Yiu Wing Wong, Ka-Fai Chow, Wai H. Lau, Wai Kit Lo, Tan Lee, Pak-Chung Ching |
EUROSPEECH | 5 |
| 1999 | Cantonese syllable recognition using neural networksabstractThis work describes a novel neural network based speech recognition system for isolated Cantonese syllables. Since Cantonese is a monosyllabic and tonal language, the recognition system is composed of two major components, namely the tone recognizer and the base syllable recognizer. The tone recognizer adopts the architecture of multilayer perceptron in which each output neuron represents a particular tone. The base syllable recognizer consists of a large number of independently trained recurrent networks, each representing a designated Cantonese syllable. An integrated recognition algorithm is developed to give the ultimate recognition results based on N-best outputs of the two subrecognizers. To demonstrate the effectiveness of the proposed methods, a speaker-dependent recognition system has been built with the vocabulary expanding progressively from 10 syllables to 200 syllables. In the case of 200 syllables, a top-1 recognition accuracy of 81.8% has been attained whilst the top-3 accuracy is 95.28. Tan Lee, Pak-Chung Ching |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | Context-dependent duration modelling for continuous speech recognition
Tan Lee, Rolf Carlson, Björn Granström |
ICSLP | 1 |
| 1998 | Isolated word recognition using modular recurrent neural networks
Tan Lee, Pak-Chung Ching, Lai-Wan Chan |
Pattern Recognit. | 1 |
| 1997 | Development of a large vocabulary speech database for CantoneseabstractThis paper describes work on developing a large vocabulary speech database for Cantonese. As a major Chinese dialect, Cantonese is spoken by tens of millions of people in Southern China and Hong Kong. It is very different from Mandarin or Putonghua in phonology, phonetics, vocabulary and grammatical structure. A speech database specially designed for Cantonese is urgently needed for the design, implementation and performance evaluation of various speech recognition systems. The proposed database contains a large number of speech utterances which include isolated syllables, polysyllabic words and phonetically rich sentences. It covers most of the intra-syllable and inter-syllable acoustic variations. Pak-Chung Ching, Ka-Fai Chow, Tan Lee, Alfred Ying Pang Ng, Lai-Wan Chan |
ICASSP | 3 |
| 1997 | A neural network based speech recognition system for isolated Cantonese syllablesabstractThis paper describes a novel design of neural network based speech recognition system for isolated Cantonese syllables. Since Cantonese is a monosyllabic and tonal language, the recognition system consists of a tone recognizer and a base syllable recognizer. The tone recognizer adopts the architecture of a multi-layer perceptron in which each output neuron represents a particular tone. The syllable recognizer contains a large number of independently trained recurrent networks, each representing a designated Cantonese syllable. Such a modular structure provides greater flexibility to expand the system vocabulary progressively by adding new syllable models. To demonstrate the effectiveness of the proposed method, a speaker-dependent recognition system has been built with the vocabulary growing from 40 syllables to 200 syllables. In the case of 200 syllables, a top-1 recognition accuracy of 81.8% has been attained and the top-3 accuracy is 95.2%. Tan Lee, Pak-Chung Ching |
ICASSP | 1 |
| 1996 | On improving discrimination capability of an RNN based recognizer
Tan Lee, Pak-Chung Ching |
ICSLP | 1 |
| 1995 | Recurrent neural networks for speech modeling and speech recognitionabstractDescribes a new method of utilizing recurrent neural networks (RNNs) for speech modeling and speech recognition. For each particular speech unit, a fully connected recurrent neural network is built such that the static and dynamic speech characteristics are represented simultaneously by a specific temporal pattern of neuron activation states. By using the temporal RNN output, an input utterance can be represented as a number of stationary speech segments, which may be related to the basic phonetic components of the speech unit. An efficient self-supervised training algorithm has been developed for the RNN speech model. The segmentation for input utterances and the statistical modeling for individual phonetic segments are performed interactively in this training process. Some experimental results are used to demonstrate how the proposed RNN speech model can be used effectively for automatic recognition of isolated speech utterances. Tan Lee, Pak-Chung Ching, Lai-Wan Chan |
ICASSP | 1 |
| 1995 | An RNN based speech recognition system with discriminative trainingabstractIn our previous work #1#, a novel method of utilizing a set of fully connected recurrent neural networks #RNNs# for speech modeling has been proposed. Despite the e#ectiveness of the RNN model in characterizing individual speech units, the system performs less satisfactorily for speech recognition due to poor discrimination between models. In this paper, an e#cient discriminative training procedure is developed for the RNN based recognition system. By using discriminative training, each RNN speech model is adjusted to reduce its distance from the designated speech unit while increase distances from the others. In addition, a duration-screening process is introduced to enhance the discriminating power of the recognition system. Speaker-dependent recognition experiments have been carried out for 1# 11 isolated Cantonese digits, 2# 58 very confusing Cantonese CV syllables, and 3# 20 English isolated words. The recognition rates attained are 90.9#, 86.7# and 93.5# respectively. I. Int... Tan Lee, Pak-Chung Ching, Lai-Wan Chan |
EUROSPEECH | 1 |
| 1995 | Tone recognition of isolated Cantonese syllablesabstractTone identification is essential for the recognition of the Chinese language, specifically far Cantonese which is well known for being very rich in tones. The paper presents an efficient method for tone recognition of isolated Cantonese syllables. Suprasegmental feature parameters are extracted from the voiced portion of a monosyllabic utterance and a three-layer feedforward neural network is used to classify these feature vectors. Using a phonologically complete vocabulary of 234 distinct syllables, the recognition accuracy for single-speaker and multispeaker is given by 89.0% and 87.6% respectively.> Tan Lee, Pak-Chung Ching, Lai-Wan Chan, Y. H. Cheng, Brian Kan-Wing Mak |
IEEE Trans. Speech Audio Process. | 1 |
| 1992 | A Node Pruning Algorithm for Backpropagation NetworksabstractBackpropagation (BP) networks are a class of artificial neural network model that has been widely used in many areas of interest. One difficulty in adopting this model is the need to predetermine a suitable network size, particularly, the number of hidden nodes. A common approach is to start with an oversized network and then the size is reduced by eliminating unnecessary nodes and links. In this paper, starting from the viewpoint of pattern classification theory, a study on the characteristics of hidden nodes in oversized networks is reported. The study shows that four categories of excessive hidden nodes can be found in an oversized network. A new node pruning algorithm to attain appropriate size BP networks is then proposed by detecting and removing those excessive nodes. Moreover, the algorithm is extended to cater for the larger class of problems, i.e. real-to-real mapping. Unlike previous works, the concept of “excessiveness” advocated here has strong indications whether a node can be removed without impairing the performance of original network and hence the proposed algorithm is useful in obtaining a network that is optimized with both the network size and performance. The effectiveness of the proposed algorithm has been demonstrated through the N-bit parity problems and the experiment in predicting the chaotic time series. Korris Fu-Lai Chung, Tan Lee |
Int. J. Neural Syst. | 2 |