Yu Ting Yeung

dblp:47/9236 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
11since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 9 since 2021
YearPublicationVenuePosition
2022 A Time Domain Progressive Learning Approach with SNR Constriction for Single-Channel Speech Enhancement and Recognition
abstract
Single-channel speech enhancement for automatic speech recognition (ASR) has been widely studied. However, most speech enhancement methods conduct over suppression and introduce distortion, which limits performance gains or even deteriorates the back-end performance. The key to solving this problem is preserving the integrity of speech while suppressing the background noises. There-fore, we propose a time domain progressive learning (TDPL) approach for speech enhancement and ASR. TDPL model consists of encoder, progressive enhancer and decoder. Both SNR-increased intermediate target with less speech distortion and clean target with better listening quality/intelligibility are learned, which are provided for ASR pre-processing and speech communication, respectively. Additionally, we also present an SNR constriction loss that is fit for TDPL to further improve ASR performance. We evaluate the proposed methods on CHiME-4 real evaluation set. The results show that the TDPL method significantly outperforms time domain speech enhancement methods and frequency domain progressive learning methods in ASR task, and the intermediate output of TDPL achieves a 36.3% relative word error rate reduction with a powerful ASR back-end without retraining. Moreover, the estimated clean output achieves certain improvement on CHiME-4 simulation evaluation set in terms of PESQ and STOI measures.
Zhaoxu Nian, Jun Du 0002, Yu Ting Yeung, Renyu Wang
ICASSP3
2022 SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training
Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang 0002, Qun Liu 0001
ICLR3
2022 reducing multilingual context confusion for end-to-end code-switching automatic speech recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Jianhua Tao 0001, Yu Ting Yeung, Liqun Deng
INTERSPEECH5
2022 Streamable Speech Representation Disentanglement and Multi-Level Prosody Modeling for Live One-Shot Voice Conversion
Haoquan Yang, Liqun Deng, Yu Ting Yeung, Nianzu Zheng
INTERSPEECH3
2022 Online Speaker Diarization with Core Samples Selection
Yanyan Yue, Jun Du 0002, Maokui He, Yu Ting Yeung, Renyu Wang
INTERSPEECH4
2022 CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis
abstract
Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762.
Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
INTERSPEECH4
2021 EditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional Fusion
abstract
This paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without causing audible degradation in speech quality and naturalness. The EditSpeech system is developed upon a neural text-to-speech (NTTS) synthesis framework. Partial inference and bidirectional fusion are proposed to effectively incorporate the contextual information related to the edited region and achieve smooth transition at both left and right boundaries. Distortion introduced to the unmodified parts of the utterance is alleviated. The EditSpeech system is developed and evaluated on English and Chinese in multi-speaker scenarios. Objective and subjective evaluation demonstrate that EditSpeech outperforms a few baseline systems in terms of low spectral distortion and preferred speech quality. Audio samples are available online for demonstration11https://daxintan-cuhk.github.io/EditSpeech/.
Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang 0002, Xiao Chen 0012, Tan Lee
ASRU3
2021 Fcl-Taco2: Towards Fast, Controllable and Lightweight Text-to-Speech Synthesis
abstract
Sequence-to-sequence (seq2seq) learning has greatly improved text-to-speech (TTS) synthesis performance, but effective implementation on resource-restricted devices remains challenging as seq2seq models are usually computationally expensive and memory intensive. To achieve fast inference speed and small model size while maintain high-quality speech, we propose FCL-taco2, a Fast, Controllable and Lightweight (FCL) TTS model based on Tacotron2. FCL-taco2 adopts a novel semi-autoregressive (SAR) mode for phoneme level based parallel mel-spectrograms generation conditioned on prosody features, leading to faster inference speed and higher prosody controllability than Tacotron2. Besides, knowledge distillation (KD) is leveraged to compress a relatively large FCL-taco2 model to its small version with minor loss of speech quality. Experimental results on English (EN) and Chinese (CN) datasets show that the small version of FCL-taco2 achieves comparable performance with Tacotron2 in terms of speech quality, while it has a 4.8× smaller footprint with 17.7× and 18.5× faster inference speeds on average for EN and CN experiments respectively. Besides, execution on mobile devices shows that the proposed model can achieve faster than real-time speech synthesis. Our code and audio samples are released1.
Disong Wang, Liqun Deng, Yang Zhang 0025, Nianzu Zheng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng
ICASSP5
2021 VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-Shot Voice Conversion
abstract
One-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement.Existing work generally ignores the correlation between different speech representations during training, which causes leakage of content information into the speaker representation and thus degrades VC performance.To alleviate this issue, we employ vector quantization (VQ) for content encoding and introduce mutual information (MI) as the correlation metric during training, to achieve proper disentanglement of content, speaker and pitch representations, by reducing their inter-dependencies in an unsupervised manner.Experimental results reflect the superiority of the proposed method in learning effective disentangled speech representations for retaining source linguistic content and intonation variations, while capturing target speaker characteristics.In doing so, the proposed approach achieves higher speech naturalness and speaker similarity than current state-of-the-art one-shot VC systems.Our code, pre-trained models and demo are available at https://github.com/Wendison/VQMIVC.
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng
Interspeech3
2021 Unsupervised Domain Adaptation for Dysarthric Speech Detection via Domain Adversarial Training and Mutual Information Minimization
abstract
Dysarthric speech detection (DSD) systems aim to detect characteristics of the neuromotor disorder from speech.Such systems are particularly susceptible to domain mismatch where the training and testing data come from the source and target domains respectively, but the two domains may differ in terms of speech stimuli, disease etiology, etc.It is hard to acquire labelled data in the target domain, due to high costs of annotating sizeable datasets.This paper makes a first attempt to formulate cross-domain DSD as an unsupervised domain adaptation (UDA) problem.We use labelled source-domain data and unlabelled target-domain data, and propose a multi-task learning strategy, including dysarthria presence classification (DPC), domain adversarial training (DAT) and mutual information minimization (MIM), which aim to learn dysarthriadiscriminative and domain-invariant biomarker embeddings.Specifically, DPC helps biomarker embeddings capture critical indicators of dysarthria; DAT forces biomarker embeddings to be indistinguishable in source and target domains; and MIM further reduces the correlation between biomarker embeddings and domain-related cues.By treating the UASPEECH and TORGO corpora respectively as the source and target domains, experiments show that the incorporation of UDA attains absolute increases of 22.2% and 20.0% respectively in utterancelevel weighted average recall and speaker-level accuracy.
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 0012, Xunying Liu, Helen M. Meng
Interspeech3
2021 Energy-Friendly Keyword Spotting System Using Add-Based Convolution
Wenchao Hu, Yu Ting Yeung, Xiao Chen 0012
Interspeech3
2020 Conv-Transformer Transducer: Low Latency, Low Frame Rate, Streamable End-to-End Speech Recognition
abstract
Transformer has achieved competitive performance against state-of-the-art end-to-end models in automatic speech recognition (ASR), and requires significantly less training time than RNN-based models.The original Transformer, with encoderdecoder architecture, is only suitable for offline ASR.It relies on an attention mechanism to learn alignments, and encodes input audio bidirectionally.The high computation cost of Transformer decoding also limits its use in production streaming systems.To make Transformer suitable for streaming ASR, we explore Transducer framework as a streamable way to learn alignments.For audio encoding, we apply unidirectional Transformer with interleaved convolution layers.The interleaved convolution layers are used for modeling future context which is important to performance.To reduce computation cost, we gradually downsample acoustic input, also with the interleaved convolution layers.Moreover, we limit the length of history context in self-attention to maintain constant computation cost for each decoding step.We show that this architecture, named Conv-Transformer Transducer, achieves competitive performance on LibriSpeech dataset (3.6% WER on test-clean) without external language models.The performance is comparable to previously published streamable Transformer Transducer and strong hybrid streaming ASR systems, and is achieved with smaller look-ahead window (140 ms), fewer parameters and lower frame rate.
Wenyong Huang, Wenchao Hu, Yu Ting Yeung, Xiao Chen 0012
INTERSPEECH3
2016 Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speech
abstract
This paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the types and amount of speech materials used for acoustical assessment were very limited. With the ASR technology, we are able to perform acoustical and linguistic analyses with a large amount of natural speech from impaired speakers. The present study is focused on Cantonese, which is a major Chinese dialect. Two representative disorders of speech production are investigated: dysphonia and aphasia. ASR experiments are carried out with continuous and spontaneous speech utterances from Cantonese-speaking patients. The results confirm the feasibility and potential of using natural speech for acoustical assessment of voice and speech disorders, and reveal the challenging issues in acoustic modeling and language modeling of pathological speech.
Tan Lee, Yuanyuan Liu 0002, Pei-Wen Huang, Jen-Tzung Chien, Wang-Kong Lam, Yu Ting Yeung, Thomas K. T. Law, Kathy Yuet-Sheung Lee, Anthony Pak-Hin Kong, Sam-Po Law
ICASSP6
2016 Exploring articulatory characteristics of Cantonese dysarthric speech using distinctive features
abstract
Dysarthria is a kind of motor speech disorder due to neurological deficits. Understanding the articulatory problems of dysarthric speakers may help to design suitable intervention strategies to improve their speech intelligibility. We have developed an automatic articulatory characteristics analysis framework based on a distinctive feature (DF) recognition. We recruited 16 Cantonese dysarthric subjects with spinocerebellar ataxia (SCA) or cerebral palsy (CP) to support our research. To the best of our knowledge, this is among the first efforts in collecting and automatically analyzing Cantonese dysarthric speech. The framework shows a close Pearson correlation to manual annotation of the subjects in most DFs and also in the average DF error rates. It indicates a potential way to describe articulatory characteristics of dysarthric speech and automatically assess it.
Ka-Ho Wong, Wing Sum Yeung, Yu Ting Yeung, Helen M. Meng
ICASSP3
2016 Predicting Severity of Voice Disorder from DNN-HMM Acoustic Posteriors
Tan Lee, Yuanyuan Liu 0002, Yu Ting Yeung, Thomas K. T. Law, Kathy Yuet-Sheung Lee
INTERSPEECH3
2015 Modeling temporal dependency for robust estimation of LP model parameters in speech enhancement
Chun Hoy Wong, Tan Lee, Yu Ting Yeung, Pak-Chung Ching
INTERSPEECH3
2015 Development of a Cantonese dysarthric speech corpus
abstract
Dysarthria is a neurogenic communication disorder affecting speech production. Significant differences in phonemic inventories and phonological patterns across the world’s languages render generalization of disordered speech patterns from one language (e.g, English) to another (e.g., Cantonese) difficult. Capitalizing on existing methods in developing Englishlanguage dysarthric speech corpora, we develop a Cantonese corpus in order to investigate articulatory and prosodic characteristics of Cantonese dysarthric speech, focusing on speaking rate and pitch and loudness control. Currently, we have collected 7.5 and 2.5 hours of speech data from 11 dysarthric subjects and 5 control speakers respectively. Our preliminary analysis reveals the characteristics of Cantonese dysarthric speech are consistent with general properties of motor speech disorders found in other languages.
Ka-Ho Wong, Yu Ting Yeung, Edwin H. Y. Chan, Patrick C. M. Wong, Gina-Anne Levow, Helen M. Meng
INTERSPEECH2
2015 Improving automatic forced alignment for dysarthric speech transcription
abstract
Dysarthria is a motor speech disorder due to neurologic deficits. The impaired movement of muscles for speech production leads to disordered speech where utterances have prolonged pause in-tervals, slow speaking rates, poor articulation of phonemes, syl-lable deletions, etc. These present challenges towards the use of speech technologies for automatic processing of dysarthric speech data. In order to address these challenges, this work be-gins by addressing the performance degradation faced in forced alignment. We perform initial alignments to locate long pauses in dysarthric speech and make use of the pause intervals as an-chor points. We apply speech recognition for word lattice out-puts for recovering the time-stamps of the words in disordered or incomplete pronunciations. By verifying the initial align-ments with word lattices, we obtain the reliably aligned seg-ments. These segments provide constraints for new alignment grammars, that can improve alignment and transcription quality. We have applied the proposed strategy to the TORGO corpus and obtained improved alignments for most dysarthric speech data, while maintaining good alignments for non-dysarthric speech data. Index Terms: automatic forced alignment, speech recognition, dysarthric speech, word lattices
Yu Ting Yeung, Ka-Ho Wong, Helen M. Meng
INTERSPEECH1
2015 Supervised Single-Microphone Multi-Talker Speech Separation with Conditional Random Fields
abstract
We apply conditional random field (CRF) for single-microphone speech separation in a supervised learning scenario. We train the parameters with mixture data in which the sources are competing with the same average signal power. Compared with factorial hidden Markov model (HMM) baselines, the CRF settings require fewer training mixture data to improve objective speech quality measures and speech recognition accuracy of the reconstructed sources, when mixing ratios of training and testing mixture data are matched. The CRF settings also handle minor mixing ratio mismatch after adjusting the gain factors of the sources with non-linear mappings inspired from the mixture-maximization model. When the mixing ratio mismatch further increases such that the speech mixture is dominated by only one source, factorial HMM finally catches up with and performs better than the CRF settings due to improved model accuracy. We also develop a convex statistical inference simplification based on linear-chain CRFs. The simplification achieves the same performance level as the original CRF settings after integrating additional observations.
Yu Ting Yeung, Tan Lee, Cheung-Chi Leung
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Large-margin conditional random fields for single-microphone speech separation
Yu Ting Yeung, Tan Lee, Cheung-Chi Leung
INTERSPEECH1
2013 Evaluation of pitch estimation algorithms on separated speech
abstract
To post-process outputs of speech separation systems with harmonic enhancement, it is normally required to estimate the fundamental frequency. This paper evaluates the performance of a few representative robust pitch estimation algorithms on speech reconstructed from two-speaker mixture signals. The separation outputs obtained by two state-of-the-art single-channel separation algorithms are used for the evaluation. A recently proposed sparsity-based pitch estimation method is applied to the separated speech and a new pitch tracking algorithm is proposed. Experimental results show that on the separated speech the proposed method consistently surpasses the others with significantly low gross error rate, which is similar to the gross error rates of the other methods on clean speech.
Feng Huang 0002, Yu Ting Yeung, Tan Lee
ICASSP2
2013 Using dynamic conditional random field on single-microphone speech separation
abstract
The use of dynamic conditional random field (DCRF) for model-based single-microphone speech separation is investigated. The speech sources are represented by acoustic state sequences from speaker-dependent acoustic models. The posterior probabilities of the source acoustic states given a speech mixture are inferred with a maximum entropy probability distribution which is represented by DCRF. The posterior probabilities are needed for minimum mean-square error estimation of the speech sources. Loopy belief propagation is applied for the inference. Averaged stochastic gradient descent and limited-memory BFGS are compared for parameter estimation. With the log-magnitude spectrum of the speech mixture as input observation, the proposed method achieves better separation performance in terms of Blind Source Separation Metrics (SDR, SAR, SIR) and PESQ than a factorial hidden Markov model baseline system in our experiments.
Yu Ting Yeung, Tan Lee, Cheung-Chi Leung
ICASSP1
2012 Integrating multiple observations for model-based single-microphone speech separation with conditional random fields
abstract
A single-microphone speech separation framework based on conditional random fields (CRFs) is proposed in this paper. Unlike factorial HMM, CRF does not have the conditional independence assumption on observations, thus different types of observations from the speech mixture can be integrated into the models through feature functions. Similar to factorial HMM, there is the statistical independence assumption on sources. Under this assumption, the two-source single-microphone speech separation problem can be expressed by two independent linear-chain CRFs. The separation problem becomes two pattern recognition problems, with respect to CRF models of the two sources. Experimental results show that by integrating initial separation outputs from factorial HMM with log power spectrum, fundamental frequency and speaker likelihoods of the mixture, CRF separation framework consistently improves the results from factorial HMM in terms of SNR, segmental SNR and PESQ.
Yu Ting Yeung, Tan Lee, Cheung-Chi Leung
ICASSP1
2008 Language modeling for speech recognition of spoken Cantonese
Yu Ting Yeung, Houwei Cao, Nengheng Zheng, Tan Lee, Pak-Chung Ching
INTERSPEECH1
2008 Prosody for Mandarin speech recognition: a comparative study of read and spontaneous speech
Yu Ting Yeung, Yao Qian, Tan Lee, Frank K. Soong
INTERSPEECH1