Kwok Chin Yuen

dblp:362/1351 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Extending Whisper for Emotion Prediction Using Word-level Pseudo Labels
abstract
This paper extends Whisper’s automatic speech recognition (ASR) capabilities to perform speech-based emotion recognition (SER) by incorporating word-level emotion classification alongside ASR output. We generate four emotion pseudo-labels (neutral, happy, sad, angry) for each word using a pretrained frame-level SER model, and Whisper is fine-tuned for joint ASR and emotion classification at the word level. Sentence-level emotion labels are masked during training to encourage the transformer to use the ASR output for word-level emotion prediction. During inference, word-level predictions are combined with sentence-level predictions through majority voting to generate the final sentence-level label. When evaluated on the IEMOCAP dataset, our method maintains Whisper’s ASR word error rate while improving the SER weighted accuracy from 74.4% to 76.4% and the unweighted average recall from 77.1% to 79.0%.
Kwok Chin Yuen, Sheng Li 0010, Jia Qi Yip, Chenhui Chu, Tatsuya Kawahara, Chng Eng Siong
ICASSP1
2025 Robust Audio Deepfake Detection using Ensemble Confidence Calibration
abstract
Model ensembles using linear interpolation are commonly employed to improve classification performance, with higher weights assigned to better-performing models in the ensemble. However, prior methods use fixed weights across all test samples, which is suboptimal as different models may perform better in different subsets of the samples, especially in out-of-domain (OOD) scenarios. This is a key challenge in Audio Deepfake Detection (ADD) due to variations between training and testing domains. To address this, we propose using EOW-Softmax, a method for modeling open-world uncertainties, to calibrate the magnitudes of OOD classification scores at the sample level. This dynamic adjustment improves ensemble predictions on OOD samples. When tested on the ASVspoof 2021 dataset, our calibrated ensemble reduced the equal error rate (EER) from 2.66% to 2.03%.
Kwok Chin Yuen, Duc-Tuan Truong, Jia Qi Yip
ICASSP1
2025 Efficient Trie-based Biasing using K-step Prediction for Rare Word Recognition
abstract
Contextual biasing improves rare word recognition of ASR models by prioritizing the output of rare words during decoding. A common approach is Trie-based biasing, which gives "bonus scores" to partial hypothesis (e.g. "Bon") that may lead to the generation of the rare word (e.g. "Bonham"). If the full word ("Bonham") isn't ultimately recognized, the system revokes those earlier bonuses. This revocation is limited to beam search and is computationally expensive, particularly for models with large decoders. To overcome these limitations, we propose adapting ASR models to look ahead and predict multiple steps at once. This avoids the revocation step entirely by better estimating whether a partial hypothesis will lead to the generation of the full rare word. By fine-tuning Whisper with only 10 hours of synthetic data, our method reduces the word error rate on the NSC Part 2 test set from 30.86% to 12.19%.
Kwok Chin Yuen, Jia Qi Yip
INTERSPEECH1
2025 Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function
abstract
Rare word recognition can be improved by adapting ASR models to synthetic data that includes these words. Further improvements can be achieved through contextual biasing, which trains and adds a biasing module into the model architecture to prioritize rare words. While training the module on synthetic rare word data is more effective than using non-rare-word data, it can lead to overfitting due to artifacts in the synthetic audio. To address this, we enhance the TCPGen-based contextual biasing approach and propose a keyword-aware loss function that additionally focuses on biased words when training biasing modules. This loss includes a masked cross-entropy term for biased word prediction and a binary classification term for detecting biased word positions. These two terms complementarily support the decoding of biased words during inference. By adapting Whisper to 10 hours of synthetic data, our method reduced the word error rate on the NSC Part 2 test set from 29.71% to 11.81%.
Kwok Chin Yuen, Jia Qi Yip, Chng Eng Siong
INTERSPEECH1
2025 Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
abstract
Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers with more samples, underrepresenting others and reducing the overall reliability of EER. Additionally, most ADD datasets lack diversity in bona fide speech, often featuring a single environment and speech style (e.g., clean read speech), limiting their ability to simulate real-world conditions. To address these challenges, we propose bona fide cross-testing, a novel evaluation framework that incorporates diverse bona fide datasets and aggregates EERs for more balanced assessments. Our approach improves robustness and interpretability compared to traditional evaluation methods. We benchmark over 150 synthesizers across nine bona fide speech types and release a new dataset to facilitate further research at https://github.com/cyaaronk/audio_deepfake_eval.
Kwok Chin Yuen, Jia Qi Yip, Chihung Chi, Kwok-Yan Lam
INTERSPEECH1
2024 Investigating ASR Error Correction with Large Language Model and Multilingual 1-best Hypotheses
Sheng Li 0010, Chen Chen 0075, Kwok Chin Yuen, Chenhui Chu, Chng Eng Siong, Hisashi Kawai
INTERSPEECH3
2024 Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
Kwok Chin Yuen, Jia Qi Yip, Chng Eng Siong
INTERSPEECH1
2024 Continual Learning With Embedding Layer Surgery and Task-Wise Beam Search Using Whisper
abstract
Current Multilingual ASR models only support a fraction of the world’s languages. Continual Learning (CL) aims to tackle this problem by adding new languages to pre-trained models while avoiding the loss of performance on existing languages, also known as Catastrophic Forgetting (CF). However, existing CL methods overlook the adaptation of the token embedding lookup table at the decoder, despite its significant contribution to CF. We propose Embedding Layer Surgery where separate copies of the token embeddings are created for each new languages, and one of the copies is selected to replace the old languages embeddings when transcribing the corresponding new language. Unfortunately, this approach means LID errors also cause incorrect ASR embedding selection. Our Task-wise Beam Search allows self-correction for such mistakes. By adapting Whisper to 10 hours of data for each of 10 unseen languages from Common Voice, results show that our method reduces the Average WER (AWER) of pre-trained languages from 14.2% to 11.9% compared with Experience Replay, without compromising the AWER of the unseen languages.
Kwok Chin Yuen, Jia Qi Yip, Chng Eng Siong
SLT1