Tom Ko

dblp:96/8762 · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
25since 2021 · last 2025
0000-0002-5324-8961ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 34 · 3 first-author · 22 since 2021
YearPublicationVenuePosition
2025 ComLoRA: A Competitive Learning Approach for Enhancing LoRA
abstract
We propose a Competitive Low-Rank Adaptation (ComLoRA) framework to address the limitations of the LoRA method, which either lacks capacity with a single rank-$r$ LoRA or risks inefficiency and overfitting with a larger rank-$Kr$ LoRA, where $K$ is an integer larger than 1. The proposed ComLoRA method initializes $K$ distinct LoRA components, each with rank $r$, and allows them to compete during training. This competition drives each LoRA component to outperform the others, improving overall model performance. The best-performing LoRA is selected based on validation metrics, ensuring that the final model outperforms a single rank-$r$ LoRA and matches the effectiveness of a larger rank-$Kr$ LoRA, all while avoiding extra computational overhead during inference. To the best of our knowledge, this is the first work to introduce and explore competitive learning in the context of LoRA optimization. The ComLoRA's code is available at https://github.com/hqsiswiliam/comlora.
Qiushi Huang, Tom Ko, Lilian Tang, Yu Zhang 0006
ICLR2
2025 HiRA: Parameter-Efficient Hadamard High-Rank Adaptation for Large Language Models
abstract
We propose Hadamard High-Rank Adaptation (HiRA), a parameter-efficient fine-tuning (PEFT) method that enhances the adaptability of Large Language Models (LLMs). While Low-rank Adaptation (LoRA) is widely used to reduce resource demands, its low-rank updates may limit its expressiveness for new tasks. HiRA addresses this by using a Hadamard product to retain high-rank update parameters, improving the model capacity. Empirically, HiRA outperforms LoRA and its variants on several tasks, with extensive ablation studies validating its effectiveness. Our code is available at https://github.com/hqsiswiliam/hira.
Qiushi Huang, Tom Ko, Zhan Zhuang, Lilian Tang, Yu Zhang 0006
ICLR2
2024 RepCodec: A Speech Representation Codec for Speech Tokenization
abstract
With recent rapid growth of large language models (LLMs), discrete speech tokenization has played an important role for injecting speech into LLMs.However, this discretization gives rise to a loss of information, consequently impairing overall performance.To improve the performance of these discrete speech tokens, we present RepCodec, a novel speech representation codec for semantic speech tokenization.In contrast to audio codecs which reconstruct the raw audio, RepCodec learns a vector quantization codebook through reconstructing speech representations from speech encoders like HuBERT or data2vec.Together, the speech encoder, the codec encoder and the vector quantization codebook form a pipeline for converting speech waveforms into semantic tokens.The extensive experiments illustrate that RepCodec, by virtue of its enhanced information retention capacity, significantly outperforms the widely used k-means clustering approach in both speech understanding and generation.Furthermore, this superiority extends across various speech encoders and languages, affirming the robustness of RepCodec.We believe our method can facilitate large language modeling research on speech processing.Our code and models are released at https://github.com/mct10/AudioDec_ct.
Zhichao Huang 0002, Chutong Meng, Tom Ko
ACL (1)3
2024 Parameter-Efficient Transfer Learning for End-to-end Speech Translation
abstract
Recently, end-to-end speech translation (ST) has gained significant attention in research, but its progress is hindered by the limited availability of labeled data. To overcome this challenge, leveraging pre-trained models for knowledge transfer in ST has emerged as a promising direction. In this paper, we propose PETL-ST, which investigates parameter-efficient transfer learning for end-to-end speech translation. Our method utilizes two lightweight adaptation techniques, namely prefix and adapter, to modulate Attention and the Feed-Forward Network, respectively, while preserving the capabilities of pre-trained models. We conduct experiments on MuST-C En-De, Es, Fr, Ru datasets to evaluate the performance of our approach. The results demonstrate that PETL-ST outperforms strong baselines, achieving superior translation quality with high parameter efficiency. Moreover, our method exhibits remarkable data efficiency and significantly improves performance in low-resource settings.
Yunlong Zhao 0004, Qianqian Dong, Tom Ko
LREC/COLING4
2024 PolyVoice: Language Models for Speech to Speech Translation
abstract
With the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks. Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate the impact of language modeling approaches in this area. In this study, we introduce PolyVoice, a language model-based framework designed for S2ST systems. Our framework comprises three decoder-only language models: a translation language model, a duration language model, and a speech synthesis language model. These language models employ different types of prompts to extract learned information effectively. By utilizing unsupervised semantic units, our framework can transfer semantic information across these models, making it applicable even to unwritten languages. We evaluate our system on Chinese $\rightarrow$ English and English $\rightarrow$ Spanish language pairs. Experimental results demonstrate that \method outperforms the state-of-the-art encoder-decoder model, producing voice-cloned speech with high translation and audio quality. Speech samples are available at https://polyvoice.github.io.
Qianqian Dong, Zhiying Huang, Qi Tian 0001, Chen Xu 0008, Tom Ko, Yunlong Zhao 0004, Tang Li 0001, Xuxin Cheng, Fengpeng Yue, Ye Bai 0001, Lu Lu 0015, Zejun Ma 0001, Yuping Wang 0005, Mingxuan Wang, Yuxuan Wang 0002
ICLR5
2024 WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
abstract
The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years, yet the limited size of existing audio-language datasets poses challenges for researchers due to the costly and time-consuming collection process. To address this data scarcity issue, we introduceWavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400 k audio clips with paired captions. We sourced audio clips and their raw descriptions from web sources and a sound event detection dataset. However, the online-harvested raw descriptions are highly noisy and unsuitable for direct use in tasks such as automated audio captioning. To overcome this issue, we propose a three-stage processing pipeline for filtering noisy data and generating high-quality captions, where ChatGPT, a large language model, is leveraged to filter and transform raw descriptions automatically. We conduct a comprehensive analysis of the characteristics of WavCaps dataset and evaluate it on multiple downstream audio-language multimodal learning tasks. The systems trained on WavCaps outperform previous state-of-the-art (SOTA) models by a significant margin. Our aspiration is for the WavCaps dataset we have proposed to facilitate research in audio-language multimodal learning and demonstrate the potential of utilizing large language models (LLMs) to enhance academic research.
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Personalized Dialogue Generation with Persona-Adaptive Attention
abstract
Persona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs.
Qiushi Huang, Yu Zhang 0006, Tom Ko, Xubo Liu 0001, Bo Wu 0018, Wenwu Wang 0001, Lilian Tang
AAAI3
2023 CTC-based Non-autoregressive Speech Translation
abstract
Chen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun, Yuhao Zhang, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao, Anxiang Ma, Jingbo Zhu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chen Xu 0008, Qingxuan Sun, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao 0001, Anxiang Ma
ACL (1)8
2023 Leveraging per Image-Token Consistency for Vision-Language Pre-training
abstract
Most existing vision-language pre-training (VLP) approaches adopt cross-modal masked language modeling (CMLM) to learn vision-language associations. However, we find that CMLM is insufficient for this purpose according to our observations: (1) Modality bias: a considerable amount of masked tokens in CMLM can be recovered with only the language information, ignoring the visual inputs. (2) Underutilization of the unmasked tokens: CMLM primarily focuses on the masked token but it cannot simultaneously leverage other tokens to learn vision-language associations. To handle those limitations, we propose EPIC (lEveraging Per Image-Token Consistency for vision-language pre-training). In EPIC, for each image-sentence pair, we mask tokens that are salient to the image (i.e., Saliency-based Masking Strategy) and replace them with alternatives sampled from a language model (i.e., Inconsistent Token Generation Procedure), and then the model is required to determine for each token in the sentence whether it is consistent with the image (i.e., Image-Token Consistency Task). The proposed EPIC method is easily combined with pre-training methods. Extensive experiments show that the combination of the EPIC method and state-of-the-art pre-training approaches, including ViLT, ALBEF, METER, and X-VLM, leads to significant improvements on downstream tasks. Our coude is released at https://github.com/gyhdog99/epic
Yunhao Gou, Tom Ko, Hansi Yang, James T. Kwok, Yu Zhang 0006, Mingxuan Wang
CVPR2
2023 Learning Retrieval Augmentation for Personalized Dialogue Generation
abstract
Personalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications.However, persona profiles, a prevalent setting in current personalized dialogue datasets, typically composed of merely four to five sentences, may not offer comprehensive descriptions of the persona about the agent, posing a challenge to generate truly personalized dialogues.To handle this problem, we propose Learning Retrieval Augmentation for Personalized DialOgue Generation (LAPDOG), which studies the potential of leveraging external knowledge for persona dialogue generation.Specifically, the proposed LAPDOG model consists of a story retriever and a dialogue generator.The story retriever uses a given persona profile as queries to retrieve relevant information from the story document, which serves as a supplementary context to augment the persona profile.The dialogue generator utilizes both the dialogue history and the augmented persona profile to generate personalized responses.For optimization, we adopt a joint training framework that collaboratively learns the story retriever and dialogue generator, where the story retriever is optimized towards desired ultimate metrics (e.g., BLEU) to retrieve content for the dialogue generator to generate personalized responses.Experiments conducted on the CONVAI2 dataset with ROCStory as a supplementary data source show that the proposed LAPDOG method substantially outperforms the baselines, indicating the effectiveness of the proposed method.The LAPDOG model code is publicly available for further exploration.
Qiushi Huang, Xubo Liu 0001, Wenwu Wang 0001, Tom Ko, Yu Zhang 0006, Lilian Tang
EMNLP5
2023 M3ST: Mix at Three Levels for Speech Translation
abstract
How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It’s well known that data augmentation is an efficient method to improve performance for many tasks by enlarging the dataset. In this paper, we propose Mix at three levels for Speech Translation (M3ST) method to increase the diversity of the augmented training corpus. Specifically, we conduct two phases of fine-tuning based on a pre-trained model using external machine translation (MT) data. In the first stage of fine-tuning, we mix the training corpus at three levels, including word level, sentence level and frame level, and fine-tune the entire model with mixed data. At the second stage of fine-tuning, we take both original speech sequences and original text sequences in parallel into the model to fine-tune the network, and use Jensen-Shannon divergence to regularize their outputs. Experiments on MuST-C speech translation benchmark and analysis show M3ST outperforms current strong baselines and achieves state-of-the-art results on eight directions with an average BLEU of 29.9.
Xuxin Cheng, Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Yuexian Zou
ICASSP4
2023 Recent Advances in Direct Speech-to-text Translation
abstract
Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges --- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.
Chen Xu 0008, Rong Ye, Qianqian Dong, Chengqi Zhao, Tom Ko, Mingxuan Wang, Tong Xiao 0001
IJCAI5
2023 Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention
abstract
Audio captioning aims to generate text descriptions of audio clips.In the real world, many objects produce similar sounds.How to accurately recognize ambiguous sounds is a major challenge for audio captioning.In this work, inspired by inherent human multimodal perception, we propose visuallyaware audio captioning, which makes use of visual information to help the description of ambiguous sounding objects.Specifically, we introduce an off-the-shelf visual encoder to extract video features and incorporate the visual features into an audio captioning system.Furthermore, to better exploit complementary audio-visual contexts, we propose an audio-visual attention mechanism that adaptively integrates audio and visual context and removes the redundant information in the latent space.Experimental results on AudioCaps, the largest audio captioning dataset, show that our proposed method achieves state-of-theart results on machine translation metrics.
Xubo Liu 0001, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang 0006, Lilian Tang, Mark D. Plumbley, Volkan Kilic, Wenwu Wang 0001
INTERSPEECH8
2023 CoBERT: Self-Supervised Speech Representation Learning Through Code Representation Learning
Chutong Meng, Junyi Ao, Tom Ko, Mingxuan Wang, Haizhou Li 0001
INTERSPEECH3
2023 GigaST: A 10, 000-hour Pseudo Speech Translation Corpus
Rong Ye, Chengqi Zhao, Tom Ko, Chutong Meng, Tao Wang 0086, Mingxuan Wang
INTERSPEECH3
2022 SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
abstract
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei
ACL (1)8
2022 Multi-View Self-Attention Based Transformer for Speaker Recognition
abstract
Initially developed for natural language processing (NLP), Transformer model is now widely used for speech processing tasks such as speaker recognition, due to its powerful sequence modeling capabilities. However, conventional self-attention mechanisms are originally designed for modeling textual sequence without considering the characteristics of speech and speaker modeling. Besides, different Transformer variants for speaker recognition have not been well studied. In this work, we propose a novel multi-view self-attention mechanism and present an empirical study of different Transformer variants with or without the proposed attention mechanism for speaker recognition. Specifically, to balance the capabilities of capturing global dependencies and modeling the locality, we propose a multi-view self-attention mechanism for speaker Transformer, in which different attention heads can attend to different ranges of the receptive field. Furthermore, we introduce and compare five Transformer variants with different network architectures, embedding locations, and pooling methods to learn speaker embeddings. Experimental results on the VoxCeleb1 and VoxCeleb2 datasets show that the proposed multi-view self-attention mechanism achieves improvement in the performance of speaker recognition, and the proposed speaker Transformer network attains excellent results compared with state-of-the-art models.
Rui Wang 0073, Junyi Ao, Shujie Liu 0001, Zhihua Wei 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006
ICASSP6
2022 Exploring Machine Speech Chain For Domain Adaptation
abstract
Machine Speech Chain integrates both end-to-end (E2E) automatic speech recognition (ASR) and neural text-to-speech (TTS) into one circle for joint training. It has been proven that it can effectively leverage a large amount of unpaired data in the spirit of data augmentation. In this paper, we explore the TTS→ASR pipeline in machine speech chain to perform domain adaptation for both E2E ASR and neural TTS models with only text data from the target domain. We conduct experiments by adapting from audiobook domain (i.e., LibriSpeech) to presentation domain (i.e., TED-LIUM). There is a relative word error rate (WER) reduction of 19.7% for the E2E ASR model on the TED-LIUM test set, and a relative WER reduction of 29.4% in synthetic speech generated by neural TTS in the presentation domain. Moreover, we observe that the gains from the proposed method and conventional adaptation methods of language models are additive.
Fengpeng Yue, Lei He 0005, Tom Ko, Yu Zhang 0006
ICASSP4
2022 Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
abstract
This paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR).Within a multi-task learning framework, we introduce two pre-training tasks for the encoderdecoder network using acoustic units, i.e., pseudo codes, derived from an offline clustering model.One is to predict the pseudo codes via masked language modeling in encoder output, like HuBERT model, while the other lets the decoder learn to reconstruct pseudo codes autoregressively instead of generating textual scripts.In this way, the decoder learns to reconstruct original speech information with codes before learning to generate correct text.Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and Hu-BERT on fine-tuning subsets of 10h and 100h.We release our code and model at https://github.com/microsoft/SpeechT5/tree/main/Speech2C.
Junyi Ao, Shujie Liu 0001, Haizhou Li 0001, Tom Ko, Li-Rong Dai 0001, Jinyu Li 0001, Yao Qian, Furu Wei
INTERSPEECH6
2022 A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis
abstract
In human speech, the attitude of a speaker cannot be fully expressed only by the textual content.It has to come along with the intonation.Declarative questions are commonly used in daily Cantonese conversations, and they are usually uttered with rising intonation.Vanilla neural text-to-speech (TTS) systems are not capable of synthesizing rising intonation for these sentences due to the loss of semantic information.Though it has become more common to complement the systems with extra language models, their performance in modeling rising intonation is not well studied.In this paper, we propose to complement the Cantonese TTS model with a BERT-based statement/question classifier.We design different training strategies and compare their performance.We conduct our experiments on a Cantonese corpus named CanTTS.Empirical results show that the separate training approach obtains the best generalization performance and feasibility.
Qibing Bai, Tom Ko, Yu Zhang 0006
INTERSPEECH2
2022 Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation
abstract
Direct Speech-to-speech translation (S2ST) has drawn more and more attention recently. The task is very challenging due to data scarcity and complex speech-to-speech mapping. In this paper, we report our recent achievements in S2ST. Firstly, we build a S2ST Transformer baseline which outperforms the original Translatotron. Secondly, we utilize the external data by pseudo-labeling and obtain a new state-of-the-art result on the Fisher English-to-Spanish test set. Indeed, we exploit the pseudo data with a combination of popular techniques which are not trivial when applied to S2ST. Moreover, we evaluate our approach on both syntactically similar (Spanish-English) and distant (English-Chinese) language pairs. Our implementation is available at https://github.com/fengpeng-yue/speech-to-speech-translation.
Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, Yu Zhang 0006
INTERSPEECH3
2022 LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT
abstract
Self-supervised speech representation learning has shown promising results in various speech processing tasks.However, the pre-trained models, e.g., HuBERT, are storage-intensive Transformers, limiting their scope of applications under lowresource settings.To this end, we propose LightHuBERT, a once-for-all Transformer compression framework, to find the desired architectures automatically by pruning structured parameters.More precisely, we create a Transformer-based supernet that is nested with thousands of weight-sharing subnets and design a two-stage distillation strategy to leverage the contextualized latent representations from HuBERT.Experiments on automatic speech recognition (ASR) and the SU-PERB benchmark show the proposed LightHuBERT enables over 10 9 architectures concerning the embedding dimension, attention dimension, head number, feed-forward network ratio, and network depth.LightHuBERT outperforms the original HuBERT on ASR and five SUPERB tasks with the Hu-BERT size, achieves comparable performance to the teacher model in most tasks with a reduction of 29% parameters, and obtains a 3.5× compression ratio in three SUPERB tasks, e.g., automatic speaker verification, keyword spotting, and intent classification, with a slight accuracy loss.The code and pre-trained models are available at https://github.com/mechanicalsea/lighthubert.
Rui Wang 0073, Qibing Bai, Junyi Ao, Zhixiang Xiong, Zhihua Wei 0001, Yu Zhang 0006, Tom Ko, Haizhou Li 0001
INTERSPEECH8
2021 A Meta-Learning Approach for User-Defined Spoken Term Classification with Varying Classes and Examples
Yangbin Chen, Tom Ko, Jianping Wang 0001
Interspeech2
2021 Token-Level Supervised Contrastive Learning for Punctuation Restoration
abstract
Punctuation is critical in understanding natural language text. Currently, most automatic speech recognition (ASR) systems do not generate punctuation, which affects the performance of downstream tasks, such as intent detection and slot filling. This gives rise to the need for punctuation restoration. Recent work in punctuation restoration heavily utilizes pre-trained language models without considering data imbalance when predicting punctuation classes. In this work, we address this problem by proposing a token-level supervised contrastive learning method that aims at maximizing the distance of representation of different punctuation marks in the embedding space. The result shows that training with token-level supervised contrastive learning obtains up to 3.2% absolute F1 improvement on the test set.
Qiushi Huang, Tom Ko, Lilian Tang, Xubo Liu 0001, Bo Wu 0018
Interspeech2
2021 Auto-KWS 2021 Challenge: Task, Datasets, and Baselines
abstract
Auto-KWS 2021 challenge calls for automated machine learning (AutoML) solutions to automate the process of applying machine learning to a customized keyword spotting task.Compared with other keyword spotting tasks, Auto-KWS challenge has the following three characteristics: 1) The challenge focuses on the problem of customized keyword spotting, where the target device can only be awakened by an enrolled speaker with his specified keyword.The speaker can use any language and accent to define his keyword.2) All dataset of the challenge is recorded in realistic environment.It is to simulate different user scenarios.3) Auto-KWS is a "code competition", where participants need to submit AutoML solutions, then the platform automatically runs the enrollment and prediction steps with the submitted code.This challenge aims at promoting the development of a more personalized and flexible keyword spotting system.Two baseline systems are provided to all participants as references.
Jingsong Wang, Qijie Shao, Wei-Wei Tu, Tom Ko, Hung-yi Lee, Lei Xie 0001
Interspeech6
2020 Prototypical Networks for Small Footprint Text-Independent Speaker Verification
abstract
Speaker verification aims to recognize target speakers with very few enrollment utterances. Conventional approaches learn a representation model to extract the speaker embeddings for verification. Recently, there are several new approaches in meta-learning which try to learn a shared metric space. Among these approaches, prototypical networks aim at learning a non-linear mapping from the input space to an embedding space with a predefined distance metric. In this paper, we investigate the use of prototypical networks in a small footprint text-independent speaker verification task. Our work is evaluated on SRE10 evaluation set. Experiments show that prototypical networks outperform the conventional method when the amount of data per training speaker is limited.
Tom Ko, Yangbin Chen, Qing Li 0001
ICASSP1
2020 MetaMix: Improved Meta-Learning with Interpolation-based Consistency Regularization
abstract
Model-Agnostic Meta-Learning (MAML) and its variants are popular few-shot classification methods. They train an initializer across a variety of sampled learning tasks (also known as episodes) such that the initialized model can adapt quickly to new ones. However, current MAML-based algorithms have limitations in forming generalizable decision boundaries. In this paper, we propose an approach called MetaMix, which generates virtual feature-target pairs within each episode to regularize the backbone models. MetaMix can be integrated with any of the MAML-based algorithms and learn the decision boundaries generalizing better to new tasks. Experiments on the mini-ImageNet, CUB, and FC100 datasets show that MetaMix improves the performance of MAML-based algorithms and achieves state-of-the-art result when integrated with Meta-Transfer Learning.
Yangbin Chen, Yun Ma 0001, Tom Ko, Jianping Wang 0001, Qing Li 0001
ICPR3
2020 An Investigation of Few-Shot Learning in Spoken Term Classification
abstract
202402 bcch
Yangbin Chen, Tom Ko, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qing Li 0001
INTERSPEECH2
2020 AutoSpeech 2020: The Second Automated Machine Learning Challenge for Speech Classification
abstract
The AutoSpeech challenge calls for automated machine learning (AutoML) solutions to automate the process of applying machine learning to speech processing tasks.These tasks, which cover a large variety of domains, will be shown to the automated system in a random order.Each time when the tasks are switched, the information of the new task will be hinted with its corresponding training set.Thus, every submitted solution should contain an adaptation routine which adapts the system to the new task.Compared to the first edition, the 2020 edition includes advances of 1) more speech tasks, 2) noisier data in each task, 3) a modified evaluation metric.This paper outlines the challenge and describe the competition protocol, datasets, evaluation metric, starting kit, and baseline systems.
Jingsong Wang, Tom Ko, Zhen Xu 0007, Xiawei Guo, Souxiang Liu, Wei-Wei Tu, Lei Xie 0001
INTERSPEECH2
2019 Mixup Learning Strategies for Text-Independent Speaker Verification
abstract
Mixup is a learning strategy that constructs additional virtual training samples from existing training samples by linearly interpolating random pairs of them. It has been shown that mixup can help avoid data memorization and thus improve model generalization. This paper investigates the mixup learning strategy in training speaker-discriminative deep neural network (DNN) for better text-independent speaker verification. In recent speaker verification systems, a DNN is usually trained to classify speakers in the training set. The DNN, at the same time, learns a low-dimensional embedding of speakers so that speaker embeddings can be generated for any speakers during evaluation. We adapted the mixup strategy to the speaker-discriminative DNN training procedure, and studied different mixup schemes, such as performing mixup on MFCC features or raw audio samples. The mixup learning strategy was evaluated on NIST SRE 2010, 2016 and SITW evaluation sets. Experimental results show consistent performance improvements both in terms of EER and DCF of up to 13% relative. We further find that mixup training also improves the DNN's speaker classification accuracy consistently without requiring any additional data sources. Copyright © 2019 ISCA
Yingke Zhu, Tom Ko, Brian Kan-Wing Mak
INTERSPEECH2
2018 Long Distance Voice Channel Diagnosis Using Deep Neural Networks
Tom Ko, Guangjian Tian
INTERSPEECH2
2018 Self-Attentive Speaker Embeddings for Text-Independent Speaker Verification
abstract
This paper introduces a new method to extract speaker embeddings from a deep neural network (DNN) for text-independent speaker verification. Usually, speaker embeddings are extracted from a speaker-classification DNN that averages the hidden vectors over the frames of a speaker; the hidden vectors produced from all the frames are assumed to be equally important. We relax this assumption and compute the speaker embedding as a weighted average of a speaker's frame-level hidden vectors, and their weights are automatically determined by a self-attention mechanism. The effect of multiple attention heads are also investigated to capture different aspects of a speaker's input speech. Finally, a PLDA classifier is used to compare pairs of embeddings. The proposed self-attentive speaker embedding system is compared with a strong DNN embedding baseline on NIST SRE 2016. We find that the self-attentive embeddings achieve superior performance. Moreover, the improvement produced by the self-attentive speaker embeddings is consistent with both short and long testing utterances. © 2018 International Speech Communication Association. All rights reserved.
Yingke Zhu, Tom Ko, David Snyder, Brian Kan-Wing Mak, Daniel Povey
INTERSPEECH2
2017 A study on data augmentation of reverberant speech for robust speech recognition
abstract
The environmental robustness of DNN-based acoustic models can be significantly improved by using multi-condition training data. However, as data collection is a costly proposition, simulation of the desired conditions is a frequently adopted strategy. In this paper we detail a data augmentation approach for far-field ASR. We examine the impact of using simulated room impulse responses (RIRs), as real RIRs can be difficult to acquire, and also the effect of adding point-source noises. We find that the performance gap between using simulated and real RIRs can be eliminated when point-source noises are added. Further we show that the trained acoustic models not only perform well in the distant-talking scenario but also provide better results in the close-talking scenario. We evaluate our approach on several LVCSR tasks which can adequately represent both scenarios.
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, Sanjeev Khudanpur
ICASSP1
2016 An empirical exploration of CTC acoustic models
abstract
The connectionist temporal classification (CTC) loss function has several interesting properties relevant for automatic speech recognition (ASR): applied on top of deep recurrent neural networks (RNNs), CTC learns the alignments between speech frames and label sequences automatically, which removes the need for pre-generated frame-level labels. CTC systems also do not require context decision trees for good performance, using context-independent (CI) phonemes or characters as targets. This paper presents an extensive exploration of CTC-based acoustic models applied to a variety of ASR tasks, including an empirical study of the optimal configuration and architectural variants for CTC. We observe that on large amounts of training data, CTC models tend to outperform state-of-the-art hybrid approach. Further experiments reveal that CTC can be readily ported to syllable-based languages, and can be enhanced by employing improved feature front-ends.
Yajie Miao, Mohammad Gowayyed, Xingyu Na, Tom Ko, Florian Metze, Alex Waibel
ICASSP4
2015 JHU ASpIRE system: Robust LVCSR with TDNNS, iVector adaptation and RNN-LMS
abstract
Multi-style training, using data which emulates a variety of possible test scenarios, is a popular approach towards robust acoustic modeling. However acoustic models capable of exploiting large amounts of training data in a comparatively short amount of training time are essential. In this paper we tackle the problem of reverberant speech recognition using 5500 hours of simulated reverberant data. We use time-delay neural network (TDNN) architecture, which is capable of tackling long-term interactions between speech and corrupting sources in reverberant environments. By sub-sampling the outputs at TDNN layers across time steps, training time is substantially reduced. Combining this with distributed-optimization we show that the TDNN can be trained in 3 days using up to 32 GPUs. Further, iVectors are used as an input to the neural network to perform instantaneous speaker and environment adaptation. Finally, recurrent neural network language models are applied to the lattices to further improve the performance. Our system is shown to provide state-of-the-art results in the IARPA ASpIRE challenge, with 26.5% WER on the dev Jest set.
Vijayaditya Peddinti, Guoguo Chen, Vimal Manohar, Tom Ko, Daniel Povey, Sanjeev Khudanpur
ASRU4
2015 Audio augmentation for speech recognition
abstract
Data augmentation is a common strategy adopted to increase the quantity of training data, avoid overfitting and improve robustness of the models. In this paper, we investigate audio-level speech augmentation methods which directly process the raw signal. The method we particularly recommend is to change the speed of the audio signal, producing 3 versions of the original signal with speed factors of 0.9, 1.0 and 1.1. The proposed technique has a low implementation cost, making it easy to adopt. We present results on 4 different LVCSR tasks with training data ranging from 100 hours to 1000 hours, to examine the effectiveness of audio augmentation in a variety of data scenarios. An average relative improvement of 4.3% was observed across the 4 tasks.
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Sanjeev Khudanpur
INTERSPEECH1
2014 Subspace Gaussian mixture model with state-dependent subspace dimensions
abstract
In recent years, under the hidden Markov modeling (HMM) framework, the use of subspace Gaussian mixture models (SGMMs) has demonstrated better recognition performance than traditional Gaussian mixture models (GMMs) in automatic speech recognition. In state-of-the-art SGMM formulation, a fixed subspace dimension is assigned to every phone states. While a constant subspace dimension is easier to implement, it may, however, lead to overfitting or underfitting of some state models as the data is usually distributed unevenly among the states. In a later extension of SGMM, states are split to sub-states with an appropriate objective function so that the problem is eased by increasing the state-specific parameters for the underfitting state. In this paper, we propose another solution and allow each sub-state to have a different subspace dimension depending on its amount of training frames so that the state-specific parameters can be robustly estimated. Experimental evaluation on the Switchboard recognition task shows that our proposed method brings improvement to the existing SGMM training procedure.
Tom Ko, Brian Kan-Wing Mak, Cheung-Chi Leung
ICASSP1
2014 Eigentrigraphemes for under-resourced languages
Tom Ko, Brian Kan-Wing Mak
Speech Commun.1
2013 Eigentriphones for Context-Dependent Acoustic Modeling
abstract
Most automatic speech recognizers employ tied-state triphone hidden Markov models (HMM), in which the corresponding triphone states of the same base phone are tied. State tying is commonly performed with the use of a phonetic regression class tree which renders robust context-dependent modeling possible by carefully balancing the amount of training data with the degree of tying. However, tying inevitably introduces quantization error: triphones tied to the same state are not distinguishable in that state. Recently we proposed a new triphone modeling approach called eigentriphone modeling in which all triphone models are, in general, distinct. The idea is to create an eigenbasis for each base phone (or phone state) and all its triphones (or triphone states) are represented as distinct points in the space spanned by the basis. We have shown that triphone HMMs trained using model-based or state-based eigentriphones perform at least as well as conventional tied-state HMMs. In this paper, we further generalize the definition of eigentriphones over clusters of acoustic units. Our experiments on TIMIT phone recognition and the Wall Street Journal 5K-vocabulary continuous speech recognition show that eigentriphones estimated from state clusters defined by the nodes in the same phonetic regression class tree used in state tying result in further performance gain.
Tom Ko, Brian Kan-Wing Mak
IEEE Trans. Speech Audio Process.1
2012 Derivation of eigentriphones by weighted principal component analysis
abstract
Last year we proposed a new acoustic modeling method called eigentriphones in which all triphones are distinct (with no tied states) so that they may be more discriminative. In our method, frequent triphones are used to derive an eigenbasis using PCA, and the infrequent triphones are then “adapted” as a linear combination of the eigenvectors which are also called eigentriphones. Although the eigentriphones method compares favorably with traditional tied-state triphones, the PCA procedure has two limitations: (1) only the frequent triphones are employed, and (2) they are considered “equal” even though some are more robust than the others. In this paper, weighted PCA is proposed to solve both problems so that all triphones-frequent and infrequent triphones-may contribute to the derivation of the eigentriphones, each at a different extent depending on its sample count. Experimental evaluation on the WSJ 5Kvocabulary speech recognition task shows that weighted PCA produces better models than simple PCA, and its performance is fairly independent of the number of eigentriphones once more than 20% of them are used. As a consequence, all triphones may be represented by fewer eigentriphones, resulting in a more compact model.
Tom Ko, Brian Kan-Wing Mak
ICASSP1
2011 Eigentriphones: A basis for context-dependent acoustic modeling
abstract
In context-dependent acoustic modeling, it is important to strike a balance between detailed modeling and data sufficiency for robust estimation of model parameters. In the past, parameter sharing or tying is one of the most common techniques to solve the problem. In recent years, another technique which may be loosely and collectively called the subspace approach tries to express a phonetic or sub-phonetic unit in terms of a small set of canonical vectors or units. In this paper, we investigate the development of an eigenbasis over the triphones and model each triphone as a point in the basis. We call the eigenvectors in the basis eigentriphones. From another perspective, we investigate the use of the eigenvoice adaptation method as a general acoustic modeling method for training triphones - especially the less frequent triphones without tying their states so that all the triphones are really distinct from each other and thus may be more discriminative. Experimental evaluation on the 5K-vocabulary HUB2 recognition task shows that a triphone HMM system trained using only eigentriphones without state tying may achieve slightly better performance than the common tied-state triphones.
Tom Ko, Brian Kan-Wing Mak
ICASSP1
2011 A Fully Automated Derivation of State-Based Eigentriphones for Triphone Modeling with No Tied States Using Regularization
abstract
Recently we proposed an alternative method called eigentriphone to solve the data insufficiency problem in triphone acoustic modeling without the need of state tying. The idea is to treat the acoustic modeling problem of infrequent triphones ("poor triphones") as an adaptation problem from the more frequent triphones ("rich triphones"): firstly, an eigenbasis is developed over the rich triphones that have sufficient training data and the eigenvectors are called eigentriphones; then the poor triphones are adapted in a fashion similar to eigenvoice adaptation. Since, in general, no states are tied in our method, all triphones (states) are distinct so that they can be more discriminative than tied-state triphones. In our previous work, the number of eigentriphones was determined in advance with a set of development data. In this paper, we investigate simply using all of them with the help of regularization to naturally penalize the less important ones. In addition, the model-based eigenbasis is replaced by three state-based eigenbases. Experimental evaluation on the WSJ 5K task shows that triphone models trained using our new eigentriphone approach without state tying perform at least as well as the common tied-state triphone models.
Tom Ko, Brian Kan-Wing Mak
INTERSPEECH1
2010 Improving speech recognition by explicit modeling of phone deletions
abstract
In a paper published by Greenberg in 1998, it was said that in conversational speech, phone deletion rate may go as high as 12% whereas syllable deletion rate is about 1%. The finding prompted a new research direction of syllable modeling for speech recognition. To date, the syllable approach has not yet fulfilled its promise. On the other hand, there were few attempts to model phone deletions explicitly in current ASR systems. In this paper, fragmented word models were derived from well-trained cross-word triphone models, and phone deletion was implemented by skip arcs for words consisting of at least four phonemes. An evaluation on CSR-II WSJ1 Hub2 5K task shows that even with this limited implementation of phone deletions in read speech, we obtained a word error rate reduction of 6.73%.
Tom Ko, Brian Kan-Wing Mak
ICASSP1
2009 Automatic estimation of decoding parameters using large-margin iterative linear programming
abstract
The decoding parameters in automatic speech recognition - grammar factor and word insertion penalty - are usually determined by performing a grid search on a development set. Recently, we cast their estimation as a convex optimization problem, and proposed a solution using an iterative linear programming algorithm. However, the solution depends on how well the development data set matches with the test set. In this paper, we further investigates an improvement on the generalization property of the solution by using large margin training within the iterative linear programming framework. Empirical evaluation on the WSJ0 5K speech recognition tasks shows that the recognition performance of the decoding parameters found by the improved algorithm using only a subset of the acoustic model training data is even better than that of the decoding parameters found by grid search on the development data, and is close to the performance of those found by grid search on the test set.
Brian Kan-Wing Mak, Tom Ko
INTERSPEECH2
2008 Min-max discriminative training of decoding parameters using iterative linear programming
abstract
In automatic speech recognition, the decoding parameters - grammar factor and word insertion penalty-are usually hand-tuned to give the best recognition performance. This paper investigates an automatic procedure to determine their values using an iterative linear programming (LP) algorithm. LP naturally implements discriminative training by mapping linear discriminants into LP constraints. A min-max cost function is also defined to get more stable and robust result. Empirical evaluations on the RM1 and WSJ0 speech recognition tasks show that decoding parameters found by the proposed algorithm are as good as those found by a brute-force grid search; their optimal values also seem to be independent of the initial values set to start the iterative LP algorithm.
Brian Kan-Wing Mak, Tom Ko
INTERSPEECH2