VLDB 2026 Research / reviewers in the wild / expert
Brian Kan-Wing Mak
dblp:70/6226 · also Brian Mak
· DBLP profile ↗
103ranked-venue papers
29as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 82 · 21 first-author · 12 since 2021Artificial intelligence and machine learning · 65 · 19 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Residual Matrix Transformers: Scaling the Size of the Residual StreamabstractThe residual stream acts as a memory bus where transformer layers both store and access features (Elhage et al., 2021). We consider changing the mechanism for retrieving and storing information in the residual stream, and replace the residual stream of the transformer with an outer product memory matrix (Kohonen, 1972, Anderson, 1972). We call this model the Residual Matrix Transformer (RMT). We find that the RMT enjoys a number of attractive properties: 1) the size of the residual stream can be scaled independently of compute and model size, improving performance, 2) the RMT can achieve the same loss as the transformer with 58% fewer FLOPS, 25% fewer parameters, and 41% fewer training tokens tokens, and 3) the RMT outperforms the transformer on downstream evaluations. We theoretically analyze the transformer and the RMT, and show that the RMT allows for more efficient scaling of the residual stream, as well as improved variance propagation properties. Brian Kan-Wing Mak, Jeffrey Flanigan |
ICML | 1 |
| 2024 | A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV NewsabstractThis paper introduces TVB-HKSL-News, a new Hong Kong sign language (HKSL) dataset collected from a TV news program over a period of 7 months. The dataset is collected to enrich resources for HKSL and support research in large-vocabulary continuous sign language recognition (SLR) and translation (SLT). It consists of 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses (for SLR) and 2,850 Chinese characters or 18K Chinese words (for SLT). One signer has 11.66 hours of sign videos and the other has 4.41 hours. One objective in building the dataset is to support the investigation of how well large-vocabulary continuous sign language recognition/translation can be done for a single signer given a (relatively) large amount of his/her training data, which could potentially lead to the development of new modeling methods. Besides, most parts of the data collection pipeline are automated with little human intervention; we believe that our collection method can be scaled up to collect more sign language data easily for SLT in the future for any sign languages if such sign-interpreted videos are available. We also run a SOTA SLR/SLT model on the dataset and get a baseline SLR word error rate of 34.08% and a baseline SLT BLEU-4 score of 23.58 for benchmarking future research on the dataset. Zhe Niu, Ronglai Zuo, Brian Kan-Wing Mak, Fangyun Wei |
LREC/COLING | 3 |
| 2024 | A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
Ronglai Zuo, Fangyun Wei, Zenggui Chen, Brian Kan-Wing Mak, Jiaolong Yang, Xin Tong 0001 |
ECCV (49) | 4 |
| 2024 | Towards Online Continuous Sign Language Recognition and TranslationabstractResearch on continuous sign language recognition (CSLR) is essential to bridge the communication gap between deaf and hearing individuals.Numerous previous studies have trained their models using the connectionist temporal classification (CTC) loss.During inference, these CTC-based models generally require the entire sign video as input to make predictions, a process known as offline recognition, which suffers from high latency and substantial memory usage.In this work, we take the first step towards online CSLR.Our approach consists of three phases: 1) developing a sign dictionary; 2) training an isolated sign language recognition model on the dictionary; and 3) employing a sliding window approach on the input sign sequence, feeding each sign clip to the optimized model for online recognition.Additionally, our online recognition model can be extended to support online translation by integrating a gloss-to-text network and can enhance the performance of any offline model.With these extensions, our online approach achieves new state-of-the-art performance on three popular benchmarks across various task settings. Ronglai Zuo, Fangyun Wei, Brian Kan-Wing Mak |
EMNLP | 3 |
| 2024 | Improving Continuous Sign Language Recognition with Consistency Constraints and Signer RemovalabstractDeep-learning-based continuous sign language recognition (CSLR) models typically consist of a visual module, a sequential module, and an alignment module. However, the effectiveness of training such CSLR backbones is hindered by limited training samples, rendering the use of a single connectionist temporal classification loss insufficient. To address this limitation, we propose three auxiliary tasks to enhance CSLR backbones. First, we enhance the visual module, which is particularly sensitive to the challenges posed by limited training samples, from the perspective of consistency. Specifically, since sign languages primarily rely on signers’ facial expressions and hand movements to convey information, we develop a keypoint-guided spatial attention module that directs the visual module to focus on informative regions, thereby ensuring spatial attention consistency. Furthermore, recognizing that the output features of both the visual and sequential modules represent the same sentence, we leverage this prior knowledge to better exploit the power of the backbone. We impose a sentence embedding consistency constraint between the visual and sequential modules, enhancing the representation power of both features. The resulting CSLR model, referred to as consistency-enhanced CSLR, demonstrates superior performance on signer-dependent datasets, where all signers appear during both training and testing. To enhance its robustness for the signer-independent setting, we propose a signer removal module based on feature disentanglement, effectively eliminating signer-specific information from the backbone. To validate the effectiveness of the proposed auxiliary tasks, we conduct extensive ablation studies. Notably, utilizing a transformer-based backbone, our model achieves state-of-the-art or competitive performance on five benchmarks, including PHOENIX-2014, PHOENIX-2014-T, PHOENIX-2014-SI, CSL, and CSL-Daily. Code and models are available at https://github.com/2000ZRL/LCSA_C2SLR_SRM. Ronglai Zuo, Brian Kan-Wing Mak |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Natural Language-Assisted Sign Language RecognitionabstractSign languages are visual languages which convey in-formation by signers' handshape, facial expression, body movement, and so forth. Due to the inherent restriction of combinations of these visual ingredients, there exist a significant number of visually indistinguishable signs (VISigns) in sign languages, which limits the recognition capacity of vision neural networks. To mitigate the problem, we propose the Natural Language-Assisted Sign Language Recognition (NLA-SLR) framework, which exploits semantic information contained in glosses (sign labels). First, for VISigns with similar semantic meanings, we propose language-aware label smoothing by generating soft labels for each training sign whose smoothing weights are computed from the normalized semantic similarities among the glosses to ease training. Second, for VISigns with distinct semantic meanings, we present an inter-modality mixup technique which blends vision and gloss features to further maximize the separability of different signs under the super-vision of blended labels. Besides, we also introduce a novel backbone, video-keypoint network, which not only models both RGB videos and human body keypoints but also derives knowledge from sign videos of different temporal receptive fields. Empirically, our method achieves state-of-the-art performance on three widely-adopted benchmarks: MSASL, WLASL, and NMFs-CSL. Codes are available at https://github.com/FangyunWeilSLRT. Ronglai Zuo, Fangyun Wei, Brian Kan-Wing Mak |
CVPR | 3 |
| 2023 | On the Audio-visual Synchronization for Lip-to-Speech SynthesisabstractMost lip-to-speech (LTS) synthesis models are trained and evaluated with the assumption that the audio-video pairs in the dataset are well synchronized. In this work, we demonstrate that commonly used audiovisual datasets such as GRID, TCD-TIMIT, and Lip2Wav can, however, have the data asynchrony issue, which will lead to inaccurate evaluation with conventional time alignment-sensitive metrics such as STOI, ESTOI, and MCD. Moreover, training an LTS model with such datasets can result in model asynchrony, meaning that the generated speech and input video are out of sync. To address these problems, we first provide a time-alignment frontend for the commonly used metrics to ensure accurate evaluation. Then, we propose a synchronized lip-to-speech (SLTS) model with an automatic synchronization mechanism (ASM) that corrects data asynchrony and penalizes model asynchrony during training. We evaluated the effectiveness of our approach on both artificial and popular audiovisual datasets. Our proposed method outperforms existing SOTA models in a variety of evaluation metrics. Zhe Niu, Brian Kan-Wing Mak |
ICCV | 2 |
| 2023 | wav2vec 2.0 ASR for Cantonese-Speaking Older Adults in a Clinical SettingabstractThe lack of large-scale speech corpora for Cantonese and older adults has impeded the academia's research of automatic speech recognition (ASR) systems for the two. On the other hand, the recent success of self-supervised speech representation learning has shown its competitiveness in low-resource ASR. This work therefore studies the application of wav2vec 2.0 ASR using monolingual and cross-lingual pre-trained models on a developing speech corpus, CU-MARVEL, which is dedicated to the automated screening of neurocognitive disorders (NCD) for Cantonese-speaking older adults in Hong Kong. We detail our data preparation procedures for creating a monolingual wav2vec 2.0 model from scratch and further pre-training a cross-lingual model. We report the performance of our wav2vec 2.0 ASR models on the said corpus and present a preliminary analysis of the relationship between the ASR performance of older adult speech and various demographic characteristics. Ranzo Huang, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2023 | Integrated and Enhanced Pipeline System to Support Spoken Language Analytics for Screening Neurocognitive Disordersabstract24th Annual Conference of the International Speech Communication Association, INTERSPEECH 2023, Dublin, Ireland, August 20-24, 2023 Helen M. Meng, Brian Kan-Wing Mak, Man-Wai Mak, Helene H. Fung, Xianmin Gong, Timothy C. Y. Kwok, Xunying Liu, Vincent C. T. Mok, Patrick C. M. Wong, Jean Woo, Xixin Wu, Ka-Ho Wong, Sean Shensheng Xu, Naijun Zheng, Ranzo Huang, Jiawen Kang 0002, Xiaoquan Ke, Junan Li, Jinchao Li |
INTERSPEECH | 2 |
| 2023 | Bayesian Self-Attentive Speaker Embeddings for Text-Independent Speaker VerificationabstractLearning effective and discriminative speaker embeddings is a crucial task in speaker verification. Usually, speaker embeddings are extracted from a speaker-classification DNN that averages the hidden vectors over all the spoken frames of a speaker; the hidden vectors produced from all the frames are assumed to be equally important. In our previous work, we relaxed this assumption and computed the speaker embedding as a weighted average of a speaker's frame-level hidden vectors, and their weights were automatically determined by a self-attention mechanism. The effect of multiple attention heads have also been investigated to capture different aspects of a speaker's input speech. One challenge for multi-head attention is the information redundancy problem. If there is no constraint during the training of multi-head attention, different heads may extract similar attentive features, leading to the attention redundancy problem. In this paper, we generalize the deterministic multi-head attention to a Bayesian attention framework, and provide a new understanding of multi-head attention from a Bayesian perspective. Under the Bayesian framework, we adopt the recently developed sampling method in optimization, which explicitly enforces the repulsiveness among the multiple heads. Systematic evaluation of the proposed Bayesian self-attentive speaker embeddings is performed on VoxCeleb and SITW evaluation sets. Significant and consistent improvements over other multi-head attention systems are achieved on all the evaluation datasets. The best Bayesian system with eight heads improves the EER by around 26% on VoxCeleb and 9% on SITW over the single-head baseline. Yingke Zhu, Brian Kan-Wing Mak |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Access on Demand: Real-time, Multi-modal Accessibility for the Deaf and Hard-of-Hearing based on Augmented RealityabstractIn this experience report, two deaf researchers with varying expertise, communication preferences, and technological skills document their experiences using Access on Demand (AoD), an Augmented Reality (AR) based accessibility application that provides on-demand real-time captioning and sign language interpretation services using the Vuzix Blade AR smart glasses. The researchers report their observations regarding using remote real-time American Sign Language (ASL) interpreting, captioning, and auto-captions offered by the AoD platform. The authors discuss the benefits and limitations of using AoD as an assistive technology device and how it would benefit the deaf community from the perspective of Deaf and Hard-of-Hearing (DHH) users. Roshan Mathew, Brian Kan-Wing Mak, Wendy Dannels |
ASSETS | 2 |
| 2022 | C2SLR: Consistency-enhanced Continuous Sign Language RecognitionabstractThe backbone of most deep-learning-based continuous sign language recognition (CSLR) models consists of a visual module, a sequential module, and an alignment module. However, such CSLR backbones are hard to be trained sufficiently with a single connectionist temporal classification loss. In this work, we propose two auxiliary constraints to enhance the CSLR backbones from the perspective of consistency. The first constraint aims to enhance the visual module, which easily suffers from the insufficient training problem. Specifically, since sign languages convey information mainly with signers' faces and hands, we insert a keypoint-guided spatial attention module into the visual module to enforce it to focus on informative regions, i.e., spatial attention consistency. Nevertheless, only enhancing the visual module may not fully exploit the power of the backbone. Motivated by that both the output features of the visual and sequential modules represent the same sentence, we further impose a sentence embedding consistency constraint between them to enhance the representation power of both the features. Experimental results over three representative backbones validate the effectiveness of the two constraints. More remarkably, with a transformer-based backbone, our model achieves state-of-the-art or competitive performance on three benchmarks, PHOENIX-2014, PHOENIX-2014-T, and CSL. Ronglai Zuo, Brian Kan-Wing Mak |
CVPR | 2 |
| 2022 | Synthesizing Near Native-accented Speech for a Non-native Speaker by Imitating the Pronunciation and Prosody of a Native SpeakerabstractThis paper investigates how to reduce foreign accent in the synthesis of native (L1) speech for a non-native (L2) speaker. We focus on two major aspects of foreign accents: mispronunciations and improper prosody (rhythm, phonemes duration, and pauses). Firstly, to reduce mispronunciations, the mel-spectrograms generated by an L2 text-to-speech (TTS) model are fed to a pre-trained speech recognizer and the mispronunciation information is fed back to the TTS model during back-propagation to help the model learn correct native mel-spectrograms. Secondly, to imitate L1 speech prosody, a recent data augmentation (DA) technique originally proposed for speaking style transfer is applied to transfer L1 speaking style to L2 speakers. The DA technique creates additional L2 speeches when L2 speakers try to imitate L1 speeches. Automatic speech recognition on native-accented speeches synthesized from nonnative speakers by the proposed method gives a lower word error rate. The speaker embeddings produced by a pre-trained speaker verifier from the original L2 speakers' speech and their synthesized speech are highly similar. Finally, subjective MOS scores on the synthesized speech show that they have good quality and reduced accentedness. Raymond Chung, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2022 | Local Context-aware Self-attention for Continuous Sign Language RecognitionabstractTransformer-based architectures are adopted in many continuous sign language recognition (CSLR) works for sequence modeling due to their strong capability of extracting global contexts. However, since vanilla self-attention (SA), the core module of Transformer, computes a weighted average over all time steps, the local temporal semantics of sign videos may not be fully exploited. In this work, we propose local context-aware self-attention (LCSA) to enhance the vanilla SA to leverage both local and global contexts. We introduce the local contexts at two different levels of model computation: score and query levels. At the score level, we modulate the attention scores explicitly with an additional Gaussian bias. At the query level, local contexts are modeled implicitly using depth-wise temporal convolutional networks (DTCNs). However, the vanilla Gaussian bias has two major shortcomings: first, its window size is fixed and needs to be fine-tuned laboriously; second, the fixed window size is common among all time steps. In this work, a dynamic Gaussian bias is further proposed to address the above issues. Experimental results on two benchmarks, PHOENIX-2014 and CSL, validate the effectiveness and superiority of our method. Ronglai Zuo, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2022 | Two-Stream Network for Sign Language Recognition and TranslationabstractSign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden representations. RGB videos, however, are raw signals with substantial visual redundancy, leading the encoder to overlook the key information for sign language understanding. To mitigate this problem and better incorporate domain knowledge, such as handshape and body movement, we introduce a dual visual encoder containing two separate streams to model both the raw videos and the keypoint sequences generated by an off-the-shelf keypoint estimator. To make the two streams interact with each other, we explore a variety of techniques, including bidirectional lateral connection, sign pyramid network with auxiliary supervision, and frame-level self-distillation. The resulting model is called TwoStream-SLR, which is competent for sign language recognition (SLR). TwoStream-SLR is extended to a sign language translation (SLT) model, TwoStream-SLT, by simply attaching an extra translation network. Experimentally, our TwoStream-SLR and TwoStream-SLT achieve state-of-the-art performance on SLR and SLT tasks across a series of datasets including Phoenix-2014, Phoenix-2014T, and CSL-Daily. Ronglai Zuo, Fangyun Wei, Yu Wu 0012, Shujie Liu 0001, Brian Kan-Wing Mak |
NeurIPS | 6 |
| 2021 | On-The-Fly Data Augmentation for Text-to-Speech Style TransferabstractRecent advanced text-to-speech (TTS) systems synthesize natural speeches. However, in many applications, it is desirable to synthesize utterances in a specific style. In this paper, we investigate synthesizing audios with three styles — news-casting, public speaking and storytelling — for a speaker who provides only neutral speech data. Firstly, considerable speech data were collected from the neutral speaker, and small amounts of speech from the wanted styles were collected from other speakers such that no speakers uttered in more than one style. All the data were used to train a basic multi-style multi-speaker TTS model. Secondly, augmented audios were created on-the-fly with the latest TTS model during its training and were used to further train the TTS model. Specifically, augmented data were created by ‘forcing’ a speaker to imitate stylish speeches of other three speakers by requiring their attention alignment matrices as similar as possible. Objective evaluation on the rhythm and pitch profile of the synthesized speech shows that the TTS model trained with our proposed data augmentation successfully transfers speech styles in these aspects. Subjective ABX evaluation also shows that stylish speeches synthesized by our proposed method are overwhelmingly preferred than those from a baseline TTS model by 40-60%. Raymond Chung, Brian Kan-Wing Mak |
ASRU | 2 |
| 2021 | A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease DetectionabstractWith the global population ageing rapidly, Alzheimer's disease (AD) is particularly prominent in older adults, which has an insidious onset followed by gradual, irreversible deterioration in cognitive domains (memory, communication, etc). Thus the detection of Alzheimer's disease is crucial for timely intervention to slow down disease progression. This paper presents a comparative study of different acoustic and linguistic features for the AD detection using various classifiers. Experimental results on ADReSS dataset reflect that the proposed models using ComParE, X-vector, Linguistics, TFIDF and BERT features are able to detect AD with high accuracy and sensitivity, and are comparable with the state-of-the-art results reported. While most previous work used manual transcripts, our results also indicate that similar or even better performance could be obtained using automatically recognized transcripts over manually collected ones. This work achieves accuracy scores at 0.67 for acoustic features and 0.88 for linguistic features on either manual or ASR transcripts on the ADReSS Challenge1test set. Jinchao Li, Jianwei Yu 0001, Zi Ye 0001, Simon Wong, Man-Wai Mak, Brian Kan-Wing Mak, Xunying Liu, Helen M. Meng |
ICASSP | 6 |
| 2021 | Non-Parallel Many-To-Many Voice Conversion by Knowledge Transfer from a Text-To-Speech ModelabstractIn this paper, we present a simple but novel framework to train a nonparallel many-to-many voice conversion (VC) model based on the encoder-decoder architecture. It is observed that an encoder-decoder text-to-speech (TTS) model and an encoder-decoder VC model have the same structure. Thus, we propose to pre-train a multi-speaker encoder-decoder TTS model and transfer knowledge from the TTS model to a VC model by (1) adopting the TTS acoustic decoder as the VC acoustic decoder, and (2) forcing the VC speech encoder to learn the same speaker-agnostic linguistic features from the TTS text encoder so as to achieve speaker disentanglement in the VC encoder output. We further control the conversion of the pitch contour from source speech to target speech, and condition the VC decoder on the converted pitch contour during inference. Subjective evaluation shows that our proposed model is able to handle VC between any speaker pairs in the training speech corpus of over 200 speakers with high naturalness and speaker similarity. Xinyuan Yu, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2020 | Stochastic Fine-Grained Labeling of Multi-state Sign Glosses for Continuous Sign Language Recognition
Zhe Niu, Brian Kan-Wing Mak |
ECCV (16) | 2 |
| 2020 | Orthogonal Training for Text-Independent Speaker VerificationabstractIn this paper we propose orthogonal training schemes to improve the effectiveness of cosine similarity measurements in text-independent speaker verification (SV) tasks. Compared to the PLDA backend, cosine similarity is simple to compute, and it does not require extra data or time to build a separate model. The use of cosine similarity measurement is also highly desirable for building end-to-end SV systems. However, the cosine similarity has an underlying assumption that the dimensions of the speaker embeddings are orthogonal, which is usually not satisfied in current SV systems. The training scheme applies singular vector decomposition (SVD) to the weight matrix of the speaker embedding extraction layer in our time delay neural network (TDNN)-based SV system, and replaces the original weight matrix by the matrix constructed from the left unitary matrix and the singular value matrix. Then the reconstructed matrix in the extraction layer is held constant and the remaining network is fine-tuned with an orthogonality regularizer. We further investigate orthogonal training from scratch, with orthogonality regularization incorporated throughout the network training. Experimental results show that our orthogonal training methods can significantly improve the system performance with a cosine similarity backend. Yingke Zhu, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2020 | Multi-Lingual Multi-Speaker Text-to-Speech Synthesis for Voice Cloning with Online Speaker EnrollmentabstractRecent studies in multi-lingual and multi-speaker text-to-speech synthesis proposed approaches that use proprietary corpora of performing artists and require fine-tuning to enroll new voices. To reduce these costs, we investigate a novel approach for generating high-quality speeches in multiple languages of speakers enrolled in their native language. In our proposed system, we introduce tone/stress embeddings which extend the language embedding to represent tone and stress information. By manipulating the tone/stress embedding input, our system can synthesize speeches in native accent or foreign accent. To support online enrollment of new speakers, we condition the Tacotron-based synthesizer on speaker embeddings derived from a pre-trained x-vector speaker encoder by transfer learning. We introduce a shared phoneme set to encourage more phoneme sharing compared with the IPA. Our MOS results demonstrate that the native speech in all languages is highly intelligible and natural. We also find L2-norm normalization and ZCA-whitening on x-vectors are helpful to improve the system stability and audio quality. We also find that the WaveNet performance is seemingly language-independent: the WaveNet model trained with any of the three supported languages in our system can be used to generate speeches in the other two languages very well. Copyright © 2020 ISCA Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2019 | Recurrent Poisson Process Unit for Speech RecognitionabstractOver the past few years, there has been a resurgence of interest in using recurrent neural network-hidden Markov model (RNN-HMM) for automatic speech recognition (ASR). Some modern recurrent network models, such as long shortterm memory (LSTM) and simple recurrent unit (SRU), have demonstrated promising results on this task. Recently, several scientific perspectives in the fields of neuroethology and speech production suggest that human speech signals may be represented in discrete point patterns involving acoustic events in the speech signal. Based on this hypothesis, it may pose some challenges for RNN-HMM acoustic modeling: firstly, it arbitrarily discretizes the continuous input into the interval features at a fixed frame rate, which may introduce discretization errors; secondly, the occurrences of such acoustic events are unknown. Furthermore, the training targets of RNN-HMM are obtained from other (inferior) models, giving rise to misalignments. In this paper, we propose a recurrent Poisson process (RPP) which can be seen as a collection of Poisson processes at a series of time intervals in which the intensity evolves according to the RNN hidden states that encode the history of the acoustic signal. It aims at allocating the latent acoustic events in continuous time. Such events are efficiently drawn from the RPP using a sampling-free solution in an analytic form. The speech signal containing latent acoustic events is reconstructed/sampled dynamically from the discretized acoustic features using linear interpolation, in which the weight parameters are estimated from the onset of these events. The above processes are further integrated into an SRU, forming our final model, called recurrent Poisson process unit (RPPU). Experimental evaluations on ASR tasks including ChiME-2, WSJ0 and WSJ0&1 demonstrate the effectiveness and benefits of the RPPU. For example, it achieves a relative WER reduction of 10.7% over state-of-the-art models on WSJ0. Hengguan Huang, Hao Wang 0014, Brian Kan-Wing Mak |
AAAI | 3 |
| 2019 | Mixup Learning Strategies for Text-Independent Speaker VerificationabstractMixup is a learning strategy that constructs additional virtual training samples from existing training samples by linearly interpolating random pairs of them. It has been shown that mixup can help avoid data memorization and thus improve model generalization. This paper investigates the mixup learning strategy in training speaker-discriminative deep neural network (DNN) for better text-independent speaker verification. In recent speaker verification systems, a DNN is usually trained to classify speakers in the training set. The DNN, at the same time, learns a low-dimensional embedding of speakers so that speaker embeddings can be generated for any speakers during evaluation. We adapted the mixup strategy to the speaker-discriminative DNN training procedure, and studied different mixup schemes, such as performing mixup on MFCC features or raw audio samples. The mixup learning strategy was evaluated on NIST SRE 2010, 2016 and SITW evaluation sets. Experimental results show consistent performance improvements both in terms of EER and DCF of up to 13% relative. We further find that mixup training also improves the DNN's speaker classification accuracy consistently without requiring any additional data sources. Copyright © 2019 ISCA Yingke Zhu, Tom Ko, Brian Kan-Wing Mak |
INTERSPEECH | 3 |
| 2018 | End-To-End Low-Resource Lip-Reading with Maxout Cnn and LstmabstractLip-reading is the task of recognizing speech solely from the visual movement of the mouth. Although recent works have demonstrated the effectiveness of convolutional neural network (CNN) and long short-term memory (LSTM) recurrent neural network in lip-reading, similar architectures under low-resource scenario have not yet been explored. Our proposed end - to-end deep learning model fuses conventional CNN and bidirectional LSTM (BLSTM) together with max-out activation units (maxout-CNN-BLSTM), and is capable of attaining a word accuracy of 87.6% on the Ouluvs2 corpus, offering an absolute improvement of 3.1 % to the previous state-of-the-art auto-encoder-BLSTM model. To the best of our knowledge, this is the first end - to-end low -resource lip-reading system that does not require any separate feature extraction stage nor pre-training phase with external data resources. This is also the first work that utilizes maxout units in both CNN and LSTM in one single deep neural network. Ivan Fung, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2018 | learning Effective Factorized Hidden Layer Bases Using Student-Teacher Training for LSTM Acoustic Model AdaptationabstractFactorized Hidden Layer (FHL) has been proposed for the adaptation of deep neural network (DNN) and Long Short-Term Memory (LSTM) based acoustic models (AMs). In FHL, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD transformation is a linear combination of rank- l matrices whereas the SD bias is a linear combination of vectors. However, the adaptation of LSTMs is challenging and often reports modest gains. In this paper, we propose to use student-teacher training to estimate more efficient FHL bases for LSTM AMs using an FHL adapted DNN as the teacher model. For both AMI IHM and AMI SDM tasks, FHL achieves 3.2% absolute improvement over the frame-level cross entropy trained LSTM baselines. Moreover, FHL results 3.0% and 3.8% absolute improvements over sequentially trained LSTM baselines for the AMI IHM and AMI SDM tasks respectively. Lahiru Samarakoon, Brian Kan-Wing Mak, Khe Chai Sim |
ICASSP | 2 |
| 2018 | Fast Derivation of Cross-lingual Document Vectors from Self-attentive Neural Machine Translation ModelabstractA universal cross-lingual representation of documents, which can capture the underlying semantics is very useful in many natural language processing tasks. In this paper, we develop a new document vectorization method which effectively selects the most salient sequential patterns from the inputs to create document vectors via a self-attention mechanism using a neural machine translation (NMT) model. The model used by our method can be trained with parallel corpora that are unrelated to the task at hand. During testing, our method will take a monolingual document and convert it into a “Neural machine Translation framework based cross-lingual Document Vector” (NTDV). NTDV has two comparative advantages. Firstly, the NTDV can be produced by the forward-pass of the encoder in the NMT, and the process is very fast and does not require any training/optimization. Secondly, our model can be conveniently adapted from a pair of existing attention-based NMT models, and the training requirement on parallel corpus can be reduced significantly. In a cross-lingual document classification task, our NTDV embeddings surpass the previous state-of-the-art performance in the English-to-German classification test, and, to our best knowledge, it also achieves the best performance among the fast decoding methods in the German-to-English classification test. © 2018 International Speech Communication Association. All rights reserved. Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2018 | Self-Attentive Speaker Embeddings for Text-Independent Speaker VerificationabstractThis paper introduces a new method to extract speaker embeddings from a deep neural network (DNN) for text-independent speaker verification. Usually, speaker embeddings are extracted from a speaker-classification DNN that averages the hidden vectors over the frames of a speaker; the hidden vectors produced from all the frames are assumed to be equally important. We relax this assumption and compute the speaker embedding as a weighted average of a speaker's frame-level hidden vectors, and their weights are automatically determined by a self-attention mechanism. The effect of multiple attention heads are also investigated to capture different aspects of a speaker's input speech. Finally, a PLDA classifier is used to compare pairs of embeddings. The proposed self-attentive speaker embedding system is compared with a strong DNN embedding baseline on NIST SRE 2016. We find that the self-attentive embeddings achieve superior performance. Moreover, the improvement produced by the self-attentive speaker embeddings is consistent with both short and long testing utterances. © 2018 International Speech Communication Association. All rights reserved. Yingke Zhu, Tom Ko, David Snyder, Brian Kan-Wing Mak, Daniel Povey |
INTERSPEECH | 4 |
| 2018 | Domain Adaptation of End-to-end Speech Recognition in Low-Resource SettingsabstractEnd-to-end automatic speech recognition (ASR) has simplified the traditional ASR system building pipeline by eliminating the need to have multiple components and also the requirement for expert linguistic knowledge for creating pronunciation dictionaries. Therefore, end-to-end ASR fits well when building systems for new domains. However, one major drawback of end-to-end ASR is that, it is necessary to have a larger amount of labeled speech in comparison to traditional methods. Therefore, in this paper, we explore domain adaptation approaches for end-to-end ASR in low-resource settings. We show that joint domain identification and speech recognition by inserting a symbol for domain at the beginning of the label sequence, factorized hidden layer adaptation and a domain-specific gating mechanism improve the performance for a low-resource target domain. Furthermore, we also show the robustness of proposed adaptation methods to an unseen domain, when only 3 hours of untranscribed data is available with improvements reporting upto 8.7% relative. Lahiru Samarakoon, Brian Kan-Wing Mak, Albert Y. S. Lam |
SLT | 2 |
| 2018 | DNN-Based Score Calibration With Multitask Learning for Noise Robust Speaker VerificationabstractThis paper proposes and investigates several deep neural network (DNN) based score compensation, transformation, and calibration algorithms for enhancing the noise robustness of i-vector speaker verification systems. Unlike conventional calibration methods where the required score shift is a linear function of SNR or log-duration, the DNN approach learns the complex relationship between the score shifts and the combination of i-vector pairs and uncalibrated scores. Furthermore, with the flexibility of DNNs, it is possible to explicitly train a DNN to recover the clean scores without having to estimate the score shifts. To alleviate the overfitting problem, multitask learning is applied to incorporate auxiliary information such as SNRs and speaker ID of training utterances into the DNN. Experiments on NIST 2012 SRE show that score calibration derived from multitask DNNs can improve the performance of the conventional score-shift approch significantly, especially under noisy conditions. Zhili Tan, Man-Wai Mak, Brian Kan-Wing Mak |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Denoised Senone I-Vectors for Robust Speaker VerificationabstractRecently, it has been shown that senone i-vectors, whose posteriors are produced by senone deep neural networks (DNNs), outperform the conventional Gaussian mixture model (GMM) i-vectors in both speaker and language recognition tasks. The success of senone i-vectors relies on the capability of the DNN to incorporate phonetic information into the i-vector extraction process. In this paper, we argue that to apply senone i-vectors in noisy environments, it is important to robustify the phonetically discriminative acoustic features and senone posteriors estimated by the DNN. To this end, we propose a deep architecture formed by stacking a deep belief network on top of a denoising autoencoder (DAE). After backpropagation fine-tuning, the network, referred to as denoising autoencoder-deep neural network (DAE-DNN), facilitates the extraction of robust phonetically-discriminitive bottleneck (BN) features and senone posteriors for i-vector extraction. We refer to the resulting i-vectors as denoised BN-based senone i-vectors. Results on NIST 2012 SRE show that senone i-vectors outperform the conventional GMM i-vectors. More interestingly, the BN features are not only phonetically discriminative, results suggest that they also contain sufficient speaker information to produce BN-based senone i-vectors that outperform the conventional senone i-vectors. This work also shows that DAE training is more beneficial to BN feature extraction than senone posterior estimation. Zhili Tan, Man-Wai Mak, Brian Kan-Wing Mak, Yingke Zhu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Unsupervised adaptation of student DNNS learned from teacher RNNS for improved ASR performanceabstractIn automatic speech recognition (ASR), adaptation techniques are used to minimize the mismatch between training and testing conditions. Many successful techniques have been proposed for deep neural network (DNN) acoustic model (AM) adaptation. Recently, recurrent neural networks (RNNs) have outperformed DNNs in ASR tasks. However, the adaptation of RNN AMs is challenging and in some cases when combined with adaptation, DNN AMs outperform adapted RNN AMs. In this paper, we combine student-teacher training and unsupervised adaptation to improve ASR performance. First, RNNs are used as teachers to train student DNNs. Then, these student DNNs are adapted in an unsupervised fashion. Experimental results on the AMI IHM and AMI SDM tasks show that student DNNs are adaptable with significant performance improvements for both frame-wise and sequentially trained systems. We also show that the combination of adapted DNNs with teacher RNNs can further improve the performance. Lahiru Samarakoon, Brian Kan-Wing Mak |
ASRU | 2 |
| 2017 | An investigation into learning effective speaker subspaces for robust unsupervised DNN adaptationabstractSubspace methods are used for deep neural network (DNN)-based acoustic model adaptation. These methods first construct a subspace and then perform the speaker adaptation as a point in the subspace. This paper aims to investigate the effectiveness of subspace methods for robust unsupervised adaptation. For the analysis, we compare two state-of-the-art subspace methods, namely, the singular value decomposition (SVD)-based bottleneck adaptation and the factorized hidden layer (FHL) adaptation. Both of these methods perform speaker adaptation as a linear combination of rank-1 bases. The main difference between the subspace construction is that FHL adaptation constructs a speaker subspace separate from the phoneme classification space while SVD-based bottleneck adaptation shares the same subspace for both the phoneme classification and the speaker adaptation. So far, no direct comparisons between these two methods are reported. In this work, we compare these two methods for their robustness to unsupervised adaptation on Aurora 4, AMI IHM and AMI SDM tasks. Our findings show that the FHL adaptation outperforms the SVD-based bottleneck adaptation especially in challenging conditions where the adaptation data is limited, or the quality of the adaptation alignments are low. Lahiru Samarakoon, Khe Chai Sim, Brian Kan-Wing Mak |
ICASSP | 3 |
| 2017 | Speeding up softmax computations in DNN-based large vocabulary speech recognition by senone weight vector selectionabstractDeep neural network has obtained significant accuracy improvement in many large vocabulary continuous speech recognition (LVCSR) tasks. Recently, it was shown that even better performance can be obtained by modeling a larger number of more discriminative senones. However, as the neural network becomes larger, the number of parameters increases greatly, resulting in greater computation cost and slower decoding process. Since in LVCSR systems, most DNN computations are done in the output softmax layer, we propose a senone weight vector selection method in this paper to speed up the DNN softmax computation while keeping the system accuracy more or less the same. We apply clustering on the weight vectors of the softmax layer and group all the senone weight vectors into several clusters. During decoding, we only compute the exact posteriors for senones in the selected clusters. For the senones in the unselected clusters, their posteriors are approximated using their cluster centers. Experimental results show that our speed-up method can reduce DNN computation time by more than 35% with negligible accuracy loss in a DNN model with 60,000 senones on Switchboard. Yingke Zhu, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2017 | To Improve the Robustness of LSTM-RNN Acoustic Models Using Higher-Order Feedback from Multiple HistoriesabstractThis paper investigates a novel multiple-history long short-Term memory (MH-LSTM) RNN acoustic model to mitigate the robustness problem of noisy outputs in the form of mis-labeled data and/or mis-Alignments. Conceptually, after an RNN is unfolded in time, the hidden units in each layer are re-Arranged into ordered sub-layers with a master sub-layer on top and a set of auxiliary sub-layers below it. Only the master sub-layer generates outputs for the next layer whereas the auxiliary sublayers run in parallel with the master sub-layer but with increasing time lags. Each sub-layer also receives higher-order feedback from a fixed number of sub-layers below it. As a result, each sub-layer maintains a different history of the input speech, and the ensemble of all the different histories lends itself to the model's robustness. The higher-order connections not only provide shorter feedback paths for error signals to propagate to the farther preceding hidden states to better model the long-Term memory, but also more feedback paths to each model parameter and smooth its update during training. Phoneme recognition results on both real TIMIT data as well as synthetic TIMIT data with noisy labels or alignments show that the new model outperforms the conventional LSTM RNN model. Hengguan Huang, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2017 | Learning Factorized Transforms for Unsupervised Adaptation of LSTM-RNN Acoustic ModelsabstractFactorized Hidden Layer (FHL) adaptation has been proposed for speaker adaptation of deep neural network (DNN) based acoustic models. In FHL adaptation, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD transformation is a linear combination of rank-1 matrices whereas the SD bias is a linear combination of vectors. Recently, the Long Short- Term Memory (LSTM) Recurrent Neural Networks (RNNs) have shown to outperform DNN acoustic models in many Automatic Speech Recognition (ASR) tasks. In this work, we investigate the effectiveness of SD transformations for LSTM-RNN acoustic models. Experimental results show that when combined with scaling of LSTM cell states' outputs, SD transformations achieve 2.3% and 2.1% absolute improvements over the baseline LSTM systems for the AMI IHM and AMI SDM tasks respectively. Lahiru Samarakoon, Brian Kan-Wing Mak, Khe Chai Sim |
INTERSPEECH | 2 |
| 2015 | The harp of light: a musical string projection mappingabstractThe Harp of Light is a creative installation that uses string projection mapping to simulate an aesthetic visual corresponding to beautiful sound of a harp. The installation mimics a harp with strings that receives projection mapping of digital visuals that is aligned to the live interaction from audiences through natural based interactive means. The main contribution of this artwork is an attempt to provoke an emotional experience through a highly realistic simulation of an acoustic harp sound and a hepta-symmetrical visual generated from natural interaction from the audiences creatively projected on strings. The set-up of this creative installation creates a fresh and colorful ambience that gives the audiences a digital look-and-feel of musical tone through the projected chords. A novel combination of various creative technologies has been used to create The Harp of Light. Goh Wen Shyan, Brian Kan-Wing Mak, Chee-Onn Wong, Tan Yee Lyn, Tey Zi Ming |
Advances in Computer Entertainment | 2 |
| 2015 | Distinct triphone acoustic modeling using deep neural networksabstractTo strike a balance between robust parameter estimation and detailed modeling, most automatic speech recognition systems are built using tied-state continuous density hidden Markov models (CDHMM). Consequently, states that are tied together in a tied-state are not distinguishable, introducing quantization errors inevitably. It has been shown that it is possible to model (almost) all distinct triphones effectively by using a basis approach; previously two methods were proposed: eigentriphone modeling and reference model weighting (RMW) in CDHMM using Gaussian-mixture states. In this paper, we investigate distinct triphone modeling under the state-of-the-art deep neural network (DNN) framework. Due to the large number of DNN model parameters, regularization is necessary. Multi-task learning (MTL) is first used to train distinct triphone states together with carefully chosen related tasks which serve as a regularizer. The RMW approach is then applied to linearly combine the neural network weight vectors of member triphones of each tied-state before the output softmax activation for each distinct triphone state. The method successfully improves phoneme recognition in TIMIT and word recognition in the Wall Street Journal task. Dongpeng Chen, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2015 | Multitask Learning of Deep Neural Networks for Low-Resource Speech RecognitionabstractWe propose a multitask learning (MTL) approach to improve low-resource automatic speech recognition using deep neural networks (DNNs) without requiring additional language resources. We first demonstrate that the performance of the phone models of a single low-resource language can be improved by training its grapheme models in parallel under the MTL framework. If multiple low-resource languages are trained together, we investigate learning a set of universal phones (UPS) as an additional task again in the MTL framework to improve the performance of the phone models of all the involved languages. In both cases, the heuristic guideline is to select a task that may exploit extra information from the training data of the primary task(s). In the first method, the extra information is the phone-to-grapheme mappings, whereas in the second method, the UPS helps to implicitly map the phones of the multiple languages among each other. In a series of experiments using three low-resource South African languages in the Lwazi corpus, the proposed MTL methods obtain significant word recognition gains when compared with single-task learning (STL) of the corresponding DNNs or ROVER that combines results from several STL-trained DNNs. Dongpeng Chen, Brian Kan-Wing Mak |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Joint acoustic modeling of triphones and trigraphemes by multi-task learning deep neural networks for low-resource speech recognitionabstractIt is well-known in machine learning that multitask learning (MTL) can help improve the generalization performance of singly learning tasks if the tasks being trained in parallel are related, especially when the amount of training data is relatively small. In this paper, we investigate the estimation of triphone acoustic models in parallel with the estimation of trigrapheme acoustic models under the MTL framework using deep neural network (DNN). As triphone modeling and trigrapheme modeling are highly related learning tasks, a better shared internal representation (the hidden layers) can be learned to improve their generalization performance. Experimental evaluation on three low-resource South African languages shows that triphone DNNs trained by the MTL approach perform significantly better than triphone DNNs that are trained by the single-task learning (STL) approach by ~3-13%. The MTL-DNN triphone models also outperform the ROVER result that combines a triphone STL-DNN and a trigrapheme STL-DNN. Dongpeng Chen, Brian Kan-Wing Mak, Cheung-Chi Leung, Sunil Sivadas |
ICASSP | 2 |
| 2014 | Subspace Gaussian mixture model with state-dependent subspace dimensionsabstractIn recent years, under the hidden Markov modeling (HMM) framework, the use of subspace Gaussian mixture models (SGMMs) has demonstrated better recognition performance than traditional Gaussian mixture models (GMMs) in automatic speech recognition. In state-of-the-art SGMM formulation, a fixed subspace dimension is assigned to every phone states. While a constant subspace dimension is easier to implement, it may, however, lead to overfitting or underfitting of some state models as the data is usually distributed unevenly among the states. In a later extension of SGMM, states are split to sub-states with an appropriate objective function so that the problem is eased by increasing the state-specific parameters for the underfitting state. In this paper, we propose another solution and allow each sub-state to have a different subspace dimension depending on its amount of training frames so that the state-specific parameters can be robustly estimated. Experimental evaluation on the Switchboard recognition task shows that our proposed method brings improvement to the existing SGMM training procedure. Tom Ko, Brian Kan-Wing Mak, Cheung-Chi Leung |
ICASSP | 2 |
| 2014 | Joint sequence training of phone and grapheme acoustic model based on multi-task learning deep neural networksabstractMulti-task learning (MTL) can be an effective way to improve the generalization performance of singly learning tasks if the tasks are related, especially when the amount of training data is small. Our previous work applied MTL to the joint training of triphone and trigrapheme acoustic models using deep neural networks (DNNs) for low-resource speech recognition. Significant recognition improvement over the performance of their DNNs trained by single-task learning (STL) was obtained. In that work, both STL-DNNs and MTL-DNNs were trained by minimizing the total frame-wise cross entropies. Since phoneme and grapheme recognition are inherently sequence classification tasks, here we study the effect of sequencediscriminative training on their joint estimation using MTLDNNs. Experimental evaluation on TIMIT phoneme recognition shows that joint sequence training outperforms frame-wise training of phone and grapheme MTL-DNNs significantly. Dongpeng Chen, Brian Kan-Wing Mak, Sunil Sivadas |
INTERSPEECH | 2 |
| 2014 | Eigentrigraphemes for under-resourced languages
Tom Ko, Brian Kan-Wing Mak |
Speech Commun. | 2 |
| 2013 | Distinct triphone modeling by reference model weightingabstractState tying effectively strikes a balance between detailed modeling and robust parameter estimation for hidden Markov models (HMMs) in automatic speech recognition. However, triphone HMMs that are tied to the same state are not distinguishable in that state. Recently we proposed the idea of distinct acoustic modeling in which no states are tied. In our novel clustered-based eigentriphone modeling method, triphones (or states) are grouped into non-overlapping clusters, from each of which, an orthogonal eigenbasis is derived using weighted PCA. Then all member triphones (or states) of a cluster are projected as distinct points onto the space spanned by its eigenvectors. In this paper, we propose a new simpler training method called reference model weighting (RMW) which removes the requirement of an orthogonal basis in eigentriphone, and directly uses a set of reference model vectors in a cluster as the basis. All member model vectors are then constrained to lie in the space spanned by these reference model vectors. The difference between eigentriphone modeling and reference model weighting is analogous to the difference between eigenvoice and reference speaker weighting in speaker adaptation. The new RMW method shows consistently better performance than eigentriphone and the baseline tied-state HMMs in WSJ0 word recognition and TIMIT phoneme recognition. Dongpeng Chen, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2013 | Eigentriphones for Context-Dependent Acoustic ModelingabstractMost automatic speech recognizers employ tied-state triphone hidden Markov models (HMM), in which the corresponding triphone states of the same base phone are tied. State tying is commonly performed with the use of a phonetic regression class tree which renders robust context-dependent modeling possible by carefully balancing the amount of training data with the degree of tying. However, tying inevitably introduces quantization error: triphones tied to the same state are not distinguishable in that state. Recently we proposed a new triphone modeling approach called eigentriphone modeling in which all triphone models are, in general, distinct. The idea is to create an eigenbasis for each base phone (or phone state) and all its triphones (or triphone states) are represented as distinct points in the space spanned by the basis. We have shown that triphone HMMs trained using model-based or state-based eigentriphones perform at least as well as conventional tied-state HMMs. In this paper, we further generalize the definition of eigentriphones over clusters of acoustic units. Our experiments on TIMIT phone recognition and the Wall Street Journal 5K-vocabulary continuous speech recognition show that eigentriphones estimated from state clusters defined by the nodes in the same phonetic regression class tree used in state tying result in further performance gain. Tom Ko, Brian Kan-Wing Mak |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Derivation of eigentriphones by weighted principal component analysisabstractLast year we proposed a new acoustic modeling method called eigentriphones in which all triphones are distinct (with no tied states) so that they may be more discriminative. In our method, frequent triphones are used to derive an eigenbasis using PCA, and the infrequent triphones are then “adapted” as a linear combination of the eigenvectors which are also called eigentriphones. Although the eigentriphones method compares favorably with traditional tied-state triphones, the PCA procedure has two limitations: (1) only the frequent triphones are employed, and (2) they are considered “equal” even though some are more robust than the others. In this paper, weighted PCA is proposed to solve both problems so that all triphones-frequent and infrequent triphones-may contribute to the derivation of the eigentriphones, each at a different extent depending on its sample count. Experimental evaluation on the WSJ 5Kvocabulary speech recognition task shows that weighted PCA produces better models than simple PCA, and its performance is fairly independent of the number of eigentriphones once more than 20% of them are used. As a consequence, all triphones may be represented by fewer eigentriphones, resulting in a more compact model. Tom Ko, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2012 | Transition probabilities are more important than we once thoughtabstractIt is generally believed that the transition probabilities in a hidden Markov model (HMM) have a limited role in the speech decoding process. In this paper, through a series of recognition experiments on Wall Street Journal (WSJ) read speech and SVitchboard (SVB) conversational telephone speech, we find that the HMM transition probabilities may be more important than we once thought. The experiments include: (1) setting or not setting all outgoing transition probabilities equal; (2) the introduction of word-final triphones and the re-estimation of their transition probabilities; (3) besides grammar factor and insertion penalty, the addition of a third decoding parameter called transition factor to scale the transition probability score during decoding. The results of the above three experiments enable us to improve the the word accuracy of the WSJ and SVB speech recognition task by 0.7% and 5.3% absolute respectively when compared to their baseline model in which all transition probabilities are simply set to 0.5. Guoli Ye, Dongpeng Chen, Brian Kan-Wing Mak |
ICASSP | 3 |
| 2011 | Eigentriphones: A basis for context-dependent acoustic modelingabstractIn context-dependent acoustic modeling, it is important to strike a balance between detailed modeling and data sufficiency for robust estimation of model parameters. In the past, parameter sharing or tying is one of the most common techniques to solve the problem. In recent years, another technique which may be loosely and collectively called the subspace approach tries to express a phonetic or sub-phonetic unit in terms of a small set of canonical vectors or units. In this paper, we investigate the development of an eigenbasis over the triphones and model each triphone as a point in the basis. We call the eigenvectors in the basis eigentriphones. From another perspective, we investigate the use of the eigenvoice adaptation method as a general acoustic modeling method for training triphones - especially the less frequent triphones without tying their states so that all the triphones are really distinct from each other and thus may be more discriminative. Experimental evaluation on the 5K-vocabulary HUB2 recognition task shows that a triphone HMM system trained using only eigentriphones without state tying may achieve slightly better performance than the common tied-state triphones. Tom Ko, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2011 | A Fully Automated Derivation of State-Based Eigentriphones for Triphone Modeling with No Tied States Using RegularizationabstractRecently we proposed an alternative method called eigentriphone to solve the data insufficiency problem in triphone acoustic modeling without the need of state tying. The idea is to treat the acoustic modeling problem of infrequent triphones ("poor triphones") as an adaptation problem from the more frequent triphones ("rich triphones"): firstly, an eigenbasis is developed over the rich triphones that have sufficient training data and the eigenvectors are called eigentriphones; then the poor triphones are adapted in a fashion similar to eigenvoice adaptation. Since, in general, no states are tied in our method, all triphones (states) are distinct so that they can be more discriminative than tied-state triphones. In our previous work, the number of eigentriphones was determined in advance with a set of development data. In this paper, we investigate simply using all of them with the help of regularization to naturally penalize the less important ones. In addition, the model-based eigenbasis is replaced by three state-based eigenbases. Experimental evaluation on the WSJ 5K task shows that triphone models trained using our new eigentriphone approach without state tying perform at least as well as the common tied-state triphone models. Tom Ko, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2010 | Improving speech recognition by explicit modeling of phone deletionsabstractIn a paper published by Greenberg in 1998, it was said that in conversational speech, phone deletion rate may go as high as 12% whereas syllable deletion rate is about 1%. The finding prompted a new research direction of syllable modeling for speech recognition. To date, the syllable approach has not yet fulfilled its promise. On the other hand, there were few attempts to model phone deletions explicitly in current ASR systems. In this paper, fragmented word models were derived from well-trained cross-word triphone models, and phone deletion was implemented by skip arcs for words consisting of at least four phonemes. An evaluation on CSR-II WSJ1 Hub2 5K task shows that even with this limited implementation of phone deletions in read speech, we obtained a word error rate reduction of 6.73%. Tom Ko, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2010 | The use of subvector quantization and discrete densities for fast GMM computation for speaker verificationabstractLast year, we showed that the computation of a GMM-UBM-based speaker verification (SV) system may be sped up by 30 times by using a high-density discrete model (HDDM) on the NIST 2002 evaluation task. The speedup was obtained using a special case of the product-code vector quantization in which each dimension is scalar-quantized in the construction of the discrete model. However, the speedup resulted in a drop of an absolute 1.5% in equal-error rate (EER). In this paper, our previous work is generalized to the use of subvector quantization (SVQ) in the construction of HDDM. For the same NIST 2002 SV task, the use of SVQ leads to an overall speedup by a factor of 8-25 with no significant loss in EER performance. Guoli Ye, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2009 | Automatic estimation of decoding parameters using large-margin iterative linear programmingabstractThe decoding parameters in automatic speech recognition - grammar factor and word insertion penalty - are usually determined by performing a grid search on a development set. Recently, we cast their estimation as a convex optimization problem, and proposed a solution using an iterative linear programming algorithm. However, the solution depends on how well the development data set matches with the test set. In this paper, we further investigates an improvement on the generalization property of the solution by using large margin training within the iterative linear programming framework. Empirical evaluation on the WSJ0 5K speech recognition tasks shows that the recognition performance of the decoding parameters found by the improved algorithm using only a subset of the acoustic model training data is even better than that of the decoding parameters found by grid search on the development data, and is close to the performance of those found by grid search on the test set. Brian Kan-Wing Mak, Tom Ko |
INTERSPEECH | 1 |
| 2009 | Fast GMM computation for speaker verification using scalar quantization and discrete densitiesabstractMost of current state-of-the-art speaker verification (SV) sys-tems use Gaussian mixture model (GMM) to represent the uni-versal background model (UBM) and the speaker models (SM). For an SV system that employs log-likelihood ratio between SM and UBM to make the decision, its computational effi-ciency is largely determined by the GMM computation. This paper attempts to speedup GMM computation by converting a continuous-density GMM to a single or a mixture of discrete densities using scalar quantization. We investigated a spectrum of such discrete models: from high-density discrete models to discrete mixture models, and their combination called high-density discrete-mixture models. For the NIST 2002 SV task, we obtained an overall speedup by a factor of 2–100 with little loss in EER performance. Index Terms: speaker verification, scalar quantization, high density discrete HMM, discrete mixture HMM Guoli Ye, Brian Kan-Wing Mak, Man-Wai Mak |
INTERSPEECH | 2 |
| 2009 | Maximum Penalized Likelihood Kernel Regression for Fast AdaptationabstractThis paper proposes a nonlinear generalization of the popularmaximum-likelihoodlinearregression(MLLR) adaptation algorithm using kernel methods. The proposed method, calledmaximumpenalizedlikelihoodkernelregressionadaptation (MPLKR), applies kernel regression with appropriate regularization to determine the affine model transform in a kernel-induced high-dimensional feature space. Although this is not the first attempt of applying kernel methods to conventional linear adaptation algorithms, unlike most of other kernelized adaptation methods such as kernel eigenvoice or kernel eigen-MLLR, MPLKR has the advantage that it is a convex optimization and its solution is always guaranteed to be globally optimal. In fact, the adapted Gaussian means can be obtained analytically by simply solving a system of linear equations. From the Bayesian perspective, MPLKR can also be considered as the kernel version ofmaximumaposteriorilinearregression(MAPLR) adaptation. Supervised and unsupervised speaker adaptation using MPLKR were evaluated on the Resource Management and Wall Street Journal 5K tasks, respectively, achieving a word error rate reduction of 23.6% and 15.5% respectively over the speaker-independently model. Brian Kan-Wing Mak, Tsz-Chung Lai, Ivor W. Tsang, James T. Kwok |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Discriminative training by iterative linear programming optimizationabstractIn this paper, we cast discriminative training problems into standard linear programming (LP) optimization. Besides being convex and having globally optimal solution(s), LP programs are well-studied with well-established solutions, and efficient LP solvers are freely available. In practice, however, one may not have complete knowledge of the feasible region since it is constructed from a limited number of competing hypotheses based on the current model - not the final model which, by definition, is not known a priori at the time of hypotheses generation. We investigate an iterative LP optimization algorithm in which an additional constraint on the parameters being optimized is further imposed. Our proposed method is evaluated on the estimation of global and state-dependent stream weights and biases of a multi-stream hidden Markov model system. Results show that the stream weights and biases found by our iterative LP optimization algorithm may give better recognition performance than the ones found by a brute-force grid search. Brian Kan-Wing Mak, Benny Ng |
ICASSP | 1 |
| 2008 | Robust speaker verification using short-time frequency with long-time window and fusion of multi-resolutionsabstractThis study presents a novel approach of feature analysis to speaker verification. There are two main contributions in this paper. First, the feature analysis of short-time frequency with long-time window (SFLW) is a compact feature for the efficiency of speaker verification. The purpose of SFLW is to take account of short-time frequency characteristics and longtime resolution at the same time. Secondly, the fusion of multi-resolutions is used for the effectiveness of robust speaker verification. The speaker verification system can be further improved using multi-resolution features. The experimental results indicate that the proposed approaches not only speed up the processing time but also improve the performance of speaker verification. Chien-Lin Huang, Bin Ma 0001, Chung-Hsien Wu 0001, Brian Kan-Wing Mak, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2008 | Min-max discriminative training of decoding parameters using iterative linear programmingabstractIn automatic speech recognition, the decoding parameters - grammar factor and word insertion penalty-are usually hand-tuned to give the best recognition performance. This paper investigates an automatic procedure to determine their values using an iterative linear programming (LP) algorithm. LP naturally implements discriminative training by mapping linear discriminants into LP constraints. A min-max cost function is also defined to get more stable and robust result. Empirical evaluations on the RM1 and WSJ0 speech recognition tasks show that decoding parameters found by the proposed algorithm are as good as those found by a brute-force grid search; their optimal values also seem to be independent of the initial values set to start the iterative LP algorithm. Brian Kan-Wing Mak, Tom Ko |
INTERSPEECH | 1 |
| 2007 | Robustness of several kernel-based fast adaptation methods on noisy LVCSRabstractWe have been investigating the use of kernel methods to im-prove conventional linear adaptation algorithms for fast adap-tation, when there are less than 10s of adaptation speech. On clean speech, we had shown that our new kernel-based adap-tation methods, namely, embedded kernel eigenvoice (eKEV) and kernel eigenspace-based MLLR (KEMLLR) outperformed their linear counterparts. In this paper, we study their unsu-pervised adaptation performance under additive and convoluted noises using the Aurora4 Corpus, with no assumption or prior knowledge of the noise type and its level. It is found that both eKEV and KEMLLR adaptation continue to outperform MAP and MLLR, and the simple reference speaker weighting (RSW) algorithm continues to perform favorably with KEMLLR. Fur-thermore, KEMLLR adaptation gives the greatest overall im-provement over the speaker-independent model by about 19%. Index Terms: fast adaptation, kernel method, kernel eigenspace-based MLLR, embedded kernel eigenvoice, MAP, Brian Kan-Wing Mak, Roger Hsiao |
INTERSPEECH | 1 |
| 2007 | A model-based estimation of phonotactic language verification performanceabstractOne of the most common approaches in language verification (LV) is the phonotactic language verification. Currently, LV performances for different languages under different environments and durations have to be compared experimentally and this can make it difficult to understand LV performances across corpora or durations. LV can be viewed as a special case of hypothesis testing such that Neyman-Pearson theorem and other information theoretic analysis are applicable. In this paper, we introduce a measure of phonotactic confusablity based on the phonotactic distribution, and make it possible to assess the difficulty of the verification problem analytically. We then propose a method of predicting LV performance. The effectiveness of the proposed approach is demonstrated on the NIST 2003 language recognition evaluation test set. Kakeung Wong, Man-Hung Siu, Brian Kan-Wing Mak |
INTERSPEECH | 3 |
| 2007 | Boosting with anti-models for automatic language identificationabstractIn this paper, we adopt the boosting framework to improve the performance of acoustic-based Gaussian mixture model (GMM) Language Identification (LID) systems. We introduce a set of low-complexity, boosted target and anti-models that are estimated from training data to improve class separation, and these models are integrated during the LID backend process. This results in a fast estimation process. Experiments were performed on the 12-language, NIST 2003 language recognition evaluation classification task using a GMM-acoustic-score- only LID system, as well as the one that combines GMM acoustic scores with sequence language model scores from GMM tokenization. Classification errors were reduced from 18.8% to 10.5% on the acoustic-score-only system, and from 11.3% to 7.8% on the combined acoustic and tokenization system. Man-Hung Siu, Herbert Gish, Brian Kan-Wing Mak |
INTERSPEECH | 4 |
| 2007 | Kernel Eigenspace-Based MLLR AdaptationabstractIn this paper, we propose an application of kernel methods for fast speaker adaptation based on kernelizing the eigenspace-based maximum-likelihood linear regression adaptation method. We call our new method "kernel eigenspace-based maximum-likelihood linear regression adaptation" (KEMLLR). In KEMLLR, speaker-dependent (SD) models are estimated from a common speaker-independent (SI) model using MLLR adaptation, and the MLLR transformation matrices are mapped to a kernel-induced high-dimensional feature space, wherein kernel principal component analysis is used to derive a set of eigenmatrices. In addition, a composite kernel is used to preserve row information in the transformation matrices. A new speaker's MLLR transformation matrix is then represented as a linear combination of the leading kernel eigenmatrices, which, though exists only in the feature space, still allows the speaker's mean vectors to be found explicitly. As a result, at the end of KEMLLR adaptation, a regular hidden Markov model (HMM) is obtained for the new speaker and subsequent speech recognition is as fast as normal HMM decoding. KEMLLR adaptation was tested and compared with other adaptation methods on the Resource Management and Wall Street Journal tasks using 5 or 10 s of adaptation speech. In both cases, KEMLLR adaptation gives the greatest improvement over the SI model with 11%-20% word error rate reduction Brian Kan-Wing Mak, Roger Hsiao |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | A Comparison of Various Adaptation Methods for Speaker Verification With Limited Enrollment DataabstractOne key factor that hinders the widespread deployment of speaker verification technologies is the requirement of long enrollment utterances to guarantee low error rate during verification. To gain user acceptance of speaker verification technologies, adaptation algorithms that can enroll speakers with short utterances are highly essential. To this end, this paper applies kernel eigenspace-based MLLR (KEMLLR) for speaker enrollment and compares its performance against three state-of-the-art model adaptation techniques: maximum a posteriori (MAP), maximum-likelihood linear regression (MLLR), and reference speaker weighting (RSW). The techniques were compared under the NIST2001 SRE framework, with enrollment data vary from 2 to 32 seconds. Experimental results show that KEMLLR is most effective for short enrollment utterances (between 2 to 4 seconds) and that MAP performs better when long utterances (32 seconds) are available.*This work was supported by the Research Grant Council of the Hong Kong SAR (Project Nos. CUHK 1/02C and PolyU 5214/04E).†Roger completed this work while he was with the Hong Kong University of Science and Technology before he left for CMU.‡This research is partially supported by the Research Grants Council of the Hong Kong SAR under the grant number CA02/03.EG04. Man-Wai Mak, Roger Hsiao, Brian Kan-Wing Mak |
ICASSP (1) | 3 |
| 2006 | Improving Reference Speaker Weighting Adaptation by the Use of Maximum-Likelihood Reference SpeakersabstractWe would like to revisit a simple fast adaptation technique called reference speaker weighting (RSW). RSW is similar to eigenvoice (EV) adaptation, and simply requires the model of a new speaker to lie on the span of a set of reference speaker vectors. In the original RSW, the reference speakers are computed through a hierarchical speaker clustering (HSC) algorithm using information such as the gender and speaking rate. We show in this paper that RSW adaptation may be improved if those training speakers that have the highest likelihoods of the adaptation data are selected as the reference speakers; we call them the maximum-likelihood (ML) reference spers.ak- When RSW adaptation was evaluated on WSJ0 using 5s of adaptation speech, the word error rate reduction can be boosted from 2.54% to 9.15% by using 10 ML reference speakers instead of reference speakers determined from HSC. Moreover, when compared with EV, MAP, MLLR, and eKEV on fast adaptation, we are surprised that the algorithmically simplest RSW technique actually gives the best performance. Brian Kan-Wing Mak, Tsz-Chung Lai, Roger Hsiao |
ICASSP (1) | 1 |
| 2006 | Fast Speaker Adaption Via Maximum Penalized Likelihood Kernel RegressionabstractMaximum likelihood linear regression (MLLR) has been a popular speaker adaptation method for many years. In this paper, we investigate a generalization of MLLR using nonlinear regression. Specifically, kernel regression is applied with appropriate regularization to determine the transformation matrix in MLLR for fast speaker adaptation. The proposed method, called maximum penalized likelihood kernel regression adaptation (MPLKR), is computationally simple and the mean vectors of the speaker adapted acoustic model can be obtained analytically by simply solving a linear system. Since no nonlinear optimization is involved, the obtained solution is always guaranteed to be globally optimal. The new adaptation method was evaluated on the resource management task with 5s and 10s of adaptation speech. Results show that MPLKR outperforms the standard MLLR method Ivor W. Tsang, James T. Kwok, Brian Kan-Wing Mak, Kai Zhang 0001, Jeffrey Junfeng Pan |
ICASSP (1) | 3 |
| 2006 | Joint Optimization of the Frequency-Domain and Time-Domain Transformations in Deriving Generalized Static and Dynamic MFCCsabstractTraditionally, static mel-frequency cepstral coefficients (MFCCs) are derived by discrete cosine transformation (DCT), and dynamic MFCCs are derived by linear regression. Their derivation may be generalized as a frequency-domain transformation of the log filter-bank energies (FBEs) followed by a time-domain transformation. In the past, these two transformations are usually estimated or optimized separately. In this letter, we consider sequences of log FBEs as a set of spectrogram images and investigate an image compression technique to jointly optimize the two transformations so that the reconstruction error of the spectrogram images is minimized; there is an efficient algorithm that solves the optimization problem. The framework allows extension to other optimization costs as well Yiu-Pong Lai, Man-Hung Siu, Brian Kan-Wing Mak |
IEEE Signal Process. Lett. | 3 |
| 2006 | Minimization of Utterance Verification Error Rate as a Constrained Optimization ProblemabstractSince utterance verification (UV) may be treated as a two-class classification problem, it may be improved with discriminative training such as minimum verification error training or minimum verification error rate training. However, since in practice, one usually has to pick a specific false-acceptance or false-rejection rate for one's system, it is more desirable to optimize UV performance at a particular operating point. In this letter, we show that further improvement can be achieved by treating UV at a specific operating point as a constrained optimization problem Man-Hung Siu, Brian Kan-Wing Mak, Wing-Hei Au |
IEEE Signal Process. Lett. | 2 |
| 2006 | Embedded kernel eigenvoice speaker adaptation and its implication to reference speaker weightingabstractRecently, we proposed an improvement to the conventional eigenvoice (EV) speaker adaptation using kernel methods. In our novel kernel eigenvoice (KEV) speaker adaptation, speaker supervectors are mapped to a kernel-induced high dimensional feature space, where eigenvoices are computed using kernel principal component analysis. A new speaker model is then constructed as a linear combination of the leading eigenvoices in the kernel-induced feature space. KEV adaptation was shown to outperform EV, MAP, and MLLR adaptation in a TIDIGITS task with less than 10 s of adaptation speech. Nonetheless, due to many kernel evaluations, both adaptation and subsequent recognition in KEV adaptation are considerably slower than conventional EV adaptation. In this paper, we solve the efficiency problem and eliminate all kernel evaluations involving adaptation or testing observations by finding an approximate pre-image of the implicit adapted model found by KEV adaptation in the feature space; we call our new method embedded kernel eigenvoice (eKEV) adaptation. eKEV adaptation is faster than KEV adaptation, and subsequent recognition runs as fast as normal HMM decoding. eKEV adaptation makes use of multidimensional scaling technique so that the resulting adapted model lies in the span of a subset of carefully chosen training speakers. It is related to the reference speaker weighting (RSW) adaptation method that is based on speaker clustering. Our experimental results on Wall Street Journal show that eKEV adaptation continues to outperform EV, MAP, MLLR, and the original RSW method. However, by adopting the way we choose the subset of reference speakers for eKEV adaptation, we may also improve RSW adaptation so that it performs as well as our eKEV adaptation. Brian Kan-Wing Mak, Roger Hsiao, Simon Ka-Lung Ho, James T. Kwok |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Kernel Eigenspace-based MLLR Adaptation Using Multiple Regression ClassesabstractRecently, we have been investigating the application of kernel methods to improve the performance of eigenvoice-based adaptation methods by exploiting possible nonlinearity in their original working space. We proposed the kernel eigenvoice adaptation (KEV), and the kernel eigenspace-based MLLR adaptation (KEMLLR). In KEMLLR, speaker-dependent MLLR transformation matrices are mapped to a kernel-induced high dimensional feature space, and kernel principal component analysis (KPCA) is used to derive a set of eigenmatrices in the feature space. A new speaker is then represented by a linear combination of the leading eigenmatrices. In this paper, we further improve KEMLLR by the use of multiple regression classes and the quasi-Newton BFGS optimization algorithm. Roger Hsiao, Brian Kan-Wing Mak |
ICASSP (1) | 2 |
| 2005 | Various Reference Speakers Determination Methods for Embedded Kernel Eigenvoice Speaker AdaptationabstractRecently, we proposed two improvements to the eigenvoice (EV) speaker adaptation using kernel methods: kernel eigenvoice (KEV) speaker adaptation, and embedded kernel eigenvoice (eKEV) speaker adaptation. In both KEV and eKEV adaptation methods, kernel eigenvoices are computed using kernel PCA, and an implicit speaker adapted model is defined as a linear combination of the leading kernel eigenvoices in the kernel-induced feature space. eKEV adaptation further finds an approximate pre-image of the implicit speaker adapted model so that all online kernel evaluations involving any acoustic vectors are eliminated during adaptation and subsequent recognition. The pre-image finding algorithm is cast as a constrained optimization problem using the distances between the expected pre-image and a set of pre-determined reference speakers as constraints. In this paper, we investigate two different ways to determine the reference speakers and the effect of their numbers on the eKEV adaptation performance. Brian Kan-Wing Mak, Simon Ka-Lung Ho |
ICASSP (1) | 1 |
| 2005 | A comparative study of two kernel eigenspace-based speaker adaptation methods on large vocabulary continuous speech recognitionabstractEigenvoice (EV) speaker adaptation has been shown effective for fast speaker adaptation when the amount of adaptation data is scarce. In the past two years, we have been investigating the application of kernel methods to improve EV speaker adaptation by exploiting possible nonlinearity in the speaker space, and two methods were proposed: embedded kernel eigenvoice (eKEV) and kernel eigenspace-based MLLR (KEMLLR). In both methods, kernel PCA is used to derive eigenvoices in the kernel-induced high-dimensional feature space, and they differ mainly in the representation of the speaker models. Both had been shown to outperform all other common adaptation methods when the amount of adaptation data is less than 10s. However, in the past, only small vocabulary speech recognition tasks were tried since we were not familiar with the behaviour of these kernelized methods. As we gain more experience, we are now ready to tackle larger vocabularies. In this paper, we show that both methods continue to outperform MAP, and MLLR when only 5s or 10s of adaptation data are available on the WSJ0 5K-vocabulary task. Compared with the speaker-independent model, the two methods reduce recognition word error rate by 13.4%-21.1%. Roger Hsiao, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2005 | High-density discrete HMM with the use of scalar quantization indexingabstractWith the advance in semiconductor memory and the availability of very large speech corpora (of hundreds to thousands of hours of speech), we would like to revisit the use of discrete hidden Markov model (DHMM) in automatic speech recognition. To estimate the discrete density in a DHMM state, the acoustic space is divided into bins and one simply count the relative amount of observations falling into each bin. With a very large speech corpus, we believe that the number of bins may be greatly increased to get a much higher density than before, and we will call the new models, the high-density discrete hidden Markov model (HDDHMM). Our HDDHMM is different from traditional DHMM in two aspects: firstly, the codebook will have a size in thousands or even tens of thousands; secondly, we propose a method based on scalar quantization indexing so that for a d-dimensional acoustic vector, the discrete codeword can be determined in O(d) time. During recognition, the state probability is reduced to an O(1) table look-up. The new HDDHMM was tested on WSJO with 5K vocabulary. Compared with a baseline 4-stream continuous density HMM system which has a WER of 9.71 %, a 4-stream HDDHMM system converted from the former achieves a WER of 11.60%, with no distance or Gaussian computation. Brian Kan-Wing Mak, Jeff Siu-Kei Au-Yeung, Yiu-Pong Lai, Man-Hung Siu |
INTERSPEECH | 1 |
| 2005 | Pruning Hidden Markov Models With Optimal Brain SurgeonabstractA method of pruning hidden Markov models (HMMs) is presented. The main purpose is to find a good HMM topology for a given task with improved generalization capability. As a side effect, the resulting model will also save memory and computation costs. The first goal falls into the active research area of model selection. From the model-theoretic research community, various measures such as Bayesian information criterion, minimum description length, minimum message length have been proposed and used with some success. In this paper, we are considering another approach in which a well-performed HMM, though perhaps oversized, is optimally pruned so that the loss in the model training cost function is minimal. The method is known as optimal brain surgeon (OBS) that has been applied to pruning neural networks (NNs) in the past. In this paper, the OBS algorithm is modified to prune HMMs. While the application of OBS to NNs is a constrained optimization problem with only equality constraints that can be solved by Lagrange multipliers, its application to HMMs requires significant modifications, resulting in a quadratic programming problem with both equality and inequality constraints. The detailed formulation of pruning an HMM with OBS is presented. It was evaluated by two experiments: one simulation using a discrete HMM, and another with continuous density HMMs trained for the TIDIGITS task. It is found that our novel OBS algorithm was able to "re-discover" the true topology of the discrete HMM in the first simulation experiment; in the second speech recognition experiment, up to about 30% of HMM transitions were successfully pruned, and yet the reduced models gave better generalization performance on unseen test data. Brian Kan-Wing Mak, Kin-Wah Chan |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Kernel Eigenvoice Speaker AdaptationabstractEigenvoice-based methods have been shown to be effective for fast speaker adaptation when only a small amount of adaptation data, say, less than 10 s, is available. At the heart of the method is principal component analysis (PCA) employed to find the most important eigenvoices. In this paper, we postulate that nonlinear PCA using kernel methods may be even more effective. The eigenvoices thus derived will be called kernel eigenvoices (KEV), and we will call our new adaptation method kernel eigenvoice speaker adaptation. However, unlike the standard eigenvoice (EV) method, an adapted speaker model found by the kernel eigenvoice method resides in the high-dimensional kernel-induced feature space, which, in general, cannot be mapped back to an exact preimage in the input speaker supervector space. Consequently, it is not clear how to obtain the constituent Gaussians of the adapted model that are needed for the computation of state observation likelihoods during the estimation of eigenvoice weights and subsequent decoding. Our solution is the use of composite kernels in such a way that state observation likelihoods can be computed using only kernel functions without the need of a speaker-adapted model in the input supervector space. In this paper, we investigate two different composite kernels for KEV adaptation: direct sum kernel and tensor product kernel. In an evaluation on the TIDIGITS task, it is found that KEV speaker adaptation using both forms of composite Gaussian kernels are equally effective, and they outperform a speaker-independent model and adapted models found by EV, MAP, or MLLR adaptation using 2.1 and 4.1 s of speech. For example, with 2.1 s of adaptation data, KEV adaptation outperforms the speaker-independent model by 27.5%, whereas EV, MAP, or MLLR adaptation are not effective at all. Brian Kan-Wing Mak, James T. Kwok, Simon Ka-Lung Ho |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Discriminative feature transformation by guided discriminative trainingabstractIn this paper, we investigate guided discriminative training in the context of improving multi-class classification problems. We are interested in applications that require improvement in the classification performance of only a subset of the classes at the possible expense of poorer classification performance of the remaining classes. However, should the classification of the remaining classes deteriorate, it is guaranteed not to be worse than the extent that the user specifies. The problem is formulated as a nonlinear programming problem, which can be translated to a unconstrained nonlinear optimization problem using the barrier method that, in turn, can be solved by the gradient descent method. To prove the concept, we apply guided discriminative training to derive an optimal linear transformation on the mel-filterbank log power spectra to improve TIMIT phoneme classification. Encouraging results are obtained. Roger Hsiao, Brian Kan-Wing Mak |
ICASSP (1) | 2 |
| 2004 | A study of various composite kernels for kernel eigenvoice speaker adaptationabstractEigenvoice-based methods have been shown to be effective for fast speaker adaptation when the amount of adaptation data is small, say, less than 10 seconds. In traditional eigenvoice (EV) speaker adaptation, linear principal component analysis (PCA) is used to derive the eigenvoices. Recently, we proposed that eigenvoices found by nonlinear kernel PCA could be more effective, and the eigenvoices thus derived were called kernel eigenvoices (KEV). One of our novelties is the use of composite kernel that makes it possible to compute state observation likelihoods via kernel functions. We investigate two different composite kernels: direct sum kernel and tensor product kernel for KEV adaptation. In an evaluation on the TIDIGITS task, it is found that KEV speaker adaptations using either form of composite kernel are equally effective, and they outperform a speaker-independent model and the adapted models from EV, MAP, or MLLR adaptation using 2.1s and 4.1s of speech. For example, with 2.1s of adaptation data, KEV adaptation outperforms the speaker-independent model by 27.5%, whereas EV, MAP, and MLLR adaptations are not effective at all. Brian Kan-Wing Mak, James T. Kwok, Simon Ka-Lung Ho |
ICASSP (1) | 1 |
| 2004 | Improving eigenspace-based MLLR adaptation by kernel PCAabstractEigenspace-based MLLR (EMLLR) adaptation has been shown effective for fast speaker adaptation. It applies the basic idea of eigenvoice adaptation, and derives a small set of eigenmatrices using principal component analysis (PCA). The MLLR adapta-tion transformation of a new speaker is then a linear combina-tion of the eigenmatrices. In this paper, we investigate the use of kernel PCA to find the eigenmatrices in the kernel-induced high dimensional feature space so as to exploit possible nonlinearity in the transformation supervector space. In addition, composite kernel is used to preserve the row information in the transfor-mation supervector which, otherwise, will be lost during the mapping to the kernel-induced feature space. We call our new method kernel eigenspace-based MLLR (KEMLLR) adaptation. On a RM adaptation task, we find that KEMLLR adaptation may reduce the word error rate of a speaker-independent model by 11%, and outperforms MLLR and EMLLR adaptation. 1. Brian Kan-Wing Mak, Roger Hsiao |
INTERSPEECH | 1 |
| 2004 | Speedup of kernel eigenvoice speaker adaptation by embedded kernel PCAabstractRecently, we proposed an improvement to the eigenvoice (EV) speaker adaptation called kernel eigenvoice (KEV) speaker adaptation. In KEV adaptation, eigenvoices are computed using kernel PCA, and a new speaker’s adapted model is implicitly computed in the kernel-induced feature space. Due to many online kernel evaluations, both adaptation and subsequent recognition of KEV adaptation are slower than EV adaptation. In this paper, we eliminate all online kernel computations by finding an approximate pre-image of the implicit adapted model found by KEV adaptation. Furthermore, the two steps of finding the implicit adapted model and its approximate pre-image are integrated by embedding the kernel PCA procedure in our new embedded kernel eigenvoice (eKEV) speaker adaptation method. When tested in an TIDIGITS task with less than 10s of adaptation speech, eKEV adaptation obtained a speedup of 6–14 times in adaptation and 136 times in recognition over KEV adaptation with 12–13 % relative improvement in recognition accuracy. 1. Brian Kan-Wing Mak, Simon Ka-Lung Ho, James T. Kwok |
INTERSPEECH | 1 |
| 2004 | Discriminative auditory-based features for robust speech recognitionabstractRecently, a new auditory-based feature extraction algorithm for robust speech recognition in noisy environments was proposed. The new features are derived by mimicking closely the human peripheral auditory process and the filters in the outer ear, middle ear, and inner ear are obtained from psychoacoustics literature with some manual adjustments. In this paper, we extend the auditory-based feature extraction algorithm and propose to further train the auditory-based filters through discriminative training. Using the data-driven approach, we optimize the filters by minimizing the subsequent recognition errors on a task. One significant contribution over similar efforts in the past (generally under the name of "discriminative feature extraction") is that we make no assumption on the parametric form of the auditory-based filters. Instead, we only require the filters to be triangular-like: the filter weights have a maximum value in the middle and then monotonically decrease to both ends. Discriminative training of these constrained auditory-based filters leads to improved performance. Furthermore, we study the combined discriminative training procedure for both feature and acoustic model parameters. Our experiments show that the best performance can be obtained in a sequential procedure under the unified framework of MCE/GPD. Brian Kan-Wing Mak, Yik-Cheung Tam, Peter Qi Li |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Discriminative training of auditory filters of different shapes for robust speech recognitionabstractThe bank-of-filters spectrum analysis model is commonly used in the extraction of acoustic features for automatic speech recognition. The most critical component in the analysis model is a bank of bandpass filters. We studied a data-driven approach to designing a bank of "optimal" filters of various shapes discriminatively so that the recognition error of a task is minimized. Three different shapes of varying degree of constraints were investigated: (1) parametric Gaussian filters; (2) non-parametric but constrained triangular-like filters; and (3) non-parametric and unconstrained free-formed filters. Filters were trained to derive the new robust auditory features proposed by the Bell Labs. In addition, both the filters (and thus the ensuing acoustic features) and the acoustic model parameters were discriminatively trained. The major result is that our proposed triangular-like filters perform at least as well as the free-formed filters and perform better than the Gaussian filters. Brian Kan-Wing Mak, Yik-Cheung Tam, Roger Hsiao |
ICASSP (2) | 1 |
| 2003 | Joint estimation of thresholds in a bi-threshold verification problem
Simon Ka-Lung Ho, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2003 | Pruning transitions in a hidden Markov model with optimal brain surgeonabstractThis paper concerns about reducing the topology of a hidden Markov model (HMM) for a given task. The purpose is two-fold: (1) to select a good model topology with improved generalization capability; and/or (2) to reduce the model complexity so as to save memory and computation costs. The rst goal falls into the active research area of model selection. From the model-theoretic research community, various measures such as Bayesian information criterion, minimum description length, minimum message length have been proposed and used with some success. In this paper, we are considering another approach in which a well-performed HMM, though perhaps oversized, is optimally pruned so that the loss in the model training cost function is minimal. The method is known as Optimal Brain Surgeon (OBS) that has been used in the neural network (NN) community. The application of OBS to NN is a constrained optimization problem; its application to HMM is more involved and it becomes a quadratic programming problem with both equality and inequality constraints. The detailed formulation is presented, and the algorithm is shown effective by an example in which HMM state transitions are pruned. The reduced model also results in better generalization performance on unseen test data. Brian Kan-Wing Mak, Kin-Wah Chan |
INTERSPEECH | 1 |
| 2003 | Eigenvoice Speaker Adaptation via Composite Kernel PCA
James T. Kwok, Brian Kan-Wing Mak, Simon Ka-Lung Ho |
NIPS | 2 |
| 2002 | Discriminative auditory features for robust speech recognitionabstractRecently, Li et al. proposed a new auditory feature for robust speech recognition in noise environments. The new feature was derived by mimicking closely the function of human auditory process. Several filters were used to model the outer ear, middle ear, and cochlea, and the initial filter parameters and shapes were obtained from crude psychoacoustics results, experience, or experiments. Although one may adjust the feature parameters by hand to get better performance, the resulting feature parameters still may not be optimal in the sense of minimal recognition errors, especially for different tasks. To further improve the auditory feature, in this paper we apply discriminative training to optimize the auditory feature parameters with some guidance from psychoacoustic evidence but otherwise in a data-driven approach so as to minimize the recognition errors. One significant contribution over similar efforts in the past, such as discriminative feature extraction, is that we make no assumption on the parametric form of the auditory filters. Instead, we only require the filters to be smooth and triangular-like as suggested by psychoacoustics research. Our approach is evaluated on the Aurora database and achieves a word error reduction of 19.2%. Brian Kan-Wing Mak, Yik-Cheung Tam, Peter Qi Li |
ICASSP | 1 |
| 2002 | An alternative approach of finding competing hypotheses for better minimum classification error trainingabstractDuring minimum-classification-error (MCE) training, competing hypotheses against the correct one are commonly derived by the N-best algorithm. One problem with the N-best algorithm is that, in practice, some misclassified data can have very large misclassification distances from the N-best competitors and fall out of the steep/trainable region of the sigmoid function, and thus cannot be utilized effectively. Although one may alleviate the problem by adjusting the shape of the sigmoid and then using an appropriate learning rate, it requires careful tuning of these training parameters. In this paper, we propose using the nearest competing hypothesis instead of the traditional N-best hypotheses for MCE training. The aim is to keep the training data as close to the trainable region as possible. Consequently, the amount of “effective” training data is increased. Furthermore, by progressively beating the nearest competitors, the training seems to be more stable. We also design an approximation algorithm based on beam search to locate the nearest competing hypothesis efficiently. We compare the performance of MCE training using 1-nearest or 1-best competing hypotheses on the Aurora database and find that the new approach (using 1-nearest hypotheses) reduces the word error rates by 5.1% and 17.8% over the latter (of using the 1-best competing hypotheses) and the official Aurora baseline respectively. Yik-Cheung Tam, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2002 | Performance of discriminatively trained auditory features on Aurora2 and Aurora3abstractThe design of acoustic models involves two main tasks: feature extraction and data modeling; and hidden Markov modeling (HMM) is commonly used in contemporary automatic speech recognition. In the past, discriminative training has been applied successfully to rene HMM parameters that are initially trained by EM algorithm. Recently, we applied discriminative training in the feature extraction process. We proposed a novel Discriminative Auditory Feature extraction method (DAF) in which lters are discriminatively trained from data. In DAF, we do not make any assumptions on the functional form of the auditory lters except that they have to be smooth and triangular-like. On the method of discriminative training, we also proposed an alternative approach to nding the competing hypotheses which we call N-nearest hypotheses (as opposed to the traditional N-best hypotheses). By applying the two new ideas and the new robust auditory features proposed by Li et al. of Bell Labs, we reduce the overall word error rate (WER) by 30.27% over ICSLP2002 Aurora2 baseline on multi-condition training. Similarly, we obtain a relative WER reduction of 38.42% over ICSLP2002 Aurora3 baseline. Brian Kan-Wing Mak, Yik-Cheung Tam |
INTERSPEECH | 1 |
| 2002 | A mathematical relationship between full-band and multiband mel-frequency cepstral coefficientsabstractRecently, it has been shown that robustness of automatic speech recognition (ASR) against band-limited additive noises may be improved by multiband ASR (MBASR) approaches. In an M-subband MBASR system, the channels in the full-band filterbank are divided into M subbands, usually of equal partitions, and subband mel-frequency cepstral coefficients (MFCCs) are computed from each filterbank partition using the discrete cosine transform. However, there is not as yet any analysis on the relationship between full-band and multiband MFCCs. In this letter, we show that the (Mj)th full-band MFCC is the sum of or difference between the Mjth multiband MFCCs multiplied by /spl radic/M. Brian Kan-Wing Mak |
IEEE Signal Process. Lett. | 1 |
| 2001 | Development of an asynchronous multi-band system for continuous speech recognitionabstractRecently, multi-band automatic speech recognition (MBASR) is proposed to combat environmental noises.In this paper, we describe the two major efforts in the development of our asynchronous MBASR system for continuous speech recognition.Firstly, we successfully introduce asynchrony among sub-bands under the HMM composition framework.An asynchrony limit of one state is found adequate -relaxing the limit further does not improve performance.Secondly, the sub-band log likelihoods are combined linearly at the frame level with weightings estimated by minimizing the string classification error (MCE) among the N-best hypotheses using simulated noisy speech.When our asynchronous MBASR system is evaluated on connected TI digits with 0db additive low-pass white noise, compared with a full-band system, (1) our synchronous sub-band system reduces the absolute string error rate (SER) and word error rate (WER) by 19.8% and 14.1% respectively; (2) the introduction of asynchrony further reduces the absolute SER (WER) by 5.2%(2.5%);(3) an estimation of sub-band weightings using N-best string MCE training gives an additional reduction of absolute SER (WER) by 19.7% (5.1%).Thus, in that test, our asynchronous MBASR system has outperformed a full-band system with a 44.7% (21.7%) reduction in absolute SER (WER).In summary, N-best MCE training can effectively emphasize the more reliable sub-band, and asynchronous recombination of sub-bands is preferred. Yik-Cheung Tam, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2001 | Rapid speaker adaptation using MLLR and subspace regression classesabstractIn recent years, various adaptation techniques for hidden Markov modeling with mixture Gaussians have been proposed, most notably MAP estimation and MLLR transformation. When the amount of adaptation data is limited, adaptation can be done by grouping similar Gaussians together to form regression classes and then transforming the Gaussians in groups. The grouping of Gaussians is often determined at the full-space level. In this paper, we propose to group the Gaussians at a finer acoustic subspace level. The motivation is that clustering at subspaces of lower dimensions results in lower distortion. Besides, as the dimension of subspace Gaussians reduces, there are fewer parameters to estimate for the subsequent MLLR transformation matrix. This is particular attractive in fast adaptation. Speaker adaptation experiments on the Resource Management task with few seconds of speech show that the use of subspace regression classes is more effective than traditional full-space regression classes. Kwok-Man Wong, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 2001 | Subspace distribution clustering hidden Markov modelabstractMost contemporary laboratory recognizers require too much memory to run, and are too slow for mass applications. One major cause of the problem is the large parameter space of their acoustic models. In this paper, we propose a new acoustic modeling methodology which we call subspace distribution clustering hidden Markov modeling (SDCHMM) with the aim of achieving much more compact acoustic models. The theory of SDCHMM is based on tying the parameters of a new unit, namely the subspace distribution, of continuous density hidden Markov models (CDHMMs). SDCHMMs can be converted from CDHMMs by projecting the distributions of the CDHMMs onto orthogonal subspaces, and then tying similar subspace distributions over all states and all acoustic models in each subspace, by exploiting the combinatorial effect of subspace distribution encoding, all original full-space distributions can be represented by combinations of a small number of subspace distribution prototypes. Consequently, there is a great reduction in the number of model parameters, and thus substantial savings in memory and computation. This renders SDCHMM very attractive in the practical implementation of acoustic models. Evaluation on the Airline Travel Information System (ATIS) task shows that in comparison to its parent CDHMM system, a converted SDCHMM system achieves seven- to 18-fold reduction in memory requirement for acoustic models, and runs 30%-60% faster without any loss of recognition accuracy. Enrico Bocchieri, Brian Kan-Wing Mak |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Direct training of subspace distribution clustering hidden Markov modelabstractIt generally takes a long time and requires a large amount of speech data to train hidden Markov models for a speech recognition task of a reasonably large vocabulary. Previously, we proposed a compact acoustic model called "subspace distribution clustering hidden Markov model" (SDCHMM) with an aim to save some of the training effort. SDCHMMs are derived from tying continuous density hidden Markov models (CDHMMs) at a finer subphonetic level, namely the subspace distributions. Experiments on the Airline Travel Information System (ATIS) task show that SDCHMMs with significantly fewer model parameters-by one to two orders of magnitude-can be converted from CDHMMs with no loss in word accuracy. With such compact acoustic models, one should be able to train SDCHMM directly from significantly less speech data (without intermediate CDHMMs). We devise a direct SDCHMM training algorithm, assuming an a priori knowledge of the subspace distribution tying structure. On the ATIS task, it is found that both a context-independent and a context-dependent speaker-independent 20-stream SDCHMM system trained with 8 min of speech perform as well as their corresponding CDHMM system trained with 105 min and 36 h of speech, respectively. Brian Kan-Wing Mak, Enrico Bocchieri |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | MAP adaptation with subspace regression classes and tyingabstractIn the hidden Markov modeling framework with mixture Gaussians, adaptation is often done by modifying the Gaussian mean vectors using MAP estimation or MLLR transformation. When the amount of adaptation data is scarce or when some speech units are unseen in the data, it is necessary to do adaptation in groups-either with regression classes of Gaussians or via vector field smoothing. In this paper, we propose to derive regression classes of subspace Gaussians for MAP adaptation. The motivation is that clustering at the finer acoustic level of subspace Gaussians of lower dimension is more effective, resulting in lower distortions and relatively fewer regression classes. Experiments in which context-dependent TIMIT HMMs are adapted to the resource management task with few minutes of speech show improvement of our subspace regression classes over traditional full-space regression classes. Kwok-Man Wong, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2000 | Pruning of state-tying tree using bayesian information criterion with multiple mixturesabstractThe use of context-dependent phonetic units together with Gaussian mixture models allows modern-day speech recognizer to build very complex and accurate acoustic models. However, because of data sparseness issue, some sharing of data across dierent triphone states is necessary. The acoustic model design is typically done in two stages, namely, designing the state-tying map and growing the number of mixtures in each tied-state. In the design of the state-tying map, single Gaussians are used to represent the data, ignoring the fact that a single Gaussian is an insucient model. In this paper, we propose a simple modication to the two-stage process by adding a third stage. In this added stage, the state-tying tree is pruned and the pruning is based on the mixture representation of the tied-states. We propose using Bayesian Information Criterion(BIC) as the criterion for this pruning and show that by adding this step, the resulting model is more compact and gives better recognition accura... Yu-Chung Chan, Man-Hung Siu, Brian Kan-Wing Mak |
INTERSPEECH | 3 |
| 2000 | Asynchrony with trained transition probabilities improves performance in multi-band speech recognitionabstractOne of the central themes in multi-band automatic speech recognition (ASR) is to devise a strategy for recombining sub-band information. This in turn raises two questions: (1) at what phonetic unit should the recombination take place? (2) How asynchronously should the sub-bands be run? Theoretically asynchronous multi-band ASR should perform at least as well as synchronous multi-band ASR. However, in the past few years, there are conicting results on the issue. In this paper, we study the asynchrony issue under the framework of HMM composition in which a model-based recombination strategy is used to recombine sub-band HMMs at the state level. We hypothesize that re-estimation of the transition probabilities is crucial for multi-band ASR (using HMM composition). Experiments on connected TI digits show that for both clean speech and noisy speech (with additive white noise of 10db), HMMs composed from sub-band HMMs in which transition probabilities are trained with Baum-Welch algorithm o... Brian Kan-Wing Mak, Yik-Cheung Tam |
INTERSPEECH | 1 |
| 2000 | Optimization of sub-band weights using simulated noisy speech in multi-band speech recognitionabstractRecently multi-band speech recognition has been proposed to improve robustness under environmental noises. One important issue is how to combine decisions from individual sub-band recognizers to arrive at a nal decision. Under the hidden Markov modeling (HMM) framework, one common approach is combining sub-band likelihoods linearly in an optimal manner so that the more reliable sub-bands are emphasized and the corrupted sub-bands are de-emphasized. In our experience, estimating the weights from clean speech is not eective as the weights are not optimal under noisy environments. In this paper, we derive the optimal weights from simulated noisy speech using discriminative training method with minimum classi cation errors (MCE) or maximum mutual information (MMI) as the cost function. The methods are evaluated on recognition of isolated TI digits. Compared with full-band recognition with noises at an SNR of 0dB, multiband recognition with MCE-derived weights reduces word errors by 45.9%... Yik-Cheung Tam, Brian Kan-Wing Mak |
INTERSPEECH | 2 |
| 1998 | Training of subspace distribution clustering hidden Markov modelabstractLevinson, Juang and Sondhi (1986), and Mak, Bocchieri, and E. Barnard (see Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop, 1997) presented novel subspace distribution clustering hidden Markov models (SDCHMMs) which can be converted from continuous density hidden Markov models (CDHMMs) by clustering subspace Gaussians in each stream over all models. Though such model conversion is simple and runs fast, it has two drawbacks: (1) it does not take advantage of the fewer model parameters in SDCHMMs-theoretically SDCHMMs may be trained with smaller amount of data; and, (2) it involves two separate optimization steps (first training CDHMMs, then clustering subspace Gaussians) and the resulting SDCHMMs are not guaranteed to be optimal. We show how SDCHMMs may be trained directly from less speech data if we have a priori knowledge of their architecture. On the ATIS task, a speaker-independent, context-independent (CI) 20-stream SDCHMM system trained using our novel SDCHMM reestimation algorithm with only 8 minutes of speech performs as well as a CDHMM system trained using conventional CDHMM reestimation algorithm with 105 minutes of speech. Brian Kan-Wing Mak, Enrico Bocchieri |
ICASSP | 1 |
| 1998 | Training of context-dependent subspace distribution clustering hidden Markov modelabstractTraining of continuous density hidden Markov models (CDHMMs) is usually time-consuming and tedious due to the large number of model parameters involved. Recently we proposed a new derivative of CDHMM, the subspace distribution clustering hidden Markov model (SDCHMM) which tie CDHMMs at the finer level of subspace distributions, resulting in many fewer model parameters. An SDCHMM training algorithm is also devised to train SDCHMMs directly from speech data without intermediate CDHMMs. On the ATIS task, speaker-independent context-independent (CI) SDCHMMs can be trained with as little as 8 minutes of speech with no loss in recognition accuracy --- a 25-fold reduction when compared with their CDHMM counterparts [1]. In this paper, we extend our novel SDCHMM training to context-dependent (CD) modeling with the assumption of various prior knowledge. Despite the 30-fold increase of model parameters in the CD ATIS CDHMMs, their equivalent CD SDCHMMs can still be estimated with a few minutes o... Brian Kan-Wing Mak, Enrico Bocchieri |
ICSLP | 1 |
| 1997 | Combining ANNs to improve phone recognitionabstractIn applying neural networks to speech recognition, one often finds that slightly different training configurations lead to significantly different networks. Thus different training sessions using different setups will likely end up in "mixed" network configurations representing different solutions in different regions of the data space. This sensitivity to the initial weights assigned, the training parameters and the training data can be used to enhance performance, using a committee of neural networks. We study various ways to combine context-dependent (CD) and context-independent (CD) neural network phone estimators to improve phone recognition. As a result, we obtain 6.3% and 2.2% increase in accuracy in phone recognition using monophones and biphones respectively. Brian Kan-Wing Mak |
ICASSP | 1 |
| 1997 | Subspace distribution clustering for continuous observation density hidden Markov modelsabstractThis paper presents an efficient approximation of the Gaussian mixture state probability density functions of continuous observation density hidden Markov models (CHMM 's). In CHMM 's, the Gaussian mixtures carry a high computational cost, which amounts to a significant fraction (e.g. 30% to 70%) of the total computation. To achieve higher computation and memory efficiency, we approximate the Gaussian mixtures by (a) decomposition into functions defined on subspaces of the feature space, and (b) clustering the resulting subspace pdf's. Intuitively, when clustering in a subspace of few dimensions, even few function codewords can provide a small distortion. Therefore, we obtain significant reduction of the total computation (up to a factor of two), and memory savings (up to a factor of twelve), without significant changes of the CHMMM 's accuracy. 1. INTRODUCTION Most of state-of-the-art speech recognition systems are based on hidden Markov models (HMM) technology. In particular, conti... Enrico Bocchieri, Brian Kan-Wing Mak |
EUROSPEECH | 2 |
| 1996 | The contribution of consonants versus vowels to word recognition in fluent speechabstractThree perceptual experiments were conducted to test the relative importance of vowels vs. consonants to recognition of fluent speech. Sentences were selected from the TIMIT corpus to obtain approximately equal numbers of vowels and consonants within each sentence and equal durations across the set of sentences. In experiments 1 and 2, subjects listened to (a) unaltered TIMIT sentences; (b) sentences in which all of the vowels were replaced by noise; or (c) sentences in which all of the consonants were replaced by noise. The subjects listened to each sentence five times, and attempted to transcribe what they heard. The results of these experiments show that recognition of words depends more upon vowels than consonants-about twice as many words are recognized when vowels are retained in the speech. The effect was observed when occurrences of [1], [r], [w], [y] [m], [n], were included in the sentences (experiment 1) or replaced by noise (experiment 2). Experiment 3 tested the hypothesis that vowel boundaries contain more information about the neighboring consonants than vice versa. Ronald A. Cole, Yonghong Yan 0002, Brian Kan-Wing Mak, Mark A. Fanty, Troy Bailey |
ICASSP | 3 |
| 1996 | Phone clustering using the bhattacharyya distance
Brian Kan-Wing Mak, Etienne Barnard |
ICSLP | 1 |
| 1995 | Tone recognition of isolated Cantonese syllablesabstractTone identification is essential for the recognition of the Chinese language, specifically far Cantonese which is well known for being very rich in tones. The paper presents an efficient method for tone recognition of isolated Cantonese syllables. Suprasegmental feature parameters are extracted from the voiced portion of a monosyllabic utterance and a three-layer feedforward neural network is used to classify these feature vectors. Using a phonologically complete vocabulary of 234 distinct syllables, the recognition accuracy for single-speaker and multispeaker is given by 89.0% and 87.6% respectively.> Tan Lee, Pak-Chung Ching, Lai-Wan Chan, Y. H. Cheng, Brian Kan-Wing Mak |
IEEE Trans. Speech Audio Process. | 5 |
| 1994 | A robust algorithm for word boundary detection in the presence of noiseabstractThe authors address the problem of automatic word boundary detection in quiet and in the presence of noise. Attention has been given to automatic word boundary detection for both additive noise and noise-induced changes in the talker's speech production (Lombard reflex). After a comparison of several automatic word boundary detection algorithms in different noisy-Lombard conditions, they propose a new algorithm that is robust in the presence of noise. This new algorithm identifies islands of reliability (essentially the portion of speech contained between the first and the last vowel) using time and frequency-based features and then, after a noise classification, applies a noise adaptive procedure to refine the boundaries. It is shown that this new algorithm outperforms the commonly used algorithm developed by Lamel (1981) et al. and several other recently developed methods. They evaluated the average recognition error rate due to word boundary detection in an HMM-based recognition system across several signal-to-noise ratios and noise conditions. The recognition error rate decreased to about 20% compared to an average of approximately 50% obtained with a modified version of the Lamel et al. algorithm.> Jean-Claude Junqua, Brian Kan-Wing Mak, Ben Reaves |
IEEE Trans. Speech Audio Process. | 2 |
| 1992 | A robust speech/non-speech detection algorithm using time and frequency-based featuresabstractThe authors address the problem of automatic endpoint detection in normal and adverse conditions. Attention has been given to automatic endpoint detection for both additive noise and noise-induced changes in the talker's speech production (Lombard reflex). After a comparison of several automatic endpoint detection algorithms in different noisy-Lombard conditions, the authors propose a new algorithm. This algorithm identifies islands of reliability (essentially the portion of speech contained between the first and last vowel) using time- and frequency-based features and then applies a noise adaptive procedure to refine the endpoints. It is shown that this algorithm outperforms the commonly used algorithm developed by Lamel et al. (1981), and several other recently developed methods.> Brian Kan-Wing Mak, Jean-Claude Junqua, Ben Reaves |
ICASSP | 1 |
| 1991 | A study of endpoint detection algorithms in adverse conditions: incidence on a DTW and HMM recognizerabstractIn this paper the performances of three recently developed end-point algorithms are evaluated and compared to the Lamel and Rosenberg's algorithm [1] based on energy levels and timing, which is enhanced by automatic threshold setting. Their performances are reported when integrated with two commonly used speech recognizers (discrete density vector quantization-based hidden Markov model (VQ-based HMM) and dynamic time warping (DTW)) in various types of noisy conditions. Accuracy was judged by agreement with hand-labeled endpoints, and by recognition rates. Results show that 1) a new noise adaptive algorithm using rms energy, zero-crossing rate, and a set of heuristics gives generally the best results at high or medium signal-to-noise ratio (>15 dB). 2) The HMM recognizer when used with this algorithm performs as well as if the endpoints were hand-labeled for clean Lombard speech; for noisy Lombard speech, depending on the type of noise used and the SNR, there is a degradation from 1% to 43% in recognition accuracy compared to manual labeling. 3) At low SNR, the algorithm based on Lamel and Rosenberg's method [1] and enhanced by automatic threshold setting gives generally better performance than the other algorithms. Jean-Claude Junqua, Ben Reaves, Brian Kan-Wing Mak |
EUROSPEECH | 3 |