Yik-Cheung Tam

dblp:72/158 · DBLP profile ↗
← Back
35ranked-venue papers
19as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 15 first-author · 2 since 2021Artificial intelligence and machine learning · 22 · 13 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 77% Reinforcement learning · 12% Vision and language · 6%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
Predicate-Guided Generation for Mathematical Reasoning · EMNLP 2025
Natural language and speech › Language models and text generation
watermarking
0.912025
VLA-Mark: A cross modal watermark for large vision-language alignment models · EMNLP 2025
Program synthesis and code generation
logic program synthesis
0.912025
Predicate-Guided Generation for Mathematical Reasoning · EMNLP 2025
Natural language and speech › Language models and text generation › text generation › poetry generation
chinese poetry generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Natural language and speech › Language models and text generation
controllable text generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Natural language and speech › Language models and text generation › text generation
poetry generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Natural language and speech › Language models and text generation
text generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.312025
Predicate-Guided Generation for Mathematical Reasoning · EMNLP 2025
Machine learning › Reinforcement learning
policy optimization
0.312025
Predicate-Guided Generation for Mathematical Reasoning · EMNLP 2025
Natural language and speech › Language models and text generation › large language model
large language model adaptation
0.222008
Correlated Bigram LSA for Unsupervised Language Model Adaptation · NIPS 2008
Bilingual-LSA Based LM Adaptation for Spoken Language Translation · ACL 2007
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.112008
Correlated Bigram LSA for Unsupervised Language Model Adaptation · NIPS 2008
Natural language and speech › Machine translation
speech translation
0.112007
Bilingual-LSA Based LM Adaptation for Spoken Language Translation · ACL 2007
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition
0.012004
Discriminative auditory-based features for robust speech recognition · IEEE Trans. Speech Audio Process. 2004
Natural language and speech › Information extraction and text analysis › topic model
latent semantic analysis
0.012008
Correlated Bigram LSA for Unsupervised Language Model Adaptation · NIPS 2008

Methods — techniques the papers use, named apart from their topics

supervised fine-tuning · 1.7predicate-aware reward · 1.7GRPO · 1.7latent variable model · 0.4encoder-decoder · 0.4variational EM · 0.1kneser-ney smoothing · 0.1bigram LSA · 0.1bilingual latent semantic analysis · 0.1discriminative training · 0.0
YearPublicationVenuePosition
2025 Predicate-Guided Generation for Mathematical Reasoning
abstract
We present Prolog-MATH, a curated corpus designed to support mathematical reasoning in large language models (LLMs) through logic programming.Each verbal math problem in the dataset is paired with a chain-of-thought explanation to generate Prolog program via a two-stage automated pipeline.In the first stage, an LLM (e.g., Deepseek-V3) predicts a set of relevant mathematical predicates that could be useful in solving the problem.In the second stage, the LLM uses these suggested predicates along with the expected answer type to generate a complete Prolog program.To improve coverage, we fine-tune an open-source LLM using supervised fine-tuning, followed by GRPO (Group Relative Policy Optimization) training to address problems that Deepseek-V3 fails to solve.To support this training, we propose a predicate-aware reward function that evaluates how well the generated solution incorporates the suggested predicates, complementing the standard binary reward.Experimental results show that: 1) Our two-stage pipeline achieves 81.3% solution coverage on the MATH training set; 2) GRPO training with the predicate-aware reward function enables a series of base models to correctly solve additional problems missed by Deepseek-V3, further increasing solution coverage to 97.4%.Data and source code can be obtained at the Github repository 1 .
Yik-Cheung Tam
EMNLP2
2025 VLA-Mark: A cross modal watermark for large vision-language alignment models
abstract
Shuliang Liu, Zheng Qi, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zheng Qi, Jesse Jiaxi Xu, He Geng, Aiwei Liu, Peijie Jiang, Yik-Cheung Tam, Xuming Hu
EMNLP10
2023 Suffix Retrieval-Augmented Language Modeling
abstract
Causal language modeling (LM) uses word history to predict the next word. BERT, on the other hand, makes use of bi-directional word information in a sentence to predict words at masked positions. While BERT is effective in sequence encoding, it is non-causal by nature and is not designed for sequence generation. In this paper, we propose a novel language model, SUffix REtrieval-Augmented LM (SUREALM), that simulates a bi-directional contextual effect in an autoregressive manner. SUREALM employs an embedding retriever to search for training sentences in a data store that share similar word history during sequence generation. In particular, the suffix portions of the retrieved sentences mimick the "future" context. We evalu-ated our proposed model on the DSTC9 spoken dialogue corpus and showed promising word perplexity reduction on the validation and test set compared to competitive baselines. Our source code is re-leased on GitHub1.
Zecheng Wang, Yik-Cheung Tam
ICASSP2
2022 Robust Unstructured Knowledge Access in Conversational Dialogue with ASR Errors
abstract
Performance of spoken language understanding (SLU) can be degraded with automatic speech recognition (ASR) errors. We propose a novel approach to improve SLU robustness by randomly corrupting clean training text with an ASR error simulator, followed by self-correcting the errors and minimizing the target classification loss in a joint manner. In the proposed error simulator, we leverage confusion networks generated from an ASR decoder without human transcriptions to generate variety of error patterns for model training. We evaluate our approach on DSTC10 challenge targeted for knowledge-grounded task-oriented conversational dialogues with ASR errors. Experimental results show effectiveness of our proposed approach, boosting the knowledge-seeking turn detection (KTD) F1 significantly from 0.9433 to 0.9904. Knowledge cluster classification is boosted from 0.7924 to 0.9333 in Recall@1. After knowledge document re-ranking, our approach shows significant improvement in all knowledge selection metrics, from 0.7358 to 0.7806 in Recall@1, from 0.8301 to 0.9333 in Recall@5, and from 0.7798 to 0.8460 in MRR@5 (Mean Reciprocal Rank) on the test set. On the recent DSTC10 evaluation, our approach demonstrates significant improvement in knowledge selection, boosting Recall@1 from 0.495 to 0.7105 compared to the official baseline. Our source code is released in GitHub1.
Yik-Cheung Tam, Jiakai Zou, Zecheng Wang, Tinglong Liao, Shuhan Yuan
ICASSP1
2020 Cluster-based beam search for pointer-generator chatbot grounded by knowledge
Yik-Cheung Tam
Comput. Speech Lang.1
2019 Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation
abstract
Rhetoric is a vital element in modern poetry, and plays an essential role in improving its aesthetics.However, to date, it has not been considered in research on automatic poetry generation.In this paper, we propose a rhetorically controlled encoder-decoder for modern Chinese poetry generation.Our model relies on a continuous latent variable as a rhetoric controller to capture various rhetorical patterns in an encoder, and then incorporates rhetoricbased mixtures while generating modern Chinese poetry.For metaphor and personification, an automated evaluation shows that our model outperforms state-of-the-art baselines by a substantial margin, while a human evaluation shows that our model generates better poems than baseline methods in terms of fluency, coherence, meaningfulness, and rhetorical aesthetics.
Zuohui Fu, Jie Cao 0010, Gerard de Melo, Yik-Cheung Tam, Cheng Niu, Jie Zhou 0016
ACL (1)5
2015 Neural network joint modeling via context-dependent projection
abstract
Neural network joint modeling (NNJM) has produced huge improvement in machine translation performance. As in standard neural network language modeling, a context-independent linear projection is applied to project a sparse input vector into a continuous representation at each word position. Because neighboring words are dependent on each other, context-independent projection may not be optimal. We propose a context-dependent linear projection approach which considers neighboring words. Experimental results showed that the proposed approach further improves NNJM by 0.5 BLEU for English-Iraqi Arabic translation in N-best rescoring. Compared to a baseline using hierarchical phrases and sparse features, NNJM with our proposed approach has achieved a 2 BLEU improvement.
Yik-Cheung Tam, Yun Lei
ICASSP1
2015 RNN-based labeled data generation for spoken language understanding
Yik-Cheung Tam, Yangyang Shi, Hunk Chen, Mei-Yuh Hwang
INTERSPEECH1
2015 Morphological Modeling for Machine Translation of English-Iraqi Arabic Spoken Dialogs
abstract
Katrin Kirchhoff, Yik-Cheung Tam, Colleen Richey, Wen Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Katrin Kirchhoff, Yik-Cheung Tam, Colleen Richey, Wen Wang 0001
HLT-NAACL2
2014 Feature fusion for high-accuracy keyword spotting
abstract
This paper assesses the role of robust acoustic features in spoken term detection (a.k.a keyword spotting - KWS) under heavily degraded channel and noise corrupted conditions. A number of noise-robust acoustic features were used, both in isolation and in combination, to train large vocabulary continuous speech recognition (LVCSR) systems, with the resulting word lattices used for spoken term detection. Results indicate that the use of robust acoustic features improved KWS performance with respect to a highly optimized state-of-the art baseline system. It has been shown that fusion of multiple systems improve KWS performance, however the number of systems that can be trained is constrained by the number of frontend features. This work shows that given a number of frontend features it is possible to train several systems by using the frontend features by themselves along with different feature fusion techniques, which provides a richer set of individual systems. Results from this work show that KWS performance can be improved compared to individual feature based systems when multiple features are fused with one another and even further when multiple such systems are combined. Finally this work shows that fusion of fused and single feature bases systems provide significant improvement in KWS performance compared to fusion of singlefeature based systems.
Vikramjit Mitra, Julien van Hout, Horacio Franco, Dimitra Vergyri, Yun Lei, Martin Graciarena, Yik-Cheung Tam, Jing Zheng 0001
ICASSP7
2014 ASR error detection using recurrent neural network language model and complementary ASR
abstract
Detecting automatic speech recognition (ASR) errors can play an important role for effective human-computer spoken dialogue system, as recognition errors can hinder accurate system understanding of user intents. Our goal is to locate errors in an utterance so that the dialogue manager can pose appropriate clarification questions to the users. We propose two approaches to improve ASR error detection: (1) using recurrent neural network language models to capture long-distance word context within and across previous utterances; (2) using a complementary ASR system. The intuition is that when two complementary ASR systems disagree on a region in an utterance, this region is most likely an error. We train a neural network predictor of errors using a variety of features. We performed experiments on both English and Iraqi Arabic ASR and observed significant improvement in error detection using the proposed methods.
Yik-Cheung Tam, Yun Lei, Jing Zheng 0001, Wen Wang 0001
ICASSP1
2014 An autoencoder with bilingual sparse features for improved statistical machine translation
abstract
Though sparse features have produced significant gains over traditional dense features in statistical machine translation, careful feature selection and feature engineering are necessary to avoid over-fitting in optimizations. However, many sparse features are highly overlapping with each other; that is, they cover the same or similar information of translational equivalence from slightly different points of view, and eventually overfit easily with only very feature training samples in given bilingual stochastic context-free grammar (SCFG) rules. We propose a natural autoencoder that maps all the discrete and overlapping sparse features for each SCFG rule into a continuous vector, so that the information encoded in sparse feature vectors becomes a dense vector that may enjoy more samples during training and avoid overfitting. Our experiments showed that for a 33-million bilingual SCFG rules statistical machine translation system, the autoencoder generalizes much better than sparse features alone using the same optimization framework.
Yik-Cheung Tam, Jing Zheng 0001
ICASSP2
2014 Bilingual Recurrent Neural Networks for improved statistical machine translation
abstract
Recurrent Neural Networks (RNN) have been successfully applied for improved speech recognition and statistical machine translation (SMT) for N-best list re-ranking. In SMT, we investigate using bilingual word-aligned sentences to train a bilingual recurrent neural network model. We employ a bag-of-word representation of a source sentence as additional input features in model training. Experimental results show that our proposed approach performs consistently better than recurrent neural network language model trained only on target-side text in terms of machine translation performance. We also investigate other input representation of a source sentence based on latent semantic analysis.
Yik-Cheung Tam
SLT2
2013 Strategies for high accuracy keyword detection in noisy channels
abstract
We present design strategies for a keyword spotting (KWS) system that operates in highly degraded channel conditions with very low signal-to-noise ratio levels. We employ a system combination approach by combining the outputs of multiple large vocabulary automatic speech recognition (LVCSR) systems, each of which employs a different system design approach targeting three different levels of information: front-end signal processing features (standard cepstra-based, noise-robust modulation and multi layer perceptron features), statistical acoustic models (gaussian mixtures models (GMM) and subspace GMMs) and keyword search strategies (word-based and phonebased). We also use keyword-aware capabilities in the system at two levels: in the LVCSR language models by assigning higher weights to n-grams with keywords in them and in LVCSR search by using a relaxed pruning threshold for keywords. The LVCSR system outputs are represented as latticebased unigram indices whose scores are fused by a logisticregression based classifier to produce the final system combination output. We present the performance of our system in the phase II evaluations of DARPA’s Robust Automatic Transcription of Speech (RATS) program for both Levantine Arabic and Farsi conversational speech corpora. Index Terms: noise-robust keyword detection, automatic speech recognition, system combination, noise robustness
Arindam Mandal, Julien van Hout, Yik-Cheung Tam, Vikramjit Mitra, Yun Lei, Jing Zheng 0001, Dimitra Vergyri, Luciana Ferrer, Martin Graciarena, Andreas Kathol, Horacio Franco
INTERSPEECH3
2012 A Hierarchical Bayesian Approach for Semi-supervised Discriminative Language Modeling
abstract
Discriminative language modeling provides a mechanism for differentiating between competing word hypotheses, which are usually ignored in traditional maximum likelihood estimation of N-gram language models. Discriminative language modeling usually requires manual transcription which can be costly and slow to obtain. On the other hand, there are vast amount of untranscribed speech data on which offline adaptation technique can be applied to generate pseudo-truth transcription as an approximation to manual transcription. Viewing manual and pseudo-truth transcriptions as two domains, we perform domain adaptation on the discriminative language models via hierarchical Bayesian, in which the domainspecific models share a common prior model. Domainspecific and prior models are then estimated jointly using training data. On the N-best list rescoring experiment, hierarchical Bayesian has yielded better recognition performance than the model trained only on manual transcription, and is robust against inferior prior.
Yik-Cheung Tam, Paul Vozila
INTERSPEECH1
2011 Unsupervised Latent Speaker Language Modeling
Yik-Cheung Tam, Paul Vozila
INTERSPEECH1
2009 Generalized Baum-Welch algorithm for discriminative training on large vocabulary continuous speech recognition system
abstract
We propose a new optimization algorithm called Generalized Baum Welch (GBW) algorithm for discriminative training on hidden Markov model (HMM). GBW is based on Lagrange relaxation on a transformed optimization problem. We show that both Baum-Welch (BW) algorithm for ML estimate of HMM parameters, and the popular extended Baum-Welch (EBW) algorithm for discriminative training are special cases of GBW.We compare the performance of GBW and EBW for Farsi large vocabulary continuous speech recognition (LVCSR).
Roger Hsiao, Yik-Cheung Tam, Tanja Schultz
ICASSP2
2009 Incorporating monolingual corpora into bilingual latent semantic analysis for crosslingual LM adaptation
abstract
The major limitation in bilingual latent semantic analysis (bLSA) is the requirement of parallel training corpora. Motivated by semi-supervised learning, we propose a clusterbased bLSA training approach to incorporate monolingual corpora. Treating each parallel document pair as centroids of the parallel document clusters, each monolingual document is associated to the closest centroid according to their topic similarity. The resulting parallel document clusters are used as constraints to enforce a one-to-one topic correspondence in variational EM. Slight performance improvement in crosslingual language model adaptation is observed compared to the baseline without monolingual corpora.
Yik-Cheung Tam, Tanja Schultz
ICASSP1
2008 The CMU-interACT 2008 Mandarin transcription system
abstract
We present our Mandarin BN/BC transcription system recently developed for the GALE07 evaluation. The system employs a 3-pass decoding strategy trained with over 1300 hours of quickly transcribed audio. We successfully apply discriminative training, dynamic unsupervised language model adaptation, and system combination techniques in our system. We furthermore achieve improvements by combining an Initial-Final system with a genre dependent phone system. On the GALE07 phase 2 retest evaluation, our system achieves a character error rate(CER) of 13.3 % on dev07 test set and 13.5 % on eval07 unsequestered test set. Our system also allows combination with other sites and in this paper, we investigate different system combination strategies which significantly improve the final recognition performance. Index Terms: Mandarin transcription system, broadcast news, broadcast conversation, GALE evaluation
Roger Hsiao, Mark C. Fuhs, Yik-Cheung Tam, Qin Jin, Tanja Schultz
INTERSPEECH3
2008 Correlated Bigram LSA for Unsupervised Language Model Adaptation
abstract
We propose using correlated bigram LSA for unsupervised LM adaptation for automatic speech recognition. The model is trained using efficient variational EM and smoothed using the proposed fractional Kneser-Ney smoothing which handles fractional counts. Our approach can be scalable to large training corpora via bootstrapping of bigram LSA from unigram LSA. For LM adaptation, unigram and bigram LSA are integrated into the background N-gram LM via marginal adaptation and linear interpolation respectively. Experimental results show that applying unigram and bigram LSA together yields 6%--8% relative perplexity reduction and 0.6% absolute character error rates (CER) reduction compared to applying only unigram LSA on the Mandarin RT04 test set. Comparing with the unadapted baseline, our approach reduces the absolute CER by 1.2%.
Yik-Cheung Tam, Tanja Schultz
NIPS1
2007 Bilingual-LSA Based LM Adaptation for Spoken Language Translation
Yik-Cheung Tam, Ian Lane, Tanja Schultz
ACL1
2007 Correlated Latent Semantic Model for Unsupervised LM Adaptation
abstract
We propose a latent Dirichlet-tree allocation (LDTA) model - a correlated latent semantic model - for unsupervised language model adaptation. The LDTA model extends the latent Dirichlet allocation (LDA) model by replacing a Dirichlet prior with a Dirichlet-tree prior over the topic proportions. Latent topics under the same subtree are expected to be more correlated than topics under different subtrees. The LDTA model falls back to the LDA model using a depth-one Dirichlet-tree, and the model fits to the variational Bayes inference framework employed in the LDA model. Empirical results show that the LDTA model has a faster training convergence than the LDA model with the same initial flat model. Experimental results show that LDTA-adapted LM performed better than LDA-adapted LM on the Mandarin RT04-eval set when the models were trained using a small text corpus, while both models had the same recognition performance when the models were trained using a big text corpus. We observed 0.4% absolute CER reduction after LM adaptation using LSA marginals.
Yik-Cheung Tam, Tanja Schultz
ICASSP (4)1
2007 Bilingual LSA-based translation lexicon adaptation for spoken language translation
abstract
We present a bilingual LSA (bLSA) framework for translation lexicon adaptation. The idea is to apply marginal adaptation on a translation lexicon so that the lexicon marginals match to in-domain marginals. In the framework of speech translation, the bLSA method transfers topic distributions from the source to the target side, such that the translation lexicon can be adapted before translation based on the source document. We evaluated the proposed approach on our Mandarin RT04 spoken language translation system. Results showed that the conditional likeli-hood on the test sentence pairs is improved significantly using an adapted translation lexicon compared to an unadapted base-line. The proposed approach showed improvement on BLEU-score in SMT. When both the target-side LM and the translation lexicon were adapted and applied simultaneously for SMT de-coding, the gain on BLEU-score was more than additive com-pared to the scenarios when the adapted models were individu-ally applied.
Yik-Cheung Tam, Tanja Schultz
INTERSPEECH1
2007 Bilingual LSA-based adaptation for statistical machine translation
Yik-Cheung Tam, Ian Lane, Tanja Schultz
Mach. Transl.1
2006 Unsupervised language model adaptation using latent semantic marginals
abstract
We integrated the Latent Dirichlet Allocation (LDA) approach, a latent semantic analysis model, into unsupervised language model adaptation framework. We adapted a background language model by minimizing the Kullback-Leibler divergence between the adapted model and the background model subject to a constraint that the marginalized unigram probability distribution of the adapted model is equal to the corresponding distribution estimated by the LDA model – the latent semantic marginals. We evaluated our approach on the RT04 Mandarin Broadcast News test set and experimented with different LM training settings. Results showed that our approach reduces the perplexity and the character error rates using supervised and unsupervised adaptation.
Yik-Cheung Tam, Tanja Schultz
INTERSPEECH1
2005 Dynamic language model adaptation using variational Bayes inference
abstract
We propose an unsupervised dynamic language model (LM) adaptation framework using long-distance latent topic mixtures.The framework employs the Latent Dirichlet Allocation model (LDA) which models the latent topics of a document collection in an unsupervised and Bayesian fashion.In the LDA model, each word is modeled as a mixture of latent topics.Varying topics within a context can be modeled by re-sampling the mixture weights of the latent topics from a prior Dirichlet distribution.The model can be trained using the variational Bayes Expectation Maximization algorithm.During decoding, mixture weights of the latent topics are adapted dynamically using the hypotheses of previously decoded utterances.In our work, the LDA model is combined with the trigram language model using linear interpolation.We evaluated the approach on the CCTV episode of the RT04 Mandarin Broadcast News test set.Results show that the proposed approach reduces the perplexity by up to 15.4% relative and the character error rate by 4.9% relative depending on the size and setup of the training set.
Yik-Cheung Tam, Tanja Schultz
INTERSPEECH1
2004 Discriminative auditory-based features for robust speech recognition
abstract
Recently, a new auditory-based feature extraction algorithm for robust speech recognition in noisy environments was proposed. The new features are derived by mimicking closely the human peripheral auditory process and the filters in the outer ear, middle ear, and inner ear are obtained from psychoacoustics literature with some manual adjustments. In this paper, we extend the auditory-based feature extraction algorithm and propose to further train the auditory-based filters through discriminative training. Using the data-driven approach, we optimize the filters by minimizing the subsequent recognition errors on a task. One significant contribution over similar efforts in the past (generally under the name of "discriminative feature extraction") is that we make no assumption on the parametric form of the auditory-based filters. Instead, we only require the filters to be triangular-like: the filter weights have a maximum value in the middle and then monotonically decrease to both ends. Discriminative training of these constrained auditory-based filters leads to improved performance. Furthermore, we study the combined discriminative training procedure for both feature and acoustic model parameters. Our experiments show that the best performance can be obtained in a sequential procedure under the unified framework of MCE/GPD.
Brian Kan-Wing Mak, Yik-Cheung Tam, Peter Qi Li
IEEE Trans. Speech Audio Process.2
2003 Discriminative training of auditory filters of different shapes for robust speech recognition
abstract
The bank-of-filters spectrum analysis model is commonly used in the extraction of acoustic features for automatic speech recognition. The most critical component in the analysis model is a bank of bandpass filters. We studied a data-driven approach to designing a bank of "optimal" filters of various shapes discriminatively so that the recognition error of a task is minimized. Three different shapes of varying degree of constraints were investigated: (1) parametric Gaussian filters; (2) non-parametric but constrained triangular-like filters; and (3) non-parametric and unconstrained free-formed filters. Filters were trained to derive the new robust auditory features proposed by the Bell Labs. In addition, both the filters (and thus the ensuing acoustic features) and the acoustic model parameters were discriminatively trained. The major result is that our proposed triangular-like filters perform at least as well as the free-formed filters and perform better than the Gaussian filters.
Brian Kan-Wing Mak, Yik-Cheung Tam, Roger Hsiao
ICASSP (2)2
2003 Training a confidence measure for a reading tutor that listens
abstract
One issue in a Reading Tutor that listens is to determine which words the student read correctly. We describe a confidence measure that uses a variety of features to estimate the probability that a word was read correctly. We trained two decision tree classifiers. The first classifier tries to fix insertion and substitution errors made by the speech decoder, while the second classifier tries to fix deletion errors. By applying the two classifiers together, we achieved a relative reduction in false alarm rate by 25.89 % while holding the miscue detection rate constant. 1.
Yik-Cheung Tam, Jack Mostow, Joseph E. Beck, Satanjeev Banerjee
INTERSPEECH1
2002 Discriminative auditory features for robust speech recognition
abstract
Recently, Li et al. proposed a new auditory feature for robust speech recognition in noise environments. The new feature was derived by mimicking closely the function of human auditory process. Several filters were used to model the outer ear, middle ear, and cochlea, and the initial filter parameters and shapes were obtained from crude psychoacoustics results, experience, or experiments. Although one may adjust the feature parameters by hand to get better performance, the resulting feature parameters still may not be optimal in the sense of minimal recognition errors, especially for different tasks. To further improve the auditory feature, in this paper we apply discriminative training to optimize the auditory feature parameters with some guidance from psychoacoustic evidence but otherwise in a data-driven approach so as to minimize the recognition errors. One significant contribution over similar efforts in the past, such as discriminative feature extraction, is that we make no assumption on the parametric form of the auditory filters. Instead, we only require the filters to be smooth and triangular-like as suggested by psychoacoustics research. Our approach is evaluated on the Aurora database and achieves a word error reduction of 19.2%.
Brian Kan-Wing Mak, Yik-Cheung Tam, Peter Qi Li
ICASSP2
2002 An alternative approach of finding competing hypotheses for better minimum classification error training
abstract
During minimum-classification-error (MCE) training, competing hypotheses against the correct one are commonly derived by the N-best algorithm. One problem with the N-best algorithm is that, in practice, some misclassified data can have very large misclassification distances from the N-best competitors and fall out of the steep/trainable region of the sigmoid function, and thus cannot be utilized effectively. Although one may alleviate the problem by adjusting the shape of the sigmoid and then using an appropriate learning rate, it requires careful tuning of these training parameters. In this paper, we propose using the nearest competing hypothesis instead of the traditional N-best hypotheses for MCE training. The aim is to keep the training data as close to the trainable region as possible. Consequently, the amount of “effective” training data is increased. Furthermore, by progressively beating the nearest competitors, the training seems to be more stable. We also design an approximation algorithm based on beam search to locate the nearest competing hypothesis efficiently. We compare the performance of MCE training using 1-nearest or 1-best competing hypotheses on the Aurora database and find that the new approach (using 1-nearest hypotheses) reduces the word error rates by 5.1% and 17.8% over the latter (of using the 1-best competing hypotheses) and the official Aurora baseline respectively.
Yik-Cheung Tam, Brian Kan-Wing Mak
ICASSP1
2002 Performance of discriminatively trained auditory features on Aurora2 and Aurora3
abstract
The design of acoustic models involves two main tasks: feature extraction and data modeling; and hidden Markov modeling (HMM) is commonly used in contemporary automatic speech recognition. In the past, discriminative training has been applied successfully to rene HMM parameters that are initially trained by EM algorithm. Recently, we applied discriminative training in the feature extraction process. We proposed a novel Discriminative Auditory Feature extraction method (DAF) in which lters are discriminatively trained from data. In DAF, we do not make any assumptions on the functional form of the auditory lters except that they have to be smooth and triangular-like. On the method of discriminative training, we also proposed an alternative approach to nding the competing hypotheses which we call N-nearest hypotheses (as opposed to the traditional N-best hypotheses). By applying the two new ideas and the new robust auditory features proposed by Li et al. of Bell Labs, we reduce the overall word error rate (WER) by 30.27% over ICSLP2002 Aurora2 baseline on multi-condition training. Similarly, we obtain a relative WER reduction of 38.42% over ICSLP2002 Aurora3 baseline.
Brian Kan-Wing Mak, Yik-Cheung Tam
INTERSPEECH2
2001 Development of an asynchronous multi-band system for continuous speech recognition
abstract
Recently, multi-band automatic speech recognition (MBASR) is proposed to combat environmental noises.In this paper, we describe the two major efforts in the development of our asynchronous MBASR system for continuous speech recognition.Firstly, we successfully introduce asynchrony among sub-bands under the HMM composition framework.An asynchrony limit of one state is found adequate -relaxing the limit further does not improve performance.Secondly, the sub-band log likelihoods are combined linearly at the frame level with weightings estimated by minimizing the string classification error (MCE) among the N-best hypotheses using simulated noisy speech.When our asynchronous MBASR system is evaluated on connected TI digits with 0db additive low-pass white noise, compared with a full-band system, (1) our synchronous sub-band system reduces the absolute string error rate (SER) and word error rate (WER) by 19.8% and 14.1% respectively; (2) the introduction of asynchrony further reduces the absolute SER (WER) by 5.2%(2.5%);(3) an estimation of sub-band weightings using N-best string MCE training gives an additional reduction of absolute SER (WER) by 19.7% (5.1%).Thus, in that test, our asynchronous MBASR system has outperformed a full-band system with a 44.7% (21.7%) reduction in absolute SER (WER).In summary, N-best MCE training can effectively emphasize the more reliable sub-band, and asynchronous recombination of sub-bands is preferred.
Yik-Cheung Tam, Brian Kan-Wing Mak
INTERSPEECH1
2000 Asynchrony with trained transition probabilities improves performance in multi-band speech recognition
abstract
One of the central themes in multi-band automatic speech recognition (ASR) is to devise a strategy for recombining sub-band information. This in turn raises two questions: (1) at what phonetic unit should the recombination take place? (2) How asynchronously should the sub-bands be run? Theoretically asynchronous multi-band ASR should perform at least as well as synchronous multi-band ASR. However, in the past few years, there are conicting results on the issue. In this paper, we study the asynchrony issue under the framework of HMM composition in which a model-based recombination strategy is used to recombine sub-band HMMs at the state level. We hypothesize that re-estimation of the transition probabilities is crucial for multi-band ASR (using HMM composition). Experiments on connected TI digits show that for both clean speech and noisy speech (with additive white noise of 10db), HMMs composed from sub-band HMMs in which transition probabilities are trained with Baum-Welch algorithm o...
Brian Kan-Wing Mak, Yik-Cheung Tam
INTERSPEECH2
2000 Optimization of sub-band weights using simulated noisy speech in multi-band speech recognition
abstract
Recently multi-band speech recognition has been proposed to improve robustness under environmental noises. One important issue is how to combine decisions from individual sub-band recognizers to arrive at a nal decision. Under the hidden Markov modeling (HMM) framework, one common approach is combining sub-band likelihoods linearly in an optimal manner so that the more reliable sub-bands are emphasized and the corrupted sub-bands are de-emphasized. In our experience, estimating the weights from clean speech is not eective as the weights are not optimal under noisy environments. In this paper, we derive the optimal weights from simulated noisy speech using discriminative training method with minimum classi cation errors (MCE) or maximum mutual information (MMI) as the cost function. The methods are evaluated on recognition of isolated TI digits. Compared with full-band recognition with noises at an SNR of 0dB, multiband recognition with MCE-derived weights reduces word errors by 45.9%...
Yik-Cheung Tam, Brian Kan-Wing Mak
INTERSPEECH1