Jeremy H. M. Wong

dblp:248/4720 · also Jeremy Heng Meng Wong · DBLP profile ↗
← Back
28ranked-venue papers
18as first author
18since 2021 · last 2025
0000-0003-3742-7510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 16 first-author · 17 since 2021Artificial intelligence and machine learning · 19 · 13 first-author · 12 since 2021
YearPublicationVenuePosition
2025 A correlation-permutation approach for speech-music encoders model merging
abstract
Creating a unified speech and music model requires expensive pre-training. Model merging can instead create a unified audio model with minimal computational expense. However, direct merging is challenging when the models are not aligned in the weight space. Motivated by Git Re-Basin, we introduce a correlation-permutation approach that aligns a music encoder’s internal layers with a speech encoder. We extend previous work to the case of merging transformer layers. The method computes a permutation matrix that maximizes the model’s features-wise cross-correlations layer by layer, enabling effective fusion of these otherwise disjoint models. The merged model retains speech capabilities through this method while significantly enhancing music performance, achieving an improvement of 14.83 points in average score compared to the linear interpolation model merging. This work allows the creation of unified audio models from independently trained encoders.
Fabian Ritter Gutierrez, Yi-Cheng Lin, Jeremy H. M. Wong, Hung-yi Lee, Chng Eng Siong, Nancy F. Chen
ASRU3
2025 ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
abstract
Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically predict music impression (MI) as well as text alignment (TA) between the prompt and the generated musical piece. This paper reports our winning system, which uses a dual-branch architecture with pre-trained MuQ and RoBERTa models as audio and text encoders. A cross-attention mechanism fuses the audio and text representations. For training, we reframe the MI and TA prediction as a classification task. To incorporate the ordinal nature of MOS scores, one-hot labels are converted to a soft distribution using a Gaussian kernel. On the official test set, a single model trained with this method achieves a system-level Spearman’s Rank Correlation Coefficient (SRCC) of 0.991 for MI and 0.952 for TA, corresponding to a relative improvement of $21.21 \%$ in MI SRCC and $31.47 \%$ in TA SRCC over the challenge baseline.
Fabian Ritter Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H. M. Wong, Nancy F. Chen, Hung-yi Lee
ASRU4
2025 Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
abstract
Current large speech language models (SpeechLLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability.
Qiongqiong Wang, Hardik Bhupendra Sailor, Jeremy H. M. Wong, Tianchi Liu 0004, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw
ASRU3
2025 Obtaining objective labels and analysing annotator subjectivity by using a Rasch model for ordinal speech processing
abstract
In datasets for subjective tasks, disagreement between annotators is often accommodated by recording annotations from multiple annotators for each datapoint. However, it is more convenient to train and evaluate models against a scalar reference. In ordinal tasks, where outputs follow a monotonic order, the standard approach of computing the scalar reference as either the mean, median, or majority vote of the multiple annotations does not consider the differing bias between groups of annotators and assumes linearity of the output. This paper proposes to compute the scalar reference using a Rasch model. This expresses differing annotator bias, avoids linear assumptions, and allows control of the set of confounding variables that should influence the reference. Demonstrations on MSP-Podcast emotion recognition and speechocean762 spoken language assessment show how to use the Rasch model to compute a scalar reference, analyse confounders in the dataset, and compare multiple trained models.
Jeremy H. M. Wong, Nancy F. Chen
ASRU1
2025 Speech in-context learning of paralinguistic tasks
abstract
In-context learning adapts a large language model to a new task, without computationally expensive parameter updates. This has previously been demonstrated for text tasks, as well as speech tasks that rely primarily on lexical information, such as recognition and translation. This paper proposes to extend this investigation to consider the ability of current open-source models to exhibit in-context learning on speech tasks that require an understanding of paralinguistic information. The tasks of stutter detection, pronunciation assessment, and speech emotion recognition are investigated. The results suggest that current open-source models already exhibit some degree of speech in-context learning on paralinguistic tasks. To more fully utilise available adaptation data, it is also proposed to overcome the finite number of in-context exemplars allowed by the model’s prompt length limit, through ensemble combination over multiple in-context learning runs that each use different exemplars.
Jeremy H. M. Wong, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw
ASRU1
2025 Diversity and complementarity of speech encoders across diverse tasks in a multi-modal large language model
abstract
A Large Language Model (LLM) can be extended to understand speech inputs by using a speech encoder to compute embeddings from the speech, which are then used with a text prompt. Diverse information is expressed in speech and a wide variety of tasks can be performed. Different speech encoders may specialise toward different information types and tasks. This complementarity can be leveraged upon by using multiple speech encoders. This paper presents a comprehensive analysis of the diversity and complementarity between open-source speech encoders, when used in a multi-modal LLM framework. Experiments identify the encoders that excel in each type of downstream task, thereby guiding future system design. The diversity between encoders is measured, showing that Whisper tends to behave more differently. Diversity between encoders is compared across tasks, showing that semantic tasks tend to yield more diverse predictions. Early and late fusion show that complementarity can yield improvements.
Jeremy H. M. Wong, Muhammad Huzaifah 0001, Hardik B. Sailor, Kye Min Tan, Bin Wang 0040, Qiongqiong Wang, Xunlong Zou, Nancy F. Chen, AiTi Aw
ASRU1
2025 Distilling a speech and music encoder with task arithmetic
Fabian Ritter Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H. M. Wong, Chng Eng Siong, Nancy F. Chen, Hung-yi Lee
INTERSPEECH4
2024 Distilling Distributional Uncertainty from a Gaussian Process
abstract
A Neural Network (NN) may exhibit overconfidence about wrong hypotheses, especially for Out-Of-Domain (OOD) inputs. A Gaussian process (GP) instead has an explainable distributional uncertainty behaviour, by predicting hypotheses with greater uncertainty for query inputs further from the training data. Previous work has shown that a NN can learn to emulate the behaviour of a GP on in-domain data. This paper expands upon this, by proposing to train a NN student to emulate the GP teacher’s distributional uncertainty behaviour on OOD data. This avoids the computational cost of using a GP at run-time, while improving the OOD confidence calibration of a NN. More accurate confidence calibration may better inform how the system should feedback to the user. Experiments on the SEP-28k-E stutter detection dataset suggest that distillation of such knowledge is feasible between these models.
Jeremy H. M. Wong, Nancy F. Chen
ICASSP1
2024 Dataset-Distillation Generative Model for Speech Emotion Recognition
Fabian Ritter Gutierrez, Kuan-Po Huang, Jeremy H. M. Wong, Dianwen Ng, Hung-yi Lee, Nancy F. Chen, Chng Eng Siong
INTERSPEECH3
2024 Semi-Supervised Learning for Robust Speech Evaluation
abstract
Speech evaluation measures a learner’s oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score distribution across proficiency levels being often imbalanced among student cohorts. Automatic scoring is thus not robust when faced with under-represented samples or out-of-distribution samples, which inevitably exist in real-world deployment scenarios. This paper proposes to address such challenges by exploiting semi-supervised pre-training and objective regularization to approximate subjective evaluation criteria. In particular, normalized mutual information is used to quantify the speech characteristics from the learner and the reference. An anchor model is trained using pseudo labels to predict the correctness of pronunciation. An interpolated loss function is proposed to minimize not only the prediction error with respect to ground-truth scores but also the divergence between two probability distributions estimated by the speech evaluation model and the anchor model. Compared to other state-of-the-art methods on a public data-set, this approach not only achieves high performance while evaluating the entire test-set as a whole, but also brings the most evenly distributed prediction error across distinct proficiency levels. Furthermore, empirical results show the model accuracy on out-of-distribution data also compares favorably with competitive baselines.
Huayun Zhang, Jeremy H. M. Wong, Geyu Lin, Nancy F. Chen
SLT2
2023 Variational Gaussian Process Data Uncertainty
abstract
A Gaussian process (GP) computes an explainable distributional uncertainty by hypothesising higher output uncertainty for query inputs far from the training inputs. However, a GP may not capture data uncertainty well. Accurate data uncertainty estimation may be important for subjective tasks, such as Spoken Language Assessment (SLA), where human expert raters may disagree on the output scores. This paper shows that a variational approximation of a GP has capacity to learn data uncertainty from the training data. However, standard training criteria tune only a scalar noise hyper-parameter toward the standard deviation of the output reference, thereby limiting the learning of this uncertainty. A training criterion is proposed to explicitly encourage the GP posterior to emulate the distribution of scores from multiple raters. Experiments on the speechocean762 SLA task show that this allows the GP to better express data uncertainty and improves the modelling of inter-rater disagreements.
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
ASRU1
2023 Distilling knowledge from Gaussian process teacher to neural network student
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
INTERSPEECH1
2023 Modelling Inter-Rater Uncertainty in Spoken Language Assessment
abstract
In a subjective task, such as Spoken Language Assessment (SLA), the reference scores provided by different human raters may vary. A collection of annotated scores from multiple raters can be interpreted as an expression of data uncertainty. Previous studies often treat SLA as classification or regression tasks, and train and evaluate models against scalar reference scores that were computed from the multiple rater scores, for example by majority voting. However, a scalar representation may not adequately capture information about uncertainty that is expressed by the multiple rater scores. This paper proposes to reformulate this subjective task as a distribution fitting problem, where the model should aim to emulate the uncertainty expressed by the multiple raters. Toward this aim, the model is trained and evaluated by computing a distance between the model's output posterior and the distribution of reference scores from the multiple raters. Different methods to infer a scalar score from the model's output posterior are also considered. This paper also proposes to improve the match between the model and the SLA task, by interpreting the model's outputs as parameters of a beta density function, to capture both uncertainty and score monotonicity. Finally, ensemble combination is investigated and a novel combination method is proposed, to marginalise out model uncertainty from the combined output distribution. These approaches are evaluated on the speechocean762 dataset and an in-house Tamil dataset.
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Variations of multi-task learning for spoken language assessment
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
INTERSPEECH1
2022 Diarisation Using Location Tracking with Agglomerative Clustering
abstract
Previous works have shown that spatial location information can be complementary to speaker embeddings for a speaker diarisation task. However, the models used often assume that speakers are fairly stationary throughout a meeting. This paper proposes to relax this assumption, by explicitly modelling the movements of speakers within an Agglomerative Hierarchical Clustering (AHC) diarisation framework. Kalman filters, which track the locations of speakers, are used to compute log-likelihood ratios that contribute to the cluster affinity computations for the AHC merging and stopping decisions. Experiments show that the proposed approach is able to yield improvements on a Microsoft rich meeting transcription task, compared to methods that do not use location information or that make stationarity assumptions.
Jeremy H. M. Wong, Igor Abramovski, Yifan Gong 0001
SLT1
2022 Joint Speaker Diarisation and Tracking in Switching State-Space Model
abstract
Speakers may move around while diarisation is being performed. When a microphone array is used, the instantaneous locations of where the sounds originated from can be estimated, and previous investigations have shown that such information can be complementary to speaker embeddings in the diarisation task. However, these approaches often assume that speakers are fairly stationary throughout a meeting. This paper relaxes this assumption, by proposing to explicitly track the movements of speakers while jointly performing diarisation within a unified model. A state-space model is proposed, where the hidden state expresses the identity of the current active speaker and the predicted locations of all speakers. The model is implemented as a particle filter. Experiments on a Microsoft rich meeting transcription task show that the proposed joint location tracking and diarisation approach is able to perform comparably with other methods that use location information.
Jeremy H. M. Wong, Yifan Gong 0001
SLT1
2021 Ensemble Combination between Different Time Segmentations
abstract
Hypothesis-level combination between multiple models can often yield gains in speech recognition. However, all models in the ensemble are usually restricted to use the same audio segmentation times. This paper proposes to generalise hypothesis-level combination, allowing the use of different audio segmentation times between the models, by splitting and re-joining the hypothesised N-best lists in time. A hypothesis tree method is also proposed to distribute hypothesis posteriors among the constituent words, to facilitate such splitting when per-word scores are not available. The approach is assessed on a Microsoft meeting transcription task, by performing combination between a streaming first-pass recognition and an offline second-pass recognition. The experimental results show that the proposed approach can yield gains when combining over different segmentation times. Furthermore, the results also show that a combination between a hybrid model and an end-to-end neural network model yields a greater improvement than a combination between two hybrid models.
Jeremy H. M. Wong, Dimitrios Dimitriadis, Ken'ichi Kumatani, Yashesh Gaur, George Polovets, Partha Parthasarathy, Eric Sun, Jinyu Li 0001, Yifan Gong 0001
ICASSP1
2021 Hidden Markov Model Diarisation with Speaker Location Information
abstract
Speaker diarisation methods often rely on speaker embeddings to cluster together the segments of audio that are uttered by the same speaker. When the audio is captured using a microphone array, it is possible to estimate the locations of where the sounds originate from. This location information may be complementary to the speaker embeddings in the diarisation processes. This report proposes to extend the Hidden Markov Model (HMM) clustering method, to enable the use of speaker location information. The HMM observation log-likelihood for the speaker location can take the form of a KL-divergence, when the speaker location is represented as a discrete posterior distribution of the probabilities that the sound originated from each possible location. Experimental results on a Microsoft rich meeting transcription task show that using speaker location information with the proposed HMM modification can yield performance improvements over using speaker embeddings alone.
Jeremy H. M. Wong, Yifan Gong 0001
ICASSP1
2020 High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
abstract
While the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion followed by a sequence discriminative training criterion, we argue that such conventional hybrid models can still be significantly improved. In this paper, we detail our recent efforts to improve conventional hybrid LSTM acoustic models for high-accuracy and low-latency automatic speech recognition. To achieve high accuracy, we use a contextual layer trajectory LSTM (cltLSTM), which decouples the temporal modeling and target classification tasks, and incorporates future context frames to get more information for accurate acoustic modeling. We further improve the training strategy with sequence-level teacher-student learning. To obtain low latency, we design a two-head cltLSTM, in which one head has zero latency and the other head has a small latency, compared to an LSTM. When trained with Microsoft's 65 thousand hours of anonymized training data and evaluated with test sets with 1.8 million words, the proposed two-head cltLSTM model with the proposed training strategy yields a 28.2% relative WER reduction over the conventional LSTM acoustic model, with a similar perceived latency.
Jinyu Li 0001, Rui Zhao 0017, Eric Sun, Jeremy H. M. Wong, Amit Das 0007, Zhong Meng, Yifan Gong 0001
ICASSP4
2020 Combination of End-to-End and Hybrid Models for Speech Recognition
abstract
Recent studies suggest that it may now be possible to construct end-to-end Neural Network (NN) models that perform on-par with, or even outperform, hybrid models in speech recognition. These models differ in their designs, and as such, may exhibit diverse and complementary error patterns. A combination between the predictions of these models may therefore yield significant gains. This paper studies the feasibility of performing hypothesis-level combination between hybrid and end-to-end NN models. The end-to-end NN models often exhibit a bias in their posteriors toward short hypotheses, and this may adversely affect Minimum Bayes’ Risk (MBR) combination methods. MBR training and length normalisation can be used to reduce this bias. Models are trained on Microsoft’s 75 thousand hours of anonymised data and evaluated on test sets with 1.8 million words. The results show that significant gains can be obtained by combining the hypotheses of hybrid and end-to-end NN models together.
Jeremy H. M. Wong, Yashesh Gaur, Rui Zhao 0017, Liang Lu 0001, Eric Sun, Jinyu Li 0001, Yifan Gong 0001
INTERSPEECH1
2019 Learning Between Different Teacher and Student Models in ASR
abstract
Teacher-student learning can be applied in automatic speech recognition for model compression and domain adaptation. This trains a student model to emulate the behaviour of a teacher model, and only the student is used to perform recognition. Depending on the application, the teacher and student may differ in their model types, complexities, input contexts, and input features. In previous works, it is often shown that learning from a strong teacher allows the student to perform better than an equivalent model trained with only the reference transcriptions. However, there has not been much investigation into whether a particular form of teacher is appropriate for the student to learn from. This paper aims to study how effectively the student is able to learn from the teacher, when differences exist between their designs. The Augmented Multi-party Interaction (AMI) meeting transcription and Multi-Genre Broadcast (MGB-3) television broadcast audio tasks are used in this analysis. Experimental results suggest that a student can effectively learn from a more complex teacher, but may struggle when it lacks input information. It is therefore important to carefully consider the design of the student for each application.
Jeremy H. M. Wong, Mark J. F. Gales, Yu Wang 0027
ASRU1
2019 Exploiting Future Word Contexts in Neural Network Language Models for Speech Recognition
abstract
Language modeling is a crucial component in a wide range of applications including speech recognition. Language models (LMs) are usually constructed by splitting a sentence into words and computing the probability of a word based on its word history. This sentence probability calculation, making use of conditional probability distributions, assumes that there is little impact from approximations used in the LMs, including the word history representations and finite training data. This motivates examining models that make use of additional information from the sentence. In this paper, future word information, in addition to the history, is used to predict the probability of the current word. For recurrent neural network LMs (RNNLMs), this information can be encapsulated in a bi-directional model. However, if used directly, this form of model is computationally expensive when trained on large quantities of data, and can be problematic when used with word lattices. This paper proposes a novel neural network language model structure, the succeeding-word RNNLM, su-RNNLM, to address these issues. Instead of using a recurrent unit to capture the complete future word contexts, a feedforward unit is used to model a fixed finite number of succeeding words. This is more efficient in training than bi-directional models and can be applied to lattice rescoring. The generated lattices can be used for downstream applications, such as confusion network decoding and keyword search. Experimental results on speech recognition and keyword spotting tasks illustrate the empirical usefulness of future word information, and the flexibility of the proposed model to represent this information.
Xie Chen 0001, Xunying Liu, Yu Wang 0027, Anton Ragni, Jeremy H. M. Wong, Mark J. F. Gales
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 General Sequence Teacher-Student Learning
abstract
In automatic speech recognition, performance gains can often be obtained by combining an ensemble of multiple models. However, this can be computationally expensive when performing recognition. Teacher-student learning alleviates this cost by training a single student model to emulate the combined ensemble behaviour. Only this student needs to be used for recognition. Previously investigated teacher-student criteria often limit the forms of diversity allowed in the ensemble, and only propagate information from the teachers to the student at the frame level. This paper addresses both of these issues by examining teacher-student learning within a sequence-level framework, and assessing the flexibility that these approaches offer. Various sequence-level teacher-student criteria are examined in this work, to propagate sequence posterior information. A training criterion based on the Kullback-Leibler (KL)-divergence between context-dependent state sequence posteriors is proposed that allows for a diversity of state cluster sets to be present in the ensemble. This criterion is shown to be an upper bound to a more general KL-divergence between word sequence posteriors, which places even fewer restrictions on the ensemble diversity, but whose gradient can be expensive to compute. These methods are evaluated on the augmented multi-party interaction (AMI) meeting transcription and MGB-3 television broadcast audio tasks.
Jeremy H. M. Wong, Mark J. F. Gales, Yu Wang 0027
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription
abstract
State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to the spoken form is complicated. However, in recent years the representational power of deep-learning based acoustic models has improved, raising interest in graphemic acoustic models for English, due to the simplicity of generating the lexicon. In this paper, phonetic and graphemic models are compared for an English Multi-Genre Broadcast transcription task. A range of acoustic models based on lattice-free MMI training are constructed using phonetic and graphemic lexicons. For this task, it is found that having a long-span temporal history reduces the difference in performance between the two forms of models. In addition, system combination is examined, using parameter smoothing and hypothesis combination. As the combination approaches become more complicated the difference between the phonetic and graphemic systems further decreases. Finally, for all configurations examined the combination of phonetic and graphemic systems yields consistent gains.
Yu Wang 0027, Xie Chen 0001, Mark J. F. Gales, Anton Ragni, Jeremy H. M. Wong
ICASSP5
2018 Sequence Teacher-Student Training of Acoustic Models for Automatic Free Speaking Language Assessment
abstract
A high performance automatic speech recognition (ASR) system is an important constituent component of an automatic language assessment system for free speaking language tests. The ASR system is required to be capable of recognising non-native spontaneous English speech and to be deployable under real-time conditions. The performance of ASR systems can often be significantly improved by leveraging upon multiple systems that are complementary, such as an ensemble. Ensemble methods, however, can be computationally expensive, often requiring multiple decoding runs, which makes them impractical for deployment. In this paper, a lattice-free implementation of sequence-level teacher-student training is used to reduce this computational cost, thereby allowing for real-time applications. This method allows a single student model to emulate the performance of an ensemble of teachers, but without the need for multiple decoding runs. Adaptations of the student model to speakers from different first languages (L1s) and grades are also explored.
Yu Wang 0027, Jeremy H. M. Wong, Mark J. F. Gales, Kate M. Knill, Anton Ragni
SLT2
2017 Multi-task ensembles with teacher-student training
abstract
Ensemble methods often yield significant gains for automatic speech recognition. One method to obtain a diverse ensemble is to separately train models with a range of context dependent targets, often implemented as state clusters. However, decoding the complete ensemble can be computationally expensive. To reduce this cost, the ensemble can be generated using a multi-task architecture. Here, the hidden layers are merged across all members of the ensemble, leaving only separate output layers for each set of targets. Previous investigations of this form of ensemble have used cross-entropy training, which is shown in this paper to produce only limited diversity between members of the ensemble. This paper extends the multi-task framework in several ways. First, the multi-task ensemble can be trained in a teacher-student fashion toward the ensemble of separate models, with the aim of increasing diversity. Second, the multi-task ensemble can be trained with a sequence discriminative criterion. Finally, a student model, with a single output layer, can be trained to emulate the combined ensemble, to further reduce the computational cost of decoding. These methods are evaluated on the Babel conversational telephone speech, AMI meeting transcription, and HUB4 English broadcast news tasks.
Jeremy H. M. Wong, Mark J. F. Gales
ASRU1
2017 Student-Teacher Training with Diverse Decision Tree Ensembles
abstract
Student-teacher training allows a large teacher model or ensemble of teachers to be compressed into a single student model, for the purpose of efficient decoding. However, current approaches in automatic speech recognition assume that the state clusters, often defined by Phonetic Decision Trees (PDT), are the same across all models. This limits the diversity that can be captured within the ensemble, and also the flexibility when selecting the complexity of the student model output. This paper examines an extension to student-teacher training that allows for the possibility of having different PDTs between teachers, and also for the student to have a different PDT from the teacher. The proposal is to train the student to emulate the logical context dependent state posteriors of the teacher, instead of the frame posteriors. This leads to a method of mapping frame posteriors from one PDT to another. This approach is evaluated on three speech recognition tasks: the Tok Pisin and Javanese low resource conversational telephone speech tasks from the IARPA Babel programme, and the HUB4 English broadcast news task.
Jeremy H. M. Wong, Mark J. F. Gales
INTERSPEECH1
2016 Sequence Student-Teacher Training of Deep Neural Networks
abstract
The performance of automatic speech recognition can often be significantly improved by combining multiple systems together. Though beneficial, ensemble methods can be computationally expensive, often requiring multiple decoding runs. An alternative approach, appropriate for deep learning schemes, is to adopt student-teacher training. Here, a student model is trained to reproduce the outputs of a teacher model, or ensemble of teachers. The standard approach is to train the student model on the frame posterior outputs of the teacher. This paper examines the interaction between student-teacher training schemes and sequence training criteria, which have been shown to yield significant performance gains over frame-level criteria. There are several possible options for integrating sequence training, including training of the ensemble and further training of the student. This paper also proposes an extension to the student-teacher framework, where the student is trained to emulate the hypothesis posterior distribution of the teacher, or ensemble of teachers. This sequence student-teacher training approach allows the benefit of student-teacher training to be directly combined with sequence training schemes. These approaches are evaluated on two speech recognition tasks: a Wall Street Journal based task and a low-resource Tok Pisin conversational telephone speech task from the IARPA Babel programme.
Jeremy H. M. Wong, Mark J. F. Gales
INTERSPEECH1