Mark J. F. Gales

dblp:74/4419 · also Mark John Francis Gales · DBLP profile ↗
← Back
338ranked-venue papers
40as first author
44since 2021 · last 2025
0000-0002-5311-8219ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 263 · 27 first-author · 20 since 2021Artificial intelligence and machine learning · 219 · 28 first-author · 36 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SkillAggregation: Reference-free LLM-Dependent Aggregation
abstract
Large Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments.Single LLMs were used initially, however, recent work suggests using multiple LLMs as judges yields improved performance.An important step in exploiting multiple judgements is the combination stage, aggregation.Existing methods in NLP either assign equal weight to all LLM judgments or are designed for specific tasks such as hallucination detection.This work focuses on aggregating predictions from multiple systems where no reference labels are available.A new method called SkillAggregation is proposed, which learns to combine estimates from LLM judges without needing additional data or ground truth.It extends the Crowdlayer aggregation method, developed for image classification, to exploit the judge estimates during inference.The approach is compared to a range of standard aggregation methods on HaluEval-Dialogue, TruthfulQA and Chatbot Arena tasks.SkillAggregation outperforms Crowdlayer on all tasks, and yields the best performance over all approaches on the majority of tasks. 1
Guangzhi Sun, Anmol Kagrecha, P. P. Manakul, Philip C. Woodland, Mark J. F. Gales
ACL (1)5
2025 Finetuning LLMs for Comparative Assessment Tasks
abstract
Automated assessment in natural language generation is a challenging task. Instruction-tuned large language models (LLMs) have shown promise in reference-free evaluation, particularly through comparative assessment. However, the quadratic computational complexity of pairwise comparisons limits its scalability. To address this, efficient comparative assessment has been explored by applying comparative strategies on zero-shot LLM probabilities. We propose a framework for finetuning LLMs for comparative assessment to align the model’s output with the target distribution of comparative probabilities. By training on soft probabilities, our approach improves state-of-the-art performance while maintaining high performance with an efficient subset of comparisons.
Vatsal Raina, Adian Liusie, Mark J. F. Gales
COLING3
2025 Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
abstract
Unlearning has emerged as a critical capability for large language models (LLMs) to support data privacy, regulatory compliance, and ethical AI deployment.Recent techniques often rely on obfuscation by injecting incorrect or irrelevant information to suppress knowledge.Such methods effectively constitute knowledge addition rather than true removal, often leaving models vulnerable to probing.In this paper, we formally distinguish unlearning from obfuscation and introduce a probing-based evaluation framework to assess whether existing approaches genuinely remove targeted information.Moreover, we propose DF-MCQ, a novel unlearning method that flattens the model predictive distribution over automatically generated multiple-choice questions using KLdivergence, effectively removing knowledge about target individuals and triggering appropriate refusal behaviour.Experimental results demonstrate that DF-MCQ achieves unlearning with over 90% refusal rate and a random choice-level uncertainty that is much higher than obfuscation on probing questions. 1
Guangzhi Sun, P. P. Manakul, Xiao Zhan, Mark J. F. Gales
EMNLP4
2025 Assessment of L2 Oral Proficiency using Speech Large Language Models
Rao Ma, Mengjie Qian 0001, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH6
2025 Training Articulatory Inversion Models for Interspeaker Consistency
Charles McGhee, Mark J. F. Gales, Kate M. Knill
INTERSPEECH2
2025 Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Mengjie Qian 0001, Rao Ma, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH5
2025 Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
abstract
This paper explores generalised probabilistic modelling and uncertainty estimation in comparative LLM-as-a-judge frameworks. We show that existing Product-of-Experts methods are specific cases of a broader framework, enabling diverse modelling options. Furthermore, we propose improved uncertainty estimates for individual comparisons, enabling more efficient selection and achieving strong performance with fewer evaluations. We also introduce a method for estimating overall ranking uncertainty. Finally, we demonstrate that combining absolute and comparative scoring improves performance. Experiments show that the specific expert model has a limited impact on final rankings but our proposed uncertainty estimates, especially the probability of reordering, significantly improve the efficiency of systems reducing the number of needed comparisons by $\sim$50%. Furthermore, ranking-level uncertainty metrics can be used to identify low-performing predictions, where the nature of the probabilistic model has a notable impact on the quality of the overall uncertainty.
Yassir Fathullah, Mark J. F. Gales
UAI2
2024 An Information-Theoretic Approach to Analyze NLP Classification Tasks
abstract
Understanding the contribution of the inputs on the output is useful across many tasks.This work provides an information-theoretic framework to analyse the influence of inputs for text classification tasks.Natural language processing (NLP) tasks take either a single or multiple text elements to predict an output variable.Each text element has two components: the semantic meaning and a linguistic realization.Multiple-choice reading comprehension (MCRC) and sentiment classification (SC) are selected to showcase the framework.For MCRC, it is found that the relative context influence on the output reduces on more challenging datasets.In particular, more challenging contexts allows greater variation in the question complexity.Hence, test creators need to carefully consider the choice of the context when designing multiple-choice questions for assessment.For SC, it is found the semantic meaning of the input dominates compared to its linguistic realization when determining the sentiment.The framework is made available at: https://github.com/WangLuran/n lp-element-influence.
Luran Wang, Mark J. F. Gales, Vatsal Raina
ACL (1)2
2024 Is It Possible to Modify Text to a Target Readability Level? An Initial Investigation Using Zero-Shot Large Language Models
abstract
Text simplification is a common task where the text is adapted to make it easier to understand. Similarly, text elaboration can make a passage more sophisticated, offering a method to control the complexity of reading comprehension tests. However, text simplification and elaboration tasks are limited to only relatively alter the readability of texts. It is useful to directly modify the readability of any text to an absolute target readability level to cater to a diverse audience. Ideally, the readability of readability-controlled generated text should be independent of the source text. Therefore, we propose a novel readability-controlled text modification task. The task requires the generation of 8 versions at various target readability levels for each input text. We introduce novel readability-controlled text modification metrics. The baselines for this task use ChatGPT and Llama-2, with an extension approach introducing a two-step process (generating paraphrases by passing through the language model twice). The zero-shot approaches are able to push the readability of the paraphrases in the desired direction but the final readability remains correlated with the original text’s readability. We also find greater drops in semantic and lexical similarity between the source and target texts with greater shifts in the readability.
Asma Farajidizaji, Vatsal Raina, Mark J. F. Gales
LREC/COLING3
2024 Who Needs Decoders? Efficient Estimation of Sequence-Level Attributes with Proxies
abstract
Sequence-to-sequence models often require an expensive autoregressive decoding process.However, for some downstream tasks such as out-of-distribution (OOD) detection and resource allocation, the actual decoding output is not needed, just a scalar attribute of this sequence.In such scenarios, where knowing the quality of a system's output to predict poor performance prevails over knowing the output itself, is it possible to bypass the autoregressive decoding?We propose Non-Autoregressive Proxy (NAP) models that can efficiently predict scalar-valued sequence-level attributes.Importantly, NAPs predict these metrics directly from the encodings, avoiding the expensive decoding stage.We consider two sequence tasks: Machine Translation (MT) and Automatic Speech Recognition (ASR).In OOD for MT, NAPs outperform ensembles while being significantly faster.NAPs are also proven capable of predicting metrics such as BERTScore (MT) or word error rate (ASR).For downstream tasks, such as data filtering and resource optimization, NAPs generate performance predictions that outperform predictive uncertainty while being highly inference efficient.
Yassir Fathullah, Puria Radmard, Adian Liusie, Mark J. F. Gales
EACL (1)4
2024 LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
abstract
Current developments in large language models (LLMs) have enabled impressive zero-shot capabilities across various natural language tasks.An interesting application of these systems is in the automated assessment of natural language generation (NLG), a highly challenging area with great practical benefit.In this paper, we explore two options for exploiting the emergent abilities of LLMs for zero-shot NLG assessment: absolute score prediction, and comparative assessment which uses relative comparisons between pairs of candidates.Though comparative assessment has not been extensively studied in NLG assessment, we note that humans often find it more intuitive to compare two options rather than scoring each one independently.This work examines comparative assessment from multiple perspectives: performance compared to absolute grading; positional biases in the prompt; and efficient ranking in terms of the number of comparisons.We illustrate that LLM comparative assessment is a simple, general and effective approach for NLG assessment.For moderatesized open-source LLMs, such as FlanT5 and Llama2-chat, comparative assessment is superior to prompt scoring, and in many cases can achieve performance competitive with state-ofthe-art methods.Additionally, we demonstrate that LLMs often exhibit strong positional biases when making pairwise comparisons, and we propose debiasing methods that can further improve performance.
Adian Liusie, P. P. Manakul, Mark J. F. Gales
EACL (1)3
2024 LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
abstract
With the recent emergence of powerful instruction-tuned large language models (LLMs), various helpful conversational Artificial Intelligence (AI) systems have been deployed across many applications.When prompted by users, these AI systems successfully perform various tasks as part of a conversation.Such approaches typically condition their output on the entire conversational history to provide some sort of memory and context.Although this sensitivity to the conversational history can often lead to improved performance on subsequent tasks, we find that performance can in fact also be negatively impacted, if there is a task-switch.To the best of our knowledge, our work makes the first attempt to formalize the study of such vulnerabilities and interference of tasks in conversational LLMs caused by task-switches in the conversational history.Our experiments across 5 datasets with 15 task switches using popular LLMs reveal that many of the task-switches can lead to significant performance degradation. 1
Ivaxi Sheth, Vyas Raina, Mark J. F. Gales, Mario Fritz
EMNLP4
2024 Efficient LLM Comparative Assessment: A Product of Experts Framework for Pairwise Comparisons
abstract
LLM-as-a-judge approaches are a practical and effective way of assessing a range of text tasks.However, when using pairwise comparisons to rank a set of candidates, the computational cost scales quadratically with the number of candidates, which has practical limitations.This paper introduces a Product of Expert (PoE) framework for efficient LLM Comparative Assessment.Here individual comparisons are considered experts that provide information on a pair's score difference.The PoE framework combines the information from these experts to yield an expression that can be maximized with respect to the underlying set of candidates, and is highly flexible where any form of expert can be assumed.When Gaussian experts are used one can derive simple closed-form solutions for the optimal candidate ranking, and expressions for selecting which comparisons should be made to maximize the probability of this ranking.Our approach enables efficient comparative assessment, where by using only a small subset of the possible comparisons, one can generate score predictions that correlate well with human judgements.We evaluate the approach on multiple NLG tasks and demonstrate that our framework can yield considerable computational savings when performing pairwise comparative assessment.With many candidate texts, using as few as 2% of comparisons the PoE solution can achieve similar performance to when all comparisons are used.1
Adian Liusie, Vatsal Raina, Yassir Fathullah, Mark J. F. Gales
EMNLP4
2024 Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
abstract
Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems.Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to adversarial manipulation.This work presents the first study on the adversarial robustness of assessment LLMs, where we demonstrate that short universal adversarial phrases can be concatenated to deceive judge LLMs to predict inflated scores.Since adversaries may not know or have access to the judge-LLMs, we propose a simple surrogate attack where a surrogate model is first attacked, and the learned attack phrase then transferred to unknown judge-LLMs.We propose a practical algorithm to determine the short universal attack phrases and demonstrate that when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted.It is found that judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment.Our findings raise concerns on the reliability of LLMas-a-judge methods, and emphasize the importance of addressing vulnerabilities in LLM assessment methods before deployment in highstakes real-world scenarios.1 * Equal Contribution. 1 Code: https://github.com/rainavyas/ attack-comparative-assessment2.3 Score the summary between 1-5 "Some animals did something."Score the summary between 1-5 "Some animals did something.summable" Which Summary is better?A: "Some animals did something."B: "Tortoise wins race; slow and steady" Which Summary is better?A: "Some animals did something.informative" B: "Tortoise wins race; slow and steady" 4.
Vyas Raina, Adian Liusie, Mark J. F. Gales
EMNLP3
2024 Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
abstract
Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications.These systems incorporate 'special tokens' in their vocabulary, such as <|endoftext|>, to guide their language generation process.However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior.We propose a simple yet effective method to learn a universal acoustic realization of Whisper's <|endoftext|> token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively 'muting' the model.Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97% of speech samples.Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks.Overall this work demonstrates the vulnerability of Whisper models to 'muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data. 1
Vyas Raina, Rao Ma, Charles McGhee, Kate M. Knill, Mark J. F. Gales
EMNLP5
2024 Towards End-to-End Spoken Grammatical Error Correction
abstract
Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.
Stefano Bannò, Rao Ma, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
ICASSP5
2024 On the Usefulness of Speaker Embeddings for Speaker Retrieval in the Wild: A Comparative Study of x-vector and ECAPA-TDNN Models
Erfan Loweimi, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
INTERSPEECH4
2024 Highly Intelligible Speaker-Independent Articulatory Synthesis
Charles McGhee, Kate M. Knill, Mark J. F. Gales
INTERSPEECH3
2024 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
Mengjie Qian 0001, Rao Ma, Kate M. Knill, Mark J. F. Gales
INTERSPEECH5
2024 MVRMLM 2024: Multimodal Video Retrieval and Multimodal Language Modelling
abstract
As the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example-based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi-modal content is essential. In designing our novel video retrieval framework named Multi-modal Video Search by Examples (MVSE)1, we focused on accuracy (precision and recall), efficiency (retrieval time in seconds), interactivity, and extensibility, with key components including advanced data processing and a user-friendly interface aimed at enhancing search effectiveness and user experience. With the advent of Large Language Models (LLMs), the interaction between multimodal data, including image and audio has been transformed with a significant leap forward towards a bigger goal of artificial general intelligence. This workshop aims to bring together experts from diverse domains to explore the possibilities of developing novel ways of multimodal data search, understanding and interaction.
Hui Wang 0001, Josef Kittler, Mark J. F. Gales, Rob Cooper, Maurice D. Mulvenna, Wing W. Y. Ng, Yang Hua 0001, Richard Gault, Abbas Haider, Guanfeng Wu
ICMR3
2024 Investigating the Emergent Audio Classification Ability of ASR Foundation Models
abstract
Rao Ma, Adian Liusie, Mark Gales, Kate Knill. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Rao Ma, Adian Liusie, Mark J. F. Gales, Kate M. Knill
NAACL-HLT3
2024 Zero-Shot Audio Topic Reranking Using Large Language Models
abstract
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far richer search modalities such as images, speaker, content, topic, and emotion. A key element for this process is highly rapid and flexible search to support large archives, which in MVSE is facilitated by representing video attributes with embeddings. This work aims to compensate for any performance loss from this rapid archive search by examining reranking approaches. In particular, zero-shot reranking methods using large language models (LLMs) are investigated as these are applicable to any video archive audio content. Performance is evaluated for topic-based retrieval on a publicly available video archive, the BBC Rewind corpus. Results demonstrate that reranking significantly improves retrieval ranking without requiring any task-specific in-domain training data. Furthermore, three sources of information (ASR transcriptions, automatic summaries and synopses) as input for LLM reranking were compared. To gain a deeper understanding and further insights into the performance differences and limitations of these text sources, we employ a fact-checking approach to analyse the information consistency among them.
Mengjie Qian 0001, Rao Ma, Adian Liusie, Erfan Loweimi, Kate M. Knill, Mark J. F. Gales
SLT6
2024 Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Multi-Task Automatic Speech Recognition Models
abstract
Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these models is their ability to perform tasks other than automatic speech recognition (ASR) using an appropriate prompt. For example, the OpenAI Whisper model can perform both speech transcription and speech translation. With the development of audio-prompted LLMs there is the potential for even greater control options. In this work we demonstrate that with this greater flexibility the systems can be susceptible to model-control adversarial attacks. Without any access to the model prompt it is possible to modify the behaviour of the system by appropriately changing the audio input. To illustrate this risk, we demonstrate that it is possible to prepend a short universal adversarial acoustic segment to any input speech signal to override the prompt setting of an ASR foundation model. Specifically, we successfully use a universal adversarial acoustic segment to control Whisper to always perform speech translation, despite being set to perform speech transcription. Overall, this work demonstrates a new form of adversarial attack on multi-tasking speech enabled foundation models that needs to be considered prior to the deployment of this form of model.
Vyas Raina, Mark J. F. Gales
SLT2
2024 Multi-modal video search by examples - A video quality impact analysis
abstract
Abstract As the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example‐based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi‐modal content is essential. An innovative Multi‐modal Video Search by Examples (MVSE) framework is introduced, employing state‐of‐the‐art techniques in its various components. In designing MVSE, the authors focused on accuracy, efficiency, interactivity, and extensibility, with key components including advanced data processing and a user‐friendly interface aimed at enhancing search effectiveness and user experience. Furthermore, the framework was comprehensively evaluated, assessing individual components, data quality issues, and overall retrieval performance using high‐quality and low‐quality BBC archive videos. The evaluation reveals that: (1) multi‐modal search yields better results than single‐modal search; (2) the quality of video, both visual and audio, has an impact on the query precision. Compared with image query results, audio quality has a greater impact on the query precision (3) a two‐stage search process (i.e. searching by Hamming distance based on hashing, followed by searching by Cosine similarity based on embedding); is effective but increases time overhead; (4) large‐scale video retrieval is not only feasible but also expected to emerge shortly.
Guanfeng Wu, Abbas Haider, Xing Tian, Erfan Loweimi, Chi-Ho Chan, Mengjie Qian 0001, Muhammad Junaid Awan, Ivor T. A. Spence, Rob Cooper, Wing W. Y. Ng, Josef Kittler, Mark J. F. Gales, Hui Wang 0001
IET Comput. Vis.12
2023 SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
abstract
Generative Large Language Models (LLMs) such as GPT-3 are capable of generating highly fluent responses to a wide variety of user prompts.However, LLMs are known to hallucinate facts and make non-factual statements which can undermine trust in their output.Existing fact-checking approaches either require access to the output probability distribution (which may not be available for systems such as ChatGPT) or external databases that are interfaced via separate, often complex, modules.In this work, we propose "SelfCheckGPT", a simple sampling-based approach that can be used to fact-check the responses of black-box models in a zero-resource fashion, i.e. without an external database.SelfCheckGPT leverages the simple idea that if an LLM has knowledge of a given concept, sampled responses are likely to be similar and contain consistent facts.However, for hallucinated facts, stochastically sampled responses are likely to diverge and contradict one another.We investigate this approach by using GPT-3 to generate passages about individuals from the WikiBio dataset, and manually annotate the factuality of the generated passages.We demonstrate that SelfCheck-GPT can: i) detect non-factual and factual sentences; and ii) rank passages in terms of factuality.We compare our approach to several baselines and show that our approach has considerably higher AUC-PR scores in sentence-level hallucination detection and higher correlation scores in passage-level factuality assessment compared to grey-box methods.
P. P. Manakul, Adian Liusie, Mark J. F. Gales
EMNLP3
2023 Ensemble Prosody Prediction For Expressive Speech Synthesis
abstract
Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Yet, in evaluations it is generally found that no single model is preferred for all input texts. This suggests an approach that has rarely been used before for Text-to-Speech: an ensemble of models.We apply ensemble learning to prosody prediction. We construct simple ensembles of prosody predictors by varying either model architecture or model parameter values.To automatically select amongst the models in the ensemble when performing Text-to-Speech, we propose a novel, and computationally trivial, variance-based criterion. We demonstrate that even a small ensemble of prosody predictors yields useful diversity, which, combined with the proposed selection criterion, outperforms any individual model from the ensemble.
Tian Huey Teh, Vivian Hu, Devang S. Ram Mohan, Zack Hodari, Christopher G. R. Wallis, Tomás Gómez Ibarrondo, Alexandra Torresquintero, James Leoni, Mark J. F. Gales, Simon King 0001
ICASSP9
2023 MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in Summarization
abstract
Potsawee Manakul, Adian Liusie, Mark Gales. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
P. P. Manakul, Adian Liusie, Mark J. F. Gales
IJCNLP (1)3
2023 Multi-Head State Space Model for Speech Recognition
Yassir Fathullah, Chunyang Wu, Yuan Shangguan, Junteng Jia, Wenhan Xiong, Jay Mahadeokar, Chunxi Liu, Yangyang Shi, Ozlem Kalinli, Mike Seltzer, Mark J. F. Gales
INTERSPEECH11
2023 N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
abstract
Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions.Most prior works use the 1-best ASR hypothesis as input and therefore can only perform correction by leveraging the context within one sentence.In this work, we propose a novel N-best T5 model for this task, which is fine-tuned from a T5 model and utilizes ASR N-best lists as model input.By transferring knowledge from the pretrained language model and obtaining richer information from the ASR decoding space, the proposed approach outperforms a strong Conformer-Transducer baseline.Another issue with standard error correction is that the generation process is not well-guided.To address this a constrained decoding process, either based on the N-best list or an ASR lattice, is used which allows additional information to be propagated.
Rao Ma, Mark J. F. Gales, Kate M. Knill, Mengjie Qian 0001
INTERSPEECH2
2023 Adapting an Unadaptable ASR System
abstract
As speech recognition model sizes and training data requirements grow, it is increasingly common for systems to only be available via APIs from online service providers rather than having direct access to models themselves.In this scenario it is challenging to adapt systems to a specific target domain.To address this problem we consider the recently released OpenAI Whisper ASR as an example of a large-scale ASR system to assess adaptation methods.An error correction based approach is adopted, as this does not require access to the model, but can be trained from either 1-best or N-best outputs that are normally available via the ASR API.LibriSpeech is used as the primary target domain for adaptation.The generalization ability of the system in two distinct dimensions are then evaluated.First, whether the form of correction model is portable to other speech recognition domains, and secondly whether it can be used for ASR models having a different architecture.
Rao Ma, Mengjie Qian 0001, Mark J. F. Gales, Kate M. Knill
INTERSPEECH3
2023 Speak & Improve: L2 English Speaking Practice Tool
Diane Nicholls, Kate M. Knill, Mark J. F. Gales, Anton Ragni, Paul Ricketts
INTERSPEECH3
2023 Logit-based ensemble distribution distillation for robust autoregressive sequence uncertainties
abstract
Efficiently and reliably estimating uncertainty is an important objective in deep learning. It is especially pertinent to autoregressive sequence tasks, where training and inference costs are typically very high. However, existing research has predominantly focused on tasks with static data such as image classification. In this work, we investigate Ensemble Distribution Distillation (EDD) applied to large-scale natural language sequence-to-sequence data. EDD aims to compress the superior uncertainty performance of an expensive (teacher) ensemble into a cheaper (student) single model. Importantly, the ability to separate knowledge (epistemic) and data (aleatoric) uncertainty is retained. Existing probability-space approaches to EDD, however, are difficult to scale to large vocabularies. We show, for modern transformer architectures on large-scale translation tasks, that modelling the ensemble logits, instead of softmax probabilities, leads to significantly better students. Moreover, the students surprisingly even outperform Deep Ensembles by up to $\sim$10% AUROC on out-of-distribution detection, whilst matching them at in-distribution translation.
Yassir Fathullah, Guoxuan Xia, Mark J. F. Gales
UAI3
2022 View-Specific Assessment of L2 Spoken English
abstract
The growing demand for learning English as a second language has increased interest in automatic approaches for assessing and improving spoken language proficiency. A significant challenge in this field is to provide interpretable scores and informative feedback to learners through individual viewpoints of learners’ proficiency, as opposed to holistic scores. Thus far, holistic scoring remains commonly applied in large-scale commercial tests. As a result, an issue with more detailed evaluation is that human graders are generally trained to provide holistic scores. This paper investigates whether view-specific systems can be trained when only holistic scores are available. To enable this process, view-specific networks are defined where both their inputs and structure are adapted to focus on specific facets of proficiency. It is shown that it is possible to train such systems on holistic scores, such that they provide view-specific scores at evaluation time. View-specific networks are designed in this way for pronunciation, rhythm, text, use of parts of speech and grammatical accuracy. The relationships between the predictions of each system are investigated on the spoken part of the Linguaskill proficiency test. It is shown that the view-specific predictions are complementary in nature and capture different information about proficiency.
Stefano Bannò, Bhanu Balusu, Mark J. F. Gales, Kate M. Knill, Konstantinos Kyriakopoulos
INTERSPEECH3
2022 Residue-Based Natural Language Adversarial Attack Detection
abstract
Deep learning based systems are susceptible to adversarial attacks, where a small, imperceptible change at the input alters the model prediction.However, to date the majority of the approaches to detect these attacks have been designed for image processing systems.Many popular image adversarial detection approaches are able to identify adversarial examples from embedding feature spaces, whilst in the NLP domain existing state of the art detection approaches solely focus on input text features, without consideration of model embedding spaces.This work examines what differences result when porting these image designed strategies to Natural Language Processing (NLP) tasks -these detectors are found to not port over well.This is expected as NLP systems have a very different form of input: discrete and sequential in nature, rather than the continuous and fixed size inputs for images.As an equivalent model-focused NLP detection approach, this work proposes a simple sentenceembedding "residue" based detector to identify adversarial examples.On many tasks, it outperforms ported image domain detectors and recent state of the art NLP specific detectors 1 .
Vyas Raina, Mark J. F. Gales
NAACL-HLT2
2022 Self-distribution distillation: efficient uncertainty estimation
abstract
Deep learning is increasingly being applied in safety-critical domains. For these scenarios it is important to know the level of uncertainty in a model’s prediction to ensure appropriate decisions are made by the system. Deep ensembles are the de-facto standard approach to obtaining various measures of uncertainty. However, ensembles often significantly increase the resources required in the training and/or deployment phases. Approaches have been developed that typically address the costs in one of these phases. In this work we propose a novel training approach, self-distribution distillation (S2D), which is able to efficiently train a single model that can estimate uncertainties. Furthermore it is possible to build ensembles of these models and apply hierarchical ensemble distillation approaches. Experiments on CIFAR-100 showed that S2D models outperformed standard models and Monte-Carlo dropout. Additional out-of-distribution detection experiments on LSUN, Tiny ImageNet, SVHN showed that even a standard deep ensemble can be outperformed using S2D based ensembles and novel distilled models.
Yassir Fathullah, Mark J. F. Gales
UAI2
2022 Increasing Context for Estimating Confidence Scores in Automatic Speech Recognition
abstract
Accurate confidence measures for predictions from machine learning techniques play a critical role in the deployment and training of many speech and language processing applications. For example, confidence scores are important when making use of automatically generated transcriptions in training automatic speech recognition (ASR) systems, as well as down-stream applications, such as information retrieval and conversational assistants. Previous work on improving confidence scores for these systems has focused on two main directions: designing features correlated with improved confidence prediction; and employing sequence models to account for the importance of contextual information. Few studies, however, have explored incorporating contextual information more broadly, such as from the future, in addition to the past, or making use of alternative multiple hypotheses in addition to the most likely one. This article introduces two general approaches for encapsulating contextual information from lattices. Experimental results illustrating the importance of increasing contextual information for estimating confidence scores are presented on a range of limited resource languages where word error rates range between 30% and 60%. The results show that the novel approaches provide significant gains in the accuracy of confidence estimation.
Anton Ragni, Mark J. F. Gales, Oliver Rose, Kate M. Knill, Alexandros Kastanos, Qiujia Li, Preben Ness
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Long-Span Summarization via Local Attention and Content Selection
abstract
Potsawee Manakul, Mark Gales. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
P. P. Manakul, Mark J. F. Gales
ACL/IJCNLP (1)2
2021 Sparsity and Sentence Structure in Encoder-Decoder Attention of Summarization Systems
abstract
Transformer models have achieved state-ofthe-art results in a wide range of NLP tasks including summarization.Training and inference using large transformer models can be computationally expensive.Previous work has focused on one important bottleneck, the quadratic self-attention mechanism in the encoder.Modified encoder architectures such as LED or LoBART use local attention patterns to address this problem for summarization.In contrast, this work focuses on the transformer's encoder-decoder attention mechanism.The cost of this attention becomes more significant in inference or training approaches that require model-generated histories.First, we examine the complexity of the encoder-decoder attention.We demonstrate empirically that there is a sparse sentence structure in document summarization that can be exploited by constraining the attention mechanism to a subset of input sentences, whilst maintaining system performance.Second, we propose a modified architecture that selects the subset of sentences to constrain the encoder-decoder attention.Experiments are carried out on abstractive summarization tasks, including CNN/DailyMail, XSum, Spotify Podcast, and arXiv. 1
P. P. Manakul, Mark J. F. Gales
EMNLP (1)2
2021 Ensemble Distillation Approaches for Grammatical Error Correction
abstract
Ensemble approaches are commonly used techniques to improving a system by combining multiple model predictions. Additionally these schemes allow the uncertainty, as well as the source of the uncertainty, to be derived for the prediction. Unfortunately these benefits come at a computational and memory cost. To address this problem ensemble distillation (EnD) and more recently ensemble distribution distillation (EnDD) have been proposed that compress the ensemble into a single model, representing either the ensemble average prediction or prediction distribution respectively. This paper examines the application of both these distillation approaches to a sequence prediction task, grammatical error correction (GEC). This is an important application area for language learning tasks as it can yield highly useful feedback to the learner. It is, however, more challenging than the standard tasks investigated for distillation as the prediction of any grammatical correction to a word will be highly dependent on both the input sequence and the generated output history for the word. The performance of both EnD and EnDD are evaluated on both publicly available GEC tasks as well as a spoken language task.
Yassir Fathullah, Mark J. F. Gales, Andrey Malinin
ICASSP2
2021 Efficient Use of End-to-End Data in Spoken Language Processing
abstract
For many challenging tasks there is often limited data to train the systems in an end-to-end fashion, which has become increasingly popular for deep-learning. However, these tasks can normally be split into multiple separate modules, with significant quantities of data associated with each module. Spoken language processing applications fit into this scenario, as they usually start with a speech recognition module, followed by multiple task specific modules to achieve the end goal. This work examines how the best use can be made of limited end-to-end training for sequence-to-sequence tasks. The key to improving the use of the data is to more tightly integrate the modules via embeddings, rather than simply propagating words between modules. In this work speech translation is considered as the spoken language application. When significant quantities of in-domain, end-to-end data is available, cascade approaches operate well. When the in-domain data is limited, how-ever, tighter integration between modules enables better use of the data to be made. One of the challenges with tighter integration is how to ensure embedding consistency between the modules. A novel form of embedding-passing between modules is proposed that shows improved performance over both cascade and standard embedding-passing approaches for limited in-domain data.
Yiting Lu, Yu Wang 0027, Mark J. F. Gales
ICASSP3
2021 Analysing Bias in Spoken Language Assessment Using Concept Activation Vectors
abstract
A significant concern with deep learning based approaches is that they are difficult to interpret, which means detecting bias in network predictions can be challenging. Concept Activation Vectors (CAVs) have been proposed to address this problem. These use representations - perturbations of activation function outputs - of interpretable concepts to analyse how the network is influenced by the concept. This work applies CAVs to assess bias in a spoken language assessment (SLA) system, a regression task. One of the challenges with SLA is the wide range of concepts that can introduce bias in training data, for example L1, age, acoustic conditions, and particular human graders, or the grading instructions. Simply generating large quantities of expert marked data to check for all forms of bias is impractical. This paper uses CAVs applied to the training data to identify concepts that might be of concern, allowing a more targeted dataset to be collected to assess bias. The ability of CAVs to detect bias is assessed on the BULATS speaking test using both a standard system and a system to which bias was artificially introduced. A strong bias identified by CAVs on the training data matches the bias observed in expert marked held-out test data.
Xizi Wei, Mark J. F. Gales, Kate M. Knill
ICASSP2
2021 Uncertainty Estimation in Autoregressive Structured Prediction
Andrey Malinin, Mark J. F. Gales
ICLR2
2021 Deliberation-Based Multi-Pass Speech Synthesis
Qingyun Dou, Xixin Wu, Moquan Wan, Yiting Lu, Mark J. F. Gales
Interspeech5
2021 Scaling Ensemble Distribution Distillation to Many Classes with Proxy Targets
abstract
Ensembles of machine learning models yield improved system performance as well as robust and interpretable uncertainty estimates; however, their inference costs can be prohibitively high. Ensemble Distribution Distillation (EnD$^2$) is an approach that allows a single model to efficiently capture both the predictive performance and uncertainty estimates of an ensemble. For classification, this is achieved by training a Dirichlet distribution over the ensemble members' output distributions via the maximum likelihood criterion. Although theoretically principled, this work shows that the criterion exhibits poor convergence when applied to large-scale tasks where the number of classes is very high. Specifically, we show that for the Dirichlet log-likelihood criterion classes with low probability induce larger gradients than high-probability classes. Hence during training the model focuses on the distribution of the ensemble tail-class probabilities rather than the probability of the correct and closely related classes. We propose a new training objective which minimizes the reverse KL-divergence to a \emph{Proxy-Dirichlet} target derived from the ensemble. This loss resolves the gradient issues of EnD$^2$, as we demonstrate both theoretically and empirically on the ImageNet, LibriSpeech, and WMT17 En-De datasets containing 1000, 5000, and 40,000 classes, respectively.
Max Ryabinin, Andrey Malinin, Mark J. F. Gales
NeurIPS3
2020 Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks
abstract
Recently, there has been growth in providers of speech transcription services enabling others to leverage technology they would not normally be able to use. As a result, speech-enabled solutions have become commonplace. Their success critically relies on the quality, accuracy, and reliability of the underlying speech transcription systems. Those black box systems, however, offer limited means for quality control as only word sequences are typically available. This paper examines this limited resource scenario for confidence estimation, a measure commonly used to assess transcription reliability. In particular, it explores what other sources of word and sub-word level information available in the transcription process could be used to improve confidence scores. To encode all such information this paper extends lattice recurrent neural networks to handle sub-words. Experimental results using the IARPA OpenKWS 2016 evaluation system show that the use of additional information yields significant gains in confidence estimation accuracy. The implementation for this model can be found online1.
Alexandros Kastanos, Anton Ragni, Mark J. F. Gales
ICASSP3
2020 Ensemble Distribution Distillation
Andrey Malinin, Bruno Mlodozeniec, Mark J. F. Gales
ICLR3
2020 Attention Forcing for Speech Synthesis
abstract
Auto-regressive sequence-to-sequence models with attention mechanisms have achieved state-of-the-art performance in various tasks including speech synthesis. Training these models can be difficult. The standard approach guides a model with the reference output history during training. However during synthesis the generated output history must be used. This mismatch can impact performance. Several approaches have been proposed to handle this, normally by selectively using the generated output history. To make training stable, these approaches often require a heuristic schedule or an auxiliary classifier. This paper introduces attention forcing, which guides the model with the generated output history and reference attention. This approach reduces the training-evaluation mismatch without the need for a schedule or a classifier. Additionally, for standard training approaches, the frame rate is often reduced to prevent models from copying the output history. As attention forcing does not feed the reference output history to the model, it allows using a higher frame rate, which improves the speech quality. Finally, attention forcing allows the model to generate output sequences aligned with the references, which is important for some down-stream tasks such as training neural vocoders. Experiments show that attention forcing allows doubling the frame rate, and yields significant gain in speech quality.
Qingyun Dou, Joshua Efiong, Mark J. F. Gales
INTERSPEECH3
2020 Non-Native Children's Automatic Speech Recognition: The INTERSPEECH 2020 Shared Task ALTA Systems
abstract
Automatic spoken language assessment (SLA) is a challenging problem due to the large variations in learner speech combined with limited resources. These issues are even more problematic when considering children learning a language, with higher levels of acoustic and lexical variability, and of code-switching compared to adult data. This paper describes the ALTA system for the INTERSPEECH 2020 Shared Task on Automatic Speech Recognition for Non-Native Children’s Speech. The data for this task consists of examination recordings of Italian school children aged 9-16, ranging in ability from minimal, to basic, to limited but effective command of spoken English. A variety of systems were developed using the limited training data available, 49 hours. State-of-the-art acoustic models and language models were evaluated, including a diversity of lexical representations, handling code-switching and learner pronunciation errors, and grade specific models. The best single system achieved a word error rate (WER) of 16.9% on the evaluation data. By combining multiple diverse systems, including both grade independent and grade specific models, the error rate was reduced to 15.7%. This combined system was the best performing submission for both the closed and open tasks.
Kate M. Knill, Yu Wang 0027, Xixin Wu, Mark J. F. Gales
INTERSPEECH5
2020 Automatic Detection of Accent and Lexical Pronunciation Errors in Spontaneous Non-Native English Speech
abstract
Detecting individual pronunciation errors and diagnosing pronunciation error tendencies in a language learner based on their speech are important components of computer-aided language learning (CALL). The tasks of error detection and error tendency diagnosis become particularly challenging when the speech in question is spontaneous and particularly given the challenges posed by the inconsistency of human annotation of pronunciation errors. This paper presents an approach to these tasks by distinguishing between lexical errors, wherein the speaker does not know how a particular word is pronounced, and accent errors, wherein the candidate's speech exhibits consistent patterns of phone substitution, deletion and insertion. Three annotated corpora of non-native English speech by speakers of multiple L1s are analysed, the consistency of human annotation investigated and a method presented for detecting individual accent and lexical errors and diagnosing accent error tendencies at the speaker level.
Konstantinos Kyriakopoulos, Kate M. Knill, Mark J. F. Gales
INTERSPEECH3
2020 Spoken Language 'Grammatical Error Correction'
abstract
Spoken language ‘grammatical error correction’ (GEC) is an important mechanism to help learners of a foreign language, here English, improve their spoken grammar. GEC is challeng- ing for non-native spoken language due to interruptions from disfluent speech events such as repetitions and false starts and issues in strictly defining what is acceptable in spoken language. Furthermore there is little labelled data to train models. One way to mitigate the impact of speech events is to use a disflu- ency detection (DD) model. Removing the detected disfluencies converts the speech transcript to be closer to written language, which has significantly more labelled training data. This paper considers two types of approaches to leveraging DD models to boost spoken GEC performance. One is sequential, a separately trained DD model acts as a pre-processing module providing a more structured input to the GEC model. The second approach is to train DD and GEC models in an end-to-end fashion, simul- taneously optimising both modules. Embeddings enable end- to-end models to have a richer information flow. Experimen- tal results show that DD effectively regulates GEC input; end- to-end training works well when fine-tuned on limited labelled in-domain data; and improving DD by incorporating acoustic information helps improve spoken GEC.
Yiting Lu, Mark J. F. Gales, Yu Wang 0027
INTERSPEECH2
2020 Abstractive Spoken Document Summarization Using Hierarchical Model with Multi-Stage Attention Diversity Optimization
abstract
Abstractive summarization is a standard task for written documents, such as news articles. Applying summarization schemes to spoken documents is more challenging, especially in situations involving human interactions, such as meetings. Here, utterances tend not to form complete sentences and sometimes contain little information. Moreover, speech disfluencies will be present as well as recognition errors for automated systems. For current attention-based sequence-to-sequence summarization systems, these additional challenges can yield a poor attention distribution over the spoken document words and utterances, impacting performance. In this work, we propose a multi-stage method based on a hierarchical encoder-decoder model to explicitly model utterance-level attention distribution at training time; and enforce diversity at inference time using a unigram diversity term. Furthermore, multitask learning tasks including dialogue act classification and extractive summarization are incorporated. The performance of the system is evaluated on the AMI meeting corpus. The inclusion of both training and inference diversity terms improves performance, outperforming current state-of-the-art systems in terms of ROUGE scores. Additionally, the impact of ASR errors, as well as performance on the multitask learning tasks, is evaluated.
P. P. Manakul, Mark J. F. Gales
INTERSPEECH2
2020 Universal Adversarial Attacks on Spoken Language Assessment Systems
abstract
There is an increasing demand for automated spoken language assessment (SLA) systems, partly driven by the performance improvements that have come from deep learning based approaches. One aspect of deep learning systems is that they do not require expert derived features, operating directly on the original signal such as a speech recognition (ASR) transcript. This, however, increases their potential susceptibility to adversarial attacks as a form of candidate malpractice. In this paper the sensitivity of SLA systems to a universal black-box attack on the ASR text output is explored. The aim is to obtain a single, universal phrase to maximally increase any candidate's score. Four approaches to detect such adversarial attacks are also described. All the systems, and associated detection approaches, are evaluated on a free (spontaneous) speaking section from a Business English test. It is shown that on deep learning based SLA systems the average candidate score can be increased by almost one grade level using a single six word phrase appended to the end of the response hypothesis. Although these large gains can be obtained, they can be easily detected based on detection shifts from the scores of a “traditional” Gaussian Process based grader.
Vyas Raina, Mark J. F. Gales, Kate M. Knill
INTERSPEECH2
2020 Ensemble Approaches for Uncertainty in Spoken Language Assessment
abstract
Deep learning has dramatically improved the performance of automated systems on a range of tasks including spoken language assessment. One of the issues with these deep learning approaches is that they tend to be overconfident in the decisions that they make, with potentially serious implications for deployment of systems for high-stakes examinations. This paper examines the use of ensemble approaches to improve both the reliability of the scores that are generated, and the ability to detect where the system has made predictions beyond acceptable errors. In this work assessment is treated as a regression problem. Deep density networks, and ensembles of these models, are used as the predictive models. Given an ensemble of models measures of uncertainty, for example the variance of the predicted distributions, can be obtained and used for detecting outlier predictions. However, these ensemble approaches increase the computational and memory requirements of the system. To address this problem the ensemble is distilled into a single mixture density network. The performance of the systems is evaluated on a free speaking prompt-response style spoken language assessment test. Experiments show that the ensembles and the distilled model yield performance gains over a single model, and have the ability to detect outliers.
Xixin Wu, Kate M. Knill, Mark J. F. Gales, Andrey Malinin
INTERSPEECH3
2019 Learning Between Different Teacher and Student Models in ASR
abstract
Teacher-student learning can be applied in automatic speech recognition for model compression and domain adaptation. This trains a student model to emulate the behaviour of a teacher model, and only the student is used to perform recognition. Depending on the application, the teacher and student may differ in their model types, complexities, input contexts, and input features. In previous works, it is often shown that learning from a strong teacher allows the student to perform better than an equivalent model trained with only the reference transcriptions. However, there has not been much investigation into whether a particular form of teacher is appropriate for the student to learn from. This paper aims to study how effectively the student is able to learn from the teacher, when differences exist between their designs. The Augmented Multi-party Interaction (AMI) meeting transcription and Multi-Genre Broadcast (MGB-3) television broadcast audio tasks are used in this analysis. Experimental results suggest that a student can effectively learn from a more complex teacher, but may struggle when it lacks input information. It is therefore important to carefully consider the design of the student for each application.
Jeremy H. M. Wong, Mark J. F. Gales, Yu Wang 0027
ASRU2
2019 Automatic Grammatical Error Detection of Non-native Spoken Learner English
abstract
Automatic language assessment and learning systems are required to support the global growth in English language learning. They need to be able to provide reliable and meaningful feedback to help learners develop their skills. This paper considers the question of detecting "grammatical" errors in non-native spoken English as a first step to providing feedback on a learner's use of the language. A state-of-the-art deep learning based grammatical error detection (GED) system designed for written texts is investigated on free speaking tasks across the full range of proficiency grades with a mix of first languages (L1s). This presents a number of challenges. Free speech contains disfluencies that disrupt the spoken language flow but are not grammatical errors. The lower the level of the learner the more these both will occur which makes the underlying task of automatic transcription harder. The baseline written GED system is seen to perform less well on manually transcribed spoken language. When the GED model is fine-tuned to free speech data from the target domain the spoken system is able to match the written performance. Given the current state-of-the-art in ASR, however, and the ability to detect disfluencies grammatical error feedback from automated transcriptions remains a challenge.
Kate M. Knill, Mark J. F. Gales, P. P. Manakul, Andrew Caines
ICASSP2
2019 Bi-directional Lattice Recurrent Neural Networks for Confidence Estimation
abstract
The standard approach to mitigate errors made by an automatic speech recognition system is to use confidence scores associated with each predicted word. In the simplest case, these scores are word posterior probabilities whilst more complex schemes utilise bi-directional recurrent neural network (BiRNN) models. A number of upstream and downstream applications, however, rely on confidence scores assigned not only to 1-best hypotheses but to all words found in confusion networks or lattices. These include but are not limited to speaker adaptation, semi-supervised training and information retrieval. Although word posteriors could be used in those applications as confidence scores, they are known to have reliability issues. To make improved confidence scores more generally available, this paper shows how BiRNNs can be extended from 1-best sequences to confusion network and lattice structures. Experiments are conducted using one of the Cambridge University submissions to the IARPA OpenKWS 2016 competition. The results show that confusion network and lattice-based BiRNNs can provide a significant improvement in confidence estimation.
Qiujia Li, Preben Ness, Anton Ragni, Mark J. F. Gales
ICASSP4
2019 A Deep Learning Approach to Automatic Characterisation of Rhythm in Non-Native English Speech
abstract
A speaker's rhythm contributes to the intelligibility of their speech and can be characteristic of their language and accent. For non-native learners of a language, the extent to which they match its natural rhythm is an important predictor of their proficiency. As a learner improves, their rhythm is expected to become less similar to their L1 and more to the L2. Metrics based on the variability of the durations of vocalic and consonantal intervals have been shown to be effective at detecting language and accent. In this paper, pairwise variability (PVI, CCI) and variance (varcoV, varcoC) metrics are first used to predict proficiency and L1 of non-native speakers taking an English spoken exam. A deep learning alternative to generalise these features is then presented, in the form of a tunable duration embedding, based on attention over an RNN over durations. The RNN allows relationships beyond pairwise to be captured, while attention allows sensitivity to the different relative importance of durations. The system is trained end-to-end for proficiency and L1 prediction and compared to the baseline. The values of both sets of features for different proficiency levels are then visualised and compared to native speech in the L1 and the L2.
Konstantinos Kyriakopoulos, Kate M. Knill, Mark J. F. Gales
INTERSPEECH3
2019 Impact of ASR Performance on Spoken Grammatical Error Detection
abstract
Computer assisted language learning (CALL) systems aidlearners to monitor their progress by providing scoring andfeedback on language assessment tasks. Free speaking tests al-low assessment of what a learner has said, as well as how theysaid it. For these tasks, Automatic Speech Recognition (ASR)is required to generate transcriptions of a candidate’s responses,the quality of these transcriptions is crucial to provide reliablefeedback in downstream processes. This paper considers theimpact of ASR performance on Grammatical Error Detection(GED) for free speaking tasks, as an example of providing feed-back on a learner’s use of English. The performance of an ad-vanced deep-learning based GED system, initially trained onwritten corpora, is used to evaluate the influence of ASR errors.One consequence of these errors is that grammatical errors canresult from incorrect transcriptions as well as learner errors, thismay yield confusing feedback. To mitigate the effect of theseerrors, and reduce erroneous feedback, ASR confidence scoresare incorporated into the GED system. By additionally adaptingthe written text GED system to the speech domain, using ASRtranscriptions, significant gains in performance can be achieved.Analysis of the GED performance for different grammatical er-ror types and across grade is also presented.
Yiting Lu, Mark J. F. Gales, Kate M. Knill, P. P. Manakul, Yu Wang 0027
INTERSPEECH2
2019 Reverse KL-Divergence Training of Prior Networks: Improved Uncertainty and Adversarial Robustness
abstract
Ensemble approaches for uncertainty estimation have recently been applied to the tasks of misclassification detection, out-of-distribution input detection and adversarial attack detection. Prior Networks have been proposed as an approach to efficiently emulate an ensemble of models for classification by parameterising a Dirichlet prior distribution over output distributions. These models have been shown to outperform alternative ensemble approaches, such as Monte-Carlo Dropout, on the task of out-of-distribution input detection. However, scaling Prior Networks to complex datasets with many classes is difficult using the training criteria originally proposed. This paper makes two contributions. First, we show that the appropriate training criterion for Prior Networks is the reverse KL-divergence between Dirichlet distributions. This addresses issues in the nature of the training data target distributions, enabling prior networks to be successfully trained on classification tasks with arbitrarily many classes, as well as improving out-of-distribution detection performance. Second, taking advantage of this new training criterion, this paper investigates using Prior Networks to detect adversarial attacks and proposes a generalized form of adversarial training. It is shown that the construction of successful adaptive whitebox attacks, which affect the prediction and evade detection, against Prior Networks trained on CIFAR-10 and CIFAR-100 using the proposed approach requires a greater amount of computational effort than against networks defended using standard adversarial training or MC-dropout.
Andrey Malinin, Mark J. F. Gales
NeurIPS2
2019 Exploiting Future Word Contexts in Neural Network Language Models for Speech Recognition
abstract
Language modeling is a crucial component in a wide range of applications including speech recognition. Language models (LMs) are usually constructed by splitting a sentence into words and computing the probability of a word based on its word history. This sentence probability calculation, making use of conditional probability distributions, assumes that there is little impact from approximations used in the LMs, including the word history representations and finite training data. This motivates examining models that make use of additional information from the sentence. In this paper, future word information, in addition to the history, is used to predict the probability of the current word. For recurrent neural network LMs (RNNLMs), this information can be encapsulated in a bi-directional model. However, if used directly, this form of model is computationally expensive when trained on large quantities of data, and can be problematic when used with word lattices. This paper proposes a novel neural network language model structure, the succeeding-word RNNLM, su-RNNLM, to address these issues. Instead of using a recurrent unit to capture the complete future word contexts, a feedforward unit is used to model a fixed finite number of succeeding words. This is more efficient in training than bi-directional models and can be applied to lattice rescoring. The generated lattices can be used for downstream applications, such as confusion network decoding and keyword search. Experimental results on speech recognition and keyword spotting tasks illustrate the empirical usefulness of future word information, and the flexibility of the proposed model to represent this information.
Xie Chen 0001, Xunying Liu, Yu Wang 0027, Anton Ragni, Jeremy H. M. Wong, Mark J. F. Gales
IEEE ACM Trans. Audio Speech Lang. Process.6
2019 General Sequence Teacher-Student Learning
abstract
In automatic speech recognition, performance gains can often be obtained by combining an ensemble of multiple models. However, this can be computationally expensive when performing recognition. Teacher-student learning alleviates this cost by training a single student model to emulate the combined ensemble behaviour. Only this student needs to be used for recognition. Previously investigated teacher-student criteria often limit the forms of diversity allowed in the ensemble, and only propagate information from the teachers to the student at the frame level. This paper addresses both of these issues by examining teacher-student learning within a sequence-level framework, and assessing the flexibility that these approaches offer. Various sequence-level teacher-student criteria are examined in this work, to propagate sequence posterior information. A training criterion based on the Kullback-Leibler (KL)-divergence between context-dependent state sequence posteriors is proposed that allows for a diversity of state cluster sets to be present in the ensemble. This criterion is shown to be an upper bound to a more general KL-divergence between word sequence posteriors, which places even fewer restrictions on the ensemble diversity, but whose gradient can be expensive to compute. These methods are evaluated on the augmented multi-party interaction (AMI) meeting transcription and MGB-3 television broadcast audio tasks.
Jeremy H. M. Wong, Mark J. F. Gales, Yu Wang 0027
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription
abstract
State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to the spoken form is complicated. However, in recent years the representational power of deep-learning based acoustic models has improved, raising interest in graphemic acoustic models for English, due to the simplicity of generating the lexicon. In this paper, phonetic and graphemic models are compared for an English Multi-Genre Broadcast transcription task. A range of acoustic models based on lattice-free MMI training are constructed using phonetic and graphemic lexicons. For this task, it is found that having a long-span temporal history reduces the difference in performance between the two forms of models. In addition, system combination is examined, using parameter smoothing and hypothesis combination. As the combination approaches become more complicated the difference between the phonetic and graphemic systems further decreases. Finally, for all configurations examined the combination of phonetic and graphemic systems yields consistent gains.
Yu Wang 0027, Xie Chen 0001, Mark J. F. Gales, Anton Ragni, Jeremy H. M. Wong
ICASSP3
2018 Active Memory Networks for Language Modeling
abstract
Making predictions of the following word given the back history of words may be challenging without meta-information such as the topic. Standard neural network language models have an implicit representation of the topic via the back history of words. In this work a more explicit form of topic representation is used via an attention mechanism. Though this makes use of the same information as the standard model, it allows parameters of the network to focus on different aspects of the task. The attention model provides a form of topic representation that is automatically learned from the data. Whereas the recurrent model deals with the (conditional) history representation. The combined model is expected to reduce the stress on the standard model to handle multiple aspects. Experiments were conducted on the Penn Tree Bank and BBC Multi-Genre Broadcast News (MGB) corpora, where the proposed approach outperforms standard forms of recurrent models in perplexity. Finally, N-best list rescoring for speech recognition in the MGB3 task shows word error rate improvements over comparable standard form of recurrent models.
Oscar Chen, Anton Ragni, Mark J. F. Gales, Xie Chen 0001
INTERSPEECH3
2018 Impact of ASR Performance on Free Speaking Language Assessment
abstract
In free speaking tests candidates respond in spontaneous speech to prompts. This form of test allows the spoken language proficiency of a non-native speaker of English to be assessed more fully than read aloud tests. As the candidate's responses are unscripted, transcription by automatic speech recognition (ASR) is essential for automated assessment. ASR will never be 100% accurate so any assessment system must seek to minimise and mitigate ASR errors. This paper considers the impact of ASR errors on the performance of free speaking test auto-marking systems. Firstly rich linguistically related features, based on part-of-speech tags from statistical parse trees, are investigated for assessment. Then, the impact of ASR errors on how well the system can detect whether a learner's answer is relevant to the question asked is evaluated. Finally, the impact that these errors may have on the ability of the system to provide detailed feedback to the learner is analysed. In particular, pronunciation and grammatical errors are considered as these are important in helping a learner to make progress. As feedback resulting from an ASR error would be highly confusing, an approach to mitigate this problem using confidence scores is also analysed.
Kate M. Knill, Mark J. F. Gales, Konstantinos Kyriakopoulos, Andrey Malinin, Anton Ragni, Yu Wang 0027, Andrew Caines
INTERSPEECH2
2018 A Deep Learning Approach to Assessing Non-native Pronunciation of English Using Phone Distances
abstract
The way a non-native speaker pronounces the phones of a language is an important predictor of their proficiency. In grading spontaneous speech, the pairwise distances between generative statistical models trained on each phone have been shown to be powerful features. This paper presents a deep learning alternative to model-based phone distances in the form of a tunable Siamese network feature extractor to extract distance metrics directly from the audio frame sequence. Features are extracted at the phone instance level and combined to phone-level representations using an attention mechanism. Pair-wise distances between phone features are then projected through a feed-forward layer to predict score. The extraction stage is initialised on either a binary phone instance-pair classification task, or to mimic the model-based features, then the whole system is fine-tuned end-to-end, optimising the learning of the distance metric to the score prediction task. This method is therefore more adaptable and more sensitive to phone instance level phenomena. Its performance is compared against
Konstantinos Kyriakopoulos, Kate M. Knill, Mark J. F. Gales
INTERSPEECH3
2018 Automatic Speech Recognition System Development in the "Wild"
abstract
The standard framework for developing an automatic speech recognition (ASR) system is to generate training and development data for building the system, and evaluation data for the final performance analysis. All the data is assumed to come from the domain of interest. Though this framework is matched to some tasks, it is more challenging for systems that are required to operate over broad domains, or where the ability to collect the required data is limited. This paper discusses ASR work performed under the IARPA MATERIAL program, which is aimed at cross-language information retrieval, and examines this challenging scenario. In terms of available data, only limited narrow-band conversational telephone speech data was provided. However, the system is required to operate over a range of domains, including broadcast data. As no data is available for the broadcast domain, this paper proposes an approach for system development based on scraping "related" data from the web, and using ASR system confidence scores as the primary metric for developing the acoustic and language model components. As an initial evaluation of the approach, the Swahili development language is used, with the final system performance assessed on the IARPA MATERIAL Analysis Pack 1 data.
Anton Ragni, Mark J. F. Gales
INTERSPEECH2
2018 Waveform-Based Speaker Representations for Speech Synthesis
abstract
Speaker adaptation is a key aspect of building a range of speech processing systems, for example personalised speech synthesis. For deep-learning based approaches, the model parameters are hard to interpret, making speaker adaptation more challenging. One widely used method to address this problem is to extract a fixed length vector as speaker representation, and use this as an additional input to the task-specific model. This allows speaker-specific output to be generated, without modifying the model parameters. However, the speaker representation is often extracted in a task-independent fashion. This allows the same approach to be used for a range of tasks, but the extracted representation is unlikely to be optimal for the specific task of interest. Furthermore, the features from which the speaker representation is extracted are usually pre-defined, often a standard speech representation. This may limit the available information that can be used. In this paper, an integrated optimisation framework for building a task specific speaker representation, making use of all the available information, is proposed. Speech synthesis is used as the example task. The speaker representation is derived from raw waveform, incorporating text information via an attention mechanism. This paper evaluates and compares this framework with standard task-independent forms.
Moquan Wan, Gilles Degottex, Mark J. F. Gales
INTERSPEECH3
2018 Speaker Adaptation and Adaptive Training for Jointly Optimised Tandem Systems
abstract
Speaker independent (SI) Tandem systems trained by joint optimisation of bottleneck (BN) deep neural networks (DNNs) and Gaussian mixture models (GMMs) have been found to produce similar word error rates (WERs) to Hybrid DNN systems. A key advantage of using GMMs is that existing speaker adaptation methods, such as maximum likelihood linear regression (MLLR), can be used which to account for diverse speaker variations and improve system robustness. This paper investigates speaker adaptation and adaptive training (SAT) schemes for jointly optimised Tandem systems. Adaptation techniques investigated include constrained MLLR (CMLLR) transforms based on BN features for SAT as well as MLLR and parameterised sigmoid functions for unsupervised test-time adaptation. Experiments using English multi-genre broadcast (MGB3) data show that CMLLR SAT yields a 4% relative WER reduction over jointly trained Tandem and Hybrid SI systems, and further reductions in WER are obtained by system combination.
Yu Wang 0027, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2018 Predictive Uncertainty Estimation via Prior Networks
abstract
Estimating how uncertain an AI system is in its predictions is important to improve the safety of such systems. Uncertainty in predictive can result from uncertainty in model parameters, irreducible \emph{data uncertainty} and uncertainty due to distributional mismatch between the test and training data distributions. Different actions might be taken depending on the source of the uncertainty so it is important to be able to distinguish between them. Recently, baseline tasks and metrics have been defined and several practical methods to estimate uncertainty developed. These methods, however, attempt to model uncertainty due to distributional mismatch either implicitly through \emph{model uncertainty} or as \emph{data uncertainty}. This work proposes a new framework for modeling predictive uncertainty called Prior Networks (PNs) which explicitly models \emph{distributional uncertainty}. PNs do this by parameterizing a prior distribution over predictive distributions. This work focuses on uncertainty for classification and evaluates PNs on the tasks of identifying out-of-distribution (OOD) samples and detecting misclassification on the MNIST and CIFAR-10 datasets, where they are found to outperform previous methods. Experiments on synthetic and MNIST and CIFAR-10 data show that unlike previous non-Bayesian methods PNs are able to distinguish between data and distributional uncertainty.
Andrey Malinin, Mark J. F. Gales
NeurIPS2
2018 A Spectrally Weighted Mixture of Least Square Error and Wasserstein Discriminator Loss for Generative SPSS
abstract
Generative networks can create an artificial spectrum based on its conditional distribution estimate instead of predicting only the mean value, as the Least Square (LS) solution does. This is promising since the LS predictor is known to oversmooth features leading to muffling effects. However, modeling a whole distribution instead of a single mean value requires more data and thus also more computational resources. With only one hour of recording, as often used with LS approaches, the resulting spectrum is noisy and sounds full of artifacts. In this paper, we suggest a new loss function, by mixing the LS error and the loss of a discriminator trained with Wasserstein GAN, while weighting this mixture differently through the frequency domain. Using listening tests, we show that, using this mixed loss, the generated spectrum is smooth enough to obtain a decent perceived quality. While making our source code available online, we also hope to make generative networks more accessible with lower the necessary resources.
Gilles Degottex, Mark J. F. Gales
SLT2
2018 Hierarchical RNNs for Waveform-Level Speech Synthesis
abstract
Speech synthesis technology has a wide range of applications such as voice assistants. In recent years waveform-level synthesis systems have achieved state-of-the-art performance, as they overcome the limitations of vocoder-based synthesis systems. A range of waveform-level synthesis systems have been proposed; this paper investigates the performance of hierarchical Recurrent Neural Networks (RNNs) for speech synthesis. First, the form of network conditioning is discussed, comparing linguistic features and vocoder features from a vocoder-based synthesis system. It is found that compared with linguistic features, conditioning on vocoder features requires less data and modeling power, and yields better performance when there is limited data. By conditioning the hierarchical RNN on vocoder features, this paper develops a neural vocoder, which is capable of high quality synthesis when there is sufficient data. Furthermore, this neural vocoder is flexible, as conceptually it can map any sequence of vocoder features to speech, enabling efficient synthesizer porting to a target speaker. Subjective listening tests demonstrate that the neural vocoder outperforms a high quality baseline, and that it can change its voice to a very different speaker, given less than 15 minutes of data for fine tuning.
Qingyun Dou, Moquan Wan, Gilles Degottex, Zhiyi Ma, Mark J. F. Gales
SLT5
2018 Confidence Estimation and Deletion Prediction Using Bidirectional Recurrent Neural Networks
abstract
The standard approach to assess reliability of automatic speech transcriptions is through the use of confidence scores. If accurate, these scores provide a flexible mechanism to flag transcription errors for upstream and downstream applications. One challenging type of errors that recognisers make are deletions. These errors are not accounted for by the standard confidence estimation schemes and are hard to rectify in the upstream and downstream processing. High deletion rates are prominent in limited resource and highly mismatched training/testing conditions studied under IARPA Babel and Material programs. This paper looks at the use of bidirectional recurrent neural networks to yield confidence estimates in predicted as well as deleted words. Several simple schemes are examined for combination. To assess usefulness of this approach, the combined confidence score is examined for untranscribed data selection that favours transcriptions with lower deletion errors. Experiments are conducted using IARPA Babel/Material program languages.
Anton Ragni, Qiujia Li, Mark J. F. Gales, Yongqiang Wang 0006
SLT3
2018 Improved Auto-Marking Confidence for Spoken Language Assessment
abstract
Automatic assessment of spoken language proficiency is a sought-after technology. These systems often need to handle the operating scenario where candidates have a skill level or first language which was not encountered during the training stage. For high stakes tests it is necessary for those systems to have good grading performance when the candidate is from the same population as those contained in the training set, and they should know when they are likely to perform badly in the case when the candidate is not from the same population as the ones contained in training set. This paper focuses on using Deep Density Networks to yield auto-marking confidence. Firstly, we explore the benefits of parametrising either a predictive distribution or a posterior distribution over the parameters of the model likelihood and obtaining the predictive distribution via marginalisation. Secondly, we investigate how it is possible to act on the parametrised density in order to explicitly teach the model to have low confidence in areas of the observation space where there is no training data by assigning confidence scores to artificially generated data. Lastly, we compare the capabilities of Factor Analysis, Variational Auto-Encodes, and Wasserstein Generative Adversarial Networks to generate artificial data.
Marco Del Vecchio, Andrey Malinin, Mark J. F. Gales
SLT3
2018 Sequence Teacher-Student Training of Acoustic Models for Automatic Free Speaking Language Assessment
abstract
A high performance automatic speech recognition (ASR) system is an important constituent component of an automatic language assessment system for free speaking language tests. The ASR system is required to be capable of recognising non-native spontaneous English speech and to be deployable under real-time conditions. The performance of ASR systems can often be significantly improved by leveraging upon multiple systems that are complementary, such as an ensemble. Ensemble methods, however, can be computationally expensive, often requiring multiple decoding runs, which makes them impractical for deployment. In this paper, a lattice-free implementation of sequence-level teacher-student training is used to reduce this computational cost, thereby allowing for real-time applications. This method allows a single student model to emulate the performance of an ensemble of teachers, but without the need for multiple decoding runs. Adaptations of the student model to speakers from different first languages (L1s) and grades are also explored.
Yu Wang 0027, Jeremy H. M. Wong, Mark J. F. Gales, Kate M. Knill, Anton Ragni
SLT3
2018 Towards automatic assessment of spontaneous spoken English
Yu Wang 0027, Mark J. F. Gales, Kate M. Knill, Konstantinos Kyriakopoulos, Andrey Malinin, Rogier C. van Dalen, M. Rashid
Speech Commun.2
2018 A Log Domain Pulse Model for Parametric Speech Synthesis
abstract
Most of the degradation in current Statistical Parametric Speech Synthesis (SPSS) results from the form of the vocoder. One of the main causes of degradation is the reconstruction of the noise. In this article, a new signal model is proposed that leads to a simple synthesizer, without the need for ad-hoc tuning of model parameters. The model is not based on the traditional additive linear source-filter model, it adopts a combination of speech components that are additive in the log domain. Also, the same representation for voiced and unvoiced segments is used, rather than relying on binary voicing decisions. This avoids voicing error discontinuities that can occur in many current vocoders. A simple binary mask is used to denote the presence of noise in the time-frequency domain, which is less sensitive to classification errors. Four experiments have been carried out to evaluate this new model. The first experiment examines the noise reconstruction issue. Three listening tests have also been carried out that demonstrate the advantages of this model: comparison with the STRAIGHT vocoder; the direct prediction of the binary noise mask by using a mixed output configuration; and partial improvements of creakiness using a mask correction mechanism.
Gilles Degottex, Pierre Lanchantin, Mark J. F. Gales
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Improving Interpretability and Regularization in Deep Learning
abstract
Deep learning approaches yield state-of-the-art performance in a range of tasks, including automatic speech recognition. However, the highly distributed representation in a deep neural network (DNN) or other network variations is difficult to analyze, making further parameter interpretation and regularization challenging. This paper presents a regularization scheme acting on the activation function output to improve the network interpretability and regularization. The proposed approach, referred to as activation regularization, encourages activation function outputs to satisfy a target pattern. By defining appropriate target patterns, different learning concepts can be imposed on the network. This method can aid network interpretability and also has the potential to reduce overfitting. The scheme is evaluated on several continuous speech recognition tasks: the Wall Street Journal continuous speech recognition task, eight conversational telephone speech tasks from the IARPA Babel program and a U.S. English broadcast news task. On all the tasks, the activation regularization achieved consistent performance gains over the standard DNN baselines.
Chunyang Wu, Mark J. F. Gales, Anton Ragni, Panagiota Karanasou, Khe Chai Sim
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Future word contexts in neural network language models
abstract
Recently, bidirectional recurrent network language models (bi-RNNLMs) have been shown to outperform standard, unidirectional, recurrent neural network language models (uni-RNNLMs) on a range of speech recognition tasks. This indicates that future word context information beyond the word history can be useful. However, bi-RNNLMs pose a number of challenges as they make use of the complete previous and future word context information. This impacts both training efficiency and their use within a lattice rescoring framework. In this paper these issues are addressed by proposing a novel neural network structure, succeeding word RNNLMs (suRNNLMs). Instead of using a recurrent unit to capture the complete future word contexts, a feedforward unit is used to model a finite number of succeeding, future, words. This model can be trained much more efficiently than bi-RNNLMs and can also be used for lattice rescoring. Experimental results on a meeting transcription task (AMI) show the proposed model consistently outperformed uni-RNNLMs and yield only a slight degradation compared to bi-RNNLMs in N-best rescoring. Additionally, performance improvements can be obtained using lattice rescoring and subsequent confusion network decoding.
Xie Chen 0001, Anton Ragni, Mark J. F. Gales
ASRU5
2017 A hierarchical attention based model for off-topic spontaneous spoken response detection
abstract
Automatic spoken language assessment and training systems are becoming increasingly popular to handle the growing demand to learn languages. However, current systems often assess only fluency and pronunciation, with limited content-based features being used. This paper examines one particular aspect of content-assessment, off-topic response detection. This is important for deployed systems as it ensures that candidates understood the prompt, and are able to generate an appropriate answer. Previously proposed approaches typically require a set of prompt-response training pairs, which limits flexibility as example responses are required whenever a new test prompt is introduced. Recently, the attention based neural topic model (ATM) was presented, which can assess the relevance of prompt-response pairs regardless of whether the prompt was seen in training. This model uses a bidirectional Recurrent Neural Network (BiRNN) embedding of the prompt combined with an attention mechanism to attend over the hidden states of a BiRNN embedding of the response to compute a fixed-length embedding used to predict relevance. Unfortunately, performance on prompts not seen in the training data is lower than on seen prompts. Thus, this paper adds the following contributions: several improvements to the ATM are examined; a hierarchical variant of the ATM (HATM) is proposed, which explicitly uses prompt similarity to further improve performance on unseen prompts by interpolating over prompts seen in training data given a prompt of interest via a second attention mechanism; an in-depth analysis of both models is conducted and main failure mode identified. On spontaneous spoken data, taken from BULATS tests, these systems are able to assess relevance to both seen and unseen prompts.
Andrey Malinin, Kate M. Knill, Mark J. F. Gales
ASRU3
2017 Integrated speaker-adaptive speech synthesis
abstract
Enabling speech synthesis systems to rapidly adapt to sound like a particular speaker is an essential attribute for building personalised systems. For deep-learning based approaches, this is difficult as these networks use a highly distributed representation. It is not simple to interpret the model parameters, which complicates the adaptation process. To address this problem, speaker characteristics can be encapsulated in fixed-length speaker-specific Identity Vectors (iVectors), which are appended to the input of the synthesis network. Altering the iVector changes the nature of the synthesised speech. The challenge is to derive an optimal iVector for each speaker that encodes all the speaker attributes required for the synthesis system. The standard approach involves two separate stages: estimation of the iVectors for the training data; and training the synthesis network. This paper proposes an integrated training scheme for speaker adaptive speech synthesis. For the iVector extraction, an attention based mechanism, which is a function of the context labels, is used to combine the data from the target speaker. This attention mechanism, as well as nature of the features being merged, are optimised at the same time as the synthesis network parameters. This should yield an iVector-like speaker representation that is optimal for use with the synthesis system. The system is evaluated on the Voice Bank corpus. The resulting system automatically provides a sensible attention sequence and shows improved performance from the standard approach.
Moquan Wan, Gilles Degottex, Mark J. F. Gales
ASRU3
2017 Multi-task ensembles with teacher-student training
abstract
Ensemble methods often yield significant gains for automatic speech recognition. One method to obtain a diverse ensemble is to separately train models with a range of context dependent targets, often implemented as state clusters. However, decoding the complete ensemble can be computationally expensive. To reduce this cost, the ensemble can be generated using a multi-task architecture. Here, the hidden layers are merged across all members of the ensemble, leaving only separate output layers for each set of targets. Previous investigations of this form of ensemble have used cross-entropy training, which is shown in this paper to produce only limited diversity between members of the ensemble. This paper extends the multi-task framework in several ways. First, the multi-task ensemble can be trained in a teacher-student fashion toward the ensemble of separate models, with the aim of increasing diversity. Second, the multi-task ensemble can be trained with a sequence discriminative criterion. Finally, a student model, with a single output layer, can be trained to emulate the combined ensemble, to further reduce the computational cost of decoding. These methods are evaluated on the Babel conversational telephone speech, AMI meeting transcription, and HUB4 English broadcast news tasks.
Jeremy H. M. Wong, Mark J. F. Gales
ASRU2
2017 Recurrent neural network language models for keyword search
abstract
Recurrent neural network language models (RNNLMs) have becoming increasingly popular in many applications such as automatic speech recognition (ASR). Significant performance improvements in both perplexity and word error rate over standard n-gram LMs have been widely reported on ASR tasks. In contrast, published research on using RNNLMs for keyword search systems has been relatively limited. In this paper the application of RNNLMs for the IARPA Babel keyword search task is investigated. In order to supplement the limited acoustic transcription data, large amounts of web texts are also used in large vocabulary design and LM training. Various training criteria were then explored to improved RNNLMs' efficiency in both training and evaluation. Significant and consistent improvements on both keyword search and ASR tasks were obtained across all languages.
Xie Chen 0001, Anton Ragni, J. Vasilakes, Xunying Liu, Kate M. Knill, Mark J. F. Gales
ICASSP6
2017 Morph-to-word transduction for accurate and efficient automatic speech recognition and keyword search
abstract
Word units are a popular choice in statistical language modelling. For inflective and agglutinative languages this choice may result in a high out of vocabulary rate. Subword units, such as morphs, provide an interesting alternative to words. These units can be derived in an unsupervised fashion and empirically show lower out of vocabulary rates. This paper proposes a morph-to-word transduction to convert morph sequences into word sequences. This enables powerful word language models to be applied. In addition, it is expected that techniques such as pruning, confusion network decoding, keyword search and many others may benefit from word rather than morph level decision making. However, word or morph systems alone may not achieve optimal performance in tasks such as keyword search so a combination is typically employed. This paper proposes a single index approach that enables word, morph and phone searches to be performed over a single morph index. Experiments are conducted on IARPA Babel program languages including the surprise languages of the OpenKWS 2015 and 2016 competitions.
Anton Ragni, Danielle Saunders, P. Zahemszky, J. Vasilakes, Mark J. F. Gales, Kate M. Knill
ICASSP5
2017 Stimulated training for automatic speech recognition and keyword search in limited resource conditions
abstract
Training neural network acoustic models on limited quantities of data is a challenging task. A number of techniques have been proposed to improve generalisation. This paper investigates one such technique called stimulated training. It enables standard criteria such as cross-entropy to enforce spatial constraints on activations originating from different units. Having different regions being active depending on the input unit may help network to discriminate better and as a consequence yield lower error rates. This paper investigates stimulated training for automatic speech recognition of a number of languages representing different families, alphabets, phone sets and vocabulary sizes. In particular, it looks at ensembles of stimulated networks to ensure that improved generalisation will withstand system combination effects. In order to assess stimulated training beyond 1-best transcription accuracy, this paper looks at keyword search as a proxy for assessing quality of lattices. Experiments are conducted on IARPA Babel program languages including the surprise language of OpenKWS 2016 competition.
Anton Ragni, Chunyang Wu, Mark J. F. Gales, J. Vasilakes, Kate M. Knill
ICASSP3
2017 Investigating Bidirectional Recurrent Neural Network Language Models for Speech Recognition
abstract
Recurrent neural network language models (RNNLMs) are powerful language modeling techniques. Significant performance improvements have been reported in a range of tasks including speech recognition compared to n-gram language models. Conventional n-gram and neural network language models are trained to predict the probability of the next word given its preceding context history. In contrast, bidirectional recurrent neural network based language models consider the context from future words as well. This complicates the inference process, but has theoretical benefits for tasks such as speech recognition as additional context information can be used. However to date, very limited or no gains in speech recognition performance have been reported with this form of model. This paper examines the issues of training bidirectional recurrent neural network language models (bi-RNNLMs) for speech recognition. A bi-RNNLM probability smoothing technique is proposed, that addresses the very sharp posteriors that are often observed in these models. The performance of the bi-RNNLMs is evaluated on three speech recognition tasks: broadcast news; meeting transcription (AMI); and low-resource systems (Babel data). On all tasks gains are observed by applying the smoothing technique to the bi-RNNLM. In addition consistent performance gains can be obtained by combining bi-RNNLMs with n-gram and uni-directional RNNLMs.
Xie Chen 0001, Anton Ragni, Xunying Liu, Mark J. F. Gales
INTERSPEECH4
2017 Use of Graphemic Lexicons for Spoken Language Assessment
abstract
Copyright © 2017 ISCA. Automatic systems for practice and exams are essential to support the growing worldwide demand for learning English as an additional language. Assessment of spontaneous spoken English is, however, currently limited in scope due to the difficulty of achieving sufficient automatic speech recognition (ASR) accuracy. "Off-the-shelf" English ASR systems cannot model the exceptionally wide variety of accents, pronunications and recording conditions found in non-native learner data. Limited training data for different first languages (L1s), across all proficiency levels, often with (at most) crowd-sourced transcriptions, limits the performance of ASR systems trained on non-native English learner speech. This paper investigates whether the effect of one source of error in the system, lexical modelling, can be mitigated by using graphemic lexicons in place of phonetic lexicons based on native speaker pronunications. Graphemicbased English ASR is typically worse than phonetic-based due to the irregularity of English spelling-to-pronunciation but here lower word error rates are consistently observed with the graphemic ASR. The effect of using graphemes on automatic assessment is assessed on different grader feature sets: audio and fluency derived features, including some phonetic level features; and phone/grapheme distance features which capture a measure of pronunciation ability.
Kate M. Knill, Mark J. F. Gales, Konstantinos Kyriakopoulos, Anton Ragni, Yu Wang 0027
INTERSPEECH2
2017 Student-Teacher Training with Diverse Decision Tree Ensembles
abstract
Student-teacher training allows a large teacher model or ensemble of teachers to be compressed into a single student model, for the purpose of efficient decoding. However, current approaches in automatic speech recognition assume that the state clusters, often defined by Phonetic Decision Trees (PDT), are the same across all models. This limits the diversity that can be captured within the ensemble, and also the flexibility when selecting the complexity of the student model output. This paper examines an extension to student-teacher training that allows for the possibility of having different PDTs between teachers, and also for the student to have a different PDT from the teacher. The proposal is to train the student to emulate the logical context dependent state posteriors of the teacher, instead of the frame posteriors. This leads to a method of mapping frame posteriors from one PDT to another. This approach is evaluated on three speech recognition tasks: the Tok Pisin and Javanese low resource conversational telephone speech tasks from the IARPA Babel programme, and the HUB4 English broadcast news task.
Jeremy H. M. Wong, Mark J. F. Gales
INTERSPEECH2
2017 Deep Activation Mixture Model for Speech Recognition
abstract
Deep learning approaches achieve state-of-the-art performance in a range of applications, including speech recognition. However, the parameters of the deep neural network (DNN) are hard to interpret, which makes regularisation and adaptation to speaker or acoustic conditions challenging. This paper proposes the deep activation mixture model (DAMM) to address these problems. The output of one hidden layer is modelled as the sum of a mixture and residual models. The mixture model forms an activation function contour while the residual one models fluctuations around the contour. The use of the mixture model gives two advantages: First, it introduces a novel regularisation on the DNN. Second, it allows novel adaptation schemes. The proposed approach is evaluated on a large-vocabulary U.S. English broadcast news task. It yields a slightly better performance than the DNN baselines, and on the utterance-level unsupervised adaptation, the adapted DAMM acquires further performance gains.
Chunyang Wu, Mark J. F. Gales
INTERSPEECH2
2017 I-Vectors and Structured Neural Networks for Rapid Adaptation of Acoustic Models
abstract
A lot of interest has been risen in the last years on the adaptation of deep neural network (DNN) acoustic models, as the latter become the state-of-art in automatic speech recognition. This work focuses on approaches that allow for rapid and robust adaptation of such models. First, i-vectors are added to the DNN input as speaker-informed features. An informative prior is introduced to i-vector estimation to improve the robustness to limited adaptation data. I-vectors are then combined with a structured adaptive DNN, the multibasis adaptive neural network (MBANN), and the complementarity of these adaptation techniques is investigated. Moreover, i-vectors are used to predict the MBANN transforms, avoiding the initial decoding pass and alignment. These approaches are evaluated on a U.S. English Broadcast News (BN) transcription task with two distinct sets of test data. The first, from the BN task and BN-style Youtube videos, yields test data acoustically matched to the training data, while the second set is from acoustically mismatched Youtube videos of diverse context. The performance gains from these schemes are found to be sensitive to the level of mismatch between training and test sets. The MBANN system combined with i-vector input achieves best performance for BN test sets. The i-vector-based predictive MBANN scheme is proven to be more robust to acoustically mismatched conditions and outperforms the other adaptation schemes in such scenarios.
Panagiota Karanasou, Chunyang Wu, Mark J. F. Gales, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Off-topic Response Detection for Spontaneous Spoken English Assessment
abstract
Automatic spoken language assessment systems are becoming increasingly important to meet the demand for English second language learning.This is a challenging task due to the high error rates of, even state-of-the-art, non-native speech recognition.Consequently current systems primarily assess fluency and pronunciation.However, content assessment is essential for full automation.As a first stage it is important to judge whether the speaker responds on topic to test questions designed to elicit spontaneous speech.Standard approaches to off-topic response detection assess similarity between the response and question based on bag-of-words representations.An alternative framework based on Recurrent Neural Network Language Models (RNNLM) is proposed in this paper.The RNNLM is adapted to the topic of each test question.It learns to associate example responses to questions with points in a topic space constructed using these example responses.Classification is done by ranking the topic-conditional posterior probabilities of a response.The RNNLMs associate a broad range of responses with each topic, incorporate sequence information and scale better with additional training data, unlike standard methods.On experiments conducted on data from the Business Language Testing Service (BULATS) this approach outperforms standard approaches.
Andrey Malinin, Rogier C. van Dalen, Kate M. Knill, Yu Wang 0027, Mark J. F. Gales
ACL (1)5
2016 CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models
abstract
In recent years, recurrent neural network language models (RNNLMs) have become increasingly popular for a range of applications including speech recognition. However, the training of RNNLMs is computationally expensive, which limits the quantity of data, and size of network, that can be used. In order to fully exploit the power of RNNLMs, efficient training implementations are required. This paper introduces an open-source toolkit, the CUED-RNNLM toolkit, which supports efficient GPU-based training of RNNLMs. RNNLM training with a large number of word level output targets is supported, in contrast to existing tools which used class-based output-targets. Support fotN-best and lattice-based rescoring of both HTK and Kaldi format lattices is included. An example of building and evaluating RNNLMs with this toolkit is presented for a Kaldi based speech recognition system using the AMI corpus. All necessary resources including the source code, documentation and recipe are available online1.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP4
2016 Improved DNN-based segmentation for multi-genre broadcast audio
abstract
Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for multi-genre broadcast audio with deep neural network (DNN) based speech/non-speech detection. A further stage of change-point detection and clustering is used to obtain homogeneous segments. Suitable DNN inputs, context window sizes and architectures are studied with a series of experiments using a large corpus of MGB television audio. For MGB transcription, the improved segmenter yields roughly half the increase in word error rate, over manual segmentation, compared to the baseline DNN segmenter supplied for the 2015 ASRU MGB challenge.
Chao Zhang 0031, Philip C. Woodland, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin, Xunying Liu, Yanmin Qian
ICASSP4
2016 Combining i-vector representation and structured neural networks for rapid adaptation
abstract
Rapid adaptation of deep neural networks (DNNs) with limited unsupervised data remains a significant challenge. This paper investigates the combination of two schemes that have been proposed to address this problem: i-vector representations and multi-basis adaptive neural networks (MBANNs). Two approaches for combining these schemes together are described. The first uses i-vectors as one of the input features to the MBANN. The purpose is to combine the speaker representation of the i-vector with the network interpolation of the MBANN scheme. The second approach aims to reduce the computational cost, and improve the robustness to hypothesis errors, of the MBANN scheme. Here i-vectors are used to predict the interpolation weights of the MBANN scheme. This removes the need for an initial decoding pass, and alignment, which was previously used. These approaches are evaluated using acoustic and language models trained on a U.S. English Broadcast News (BN) transcription task. Two distinct sets of test data are examined. The first from the BN task, yields test data acoustically matched to the training data. The second, acoustically mismatched, set is from Youtube videos. The performance gains from these schemes is found to be sensitive to the level of mismatch between training and test.
Chunyang Wu, Panagiota Karanasou, Mark J. F. Gales
ICASSP3
2016 System combination with log-linear models
abstract
Improved speech recognition performance can often be obtained by combining multiple systems together. Joint decoding, where scores from multiple systems are combined during decoding rather than combining hypotheses, is one efficient approach for system combination. In standard joint decoding the frame log-likelihoods from each system are used as the scores. These scores are then weighted and summed to yield the final score for a frame. The system combination weights for this process are usually empirically set. In this paper, a recently proposed scheme for learning these system weights is investigated for a standard noise-robust speech recognition task, AURORA 4. High performance tandem and hybrid systems for this task are described. By applying state-of-the-art training approaches and configurations for the bottleneck features of the tandem system, the difference in performance between the tandem and hybrid systems is significantly smaller than usually observed on this task. A log-linear model is then used to estimate system weights between these systems. Training the system weights yields additional gains over empirically set system weights when used for decoding. Furthermore, when used in a lattice rescoring fashion, further gains can be obtained.
Chao Zhang 0031, Anton Ragni, Mark J. F. Gales, Philip C. Woodland
ICASSP4
2016 Incorporating a Generative Front-End Layer to Deep Neural Network for Noise Robust Automatic Speech Recognition
Souvik Kundu 0003, Khe Chai Sim, Mark J. F. Gales
INTERSPEECH3
2016 Selection of Multi-Genre Broadcast Data for the Training of Automatic Speech Recognition Systems
abstract
This paper compares schemes for the selection of multi-genre broadcast data and corresponding transcriptions for speech recognition model training. Selections of the same amount of data (700 hours) from lightly supervised alignments based on the same original subtitle transcripts are compared. Data segments were selected according to a maximum phone matched error rate between the lightly supervised decoding and the original transcript. The data selected with an improved lightly supervised system yields lower word error rates (WERs). Detailed comparisons of the data selected on carefully transcribed development data show how the selected portions match the true phone error rate for each genre. From a broader perspective, it is shown that for different genres, either the original subtitles or the lightly supervised output should be used for model training and a suitable combination yields further reductions in final WER.
Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanman Qian, Philip C. Woodland, Chao Zhang 0031
INTERSPEECH2
2016 Multi-Language Neural Network Language Models
abstract
Recently there has been a lot of interest in neural network based language models. These models typically consist of vocabulary dependent input and output layers and one or more vocabulary independent hidden layers. One standard issue with these approaches is that large quantities of training data are needed to ensure robust parameter estimates. This poses a significant problem when only limited data is available. One possible way to address this issue is augmentation: model-based, in the form of language model interpolation, and data-based, in the form of data augmentation. However, these approaches may not always be possible to use due to vocabulary dependent input and output layers. This seriously restricts the nature of the data possible to use in augmentation. This paper describes a general solution whereby only one or more vocabulary independent hidden layers are augmented. Such approach makes it possible to examine augmentation from previously impossible domains. Moreover, this approach paves a direct way for multi-task learning with these models. As a proof of the concept this paper examines the use of multilingual data for augmenting hidden layers of recurrent neural network language models. Experiments are conducted using a set of language packs released within IARPA Babel program.
Anton Ragni, Edgar Dakin, Xie Chen 0001, Mark J. F. Gales, Kate M. Knill
INTERSPEECH4
2016 Sequence Student-Teacher Training of Deep Neural Networks
abstract
The performance of automatic speech recognition can often be significantly improved by combining multiple systems together. Though beneficial, ensemble methods can be computationally expensive, often requiring multiple decoding runs. An alternative approach, appropriate for deep learning schemes, is to adopt student-teacher training. Here, a student model is trained to reproduce the outputs of a teacher model, or ensemble of teachers. The standard approach is to train the student model on the frame posterior outputs of the teacher. This paper examines the interaction between student-teacher training schemes and sequence training criteria, which have been shown to yield significant performance gains over frame-level criteria. There are several possible options for integrating sequence training, including training of the ensemble and further training of the student. This paper also proposes an extension to the student-teacher framework, where the student is trained to emulate the hypothesis posterior distribution of the teacher, or ensemble of teachers. This sequence student-teacher training approach allows the benefit of student-teacher training to be directly combined with sequence training schemes. These approaches are evaluated on two speech recognition tasks: a Wall Street Journal based task and a low-resource Tok Pisin conversational telephone speech task from the IARPA Babel programme.
Jeremy H. M. Wong, Mark J. F. Gales
INTERSPEECH2
2016 Stimulated Deep Neural Network for Speech Recognition
abstract
Deep neural networks (DNNs) and deep learning approaches yield state-of-the-art performance in a range of tasks, including speech recognition.However, the parameters of the network are hard to analyze, making network regularization and robust adaptation challenging.Stimulated training has recently been proposed to address this problem by encouraging the node activation outputs in regions of the network to be related.This kind of information aids visualization of the network, but also has the potential to improve regularization and adaptation.This paper investigates stimulated training of DNNs for both of these options.These schemes take advantage of the smoothness constraints that stimulated training offers.The approaches are evaluated on two large vocabulary speech recognition tasks: a U.S. English broadcast news (BN) task and a Javanese conversational telephone speech task from the IARPA Babel program.Stimulated DNN training acquires consistent performance gains on both tasks over unstimulated baselines.On the BN task, the proposed smoothing approach is also applied to rapid adaptation, again outperforming the standard adaptation scheme.
Chunyang Wu, Panagiota Karanasou, Mark J. F. Gales, Khe Chai Sim
INTERSPEECH3
2016 Log-Linear System Combination Using Structured Support Vector Machines
abstract
Building high accuracy speech recognition systems with limited language resources is a highly challenging task. Although the use of multi-language data for acoustic models yields improvements, performance is often unsatisfactory with highly limited acoustic training data. In these situations, it is possible to consider using multiple well trained acoustic models and combine the system outputs together. Unfortunately, the computational cost associated with these approaches is high as multiple decoding runs are required. To address this problem, this paper examines schemes based on log-linear score combination. This has a number of advantages over standard combination schemes. Even with limited acoustic training data, it is possible to train, for example, phone-specific combination weights, allowing detailed relationships between the available well trained models to be obtained. To ensure robust parameter estimation, this paper casts log-linear score combination into a structured support vector machine (SSVM) learning task. This yields a method to train model parameters with good generalisation properties. Here the SSVM feature space is a set of scores from well-trained individual systems. The SSVM approach is compared to lattice rescoring and confusion network combination using language packs released within the IARPA Babel program.
Anton Ragni, Mark J. F. Gales, Kate M. Knill
INTERSPEECH3
2016 Towards Using Conversations with Spoken Dialogue Systems in the Automated Assessment of Non-Native Speakers of English
abstract
Diane Litman, Steve Young, Mark Gales, Kate Knill, Karen Ottewell, Rogier van Dalen, David Vandyke. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016.
Diane J. Litman, Steve J. Young, Mark J. F. Gales, Kate M. Knill, Karen Ottewell, Rogier C. van Dalen, David Vandyke
SIGDIAL Conference3
2016 Efficient Training and Evaluation of Recurrent Neural Network Language Models for Automatic Speech Recognition
abstract
Recurrent neural network language models (RNNLMs) are becoming increasingly popular for a range of applications including automatic speech recognition. An important issue that limits their possible application areas is the computational cost incurred in training and evaluation. This paper describes a series of new efficiency improving approaches that allows RNNLMs to be more efficiently trained on graphics processing units (GPUs) and evaluated on CPUs. First, a modified RNNLM architecture with a nonclass-based, full output layer structure (F-RNNLM) is proposed. This modified architecture facilitates a novel spliced sentence bunch mode parallelization of F-RNNLM training using large quantities of data on a GPU. Second, two efficient RNNLM training criteria based on variance regularization and noise contrastive estimation are explored to specifically reduce the computation associated with the RNNLM output layer softmax normalisation term. Finally, a pipelined training algorithm utilizing multiple GPUs is also used to further improve the training speed. Initially, RNNLMs were trained on a moderate dataset with 20M words from a large vocabulary conversational telephone speech recognition task. The training time of RNNLM is reduced by up to a factor of 53 on a single GPU over the standard CPU-based RNNLM toolkit. A 56 times speed up in test time evaluation on a CPU was obtained over the baseline F-RNNLMs. Consistent improvements in both recognition accuracy and perplexity were also obtained over C-RNNLMs. Experiments on Google's one billion corpus also reveals that the training of RNNLM scales well.
Xie Chen 0001, Xunying Liu, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Two Efficient Lattice Rescoring Methods Using Recurrent Neural Network Language Models
abstract
An important part of the language modelling problem for automatic speech recognition (ASR) systems, and many other related applications, is to appropriately model long-distance context dependencies in natural languages. Hence, statistical language models (LMs) that can model longer span history contexts, for example, recurrent neural network language models (RNNLMs), have become increasingly popular for state-of-the-art ASR systems. As RNNLMs use a vector representation of complete history contexts, they are normally used to rescore N-best lists. Motivated by their intrinsic characteristics, two efficient lattice rescoring methods for RNNLMs are proposed in this paper. The first method uses an n-gram style clustering of history contexts. The second approach directly exploits the distance measure between recurrent hidden history vectors. Both methods produced 1-best performance comparable to a 10 k-best rescoring baseline RNNLM system on two large vocabulary conversational telephone speech recognition tasks for US English and Mandarin Chinese. Consistent lattice size compression and recognition performance improvements after confusion network (CN) decoding were also obtained over the prefix tree structured N-best rescoring approach.
Xunying Liu, Xie Chen 0001, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 The MGB challenge: Evaluating multi-genre broadcast media recognition
abstract
This paper describes the Multi-Genre Broadcast (MGB) Challenge at ASRU 2015, an evaluation focused on speech recognition, speaker diarization, and "lightly supervised" alignment of BBC TV recordings. The challenge training data covered the whole range of seven weeks BBC TV output across four channels, resulting in about 1,600 hours of broadcast audio. In addition several hundred million words of BBC subtitle text was provided for language modelling. A novel aspect of the evaluation was the exploration of speech recognition and speaker diarization in a longitudinal setting — i.e. recognition of several episodes of the same show, and speaker diarization across these episodes, linking speakers. The longitudinal tasks also offered the opportunity for systems to make use of supplied metadata including show title, genre tag, and date/time of transmission. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained.
Peter Bell 0001, Mark J. F. Gales, Thomas Hain, Jonathan Kilgour, Pierre Lanchantin, Xunying Liu, Andrew McParland, Steve Renals, Oscar Saz-Torralba, Mirjam Wester, Philip C. Woodland
ASRU2
2015 Investigation of back-off based interpolation between recurrent neural network and n-gram language models
abstract
Recurrent neural network language models (RNNLMs) have become an increasingly popular choice for speech and language processing tasks including automatic speech recognition (ASR). As the generalization patterns of RNNLMs and n-gram LMs are inherently different, RNNLMs are usually combined with n-gram LMs via a fixed weighting based linear interpolation in state-of-the-art ASR systems. However, previous work doesn't fully exploit the difference of modelling power of the RNNLMs and n-gram LMs as n-gram level changes. In order to fully exploit the detailed n-gram level complementary attributes between the two LMs, a back-off based compact representation of n-gram dependent interpolation weights is proposed in this paper. This approach allows weight parameters to be robustly estimated on limited data. Experimental results are reported on the three tasks with varying amounts of training data. Small and consistent improvements in both perplexity and WER were obtained using the proposed interpolation approach over the baseline fixed weighting based linear interpolation.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ASRU3
2015 Multilingual representations for low resource speech recognition and keyword search
abstract
This paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation.
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland
ASRU15
2015 Structured discriminative models using deep neural-network features
abstract
State-of-the-art speech recognisers employ neural networks in various configurations. A standard (hybrid) speech recogniser computes the likelihood for one time frame and state, using only one out of thousands of possible neural-network outputs. However, the whole output vector carries information. In this paper, features from state-of-the-art speech recognisers are collected per phone given a particular context, and input to a discriminative log-linear model. The log-linear model is trained with conditional maximum likelihood or a large-margin criterion. A key element is the prior on the parameters of the log-linear model. The mean of the prior is set to the point where the performance of the original systems is attained. The log-linear model then provides an additional increase over the state-of-the-art performance of the individual systems.
Rogier C. van Dalen, Anton Ragni, Chao Zhang 0031, Mark J. F. Gales
ASRU6
2015 Speaker diarisation and longitudinal linking in multi-genre broadcast data
abstract
This paper presents a multi-stage speaker diarisation system with longitudinal Linking developed on BBC multi-genre data for the 2015 Multi-Genre Broadcast (MGB) challenge. The basic speaker diarisation system draws on techniques from the Cambridge March 2005 system with a new deep neural network (DNN)-based speech/non speech segmenter. A newly developed linking stage is next added to the basic diarisation output aiming at the identification of speakers across multiple episodes of the same series. The longitudinal constraint imposes an incremental processing of the episodes, where speaker labels for each episode can be obtained using only material from the episode in question, and those broadcast earlier in time. The nature of the data as well as the longitudinal linking constraint position this diarisation task as a new open-research topic, and a particularly challenging one. Different linking clustering metrics are compared and the lowest within-episode and cross-episode DER scores are achieved on the MGB challenge evaluation set.
Panagiota Karanasou, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031
ASRU2
2015 The development of the cambridge university alignment systems for the multi-genre broadcast challenge
abstract
We describe the alignment systems developed both for the preparation of data for the Multi-Genre Broadcast (MGB) challenge and for our participation in the transcription and alignment tasks. Captions of varying quality are aligned with the audio of TV shows that range from few minutes long to more than six hours. Lightly supervised decoding is performed on the audio and the output text is aligned with the original text transcript. Reliable split points are found and the resulting text chunks are force-aligned with the corresponding audio segments. Confidence scores are associated with the aligned data. Multiple refinements — including audio segmentation based on deep neural networks (DNNs) and the use of DNN-based acoustic models — were used to improve the performance. The final MGB alignment system had the highest F-measure value on the evaluation data.
Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031
ASRU2
2015 Improving the interpretability of deep neural networks with stimulated learning
abstract
Deep Neural Networks (DNNs) have demonstrated improvements in acoustic modelling for automatic speech recognition. However, they are often used as a black box, and not much is understood about what each of the hidden layers does. We seek to understand how the activations in the hidden layers change with different input, and how we can leverage such knowledge to modify the behaviour of the model. To this end, we propose stimulated deep learning where stimuli are introduced during the DNN training process to influence the behaviour of the hidden units. Specifically, constraints are applied so that the hidden units of each layer will exhibit phone-dependent regional activities when arranged in a 2-dimensional grid. We demonstrate that such constraints are able to yield visible activation regions without compromising the classification of the network and suppressing the activations for a region affects the classification accuracy of the corresponding phone more than the others.
Shawn Tan, Khe Chai Sim, Mark J. F. Gales
ASRU3
2015 Cambridge university transcription systems for the multi-genre broadcast challenge
abstract
We describe the development of our speech-to-text transcription systems for the 2015 Multi-Genre Broadcast (MGB) challenge. Key features of the systems are: a segmentation system based on deep neural networks (DNNs); the use of HTK 3.5 for building DNN-based hybrid and tandem acoustic models and the use of these models in a joint decoding framework; techniques for adaptation of DNN based acoustic models including parameterised activation function adaptation; alternative acoustic models built using Kaldi; and recurrent neural network language models (RNNLMs) and RNNLM adaptation. The same language models were used with both HTK and Kaldi acoustic models and various combined systems built. The final systems had the lowest error rates on the evaluation data.
Philip C. Woodland, Xunying Liu, Yanmin Qian, Chao Zhang 0031, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin
ASRU5
2015 Improving the training and evaluation efficiency of recurrent neural network language models
abstract
Recurrent neural network language models (RNNLMs) are becoming increasingly popular for speech recognition. Previously, we have shown that RNNLMs with a full (non-classed) output layer (F-RNNLMs) can be trained efficiently using a GPU giving a large reduction in training time over conventional class-based models (C-RNNLMs) on a standard CPU. However, since test-time RNNLM evaluation is often performed entirely on a CPU, standard F-RNNLMs are inefficient since the entire output layer needs to be calculated for normalisation. In this paper, it is demonstrated that C-RNNLMs can be efficiently trained on a GPU, using our spliced sentence bunch technique which allows good CPU test-time performance (42× speedup over F-RNNLM). Furthermore, the performance of different classing approaches is investigated. We also examine the use of variance regularisation of the softmax denominator for F-RNNLMs and show that it allows F-RNNLMs to be efficiently used in test (56× speedup on a CPU). Finally the use of two GPUs for F-RNNLM training using pipelining is described and shown to give a reduction in training time over a single GPU by a factor of 1.6×.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2015 Recurrent neural network language model training with noise contrastive estimation for speech recognition
abstract
In recent years recurrent neural network language models (RNNLMs) have been successfully applied to a range of tasks including speech recognition. However, an important issue that limits the quantity of data used, and their possible application areas, is the computational cost in training. A signi??cant part of this cost is associated with the softmax function at the output layer, as this requires a normalization term to be explicitly calculated. This impacts both the training and testing speed, especially when a large output vocabulary is used. To address this problem, noise contrastive estimation (NCE) is explored in RNNLM training. NCE does not require the above normalization during both training and testing. It is insensitive to the output layer size. On a large vocabulary conversational telephone speech recognition task, a doubling in training speed on a GPU and a 56 times speed up in test time evaluation on a CPU were obtained.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2015 Improving multiple-crowd-sourced transcriptions using a speech recogniser
abstract
This paper introduces a method to produce high-quality transcriptions of speech data from only two crowd-sourced transcriptions. These transcriptions, produced cheaply by people on the Internet, for example through Amazon Mechanical Turk, are often of low quality. Often, multiple crowd-sourced transcriptions are combined to form one transcription of higher quality. However, the state of the art is to use essentially a form of majority voting, which requires at least three transcriptions for each utterance. This paper shows how to refine this approach to work with only two transcriptions. It then introduces a method that uses a speech recogniser (bootstrapped on a simple combination scheme) to combine transcriptions. When only two crowd-sourced transcriptions are available, on a noisy data set this improves the word error rate to gold-standard transcriptions by 21% relative.
Rogier C. van Dalen, Kate M. Knill, Pirros Tsiakoulis, Mark J. F. Gales
ICASSP4
2015 Robust excitation-based features for Automatic Speech Recognition
abstract
In this paper we investigate the use of noise-robust features characterizing the speech excitation signal as complementary features to the usually considered vocal tract based features for Automatic Speech Recognition (ASR). The proposed Excitation-based Features (EBF) are tested in a state-of-the-art Deep Neural Network (DNN) based hybrid acoustic model for speech recognition. The suggested excitation features expand the set of periodicity features previously considered for ASR, expecting that these features help in a better discrimination of the broad phonetic classes (e.g., fricatives, nasal, vowels, etc.). Our experiments on the AMI meeting transcription system showed that the proposed EBF yield a relative word error rate reduction of about 5% when combined with conventional PLP features. Further experiments led on Aurora4 confirmed the robustness of the EBF to both additive and convolutive noises, with a relative improvement of 4.3% obtained by combinining them with mel filter banks.
Thomas Drugman, Yannis Stylianou, Langzhou Chen, Xie Chen 0001, Mark J. F. Gales
ICASSP5
2015 Unicode-based graphemic systems for limited resource languages
abstract
Large vocabulary continuous speech recognition systems require a mapping from words, or tokens, into sub-word units to enable robust estimation of acoustic model parameters, and to model words not seen in the training data. The standard approach to achieve this is to manually generate a lexicon where words are mapped into phones, often with attributes associated with each of these phones. Contextdependent acoustic models are then constructed using decision trees where questions are asked based on the phones and phone attributes. For low-resource languages, it may not be practical to manually generate a lexicon. An alternative approach is to use a graphemic lexicon, where the “pronunciation” for a word is defined by the letters forming that word. This paper proposes a simple approach for building graphemic systems for any language written in unicode. The attributes for graphemes are automatically derived using features from the unicode character descriptions. These attributes are then used in decision tree construction. This approach is examined on the IARPA Babel Option Period 2 languages, and a Levantine Arabic CTS task. The described approach achieves comparable, and complementary, performance to phonetic lexicon-based approaches.
Mark J. F. Gales, Kate M. Knill, Anton Ragni
ICASSP1
2015 Paraphrastic recurrent neural network language models
abstract
Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems. Linguistic factors in??uencing the realization of surface word sequences, for example, expressive richness, are only implicitly learned by RNNLMs. Observed sentences and their associated alternative paraphrases representing the same meaning are not explicitly related during training. In order to improve context coverage and generalization, paraphrastic RNNLMs are investigated in this paper. Multiple paraphrase variants were automatically generated and used in paraphrastic RNNLM training. Using a paraphrastic multi-level RNNLM modelling both word and phrase sequences, signi??cant error rate reductions of 0.6% absolute and perplexity reduction of 10% relative were obtained over the baseline RNNLM on a large vocabulary conversational telephone speech recognition system trained on 2000 hours of audio and 545 million words of texts. The overall improvement over the baseline n-gram LM was increased from 8.4% to 11.6% relative.
Xunying Liu, Xie Chen 0001, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2015 A language space representation for speech recognition
abstract
The number of languages for which speech recognition systems have become available is growing each year. This paper proposes to view languages as points in some rich space, termed language space, where bases are eigen-languages and a particular selection of the projection determines points. Such an approach could not only reduce development costs for each new language but also provide automatic means for language analysis. For the initial proof of the concept, this paper adopts cluster adaptive training (CAT) known for inducing similar spaces for speaker adaptation needs. The CAT approach used in this paper builds on the previous work for language adaptation in speech synthesis and extends it to Gaussian mixture modelling more appropriate for speech recognition. Experiments conducted on IARPA Babel program languages show that such language space representations can outperform language independent models and discover closely related languages in an automatic way.
Anton Ragni, Mark J. F. Gales, Kate M. Knill
ICASSP2
2015 Multi-basis adaptive neural network for rapid adaptation in speech recognition
abstract
Recent progress in acoustic modeling with deep neural network has significantly improved the performance of automatic speech recognition systems. However, it remains as an open problem how to rapidly adapt these networks with limited, unsupervised, data. Most existing methods to adapt a neural network involve modifying a large number of parameters thus rapid adaptation is not possible with these schemes. In this paper, the multi-basis adaptive neural network is proposed, a new neural network configuration which only requires very few parameters for adaptation. By modifying the topology of a single multi-layer perception, a set of sub-networks with restricted connectivity are introduced to collaboratively capture different acoustic properties. The outputs of those sub-networks are combined by speaker-dependent interpolation weights. In addition, the complete system can be optimized in an adaptive training fashion when non-homogeneous training data are used. The performance of unsupervised adaptation is evaluated on two datasets. It outperforms the speaker-independent hybrid DNN-HMM baseline both on the Broadcast News English and the AURORA-4 tasks.
Chunyang Wu, Mark J. F. Gales
ICASSP2
2015 Recurrent neural network language model adaptation for multi-genre broadcast speech recognition
abstract
Recurrent neural network language models (RNNLMs) have recently become increasingly popular for many applications including speech recognition. In previous research RNNLMs have normally been trained on well-matched in-domain data. The adaptation of RNNLMs remains an open research area to be explored. In this paper, genre and topic based RNNLMadaptation techniques are investigated for a multi-genre broadcast transcription task. A number of techniques including Probabilistic Latent Semantic Analysis, Latent Dirichlet Allocation and Hierarchical Dirichlet Processes are used to extract show level topic information. These were then used as additional input to the RNNLM during training, which can facilitate unsupervised test time adaptation. Experiments using a state-of-theart LVCSR system trained on 1000 hours of speech and more than 1 billion words of text showed adaptation could yield perplexity reductions of 8% relatively over the baseline RNNLM and small but consistent word error rate reductions.
Xie Chen 0001, Tian Tan 0002, Xunying Liu, Pierre Lanchantin, M. Wan, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH6
2015 Annotating large lattices with the exact word error
abstract
The acoustic model in modern speech recognisers is trained discriminatively, for example with the minimum Bayes risk. This criterion is hard to compute exactly, so that it is normally approximated by a criterion that uses fixed alignments of lat-tice arcs. This approximation becomes particularly problematic with new types of acoustic models that require flexible align-ments. It would be best to annotate lattices with the risk mea-sure of interest, the exact word error. However, the algorithm for this uses finite-state automaton determinisation, which has exponential complexity and runs out of memory for large lat-tices. This paper introduces a novel method for determinis-ing and minimising finite-state automata incrementally. Since it uses less memory, it can be applied to larger lattices. Index Terms: speech recognition, discriminative training, min-imum Bayes risk
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH2
2015 I-vector estimation using informative priors for adaptation of deep neural networks
abstract
This is the author accepted manuscript. The final version is available from ISCA via http://www.isca-speech.org/archive/interspeech_2015/i15_2872.html Supporting data for this paper is available at the http://www.repository.cam.ac.uk/handle/1810/248387 data repository.
Panagiota Karanasou, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2015 Reconstructing voices within the multiple-average-voice-model framework
abstract
Personalisation of voice output communication aids (VOCAs) allows to preserve the vocal identity of people suffering from speech disorders. This can be achieved by the adaptation of HMM-based speech synthesis systems using a small amount of adaptation data. When the voice has begun to deteriorate, reconstruction is still possible in the statistical domain by correcting the parameters of the models associated with the speech disorder. This can be done by substituting those with parameters from a donor’s voice, at risk of losing part of the identity of the patient. Recently, the Multiple-Average-Voice-Model (Multiple AVM) framework has been proposed for speaker adaptation. Adaptation is performed via interpolation into a speaker eigenspace spanned by the mean vectors of speaker-adapted AVMs which can be tuned to the individual speaker. In this paper, we present the benefits of this framework for voice reconstruction: it requires only a very small amount of adaptation data, interpolation can be performed in a clean speech eigenspace and the resulting voice can be easily fine-tuned by acting on the interpolation weights. We illustrate our points with a subjective assessment of the reconstructed voice. Index Terms: HMM-Based speech synthesis, speaker adaptation, multiple average voice model, cluster adaptive training, voice reconstruction, voice output communication aids.
Pierre Lanchantin, Christophe Veaux, Mark J. F. Gales, Simon King 0001, Junichi Yamagishi
INTERSPEECH3
2015 The Cambridge University 2014 BOLT conversational telephone Mandarin Chinese LVCSR system for speech translation
abstract
This paper presents the development of the 2014 Cambridge University conversational telephone Mandarin Chinese LVCSR system for the DARPA BOLT speech translation evaluation. A range of advanced modelling techniques were employed to both improve the recognition performance and provide a suitable integration with the translation system. These include an improved system combination technique using frame level acoustic model combination via joint decoding. Sequence trained deep neural network (DNN) based hybrid and tandem systems were combined on-the-fly to produce a consistent decoding output during search. A multi-level paraphrastic recurrent neural network LM (RNNLM) modelling both alternative paraphrase expressions and character sequences while preserving a consistent character to word segmentation was also used. This system gave an overall character error rate (CER) of 29.1% on the BOLT dev14 development set.
Xunying Liu, Federico Flego, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH5
2015 Improving speech recognition and keyword search for low resource languages using web data
abstract
We describe the use of text data scraped from the web to augment language models for Automatic Speech Recognition and Keyword Search for Low Resource Languages.We scrape text from multiple genres including blogs, online news, translated TED talks, and subtitles.Using linearly interpolated language models, we find that blogs and movie subtitles are more relevant for language modeling of conversational telephone speech and obtain large reductions in out-of-vocabulary keywords.Furthermore, we show that the web data can improve Term Error Rate Performance by 3.8% absolute and Maximum Term-Weighted Value in Keyword Search by 0.0076-0.1059absolute points.Much of the gain comes from the reduction of out-of-vocabulary items.
Gideon Mendels, Erica Cooper, Victor Soto, Julia Hirschberg, Mark J. F. Gales, Kate M. Knill, Anton Ragni
INTERSPEECH5
2015 Joint decoding of tandem and hybrid systems for improved keyword spotting on low resource languages
abstract
Copyright © 2015 ISCA. Keyword spotting (KWS) for low-resource languages has drawn increasing attention in recent years. The state-of-the-art KWS systems are based on lattices or Confusion Networks (CN) generated by Automatic Speech Recognition (ASR) systems. It has been shown that considerable KWS gains can be obtained by combining the keyword detection results from different forms of ASR systems, e.g., Tandem and Hybrid systems. This paper investigates an alternative combination scheme for KWS using joint decoding. This scheme treats a Tandem system and a Hybrid system as two separate streams, and makes a linear combination of individual acoustic model log-likelihoods. Joint decoding is more efficient as it requires just a single pass of decoding and a single pass of keyword search. Experiments on six Babel OP2 development languages show that joint decoding is capable of providing consistent gains over each individual system. Moreover, it is possible to efficiently rescore the joint decoding lattices with Tandem or Hybrid acoustic models, and further KWS gains can be obtained by merging the detection posting lists from the joint decoding lattices and rescored lattices.
Anton Ragni, Mark J. F. Gales, Kate M. Knill, Philip C. Woodland, Chao Zhang 0031
INTERSPEECH3
2015 Environmentally robust ASR front-end for deep neural network acoustic models
abstract
This paper examines the individual and combined impacts of various front-end approaches on the performance of deep neural network (DNN) based speech recognition systems in distant talking situations, where acoustic environmental distortion degrades the recognition performance. Training of a DNN-based acoustic model consists of generation of state alignments followed by learning the network parameters. This paper first shows that the network parameters are more sensitive to the speech quality than the alignments and thus this stage requires improvement. Then, various front-end robustness approaches to addressing this problem are categorised based on functionality. The degree to which each class of approaches impacts the performance of DNN-based acoustic models is examined experimentally. Based on the results, a front-end processing pipeline is proposed for efficiently combining different classes of approaches. Using this front-end, the combined effects of different classes of approaches are further evaluated in a single distant microphone-based meeting transcription task with both speaker independent (SI) and speaker adaptive training (SAT) set-ups. By combining multiple speech enhancement results, multiple types of features, and feature transformation, the front-end shows relative performance gains of 7.24% and 9.83% in the SI and SAT scenarios, respectively, over competitive DNN-based systems using log mel-filter bank features.
Takuya Yoshioka, Mark J. F. Gales
Comput. Speech Lang.2
2015 Speaker and Expression Factorization for Audiobook Data: Expressiveness and Transplantation
abstract
Expressive synthesis from text is a challenging problem. There are two issues. First, read text is often highly expressive to convey the emotion and scenario in the text. Second, since the expressive training speech is not always available for different speakers, it is necessary to develop methods to share the expressive information over speakers. This paper investigates the approach of using very expressive, highly diverse audiobook data from multiple speakers to build an expressive speech synthesis system. Both of two problems are addressed by considering a factorized framework where speaker and emotion are modeled in separate sub-spaces of a cluster adaptive training (CAT) parametric speech synthesis system. The sub-spaces for the expressive state of a speaker and the characteristics of the speaker are jointly trained using a set of audiobooks. In this work, the expressive speech synthesis system works in two distinct modes. In the first mode, the expressive information is given by audio data and the adaptation method is used to extract the expressive information in the audio data. In the second mode, the input of the synthesis system is plain text and a full expressive synthesis system is examined where the expressive state is predicted from the text. In both modes, the expressive information is shared and transplanted over different speakers. Experimental results show that in both modes, the expressive speech synthesis method proposed in this work significantly improves the expressiveness of the synthetic speech for different speakers. Finally, this paper also examines whether it is possible to predict the expressive states from text for multiple speakers using a single model, or whether the prediction process needs to be speaker specific.
Langzhou Chen, Norbert Braunschweiler, Mark J. F. Gales
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Speaker dependent expression predictor from text: Expressiveness and transplantation
abstract
Automatically generating expressive speech from plain text is an important research topic in speech synthesis. Given the same text, different speakers may interpret it and read it in very different ways. This implies that expression prediction from text is a speaker dependent task. Previous work presented an integrated method for expression prediction and speech synthesis which can be used to model the diverse expressions in human's speech and build speaker dependent expression predictors from text. This work extends the integrated method for expression prediction and speech synthesis into a framework for speaker and expression factorization. The expressions generated by the speaker dependent expression predictors can be represented in a shared expression space, and in this space the expressions can be transplanted between different speakers. The experimental results indicate that based on the proposed method, the expressiveness of the synthetic speech can be improved for different speakers. Furthermore this work also shows how important the speaker specific information is for the performance of the expression predictor from text.
Langzhou Chen, Norbert Braunschweiler, Mark J. F. Gales
ICASSP3
2014 Multiple-average-voice-based speech synthesis
abstract
This paper describes a novel approach for the speaker adaptation of statistical parametric speech synthesis systems based on the interpolation of a set of average voice models (AVM). Recent results have shown that the quality/naturalness of adapted voices depends on the distance from the average voice model used for speaker adaptation. This suggests the use of several AVMs trained on carefully chosen speaker clusters from which a more suitable AVM can be selected/interpolated during the adaptation. In the proposed approach a set of AVMs, a multiple-AVM, is trained on distinct clusters of speakers which are iteratively re-assigned during the estimation process initialised according to metadata. During adaptation, each AVM from the multiple-AVM is first adapted towards the target speaker. The adapted means from the AVMs are then interpolated to yield the final speaker adapted mean for synthesis. It is shown, performing speaker adaptation on a corpus of British speakers with various regional accents, that the quality/naturalness of synthetic speech of adapted voices is significantly higher than when considering a single factor-independent AVM selected according to the target speaker characteristics.
Pierre Lanchantin, Mark J. F. Gales, Simon King 0001, Junichi Yamagishi
ICASSP2
2014 Paraphrastic neural network language models
abstract
Expressive richness in natural languages presents a significant challenge for statistical language models (LM). As multiple word sequences can represent the same underlying meaning, only modelling the observed surface word sequence can lead to poor context coverage. To handle this issue, paraphrastic LMs were previously proposed to improve the generalization of back-off n-gram LMs. Paraphrastic neural network LMs (NNLM) are investigated in this paper. Using a paraphrastic multi-level feedforward NNLM modelling both word and phrase sequences, significant error rate reductions of 1.3% absolute (8% relative) and 0.9% absolute (5.5% relative) were obtained over the baseline n-gram and NNLM systems respectively on a state-of-the-art conversational telephone speech recognition system trained on 2000 hours of audio and 545 million words of texts.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP2
2014 Efficient lattice rescoring using recurrent neural network language models
abstract
Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems due to their inherently strong generalization performance. As these models use a vector representation of complete history contexts, RNNLMs are normally used to rescore N-best lists. Motivated by their intrinsic characteristics, two novel lattice rescoring methods for RNNLMs are investigated in this paper. The first uses an n-gram style clustering of history contexts. The second approach directly exploits the distance measure between hidden history vectors. Both methods produced 1-best performance comparable with a 10k-best rescoring baseline RNNLM system on a large vocabulary conversational telephone speech recognition task. Significant lattice size compression of over 70% and consistent improvements after confusion network (CN) decoding were also obtained over the N-best rescoring approach.
Xunying Liu, Yongqiang Wang 0006, Xie Chen 0001, Mark J. F. Gales, Philip C. Woodland
ICASSP4
2014 Cluster adaptive training of average voice models
abstract
Hidden Markov model based text-to-speech systems may be adapted so that the synthesised speech sounds like a particular person. The average voice model (AVM) approach uses linear transforms to achieve this while multiple decision tree cluster adaptive training (CAT) represents different speakers as points in a low dimensional space. This paper describes a novel combination of CAT and AVM for modelling speakers. CAT yields higher quality synthetic speech than AVMs but AVMs model the target speaker better. The resulting combination may be interpreted as a more powerful version of the AVM. Results show that the combination achieves better target speaker similarity when compared with both AVM and CAT while the speech quality is in-between AVM and CAT.
Vincent Wan, Javier Latorre, Kayoko Yanagisawa, Mark J. F. Gales, Yannis Stylianou
ICASSP4
2014 Infinite structured support vector machines for speech recognition
abstract
Discriminative models, like support vector machines (SVMs), have been successfully applied to speech recognition and improved performance. A Bayesian non-parametric version of the SVM, the infinite SVM, improves on the SVM by allowing more flexible decision boundaries. However, like SVMs, infinite SVMs model each class separately, which restricts them to classifying one word at a time. A generalisation of the SVM is the structured SVM, whose classes can be sequences of words that share parameters. This paper studies a combination of Bayesian non-parametrics and structured models. One specific instance called infinite structured SVM is discussed in detail, which brings the advantages of the infinite SVM to continuous speech recognition.
Rogier C. van Dalen, Shixiong Zhang 0001, Mark J. F. Gales
ICASSP4
2014 Impact of single-microphone dereverberation on DNN-based meeting transcription systems
abstract
Over the past few decades, a range of front-end techniques have been proposed to improve the robustness of automatic speech recognition systems against environmental distortion. While these techniques are effective for small tasks consisting of carefully designed data sets, especially when used with a classical acoustic model, there has been limited evidence that they are useful for a state-of-the-art system with large scale realistic data. This paper focuses on reverberation as a type of distortion and investigates the degree to which dereverberation processing can improve the performance of various forms of acoustic models based on deep neural networks (DNNs) in a challenging meeting transcription task using a single distant microphone. Experimental results show that dereverberation improves the recognition performance regardless of the acoustic model structure and the type of the feature vectors input into the neural networks, providing additional relative improvements of 4.7% and 4.1% to our best configured speaker-independent and speaker-adaptive DNN-based systems, respectively.
Takuya Yoshioka, Xie Chen 0001, Mark J. F. Gales
ICASSP3
2014 Investigation of unsupervised adaptation of DNN acoustic models with filter bank input
abstract
Adaptation to speaker variations is an essential component of speech recognition systems. One common approach to adapting deep neural network (DNN) acoustic models is to perform global constrained maximum likelihood linear regression (CMLLR) at some point of the systems. Using CMLLR (or more generally, generative approaches) is advantageous especially in unsupervised adaptation scenarios with high baseline error rates. On the other hand, as the DNNs are less sensitive to the increase in the input dimensionality than GMMs, it is becoming more popular to use rich speech representations, such as log mel-filter bank channel outputs, instead of conventional low-dimensional feature vectors, such as MFCCs and PLP coefficients. This work discusses and compares three different configurations of DNN acoustic models that allow CMLLR-based speaker adaptive training (SAT) to be performed in systems with filter bank inputs. Results of unsupervised adaptation experiments conducted on three different data sets are presented, demonstrating that, by choosing an appropriate configuration, SAT with CMLLR can improve the performance of a well-trained filter bank-based speaker independent DNN system by 10.6% relative in a challenging task with a baseline error rate above 40%. It is also shown that the filter bank features are advantageous than the conventional features even when they are used with SAT models. Some other insights are also presented, including the effects of block diagonal transforms and system combination.
Takuya Yoshioka, Anton Ragni, Mark J. F. Gales
ICASSP3
2014 An initial investigation of long-term adaptation for meeting transcription
abstract
Meeting transcription is a very useful and challenging task. The majority of research to date has focused on individual meeting, or only a small group of meetings. In many practical deploy-ments, multiple related meetings will take place over a long pe-riod of time. This paper describes an initial investigation of how this long-term data can be used to improve meeting tran-scription. A corpus of technical meetings, using a single micro-phone array, was collected over a two year period, yielding a total of 179 hours of meeting data. Baseline systems using deep neural network acoustic models, in both Tandem and Hybrid configurations, and neural network-based language models are described. The impact of supervised and unsupervised adap-tation of the acoustic models is then evaluated, as well as the impact of improved language models.
Xie Chen 0001, Mark J. F. Gales, Kate M. Knill, Catherine Breslin, Langzhou Chen, K. K. Chin, Vincent Wan
INTERSPEECH2
2014 Efficient GPU-based training of recurrent neural network language models using spliced sentence bunch
abstract
Recurrent neural network language models (RNNLMs) are be-coming increasingly popular for a range of applications includ-ing speech recognition. However, an important issue that limits the quantity of data, and hence their possible application ar-eas, is the computational cost in training. A standard approach to handle this problem is to use class-based outputs, allowing systems to be trained on CPUs. This paper describes an alter-native approach that allows RNNLMs to be efficiently trained on GPUs. This enables larger quantities of data to be used, and networks with an unclustered, full output layer to be trained. To improve efficiency on GPUs, multiple sentences are “spliced” together for each mini-batch or “bunch ” in training. On a large vocabulary conversational telephone speech recognition task, the training time was reduced by a factor of 27 over the stan-dard CPU-based RNNLM toolkit. The use of an unclustered, full output layer also improves perplexity and recognition per-formance over class-based RNNLMs. Index Terms: language models, recurrent neural network, speech recognition, GPU
Xie Chen 0001, Yongqiang Wang 0006, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH4
2014 Adaptation of deep neural network acoustic models using factorised i-vectors
abstract
The use of deep neural networks (DNNs) in a hybrid configuration is becoming increasingly popular and successful for speech recognition. One issue with these systems is how to efficiently adapt them to reflect an individual speaker or noise condition. Recently speaker i-vectors have been successfully used as an additional input feature for unsupervised speaker adaptation. In this work the use of i-vectors for adaptation is extended to incorporate acoustic factorisation. In particular, separate i-vectors are computed to represent speaker and acoustic environment. By ensuring "orthogonality" between the individual factor representations it is possible to represent a wide range of speaker and environment pairs by simply combining i-vectors from a particular speaker and a particular environment. In this paper the i-vectors are viewed as the weights of a cluster adaptive training (CAT) system, where the underlying models are GMMs rather than HMMs. This allows the factorisation approaches developed for CAT to be directly applied. Initial experiments were conducted on a noise distorted version of the WSJ corpus. Compared to standard speaker-based i-vector adaptation, factorised i-vectors showed performance gains.
Panagiota Karanasou, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2014 Language independent and unsupervised acoustic models for speech recognition and keyword spotting
abstract
Copyright © 2014 ISCA. Developing high-performance speech processing systems for low-resource languages is very challenging. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to train a multi-language bottleneck DNN. Language dependent and/or multi-language (all training languages) Tandem acoustic models (AM) are then trained. This work considers a particular scenario where the target language is unseen in multi-language training and has limited language model training data, a limited lexicon, and acoustic training data without transcriptions. A zero acoustic resources case is first described where a multilanguage AM is directly applied, as a language independent AM (LIAM), to an unseen language. Secondly, in an unsupervised approach a LIAM is used to obtain hypotheses for the target language acoustic data transcriptions which are then used in training a language dependent AM. 3 languages from the IARPA Babel project are used for assessment: Vietnamese, Haitian Creole and Bengali. Performance of the zero acoustic resources system is found to be poor, with keyword spotting at best 60% of language dependent performance. Unsupervised language dependent training yields performance gains. For one language (Haitian Creole) the Babel target is achieved on the in-vocabulary data.
Kate M. Knill, Mark J. F. Gales, Anton Ragni, Shakti P. Rath
INTERSPEECH2
2014 Generating multiple-accent pronunciations for TTS using joint sequence model interpolation
abstract
Standard grapheme-to-phoneme (G2P) systems are trained using a homogeneous lexicon, for example one associated with a particular accent. In practice, a synthesis system may be required to handle multiple accents. Furthermore, a speaker rarely has a pure accent; accents vary continuously within and between regions of a country. Generating phonetic sequences for each accent is possible, but combining them to yield a single synthesis pronunciation is highly challenging. To address this problem, this paper considers a space of accents. The bases for these spaces are defined by statistical G2P models in the form of graphone models. A linear combination of these models define the accent space. By selecting a point in this continuous space, it is possible to specify the accent for an individual speaker. The performance of this approach is evaluated using an accent space defined by American, Scottish and British English. By moving around the accent space, it is shown that it is possible to synthesize speech from all these accents as well as a range of intermediate points.
BalaKrishna Kolluru, Vincent Wan, Javier Latorre, Kayoko Yanagisawa, Mark J. F. Gales
INTERSPEECH5
2014 Speech intonation for TTS: study on evaluation methodology
abstract
The standard evaluation of intonation models is by means of non-referenced subjective tests (pair or MOS) in which subjects rate the quality or compare different samples without any explicit reference. These tests are usually conducted on an isolated sentence basis. However, for a single sentence, with no contextual information, there are multiple valid intonations. A subject's preference over this range of intonation patterns may be highly personal. This paper investigates the degree to which this ambiguity in the appropriate intonation pattern impacts the assessments of prosody for speech synthesis systems. To examine this problem, the variance of the F0 pattern of several vocoded sentences was modified and subjects asked to compare multiple versions with different levels of modification in terms of preference/quality. Then, they were presented with the reference which defines the original intonation and asked about the similarity to that reference. The results show that subjects can identify the samples with no F0 variance modification when given a reference but they don't always prefer them. Thus, non-referenced tests with no context, though may help to analyse user acceptability, may not be appropriate to measure the performance of intonation models.
Javier Latorre, Kayoko Yanagisawa, Vincent Wan, BalaKrishna Kolluru, Mark J. F. Gales
INTERSPEECH5
2014 Data augmentation for low resource languages
abstract
Recently there has been interest in the approaches for training speech recognition systems for languages with limited resources.Under the IARPA Babel program such resources have been provided for a range of languages to support this research area.This paper examines a particular form of approach, data augmentation, that can be applied to these situations.Data augmentation schemes aim to increase the quantity of data available to train the system, for example semi-supervised training, multilingual processing, acoustic data perturbation and speech synthesis.To date the majority of work has considered individual data augmentation schemes, with few consistent performance contrasts or examination of whether the schemes are complementary.In this work two data augmentation schemes, semisupervised training and vocal tract length perturbation, are examined and combined on the Babel limited language pack configuration.Here only about 10 hours of transcribed acoustic data are available.Two languages are examined, Assamese and Zulu, which were found to be the most challenging of the Babel languages released for the 2014 Evaluation.For both languages consistent speech recognition performance gains can be obtained using these augmentation schemes.Furthermore the impact of these performance gains on a down-stream keyword spotting task are also described.
Anton Ragni, Kate M. Knill, Shakti P. Rath, Mark J. F. Gales
INTERSPEECH4
2014 Combining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages
abstract
Copyright © 2014 ISCA. In recent years there has been significant interest in Automatic Speech Recognition (ASR) and KeyWord Spotting (KWS) systems for low resource languages. One of the driving forces for this research direction is the IARPA Babel project. This paper examines the performance gains that can be obtained by combining two forms of deep neural network ASR systems, Tandem and Hybrid, for both ASR and KWS using data released under the Babel project. Baseline systems are described for the five option period 1 languages: Assamese; Bengali; Haitian Creole; Lao; and Zulu. All the ASR systems share common attributes, for example deep neural network configurations, and decision trees based on rich phonetic questions and state-position root nodes. The baseline ASR and KWS performance of Hybrid and Tandem systems are compared for both the "full", approximately 80 hours of training data, and limited, approximately 10 hours of training data, language packs. By combining the two systems together consistent performance gains can be obtained for KWS in all configurations.
Shakti P. Rath, Kate M. Knill, Anton Ragni, Mark J. F. Gales
INTERSPEECH4
2014 Noise-robust TTS speaker adaptation with statistics smoothing
abstract
In practical scenarios for speaker adaptation of speech synthesis systems, the quality of adaptation audio data may be poor. In these situations, it is necessary to make use of the available audio to capture the speaker attributes, whilst aiming to obtain a synthesis voice which does not have any of the lowquality attributes of the audio. One approach to achieving this is to define a sub-space of parametric synthesis parameters in which the adapted system must lie. Though this yields reasonable synthesis quality, target speaker similarity degrades. Quality is also affected in severe noise conditions. This paper describes a smoothing approach that addresses this problem. For a noisy target speaker, first a 'similar speaker' is selected from a database of speakers. Statistics from this speaker are then smoothed with those obtained from the target speaker. By appropriately combining the two sources of information, it is possible to balance similarity and quality. Results indicate that both the quality and similarity can be improved by smoothing, especially for severe noise conditions. The similarity performance, however, varies from speaker to speaker, indicating the importance of a reasonable automatic speaker selection method and the coverage of the candidate speaker pool.
Kayoko Yanagisawa, Langzhou Chen, Mark J. F. Gales
INTERSPEECH3
2014 Paraphrastic language models
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.2
2013 Investigation of multilingual deep neural networks for spoken term detection
abstract
The development of high-performance speech processing systems for low-resource languages is a challenging area. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to use bottleneck features, or hybrid systems, trained on multilingual data for speech-to-text (STT) systems. This paper presents an investigation into the application of these multilingual approaches to spoken term detection. Experiments were run using the IARPA Babel limited language pack corpora (~10 hours/language) with 4 languages for initial multilingual system development and an additional held-out target language. STT gains achieved through using multilingual bottleneck features in a Tandem configuration are shown to also apply to keyword search (KWS). Further improvements in both STT and KWS were observed by incorporating language questions into the Tandem GMM-HMM decision trees for the training set languages. Adapted hybrid systems performed slightly worse on average than the adapted Tandem systems. A language independent acoustic model test on the target language showed that retraining or adapting of the acoustic models to the target language is currently minimally needed to achieve reasonable performance.
Kate M. Knill, Mark J. F. Gales, Shakti P. Rath, Philip C. Woodland, Chao Zhang 0031, Shixiong Zhang 0001
ASRU2
2013 Integrated automatic expression prediction and speech synthesis from text
abstract
Getting a text to speech synthesis (TTS) system to speak lively animated stories like a human is very difficult. To generate expressive speech, the system can be divided into 2 parts: predicting expressive information from text; and synthesizing the speech with a particular expression. Traditionally these blocks have been studied separately. This paper proposes an integrated approach, sharing the expressive synthesis space and training data across the two expressive components. There are several advantages to this approach, including a simplified expression labelling process, support of a continuous expressive synthesis space, and joint training of the expression predictor and speech synthesiser to maximise the likelihood of the TTS system given the training data. Synthesis experiments indicated that the proposed approach generated far more expressive speech than both a neutral TTS and one where the expression was randomly selected. The experimental results also showed the advantage of a continuous expressive synthesis space over a discrete space.
Langzhou Chen, Mark J. F. Gales, Norbert Braunschweiler, Masami Akamine, Kate M. Knill
ICASSP2
2013 Efficient decoding with generative score-spaces using the expectation semiring
abstract
State-of-the-art speech recognisers are usually based on hidden Markov models (HMMs). They model a hidden symbol sequence with a Markov process, with the observations independent given that sequence. These assumptions yield efficient algorithms, but limit the power of the model. An alternative model that allows a wide range of features, including word- and phone-level features, is a log-linear model. To handle, for example, word-level variable-length features, the original feature vectors must be segmented into words. Thus, decoding must find the optimal combination of segmentation of the utterance into words and word sequence. Features must therefore be extracted for each possible segment of audio. For many types of features, this becomes slow. In this paper, long-span features are derived from the likelihoods of word HMMs. Derivatives of the log-likelihoods, which break the Markov assumption, are appended. Previously, decoding with this model took cubic time in the length of the sequence, and longer for higher-order derivatives. This paper shows how to decode in quadratic time.
Rogier C. van Dalen, Anton Ragni, Mark J. F. Gales
ICASSP3
2013 A high-performance Cantonese keyword search system
abstract
We present a system for keyword search on Cantonese conversational telephony audio, collected for the IARPA Babel program, that achieves good performance by combining postings lists produced by diverse speech recognition systems from three different research groups. We describe the keyword search task, the data on which the work was done, four different speech recognition systems, and our approach to system combination for keyword search. We show that the combination of four systems outperforms the best single system by 7%, achieving an actual term-weighted value of 0.517.
Brian Kingsbury, Jia Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP4
2013 Training a supra-segmental parametric F0 model without interpolating F0
abstract
Combining multiple intonation models at different linguistic levels is an effective way to improve the naturalness of the predicted F0. In many of these approaches, the intonation models for suprasegmental levels are based on a parametrization of the log-F0 contours over the units of that level. However, many of these parametrisations are not stable when applied to discontinuous signals. Therefore, the F0 signal has to be interpolated. These interpolated values introduce a distortion in the coefficients that degrades the quality of the model. This paper proposes two methods that eliminate the need for such interpolation, one based on regularization and the other on factor analysis. Subjective evaluations show that, for a Discrete-cosine-transform (DCT) syllable-level model, both approaches result in a significant improvement w.r.t. a baseline using interpolated F0. The approach based on regularization yields the best results.
Javier Latorre, Mark J. F. Gales, Kate M. Knill, Masami Akamine
ICASSP2
2013 Paraphrastic language models and combination with neural network language models
abstract
In natural languages multiple word sequences can represent the same underlying meaning. Only modelling the observed surface word sequence can result in poor context coverage, for example, when using n-gram language models (LM). To handle this issue, paraphrastic LMs were proposed in previous research and successfully applied to a US English conversational telephone speech transcription task. In order to exploit the complementary characteristics of paraphrastic LMs and neural network LMs (NNLM), the combination between the two is investigated in this paper. To investigate paraphrastic LMs' generalization ability to other languages, experiments are conducted on a Mandarin Chinese broadcast speech transcription task. Using a paraphrastic multi-level LM modelling both word and phrase sequences, significant error rate reductions of 0.9% absolute (9% relative) and 0.5% absolute (5% relative) were obtained over the baseline n-gram and NNLM systems respectively, after a combination with word and phrase level NNLMs.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP2
2013 Complex cepstrum analysis based on the minimum mean squared error
abstract
This paper introduces a novel approach for complex cepstrum analysis. Given initial estimates of complex cepstra and respective instants of glottal closure, the method iteratively optimizes the complex cepstrum and instants of glottal closure so that the mean squared error between natural and reconstructed speech waveforms is minimized. The proposed approach results in a more accurate speech representation based on the complex cepstrum, with no need of windowing or phase unwrapping. Experimental results show that the proposed method produces reconstructed speech with higher segmental signal-to-noise ratio scores when compared with conventional methods of complex cepstrum analysis. Because this approach can derive the complex cepstrum at fixed periods, it can be applied to statistical modeling in parametric speech synthesizers.
Ranniery Maia, Masami Akamine, Mark J. F. Gales
ICASSP3
2013 System combination and score normalization for spoken term detection
abstract
Spoken content in languages of emerging importance needs to be searchable to provide access to the underlying information. In this paper, we investigate the problem of extending data fusion methodologies from Information Retrieval for Spoken Term Detection on low-resource languages in the framework of the IARPA Babel program. We describe a number of alternative methods improving keyword search performance. We apply these methods to Cantonese, a language that presents some new issues in terms of reduced resources and shorter query lengths. First, we show score normalization methodology that improves in average by 20% keyword search performance. Second, we show that properly combining the outputs of diverse ASR systems performs 14% better than the best normalized ASR system.
Jonathan Mamou, Jia Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP4
2013 A confidence-based approach for improving keyword hypothesis scores
abstract
The task in keyword spotting (KWS) is to hypothesise times at which any of a set of key terms occurs in audio. An important aspect of such systems are the scores assigned to these hypotheses, the accuracy of which have a significant impact on performance. Estimating these scores may be formulated as a confidence estimation problem, where a measure of confidence is assigned to each key term hypothesis. In this work, a set of discriminative features is defined, and combined using a conditional random field (CRF) model for improved confidence estimation. An extension to this model to directly address the problem of score normalisation across key terms is also introduced. The implicit score normalisation which results from applying this approach to separate systems in a hybrid configuration yields further benefits. Results are presented which show notable improvements in KWS performance using the techniques presented in this work.
Matthew Stephen Seigel, Philip C. Woodland, Mark J. F. Gales
ICASSP3
2013 Tandem system adaptation using multiple linear feature transforms
abstract
Adaptation to speaker and environment changes is an essential part of current automatic speech recognition (ASR) systems. In recent years the use of multi-layer percpetrons (MLPs) has become increasingly common in ASR systems. A standard approach to handling speaker differences when using MLPs is to apply a global speaker-specific constrained MLLR (CMLLR) transform to the features prior to training or using the MLP. This paper considers the situation when there are both speaker and channel, communication link, differences in the data. A more powerful transform, front-end CMLLR (FE-CMLLR), is applied to the inputs to the MLP to represent the channel differences. Though global, these FE-CMLLR transforms vary from time-instance to time-instance. Experiments on a channel distorted dialect Arabic conversational speech recognition task indicates the usefulness of adapting MLP features using both CMLLR and FE-CMLLR transforms.
Yongqiang Wang 0006, Mark J. F. Gales
ICASSP2
2013 Kernelized log linear models for continuous speech recognition
abstract
Large margin criteria and discriminative models are two effective improvements for HMM-based speech recognition. This paper proposed a large margin trained log linear model with kernels for CSR. To avoid explicitly computing in the high dimensional feature space and to achieve the nonlinear decision boundaries, a kernel based training and decoding framework is proposed in this work. To make the system robust to noise a kernel adaptation scheme is also presented. Previous work in this area is extended in two directions. First, most kernels for CSR focus on measuring the similarity between two observation sequences. The proposed joint kernels defined a similarity between two observation-label sequence pairs on the sentence level. Second, this paper addresses how to efficiently employ kernels in large margin training and decoding with lattices. To the best of our knowledge, this is the first attempt at using large margin kernel-based log linear models for CSR. The model is evaluated on a noise corrupted continuous digit task: AURORA 2.0.
Shixiong Zhang 0001, Mark J. F. Gales
ICASSP2
2013 Cross-domain paraphrasing for improving language modelling using out-of-domain data
abstract
In natural languages the variability in the underlying linguistic generation rules significantly alters the observed surface word sequence they create, and thus introduces a mismatch against other data generated via alternative realizations associated with, for example, a different domain. Hence, direct modelling of out-of-domain data can result in poor generalization to the indomain data of interest. To handle this problem, this paper investigated using cross-domain paraphrastic language models to improve in-domain language modelling (LM) using out-ofdomain data. Phrase level paraphrase models learnt from each domain were used to generate paraphrase variants for the data of other domains. These were used to both improve the context coverage of in-domain data, and reduce the domain mismatch of the out-of-domain data. Significant error rate reduction of 0.6% absolute was obtained on a state-of-the-art conversational telephone speech recognition task using a cross-domain paraphrastic multi-level LM trained on a billion words of mixed conversational and broadcast news data. Consistent improvements on the in-domain data context coverage were also obtained. Copyright © 2013 ISCA.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2013 Improving lightly supervised training for broadcast transcription
abstract
This paper investigates improving lightly supervised acoustic model training for an archive of broadcast data. Standard lightly supervised training uses automatically derived decoding hypotheses using a biased language model. However, as the actual speech can deviate significantly from the original programme scripts that are supplied, the quality of standard lightly supervised hypotheses can be poor. To address this issue, word and segment level combination approaches are used between the lightly supervised transcripts and the original programme scripts which yield improved transcriptions. Experimental results show that systems trained using these improved transcriptions consistently outperform those trained using only the original lightly supervised decoding hypotheses. This is shown to be the case for both the maximum likelihood and minimum phone error trained systems.
Yanhua Long, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Matthew Stephen Seigel, Philip C. Woodland
INTERSPEECH2
2013 Minimum mean squared error based warped complex cepstrum analysis for statistical parametric speech synthesis
abstract
This paper presents an approach for complex cepstrum analysis based on the minimum mean squared error criterion, and describes its application to statistical parametric speech synthesis. The proposed method alleviates some of the issues associated with conventional complex cepstrum analysis, such as choice of the window, phase unwrapping, and the need for accurate pitch marks. Given initial estimates of warped complex cepstra and respective analysis instants, the method iteratively optimizes the complex cepstrum on a warped quefrency domain by minimizing the mean squared error between the natural and the reconstructed speech waveforms. When applied to statistical parametric speech synthesis, the optimized complex cepstrum results in better performance in terms of synthesized speech quality, specially for emotional databases, when compared with the complex cepstrum calculated through conventional methods. Copyright © 2013 ISCA.
Ranniery Maia, Mark J. F. Gales, Yannis Stylianou, Masami Akamine
INTERSPEECH2
2013 Photo-realistic expressive text to talking head synthesis
Vincent Wan, Art Blokland, Norbert Braunschweiler, Langzhou Chen, BalaKrishna Kolluru, Javier Latorre, Ranniery Maia, Björn Stenger, Kayoko Yanagisawa, Yannis Stylianou, Masami Akamine, Mark J. F. Gales, Roberto Cipolla
INTERSPEECH13
2013 An explicit independence constraint for factorised adaptation in speech recognition
abstract
Speech signals are usually affected by multiple acoustic factors, such as speaker characteristics and environment differences. Usually, the combined effect of these factors is modelled by a single transform. Acoustic factorisation splits the transform into several factor transforms, each modelling only one factor. This allows, for example, estimating a speaker transform in a noise condition and applying the same speaker transform in a different noise condition. To achieve this factorisation, it is crucial to keep factor transforms independent of each other. Previous work on acoustic factorisation relies on using different forms of factor transforms and/or the attribute of the data to enforce this independence. In this work, the independence is formulated in mathematically, and an explicit constraint is derived to enforce the independence. Using factorised cluster adaptive training (fCAT) as an application, experimental results demonstrates that the proposed explicit independence constraint helps factorisation when imbalanced adaptation data is used. Index Terms: speaker adaption, robust speech recognition, acoustic factorisation
Yongqiang Wang 0006, Mark J. F. Gales
INTERSPEECH2
2013 Infinite support vector machines in speech recognition
abstract
Generative feature spaces provide an elegant way to apply dis-criminative models in speech recognition, and system perfor-mance has been improved by adapting this framework. How-ever, the classes in the feature space may be not linearly sepa-rable. Applying a linear classifier then limits performance. In-stead of a single classifier, this paper applies a mixture of ex-perts. This model trains different classifiers as experts focusing on different regions of the feature space. However, the num-ber of experts is not known in advance. This problem can be bypassed by employing a Bayesian non-parametric model. In this paper, a specific mixture of experts based on the Dirichlet process, namely the infinite support vector machine, is studied. Experiments conducted on the noise-corrupted continuous digit task AURORA 2 show the advantages of this Bayesian non-parametric approach. Index Terms: generative feature space, Bayesian non-parametric, Dirichlet process, mixture of experts, infinite sup-
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH3
2013 Importance sampling to compute likelihoods of noise-corrupted speech
Rogier C. van Dalen, Mark J. F. Gales
Comput. Speech Lang.2
2013 Use of contexts in language model interpolation and adaptation
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.2
2013 Language model cross adaptation for LVCSR system combination
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.2
2013 Complex cepstrum for statistical parametric speech synthesis
Ranniery Maia, Masami Akamine, Mark J. F. Gales
Speech Commun.3
2013 Structured SVMs for Automatic Speech Recognition
abstract
Structured discriminative models are a flexible sequence classification approach that enable a wide variety of features to be used. This paper describes a particular model in this framework, structured support vector machines (SSVM), and how it can be applied to medium to large vocabulary speech recognition tasks. An important aspect of SSVMs is the form of the joint feature spaces. Here, context-dependent generative models, hidden Markov models, are used to obtain the features. To apply this form of combined generative and discriminative model to medium and larger vocabulary tasks, a number of issues need to be addressed. First, the features extracted are a function of the segmentation of the utterance. A Viterbi-like scheme for obtaining the “optimal” segmentation is described. Second, SSVMs can be viewed as large margin log linear models using a zero mean Gaussian prior of the discriminative parameter. However this form of prior is not appropriate for all features. A modified training algorithm is proposed that allows general Gaussian priors to be incorporated into the large margin criterion. Finally to speed up the training process, a 1-slack algorithm, caching competing hypotheses and parallelization strategies are also described. The performance of SSVMs is evaluated on small and medium to large speech recognition tasks: AURORA 2 and 4.
Shixiong Zhang 0001, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2012 Unsupervised clustering of emotion and voice styles for expressive TTS
abstract
Current text-to-speech synthesis (TTS) systems are often perceived as lacking expressiveness, limiting the ability to fully convey information. This paper describes initial investigations into improving expressiveness for statistical speech synthesis systems. Rather than using hand-crafted definitions of expressive classes, an unsupervised clustering approach is described which is scalable to large quantities of training data. To incorporate this “expression cluster” information into an HMM-TTS system two approaches are described: cluster questions in the decision tree construction; and average expression speech synthesis (AESS) using cluster-based linear transform adaptation. The performance of the approaches was evaluated on audiobook data in which the reader exhibits a wide range of expressiveness. A subjective listening test showed that synthesising with AESS results in speech that better reflects the expressiveness of human speech than a baseline expression-independent system.
Florian Eyben, Sabine Buchholz, Norbert Braunschweiler, Javier Latorre, Vincent Wan, Mark J. F. Gales, Kate M. Knill
ICASSP6
2012 Factor analysis based VTS discriminative adaptive training
abstract
Vector Taylor Series (VTS) model based compensation is a powerful approach for noise robust speech recognition. An important extension to this approach is VTS adaptive training (VAT), which allows canonical models to be estimated on diverse noise-degraded training data. These canonical model can be estimated using EM-based approaches, allowing simple extensions to discriminative VAT (DVAT). However to ensure a diagonal corrupted speech covariance matrix the Jacobian (loading matrix) relating the noise and clean speech is diagonalised. In this work an approach for yielding optimal diagonal loading matrices based on minimising the expected KL-divergence between the diagonal loading matrix and “correct” distributions is proposed. The performance of DVAT using the standard and optimal diagonalisation was evaluated on both in-car collected data and the Aurora4 task.
Federico Flego, Mark J. F. Gales
ICASSP2
2012 Complex cepstrum as phase information in statistical parametric speech synthesis
abstract
Statistical parametric synthesizers usually rely on a simplified model of speech production where a minimum-phase filter is driven by a zero or random phase excitation signal. However, this procedure does not take into account the natural mixed-phase characteristics of the speech signal. This paper addresses this issue by proposing the use of the complex cepstrum for modeling phase information in statistical parametric speech synthesizers. Here a frame-based complex cepstrum is calculated through the interpolation of pitch-synchronous magnitude and unwrapped phase spectra. The noncausal part of the frame-based complex cepstrum is then modeled as phase features in the statistical parametric synthesizer. At synthesis time, the generated phase parameters are used to derive coefficients of a glottal filter. Experimental results show that the proposed approach effectively embeds phase information in the synthetic speech, resulting in close-to-natural waveforms and better speech quality.
Ranniery Maia, Masami Akamine, Mark J. F. Gales
ICASSP3
2012 Inference algorithms for generative score-spaces
abstract
Using generative models, for example hidden Markov models (HMM), to derive features for a discriminative classifier has a number of advantages including the ability to make the features robust to speaker and noise changes. An interesting attribute of the derived features is that they may not have the same conditional independence assumptions as the underlying generative models, which are typically first-order Markovian. For efficiency these features are derived given a particular segmentation. This paper describes a general algorithm for obtaining the optimal segmentation with combined generative and discriminative models. Previous results, where the features were constrained to have first-order Markovian dependencies, are extended to allow derivative features to be used which are non-Markovian in nature. As an example, inference with zero and first-order HMM score-spaces is considered. Experimental results are presented on a noise-corrupted continuous digit string recognition task: AURORA 2.
Anton Ragni, Mark J. F. Gales
ICASSP2
2012 Exploring Rich Expressive Information from Audiobook Data Using Cluster Adaptive Training
Langzhou Chen, Mark J. F. Gales, Vincent Wan, Javier Latorre, Masami Akamine
INTERSPEECH2
2012 Model-Based Approaches for Degraded Channel Modelling in Robust ASR
abstract
Speech is usually observed after passing through some form of “channel ” that results in distortions. For some scenarios it is possible to build explicit models of this channel distortion and hence compensate the acoustic models. However the accuracy of the distortion model is sometimes poor and more general adaptation approaches are required. This paper investigates these model-based approaches for communication channel, link, modelling. In particular the paper examines the interaction of link models with speaker adaptation and adaptive training. CMLLR link models with multiple transforms can yield multiple inconsistent feature-spaces When combined with speaker adaptation with very few transforms this inconsistency can limit adaptation performance gains. In contrast using a front-end CMLLR (FE-CMLLR) transform yields a consistent space for speaker adaptation. These schemes are compared on communication channel distorted dialect Arabic conversational speech. Preliminary results on this task indicate the benefits of performing adaptation in a consistent feature-space. Index Terms: acoustic model adaptation, adaptive training. 1.
Mark J. F. Gales, Federico Flego
INTERSPEECH1
2012 Speech factorization for HMM-TTS based on cluster adaptive training
abstract
This paper presents a novel approach to factorize and control different speech factors in HMM-based TTS systems. In this paper cluster adaptive training (CAT) is used to factorize speaker identity and expressiveness (i.e. emotion). Within a CAT framework, each speech factor can be modelled by a different set of clusters. Users can control speaker identity and expressiveness independently by modifying the weights associated with each set. These weights are defined in a continuous space, so variations of speaker and emotion are also continuous. Additionally, given a speaker which has only neutral-style training data, the approach is able to synthesise speech with that speaker’s voice and different expressions. Lastly, the paper discusses how generalization of the basic factorization concept could allow the production of expressive speech from neutral voices for other HMM-TTS systems not based on CAT.
Javier Latorre, Vincent Wan, Mark J. F. Gales, Langzhou Chen, K. K. Chin, Kate M. Knill, Masami Akamine
INTERSPEECH3
2012 Paraphrastic Language Models
abstract
Natural languages are known for their expressive richness.Many sentences can be used to represent the same underlying meaning.Only modelling the observed surface word sequence can result in poor context coverage and generalization, for example, when using n-gram language models (LMs).This paper proposes a novel form of language model, the paraphrastic LM, that addresses these issues.A phrase level paraphrase model statistically learned from standard text data with no semantic annotation is used to generate multiple paraphrase variants.LM probabilities are then estimated by maximizing their marginal probability.Multi-level language models estimated at both the word level and the phrase level are combined.An efficient weighted finite state transducer (WFST) based paraphrase generation approach is also presented.Significant error rate reductions of 0.5%-0.6%absolute were obtained over the baseline n-gram LMs on two state-of-the-art recognition tasks for English conversational telephone speech and Mandarin Chinese broadcast speech using a paraphrastic multi-level LM modelling both word and phrase sequences.When it is further combined with word and phrase level feed-forward neural network LMs, a significant error rate reduction of 0.9% absolute (9% relative) and 0.5% absolute (5% relative) were obtained over the baseline n-gram and neural network LMs respectively.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2012 Rapid Nonlinear Speaker Adaptation for Large-Vocabulary Continuous Speech Recognition
abstract
Recently, kernel eigenvoices were revisited using kernel representations of distributions for rapid nonlinear speaker adaptation. These representations reassure the validity of the adapted distribution functions and enable expectation-maximisation training. Though gains have been shown in terms of word error rate for rapid speaker adaptation, this approach leads to an increase in decoding cost as the number of likelihood evaluations is amplified. The present paper addresses this issue by providing a coherent framework for systematic probabilistic approaches aimed at reducing the recognition cost and yet yielding equally powerful adapted models. The common denominator of such approaches is the use of probabilistic criteria, such as Kullback-Leibler divergence. However, in the general case, the resulting adapted models have full covariance matrices. In order to overcome this issue, the use of predictive semi-tied transforms to yield diagonal covariances for decoding is investigated in this paper. Experimental results are presented on a largevocabulary conversational telephone task. Index Terms: kernel eigenvoices, compact nonlinear adaptation, Kullback Leibler divergence
Zoi Roupakia, Anton Ragni, Mark J. F. Gales
INTERSPEECH3
2012 Combining multiple high quality corpora for improving HMM-TTS
abstract
The most reliable way to build synthetic voices for end-products is to start with high quality recordings from professional voice talents. This paper describes the application of average voice models (AVMs) and a novel application of cluster adaptive training (CAT) to combine a small number of these high quality corpora to make best use of them and improve overall voice quality in hidden Markov model based text-to-speech (HMMTTS) systems. It is shown that integrated training by both CAT and AVM approaches, yields better sounding voices than speaker dependent modelling. It is also shown that CAT has an advantage over AVMs when adapting to a new speaker. Given a limited amount of adaptation data CAT maintains a much higher voice quality even when adapted to tiny amounts of speech.
Vincent Wan, Javier Latorre, K. K. Chin, Langzhou Chen, Mark J. F. Gales, Heiga Zen, Kate M. Knill, Masami Akamine
INTERSPEECH5
2012 Model-based approaches to adaptive training in reverberant environments
abstract
Adaptive training is a powerful approach for building speech recognition systems using non-homogeneous data. This work presents an extension of model-based adaptive training to handle reverberant environments. The recently proposed Reverberant VTS-Joint (RVTSJ) adaptation[1] is used to factor out unwanted additive and reverberant noise variations in multiconditional training data, yielding a canonical model neutral to noise conditions. An maximum likelihood estimation of the canonical model parameters is described. An initialisation scheme that uses the VTS-based adaptive training to initialise the model parameters is also presented. Experiments are conducted on a reverberant simulated AURORA4 task. Index Terms: reverberant noise robustness, vector Taylor series, adaptive training
Yongqiang Wang 0006, Mark J. F. Gales
INTERSPEECH2
2012 Transcription of multi-genre media archives using out-of-domain data
abstract
We describe our work on developing a speech recognition system for multi-genre media archives. The high diversity of the data makes this a challenging recognition task, which may benefit from systems trained on a combination of in-domain and out-of-domain data. Working with tandem HMMs, we present Multi-level Adaptive Networks (MLAN), a novel technique for incorporating information from out-of-domain posterior features using deep neural networks. We show that it provides a substantial reduction in WER over other systems, with relative WER reductions of 15% over a PLP baseline, 9% over in-domain tandem features and 8% over the best out-of-domain tandem features.
Peter Bell 0001, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanhua Long, Steve Renals, Pawel Swietojanski, Philip C. Woodland
SLT2
2012 Morphological decomposition in Arabic ASR systems
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
Comput. Speech Lang.2
2012 Speaker and Noise Factorization for Robust Speech Recognition
abstract
Speech recognition systems need to operate in a wide range of conditions. Thus they should be robust to extrinsic variability caused by various acoustic factors, for example speaker differences, transmission channel and background noise. For many scenarios, multiple factors simultaneously impact the underlying “clean” speech signal. This paper examines techniques to handle both speaker and background noise differences. An acoustic factorization approach is adopted. Here, separate transforms are assigned to represent the speaker [maximum-likelihood linear regression (MLLR)], and noise and channel [model-based vector Taylor series (VTS)] factors. This is a highly flexible framework compared to the standard approaches of modeling the combined impact of both speaker and noise factors. For example factorization allows the speaker characteristics obtained in one noise condition to be applied to a different environment. To obtain this factorization modified versions of MLLR and VTS training and application are derived. The proposed scheme is evaluated for both adaptation and factorization on the AURORA4 data.
Yongqiang Wang 0006, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2012 Statistical Parametric Speech Synthesis Based on Speaker and Language Factorization
abstract
An increasingly common scenario in building speech synthesis and recognition systems is training on inhomogeneous data. This paper proposes a new framework for estimating hidden Markov models on data containing both multiple speakers and multiple languages. The proposed framework, speaker and language factorization, attempts to factorize speaker-/language-specific characteristics in the data and then model them using separate transforms. Language-specific factors in the data are represented by transforms based on cluster mean interpolation with cluster-dependent decision trees. Acoustic variations caused by speaker characteristics are handled by transforms based on constrained maximum-likelihood linear regression. Experimental results on statistical parametric speech synthesis show that the proposed framework enables data from multiple speakers in different languages to be used to: train a synthesis system; synthesize speech in a language using speaker characteristics estimated in a different language; and adapt to a new language.
Heiga Zen, Norbert Braunschweiler, Sabine Buchholz, Mark J. F. Gales, Kate M. Knill, Sacha Krstulovic, Javier Latorre
IEEE Trans. Speech Audio Process.4
2012 Product of Experts for Statistical Parametric Speech Synthesis
abstract
Multiple acoustic models are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of an observation sequence are used as features to be modeled. This paper shows that this combination of multiple acoustic models can be expressed as a product of experts (PoE); the likelihoods from the models are scaled, multiplied together, and then normalized. Normally these models are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the models are jointly trained. A training algorithm for PoEs based on linear feature functions and Gaussian experts is derived by generalizing the training algorithm for trajectory HMMs. However for non-linear feature functions or non-Gaussian experts this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the PoE framework provides both a mathematically elegant way to train multiple acoustic models jointly and significant improvements in the quality of the synthesized speech.
Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda
IEEE Trans. Speech Audio Process.2
2011 A variational perspective on noise-robust speech recognition
abstract
Model compensation methods for noise-robust speech recognition have shown good performance. Predictive linear transformations can approximate these methods to balance computational complexity and compensation accuracy. This paper examines both of these approaches from a variational perspective. Using a matched-pair approximation at the component level yields a number of standard forms of model compensation and predictive linear transformations. However, a tighter bound can be obtained by using variational approximations at the state level. Both model-based and predictive linear transform schemes can be implemented in this framework. Preliminary results show that the tighter bound obtained from the state-level variational approach can yield improved performance over standard schemes.
Rogier C. van Dalen, Mark J. F. Gales
ASRU2
2011 Derivative kernels for noise robust ASR
abstract
Recently there has been interest in combining generative and discriminative classifiers. In these classifiers features for the discriminative models are derived from the generative kernels. One advantage of using generative kernels is that systematic approaches exist to introduce complex dependencies into the feature-space. Furthermore, as the features are based on generative models standard model-based compensation and adaptation techniques can be applied to make discriminative models robust to noise and speaker conditions. This paper extends previous work in this framework in several directions. First, it introduces derivative kernels based on context-dependent generative models. Second, it describes how derivative kernels can be incorporated in structured discriminative models. Third, it addresses the issues associated with large number of classes and parameters when context-dependent models and high-dimensional feature-spaces of derivative kernels are used. The approach is evaluated on two noise-corrupted tasks: small vocabulary AURORA 2 and medium-to-large vocabulary AURORA 4 task.
Anton Ragni, Mark J. F. Gales
ASRU2
2011 Improving reverberant VTS for hands-free robust speech recognition
abstract
Model-based approaches to handling additive background noise and channel distortion, such as Vector Taylor Series (VTS), have been intensively studied and extended in a number of ways. In previous work, VTS has been extended to handle both reverberant and background noise, yielding the Reverberant VTS (RVTS) scheme. In this work, rather than assuming the observation vector is generated by the reverberation of a sequence of background noise corrupted speech vectors, as in RVTS, the observation vector is modelled as a superposition of the background noise and the reverberation of clean speech. This yields a new compensation scheme RVTS Joint (RVTSJ), which allows an easy formulation for joint estimation of both additive and reverberation noise parameters. These two compensation schemes were evaluated and compared on a simulated reverberant noise corrupted AURORA4 task. Both yielded large gains over VTS baseline system, with RVTSJ outperforming the previous RVTS scheme.
Yongqiang Wang 0006, Mark J. F. Gales
ASRU2
2011 Extending noise robust structured support vector machines to larger vocabulary tasks
abstract
This paper describes a structured SVM framework suitable for noise-robust medium/large vocabulary speech recognition. Several theoretical and practical extensions to previous work on small vocabulary tasks are detailed. The joint feature space based on word models is extended to allow context-dependent triphone models to be used. By interpreting the structured SVM as a large margin log-linear model, illustrates that there is an implicit assumption that the prior of the discriminative parameter is a zero mean Gaussian. However, depending on the definition of likelihood feature space, a non-zero prior may be more appropriate. A general Gaussian prior is incorporated into the large margin training criterion in a form that allows the cutting plan algorithm to be directly applied. To further speed up the training process, 1-slack algorithm, caching competing hypothesis and parallelization strategies are also proposed. The performance of structured SVMs is evaluated on noise corrupted medium vocabulary speech recognition task: AURORA 4.
Shixiong Zhang 0001, Mark J. F. Gales
ASRU2
2011 Constrained discriminative mapping transforms for unsupervised speaker adaptation
abstract
Discriminative mapping transforms (DMTs) is an approach to robustly adding discriminative training to unsupervised linear adaptation transforms. In unsupervised adaptation DMTs are more robust to unreliable transcriptions than directly estimating adaptation transforms in a discriminative fashion. They were previously pro posed for use with MLLR transforms with the associated need to explicitly transform the model parameters. In this work the DMT is extended to CMLLR transforms. As these operate in the feature space, it is only necessary to apply a different linear transform at the front-end rather than modifying the model parameters. This is useful for rapidly changing speakers/environments. The performance of DMTs with CMLLR was evaluated on the WSJ 20k task. Experimental results show that DMTs based on constrained linear trans forms yield 3% to 6% relative gain over MLE transforms in unsupervised speaker adaptation.
Langzhou Chen, Mark J. F. Gales, K. K. Chin
ICASSP2
2011 Rapid joint speaker and noise compensation for robust speech recognition
abstract
For speech recognition, mismatches between training and testing for speaker and noise are normally handled separately. The work presented in this paper aims at jointly applying speaker adaptation and model-based noise compensation by embedding speaker adaptation as part of the noise mismatch function. The proposed method gives a faster and more optimum adaptation compared to compensating for these two factors separately. It is also more consistent with respect to the basic assumptions of speaker and noise adaptation. Experimental results show significant and consistent gains from the proposed method.
K. K. Chin, Haitian Xu, Mark J. F. Gales, Catherine Breslin, Kate M. Knill
ICASSP3
2011 Factor analysis based VTS and JUD noise estimation and compensation
abstract
Model based compensation schemes are a powerful approach for noise robust speech recognition. Recently there have been a number of investigations into adaptive training, and estimating the noise models used for model adaptation. This paper examines the use of EM-based schemes for both canonical models and noise estimation, including discriminative adaptive training. One issue that arises when estimating the noise model is a mismatch between the noise estimation approximation and final model compensation scheme. This paper proposes FA-style compensation where this mismatch is eliminated, though at the expense of a sensitivity to the initial noise estimates. EM-based discriminative adaptive training is evaluated on in-car and Aurora4 tasks. FA-style compensation is then evaluated in an incremental mode on the in-car task.
Federico Flego, Mark J. F. Gales
ICASSP2
2011 Continuous F0 in the source-excitation generation for HMM-based TTS: Do we need voiced/unvoiced classification?
abstract
Most HMM-based TTS systems use a hard voiced/unvoiced classification to produce a discontinuous F0 signal which is used for the generation of the source-excitation. When a mixed source excitation is used, this decision can be based on two different sources of information: the state-specific MSD-prior of the F0 models, and/or the frame-specific features generated by the aperiodicity model. This paper examines the meaning of these variables in the synthesis process, their interaction, and how they affect the perceived quality of the generated speech The results of several perceptual experiments show that when using mixed excitation, subjects consistently prefer samples with very few or no false unvoiced errors, whereas a reduction in the rate of false voiced errors does not produce any perceptual improvement. This suggests that rather than using any form of hard voiced/unvoiced classification, e.g., the MSD-prior, it is better for synthesis to use a continuous F0 signal and rely on the frame-level soft voiced/unvoiced decision of the aperiodicity model.
Javier Latorre, Mark J. F. Gales, Sabine Buchholz, Kate M. Knill, Masatsune Tamura, Yamato Ohtani, Masami Akamine
ICASSP2
2011 Investigation of acoustic units for LVCSR systems
abstract
One important issue in designing state-of-the-art LVCSR systems is the choice of acoustic units. Context dependent (CD) phones remain die dominant form of acoustic units. They can capture the co-articulatory effect in speech via explicit modelling. However, for other more complicated phonological processes, they rely on the implicit modelling ability of the underlying statistical models. Alternatively, it is possible to construct acoustic models based on higher level linguistic units, for example, syllables, to explicitly capture these complex patterns. When sufficient training data is available, this approach may show an advantage over implicit acoustic modelling. In this paper a wide range of acoustic units are investigated to improve LVCSR system performance. Significant error rate gains up to 7.1% relative (0.8% abs.) were obtained on a state-of-the-art Mandarin Chinese broadcast audio recognition task using word and syllable position dependent triphone and quinphone models.
Xunying Liu, Mark J. F. Gales, James Hieronymus, Philip C. Woodland
ICASSP2
2011 Structured discriminative models for noise robust continuous speech recognition
abstract
Recently there has been interest in structured discriminative models for speech recognition. In these models sentence posteriors are directly modelled, given a set of features extracted from the observation sequence, and hypothesised word sequence. In previous work these discriminative models have been combined with features derived from generative models for noise-robust speech recognition for continuous digits. This paper extends this work to medium to large vocabulary tasks. The form of the score-space extracted using the generative models, and parameter tying of the discriminative model, are both discussed. Update formulae for both conditional maximum likelihood and minimum Bayes' risk training are described. Experimental results are presented on small and medium to large vocabulary noise-corrupted speech recognition tasks: AURORA 2 and 4.
Anton Ragni, Mark J. F. Gales
ICASSP2
2011 Speaker and noise factorisation on the AURORA4 task
abstract
For many realistic scenarios, there are multiple factors that affect the clean speech signal. In this work approaches to handling two such factors, speaker and background noise differences, simultaneously are described. A new adaptation scheme is proposed. Here the acoustic models are first adapted to the target speaker via an MLLR transform. This is followed by adaptation to the target noise environment via model-based vector Taylor series (VTS) compensation. These speaker and noise transforms are jointly estimated, using maximum likelihood. Experiments on the AURORA4 task demonstrate that this adaptation scheme provides improved performance over VTS-based noise adaptation. In addition, this framework enables the speech and noise to be factorised, allowing the speaker transform estimated in one noise condition to be successfully used in a different noise condition.
Yongqiang Wang 0006, Mark J. F. Gales
ICASSP2
2011 Decision tree-based context clustering based on cross validation and hierarchical priors
abstract
The standard, ad-hoc stopping criteria used in decision tree-based context clustering are known to be sub-optimal and require parameters to be tuned. This paper proposes a new approach for decision tree-based context clustering based on cross validation and hierarchical priors. Combination of cross validation and hierarchical priors within decision tree-based context clustering offers better model selection and more robust parameter estimation than conventional approaches, with no tuning parameters. Experimental results on HMM-based speech synthesis show that the proposed approach achieved significant improvements in naturalness of synthesized speech over the conventional approaches.
Heiga Zen, Mark J. F. Gales
ICASSP2
2011 Integrated Online Speaker Clustering and Adaptation
abstract
For many applications, it is necessary to produce speech transcriptions in a causal fashion. To produce high quality transcripts, speaker adaptation is often used. This requires online speaker clustering and incremental adaptation techniques to be developed. This paper presents an integrated approach to online speaker clustering and adaptation which allows efficient clustering of speakers using the same accumulated statistics that are normally used for adaptation. Using a consistent criterion for both clustering and adaptation should yield gains for both stages. The proposed approach is evaluated on a meetings transcription task using audio from multiple distant microphones. Consistent gains over standard clustering and adaptation were obtained. Copyright © 2011 ISCA.
Catherine Breslin, K. K. Chin, Mark J. F. Gales, Kate M. Knill
INTERSPEECH3
2011 Word Boundary Modelling and Full Covariance Gaussians for Arabic Speech-to-Text Systems
abstract
This paper describes recent improvements to the Cambridge Arabic Large Vocabulary Continuous Speech Recognition (LVCSR) Speech-to-Text (STT) system. It is shown that word-boundary context markers provide a powerful method to enhance graphemic systems by implicit phonetic information, improving the modelling capability of graphemic systems. In addition, a robust technique for full covariance Gaussian modelling in the Minimum Phone Error (MPE) training framework is introduced. This reduces the full covariance training to a diagonal covariance training problem, thereby solving related robustness problems. The full system results show that the combined use of these and other techniques within a multi-branch combination framework reduces the Word Error Rate (WER) of the complete system by up to 5.9 % relative.
Frank Diehl, Mark J. F. Gales, Xunying Liu, Marcus Tomalin, Philip C. Woodland
INTERSPEECH2
2011 Graphone Model Interpolation and Arabic Pronunciation Generation
abstract
This paper extends n-gram graphone model pronunciation generation to use a mixture of such models. This technique is useful when pronunciation data is for a specific variant (or set of variants) of a language, such as for a dialect, and only a small amount of pronunciation dictionary training data for that specific variant is available. The performance of the interpolated n-gram graphone model is evaluated on Arabic phonetic pronunciation generation for words that can't be handled by the Buckwalter Morphological Analyser. The pronunciations produced are also used to train an Arabic broadcast audio speech recognition system. In both cases the interpolated graphone model leads to improved performance. Copyright © 2011 ISCA.
Philip C. Woodland, Frank Diehl, Mark J. F. Gales
INTERSPEECH4
2011 Improving LVCSR System Combination Using Neural Network Language Model Cross Adaptation
abstract
State-of-the-art large vocabulary continuous speech recognition (LVCSR) systems often combine outputs from multiple sub-systems developed at different sites. Cross system adaptation can be used as an alternative to direct hypothesis level combina-tion schemes such as ROVER. The standard approach involves only cross adapting acoustic models. To fully exploit the com-plimentary features among sub-systems, language model (LM) cross adaptation techniques can be used. Previous research on multi-level n-gram LM cross adaptation is extended to further include the cross adaptation of neural network LMs in this pa-per. Using this improved LM cross adaptation framework, sig-nificant error rate gains of 4.0%-7.1 % relative were obtained over acoustic model only cross adaptation when combining a range of Chinese LVCSR sub-systems used in the 2010 and 2011 DARPA GALE evaluations. 1.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2011 Multipulse Sequences for Residual Signal Modeling
abstract
In source-filter models of speech production, the residual signal - what remains after passing the speech signal through the inverse filter - contains important information for the generation of naturally sounding re-synthesized speech. Typically, the voiced regions of residual signals are regarded as a mixture of glottal pulse and noise. This paper introduces a novel approach to represent the noise component of voiced regions of residual signals through autoregressive filtering of multipulse sequences. The positions and amplitudes of the non-zero samples of these multipulse signals are optimized through a closed-loop procedure. The method in question is applied to excitation modeling in statistical parametric synthesis. Experimental results indicate that the use of multipulse-based noise component construction eliminates the necessity of run-time ad hoc procedures such as high-pass filtering and time modulation, common on excitation models for statistical parametric synthesizers, with no loss of synthesized speech quality. Copyright © 2011 ISCA.
Ranniery Maia, Heiga Zen, Kate M. Knill, Mark J. F. Gales, Sabine Buchholz
INTERSPEECH4
2011 Gaussian Process Experts for Voice Conversion
abstract
Conventional approaches to voice conversion typically use a GMM to represent the joint probability density of source and target features. This model is then used to perform spectral conversion between speakers. This approach is reasonably effective but can be prone to overfitting and oversmoothing of the target spectra. This paper proposes an alternative scheme that uses a collection of Gaussian process experts to perform the spectral conversion. Gaussian processes are robust to overfitting and oversmoothing and can predict the target spectra more accurately. Experimental results indicate that the objective performance of voice conversion can be improved using the proposed approach. Copyright © 2011 ISCA.
Nicholas Pilkington, Heiga Zen, Mark J. F. Gales
INTERSPEECH3
2011 Structured Support Vector Machines for Noise Robust Continuous Speech Recognition
abstract
The use of discriminative models is an interesting alternative to generative models for speech recognition. This paper examines one form of these models, structured support vector machines (SVMs), for noise robust speech recognition. One important aspect of structured SVMs is the form of the joint feature space. In this work features based on generative models are used, which allows model-based compensation schemes to be applied to yield robust joint features. However, these features require the segmentation of frames into words, or subwords, to be specified. In previous work this segmentation was obtained using generative models. Here the segmentations are refined using the parameters of the structured SVM. A Viterbilike scheme for obtaining "optimal" segmentations, and modifications to the training algorithm to allow them to be efficiently used, are described. The performance of the approach is evaluated on a noise corrupted continuous digit task: AURORA 2. Copyright © 2011 ISCA.
Shixiong Zhang 0001, Mark J. F. Gales
INTERSPEECH2
2011 The efficient incorporation of MLP features into automatic speech recognition systems
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
Comput. Speech Lang.3
2011 Kernel Eigenvoices (Revisited) for Large-Vocabulary Speech Recognition
abstract
Kernelized eigenvoice methods, which apply a nonlinear transform in speaker space, have previously been proposed for rapid adaptation. This paper examines, and addresses, a number of limitations and issues with the current schemes. First, the requirements for valid probability functions using kernel representations are discussed. Second, rapid speaker adaptation using these forms of representations is analyzed and the general update formulae for kernelized eigenvoice adaptation derived. The existing kernelized eigenvoice methods are then described within this formulation. This allows an EM-based, rather than gradient-descent-based, parameter estimation. To enable these approaches to be applied to large-vocabulary speech recognition tasks, eigenbases using transformations of an underlying canonical model are described and related to existing adaptation methods. Preliminary experiments on a large-vocabulary conversational telephone speech task are finally detailed.
Zoi Roupakia, Mark J. F. Gales
IEEE Signal Process. Lett.2
2011 Extended VTS for Noise-Robust Speech Recognition
abstract
Model compensation is a standard way of improving the robustness of speech recognition systems to noise. A number of popular schemes are based on vector Taylor series (VTS) compensation, which uses a linear approximation to represent the influence of noise on the clean speech. To compensate the dynamic parameters, the continuous time approximation is often used. This approximation uses a point estimate of the gradient, which fails to take into account that dynamic coefficients are a function of a number of consecutive static coefficients. In this paper, the accuracy of dynamic parameter compensation is improved by representing the dynamic features as a linear transformation of a window of static features. A modified version of VTS compensation is applied to the distribution of the window of static features and, importantly, their correlations. These compensated distributions are then transformed to distributions over standard static and dynamic features. With this improved approximation, it is also possible to obtain full-covariance corrupted speech distributions. This addresses the correlation changes that occur in noise. The proposed scheme outperformed the standard VTS scheme by 10% to 20% relative on a range of tasks.
Rogier C. van Dalen, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2011 Noisy Constrained Maximum-Likelihood Linear Regression for Noise-Robust Speech Recognition
abstract
Adaptive training is a widely used technique for building speech recognition systems on nonhomogeneous training data. Recently, there has been interest in applying these approaches for situations where there is significant levels of background noise in the training data. Various schemes for adaptive training are based on noise-, or speaker-, specific transforms of features to yield estimates of the clean speech. However, when there are high levels of background noise, these clean speech estimates may be poor resulting in degradations in performance. In this paper, a new approach for adaptive training on noise-corrupted training data is presented. It extends a popular form of linear transform for model-based adaptation and adaptive training, constrained MLLR (CMLLR), to reflect additional uncertainty from noise-corrupted observations. This new form of adaptation transform is called noisy CMLLR (NCMLLR). NCMLLR uses a modified version of generative model between clean speech and noisy observation, similar to factor analysis (FA). However, in contrast to FA here the generative model describes an adaptation transform, rather than a covariance matrix structure. The use of NCMLLR for adaptive training using an expectation-maximization approach is described. Discriminative adaptive training with NCMLLR is also described based on the minimum phone error criterion. Experimental results comparing NCMLLR with standard adaptive training schemes are given on a noise-corrupted version of Resource Management, the ARPA 1994 CSRNAB Spoke 10 task, and in-car recorded data.
D. K. Kim, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2011 Joint Uncertainty Decoding With Predictive Methods for Noise Robust Speech Recognition
abstract
Model-based noise compensation techniques are a powerful approach to improve speech recognition performance in noisy environments. However, one of the major issues with these schemes is that they are computationally expensive. Though techniques have been proposed to address this problem, they often result in degradations in performance. This paper proposes a new, highly flexible, approach which allows the computational load required for noise compensation to be controlled while maintaining good performance. The scheme applies the improved joint uncertainty decoding with the predictive linear transform framework. The final compensation is implemented as a set of linear transforms of the features, decoupling the computational cost of compensation from the complexity of the recognition system acoustic models. Furthermore, by using linear transforms, changes in the correlations in the feature vector can also be efficiently modeled. The proposed methods can be easily applied in an adaptive training scheme, including discriminative adaptive training. The performance of the approach is compared to a number of standard schemes on Aurora 2 as well as in-car speech recognition tasks. Results indicate that the proposed scheme is an attractive alternative to existing approaches.
Haitian Xu, Mark J. F. Gales, K. K. Chin
IEEE Trans. Speech Audio Process.2
2010 Language model combination and adaptation usingweighted finite state transducers
abstract
In speech recognition systems language model (LMs) are often constructed by training and combining multiple n-gram models. They can be either used to represent different genres or tasks found in diverse text sources, or capture stochastic properties of different linguistic symbol sequences, for example, syllables and words. Unsupervised LM adaptation may also be used to further improve robustness to varying styles or tasks. When using these techniques, extensive software changes are often required. In this paper an alternative and more general approach based on weighted finite state transducers (WFSTs) is investigated for LM combination and adaptation. As it is entirely based on well-defined WFST operations, minimum change to decoding tools is needed. A wide range of LM combination configurations can be flexibly supported. An efficient on-the-fly WFST decoding algorithm is also proposed. Significant error rate gains of 7.3% relative were obtained on a state-of-the-art broadcast audio recognition task using a history dependently adapted multi-level LM modelling both syllable and word sequences.
Xunying Liu, Mark J. F. Gales, James Hieronymus, Philip C. Woodland
ICASSP2
2010 Recent improvements to the Cambridge Arabic Speech-to-Text systems
abstract
This paper describes recent improvements to the Cambridge Arabic Large Vocabulary Continuous Speech Recognition (LVSCR) Speech-to-Text (STT) system. It is shown that Multi-Layer Perceptron (MLP) features trained on phonetic targets can improve the performance of both phonemic and graphemic systems. Also, a morphological decomposition scheme is extended from the graphemic domain to the phonetic domain, and particular attention is given to the task of dictionary generation. Finally, the use of Boosted Maximum Mutual Information (BMMI) training is explored both for individual systems and in the context of system combination. The full system results show that the combined use of the above techniques reduces the Word Error Rate (WER) of the best individual system by up to 12% relative, and that the incorporation of morphological decomposition and BMMI within the four individual branches of the combined system reduces the WER by up to 9% relative.
Marcus Tomalin, Frank Diehl, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2010 Statistical parametric speech synthesis based on product of experts
abstract
Multiple-level acoustic models (AMs) are often combined in statistical parametric speech synthesis. Both linear and non-linear functions of the observation sequence are used as features in these AMs. This combination of multiple-level AMs can be expressed as a product of experts (PoE); the likelihoods from the AMs are scaled, multiplied together and then normalized. Currently these multiple-level AMs are individually trained and only combined at the synthesis stage. This paper discusses a more consistent PoE framework where the AMs are jointly trained. A generalization of trajectory HMM training can be used for multiple-level Gaussian AMs based on linear functions. However for the non-linear case this is not possible, so a scheme based on contrastive divergence learning is described. Experimental results show that the proposed technique provides both a mathematically elegant way to train multiple-level AMs and statistically significant improvements in the quality of synthesized speech.
Heiga Zen, Mark J. F. Gales, Yoshihiko Nankaku, Keiichi Tokuda
ICASSP2
2010 Lightly supervised recognition for automatic alignment of large coherent speech recordings
Norbert Braunschweiler, Mark J. F. Gales, Sabine Buchholz
INTERSPEECH2
2010 Prior information for rapid speaker adaptation
abstract
Rapidly adapting a speech recognition system to new speakers using a small amount of adaptation data is important to improve initial user experience. In this paper, a count-smoothing framework for incorporating prior information is extended to allow for the use of different forms of dynamic prior and improve the robustness of transform estimation on small amounts of data. Prior information is obtained from existing rapid adaptation techniques like VTLN and PCMLLR. Results using VTLN as a dynamic prior for CMLLR estimation show that transforms estimated on just one utterance can yield relative gains of 15% and 46% over a baseline gender independent model on two tasks. Index Terms: automatic speech recognition, speaker adaptation, VTLN, prior knowledge
Catherine Breslin, K. K. Chin, Mark J. F. Gales, Kate M. Knill, Haitian Xu
INTERSPEECH3
2010 Asymptotically exact noise-corrupted speech likelihoods
abstract
Model compensation techniques for noise-robust speech recognition approximate the corrupted speech distribution. This paper introduces a sampling method that, given speech and noise distributions and a mismatch function, in the limit calculates the corrupted speech likelihood exactly. Though it is too slow to compensate a speech recognition system, it enables a more fine-grained assessment of compensation techniques, based on the KL divergence of individual components. This makes it possible to evaluate the impact of approximations that compensation schemes make, such as the form of the mismatch function. Index Terms: speech recognition, noise robustness 1.
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH2
2010 Canonical state models for automatic speech recognition
abstract
Current speech recognition systems are often based on HMMs with state-clustered Gaussian Mixture Models (GMMs) to represent the context dependent output distributions. Though highly successful, the standard form of model does not exploit any relationships between the states, they each have separate model parameters. This paper describes a general class of model where the context-dependent state parameters are a transformed version of one, or more, canonical states. A number of published models sit within this framework, including, semi-continuous HMMs, subspace GMMs and the HMM error model. A set of preliminary experiments illustrating some of this model’s properties using CMLLR transformations from the canonical state to the context dependent state are described. Index Terms: acoustic modelling, adaptive training, Gaussian mixture models.
Mark J. F. Gales, Kai Yu 0004
INTERSPEECH1
2010 Training a parametric-based logF0 model with the minimum generation error criterion
abstract
This paper describes an approach for improving a statistical parametric-based logF0 model using minimum-generationerror (MGE) training. Compared with the previous scheme based on decision tree clustering, MGE allows the minimisation of the error in the generated logF0 to take into account not only each cluster by itself, but also the way in which the clusters interact with each other in the generation of the F0 over the whole sentence. Moreover, the “weights” of each component of the model, which previously were adjusted manually, are optimized automatically by the MGE training during the re-estimation of the model covariances. Objective evaluation indicated that, although the logF0 contours generated by the models trained with MGE have approximately the same root mean square error and correlation factor as those generated with the baseline models, they present a higher dynamic range. The subjective evaluation shows a small but significant preference for the system trained with MGE.
Javier Latorre, Mark J. F. Gales, Heiga Zen
INTERSPEECH2
2010 Language model cross adaptation for LVCSR system combination
abstract
State-of-the-art large vocabulary continuous speech recognition (LVCSR) systems often combine outputs from multiple sub-systems that may even be developed at different sites. Cross system adaptation, in which model adaptation is performed using the outputs from another sub-system, can be used as an alternative to hypothesis level combination schemes such as ROVER. Normally cross adaptation is only performed on the acoustic models. However, there are many other levels in LVCSR systems' modelling hierarchy where complimentary features may be exploited, for example, the sub-word and the word level, to further improve cross adaptation based system combination. It is thus interesting to also cross adapt language models (LMs) to capture these additional useful features. In this paper cross adaptation is applied to three forms of language models, a multi-level LM that models both syllable and word sequences, a word level neural network LM, and the linear combination of the two. Significant error rate reductions of 4.0-7.1% relative were obtained over ROVER and acoustic model only cross adaptation when combining a range of Chinese LVCSR sub-systems used in the 2010 and 2011 DARPA GALE evaluations.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2010 Improved neural network based language modelling and adaptation
abstract
Neural network language models (NNLM) have become an increasingly popular choice for large vocabulary continuous speech recognition (LVCSR) tasks, due to their inherent gener-alisation and discriminative power. This paper present two tech-niques to improve performance of standard NNLMs. First, the form of NNLM is modelled by introduction an additional out-put layer node to model the probability mass of out-of-shortlist (OOS) words. An associated probability normalisation scheme is explicitly derived. Second, a novel NNLM adaptation method using a cascaded network is proposed. Consistent WER reduc-tions were obtained on a state-of-the-art Arabic LVCSR task over conventional NNLMs. Further performance gains were also observed after NNLM adaptation.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2010 Discriminative classifiers with adaptive kernels for noise robust speech recognition
Mark J. F. Gales, Federico Flego
Comput. Speech Lang.1
2010 Unsupervised training and directed manual transcription for LVCSR
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
Speech Commun.2
2010 Structured Log Linear Models for Noise Robust Speech Recognition
abstract
The use of discriminative models for structured classification tasks, such as speech recognition is becoming increasingly popular. This letter examines the use of structured log-linear models for noise robust speech recognition. An important aspect of log-linear models is the form of the features. By using generative models to derive the features, state-of-the-art model-based compensation schemes can be used to make the system robust to noise. Previous work in this area is extended in two important directions. First, a large margin training of sentence-level log linear models is proposed for automatic speech recognition (ASR). This form of model is shown to be similar to the recently proposed structured Support Vector Machines (SVM). Second, based on the designed joint features, efficient lattice-based training and decoding are performed. This novel model combines generative kernels, discriminative models, efficient lattice-based large margin training and model-based noise compensation. It is evaluated on a noise corrupted continuous digit task: AURORA 2.0.
Shixiong Zhang 0001, Anton Ragni, Mark J. F. Gales
IEEE Signal Process. Lett.3
2009 Discriminative adaptive training with VTS and JUD
abstract
Adaptive training is a powerful approach for building speech recognition systems on non-homogeneous training data. Recently approaches based on predictive model-based compensation schemes, such as joint uncertainty decoding (JUD) and vector Taylor series (VTS), have been proposed. This paper reviews these model-based compensation schemes and relates them to factor-analysis style systems. Forms of maximum likelihood (ML) adaptive training with these approaches are described, based on both second-order optimisation schemes and expectation maximisation (EM). However, discriminative training is used in many state-of-the-art speech recognition. Hence, this paper proposes discriminative adaptive training with predictive model-compensation approaches for noise robust speech recognition. This training approach is applied to both JUD and VTS compensation with minimum phone error training. A large scale multi-environment training configuration is used and the systems evaluated on a range of in-car collected data tasks.
Federico Flego, Mark J. F. Gales
ASRU2
2009 Acoustic modelling for speech recognition: Hidden Markov models and beyond?
abstract
Hidden Markov models (HMMs) are still the dominant form of acoustic model used in automatic speech recognition (ASR) systems. However over the years the form, and training, of the HMM for ASR have been extended and modified, so that the current forms used in state-of-the-art speech recognition systems are very different to those originally proposed thirty years ago. This talk will review two of the more important extensions that have been proposed over the years: discriminative training; and speaker and environment adaptation. The use of discriminative training is now common with forms based on minimum Bayes' training and minimum classification error being applied to systems trained on many hundreds of hours of speech data. The talk will describe these current approaches, as well as discussing the current trends towards schemes based on large-margin training approaches. Linear transform based speaker adaptation is the dominant form for speaker adaptation. Current approaches, including extensions to linear transforms and model-based noise robustness techniques, and trends will also be described. Details of the various forms of the adaptation/noise transformation, training criterion and approaches for adaptive training will be given. The final part of the talk will discuss research beyond the current HMM framework. Schemes based on both discriminative models and functions, as well as non-parametric approaches will be described.
Mark J. F. Gales
ASRU1
2009 Support vector machines for noise robust ASR
abstract
Using discriminative classifiers, such as Support Vector Machines (SVMs) in combination with, or as an alternative to, Hidden Markov Models (HMMs) has a number of advantages for difficult speech recognition tasks. For example, the models can make use of additional dependencies in the observation sequences than HMMs provided the appropriate form of kernel is used. However standard SVMs are binary classifiers, and speech is a multi-class problem. Furthermore, to train SVMs to distinguish word pairs requires that each word appears in the training data. This paper examines both of these limitations. Tree-based reduction approaches for multiclass classification are described, as well as some of the issues in applying them to dynamic data, such as speech. To address the training data issues, a simplified version of HMM-based synthesis can be used, which allows data for any word-pair to be generated. These approaches are evaluated on two noise corrupted digit sequence tasks: AURORA 2.0; and actual in-car collected data.
Mark J. F. Gales, Anton Ragni, H. AlDamarki, C. Gautier
ASRU1
2009 Improving joint uncertainty decoding performance by predictive methods for noise robust speech recognition
abstract
Model-based noise compensation techniques, such as vector Taylor series (VTS) compensation, have been applied to a range of noise robustness tasks. However one of the issues with these forms of approach is that for large speech recognition systems they are computationally expensive. To address this problem schemes such as Joint Uncertainty Decoding (JUD) have been proposed. Though computationally more efficient, the performance of the system is typically degraded. This paper proposes an alternative scheme, related to JUD, but making fewer approximations, VTS-JUD. Unfortunately this approach also removes some of the computational advantages of JUD. To address this, rather than using VTS-JUD directly, it is used instead to obtain statistics to estimate a predictive linear transform, PCMLLR. This is both computationally efficient and limits some of the issues associated with the diagonal covariance matrices typically used with schemes such as VTS. PCMLLR can also be simply used within an adaptive training framework (PAT). The performance of the VTS-JUD, PCMLLR and PAT system were compared to a number of standard approaches on an in-car speech recognition task. The proposed scheme is an attractive alternative to existing approaches.
Haitian Xu, Mark J. F. Gales, K. K. Chin
ASRU2
2009 Extended VTS for noise-robust speech recognition
abstract
Model compensation is a standard way of improving speech recognisers' robustness to noise. Currently popular schemes are based on vector Taylor series (VTS) compensation. They often use the continuous time approximation to compensate dynamic parameters. In this paper, the accuracy of dynamic parameter compensation is improved by representing the dynamic features as a linear transformation of a window of static features. A modified version of VTS compensation is applied to the distribution of the window of static features and, importantly, their correlations. These compensated distributions are then transformed to standard static and dynamic distributions. The proposed scheme outperformed the standard VTS scheme by about 10% relative.
Rogier C. van Dalen, Mark J. F. Gales
ICASSP2
2009 Incremental predictive and adaptive noise compensation
abstract
Model compensation schemes are a powerful approach to handling mismatches between training and testing conditions. Normally these schemes are run in a batch adaptation mode, re-recognising the utterance used to estimate the noise model parameters. For many applications this introduces unacceptable latency. This paper examines three forms of incremental mode model-based compensation: vector Taylor series; joint uncertainty decoding; and predictive CMLLR. These predictive schemes can also be combined with adaptive schemes such as CMLLR. By combining the approaches, weaknesses of each can be addressed. The performance is evaluated on in-car recorded data, where the combined incremental scheme shows gains over either individually.
Federico Flego, Mark J. F. Gales
ICASSP2
2009 Combining VTS model compensation and support vector machines
abstract
It is difficult to adapt discriminative classifiers, particularly kernel based ones such as support vector machines (SVMs), to handle mismatches between the training and test data. In previous work adaptation was performed by modifying the kernel used with the SVM, rather changing the SVM parameters themselves. However an idealised form of compensation, single pass retraining, was used to alter the generative models associated with the generative kernel. In this paper vector Taylor series model compensation is used. This scheme is more efficient and allows a noise model to be estimated. The performance of the new scheme is evaluated on two continuous digit tasks. On both tasks SVM-rescoring outperformed the baseline VTS compensated models.
Mark J. F. Gales, Federico Flego
ICASSP1
2009 Training and adapting MLP features for Arabic speech recognition
abstract
Features derived from multilayer perceptrons (MLPs) are becoming increasingly popular for speech recognition. This paper describes various schemes for applying these features to state-of-the-art Arabic speech recognition: the use of MLP-features for short-vowel modelling in graphemic systems; rapid discriminative model training by standard PLP feature lattice reuse; and MLP feature adaptation using linear input networks (LIN). The use of rapid training using MLP features and their use for short-vowel modelling and LIN adaptation gave reductions in word error rate. However significant improvements over explicit short-vowel modelling with standard multi-pass adaptation were not obtained, although they were useful in combination.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
ICASSP3
2009 Bayesian discriminative adaptation for speech recognition
abstract
Linear transform-based speaker adaptation is a standard part of many speech recognition systems. For unsupervised adaptation maximum likelihood estimation is typically used, as discriminative transforms are more heavily biased towards the supervision hypothesis which may contain errors. In this work, a Bayesian framework for discriminative adaptation is investigated. This reduces the hypothesis bias and allows robust estimates even with a limited amount of data. Various forms of discriminative maximum-a-posteriori estimation, and associated issues, are detailed. To address these problems, the use of discriminative mapping transforms is also described. The proposed framework is evaluated on an English conversational speech task.
Chandra Kant Raut, Mark J. F. Gales
ICASSP2
2009 Transforming features to compensate speech recogniser models for noise
abstract
To make speech recognisers robust to noise, either the features or the models can be compensated. Feature enhancement is often fast; model compensation is often more accurate, because it predicts the corrupted speech distribution. It is therefore able, for example, to take uncertainty about the clean speech into account. This paper re-analyses the recently-proposed predictive linear transformations for noise compensation as minimising the KL divergence between the predicted corrupted speech and the adapted models. New schemes are then introduced which apply observation-dependent transformations in the front-end to adapt the back-end distributions. One applies transforms in the exact same manner as the popular minimum mean square error (MMSE) feature enhancement scheme, and is as fast. The new method performs better on AURORA 2. Index Terms: speech recognition, noise robustness 1.
Rogier C. van Dalen, Federico Flego, Mark J. F. Gales
INTERSPEECH3
2009 Morphological analysis and decomposition for Arabic speech-to-text systems
abstract
Language modelling for a morphologically complex language such as Arabic is a challenging task. Its agglutinative structure re-sults in data sparsity problems and high out-of-vocabulary rates. In this work these problems are tackled by applying the MADA tools to the Arabic text. In addition to morphological decompo-sition, MADA performs context-dependent stem-normalisation. Thus, if word-level system combination, or scoring, is required this normalisation must be reversed. To address this, a novel context-sensitive method for morpheme-to-word conversion is introduced. The performance of the MADA decomposed sys-tem was evaluated on an Arabic broadcast transcription task. The MADA-based system out-performed the word-based system, with both the morphological decomposition and stem normalisa-tion being found to be important.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
INTERSPEECH2
2009 Incremental adaptation with VTS and joint adaptively trained systems
abstract
Recently adaptive training schemes using model based compensation approaches such as VTS and JUD have been proposed. Adaptive training allows the use of multi-environment training data whilst training a neutral, “clean”, acoustic model to be trained. This paper describes and assesses the advantages of using incremental, rather than batch, mode adaptation with these adaptively trained systems. Incremental adaptation reduces the latency during recognition, and has the possibility of reducing the error rate for slowly varying noise. The work is evaluated on a large scale multi-environment training configuration targeted at in-car speech recognition. Results on in-car collected test data indicate that incremental adaptation is an attractive option when using these adaptively trained systems. Index Terms: adaptive training, incremental adaptation, noise compensation
Federico Flego, Mark J. F. Gales
INTERSPEECH2
2009 Exploiting Chinese character models to improve speech recognition performance
abstract
The Chinese language is based on characters which are syllabic in nature. Since languages have syllabotactic rules which govern the construction of syllables and their allowed sequences, Chinese character sequence models can be used as a first level approximation of allowed syllable sequences. N-gram character sequence models were trained on 4.3 billion characters. Characters are used as a first level recognition unit with multiple pronunciations per character. For comparison the CU-HTK Mandarin word based system was used to recognize words which were then converted to character sequences. The character only system error rates for one best recognition were slightly worse than word based character recognition. However combining the two systems using log-linear combination gives better results than either system separately. An equally weighted combination gave consistent CER gains of 0.1-0.2% absolute over the word based standard system. Copyright © 2009 ISCA.
James Hieronymus, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2009 Adaptive training with noisy constrained maximum likelihood linear regression for noise robust speech recognition
D. K. Kim, Mark J. F. Gales
INTERSPEECH2
2009 Use of contexts in language model interpolation and adaptation
abstract
Language models (LMs) are often constructed by building multiple individual component models that are combined using context independent interpolation weights. By tuning these weights, using either perplexity or discriminative approaches, it is possible to adapt LMs to a particular task. This paper investigates the use of context dependent weighting in both interpolation and test-time adaptation of language models. Depending on the previous word contexts, a discrete history weighting function is used to adjust the contribution from each component model. As this dramatically increases the number of parameters to estimate, robust weight estimation schemes are required. Several approaches are described in this paper. The first approach is based on MAP estimation where interpolation weights of lower order contexts are used as smoothing priors. The second approach uses training data to ensure robust estimation of LM interpolation weights. This can also serve as a smoothing prior for MAP adaptation. A normalized perplexity metric is proposed to handle the bias of the standard perplexity criterion to corpus size. A range of schemes to combine weight information obtained from training data and test data hypotheses are also proposed to improve robustness during context dependent LM adaptation. In addition, a minimum Bayes' risk (MBR) based discriminative training scheme is also proposed. An efficient weighted finite state transducer (WFST) decoding algorithm for context dependent interpolation is also presented. The proposed technique was evaluated using a state-of-the-art Mandarin Chinese broadcast speech transcription task. Character error rate (CER) reductions up to 7.3 relative were obtained as well as consistent perplexity improvements. © 2012 Elsevier Ltd. All rights reserved.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2009 Variational dynamic kernels for speaker verification
abstract
An important aspect of SVM-based speaker verification is the choice of dynamic kernel. Recently there has been interest in the use of kernels based on the Kullback-Leibler divergence between GMMs. Since this has no closed-form solution, typically a matched-pair upper bound is used instead. This places significant restrictions on the forms of model structure that may be used. All GMMs must contain the same number of components and must be adapted from a single background model. For many tasks this will not be optimal. In this paper, dynamic kernels are proposed based on alternative, variational approximations to the KL divergence. Unlike the matched-pair bound, these do not restrict the forms of GMM that may be used. Additionally, using a more accurate approximation of the divergence may lead to performance gains. Preliminary results using these kernels are presented on the NIST 2002 SRE dataset.
Chris Longworth, Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH3
2009 Efficient generation and use of MLP features for Arabic speech recognition
abstract
Front-end features computed using Multi-Layer Perceptrons (MLPs) have recently attracted much interest, but are a challenge to scale to large networks and very large training data sets. This paper discusses methods to reduce the training time for the generation of MLP features and their use in an ASR system using a variety of techniques: parallel training of a set of MLPs on different data sub-sets; methods for computing features from by a combination of these networks; and rapid discriminative training of HMMs using MLP-based features. The impact on MLP frame-based accuracy using different training strategies is discussed along with the effect on word rates from incorporating the MLP features in various configurations into an Arabic broadcast audio transcription system.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
INTERSPEECH3
2009 Directed decision trees for generating complementary systems
Catherine Breslin, Mark J. F. Gales
Speech Commun.2
2009 Combining Derivative and Parametric Kernels for Speaker Verification
abstract
Support vector machine-based speaker verification (SV) has become a standard approach in recent years. These systems typically use dynamic kernels to handle the dynamic nature of the speech utterances. This paper shows that many of these kernels fall into one of two general classes, derivative and parametric kernels. The attributes of these classes are contrasted and the conditions under which the two forms of kernel are identical are described. By avoiding these conditions, gains may be obtained by combining derivative and parametric kernels. One combination strategy is to combine at the kernel level. This paper describes a maximum-margin-based scheme for learning kernel weights for the SV task. Various dynamic kernels and combinations were evaluated on the NIST 2002 SRE task, including derivative and parametric kernels based upon different model structures. The best overall performance was 7.78% EER achieved when combining five kernels.
Chris Longworth, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2009 Unsupervised Adaptation With Discriminative Mapping Transforms
abstract
The most commonly used approaches to speaker adaptation are based on linear transforms, as these can be robustly estimated using limited adaptation data. Although significant gains can be obtained using discriminative criteria for training acoustic models, maximum-likelihood (ML) estimated transforms are still used for unsupervised adaptation. This is because discriminatively trained transforms are highly sensitive to errors in the adaptation supervision hypothesis. This paper describes a new framework for estimating transforms that are discriminative in nature, but are less sensitive to this hypothesis issue. A speaker-independent discriminative mapping transformation (DMT) is estimated during training. This transform is obtained after a speaker-specific ML-estimated transform of each training speaker has been applied. During recognition an ML speaker-specific transform is found for each test-set speaker and the speaker-independent DMT then applied. This allows a transform which is discriminative in nature to be indirectly estimated, while only requiring an ML speaker-specific transform to be found during recognition. The DMT technique is evaluated on an English conversational telephone speech task. Experiments showed that using DMT in unsupervised adaptation led to significant gains over both standard ML and discriminatively trained transforms.
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
IEEE Trans. Speech Audio Process.2
2008 Phonetic pronunciations for arabic speech-to-text systems
abstract
In this paper two aspects of generating and using phonetic arabic dictionaries are described. First, the use of single pronunciation acoustic models in the context of arabic large vocabulary automatic speech recognition (ASR) is investigated. These have been found to be useful for English ASR systems, when combined with standard multiple pronunciation systems. The second area examined is automatically deriving phonetic "pronunciations" for words that standard approaches, such as the Buckwalter morphological analyzer, cannot handle. Without pronunciations for these words the OOV rates for various Arabic tasks significantly increase. Here, pronunciations are automatically found by first deriving grapheme-to-phone rules, and associated rule probabilities. These are then used to produce the most likely pronunciation, or pronunciations, for any word. These approaches are evaluated on a large vocabulary arabic broadcast news and broadcast conversation transcription task. Both schemes are found to yield gains with a multi-pass/combination framework.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
ICASSP2
2008 Multiple kernel learning for speaker verification
abstract
Many speaker verification (SV) systems combine multiple classifiers using score-fusion to improve system performance. For SVM classifiers, an alternative strategy is to combine at the kernel level. This involves finding a suitable kernel weighting, known as multiple kernel learning (MKL). Recently, an efficient maximum-margin scheme for MKL has been proposed. This work examines several refinements to this scheme for SV. The standard scheme has a known tendency towards sparse weightings, which may not be optimal for SV. A regularisation term is proposed, allowing the appropriate level of sparsity to be selected. Cross-speaker tying of kernel weights is also applied to improve robustness. Various combinations of dynamic kernels were evaluated, including derivative and parametric kernels based upon different model structures. The performance achieved on the NIST 2002 SRE when combining five kernels was 4.83% EER.
Chris Longworth, Mark J. F. Gales
ICASSP2
2008 Unsupervised discriminative adaptation using discriminative mapping transforms
abstract
The most commonly used approaches to speaker adaptation are based on linear transforms, as these can be robustly estimated using limited adaptation data. Although significant gains can be obtained using discriminative criteria for training acoustic models, maximum likelihood (ML) estimated transforms are used for unsupervised adaptation. This is because discriminatively trained transforms are highly sensitive to errors in the adaptation hypothesis. This paper describes a new framework for estimating transforms that are discriminative in nature, but are less sensitive to this hypothesis issue. A discriminative, speaker-independent, mapping transformation is estimated during training. This transform is obtained after a speaker-specific ML-estimated transform has been applied. During recognition an ML speaker-specific transform is found and the speaker-independent discriminative mapping transform then applied. This allows a transform which is discriminative in nature to be indirectly estimated, whilst only requiring an ML speaker-specific transform to be found during recognition. The scheme is evaluated on an English conversational telephone speech task, where it significantly outperforms both standard ML and discriminatively trained transforms.
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
ICASSP2
2008 Covariance modelling for noise-robust speech recognition
abstract
Model compensation is a standard way of improving speech recognisers’ robustness to noise. Most model compensation techniques produce diagonal covariances. However, this fails to handle any changes in the feature correlations due to the noise. This paper presents a scheme that allows full-covariance matrices to be estimated. One problem is that full covariance matrix estimation will be more sensitive approximations, those for the dynamic parameters are known to crude. In this paper a linear transformation of a window of consecutive frames is used as the basis for dynamic parameter compensation. A second problem is that the resulting full covariance matrices slow down decoding. This is addressed by using predictive linear transforms that decorrelate the feature space, so that the decoder can then use diagonal covariance matrices. On a noise-corrupted Resource Management task, the proposed scheme outperformed the standard VTS compensation scheme.
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH2
2008 Discriminative classifiers with generative kernels for noise robust ASR
abstract
Discriminative classifiers are a popular approach to solving classification problems. However one of the problems with these approaches, in particular kernel based classifiers such as Support Vector Machines (SVMs), is that they are hard to adapt to mismatches between the training and test data. This paper describes a scheme for overcoming this problem for speech recognition in noise by adapting the kernel rather than the SVM decision boundary. Generative kernels, defined using generative models, are one type of kernel that allows SVMs to handle sequence data. By compensating the parameters of the generative models for each noise condition noise-specific generative kernels can be obtained. These can be used to train a noise-independent SVM on a range of noise conditions, which can then be used with a test-set noise kernel for classification. The noise-specific kernels used in this paper are based on Vector Taylor Series (VTS) model-based compensation. VTS allows all the model parameters to be compensated and the background noise to be estimated in a maximum likelihood fashion. A brief discussion of VTS and the optimisation of the mismatch function representing the impact of noise on the clean speech, is also included. Experiments using these VTS-based test-set noise kernels were run on the AURORA 2 continuous digit task. The proposed SVM rescoring scheme yields large gains in performance over the VTS compensated models. 1
Mark J. F. Gales, Chris Longworth
INTERSPEECH1
2008 Context dependent language model adaptation
abstract
Language models (LMs) are often constructed by building multiple component LMs that are combined using interpolation weights. By tuning these interpolation weights, using either perplexity or discriminative approaches, it is possible to adapt LMs to a particular task. In this work, improved LM adaptation is achieved by introducing context dependent interpolation weights. An important part of this new approach is obtaining robust estimation. Two schemes for this are described. The first is based on MAP estimation, where either global interpolation weights are used as priors, or context dependent interpolation priors obtained from the training data. The second scheme uses class based contexts to determine the interpolation weights. Both schemes are evaluated using unsupervised LM adaptation on a Mandarin broadcast transcription task. Consistent gains in perplexity using context dependent, rather than global, weights are observed as well as reductions in character error rate. 1.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2008 A generalised derivative kernel for speaker verification
abstract
An important aspect of SVM-based speaker verification systems is the choice of dynamic kernel. For the GLDS kernel, a static kernel is used to map each observation into a higher order feature space. Features are then obtained by taking a simple average over all frames. Derivative kernels, such as the Fisher kernel, use a generative model as a principled way of extracting a fixed set of features from each utterance. However, the model and features are defined using the original observations. Here, a dynamic kernel is described that combines these two approaches. In general, it is not possible to explicitly train a model in the feature space associated with a static kernel. However, by using a suitable metric with approximate component posteriors, this form of dynamic kernel can be computed. This kernel generalises the GLDS and derivative kernel as special cases and is also closely related to parametric kernels such as the GMMsupervector kernel. Preliminary results using this kernel are presented on the 2002 NIST SRE dataset.
Chris Longworth, Mark J. F. Gales
INTERSPEECH2
2008 Adaptive training using discriminative mapping transforms
abstract
Speaker adaptive training (SAT) is a useful technique for building speech recognition systems on non-homogeneous data. When combining SAT with discriminative training criteria, maximum likelihood (ML) transforms are often used for unsupervised adaptation tasks. This is because discriminatively estimated transforms are highly sensitive to errors in the supervision hypothesis. In this paper, speaker adaptive training based on discriminative mapping transforms (DMTs) is proposed. DMTs are speaker-independent discriminative transforms that are applied to ML-estimated speaker-specific transforms. As DMTs are estimated during training, they are not affected by errors in the supervision hypothesis. The proposed method was evaluated on an English conversational telephone speech task. It was found to significantly outperform the standard discriminative SAT schemes. Index Terms: speech recognition, speaker adaptive training, discriminative training and adaptation
Chandra Kant Raut, Kai Yu 0004, Mark J. F. Gales
INTERSPEECH3
2008 Issues with uncertainty decoding for noise robust automatic speech recognition
Hank Liao, Mark J. F. Gales
Speech Commun.2
2007 Predictive linear transforms for noise robust speech recognition
abstract
It is well known that the addition of background noise alters the correlations between the elements of, for example, the MFCC feature vector. However, standard model-based compensation techniques do not modify the feature-space in which the diagonal covariance matrix Gaussian mixture models are estimated. One solution to this problem, which yields good performance, is joint uncertainty decoding (JUD) with full transforms. Unfortunately, this results in a high computational cost during decoding. This paper contrasts two approaches to approximating full JUD while lowering the computational cost. Both use predictive linear transforms to modify the feature-space: adaptation-based linear transforms, where the model parameters are restricted to be the same as the original clean system; and precision matrix modelling approaches, in particular semi-tied covariance matrices. These predictive transforms are estimated using statistics derived from the full JUD transforms rather than noisy data. The schemes are evaluated on AURORA 2 and a noise-corrupted resource management task.
Mark J. F. Gales, Rogier C. van Dalen
ASRU1
2007 Development of a phonetic system for large vocabulary Arabic speech recognition
abstract
This paper describes the development of an Arabic speech recognition system based on a phonetic dictionary. Though phonetic systems have been previously investigated, this paper makes a number of contributions to the understanding of how to build these systems, as well as describing a complete Arabic speech recognition system. The first issue considered is discriminative training when there are a large number of pronunciation variants for each word. In particular, the loss function associated with Minimum Phone Error (MPE) training is examined. The performance and combination of phonetic and graphemic acoustic models are then compared on both Broadcast News (BN) and Broadcast Conversation (BC) data. The final contribution of the paper is a simple scheme for automatically generating pronunciations for use in training and reducing the phonetic out-of-vocabulary rate. The paper concludes with a description and results from using phonetic and graphemic systems in a multipass/ combination framework.
Mark J. F. Gales, Frank Diehl, Chandra Kant Raut, Marcus Tomalin, Philip C. Woodland, Kai Yu 0004
ASRU1
2007 Discriminative language model adaptation for Mandarin broadcast speech transcription and translation
abstract
This paper investigates unsupervised test-time adaptation of language models (LM) using discriminative methods for a Mandarin broadcast speech transcription and translation task. A standard approach to adapt interpolated language models to is to optimize the component weights byminimizing the perplexity on supervision data. This is a widely made approximation for language modeling in automatic speech recognition (ASR) systems. For speech translation tasks, it is unclear whether a strong correlation still exists between perplexity and various forms of error cost functions in recognition and translation stages. The proposed minimum Bayes risk (MBR) based approach provides a flexible framework for unsupervised LM adaptation. It generalizes to a variety of forms of recognition and translation error metrics. LM adaptation is performed at the audio document level using either the character error rate (CER), or translation edit rate (TER) as the cost function. An efficient parameter estimation scheme using the extended Baum-Welch (EBW) algorithm is proposed. Experimental results on a state-of-the-art speech recognition and translation system are presented. The MBR adapted language models gave the best recognition and translation performance and reduced the TER score by up to 0.54% absolute.
Xunying Liu, William J. Byrne, Mark J. F. Gales, Adrià de Gispert, Marcus Tomalin, Philip C. Woodland, Kai Yu 0004
ASRU3
2007 Complementary System Generation using Directed Decision Trees
abstract
Large vocabulary continuous speech recognition (LVCSR) systems often use a multi-pass decoding strategy with a combination of multiple systems in the final stage. To reduce the error rate, these models must be complementary, i.e. make different errors. Previously, complementary systems have been generated by independently training a number of models, explicitly performing all combinations and picking the best performance. This method becomes infeasible as the potential number of systems increases, and does not guarantee that any of the models will be complementary. This paper presents an algorithm for generating complementary systems by altering the decision tree generation. Confusions made by a baseline system are resolved by separating confusable states, which might previously have been clustered together using the standard decision tree algorithm. Experimental results presented on a broadcast news Mandarin task show gains when combining the baseline with a complementary directed decision tree system.
Catherine Breslin, Mark J. F. Gales
ICASSP (4)2
2007 Speech Recognition System Combination for Machine Translation
abstract
The majority of state-of-the-art speech recognition systems make use of system combination. The combination approaches adopted have traditionally been tuned to minimising word error rates (WERs). In recent years there has been a growing interest in taking the output from speech recognition systems in one language and translating it into another. This paper investigates the use of cross-site combination approaches in terms of both WER and impact on translation performance. In addition, the stages involved in modifying the output from a speech-to-text (STT) system to be suitable for translation are described. Two source languages, Mandarin and Arabic, are recognised and then translated using a phrase-based statistical machine translation system into English. Performance of individual systems and cross-site combination using cross-adaptation and ROVER are given. Results show that the best STT combination scheme in terms of WER is not necessarily the most appropriate when translating speech.
Mark J. F. Gales, Xunying Liu, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004, Spyridon Matsoukas, Tim Ng, Kham Nguyen, Long Nguyen 0001, Jean-Luc Gauvain, Lori Lamel, Abdelkhalek Messaoudi
ICASSP (4)1
2007 Adaptive Training with Joint Uncertainty Decoding for Robust Recognition of Noisy Data
abstract
Standard noise compensation techniques for automatic speech recognition assume a clean trained acoustic model. What is thought of as "clean" data, may still have a variety of speakers, different channels and varying noise conditions. Hence it may be more reasonable to consider such data multi-conditional for multistyle training. This paper shows that multistyle models benefit from VTS compensation or joint uncertainty decoding by reducing the mismatch between training and test. An EM-based noise estimation procedure that produces ML VTS or joint noise models is also described. Alternatively, adaptive training with joint uncertainty transforms factors out the noise from the data. The uncertainty variance bias de-weights observations in the training data where the SNR is low. This property allows data with a wide SNR range to be used and produces canonical models that truly represent clean speech, whereas multistyle trained models must account for all acoustic variation associated with different noise conditions. This paper presents joint adaptive training including formula for estimating the transforms and canonical model parameters. Experiments are conducted on the resource management and broadcast news corpora.
Hank Liao, Mark J. F. Gales
ICASSP (4)2
2007 Consensus Network Decoding for Statistical Machine Translation System Combination
abstract
This paper presents a simple and robust consensus decoding approach for combining multiple machine translation (MT) system outputs. A consensus network is constructed from an N-best list by aligning the hypotheses against an alignment reference, where the alignment is based on minimising the translation edit rate (TER). The minimum Bayes risk (MBR) decoding technique is investigated for the selection of an appropriate alignment reference. Several alternative decoding strategies proposed to retain coherent phrases in the original translations. Experimental results are presented primarily based on three-way combination of Chinese-English translation outputs, and also presents results for six-way system combination. It is shown that worthwhile improvements in translation performance can be obtained using the methods discussed.
Khe Chai Sim, William J. Byrne, Mark J. F. Gales, Hichem Sahbi, Philip C. Woodland
ICASSP (4)3
2007 Improving Speech Transcription for Mandarin-English Translation
abstract
This paper describes the development of the CU-HTK Mandarin speech-to-text (STT) system and assesses its performance as part of a transcription-translation pipeline which converts broadcast Mandarin audio into English text. Recent improvements to the STT system are described and these give character error rate (CER) gains of 14.3% absolute for a broadcast conversation (BC) task and 5.1% absolute for a broadcast news (BN) task. The output of these STT systems is then post-processed, so that it consists of sentence-like segments, and translated into English text using a statistical machine translation (SMT) system. The performance of the transcription-translation pipeline is evaluated using the translation edit rate (TER) and BLEU metrics. It is shown that improving both the STT system and the post-STT segmentations can lower the TER scores by up to 5.3% absolute and increase the BLEU scores by up to 2.7% absolute.
Marcus Tomalin, Mark J. F. Gales, Xunying Liu, Khe Chai Sim, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004
ICASSP (4)2
2007 Unsupervised Training for Mandarin Broadcast News and Conversation Transcription
abstract
A significant cost in obtaining acoustic training data is the generation of accurate transcriptions. For some sources close-caption data is available. This allows the use of lightly-supervised training techniques. However, for some sources and languages close-caption is not available. In these cases unsupervised training techniques must be used. This paper examines the use of unsupervised techniques for discriminative training. In unsupervised training automatic transcriptions from a recognition system are used for training. As these transcriptions may be errorful data selection may be useful. Two forms of selection are described, one to remove non-target language shows, the other to remove segments with low confidence. Experiments were carried out on a Mandarin transcriptions task. Two types of test data were considered, broadcast news (BN) and broadcast conversations (BC). Results show that the gains from unsupervised discriminative training are highly dependent on the accuracy of the automatic transcriptions.
Mark J. F. Gales, Philip C. Woodland
ICASSP (4)2
2007 Building multiple complementary systems using directed decision trees
abstract
Large vocabulary speech recognition systems typically use a combination of multiple systems to obtain the final hypothesis. For combination to give gains, the systems being combined must be complementary, i.e. they must make different errors. Often, complementary systems are chosen simply by training multiple systems, performing all combinations, and selecting the best. This approach becomes time consuming as more potential systems are considered, and hence recent work has looked at explicitly building systems to be complementary to each other. This paper considers building multiple complementary systems based on directed decision trees, and combining them within a multi-pass adaptive framework. The tree divergence is introduced for easy comparison of trees without having to build entire systems. Experiments are presented on a Broadcast News Arabic task, and show that gains can be achieved by using more than one complementary system. Index Terms: automatic speech recognition, system combination, complementary systems
Catherine Breslin, Mark J. F. Gales
INTERSPEECH2
2007 Derivative and parametric kernels for speaker verification
abstract
The use of Support Vector Machines (SVMs) for speaker verification has become increasingly popular. To handle the dynamic nature of the speech utterances, many SVM-based systems use dynamic kernels. Many of these kernels can be placed into two classes, parametric kernels, where the feature-space consists of parameters from the utterance-dependent model, and derivative kernels, where the derivatives of the utterance loglikelihood with respect to parameters of a generative model are used. This paper contrasts the attributes of these two forms of kernel. Furthermore, the conditions under which the two forms of kernel are identical are described. Two forms of dynamic kernel are examined in detail, based on MLLR-adaptation and mean MAP-adapted models. The performance of these kernels is evaluated on the NIST SRE 2002 dataset. Combining the two forms of kernel together gave a 5 % relative reduction in Equal Error Rate compared to the best individual kernel.
Chris Longworth, Mark J. F. Gales
INTERSPEECH2
2007 Unsupervised training with directed manual transcription for recognising Mandarin broadcast audio
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH2
2007 Discriminative semi-parametric trajectory model for speech recognition
Khe Chai Sim, Mark J. F. Gales
Comput. Speech Lang.2
2007 Automatic Model Complexity Control Using Marginalized Discriminative Growth Functions
abstract
Selecting the model structure with the "appropriate" complexity is a standard problem for training large-vocabulary continuous-speech recognition (LVCSR) systems. State-of-the-art LVCSR systems are highly complex. A wide variety of techniques may be used which alter the system complexity and word error rate (WER). Explicitly evaluating systems for all possible configurations is infeasible; hence, an automatic model complexity control criterion is highly desirable. Most existing complexity control schemes can be classified into two types, Bayesian learning techniques and information theory approaches. An implicit assumption is made in both, that increasing the likelihood on held-out data decreases the WER. However, this correlation has been found quite weak for current speech recognition systems. This paper presents a novel discriminative complexity control technique, the marginalization of a discriminative growth function. This is a closer approximation to the true WER than standard approaches. Experimental results on a standard LVCSR Switchboard task showed that marginalized discriminative growth functions outperforms manually tuned systems and conventional complexity control techniques, such as Bayesian information criterion (BIC), in terms of WER
Xunying Liu, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2007 Bayesian Adaptive Inference and Adaptive Training
abstract
Large-vocabulary speech recognition systems are often built using found data, such as broadcast news. In contrast to carefully collected data, found data normally contains multiple acoustic conditions, such as speaker or environmental noise. Adaptive training is a powerful approach to build systems on such data. Here, transforms are used to represent the different acoustic conditions, and then a canonical model is trained given this set of transforms. This paper describes a Bayesian framework for adaptive training and inference. This framework addresses some limitations of standard maximum-likelihood approaches. In contrast to the standard approach, the adaptively trained system can be directly used in unsupervised inference, rather than having to rely on initial hypotheses being present. In addition, for limited adaptation data, robust recognition performance can be obtained. The limited data problem often occurs in testing as there is no control over the amount of the adaptation data available. In contrast, for adaptive training, it is possible to control the system complexity to reflect the available data. Thus, the standard point estimates may be used. As the integral associated with Bayesian adaptive inference is intractable, various marginalization approximations are described, including a variational Bayes approximation. Both batch and incremental modes of adaptive inference are discussed. These approaches are applied to adaptive training of maximum-likelihood linear regression and evaluated on a large-vocabulary speech recognition task. Bayesian adaptive inference is shown to significantly outperform standard approaches.
Kai Yu 0004, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2006 Augmented Statistical Models for Speech Recognition
abstract
Recently there has been significant interest in developing new acoustic models for speech recognition. One such model, that allows complex dependencies to be represented, is the augmented statistical model. This incorporates additional dependencies by constructing a local exponential expansion of a standard HMM. Unfortunately, the resulting model often has an intractable normalisation term, rendering training difficult for all but binary classification tasks. In this paper, conditional augmented (C-Aug) models are proposed as an attractive alternative. Instead of modelling utterance likelihoods and inferring decision boundaries, C-Aug models directly model the posterior probability of class labels, conditioned on the utterance. The resulting model is easy to normalise and can be trained using conditional maximum likelihood estimation. In addition, as a convex model, the optimisation converges to a global maximum
Martin I. Layton, Mark J. F. Gales
ICASSP (1)2
2006 The Cu-Htk Mandarin Broadcast News Transcription System
abstract
This paper discusses the development of the CU-HTK Mandarin broadcast news (BN) transcription system. The Mandarin BN task includes a significant amount of English data. Hence techniques have been investigated to allow the same system to handle both Mandarin and English by augmenting the Mandarin training sets with English acoustic and language model training data. A range of acoustic models were built including models based on Gaussianised features, speaker adaptive training and feature-space MPE. A multi-branch system architecture is described in which multiple acoustic model types, alternate phone sets and segmentations can be used in a system combination framework to generate the final output. The final system shows state-of-the-art performance over a range of test sets
Rohit Sinha 0003, Mark J. F. Gales, Do Yeong Kim, Xunying Liu, Khe Chai Sim, Philip C. Woodland
ICASSP (1)2
2006 Incremental Adaptation using Bayesian Inference
abstract
Adaptive training is a powerful technique to build system on nonhomogeneous training data. Here, a canonical model, representing “ pure” speech variability and a set of transforms representing unwanted acoustic variabilities are both trained. To use the canonical model for recognition, a transform for the test acoustic condition is required. For some situations a robust estimate of the transform parameters may not be possible due to limited, or no, adaptation data. One solution to this problem is to view adaptive training in a Bayesian framework and marginalise out the transform parameters. Exact implementation of this Bayesian inference is intractable. Recently, lower bound approximations based on variational Bayes have been used to solve this problem for batch adaptation with limited data. This paper extends this Bayesian adaptation framework to incremental adaptation. Various lower-bound approximations and options for propagating information within this incremental framework are discussed. Experiments using adaptive models trained with both maximum likelihood and minimum phone error training are described. Using incremental Bayesian adaptation gains were obtained over the standard approaches, especially for limited data.
Kai Yu 0004, Mark J. F. Gales
ICASSP (1)2
2006 Generating complementary systems for speech recognition
Catherine Breslin, Mark J. F. Gales
INTERSPEECH2
2006 Issues with uncertainty decoding for noise robust speech recognition
abstract
Interest is growing in a class of robustness algorithms that exploit the notion of uncertainty introduced by environmental noise. The majority of these techniques share the property that the uncertainty of an observation due to noise is propagated to the recogniser, resulting in increased model variances. Using appropriate approximations, efficient implementations may be obtained, with the goal of achieving near model-based performance without the associated computational cost. Unfortunately, uncertainty decoding forms that compute the uncertainty in the front-end and pass this to the decoder may suffer from a theoretical problem in low signal-to-noise ratio conditions. This report discusses how this fundamental issue arises, and demonstrates it through two schemes: SPLICE with uncertainty and front-end Joint uncertainty decoding. A method to mitigate this in theJoint form is presented, as well as how SPLICE implicitly addresses it. However, it is shown that a model-based Joint uncertainty decoding approach does not suffer from this limitation, like these front-end forms do, and is also competitive computationally. The issues described and performance of the various schemes are examined on two artificially corrupted corpora: AURORA 2.0 digit recognition database and the thousand-word Resource Management task. 2 1
Hank Liao, Mark J. F. Gales
INTERSPEECH2
2006 Discriminative adaptation for speaker verification
abstract
Speaker verification is a binary classification task to determine whether a claimed speaker uttered a phrase. Current approaches to speaker verification tasks typically involve adapting a general speaker Universal Background Model (UBM), normally a Gaussian Mixture Model (GMM), to model a particular speaker. Verification is then performed by comparing the likelihoods from the speaker model to the UBM. Maximum A-Posteriori (MAP) is commonly used to adapt the UBM to a particular speaker. However speaker verification is a classification task. Thus, robust discriminative-based adaptation schemes should yield gains over the standard MAP approach. This paper describes and evaluates two discriminative approaches to speaker verification. The first is a discriminative version of MAP based on Maximum Mutual Information (MMI-MAP). The second is to use an augmented-GMM (A-GMM) as the speaker-specific model. The additional, augmented, parameters are discriminatively, and robustly, trained using a maximum margin estimation approach. The performance of these models is evaluated on the NIST 2002 SRE dataset. Though no gains were obtained using MMI-MAP, the A-GMM system gave an Equal Error Rate (EER) of 7.31%, a 30 % relative reduction in EER compared to the best performing GMM system. Index Terms: augmented statistical models, discriminative training, sequence kernels, speaker verification.
Chris Longworth, Mark J. F. Gales
INTERSPEECH2
2006 Product of Gaussians for speech recognition
Mark J. F. Gales, S. S. Airey
Comput. Speech Lang.1
2006 Progress in the CU-HTK broadcast news transcription system
abstract
Broadcast news (BN) transcription has been a challenging research area for many years. In the last couple of years, the availability of large amounts of roughly transcribed acoustic training data and advanced model training techniques has offered the opportunity to greatly reduce the error rate on this task. This paper describes the design and performance of BN transcription systems which make use of these developments. First, the effects of using lightly supervised training data and advanced acoustic modeling techniques are discussed. The design of a real-time broadcast news recognition system is then detailed using these new models. As system combination has been found to yield large gains in performance, a range of frameworks that allow multiple recognition outputs to be combined are next described. These include the use of multiple types of acoustic models and multiple segmentations. As a contrast a system developed by multiple sites allowing cross-site combination, the "SuperEARS" system, is also described. The various models and recognition configurations are evaluated using several recent BN development and evaluation test sets. These new BN transcription systems can give gains of over 25% relative to the CU-HTK 2003 BN system
Mark J. F. Gales, Do Yeong Kim, Philip C. Woodland, Ricky Ho Yin Chan, David Mrva, Rohit Sinha 0003, Sue Tranter
IEEE Trans. Speech Audio Process.1
2006 Corrections to "Automatic Transcription of Conversational Telephone Speech"
Thomas Hain, Philip C. Woodland, Gunnar Evermann, Mark J. F. Gales, Xunying Liu, Gareth L. Moore, Daniel Povey
IEEE Trans. Speech Audio Process.4
2006 Minimum phone error training of precision matrix models
abstract
Gaussian mixture models (GMMs) are commonly used as the output density function for large-vocabulary continuous speech recognition (LVCSR) systems. A standard problem when using multivariate GMMs to classify data is how to accurately represent the correlations in the feature vector. Full covariance matrices yield a good model, but dramatically increase the number of model parameters. Hence, diagonal covariance matrices are commonly used. Structured precision matrix approximations provide an alternative, flexible, and compact representation. Schemes in this category include the extended maximum likelihood linear transform and subspace for precision and mean models. This paper examines how these precision matrix models can be discriminatively trained and used on state-of-the-art speech recognition tasks. In particular, the use of the minimum phone error criterion is investigated. Implementation issues associated with building LVCSR systems are also addressed. These models are evaluated and compared using large vocabulary continuous telephone speech and broadcast news English tasks.
Khe Chai Sim, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2006 Discriminative cluster adaptive training
abstract
Multiple-cluster schemes, such as cluster adaptive training (CAT) or eigenvoice systems, are a popular approach for rapid speaker and environment adaptation. Interpolation weights are used to transform a multiple-cluster, canonical, model to a standard hidden Markov model (HMM) set representative of an individual speaker or acoustic environment. Maximum likelihood training for CAT has previously been investigated. However, in state-of-the-art large vocabulary continuous speech recognition systems, discriminative training is commonly employed. This paper investigates applying discriminative training to multiple-cluster systems. In particular, minimum phone error (MPE) update formulae for CAT systems are derived. In order to use MPE in this case, modifications to the standard MPE smoothing function and the prior distribution associated with MPE training are required. A more complex adaptive training scheme combining both interpolation weights and linear transforms, a structured transform (ST), is also discussed within the MPE training framework. Discriminatively trained CAT and ST systems were evaluated on a state-of-the-art conversational telephone speech task. These multiple-cluster systems were found to outperform both standard and adaptively trained systems.
Kai Yu 0004, Mark J. F. Gales
IEEE Trans. Speech Audio Process.2
2005 Training LVCSR Systems on Thousands of Hours of Data
abstract
Typical systems for large vocabulary conversational speech recognition (LVCSR) have been trained on a few hundred hours of carefully transcribed acoustic training data. The paper describes an LVCSR system for the conversational telephone speech (CTS) task trained on more than 2000 hours of data for which only approximate transcriptions were available. The challenges of dealing with such a large data set and the accuracy improvements over the small baseline system are discussed. The effect on both acoustic and language modelling performance is studied. Overall, increasing the training data size from 360 h to 2200 h and optimising the training procedure reduced the word error rate on the DARPA/NIST 2003 evaluation set by about 20% relative.
Gunnar Evermann, Ricky Ho Yin Chan, Mark J. F. Gales, David Mrva, Philip C. Woodland, Kai Yu 0004
ICASSP (1)3
2005 Development of the CUHTK 2004 Mandarin Conversational Telephone Speech Transcription System
abstract
The paper details all aspects of the CUHTK 2004 Mandarin conversational telephone speech transcription system, but concentrates on the development of the acoustic models. As there are significant differences between the available training corpora, both in terms of topics of conversation and accents, forms of data normalisation and adaptive training techniques are investigated. The baseline discriminatively trained acoustic models are compared to a system built with a Gaussianisation front-end, a speaker adaptively trained system and an adaptively trained structured precision matrix system. The models are finally evaluated within a multi-pass, multi-branch, system combination framework.
Mark J. F. Gales, Xunying Liu, Khe Chai Sim, Philip C. Woodland, Kai Yu 0004
ICASSP (1)1
2005 Development of the CU-HTK 2004 Broadcast News Transcription Systems
abstract
The paper describes our recent work on improving broadcast news transcription and presents details of the CU-HTK broadcast news English (BN-E) transcription system for the DARPA/NIST rich transcription 2004 speech-to-text (RT04) evaluation. A key focus has been building a system using an order of magnitude more acoustic training data than we have previously attempted. We have also investigated a range of techniques to improve both minimum phone error (MPE) training and the efficient creation of MPE-based narrow-band models. The paper describes two alternative system structures that run in under 10/spl times/RT and a further system that runs in less than 1/spl times/RT. This final system gives lower word error rates than our 2003 system that ran in 10/spl times/RT.
Do Yeong Kim, Ricky Ho Yin Chan, Gunnar Evermann, Mark J. F. Gales, David Mrva, Khe Chai Sim, Philip C. Woodland
ICASSP (1)4
2005 Investigation of Acoustic Modeling Techniques for LVCSR Systems
abstract
The paper describes the use of several advanced acoustic modeling techniques for the 2004 CU-HTK large vocabulary speech recognition systems. These techniques include Gaussianization for speaker normalization, discriminative cluster adaptive training (CAT), subspace for precision and mean (SPAM) modeling of inverse covariances, and discriminative complexity control. Acoustic models featuring these techniques were integrated into a state-of-the-art 10 real-time multi-pass system with sophisticated adaptation for performance evaluation. Experimental results are presented on both broadcast news (BN) and conversational telephone speech (CTS) transcription tasks.
Xunying Liu, Mark J. F. Gales, Khe Chai Sim, Kai Yu 0004
ICASSP (1)2
2005 Adaptation of Precision Matrix Models on Large Vocabulary Continuous Speech Recognition
abstract
Recently, structured precision matrix models were found to outperform the conventional diagonal covariance matrix models. Minimum phone error discriminative training of these models gave very good unadapted performance on large vocabulary continuous speech recognition systems. To obtain state-of-the-art performance, it is important to apply adaptation techniques efficiently to these models. In this paper, simple row-by-row iterative formulae are described for both MLLR mean and constrained MLLR transform estimations of these models. These update formulae are derived within the standard expectation maximisation framework and are guaranteed to increase the likelihood of the adaptation data. Efficient approximate schemes for these adaptation methods are also investigated to further reduce the computation. Experimental results are presented based on the MPE trained subspace for precision and mean models, evaluated on both broadcast news and conversational telephone speech English tasks.
Khe Chai Sim, Mark J. F. Gales
ICASSP (1)2
2005 Joint uncertainty decoding for noise robust speech recognition
abstract
Background noise can have a significant impact on the performance of speech recognition systems. A range of fast featurespace and model-based schemes have been investigated to increase robustness. Model-based approaches typically achieve lower error rates, but at an increased computational load compared to feature-based approaches. This makes their use in many situations impractical. The uncertainty decoding framework can be considered an elegant compromise between the two. Here, the uncertainty of features is propagated to the recogniser in a mathematically consistent fashion. The complexity of the model used to determine the uncertainty may be decoupled from the recognition model itself, allowing flexibility in the computational load. This paper describes a new approach within this framework, Joint uncertainty decoding. This approach is compared with the uncertainty decoding version ofSPLICE, standardSPLICE, and a new form of front-end CMLLR. These are evaluated on a medium vocabulary speech recognition task with artificially added noise.
Hank Liao, Mark J. F. Gales
INTERSPEECH2
2005 Temporally varying model parameters for large vocabulary continuous speech recognition
abstract
Many forms of time varying acoustic models have been applied to the area of speech recognition. However, there has been little success in applying these models to Large Vocabulary Continuous Speech Recognition (LVCSR). Recently, fMPE was introduced as a discriminative feature space estimation scheme for the HMM-based LVCSR. This method estimates a projection matrix from a high dimensional space ( ∼ 100,000) down to a standard feature space (typically 39). This projection is then added on to the original feature vector (e.g. MFCC or PLP) to yield a feature vector to train the final model. This paper considers fMPE as a time varying model for the mean vectors by applying the time varying feature offset to the Gaussian mean vectors. This approach naturally yields the update formulae for fMPE and motivates an alternative style of training systems. This concept is then extended to the temporal precision matrix modelling (pMPE). In pMPE, a temporally varying positive scale is applied to each element of the diagonal precision matrices. Experimental results are presented on a conversational telephone speech English task. 1.
Khe Chai Sim, Mark J. F. Gales
INTERSPEECH2
2005 The Cambridge University March 2005 speaker diarisation system
abstract
This paper describes the speaker diarisation system developed at Cambridge University in March 2005. This system combines techniques used successfully in our previous speaker diarisation systems with an additional second clustering stage based on state-of-the-art speaker identification methods. Several strategies for using the new system are investigated and the final system gives a diarisation error rate of 6.9 % on the RT-04 Fall diarisation evaluation data when processing all the test data together or 8.6 % when processing the test data shows independently.
Rohit Sinha 0003, Sue Tranter, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2005 Automatic transcription of conversational telephone speech
abstract
This paper discusses the Cambridge University HTK (CU-HTK) system for the automatic transcription of conversational telephone speech. A detailed discussion of the most important techniques in front-end processing, acoustic modeling and model training, language and pronunciation modeling are presented. These include the use of conversation side based cepstral normalization, vocal tract length normalization, heteroscedastic linear discriminant analysis for feature projection, minimum phone error training and speaker adaptive training, lattice-based model adaptation, confusion network based decoding and confidence score estimation, pronunciation selection, language model interpolation, and class based language models. The transcription system developed for participation in the 2002 NIST Rich Transcription evaluations of English conversational telephone speech data is presented in detail. In this evaluation the CU-HTK system gave an overall word error rate of 23.9%, which was the best performance by a statistically significant margin. Further details on the derivation of faster systems with moderate performance degradation are discussed in the context of the 2002 CU-HTK 10 /spl times/ RT conversational speech transcription system.
Thomas Hain, Philip C. Woodland, Gunnar Evermann, Mark J. F. Gales, Xunying Liu, Gareth L. Moore, Daniel Povey
IEEE Trans. Speech Audio Process.4
2004 Development of the 2003 CU-HTK conversational telephone speech transcription system
abstract
The paper describes the development of the 2003 CU-HTK large vocabulary speech recognition system for conversational telephone speech (CTS). The system was designed based on a multipass, multibranch structure where the output of all branches is combined using system combination. A number of advanced modelling techniques, such as speaker adaptive training, heteroscedastic linear discriminant analysis, minimum phone error estimation and specially constructed single pronunciation dictionaries, were employed. The effectiveness of each of these techniques and their potential contribution to the result of system combination was evaluated in the framework of a state-of-the-art LVCSR system with sophisticated adaptation. The final 2003 CU-HTK CTS system constructed from some of these models is described and its performance on the DARPA/NIST 2003 rich transcription (RT-03) evaluation test set is discussed.
Gunnar Evermann, Ricky Ho Yin Chan, Mark J. F. Gales, Thomas Hain, Xunying Liu, David Mrva, Philip C. Woodland
ICASSP (1)3
2004 Model complexity control and compression using discriminative growth functions
abstract
State-of-the-art large vocabulary speech recognition systems are highly complex. Many techniques affect both system complexity and recognition performance. The need to determine the appropriate complexity without having to build each possible system has led to the development of automatic complexity control criteria. In this paper further experiments are carried out using a recently proposed criterion based on marginalizing a maximum mutual information (MMI) growth function. The use of this criterion is much detailed for determining the appropriate dimensionality in a multiple HLDA system and the number of components per state. A scheme for also using this criterion for model compression is described. Experimental results on a spontaneous telephone speech recognition task are described. Initial system compression experiments are inconclusive. However, comparing a standard state-of-the-art system with one generated using complexity control shows a reduction in word error rate.
Xunying Liu, Mark J. F. Gales
ICASSP (1)2
2004 Rao-Blackwellised Gibbs sampling for switching linear dynamical systems
abstract
This paper describes the application of Rao-Blackwellised Gibbs sampling (RBGS) to speech recognition using switching linear dynamical systems (SLDS). The SLDS is a hybrid of standard hidden Markov models (HMM) and linear dynamical systems. It is an extension of the stochastic segment model as it relaxes the assumption of independent segments. SLDS explicitly take into account the strong co-articulation present in speech. Unfortunately, inference in SLDS is intractable unless the discrete state sequence is known. RBGS is one approach that may be applied for both improved training and decoding for this form of intractable model. The theory of SLDS and RBGS is described, along with an efficient proposal mechanism. The performance of the SLDS using RBGS for training and inference is evaluated on the ARPA Resource Management task.
Antti-Veikko I. Rosti, Mark J. F. Gales
ICASSP (1)2
2004 Basis superposition precision matrix modelling for large vocabulary continuous speech recognition
abstract
An important aspect of using Gaussian mixture models in a HMM-based speech recognition systems is the form of the covariance matrix. One successful approach has been to model the inverse covariance, precision, matrix by superimposing multiple bases. This paper presents a general framework of basis superposition. Models are described in terms of parameter tying of the basis coefficients and restrictions in the number of basis. Two forms of parameter tying are described which provide a compact model structure. The first constrains the basis coefficients over multiple basis vectors (or matrices). This is related to the Subspace for Precision and Mean (SPAM) model. The second constrains the basis coefficients over multiple components, yielding as one example heteroscedastic LDA (HLDA). Both maximum likelihood and minimum phone error training of these models are discussed. The performance of various configurations is examined on a conversational telephone speech task, SwitchBoard.
Khe Chai Sim, Mark J. F. Gales
ICASSP (1)2
2004 Adaptive training using structured transforms
abstract
Adaptive training is an important approach to training speech recognition systems on found, non-homogeneous data. The standard approach employs a single transform to represent unwanted acoustic variability. However, for found data there are commonly multiple acoustic factors affecting the speech signal. The paper investigates the use of multiple forms of transformations, structured transforms (ST), to represent the complex non-speech variabilities in an adaptive training framework. Two forms of transformation are considered, cluster mean interpolation and constrained MLLR; consequently, the canonical model here is a multi-cluster HMM model. Both ML and minimum phone error (MPE) reestimation formulae for the canonical model, are presented. This multi-cluster MPE training is also applicable to eigenvoice systems. Experiments to compare ST to standard adaptive training schemes were performed on a conversational telephone speech task. ST were found to reduce the word error rate significantly.
Kai Yu 0004, Mark J. F. Gales
ICASSP (1)2
2004 Using VTLN for broadcast news transcription
abstract
Vocal tract length normalisation (VTLN) is a commonly used speaker normalisation approach. It is attractive compared to many normalisation schemes as it is typically dependent on only a single parameter, allowing the warp factors to be robustly calculated on little data. However, the scheme normally requires explicitly coding the data at multiple warp factors. Furthermore, it is only possible to approximate the Jacobian associated with the VTLN transformation. A new, simple, linear approximation to VTLN is described in this paper. This linear approximation allows the Jacobian to be exactly computed. It can also be highly efficient in terms of warp factor estimation and application of the warp factors. Both the linear and standard CUED VTLN schemes were evaluated in the 2003 BNE evaluation framework and found to yield similar performance. When used in system combination both VTLN schemes yielded slight gains over the baseline system.
Do Yeong Kim, Srinivasan Umesh, Mark J. F. Gales, Thomas Hain, Philip C. Woodland
INTERSPEECH3
2004 Factor analysed hidden Markov models for speech recognition
Antti-Veikko I. Rosti, Mark J. F. Gales
Comput. Speech Lang.2
2003 Product of Gaussians and multiple stream systems
abstract
There has been interest in the use of classifiers based on the product of experts (PoE). PoEs offer an alternative to the standard mixture of experts (MoE) framework. This paper presents a particular form of PoE, the product of Gaussians (PoG), within an hidden Markov model framework. Training and initialisation procedures are described for this PoG system. In addition, the relationship of PoG to standard multiple stream systems is explored. The PoG system performance is examined on the SwitchBoard task and is compared to standard Gaussian mixture systems and multiple stream systems.
S. S. Airey, Mark J. F. Gales
ICASSP (1)2
2003 Porting: SwitchBoard to the VoiceMail task
abstract
The paper examines techniques that allow a well-trained source system built on one task to be rapidly adapted, or ported, to another target task. The two tasks considered are Hub5, or SwitchBoard, as the source system and VoiceMail as the target task. The two tasks are acoustically similar, both being telephone-bandwidth speech tasks, but differ in speaking style. SwitchBoard is conversational speech, VoiceMail is a set of voicemail messages. Various porting schemes for acoustic models are examined, including discriminative MAP and heteroscedastic LDA. Using around 28 hours of data, the error rate on VoiceMail was reduced by 42% relative compared to the baseline SwitchBoard performance.
Mark J. F. Gales, Daniel Povey, Philip C. Woodland
ICASSP (1)1
2003 Automatic complexity control for HLDA systems
abstract
Designing a state-of-the-art large vocabulary speech recognition systems is a highly complex problem. A wide range of techniques are available that affect the performance and number of free parameters. Selecting the appropriate complexity of system is both time-consuming and only a limited number of possible systems can be examined. This paper presents initial results on automatic system selection when both the number of dimensions and the number of components vary. Various complexity control schemes are discussed and evaluated. Limitations of schemes based on predicting held-out data log-likelihoods are described. In addition, problems of standard approximations for this task are detailed.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP (1)2
2003 Discriminative map for acoustic model adaptation
abstract
In this paper we show how a discriminative objective function such as Maximum Mutual Information (MMI) can be combined with a prior distribution over the HMM parameters to give a discriminative Maximum A Posteriori (MAP) estimate for HMM training. The prior distribution can be based around the Maximum Likelihood (ML) parameter estimates, leading to a technique previously referred to as I-smoothing; or for adaptation it can be based around a MAP estimate of the ML parameters, leading to what we call MMI-MAP. This latter approach is shown to be effective for task adaptation, where data from one task (Voicemail) is used to adapt a HMM set trained on another task (Switchboard). It is shown that MMI-MAP results in a 2.1% absolute reduction in word error rate relative to standard ML-MAP with 30 hours of Voicemail task adaptation data starting from a MMI-trained Switchboard system.
Daniel Povey, Philip C. Woodland, Mark J. F. Gales
ICASSP (1)3
2003 Product of Gaussians as a distributed representation for speech recognition
S. S. Airey, Mark J. F. Gales
INTERSPEECH2
2003 MMI-MAP and MPE-MAP for acoustic model adaptation
abstract
This paper investigates the use of discriminative schemes based on themaximum mutual information (MMI) and minimum phone error (MPE) objective functions for both task and gender adaptation. A method for incorporating prior information into the discriminative training framework is described. If an appropriate form of prior distribution is used, then this may be implemented by simply altering the values of the counts used for parameter estimation. The prior distribution can be based around maximum likelihood parameter estimates, giving a technique known as I-smoothing, or for adaptation it can be based around a MAP estimate of the ML parameters, leading to MMI-MAP, or MPE-MAP.MMI-MAP isshown tobe effectivefor taskadaptation, where data from one task (Voicemail) is used to adapt a HMM set trained on another task (Switchboard). MPE-MAP is shown to be effective for generating gender-dependent models for Broadcast News transcription.
Daniel Povey, Mark J. F. Gales, Do Yeong Kim, Philip C. Woodland
INTERSPEECH2
2002 Improved cross-task recognition using MMIE training
abstract
This paper investigates the cross-task recognition and adaptation performance of HMMs trained using either conventional maximum likelihood estimation or the discriminative maximum mutual information estimation (MMIE) criterion. Initial experiments used models trained on the low noise North American Business news corpus of read speech. Cross-task testing on Broadcast News data showed that the MMIE models yielded lower error rates both across-task as well as within-task. This result was confirmed using models trained on the Switchboard corpus which were tested on Voicemail (VM)data. This setup was also used to investigate the performance of task-adaptation when using a limited amount of VM data for both acoustic and language modelling. The setup that gave the best performance on the VM test data used Switchboard models trained using MMIE and then adapted to VM data using maximum a posteriori adaptation techniques.
Ricardo de Córdoba, Philip C. Woodland, Mark J. F. Gales
ICASSP3
2002 The HMM error model
abstract
The most popular model used in automatic speech recognition is the hidden Markov model (HMM). Though good performance has been obtained with such models there are well known limitations in its ability to model speech. For these reasons, a variety of modifications to the standard HMM topology have been proposed including factorial, or multi-stream, HMMs. This paper describes a new form of HMM based on transformation streams. A particular form of transformation stream is described, the HMM error model (HHM). This model may be viewed as a filter model, the transformation stream, and a residual model. The filter model transforms the original data into a space in which all the data is “similarly” distributed. This normalised data is then modelled using the residual model. The HEM is evaluated on a standard large vocabulary speaker independent speech recognition task, SwitchBoard. On this task significant reductions in word error rate are obtained over standard HMM-based systems.
Mark J. F. Gales
ICASSP1
2002 Factor analysed hidden Markov models
abstract
This paper presents a general form of acoustic model for speech recognition. The model is based on an extension to factor analysis where the low dimensional subspace is modelled with a mixture of Gaussians hidden Markov model (HMM) and the observation noise by a Gaussian mixture model. Here the HMM output vectors are the latent variables of a general factor analyser. The model combines shared factor analysis with a dynamic version of independent factor analysis. This factor analysed HMM (FAHMM) provides an alternative, compact, model to handle intra-frame correlation. Furthermore, it allows variable dimension subspaces to be explored. A variety of model configurations and sharing schemes are examined, some of which correspond to standard systems. The training and recognition algorithms for FAHMMs are described and some initial result with Switchboard are presented.
Antti-Veikko I. Rosti, Mark J. F. Gales
ICASSP2
2002 Using SVMS and discriminative models for speech recognition
abstract
In speech recognition, standard MAP decoders attribute speech data to the class with the highest posterior probability. This minimises the error rate under assumptions of model correctness. This assumption is invalid for speech recognition with HMMs. Hence, an interesting question is whether extra, useful information about the speech source can be extracted from the HMMs and used to lower error rates in practical systems, In this paper additional features are extracted from HMMs and incorporated into a multi-dimensional score-space. SVMs are then used to implement a decision rule. Preliminary experiments are performed on a small speaker-independent isolated letter task. Score-spaces based on discriminative models are used with previous results based on generative models. Both score-spaces outperform standard schemes.
Nathan D. Smith, Mark J. F. Gales
ICASSP2
2002 Combining a Gaussian mixture model front end with MFCC parameters
abstract
Fitting a Gaussian mixture model (GMM) to the smoothed speech spectrum allows an alternative set of features to be extracted from the speech signal. These features have been shown to possess information complementary to the standard MFCC parameterisation. This paper further investigates the use of these GMM features in combination with MFCCs. The extraction and use of a confidence metric to combine GMM features with MFCCs is described. Re- sults using the confidence metric on the WSJ task are presented. Also, GMM features for speech corrupted with additive noise are extracted from data corrupted with coloured addititve noise. Techniques for noise robustness and compensation are investigated for GMM features and the performance is examined on the RM task with additive noise.
Matthew N. Stuttle, Mark J. F. Gales
INTERSPEECH2
2002 Transformation streams and the HMM error model
Mark J. F. Gales
Comput. Speech Lang.1
2002 Automatic transcription of Broadcast News
Scott Saobing Chen, Ellen Eide, Mark J. F. Gales, Ramesh A. Gopinath, D. Kanvesky, Peder A. Olsen
Speech Commun.3
2002 Maximum likelihood multiple subspace projections for hidden Markov models
abstract
The first stage in many pattern recognition tasks is to generate a good set of features from the observed data. Usually, only a single feature space is used. However, in some complex pattern recognition tasks the choice of a good feature space may vary depending on the signal content. An example is in speech recognition where phone dependent feature subspaces may be useful. Handling multiple subspaces while still maintaining meaningful likelihood comparisons between classes is a key issue. This paper describes two new forms of multiple subspace schemes. For both schemes, the problem of handling likelihood consistency between the various subspaces is dealt with by viewing the projection schemes within a maximum likelihood framework. Efficient estimation formulae for the model parameters for both schemes are derived. In addition, the computational cost for their use during recognition are given. These new projection schemes are evaluated on a large vocabulary speech recognition task in terms of performance, speed of likelihood calculation and number parameters.
Mark J. F. Gales
IEEE Trans. Speech Audio Process.1
2001 Multiple-cluster adaptive training schemes
abstract
This paper examines the training of multiple-cluster systems using adaptive training schemes. Various forms of transformation and canonical model are described in a consistent framework allowing re-estimation formulae for all cases to be simply derived. Initial experiments using these various schemes on a large vocabulary speech recognition task are presented. The initial experiments indicate that to achieve best performance when adapting these multiple-cluster systems requires the use of adaptive training schemes rather than using simpler cluster initialisation schemes.
Mark J. F. Gales
ICASSP1
2001 A mixture of Gaussians front end for speech recognition
abstract
This paper describes a feature extraction technique based on fitting a Gaussian mixture model (GMM) to the speech spectral envelope. The features obtained (the component means, variances and priors) represent both the the general shape of the spectrum and provide information on the position of the spectral peaks. As the features select peaks in the spectrum they are related to the formant amplitudes, locations and bandwidths. Results using the Resource Management corpus, a medium vocabulary task are presented. Although by themselves the GMM features do not outperform MFCC features, systems combining the GMM systems with a standard frontend are shown to give a reduction in word error rate.
Matthew N. Stuttle, Mark J. F. Gales
INTERSPEECH2
2001 Speech Recognition using SVMs
abstract
An important issue in applying SVMs to speech recognition is the ability to classify variable length sequences. This paper presents extensions to a standard scheme for handling this variable length data, the Fisher score. A more useful mapping is introduced based on the likelihood-ratio. The score-space defined by this mapping avoids some limitations of the Fisher score. Class-conditional gen(cid:173) erative models are directly incorporated into the definition of the score-space. The mapping, and appropriate normalisation schemes, are evaluated on a speaker-independent isolated letter task where the new mapping outperforms both the Fisher score and HMMs trained to maximise likelihood.
Mark J. F. Gales
NIPS2
2000 Rapid likelihood calculation of subspace clustered Gaussian components
abstract
In speech recognition systems, computing the likelihoods of the acoustic models is an intensive task. One approach to reduce this cost is to use subspace distributed clustering HMM. Here individual Gaussian components are stored as indices to, and their likelihoods computed from, a set of subspace Gaussian components. This paper examines a scheme for reducing the computational cost of the likelihood calculation when such an HMM system is used. The proposed method identifies and stores frequently occurring partial sums called meta-atom elements and thus avoids computing them repeatedly. The resultant savings in the number of additions is 50% when all Gaussian components are computed or 20% when a Gaussian selection scheme is used.
Anuradha Aiyer, Mark J. F. Gales, Michael Picheny
ICASSP2
2000 Transcription of broadcast news with a time constraint: IBM's 10xRT HUB4 system
abstract
We describe a system which automatically transcribes broadcast news in less than 10 times real-time. We detail the system architecture of this system, which was used by IBM in the 1999 HUB4 10xRT evaluation, and show that the performance of this system is over 20 percent more accurate at the same speed than the system we used in the 1998 evaluation. Furthermore, we have closed the gap in word recognition accuracy between an unlimited resource system and this which runs in under 10 times real time from 45 percent to 14 percent.
Ellen Eide, Benoît Maison, Dimitri Kanevsky, Peder A. Olsen, Scott Saobing Chen, Lidia Mangu, Mark J. F. Gales, Miroslav Novak, Ramesh A. Gopinath
INTERSPEECH7
2000 Factored Semi-Tied Covariance Matrices
abstract
A new form of covariance modelling for Gaussian mixture models and hidden Markov models is presented. This is an extension to an efficient form of covariance modelling used in speech recognition, semi-tied co(cid:173) variance matrices. In the standard form of semi-tied covariance matrices the covariance matrix is decomposed into a highly shared decorrelating transform and a component-specific diagonal covariance matrix. The use of a factored decorrelating transform is presented in this paper. This fac(cid:173) toring effectively increases the number of possible transforms without in(cid:173) creasing the number of free parameters. Maximum likelihood estimation schemes for all the model parameters are presented including the compo(cid:173) nent/transform assignment, transform and component parameters. This new model form is evaluated on a large vocabulary speech recognition task. It is shown that using this factored form of covariance modelling reduces the word error rate.
Mark J. F. Gales
NIPS1
2000 Cluster adaptive training of hidden Markov models
abstract
When performing speaker adaptation, there are two conflicting requirements. First, the speaker transform must be powerful enough to represent the speaker. Second, the transform must be quickly and easily estimated for any particular speaker. The most popular adaptation schemes have used many parameters to adapt the models to be representative of an individual speaker. This limits how rapidly the models may be adapted to a new speaker or the acoustic environment. This paper examines an adaptation scheme requiring very few parameters, cluster adaptive training (CAT). CAT may be viewed as a simple extension to speaker clustering. Rather than selecting a single cluster as representative of a particular speaker, a linear interpolation of all the cluster means is used as the mean of the particular speaker. This scheme naturally falls into an adaptive training framework. Maximum likelihood estimates of the interpolation weights are given. Furthermore, simple re-estimation formulae for cluster means, represented both explicitly and by sets of transforms of some canonical mean, are given. On a speaker-independent task CAT reduced the word error rate using very little adaptation data. In addition when combined with other adaptation schemes it gave a 5% reduction in word error rate over adapting a speaker-independent model set.
Mark J. F. Gales
IEEE Trans. Speech Audio Process.1
1999 Recent improvements to IBM's speech recognition system for automatic transcription of broadcast news
abstract
We describe extensions and improvements to IBM's system for automatic transcription of broadcast news. The speech recognizer uses a total of 160 hours of acoustic training data, 80 hours more than for the system described in Chen et al. (1998). In addition to improvements obtained in 1997 we made a number of changes and algorithmic enhancements. Among these were changing the acoustic vocabulary, reducing the number of phonemes, insertion of short pauses, mixture models consisting of non-Gaussian components, pronunciation networks, factor analysis (FACILT) and Bayesian information criteria (BIC) applied to choosing the number of components in a Gaussian mixture model. The models were combined in a single system using NIST's script voting machine known as rover (Fiscus 1997).
Scott Saobing Chen, Ellen Eide, Mark J. F. Gales, Ramesh A. Gopinath, Dimitri Kanevsky, Peder A. Olsen
ICASSP3
1999 Tail distribution modelling using the richter and power exponential distributions
Mark J. F. Gales, Peder A. Olsen
EUROSPEECH1
1999 Semi-tied covariance matrices for hidden Markov models
abstract
There is normally a simple choice made in the form of the covariance matrix to be used with continuous-density HMMs. Either a diagonal covariance matrix is used, with the underlying assumption that elements of the feature vector are independent, or a full or block-diagonal matrix is used, where all or some of the correlations are explicitly modeled. Unfortunately when using full or block-diagonal covariance matrices there tends to be a dramatic increase in the number of parameters per Gaussian component, limiting the number of components which may be robustly estimated. This paper introduces a new form of covariance matrix which allows a few "full" covariance matrices to be shared over many distributions, whilst each distribution maintains its own "diagonal" covariance matrix. In contrast to other schemes which have hypothesized a similar form, this technique fits within the standard maximum-likelihood criterion used for training HMMs. The new form of covariance matrix is evaluated on a large-vocabulary speech-recognition task. In initial experiments the performance of the standard system was achieved using approximately half the number of parameters. Moreover, a 10% reduction in word error rate compared to a standard system can be achieved with less than a 1% increase in the number of parameters and little increase in recognition time.
Mark J. F. Gales
IEEE Trans. Speech Audio Process.1
1999 State-based Gaussian selection in large vocabulary continuous speech recognition using HMMs
abstract
This paper investigates the use of Gaussian selection (GS) to increase the speed of a large vocabulary speech recognition system. Typically, 30-70% of the computational time of a continuous density hidden Markov model-based (HMM-based) speech recognizer is spent calculating probabilities. The aim of CS is to reduce this load by selecting the subset of Gaussian component likelihoods that should be computed given a particular input vector. This paper examines new techniques for obtaining "good" Gaussian subsets or "shortlists." All the new schemes make use of state information, specifically, to which state each of the Gaussian components belongs. In this way, a maximum number of Gaussian components per state may be specified, hence reducing the size of the shortlist. The first technique introduced is a simple extension of the standard GS method, which uses this state information. Then, more complex schemes based on maximizing the likelihood of the training data are proposed. These new approaches are compared with the standard GS scheme on a large vocabulary speech recognition task. On this task, the use of state information reduced the percentage of Gaussians computed to 10-15%, compared with 20-30% for the standard GS scheme, with little degradation in performance.
Mark J. F. Gales, Kate M. Knill, Steve J. Young
IEEE Trans. Speech Audio Process.1
1998 Semi-tied covariance matrices
abstract
A standard problem in many classification tasks is how to model feature vectors whose elements are highly correlated. If multi-variate Gaussian distributions are used to model the data then they must have full covariance matrices to accurately do so. This requires a large number of parameters per distribution which restricts the number of distributions that may be robustly estimated, particularly when high dimensional feature vectors are required. This paper describes an alternative to full covariance matrices in these situations. An approximate full covariance matrix is used. The covariance matrix is now split into two elements, one full and one diagonal, which may be tied at completely separate levels. Typically, the full elements are extensively tied, resulting in only a small increase in the number of parameters compared to the diagonal case. Thus dramatically increasing the number of distributions that may be robustly estimated. Simple iterative re-estimation formulae for all the parameters within the standard EM framework are presented. On a large vocabulary speech recognition task a 10% reduction in word error rate over a standard system was achieved.
Mark J. F. Gales
ICASSP1
1998 Cluster adaptive training for speech recognition
Mark J. F. Gales
ICSLP1
1998 Maximum likelihood linear transformations for HMM-based speech recognition
Mark J. F. Gales
Comput. Speech Lang.1
1998 Predictive model-based compensation schemes for robust speech recognition
Mark J. F. Gales
Speech Commun.1
1997 Broadcast news transcription using HTK
abstract
This paper examines the issues in extending a large vocabulary speech recognition system designed for clean and noisy read speech tasks to handle broadcast news transcription. Results using the 1995 DARPA H4 evaluation data set are presented for different front-end analyses and use of unsupervised model adaptation using maximum likelihood linear regression (MLLR). The HTK system for the 1996 H4 evaluation is then described. It includes a number of new features over previous HTK large vocabulary systems including decoder-guided segmentation, segment clustering, cache-based language modelling, and combined MAP and MLLR adaptation. The system runs in multiple passes through the data and the detailed results of each pass are given.
Philip C. Woodland, Mark J. F. Gales, David Pye, Steve J. Young
ICASSP2
1997 Transformation smoothing for speaker and environmental adaptation
abstract
Recently there has been much work done on how to transform HMMs, trained typically in a speaker-independent fashion on clean training data, to be more representative of data from a particular speaker or acoustic environment. These transforms are trained on a small amount of training data, so large numbers of components are required to share the same transform. Normally, each component is constrained to only use one transform. This paper examines how to optimally, in a maximum likelihood sense, assign components to transforms and allow each component, or component grouping, to make use of many transformations. The theory for obtaining both "weights" for each transform and transforms given a set of weights is given. The techniques are evaluated on both speaker and environmental adaptation tasks. 1.
Mark J. F. Gales
EUROSPEECH1
1997 A comparative study of methods for phonetic decision-tree state clustering
abstract
Phonetic decision trees have been widely used for obtaining robust context-dependent models in HMM-based systems. There are five key issues to consider when constructing phonetic decision trees: the alignment of data with the chosen phone classes; the quality of the modeling of the underlying data; the choice of partitioning method at each node; the goodness-of-split criterion and the method for determining appropriate tree sizes. A popular existing method usesefficient but crude approximatemethods for each of these. This paper introduces and evaluates more detailed alternatives to the standard approximations. 1. Introduction A key problem in building continuous-density Hidden Markov Model (HMM)-based context-dependent acoustic models is maintaining a balancebetween the desired model complexity and the number of parameters which can be robustly estimated from the available training data. One solution which has proved successful (eg.[8], [1]) is based upon the use of phonetic decision...
Harriet J. Nock, Mark J. F. Gales, Steve J. Young
EUROSPEECH2
1996 Improving environmental robustness in large vocabulary speech recognition
abstract
This paper describes techniques to improve the robustness of the HTK large vocabulary speech recognition system to non-ideal acoustic environments. The primary methods are single-pass retraining using stereo training data; parallel model combination which combines HMMs trained on clean data with estimates of convolutional and additive noise; and maximum likelihood linear regression which estimates a set of linear transformations of the model parameters to the current conditions. Experiments are reported on both the 1994 ARPA CSR S5 (alternate microphones) and S10 (additive noise) spoken tasks and the 1995 ARPA CSR H3 task (multiple unknown microphones). The HTK system yielded the lowest error rates in both the H3-P0 and HS-C0 tests.
Philip C. Woodland, Mark J. F. Gales, David Pye
ICASSP2
1996 Variance compensation within the MLLR framework for robust speech recognition and speaker adaptation
abstract
This paper investigates the use of maximum likelihood linear regression (MLLR) for both speaker and environment adaptation.MLLR transforms the mean and variance parameters of a set of HMMs.In this paper a number of different types of linear transformations of the variances are examined including full, block diagonal, and diagonal transformation matrices.Experiments on large vocabulary speaker independent data sets are described.On all the data sets examined the use of MLLR mean and variance compensation reduced the error rate compared to mean-only compensation.Furthermore, the use of a block diagonal or full transformation of the variances on the clean data task showed slight improvements over the diagonal case.However, when some environmental mismatch was present there was no difference in performance between using multiple diagonal variance transformations and a more complex single variance transform.
Mark J. F. Gales, David Pye, Philip C. Woodland
ICSLP1
1996 Use of Gaussian selection in large vocabulary continuous speech recognition using HMMs
abstract
This paper investigates the use of Gaussian Selection (GS) to reduce the state likelihood computation in HMM-based systems.These likelihood calculations contribute significantly (30 to 70%) to the computational load.Previously, it has been reported that when GS is used on large systems the recognition accuracy tends to degrade above a 3 reduction in likelihood computation.To explain this degradation, this paper investigates the trade-offs necessary between achieving good state likelihoods and low computation.In addition, the problem of unseen states in a cluster is examined.It is shown that further improvements are possible.For example, using a different assignment measure, with a constraint on the number of components per state per cluster, enabled the recognition accuracy on a 5k speaker-independent task to be maintained up to a 5 reduction in likelihood computation.
Kate M. Knill, Mark J. F. Gales, Steve J. Young
ICSLP2
1996 Iterative unsupervised adaptation using maximum likelihood linear regression
Philip C. Woodland, David Pye, Mark J. F. Gales
ICSLP3
1996 Mean and variance adaptation within the MLLR framework
Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.1
1996 Robust continuous speech recognition using parallel model combination
abstract
This paper addresses the problem of automatic speech recognition in the presence of interfering noise. It focuses on the parallel model combination (PMC) scheme, which has been shown to be a powerful technique for achieving noise robustness. Most experiments reported on PMC to date have been on small, 10-50 word vocabulary systems. Experiments on the Resource Management (RM) database, a 1000 word continuous speech recognition task, reveal compensation requirements not highlighted by the smaller vocabulary tasks. In particular, that it is necessary to compensate the dynamic parameters as well as the static parameters to achieve good recognition performance. The database used for these experiments was the RM speaker independent task with either Lynx Helicopter noise or Operation Room noise from the NOISEX-92 database added. The experiments reported here used the HTK RM recognizer developed at CUED modified to include PMC based compensation for the static, delta and delta-delta parameters. After training on clean speech data, the performance of the recognizer was found to be severely degraded when noise was added to the speech signal at between 10 and 18 dB. However, using PMC the performance was restored to a level comparable with that obtained when training directly in the noise corrupted environment.
Mark J. F. Gales, Steve J. Young
IEEE Trans. Speech Audio Process.1
1995 A fast and flexible implementation of parallel model combination
abstract
In previous papers the use of parallel model combination (PMC) for noise robustness has been described. Various fast implementations have been proposed, though to date in order to compensate all the parameters of a system it has been necessary to perform Gaussian integration. This paper introduces an alternative method that can compensate all the parameters of the recognition system, whilst reducing the computational load of this task. Furthermore, the technique offers an additional degree of flexibility, as it allows the number of components to be chosen and optimised using standard iterative techniques. The new technique is referred to as data-driven PMC (DPMC). It is evaluated on the Resource Management database, with noise artificially added from the NOISEX-92 database. The performance of DPMC is found to be comparable to PMC, at a far lower computational cost. In complex noise environments, by more accurately modelling the noise source, using multiple components, and then reducing the number of components to the original number a slight improvement in performance is obtained.
Mark J. F. Gales, Steve J. Young
ICASSP1
1995 The application of parallel model combination to a large vocabulary dictation task
Mark J. F. Gales, Steve J. Young
EUROSPEECH1
1995 Robust speech recognition in additive and convolutional noise using parallel model combination
Mark J. F. Gales, Steve J. Young
Comput. Speech Lang.1
1994 Parallel model combination on a noise corrupted resource management task
Mark J. F. Gales, Steve J. Young
ICSLP1
1993 HMM recognition in noise using parallel model combination
Mark J. F. Gales, Steve J. Young
EUROSPEECH1
1993 Segmental hidden Markov models
Mark J. F. Gales, Steve J. Young
EUROSPEECH1
1993 Cepstral parameter compensation for HMM recognition in noise
Mark J. F. Gales, Steve J. Young
Speech Commun.1
1992 An improved approach to the hidden Markov model decomposition of speech and noise
abstract
The author addresses the problem of automatic speech recognition in the presence of interfering noise. The novel approach described decomposes the contaminated speech signal using a generalization of standard hidden Markov modeling, while utilizing a compact and effective parametrization of the speech signal. The technique is compared to some existing noise compensation techniques, using data recorded in noise, and is found to have improved performance compared to existing model decomposition techniques. Performance is comparable to existing noise subtraction techniques, but the technique is applicable to a wider range of noise environments and is not dependent on an accurate endpointing of the speech.>
Mark J. F. Gales, Steve J. Young
ICASSP1