Kate M. Knill

dblp:06/9262 · also Katherine M. Knill · DBLP profile ↗
← Back
76ranked-venue papers
9as first author
19since 2021 · last 2025
0000-0003-1292-2769ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 65 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 55 · 7 first-author · 16 since 2021
YearPublicationVenuePosition
2025 Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
abstract
Chain-of-Thought (CoT) prompting is a widely used method to improve the reasoning capability of Large Language Models (LLMs). More recently, CoT has been leveraged in Knowledge Distillation (KD) to transfer reasoning capability from a larger LLM to a smaller one. This paper examines the role of CoT in distilling the reasoning capability from larger LLMs to smaller LLMs using white-box KD, analyzing its effectiveness in improving the performance of the distilled models for various natural language reasoning and understanding tasks. We conduct white-box KD experiments using LLMs from the Qwen and Llama2 families, employing CoT data from the CoT-Collection dataset. The distilled models are then evaluated on natural language reasoning and understanding tasks from the BIG-Bench-Hard (BBH) benchmark, which presents complex challenges for smaller LLMs. Experimental results demonstrate the role of CoT in improving white-box KD effectiveness, enabling the distilled models to achieve better average performance in natural language reasoning and understanding tasks from BBH.
Cong-Thanh Do, Rama Sanand Doddipatla, Kate M. Knill
INLG3
2025 Assessment of L2 Oral Proficiency using Speech Large Language Models
Rao Ma, Mengjie Qian 0001, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH5
2025 Training Articulatory Inversion Models for Interspeaker Consistency
Charles McGhee, Mark J. F. Gales, Kate M. Knill
INTERSPEECH3
2025 Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Mengjie Qian 0001, Rao Ma, Stefano Bannò, Kate M. Knill, Mark J. F. Gales
INTERSPEECH4
2024 Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
abstract
Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications.These systems incorporate 'special tokens' in their vocabulary, such as <|endoftext|>, to guide their language generation process.However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior.We propose a simple yet effective method to learn a universal acoustic realization of Whisper's <|endoftext|> token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively 'muting' the model.Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97% of speech samples.Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks.Overall this work demonstrates the vulnerability of Whisper models to 'muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data. 1
Vyas Raina, Rao Ma, Charles McGhee, Kate M. Knill, Mark J. F. Gales
EMNLP4
2024 Towards End-to-End Spoken Grammatical Error Correction
abstract
Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.
Stefano Bannò, Rao Ma, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
ICASSP4
2024 On the Usefulness of Speaker Embeddings for Speaker Retrieval in the Wild: A Comparative Study of x-vector and ECAPA-TDNN Models
Erfan Loweimi, Mengjie Qian 0001, Kate M. Knill, Mark J. F. Gales
INTERSPEECH3
2024 Highly Intelligible Speaker-Independent Articulatory Synthesis
Charles McGhee, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2024 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
Mengjie Qian 0001, Rao Ma, Kate M. Knill, Mark J. F. Gales
INTERSPEECH4
2024 Investigating the Emergent Audio Classification Ability of ASR Foundation Models
abstract
Rao Ma, Adian Liusie, Mark Gales, Kate Knill. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Rao Ma, Adian Liusie, Mark J. F. Gales, Kate M. Knill
NAACL-HLT4
2024 Zero-Shot Audio Topic Reranking Using Large Language Models
abstract
Multimodal Video Search by Examples (MVSE) investigates using video clips as the query term for information retrieval, rather than the more traditional text query. This enables far richer search modalities such as images, speaker, content, topic, and emotion. A key element for this process is highly rapid and flexible search to support large archives, which in MVSE is facilitated by representing video attributes with embeddings. This work aims to compensate for any performance loss from this rapid archive search by examining reranking approaches. In particular, zero-shot reranking methods using large language models (LLMs) are investigated as these are applicable to any video archive audio content. Performance is evaluated for topic-based retrieval on a publicly available video archive, the BBC Rewind corpus. Results demonstrate that reranking significantly improves retrieval ranking without requiring any task-specific in-domain training data. Furthermore, three sources of information (ASR transcriptions, automatic summaries and synopses) as input for LLM reranking were compared. To gain a deeper understanding and further insights into the performance differences and limitations of these text sources, we employ a fact-checking approach to analyse the information consistency among them.
Mengjie Qian 0001, Rao Ma, Adian Liusie, Erfan Loweimi, Kate M. Knill, Mark J. F. Gales
SLT5
2024 Entity Resolution in Situated Dialog With Unimodal and Multimodal Transformers
abstract
In this work we address the entity resolution task for situated multimodal dialog investigating how a unimodal approach, which uses only textual information as input (representing visual attributes as text), compares to a multimodal system, which processes both text and visual information. We analyze two of the top performing models presented in the Tenth Dialog Systems Technology Challenge and propose modifications that enhance their performance on the multimodal coreference resolution task. We evaluate these approaches on in- and out-of-domain settings by training the models on the fashion domain and testing on the furniture domain, and vice-versa, to assess the generalizability of the models. Through systematic analysis, we show that while both systems achieve similar performance on in-domain scenarios, the multimodal system generalizes better to out-of-domain settings. A combination strategy of enhanced unimodal and multimodal systems achieves F1 = 0.80 (5% absolute gain compared to the best performing system). Finally, human performance on the same task is evaluated on a small subset, suggesting that the performance of the current automatic models is on par with people on this task.
Alejandro Santorum Varela, Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla, Kate M. Knill
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
abstract
Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions.Most prior works use the 1-best ASR hypothesis as input and therefore can only perform correction by leveraging the context within one sentence.In this work, we propose a novel N-best T5 model for this task, which is fine-tuned from a T5 model and utilizes ASR N-best lists as model input.By transferring knowledge from the pretrained language model and obtaining richer information from the ASR decoding space, the proposed approach outperforms a strong Conformer-Transducer baseline.Another issue with standard error correction is that the generation process is not well-guided.To address this a constrained decoding process, either based on the N-best list or an ASR lattice, is used which allows additional information to be propagated.
Rao Ma, Mark J. F. Gales, Kate M. Knill, Mengjie Qian 0001
INTERSPEECH3
2023 Adapting an Unadaptable ASR System
abstract
As speech recognition model sizes and training data requirements grow, it is increasingly common for systems to only be available via APIs from online service providers rather than having direct access to models themselves.In this scenario it is challenging to adapt systems to a specific target domain.To address this problem we consider the recently released OpenAI Whisper ASR as an example of a large-scale ASR system to assess adaptation methods.An error correction based approach is adopted, as this does not require access to the model, but can be trained from either 1-best or N-best outputs that are normally available via the ASR API.LibriSpeech is used as the primary target domain for adaptation.The generalization ability of the system in two distinct dimensions are then evaluated.First, whether the form of correction model is portable to other speech recognition domains, and secondly whether it can be used for ASR models having a different architecture.
Rao Ma, Mengjie Qian 0001, Mark J. F. Gales, Kate M. Knill
INTERSPEECH4
2023 Speak & Improve: L2 English Speaking Practice Tool
Diane Nicholls, Kate M. Knill, Mark J. F. Gales, Anton Ragni, Paul Ricketts
INTERSPEECH2
2022 View-Specific Assessment of L2 Spoken English
abstract
The growing demand for learning English as a second language has increased interest in automatic approaches for assessing and improving spoken language proficiency. A significant challenge in this field is to provide interpretable scores and informative feedback to learners through individual viewpoints of learners’ proficiency, as opposed to holistic scores. Thus far, holistic scoring remains commonly applied in large-scale commercial tests. As a result, an issue with more detailed evaluation is that human graders are generally trained to provide holistic scores. This paper investigates whether view-specific systems can be trained when only holistic scores are available. To enable this process, view-specific networks are defined where both their inputs and structure are adapted to focus on specific facets of proficiency. It is shown that it is possible to train such systems on holistic scores, such that they provide view-specific scores at evaluation time. View-specific networks are designed in this way for pronunciation, rhythm, text, use of parts of speech and grammatical accuracy. The relationships between the predictions of each system are investigated on the spoken part of the Linguaskill proficiency test. It is shown that the view-specific predictions are complementary in nature and capture different information about proficiency.
Stefano Bannò, Bhanu Balusu, Mark J. F. Gales, Kate M. Knill, Konstantinos Kyriakopoulos
INTERSPEECH4
2022 Increasing Context for Estimating Confidence Scores in Automatic Speech Recognition
abstract
Accurate confidence measures for predictions from machine learning techniques play a critical role in the deployment and training of many speech and language processing applications. For example, confidence scores are important when making use of automatically generated transcriptions in training automatic speech recognition (ASR) systems, as well as down-stream applications, such as information retrieval and conversational assistants. Previous work on improving confidence scores for these systems has focused on two main directions: designing features correlated with improved confidence prediction; and employing sequence models to account for the importance of contextual information. Few studies, however, have explored incorporating contextual information more broadly, such as from the future, in addition to the past, or making use of alternative multiple hypotheses in addition to the most likely one. This article introduces two general approaches for encapsulating contextual information from lattices. Experimental results illustrating the importance of increasing contextual information for estimating confidence scores are presented on a range of limited resource languages where word error rates range between 30% and 60%. The results show that the novel approaches provide significant gains in the accuracy of confidence estimation.
Anton Ragni, Mark J. F. Gales, Oliver Rose, Kate M. Knill, Alexandros Kastanos, Qiujia Li, Preben Ness
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Analysing Bias in Spoken Language Assessment Using Concept Activation Vectors
abstract
A significant concern with deep learning based approaches is that they are difficult to interpret, which means detecting bias in network predictions can be challenging. Concept Activation Vectors (CAVs) have been proposed to address this problem. These use representations - perturbations of activation function outputs - of interpretable concepts to analyse how the network is influenced by the concept. This work applies CAVs to assess bias in a spoken language assessment (SLA) system, a regression task. One of the challenges with SLA is the wide range of concepts that can introduce bias in training data, for example L1, age, acoustic conditions, and particular human graders, or the grading instructions. Simply generating large quantities of expert marked data to check for all forms of bias is impractical. This paper uses CAVs applied to the training data to identify concepts that might be of concern, allowing a more targeted dataset to be collected to assess bias. The ability of CAVs to detect bias is assessed on the BULATS speaking test using both a standard system and a system to which bias was artificially introduced. A strong bias identified by CAVs on the training data matches the bias observed in expert marked held-out test data.
Xizi Wei, Mark J. F. Gales, Kate M. Knill
ICASSP3
2021 ETLT 2021: Shared Task on Automatic Speech Recognition for Non-Native Children's Speech
Roberto Gretter, Marco Matassoni, Daniele Falavigna, A. Misra, Chee Wee Leong, Kate M. Knill
Interspeech6
2020 Grammatical error detection in transcriptions of spoken English
abstract
We describe the collection of transcription corrections and grammatical error annotations for the CROWDED Corpus of spoken English monologues on business topics.The corpus recordings were crowdsourced from native speakers of English and learners of English with German as their first language.The new transcriptions and annotations are obtained from different crowdworkers: we analyse the 1108 new crowdworker submissions and propose that they can be used for automatic transcription post-editing and grammatical error correction for speech.To further explore the data we train grammatical error detection models with various configurations including pretrained and contextual word representations as input, additional features and auxiliary objectives, and extra training data from written error-annotated corpora.We find that a model concatenating pre-trained and contextual word representations as input performs best, and that additional information does not lead to further performance gains.
Andrew Caines, Christian Bentz, Kate M. Knill, Marek Rei, Paula Buttery
COLING3
2020 Non-Native Children's Automatic Speech Recognition: The INTERSPEECH 2020 Shared Task ALTA Systems
abstract
Automatic spoken language assessment (SLA) is a challenging problem due to the large variations in learner speech combined with limited resources. These issues are even more problematic when considering children learning a language, with higher levels of acoustic and lexical variability, and of code-switching compared to adult data. This paper describes the ALTA system for the INTERSPEECH 2020 Shared Task on Automatic Speech Recognition for Non-Native Children’s Speech. The data for this task consists of examination recordings of Italian school children aged 9-16, ranging in ability from minimal, to basic, to limited but effective command of spoken English. A variety of systems were developed using the limited training data available, 49 hours. State-of-the-art acoustic models and language models were evaluated, including a diversity of lexical representations, handling code-switching and learner pronunciation errors, and grade specific models. The best single system achieved a word error rate (WER) of 16.9% on the evaluation data. By combining multiple diverse systems, including both grade independent and grade specific models, the error rate was reduced to 15.7%. This combined system was the best performing submission for both the closed and open tasks.
Kate M. Knill, Yu Wang 0027, Xixin Wu, Mark J. F. Gales
INTERSPEECH1
2020 Automatic Detection of Accent and Lexical Pronunciation Errors in Spontaneous Non-Native English Speech
abstract
Detecting individual pronunciation errors and diagnosing pronunciation error tendencies in a language learner based on their speech are important components of computer-aided language learning (CALL). The tasks of error detection and error tendency diagnosis become particularly challenging when the speech in question is spontaneous and particularly given the challenges posed by the inconsistency of human annotation of pronunciation errors. This paper presents an approach to these tasks by distinguishing between lexical errors, wherein the speaker does not know how a particular word is pronounced, and accent errors, wherein the candidate's speech exhibits consistent patterns of phone substitution, deletion and insertion. Three annotated corpora of non-native English speech by speakers of multiple L1s are analysed, the consistency of human annotation investigated and a method presented for detecting individual accent and lexical errors and diagnosing accent error tendencies at the speaker level.
Konstantinos Kyriakopoulos, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2020 Universal Adversarial Attacks on Spoken Language Assessment Systems
abstract
There is an increasing demand for automated spoken language assessment (SLA) systems, partly driven by the performance improvements that have come from deep learning based approaches. One aspect of deep learning systems is that they do not require expert derived features, operating directly on the original signal such as a speech recognition (ASR) transcript. This, however, increases their potential susceptibility to adversarial attacks as a form of candidate malpractice. In this paper the sensitivity of SLA systems to a universal black-box attack on the ASR text output is explored. The aim is to obtain a single, universal phrase to maximally increase any candidate's score. Four approaches to detect such adversarial attacks are also described. All the systems, and associated detection approaches, are evaluated on a free (spontaneous) speaking section from a Business English test. It is shown that on deep learning based SLA systems the average candidate score can be increased by almost one grade level using a single six word phrase appended to the end of the response hypothesis. Although these large gains can be obtained, they can be easily detected based on detection shifts from the scores of a “traditional” Gaussian Process based grader.
Vyas Raina, Mark J. F. Gales, Kate M. Knill
INTERSPEECH3
2020 Ensemble Approaches for Uncertainty in Spoken Language Assessment
abstract
Deep learning has dramatically improved the performance of automated systems on a range of tasks including spoken language assessment. One of the issues with these deep learning approaches is that they tend to be overconfident in the decisions that they make, with potentially serious implications for deployment of systems for high-stakes examinations. This paper examines the use of ensemble approaches to improve both the reliability of the scores that are generated, and the ability to detect where the system has made predictions beyond acceptable errors. In this work assessment is treated as a regression problem. Deep density networks, and ensembles of these models, are used as the predictive models. Given an ensemble of models measures of uncertainty, for example the variance of the predicted distributions, can be obtained and used for detecting outlier predictions. However, these ensemble approaches increase the computational and memory requirements of the system. To address this problem the ensemble is distilled into a single mixture density network. The performance of the systems is evaluated on a free speaking prompt-response style spoken language assessment test. Experiments show that the ensembles and the distilled model yield performance gains over a single model, and have the ability to detect outliers.
Xixin Wu, Kate M. Knill, Mark J. F. Gales, Andrey Malinin
INTERSPEECH2
2019 Automatic Grammatical Error Detection of Non-native Spoken Learner English
abstract
Automatic language assessment and learning systems are required to support the global growth in English language learning. They need to be able to provide reliable and meaningful feedback to help learners develop their skills. This paper considers the question of detecting "grammatical" errors in non-native spoken English as a first step to providing feedback on a learner's use of the language. A state-of-the-art deep learning based grammatical error detection (GED) system designed for written texts is investigated on free speaking tasks across the full range of proficiency grades with a mix of first languages (L1s). This presents a number of challenges. Free speech contains disfluencies that disrupt the spoken language flow but are not grammatical errors. The lower the level of the learner the more these both will occur which makes the underlying task of automatic transcription harder. The baseline written GED system is seen to perform less well on manually transcribed spoken language. When the GED model is fine-tuned to free speech data from the target domain the spoken system is able to match the written performance. Given the current state-of-the-art in ASR, however, and the ability to detect disfluencies grammatical error feedback from automated transcriptions remains a challenge.
Kate M. Knill, Mark J. F. Gales, P. P. Manakul, Andrew Caines
ICASSP1
2019 A Deep Learning Approach to Automatic Characterisation of Rhythm in Non-Native English Speech
abstract
A speaker's rhythm contributes to the intelligibility of their speech and can be characteristic of their language and accent. For non-native learners of a language, the extent to which they match its natural rhythm is an important predictor of their proficiency. As a learner improves, their rhythm is expected to become less similar to their L1 and more to the L2. Metrics based on the variability of the durations of vocalic and consonantal intervals have been shown to be effective at detecting language and accent. In this paper, pairwise variability (PVI, CCI) and variance (varcoV, varcoC) metrics are first used to predict proficiency and L1 of non-native speakers taking an English spoken exam. A deep learning alternative to generalise these features is then presented, in the form of a tunable duration embedding, based on attention over an RNN over durations. The RNN allows relationships beyond pairwise to be captured, while attention allows sensitivity to the different relative importance of durations. The system is trained end-to-end for proficiency and L1 prediction and compared to the baseline. The values of both sets of features for different proficiency levels are then visualised and compared to native speech in the L1 and the L2.
Konstantinos Kyriakopoulos, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2019 Impact of ASR Performance on Spoken Grammatical Error Detection
abstract
Computer assisted language learning (CALL) systems aidlearners to monitor their progress by providing scoring andfeedback on language assessment tasks. Free speaking tests al-low assessment of what a learner has said, as well as how theysaid it. For these tasks, Automatic Speech Recognition (ASR)is required to generate transcriptions of a candidate’s responses,the quality of these transcriptions is crucial to provide reliablefeedback in downstream processes. This paper considers theimpact of ASR performance on Grammatical Error Detection(GED) for free speaking tasks, as an example of providing feed-back on a learner’s use of English. The performance of an ad-vanced deep-learning based GED system, initially trained onwritten corpora, is used to evaluate the influence of ASR errors.One consequence of these errors is that grammatical errors canresult from incorrect transcriptions as well as learner errors, thismay yield confusing feedback. To mitigate the effect of theseerrors, and reduce erroneous feedback, ASR confidence scoresare incorporated into the GED system. By additionally adaptingthe written text GED system to the speech domain, using ASRtranscriptions, significant gains in performance can be achieved.Analysis of the GED performance for different grammatical er-ror types and across grade is also presented.
Yiting Lu, Mark J. F. Gales, Kate M. Knill, P. P. Manakul, Yu Wang 0027
INTERSPEECH3
2019 Splash: Speech and Language Assessment in Schools and Homes
Avin Miwardelli, Ian Gallagher, Jenny Gibson, Napoleon Katsos, Kate M. Knill, Helena Wood
INTERSPEECH5
2018 Impact of ASR Performance on Free Speaking Language Assessment
abstract
In free speaking tests candidates respond in spontaneous speech to prompts. This form of test allows the spoken language proficiency of a non-native speaker of English to be assessed more fully than read aloud tests. As the candidate's responses are unscripted, transcription by automatic speech recognition (ASR) is essential for automated assessment. ASR will never be 100% accurate so any assessment system must seek to minimise and mitigate ASR errors. This paper considers the impact of ASR errors on the performance of free speaking test auto-marking systems. Firstly rich linguistically related features, based on part-of-speech tags from statistical parse trees, are investigated for assessment. Then, the impact of ASR errors on how well the system can detect whether a learner's answer is relevant to the question asked is evaluated. Finally, the impact that these errors may have on the ability of the system to provide detailed feedback to the learner is analysed. In particular, pronunciation and grammatical errors are considered as these are important in helping a learner to make progress. As feedback resulting from an ASR error would be highly confusing, an approach to mitigate this problem using confidence scores is also analysed.
Kate M. Knill, Mark J. F. Gales, Konstantinos Kyriakopoulos, Andrey Malinin, Anton Ragni, Yu Wang 0027, Andrew Caines
INTERSPEECH1
2018 A Deep Learning Approach to Assessing Non-native Pronunciation of English Using Phone Distances
abstract
The way a non-native speaker pronounces the phones of a language is an important predictor of their proficiency. In grading spontaneous speech, the pairwise distances between generative statistical models trained on each phone have been shown to be powerful features. This paper presents a deep learning alternative to model-based phone distances in the form of a tunable Siamese network feature extractor to extract distance metrics directly from the audio frame sequence. Features are extracted at the phone instance level and combined to phone-level representations using an attention mechanism. Pair-wise distances between phone features are then projected through a feed-forward layer to predict score. The extraction stage is initialised on either a binary phone instance-pair classification task, or to mimic the model-based features, then the whole system is fine-tuned end-to-end, optimising the learning of the distance metric to the score prediction task. This method is therefore more adaptable and more sensitive to phone instance level phenomena. Its performance is compared against
Konstantinos Kyriakopoulos, Kate M. Knill, Mark J. F. Gales
INTERSPEECH2
2018 Sequence Teacher-Student Training of Acoustic Models for Automatic Free Speaking Language Assessment
abstract
A high performance automatic speech recognition (ASR) system is an important constituent component of an automatic language assessment system for free speaking language tests. The ASR system is required to be capable of recognising non-native spontaneous English speech and to be deployable under real-time conditions. The performance of ASR systems can often be significantly improved by leveraging upon multiple systems that are complementary, such as an ensemble. Ensemble methods, however, can be computationally expensive, often requiring multiple decoding runs, which makes them impractical for deployment. In this paper, a lattice-free implementation of sequence-level teacher-student training is used to reduce this computational cost, thereby allowing for real-time applications. This method allows a single student model to emulate the performance of an ensemble of teachers, but without the need for multiple decoding runs. Adaptations of the student model to speakers from different first languages (L1s) and grades are also explored.
Yu Wang 0027, Jeremy H. M. Wong, Mark J. F. Gales, Kate M. Knill, Anton Ragni
SLT4
2018 Towards automatic assessment of spontaneous spoken English
Yu Wang 0027, Mark J. F. Gales, Kate M. Knill, Konstantinos Kyriakopoulos, Andrey Malinin, Rogier C. van Dalen, M. Rashid
Speech Commun.3
2017 A hierarchical attention based model for off-topic spontaneous spoken response detection
abstract
Automatic spoken language assessment and training systems are becoming increasingly popular to handle the growing demand to learn languages. However, current systems often assess only fluency and pronunciation, with limited content-based features being used. This paper examines one particular aspect of content-assessment, off-topic response detection. This is important for deployed systems as it ensures that candidates understood the prompt, and are able to generate an appropriate answer. Previously proposed approaches typically require a set of prompt-response training pairs, which limits flexibility as example responses are required whenever a new test prompt is introduced. Recently, the attention based neural topic model (ATM) was presented, which can assess the relevance of prompt-response pairs regardless of whether the prompt was seen in training. This model uses a bidirectional Recurrent Neural Network (BiRNN) embedding of the prompt combined with an attention mechanism to attend over the hidden states of a BiRNN embedding of the response to compute a fixed-length embedding used to predict relevance. Unfortunately, performance on prompts not seen in the training data is lower than on seen prompts. Thus, this paper adds the following contributions: several improvements to the ATM are examined; a hierarchical variant of the ATM (HATM) is proposed, which explicitly uses prompt similarity to further improve performance on unseen prompts by interpolating over prompts seen in training data given a prompt of interest via a second attention mechanism; an in-depth analysis of both models is conducted and main failure mode identified. On spontaneous spoken data, taken from BULATS tests, these systems are able to assess relevance to both seen and unseen prompts.
Andrey Malinin, Kate M. Knill, Mark J. F. Gales
ASRU2
2017 Recurrent neural network language models for keyword search
abstract
Recurrent neural network language models (RNNLMs) have becoming increasingly popular in many applications such as automatic speech recognition (ASR). Significant performance improvements in both perplexity and word error rate over standard n-gram LMs have been widely reported on ASR tasks. In contrast, published research on using RNNLMs for keyword search systems has been relatively limited. In this paper the application of RNNLMs for the IARPA Babel keyword search task is investigated. In order to supplement the limited acoustic transcription data, large amounts of web texts are also used in large vocabulary design and LM training. Various training criteria were then explored to improved RNNLMs' efficiency in both training and evaluation. Significant and consistent improvements on both keyword search and ASR tasks were obtained across all languages.
Xie Chen 0001, Anton Ragni, J. Vasilakes, Xunying Liu, Kate M. Knill, Mark J. F. Gales
ICASSP5
2017 Morph-to-word transduction for accurate and efficient automatic speech recognition and keyword search
abstract
Word units are a popular choice in statistical language modelling. For inflective and agglutinative languages this choice may result in a high out of vocabulary rate. Subword units, such as morphs, provide an interesting alternative to words. These units can be derived in an unsupervised fashion and empirically show lower out of vocabulary rates. This paper proposes a morph-to-word transduction to convert morph sequences into word sequences. This enables powerful word language models to be applied. In addition, it is expected that techniques such as pruning, confusion network decoding, keyword search and many others may benefit from word rather than morph level decision making. However, word or morph systems alone may not achieve optimal performance in tasks such as keyword search so a combination is typically employed. This paper proposes a single index approach that enables word, morph and phone searches to be performed over a single morph index. Experiments are conducted on IARPA Babel program languages including the surprise languages of the OpenKWS 2015 and 2016 competitions.
Anton Ragni, Danielle Saunders, P. Zahemszky, J. Vasilakes, Mark J. F. Gales, Kate M. Knill
ICASSP6
2017 Stimulated training for automatic speech recognition and keyword search in limited resource conditions
abstract
Training neural network acoustic models on limited quantities of data is a challenging task. A number of techniques have been proposed to improve generalisation. This paper investigates one such technique called stimulated training. It enables standard criteria such as cross-entropy to enforce spatial constraints on activations originating from different units. Having different regions being active depending on the input unit may help network to discriminate better and as a consequence yield lower error rates. This paper investigates stimulated training for automatic speech recognition of a number of languages representing different families, alphabets, phone sets and vocabulary sizes. In particular, it looks at ensembles of stimulated networks to ensure that improved generalisation will withstand system combination effects. In order to assess stimulated training beyond 1-best transcription accuracy, this paper looks at keyword search as a proxy for assessing quality of lattices. Experiments are conducted on IARPA Babel program languages including the surprise language of OpenKWS 2016 competition.
Anton Ragni, Chunyang Wu, Mark J. F. Gales, J. Vasilakes, Kate M. Knill
ICASSP5
2017 Use of Graphemic Lexicons for Spoken Language Assessment
abstract
Copyright © 2017 ISCA. Automatic systems for practice and exams are essential to support the growing worldwide demand for learning English as an additional language. Assessment of spontaneous spoken English is, however, currently limited in scope due to the difficulty of achieving sufficient automatic speech recognition (ASR) accuracy. "Off-the-shelf" English ASR systems cannot model the exceptionally wide variety of accents, pronunications and recording conditions found in non-native learner data. Limited training data for different first languages (L1s), across all proficiency levels, often with (at most) crowd-sourced transcriptions, limits the performance of ASR systems trained on non-native English learner speech. This paper investigates whether the effect of one source of error in the system, lexical modelling, can be mitigated by using graphemic lexicons in place of phonetic lexicons based on native speaker pronunications. Graphemicbased English ASR is typically worse than phonetic-based due to the irregularity of English spelling-to-pronunciation but here lower word error rates are consistently observed with the graphemic ASR. The effect of using graphemes on automatic assessment is assessed on different grader feature sets: audio and fluency derived features, including some phonetic level features; and phone/grapheme distance features which capture a measure of pronunciation ability.
Kate M. Knill, Mark J. F. Gales, Konstantinos Kyriakopoulos, Anton Ragni, Yu Wang 0027
INTERSPEECH1
2016 Off-topic Response Detection for Spontaneous Spoken English Assessment
abstract
Automatic spoken language assessment systems are becoming increasingly important to meet the demand for English second language learning.This is a challenging task due to the high error rates of, even state-of-the-art, non-native speech recognition.Consequently current systems primarily assess fluency and pronunciation.However, content assessment is essential for full automation.As a first stage it is important to judge whether the speaker responds on topic to test questions designed to elicit spontaneous speech.Standard approaches to off-topic response detection assess similarity between the response and question based on bag-of-words representations.An alternative framework based on Recurrent Neural Network Language Models (RNNLM) is proposed in this paper.The RNNLM is adapted to the topic of each test question.It learns to associate example responses to questions with points in a topic space constructed using these example responses.Classification is done by ranking the topic-conditional posterior probabilities of a response.The RNNLMs associate a broad range of responses with each topic, incorporate sequence information and scale better with additional training data, unlike standard methods.On experiments conducted on data from the Business Language Testing Service (BULATS) this approach outperforms standard approaches.
Andrey Malinin, Rogier C. van Dalen, Kate M. Knill, Yu Wang 0027, Mark J. F. Gales
ACL (1)3
2016 Multi-Language Neural Network Language Models
abstract
Recently there has been a lot of interest in neural network based language models. These models typically consist of vocabulary dependent input and output layers and one or more vocabulary independent hidden layers. One standard issue with these approaches is that large quantities of training data are needed to ensure robust parameter estimates. This poses a significant problem when only limited data is available. One possible way to address this issue is augmentation: model-based, in the form of language model interpolation, and data-based, in the form of data augmentation. However, these approaches may not always be possible to use due to vocabulary dependent input and output layers. This seriously restricts the nature of the data possible to use in augmentation. This paper describes a general solution whereby only one or more vocabulary independent hidden layers are augmented. Such approach makes it possible to examine augmentation from previously impossible domains. Moreover, this approach paves a direct way for multi-task learning with these models. As a proof of the concept this paper examines the use of multilingual data for augmenting hidden layers of recurrent neural network language models. Experiments are conducted using a set of language packs released within IARPA Babel program.
Anton Ragni, Edgar Dakin, Xie Chen 0001, Mark J. F. Gales, Kate M. Knill
INTERSPEECH5
2016 Log-Linear System Combination Using Structured Support Vector Machines
abstract
Building high accuracy speech recognition systems with limited language resources is a highly challenging task. Although the use of multi-language data for acoustic models yields improvements, performance is often unsatisfactory with highly limited acoustic training data. In these situations, it is possible to consider using multiple well trained acoustic models and combine the system outputs together. Unfortunately, the computational cost associated with these approaches is high as multiple decoding runs are required. To address this problem, this paper examines schemes based on log-linear score combination. This has a number of advantages over standard combination schemes. Even with limited acoustic training data, it is possible to train, for example, phone-specific combination weights, allowing detailed relationships between the available well trained models to be obtained. To ensure robust parameter estimation, this paper casts log-linear score combination into a structured support vector machine (SSVM) learning task. This yields a method to train model parameters with good generalisation properties. Here the SSVM feature space is a set of scores from well-trained individual systems. The SSVM approach is compared to lattice rescoring and confusion network combination using language packs released within the IARPA Babel program.
Anton Ragni, Mark J. F. Gales, Kate M. Knill
INTERSPEECH4
2016 Towards Using Conversations with Spoken Dialogue Systems in the Automated Assessment of Non-Native Speakers of English
abstract
Diane Litman, Steve Young, Mark Gales, Kate Knill, Karen Ottewell, Rogier van Dalen, David Vandyke. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016.
Diane J. Litman, Steve J. Young, Mark J. F. Gales, Kate M. Knill, Karen Ottewell, Rogier C. van Dalen, David Vandyke
SIGDIAL Conference4
2015 Multilingual representations for low resource speech recognition and keyword search
abstract
This paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation.
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland
ASRU16
2015 Improving multiple-crowd-sourced transcriptions using a speech recogniser
abstract
This paper introduces a method to produce high-quality transcriptions of speech data from only two crowd-sourced transcriptions. These transcriptions, produced cheaply by people on the Internet, for example through Amazon Mechanical Turk, are often of low quality. Often, multiple crowd-sourced transcriptions are combined to form one transcription of higher quality. However, the state of the art is to use essentially a form of majority voting, which requires at least three transcriptions for each utterance. This paper shows how to refine this approach to work with only two transcriptions. It then introduces a method that uses a speech recogniser (bootstrapped on a simple combination scheme) to combine transcriptions. When only two crowd-sourced transcriptions are available, on a noisy data set this improves the word error rate to gold-standard transcriptions by 21% relative.
Rogier C. van Dalen, Kate M. Knill, Pirros Tsiakoulis, Mark J. F. Gales
ICASSP2
2015 Unicode-based graphemic systems for limited resource languages
abstract
Large vocabulary continuous speech recognition systems require a mapping from words, or tokens, into sub-word units to enable robust estimation of acoustic model parameters, and to model words not seen in the training data. The standard approach to achieve this is to manually generate a lexicon where words are mapped into phones, often with attributes associated with each of these phones. Contextdependent acoustic models are then constructed using decision trees where questions are asked based on the phones and phone attributes. For low-resource languages, it may not be practical to manually generate a lexicon. An alternative approach is to use a graphemic lexicon, where the “pronunciation” for a word is defined by the letters forming that word. This paper proposes a simple approach for building graphemic systems for any language written in unicode. The attributes for graphemes are automatically derived using features from the unicode character descriptions. These attributes are then used in decision tree construction. This approach is examined on the IARPA Babel Option Period 2 languages, and a Levantine Arabic CTS task. The described approach achieves comparable, and complementary, performance to phonetic lexicon-based approaches.
Mark J. F. Gales, Kate M. Knill, Anton Ragni
ICASSP2
2015 A language space representation for speech recognition
abstract
The number of languages for which speech recognition systems have become available is growing each year. This paper proposes to view languages as points in some rich space, termed language space, where bases are eigen-languages and a particular selection of the projection determines points. Such an approach could not only reduce development costs for each new language but also provide automatic means for language analysis. For the initial proof of the concept, this paper adopts cluster adaptive training (CAT) known for inducing similar spaces for speaker adaptation needs. The CAT approach used in this paper builds on the previous work for language adaptation in speech synthesis and extends it to Gaussian mixture modelling more appropriate for speech recognition. Experiments conducted on IARPA Babel program languages show that such language space representations can outperform language independent models and discover closely related languages in an automatic way.
Anton Ragni, Mark J. F. Gales, Kate M. Knill
ICASSP3
2015 Improving speech recognition and keyword search for low resource languages using web data
abstract
We describe the use of text data scraped from the web to augment language models for Automatic Speech Recognition and Keyword Search for Low Resource Languages.We scrape text from multiple genres including blogs, online news, translated TED talks, and subtitles.Using linearly interpolated language models, we find that blogs and movie subtitles are more relevant for language modeling of conversational telephone speech and obtain large reductions in out-of-vocabulary keywords.Furthermore, we show that the web data can improve Term Error Rate Performance by 3.8% absolute and Maximum Term-Weighted Value in Keyword Search by 0.0076-0.1059absolute points.Much of the gain comes from the reduction of out-of-vocabulary items.
Gideon Mendels, Erica Cooper, Victor Soto, Julia Hirschberg, Mark J. F. Gales, Kate M. Knill, Anton Ragni
INTERSPEECH6
2015 Joint decoding of tandem and hybrid systems for improved keyword spotting on low resource languages
abstract
Copyright © 2015 ISCA. Keyword spotting (KWS) for low-resource languages has drawn increasing attention in recent years. The state-of-the-art KWS systems are based on lattices or Confusion Networks (CN) generated by Automatic Speech Recognition (ASR) systems. It has been shown that considerable KWS gains can be obtained by combining the keyword detection results from different forms of ASR systems, e.g., Tandem and Hybrid systems. This paper investigates an alternative combination scheme for KWS using joint decoding. This scheme treats a Tandem system and a Hybrid system as two separate streams, and makes a linear combination of individual acoustic model log-likelihoods. Joint decoding is more efficient as it requires just a single pass of decoding and a single pass of keyword search. Experiments on six Babel OP2 development languages show that joint decoding is capable of providing consistent gains over each individual system. Moreover, it is possible to efficiently rescore the joint decoding lattices with Tandem or Hybrid acoustic models, and further KWS gains can be obtained by merging the detection posting lists from the joint decoding lattices and rescored lattices.
Anton Ragni, Mark J. F. Gales, Kate M. Knill, Philip C. Woodland, Chao Zhang 0031
INTERSPEECH4
2014 An initial investigation of long-term adaptation for meeting transcription
abstract
Meeting transcription is a very useful and challenging task. The majority of research to date has focused on individual meeting, or only a small group of meetings. In many practical deploy-ments, multiple related meetings will take place over a long pe-riod of time. This paper describes an initial investigation of how this long-term data can be used to improve meeting tran-scription. A corpus of technical meetings, using a single micro-phone array, was collected over a two year period, yielding a total of 179 hours of meeting data. Baseline systems using deep neural network acoustic models, in both Tandem and Hybrid configurations, and neural network-based language models are described. The impact of supervised and unsupervised adap-tation of the acoustic models is then evaluated, as well as the impact of improved language models.
Xie Chen 0001, Mark J. F. Gales, Kate M. Knill, Catherine Breslin, Langzhou Chen, K. K. Chin, Vincent Wan
INTERSPEECH3
2014 Language independent and unsupervised acoustic models for speech recognition and keyword spotting
abstract
Copyright © 2014 ISCA. Developing high-performance speech processing systems for low-resource languages is very challenging. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to train a multi-language bottleneck DNN. Language dependent and/or multi-language (all training languages) Tandem acoustic models (AM) are then trained. This work considers a particular scenario where the target language is unseen in multi-language training and has limited language model training data, a limited lexicon, and acoustic training data without transcriptions. A zero acoustic resources case is first described where a multilanguage AM is directly applied, as a language independent AM (LIAM), to an unseen language. Secondly, in an unsupervised approach a LIAM is used to obtain hypotheses for the target language acoustic data transcriptions which are then used in training a language dependent AM. 3 languages from the IARPA Babel project are used for assessment: Vietnamese, Haitian Creole and Bengali. Performance of the zero acoustic resources system is found to be poor, with keyword spotting at best 60% of language dependent performance. Unsupervised language dependent training yields performance gains. For one language (Haitian Creole) the Babel target is achieved on the in-vocabulary data.
Kate M. Knill, Mark J. F. Gales, Anton Ragni, Shakti P. Rath
INTERSPEECH1
2014 Data augmentation for low resource languages
abstract
Recently there has been interest in the approaches for training speech recognition systems for languages with limited resources.Under the IARPA Babel program such resources have been provided for a range of languages to support this research area.This paper examines a particular form of approach, data augmentation, that can be applied to these situations.Data augmentation schemes aim to increase the quantity of data available to train the system, for example semi-supervised training, multilingual processing, acoustic data perturbation and speech synthesis.To date the majority of work has considered individual data augmentation schemes, with few consistent performance contrasts or examination of whether the schemes are complementary.In this work two data augmentation schemes, semisupervised training and vocal tract length perturbation, are examined and combined on the Babel limited language pack configuration.Here only about 10 hours of transcribed acoustic data are available.Two languages are examined, Assamese and Zulu, which were found to be the most challenging of the Babel languages released for the 2014 Evaluation.For both languages consistent speech recognition performance gains can be obtained using these augmentation schemes.Furthermore the impact of these performance gains on a down-stream keyword spotting task are also described.
Anton Ragni, Kate M. Knill, Shakti P. Rath, Mark J. F. Gales
INTERSPEECH2
2014 Combining tandem and hybrid systems for improved speech recognition and keyword spotting on low resource languages
abstract
Copyright © 2014 ISCA. In recent years there has been significant interest in Automatic Speech Recognition (ASR) and KeyWord Spotting (KWS) systems for low resource languages. One of the driving forces for this research direction is the IARPA Babel project. This paper examines the performance gains that can be obtained by combining two forms of deep neural network ASR systems, Tandem and Hybrid, for both ASR and KWS using data released under the Babel project. Baseline systems are described for the five option period 1 languages: Assamese; Bengali; Haitian Creole; Lao; and Zulu. All the ASR systems share common attributes, for example deep neural network configurations, and decision trees based on rich phonetic questions and state-position root nodes. The baseline ASR and KWS performance of Hybrid and Tandem systems are compared for both the "full", approximately 80 hours of training data, and limited, approximately 10 hours of training data, language packs. By combining the two systems together consistent performance gains can be obtained for KWS in all configurations.
Shakti P. Rath, Kate M. Knill, Anton Ragni, Mark J. F. Gales
INTERSPEECH2
2013 Investigation of multilingual deep neural networks for spoken term detection
abstract
The development of high-performance speech processing systems for low-resource languages is a challenging area. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to use bottleneck features, or hybrid systems, trained on multilingual data for speech-to-text (STT) systems. This paper presents an investigation into the application of these multilingual approaches to spoken term detection. Experiments were run using the IARPA Babel limited language pack corpora (~10 hours/language) with 4 languages for initial multilingual system development and an additional held-out target language. STT gains achieved through using multilingual bottleneck features in a Tandem configuration are shown to also apply to keyword search (KWS). Further improvements in both STT and KWS were observed by incorporating language questions into the Tandem GMM-HMM decision trees for the training set languages. Adapted hybrid systems performed slightly worse on average than the adapted Tandem systems. A language independent acoustic model test on the target language showed that retraining or adapting of the acoustic models to the target language is currently minimally needed to achieve reasonable performance.
Kate M. Knill, Mark J. F. Gales, Shakti P. Rath, Philip C. Woodland, Chao Zhang 0031, Shixiong Zhang 0001
ASRU1
2013 Integrated automatic expression prediction and speech synthesis from text
abstract
Getting a text to speech synthesis (TTS) system to speak lively animated stories like a human is very difficult. To generate expressive speech, the system can be divided into 2 parts: predicting expressive information from text; and synthesizing the speech with a particular expression. Traditionally these blocks have been studied separately. This paper proposes an integrated approach, sharing the expressive synthesis space and training data across the two expressive components. There are several advantages to this approach, including a simplified expression labelling process, support of a continuous expressive synthesis space, and joint training of the expression predictor and speech synthesiser to maximise the likelihood of the TTS system given the training data. Synthesis experiments indicated that the proposed approach generated far more expressive speech than both a neutral TTS and one where the expression was randomly selected. The experimental results also showed the advantage of a continuous expressive synthesis space over a discrete space.
Langzhou Chen, Mark J. F. Gales, Norbert Braunschweiler, Masami Akamine, Kate M. Knill
ICASSP5
2013 A high-performance Cantonese keyword search system
abstract
We present a system for keyword search on Cantonese conversational telephony audio, collected for the IARPA Babel program, that achieves good performance by combining postings lists produced by diverse speech recognition systems from three different research groups. We describe the keyword search task, the data on which the work was done, four different speech recognition systems, and our approach to system combination for keyword search. We show that the combination of four systems outperforms the best single system by 7%, achieving an actual term-weighted value of 0.517.
Brian Kingsbury, Jia Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP5
2013 Training a supra-segmental parametric F0 model without interpolating F0
abstract
Combining multiple intonation models at different linguistic levels is an effective way to improve the naturalness of the predicted F0. In many of these approaches, the intonation models for suprasegmental levels are based on a parametrization of the log-F0 contours over the units of that level. However, many of these parametrisations are not stable when applied to discontinuous signals. Therefore, the F0 signal has to be interpolated. These interpolated values introduce a distortion in the coefficients that degrades the quality of the model. This paper proposes two methods that eliminate the need for such interpolation, one based on regularization and the other on factor analysis. Subjective evaluations show that, for a Discrete-cosine-transform (DCT) syllable-level model, both approaches result in a significant improvement w.r.t. a baseline using interpolated F0. The approach based on regularization yields the best results.
Javier Latorre, Mark J. F. Gales, Kate M. Knill, Masami Akamine
ICASSP3
2013 System combination and score normalization for spoken term detection
abstract
Spoken content in languages of emerging importance needs to be searchable to provide access to the underlying information. In this paper, we investigate the problem of extending data fusion methodologies from Information Retrieval for Spoken Term Detection on low-resource languages in the framework of the IARPA Babel program. We describe a number of alternative methods improving keyword search performance. We apply these methods to Cantonese, a language that presents some new issues in terms of reduced resources and shorter query lengths. First, we show score normalization methodology that improves in average by 20% keyword search performance. Second, we show that properly combining the outputs of diverse ASR systems performs 14% better than the best normalized ASR system.
Jonathan Mamou, Jia Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP6
2012 Unsupervised clustering of emotion and voice styles for expressive TTS
abstract
Current text-to-speech synthesis (TTS) systems are often perceived as lacking expressiveness, limiting the ability to fully convey information. This paper describes initial investigations into improving expressiveness for statistical speech synthesis systems. Rather than using hand-crafted definitions of expressive classes, an unsupervised clustering approach is described which is scalable to large quantities of training data. To incorporate this “expression cluster” information into an HMM-TTS system two approaches are described: cluster questions in the decision tree construction; and average expression speech synthesis (AESS) using cluster-based linear transform adaptation. The performance of the approaches was evaluated on audiobook data in which the reader exhibits a wide range of expressiveness. A subjective listening test showed that synthesising with AESS results in speech that better reflects the expressiveness of human speech than a baseline expression-independent system.
Florian Eyben, Sabine Buchholz, Norbert Braunschweiler, Javier Latorre, Vincent Wan, Mark J. F. Gales, Kate M. Knill
ICASSP7
2012 Speech factorization for HMM-TTS based on cluster adaptive training
abstract
This paper presents a novel approach to factorize and control different speech factors in HMM-based TTS systems. In this paper cluster adaptive training (CAT) is used to factorize speaker identity and expressiveness (i.e. emotion). Within a CAT framework, each speech factor can be modelled by a different set of clusters. Users can control speaker identity and expressiveness independently by modifying the weights associated with each set. These weights are defined in a continuous space, so variations of speaker and emotion are also continuous. Additionally, given a speaker which has only neutral-style training data, the approach is able to synthesise speech with that speaker’s voice and different expressions. Lastly, the paper discusses how generalization of the basic factorization concept could allow the production of expressive speech from neutral voices for other HMM-TTS systems not based on CAT.
Javier Latorre, Vincent Wan, Mark J. F. Gales, Langzhou Chen, K. K. Chin, Kate M. Knill, Masami Akamine
INTERSPEECH6
2012 Combining multiple high quality corpora for improving HMM-TTS
abstract
The most reliable way to build synthetic voices for end-products is to start with high quality recordings from professional voice talents. This paper describes the application of average voice models (AVMs) and a novel application of cluster adaptive training (CAT) to combine a small number of these high quality corpora to make best use of them and improve overall voice quality in hidden Markov model based text-to-speech (HMMTTS) systems. It is shown that integrated training by both CAT and AVM approaches, yields better sounding voices than speaker dependent modelling. It is also shown that CAT has an advantage over AVMs when adapting to a new speaker. Given a limited amount of adaptation data CAT maintains a much higher voice quality even when adapted to tiny amounts of speech.
Vincent Wan, Javier Latorre, K. K. Chin, Langzhou Chen, Mark J. F. Gales, Heiga Zen, Kate M. Knill, Masami Akamine
INTERSPEECH7
2012 Statistical Parametric Speech Synthesis Based on Speaker and Language Factorization
abstract
An increasingly common scenario in building speech synthesis and recognition systems is training on inhomogeneous data. This paper proposes a new framework for estimating hidden Markov models on data containing both multiple speakers and multiple languages. The proposed framework, speaker and language factorization, attempts to factorize speaker-/language-specific characteristics in the data and then model them using separate transforms. Language-specific factors in the data are represented by transforms based on cluster mean interpolation with cluster-dependent decision trees. Acoustic variations caused by speaker characteristics are handled by transforms based on constrained maximum-likelihood linear regression. Experimental results on statistical parametric speech synthesis show that the proposed framework enables data from multiple speakers in different languages to be used to: train a synthesis system; synthesize speech in a language using speaker characteristics estimated in a different language; and adapt to a new language.
Heiga Zen, Norbert Braunschweiler, Sabine Buchholz, Mark J. F. Gales, Kate M. Knill, Sacha Krstulovic, Javier Latorre
IEEE Trans. Speech Audio Process.5
2011 Rapid joint speaker and noise compensation for robust speech recognition
abstract
For speech recognition, mismatches between training and testing for speaker and noise are normally handled separately. The work presented in this paper aims at jointly applying speaker adaptation and model-based noise compensation by embedding speaker adaptation as part of the noise mismatch function. The proposed method gives a faster and more optimum adaptation compared to compensating for these two factors separately. It is also more consistent with respect to the basic assumptions of speaker and noise adaptation. Experimental results show significant and consistent gains from the proposed method.
K. K. Chin, Haitian Xu, Mark J. F. Gales, Catherine Breslin, Kate M. Knill
ICASSP5
2011 Continuous F0 in the source-excitation generation for HMM-based TTS: Do we need voiced/unvoiced classification?
abstract
Most HMM-based TTS systems use a hard voiced/unvoiced classification to produce a discontinuous F0 signal which is used for the generation of the source-excitation. When a mixed source excitation is used, this decision can be based on two different sources of information: the state-specific MSD-prior of the F0 models, and/or the frame-specific features generated by the aperiodicity model. This paper examines the meaning of these variables in the synthesis process, their interaction, and how they affect the perceived quality of the generated speech The results of several perceptual experiments show that when using mixed excitation, subjects consistently prefer samples with very few or no false unvoiced errors, whereas a reduction in the rate of false voiced errors does not produce any perceptual improvement. This suggests that rather than using any form of hard voiced/unvoiced classification, e.g., the MSD-prior, it is better for synthesis to use a continuous F0 signal and rely on the frame-level soft voiced/unvoiced decision of the aperiodicity model.
Javier Latorre, Mark J. F. Gales, Sabine Buchholz, Kate M. Knill, Masatsune Tamura, Yamato Ohtani, Masami Akamine
ICASSP4
2011 Integrated Online Speaker Clustering and Adaptation
abstract
For many applications, it is necessary to produce speech transcriptions in a causal fashion. To produce high quality transcripts, speaker adaptation is often used. This requires online speaker clustering and incremental adaptation techniques to be developed. This paper presents an integrated approach to online speaker clustering and adaptation which allows efficient clustering of speakers using the same accumulated statistics that are normally used for adaptation. Using a consistent criterion for both clustering and adaptation should yield gains for both stages. The proposed approach is evaluated on a meetings transcription task using audio from multiple distant microphones. Consistent gains over standard clustering and adaptation were obtained. Copyright © 2011 ISCA.
Catherine Breslin, K. K. Chin, Mark J. F. Gales, Kate M. Knill
INTERSPEECH4
2011 Multipulse Sequences for Residual Signal Modeling
abstract
In source-filter models of speech production, the residual signal - what remains after passing the speech signal through the inverse filter - contains important information for the generation of naturally sounding re-synthesized speech. Typically, the voiced regions of residual signals are regarded as a mixture of glottal pulse and noise. This paper introduces a novel approach to represent the noise component of voiced regions of residual signals through autoregressive filtering of multipulse sequences. The positions and amplitudes of the non-zero samples of these multipulse signals are optimized through a closed-loop procedure. The method in question is applied to excitation modeling in statistical parametric synthesis. Experimental results indicate that the use of multipulse-based noise component construction eliminates the necessity of run-time ad hoc procedures such as high-pass filtering and time modulation, common on excitation models for statistical parametric synthesizers, with no loss of synthesized speech quality. Copyright © 2011 ISCA.
Ranniery Maia, Heiga Zen, Kate M. Knill, Mark J. F. Gales, Sabine Buchholz
INTERSPEECH3
2010 Prior information for rapid speaker adaptation
abstract
Rapidly adapting a speech recognition system to new speakers using a small amount of adaptation data is important to improve initial user experience. In this paper, a count-smoothing framework for incorporating prior information is extended to allow for the use of different forms of dynamic prior and improve the robustness of transform estimation on small amounts of data. Prior information is obtained from existing rapid adaptation techniques like VTLN and PCMLLR. Results using VTLN as a dynamic prior for CMLLR estimation show that transforms estimated on just one utterance can yield relative gains of 15% and 46% over a baseline gender independent model on two tasks. Index Terms: automatic speech recognition, speaker adaptation, VTLN, prior knowledge
Catherine Breslin, K. K. Chin, Mark J. F. Gales, Kate M. Knill, Haitian Xu
INTERSPEECH4
2010 A comparison of pronunciation modeling approaches for HMM-TTS
Gabriel Webster, Sacha Krstulovic, Kate M. Knill
INTERSPEECH3
2009 Compression techniques applied to multiple speech recognition systems
abstract
Speech recognition systems typically contain many Gaussian distributions, and hence a large number of parameters. This makes them both slow to decode speech, and large to store. Techniques have been proposed to decrease the number of parameters. One approach is to share parameters between multiple Gaussians, thus reducing the total number of parameters and allowing for shared likelihood calculation. Gaussian tying and subspace clustering are two related techniques which take this approach to system compression. These techniques can decrease the number of parameters with no noticeable drop in performance for single systems. However, multiple acoustic models are often used in real speech recognition systems. This paper considers the application of Gaussian tying and subspace compression to multiple systems. Results show that two speech recognition systems can be modelled using the same number of Gaussians as just one system, with little effect on individual system performance. Copyright © 2009 ISCA.
Catherine Breslin, Matthew N. Stuttle, Kate M. Knill
INTERSPEECH3
2009 Improved language modelling using bag of word pairs
abstract
The bag-of-words (BoW) method has been used widely in language modelling and information retrieval. A document is expressed as a group of words disregarding the grammar and the order of word information. A typical BoW method is latent semantic analysis (LSA), which maps the words and documents onto the vectors in LSA space. In this paper, the concept of BoW is extended to Bag-of-Word Pairs (BoWP), which expresses the document as a group of word pairs. Using word pairs as a unit, the system can capture more complex semantic information than BoW. Under the LSA framework, the BoWP system is shown to improve both perplexity and word error rate (WER) compared to a BoW system. Copyright © 2009 ISCA.
Langzhou Chen, K. K. Chin, Kate M. Knill
INTERSPEECH3
2006 Comparison of the ITU-t p.85 standard to other methods for the evaluation of text-to-speech systems
abstract
Evaluation of TTS systems is essential to assess performance. The ITU-T P.85 standard was introduced in 1994 to assess the overall quality of speech synthesis systems. However it has not been widely accepted or used. This paper compares the ITU test to more commonly used tests for intelligibility (semantically unpredictable sentences (SUS)) and naturalness (mean opinion score based). The aim of this research was to determine if the ITU test can provide a better performance measure and/or supplementary information to help evaluate TTS systems. Index Terms: speech synthesis, evaluation
Dmitry Sityaev, Kate M. Knill, Tina Burrows
INTERSPEECH2
2005 Combining models of prosodic phrasing and pausing
abstract
This paper describes two approaches to assigning prosodic phrase structure and pauses to text and investigates the impact of errors in the assignments for different granularities of prosodic phrase structure. One approach uses a cascaded combination of models trained separately for prediction of prosodic phrase structure and pauses and the other uses a model trained for the joint prediction task directly. Objective measurements show similar performance for both approaches while perceptual evaluations show a slight preference for an optimised cascaded combination of prosodic phrase structure and pause models using a single-level encoding of prosodic phrase structure.
Tina Burrows, Peter Jackson, Kate M. Knill, Dmitry Sityaev
INTERSPEECH3
2005 A comparison of methods for speaker-dependent pronunciation tuning for text-to-speech synthesis
abstract
Unit-based text-to-speech (TTS) systems typically use a set of speech recordings that have been phonetically transcribed to create a large set of phonetic units. During synthesis, pronunciations for input text are generated and used to guide the selection of a sequence of phonetic units. The style of these system pronunciations must match the style of the phonetic transcriptions of the recorded speech database in order to maximize the quality of the synthesized speech. Furthermore, since different speakers have different speech characteristics, supporting multiple speakers for a single language generally requires applying a speaker-dependent mapping to speaker-independent pronunciations. This paper investigates three automatic methods for this process of speaker-dependent pronunciation tuning: word N-grams, decision trees, and transformation-based learning. Transformation-based learning achieved the best results, lowering the phone error rate of the text pronunciations compared to the speech transcriptions by 26% over the error rate of the unmodified text transcriptions.
Gabriel Webster, Tina Burrows, Kate M. Knill
INTERSPEECH3
1999 Low-cost implementation of open set keyword spotting
Kate M. Knill, Steve J. Young
Comput. Speech Lang.1
1999 State-based Gaussian selection in large vocabulary continuous speech recognition using HMMs
abstract
This paper investigates the use of Gaussian selection (GS) to increase the speed of a large vocabulary speech recognition system. Typically, 30-70% of the computational time of a continuous density hidden Markov model-based (HMM-based) speech recognizer is spent calculating probabilities. The aim of CS is to reduce this load by selecting the subset of Gaussian component likelihoods that should be computed given a particular input vector. This paper examines new techniques for obtaining "good" Gaussian subsets or "shortlists." All the new schemes make use of state information, specifically, to which state each of the Gaussian components belongs. In this way, a maximum number of Gaussian components per state may be specified, hence reducing the size of the shortlist. The first technique introduced is a simple extension of the standard GS method, which uses this state information. Then, more complex schemes based on maximizing the likelihood of the training data are proposed. These new approaches are compared with the standard GS scheme on a large vocabulary speech recognition task. On this task, the use of state information reduced the percentage of Gaussians computed to 10-15%, compared with 20-30% for the standard GS scheme, with little degradation in performance.
Mark J. F. Gales, Kate M. Knill, Steve J. Young
IEEE Trans. Speech Audio Process.2
1996 Fast implementation methods for Viterbi-based word-spotting
abstract
This paper explores methods of increasing the speed of a Viterbi-based word-spotting system for audio document retrieval. Fast processing is essential since the user expects to receive the results of a keyword search many times faster than the actual length of the speech. A number of computational short-cuts to the standard Viterbi word-spotter are presented. These are based on exploiting the background Viterbi phone recognition path that is computed to provide a normalisation base. An initial approximation using the phone transition boundaries reduces the retrieval time by a factor of 5, while achieving a slight improvement in word-spotting performance. To further reduce retrieval time, pattern matching, feature selection, and Gaussian selection techniques are applied to this approximate pass to give a total /spl times/50 increase in speed with little loss in performance. In addition, a low memory requirement means that these approaches can be implemented on any platform, including hand-held devices.
Kate M. Knill, Steve J. Young
ICASSP1
1996 Use of Gaussian selection in large vocabulary continuous speech recognition using HMMs
abstract
This paper investigates the use of Gaussian Selection (GS) to reduce the state likelihood computation in HMM-based systems.These likelihood calculations contribute significantly (30 to 70%) to the computational load.Previously, it has been reported that when GS is used on large systems the recognition accuracy tends to degrade above a 3 reduction in likelihood computation.To explain this degradation, this paper investigates the trade-offs necessary between achieving good state likelihoods and low computation.In addition, the problem of unseen states in a cluster is examined.It is shown that further improvements are possible.For example, using a different assignment measure, with a constraint on the number of components per state per cluster, enabled the recognition accuracy on a 5k speaker-independent task to be maintained up to a 5 reduction in likelihood computation.
Kate M. Knill, Mark J. F. Gales, Steve J. Young
ICSLP1
1992 A novel orthogonal set adaptive line enhancer tuned with fourth-order cumulants
abstract
A novel adaptive line enhancer (ALE) structure is introduced. The adaptive filter has an autoregressive moving average (ARMA) structure which is based on classical Laguerre orthogonal functions. The frequencies and radii of the poles within the orthogonal set admit tuning. Tuning of these terms involves the use of higher-order statistics. Indeed the properties of fourth-order cumulants are exploited to aid harmonic retrieval of the sinusoids embedded in the input signal. Furthermore, adaptation of the feedforward filter parameters is straightforward since they are linearly related to the filter output. Advantages of stability, short filter length, and robustness in the presence of noise are expected. Simulation results are included to show typical performance.>
Anthony G. Constantinides, Kate M. Knill, Jonathon A. Chambers
ICASSP2