EDBT 2026 Demo / reviewers in the wild / expert
Thomas Hain
dblp:46/1798
· DBLP profile ↗
162ranked-venue papers
14as first author
41since 2021 · last 2026
0000-0003-0939-3464ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 11 first-author · 33 since 2021Artificial intelligence and machine learning · 113 · 10 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difference in Task Performance on Sparse Speech RepresentationsabstractLearning speech representations that are useful for a variety of downstream tasks has received considerable attention, due to the outstanding properties of Self-Supervised Learning (SSL) trained models.Despite advancements in modelling methods, understanding the difference in task performance on representations is limited.Mainly motivated by the nofree-lunch theorem and speech production, this work investigates changes in task performance in sparse speech representations, providing interpretability analysis under the Information Bottleneck (IB) framework.Autoencoders with varying sparsity levels were trained using three SSL features, and evaluated on six tasks of SUPERB: Speech Enhancement (SE), Speaker Identification (SID), Speech Emotion Recognition (SER), Phone Recognition (PR), Automatic Speech Recognition (ASR) and Slot Filling (SF).Experiments show that: 1) different tasks manifest different degrees of sensitivity to the sparsity levels; 2) the optimal sparsity level for task performance varies; 3) the choice of SSL features has a limited impact on most tasks but with an exception of PR; 4) overall PR and ASR require more preservation of relevant information about the labels, while SID and SER demand more compression of irrelevant information, where the input quality can shift this trade-off to some degree.These findings can contribute to the design of a universal sparse speech representation learner. Wenjie Peng, Thomas Hain |
ACL (1) | 3 |
| 2026 | CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech
Brian Yan, Qingzheng Wang, Matthew Wiesner, Anuj Diwan, Olga Iakovenko, Alexander Polok, Injy Hamed, Shuichiro Shimizu, Iris Emerman, Thomas Hain, David R. Mortensen, Peter Viechnicki, Shinji Watanabe 0001 |
LREC | 10 |
| 2025 | Emphasis Sensitivity in Speech RepresentationsabstractThis work investigates whether modern speech models are sensitive to prosodic emphasis-whether they encode emphasized and neutral words in systematically different ways. Prior work typically relies on isolated acoustic correlates (e.g., pitch, duration) or label prediction, both of which miss the relational structure of emphasis. This paper proposes a residualbased framework, defining emphasis as the difference between paired neutral and emphasized word representations. Analysis on self-supervised speech models shows that these residuals correlate strongly with duration changes and perform poorly at word identity prediction, indicating a structured, relational encoding of prosodic emphasis. In ASR fine-tuned models, residuals occupy a subspace up to 50% more compact than in pre-trained models, further suggesting that emphasis is encoded as a consistent, lowdimensional transformation that becomes more structured with task-specific learning. Shaun Cassini, Thomas Hain, Anton Ragni |
ASRU | 2 |
| 2025 | Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and TextabstractWord error rate (WER) estimation aims to evaluate the quality of an automatic speech recognition (ASR) system’s output without requiring ground-truth labels. This task has gained increasing attention as advanced ASR systems are trained on large amounts of data. In this context, the computational efficiency of a WER estimator becomes essential in practice. However, previous works have not prioritised this aspect. In this paper, a Fast estimator for WER (Fe-WER) is introduced, utilizing average pooling over self-supervised learning representations for speech and text. Our results demonstrate that FeWER outperformed a baseline relatively by 14.10% in root mean square error and 1.22% in Pearson correlation coefficient on TedLium3. Moreover, a comparative analysis of the distributions of target WER and WER estimates was conducted, including an examination of the average values per speaker. Lastly, the inference speed was approximately 3.4 times faster in the real-time factor. Chanho Park 0003, Chengsong Lu, Thomas Hain |
ICASSP | 4 |
| 2025 | Towards a Unified Benchmark for Arabic Pronunciation Assessment: Qur'anic Recitation as Case Study
Yassine El Kheir, Omnia Ibrahim, Amit Meghanani, Nada Almarwani, Hawau Olamide Toyin, Sadeen Alharbi, Modar Alfadly, Lamya Alkanhal, Ibrahim Selim, Shehab Elbatal, Salima Mdhaffar, Thomas Hain, Yasser Hifny, Mostafa Shahin, Ahmed Ali 0002 |
INTERSPEECH | 12 |
| 2025 | Semi-Supervised Learning for Automatic Speech Recognition with Word Error Rate Estimation and Targeted Domain Data Selection
Chanho Park 0003, Thomas Hain |
INTERSPEECH | 2 |
| 2024 | Automatic Speech Recognition System-Independent Word Error Rate EstimationabstractWord error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition (ASR) systems. In many applications, it is of interest to estimate WER given a pair of a speech utterance and a transcript. Previous work on WER estimation focused on building models that are trained with a specific ASR system in mind (referred to as ASR system-dependent). These are also domain-dependent and inflexible in real-world applications. In this paper, a hypothesis generation method for ASR System-Independent WER estimation (SIWE) is proposed. In contrast to prior work, the WER estimators are trained using data that simulates ASR system output. Hypotheses are generated using phonetically similar or linguistically more likely alternative words. In WER estimation experiments, the proposed method reaches a similar performance to ASR system-dependent WER estimators on in-domain data and achieves state-of-the-art performance on out-of-domain data. On the out-of-domain data, the SIWE model outperformed the baseline estimators in root mean square error and Pearson correlation coefficient by relative 17.58% and 18.21%, respectively, on Switchboard and CALLHOME. The performance was further improved when the WER of the training set was close to the WER of the evaluation dataset. Chanho Park 0003, Thomas Hain |
LREC/COLING | 3 |
| 2024 | Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech RepresentationsabstractAcoustic word embeddings (AWEs) are vector representations of spoken words.An effective method for obtaining AWEs is the Correspondence Auto-Encoder (CAE).In the past, the CAE method has been associated with traditional MFCC features.Representations obtained from self-supervised learning (SSL)-based speech models such as HuBERT, Wav2vec2, etc., are outperforming MFCC in many downstream tasks.However, they have not been well studied in the context of learning AWEs.This work explores the effectiveness of CAE with SSL-based speech representations to obtain improved AWEs.Additionally, the capabilities of SSL-based speech models are explored in cross-lingual scenarios for obtaining AWEs.Experiments are conducted on five languages: Polish, Portuguese, Spanish, French, and English.HuBERT-based CAE model achieves the best results for word discrimination in all languages, despite Hu-BERT being pre-trained on English only.Also, the HuBERT-based CAE model works well in cross-lingual settings.It outperforms MFCCbased CAE models trained on the target languages when trained on one source language and tested on target languages. Amit Meghanani, Thomas Hain |
EACL (1) | 2 |
| 2024 | Methods of Automatic Matrix Language Determination for Code-Switched SpeechabstractCode-switching (CS) is the process of speakers interchanging between two or more languages which in the modern world becomes increasingly common.In order to better describe CS speech the Matrix Language Frame (MLF) theory introduces the concept of a Matrix Language, which is the language that provides the grammatical structure for a CS utterance.In this work the MLF theory was used to develop systems for Matrix Language Identity (MLID) determination.The MLID of English/Mandarin and English/Spanish CS text and speech was compared to acoustic language identity (LID), which is a typical way to identify a language in monolingual utterances.MLID predictors from audio show higher correlation with the textual principles than LID in all cases while also outperforming LID in an MLID recognition task based on F1 macro (60%) and correlation score (0.38).This novel approach has identified that non-English languages (Mandarin and Spanish) are preferred over the English language as the ML contrary to the monolingual choice of LID. 1 https://biling.talkbank.org/access/Bangor/Miami.html 2 LDC97T14 3 LDC96T16 4 LDC96T17 5794 Table 3: Examples of applying the principles.Utterance baseline ML P1.1 ML P1.2 ML P2 i thought all trains 都是via jurongeast 去到pasirris en en en en but 他蛮zai 的right en zh en en but 我的parents 都没有sponsor 我 zh zh zh en 还有chicken noodles en en en zhTable 4: CS dataset splits. Olga Iakovenko, Thomas Hain |
EMNLP | 2 |
| 2024 | Progressive Unsupervised Domain Adaptation for ASR Using Ensemble Models and Multi-Stage TrainingabstractIn Automatic Speech Recognition (ASR), teacher-student (T/S) training has shown to perform well for domain adaptation with small amount of training data. However, adaption without ground-truth labels is still challenging. A previous study has shown the effectiveness of using ensemble teacher models in T/S training for unsupervised domain adaptation (UDA) but its performance still lags behind compared to the model trained on in-domain data. This paper proposes a method to yield better UDA by training multistage students with ensemble teacher models. Initially, multiple teacher models are trained on labelled data from read and meeting domains. These teachers are used to train a student model on unlabelled out-of-domain telephone speech data. To improve the adaptation, subsequent student models are trained sequentially considering previously trained model as their teacher. Experiments are conducted with three teachers trained on AMI, WSJ and LibriSpeech and three stages of students on SwitchBoard data. Results shown on eval00 test set show significant WER improvement with multi-stage training with an absolute gain of 9.8%, 7.7% and 3.3% at each stage. Rehan Ahmad, Thomas Hain |
ICASSP | 3 |
| 2024 | Multi-CMGAN+/+: Leveraging Multi-Objective Speech Quality Metric Prediction for Speech EnhancementabstractNeural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on artificially created labelled training data such that the neural model can be trained using intrusive loss functions which compare the output of the model with clean reference speech. Performance of such systems when enhancing real-world audio often suffers relative to their performance on simulated test data. In this work, a non-intrusive multi-metric prediction approach is introduced, wherein a model trained on artificial labelled data using inference of an adversarially trained metric prediction neural network. The proposed approach shows improved performance versus state-of-the-art systems on the recent CHiME-7 challenge unsupervised domain adaptation speech enhancement (UDASE) task evaluation sets. Index Terms: speech enhancement, model generalisation, generative adversarial networks, conformer, metric prediction George Close, William Ravenscroft, Thomas Hain, Stefan Goetze |
ICASSP | 3 |
| 2024 | SCORE: Self-Supervised Correspondence Fine-Tuning for Improved Content RepresentationsabstractThere is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used for robust performance on various downstream tasks by fine-tuning on the labelled data. This work presents a cost-effective SSFT method named Self-supervised Correspondence (SCORE) fine-tuning to adapt the SSL speech representations for content-related tasks. The proposed method uses a correspondence training strategy, aiming to learn similar representations from perturbed speech and original speech. Commonly used data augmentation techniques for content-related tasks (ASR) are applied to obtain perturbed speech. SCORE fine-tuned HuBERT outperforms the vanilla HuBERT on SUPERB benchmark with only a few hours of fine-tuning (< 5 hrs) on a single GPU for automatic speech recognition, phoneme recognition, and query-by-example tasks, with relative improvements of 1.09%, 3.58%, and 12.65%, respectively. SCORE provides competitive results with the recently proposed SSFT method SPIN, using only 1/3 of the processed speech compared to SPIN. Amit Meghanani, Thomas Hain |
ICASSP | 2 |
| 2024 | Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory ModelsabstractNeural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work combines the use of Whisper ASR decoder layer representations as neural network input features with an exemplar-based, psychologically motivated model of human memory to predict human intelligibility ratings for hearing-aid users. Substantial performance improvement over an established intrusive HASPI baseline system is found, including on enhancement systems and listeners unseen in the training data, with a root mean squared error of 25.3 compared with the baseline of 28.7. Rhiannon Mogridge, George Close, Robert Sutherland, Thomas Hain, Jon Barker, Stefan Goetze, Anton Ragni |
ICASSP | 4 |
| 2024 | Combining Conformer and Dual-Path-Transformer Networks for Single Channel Noisy Reverberant Speech SeparationabstractSeparation of overlapping speakers remains an active area of speech technology research. Many deep neural network (DNN) separation models propose modelling local and global temporal context separately using alternating DNN layers. Two such models are SepFormer and TD-Conformer. The largest configurations of each have comparable computational cost and similar performance; with SepFormer performing better on anechoic data and TD-Conformer yielding better results on noisy reverberant data. This work combines these two model types to gain insights into how their computational characteristics affect their performance. The generalization benefits of the larger model size of the conformer layers are demonstrated both on the WHAMR and the out-of-domain far-field evaluation set MC-WSJ-AV across a number of evaluation metrics. The proposed model is able to achieve 22.1 dB and 14.7 dB average scale-invariant signal-to-distortion ratio (SISDR) improvement when trained and evaluated on WSJ0-2Mix and WHAMR, respectively. The model trained using WHAMR is able to achieve 4.3 dB average SISDR improvement on the out-of-domain MC-WSJ-AV dataset. William Ravenscroft, Stefan Goetze, Thomas Hain |
ICASSP | 3 |
| 2024 | EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Ziyang Ma 0001, Hezhao Zhang, Zhisheng Zheng, Xiquan Li, Jiaxin Ye, Xie Chen 0001, Thomas Hain |
INTERSPEECH | 9 |
| 2024 | LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
Amit Meghanani, Thomas Hain |
INTERSPEECH | 2 |
| 2024 | Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech RecognitionabstractOne solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals.Commonly, the separator produces artefacts which often degrade ASR performance.Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks.This is often not viable for training on real-world in-domain audio where reference transcript information is not always available.This paper proposes a transcription-free method for joint training using only audio signals.The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT).The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI). William Ravenscroft, George Close, Stefan Goetze, Thomas Hain, Mohammad Soleymanpour, Anurag Chowdhury, Mark C. Fuhs |
INTERSPEECH | 4 |
| 2023 | MUST: A Multilingual Student-Teacher Learning Approach for Low-Resource Speech RecognitionabstractStudent-teacher learning or knowledge distillation (KD) has been previously used to address data scarcity issue for training of speech recognition (ASR) systems. However, a limitation of KD training is that the student model classes must be a proper or improper subset of the teacher model classes. It prevents distillation from even acoustically similar languages if the character sets are not same. In this work, the aforementioned limitation is addressed by proposing a MUltilingual Student-Teacher (MUST) learning which exploits a posteriors mapping approach. A pre-trained mapping model is used to map posteriors from a teacher language to the student language ASR. These mapped posteriors are used as soft labels for KD learning. Various teacher ensemble schemes are experimented to train an ASR model for low-resource languages. A model trained with MUST learning reduces relative character error rate (CER) up to $9.5 \%$ in comparison with a baseline monolingual ASR. Rehan Ahmad, Thomas Hain |
ASRU | 3 |
| 2023 | Simulation of Teacher-Learner Interaction in English Language Pronunciation LearningabstractSecond language (L2) learning is a complex process that is difficult to model. This work aims to develop a computational model of the teacher-learner interaction as used for L2 learning. The teacher model simulates a native English speaker, which uses repetition as a teaching strategy, while the learner model simulates a native Chinese speaker at an early stage of L2 English learning. Joint simulation may allow valuable insights into the entire learning process. In this study, speakers from the speechocean762 corpus were enlisted, using a word list that includes phonemes known to pose difficulties for Chinese speakers. The similarity between the output of the learning process and real learner data is evaluated using MCD, PPG, and wav2vec 2.0 distortion measures. The results indicate that the similarity between the process output and real learners with low proficiency is higher compared to that with real learners with high proficiency. Elaf Islam, Thomas Hain, Protima Nomo Sudro |
ASRU | 2 |
| 2023 | Deriving Translational Acoustic Sub-Word EmbeddingsabstractThere is a growing interest in understanding the representational geometry of acoustic word embeddings (AWEs), which are fixed-dimensional representations of spoken words. However, not much research has been conducted on acoustic sub-word embeddings (ASWEs), which can provide a better understanding of the AWE space. This work focuses on decomposing AWEs to obtain ASWEs while retaining the ability to reconstruct AWEs by translating ASWEs in the embedding space, under constrained settings. Initially, high-quality AWEs are obtained with an Average Precision (AP) score of 0.97 on the word discrimination task. Subsequently, ASWEs are derived through the decomposition of AWEs. Three adapted versions of the AP metric, utilized for evaluating the quality of the derived ASWEs and their translational properties, are proposed. The results demonstrate that the derived ASWEs exhibit high quality, and the reconstruction of AWEs from the ASWEs is achievable by translating them in the embedding space. Amit Meghanani, Thomas Hain |
ASRU | 2 |
| 2023 | On Time Domain Conformer Models for Monaural Speech Separation in Noisy Reverberant Acoustic EnvironmentsabstractSpeech separation remains an important topic for multispeaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched for speech separation. Most recent state-of-the-art (SOTA) separation models have been time-domain audio separation networks (TasNets). A number of successful models have made use of dual-path (DP) networks which sequentially process local and global information. Time domain conformers (TD-Conformers) are an analogue of the DP approach in that they also process local and global context sequentially but have a different time complexity function. It is shown that for realistic shorter signal lengths, conformers are more efficient when controlling for feature dimension. Subsampling layers are proposed to further improve computational efficiency. The best TD-Conformer achieves $14.6 \mathrm{~dB}$ and $21.2 \mathrm{~dB}$ SISDR improvement on the WHAMR and WSJO2Mix benchmarks, respectively. William Ravenscroft, Stefan Goetze, Thomas Hain |
ASRU | 3 |
| 2023 | Towards Domain Generalisation in ASR with Elitist Sampling and Ensemble Knowledge DistillationabstractKnowledge distillation (KD) has widely been used for model compression and domain adaptation for speech applications. In the presence of multiple teachers, knowledge can easily be transferred to the student by averaging the models output. However, previous research shows that the student do not adapt well with such combination. This paper propose to use an elitist sampling strategy at the output of ensemble teacher models to select the best-decoded utterance generated by completely out-of-domain teacher models for generalizing unseen domain. The teacher models are trained on AMI, LibriSpeech and WSJ while the student is adapted for the Switchboard data. The results show that with the selection strategy based on the individual model’s posteriors the student model achieves a better WER compared to all the teachers and baselines with a minimum absolute improvement of about 8.4%. Furthermore, an insights on the model adaptation with out-of-domain data has also been studied via correlation analysis. Rehan Ahmad, Md Asif Jalal, Anna Ollerenshaw, Thomas Hain |
ICASSP | 5 |
| 2023 | Perceive and Predict: Self-Supervised Speech Representation Based Loss Functions for Speech EnhancementabstractRecent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final outputs of self supervised speech representation models, rather than the earlier feature encodings. The use of self supervised representations in such a way is often not fully motivated. In this work it is shown that the distance between the feature encodings of clean and noisy speech correlate strongly with psychoacoustically motivated measures of speech quality and intelligibility, as well as with human Mean Opinion Score (MOS) ratings. Experiments using this distance as a loss function are performed and improved performance over the use of STFT spectrogram distance based loss as well as other common loss functions from speech enhancement literature is demonstrated using objective measures such as perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI). George Close, William Ravenscroft, Thomas Hain, Stefan Goetze |
ICASSP | 3 |
| 2023 | Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech SeparationabstractSpeech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One such class of models known as temporal convolutional networks (TCNs) has shown promising results for speech separation tasks. A limitation of these models is that they have a fixed receptive field (RF). Recent research in speech dereverberation has shown that the optimal RF of a TCN varies with the reverberation characteristics of the speech signal. In this work deformable convolution is proposed as a solution to allow TCN models to have dynamic RFs that can adapt to various reverberation times for reverberant speech separation. The proposed models are capable of achieving an 11.1 dB average scale-invariant signal-to-distortion ratio (SISDR) improvement over the input signal on the WHAMR benchmark. A relatively small deformable TCN model of 1.3M parameters is proposed which gives comparable separation performance to larger and more computationally complex models. William Ravenscroft, Stefan Goetze, Thomas Hain |
ICASSP | 3 |
| 2023 | Domain Adaptive Self-supervised Training of Automatic Speech Recognition
Cong-Thanh Do, Rama Sanand Doddipatla, Mohan Li, Thomas Hain |
INTERSPEECH | 4 |
| 2023 | Learning Cross-lingual Mappings for Data Augmentation to Improve Low-Resource Speech Recognitionabstractaugmentation to improve low-resource speech recognition. Thomas Hain |
INTERSPEECH | 2 |
| 2022 | Unsupervised Data Selection for Speech Recognition with Contrastive Loss RatiosabstractThis paper proposes an unsupervised data selection method by using a submodular function based on contrastive loss ratios of target and training data sets. A model using a contrastive loss function is trained on both sets. Then the ratio of frame-level losses for each model is used by a submodular function. By using the submodular function, a training set for automatic speech recognition matching the target data set is selected. Experiments show that models trained on the data sets selected by the proposed method outperform the selection method based on log-likelihoods produced by GMM-HMM models, in terms of word error rate (WER). When selecting a fixed amount, e.g. 10 hours of data, the difference between the results of two methods on Tedtalks was 20.23% WER relative. The method can also be used to select data with the aim of minimising negative transfer, while maintaining or improving on performance of models trained on the whole training set. Results show that the WER on the WSJCAM0 data set was reduced by 6.26% relative when selecting 85% from the whole data set. Chanho Park 0003, Rehan Ahmad, Thomas Hain |
ICASSP | 3 |
| 2022 | A Model for Assessor Bias in Automatic Pronunciation AssessmentabstractIn pronunciation assessment, the assessor’s perception is influenced by a particular pronunciation template. This assessor may hold a bias towards certain variations in pronunciation which do not necessarily impact communication, yet they may be penalized during the assessment. This work proposes a model for pronunciation assessment as the combination of an assessor independent (A) and an assessor specific (B) component. The latter could be interpreted as the assessor bias. The resulting assessment function was implemented as a dual model trained to detect mispronounced speech segments. The models incorporate Long-Short Memory and saliency region selection using attention. An experiment was performed using recordings from young Dutch learners of English as second language, which were annotated for mispronunciation by three trained phoneticians (a1, a2, a3). The models combined were able to detect mispronunciations given the assessor identity achieving F1 scores of 0.77, 0.68 and 0.86 for a1, a2, a3 respectively on the Train set and 0.66, 0.53 and 0.81 on the Test set. Additionally, the attention weights of the B model were able to illustrate disagreements between assessors related to the bias. Jose Antonio Lopez Saenz, Thomas Hain |
ICASSP | 2 |
| 2022 | Non-intrusive Speech Intelligibility Metric Prediction for Hearing Impaired IndividualsabstractThis paper proposes neural models to predict Speech Intelligibility (SI),both by prediction of established SI metrics and of human speech recognition (HSR) on the 1st Clarity Prediction Challenge. Both intrusive and non-intrusive predictors for intrusive SI metrics are trained, then fine tuned on the HSR ground truth. Results are reported on a number of SI metrics, and the model choice for the Clarity challenge submission is explained. Additionally, the relationship between the SI scores in the data and commonly used signal processing metrics which approximate SI are analysed, and some issues emerging from this relationship discussed. It is found that intrusive neural predictors of SI metrics when finetuned on the true HSR scores outperform the non neural challenge baseline. George Close, Samuel Schmück, Stefan Goetze, Thomas Hain |
INTERSPEECH | 4 |
| 2022 | Investigating the Impact of Crosslingual Acoustic-Phonetic Similarities on Multilingual Speech RecognitionabstractMultilingual speech recognition systems mostly benefit low resource languages but suffer degradation in the performance of several languages relative to their monolingual counterparts. Limited studies have focused on understanding the languages behaviour in the multilingual speech recognition setups. In this paper, a novel data-driven approach is proposed to investigate the cross-lingual acoustic-phonetic similarities. This technique measures the similarities between posterior distributions from various monolingual acoustic models against a target speech signal. Deep neural networks are trained as mapping networks to transform the distributions from different acoustic models into a directly comparable form. The analysis observes that the languages ‘closeness' can not be truly estimated by the volume of overlapping phonemes set. Entropy analysis of the proposed mapping networks exhibits that a language with lesser overlap can be more amenable to cross-lingual transfer, and hence more beneficial in the multilingual setup. Finally, the proposed posterior transformation approach is leveraged to fuse monolingual models for a target language. A relative improvement of ∼8% over monolingual counterpart is achieved. Thomas Hain |
INTERSPEECH | 2 |
| 2022 | Non-Linear Pairwise Language Mappings for Low-Resource Multilingual Acoustic Model FusionabstractMultilingual speech recognition has drawn significant attention as an effective way to compensate data scarcity for low-resource languages. End-to-end (e2e) modelling is preferred over conventional hybrid systems, mainly because of no lexicon requirement. However, hybrid DNN-HMMs still outperform e2e models in limited data scenarios. Furthermore, the problem of manual lexicon creation has been alleviated by publicly available trained models of grapheme-to-phoneme (G2P) and text to IPA transliteration for a lot of languages. In this paper, a novel approach of hybrid DNN-HMM acoustic models fusion is proposed in a multilingual setup for the low-resource languages. Posterior distributions from different monolingual acoustic models, against a target language speech signal, are fused together. A separate regression neural network is trained for each source-target language pair to transform posteriors from source acoustic model to the target language. These networks require very limited data as compared to the ASR training. Posterior fusion yields a relative gain of 14.65% and 6.5% when compared with multilingual and monolingual baselines respectively. Cross-lingual model fusion shows that the comparable results can be achieved without using posteriors from the language dependent ASR. Darshan Adiga Haniya Narayana, Thomas Hain |
INTERSPEECH | 3 |
| 2022 | Automatic detection of behavioural codes in team interactions
Madina Hasan, Nicholas Jefferson, Thomas Hain, Jeremy Dawson |
Comput. Speech Lang. | 3 |
| 2021 | Attention Based Model for Segmental Pronunciation Error DetectionabstractThe Goodness of Pronunciation (GOP) algorithm is one well-established method of pronunciation assessment dependent on precise phoneme segment boundaries. The alignment process of canonical pronunciations to obtain these boundaries is prone to errors. To overcome this issue, the present paper proposes a combination of Bidirectional Long-Short Memory with saliency region selection with attention weights to provide estimates of mispronunciations without the need of phoneme boundaries. Three output and annotation configurations were used to train the model which was then assessed against a GOP baseline in the task of detecting segments with mispronunciations. The experiments were conducted using data from young Dutch learners of English, annotated by three trained phoneticians. The proposed model outperformed the GOP baseline. It was also found that the attention weights aligned themselves with the phoneme labels, helping detect mispronunciations and allowing further interpretation of the internal representation of the model. Jose Antonio Lopez Saenz, Md Asif Jalal, Rosanna Milner, Thomas Hain |
ASRU | 4 |
| 2021 | Improving Audio Anomalies Recognition Using Temporal Convolutional Attention NetworksabstractAnomalous audio in speech recordings is often caused by speaker voice distortion, external noise, or even electric interferences. These obstacles have become a serious problem in some fields, such as high-quality dubbing and speech processing. In this paper, a novel approach using a temporal convolutional attention network (TCAN) is proposed to tackle this problem. The use of temporal conventional network (TCN) can capture long range patterns using a hierarchy of temporal convolutional filters. To enhance the ability to tackle audio anomalies in different acoustic conditions, an attention mechanism is used in TCN, where a self-attention block is added after each temporal convolutional layer. This aims to high-light the target related features and mitigate the interferences from irrelevant information. To evaluate the performance of the proposed model, audio recordings are collected from the TIMIT dataset, and are then changed by adding five different types of audio distortions: gaussian noise, magnitude drift, random dropout, reduction of temporal resolution, and time warping. Distortions are mixed at different signal-to-noise ratios (SNRs) (5dB, 10dB, 15dB, 20dB, 25dB, 30dB). The experimental results show that the use of proposed model can yield better classification performances than some strong baseline methods, such as the LSTM and TCN based models, by approximate 3~ 10% relative improvements. Qiang Huang 0008, Thomas Hain |
ICASSP | 2 |
| 2021 | Towards Low-Resource Stargan Voice Conversion Using Weight Adaptive Instance NormalizationabstractMany-to-many voice conversion with non-parallel training data has seen significant progress in recent years. It is challenging because of lacking of ground truth parallel data. StarGAN-based models have gained attentions because of their efficiency and effective. However, most of the StarGAN-based works only focused on small number of speakers and large amount of training data. In this work, we aim at improving the data efficiency of the model and achieving a many-to-many non-parallel StarGAN-based voice conversion for a relatively large number of speakers with limited training samples. In order to improve data efficiency, the proposed model uses a speaker encoder for extracting speaker embeddings and weight adaptive instance normalization (W-AdaIN) layers. Experiments are conducted with 109 speakers under two low-resource situations, where the number of training samples is 20 and 5 per speaker. An objective evaluation shows the proposed model outperforms baseline methods significantly. Furthermore, a subjective evaluation shows that, for both naturalness and similarity, the proposed model outperforms baseline method. Yanpei Shi, Thomas Hain |
ICASSP | 3 |
| 2021 | Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech RecognitionabstractThis paper proposes an adaptation method for end-to-end speech recognition. In this method, multiple automatic speech recognition (ASR) 1-best hypotheses are integrated in the computation of the connectionist temporal classification (CTC) loss function. The integration of multiple ASR hypotheses helps alleviating the impact of errors in the ASR hypotheses to the computation of the CTC loss when ASR hypotheses are used. When being applied in semi-supervised adaptation scenarios where part of the adaptation data do not have labels, the CTC loss of the proposed method is computed from different ASR 1-best hypotheses obtained by decoding the unlabeled adaptation data. Experiments are performed in clean and multi-condition training scenarios where the CTC-based end-to-end ASR systems are trained on Wall Street Journal (WSJ) clean training data and CHiME-4 multi-condition training data, respectively, and tested on Aurora-4 test data. The proposed adaptation method yields 6.6% and 5.8% relative word error rate (WER) reductions in clean and multi-condition training scenarios, respectively, compared to a baseline system which is adapted with part of the adaptation data having manual transcriptions using back-propagation fine-tuning. Cong-Thanh Do, Rama Sanand Doddipatla, Thomas Hain |
ICASSP | 3 |
| 2021 | Insights on Neural Representations for End-to-End Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) models aim to learn a generalised speech representation. However, there are limited tools available to understand the internal functions and the effect of hierarchical dependencies within the model architecture. It is crucial to understand the correlations between the layer-wise representations, to derive insights on the relationship between neural representations and performance. Previous investigations of network similarities using correlation analysis techniques have not been explored for End-to-End ASR models. This paper analyses and explores the internal dynamics between layers during training with CNN, LSTM and Transformer based approaches using Canonical correlation analysis (CCA) and centered kernel alignment (CKA) for the experiments. It was found that neural representations within CNN layers exhibit hierarchical correlation dependencies as layer depth increases but this is mostly limited to cases where neural representation correlates more closely. This behaviour is not observed in LSTM architecture, however there is a bottom-up pattern observed across the training process, while Transformer encoder layers exhibit irregular coefficiency correlation as neural depth increases. Altogether, these results provide new insights into the role that neural architectures have upon speech recognition performance. More specifically, these techniques can be used as indicators to build better performing speech recognition models. Anna Ollerenshaw, Md Asif Jalal, Thomas Hain |
Interspeech | 3 |
| 2021 | WINVC: One-Shot Voice Conversion with Weight Adaptive Instance Normalization
Shengjie Huang, Yanyan Xu 0001, Dengfeng Ke, Thomas Hain |
PRICAI (2) | 5 |
| 2021 | Contextual Joint Factor Acoustic EmbeddingsabstractEmbedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic context are proposed. The first approach is a contextual joint factor synthesis encoder, where the encoder in an encoder/decoder framework is trained to extract joint factors from surrounding audio frames to best generate the target output. The second approach is a contextual joint factor analysis encoder, where the encoder is trained to analyse joint factors from the source signal that correlates best with the neighbouring audio. To evaluate the effectiveness of our approaches compared to prior work, two tasks are conducted-phone classification and speaker recognition - and test on different TIMIT data sets. Experimental results show that one of the proposed approaches outperforms phone classification baselines, yielding a classification accuracy of 74.1%. When using additional out-of-domain data for training, an additional 3% improvements can be obtained, for both for phone classification and speaker recognition tasks. Yanpei Shi, Thomas Hain |
SLT | 2 |
| 2021 | Supervised Speaker Embedding De-Mixing in Two-Speaker EnvironmentabstractSeparating different speaker properties from a multi-speaker environment is challenging. Instead of separating a two-speaker signal in signal space like speech source separation, a speaker embedding de-mixing approach is proposed. The proposed approach separates different speaker properties from a two-speaker signal in embedding space. The proposed approach contains two steps. In step one, the clean speaker embeddings are learned and collected by a residual TDNN based network. In step two, the two-speaker signal and the embedding of one of the speakers are both input to a speaker embedding de-mixing network. The de-mixing network is trained to generate the embedding of the other speaker by reconstruction loss. Speaker identification accuracy and the cosine similarity score between the clean embeddings and the de-mixed embeddings are used to evaluate the quality of the obtained embeddings. Experiments are done in two kind of data: artificial augmented two-speaker data (TIMIT) and real world recording of two-speaker data (MC-WSJ). Six different speaker embedding de-mixing architectures are investigated. Comparing with the performance on the clean speaker embeddings, the obtained results show that one of the proposed architectures obtained close performance, reaching 96.9% identification accuracy and 0.89 cosine similarity. Yanpei Shi, Thomas Hain |
SLT | 2 |
| 2021 | H-VECTORS: Improving the robustness in utterance-level speaker embeddings using a hierarchical attention model
Yanpei Shi, Qiang Huang 0008, Thomas Hain |
Neural Networks | 3 |
| 2020 | H-Vectors: Utterance-Level Speaker Embedding Using a Hierarchical Attention ModelabstractIn this paper, a hierarchical attention network is proposed to generate utterance-level embeddings (H-vectors) for speaker identification and verification. Since different parts of an utterance may have different contributions to speaker identities, the use of hierarchical structure aims to learn speaker related information locally and globally. In the proposed approach, frame-level encoder and attention are applied on segments of an input utterance and generate individual segment vectors. Then, segment level attention is applied on the segment vectors to construct an utterance representation. To evaluate the effectiveness of the proposed approach, the data of the NIST SRE2008 Part1 is used for training, and two datasets, the Switchboard Cellular (Part1) and the CallHome American English Speech, are used to evaluate the quality of extracted utterance embeddings on speaker identification and verification tasks. In comparison with two baselines, X-vectors and X-vectors+Attention, the obtained results show that the use of H-vectors can achieve a significantly better performance. Furthermore, the learned utterance-level embeddings are more discriminative than the two baselines when mapped into a 2D space using t-SNE. Yanpei Shi, Qiang Huang 0008, Thomas Hain |
ICASSP | 3 |
| 2020 | Exploration of Audio Quality Assessment and Anomaly Localisation Using Attention ModelsabstractMany applications of speech technology require more and more audio data.Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications.However, effective and high performing assessment remains a challenging task without a clean reference.In this paper, a novel model for audio quality assessment is proposed by jointly using bidirectional long short-term memory and an attention mechanism.The former is to mimic a human auditory perception ability to learn information from a recording, and the latter is to further discriminate interferences from desired signals by highlighting target related features.To evaluate our proposed approach, the TIMIT dataset is used and augmented by mixing with various natural sounds.In our experiments, two tasks are explored.The first task is to predict an utterance quality score, and the second is to identify where an anomalous distortion takes place in a recording.The obtained results show that the use of our proposed approach outperforms a strong baseline method and gains about 5% improvements after being measured by three metrics, Linear Correlation Coefficient and Spearman's Rank Correlation Coefficient, and F1. Qiang Huang 0008, Thomas Hain |
INTERSPEECH | 2 |
| 2020 | Unsupervised Acoustic Unit Representation Learning for Voice Conversion Using WaveNet Auto-EncodersabstractThis is a repository copy of Unsupervised acoustic unit representation learning for voice conversion using WaveNet auto-encoders. Thomas Hain |
INTERSPEECH | 2 |
| 2020 | Empirical Interpretation of Speech Emotion Perception with Attention Based Model for Speech Emotion RecognitionabstractSpeech emotion recognition is essential for obtaining emotional intelligence which affects the understanding of context and meaning of speech. Harmonically structured vowel and consonant sounds add indexical and linguistic cues in spoken information. Previous research argued whether vowel sound cues were more important in carrying the emotional context from a psychological and linguistic point of view. Other research also claimed that emotion information could exist in small overlapping acoustic cues. However, these claims are not corroborated in computational speech emotion recognition systems. In this research, a convolution-based model and a long-short-term memory-based model, both using attention, are applied to investigate these theories of speech emotion on computational models. The role of acoustic context and word importance is demonstrated for the task of speech emotion recognition. The IEMOCAP corpus is evaluated by the proposed models, and 80.1% unweighted accuracy is achieved on pure acoustic data which is higher than current state-of-the-art models on this task. The phones and words are mapped to the attention vectors and it is seen that the vowel sounds are more important for defining emotion acoustic cues than the consonants, and the model can assign word importance based on acoustic context. Md Asif Jalal, Rosanna Milner, Thomas Hain |
INTERSPEECH | 3 |
| 2020 | Removing Bias with Residual Mixture of Multi-View Attention for Speech Emotion RecognitionabstractSpeech emotion recognition is essential for obtaining emotional intelligence which affects the understanding of context and meaning of speech. The fundamental challenges of speech emotion recognition from a machine learning standpoint is to extract patterns which carry maximum correlation with the emotion information encoded in this signal, and to be as insensitive as possible to other types of information carried by speech. In this paper, a novel recurrent residual temporal context modelling framework is proposed. The framework includes mixture of multi-view attention smoothing and high dimensional feature projection for context expansion and learning feature representations. The framework is designed to be robust to changes in speaker and other distortions, and it provides state-of-the-art results for speech emotion recognition. Performance of the proposed approach is compared with a wide range of current architectures in a standard 4-class classification task on the widely used IEMOCAP corpus. A significant improvement of 4% unweighted accuracy over state-of-the-art systems is observed. Additionally, the attention vectors have been aligned with the input segments and plotted at two different attention levels to demonstrate the effectiveness. Md Asif Jalal, Rosanna Milner, Thomas Hain, Roger K. Moore |
INTERSPEECH | 3 |
| 2020 | Multilingual Speech Recognition Using Language-Specific Phoneme Recognition as Auxiliary Task for Indian Languages
Hardik B. Sailor, Thomas Hain |
INTERSPEECH | 2 |
| 2020 | Speaker Re-Identification with Speaker Dependent Speech EnhancementabstractWhile the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments.Here speech enhancement methods have traditionally allowed improved performance.The recent works have shown that adapting speech enhancement can lead to further gains.This paper introduces a novel approach that cascades speech enhancement and speaker recognition.In the first step, a speaker embedding vector is generated , which is used in the second step to enhance the speech quality and re-identify the speakers.Models are trained in an integrated framework with joint optimisation.The proposed approach is evaluated using the Voxceleb1 dataset, which aims to assess speaker recognition in real world situations.In addition three types of noise at different signal-noise-ratios were added for this work.The obtained results show that the proposed approach using speaker dependent speech enhancement can yield better speaker recognition and speech enhancement performances than two baselines in various noise conditions. Yanpei Shi, Qiang Huang 0008, Thomas Hain |
INTERSPEECH | 3 |
| 2020 | Weakly Supervised Training of Hierarchical Attention Networks for Speaker IdentificationabstractIdentifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task.In this paper, a hierarchical attention network is proposed to solve a weakly labelled speaker identification problem.The use of a hierarchical structure, consisting of a frame-level encoder and a segment-level encoder, aims to learn speaker related information locally and globally.Speech streams are segmented into fragments.The frame-level encoder with attention learns features and highlights the target related frames locally, and output a fragment based embedding.The segment-level encoder works with a second attention layer to emphasize the fragments probably related to target speakers.The global information is finally collected from segment-level module to predict speakers via a classifier.To evaluate the effectiveness of the proposed approach, artificial datasets based on Switchboard Cellular part1 (SWBC) and Voxceleb1 are constructed in two conditions, where speakers' voices are overlapped and not overlapped.Comparing to two baselines the obtained results show that the proposed approach can achieve better performances.Moreover, further experiments are conducted to evaluate the impact of utterance segmentation.The results show that a reasonable segmentation can slightly improve identification performances. Yanpei Shi, Qiang Huang 0008, Thomas Hain |
INTERSPEECH | 3 |
| 2020 | Uncertainty-Aware Machine Support for Paper Reviewing on the Interspeech 2019 Submission CorpusabstractThe evaluation of scientific submissions through peer review is both the most fundamental component of the publication process, as well as the most frequently criticised and questioned.Academic journals and conferences request reviews from multiple reviewers per submission, which an editor, or area chair aggregates into the final acceptance decision.Reviewers are often in disagreement due to varying levels of domain expertise, confidence, levels of motivation, as well as due to the heavy workload and the different interpretations by the reviewers of the score scale.Herein, we explore the possibility of a computational decision support tool for the editor, based on Natural Language Processing, that offers an additional aggregated recommendation.We provide a comparative study of state-of-the-art text modelling methods on the newly crafted, largest review dataset of its kind based on Interspeech 2019, and we are the first to explore uncertainty-aware methods (soft labels, quantile regression) to address the subjectivity inherent in this problem. Lukas Stappen, Georgios Rizos, Madina Hasan, Thomas Hain, Björn W. Schuller |
INTERSPEECH | 4 |
| 2019 | Spatio-Temporal Context Modelling for Speech Emotion ClassificationabstractSpeech emotion recognition (SER) is a requisite for emotional intelligence that affects the understanding of speech. One of the most crucial tasks is to obtain patterns having a maximum correlation for the emotion classification task from the speech signal while being invariant to the changes in frequency, time and other external distortions. Therefore, learning emotional contextual feature representation independent of speaker and environment is essential. In this paper, a novel spatiotemporal context modelling framework for robust SER is proposed to learn feature representation by using acoustic context expansion with high dimensional feature projection. The framework uses a deep convolutional neural network (CNN) and self-attention network. The CNNs combine spatiotemporal features. The attention network produces high dimensional task-specific features and combines these features for context modelling, which altogether provides a state-of-the-art technique for classifying the extracted patterns for speech emotion. Speech emotion is a categorical perception representing discrete sensory events. The proposed approach is compared with a wide range of architectures on the RAVDESS and IEMOCAP corpora for 8-class and 4-class emotion classification tasks and remarkable gain over state-of-the-art systems are obtained, absolutely 15%, 10% respectively. Md Asif Jalal, Roger K. Moore, Thomas Hain |
ASRU | 3 |
| 2019 | A Cross-Corpus Study on Speech Emotion RecognitionabstractFor speech emotion datasets, it has been difficult to acquire large quantities of reliable data and acted emotions may be over the top compared to less expressive emotions displayed in everyday life. Lately, larger datasets with natural emotions have been created. Instead of ignoring smaller, acted datasets, this study investigates whether information learnt from acted emotions is useful for detecting natural emotions. Cross-corpus research has mostly considered cross-lingual and even cross-age datasets, and difficulties arise from different methods of annotating emotions causing a drop in performance. To be consistent, four adult English datasets covering acted, elicited and natural emotions are considered. A state-of-the-art model is proposed to accurately investigate the degradation of performance. The system involves a bi-directional LSTM with an attention mechanism to classify emotions across datasets. Experiments study the effects of training models in a cross-corpus and multi-domain fashion and results show the transfer of information is not successful. Out-of-domain models, followed by adapting to the missing dataset, and domain adversarial training (DAT) are shown to be more suitable to generalising to emotions across datasets. This shows positive information transfer from acted datasets to those with more natural emotions and the benefits from training on different corpora. Rosanna Milner, Md Asif Jalal, Raymond W. M. Ng, Thomas Hain |
ASRU | 4 |
| 2019 | Unsupervised Adaptation of Acoustic Models for ASR Using Utterance-Level Embeddings from Squeeze and Excitation NetworksabstractThis paper proposes the adaptation of neural network-based acoustic models using a Squeeze-and-Excitation (SE) network for automatic speech recognition (ASR). In particular, this work explores to use the SE network to learn utterance-level embeddings. The acoustic modelling is performed using Light Gated Recurrent Units (LiGRU). The utterance embed-dings are learned from hidden unit activations jointly with LiGRU and used to scale respective activations of hidden layers in the LiGRU network. The advantage of such approach is that it does not require domain labels, such as speakers and noise to be known in order to perform the adaptation, thereby providing unsupervised adaptation. The global average and attentive pooling are applied on hidden units to extract utterance-level information that represents the speakers and acoustic conditions. ASR experiments were carried out on the TIMIT and Aurora 4 corpora. The proposed model achieves better performance on both the datasets compared to their respective baselines with relative improvements of 5.59% and 5.54% for TIMIT and Aurora 4 database, respectively. These experiments show the potential of using the conditioning information learned via utterance embeddings in the SE network to adapt acoustic models for speakers, noise, and other acoustic conditions. Hardik B. Sailor, Salil Deena, Md Asif Jalal, Rasa Lileikyte, Thomas Hain |
ASRU | 5 |
| 2019 | Latent Dirichlet Allocation Based Acoustic Data Selection for Automatic Speech RecognitionabstractSelecting in-domain data from a large pool of diverse and out-of-domain data is a non-trivial problem. In most cases simply using all of the available data will lead to sub-optimal and in some cases even worse performance compared to carefully selecting a matching set. This is true even for data-inefficient neural models. Acoustic Latent Dirichlet Allocation (aLDA) is shown to be useful in a variety of speech technology related tasks, including domain adaptation of acoustic models for automatic speech recognition and entity labeling for information retrieval. In this paper we propose to use aLDA as a data similarity criterion in a data selection framework. Given a large pool of out-of-domain and potentially mismatched data, the task is to select the best-matching training data to a set of representative utterances sampled from a target domain. Our target data consists of around 32 hours of meeting data (both far-field and close-talk) and the pool contains 2k hours of meeting, talks, voice search, dictation, command-and-control, audio books, lectures, generic media and telephony speech data. The proposed technique for training data selection, significantly outperforms random selection, posterior-based selection as well as using all of the available data. Mortaza Doulaty, Thomas Hain |
INTERSPEECH | 2 |
| 2019 | Detecting Mismatch Between Speech and Transcription Using Cross-Modal Attention
Qiang Huang 0008, Thomas Hain |
INTERSPEECH | 2 |
| 2019 | Learning Temporal Clusters Using Capsule Routing for Speech Emotion RecognitionabstractEmotion recognition from speech plays a significant role in adding emotional intelligence to machines and making human-machine interaction more natural. One of the key challenges from machine learning standpoint is to extract patterns which bear maximum correlation with the emotion information encoded in this signal while being as insensitive as possible to other types of information carried by speech. In this paper, we propose a novel temporal modelling framework for robust emotion classification using bidirectional long short-term memory network (BLSTM), CNN and Capsule networks. The BLSTM deals with the temporal dynamics of the speech signal by effectively representing forward/backward contextual information while the CNN along with the dynamic routing of the Capsule net learn temporal clusters which altogether provide a state-of-the-art technique for classifying the extracted patterns. The proposed approach was compared with a wide range of architectures on the FAU-Aibo and RAVDESS corpora and remarkable gain over state-of-the-art systems were obtained. For FAO-Aibo and RAVDESS 77.6% and 56.2% accuracy was achieved, respectively, which is 3% and 14% (absolute) higher than the best-reported result for the respective tasks. Md Asif Jalal, Erfan Loweimi, Roger K. Moore, Thomas Hain |
INTERSPEECH | 4 |
| 2019 | System-independent ASR error detection and classification using Recurrent Neural Network
Rahhal Errattahi, Asmaa El Hannani, Thomas Hain, Hassan Ouahmane |
Comput. Speech Lang. | 3 |
| 2019 | Recurrent Neural Network Language Model Adaptation for Multi-Genre Broadcast Speech Recognition and AlignmentabstractRecurrent neural network language models (RNNLMs) generally outperform n-gram language models when used in automatic speech recognition (ASR). Adapting RNNLMs to new domains is an open problem and current approaches can be categorised as either feature-based or model based. In feature-based adaptation, the input to the RNNLM is augmented with auxiliary features whilst model-based adaptation includes model fine-tuning and the introduction of adaptation layer(s) in the network. In this paper, the properties of both types of adaptation are investigated on multi-genre broadcast speech recognition. Existing techniques for both types of adaptation are reviewed and the proposed techniques for model-based adaptation, namely the linear hidden network adaptation layer and the K-component adaptive the RNNLM, are investigated. Moreover, new features derived from the acoustic domain are investigated for the RNNLM adaptation. The contributions of this paper include two hybrid adaptation techniques: the fine-tuning of feature-based RNNLMs and a feature-based adaptation layer. Moreover, the semi-supervised adaptation of RNNLMs using genre information is also proposed. The ASR systems were trained using 700 h of multi-genre broadcast speech. The gains obtained when using the RNNLM adaptation techniques proposed in this paper are consistent when using RNNLMs trained on an in-domain set of 10M words and on a combination of in-domain and out-of-domain sets of 660 M words, with approx. 10% perplexity and 2% relative word error rate improvements on a 28.3 h. test set. The best RNNLM adaptation techniques for ASR are also evaluated on a lightly supervised alignment of subtitles task for the same data, where the use of RNNLM adaptation leads to an absolute increase in the F-measure of 0.5%. Salil Deena, Madina Hasan, Mortaza Doulaty, Oscar Saz-Torralba, Thomas Hain |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Exploring the Use of Group Delay for Generalised VTS Based Noise CompensationabstractIn earlier work we studied the effect of statistical normalisation for phase-based features and observed it leads to a significant robustness improvement. This paper explores the extension of the generalised Vector Taylor Series (gVTS) noise compensation approach to the group delay (GD) domain. We discuss the problems it presents, propose some solutions and derive the corresponding formulae. Furthermore, the effects of additive and channel noise in the GD domain were studied. It was observed that the GD of the noisy observation is a convex combination of the GDs of the clean signal and the additive noise and also in the expected sense, channel GD tends to zero. Experiments on Aurora-4 showed that, despite training only on the clean speech, the proposed features provide average WER reductions of 0.8% absolute and 4.1% relative compared to an MFCC-based system trained on the multi-style data. Combining the gVTS with a bottleneck DNN-based system led to average absolute (relative) WER improvements of 6.0% (23.5%) when training on clean data and 2.5% (13.8%) when using multi-style training with additive noise. Erfan Loweimi, Jon Barker, Thomas Hain |
ICASSP | 3 |
| 2018 | On the Usefulness of the Speech Phase Spectrum for Pitch ExtractionabstractMost frequency domain techniques for pitch extraction such as cepstrum, harmonic product spectrum (HPS) and summation residual harmonics (SRH) operate on the magnitude spectrum and turn it into a function in which the fundamental frequency emerges as argmax. In this paper, we investigate the extension of these three techniques to the phase and group delay (GD) domains. Our extensions exploit the observation that the bin at which F (magnitude) becomes maximum, for some monotonically increasing function F, is equivalent to bin at which F (phase) has maximum negative slope and F (group delay) has the maximum value. To extract the pitch track from speech phase spectrum, these techniques were coupled with the source-filter model in the phase domain that we proposed in earlier publications and a novel voicing detection algorithm proposed here. The accuracy and robustness of the phase-based pitch extraction techniques are illustrated and compared with their magnitude-based counterparts using six pitch evaluation metrics. On average, it is observed that the phase spectrum can be successfully employed in pitch tracking with comparable accuracy and robustness to the speech magnitude spectrum. Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 3 |
| 2018 | Improved Acoustic Modelling for Automatic Literacy Assessment of ChildrenabstractAutomatic literacy assessment of children is a complex task that normally requires carefully annotated data. This paper focuses on a system for the assessment of reading skills, aiming to detection of a range of fluency and pronunciation errors. Naturally, reading is a prompted task, and thereby the acquisition of training data for acoustic modelling should be straightforward. However, given the prominence of errors in the training set and the importance of labelling them in the transcription, a lightly supervised approach to acoustic modelling has better chances of success. A method based on weighted finite state transducers is proposed, to model specific prompt corrections, such as repetitions, substitutions, and deletions, as observed in real recordings. Iterative cycles of lightly-supervised training are performed in which decoding improves the transcriptions and the derived models. Improvements are due to increasing accuracy in phone-to-sound alignment and in the training data selection. The effectiveness of the proposed methods for rela-belling and acoustic modelling is assessed through experiemnts on the CHOREC corpus, in terms of sequence error rate and alignment accuracy. Improvements over the baseline of up to 60% and 23.3% respectively are observed. Mauro Nicolao, Michiel Sanders, Thomas Hain |
INTERSPEECH | 3 |
| 2018 | Improving ASR Error Detection with RNNLM AdaptationabstractApplications of automatic speech recognition (ASR) such as broadcast transcription and dialog systems, can be helped by the ability to detect errors in the ASR output. The field of ASR error detection has emerged as a way to detect and subsequently correct ASR errors. The most common approach for ASR error detection is features-based, where a set of features are extracted from the ASR output and used to train a classifier to predict correct/incorrect labels.Language models (LMs), either from the ASR decoder or externally trained, can be used to provide features to an ASR error detection system, through scores computed on the ASR output. Recently, recurrent neural network language models (RNNLMs) features were proposed for ASR error detection with improvements to the classification rate, thanks to their ability to model longer-range context.RNNLM adaptation, through the introduction of auxiliary features that encode domain, has been shown to improve ASR performance. This work investigates whether RNNLM adaptation techniques can also improve ASR error detection performance in the context of multi-genre broadcast ASR. The results show that an overall improvement of about 1% in the F-measure can be achieved using adapted RNNLM features. Rahhal Errattahi, Salil Deena, Asmaa El Hannani, Hassan Ouahmane, Thomas Hain |
SLT | 5 |
| 2018 | Lightly supervised alignment of subtitles on multi-genre broadcastsabstractAbstract This paper describes a system for performing alignment of subtitles to audio on multigenre broadcasts using a lightly supervised approach. Accurate alignment of subtitles plays a substantial role in the daily work of media companies and currently still requires large human effort. Here, a comprehensive approach to performing this task in an automated way using lightly supervised alignment is proposed. The paper explores the different alternatives to speech segmentation, lightly supervised speech recognition and alignment of text streams. The proposed system uses lightly supervised decoding to improve the alignment accuracy by performing language model adaptation using the target subtitles. The system thus built achieves the third best reported result in the alignment of broadcast subtitles in the Multi–Genre Broadcast (MGB) challenge, with an F1 score of 88.8%. This system is available for research and other non–commercial purposes through webASR, the University of Sheffield’s cloud–based speech technology web service. Taking as inputs an audio file and untimed subtitles, webASR can produce timed subtitles in multiple formats, including TTML, WebVTT and SRT. Oscar Saz-Torralba, Salil Deena, Mortaza Doulaty, Madina Hasan, Bilal Khaliq, Rosanna Milner, Raymond W. M. Ng, Julia Olcoz, Thomas Hain |
Multim. Tools Appl. | 9 |
| 2017 | Exploring the use of acoustic embeddings in neural machine translationabstractNeural Machine Translation (NMT) has recently demonstrated improved performance over statistical machine translation and relies on an encoder-decoder framework for translating text from source to target. The structure of NMT makes it amenable to add auxiliary features, which can provide complementary information to that present in the source text. In this paper, auxiliary features derived from accompanying audio, are investigated for NMT and are compared and combined with text-derived features. These acoustic embeddings can help resolve ambiguity in the translation, thus improving the output. The following features are experimented with: Latent Dirichlet Allocation (LDA) topic vectors and GMM subspace i-vectors derived from audio. These are contrasted against: skip-gram/Word2Vec features and LDA features derived from text. The results are encouraging and show that acoustic information does help with NMT, leading to an overall 3.3% relative improvement in BLEU scores. Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
ASRU | 5 |
| 2017 | Statistical normalisation of phase-based feature representation for robust speech recognitionabstractIn earlier work we have proposed a source-filter decomposition of speech through phase-based processing. The decomposition leads to novel speech features that are extracted from the filter component of the phase spectrum. This paper analyses this spectrum and the proposed representation by evaluating statistical properties at various points along the parametrisation pipeline. We show that speech phase spectrum has a bell-shaped distribution which is in contrast to the uniform assumption that is usually made. It is demonstrated that the uniform density (which implies that the corresponding sequence is least-informative) is an artefact of the phase wrapping and not an original characteristic of this spectrum. In addition, we extend the idea of statistical normalisation usually applied for the magnitudebased features into the phase domain. Based on the statistical structure of the phase-based features, which is shown to be super-gaussian in the clean condition, three normalisation schemes, namely, Gaussianisation, Laplacianisation and table-based histogram equalisation have been applied for improving the robustness. Speech recognition experiments using Aurora-2 show that applying an optimal normalisation scheme at the right stage of the feature extraction process can produce average relative WER reductions of up to 18.6% across the 0-20 dB SNR conditions. Erfan Loweimi, Jon Barker, Thomas Hain |
ICASSP | 3 |
| 2017 | DNN approach to speaker diarisation using speaker channelsabstractSpeaker diarisation addresses the question of “who speaks when” in audio recordings, and has been studied extensively in the context of tasks such as broadcast news, meetings, etc. Performing diarisation on individual headset microphone (IHM) channels is sometimes assumed to easily give the desired output of speaker labelled segments with timing information. However, it is shown that given imperfect data, such as speaker channels with heavy crosstalk and overlapping speech, this is not the case. Deep neural networks (DNNs) can be trained on features derived from the concatenation of speaker channel features to detect which is the correct channel for each frame. Crosstalk features can be calculated and DNNs trained with or without overlapping speech to combat problematic data. A simple frame decision metric of counting occurrences is investigated as well as adding a bias against selecting nonspeech for a frame. Finally, two different scoring setups are applied to both datasets. The stricter SHEF setup finds diarisation error rates (DER) of 9.2% on TBL and 23.2% on RT07 while the NIST setup achieves 5.7% and 15.1% respectively. Rosanna Milner, Thomas Hain |
ICASSP | 2 |
| 2017 | Shefce: A Cantonese-English bilingual speech corpus for pronunciation assessmentabstractThis paper introduces the development of ShefCE: a Cantonese-English bilingual speech corpus from L2 English speakers in Hong Kong. Bilingual parallel recording materials were chosen from TED online lectures. Script selection were carried out according to bilingual consistency (evaluated using a machine translation system) and the distribution balance of phonemes. 31 undergraduate to postgraduate students in Hong Kong aged 20-30 were recruited and recorded a 25-hour speech corpus (12 hours in Cantonese and 13 hours in English). Baseline phoneme/syllable recognition systems were trained on background data with and without the ShefCE training data. The final syllable error rate (SER) for Cantonese is 17.3% and final phoneme error rate (PER) for English is 34.5%. The automatic speech recognition performance on English showed a significant mismatch when applying L1 models on L2 data, suggesting the need for explicit accent adaptation. ShefCE and the corresponding baseline models will be made openly available for academic research. Raymond W. M. Ng, Alvin C. M. Kwan, Tan Lee, Thomas Hain |
ICASSP | 4 |
| 2017 | Semi-Supervised Adaptation of RNNLMs by Fine-Tuning with Domain-Specific Auxiliary FeaturesabstractRecurrent neural network language models (RNNLMs) can be augmented with auxiliary features, which can provide an extra modality on top of the words. It has been found that RNNLMs perform best when trained on a large corpus of generic text and then fine-tuned on text corresponding to the sub-domain for which it is to be applied. However, in many cases the auxiliary features are available for the sub-domain text but not for the generic text. In such cases, semi-supervised techniques can be used to infer such features for the generic text data such that the RNNLM can be trained and then fine-tuned on the available in-domain data with corresponding auxiliary features. \n \nIn this paper, several novel approaches are investigated for dealing with the semi-supervised adaptation of RNNLMs with auxiliary features as input. These approaches include: using zero features during training to mask the weights of the feature sub-network; adding the feature sub-network only at the time of fine-tuning; deriving the features using a parametric model and; back-propagating to infer the features on the generic text. These approaches are investigated and results are reported both in terms of PPL and WER on a multi-genre broadcast ASR task. Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
INTERSPEECH | 5 |
| 2017 | Channel Compensation in the Generalised Vector Taylor Series Approach to Robust ASRabstractVector Taylor Series (VTS) is a powerful technique for robust ASR but, in its standard form, it can only be applied to log-filter bank and MFCC features. In earlier work, we presented a generalised VTS (gVTS) that extends the applicability of VTS to front-ends which employ a power transformation non-linearity. gVTS was shown to provide performance improvements in both clean and additive noise conditions. This paper makes two novel contributions. Firstly, while the previous gVTS formulation assumed that noise was purely additive, we now derive gVTS formulae for the case of speech in the presence of both additive noise and channel distortion. Second, we propose a novel iterative method for estimating the channel distortion which utilises gVTS itself and converges after a few iterations. Since the new gVTS blindly assumes the existence of both additive noise and channel effects, it is important not to introduce extra distortion when either are absent. Experimental results conducted on LVCSR Aurora-4 database show that the new formulation passes this test. In the presence of channel noise only, it provides relative WER reductions of up to 30% and 26%, compared with previous gVTS and multi-style training with cepstral mean normalisation, respectively. Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 3 |
| 2017 | Robust Source-Filter Separation of Speech Signal in the Phase DomainabstractIn earlier work we proposed a framework for speech source-filter separation that employs phase-based signal processing. This paper presents a further theoretical investigation of the model and optimisations that make the filter and source representations less sensitive to the effects of noise and better matched to downstream processing. To this end, first, in computing the Hilbert transform, the log function is replaced by the generalised logarithmic function. This introduces a tuning parameter that adjusts both the dynamic range and distribution of the phase-based representation. Second, when computing the group delay, a more robust estimate for the derivative is formed by applying a regression filter instead of using sample differences. The effectiveness of these modifications is evaluated in clean and noisy conditions by considering the accuracy of the fundamental frequency extracted from the estimated source, and the performance of speech recognition features extracted from the estimated filter. In particular, the proposed filter-based front-end reduces Aurora-2 WERs by 6.3% (average 0-20 dB) compared with previously reported results. Furthermore, when tested in a LVCSR task (Aurora-4) the new features resulted in 5.8% absolute WER reduction compared to MFCCs without performance loss in the clean/matched condition. Erfan Loweimi, Jon Barker, Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 4 |
| 2017 | Unsupervised crosslingual adaptation of tokenisers for spoken language recognitionabstractPhone tokenisers are used in spoken language recognition (SLR) to obtain elementary phonetic information. We present a study on the use of deep neural network tokenisers. Unsupervised crosslingual adaptation was performed to adapt the baseline tokeniser trained on English conversational telephone speech data to different languages. Two training and adaptation approaches, namely cross-entropy adaptation and state-level minimum Bayes risk adaptation, were tested in a bottleneck i-vector and a phonotactic SLR system. The SLR systems using the tokenisers adapted to different languages were combined using score fusion, giving 7–18% reduction in minimum detection cost function (minDCF) compared with the baseline configurations without adapted tokenisers. Analysis of results showed that the ensemble tokenisers gave diverse representation of phonemes , thus bringing complementary effects when SLR systems with different tokenisers were combined. SLR performance was also shown to be related to the quality of the adapted tokenisers. Raymond W. M. Ng, Mauro Nicolao, Thomas Hain |
Comput. Speech Lang. | 3 |
| 2017 | Acoustic adaptation to dynamic background conditions with asynchronous transformationsabstractThis paper proposes a framework for performing adaptation to complex and non-stationary background conditions in Automatic Speech Recognition (ASR) by means of asynchronous Constrained Maximum Likelihood Linear Regression (aCMLLR) transforms and asynchronous Noise Adaptive Training (aNAT). The proposed method aims to apply the feature transform that best compensates the background for every input frame. The implementation is done with a new Hidden Markov Model (HMM) topology that expands the usual left-to-right HMM into parallel branches adapted to different background conditions and permits transitions among them. Using this, the proposed adaptation does not require ground truth or previous knowledge about the background in each frame as it aims to maximise the overall log-likelihood of the decoded utterance. The proposed aCMLLR transforms can be further improved by retraining models in an aNAT fashion and by using speaker-based MLLR transforms in cascade for an efficient modelling of background effects and speaker. An initial evaluation in a modified version of the WSJCAM0 corpus incorporating 7 different background conditions provides a benchmark in which to evaluate the use of aCMLLR transforms. A relative reduction of 40.5% in Word Error Rate (WER) was achieved by the combined use of aCMLLR and MLLR in cascade. Finally, this selection of techniques was applied in the transcription of multi-genre media broadcasts, where the use of aNAT training, aCMLLR transforms and MLLR transforms provided a relative improvement of 2–3%. Oscar Saz-Torralba, Thomas Hain |
Comput. Speech Lang. | 2 |
| 2016 | Automatic speech recognition errors detection using supervised learning techniquesabstractOver the last years, many advances have been made in the field of Automatic Speech Recognition (ASR). However, the persistent presence of ASR errors is limiting the widespread adoption of speech technology in real life applications. This motivates the attempts to find alternative techniques to automatically detect and correct ASR errors, which can be very effective and especially when the user does not have access to tune the features, the models or the decoder of the ASR system or when the transcription serves as input to downstream systems like machine translation, information retrieval, and question answering. In this paper, we present an ASR errors detection system targeted towards substitution and insertion errors. The proposed system is based on supervised learning techniques and uses input features deducted only from the ASR output words and hence should be usable with any ASR system. Applying this system on TV program transcription data leads to identify 40.30% of the recognition errors generated by the ASR system. Rahhal Errattahi, Asmaa El Hannani, Hassan Ouahmane, Thomas Hain |
AICCSA | 4 |
| 2016 | Segment-oriented evaluation of speaker diarisation performanceabstractHigh performance diarisation is a necessity for a variety of applications, and the task has been studied extensively in the context of broadcast news and meeting processing. Upon introduction of the task in NIST led evaluations, diarisation error rate (DER) was introduced as the standard metric for evaluation, and it has been consistently used to compare systems ever since. DER is a frame based metric that does not penalise for producing many short segments. However, practical systems that require diarisation input are typically not able to cope well with such artefacts. In this paper we illustrate the need for an alternative metric focussing on segments, instead of duration or boundaries only. We propose a segment based F-measure, which specifically addresses issues such as reference errors, matching start and end boundaries, and speaker pairing. The performance of the metric is analysed in the context of state-of-the-art systems and compared with other existing metrics. It is shown to give a deeper insight into the segmentation quality over the standard metrics, and thus better value for to understand impact on follow on tasks such as ASR. Rosanna Milner, Thomas Hain |
ICASSP | 2 |
| 2016 | Groupwise learning for ASR k-best list reranking in spoken language translationabstractQuality estimation models are used to predict the quality of the output from a spoken language translation (SLT) system. When these scores are used to rerank a k-best list, the rank of the scores is more important than their absolute values. This paper proposes groupwise learning to model this rank. Groupwise features were constructed by grouping pairs, triplets or M-plets among the ASR k-best outputs of the same sentence. Regression and classification models were learnt and a score combination strategy was used to predict the rank among the k-best list. Regression models with pairwise features give a bigger gain over other model and feature constructions. Groupwise learning is robust to sentences with different ASR-confidence. This technique is also complementary to linear discriminant analysis feature projection. An overall BLEU score improvement of 0.80 was achieved on an in-domain English-to-French SLT task. Raymond W. M. Ng, Kashif Shah, Lucia Specia, Thomas Hain |
ICASSP | 4 |
| 2016 | Colloquialising Modern Standard Arabic Text for Improved Speech RecognitionabstractModern standard Arabic (MSA) is the official language of spoken and written Arabic media. Colloquial Arabic (CA) is the set of spoken variants of modern Arabic that exist in the form of regional dialects. CA is used in informal and everyday conversations while MSA is formal communication. An Arabic speaker switches between the two variants according to the situation. Developing an automatic speech recognition system always requires a large collection of transcribed speech or text, and for CA dialects this is an issue. CA has limited textual resources because it exists only as a spoken language, without a standardised written form unlike MSA. This paper focuses on the data sparsity issue in CA textual resources and proposes a strategy to emulate a native speaker in colloquialising MSA to be used in CA language models (LMs) by use of a machine translation (MT) framework. The empirical results in Levantine CA show that using LMs estimated from colloquialised MSA data outperformed MSA LMs with a perplexity reduction up to 68% relative. In addition, interpolating colloquialised MSA LMs with a CA LMs improved speech recognition performance by 4% relative. Sarah Al-Shareef, Thomas Hain |
INTERSPEECH | 2 |
| 2016 | Improving Generalisation to New Speakers in Spoken Dialogue State TrackingabstractUsers with disabilities can greatly benefit from personalised voice-enabled environmental-control interfaces, but for users with speech impairments (e.g. dysarthria) poor ASR performance poses a challenge to successful dialogue. Statistical dialogue management has shown resilience against high ASR error rates, hence making it useful to improve the performance of these interfaces. However, little research was devoted to dialogue management personalisation to specific users so far. Recently, data driven discriminative models have been shown to yield the best performance in dialogue state tracking (the inference of the user goal from the dialogue history). However, due to the unique characteristics of each speaker, training a system for a new user when user specific data is not available can be challenging due to the mismatch between training and working conditions. This work investigates two methods to improve the performance with new speakers of a LSTM-based personalised state tracker: The use of speaker specific acoustic and ASRrelated features; and dropout regularisation. It is shown that in an environmental control system for dysarthric speakers, the combination of both techniques yields improvements of 3.5% absolute in state tracking accuracy. Further analysis explores the effect of using different amounts of speaker specific data to train the tracking system. Iñigo Casanueva, Thomas Hain, Phil D. Green |
INTERSPEECH | 2 |
| 2016 | Combining Feature and Model-Based Adaptation of RNNLMs for Multi-Genre Broadcast Speech RecognitionabstractRecurrent neural network language models (RNNLMs) have consistently outperformed n-gram language models when used in automatic speech recognition (ASR). This is because RNNLMs provide robust parameter estimation through the use of a continuous-space representation of words, and can generally model longer context dependencies than n-grams. The adaptation of RNNLMs to new domains remains an active research area and the two main approaches are: feature-based adaptation, where the input to the RNNLM is augmented with auxiliary features; and model-based adaptation, which includes model fine-tuning and introduction of adaptation layer(s) in the network. This paper explores the properties of both types of adaptation on multi-genre broadcast speech recognition. Two hybrid adaptation techniques are proposed, namely the finetuning of feature-based RNNLMs and the use of a feature-based adaptation layer. A method for the semi-supervised adaptation of RNNLMs, using topic model-based genre classification, is also presented and investigated. The gains obtained with RNNLM adaptation on a system trained on 700h. of speech are consistent using both RNNLMs trained on a small (10Mwords) and large set (660M words), with 10% perplexity and 2% word error rate improvements on a 28:3h. test set. Salil Deena, Madina Hasan, Mortaza Doulaty, Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 5 |
| 2016 | Automatic Genre and Show Identification of Broadcast MediaabstractHuge amounts of digital videos are being produced and broadcast every day, leading to giant media archives. Effective techniques are needed to make such data accessible further. Automatic meta-data labelling of broadcast media is an essential task for multimedia indexing, where it is standard to use multi-modal input for such purposes. This paper describes a novel method for automatic detection of media genre and show identities using acoustic features, textual features or a combination thereof. Furthermore the inclusion of available meta-data, such as time of broadcast, is shown to lead to very high performance. Latent Dirichlet Allocation is used to model both acoustics and text, yielding fixed dimensional representations of media recordings that can then be used in Support Vector Machines based classification. Experiments are conducted on more than 1200 hours of TV broadcasts from the British Broadcasting Corporation (BBC), where the task is to categorise the broadcasts into 8 genres or 133 show identities. On a 200-hour test set, accuracies of 98.6% and 85.7% were achieved for genre and show identification respectively, using a combination of acoustic and textual features with meta-data. Mortaza Doulaty, Oscar Saz-Torralba, Raymond W. M. Ng, Thomas Hain |
INTERSPEECH | 4 |
| 2016 | webASR 2 - Improved Cloud Based Speech TechnologyabstractThis paper presents the most recent developments of the \nwebASR service (www.webasr.org), the world’s first web– \nbased fully functioning automatic speech recognition platform \nfor scientific use. Initially released in 2008, the functionalities \nof webASR have recently been expanded with 3 main goals in \nmind: Facilitate access through a RESTful architecture, that allows \nfor easy use through either the web interface or an API; allow \nthe use of input metadata when available by the user to improve \nsystem performance; and increase the coverage of available \nsystems beyond speech recognition. Several new systems \nfor transcription, diarisation, lightly supervised alignment and \ntranslation are currently available through webASR. The results \nin a series of well–known benchmarks (RT’09, IWSLT’12 and \nMGB’15 evaluations) show how these webASR systems provides \nstate–of–the–art performances across these tasks Thomas Hain, Jeremy Christian, Oscar Saz-Torralba, Salil Deena, Madina Hasan, Raymond W. M. Ng, Rosanna Milner, Mortaza Doulaty, Yulan Liu |
INTERSPEECH | 1 |
| 2016 | The Sheffield Wargame Corpus - Day Two and Day ThreeabstractImproving the performance of distant speech recognition is of considerable current interest, driven by a desire to bring speech recognition into people’s homes. Standard approaches to this task aim to enhance the signal prior to recognition, typically using beamforming techniques on multiple channels. Only few real-world recordings are available that allow experimentation with such techniques. This has become even more pertinent with recent works with deep neural networks aiming to learn beamforming from data. Such approaches require large multi-channel training sets, ideally with location annotation for moving speakers, which is scarce in existing corpora. This paper presents a freely available and new extended corpus of English speech recordings in a natural setting, with moving speakers. The data is recorded with diverse microphone arrays, and uniquely, with ground truth location tracking. It extends the 8.0 hour Sheffield Wargames Corpus released in Interspeech 2013, with a further 16.6 hours of fully annotated data, including 6.1 hours of female speech to improve gender bias. Additional blog-based language model data is provided alongside, as well as a Kaldi baseline system. Results are reported with a standard Kaldi configuration, and a baseline meeting recognition system. Yulan Liu, Charles Fox, Madina Hasan, Thomas Hain |
INTERSPEECH | 4 |
| 2016 | Use of Generalised Nonlinearity in Vector Taylor Series Noise Compensation for Robust Speech RecognitionabstractDesigning good normalisation to counter the effect of environmental distortions is one of the major challenges for automatic speech recognition (ASR). The Vector Taylor series (VTS) method is a powerful and mathematically well principled technique that can be applied to both the feature and model domains to compensate for both additive and convolutional noises. One of the limitations of this approach, however, is that it is tied to MFCC (and log-filterbank) features and does not extend to other representations such as PLP, PNCC and phase-based front-ends that use power transformation rather than log compression. This paper aims at broadening the scope of the VTS method by deriving a new formulation that assumes a power transformation is used as the non-linearity during feature extraction. It is shown that the conventional VTS, in the log domain, is a special case of the new extended framework. In addition, the new formulation introduces one more degree of freedom which makes it possible to tune the algorithm to better fit the data to the statistical requirements of the ASR back-end. Compared with MFCC and conventional VTS, the proposed approach provides up to 12.2% and 2.0% absolute performance improvements on average, in Aurora-4 tasks, respectively. Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 3 |
| 2016 | DNN-Based Speaker Clustering for Speaker DiarisationabstractSpeaker diarisation, the task of answering "who spoke when?", is often considered to consist of three independent stages: speech activity detection, speaker segmentation and speaker clustering. These represent the separation of speech and nonspeech, the splitting into speaker homogeneous speech segments, followed by grouping together those which belong to the same speaker. This paper is concerned with speaker clustering, which is typically performed by bottom-up clustering using the Bayesian information criterion (BIC). We present a novel semi-supervised method of speaker clustering based on a deep neural network (DNN) model. A speaker separation DNN trained on independent data is used to iteratively relabel the test data set. This is achieved by reconfiguration of the output layer, combined with fine tuning in each iteration. A stopping criterion involving posteriors as confidence scores is investigated. Results are shown on a meeting task (RT07) for single distant microphones and compared with standard diarisation approaches. The new method achieves a diarisation error rate (DER) of 14.8%, compared to a baseline of 19.9%. Rosanna Milner, Thomas Hain |
INTERSPEECH | 2 |
| 2016 | Combining Weak Tokenisers for Phonotactic Language Recognition in a Resource-Constrained SettingabstractIn the phonotactic approach for language recognition, a phone \ntokeniser is normally used to transform the audio signal into \nacoustic tokens. The language identity of the speech is modelled \nby the occurrence statistics of the decoded tokens. The \nperformance of this approach depends heavily on the quality of \nthe audio tokeniser. A high-quality tokeniser in matched condition \nis not always available for a language recognition task. \nThis study investigated into the performance of a phonotactic \nlanguage recogniser in a resource-constrained setting, following \nNIST LRE 2015 specification. An ensemble of phone tokenisers \nwas constructed by applying unsupervised sequence training \non different target languages followed by a score-based fusion. \nThis method gave 5−7% relative performance improvement to \nbaseline system on LRE 2015 eval set. This gain was retained \nwhen the ensemble phonotactic system was further fused with \nan acoustic iVector system Raymond W. M. Ng, Bhusan Chettri, Thomas Hain |
INTERSPEECH | 3 |
| 2016 | Error Correction in Lightly Supervised Alignment of Broadcast SubtitlesabstractThis paper presents a range of error correction techniques aimed \nat improving the accuracy of a lightly supervised alignment task \nfor broadcast subtitles. Lightly supervised approaches are frequently \nused in the multimedia domain, either for subtitling \npurposes or for providing a more reliable source for training \nspeech–based systems. The proposed methods focus on directly \ncorrecting of the alignment output using different techniques to \ninfer word insertions and words with inaccurate time boundaries. \nThe features used by the classification models are the \noutputs from the alignment system, such as confidence measures, \nand word or segment duration. Experiments in this paper \nare based on broadcast material provided by the BBC to the \nMulti–Genre Broadcast (MGB) challenge participants. Results, \nshow that the order alignment F–measure improves up to 2.6% \nabsolute (15.8% relative) when combining insertion and word– \nboundary correction Julia Olcoz, Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 3 |
| 2016 | The OpenCourseWare Metadiscourse (OCWMD) Corpus
Ghada AlHarbi, Thomas Hain |
LREC | 2 |
| 2016 | A Framework for Collecting Realistic Recordings of Dysarthric Speech - the homeService Corpus
Mauro Nicolao, Heidi Christensen, Stuart P. Cunningham, Phil D. Green, Thomas Hain |
LREC | 5 |
| 2016 | Using phone features to improve dialogue state tracking generalisation to unseen statesabstractThe generalisation of dialogue state tracking to unseen dialogue states can be very challenging.In a slot-based dialogue system, dialogue states lie in discrete space where distances between states cannot be computed.Therefore, the model parameters to track states unseen in the training data can only be estimated from more general statistics, under the assumption that every dialogue state will have the same underlying state tracking behaviour.However, this assumption is not valid.For example, two values, whose associated concepts have different ASR accuracy, may have different state tracking performance.Therefore, if the ASR performance of the concepts related to each value can be estimated, such estimates can be used as general features.The features will help to relate unseen dialogue states to states seen in the training data with similar ASR performance.Furthermore, if two phonetically similar concepts have similar ASR performance, the features extracted from the phonetic structure of the concepts can be used to improve generalisation.In this paper, ASR and phonetic structurerelated features are used to improve the dialogue state tracking generalisation to unseen states of an environmental control system developed for dysarthric speakers. Iñigo Casanueva, Thomas Hain, Mauro Nicolao, Phil D. Green |
SIGDIAL Conference | 2 |
| 2015 | The MGB challenge: Evaluating multi-genre broadcast media recognitionabstractThis paper describes the Multi-Genre Broadcast (MGB) Challenge at ASRU 2015, an evaluation focused on speech recognition, speaker diarization, and "lightly supervised" alignment of BBC TV recordings. The challenge training data covered the whole range of seven weeks BBC TV output across four channels, resulting in about 1,600 hours of broadcast audio. In addition several hundred million words of BBC subtitle text was provided for language modelling. A novel aspect of the evaluation was the exploration of speech recognition and speaker diarization in a longitudinal setting — i.e. recognition of several episodes of the same show, and speaker diarization across these episodes, linking speakers. The longitudinal tasks also offered the opportunity for systems to make use of supplied metadata including show title, genre tag, and date/time of transmission. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained. Peter Bell 0001, Mark J. F. Gales, Thomas Hain, Jonathan Kilgour, Pierre Lanchantin, Xunying Liu, Andrew McParland, Steve Renals, Oscar Saz-Torralba, Mirjam Wester, Philip C. Woodland |
ASRU | 3 |
| 2015 | Latent Dirichlet Allocation based organisation of broadcast media archives for deep neural network adaptationabstractThis paper presents a new method for the discovery of latent domains in diverse speech data, for the use of adaptation of Deep Neural Networks (DNNs) for Automatic Speech Recognition. Our work focuses on transcription of multi-genre broadcast media, which is often only categorised broadly in terms of high level genres such as sports, news, documentary, etc. However, in terms of acoustic modelling these categories are coarse. Instead, it is expected that a mixture of latent domains can better represent the complex and diverse behaviours within a TV show, and therefore lead to better and more robust performance. We propose a new method, whereby these latent domains are discovered with Latent Dirichlet Allocation, in an unsupervised manner. These are used to adapt DNNs using the Unique Binary Code (UBIC) representation for the LDA domains. Experiments conducted on a set of BBC TV broadcasts, with more than 2,000 shows for training and 47 shows for testing, show that the use of LDA-UBIC DNNs reduces the error up to 13% relative compared to the baseline hybrid DNN models. Mortaza Doulaty, Oscar Saz-Torralba, Raymond W. M. Ng, Thomas Hain |
ASRU | 4 |
| 2015 | The 2015 sheffield system for longitudinal diarisation of broadcast mediaabstractSpeaker diarisation is the task of answering "who spoke when" within a multi-speaker audio recording. Diarisation of broadcast media typically operates on individual television shows, and is a particularly difficult task, due to a high number of speakers and challenging background conditions. Using prior knowledge, such as that from previous shows in a series, can improve performance. Longitudinal diarisation allows to use knowledge from previous audio files to improve performance, but requires finding matching speakers across consecutive files. This paper describes the University of Sheffield system for participation in the 2015 Multi-Genre Broadcast (MGB) challenge. The challenge required longitudinal diarisation of data from BBC archives, under very constrained resource settings. Our system consists of three main stages: speech activity detection using DNNs with novel adaptation and decoding methods; speaker segmentation and clustering, with adaptation of the DNN-based clustering models; and finally speaker linking to match speakers across shows. The final result on the development set of 19 shows from five different television series provided a Diarisation Error Rate of 50.77% in the diarisation and linking task. Rosanna Milner, Oscar Saz-Torralba, Salil Deena, Mortaza Doulaty, Raymond W. M. Ng, Thomas Hain |
ASRU | 6 |
| 2015 | The 2015 sheffield system for transcription of Multi-Genre Broadcast mediaabstractWe describe the University of Sheffield system for participation in the 2015 Multi-Genre Broadcast (MGB) challenge task of transcribing multi-genre broadcast shows. Transcription was one of four tasks proposed in the MGB challenge, with the aim of advancing the state of the art of automatic speech recognition, speaker diarisation and automatic alignment of subtitles for broadcast media. Four topics are investigated in this work: Data selection techniques for training with unreliable data, automatic speech segmentation of broadcast media shows, acoustic modelling and adaptation in highly variable environments, and language modelling of multi-genre shows. The final system operates in multiple passes, using an initial unadapted decoding stage to refine segmentation, followed by three adapted passes: a hybrid DNN pass with input features normalised by speaker-based cepstral normalisation, another hybrid stage with input features normalised by speaker feature-MLLR transformations, and finally a bottleneck-based tandem stage with noise and speaker factorisation. The combination of these three system outputs provides a final error rate of 27.5% on the official development set, consisting of 47 multi-genre shows. Oscar Saz-Torralba, Mortaza Doulaty, Salil Deena, Rosanna Milner, Raymond W. M. Ng, Madina Hasan, Yulan Liu, Thomas Hain |
ASRU | 8 |
| 2015 | Using Topic Segmentation Models for the Automatic Organisation of MOOCs resources
Ghada AlHarbi, Thomas Hain |
EDM | 2 |
| 2015 | An investigation into speaker informed DNN front-end for LVCSRabstractDeep Neural Network (DNN) has become a standard method in many ASR tasks. Recently there is considerable interest in “informed training” of DNNs, where DNN input is augmented with auxiliary codes, such as i-vectors, speaker codes, speaker separation bottleneck (SSBN) features, etc. This paper compares different speaker informed DNN training methods in LVCSR task. We discuss mathematical equivalence between speaker informed DNN training and “bias adaptation” which uses speaker dependent biases, and give detailed analysis on influential factors such as dimension, discrimination and stability of auxiliary codes. The analysis is supported by experiments on a meeting recognition task using bottleneck feature based system. Results show that i-vector based adaptation is also effective in bottleneck feature based system (not just hybrid systems). However all tested methods show poor generalisation to unseen speakers. We introduce a system based on speaker classification followed by speaker adaptation of biases, which yields equivalent performance to an i-vector based system with 10.4% relative improvement over baseline on seen speakers. The new approach can serve as a fast alternative especially for short utterances. Yulan Liu, Panagiota Karanasou, Thomas Hain |
ICASSP | 3 |
| 2015 | Quality estimation for asr k-best list rescoring in spoken language translationabstractSpoken language translation (SLT) combines automatic speech recognition (ASR) and machine translation (MT). During the decoding stage, the best hypothesis produced by the ASR system may not be the best input candidate to the MT system, but making use of multiple sub-optimal ASR results in SLT has been shown to be too complex computationally. This paper presents a method to rescore the k-best ASR output such as to improve translation quality. A translation quality estimation model is trained on a large number of features which aim to capture complementary information from both ASR and MT on translation difficulty and adequacy, as well as syntactic properties of the SLT inputs and outputs. Based on the predicted quality score, the ASR hypotheses are rescored before they are fed to the MT system. ASR confidence is found to be crucial in guiding the rescoring step. In an English-to-French speech-to-text translation task, the coupling of ASR and MT systems led to an increase of 0.5 BLEU points in translation quality. Raymond W. M. Ng, Kashif Shah, Wilker Aziz, Lucia Specia, Thomas Hain |
ICASSP | 5 |
| 2015 | Automatic assessment of English learner pronunciation using discriminative classifiersabstractThis paper presents a novel system for automatic assessment of pronunciation quality of English learner speech, based on deep neural network (DNN) features and phoneme specific discriminative classifiers. DNNs trained on a large corpus of native and non-native learner speech are used to extract phoneme posterior probabilities. A part of the corpus includes per phone teacher annotations, which allows training of two Gaussian Mixture Models (GMM), representing correct pronunciations and typical error patterns. The likelihood ratio is then obtained for each observed phone. Several models were evaluated on a large corpus of English-learning students, with a variety of skill levels, and aged 13 upwards. The cross-correlation of the best system and average human annotator reference scores is 0.72, with miss and false alarm rate around 19%. Automatic assessment is 81.6% correct with a high degree of confidence. The new approach significantly outperforms spectral distance based baseline systems. Mauro Nicolao, Amy V. Beeston, Thomas Hain |
ICASSP | 3 |
| 2015 | Data-selective transfer learning for multi-domain speech recognitionabstractNegative transfer in training of acoustic models for automatic speech recognition has been reported in several contexts such as domain change or speaker characteristics. This paper proposes a novel technique to overcome negative transfer by efficient selection of speech data for acoustic model training. Here data is chosen on relevance for a specific target. A submodular function based on likelihood ratios is used to determine how acoustically similar each training utterance is to a target test set. The approach is evaluated on a wide-domain data set, covering speech from radio and TV broadcasts, telephone conversations, meetings, lectures and read speech. Experiments demonstrate that the proposed technique both finds relevant data and limits negative transfer. Results on a 6--hour test set show a relative improvement of 4% with data selection over using all data in PLP based models, and 2% with DNN features. Mortaza Doulaty, Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 3 |
| 2015 | Unsupervised domain discovery using latent dirichlet allocation for acoustic modelling in speech recognitionabstractSpeech recognition systems are often highly domain dependent, a fact widely reported in the literature. However the concept of domain is complex and not bound to clear criteria. Hence it is often not evident if data should be considered to be out-of-domain. While both acoustic and language models can be domain specific, work in this paper concentrates on acoustic modelling. We present a novel method to perform unsupervised discovery of domains using Latent Dirichlet Allocation (LDA) modelling. Here a set of hidden domains is assumed to exist in the data, whereby each audio segment can be considered to be a weighted mixture of domain properties. The classification of audio segments into domains allows the creation of domain specific acoustic models for automatic speech recognition. Experiments are conducted on a dataset of diverse speech data covering speech from radio and TV broadcasts, telephone conversations, meetings, lectures and read speech, with a joint training set of 60 hours and a test set of 6 hours. Maximum A Posteriori (MAP) adaptation to LDA based domains was shown to yield relative Word Error Rate (WER) improvements of up to 16% relative, compared to pooled training, and up to 10%, compared with models adapted with human-labelled prior domain knowledge. Mortaza Doulaty, Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 3 |
| 2015 | Noise-matched training of CRF based sentence end detection models
Madina Hasan, Rama Sanand Doddipatla, Thomas Hain |
INTERSPEECH | 3 |
| 2015 | Source-filter separation of speech signal in the phase domainabstractDeconvolution of the speech excitation (source) and vocal tract (filter) components through log-magnitude spectral processing is well-established and has led to the well-known cepstral features used in a multitude of speech processing tasks. This paper presents a novel source-filter decomposition based on processing in the phase domain. We show that separation between source and filter in the log-magnitude spectra is far from perfect, leading to loss of vital vocal tract information. It is demonstrated that the same task can be better performed by trend and fluctuation analysis of the phase spectrum of the minimum-phase component of speech, which can be computed via the Hilbert transform. Trend and fluctuation can be separated through low-pass filtering of the phase, using additivity of vocal tract and source in the phase domain. This results in separated signals which have a clear relation to the vocal tract and excitation components. The effectiveness of the method is put to test in a speech recognition task. The vocal tract component extracted in this way is used as the basis of a feature extraction algorithm for speech recognition on the Aurora-2 database. The recognition results shows upto 8.5% absolute improvement in comparison with MFCC features on average (0-20dB). Erfan Loweimi, Jon Barker, Thomas Hain |
INTERSPEECH | 3 |
| 2015 | A study on the stability and effectiveness of features in quality estimation for spoken language translationabstractA quality estimation (QE) approach informed with machine translation (MT) and speech recognition (ASR) features has recently shown to improve the performance of a spoken language translation (SLT) system in an in-domain scenario. When domain mismatch is progressively introduced in the MT and ASR systems, the SLT system’s performance naturally degrades. The use of QE to improve SLT performance has not been studied in this context. In this paper we investigate the effectiveness of QE under this setting. Our experiments showed that across moderate levels of domain mismatches, QE led to consistent translation improvements of around 0.4 in BLEU score. The QE system relies on 116 features derived from the ASR and MT system input and output. Feature analysis was conducted to understand the information sources contributing the most to performance improvements. LDA dimension reduction was used to summarise effective features into sets as small as 3 without affecting the SLT performance. By inspecting the principal components, eight features including the acoustic model scores and count-based word statistics on the bilingual text were found to be critically important, leading to a further boost of around 0.1 BLEU score over the full set of features. These findings provide interesting possibilities for further work by incorporating the effective QE features in SLT system training or decoding. Raymond W. M. Ng, Kashif Shah, Lucia Specia, Thomas Hain |
INTERSPEECH | 4 |
| 2015 | Knowledge transfer between speakers for personalised dialogue managementabstractModel-free reinforcement learning has been shown to be a promising data driven approach for automatic dialogue policy optimization, but a relatively large amount of dialogue interactions is needed before the system reaches reasonable performance.Recently, Gaussian process based reinforcement learning methods have been shown to reduce the number of dialogues needed to reach optimal performance, and pre-training the policy with data gathered from different dialogue systems has further reduced this amount.Following this idea, a dialogue system designed for a single speaker can be initialised with data from other speakers, but if the dynamics of the speakers are very different the model will have a poor performance.When data gathered from different speakers is available, selecting the data from the most similar ones might improve the performance.We propose a method which automatically selects the data to transfer by defining a similarity measure between speakers, and uses this measure to weight the influence of the data from each speaker in the policy model.The methods are tested by simulating users with different severities of dysarthria interacting with a voice enabled environmental control system. Iñigo Casanueva, Thomas Hain, Heidi Christensen, Ricard Marxer, Phil D. Green |
SIGDIAL Conference | 2 |
| 2014 | Using neural network front-ends on far field multiple microphones based speech recognitionabstractThis paper presents an investigation of far field speech recognition using beamforming and channel concatenation in the context of Deep Neural Network (DNN) based feature extraction. While speech enhancement with beamforming is attractive, the algorithms are typically signal-based with no information about the special properties of speech. A simple alternative to beamforming is concatenating multiple channel features. Results presented in this paper indicate that channel concatenation gives similar or better results. On average the DNN front-end yields a 25% relative reduction in Word Error Rate (WER). Further experiments aim at including relevant information in training adapted DNN features. Augmenting the standard DNN input with the bottleneck feature from a Speaker Aware Deep Neural Network (SADNN) shows a general advantage over the standard DNN based recognition system, and yields additional improvements for far field speech recognition. Yulan Liu, Pengyuan Zhang, Thomas Hain |
ICASSP | 3 |
| 2014 | Using contextual information in joint factor eigenspace MLLR for speech recognition in diverse scenariosabstractThis paper presents a new approach for rapid adaptation in the presence of highly diverse scenarios that takes advantage of information describing the input signals. We introduce a new method for joint factorisation of the background and the speaker in an eigenspace MLLR framework: Joint Factor Eigenspace MLLR (JFEMLLR). We further propose to use contextual information describing the speaker and background, such as tags or more complex metadata, to provide an immediate estimation of the best MLLR transformation for the utterance. This provides instant adaptation, since it does not require any transcription from a previous decoding stage. Evaluation in a highly diverse Automatic Speech Recognition (ASR) task, a modified version of WSJCAM0, yields an improvement of 26.9% over the baseline, which is an extra 1.2% reduction over two-pass MLLR adaptation. Oscar Saz-Torralba, Thomas Hain |
ICASSP | 2 |
| 2014 | Adaptive speech recognition and dialogue management for users with speech disordersabstractSpoken control interfaces are very attractive to people with severe physical disabilities who often also have a type of speech disorder known as dysarthria. This condition is known to decrease the accuracy of automatic speech recognisers (ASRs) especially for users with moderate to severe dysathria. In this paper we investigate how applying probabilistic dialogue management (DM) techniques can improve interaction performance of an environmental control system for such users. The effect of having access to different amounts of adaptation data, as well as using different vocabulary size for speakers of different intelligibilities is investigated. We explore the effect of adapting the DM models as the ASR performance increases, such as is the case in systems where more adaptation data is collected through system use. Improvements compared to a non-probabilistic DM baseline are seen both in terms of dialogue length and success rate, 9% and 25% mean relative improvement respectively. Looking at just the more severe dysarthric speakers these numbers rise 25% and 75% mean relative improvement. These improvements are higher when the ASR data adaptation amount is small. Further results show that a DM trained on data from multiple speakers outperform a DM trained on data from a single speaker. Iñigo Casanueva, Heidi Christensen, Thomas Hain, Phil D. Green |
INTERSPEECH | 3 |
| 2014 | Speaker dependent bottleneck layer training for speaker adaptation in automatic speech recognitionabstractSpeaker adaptation of deep neural networks (DNN) is difficult, and most commonly performed by changes to the input of the DNNs. Here we propose to learn discriminative feature transformations to obtain speaker normalised bottleneck (BN) features. This is achieved by interpreting the final two hidden layers as speaker specific matrix transformations. The hidden layer weights are updated with data from a specific speaker to learn speaker-dependent discriminative feature transformations. Such simple implementation lends itself to rapid adaptation and flexibility to be used in Speaker Adaptive Training (SAT) frameworks. The performance of this approach is evaluated on a meeting recognition task, using the official NIST RT’07 and RT’09 evaluation test sets. Supervised adaptation of the BN layer shows similar performance to the application of supervised CMLLR as a global transformation, and the combination of these appears to be additive. In unsupervised mode, CMLLR adaptation only yields 3.4% and 2.5% relative word error rate (WER) improvement, on the RT’07 and RT’09 respectively, where the baselines include speaker based cepstral mean and variance normalisation. The combined CMLLR and BN layer speaker adaptation yields a relative WER gain of 4.5% and 4.2% respectively. SAT style BN layer adaptation is attempted and combined with conventional CMLLR SAT, to show that it provides a relative gain of 1.43% and 2.02% on the RT’07 and RT’09 data sets respectively when compared with CMLLR SAT. While the overall gain from BN layer adaptation is small, the results are found to be statistically significant on both the test sets. Index Terms: Deep neural networks, bottleneck features, speaker adaptation, automatic speech recognition. Rama Sanand Doddipatla, Madina Hasan, Thomas Hain |
INTERSPEECH | 3 |
| 2014 | Extending Limabeam with discrimination and coarse gradientsabstractLimabeam is an approach to multi-microphone array processing for ASR which makes minimal assumptions about system geometry, instead searching for filters to maximise output likelihoods under a speech model. The first results of Limabeam on the AMI meeting corpus are given, then two extensions of the algorithm for this corpus. First, it is shown that the original local gradient following sticks in local minima, and a coarser gradient is used. Second, a new discriminative objective function is provided to handle mismatched silence models. The extensions are based on examination of 2D receptive fields and 2D likelihood maps which are novel near-field analogs of radial beamformer response patterns, but do not show radial symmetry and have many local minima. The extended Limabeam improves WER on TDOA baselines on the AMI corpus, by 1% rel. when both are adapted with decodes and by 19% rel. when both adapted with ground truth. Charles Fox, Thomas Hain |
INTERSPEECH | 2 |
| 2014 | Multi-pass sentence-end detection of lecture speech
Madina Hasan, Rama Sanand Doddipatla, Thomas Hain |
INTERSPEECH | 3 |
| 2014 | Automatic selection of speakers for improved acoustic modelling: recognition of disordered speech with sparse dataabstractThe automatic recognition of disordered speech is a domain that is characterised by limited amounts of training data for each speaker and large intra- and inter-speaker variations. This paper is concerned with how best to train an acoustic models in these circumstances; in particular, we look at how to select data for a background model from a pool of speakers for a given target speaker. We show that rather than including data from all available speakers (the standard approach in the typical speech domain), significantly better accuracy can be achieved by carefully selecting which speakers should contribute. Different methods based on measuring acoustic closeness between speakers and ranking them accordingly are investigated, and on the UASpeech isolated word recognition task, we achieve a 11.5% relative improvement compared to the baseline which uses data from all speakers. Accuracies for speakers with moderate to severe impairments are shown to improve the most with one speaker classed as having `low' intelligibility gaining a 60% relative improvement in accuracy. Heidi Christensen, Iñigo Casanueva, Stuart P. Cunningham, Phil D. Green, Thomas Hain |
SLT | 5 |
| 2014 | Background-tracking acoustic features for genre identification of broadcast showsabstractThis paper presents a novel method for extracting acoustic features that characterise the background environment in audio recordings. These features are based on the output of an alignment that fits multiple parallel background-based Constrained Maximum Likelihood Linear Regression transformations asynchronously to the input audio signal. With this setup, the resulting features can track changes in the audio background like appearance and disappearance of music, applause or laughter, independently of the speakers in the foreground of the audio. The ability to provide this type of acoustic description in audiovisual data has many potential applications, including automatic classification of broadcast archives or improving automatic transcription and subtitling. In this paper, the performance of these features in a genre identification task in a set of 332 BBC shows is explored. The proposed background-tracking features outperform short-term Perceptual Linear Prediction features in this task using Gaussian Mixture Model classifiers (62% vs 72% accuracy). The use of more complex classifiers, Hidden Markov Models and Support Vector Machines, increases the performance of the system with the novel background-tracking features to 79% and 81% in accuracy respectively. Oscar Saz-Torralba, Mortaza Doulaty, Thomas Hain |
SLT | 3 |
| 2014 | Semi-supervised DNN training in meeting recognitionabstractTraining acoustic models for ASR requires large amounts of labelled data which is costly to obtain. Hence it is desirable to make use of unlabelled data. While unsupervised training can give gains for standard HMM training, it is more difficult to make use of unlabelled data for discriminative models. This paper explores semi-supervised training of Deep Neural Networks (DNN) in a meeting recognition task. We first analyse the impact of imperfect transcription on the DNN and the ASR performance. As labelling error is the source of the problem, we investigate two options available to reduce that: selecting data with fewer errors, and changing the dependence on noise by reducing label precision. Both confidence based data selection and label resolution change are explored in the context of two scenarios of matched and unmatched unlabelled data. We introduce improved DNN based confidence score estimators and show their performance on data selection for both scenarios. Confidence score based data selection was found to yield up to 14.6% relative WER reduction, while better balance between label resolution and recognition hypothesis accuracy allowed further WER reductions by 16.6% relative in the mismatched scenario. Pengyuan Zhang, Yulan Liu, Thomas Hain |
SLT | 3 |
| 2014 | Capitalising on North American speech resources for the development of a South African English large vocabulary speech recognition system
Herman Kamper, Febe de Wet, Thomas Hain, Thomas Niesler |
Comput. Speech Lang. | 3 |
| 2013 | Lightly supervised learning from a damaged natural speech corpusabstractLarge corpora of transcribed speech are rare and expensive to acquire, but valuable for ASR systems. Of current research interest are corpora of natural speech, i.e. far-field recordings of multiple speakers in noisy environments. In the big data era there are many speech transcriptions collected for purposes other than ASR, which omit features required by typical ASR systems such as timing information. If we could recover training data from such `found' corpora this would open up large new resources for ASR research. We present a case study for this type of data recovery - becoming known as `lightly supervised learning' - for a highly damaged corpus called Family Life. We use a novel comparison of a parallel decode and forced audio alignment to iteratively select and grow good data. Family Life also has unusual data mislabelling problems which can be addressed by an integrated tfidf approach. These methods reduce WER on the corpus from 83.0 to 57.2. We also discuss a probabilistic loose string alignment approach which removes untranscribed `icebreaker' speech. Charles Fox, Thomas Hain |
ICASSP | 2 |
| 2013 | Adaptation of lecture speech recognition system with machine translation outputabstractIn spoken language translation, integration of the ASR and MT components is critical for good performance. In this paper, we consider the recognition setting where a text translation of each utterance is also available. We present experiments with different ASR system adaptation techniques to exploit MT system outputs. In particular, N-best MT outputs are represented as an utterance-specific language model, which are then used to rescore ASR lattices. We show that this method improves significantly over ASR alone, resulting in an absolute WER reduction of more than 6% for both indomain and out-of-domain acoustic models. Raymond W. M. Ng, Thomas Hain, Trevor Cohn |
ICASSP | 2 |
| 2013 | Combining in-domain and out-of-domain speech data for automatic recognition of disordered speechabstractRecently there has been increasing interest in ways of using out-of-domain (OOD) data to improve automatic speech recognition performance in domains where only limited data is available. This paper focuses on one such domain, namely that of disordered speech for which only very small databases exist, but where normal speech can be considered OOD. Standard approaches for handling small data domains use adaptation from OOD models into the target domain, but here we investigate an alternative approach with its focus on the feature extraction stage: OOD data is used to train feature-generating deep belief neural networks. Using AMI meeting and TED talk datasets, we investigate various tandem-based speaker independent systems as well as maximum a posteriori adapted speaker dependent systems. Results on the UAspeech isolated word task of disordered speech are very promising with our overall best system (using a combination of AMI and TED data) giving a correctness of 62.5 an increase of 15% on previously best published results based on conventional model adaptation. We show that the relative benefit of using OOD data varies considerably from speaker to speaker and is only loosely correlated with the severity of a speaker's impairments. Heidi Christensen, Magda B. Aniol, Peter Bell 0001, Phil D. Green, Thomas Hain, Simon King 0001, Pawel Swietojanski |
INTERSPEECH | 5 |
| 2013 | Learning speaker-specific pronunciations of disordered speechabstractOne of the main clinical applications of speech technology is in voice-enabled assistive technology for people with disordered speech. Progress in this area is hampered by a sparseness in suitable data and recent research have focused on ways of incorporating knowledge about typical (i.e., un-impaired) speech through the use of e.g., deep belief neural networks. This paper presents a new way of using deep belief neural networks trained on typical speech, namely to improve pronunciations for individual speakers. Analysis of the posterior probabilities show a clear correlation between measured pronunciation ‘disorderedness’ and the overall speech recognition performance of the full system. Based on this, we propose a method to use deep belief network outputs to i) identify which words are pronounced differently than what would be expected from a typical pronunciation, and ii) subsequently generate new pronunciations. We investigate different methods for pronunciation generation as well as what is the best way of using the modified pronunciations to inform the system development stages. Using the UAspeech database of disordered speech, we demonstrate improvement in average accuracy of 69.76% to 70.51%, with some speakers showing individual improvements of up to 10%. Heidi Christensen, Phil D. Green, Thomas Hain |
INTERSPEECH | 3 |
| 2013 | The sheffield wargames corpusabstractRecognition of speech in natural environments is a challenging task, even more so if this involves conversations between sev-eral speakers. Work on meeting recognition has addressed some of the significant challenges, mostly targeting formal, business style meetings where people are mostly in a static position in a room. Only limited data is available that contains high qual-ity near and far field data from real interactions between par-ticipants. In this paper we present a new corpus for research on speech recognition, speaker tracking and diarisation, based on recordings of native speakers of English playing a table-top wargame. The Sheffield Wargames Corpus comprises 7 hours of data from 10 recording sessions, obtained from 96 micro-phones, 3 video cameras and, most importantly, 3D location data provided by a sensor tracking system. The corpus repre-sents a unique resource, that provides for the first time location tracks (1.3Hz) of speakers that are constantly moving and talk-ing. The corpus is available for research purposes, and includes annotated development and evaluation test sets. Baseline results for close-talking and far field sets are included in this paper. 1. Charles Fox, Yulan Liu, Erich Zwyssig, Thomas Hain |
INTERSPEECH | 4 |
| 2013 | Asynchronous factorisation of speaker and background with feature transforms in speech recognitionabstractThis paper presents a novel approach to separate the effects of speaker and background conditions by application of featuretransform based adaptation for Automatic Speech Recognition (ASR).So far factorisation has been shown to yield improvements in the case of utterance-synchronous environments.In this paper we show successful separation of conditions asynchronous with speech, such as background music.Our work takes account of the asynchronous nature of the background, by estimation of condition-specific Constrained Maximum Likelihood Linear Regression (CMLLR) transforms.In addition, speaker adaptation is performed, allowing to factorise speaker and background effects.Equally, background transforms are used asynchronously in the decoding process, using a modified Hidden Markov Model (HMM) topology which applies the optimal transform for each frame.Experimental results are presented on the WSJCAM0 corpus of British English speech, modified to contain controlled sections of background music.This addition of music degrades the baseline Word Error Rate (WER) from 10.1% to 26.4%.While synchronous factorisation with CMLLR transforms provides 28% relative improvement in WER over the baseline, our asynchronous approach increases this reduction to 33%. Oscar Saz-Torralba, Thomas Hain |
INTERSPEECH | 2 |
| 2012 | Application of SVM-based correctness predictions to unsupervised discriminative speaker adaptationabstractThe effectiveness of unsupervised speaker adaptation is typically limited by errors in the estimated transcription of the adaptation data. Previous work has mitigated this negative effect by using only those sections of the adaptation data which are transcribed with relatively high confidence. In this work, phoneme correctness predictions are integrated into a discriminative unsupervised speaker adaptation procedure. Significant accuracy improvements (over the equivalent likelihood-based technique) are observed when using discriminative unsupervised speaker adaptation in combination with support vector machines to predict phoneme correctness. Matthew Gibson, Thomas Hain |
ICASSP | 2 |
| 2012 | CRF-based Diacritisation of Colloquial Arabic for Automatic Speech Recognition
Sarah Al-Shareef, Thomas Hain |
INTERSPEECH | 2 |
| 2012 | A comparative study of adaptive, automatic recognition of disordered speechabstractSpeech-driven assistive technology can be an attractive alternative to conventional interfaces for people with physical disabilities. However, often the lack of motor-control of the speech articulators results in disordered speech, as condition known as dysarthria. Dysarthric speakers can generally not obtain satisfactory performances with off-the-shelf automatic speech recognition (ASR) products and disordered speech ASR is an increasingly active research area. Sparseness of suitable data is a big challenge. The experiments described here use UAspeech, one of the largest dysarthric databases available, which is still easily an order of magnitude smaller than typical speech databases. This study investigates how far fundamental training and adaptation techniques developed in the LVCSR community can take us. A variety of ASR systems using maximum likelihood and MAP adaptation strategies are established with all speakers obtaining significant improvements compared to the baseline system regardless of the severity of their condition. The best systems show on average 34% relative improvement on known published results. An analysis of the correlation between intelligibility of the speaker and the type of system which would represent an optimal operating point in terms of performance shows that for severely dysarthric speakers, the exact choice of system configuration is more critical than for speakers with less disordered speech. Heidi Christensen, Stuart P. Cunningham, Charles Fox, Phil D. Green, Thomas Hain |
INTERSPEECH | 5 |
| 2012 | Supervised and unsupervised Web-based language model domain adaptationabstractDomain language model adaptation consists in re-estimating probabilities of a baseline LM in order to better match the specifics of a given broad topic of interest. To do so, a common strategy is to retrieve adaptation texts from the Web based on a given domain-representative seed text. In this paper, we study how the selection of this seed text influences the adaptation process and the performances of resulting adapted language models in automatic speech recognition. More precisely, the goal of this original study is to analyze the differences of our Web-based adaptation approach between the supervised case, in which the seed text is manually generated, and the unsupervised case, where the seed text is given by an automatic transcript. Experiments were carried out on data sourced from a real-world use case, more specifically, videos produced for a university YouTube channel. Results show that our approach is quite robust since the unsupervised adaptation provides similar performance to the supervised case in terms of the overall perplexity and word error rate. Gwénolé Lecorvé, John Dines, Thomas Hain, Petr Motlícek |
INTERSPEECH | 3 |
| 2012 | An alignment matching method to explore pseudosyllable properties across different corpora
Raymond W. M. Ng, Thomas Hain, Keikichi Hirose |
INTERSPEECH | 2 |
| 2012 | Automatic transcription of academic lectures from diverse disciplinesabstractIn a multimedia world it is now common to record professional presentations, on video or with audio only. Such recordings include talks and academic lectures, which are becoming a valuable resource for students and professionals alike. However, organising such material from a diverse set of disciplines seems to be not an easy task. One way to address this problem is to build an Automatic Speech Recognition (ASR) system in order to use its output for analysing such materials. In this work ASR results for lectures from diverse sources are presented. The work is based on a new collection of data, obtained by the Liberated Learning Consortium (LLC). The study's primary goals are two-fold: first to show variability across disciplines from an ASR perspective, and how to choose sources for the construction of language models (LMs); second, to provide an analysis of the lecture transcription for automatic determination of structures in lecture discourse. In particular, we investigate whether there are properties common to lectures from different disciplines. This study focuses on textual features. Lectures are multimodal experiences - it is not clear whether textual features alone are sufficient for the recognition of such common elements, or other features, e.g. acoustic features such as the speaking rate, are needed. The results show that such common properties are retained across disciplines even on ASR output with a Word Error Rate (WER) of 30%. Ghada AlHarbi, Thomas Hain |
SLT | 2 |
| 2012 | Correctness-Adjusted Unsupervised Discriminative Acoustic Model AdaptationabstractUnsupervised acoustic model adaptation for large vocabulary speech recognition is typically accomplished by using an estimated transcription of the adaptation data. The effectiveness of the technique is limited by errors in the estimated transcription. Previous work has mitigated this negative effect by using only those sections of the adaptation data which are transcribed with relatively high confidence. In this work, phoneme correctness predictions are integrated into a discriminative unsupervised acoustic model adaptation procedure. Small but significant performance improvements (over the equivalent maximum likelihood adaptation technique) are observed when using unsupervised discriminative adaptation in combination with support vector machines to predict phoneme correctness. Matthew Gibson, Thomas Hain |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Transcribing Meetings With the AMIDA SystemsabstractIn this paper, we give an overview of the AMIDA systems for transcription of conference and lecture room meetings. The systems were developed for participation in the Rich Transcription evaluations conducted by the National Institute for Standards and Technology in the years 2007 and 2009 and can process close talking and far field microphone recordings. The paper first discusses fundamental properties of meeting data with special focus on the AMI/AMIDA corpora. This is followed by a description and analysis of improved processing and modeling, with focus on techniques specifically addressing meeting transcription issues such as multi-room recordings or domain variability. In 2007 and 2009, two different strategies of systems building were followed. While in 2007 we used our traditional style system design based on cross adaptation, the 2009 systems were constructed semi-automatically, supported by improved decoders and a new method for system representation. Overall these changes gave a 6%-13% relative reduction in word error rate compared to our 2007 results while at the same time requiring less training material and reducing the real-time factor by five times. The meeting transcription systems are available at www.webasr.org. Thomas Hain, Lukás Burget, John Dines, Philip N. Garner, Frantisek Grézl, Asmaa El Hannani, Marijn Huijbregts, Martin Karafiát, Mike Lincoln, Vincent Wan |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | An Investigation in Speech Recognition for Colloquial ArabicabstractThis paper describes a study of grapheme-based speech recognition for colloquial Arabic. An investigation of language and acoustic model configurations is carried out to illustrate the differences between colloquial and modern standard Arabic (MSA) on the example of Levantine telephone conversations. The study defines extensive and carefully crafted data sets for different dialects and studies their overlap with MSA sources. The use of grapheme models is re-investigated, and alternative configuration for acoustic models to correct obvious shortcomings are tested. The recognition performance was analyzed on two levels: corpus-level and dialect-level. In addition modifications of dictionaries to allow better specification of sound patterns is explored. Overall the experiments highlight the need for higher level information on acoustic model selection. Sarah Al-Shareef, Thomas Hain |
INTERSPEECH | 2 |
| 2011 | Cross-Language Phone Recognition when the Target Language Phoneme Inventory is not KnownabstractCross-language speech recognition often assumes a certain amount of knowledge about the target language. However, there are hundreds of languages where not even the phoneme inven-tory is known. In the work reported here, phone recognisers are evaluated on a cross-language task with minimum target knowl-edge. A phonetic distance measure is introduced for the evalua-tion, allowing a distance to be calculated between any utterance of any language. This has a number of spin-off applications such as allophone detection, a phone-based ROVER approach to recognition, and cross-language forced alignment. Results show that some of these novel approaches will be of immediate use in characterising languages where there is little phonological knowledge. Index Terms: universal phone recognition, cross-language speech recognition, forced alignment, under-resourced lan-guages 1. Timothy Kempton, Roger K. Moore, Thomas Hain |
INTERSPEECH | 3 |
| 2011 | An Analysis of Automatic Speech Recognition with Multiple MicrophonesabstractAutomatic speech recognition in real world situations often requires the use of microphones distant from speaker’s mouth. One or several microphones are placed in the surroundings to capture many versions of the original signal. Recognition with a single far field microphone yields considerably poorer performance than with person-mounted devices (headset, lapel), with the main causes being reverberation and noise. Acoustic beamforming techniques allow significant improvements over the use of a single microphone, although the overall performance still remains well above the close-talking results. In this paper we investigate the use of beam-forming in the context of speaker movement, together with commonly used adaptation techniques and compare against a naive multi-stream approach. We show that even such a simple approach can yield equivalent results to beam-forming, allowing for far more powerful integration of multiple microphone sources in ASR systems. Davide Marino, Thomas Hain |
INTERSPEECH | 2 |
| 2011 | Extending Audio Notetaker to Browse WebASR Transcriptions
Roger C. F. Tucker, Dan Fry, Vincent Wan, Stuart N. Wrigley, Thomas Hain |
INTERSPEECH | 5 |
| 2011 | Web-Based Automatic Speech Recognition Service - webASRabstractA state-of-the-art automatic speech recognition (ASR) system was developed as part of the AMIDA project whose core domain was the transcription of small to medium sized meetings. The system has performed well in recent NIST evaluations (RT’07 and RT’09). This research-grade ASR system has now been made available as a free web service (webASR) targeting non-commercial researchers. Access to the service is via and standard browser-based interface as well as an API. The service provides the facility to upload audio recordings which are then processed by the ASR system to produce a word-level transcript. Such transcripts are available in a range of formats to suite different needs and technical expertise. The API allows the core webASR functionality to be integrated seamlessly into applications and services. Detailed descriptions of the system design and user interface are provided. Index Terms: speech recognition, interface, api, transcription, web service Stuart N. Wrigley, Thomas Hain |
INTERSPEECH | 2 |
| 2011 | Making an Automatic Speech Recognition Service Freely Available on the Web
Stuart N. Wrigley, Thomas Hain |
INTERSPEECH | 2 |
| 2010 | The AMIDA 2009 meeting transcription systemabstractWe present the AMIDA 2009 system for participation in the NIST RT’2009 STT evaluations. Systems for close-talking, far field and speaker attributed STT conditions are described. Im- provements to our previous systems are: segmentation and diar- isation; stacked bottle-neck posterior feature extraction; fMPE training of acoustic models; adaptation on complete meetings; improvements to WFST decoding; automatic optimisation of decoders and system graphs. Overall these changes gave a 6- 13% relative reduction in word error rate while at the same time reducing the real-time factor by a factor of five and using con- siderably less data for acoustic model training. Thomas Hain, Lukás Burget, John Dines, Philip N. Garner, Asmaa El Hannani, Marijn Huijbregts, Martin Karafiát, Mike Lincoln, Vincent Wan |
INTERSPEECH | 1 |
| 2010 | Automatic Optimization of Speech Decoder ParametersabstractModern speech decoders are complex with potentially a large number of parameters that allow tuning for performance and speed. In this paper we investigate methods for automatic optimization of such parameters. The objective is to find the optimal configuration that yields minimal search errors for any real-time factor. We propose a solution for this multiobjective optimization problem based on automatic tracking of that optimal curve. Two cost functions for tracking are investigated as well as techniques to enhance stability. Experiments, conducted using the large vocabulary speech decoder HDecode from the Hidden Markov Model toolkit, show on a large test set of conversation telephone speech that with modest computational cost optimal performance curves for specific decoders and data types can be obtained. Careful selection of the cost function allows a further reduction of computational cost by 55%. As no prior knowledge about the interpretation of the parameters is used, the proposed method is applicable to other decoders. Asmaa El Hannani, Thomas Hain |
IEEE Signal Process. Lett. | 2 |
| 2010 | Error Approximation and Minimum Phone Error Acoustic Model EstimationabstractMinimum phone error (MPE) acoustic parameter estimation involves calculation of edit distances (errors) between correct and incorrect hypotheses. In the context of large-vocabulary continuous-speech recognition, this error calculation becomes prohibitively expensive and so errors are approximated. This paper introduces a novel error approximation technique. Analysis shows that this approximation yields a higher correlation to the Levenshtein error metric than a previously used approximation. Experimental evaluations on a large-vocabulary recognition task demonstrate that the novel approximation also delivers significant performance improvements over the previously used approximation when applied to MPE acoustic model estimation. Matt Gibson 0002, Thomas Hain |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Real-time ASR from meetingsabstractThe AMI(DA) system is a meeting room speech recognition system that has been developed and evaluated in the context of the NIST Rich Text (RT) evaluations. Recently, the "Distant Access" requirements of the AMIDA project have necessitated that the system operate in real-time. Another more difficult requirement is that the system fit into a live meeting transcription scenario. We describe an infrastructure that has allowed the AMI(DA) system to evolve into one that fulfils these extra requirements. We emphasise the components that address the live and real-time aspects. Philip N. Garner, John Dines, Thomas Hain, Asmaa El Hannani, Martin Karafiát, Danil Korchagin, Mike Lincoln, Vincent Wan, Le Zhang 0002 |
INTERSPEECH | 3 |
| 2008 | Automatic speech recognition for scientific purposes - webASRabstractWe present ‘webASR’, an online interface to our state-of-the-art automatic speech recognition (ASR) systems. It aims to provide the wider scientific research community with an interface to speech transcription for domains and applications where the generation of such transcripts was not previously feasible. The webASR interface allows the upload of audio files and, in turn, the download of automatically generated ASR transcripts. Depending upon the specification given for an audio file, the system will transcribe using an appropriate speech recogniser chosen from one of the many available, such as a NIST RT evaluation system. The transcripts will be available for download after processing. Thomas Hain, Asmaa El Hannani, Stuart N. Wrigley, Vincent Wan |
INTERSPEECH | 1 |
| 2008 | Discrimininative training of narrow band - wide band adapted systems for meeting recognitionabstractThe amount of training data has a crucial effect on the accuracy of HMM based meeting recognition systems. One of the largest collections of speech data is conversational telephone speech which was found to match speech in meetings well. However it is naturally recorded with limited bandwidth. In previous work we presented a scheme that allows to transform wide-band meeting data into the same space for improved model training. In this paper we focused on integration of discriminative adaptation into this scheme. This integration is not straightforward and we present the complexity of this process. The models are tested on the NIST RT’05 meeting evaluation where a relative reduction in word error rate of 5.6% against non-adapted meeting system was achieved. Martin Karafiát, Lukás Burget, Thomas Hain, Jan Cernocký |
INTERSPEECH | 3 |
| 2008 | Bob: A lexicon and pronunciation dictionary generatorabstractThis paper presents Bob, a tool for managing lexicons and generating pronunciation dictionaries for automatic speech recognition systems. It aims to maintain a high level of consistency between lexicons and language modelling corpora by managing the text normalisation and lexicon generation processes in a single dedicated package. It also aims to maintain consistent pronunciation dictionaries by generating pronunciation hypotheses automatically and aiding their verification. The tool's design and functionality are described. Also two case studies highlighting the importance of consistency and illustrating the use of the tool are reported. Vincent Wan, John Dines, Asmaa El Hannani, Thomas Hain |
SLT | 4 |
| 2007 | Recognition and understanding of meetings the AMI and AMIDA projectsabstractThe AMI and AMIDA projects are concerned with the recognition and interpretation of multiparty meetings. Within these projects we have: developed an infrastructure for recording meetings using multiple microphones and cameras; released a 100 hour annotated corpus of meetings; developed techniques for the recognition and interpretation of meetings based primarily on speech recognition and computer vision; and developed an evaluation framework at both component and system levels. In this paper we present an overview of these projects, with an emphasis on speech recognition and content extraction. Steve Renals, Thomas Hain, Hervé Bourlard |
ASRU | 2 |
| 2007 | The AMI System for the Transcription of Speech in MeetingsabstractThis paper describes the AMI transcription system for speech in meetings developed in collaboration by five research groups. The system includes generic techniques such as discriminative and speaker adaptive training, vocal tract length normalisation, heteroscedastic linear discriminant analysis, maximum likelihood linear regression, and phone posterior based features, as well as techniques specifically designed for meeting data. These include segmentation and cross-talk suppression, beam-forming, domain adaptation, Web-data collection, and channel adaptive training. The system was improved by more than 20% relative in word error rate compared to our previous system and was used in the NIST RT106 evaluations where it was found to yield competitive performance. Thomas Hain, Vincent Wan, Lukás Burget, Martin Karafiát, John Dines, Jithendra Vepa, Giulia Garau, Mike Lincoln |
ICASSP (4) | 1 |
| 2007 | Temporal masking for unsupervised minimum Bayes risk speaker adaptationabstractThe minimum Bayes risk (MBR) criterion has previously been applied to the task of speaker adaptation in large vocabulary continuous speech recognition. The success of unsupervised MBR speaker adaptation, however, has been limited by the accuracy of the estimated transcription of the acoustic data. This paper addresses this issue not by improving the accuracy of the estimated transcription but via temporal masking of its erroneous regions. 1. Matthew Gibson, Thomas Hain |
INTERSPEECH | 2 |
| 2007 | Application of CMLLR in narrow band wide band adapted systems
Martin Karafiát, Lukás Burget, Jan Cernocký, Thomas Hain |
INTERSPEECH | 4 |
| 2006 | Strategies for Language Model Web-Data CollectionabstractThis paper presents an analysis of the use of textual information collected from the Internet via a search engine for the purpose of building domain specific language models. A framework to analyse the effect of search query formulation on the resulting Web-data language model performance in an evaluation is developed. The framework gives rise to improved methods of selecting n-gram search engine queries, which return documents that make better domain specific language models Vincent Wan, Thomas Hain |
ICASSP (1) | 2 |
| 2006 | The segmentation of multi-channel meeting recordings for automatic speech recognitionabstractOne major research challenge in the domain of the analysis of meeting room data is the automatic transcription of what is spoken during meetings, a task which has gained considerable attention within the ASR research community through the NIST rich transcription evaluations conducted over the last three years. One of the major difficulties in carrying out automatic speech recognition (ASR) on this data is dealing with the challenging recording environment, which has instigated the development of novel audio pre-processing approaches. In this paper we present a system for the automatic segmentation of multiple-channel individual headset microphone (IHM) meeting recordings for automatic speech recognition. The system relies on an MLP classifier trained from several meeting room corpora to identify speech/non-speech segments of the recordings. We give a detailed analysis of the segmentation performance for a number of system configurations, with our best system achieving ASR performance on automatically generated segments within 1.3\% (3.7\% relative) of a manual segmentation of the data. John Dines, Jithendra Vepa, Thomas Hain |
INTERSPEECH | 3 |
| 2006 | Hypothesis spaces for minimum Bayes risk training in large vocabulary speech recognitionabstractThe Minimum Bayes Risk (MBR) framework has been a successful strategy for the training of hidden Markov models for large vocabulary speech recognition. Practical implementations of MBR must select an appropriate hypothesis space and loss function. The set of word sequences and a word-based Levenshtein distance may be assumed to be the optimal choice but use of phoneme-based criteria appears to be more successful. This paper compares the use of different hypothesis spaces and loss functions defined using the system constituents of word, phone, physical triphone, physical state and physical mixture component. For practical reasons the competing hypotheses are constrained by sampling. The impact of the sampling technique on the performance of MBR training is also examined. Index Terms: discriminative training, Minimum Bayes Risk. 1. Matthew Gibson, Thomas Hain |
INTERSPEECH | 2 |
| 2006 | Automatic speech recognition experiments with articulatory dataabstractIn this paper we investigate the use of articulatory data for speech recognition. Recordings of the articulatory movements originate from the MOCHA corpus, a database which contains speech, EGG, EMA and EPG recordings. It was found that in a Hidden Markov Model (HMM) based recognition framework careful processing of these signals can yield significantly better performance than that obtained by decoding of the acoustic signals. We present detailed results on the processing of the signals and the associated performance of monophone and triphone systems. Experimental evidence shows that acoustic-signal-to-word mappings and articulatory-signal-to-word mappings are equally complex. However, for the latter, evidence of short-comings of standard HMM based modelling is visible and should be addressed in future systems. Esmeralda Uraga, Thomas Hain |
INTERSPEECH | 2 |
| 2006 | Corrections to "Automatic Transcription of Conversational Telephone Speech"
Thomas Hain, Philip C. Woodland, Gunnar Evermann, Mark J. F. Gales, Xunying Liu, Gareth L. Moore, Daniel Povey |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Applying vocal tract length normalization to meeting recordingsabstractVocal Tract Length Normalisation (VTLN) is a commonly used technique to normalise for inter-speaker variability. It is based on the speaker-specific warping of the frequency axis, parameterised by a scalar warp factor. This factor is typically estimated using maximum likelihood. We discuss how VTLN may be applied to multiparty conversations, reporting a substantial decrease in word error rate in experiments using the ICSI meetings corpus. We investigate the behaviour of the VTLN warping factor and show that a stable estimate is not obtained. Instead it appears to be influenced by the context of the meeting, in particular the current conversational partner. These results are consistent with predictions made by the psycholinguistic interactive alignment account of dialogue, when applied at the acoustic and phonological levels. Giulia Garau, Steve Renals, Thomas Hain |
INTERSPEECH | 3 |
| 2005 | Transcription of conference room meetings: an investigationabstractThe automatic processing of speech collected in conference style meetings has attracted considerable interest with several large scale projects devoted to this area. In this paper we explore the use of various meeting corpora for the purpose of automatic speech recognition. In particular we investigate the similarity of these resources and how to efficiently use them in the construction of a meeting transcription system. The analysis shows distinctive features for each resource. However the benefit in pooling data and hence the similarity seems sufficient to speak of a generic conference meeting domain . In this context this paper also presents work on development for the AMI meeting transcription system, a joint effort by seven sites working on the AMI (augmented multi-party interaction) project. Thomas Hain, John Dines, Giulia Garau, Martin Karafiát, Darren Moore, Vincent Wan, Roeland Ordelman, Steve Renals |
INTERSPEECH | 1 |
| 2005 | Implicit modelling of pronunciation variation in automatic speech recognition
Thomas Hain |
Speech Commun. | 1 |
| 2005 | Automatic transcription of conversational telephone speechabstractThis paper discusses the Cambridge University HTK (CU-HTK) system for the automatic transcription of conversational telephone speech. A detailed discussion of the most important techniques in front-end processing, acoustic modeling and model training, language and pronunciation modeling are presented. These include the use of conversation side based cepstral normalization, vocal tract length normalization, heteroscedastic linear discriminant analysis for feature projection, minimum phone error training and speaker adaptive training, lattice-based model adaptation, confusion network based decoding and confidence score estimation, pronunciation selection, language model interpolation, and class based language models. The transcription system developed for participation in the 2002 NIST Rich Transcription evaluations of English conversational telephone speech data is presented in detail. In this evaluation the CU-HTK system gave an overall word error rate of 23.9%, which was the best performance by a statistically significant margin. Further details on the derivation of faster systems with moderate performance degradation are discussed in the context of the 2002 CU-HTK 10 /spl times/ RT conversational speech transcription system. Thomas Hain, Philip C. Woodland, Gunnar Evermann, Mark J. F. Gales, Xunying Liu, Gareth L. Moore, Daniel Povey |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Development of the 2003 CU-HTK conversational telephone speech transcription systemabstractThe paper describes the development of the 2003 CU-HTK large vocabulary speech recognition system for conversational telephone speech (CTS). The system was designed based on a multipass, multibranch structure where the output of all branches is combined using system combination. A number of advanced modelling techniques, such as speaker adaptive training, heteroscedastic linear discriminant analysis, minimum phone error estimation and specially constructed single pronunciation dictionaries, were employed. The effectiveness of each of these techniques and their potential contribution to the result of system combination was evaluated in the framework of a state-of-the-art LVCSR system with sophisticated adaptation. The final 2003 CU-HTK CTS system constructed from some of these models is described and its performance on the DARPA/NIST 2003 rich transcription (RT-03) evaluation test set is discussed. Gunnar Evermann, Ricky Ho Yin Chan, Mark J. F. Gales, Thomas Hain, Xunying Liu, David Mrva, Philip C. Woodland |
ICASSP (1) | 4 |
| 2004 | Using VTLN for broadcast news transcriptionabstractVocal tract length normalisation (VTLN) is a commonly used speaker normalisation approach. It is attractive compared to many normalisation schemes as it is typically dependent on only a single parameter, allowing the warp factors to be robustly calculated on little data. However, the scheme normally requires explicitly coding the data at multiple warp factors. Furthermore, it is only possible to approximate the Jacobian associated with the VTLN transformation. A new, simple, linear approximation to VTLN is described in this paper. This linear approximation allows the Jacobian to be exactly computed. It can also be highly efficient in terms of warp factor estimation and application of the warp factors. Both the linear and standard CUED VTLN schemes were evaluated in the 2003 BNE evaluation framework and found to yield similar performance. When used in system combination both VTLN schemes yielded slight gains over the baseline system. Do Yeong Kim, Srinivasan Umesh, Mark J. F. Gales, Thomas Hain, Philip C. Woodland |
INTERSPEECH | 4 |
| 2001 | New features in the CU-HTK system for transcription of conversational telephone speechabstractDiscusses new features integrated into the Cambridge University HTK (CU-HTK) system for the transcription of conversational telephone speech. Major improvements have been achieved by the use of maximum mutual information estimation in training as well as maximum likelihood estimation; the use of a full variance transform for adaptation; the inclusion of unigram pronunciation probabilities; and word-level posterior probability estimation using confusion networks for use in minimum word error rate decoding, confidence score estimation and system combination. Improvements are demonstrated via performance on the NIST March 2000 evaluation of English conversational telephone speech transcription (Hub5E). In this evaluation the CU-HTK system gave an overall word error rate of 25.4%, which was the best performance by a statistically significant margin. Thomas Hain, Philip C. Woodland, Gunnar Evermann, Daniel Povey |
ICASSP | 1 |
| 2000 | Modelling sub-phone insertions and deletions in continuous speech recognition
Thomas Hain, Philip C. Woodland |
INTERSPEECH | 1 |
| 1999 | The 1998 HTK system for transcription of conversational telephone speechabstractThis paper describes the 1998 HTK large vocabulary speech recognition system for conversational telephone speech as used in the NIST 1998 Hub5E evaluation. Front-end and language modelling experiments conducted using various training and test sets from both the Switchboard and Callhome English corpora are presented. Our complete system includes reduced bandwidth analysis, side-based cepstral feature normalisation, vocal tract length normalisation (VTLN), triphone and quinphone hidden Markov models (HMMs) built using speaker adaptive training (SAT), maximum likelihood linear regression (MLLR) speaker adaptation and a confidence score based system combination. A detailed description of the complete system together with experimental results for each stage of our multi-pass decoding scheme is presented. The word error rate obtained is almost 20% better than our 1997 system on the development set. Thomas Hain, Philip C. Woodland, Thomas Niesler, Edward W. D. Whittaker |
ICASSP | 1 |
| 1999 | Dynamic HMM selection for continuous speech recognitionabstractIn this paper we propose a dynamic model selection technique based on hidden model sequences (HMS). HMS modelling assumes, that not only the actual state sequence is unknown, but also the model sequence given a particular sentence. This allows more than one model to be used for a particular phone in a certain context. The most appropriate model is determined locally rather than a priori globally by the acoustic probability of that model together with a probability that this model is produced in a particular phone (or model) context. Experiments on the Resource Management corpus show significant improvements in word error rate over phonetically model-- and state--tied triphone hidden Markov models (HMMs). Initial results on the Switchboard corpus also show improvements on a much more difficult task. 1. INTRODUCTION In HMM-based continuous speech recognition ideally each possible sentence would be modelled with a separate Markov model. Since the set of possible sentences is far too larg... Thomas Hain, Philip C. Woodland |
EUROSPEECH | 1 |
| 1999 | Improvements in accuracy and speed in the HTK broadcast news transcription system
Philip C. Woodland, J. J. Odell, Thomas Hain, Gareth L. Moore, Thomas Niesler, Andreas Tuerk, Edward W. D. Whittaker |
EUROSPEECH | 3 |
| 1998 | Experiments in broadcast news transcriptionabstractThis paper presents the development of the HTK broadcast news transcription system. Previously we have used data type specific modelling based on adapted Wall Street Journal trained HMMs. However, we are now experimenting with data for which no manual pre-classification or segmentation is available and therefore automatic techniques are required and compatible acoustic modelling strategies adopted. An approach for automatic audio segmentation and classification is described and evaluated as well as extensions to our previous work on segment clustering. A number of recognition experiments are presented that compare datatype specific and non-specific models; differing amounts of training data; the use of gender-dependent modelling and the effects of automatic data-type classification. It is shown that robust segmentation into a small number of audio types is possible and that models trained on a wide variety of data types can yield good performance. Philip C. Woodland, Thomas Hain, Sue Tranter, Thomas Niesler, Andreas Tuerk, Steve J. Young |
ICASSP | 2 |
| 1998 | Segmentation and classification of broadcast news audioabstractBroadcast news contains a wide variety of different speakers and audio conditions (channel and background noise). This paper describes a segmentation, gender detection and audio classification scheme and presents experimental results on the DARPA 1997 broadcast news evaluation set. Thomas Hain, Philip C. Woodland |
ICSLP | 1 |
| 1994 | On the convergence of fractal transformsabstractThis paper reports on investigations concerning the convergence of fractal transforms for signal modelling. Convergence is essential for the functionality of fractal based coding schemes. The coding process is described as non-linear transformation in the finite-dimensional vector space. Using spectral theory, a necessary and sufficient condition for the contractivity is derived from the eigenvalues of a special linear operator. In the same way some constraints for the choice of the encoding parameters are deduced which are less strict than those imposed so far. The proposed contractivity measure can be calculated directly from the transformation parameters during the encoding process. For complex encoding schemes the calculation of the eigenvalues may be infeasible. For those cases a contractivity criterion derived from the norm of the operator is suggested.> Bemd Hurtgen, Thomas Hain |
ICASSP (5) | 2 |