Philip C. Woodland

dblp:42/153 · DBLP profile ↗
← Back
235ranked-venue papers
15as first author
49since 2021 · last 2026
0000-0001-9069-0225ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 188 · 13 first-author · 31 since 2021Artificial intelligence and machine learning · 130 · 5 first-author · 30 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Protecting Bystander Privacy via Selective Hearing in Audio LLMs
abstract
Audio Large language models (LLMs) are increasingly deployed in the real world, where they inevitably capture speech from unintended nearby bystanders, raising privacy risks that existing benchmarks and defences did not consider. We introduce SH-Bench, the first benchmark designed to evaluate selective hearing: a model's ability to attend to an intended main speaker while refusing to process or reveal information about incidental bystander speech. SH-Bench contains 3,968 multi-speaker audio mixtures, including both real-world and synthetic scenarios, paired with 77k multiple-choice questions that probe models under general and selective operating modes. In addition, we propose Selective Efficacy (SE), a novel metric capturing both multi-speaker comprehension and bystander-privacy protection. Our evaluation of state-of-the-art open-source and proprietary LLMs reveals substantial bystander privacy leakage, with strong audio understanding failing to translate into selective protection of bystander privacy. To mitigate this gap, we also present Bystander Privacy Fine-Tuning (BPFT), a novel training pipeline that teaches models to refuse bystander-related queries without degrading main-speaker comprehension. We show that BPFT yields substantial gains, achieving an absolute 47% higher bystander accuracy under selective mode and an absolute 16% higher SE compared to Gemini 2.5 Pro, which is the best audio LLM without BPFT. Together, SH-Bench and BPFT provide the first systematic framework for measuring and improving bystander privacy in audio LLMs.
Xiao Zhan, Guangzhi Sun, Jose M. Such, Philip C. Woodland
ACL (1)4
2025 SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
abstract
Simultaneous speech translation (SST) outputs translations in parallel with streaming speech input, balancing translation quality and latency.While large language models (LLMs) have been extended to handle the speech modality, streaming remains challenging as speech is prepended as a prompt for the entire generation process.To unlock LLM streaming capability, this paper proposes SimulS2S-LLM, which trains speech LLMs offline and employs a test-time policy to guide simultaneous inference.SimulS2S-LLM alleviates the mismatch between training and inference by extracting boundary-aware speech prompts that allows it to be better matched with text input data.SimulS2S-LLM achieves simultaneous speech-to-speech translation (Simul-S2ST) by predicting discrete output speech tokens and then synthesising output speech using a pretrained vocoder.An incremental beam search is designed to expand the search space of speech token prediction without increasing latency.Experiments on the CVSS speech data show that SimulS2S-LLM offers a better translation quality-latency trade-off than existing methods that use the same training data, such as improving ASR-BLEU scores by 3 points at similar latency.
Keqi Deng, Xie Chen 0001, Philip C. Woodland
ACL (1)4
2025 SkillAggregation: Reference-free LLM-Dependent Aggregation
abstract
Large Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments.Single LLMs were used initially, however, recent work suggests using multiple LLMs as judges yields improved performance.An important step in exploiting multiple judgements is the combination stage, aggregation.Existing methods in NLP either assign equal weight to all LLM judgments or are designed for specific tasks such as hallucination detection.This work focuses on aggregating predictions from multiple systems where no reference labels are available.A new method called SkillAggregation is proposed, which learns to combine estimates from LLM judges without needing additional data or ground truth.It extends the Crowdlayer aggregation method, developed for image classification, to exploit the judge estimates during inference.The approach is compared to a range of standard aggregation methods on HaluEval-Dialogue, TruthfulQA and Chatbot Arena tasks.SkillAggregation outperforms Crowdlayer on all tasks, and yields the best performance over all approaches on the majority of tasks. 1
Guangzhi Sun, Anmol Kagrecha, P. P. Manakul, Philip C. Woodland, Mark J. F. Gales
ACL (1)4
2025 DNCASR: End-to-End Training for Speaker-Attributed ASR
abstract
This paper introduces DNCASR, a novel endto-end trainable system designed for joint neural speaker clustering and automatic speech recognition (ASR), enabling speaker-attributed transcription of long multi-party meetings.DNCASR uses two separate encoders to independently encode global speaker characteristics and local waveform information, along with two linked decoders to generate speaker-attributed transcriptions.The use of linked decoders allows the entire system to be jointly trained under a unified loss function.By employing a serialised training approach, DNCASR effectively addresses overlapping speech in real-world meetings, where the link improves the prediction of speaker indices in overlapping segments.Experiments on the AMI-MDM meeting corpus demonstrate that the jointly trained DNCASR outperforms a parallel system that does not have links between the speaker and ASR decoders.Using cpWER to measure the speaker-attributed word error rate, DNCASR achieves a 9.0% relative reduction on the AMI-MDM Eval set.
Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland
ACL (1)3
2025 Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
abstract
While large language models (LLMs) have been applied to automatic speech recognition (ASR), the task of making the model streamable remains a challenge. This paper proposes a novel model architecture, Transducer-Llama, that integrates LLMs into a Factorized Transducer (FT) model, naturally enabling streaming capabilities. Furthermore, given that the large vocabulary of LLMs can cause data sparsity issue and increased training costs for spoken language systems, this paper introduces an efficient vocabulary adaptation technique to align LLMs with speech system vocabularies. The results show that directly optimizing the FT model with a strong pre-trained LLM-based predictor using the RNN-T loss yields some but limited improvements over a smaller pre-trained LM predictor. Therefore, this paper proposes a weak-to-strong LM swap strategy, using a weak LM predictor during RNN-T loss training and then replacing it with a strong LLM. After LM replacement, the minimum word error rate (MWER) loss is employed to finetune the integration of the LLM predictor with the Transducer-Llama model. Experiments on the LibriSpeech and large-scale multi-lingual LibriSpeech corpora show that the proposed streaming Transducer-Llama approach gave a 17% relative WER reduction (WERR) over a strong FT baseline and a 32% WERR over an RNN-T baseline.
Keqi Deng, Jinxi Guo, Yingyi Ma, Niko Moritz, Philip C. Woodland, Ozlem Kalinli, Mike Seltzer
ICASSP5
2025 CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
abstract
Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety benchmarks often focus solely on the refusal of individual problematic queries, which overlooks the importance of the context where the query occurs and may cause undesired refusal of queries under safe contexts that diminish user experience. Addressing this gap, we introduce CASE-Bench, a Context-Aware SafEty Benchmark that integrates context into safety assessments of LLMs. CASE-Bench assigns distinct, formally described contexts to categorized queries based on Contextual Integrity theory. Additionally, in contrast to previous studies which mainly rely on majority voting from just a few annotators, we recruited a sufficient number of annotators necessary to ensure the detection of statistically significant differences among the experimental conditions based on power analysis. Our extensive analysis using CASE-Bench on various open-source and commercial LLMs reveals a substantial and significant influence of context on human judgments ($p<$0.0001 from a z-test), underscoring the necessity of context in safety evaluations. We also identify notable mismatches between human judgments and LLM responses, particularly in commercial models within safe contexts. Code and data used in the paper are available at https://anonymous.4open.science/r/CASEBench-D5DB.
Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose M. Such
ICML4
2025 Wav2Prompt: End-to-End Speech Prompt Learning and Task-based Fine-tuning for Text-based LLMs
abstract
Keqi Deng, Guangzhi Sun, Phil Woodland. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Keqi Deng, Guangzhi Sun, Philip C. Woodland
NAACL (Long Papers)3
2025 Multi-head Temporal Latent Attention
abstract
While Transformer self-attention offers strong parallelism, the Key-Value (KV) cache grows linearly with sequence length and becomes a bottleneck for inference efficiency. Multi-head latent attention was recently developed to compress the KV cache into a low-rank latent space. This paper proposes Multi-head Temporal Latent Attention (MTLA), which further reduces the KV cache size along the temporal dimension, greatly lowering the memory footprint of self-attention inference. MTLA employs a hyper-network to dynamically merge temporally adjacent KV cache vectors. To address the mismatch between the compressed KV cache and processed sequence lengths, a stride-aware causal mask is proposed to ensure efficient parallel training and consistency with inference behaviour. Experiments across tasks, including speech translation, speech recognition, speech understanding and text summarisation, demonstrate that MTLA achieves competitive performance compared to standard Multi-Head Attention (MHA), while greatly improving inference speed and GPU memory usage. For example, on a English-German speech translation task, MTLA achieves a 5.3$\times$ speedup and a reduction in GPU memory usage by a factor of 8.3 compared to MHA, while maintaining translation quality.
Keqi Deng, Philip C. Woodland
NeurIPS2
2025 Knowledge-aware audio-grounded generative slot filling for limited annotated data
abstract
Manually annotating fine-grained slot-value labels for task-oriented dialogue (ToD) systems is an expensive and time-consuming endeavour. This motivates research into slot-filling methods that operate with limited amounts of labelled data. Moreover, the majority of current work on ToD is based solely on text as the input modality, neglecting the additional challenges of imperfect automatic speech recognition (ASR) when working with spoken language. In this work, we propose a Knowledge-Aware Audio-Grounded generative slot filling framework, termed KA2G, that focuses on few-shot and zero-shot slot filling for ToD with speech input. KA2G achieves robust and data-efficient slot filling for speech-based ToD by (1) framing it as a text generation task, (2) grounding text generation additionally in the audio modality, and (3) conditioning on available external knowledge ( e.g. a predefined list of possible slot values). We show that combining both modalities within the KA2G framework improves the robustness against ASR errors. Further, the knowledge-aware slot-value generator in KA2G, implemented via a pointer generator mechanism, particularly benefits few-shot and zero-shot learning. Experiments, conducted on the standard speech-based single-turn SLURP dataset and a multi-turn dataset extracted from a commercial ToD system , display strong and consistent gains over prior work, especially in few-shot and zero-shot setups.
Guangzhi Sun, Chao Zhang 0031, Ivan Vulic, Pawel Budzianowski, Philip C. Woodland
Comput. Speech Lang.5
2024 Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
abstract
While the neural transducer is popular for online speech recognition, simultaneous speech translation (SST) requires both streaming and re-ordering capabilities.This paper presents the LS-Transducer-SST, a label-synchronous neural transducer for SST, which naturally possesses these two properties.The LS-Transducer-SST dynamically decides when to emit translation tokens based on an Autoregressive Integrate-and-Fire (AIF) mechanism.A latency-controllable AIF is also proposed, which can control the quality-latency tradeoff either only during decoding, or it can be used in both decoding and training.The LS-Transducer-SST can naturally utilise monolingual text-only data via its prediction network which helps alleviate the key issue of data sparsity for E2E SST.During decoding, a chunkbased incremental joint decoding technique is designed to refine and expand the search space.Experiments on the Fisher-CallHome Spanish (Es-En) and MuST-C En-De data show that the LS-Transducer-SST gives a better qualitylatency trade-off than existing popular methods.For example, the LS-Transducer-SST gives a 3.1/2.9point BLEU increase (Es-En/En-De) relative to CAAT at a similar latency and a 1.4 s reduction in average lagging latency with similar BLEU scores relative to Wait-k.
Keqi Deng, Philip C. Woodland
ACL (1)2
2024 Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution Estimation
abstract
Wen Wu, Bo Li, Chao Zhang, Chung-Cheng Chiu, Qiujia Li, Junwen Bai, Tara Sainath, Phil Woodland. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wen Wu 0007, Bo Li 0028, Chao Zhang 0031, Chung-Cheng Chiu, Qiujia Li, Junwen Bai, Tara N. Sainath, Philip C. Woodland
ACL (1)8
2024 FastInject: Injecting Unpaired Text Data into CTC-Based ASR Training
abstract
Recently, connectionist temporal classification (CTC)-based end-to-end (E2E) automatic speech recognition (ASR) models have achieved impressive results, especially with the development of self-supervised learning. However, E2E ASR models trained on paired speech-text data often suffer from domain shifts from training to testing. To alleviate this issue, this paper proposes a flat-start joint training method, named FastInject, which efficiently injects multi-domain unpaired text data into CTC-based ASR training. To maintain training efficiency, text units are pre-upsampled, and their representations are fed into the CTC model along with speech features. To bridge the modality gap between speech and text, an attention-based modality matching mechanism (AM3) is proposed, which retains the E2E flat-start training. Experiments show that the proposed FastInject gave a 22% relative WER reduction (WERR) for intra-domain Librispeech-100h data and 20% relative WERR on out-of-domain test sets.
Keqi Deng, Philip C. Woodland
ICASSP2
2024 Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
abstract
Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-efficient finetuning (PEFT) for SER. Various PEFT adaptors are systematically studied for both classification of discrete emotion categories and prediction of dimensional emotional attributes. The results demonstrate that the combination of PEFT methods surpasses full finetuning with a significant reduction in the number of trainable parameters. Furthermore, a two-stage adaptation strategy is proposed to adapt models trained on acted emotion data, which is more readily available, to make the model more adept at capturing natural emotional expressions. Both intra- and cross-corpus experiments validate the efficacy of the proposed approach in enhancing the performance on both the source and target domains.
Nineli Lashkarashvili, Wen Wu 0007, Guangzhi Sun, Philip C. Woodland
ICASSP4
2024 Confidence Estimation for Automatic Detection of Depression and Alzheimer's Disease Based on Clinical Interviews
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH3
2024 SOT Triggered Neural Clustering for Speaker Attributed ASR
Xianrui Zheng, Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH4
2024 An Improved Empirical Fisher Approximation for Natural Gradient Descent
abstract
Approximate Natural Gradient Descent (NGD) methods are an important family of optimisers for deep learning models, which use approximate Fisher information matrices to pre-condition gradients during training. The empirical Fisher (EF) method approximates the Fisher information matrix empirically by reusing the per-sample gradients collected during back-propagation. Despite its ease of implementation, the EF approximation has its theoretical and practical limitations. This paper investigates the *inversely-scaled projection* issue of EF, which is shown to be a major cause of its poor empirical approximation quality. An improved empirical Fisher (iEF) method is proposed to address this issue, which is motivated as a generalised NGD method from a loss reduction perspective, meanwhile retaining the practical convenience of EF. The exact iEF and EF methods are experimentally evaluated using practical deep learning setups, including widely-used setups for parameter-efficient fine-tuning of pre-trained models (T5-base with LoRA and Prompt-Tuning on GLUE tasks, and ViT with LoRA for CIFAR100). Optimisation experiments show that applying exact iEF directly as an optimiser provides strong convergence and generalisation. It achieves the best test performance and the lowest training loss for the majority of the tasks, even when compared to well-tuned AdamW/Adafactor baselines. Additionally, under a novel empirical evaluation framework, the proposed iEF method shows consistently better approximation quality to exact Natural Gradient updates than both the EF and the more expensive sampled Fisher methods, meanwhile demonstrating the superior property of being robust to the choice of damping across tasks and training stages. Improving existing approximate NGD optimisers with iEF is expected to lead to better convergence and robustness. Furthermore, the iEF method also serves as a better approximation method to the Fisher information matrix itself, which enables the improvement of a variety of Fisher-based methods, not limited to the scope of optimisation.
Wenyi Yu, Chao Zhang 0031, Philip C. Woodland
NeurIPS4
2024 Automatic Time Alignment Generation For End-to-End ASR Using Acoustic Probability Modelling
abstract
End-to-end trainable (E2E) automatic speech recognition (ASR) models can achieve low error rates, but unlike hidden Markov model (HMM)-based systems they cannot naturally provide reliable timestamps. In this paper, an extra alignment module is proposed to enable pre-trained E2E ASR models to find accurate time-stamps by exploiting the connections between E2E and conventional HMMbased systems, which consists of a voice activity detection (VAD) and an acoustic scorer (AS). Full-sum training was adopted and the Viterbi algorithm used for alignment generation with HMM-like acoustic probabilities. Experiments using different E2E backbones and acoustic encoders with different output frame rates showed that for a jointly trained VAD and AS, the proposed methods can achieve high in-domain VAD performance and 99% alignment accuracy with a 200 ms collar for recognised words, with an average mismatch between 20 ms and 50 ms.
Dongcheng Jiang, Chao Zhang 0031, Philip C. Woodland
SLT3
2024 Decoupled structure for improved adaptability of end-to-end models
Keqi Deng, Philip C. Woodland
Speech Commun.2
2024 Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
abstract
Although end-to-end (E2E) automatic speech recognition (ASR) has shown state-of-the-art recognition accuracy, it tends to be implicitly biased towards the training data distribution which can degrade generalisation. This paper proposes a label-synchronous neural transducer (LS-Transducer), which provides a natural approach to domain adaptation based on text-only data. The LS-Transducer extracts a label-level encoder representation before combining it with the prediction network output. Since blank tokens are no longer needed, the prediction network performs as a standard language model, which can be easily adapted using text-only data. An Auto-regressive Integrate-and-Fire (AIF) mechanism is proposed to generate the label-level encoder representation while retaining low latency operation that can be used for streaming. In addition, a streaming joint decoding method is designed to improve ASR accuracy while retaining synchronisation with AIF. Experiments show that compared to standard neural transducers, the proposed LS-Transducer gave a 12.9% relative WER reduction (WERR) for intra-domain LibriSpeech data, as well as 21.4% and 24.6% relative WERRs on cross-domain TED-LIUM 2 and AESRC2020 data with an adapted prediction network.
Keqi Deng, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Graph Neural Networks for Contextual ASR With the Tree-Constrained Pointer Generator
abstract
Incorporating biasing words obtained through contextual knowledge is paramount in automatic speech recognition (ASR) applications. This paper proposes an innovative method for achieving end-to-end contextual ASR using graph neural network (GNN) encodings based on the tree-constrained pointer generator method. GNN node encodings facilitate lookahead for future word pieces in the process of ASR decoding at each tree node by incorporating information about all word pieces on the tree branches rooted from it. This results in a more precise prediction of the generation probability of the biasing words. The study explores three GNN encoding techniques: namely the tree recursive neural network (Tree-RNN), the graph convolutional network (GCN), and GraphSAGE, along with different combinations of the complementary GCN and GraphSAGE structures. The performance of the systems was evaluated using both Librispeech and the AMI corpus with a visual-grounded contextual ASR pipeline. The findings indicate that using GNN encodings achieved consistent and significant reductions in word error rate (WER), particularly for words that are rare or have not been seen during the training process. Notably, on LibriSpeech test sets, the combined GNN proposed in this paper achieved a 20% relative rare word error rate reduction compared to Tree-RNN, 30%-40% compared to standard TCPGen and 60% compared to standard ASR systems without TCPGen.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
abstract
In automatic emotion recognition (AER), labels assigned by different human annotators to the same utterance are often inconsistent due to the inherent complexity of emotion and the subjectivity of perception.Though deterministic labels generated by averaging or voting are often used as the ground truth, it ignores the intrinsic uncertainty revealed by the inconsistent labels.This paper proposes a Bayesian approach, deep evidential emotion regression (DEER), to estimate the uncertainty in emotion attributes.Treating the emotion attribute labels of an utterance as samples drawn from an unknown Gaussian distribution, DEER places an utterance-specific normal-inverse gamma prior over the Gaussian likelihood and predicts its hyper-parameters using a deep neural network model.It enables a joint estimation of emotion attributes along with the aleatoric and epistemic uncertainties.AER experiments on the widely used MSP-Podcast and IEMOCAP datasets showed DEER produced state-of-theart results for both the mean values and the distribution of emotion attributes 1 .
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland
ACL (1)3
2023 Adaptable End-to-End ASR Models Using Replaceable Internal LMs and Residual Softmax
abstract
End-to-end (E2E) automatic speech recognition (ASR) implicitly learns the token sequence distribution of paired audio-transcript training data. However, it still suffers from domain shifts from training to testing, and domain adaptation is still challenging. To alleviate this problem, this paper designs a replaceable internal language model (RILM) method, which makes it feasible to directly replace the internal language model (LM) of E2E ASR models with a target-domain LM in the decoding stage when a domain shift is encountered. Furthermore, this paper proposes a residual softmax (R-softmax) that is designed for CTC-based E2E ASR models to adapt to the target domain without re-training during inference. For E2E ASR models trained on the LibriSpeech corpus, experiments showed that the proposed methods gave a 2.6% absolute WER reduction on the Switchboard data and a 1.0% WER reduction on the AESRC2020 corpus while maintaining intra-domain ASR results.
Keqi Deng, Philip C. Woodland
ICASSP2
2023 Spectral Clustering-Aware Learning of Embeddings for Speaker Diarisation
abstract
In speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning of embeddings (SCALE) to address the mismatch. Specifically, besides an angular prototypical (AP) loss, SCALE uses a novel affinity matrix loss which directly minimises the error between the affinity matrix estimated from speaker embeddings and the reference. SCALE also includes p-percentile thresholding and Gaussian blur as two important hyper-parameters for spectral clustering in training. Experiments on the AMI dataset showed that speaker embeddings obtained with SCALE achieved over 50% relative speaker error rate reductions using oracle segmentation, and over 30% relative diarisation error rate reductions using automatic segmentation when compared to a strong baseline with the AP-loss-based speaker embeddings.
Evonne P. C. Lee, Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
ICASSP4
2023 Self-Supervised Learning-Based Source Separation for Meeting Data
abstract
Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel source separation, most studies have focused on simulated setups. In this paper, seven SSL models were compared on both simulated and real-world corpora. Then, we propose to integrate the best-performing model WavLM into an automatic transcription system through a novel iterative source selection method. To improve real-world performance, time-domain unsupervised mixture invariant training was adapted to the time-frequency domain. Experiments showed that in the transcription system when source separation was inserted before an ASR model fine-tuned on separated speech, absolute reductions of 1.9% and 1.5% in concatenated minimum-permutation word error rate for an unknown number of speakers (cpWER-us) were observed on the AMI dev and test sets.
Yuang Li, Xianrui Zheng, Philip C. Woodland
ICASSP3
2023 End-to-End Spoken Language Understanding with Tree-Constrained Pointer Generator
abstract
End-to-end spoken language understanding (SLU) suffers from the long-tail word problem. This paper exploits contextual biasing, a technique to improve the speech recognition of rare words, in end-to-end SLU systems. Specifically, a tree-constrained pointer generator (TCPGen), a powerful and efficient biasing model component, is studied, which leverages a slot shortlist with corresponding entities to extract biasing lists. Meanwhile, to bias the SLU model output slot distribution, a slot probability biasing (SPB) mechanism is proposed to calculate a slot distribution from TCPGen. Experiments on the SLURP dataset showed consistent SLU-F1 improvements using TCPGen and SPB, especially on unseen entities. On a new split by holding out 5 slot types for the test, TCPGen with SPB achieved zero-shot learning with an SLU-F1 score over 50% compared to baselines which can not deal with it. In addition to slot filling, the intent classification accuracy was also improved.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2023 Self-Supervised Representations in Speech-Based Depression Detection
abstract
This paper proposes handling training data sparsity in speech-based automatic depression detection (SDD) using foundation models pre-trained with self-supervised learning (SSL). An analysis of SSL representations derived from different layers of pre-trained foundation models is first presented for SDD, which provides insight to suitable indicator for depression detection. Knowledge transfer is then performed from automatic speech recognition (ASR) and emotion recognition to SDD by fine-tuning the foundation models. Results show that the uses of oracle and ASR transcriptions yield similar SDD performance when the hidden representations of the ASR model is incorporated along with the ASR textual information. By integrating representations from multiple foundation models, state-of-the-art SDD results based on real ASR were achieved on the DAIC-WOZ dataset.
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2023 A Neural Time Alignment Module for End-to-End Automatic Speech Recognition
Dongcheng Jiang, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH3
2023 Biased Self-supervised Learning for ASR
Florian Kreyssig, Yangyang Shi, Jinxi Guo, Leda Sari, Abdel-rahman Mohamed, Philip C. Woodland
INTERSPEECH6
2023 Can Contextual Biasing Remain Effective with Whisper and GPT-2?
Guangzhi Sun, Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH4
2023 Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
abstract
Although automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating AER with automatic speech recognition (ASR) and speaker diarisation (SD) in a jointly-trained system. Distinct output layers are built for four sub-tasks including AER, ASR, voice activity detection and speaker classification based on a shared encoder. Taking the audio of a conversation as input, the integrated system finds all speech segments and transcribes the corresponding emotion classes, word sequences, and speaker identities. Two metrics are proposed to evaluate AER performance with automatic segmentation based on time-weighted emotion and speaker classification errors. Results on the IEMOCAP dataset show that the proposed system consistently outperforms two baselines with separately trained single-task systems on AER, ASR and SD.
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH3
2023 Combining hybrid DNN-HMM ASR systems with attention-based models using lattice rescoring
abstract
The traditional hybrid deep neural network (DNN)–hidden Markov model (HMM) system and attention-based encoder–decoder (AED) model are both commonly used automatic speech recognition (ASR) approaches with distinct characteristics and advantages. While hybrid systems are per-frame-based and highly modularised to leverage external phonetic and linguistic knowledge, AED models operate on a per-label basis and jointly learn the acoustic and language information using a single model in an end-to-end trainable fashion. In this paper, we propose combining these two approaches in a two-pass rescoring framework. The first-pass uses hybrid ASR systems to facilitate streaming and controllable ASR, and the second-pass re-scores the N-best hypotheses or lattices produced by the first-pass hybrid DNN-HMM system with AED models. We also propose an improved algorithm for lattice rescoring with AED models. Experiments show the combined two-pass systems achieve competitive performance without using extra speech or text data on two standard ASR tasks. For the 80-hour AMI IHM dataset, the combined system has a 13.7% word error rate (WER) on the evaluation set and is up to a 29% relative WER reduction over the individual systems. For the 300-hour Switchboard dataset, the WERs of the combined system are 5.7% and 12.1% on Switchboard and CallHome subsets of Hub5’00, and 13.2% and 7.6% on Switchboard Cellular and Fisher subsets of RT03, and are up to a 33% relative reduction in WER over the individual systems.
Qiujia Li, Chao Zhang 0031, Philip C. Woodland
Speech Commun.3
2023 Estimating the Uncertainty in Emotion Class Labels With Utterance-Specific Dirichlet Priors
abstract
Emotion recognition is a key attribute for artificial intelligence systems that need to naturally interact with humans. However, the task definition is still an open problem due to the inherent ambiguity of emotions. In this paper, a novel Bayesian training loss based on per-utterance Dirichlet prior distributions is proposed for verbal emotion recognition, which models the uncertainty in one-hot labels created when human annotators assign the same utterance to different emotion classes. An additional metric is used to evaluate the performance by detecting test utterances with high labelling uncertainty. This removes a major limitation that emotion classification systems only consider utterances with labels where the majority of annotators agree on the emotion class. Furthermore, a frequentist approach is studied to leverage the continuous-valued “soft” labels obtained by averaging the one-hot labels. We propose a two-branch model structure for emotion classification on a per-utterance basis, which achieves state-of-the-art classification results on the widely used IEMOCAP dataset. Based on this, uncertainty estimation experiments were performed. The best performance in terms of the area under the precision-recall curve when detecting utterances with high uncertainty was achieved by interpolating the Bayesian training loss with the Kullback-Leibler divergence training loss for the soft labels. The generality of the proposed approach was verified using the MSP-Podcast dataset which yielded the same pattern of results.
Wen Wu 0007, Chao Zhang 0031, Xixin Wu, Philip C. Woodland
IEEE Trans. Affect. Comput.4
2023 Minimising Biasing Word Errors for Contextual ASR With the Tree-Constrained Pointer Generator
abstract
Contextual knowledge is essential for reducing speech recognition errors on high-valued long-tail words. This paper proposes a novel tree-constrained pointer generator (TCPGen) component that enables end-to-end ASR models to bias towards a list of long-tail words obtained using external contextual information. With only a small overhead in memory use and computation cost, TCPGen can structure thousands of biasing words efficiently into a symbolic prefix-tree, and creates a neural shortcut between the tree and the final ASR output to facilitate the recognition of the biasing words. To enhance TCPGen, we further propose a novel minimum biasing word error (MBWE) loss that directly optimises biasing word errors during training, along with a biasing-word-driven language model discounting (BLMD) method during the test. All contextual ASR systems were evaluated on the public Librispeech audiobook corpus and the data from the dialogue state tracking challenges (DSTC) with the biasing lists extracted from the dialogue-system ontology. Consistent word error rate (WER) reductions were achieved with TCPGen, which were particularly significant on the biasing words with around 40% relative reductions in the recognition error rates. MBWE and BLMD further improved the effectiveness of TCPGen, and achieved more significant WER reductions on the biasing words. TCPGen also achieved zero-shot learning of words not in the audio training set with large WER reductions on the out-of-vocabulary words in the biasing list.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition
abstract
As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research has shown that model-based confidence estimators have a significant advantage over using the output softmax probabilities. If the input data to the speech recogniser is from mismatched acoustic and linguistic conditions, the ASR performance and the corresponding confidence estimators may exhibit severe degradation. Since confidence models are often trained on the same in-domain data as the ASR, generalising to out-of-domain (OOD) scenarios is challenging. By keeping the ASR model untouched, this paper proposes two approaches to improve the model-based confidence estimators on OOD data: using pseudo transcriptions and an additional OOD language model. With an ASR model trained on LibriSpeech, experiments show that the proposed methods can greatly improve the confidence metrics on TED-LIUM and Switchboard datasets while preserving in-domain performance. Furthermore, the improved confidence estimators are better calibrated on OOD data and can provide a much more reliable criterion for data selection.
Qiujia Li, Yu Zhang 0033, David Qiu, Yanzhang He, Liangliang Cao, Philip C. Woodland
ICASSP6
2022 Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-Trained Models
abstract
Self-supervised pre-training is an effective approach to leveraging a large amount of unlabelled data to reduce word error rates (WERs) of automatic speech recognition (ASR) systems. Since it is impractical to use large pre-trained models for many real-world ASR applications, it is desirable to have a much smaller model while retaining the performance of the pre-trained model. In this paper, we propose a simple knowledge distillation (KD) loss function for neural transducers that focuses on the one-best path in the output probability lattice under both streaming and non-streaming setups, which allows a small student model to approach the performance of the large pre-trained teacher model. Experiments on the LibriSpeech dataset show that despite being 10 times smaller than the teacher model, the proposed loss results in relative WER reductions (WERRs) of 11.5% and 6.8% on the test-other set for non-streaming and streaming student models compared to the baseline transducers trained without KD using the labelled 100-hour clean data. With an additional 860 hours of unlabelled data for KD, the WERRs increase to 48.2% and 38.5% for non-streaming and streaming students. If language model shallow fusion is used for producing distillation targets, a further improvement in the student model is observed.
Qiujia Li, Philip C. Woodland
ICASSP3
2022 Tree-constrained Pointer Generator with Graph Neural Network Encodings for Contextual Speech Recognition
abstract
Incorporating biasing words obtained as contextual knowledge is critical for many automatic speech recognition (ASR) applications.This paper proposes the use of graph neural network (GNN) encodings in a tree-constrained pointer generator (TCP-Gen) component for end-to-end contextual ASR.By encoding the biasing words in the prefix-tree with a tree-based GNN, lookahead for future wordpieces in end-to-end ASR decoding is achieved at each tree node by incorporating information about all wordpieces on the tree branches rooted from it, which allows a more accurate prediction of the generation probability of the biasing words.Systems were evaluated on the Librispeech corpus using simulated biasing tasks, and on the AMI corpus by proposing a novel visual-grounded contextual ASR pipeline that extracts biasing words from slides alongside each meeting.Results showed that TCPGen with GNN encodings achieved about a further 15% relative WER reduction on the biasing words compared to the original TCPGen, with a negligible increase in the computation cost for decoding.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH3
2022 Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription
abstract
Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a tandem multitask training (TMT) method is proposed to fine-tune W2V2. For speaker diarisation, the tasks of voice activity detection (VAD) and speaker classification (SC) are required, and connectionist temporal classification (CTC) is used for ASR. The multitask framework implements VAD, SC, and ASR using an early layer, middle layer, and late layer of W2V2, which coincides with the order of segmenting the audio with VAD, clustering the segments based on speaker embeddings, and transcribing each segment with ASR. Experimental results on the augmented multi-party (AMI) dataset showed that using different W2V2 layers for VAD, SC, and ASR from the earlier to later layers for TMT not only saves computational cost, but also reduces diarisation error rates (DERs). Joint fine-tuning of VAD, SC, and ASR yielded 16%/17% relative reductions of DER with manual/automatic segmentation respectively, and consistent reductions in speaker attributed word error rate, compared to the baseline with separately fine-tuned models.
Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH3
2022 Distribution-Based Emotion Recognition in Conversation
abstract
Automatic emotion recognition in conversation (ERC) is crucial for emotion-aware conversational artificial intelligence. This paper proposes a distribution-based framework that formulates ERC as a sequence-to-sequence problem for emotion distribution estimation. The inherent ambiguity of emotions and the subjectivity of human perception lead to disagreements in emotion labels, which is handled naturally in our framework from the perspective of uncertainty estimation in emotion distributions. A Bayesian training loss is introduced to improve the uncertainty estimation by conditioning each emotional state on an utterance-specific Dirichlet prior distribution. Experimental results on the IEMOCAP dataset show that ERC outperformed the single-utterance-based system, and the proposed distribution-based ERC methods have not only better classification accuracy, but also show improved uncertainty estimation.
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland
SLT3
2021 Tree-Constrained Pointer Generator for End-to-End Contextual Speech Recognition
abstract
Contextual knowledge is important for real-world automatic speech recognition (ASR) applications. In this paper, a novel tree-constrained pointer generator (TCPGen) component is proposed that incorpo-rates such knowledge as a list of biasing words into both attention-based encoder-decoder and transducer end-to-end ASR models in a neural-symbolic way. TCPGen structures the biasing words into an efficient prefix tree to serve as its symbolic input and creates a neu-ral shortcut between the tree and the final ASR output distribution to facilitate recognising biasing words during decoding. Systems were trained and evaluated on the Librispeech corpus where biasing words were extracted at the scales of an utterance, a chapter, or a book to simulate different application scenarios. Experimental results showed that TCPGen consistently improved word error rates (WERs) compared to the baselines, and in particular, achieved sig-nificant WER reductions on the biasing words. TCPGen is highly efficient: it can handle 5,000 biasing words and distractors and only add a small overhead to memory use and computation cost.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
ASRU3
2021 Adapting GPT, GPT-2 and BERT Language Models for Speech Recognition
abstract
Language models (LMs) pre-trained on massive amounts of text, in particular bidirectional encoder representations from Transformers (BERT), generative pre-training (GPT), and GPT-2, have become a key technology for many natural language processing tasks. In this paper, we present results using fine-tuned GPT, GPT-2, and their combination for automatic speech recognition (ASR). Unlike unidirectional LM GPT and GPT-2, BERT is bidirectional whose direct product of the output probabilities is no longer a valid language prior probability. A conversion method is proposed to compute the correct language prior probability based on bidirectional LM outputs in a mathematically exact way. Experimental results on the widely used AMI and Switchboard ASR tasks showed that the combination of the fine-tuned GPT and GPT-2 outperformed the combination of three neural LMs with different architectures trained from scratch on the in-domain text by up to a 12% relative word error rate reduction (WERR). Furthermore, on the AMI corpus, the proposed conversion for language prior probabilities enables BERT to obtain an extra 3% relative WERR, and the combination of BERT, GPT and GPT-2 results in further improvements.
Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland
ASRU3
2021 Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition
abstract
For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based automatic speech recognition (ASR) systems, confidence scores can be reliably obtained from word posteriors in decoding lattices. However, for an ASR system with an auto-regressive decoder, such as an attention-based sequence-to-sequence model, computing word posteriors is difficult. An obvious alternative is to use the decoder softmax probability as the model confidence. In this paper, we first examine how some commonly used regularisation methods influence the softmax-based confidence scores and study the overconfident behaviour of end-to-end models. Then we propose a lightweight and effective approach named confidence estimation module (CEM) on top of an existing end-to-end ASR model. Experiments on LibriSpeech show that CEM can mitigate the overconfidence problem and can produce more reliable confidence scores with and without shallow fusion of a language model. Further analysis shows that CEM generalises well to speech from a moderately mismatched domain and can potentially improve downstream tasks such as semi-supervised learning.
Qiujia Li, David Qiu, Yu Zhang 0033, Bo Li 0028, Yanzhang He, Philip C. Woodland, Liangliang Cao, Trevor Strohman
ICASSP6
2021 Transformer Language Models with LSTM-Based Cross-Utterance Information Representation
abstract
The effective incorporation of cross-utterance information has the potential to improve language models (LMs) for automatic speech recognition (ASR). To extract more powerful and robust cross-utterance representations for the Transformer LM (TLM), this paper proposes the R-TLM which uses hidden states in a long short-term memory (LSTM) LM. To encode the cross-utterance information, the R-TLM incorporates an LSTM module together with a segment-wise recurrence in some of the Transformer blocks. In addition to the LSTM module output, a shortcut connection using a fusion layer which bypasses the LSTM module is also investigated. The proposed system was evaluated on the AMI meeting corpus, the Eval2000 and the RT03 telephone conversation evaluation sets. The best R-TLM achieved 0.9%, 0.6% and 0.8% absolute WER reductions over the single-utterance TLM baseline, and 0.5%, 0.3%, 0.2% absolute WER reductions over a strong cross-utterance TLM baseline on the AMI evaluation set, Eval2000 and RT03 respectively. Improvements on Eval2000 and RT03 were further supported by significance tests. R-TLMs were found to have better LM scores on words where recognition errors are more likely to occur. The R-TLM WER can be further reduced by interpolation with an LSTM-LM.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2021 Content-Aware Speaker Embeddings for Speaker Diarisation
abstract
Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware speaker embeddings (CASE) approach is proposed, which extends the input of the speaker classifier to include not only acoustic features but also their corresponding speech content, via phone, character, and word embeddings. Compared to alternative methods that leverage similar information, such as multitask or adversarial training, CASE factorises automatic speech recognition (ASR) from speaker recognition to focus on modelling speaker characteristics and correlations with the corresponding content units to derive more expressive representations. CASE is evaluated for speaker re-clustering with a realistic speaker diarisation setup using the AMI meeting transcription dataset, where the content information is obtained by performing ASR based on an automatic segmentation. Experimental results showed that CASE achieved a 17.8% relative speaker error rate reduction over conventional methods.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
ICASSP4
2021 Emotion Recognition by Fusing Time Synchronous and Time Asynchronous Representations
abstract
In this paper, a novel two-branch neural network model structure is proposed for multimodal emotion recognition, which consists of a time synchronous branch (TSB) and a time asynchronous branch (TAB). To capture correlations between each word and its acoustic realisation, the TSB combines speech and text modalities at each input window frame and then uses pooling across time to form a single embedding vector. The TAB, by contrast, provides cross-utterance information by integrating sentence text embeddings from a number of context utterances into another embedding vector. The final emotion classification uses both the TSB and the TAB embeddings. Experimental results on the IEMOCAP dataset demonstrate that the two-branch structure achieves state-of-the-art results in 4-way classification with all common test setups. When using automatic speech recognition (ASR) output instead of manually transcribed reference text, it is shown that the cross-utterance information considerably improves robustness against ASR errors. Furthermore, by incorporating an extra class for all the other emotions, the final 5-way classification system with ASR hypotheses can be viewed as a prototype for more realistic emotion recognition systems.
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2021 Variable Frame Rate Acoustic Models Using Minimum Error Reinforcement Learning
Dongcheng Jiang, Chao Zhang 0031, Philip C. Woodland
Interspeech3
2021 Residual Energy-Based Models for End-to-End Speech Recognition
abstract
End-to-end models with auto-regressive decoders have shown impressive results for automatic speech recognition (ASR). These models formulate the sequence-level probability as a product of the conditional probabilities of all individual tokens given their histories. However, the performance of locally normalised models can be sub-optimal because of factors such as exposure bias. Consequently, the model distribution differs from the underlying data distribution. In this paper, the residual energy-based model (R-EBM) is proposed to complement the auto-regressive ASR model to close the gap between the two distributions. Meanwhile, R-EBMs can also be regarded as utterance-level confidence estimators, which may benefit many downstream tasks. Experiments on a 100hr LibriSpeech dataset show that R-EBMs can reduce the word error rates (WERs) by 8.2%/6.7% while improving areas under precision-recall curves of confidence scores by 12.6%/28.4% on test-clean/test-other sets. Furthermore, on a state-of-the-art model using self-supervised learning (wav2vec 2.0), R-EBMs still significantly improves both the WER and confidence estimation performance.
Qiujia Li, Yu Zhang 0033, Bo Li 0028, Liangliang Cao, Philip C. Woodland
Interspeech5
2021 Discriminative Neural Clustering for Speaker Diarisation
abstract
In this paper, we propose Discriminative Neural Clustering (DNC) that formulates data clustering with a maximum number of clusters as a supervised sequence-to-sequence learning problem. Com-pared to traditional unsupervised clustering algorithms, DNC learns clustering patterns from training data without requiring an explicit definition of a similarity measure. An implementation of DNC based on the Transformer architecture is shown to be effective on a speaker diarisation task using the challenging AMI dataset. Since AMI contains only 147 complete meetings as individual input sequences, data scarcity is a significant issue for training a Transformer model for DNC. Accordingly, this paper proposes three data augmentation schemes: sub-sequence randomisation, input vector randomisation, and Diaconis augmentation, which generates new data samples by rotating the entire input sequence of L2-normalised speaker embeddings. Experimental results on AMI show that DNC achieves a reduction in speaker error rate (SER) of 29.4% relative to spectral clustering.
Qiujia Li, Florian Kreyssig, Chao Zhang 0031, Philip C. Woodland
SLT4
2021 A distributed optimisation framework combining natural gradient with Hessian-free for discriminative sequence training
Adnan Haider, Chao Zhang 0031, Florian Kreyssig, Philip C. Woodland
Neural Networks4
2021 Combination of deep speaker embeddings for diarisation
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
Neural Networks3
2020 Improved Large-Margin Softmax Loss for Speaker Diarisation
abstract
Speaker diarisation systems nowadays use embeddings generated from speech segments in a bottleneck layer, which are needed to be discriminative for unseen speakers. It is well-known that large-margin training can improve the generalisation ability to unseen data, and its use in such open-set problems has been widespread. Therefore, this paper introduces a general approach to the large-margin softmax loss without any approximations to improve the quality of speaker embeddings for diarisation. Furthermore, a novel and simple way to stabilise training, when large-margin softmax is used, is proposed. Finally, to combat the effect of overlapping speech, different training margins are used to reduce the negative effect overlapping speech has on creating discriminative embeddings. Experiments on the AMI meeting corpus show that the use of large-margin softmax significantly improves the speaker error rate (SER). By using all hyper parameters of the loss in a unified way, further improvements were achieved which reached a relative SER reduction of 24.6% over the baseline. However, by training overlapping and single speaker speech samples with different margins, the best result was achieved, giving overall a 29.5% SER reduction relative to the baseline.
Yassir Fathullah, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2020 Cosine-Distance Virtual Adversarial Training for Semi-Supervised Speaker-Discriminative Acoustic Embeddings
abstract
In this paper, we propose a semi-supervised learning (SSL) technique for training deep neural networks (DNNs) to generate speaker-discriminative acoustic embeddings (speaker embeddings). Obtaining large amounts of speaker recognition train-ing data can be difficult for desired target domains, especially under privacy constraints. The proposed technique reduces requirements for labelled data by leveraging unlabelled data. The technique is a variant of virtual adversarial training (VAT) [1] in the form of a loss that is defined as the robustness of the speaker embedding against input perturbations, as measured by the cosine-distance. Thus, we term the technique cosine-distance virtual adversarial training (CD-VAT). In comparison to many existing SSL techniques, the unlabelled data does not have to come from the same set of classes (here speakers) as the labelled data. The effectiveness of CD-VAT is shown on the 2750+ hour VoxCeleb data set, where on a speaker verification task it achieves a reduction in equal error rate (EER) of 11.1% relative to a purely supervised baseline. This is 32.5% of the improvement that would be achieved from supervised training if the speaker labels for the unlabelled data were available.
Florian Kreyssig, Philip C. Woodland
INTERSPEECH2
2019 Integrating Source-Channel and Attention-Based Sequence-to-Sequence Models for Speech Recognition
abstract
This paper proposes a novel automatic speech recognition (ASR) framework called Integrated Source-Channel and Attention (ISCA) that combines the advantages of traditional systems based on the noisy source-channel model (SC) and end-to-end style systems using attention-based sequence-to-sequence models. The traditional SC system framework includes hidden Markov models and connectionist temporal classification (CTC) based acoustic models, language models (LMs), and a decoding procedure based on a lexicon, whereas the end-to-end style attention-based system jointly models the whole process with a single model. By rescoring the hypotheses produced by traditional systems using end-to-end style systems based on an extended noisy source-channel model, ISCA allows structured knowledge to be easily incorporated via the SC-based model while exploiting the complementarity of the attention-based model. Experiments on the AMI meeting corpus show that ISCA is able to give a relative word error rate reduction up to 21% over an individual system, and by 13% over an alternative method which also involves combining CTC and attention-based models.
Qiujia Li, Chao Zhang 0031, Philip C. Woodland
ASRU3
2019 Speaker Diarisation Using 2D Self-attentive Combination of Embeddings
abstract
Speaker diarisation systems often cluster audio segments using speaker embeddings such as i-vectors and d-vectors. Since different types of embeddings are often complementary, this paper proposes a generic framework to improve performance by combining them into a single embedding, referred to as a c-vector. This combination uses a 2-dimensional (2D) self-attentive structure, which extends the standard self-attentive layer by averaging not only across time but also across different types of embeddings. Two types of 2D self-attentive structure studied in this paper are simultaneous combination and consecutive combination, which adopt single and multiple self-attentive layers respectively. The penalty term in the original self-attentive layer, which is jointly minimised with the objective function to encourage diversity of annotation vectors, is also modified to obtain not only different local peaks but also the overall trends in the multiple annotation vectors. Experiments on the AMI meeting corpus show that our modified penalty term improves the d-vector relative speaker error rate (SER) by 6% and 21% for d-vector systems, and a 10% further relative SER reduction can be obtained using the c-vector from our best 2D self-attentive structure.
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2019 PyHTK: Python Library and ASR Pipelines for HTK
abstract
This paper describes PyHTK, which is a Python-based library and associated pipeline to facilitate the construction of large-scale complex automatic speech recognition (ASR) systems using the hidden Markov model toolkit (HTK). PyHTK can be used to generate sophisticated artificial neural network (ANN) models with versatile architectures by converting a compact configuration file defining the ANN, into the form used by HTK tools, as well as supporting a range of capabilities to train and test ANN models. The ASR pipeline is divided into multiple steps, which can be arranged and customised for different ASR data sets, and allows for both step-by-step and fully automatic end-to-end operation. PyHTK is integrated with HTK 3.5.1 which includes an expanded range of ANN layer types and very flexible ways to connect them, together with capabilities for ASR training and testing. Some example systems are included to illustrate the flexibility and performance achievable.
Chao Zhang 0031, Florian Kreyssig, Qiujia Li, Philip C. Woodland
ICASSP4
2019 Multi-Span Acoustic Modelling Using Raw Waveform Signals
abstract
Traditional automatic speech recognition (ASR) systems often use an acoustic model (AM) built on handcrafted acoustic features, such as log Mel-filter bank (FBANK) values. Recent studies found that AMs with convolutional neural networks (CNNs) can directly use the raw waveform signal as input. Given sufficient training data, these AMs can yield a competitive word error rate (WER) to those built on FBANK features. This paper proposes a novel multi-span structure for acoustic modelling based on the raw waveform with multiple streams of CNN input layers, each processing a different span of the raw waveform signal. Evaluation on both the single channel CHiME4 and AMI data sets show that multi-span AMs give a lower WER than FBANK AMs by an average of about 5% (relative). Analysis of the trained multi-span model reveals that the CNNs can learn filters that are rather different to the log Mel filters. Furthermore, the paper shows that a widely used single span raw waveform AM can be improved by using a smaller CNN kernel size and increased stride to yield improved WERs.
Patrick von Platen, Chao Zhang 0031, Philip C. Woodland
INTERSPEECH3
2018 Improved Tdnns Using Deep Kernels and Frequency Dependent Grid-RNNS
abstract
Time delay neural networks (TDNNs) are an effective acoustic model for large vocabulary speech recognition. The strength of the model can be attributed to its ability to effectively model long temporal contexts. However, current TDNN models are relatively shallow, which limits the modelling capability. This paper proposes a method of increasing the network depth by deepening the kernel used in the TDNN temporal convolutions. The best performing kernel consists of three fully connected layers with a residual (ResNet) connection from the output of the first to the output of the third. The addition of spectro-temporal processing as the input to the TDNN in the form of a convolutional neural network (CNN) and a newly designed Grid-RNN was investigated. The Grid-RNN strongly outperforms a CNN if different sets of parameters for different frequency bands are used and can be further enhanced by using a bi-directional Grid-RNN. Experiments using the multi-genre broadcast (MGB3) English data (275h) show that deep kernel TDNNs reduces the word error rate (WER) by 6% relative and when combined with the frequency dependent Grid-Rnn gives a relative WER reduction of 9%.
Florian Kreyssig, Chao Zhang 0031, Philip C. Woodland
ICASSP3
2018 High Order Recurrent Neural Networks for Acoustic Modelling
abstract
Vanishing long-term gradients are a major issue in training standard recurrent neural networks (RNNs), which can be alleviated by long short-term memory (LSTM) models with memory cells. However, the extra parameters associated with the memory cells mean an LSTM layer has four times as many parameters as an RNN with the same hidden vector size. This paper addresses the vanishing gradient problem using a high order RNN (HORNN) which has additional connections from multiple previous time steps. Speech recognition experiments using British English multi-genre broadcast (MGB3) data showed that the proposed HORNN architectures for rectified linear unit and sigmoid activation functions reduced word error rates (WER) by 4.2% and 6.3% over the corresponding RNNs, and gave similar WERs to a (projected) LSTM while using only 20%-50% of the recurrent layer parameters and computation.
Chao Zhang 0031, Philip C. Woodland
ICASSP2
2018 Combining Natural Gradient with Hessian Free Methods for Sequence Training
abstract
This paper presents a new optimisation approach to train Deep Neural Networks (DNNs) with discriminative sequence criteria.At each iteration, the method combines information from the Natural Gradient (NG) direction with local curvature information of the error surface that enables better paths on the parameter manifold to be traversed.The method has been applied within a Hessian Free (HF) style optimisation framework to sequence train both standard fully-connected DNNs and Time Delay Neural Networks as speech recognition acoustic models.The efficacy of the method is shown using experiments on a Multi-Genre Broadcast (MGB) transcription task and neural networks using sigmoid and ReLU activation functions have been investigated.It is shown that for the same number of updates this proposed approach achieves larger reductions in the word error rate (WER) than both NG and HF, and also leads to a lower WER than standard stochastic gradient descent.
Adnan Haider, Philip C. Woodland
INTERSPEECH2
2018 Speaker Adaptation and Adaptive Training for Jointly Optimised Tandem Systems
abstract
Speaker independent (SI) Tandem systems trained by joint optimisation of bottleneck (BN) deep neural networks (DNNs) and Gaussian mixture models (GMMs) have been found to produce similar word error rates (WERs) to Hybrid DNN systems. A key advantage of using GMMs is that existing speaker adaptation methods, such as maximum likelihood linear regression (MLLR), can be used which to account for diverse speaker variations and improve system robustness. This paper investigates speaker adaptation and adaptive training (SAT) schemes for jointly optimised Tandem systems. Adaptation techniques investigated include constrained MLLR (CMLLR) transforms based on BN features for SAT as well as MLLR and parameterised sigmoid functions for unsupervised test-time adaptation. Experiments using English multi-genre broadcast (MGB3) data show that CMLLR SAT yields a 4% relative WER reduction over jointly trained Tandem and Hybrid SI systems, and further reductions in WER are obtained by system combination.
Yu Wang 0027, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH4
2018 Semi-tied Units for Efficient Gating in LSTM and Highway Networks
abstract
Gating is a key technique used for integrating information from multiple sources by long short-term memory (LSTM) models and has recently also been applied to other models such as the highway network.Although gating is powerful, it is rather expensive in terms of both computation and storage as each gating unit uses a separate full weight matrix.This issue can be severe since several gates can be used together in e.g. an LSTM cell.This paper proposes a semi-tied unit (STU) approach to solve this efficiency issue, which uses one shared weight matrix to replace those in all the units in the same layer.The approach is termed "semi-tied" since extra parameters are used to separately scale each of the shared output values.These extra scaling factors are associated with the network activation functions and result in the use of parameterised sigmoid, hyperbolic tangent, and rectified linear unit functions.Speech recognition experiments using British English multi-genre broadcast data showed that using STUs can reduce the calculation and storage cost by a factor of three for highway networks and four for LSTMs, while giving similar word error rates to the original models.
Chao Zhang 0031, Philip C. Woodland
INTERSPEECH2
2017 Sequence training of DNN acoustic models with natural gradient
abstract
Deep Neural Network (DNN) acoustic models often use discriminative sequence training that optimises an objective function that better approximates the word error rate (WER) than frame-based training. Sequence training is normally implemented using Stochastic Gradient Descent (SGD) or Hessian Free (HF) training. This paper proposes an alternative batch style optimisation framework that employs a Natural Gradient (NG) approach to traverse through the parameter space. By correcting the gradient according to the local curvature of the KL-divergence, the NG optimisation process converges more quickly than HF. Furthermore, the proposed NG approach can be applied to any sequence discriminative training criterion. The efficacy of the NG method is shown using experiments on a Multi-Genre Broadcast (MGB) transcription task that demonstrates both the computational efficiency and the accuracy of the resulting DNN models.
Adnan Haider, Philip C. Woodland
ASRU2
2017 Joint optimisation of tandem systems using Gaussian mixture density neural network discriminative sequence training
abstract
The use of deep neural networks (DNNs) for feature extraction and Gaussian mixture models (GMMs) for acoustic modelling is often termed a tandem system configuration and can be viewed as a Gaussian mixture density neural network (MDNN). Compared to the direct use of DNN output probabilities in the acoustic model, the tandem approach suffers from a major weakness in that the feature extraction stage and the final acoustic models are optimised separately. This paper proposes a joint optimisation approach to all the stages of the tandem acoustic model by using MDNN discriminative sequence training. A set of techniques is used to improve the training performance and stability. Experiments using the multi-genre broadcast (MGB) English data show that the proposed method produced a 6% relative lower word error rate (WER) than that of a traditional discriminatively trained tandem system. The resulting jointly optimised tandem systems are comparable in WER to hybrid DNN systems optimised using discriminative sequence training with the same number of parameters.
Chao Zhang 0031, Philip C. Woodland
ICASSP2
2017 Relating dynamic brain states to dynamic machine states: Human and machine solutions to the speech recognition problem
abstract
There is widespread interest in the relationship between the neurobiological systems supporting human cognition and emerging computational systems capable of emulating these capacities. Human speech comprehension, poorly understood as a neurobiological process, is an important case in point. Automatic Speech Recognition (ASR) systems with near-human levels of performance are now available, which provide a computationally explicit solution for the recognition of words in continuous speech. This research aims to bridge the gap between speech recognition processes in humans and machines, using novel multivariate techniques to compare incremental 'machine states', generated as the ASR analysis progresses over time, to the incremental 'brain states', measured using combined electro- and magneto-encephalography (EMEG), generated as the same inputs are heard by human listeners. This direct comparison of dynamic human and machine internal states, as they respond to the same incrementally delivered sensory input, revealed a significant correspondence between neural response patterns in human superior temporal cortex and the structural properties of ASR-derived phonetic models. Spatially coherent patches in human temporal cortex responded selectively to individual phonetic features defined on the basis of machine-extracted regularities in the speech to lexicon mapping process. These results demonstrate the feasibility of relating human and ASR solutions to the problem of speech recognition, and suggest the potential for further studies relating complex neural computations in human speech comprehension to the rapidly evolving ASR systems that address the same problem domain.
Cai Wingfield, Li Su 0002, Xunying Liu, Chao Zhang 0031, Philip C. Woodland, Andrew Thwaites, Elisabeth Fonteneau, William D. Marslen-Wilson
PLoS Comput. Biol.5
2017 I-Vectors and Structured Neural Networks for Rapid Adaptation of Acoustic Models
abstract
A lot of interest has been risen in the last years on the adaptation of deep neural network (DNN) acoustic models, as the latter become the state-of-art in automatic speech recognition. This work focuses on approaches that allow for rapid and robust adaptation of such models. First, i-vectors are added to the DNN input as speaker-informed features. An informative prior is introduced to i-vector estimation to improve the robustness to limited adaptation data. I-vectors are then combined with a structured adaptive DNN, the multibasis adaptive neural network (MBANN), and the complementarity of these adaptation techniques is investigated. Moreover, i-vectors are used to predict the MBANN transforms, avoiding the initial decoding pass and alignment. These approaches are evaluated on a U.S. English Broadcast News (BN) transcription task with two distinct sets of test data. The first, from the BN task and BN-style Youtube videos, yields test data acoustically matched to the training data, while the second set is from acoustically mismatched Youtube videos of diverse context. The performance gains from these schemes are found to be sensitive to the level of mismatch between training and test sets. The MBANN system combined with i-vector input achieves best performance for BN test sets. The i-vector-based predictive MBANN scheme is proven to be more robust to acoustically mismatched conditions and outperforms the other adaptation schemes in such scenarios.
Panagiota Karanasou, Chunyang Wu, Mark J. F. Gales, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models
abstract
In recent years, recurrent neural network language models (RNNLMs) have become increasingly popular for a range of applications including speech recognition. However, the training of RNNLMs is computationally expensive, which limits the quantity of data, and size of network, that can be used. In order to fully exploit the power of RNNLMs, efficient training implementations are required. This paper introduces an open-source toolkit, the CUED-RNNLM toolkit, which supports efficient GPU-based training of RNNLMs. RNNLM training with a large number of word level output targets is supported, in contrast to existing tools which used class-based output-targets. Support fotN-best and lattice-based rescoring of both HTK and Kaldi format lattices is included. An example of building and evaluating RNNLMs with this toolkit is presented for a Kaldi based speech recognition system using the AMI corpus. All necessary resources including the source code, documentation and recipe are available online1.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP5
2016 Improved DNN-based segmentation for multi-genre broadcast audio
abstract
Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for multi-genre broadcast audio with deep neural network (DNN) based speech/non-speech detection. A further stage of change-point detection and clustering is used to obtain homogeneous segments. Suitable DNN inputs, context window sizes and architectures are studied with a series of experiments using a large corpus of MGB television audio. For MGB transcription, the improved segmenter yields roughly half the increase in word error rate, over manual segmentation, compared to the baseline DNN segmenter supplied for the 2015 ASRU MGB challenge.
Chao Zhang 0031, Philip C. Woodland, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin, Xunying Liu, Yanmin Qian
ICASSP3
2016 System combination with log-linear models
abstract
Improved speech recognition performance can often be obtained by combining multiple systems together. Joint decoding, where scores from multiple systems are combined during decoding rather than combining hypotheses, is one efficient approach for system combination. In standard joint decoding the frame log-likelihoods from each system are used as the scores. These scores are then weighted and summed to yield the final score for a frame. The system combination weights for this process are usually empirically set. In this paper, a recently proposed scheme for learning these system weights is investigated for a standard noise-robust speech recognition task, AURORA 4. High performance tandem and hybrid systems for this task are described. By applying state-of-the-art training approaches and configurations for the bottleneck features of the tandem system, the difference in performance between the tandem and hybrid systems is significantly smaller than usually observed on this task. A log-linear model is then used to estimate system weights between these systems. Training the system weights yields additional gains over empirically set system weights when used for decoding. Furthermore, when used in a lattice rescoring fashion, further gains can be obtained.
Chao Zhang 0031, Anton Ragni, Mark J. F. Gales, Philip C. Woodland
ICASSP5
2016 DNN speaker adaptation using parameterised sigmoid and ReLU hidden activation functions
abstract
This paper investigates the use of parameterised sigmoid and rectified linear unit (ReLU) hidden activation functions in deep neural network (DNN) speaker adaptation. The sigmoid and ReLU parameterisation schemes from a previous study for speaker independent (SI) training are used. An adaptive linear factor associated with each sigmoid or ReLU hidden unit is used to scale the unit output value and create a speaker dependent (SD) model. Hence, DNN adaptation becomes re-weighting the importance of different hidden units for every speaker. This adaptation scheme is applied to both hybrid DNN acoustic modelling and DNN-based bottleneck (BN) feature extraction. Experiments using multi-genre British English television broadcast data show that the technique is effective in both directly adapting DNN acoustic models and the BN features, and combines well with other DNN adaptation techniques. Reductions in word error rate are consistently obtained using parameterised sigmoid and ReLU activation function for multiple hidden layer adaptation.
Chao Zhang 0031, Philip C. Woodland
ICASSP2
2016 Selection of Multi-Genre Broadcast Data for the Training of Automatic Speech Recognition Systems
abstract
This paper compares schemes for the selection of multi-genre broadcast data and corresponding transcriptions for speech recognition model training. Selections of the same amount of data (700 hours) from lightly supervised alignments based on the same original subtitle transcripts are compared. Data segments were selected according to a maximum phone matched error rate between the lightly supervised decoding and the original transcript. The data selected with an improved lightly supervised system yields lower word error rates (WERs). Detailed comparisons of the data selected on carefully transcribed development data show how the selected portions match the true phone error rate for each genre. From a broader perspective, it is shown that for different genres, either the original subtitles or the lightly supervised output should be used for model training and a suitable combination yields further reductions in final WER.
Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanman Qian, Philip C. Woodland, Chao Zhang 0031
INTERSPEECH7
2016 Very deep convolutional neural networks for robust speech recognition
abstract
This paper describes the extension and optimisation of our previous work on very deep convolutional neural networks (CNNs) for effective recognition of noisy speech in the Aurora 4 task. The appropriate number of convolutional layers, the sizes of the filters, pooling operations and input feature maps are all modified: the filter and pooling sizes are reduced and dimensions of input feature maps are extended to allow adding more convolutional layers. Furthermore appropriate input padding and input feature map selection strategies are developed. In addition, an adaptation framework using joint training of very deep CNN with auxiliary features i-vector and fMLLR features is developed. These modifications give substantial word error rate reductions over the standard CNN used as baseline. Finally the very deep CNN is combined with an LSTM-RNN acoustic model and it is shown that state-level weighted log likelihood score combination in a joint acoustic model decoding scheme is very effective. On the Aurora 4 task, the very deep CNN achieves a WER of 8.81%, further 7.99% with auxiliary feature joint training, and 7.09% with LSTM-RNN joint decoding.
Yanmin Qian, Philip C. Woodland
SLT2
2016 Efficient Training and Evaluation of Recurrent Neural Network Language Models for Automatic Speech Recognition
abstract
Recurrent neural network language models (RNNLMs) are becoming increasingly popular for a range of applications including automatic speech recognition. An important issue that limits their possible application areas is the computational cost incurred in training and evaluation. This paper describes a series of new efficiency improving approaches that allows RNNLMs to be more efficiently trained on graphics processing units (GPUs) and evaluated on CPUs. First, a modified RNNLM architecture with a nonclass-based, full output layer structure (F-RNNLM) is proposed. This modified architecture facilitates a novel spliced sentence bunch mode parallelization of F-RNNLM training using large quantities of data on a GPU. Second, two efficient RNNLM training criteria based on variance regularization and noise contrastive estimation are explored to specifically reduce the computation associated with the RNNLM output layer softmax normalisation term. Finally, a pipelined training algorithm utilizing multiple GPUs is also used to further improve the training speed. Initially, RNNLMs were trained on a moderate dataset with 20M words from a large vocabulary conversational telephone speech recognition task. The training time of RNNLM is reduced by up to a factor of 53 on a single GPU over the standard CPU-based RNNLM toolkit. A 56 times speed up in test time evaluation on a CPU was obtained over the baseline F-RNNLMs. Consistent improvements in both recognition accuracy and perplexity were also obtained over C-RNNLMs. Experiments on Google's one billion corpus also reveals that the training of RNNLM scales well.
Xie Chen 0001, Xunying Liu, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 Two Efficient Lattice Rescoring Methods Using Recurrent Neural Network Language Models
abstract
An important part of the language modelling problem for automatic speech recognition (ASR) systems, and many other related applications, is to appropriately model long-distance context dependencies in natural languages. Hence, statistical language models (LMs) that can model longer span history contexts, for example, recurrent neural network language models (RNNLMs), have become increasingly popular for state-of-the-art ASR systems. As RNNLMs use a vector representation of complete history contexts, they are normally used to rescore N-best lists. Motivated by their intrinsic characteristics, two efficient lattice rescoring methods for RNNLMs are proposed in this paper. The first method uses an n-gram style clustering of history contexts. The second approach directly exploits the distance measure between recurrent hidden history vectors. Both methods produced 1-best performance comparable to a 10 k-best rescoring baseline RNNLM system on two large vocabulary conversational telephone speech recognition tasks for US English and Mandarin Chinese. Consistent lattice size compression and recognition performance improvements after confusion network (CN) decoding were also obtained over the prefix tree structured N-best rescoring approach.
Xunying Liu, Xie Chen 0001, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 The MGB challenge: Evaluating multi-genre broadcast media recognition
abstract
This paper describes the Multi-Genre Broadcast (MGB) Challenge at ASRU 2015, an evaluation focused on speech recognition, speaker diarization, and "lightly supervised" alignment of BBC TV recordings. The challenge training data covered the whole range of seven weeks BBC TV output across four channels, resulting in about 1,600 hours of broadcast audio. In addition several hundred million words of BBC subtitle text was provided for language modelling. A novel aspect of the evaluation was the exploration of speech recognition and speaker diarization in a longitudinal setting — i.e. recognition of several episodes of the same show, and speaker diarization across these episodes, linking speakers. The longitudinal tasks also offered the opportunity for systems to make use of supplied metadata including show title, genre tag, and date/time of transmission. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained.
Peter Bell 0001, Mark J. F. Gales, Thomas Hain, Jonathan Kilgour, Pierre Lanchantin, Xunying Liu, Andrew McParland, Steve Renals, Oscar Saz-Torralba, Mirjam Wester, Philip C. Woodland
ASRU11
2015 Investigation of back-off based interpolation between recurrent neural network and n-gram language models
abstract
Recurrent neural network language models (RNNLMs) have become an increasingly popular choice for speech and language processing tasks including automatic speech recognition (ASR). As the generalization patterns of RNNLMs and n-gram LMs are inherently different, RNNLMs are usually combined with n-gram LMs via a fixed weighting based linear interpolation in state-of-the-art ASR systems. However, previous work doesn't fully exploit the difference of modelling power of the RNNLMs and n-gram LMs as n-gram level changes. In order to fully exploit the detailed n-gram level complementary attributes between the two LMs, a back-off based compact representation of n-gram dependent interpolation weights is proposed in this paper. This approach allows weight parameters to be robustly estimated on limited data. Experimental results are reported on the three tasks with varying amounts of training data. Small and consistent improvements in both perplexity and WER were obtained using the proposed interpolation approach over the baseline fixed weighting based linear interpolation.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ASRU4
2015 Multilingual representations for low resource speech recognition and keyword search
abstract
This paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation.
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland
ASRU19
2015 Speaker diarisation and longitudinal linking in multi-genre broadcast data
abstract
This paper presents a multi-stage speaker diarisation system with longitudinal Linking developed on BBC multi-genre data for the 2015 Multi-Genre Broadcast (MGB) challenge. The basic speaker diarisation system draws on techniques from the Cambridge March 2005 system with a new deep neural network (DNN)-based speech/non speech segmenter. A newly developed linking stage is next added to the basic diarisation output aiming at the identification of speakers across multiple episodes of the same series. The longitudinal constraint imposes an incremental processing of the episodes, where speaker labels for each episode can be obtained using only material from the episode in question, and those broadcast earlier in time. The nature of the data as well as the longitudinal linking constraint position this diarisation task as a new open-research topic, and a particularly challenging one. Different linking clustering metrics are compared and the lowest within-episode and cross-episode DER scores are achieved on the MGB challenge evaluation set.
Panagiota Karanasou, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031
ASRU7
2015 The development of the cambridge university alignment systems for the multi-genre broadcast challenge
abstract
We describe the alignment systems developed both for the preparation of data for the Multi-Genre Broadcast (MGB) challenge and for our participation in the transcription and alignment tasks. Captions of varying quality are aligned with the audio of TV shows that range from few minutes long to more than six hours. Lightly supervised decoding is performed on the audio and the output text is aligned with the original text transcript. Reliable split points are found and the resulting text chunks are force-aligned with the corresponding audio segments. Confidence scores are associated with the aligned data. Multiple refinements — including audio segmentation based on deep neural networks (DNNs) and the use of DNN-based acoustic models — were used to improve the performance. The final MGB alignment system had the highest F-measure value on the evaluation data.
Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031
ASRU7
2015 Cambridge university transcription systems for the multi-genre broadcast challenge
abstract
We describe the development of our speech-to-text transcription systems for the 2015 Multi-Genre Broadcast (MGB) challenge. Key features of the systems are: a segmentation system based on deep neural networks (DNNs); the use of HTK 3.5 for building DNN-based hybrid and tandem acoustic models and the use of these models in a joint decoding framework; techniques for adaptation of DNN based acoustic models including parameterised activation function adaptation; alternative acoustic models built using Kaldi; and recurrent neural network language models (RNNLMs) and RNNLM adaptation. The same language models were used with both HTK and Kaldi acoustic models and various combined systems built. The final systems had the lowest error rates on the evaluation data.
Philip C. Woodland, Xunying Liu, Yanmin Qian, Chao Zhang 0031, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin
ASRU1
2015 Improving the training and evaluation efficiency of recurrent neural network language models
abstract
Recurrent neural network language models (RNNLMs) are becoming increasingly popular for speech recognition. Previously, we have shown that RNNLMs with a full (non-classed) output layer (F-RNNLMs) can be trained efficiently using a GPU giving a large reduction in training time over conventional class-based models (C-RNNLMs) on a standard CPU. However, since test-time RNNLM evaluation is often performed entirely on a CPU, standard F-RNNLMs are inefficient since the entire output layer needs to be calculated for normalisation. In this paper, it is demonstrated that C-RNNLMs can be efficiently trained on a GPU, using our spliced sentence bunch technique which allows good CPU test-time performance (42× speedup over F-RNNLM). Furthermore, the performance of different classing approaches is investigated. We also examine the use of variance regularisation of the softmax denominator for F-RNNLMs and show that it allows F-RNNLMs to be efficiently used in test (56× speedup on a CPU). Finally the use of two GPUs for F-RNNLM training using pipelining is described and shown to give a reduction in training time over a single GPU by a factor of 1.6×.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP4
2015 Recurrent neural network language model training with noise contrastive estimation for speech recognition
abstract
In recent years recurrent neural network language models (RNNLMs) have been successfully applied to a range of tasks including speech recognition. However, an important issue that limits the quantity of data used, and their possible application areas, is the computational cost in training. A signi??cant part of this cost is associated with the softmax function at the output layer, as this requires a normalization term to be explicitly calculated. This impacts both the training and testing speed, especially when a large output vocabulary is used. To address this problem, noise contrastive estimation (NCE) is explored in RNNLM training. NCE does not require the above normalization during both training and testing. It is insensitive to the output layer size. On a large vocabulary conversational telephone speech recognition task, a doubling in training speed on a GPU and a 56 times speed up in test time evaluation on a CPU were obtained.
Xie Chen 0001, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP4
2015 Paraphrastic recurrent neural network language models
abstract
Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems. Linguistic factors in??uencing the realization of surface word sequences, for example, expressive richness, are only implicitly learned by RNNLMs. Observed sentences and their associated alternative paraphrases representing the same meaning are not explicitly related during training. In order to improve context coverage and generalization, paraphrastic RNNLMs are investigated in this paper. Multiple paraphrase variants were automatically generated and used in paraphrastic RNNLM training. Using a paraphrastic multi-level RNNLM modelling both word and phrase sequences, signi??cant error rate reductions of 0.6% absolute and perplexity reduction of 10% relative were obtained over the baseline RNNLM on a large vocabulary conversational telephone speech recognition system trained on 2000 hours of audio and 545 million words of texts. The overall improvement over the baseline n-gram LM was increased from 8.4% to 11.6% relative.
Xunying Liu, Xie Chen 0001, Mark J. F. Gales, Philip C. Woodland
ICASSP4
2015 Recurrent neural network language model adaptation for multi-genre broadcast speech recognition
abstract
Recurrent neural network language models (RNNLMs) have recently become increasingly popular for many applications including speech recognition. In previous research RNNLMs have normally been trained on well-matched in-domain data. The adaptation of RNNLMs remains an open research area to be explored. In this paper, genre and topic based RNNLMadaptation techniques are investigated for a multi-genre broadcast transcription task. A number of techniques including Probabilistic Latent Semantic Analysis, Latent Dirichlet Allocation and Hierarchical Dirichlet Processes are used to extract show level topic information. These were then used as additional input to the RNNLM during training, which can facilitate unsupervised test time adaptation. Experiments using a state-of-theart LVCSR system trained on 1000 hours of speech and more than 1 billion words of text showed adaptation could yield perplexity reductions of 8% relatively over the baseline RNNLM and small but consistent word error rate reductions.
Xie Chen 0001, Tian Tan 0002, Xunying Liu, Pierre Lanchantin, M. Wan, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH7
2015 I-vector estimation using informative priors for adaptation of deep neural networks
abstract
This is the author accepted manuscript. The final version is available from ISCA via http://www.isca-speech.org/archive/interspeech_2015/i15_2872.html Supporting data for this paper is available at the http://www.repository.cam.ac.uk/handle/1810/248387 data repository.
Panagiota Karanasou, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2015 The Cambridge University 2014 BOLT conversational telephone Mandarin Chinese LVCSR system for speech translation
abstract
This paper presents the development of the 2014 Cambridge University conversational telephone Mandarin Chinese LVCSR system for the DARPA BOLT speech translation evaluation. A range of advanced modelling techniques were employed to both improve the recognition performance and provide a suitable integration with the translation system. These include an improved system combination technique using frame level acoustic model combination via joint decoding. Sequence trained deep neural network (DNN) based hybrid and tandem systems were combined on-the-fly to produce a consistent decoding output during search. A multi-level paraphrastic recurrent neural network LM (RNNLM) modelling both alternative paraphrase expressions and character sequences while preserving a consistent character to word segmentation was also used. This system gave an overall character error rate (CER) of 29.1% on the BOLT dev14 development set.
Xunying Liu, Federico Flego, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH6
2015 Joint decoding of tandem and hybrid systems for improved keyword spotting on low resource languages
abstract
Copyright © 2015 ISCA. Keyword spotting (KWS) for low-resource languages has drawn increasing attention in recent years. The state-of-the-art KWS systems are based on lattices or Confusion Networks (CN) generated by Automatic Speech Recognition (ASR) systems. It has been shown that considerable KWS gains can be obtained by combining the keyword detection results from different forms of ASR systems, e.g., Tandem and Hybrid systems. This paper investigates an alternative combination scheme for KWS using joint decoding. This scheme treats a Tandem system and a Hybrid system as two separate streams, and makes a linear combination of individual acoustic model log-likelihoods. Joint decoding is more efficient as it requires just a single pass of decoding and a single pass of keyword search. Experiments on six Babel OP2 development languages show that joint decoding is capable of providing consistent gains over each individual system. Moreover, it is possible to efficiently rescore the joint decoding lattices with Tandem or Hybrid acoustic models, and further KWS gains can be obtained by merging the detection posting lists from the joint decoding lattices and rescored lattices.
Anton Ragni, Mark J. F. Gales, Kate M. Knill, Philip C. Woodland, Chao Zhang 0031
INTERSPEECH5
2015 Parameterised sigmoid and reLU hidden activation functions for DNN acoustic modelling
abstract
The form of hidden activation functions has been always an important issue in deep neural network (DNN) design. The most common choices for acoustic modelling are the standard Sigmoid and rectified linear unit (ReLU), which are normally used with fixed function shapes and no adaptive parameters. Recently, there have been several papers that have studied the use of parameterised activation functions for both computer vision and speaker adaptation tasks. In this paper, we investigate generalised forms of both Sigmoid and ReLU with learnable parameters, as well as their integration with the standard DNN acoustic model training process. Experiments using conversational telephone speech (CTS) Mandarin data, result in an average of 3.4% and 2.0% relative word error rate (WER) reduction with Sigmoid and ReLU parameterisations.
Chao Zhang 0031, Philip C. Woodland
INTERSPEECH2
2015 A general artificial neural network extension for HTK
abstract
This paper describes the recently developed artificial neural network (ANN) modules in HTK hidden Markov model toolkit, which enables ANN models with very general feed-forward architectures to be used for either acoustic modelling or feature extraction. The HTK ANN extension includes many recent ANN-based speech processing techniques, such as sequence training, model stacking, speaker adaptation, and parameterised activation functions. The implementation allows efficient training by supporting GPUs and various types of data cache. The ANN modules are fully integrated into the rest of the HTK toolkit, which allows existing GMM-HMM methods to be easily used in the ANN-HMM framework. Speech recognition results on a 300 hours DARPA BOLT conversational Mandarin task show that HTK can produce tandem and hybrid systems with state-of-the-art performance on this very challenging task. Furthermore, the flexibility of the implementation is illustrated using demo systems for a Wall Street Journal (WSJ) task. The HTK ANN extension is planned for release in HTK version 3.5.
Chao Zhang 0031, Philip C. Woodland
INTERSPEECH2
2014 Paraphrastic neural network language models
abstract
Expressive richness in natural languages presents a significant challenge for statistical language models (LM). As multiple word sequences can represent the same underlying meaning, only modelling the observed surface word sequence can lead to poor context coverage. To handle this issue, paraphrastic LMs were previously proposed to improve the generalization of back-off n-gram LMs. Paraphrastic neural network LMs (NNLM) are investigated in this paper. Using a paraphrastic multi-level feedforward NNLM modelling both word and phrase sequences, significant error rate reductions of 1.3% absolute (8% relative) and 0.9% absolute (5.5% relative) were obtained over the baseline n-gram and NNLM systems respectively on a state-of-the-art conversational telephone speech recognition system trained on 2000 hours of audio and 545 million words of texts.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2014 Efficient lattice rescoring using recurrent neural network language models
abstract
Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems due to their inherently strong generalization performance. As these models use a vector representation of complete history contexts, RNNLMs are normally used to rescore N-best lists. Motivated by their intrinsic characteristics, two novel lattice rescoring methods for RNNLMs are investigated in this paper. The first uses an n-gram style clustering of history contexts. The second approach directly exploits the distance measure between hidden history vectors. Both methods produced 1-best performance comparable with a 10k-best rescoring baseline RNNLM system on a large vocabulary conversational telephone speech recognition task. Significant lattice size compression of over 70% and consistent improvements after confusion network (CN) decoding were also obtained over the N-best rescoring approach.
Xunying Liu, Yongqiang Wang 0006, Xie Chen 0001, Mark J. F. Gales, Philip C. Woodland
ICASSP5
2014 Detecting deletions in ASR output
abstract
In this work, the novel task of detecting deletions within automatic speech recognition (ASR) system output is investigated. Deletion-informed confidence estimation is proposed as an approach which simultaneously yields a confidence score in a word being correct, as well as a deletion confidence score which indicates whether a deletion is likely to occur in the output. The sequential nature of conditional random field (CRF) models is exploited as a means through which this can be achieved. It is shown that this sequence structure is crucial in yielding useful deletion detection scores, with an equivalent non-sequential model proven to be unsuitable for the task. The deletion-informed confidence estimation approach is also shown to outperform one where deletion confidence scores are estimated as a classification task separate from that of overall confidence estimation.
Matthew Stephen Seigel, Philip C. Woodland
ICASSP2
2014 Direct sub-word confidence estimation with hidden-state conditional random fields
abstract
The estimation of accurate confidence scores for sub-word-level units within automatic speech recognition (ASR) system transcriptions is investigated in this work. This is achieved through the application of linear-chain and hidden-state conditional random field (CRF) models to the task. A method for evaluating the significance of results quoted in terms of the normalised cross entropy (NCE) is also introduced. Instead of using sub-word-level information to improve wordlevel confidence scores, sub-word and word-level predictor features are combined to improve the accuracy of confidence scores in each sub-word being correct. The use of CRFs to model transitions between consecutive correct/incorrect sub-words yields large performance improvements. The scale of these gains is shown to increase further with the application of hidden-state CRFs. This is attributed to the fact that the hidden states make it possible for longer-span runs of consecutive correct/incorrect sub-words to be modelled, with these runs also not being constrained by word-level boundaries.
Matthew Stephen Seigel, Philip C. Woodland
ICASSP2
2014 Standalone training of context-dependent deep neural network acoustic models
abstract
Recently, context-dependent (CD) deep neural network (DNN) hidden Markov models (HMMs) have been widely used as acoustic models for speech recognition. However, the standard method to build such models requires target training labels from a system using HMMs with Gaussian mixture model output distributions (GMM-HMMs). In this paper, we introduce a method for training state-of-the-art CD-DNN-HMMs without relying on such a pre-existing system. We achieve this in two steps: build a context-independent (CI) DNN iteratively with word transcriptions, and then cluster the equivalent output distributions of the untied CD-DNN HMM states using the decision tree based state tying approach. Experiments have been performed on the Wall Street Journal corpus and the resulting system gave comparable word error rates (WER) to CD-DNNs built based on GMM-HMM alignments and state-clustering.
Chao Zhang 0031, Philip C. Woodland
ICASSP2
2014 Efficient GPU-based training of recurrent neural network language models using spliced sentence bunch
abstract
Recurrent neural network language models (RNNLMs) are be-coming increasingly popular for a range of applications includ-ing speech recognition. However, an important issue that limits the quantity of data, and hence their possible application ar-eas, is the computational cost in training. A standard approach to handle this problem is to use class-based outputs, allowing systems to be trained on CPUs. This paper describes an alter-native approach that allows RNNLMs to be efficiently trained on GPUs. This enables larger quantities of data to be used, and networks with an unclustered, full output layer to be trained. To improve efficiency on GPUs, multiple sentences are “spliced” together for each mini-batch or “bunch ” in training. On a large vocabulary conversational telephone speech recognition task, the training time was reduced by a factor of 27 over the stan-dard CPU-based RNNLM toolkit. The use of an unclustered, full output layer also improves perplexity and recognition per-formance over class-based RNNLMs. Index Terms: language models, recurrent neural network, speech recognition, GPU
Xie Chen 0001, Yongqiang Wang 0006, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH5
2014 Adaptation of deep neural network acoustic models using factorised i-vectors
abstract
The use of deep neural networks (DNNs) in a hybrid configuration is becoming increasingly popular and successful for speech recognition. One issue with these systems is how to efficiently adapt them to reflect an individual speaker or noise condition. Recently speaker i-vectors have been successfully used as an additional input feature for unsupervised speaker adaptation. In this work the use of i-vectors for adaptation is extended to incorporate acoustic factorisation. In particular, separate i-vectors are computed to represent speaker and acoustic environment. By ensuring "orthogonality" between the individual factor representations it is possible to represent a wide range of speaker and environment pairs by simply combining i-vectors from a particular speaker and a particular environment. In this paper the i-vectors are viewed as the weights of a cluster adaptive training (CAT) system, where the underlying models are GMMs rather than HMMs. This allows the factorisation approaches developed for CAT to be directly applied. Initial experiments were conducted on a noise distorted version of the WSJ corpus. Compared to standard speaker-based i-vector adaptation, factorised i-vectors showed performance gains.
Panagiota Karanasou, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH4
2014 Paraphrastic language models
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.3
2013 Investigation of multilingual deep neural networks for spoken term detection
abstract
The development of high-performance speech processing systems for low-resource languages is a challenging area. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to use bottleneck features, or hybrid systems, trained on multilingual data for speech-to-text (STT) systems. This paper presents an investigation into the application of these multilingual approaches to spoken term detection. Experiments were run using the IARPA Babel limited language pack corpora (~10 hours/language) with 4 languages for initial multilingual system development and an additional held-out target language. STT gains achieved through using multilingual bottleneck features in a Tandem configuration are shown to also apply to keyword search (KWS). Further improvements in both STT and KWS were observed by incorporating language questions into the Tandem GMM-HMM decision trees for the training set languages. Adapted hybrid systems performed slightly worse on average than the adapted Tandem systems. A language independent acoustic model test on the target language showed that retraining or adapting of the acoustic models to the target language is currently minimally needed to achieve reasonable performance.
Kate M. Knill, Mark J. F. Gales, Shakti P. Rath, Philip C. Woodland, Chao Zhang 0031, Shixiong Zhang 0001
ASRU4
2013 A high-performance Cantonese keyword search system
abstract
We present a system for keyword search on Cantonese conversational telephony audio, collected for the IARPA Babel program, that achieves good performance by combining postings lists produced by diverse speech recognition systems from three different research groups. We describe the keyword search task, the data on which the work was done, four different speech recognition systems, and our approach to system combination for keyword search. We show that the combination of four systems outperforms the best single system by 7%, achieving an actual term-weighted value of 0.517.
Brian Kingsbury, Jia Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP13
2013 Paraphrastic language models and combination with neural network language models
abstract
In natural languages multiple word sequences can represent the same underlying meaning. Only modelling the observed surface word sequence can result in poor context coverage, for example, when using n-gram language models (LM). To handle this issue, paraphrastic LMs were proposed in previous research and successfully applied to a US English conversational telephone speech transcription task. In order to exploit the complementary characteristics of paraphrastic LMs and neural network LMs (NNLM), the combination between the two is investigated in this paper. To investigate paraphrastic LMs' generalization ability to other languages, experiments are conducted on a Mandarin Chinese broadcast speech transcription task. Using a paraphrastic multi-level LM modelling both word and phrase sequences, significant error rate reductions of 0.9% absolute (9% relative) and 0.5% absolute (5% relative) were obtained over the baseline n-gram and NNLM systems respectively, after a combination with word and phrase level NNLMs.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2013 System combination and score normalization for spoken term detection
abstract
Spoken content in languages of emerging importance needs to be searchable to provide access to the underlying information. In this paper, we investigate the problem of extending data fusion methodologies from Information Retrieval for Spoken Term Detection on low-resource languages in the framework of the IARPA Babel program. We describe a number of alternative methods improving keyword search performance. We apply these methods to Cantonese, a language that presents some new issues in terms of reduced resources and shorter query lengths. First, we show score normalization methodology that improves in average by 20% keyword search performance. Second, we show that properly combining the outputs of diverse ASR systems performs 14% better than the best normalized ASR system.
Jonathan Mamou, Jia Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland
ICASSP13
2013 A confidence-based approach for improving keyword hypothesis scores
abstract
The task in keyword spotting (KWS) is to hypothesise times at which any of a set of key terms occurs in audio. An important aspect of such systems are the scores assigned to these hypotheses, the accuracy of which have a significant impact on performance. Estimating these scores may be formulated as a confidence estimation problem, where a measure of confidence is assigned to each key term hypothesis. In this work, a set of discriminative features is defined, and combined using a conditional random field (CRF) model for improved confidence estimation. An extension to this model to directly address the problem of score normalisation across key terms is also introduced. The implicit score normalisation which results from applying this approach to separate systems in a hybrid configuration yields further benefits. Results are presented which show notable improvements in KWS performance using the techniques presented in this work.
Matthew Stephen Seigel, Philip C. Woodland, Mark J. F. Gales
ICASSP2
2013 Cross-domain paraphrasing for improving language modelling using out-of-domain data
abstract
In natural languages the variability in the underlying linguistic generation rules significantly alters the observed surface word sequence they create, and thus introduces a mismatch against other data generated via alternative realizations associated with, for example, a different domain. Hence, direct modelling of out-of-domain data can result in poor generalization to the indomain data of interest. To handle this problem, this paper investigated using cross-domain paraphrastic language models to improve in-domain language modelling (LM) using out-ofdomain data. Phrase level paraphrase models learnt from each domain were used to generate paraphrase variants for the data of other domains. These were used to both improve the context coverage of in-domain data, and reduce the domain mismatch of the out-of-domain data. Significant error rate reduction of 0.6% absolute was obtained on a state-of-the-art conversational telephone speech recognition task using a cross-domain paraphrastic multi-level LM trained on a billion words of mixed conversational and broadcast news data. Consistent improvements on the in-domain data context coverage were also obtained. Copyright © 2013 ISCA.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2013 Improving lightly supervised training for broadcast transcription
abstract
This paper investigates improving lightly supervised acoustic model training for an archive of broadcast data. Standard lightly supervised training uses automatically derived decoding hypotheses using a biased language model. However, as the actual speech can deviate significantly from the original programme scripts that are supplied, the quality of standard lightly supervised hypotheses can be poor. To address this issue, word and segment level combination approaches are used between the lightly supervised transcripts and the original programme scripts which yield improved transcriptions. Experimental results show that systems trained using these improved transcriptions consistently outperform those trained using only the original lightly supervised decoding hypotheses. This is shown to be the case for both the maximum likelihood and minimum phone error trained systems.
Yanhua Long, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Matthew Stephen Seigel, Philip C. Woodland
INTERSPEECH6
2013 Use of contexts in language model interpolation and adaptation
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.3
2013 Language model cross adaptation for LVCSR system combination
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.3
2012 Complementary Phone Error Training
abstract
This paper introduces a novel method for the training of a complementary acoustic model with respect to set of given acoustic models. The method is based upon an extension of the Minimum Phone Error (MPE) criterion and aims at producing a model that makes complementary phone errors to those already trained. The technique is therefore called Complementary Phone Error (CPE) training. The method is evaluated using an Arabic large vocabulary continuous speech recognition task. Reductions in word error rate (WER) after combination with a CPE-trained system were obtained with up to 0.7% absolute for a system trained on 172 hours of acoustic data and up to 0.2% absolute for the final system trained on nearly 2000 hours of Arabic data.
Frank Diehl, Philip C. Woodland
INTERSPEECH2
2012 Paraphrastic Language Models
abstract
Natural languages are known for their expressive richness.Many sentences can be used to represent the same underlying meaning.Only modelling the observed surface word sequence can result in poor context coverage and generalization, for example, when using n-gram language models (LMs).This paper proposes a novel form of language model, the paraphrastic LM, that addresses these issues.A phrase level paraphrase model statistically learned from standard text data with no semantic annotation is used to generate multiple paraphrase variants.LM probabilities are then estimated by maximizing their marginal probability.Multi-level language models estimated at both the word level and the phrase level are combined.An efficient weighted finite state transducer (WFST) based paraphrase generation approach is also presented.Significant error rate reductions of 0.5%-0.6%absolute were obtained over the baseline n-gram LMs on two state-of-the-art recognition tasks for English conversational telephone speech and Mandarin Chinese broadcast speech using a paraphrastic multi-level LM modelling both word and phrase sequences.When it is further combined with word and phrase level feed-forward neural network LMs, a significant error rate reduction of 0.9% absolute (9% relative) and 0.5% absolute (5% relative) were obtained over the baseline n-gram and neural network LMs respectively.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2012 Using Sub-word-level Information for Confidence Estimation with Conditional Random Field Models
abstract
The task of word-level confidence estimation (CE) for automatic speech recognition (ASR) systems stands to benefit from the combination of suitably defined input features from multiple information sources. However, the information sources of interest may not necessarily operate at the same level of granularity as the underlying ASR system. The research described here builds on previous work on confidence estimation for ASR systems using features extracted from word-level recognition lattices, by incorporating information at the sub-word level. Furthermore, the use of Conditional Random Fields (CRFs) with hidden states is investigated as a technique to combine information for word-level CE. Performance improvements are shown using the sub-word-level information in linear-chain CRFs with appropriately engineered feature functions, as well as when applying the hidden-state CRF model at the word level.
Matthew Stephen Seigel, Philip C. Woodland
INTERSPEECH2
2012 Transcription of multi-genre media archives using out-of-domain data
abstract
We describe our work on developing a speech recognition system for multi-genre media archives. The high diversity of the data makes this a challenging recognition task, which may benefit from systems trained on a combination of in-domain and out-of-domain data. Working with tandem HMMs, we present Multi-level Adaptive Networks (MLAN), a novel technique for incorporating information from out-of-domain posterior features using deep neural networks. We show that it provides a substantial reduction in WER over other systems, with relative WER reductions of 15% over a PLP baseline, 9% over in-domain tandem features and 8% over the best out-of-domain tandem features.
Peter Bell 0001, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanhua Long, Steve Renals, Pawel Swietojanski, Philip C. Woodland
SLT8
2012 Morphological decomposition in Arabic ASR systems
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
Comput. Speech Lang.4
2011 Investigation of acoustic units for LVCSR systems
abstract
One important issue in designing state-of-the-art LVCSR systems is the choice of acoustic units. Context dependent (CD) phones remain die dominant form of acoustic units. They can capture the co-articulatory effect in speech via explicit modelling. However, for other more complicated phonological processes, they rely on the implicit modelling ability of the underlying statistical models. Alternatively, it is possible to construct acoustic models based on higher level linguistic units, for example, syllables, to explicitly capture these complex patterns. When sufficient training data is available, this approach may show an advantage over implicit acoustic modelling. In this paper a wide range of acoustic units are investigated to improve LVCSR system performance. Significant error rate gains up to 7.1% relative (0.8% abs.) were obtained on a state-of-the-art Mandarin Chinese broadcast audio recognition task using word and syllable position dependent triphone and quinphone models.
Xunying Liu, Mark J. F. Gales, James Hieronymus, Philip C. Woodland
ICASSP4
2011 Word Boundary Modelling and Full Covariance Gaussians for Arabic Speech-to-Text Systems
abstract
This paper describes recent improvements to the Cambridge Arabic Large Vocabulary Continuous Speech Recognition (LVCSR) Speech-to-Text (STT) system. It is shown that word-boundary context markers provide a powerful method to enhance graphemic systems by implicit phonetic information, improving the modelling capability of graphemic systems. In addition, a robust technique for full covariance Gaussian modelling in the Minimum Phone Error (MPE) training framework is introduced. This reduces the full covariance training to a diagonal covariance training problem, thereby solving related robustness problems. The full system results show that the combined use of these and other techniques within a multi-branch combination framework reduces the Word Error Rate (WER) of the complete system by up to 5.9 % relative.
Frank Diehl, Mark J. F. Gales, Xunying Liu, Marcus Tomalin, Philip C. Woodland
INTERSPEECH5
2011 Graphone Model Interpolation and Arabic Pronunciation Generation
abstract
This paper extends n-gram graphone model pronunciation generation to use a mixture of such models. This technique is useful when pronunciation data is for a specific variant (or set of variants) of a language, such as for a dialect, and only a small amount of pronunciation dictionary training data for that specific variant is available. The performance of the interpolated n-gram graphone model is evaluated on Arabic phonetic pronunciation generation for words that can't be handled by the Buckwalter Morphological Analyser. The pronunciations produced are also used to train an Arabic broadcast audio speech recognition system. In both cases the interpolated graphone model leads to improved performance. Copyright © 2011 ISCA.
Philip C. Woodland, Frank Diehl, Mark J. F. Gales
INTERSPEECH2
2011 Improving LVCSR System Combination Using Neural Network Language Model Cross Adaptation
abstract
State-of-the-art large vocabulary continuous speech recognition (LVCSR) systems often combine outputs from multiple sub-systems developed at different sites. Cross system adaptation can be used as an alternative to direct hypothesis level combina-tion schemes such as ROVER. The standard approach involves only cross adapting acoustic models. To fully exploit the com-plimentary features among sub-systems, language model (LM) cross adaptation techniques can be used. Previous research on multi-level n-gram LM cross adaptation is extended to further include the cross adaptation of neural network LMs in this pa-per. Using this improved LM cross adaptation framework, sig-nificant error rate gains of 4.0%-7.1 % relative were obtained over acoustic model only cross adaptation when combining a range of Chinese LVCSR sub-systems used in the 2010 and 2011 DARPA GALE evaluations. 1.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2011 Combining Information Sources for Confidence Estimation with CRF Models
abstract
Obtaining accurate confidence measures for automatic speech recognition (ASR) transcriptions is an important task which stands to benefit from the use of multiple information sources. This paper investigates the application of conditional random field (CRF) models as a principled technique for combining multiple features from such sources. A novel method for combining suitably defined features is presented, allowing for confidence annotation using lattice-based features of hypotheses other than the lattice 1-best. The resulting framework is applied to different stages of a state-of-the-art large vocabulary speech recognition pipeline, and consistent improvements are shown over a sophisticated baseline system. Copyright © 2011 ISCA.
Matthew Stephen Seigel, Philip C. Woodland
INTERSPEECH2
2011 The efficient incorporation of MLP features into automatic speech recognition systems
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
Comput. Speech Lang.5
2010 Language model combination and adaptation usingweighted finite state transducers
abstract
In speech recognition systems language model (LMs) are often constructed by training and combining multiple n-gram models. They can be either used to represent different genres or tasks found in diverse text sources, or capture stochastic properties of different linguistic symbol sequences, for example, syllables and words. Unsupervised LM adaptation may also be used to further improve robustness to varying styles or tasks. When using these techniques, extensive software changes are often required. In this paper an alternative and more general approach based on weighted finite state transducers (WFSTs) is investigated for LM combination and adaptation. As it is entirely based on well-defined WFST operations, minimum change to decoding tools is needed. A wide range of LM combination configurations can be flexibly supported. An efficient on-the-fly WFST decoding algorithm is also proposed. Significant error rate gains of 7.3% relative were obtained on a state-of-the-art broadcast audio recognition task using a history dependently adapted multi-level LM modelling both syllable and word sequences.
Xunying Liu, Mark J. F. Gales, James Hieronymus, Philip C. Woodland
ICASSP4
2010 Recent improvements to the Cambridge Arabic Speech-to-Text systems
abstract
This paper describes recent improvements to the Cambridge Arabic Large Vocabulary Continuous Speech Recognition (LVSCR) Speech-to-Text (STT) system. It is shown that Multi-Layer Perceptron (MLP) features trained on phonetic targets can improve the performance of both phonemic and graphemic systems. Also, a morphological decomposition scheme is extended from the graphemic domain to the phonetic domain, and particular attention is given to the task of dictionary generation. Finally, the use of Boosted Maximum Mutual Information (BMMI) training is explored both for individual systems and in the context of system combination. The full system results show that the combined use of the above techniques reduces the Word Error Rate (WER) of the best individual system by up to 12% relative, and that the incorporation of morphological decomposition and BMMI within the four individual branches of the combined system reduces the WER by up to 9% relative.
Marcus Tomalin, Frank Diehl, Mark J. F. Gales, Philip C. Woodland
ICASSP5
2010 Language model cross adaptation for LVCSR system combination
abstract
State-of-the-art large vocabulary continuous speech recognition (LVCSR) systems often combine outputs from multiple sub-systems that may even be developed at different sites. Cross system adaptation, in which model adaptation is performed using the outputs from another sub-system, can be used as an alternative to hypothesis level combination schemes such as ROVER. Normally cross adaptation is only performed on the acoustic models. However, there are many other levels in LVCSR systems' modelling hierarchy where complimentary features may be exploited, for example, the sub-word and the word level, to further improve cross adaptation based system combination. It is thus interesting to also cross adapt language models (LMs) to capture these additional useful features. In this paper cross adaptation is applied to three forms of language models, a multi-level LM that models both syllable and word sequences, a word level neural network LM, and the linear combination of the two. Significant error rate reductions of 4.0-7.1% relative were obtained over ROVER and acoustic model only cross adaptation when combining a range of Chinese LVCSR sub-systems used in the 2010 and 2011 DARPA GALE evaluations.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2010 Improved neural network based language modelling and adaptation
abstract
Neural network language models (NNLM) have become an increasingly popular choice for large vocabulary continuous speech recognition (LVCSR) tasks, due to their inherent gener-alisation and discriminative power. This paper present two tech-niques to improve performance of standard NNLMs. First, the form of NNLM is modelled by introduction an additional out-put layer node to model the probability mass of out-of-shortlist (OOS) words. An associated probability normalisation scheme is explicitly derived. Second, a novel NNLM adaptation method using a cascaded network is proposed. Consistent WER reduc-tions were obtained on a state-of-the-art Arabic LVCSR task over conventional NNLMs. Further performance gains were also observed after NNLM adaptation.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH4
2010 Unsupervised training and directed manual transcription for LVCSR
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
Speech Commun.4
2009 Training and adapting MLP features for Arabic speech recognition
abstract
Features derived from multilayer perceptrons (MLPs) are becoming increasingly popular for speech recognition. This paper describes various schemes for applying these features to state-of-the-art Arabic speech recognition: the use of MLP-features for short-vowel modelling in graphemic systems; rapid discriminative model training by standard PLP feature lattice reuse; and MLP feature adaptation using linear input networks (LIN). The use of rapid training using MLP features and their use for short-vowel modelling and LIN adaptation gave reductions in word error rate. However significant improvements over explicit short-vowel modelling with standard multi-pass adaptation were not obtained, although they were useful in combination.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
ICASSP5
2009 Morphological analysis and decomposition for Arabic speech-to-text systems
abstract
Language modelling for a morphologically complex language such as Arabic is a challenging task. Its agglutinative structure re-sults in data sparsity problems and high out-of-vocabulary rates. In this work these problems are tackled by applying the MADA tools to the Arabic text. In addition to morphological decompo-sition, MADA performs context-dependent stem-normalisation. Thus, if word-level system combination, or scoring, is required this normalisation must be reversed. To address this, a novel context-sensitive method for morpheme-to-word conversion is introduced. The performance of the MADA decomposed sys-tem was evaluated on an Arabic broadcast transcription task. The MADA-based system out-performed the word-based system, with both the morphological decomposition and stem normalisa-tion being found to be important.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
INTERSPEECH4
2009 Exploiting Chinese character models to improve speech recognition performance
abstract
The Chinese language is based on characters which are syllabic in nature. Since languages have syllabotactic rules which govern the construction of syllables and their allowed sequences, Chinese character sequence models can be used as a first level approximation of allowed syllable sequences. N-gram character sequence models were trained on 4.3 billion characters. Characters are used as a first level recognition unit with multiple pronunciations per character. For comparison the CU-HTK Mandarin word based system was used to recognize words which were then converted to character sequences. The character only system error rates for one best recognition were slightly worse than word based character recognition. However combining the two systems using log-linear combination gives better results than either system separately. An equally weighted combination gave consistent CER gains of 0.1-0.2% absolute over the word based standard system. Copyright © 2009 ISCA.
James Hieronymus, Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH4
2009 Use of contexts in language model interpolation and adaptation
abstract
Language models (LMs) are often constructed by building multiple individual component models that are combined using context independent interpolation weights. By tuning these weights, using either perplexity or discriminative approaches, it is possible to adapt LMs to a particular task. This paper investigates the use of context dependent weighting in both interpolation and test-time adaptation of language models. Depending on the previous word contexts, a discrete history weighting function is used to adjust the contribution from each component model. As this dramatically increases the number of parameters to estimate, robust weight estimation schemes are required. Several approaches are described in this paper. The first approach is based on MAP estimation where interpolation weights of lower order contexts are used as smoothing priors. The second approach uses training data to ensure robust estimation of LM interpolation weights. This can also serve as a smoothing prior for MAP adaptation. A normalized perplexity metric is proposed to handle the bias of the standard perplexity criterion to corpus size. A range of schemes to combine weight information obtained from training data and test data hypotheses are also proposed to improve robustness during context dependent LM adaptation. In addition, a minimum Bayes' risk (MBR) based discriminative training scheme is also proposed. An efficient weighted finite state transducer (WFST) decoding algorithm for context dependent interpolation is also presented. The proposed technique was evaluated using a state-of-the-art Mandarin Chinese broadcast speech transcription task. Character error rate (CER) reductions up to 7.3 relative were obtained as well as consistent perplexity improvements. © 2012 Elsevier Ltd. All rights reserved.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2009 Efficient generation and use of MLP features for Arabic speech recognition
abstract
Front-end features computed using Multi-Layer Perceptrons (MLPs) have recently attracted much interest, but are a challenge to scale to large networks and very large training data sets. This paper discusses methods to reduce the training time for the generation of MLP features and their use in an ASR system using a variety of techniques: parallel training of a set of MLPs on different data sub-sets; methods for computing features from by a combination of these networks; and rapid discriminative training of HMMs using MLP-based features. The impact on MLP frame-based accuracy using different training strategies is discussed along with the effect on word rates from incorporating the MLP features in various configurations into an Arabic broadcast audio transcription system.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
INTERSPEECH5
2009 Unsupervised Adaptation With Discriminative Mapping Transforms
abstract
The most commonly used approaches to speaker adaptation are based on linear transforms, as these can be robustly estimated using limited adaptation data. Although significant gains can be obtained using discriminative criteria for training acoustic models, maximum-likelihood (ML) estimated transforms are still used for unsupervised adaptation. This is because discriminatively trained transforms are highly sensitive to errors in the adaptation supervision hypothesis. This paper describes a new framework for estimating transforms that are discriminative in nature, but are less sensitive to this hypothesis issue. A speaker-independent discriminative mapping transformation (DMT) is estimated during training. This transform is obtained after a speaker-specific ML-estimated transform of each training speaker has been applied. During recognition an ML speaker-specific transform is found for each test-set speaker and the speaker-independent DMT then applied. This allows a transform which is discriminative in nature to be indirectly estimated, while only requiring an ML speaker-specific transform to be found during recognition. The DMT technique is evaluated on an English conversational telephone speech task. Experiments showed that using DMT in unsupervised adaptation led to significant gains over both standard ML and discriminatively trained transforms.
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
IEEE Trans. Speech Audio Process.3
2008 Phonetic pronunciations for arabic speech-to-text systems
abstract
In this paper two aspects of generating and using phonetic arabic dictionaries are described. First, the use of single pronunciation acoustic models in the context of arabic large vocabulary automatic speech recognition (ASR) is investigated. These have been found to be useful for English ASR systems, when combined with standard multiple pronunciation systems. The second area examined is automatically deriving phonetic "pronunciations" for words that standard approaches, such as the Buckwalter morphological analyzer, cannot handle. Without pronunciations for these words the OOV rates for various Arabic tasks significantly increase. Here, pronunciations are automatically found by first deriving grapheme-to-phone rules, and associated rule probabilities. These are then used to produce the most likely pronunciation, or pronunciations, for any word. These approaches are evaluated on a large vocabulary arabic broadcast news and broadcast conversation transcription task. Both schemes are found to yield gains with a multi-pass/combination framework.
Frank Diehl, Mark J. F. Gales, Marcus Tomalin, Philip C. Woodland
ICASSP4
2008 Unsupervised discriminative adaptation using discriminative mapping transforms
abstract
The most commonly used approaches to speaker adaptation are based on linear transforms, as these can be robustly estimated using limited adaptation data. Although significant gains can be obtained using discriminative criteria for training acoustic models, maximum likelihood (ML) estimated transforms are used for unsupervised adaptation. This is because discriminatively trained transforms are highly sensitive to errors in the adaptation hypothesis. This paper describes a new framework for estimating transforms that are discriminative in nature, but are less sensitive to this hypothesis issue. A discriminative, speaker-independent, mapping transformation is estimated during training. This transform is obtained after a speaker-specific ML-estimated transform has been applied. During recognition an ML speaker-specific transform is found and the speaker-independent discriminative mapping transform then applied. This allows a transform which is discriminative in nature to be indirectly estimated, whilst only requiring an ML speaker-specific transform to be found during recognition. The scheme is evaluated on an English conversational telephone speech task, where it significantly outperforms both standard ML and discriminatively trained transforms.
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
ICASSP3
2008 Context dependent language model adaptation
abstract
Language models (LMs) are often constructed by building multiple component LMs that are combined using interpolation weights. By tuning these interpolation weights, using either perplexity or discriminative approaches, it is possible to adapt LMs to a particular task. In this work, improved LM adaptation is achieved by introducing context dependent interpolation weights. An important part of this new approach is obtaining robust estimation. Two schemes for this are described. The first is based on MAP estimation, where either global interpolation weights are used as priors, or context dependent interpolation priors obtained from the training data. The second scheme uses class based contexts to determine the interpolation weights. Both schemes are evaluated using unsupervised LM adaptation on a Mandarin broadcast transcription task. Consistent gains in perplexity using context dependent, rather than global, weights are observed as well as reductions in character error rate. 1.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2008 MPE-based discriminative linear transforms for speaker adaptation
Philip C. Woodland
Comput. Speech Lang.2
2007 Development of a phonetic system for large vocabulary Arabic speech recognition
abstract
This paper describes the development of an Arabic speech recognition system based on a phonetic dictionary. Though phonetic systems have been previously investigated, this paper makes a number of contributions to the understanding of how to build these systems, as well as describing a complete Arabic speech recognition system. The first issue considered is discriminative training when there are a large number of pronunciation variants for each word. In particular, the loss function associated with Minimum Phone Error (MPE) training is examined. The performance and combination of phonetic and graphemic acoustic models are then compared on both Broadcast News (BN) and Broadcast Conversation (BC) data. The final contribution of the paper is a simple scheme for automatically generating pronunciations for use in training and reducing the phonetic out-of-vocabulary rate. The paper concludes with a description and results from using phonetic and graphemic systems in a multipass/ combination framework.
Mark J. F. Gales, Frank Diehl, Chandra Kant Raut, Marcus Tomalin, Philip C. Woodland, Kai Yu 0004
ASRU5
2007 Discriminative language model adaptation for Mandarin broadcast speech transcription and translation
abstract
This paper investigates unsupervised test-time adaptation of language models (LM) using discriminative methods for a Mandarin broadcast speech transcription and translation task. A standard approach to adapt interpolated language models to is to optimize the component weights byminimizing the perplexity on supervision data. This is a widely made approximation for language modeling in automatic speech recognition (ASR) systems. For speech translation tasks, it is unclear whether a strong correlation still exists between perplexity and various forms of error cost functions in recognition and translation stages. The proposed minimum Bayes risk (MBR) based approach provides a flexible framework for unsupervised LM adaptation. It generalizes to a variety of forms of recognition and translation error metrics. LM adaptation is performed at the audio document level using either the character error rate (CER), or translation edit rate (TER) as the cost function. An efficient parameter estimation scheme using the extended Baum-Welch (EBW) algorithm is proposed. Experimental results on a state-of-the-art speech recognition and translation system are presented. The MBR adapted language models gave the best recognition and translation performance and reduced the TER score by up to 0.54% absolute.
Xunying Liu, William J. Byrne, Mark J. F. Gales, Adrià de Gispert, Marcus Tomalin, Philip C. Woodland, Kai Yu 0004
ASRU6
2007 Speech Recognition System Combination for Machine Translation
abstract
The majority of state-of-the-art speech recognition systems make use of system combination. The combination approaches adopted have traditionally been tuned to minimising word error rates (WERs). In recent years there has been a growing interest in taking the output from speech recognition systems in one language and translating it into another. This paper investigates the use of cross-site combination approaches in terms of both WER and impact on translation performance. In addition, the stages involved in modifying the output from a speech-to-text (STT) system to be suitable for translation are described. Two source languages, Mandarin and Arabic, are recognised and then translated using a phrase-based statistical machine translation system into English. Performance of individual systems and cross-site combination using cross-adaptation and ROVER are given. Results show that the best STT combination scheme in terms of WER is not necessarily the most appropriate when translating speech.
Mark J. F. Gales, Xunying Liu, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004, Spyridon Matsoukas, Tim Ng, Kham Nguyen, Long Nguyen 0001, Jean-Luc Gauvain, Lori Lamel, Abdelkhalek Messaoudi
ICASSP (4)4
2007 Consensus Network Decoding for Statistical Machine Translation System Combination
abstract
This paper presents a simple and robust consensus decoding approach for combining multiple machine translation (MT) system outputs. A consensus network is constructed from an N-best list by aligning the hypotheses against an alignment reference, where the alignment is based on minimising the translation edit rate (TER). The minimum Bayes risk (MBR) decoding technique is investigated for the selection of an appropriate alignment reference. Several alternative decoding strategies proposed to retain coherent phrases in the original translations. Experimental results are presented primarily based on three-way combination of Chinese-English translation outputs, and also presents results for six-way system combination. It is shown that worthwhile improvements in translation performance can be obtained using the methods discussed.
Khe Chai Sim, William J. Byrne, Mark J. F. Gales, Hichem Sahbi, Philip C. Woodland
ICASSP (4)5
2007 Improving Speech Transcription for Mandarin-English Translation
abstract
This paper describes the development of the CU-HTK Mandarin speech-to-text (STT) system and assesses its performance as part of a transcription-translation pipeline which converts broadcast Mandarin audio into English text. Recent improvements to the STT system are described and these give character error rate (CER) gains of 14.3% absolute for a broadcast conversation (BC) task and 5.1% absolute for a broadcast news (BN) task. The output of these STT systems is then post-processed, so that it consists of sentence-like segments, and translated into English text using a statistical machine translation (SMT) system. The performance of the transcription-translation pipeline is evaluated using the translation edit rate (TER) and BLEU metrics. It is shown that improving both the STT system and the post-STT segmentations can lower the TER scores by up to 5.3% absolute and increase the BLEU scores by up to 2.7% absolute.
Marcus Tomalin, Mark J. F. Gales, Xunying Liu, Khe Chai Sim, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004
ICASSP (4)7
2007 Unsupervised Training for Mandarin Broadcast News and Conversation Transcription
abstract
A significant cost in obtaining acoustic training data is the generation of accurate transcriptions. For some sources close-caption data is available. This allows the use of lightly-supervised training techniques. However, for some sources and languages close-caption is not available. In these cases unsupervised training techniques must be used. This paper examines the use of unsupervised techniques for discriminative training. In unsupervised training automatic transcriptions from a recognition system are used for training. As these transcriptions may be errorful data selection may be useful. Two forms of selection are described, one to remove non-target language shows, the other to remove segments with low confidence. Experiments were carried out on a Mandarin transcriptions task. Two types of test data were considered, broadcast news (BN) and broadcast conversations (BC). Results show that the gains from unsupervised discriminative training are highly dependent on the accuracy of the automatic transcriptions.
Mark J. F. Gales, Philip C. Woodland
ICASSP (4)3
2007 Unsupervised training with directed manual transcription for recognising Mandarin broadcast audio
Kai Yu 0004, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH3
2006 The Cu-Htk Mandarin Broadcast News Transcription System
abstract
This paper discusses the development of the CU-HTK Mandarin broadcast news (BN) transcription system. The Mandarin BN task includes a significant amount of English data. Hence techniques have been investigated to allow the same system to handle both Mandarin and English by augmenting the Mandarin training sets with English acoustic and language model training data. A range of acoustic models were built including models based on Gaussianised features, speaker adaptive training and feature-space MPE. A multi-branch system architecture is described in which multiple acoustic model types, alternate phone sets and segmentations can be used in a system combination framework to generate the final output. The final system shows state-of-the-art performance over a range of test sets
Rohit Sinha 0003, Mark J. F. Gales, Do Yeong Kim, Xunying Liu, Khe Chai Sim, Philip C. Woodland
ICASSP (1)6
2006 Discriminatively Trained Gaussian Mixture Models for Sentence Boundary Detection
abstract
This paper compares the performance of two types of prosodic feature models (PFMs) in a sentence boundary detection task. Specifically, systems are compared that use discriminatively trained Gaussian mixture models (MMI-GMMs) and CART-style decision trees (CDT-PFMs), along with task-specific language models, in a lattice-based decoding framework in order automatically to insert slash unit (SU) boundaries into automatic speech recognition (ASR) transcriptions of input audio files. It is shown that a system which uses MMI-GMMs performs as well as a system that uses conventional CDT-PFMs. In addition, it is shown that, when the CDT-PFM and MMI-GMM systems are combined by taking weighted averages of their respective probability streams, error rate improvements of up to 0.8% abs over the CDT-PFM baseline can be obtained for four different test sets
Marcus Tomalin, Philip C. Woodland
ICASSP (1)2
2006 Unsupervised language model adaptation for Mandarin broadcast conversation transcription
abstract
1. Introduction This paper investigates language model adaptation in a speechrecognition setting with a mix of data of two mismatching styles.
David Mrva, Philip C. Woodland
INTERSPEECH2
2006 Progress in the CU-HTK broadcast news transcription system
abstract
Broadcast news (BN) transcription has been a challenging research area for many years. In the last couple of years, the availability of large amounts of roughly transcribed acoustic training data and advanced model training techniques has offered the opportunity to greatly reduce the error rate on this task. This paper describes the design and performance of BN transcription systems which make use of these developments. First, the effects of using lightly supervised training data and advanced acoustic modeling techniques are discussed. The design of a real-time broadcast news recognition system is then detailed using these new models. As system combination has been found to yield large gains in performance, a range of frameworks that allow multiple recognition outputs to be combined are next described. These include the use of multiple types of acoustic models and multiple segmentations. As a contrast a system developed by multiple sites allowing cross-site combination, the "SuperEARS" system, is also described. The various models and recognition configurations are evaluated using several recent BN development and evaluation test sets. These new BN transcription systems can give gains of over 25% relative to the CU-HTK 2003 BN system
Mark J. F. Gales, Do Yeong Kim, Philip C. Woodland, Ricky Ho Yin Chan, David Mrva, Rohit Sinha 0003, Sue Tranter
IEEE Trans. Speech Audio Process.3
2006 Corrections to "Automatic Transcription of Conversational Telephone Speech"
Thomas Hain, Philip C. Woodland, Gunnar Evermann, Mark J. F. Gales, Xunying Liu, Gareth L. Moore, Daniel Povey
IEEE Trans. Speech Audio Process.2
2005 Training LVCSR Systems on Thousands of Hours of Data
abstract
Typical systems for large vocabulary conversational speech recognition (LVCSR) have been trained on a few hundred hours of carefully transcribed acoustic training data. The paper describes an LVCSR system for the conversational telephone speech (CTS) task trained on more than 2000 hours of data for which only approximate transcriptions were available. The challenges of dealing with such a large data set and the accuracy improvements over the small baseline system are discussed. The effect on both acoustic and language modelling performance is studied. Overall, increasing the training data size from 360 h to 2200 h and optimising the training procedure reduced the word error rate on the DARPA/NIST 2003 evaluation set by about 20% relative.
Gunnar Evermann, Ricky Ho Yin Chan, Mark J. F. Gales, David Mrva, Philip C. Woodland, Kai Yu 0004
ICASSP (1)6
2005 Development of the CUHTK 2004 Mandarin Conversational Telephone Speech Transcription System
abstract
The paper details all aspects of the CUHTK 2004 Mandarin conversational telephone speech transcription system, but concentrates on the development of the acoustic models. As there are significant differences between the available training corpora, both in terms of topics of conversation and accents, forms of data normalisation and adaptive training techniques are investigated. The baseline discriminatively trained acoustic models are compared to a system built with a Gaussianisation front-end, a speaker adaptively trained system and an adaptively trained structured precision matrix system. The models are finally evaluated within a multi-pass, multi-branch, system combination framework.
Mark J. F. Gales, Xunying Liu, Khe Chai Sim, Philip C. Woodland, Kai Yu 0004
ICASSP (1)5
2005 Development of the CU-HTK 2004 Broadcast News Transcription Systems
abstract
The paper describes our recent work on improving broadcast news transcription and presents details of the CU-HTK broadcast news English (BN-E) transcription system for the DARPA/NIST rich transcription 2004 speech-to-text (RT04) evaluation. A key focus has been building a system using an order of magnitude more acoustic training data than we have previously attempted. We have also investigated a range of techniques to improve both minimum phone error (MPE) training and the efficient creation of MPE-based narrow-band models. The paper describes two alternative system structures that run in under 10/spl times/RT and a further system that runs in less than 1/spl times/RT. This final system gives lower word error rates than our 2003 system that ran in 10/spl times/RT.
Do Yeong Kim, Ricky Ho Yin Chan, Gunnar Evermann, Mark J. F. Gales, David Mrva, Khe Chai Sim, Philip C. Woodland
ICASSP (1)7
2005 Structural metadata research in the EARS program
abstract
Both human and automatic processing of speech require recognition of more than just words. In this paper we provide a brief overview of research on structural metadata extraction in the DARPA EARS rich transcription program. Tasks include detection of sentence boundaries, filler words, and disfluencies. Modeling approaches combine lexical, prosodic, and syntactic information, using various modeling techniques for knowledge source integration. The performance of these methods is evaluated by task, by data source (broadcast news versus spontaneous telephone conversations) and by whether transcriptions come from humans or from an (errorful) automatic speech recognizer. A representative sample of results shows that combining multiple knowledge sources (words, prosody, syntactic information) is helpful, that prosody is more helpful for news speech than for conversational speech, that word errors significantly impact performance, and that discriminative models generally provide benefit over maximum likelihood models. Important remaining issues, both technical and programmatic, are also discussed.
Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Barbara Peskin, Jeremy Ang, Dustin Hillard, Mari Ostendorf, Marcus Tomalin, Philip C. Woodland, Mary P. Harper
ICASSP (5)9
2005 The Cambridge University March 2005 speaker diarisation system
abstract
This paper describes the speaker diarisation system developed at Cambridge University in March 2005. This system combines techniques used successfully in our previous speaker diarisation systems with an additional second clustering stage based on state-of-the-art speaker identification methods. Several strategies for using the new system are investigated and the final system gives a diarisation error rate of 6.9 % on the RT-04 Fall diarisation evaluation data when processing all the test data together or 8.6 % when processing the test data shows independently.
Rohit Sinha 0003, Sue Tranter, Mark J. F. Gales, Philip C. Woodland
INTERSPEECH4
2005 Automatic transcription of conversational telephone speech
abstract
This paper discusses the Cambridge University HTK (CU-HTK) system for the automatic transcription of conversational telephone speech. A detailed discussion of the most important techniques in front-end processing, acoustic modeling and model training, language and pronunciation modeling are presented. These include the use of conversation side based cepstral normalization, vocal tract length normalization, heteroscedastic linear discriminant analysis for feature projection, minimum phone error training and speaker adaptive training, lattice-based model adaptation, confusion network based decoding and confidence score estimation, pronunciation selection, language model interpolation, and class based language models. The transcription system developed for participation in the 2002 NIST Rich Transcription evaluations of English conversational telephone speech data is presented in detail. In this evaluation the CU-HTK system gave an overall word error rate of 23.9%, which was the best performance by a statistically significant margin. Further details on the derivation of faster systems with moderate performance degradation are discussed in the context of the 2002 CU-HTK 10 /spl times/ RT conversational speech transcription system.
Thomas Hain, Philip C. Woodland, Gunnar Evermann, Mark J. F. Gales, Xunying Liu, Gareth L. Moore, Daniel Povey
IEEE Trans. Speech Audio Process.2
2004 Improving broadcast news transcription by lightly supervised discriminative training
abstract
We present our experiments on lightly supervised discriminative training with large amounts of broadcast news data for which only closed caption transcriptions are available (TDT data). In particular, we use language models biased to the closed-caption transcripts to recognise the audio data, and the recognised transcripts are then used as the training transcriptions for acoustic model training. A range of experiments that use maximum likelihood (ML) training as well as discriminative training based on either maximum mutual information (MMI) or minimum phone error (MPE) are presented. In a 5xRT broadcast news transcription system that includes adaptation, it is shown that reductions in word error rate (WER) in the range of 1% absolute can be achieved. Finally, some experiments on training data selection are presented to compare different methods of "filtering" the transcripts.
Ricky Ho Yin Chan, Philip C. Woodland
ICASSP (1)2
2004 Development of the 2003 CU-HTK conversational telephone speech transcription system
abstract
The paper describes the development of the 2003 CU-HTK large vocabulary speech recognition system for conversational telephone speech (CTS). The system was designed based on a multipass, multibranch structure where the output of all branches is combined using system combination. A number of advanced modelling techniques, such as speaker adaptive training, heteroscedastic linear discriminant analysis, minimum phone error estimation and specially constructed single pronunciation dictionaries, were employed. The effectiveness of each of these techniques and their potential contribution to the result of system combination was evaluated in the framework of a state-of-the-art LVCSR system with sophisticated adaptation. The final 2003 CU-HTK CTS system constructed from some of these models is described and its performance on the DARPA/NIST 2003 rich transcription (RT-03) evaluation test set is discussed.
Gunnar Evermann, Ricky Ho Yin Chan, Mark J. F. Gales, Thomas Hain, Xunying Liu, David Mrva, Philip C. Woodland
ICASSP (1)8
2004 Generating and evaluating segmentations for automatic speech recognition of conversational telephone speech
abstract
Speech recognition systems for conversational telephone speech require the audio data to be automatically divided into regions of speech and non-speech. The quality of this audio segmentation affects the recognition accuracy. This paper describes several approaches to segmentation and compares the resulting recogniser performance. It is shown that using Gaussian mixture models outperforms an energy-detection method and using the output from the speech recogniser itself increases performance further. An upper bound on possible performance was obtained when deriving a segmentation from a forced alignment of the reference words and this outperformed using manually marked word times. Finally the correlation between an appropriately defined segmentation score and WER is shown to be over 0.95 across three data sets, suggesting that segmentations can be evaluated directly without the need for full decoding runs.
Sue Tranter, Kai Yu 0004, Gunnar Evermann, Philip C. Woodland
ICASSP (1)4
2004 MPE-based discriminative linear transform for speaker adaptation
abstract
We present a discriminative method for speaker adaptation, where the minimum phone error (MPE) criterion is used to estimate the discriminative linear transforms (DLTs), including both mean and diagonal variance transforms. The I-smoothing technique is essential to improve the generalization of DLTs. Experiments on supervised adaptation for non-native speakers on the North American Business (NAB) Spoke 3 task show that MPE-based DLT outperforms both MLLR and a previously proposed discriminative method for transform estimation. Preliminary experiments on unsupervised DLT estimation are also reported for conversational telephone speech transcription.
Philip C. Woodland
ICASSP (1)2
2004 Using VTLN for broadcast news transcription
abstract
Vocal tract length normalisation (VTLN) is a commonly used speaker normalisation approach. It is attractive compared to many normalisation schemes as it is typically dependent on only a single parameter, allowing the warp factors to be robustly calculated on little data. However, the scheme normally requires explicitly coding the data at multiple warp factors. Furthermore, it is only possible to approximate the Jacobian associated with the VTLN transformation. A new, simple, linear approximation to VTLN is described in this paper. This linear approximation allows the Jacobian to be exactly computed. It can also be highly efficient in terms of warp factor estimation and application of the warp factors. Both the linear and standard CUED VTLN schemes were evaluated in the 2003 BNE evaluation framework and found to yield similar performance. When used in system combination both VTLN schemes yielded slight gains over the baseline system.
Do Yeong Kim, Srinivasan Umesh, Mark J. F. Gales, Thomas Hain, Philip C. Woodland
INTERSPEECH5
2004 A PLSA-based language model for conversational telephone speech
abstract
This paper describes experiments with a PLSA-based language model for conversational telephone speech. This model uses a long-range history and exploits topic information in the test text to adjust probabilities of test words. The PLSA-based model was found to lower test set perplexity over a traditional word+class-based 4-gram by 13% (optimistic estimate using a reference transcript as history) or by 6% (realistic estimate using recognised transcript as history). Moreover, this paper introduces a use of confidence scores to weight words in the history, a weight of the prior topic distribution and a way of calculating perplexity that accounts for recognition errors in the model context.
David Mrva, Philip C. Woodland
INTERSPEECH2
2004 Automatic capitalisation generation for speech input
Philip C. Woodland
Comput. Speech Lang.2
2003 Porting: SwitchBoard to the VoiceMail task
abstract
The paper examines techniques that allow a well-trained source system built on one task to be rapidly adapted, or ported, to another target task. The two tasks considered are Hub5, or SwitchBoard, as the source system and VoiceMail as the target task. The two tasks are acoustically similar, both being telephone-bandwidth speech tasks, but differ in speaking style. SwitchBoard is conversational speech, VoiceMail is a set of voicemail messages. Various porting schemes for acoustic models are examined, including discriminative MAP and heteroscedastic LDA. Using around 28 hours of data, the error rate on VoiceMail was reduced by 42% relative compared to the baseline SwitchBoard performance.
Mark J. F. Gales, Daniel Povey, Philip C. Woodland
ICASSP (1)4
2003 Automatic complexity control for HLDA systems
abstract
Designing a state-of-the-art large vocabulary speech recognition systems is a highly complex problem. A wide range of techniques are available that affect the performance and number of free parameters. Selecting the appropriate complexity of system is both time-consuming and only a limited number of possible systems can be examined. This paper presents initial results on automatic system selection when both the number of dimensions and the number of components vary. Various complexity control schemes are discussed and evaluated. Limitations of schemes based on predicting held-out data log-likelihoods are described. In addition, problems of standard approximations for this task are detailed.
Xunying Liu, Mark J. F. Gales, Philip C. Woodland
ICASSP (1)3
2003 Discriminative map for acoustic model adaptation
abstract
In this paper we show how a discriminative objective function such as Maximum Mutual Information (MMI) can be combined with a prior distribution over the HMM parameters to give a discriminative Maximum A Posteriori (MAP) estimate for HMM training. The prior distribution can be based around the Maximum Likelihood (ML) parameter estimates, leading to a technique previously referred to as I-smoothing; or for adaptation it can be based around a MAP estimate of the ML parameters, leading to what we call MMI-MAP. This latter approach is shown to be effective for task adaptation, where data from one task (Voicemail) is used to adapt a HMM set trained on another task (Switchboard). It is shown that MMI-MAP results in a 2.1% absolute reduction in word error rate relative to standard ML-MAP with 30 hours of Voicemail task adaptation data starting from a MMI-trained Switchboard system.
Daniel Povey, Philip C. Woodland, Mark J. F. Gales
ICASSP (1)2
2003 MMI-MAP and MPE-MAP for acoustic model adaptation
abstract
This paper investigates the use of discriminative schemes based on themaximum mutual information (MMI) and minimum phone error (MPE) objective functions for both task and gender adaptation. A method for incorporating prior information into the discriminative training framework is described. If an appropriate form of prior distribution is used, then this may be implemented by simply altering the values of the counts used for parameter estimation. The prior distribution can be based around maximum likelihood parameter estimates, giving a technique known as I-smoothing, or for adaptation it can be based around a MAP estimate of the ML parameters, leading to MMI-MAP, or MPE-MAP.MMI-MAP isshown tobe effectivefor taskadaptation, where data from one task (Voicemail) is used to adapt a HMM set trained on another task (Switchboard). MPE-MAP is shown to be effective for generating gender-dependent models for Broadcast News transcription.
Daniel Povey, Mark J. F. Gales, Do Yeong Kim, Philip C. Woodland
INTERSPEECH4
2003 Language modelling for Russian and English using words and classes
Edward W. D. Whittaker, Philip C. Woodland
Comput. Speech Lang.2
2003 Erratum: Language modelling for Russian and English using words and classes [Computer Speech and Language 17 (2003) 87-104]
Edward W. D. Whittaker, Philip C. Woodland
Comput. Speech Lang.2
2003 A combined punctuation generation and speech recognition system and its performance enhancement using prosody
Philip C. Woodland
Speech Commun.2
2002 Improved cross-task recognition using MMIE training
abstract
This paper investigates the cross-task recognition and adaptation performance of HMMs trained using either conventional maximum likelihood estimation or the discriminative maximum mutual information estimation (MMIE) criterion. Initial experiments used models trained on the low noise North American Business news corpus of read speech. Cross-task testing on Broadcast News data showed that the MMIE models yielded lower error rates both across-task as well as within-task. This result was confirmed using models trained on the Switchboard corpus which were tested on Voicemail (VM)data. This setup was also used to investigate the performance of task-adaptation when using a limited amount of VM data for both acoustic and language modelling. The setup that gave the best performance on the VM test data used Switchboard models trained using MMIE and then adapted to VM data using maximum a posteriori adaptation techniques.
Ricardo de Córdoba, Philip C. Woodland, Mark J. F. Gales
ICASSP2
2002 Implementation of automatic capitalisation generation systems for speech input
abstract
In this paper, two systems are proposed for the task of capitalisation generation. The first system is a slightly modified speech recogniser. In this system, every word in the vocabulary is duplicated: once in a decapitalised form and again in capitalised forms. In addition, the language model is re-trained on mixed case texts. The other system is based on Named Entity (NE) recognition and punctuation generation, since most capitalised words are first words in sentences or NE words. Both systems are compared for speech input. The system based on NE recognition and punctuation generation shows better results in Word Error Rate (WER) and in F-measure than the system modified from the speech recogniser.
Philip C. Woodland
ICASSP2
2002 Minimum Phone Error and I-smoothing for improved discriminative training
abstract
In this paper we introduce the Minimum Phone Error (MPE) and Minimum Word Error (MWE) criteria for the discriminative training of HMM systems. The MPE/MWE criteria are smoothed approximations to the phone or word error rate respectively. We also discuss I-smoothing which is a novel technique for smoothing discriminative training criteria using statistics for maximum likelihood estimation (MLE). Experiments have been performed on the Switchboard/Call Home corpora of telephone conversations with up to 265 hours of training data. It is shown that for the maximum mutual information estimation (MMIE) criterion, I-smoothing reduces the word error rate (WER) by 0.4% absolute over the MMIE baseline. The combination of MPE and I-smoothing gives an improvement of 1 % over MMIE and a total reduction in WER of 4.8% absolute over the original MLE system.
Daniel Povey, Philip C. Woodland
ICASSP2
2002 Maximum mutual information training of hidden Markov models with vector linear predictors
K. K. Chin, Philip C. Woodland
INTERSPEECH2
2002 Cluster identification for speaker-environment tracking
J. T. Wickramaratna, Philip C. Woodland
INTERSPEECH2
2002 Large scale discriminative training of hidden Markov models for speech recognition
Philip C. Woodland, Daniel Povey
Comput. Speech Lang.1
2002 The development of the HTK Broadcast News transcription system: An overview
Philip C. Woodland
Speech Commun.1
2001 New features in the CU-HTK system for transcription of conversational telephone speech
abstract
Discusses new features integrated into the Cambridge University HTK (CU-HTK) system for the transcription of conversational telephone speech. Major improvements have been achieved by the use of maximum mutual information estimation in training as well as maximum likelihood estimation; the use of a full variance transform for adaptation; the inclusion of unigram pronunciation probabilities; and word-level posterior probability estimation using confusion networks for use in minimum word error rate decoding, confidence score estimation and system combination. Improvements are demonstrated via performance on the NIST March 2000 evaluation of English conversational telephone speech transcription (Hub5E). In this evaluation the CU-HTK system gave an overall word error rate of 25.4%, which was the best performance by a statistically significant margin.
Thomas Hain, Philip C. Woodland, Gunnar Evermann, Daniel Povey
ICASSP2
2001 Improved discriminative training techniques for large vocabulary continuous speech recognition
abstract
Investigates the use of discriminative training techniques for large vocabulary speech recognition with training datasets up to 265 hours. Techniques for improving lattice-based maximum mutual information estimation (MMIE) training are described and compared to frame discrimination (FD). An objective function which is an interpolation of MMIE and standard maximum likelihood estimation (MLE) is also discussed. Experimental results on both the Switchboard and North American Business News tasks show that MMIE training can yield significant performance improvements over standard MLE even for the most complex speech recognition problems with very large training sets.
Daniel Povey, Philip C. Woodland
ICASSP2
2001 Improvements in linear transform based speaker adaptation
abstract
Presents three forms of linear transform based speaker adaptation that can give better performance than standard maximum likelihood linear regression (MLLR) adaptation. For unsupervised adaptation, a lattice-based technique is introduced which is compared to MLLR using confidence scores. For supervised adaptation, estimation of the adaptation matrices using the maximum mutual information criterion is discussed which leads to the MMILR approach. Recognition experiments show that lattice MLLR can reduce word error rates on a Switchboard task by 1.4% absolute. For recognition of non-native speech from the Wall Street Journal database, a reduction in word error rate of 10-16% relative was obtained using MMILR compared to standard MLLR.
Luís Felipe Uebel, Philip C. Woodland
ICASSP2
2001 Efficient class-based language modelling for very large vocabularies
abstract
Investigates the perplexity and word error rate performance of two different forms of class model and the respective data-driven algorithms for obtaining automatic word classifications. The computational complexity of the algorithm for the 'conventional' two-sided class model is found to be unsuitable for very large vocabularies (>100k) or large numbers of classes (>2000). A one-sided class model is therefore investigated and the complexity of its algorithm is found to be substantially less in such situations. Perplexity results are reported on both English and Russian data. For the latter both 65k and 430k vocabularies are used. Lattice rescoring experiments are also performed on an English language broadcast news task. These experimental results show that both models, when interpolated with a word model, perform similarly well. Moreover, classifications are obtained for the one-sided model in a fraction of the time required by the two-sided model, especially for very large vocabularies.
Edward W. D. Whittaker, Philip C. Woodland
ICASSP2
2001 The use of prosody in a combined system for punctuation generation and speech recognition
abstract
In this paper, we discuss a combined system for punctuation generation and speech recognition. This system incorporates prosodic information with acoustic and language model information. Experiments are conducted for both the reference transcriptions and speech recogniser outputs. For the reference transcription case, prosodic information is shown to be more useful than language model information. When these information sources are combined, we can obtain an F-measure of up to 0.7830 for punctuation recognition. A few straightforward...
Philip C. Woodland
INTERSPEECH2
2000 Large vocabulary decoding and confidence estimation using word posterior probabilities
abstract
The paper investigates the estimation of word posterior probabilities based on word lattices and presents applications of these posteriors in a large vocabulary speech recognition system. A novel approach to integrating these word posterior probability distributions into a conventional Viterbi decoder is presented. The problem of the robust estimation of confidence scores from word posteriors is examined and a method based on decision trees is suggested. The effectiveness of these techniques is demonstrated on the broadcast news and the conversational telephone speech corpora where improvements both in terms of word error rate and normalised cross entropy were achieved compared to the baseline HTK evaluation systems.
Gunnar Evermann, Philip C. Woodland
ICASSP2
2000 A method for direct audio search with applications to indexing and retrieval
abstract
A technique for searching audio data to find an exact match for a given piece of cue-audio is described. The method uses a cepstral parameterisation of the audio and a covariance-based distance metric to quickly locate direct repeats. Results on data from ABC news broadcasts show that the method can successfully locate matches several hundred times faster than real-time and requires less than a second of cue-audio. By applying the match recursively to the data, repeated sections of audio, which nearly always correspond to non-news items such as commercials and theme-music, can be identified. Experiments show that the application of the technique can also lead to improved information retrieval using automatically transcribed broadcast data.
Sue Tranter, Philip C. Woodland
ICASSP2
2000 Modelling sub-phone insertions and deletions in continuous speech recognition
Thomas Hain, Philip C. Woodland
INTERSPEECH2
2000 A rule-based named entity recognition system for speech input
abstract
In this paper, we propose a rule based (transformation based) named entity recognition system which uses the Brill rule inference approach. To measure its performance, we compare the performance of the rule-based system and IdentiFinder, one of the most successful stochastic systems. In the baseline case (no punctuation and no capitalisation), both systems show almost equal performance. They also have similar performance in the case of additional information such as punctuation, capitalisation and name lists. The performance of both systems degrade linearly with added speech recognition errors, and their rates of degradation are almost equal. These results show that automatic rule inference is a viable alternative to the HMM-based approach to named entity recognition, but it retains the advantages of a rule-based approach.
Philip C. Woodland
INTERSPEECH2
2000 Particle-based language modelling
abstract
This paper investigates the use of particle (sub-word) N-grams for language modelling. One linguistics-based and two datadriven algorithms are presented and evaluated in terms of perplexity for Russian and English. Interpolating word trigram and particle 6-gram models gives up to a 7.5% perplexity reduction over the baseline word trigram model for Russian. Lattice rescoring experiments are also performed on 1997 DARPA Hub4 evaluation lattices where the interpolated model gives a 0.4% absolute reduction in word error rate over the baseline word trigram model. 1. INTRODUCTION Most of the current approaches to language modelling for speech recognition tend to use words, or classes of words, as the modelling units. Words are a logical choice, since it is ultimately words that are to be output by a speech recognition system, but they are not necessarily the best units for capturing dependencies in a text. The optimal set of units will inevitably depend on the language, the sparsity of the ...
Edward W. D. Whittaker, Philip C. Woodland
INTERSPEECH2
2000 The Cambridge University multimedia document retrieval demo system
abstract
No abstract available.
Andreas Tuerk, Sue Tranter, Pierre Jourlin, Karen Spärck Jones, Philip C. Woodland
SIGIR5
2000 Effects of out of vocabulary words in spoken document retrieval
abstract
The effects of out-of-vocabulary (OOV) items in spoken document retrieval (SDR) are investigated. Several sets of transcriptions were created for the TREC-8 SDR task using a speech recognition system varying the vocabulary sizes and OOV rates, and the relative retrieval performance measured. The effects of OOV terms on a simple baseline IR system and on more sophisticated retrieval systems are described. The use of a parallel corpus for query and document expansion is found to be especially beneficial, and with this data set, good retrieval performance can be achieved even for fairly high OOV rates.
Philip C. Woodland, Sue Tranter, Pierre Jourlin, Karen Spärck Jones
SIGIR1
2000 Spoken document representations for probabilistic retrieval
Pierre Jourlin, Sue Tranter, Karen Spärck Jones, Philip C. Woodland
Speech Commun.4
1999 The 1998 HTK system for transcription of conversational telephone speech
abstract
This paper describes the 1998 HTK large vocabulary speech recognition system for conversational telephone speech as used in the NIST 1998 Hub5E evaluation. Front-end and language modelling experiments conducted using various training and test sets from both the Switchboard and Callhome English corpora are presented. Our complete system includes reduced bandwidth analysis, side-based cepstral feature normalisation, vocal tract length normalisation (VTLN), triphone and quinphone hidden Markov models (HMMs) built using speaker adaptive training (SAT), maximum likelihood linear regression (MLLR) speaker adaptation and a confidence score based system combination. A detailed description of the complete system together with experimental results for each stage of our multi-pass decoding scheme is presented. The word error rate obtained is almost 20% better than our 1997 system on the development set.
Thomas Hain, Philip C. Woodland, Thomas Niesler, Edward W. D. Whittaker
ICASSP2
1999 The Cambridge University spoken document retrieval system
abstract
This paper describes the spoken document retrieval system that we have been developing and assesses its performance using automatic transcriptions of about 50 hours of broadcast news data. The recognition engine is based on the HTK broadcast news transcription system and the retrieval engine is based on the techniques developed at City University. The retrieval performance over a wide range of speech transcription error rates is presented and a number of recognition error metrics that more accurately reflect the impact of transcription errors on retrieval accuracy are defined and computed. The results demonstrate the importance of high accuracy automatic transcription. The final system is currently being evaluated on the 1998 TREC-7 spoken document retrieval task.
Sue Tranter, Pierre Jourlin, Gareth L. Moore, Karen Spärck Jones, Philip C. Woodland
ICASSP5
1999 Frame discrimination training for HMMs for large vocabulary speech recognition
abstract
This paper describes the application of a discriminative HMM parameter estimation technique called frame discrimination (FD), to medium and large vocabulary continuous speech recognition. Previous work has shown that FD training can give better results than maximum mutual information (MMI) training for small tasks. The use of FD for much larger tasks required the development of a technique to be able to rapidly find the most likely set of Gaussians for each frame in the system. Experiments on the resource management and North American business tasks show that FD training can give comparable improvements to MMI, but is less computationally intensive.
Daniel Povey, Philip C. Woodland
ICASSP2
1999 Dynamic HMM selection for continuous speech recognition
abstract
In this paper we propose a dynamic model selection technique based on hidden model sequences (HMS). HMS modelling assumes, that not only the actual state sequence is unknown, but also the model sequence given a particular sentence. This allows more than one model to be used for a particular phone in a certain context. The most appropriate model is determined locally rather than a priori globally by the acoustic probability of that model together with a probability that this model is produced in a particular phone (or model) context. Experiments on the Resource Management corpus show significant improvements in word error rate over phonetically model-- and state--tied triphone hidden Markov models (HMMs). Initial results on the Switchboard corpus also show improvements on a much more difficult task. 1. INTRODUCTION In HMM-based continuous speech recognition ideally each possible sentence would be modelled with a separate Markov model. Since the set of possible sentences is far too larg...
Thomas Hain, Philip C. Woodland
EUROSPEECH2
1999 An investigation into vocal tract length normalisation
Luís Felipe Uebel, Philip C. Woodland
EUROSPEECH2
1999 Improvements in accuracy and speed in the HTK broadcast news transcription system
Philip C. Woodland, J. J. Odell, Thomas Hain, Gareth L. Moore, Thomas Niesler, Andreas Tuerk, Edward W. D. Whittaker
EUROSPEECH1
1999 Improving Retrieval on Imperfect Speech Transcriptions (poster abstract)
abstract
No abstract available.
Pierre Jourlin, Sue Tranter, Karen Spärck Jones, Philip C. Woodland
SIGIR4
1999 A hidden Markov-model-based trainable speech synthesizer
Robert E. Donovan, Philip C. Woodland
Comput. Speech Lang.2
1999 Variable-length categoryn-gram language models
Thomas Niesler, Philip C. Woodland
Comput. Speech Lang.2
1998 The use of accent-specific pronunciation dictionaries in acoustic model training
abstract
Speech recognition systems are increasingly being built to cover an ever wider range of speaker accents. However, electronically available pronunciation dictionaries (PDs) specific to these accents often do not exist and would be time consuming and expensive to build by hand. This paper explores the use of pronunciation modelling for the synthesis of accent-specific PDs directly from acoustic data, and their use in acoustic model training. It is shown that this is particularly effective when the amount of acoustic data from the new accent region is insufficient to build a new recogniser, and it is necessary to retrain an existing system: a further 15% reduction in word error rate can be achieved over and above the 20% reduction resulting from acoustic model retraining alone. This paper also presents an empirical evaluation of an American English PD which has been synthesised from a British English PD.
Jason J. Humphries, Philip C. Woodland
ICASSP2
1998 Comparison of part-of-speech and automatically derived category-based language models for speech recognition
abstract
This paper compares various category-based language models when used in conjunction with a word-based trigram by means of linear interpolation. Categories corresponding to parts-of-speech as well as automatically clustered groupings are considered. The category-based model employs variable-length n-grams and permits each word to belong to multiple categories. Relative word error rate reductions of between 2 and 7% over the baseline are achieved in N-best rescoring experiments on the Wall Street Journal corpus. The largest improvement is obtained with a model using automatically determined categories. Perplexities continue to decrease as the number of different categories is increased, but improvements in the word error rate reach an optimum.
Thomas Niesler, Edward W. D. Whittaker, Philip C. Woodland
ICASSP3
1998 Experiments in broadcast news transcription
abstract
This paper presents the development of the HTK broadcast news transcription system. Previously we have used data type specific modelling based on adapted Wall Street Journal trained HMMs. However, we are now experimenting with data for which no manual pre-classification or segmentation is available and therefore automatic techniques are required and compatible acoustic modelling strategies adopted. An approach for automatic audio segmentation and classification is described and evaluated as well as extensions to our previous work on segment clustering. A number of recognition experiments are presented that compare datatype specific and non-specific models; differing amounts of training data; the use of gender-dependent modelling and the effects of automatic data-type classification. It is shown that robust segmentation into a small number of audio types is possible and that models trained on a wide variety of data types can yield good performance.
Philip C. Woodland, Thomas Hain, Sue Tranter, Thomas Niesler, Andreas Tuerk, Steve J. Young
ICASSP1
1998 Segmentation and classification of broadcast news audio
abstract
Broadcast news contains a wide variety of different speakers and audio conditions (channel and background noise). This paper describes a segmentation, gender detection and audio classification scheme and presents experimental results on the DARPA 1997 broadcast news evaluation set.
Thomas Hain, Philip C. Woodland
ICSLP2
1998 Speaker clustering using direct maximisation of the MLLR-adapted likelihood
abstract
In this paper speaker clustering schemes are investigated in the context of improving unsupervised adaptation for broadcast news transcription. The various techniques are presented within a framework of top-down split-and-merge clustering. Since these schemes are to be used for MLLRbased adaptation, a natural evaluation metric for clustering is the increase in data likelihood from adaptation. Two types of cluster splitting criteria have been used. The first minimises a covariance-based distance measure and for the second we introduce a two-step E-M type procedure to form clusters which directly maximise the likelihood of the adapted data. It is shown that the direct maximisation technique produces a higher data likelihood and also gives a reduction in word error rate.
Sue Tranter, Philip C. Woodland
ICSLP2
1998 Comparison of language modelling techniques for Russian and English
abstract
In this paper the main differences between language modelling of Russian and English are examined. A Russian corpus and a comparable English corpus are described. The effects of high inflectionality in Russian and the relationship between the outof -vocabulary rate and vocabulary size are investigated. Standard word and class N-gram language modelling techniques are applied to the two corpora and perplexity results are reported. A novel approach to the modelling of inflected languages is proposed and its efficacy compared with the other techniques. 1. INTRODUCTION Much work has been conducted in recent years on language modelling techniques for speech recognition of English. In contrast, less commercially attractive yet widely spoken languages like Russian have received comparatively little attention in the literature (the first reported large-vocabulary recogniser for Russian appeared only recently[3]). Moreover, there are important difficulties with modelling Russian which are also...
Edward W. D. Whittaker, Philip C. Woodland
ICSLP2
1997 Modelling word-pair relations in a category-based language model
abstract
A new technique for modelling word occurrence correlations within a word-category based language model is presented. Empirical observations indicate that the conditional probability of a word given its category, rather than maintaining the constant value normally assumed, exhibits an exponential decay towards a constant as a function of an appropriately defined measure of separation between the correlated words. Consequently, a functional dependence of the probability upon this separation is postulated, and methods for determining both the related word pairs as well as the function parameters are developed. Experiments using the LOB, Switchboard and Wall Street Journal corpora indicate that this formulation captures the transient nature of the conditional probability effectively, and leads to reductions in perplexity of between 8 and 22%, where the largest improvements are delivered by correlations of words with themselves (self-triggers), and the reductions increase with the size of the training corpus.
Thomas Niesler, Philip C. Woodland
ICASSP2
1997 Experiments in speaker normalisation and adaptation for large vocabulary speech recognition
abstract
This paper examines techniques for speaker normalisation and adaptation that are applied in training with the aim of removing some of the variability from the speaker independent models. Two techniques are examined: vocal tract normalisation (VTN) which estimates a single "vocal tract length" parameter for each speaker and then modifies the speech parameterisation accordingly and speaker adaptive training (SAT) which estimates Gaussian mean and variance parameters jointly with a speaker specific set of maximum likelihood linear regression (MLLR) based transformations. It is shown that VTN is effective for both clean speech and mismatched conditions and that the further improvements obtained by applying MLLR in testing are essentially additive. Detailed results from the use of SAT show that worthwhile improvements over using MLLR with standard speaker independent models are obtained.
David Pye, Philip C. Woodland
ICASSP2
1997 Broadcast news transcription using HTK
abstract
This paper examines the issues in extending a large vocabulary speech recognition system designed for clean and noisy read speech tasks to handle broadcast news transcription. Results using the 1995 DARPA H4 evaluation data set are presented for different front-end analyses and use of unsupervised model adaptation using maximum likelihood linear regression (MLLR). The HTK system for the 1996 H4 evaluation is then described. It includes a number of new features over previous HTK large vocabulary systems including decoder-guided segmentation, segment clustering, cache-based language modelling, and combined MAP and MLLR adaptation. The system runs in multiple passes through the data and the detailed results of each pass are given.
Philip C. Woodland, Mark J. F. Gales, David Pye, Steve J. Young
ICASSP1
1997 Using accent-specific pronunciation modelling for improved large vocabulary continuous speech recognition
abstract
A method of modelling accent-specific pronunciation variations is presented. Speech from an unseen accent group is phonetically transcribed such that pronunciation variations may be derived. These context-dependent variations are clustered in decision trees which are used as a model of the pronunciation variation associated with this new accent group. The trees are then used to build a new pronunciation dictionary for use during the recognition process. Experiments are presented, based on Wall Street Journal and WSJCAM0 corpora, for the recognition of American speakers using a British English recogniser. Speaker independent as well as speaker dependent adaptation scenarios are presented, giving up to 20% reduction in word error rate. A linguistic analysis of the pronunciation model is presented and finally the technique is combined with maximum likelihood linear regression, a well proven acoustic adaptation technique, yielding further improvement.
Jason J. Humphries, Philip C. Woodland
EUROSPEECH2
1997 Combined Bayesian and predictive techniques for rapid speaker adaptation of continuous density hidden Markov models
S. M. Ahadi, Philip C. Woodland
Comput. Speech Lang.2
1997 Multilingual large vocabulary speech recognition: the European SQALE project
Steve J. Young, Martine Adda-Decker, Xavier L. Aubert, Christian Dugast, Jean-Luc Gauvain, Dan J. Kershaw, Lori Lamel, David A. van Leeuwen, David Pye, Anthony J. Robinson, Herman J. M. Steeneken, Philip C. Woodland
Comput. Speech Lang.12
1997 MMIE training of large vocabulary recognition systems
V. Valtchev, J. J. Odell, Philip C. Woodland, Steve J. Young
Speech Commun.3
1996 A variable-length category-based n-gram language model
abstract
A language model based on word-category n-grams and ambiguous category membership with n increased selectively to trade compactness for performance is presented. The use of categories leads intrinsically to a compact model with the ability to generalise to unseen word sequences, and diminishes the sparseness of the training data, thereby making larger n feasible. The language model implicitly involves a statistical tagging operation, which may be used explicitly to assign category assignments to untagged text. Experiments on the LOB corpus show the optimal model-building strategy to yield improved results with respect to conventional n-gram methods, and when used as a tagger, the model is seen to perform well in relation to a standard benchmark.
Thomas Niesler, Philip C. Woodland
ICASSP2
1996 Lattice-based discriminative training for large vocabulary speech recognition
abstract
This paper describes a framework for optimising the parameters of a continuous density HMM-based large vocabulary recognition system using a maximum mutual information estimation (MMIE) criterion. To limit the computational complexity arising from the need to find confusable speech segments in the large search space of alternative utterance hypotheses, word lattices generated from the training data are used. Experiments are presented on the Wall Street journal database using up to 66 hours of training data. These show that lattices combined with an improved estimation algorithm makes MMIE training practicable even for very complex recognition systems and large training sets. Furthermore, experimental results show that MMIE training can yield useful increases in recognition accuracy.
V. Valtchev, J. J. Odell, Philip C. Woodland, Steve J. Young
ICASSP3
1996 Improving environmental robustness in large vocabulary speech recognition
abstract
This paper describes techniques to improve the robustness of the HTK large vocabulary speech recognition system to non-ideal acoustic environments. The primary methods are single-pass retraining using stereo training data; parallel model combination which combines HMMs trained on clean data with estimates of convolutional and additive noise; and maximum likelihood linear regression which estimates a set of linear transformations of the model parameters to the current conditions. Experiments are reported on both the 1994 ARPA CSR S5 (alternate microphones) and S10 (additive noise) spoken tasks and the 1995 ARPA CSR H3 task (multiple unknown microphones). The HTK system yielded the lowest error rates in both the H3-P0 and HS-C0 tests.
Philip C. Woodland, Mark J. F. Gales, David Pye
ICASSP1
1996 Variance compensation within the MLLR framework for robust speech recognition and speaker adaptation
abstract
This paper investigates the use of maximum likelihood linear regression (MLLR) for both speaker and environment adaptation.MLLR transforms the mean and variance parameters of a set of HMMs.In this paper a number of different types of linear transformations of the variances are examined including full, block diagonal, and diagonal transformation matrices.Experiments on large vocabulary speaker independent data sets are described.On all the data sets examined the use of MLLR mean and variance compensation reduced the error rate compared to mean-only compensation.Furthermore, the use of a block diagonal or full transformation of the variances on the clean data task showed slight improvements over the diagonal case.However, when some environmental mismatch was present there was no difference in performance between using multiple diagonal variance transformations and a more complex single variance transform.
Mark J. F. Gales, David Pye, Philip C. Woodland
ICSLP3
1996 Using accent-specific pronunciation modelling for robust speech recognition
abstract
A method of modelling accent-specific pronunciation variations is presented.Speech from an unseen accent group is phonetically transcribed such that pronunciation variations may be derived.These context-dependent variations are clustered in a decision tree which is used as a model of the pronunciation variation associated with this new accent group.The tree is then used to build a new pronunciation dictionary for use during the recognition process.Experiments are presented for the recognition of Lancashire & Yorkshire accented speech using a recognizer trained on London & South East England speakers.The results show that the addition of accent-specific pronunciations can reduce the error rate by almost 20% for cross accent recognition.It is also shown that worthwhile gains in performance can be obtained using only a small amount of accent-specific data.
Jason J. Humphries, Philip C. Woodland, David J. B. Pearce
ICSLP2
1996 Combination of word-based and category-based language models
abstract
A language model combining word-based and category-based ngrams within a backoff framework is presented.Word n-grams conveniently capture sequential relations between particular words, while the category-model, which is based on part-of-speech classific ations and allows ambiguous category membership, is able to generalise to unseen word sequences and therefore appropriate in backoff situations.Experiments on the LOB, Switchboard and WSJ0 corpora demonstrate that the technique greatly improves language model perplexities for sparse training sets, and offers significantly improved complexity versus performance tradeoffs when compared with standard trigram models.
Thomas Niesler, Philip C. Woodland
ICSLP2
1996 Discriminative optimisation of large vocabulary recognition systems
abstract
This paper describes a framework for optimising the structure and parameters of a continuous density HMM-based large vocabulary recognition system using the Maximum Mutual Information Estimation (MMIE) criterion.To reduce the computational complexity of the MMIE training algorithm, confusable segments of speech are identified and stored as word lattices of alternative utterance hypotheses.An iterative mixture splitting procedure is also employed to adjust the number of mixture components in each state during training such that the optimal balance between number of parameters and available training data is achieved.Experiments are presented on various test sets from the Wall Street Journal database using the full SI-284 training set.These show that the use of lattices makes MMIE training practicable for very complex recognition systems and large training sets.Furthermore, experimental results demonstrate that MMIE optimisation of system structure and parameters can yield useful increases in recognition accuracy.
V. Valtchev, Philip C. Woodland, Steve J. Young
ICSLP2
1996 Iterative unsupervised adaptation using maximum likelihood linear regression
Philip C. Woodland, David Pye, Mark J. F. Gales
ICSLP1
1996 Mean and variance adaptation within the MLLR framework
Mark J. F. Gales, Philip C. Woodland
Comput. Speech Lang.2
1995 Rapid speaker adaptation using model prediction
abstract
A key issue in speaker adaptation is gaining the maximum information from a limited amount of adaptation data. In particular it is important that observations of parameters of (context-dependent) HMMs not occurring in the adaptation data can be updated. In the regression-based model prediction (RMP) approach, sets of speaker-independent linear relationships between different parameters in the HMM set are found from training data. During adaptation, distributions with sufficient adaptation data are used to update the parameters of poorly adapted models using these pre-computed regression-based relationships. The method used Bayesian techniques to combine parameter estimates from different sources. Evaluation on the ARPA Resource Management corpus gave a worthwhile reduction in error rate with just a single adaptation sentence, and that RMP consistently outperforms MAP estimation with the same amount of adaptation data.
S. M. Ahadi, Philip C. Woodland
ICASSP2
1995 Automatic speech synthesiser parameter estimation using HMMs
abstract
This paper presents a new approach to speech synthesis which uses a set of decision tree state clustered triphone HMMs to automatically segment a single speaker speech database into sub-word units suitable for use in a synthesiser. Parameters are then obtained for each of these sub-word units from the segmented database, enabling a basic synthesis system to be constructed. This automatic generation of synthesis parameters means that the system can easily be retrained on a new speaker, whose voice it then mimics. It also means that a very large number of sub-word units can be used, which enables more precise context modelling than was previously possible.
Robert E. Donovan, Philip C. Woodland
ICASSP2
1995 The 1994 HTK large vocabulary speech recognition system
abstract
This paper describes recent work on the HTK large vocabulary speech recognition system. The system uses tied-state cross-word context-dependent mixture Gaussian HMMs and a dynamic network decoder that can operate in a single pass. In the last year the decoder has been extended to produce word lattices to allow flexible and efficient system development, as well as multi-pass operation for use with computationally expensive acoustic and/or language models. The system vocabulary can now be up to 65 k words, the final acoustic models have been extended to be sensitive to more acoustic context (quinphones), a 4-gram language model has been used and unsupervised incremental speaker adaptation incorporated. The resulting system gave the lowest error rates on both the H1-P0 and H1-C1 hub tasks in the November 1994 ARPA CSR evaluation.
Philip C. Woodland, Chris Leggetter, J. J. Odell, V. Valtchev, Steve J. Young
ICASSP1
1995 Improvements in an HMM-based speech synthesiser
Robert E. Donovan, Philip C. Woodland
EUROSPEECH2
1995 Flexible speaker adaptation for large vocabulary speech recognition
Chris Leggetter, Philip C. Woodland
EUROSPEECH2
1995 Large vocabulary multilingual speech recognition using HTK
David Pye, Philip C. Woodland, Steve J. Young
EUROSPEECH2
1995 Maximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models
Chris Leggetter, Philip C. Woodland
Comput. Speech Lang.2
1994 Large vocabulary continuous speech recognition using HTK
abstract
HTK is a portable software toolkit for building speech recognition systems using continuous density hidden Markov models developed by the Cambridge University Speech Group. One particularly successful type of system uses mixture density tied-state triphones. We have used this technique for the 5 k/20 k word ARPA Wall Street Journal (WSJ) task. We have extended our approach from using word-internal gender independent modelling to use decision tree based state clustering, cross-word triphones and gender dependent models. Our current systems can be run with either bigram or trigram language models using a single pass dynamic network decoder. Systems based on these techniques were included in the November 1993 ARPA WSJ evaluation, and gave the lowest error rate reported on the 5 k word bigram, 5 k word trigram and 20 k word bigram "hub" tests and the second lowest error rate on the 20 k word trigram "hub" test.>
Philip C. Woodland, J. J. Odell, V. Valtchev, Steve J. Young
ICASSP (2)1
1994 Modelling syllable characteristics to improve a large vocabulary continuous speech recogniser
Philip C. Woodland
ICSLP2
1994 Speaker adaptation of continuous density HMMs using multivariate linear regression
Chris Leggetter, Philip C. Woodland
ICSLP2
1994 Recognition ********* a dynamic network decoder design for large vocabulary speech recognition
V. Valtchev, J. J. Odell, Philip C. Woodland, Steve J. Young
ICSLP3
1994 State clustering in hidden Markov model-based continuous speech recognition
Steve J. Young, Philip C. Woodland
Comput. Speech Lang.2
1994 Spontaneous speech recognition for the credit card corpus using the HTK toolkit
abstract
This paper describes a speech recognition system for the credit card corpus which was provided as a baseline for the Summer Workshop on Robust Speech Processing held at the Rutgers CAIP Center in July/August 1993. The system was built using a portable HMM toolkit called HTK. This is described, along with details of the training and testing methods used. Benchmark performance results are presented for mixture Gaussian monophones and tied-state triphones using four different language models. The overall accuracy levels achieved were very low, indicating that this is a very difficult recognition task. The paper concludes with a discussion of some specific problems encountered with the Credit Card data and suggestions for future work.>
Steve J. Young, Philip C. Woodland, William J. Byrne
IEEE Trans. Speech Audio Process.2
1993 A wave digital filter model of the entire auditory periphery
Christian Giguère, Philip C. Woodland
ICASSP (2)2
1993 Exploiting variable-width features in large vocabulary speech recognition
Philip C. Woodland
ICASSP (2)2
1993 Using relative duration in large vocabulary speech recognition
Philip C. Woodland
EUROSPEECH2
1993 Hidden Markov models using shared vector linear predictors
B. A. Maxwell, Philip C. Woodland
EUROSPEECH2
1993 The HTK tied-state continuous speech recogniser
abstract
HTK is a portable software toolkit for developing systems using continuous density hidden Markov models developed by the Cambridge University Speech Group. This paper describes speech recognition experiments using HTK based systems for the DARPA Resource Management (RM) task. In particular good performance is obtained using a tied-state triphone based multiple mixture approach. This system was used in the final DARPA RM evaluation (September 1992) and was found to perform at a similar level to the main DARPA systems, and yet be efficient in terms of the total of parameters and computational load. The results for that system are given along with some recent experiments that investigated the use of male-female modelling in a tied-state HMM system. Keywords: Hidden Markov Models, Resource Management, State Clustering, HTK. 1. INTRODUCTION This paper describes the use of HTK (HMM toolkit) in building speech recognisers for the DARPA Resource Management (RM) task. HTK is a software toolki...
Philip C. Woodland, Steve J. Young
EUROSPEECH1
1993 The use of state tying in continuous speech recognition
Steve J. Young, Philip C. Woodland
EUROSPEECH2
1992 Hidden Markov models using vector linear prediction and discriminative output distributions
abstract
HMMs model signal dynamics rather poorly. Modeling accuracy can be improved by adding vector linear predictors to each state in order to predict the value of the current observation based on correlations with nearby observations. A vector linear predictive HMM is discussed and re-estimation formulae for the predictor parameters presented. Multiple speaker recognition experiments on a 104 talker British English E-set database were performed to test the method on a difficult speech recognition task. It was found that a baseline test set error rate of 5.6% improved to 4.1% using a single diagonal predictor. Further improvements in performance along with a reduction in computation were obtained by using the method of discriminative output distributions on the prediction error. This resulted in a best test set error rate of 2.8% from a system that required only half the computation of the baseline.>
Philip C. Woodland
ICASSP1
1991 Optimising hidden Markov models using discriminative output distributions
abstract
Models similar to Doddington's (1989, 1990) hidden Markov models (HMMs) that use phonetically sensitive discriminants are discussed. In this style of HMM, each state models a subspace of the overall acoustic vector; the subspace is chosen to increase discrimination between the in-class and potentially confusable out-of-class utterances. The theoretical basis is presented and various aspects of using these models are discussed, such as the method of gathering confusion statistics; obtaining the correct normalization for the subspace Gaussian distribution and the effects of this term; and the computational requirements for the method. A large number of experiments on a 104 talker British English E-set database were performed that illustrate the utility of the method on a difficult speech recognition task. The experiments give a best speaker-independent error rate 7.9%, and a best multiple speaker error rate of 3.8%.>
Philip C. Woodland, David R. Cole
ICASSP1
1990 An experimental comparison of connectionist and conventional classification systems on natural data
Philip C. Woodland, S. G. Smyth
Speech Commun.1