VLDB 2026 Research / reviewers in the wild / expert
Steve Renals
dblp:33/3792 · also Stephen Renals
· DBLP profile ↗
202ranked-venue papers
20as first author
15since 2021 · last 2023
0000-0002-8790-3389ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 167 · 15 first-author · 12 since 2021Artificial intelligence and machine learning · 126 · 10 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Phonetic Error Analysis Beyond Phone Error RateabstractIn this paper, we analyse the performance of the TIMIT-based phone recognition systems beyond the overall phone error rate (PER) metric. We consider three broad phonetic classes (BPCs): {affricate, diphthong, fricative, nasal, plosive, semi-vowel, vowel, silence}, {consonant, vowel, silence} and {voiced, unvoiced, silence} and, calculate the contribution of each phonetic class in terms of the substitution, deletion, insertion and PER. Furthermore, for each BPC we investigate the following: evolution of PER during training, effect of noise (NTIMIT), importance of different spectral subbands (1, 2, 4, and 8 kHz), usefulness of bidirectional vs unidirectional sequential modelling, transfer learning from WSJ and regularisation via monophones. In addition, we construct a confusion matrix for each BPC and analyse the confusions via dimensionality reduction to 2D at the input (acoustic features) and output (logits) levels of the acoustic model. We also compare the performance and confusion matrices of the BLSTM-based hybrid baseline system with those of the GMM-HMM based hybrid, Conformer and wav2vec 2.0 based end-to-end phone recognisers. Finally, the relationship of the unweighted and weighted PERs with the broad phonetic class priors is studied for both the hybrid and end-to-end systems. Erfan Loweimi, Andrea Carmantini, Peter Bell 0001, Steve Renals, Zoran Cvetkovic |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Multi-Stream Acoustic Modelling Using Raw Real and Imaginary Parts of the Fourier TransformabstractIn this paper, we investigate multi-stream acoustic modelling using the raw real and imaginary parts of the Fourier transform of speech signals. Using the raw magnitude spectrum, or features derived from it, as a proxy for the real and imaginary parts leads to irreversible information loss and suboptimal information fusion. We discuss and quantify the importance of such information in terms of speech quality and intelligibility. In the proposed framework, the real and imaginary parts are treated as two streams of information, pre-processed via separate convolutional networks, and then combined at an optimal level of abstraction, followed by further post-processing via recurrent and fully-connected layers. The optimal level of information fusion in various architectures, training dynamics in terms of cross-entropy loss, frame classification accuracy and WER as well as the shape and properties of the filters learned in the first convolutional layer of single- and multi-stream models are analysed. We investigated the effectiveness of the proposed systems in various tasks: TIMIT/NTIMIT (phone recognition), Aurora-4 (noise robustness), WSJ (read speech), AMI (meeting) and TORGO (dysarthric speech). Across all tasks we achieved competitive performance: in Aurora-4, down to 4.6% WER on average, in WSJ down to 4.6% and 6.2% WERs for Eval-92 and Eval-93, for Dev/Eval sets of the AMI-IHM down to 23.3%/23.8% WERs and in the AMI-SDM down to 43.7%/47.6% WERs have been achieved. In TORGO, for dysarthric and typical speech we achieved down to 31.7% and 10.2% WERs, respectively. Erfan Loweimi, Zhengjun Yue, Peter Bell 0001, Steve Renals, Zoran Cvetkovic |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Investigating the contribution of speaker attributes to speaker separability using disentangled speaker representationsabstractDeep speaker embeddings have been shown to encode a wide variety of attributes relating to a speaker. The aim of this work is to separate out some of these attributes in the embedding space, disentangling these sources of speaker variation into subsets of the embedding dimensions. This is achieved modifying the training procedure of a typical speaker embedding network, which is typically only trained to classify speakers. This work instead adds pairs of attribute specific task heads to operate on complementary subsets of the speaker embedding dimensions. While specific dimensions are encouraged to encode an attribute, for example gender, the other dimensions are penalized for containing this information using an adversarial loss. We show that this method is effective in factorizing out multiple attributes in the embedding space, successfully disentangling gender, nationality and age. Using the disentangled representations, we investigate how much removing this information impacts speaker verification and diarization performance, showing that gender is a significant source of separation in the deep speaker embedding space, with nationality and age also contributing to a lesser degree. Chau Luu, Steve Renals, Peter Bell 0001 |
INTERSPEECH | 2 |
| 2022 | Towards Robust Waveform-Based Acoustic ModelsabstractWe study the problem of learning robust acoustic models in adverse environments, characterized by a significant mismatch between training and test conditions. This problem is of paramount importance for the deployment of speech recognition systems that need to perform well in unseen environments. First, we characterize data augmentation theoretically as an instance of vicinal risk minimization, which aims at improving risk estimates during training by replacing the delta functions that define the empirical density over the input space with an approximation of the marginal population density in the vicinity of the training samples. More specifically, we assume that local neighborhoods centered at training samples can be approximated using a mixture of Gaussians, and demonstrate theoretically that this can incorporate robust inductive bias into the learning process. We then specify the individual mixture components implicitly via data augmentation schemes, designed to address common sources of spurious correlations in acoustic models. To avoid potential confounding effects on robustness due to information loss, which has been associated with standard feature extraction techniques (e.g.,fbankandmfccfeatures), we focus on the waveform-based setting. Our empirical results show that the approach can generalize to unseen noise conditions, with 150% relative improvement in out-of-distribution generalization compared to training using the standard risk minimization principle. Moreover, the results demonstrate competitive performance relative to models learned using a training sample designed to match the acoustic conditions characteristic of test utterances. Dino Oglic, Zoran Cvetkovic, Peter Sollich, Steve Renals, Bin Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Leveraging Linguistic Knowledge for Accent Robustness of End-to-End ModelsabstractAcoustic models are susceptible to the difference in acoustic characteristics between the training distribution and test distributions. Accent variability is a challenging source of variability, and the variations within one accent often do not generalize to others. Consequently, end-to-end models that have only transcriptions as linguistic information need high amounts of data to learn how different accents realize their sounds. To aid with recognition of accented speech, we make use of an accent independent abstraction of phonemes, often called metaphonemes. We force our models to learn hidden representations that are correlated to metaphonemes using multi-task training. Our aim is to obtain a model that is more robust to accented speech and, can, at the same time, adapt faster to different accents through the learned structure. Our experiments on the Common Voice corpus show better generalization when making use of this additional linguistic information, with a word error rate reduction of up to 12.6% when compared to the baseline. Furthermore, the relative improvement when adapting an existing model by making use of the metaphonemes is higher than using Byte Pair Encodings alone. Andrea Carmantini, Steve Renals, Peter Bell 0001 |
ASRU | 2 |
| 2021 | Speech Acoustic Modelling from Raw Phase SpectrumabstractMagnitude spectrum-based features are the most widely employed front-ends for acoustic modelling in automatic speech recognition (ASR) systems. In this paper, we investigate the possibility and efficacy of acoustic modelling using the raw short-time phase spectrum. In particular, we study the usefulness of the raw wrapped, unwrapped and minimum-phase phase spectra as well as the phase of the source and filter components for acoustic modelling. Furthermore, we explore the effectiveness of simultaneous deployment of the vocal tract and excitation components of the raw phase spectrum using multi-head CNNs and investigate multiple information fusion schemes. This paves the way for developing an effective phase-based multi-stream information processing systems for speech recognition. The performance, even for wrapped phase with a noise-like shape, is comparable to or better than the magnitude-based classic features, and up to 4.8% WER has been achieved in the WSJ (Eval-92) task. Erfan Loweimi, Zoran Cvetkovic, Peter Bell 0001, Steve Renals |
ICASSP | 4 |
| 2021 | Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower LayersabstractAlthough the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in general, freezing the trained feature extractor (the lower layers) and retraining the classifier (the upper layers) on the same dataset leads to worse performance. In this paper, for the first time, we show that the frozen classifier is transferable within the same dataset. We develop a novel top-down training method which can be viewed as an algorithm for searching for high-quality classifiers. We tested this method on automatic speech recognition (ASR) tasks and language modelling tasks. The proposed method consistently improves recurrent neural network ASR models on Wall Street Journal, self-attention ASR models on Switchboard, and AWD-LSTM language models on WikiText-2. Shucong Zhang, Cong-Thanh Do, Rama Sanand Doddipatla, Erfan Loweimi, Peter Bell 0001, Steve Renals |
ICASSP | 6 |
| 2021 | Speech Acoustic Modelling Using Raw Source and Filter ComponentsabstractSource-filter modelling is among the fundamental techniques in speech processing with a wide range of applications. In acoustic modelling, features such as MFCC and PLP which parametrise the filter component are widely employed. In this paper, we investigate the efficacy of building acoustic models from the raw filter and source components. The raw magnitude spectrum, as the primary information stream, is decomposed into the excitation and vocal tract information streams via cepstral liftering. Then, acoustic models are built via multi-head CNNs which, among others, allow for processing each individual stream via a sequence of bespoke transforms and fusing them at an optimal level of abstraction. We discuss the possible advantages of such information factorisation and recombination, investigate the dynamics of these models and explore the optimal fusion level. Furthermore, we illustrate the CNN’s learned filters and provide some interpretation for the captured patterns. The proposed approach with optimal fusion scheme results in up to 14% and 7% relative WER reduction in WSJ and Aurora-4 tasks. Erfan Loweimi, Zoran Cvetkovic, Peter Bell 0001, Steve Renals |
Interspeech | 4 |
| 2021 | Leveraging Speaker Attribute Information Using Multi Task Learning for Speaker Verification and DiarizationabstractDeep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic aspects that make up a speaker's identity, whilst being robust to non-speaker acoustic variation. Deep speaker embeddings are normally trained discriminatively, predicting speaker identity labels on the training data. We hypothesise that additionally predicting speaker-related auxiliary variables -- such as age and nationality -- may yield representations that are better able to generalise to unseen speakers. We propose a framework for making use of auxiliary label information, even when it is only available for speech corpora mismatched to the target application. On a test set of US Supreme Court recordings, we show that by leveraging two additional forms of speaker attribute information derived respectively from the matched training data, and VoxCeleb corpus, we improve the performance of our deep speaker embeddings for both verification and diarization tasks, achieving a relative improvement of 26.2% in DER and 6.7% in EER compared to baselines using speaker labels only. This improvement is obtained despite the auxiliary labels having been scraped from the web and being potentially noisy. Chau Luu, Peter Bell 0001, Steve Renals |
Interspeech | 3 |
| 2021 | Silent versus Modal Multi-Speaker Speech Recognition from Ultrasound and VideoabstractWe investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evaluate on matched test sets of two speaking modes: silent and modal speech. We observe that silent speech recognition from imaging data underperforms compared to modal speech recognition, likely due to a speaking-mode mismatch between training and testing. We improve silent speech recognition performance using techniques that address the domain mismatch, such as fMLLR and unsupervised model adaptation. We also analyse the properties of silent and modal speech in terms of utterance duration and the size of the articulatory space. To estimate the articulatory space, we compute the convex hull of tongue splines, extracted from ultrasound tongue images. Overall, we observe that the duration of silent speech is longer than that of modal speech, and that silent speech covers a smaller articulatory space than modal speech. Although these two properties are statistically significant across speaking modes, they do not directly correlate with word error rates from speech recognition. Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals |
Interspeech | 4 |
| 2021 | Stochastic Attention Head Removal: A Simple and Effective Method for Improving Transformer Based ASR ModelsabstractRecently, Transformer based models have shown competitive automatic speech recognition (ASR) performance. One key factor in the success of these models is the multi-head attention mechanism. However, for trained models, we have previously observed that many attention matrices are close to diagonal, indicating the redundancy of the corresponding attention heads. We have also found that some architectures with reduced numbers of attention heads have better performance. Since the search for the best structure is time prohibitive, we propose to randomly remove attention heads during training and keep all attention heads at test time, thus the final model is an ensemble of models with different architectures. The proposed method also forces each head independently learn the most useful patterns. We apply the proposed method to train Transformer based and Convolution-augmented Transformer (Conformer) based ASR models. Our method gives consistent performance gains over strong baselines on the Wall Street Journal, AISHELL, Switchboard and AMI datasets. To the best of our knowledge, we have achieved state-of-the-art end-to-end Transformer based model performance on Switchboard and AMI. Shucong Zhang, Erfan Loweimi, Peter Bell 0001, Steve Renals |
Interspeech | 4 |
| 2021 | Tal: A Synchronised Multi-Speaker Corpus of Ultrasound Tongue Imaging, Audio, and Lip VideosabstractWe present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording sessions of one professional voice talent, a male native speaker of English; TaL80 is a set of recording sessions of 81 native speakers of English without voice talent experience. Overall, the corpus contains 24 hours of parallel ultrasound, video, and audio data, of which approximately 13.5 hours are speech. This paper describes the corpus and presents benchmark results for the tasks of speech recognition, speech synthesis (articulatory-to-acoustic mapping), and automatic synchronisation of ultrasound to audio. The TaL corpus is publicly available under the CC BY-NC 4.0 license. Manuel Sam Ribeiro, Jennifer Sanger, Jing-Xuan Zhang, Aciel Eshky, Alan Wrench, Korin Richmond, Steve Renals |
SLT | 7 |
| 2021 | On The Usefulness of Self-Attention for Automatic Speech Recognition with TransformersabstractSelf-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition results. However, we note the range of the learned context increases from the lower to upper self-attention layers, whilst acoustic events often happen within short time spans in a left-to-right order. This leads to a question: for speech recognition, is a global view of the entire sequence useful for the upper self-attention encoder layers in Transformers? To investigate this, we train models with lower self-attention/upper feed-forward layers encoders on Wall Street Journal and Switchboard. Compared to baseline Transformers, no performance drop but minor gains are observed. We further developed a novel metric of the diagonality of attention matrices and found the learned diagonality indeed increases from the lower to upper encoder self-attention layers. We conclude the global view is unnecessary in training upper encoder layers. Shucong Zhang, Erfan Loweimi, Peter Bell 0001, Steve Renals |
SLT | 4 |
| 2021 | Automatic audiovisual synchronisation for ultrasound tongue imagingabstractUltrasound tongue imaging is used to visualise the intra-oral articulators during speech production. It is utilised in a range of applications, including speech and language therapy and phonetics research. Ultrasound and speech audio are recorded simultaneously, and in order to correctly use this data, the two modalities should be correctly synchronised. Synchronisation is achieved using specialised hardware at recording time, but this approach can fail in practice resulting in data of limited usability. In this paper, we address the problem of automatically synchronising ultrasound and audio after data collection. We first investigate the tolerance of expert ultrasound users to synchronisation errors in order to find the thresholds for error detection. We use these thresholds to define accuracy scoring boundaries for evaluating our system. We then describe our approach for automatic synchronisation, which is driven by a self-supervised neural network, exploiting the correlation between the two signals to synchronise them. We train our model on data from multiple domains with different speaker characteristics, different equipment, and different recording environments, and achieve an accuracy >92.4% on held-out in-domain data. Finally, we introduce a novel resource, the Cleft dataset, which we gathered with a new clinical subgroup and for which hardware synchronisation proved unreliable. We apply our model to this out-of-domain data, and evaluate its performance subjectively with expert users. Results show that users prefer our model’s output over the original hardware output 79.3% of the time. Our results demonstrate the strength of our approach and its ability to generalise to data from new domains. Aciel Eshky, Joanne Cleland, Manuel Sam Ribeiro, Eleanor Sugden, Korin Richmond, Steve Renals |
Speech Commun. | 6 |
| 2021 | Exploiting ultrasound tongue imaging for the automatic detection of speech articulation errors
Manuel Sam Ribeiro, Joanne Cleland, Aciel Eshky, Korin Richmond, Steve Renals |
Speech Commun. | 5 |
| 2020 | Cross Lingual Transfer Learning for Zero-Resource Domain AdaptationabstractWe propose a method for zero-resource domain adaptation of DNN acoustic models, for use in low-resource situations where the only in-language training data available may be poorly matched to the intended target domain. Our method uses a multi-lingual model in which several DNN layers are shared between languages. This architecture enables domain adaptation transforms learned for one well-resourced language to be applied to an entirely different low- resource language. First, to develop the technique we use English as a well-resourced language and take Spanish to mimic a low-resource language. Experiments in domain adaptation between the conversational telephone speech (CTS) domain and broadcast news (BN) domain demonstrate a 29% relative WER improvement on Spanish BN test data by using only English adaptation data. Second, we demonstrate the effectiveness of the method for low-resource languages with a poor match to the well-resourced language. Even in this scenario, the proposed method achieves relative WER improvements of 18-27% by using solely English data for domain adaptation. Compared to other related approaches based on multi-task and multi-condition training, the proposed method is able to better exploit well-resource language data for improved acoustic modelling of the low-resource target domain. Alberto Abad, Peter Bell 0001, Andrea Carmantini, Steve Renals |
ICASSP | 4 |
| 2020 | Channel Adversarial Training for Speaker Verification and DiarizationabstractPrevious work has encouraged domain-invariance in deep speaker embedding by adversarially classifying the dataset or labelled environment to which the generated features belong. We propose a training strategy which aims to produce features that are invariant at the granularity of the recording or channel, a finer grained objective than dataset- or environment-invariance. By training an adversary to predict whether pairs of same-speaker embeddings belong to the same recording in a Siamese fashion, learned features are discouraged from utilizing channel information that may be speaker discriminative during training. Experiments for verification on VoxCeleb and diarization and verification on CALLHOME show promising improvements over a strong baseline in addition to outperforming a dataset-adversarial model. The VoxCeleb model in particular performs well, achieving a 4% relative improvement in EER over a Kaldi baseline, while using a similar architecture and less training data. Chau Luu, Peter Bell 0001, Steve Renals |
ICASSP | 3 |
| 2020 | Multi-Scale Octave Convolutions for Robust Speech RecognitionabstractWe propose a multi-scale octave convolution layer to learn robust speech representations efficiently. Octave convolutions were introduced by Chen et al [1] in the computer vision field to reduce the spatial redundancy of the feature maps by decomposing the output of a convolutional layer into feature maps at two different spatial resolutions, one octave apart. This approach improved the efficiency as well as the accuracy of the CNN models. The accuracy gain was attributed to the enlargement of the receptive field in the original input space. We argue that octave convolutions likewise improve the robustness of learned representations due to the use of average pooling in the lower resolution group, acting as a low-pass filter. We test this hypothesis by evaluating on two noisy speech corpora - Aurora-4 and AMI. We extend the octave convolution concept to multiple resolution groups and multiple octaves. To evaluate the robustness of the inferred representations, we report the similarity between clean and noisy encodings using an affine projection loss as a proxy robustness measure. The results show that proposed method reduces the WER by up to 6.6% relative for Aurora-4 and 3.6% for AMI, while improving the computational efficiency of the CNN acoustic models. Joanna Rownicka, Peter Bell 0001, Steve Renals |
ICASSP | 3 |
| 2020 | Learning Noise Invariant Features Through Transfer Learning For Robust End-to-End Speech RecognitionabstractEnd-to-end models yield impressive speech recognition results on clean datasets while having inferior performance on noisy datasets. To address this, we propose transfer learning from a clean dataset (WSJ) to a noisy dataset (CHiME4) for connectionist temporal classification models. We argue that the clean classifier (the upper layers of a neural network trained on clean data) can force the feature extractor (the lower layers) to learn the underlying noise invariant patterns in the noisy dataset. While training on the noisy dataset, the clean classifier is either frozen or trained with a small learning rate. The feature extractor is trained with no learning rate re-scaling. The proposed method gives up to 15.5% relative character error rate (CER) reduction compared to models trained only on CHiME-4. Furthermore, we use the test sets of Aurora-4 to perform evaluation on unseen noisy conditions. Our method has significantly lower CERs (11.3% relative on average) on all 14 Aurora-4 test sets compared to the conventional transfer learning method (no learning rate rescale for any layer), indicating our method enables the model to learn noise invariant features. Shucong Zhang, Cong-Thanh Do, Rama Sanand Doddipatla, Steve Renals |
ICASSP | 4 |
| 2020 | Word Error Rate Estimation Without ASR Output: e-WER2abstractMeasuring the performance of automatic speech recognition (ASR) systems requires manually transcribed data in order to compute the word error rate (WER), which is often time-consuming and expensive. In this paper, we continue our effort in estimating WER using acoustic, lexical and phonotactic features. Our novel approach to estimate the WER uses a multistream end-to-end architecture. We report results for systems using internal speech decoder features (glass-box), systems without speech decoder features (black-box), and for systems without having access to the ASR system (no-box). The no-box system learns joint acoustic-lexical representation from phoneme recognition results along with MFCC acoustic features to estimate WER. Considering WER per sentence, our no-box system achieves 0.56 Pearson correlation with the reference evaluation and 0.24 root mean square error (RMSE) across 1,400 sentences. The estimated overall WER by e-WER2 is 30.9% for a three hours test set, while the WER computed using the reference transcriptions was 28.5%. Ahmed Ali 0002, Steve Renals |
INTERSPEECH | 2 |
| 2020 | Deep Scattering Power Spectrum Features for Robust Speech RecognitionabstractDeep scattering spectrum consists of a cascade of wavelet transforms and modulus non-linearity. It generates features of different orders, with the first order coefficients approximately equal to the Mel-frequency cepstrum, and higher order coefficients recovering information lost at lower levels. We investigate the effect of including the information recovered by higher order coefficients on the robustness of speech recognition. To that end, we also propose a modification to the original scattering transform tailored for noisy speech. In particular, instead of the modulus non-linearity we opt to work with power coefficients and, therefore, use the squared modulus non-linearity. We quantify the robustness of scattering features using the word error rates of acoustic models trained on clean speech and evaluated using sets of utterances corrupted with different noise types. Our empirical results show that the second order scattering power spectrum coefficients capture invariants relevant for noise robustness and that this additional information improves generalization to unseen noise conditions (almost 20% relative error reduction on aurora 4). This finding can have important consequences on speech recognition systems that typically discard the second order information and keep only the first order features (known for emulating mfcc and fbank values) when representing speech. Neethu M. Joy, Dino Oglic, Zoran Cvetkovic, Peter Bell 0001, Steve Renals |
INTERSPEECH | 5 |
| 2020 | On the Robustness and Training Dynamics of Raw Waveform ModelsabstractWe investigate the robustness and training dynamics of raw waveform acoustic models for automatic speech recognition (ASR). It is known that the first layer of such models learn a set of filters, performing a form of time-frequency analysis. This layer is liable to be under-trained owing to gradient vanishing, which can negatively affect the network performance. Through a set of experiments on TIMIT, Aurora-4 and WSJ datasets, we investigate the training dynamics of the first layer by measuring the evolution of its average frequency response over different epochs. We demonstrate that the network efficiently learns an optimal set of filters with a high spectral resolution and the dynamics of the first layer highly correlates with the dynamics of the cross entropy (CE) loss and word error rate (WER). In addition, we study the robustness of raw waveform models in both matched and mismatched conditions. The accuracy of these models is found to be comparable to, or better than, their MFCC-based counterparts in matched conditions and notably improved by using a better alignment. The role of raw waveform normalisation was also examined and up to 4.3% absolute WER reduction in mismatched conditions was achieved. Erfan Loweimi, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2020 | Raw Sign and Magnitude Spectra for Multi-Head Acoustic ModellingabstractIn this paper we investigate the usefulness of the sign spectrum and its combination with the raw magnitude spectrum in acoustic modelling for automatic speech recognition (ASR). The sign spectrum is a sequence of ±1s, capturing one bit of the phase spectrum. It encodes information overlooked by the magnitude spectrum enabling unique signal characterisation and reconstruction. In particular, we demonstrate it carries information related to the temporal structure of the signal as well as the speech’s source component. Furthermore, we investigate the usefulness of combining it with the raw magnitude spectrum via multi-head CNNs at different fusion levels for ASR. While information-wise these two streams of information are together equivalent to the raw waveform signal the overall performance is noticeably higher than raw waveform and classic features such as MFCC and filterbank. This has been observed and verified in TIMIT, NTIMT, Aurora-4 and WSJ tasks and up to 14.5% relative WER reduction has been achieved. Erfan Loweimi, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2020 | A Deep 2D Convolutional Network for Waveform-Based Speech RecognitionabstractDue to limited computational resources, acoustic models of early automatic speech recognition ( asr) systems were built in low-dimensional feature spaces that incur considerable information loss at the outset of the process. Several comparative studies of automatic and human speech recognition suggest that this information loss can adversely affect the robustness of asr systems. To mitigate that and allow for learning of robust models, we propose a deep 2 d convolutional network in the waveform domain. The first layer of the network decomposes waveforms into frequency sub-bands, thereby representing them in a structured high-dimensional space. This is achieved by means of a parametric convolutional block defined via cosine modulations of compactly supported windows. The next layer embeds the waveform in an even higher-dimensional space of high-resolution spectro-temporal patterns, implemented via a 2 d convolutional block. This is followed by a gradual compression phase that selects most relevant spectro-temporal patterns using wide-pass 2 d filtering. Our results show that the approach significantly outperforms alternative waveform-based models on both noisy and spontaneous conversational speech (24% and 11% relative error reduction, respectively). Moreover, this study provides empirical evidence that learning directly from the waveform domain could be more effective than learning using hand-crafted features. Dino Oglic, Zoran Cvetkovic, Peter Bell 0001, Steve Renals |
INTERSPEECH | 4 |
| 2020 | European Language Grid: An OverviewabstractWith 24 official EU and many additional languages, multilingualism in Europe and an inclusive Digital Single Market can only be enabled through Language Technologies (LTs). European LT business is dominated by hundreds of SMEs and a few large players. Many are world-class, with technologies that outperform the global players. However, European LT business is also fragmented – by nation states, languages, verticals and sectors, significantly holding back its impact. The European Language Grid (ELG) project addresses this fragmentation by establishing the ELG as the primary platform for LT in Europe. The ELG is a scalable cloud platform, providing, in an easy-to-integrate way, access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. Once fully operational, it will enable the commercial and non-commercial European LT community to deposit and upload their technologies and data sets into the ELG, to deploy them through the grid, and to connect with other resources. The ELG will boost the Multilingual Digital Single Market towards a thriving European LT community, creating new jobs and opportunities. Furthermore, the ELG project organises two open calls for up to 20 pilot projects. It also sets up 32 national competence centres and the European LT Council for outreach and coordination purposes. Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitrios Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, Jan Hajic 0001, Jana Hamrlová, Lukás Kacena, Khalid Choukri, Victoria Arranz, Andrejs Vasiljevs, Orians Anvari, Andis Lagzdins, Julija Melnika, Gerhard Backfried, Erinç Dikici, Miroslav Jánosík, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuél Gómez-Pérez, Andrés García-Silva, Cristian Berrio, Ulrich Germann, Steve Renals, Ondrej Klejch |
LREC | 35 |
| 2019 | The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic SpeechabstractThis paper describes the fifth edition of the Multi-Genre Broadcast Challenge (MGB-5), an evaluation focused on Arabic speech recognition and dialect identification. MGB-5 extends the previous MGB-3 challenge in two ways: first it focuses on Moroccan Arabic speech recognition; second the granularity of the Arabic dialect identification task is increased from 5 dialect classes to 17, by collecting data from 17 Arabic speaking countries. Both tasks use YouTube recordings to provide a multi-genre multi-dialectal challenge in the wild. Moroccan speech transcription used about 13 hours of transcribed speech data, split across training, development, and test sets, covering 7-genres: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). The fine-grained Arabic dialect identification data was collected from known YouTube channels from 17 Arabic countries. 3,000 hours of this data was released for training, and 57 hours for development and testing. The dialect identification data was divided into three sub-categories based on the segment duration: short (under 5 s), medium (5-20 s), and long (>20 s). Overall, 25 teams registered for the challenge, and 9 teams submitted systems for the two tasks. We outline the approaches adopted in each system and summarize the evaluation results. Ahmed Ali 0002, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri |
ASRU | 7 |
| 2019 | Acoustic Model Adaptation from Raw Waveforms with SincnetabstractRaw waveform acoustic modelling has recently gained interest due to neural networks' ability to learn feature extraction, and the potential for finding better representations for a given scenario than hand-crafted features. SincNet has been proposed to reduce the number of parameters required in raw-waveform modelling, by restricting the filter functions, rather than having to learn every tap of each filter. We study the adaptation of the SincNet filter parameters from adults' to children's speech, and show that the parameterisation of the SincNet layer is well suited for adaptation in practice: we can efficiently adapt with a very small number of parameters, producing error rates comparable to techniques using orders of magnitude more parameters. Joachim Fainberg, Ondrej Klejch, Erfan Loweimi, Peter Bell 0001, Steve Renals |
ASRU | 5 |
| 2019 | Speaker Adaptive Training Using Model Agnostic Meta-LearningabstractSpeaker adaptive training (SAT) of neural network acoustic models learns models in a way that makes them more suitable for adaptation to test conditions. Conventionally, model-based speaker adaptive training is performed by having a set of speaker dependent parameters that are jointly optimised with speaker independent parameters in order to remove speaker variation. However, this does not scale well if all neural network weights are to be adapted to the speaker. In this paper we formulate speaker adaptive training as a meta-learning task, in which an adaptation process using gradient descent is encoded directly into the training of the model. We compare our approach with test-only adaptation of a standard baseline model and a SAT-LHUC model with a learned speaker adaptation schedule and demonstrate that the meta-learning approach achieves comparable results. Ondrej Klejch, Joachim Fainberg, Peter Bell 0001, Steve Renals |
ASRU | 4 |
| 2019 | Embeddings for DNN Speaker Adaptive TrainingabstractIn this work, we investigate the use of embeddings for speaker-adaptive training of DNNs (DNN-SAT) focusing on a small amount of adaptation data per speaker. DNN-SAT can be viewed as learning a mapping from each embedding to transformation parameters that are applied to the shared parameters of the DNN. We investigate different approaches to applying these transformations, and find that with a good training strategy, a multi-layer adaptation network applied to all hidden layers is no more effective than a single linear layer acting on the embeddings to transform the input features. In the second part of our work, we evaluate different embed-dings (i-vectors, x-vectors and deep CNN embeddings) in an additional speaker recognition task in order to gain insight into what should characterize an embedding for DNN-SAT. We find the performance for speaker recognition of a given representation is not correlated with its ASR performance; in fact, ability to capture more speech attributes than just speaker identity was the most important characteristic of the embed-dings for efficient DNN-SAT ASR. Our best models achieved relative WER gains of 4% and 9% over DNN baselines using speaker-level cepstral mean normalisation (CMN), and a fully speaker-independent model, respectively. Joanna Rownicka, Peter Bell 0001, Steve Renals |
ASRU | 3 |
| 2019 | On the Usefulness of Statistical Normalisation of Bottleneck Features for Speech RecognitionabstractDNNs play a major role in the state-of-the-art ASR systems. They can be used for extracting features and building probabilistic models for acoustic and language modelling. Despite their huge practical success, the level of theoretical understanding has remained shallow. This paper investigates DNNs from a statistical standpoint. In particular, the effect of activation functions on the distribution of the pre-activations and activations is investigated and discussed from both analytic and empirical viewpoints. This study, among others, shows that the pre-activation density in the bottleneck layer can be well fitted with a diagonal GMM with a few Gaussians and how and why the ReLU activation function promotes sparsity. Motivated by the statistical properties of the pre-activations, the usefulness of statistical normalisation of bottleneck features was also investigated. To this end, methods such as mean(-variance) normalisation, Gaussianisation, and histogram equalisation (HEQ) were employed and up to 2% (absolute) WER reduction achieved in the Aurora-4 task. Erfan Loweimi, Peter Bell 0001, Steve Renals |
ICASSP | 3 |
| 2019 | Speaker-independent Classification of Phonetic Segments from Raw Ultrasound in Child SpeechabstractUltrasound tongue imaging (UTI) provides a convenient way to visualize the vocal tract during speech production. UTI is increasingly being used for speech therapy, making it important to develop automatic methods to assist various time-consuming manual tasks currently performed by speech therapists. A key challenge is to generalize the automatic processing of ultrasound tongue images to previously unseen speakers. In this work, we investigate the classification of phonetic segments (tongue shapes) from raw ultrasound recordings under several training scenarios: speaker-dependent, multi-speaker, speaker-independent, and speaker-adapted. We observe that models underperform when applied to data from speakers not seen at training time. However, when provided with minimal additional speaker information, such as the mean ultrasound frame, the models generalize better to unseen speakers. Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals |
ICASSP | 4 |
| 2019 | Windowed Attention Mechanisms for Speech RecognitionabstractThe usual attention mechanisms used for encoder-decoder models do not constrain the relationship between input and output sequences to be monotonic. To address this we explore windowed attention mechanisms which restrict attention to a block of source hidden states. Rule-based windowing restricts attention to a (typically large) fixed-length window. The performance of such methods is poor if the window size is small. In this paper, we propose a fully-trainable windowed attention and provide a detailed analysis on the factors which affect the performance of such an attention mechanism. Compared to the rule-based window methods, the learned window size is significantly smaller yet the model's performance is competitive. On the TIMIT corpus this approach has resulted in a 17% (relative) performance improvement over the traditional attention model. Our model also yields comparable accuracies to the joint CTC-attention model on the Wall Street Journal corpus. Shucong Zhang, Erfan Loweimi, Peter Bell 0001, Steve Renals |
ICASSP | 4 |
| 2019 | Untranscribed Web Audio for Low Resource Speech RecognitionabstractSpeech recognition models are highly susceptible to mismatch in the acoustic and language domains between the training and the evaluation data. For low resource languages, it is difficult to obtain transcribed speech for target domains, while untranscribed data can be collected with minimal effort. Recently, a method applying lattice-free maximum mutual information (LF-MMI) to untranscribed data has been found to be effective for semi-supervised training. However, weaker initial models and domain mismatch can result in high deletion rates for the semi-supervised model. Therefore, we propose a method to force the base model to overgenerate possible transcriptions, relying on the ability of LF-MMI to deal with uncertainty. On data from the IARPA MATERIAL programme, our new semi-supervised method outperforms the standard semisupervised method, yielding significant gains when adapting for mismatched bandwidth and domain. Andrea Carmantini, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2019 | Synchronising Audio and Ultrasound by Learning Cross-Modal EmbeddingsabstractAudiovisual synchronisation is the task of determining the time offset between speech audio and a video recording of the articulators. In child speech therapy, audio and ultrasound videos of the tongue are captured using instruments which rely on hardware to synchronise the two modalities at recording time. Hardware synchronisation can fail in practice, and no mechanism exists to synchronise the signals post hoc. To address this problem, we employ a two-stream neural network which exploits the correlation between the two modalities to find the offset. We train our model on recordings from 69 speakers, and show that it correctly synchronises 82.9% of test utterances from unseen therapy sessions and unseen speakers, thus considerably reducing the number of utterances to be manually synchronised. An analysis of model performance on the test utterances shows that directed phone articulations are more difficult to automatically synchronise compared to utterances containing natural variation in speech such as words, sentences, or conversations. Aciel Eshky, Manuel Sam Ribeiro, Korin Richmond, Steve Renals |
INTERSPEECH | 4 |
| 2019 | Lattice-Based Lightly-Supervised Acoustic Model TrainingabstractIn the broadcast domain there is an abundance of related text data and partial transcriptions, such as closed captions and subtitles. This text data can be used for lightly supervised training, in which text matching the audio is selected using an existing speech recognition model. Current approaches to light supervision typically filter the data based on matching error rates between the transcriptions and biased decoding hypotheses. In contrast, semi-supervised training does not require matching text data, instead generating a hypothesis using a background language model. State-of-the-art semi-supervised training uses lattice-based supervision with the lattice-free MMI (LF-MMI) objective function. We propose a technique to combine inaccurate transcriptions with the lattices generated for semi-supervised training, thus preserving uncertainty in the lattice where appropriate. We demonstrate that this combined approach reduces the expected error rates over the lattices, and reduces the word error rate (WER) on a broadcast task. Joachim Fainberg, Ondrej Klejch, Steve Renals, Peter Bell 0001 |
INTERSPEECH | 3 |
| 2019 | On Learning Interpretable CNNs with Parametric Modulated Kernel-Based FiltersabstractWe investigate the problem of direct waveform modelling using parametric kernel-based filters in a convolutional neural network (CNN) framework, building on SincNet, a CNN employing the cardinal sine (sinc) function to implement learnable bandpass filters. To this end, the general problem of learning a filterbank consisting of modulated kernel-based baseband filters is studied. Compared to standard CNNs, such models have fewer parameters, learn faster, and require less training data. They are also more amenable to human interpretation, paving the way to embedding some perceptual prior knowledge in the architecture. We have investigated the replacement of the rectangular filters of SincNet with triangular, gammatone and Gaussian filters, resulting in higher model flexibility and a reduction to the phone error rate. We also explore the properties of the learned filters learned for TIMIT phone recognition from both perceptual and statistical standpoints. We find that the filters in the first layer, which directly operate on the waveform, are in accord with the prior knowledge utilised in designing and engineering standard filters such as mel-scale triangular filters. That is, the networks learn to pay more attention to perceptually significant spectral neighbourhoods where the data centroid is located, and the variance and Shannon entropy are highest. Erfan Loweimi, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2019 | Ultrasound Tongue Imaging for Diarization and Alignment of Child Speech Therapy SessionsabstractWe investigate the automatic processing of child speech therapy sessions using ultrasound visual biofeedback, with a specific focus on complementing acoustic features with ultrasound images of the tongue for the tasks of speaker diarization and time-alignment of target words. For speaker diarization, we propose an ultrasound-based time-domain signal which we call estimated tongue activity. For word-alignment, we augment an acoustic model with low-dimensional representations of ultrasound images of the tongue, learned by a convolutional neural network. We conduct our experiments using the Ultrasuite repository of ultrasound and speech recordings for child speech therapy sessions. For both tasks, we observe that systems augmented with ultrasound data outperform corresponding systems using only the audio signal. Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals |
INTERSPEECH | 4 |
| 2019 | Trainable Dynamic Subsampling for End-to-End Speech RecognitionabstractJointly optimised attention-based encoder-decoder models have yielded impressive speech recognition results. The recurrent neural network (RNN) encoder is a key component in such models – it learns the hidden representations of the inputs.However, it is difficult for RNNs to model the long sequences characteristic of speech recognition. To address this, subsampling between stacked recurrent layers of the encoder is commonly employed. This method reduces the length of the input sequence and leads to gains in accuracy. However, static subsampling may both include redundant information and miss relevant information. We propose using a dynamic subsampling RNN (dsRNN) encoder. Unlike a statically subsampled RNN encoder, the dsRNN encoder can learn to skip redundant frames. Furthermore, the skip ratio may vary at different stages of training, thus allowing the encoder to learn the most relevant information for each epoch. Although the dsRNN is unidirectional, it yields lower phone error rates (PERs) than a bidirectional RNN on TIMIT. The dsRNN encoder has a 16.8% PER on the TIMIT test set, a considerable improvement over static subsampling methods used with unidirectional and bidirectional RNN encoders (23.5% and 20.4% PER respectively). Shucong Zhang, Erfan Loweimi, Yumo Xu, Peter Bell 0001, Steve Renals |
INTERSPEECH | 5 |
| 2018 | Dynamic Evaluation of Neural Sequence ModelsabstractWe explore dynamic evaluation, where sequence models are adapted to the recent sequence history using gradient descent, assigning higher probabilities to re-occurring sequential patterns. We develop a dynamic evaluation approach that outperforms existing adaptation approaches in our comparisons. We apply dynamic evaluation to outperform all previous word-level perplexities on the Penn Treebank and WikiText-2 datasets (achieving 51.1 and 44.3 respectively) and all previous character-level cross-entropies on the text8 and Hutter Prize datasets (achieving 1.19 bits/char and 1.08 bits/char respectively). Ben Krause, Emmanuel Kahembwe, Iain Murray 0001, Steve Renals |
ICML | 4 |
| 2018 | Analyzing Deep CNN-Based Utterance Embeddings for Acoustic Model AdaptationabstractWe explore why deep convolutional neural networks (CNNs) with small two-dimensional kernels, primarily used for modeling spatial relations in images, are also effective in speech recognition. We analyze the representations learned by deep CNNs and compare them with deep neural network (DNN) representations and i-vectors, in the context of acoustic model adaptation. To explore whether interpretable information can be decoded from the learned representations we evaluate their ability to discriminate between speakers, acoustic conditions, noise type, and gender using the Aurora-4 dataset. We extract both whole model embeddings (to capture the information learned across the whole network) and layer-specific embeddings which enable understanding of the flow of information across the network. We also use learned representations as the additional input for a time-delay neural network (TDNN) for the Aurora-4 and MGB-3 English datasets. We find that deep CNN embeddings outperform DNN embeddings for acoustic model adaptation and auxiliary features based on deep CNN embeddings result in similar word error rates to i-vectors. Joanna Rownicka, Peter Bell 0001, Steve Renals |
SLT | 3 |
| 2017 | WERD: Using social text spelling variants for evaluating dialectal speech recognitionabstractWe study the problem of evaluating automatic speech recognition (ASR) systems that target dialectal speech input. A major challenge in this case is that the orthography of dialects is typically not standardized. From an ASR evaluation perspective, this means that there is no clear gold standard for the expected output, and several possible outputs could be considered correct according to different human annotators, which makes standard word error rate (WER) inadequate as an evaluation metric. Such a situation is typical for machine translation (MT), and thus we borrow ideas from an MT evaluation metric, namely TERp, an extension of translation error rate which is closely-related to WER. In particular, in the process of comparing a hypothesis to a reference, we make use of spelling variants for words and phrases, which we mine from Twitter in an unsupervised fashion. Our experiments with evaluating ASR output for Egyptian Arabic, and further manual analysis, show that the resulting WERd (i.e., WER for dialects) metric, a variant of TERp, is more adequate than WER for evaluating dialectal ASR. Ahmed Ali 0002, Preslav Nakov, Peter Bell 0001, Steve Renals |
ASRU | 4 |
| 2017 | Speech recognition challenge in the wild: Arabic MGB-3abstractThis paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news recordings from Aljazeera Arabic TV programs, MGB-3 emphasises dialectal Arabic using a multi-genre collection of Egyptian YouTube videos. Seven genres were used for the data collection: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). A total of 16 hours of videos, split evenly across the different genres, were divided into adaptation, development and evaluation data sets. The Arabic MGB-Challenge comprised two tasks: A) Speech transcription, evaluated on the MGB-3 test set, along with the 10 hour MGB-2 test set to report progress on the MGB-2 evaluation; B) Arabic dialect identification, introduced this year in order to distinguish between four major Arabic dialects - Egyptian, Levantine, North African, Gulf, as well as Modern Standard Arabic. Two hours of audio per dialect were released for development and a further two hours were used for evaluation. For dialect identification, both lexical features and i-vector bottleneck features were shared with participants in addition to the raw audio recordings. Overall, thirteen teams submitted ten systems to the challenge. We outline the approaches adopted in each system, and summarise the evaluation results. Ahmed Ali 0002, Stephan Vogel, Steve Renals |
ASRU | 3 |
| 2017 | Simplifying very deep convolutional neural network architectures for robust speech recognitionabstractVery deep convolutional neural networks (VDCNNs) have been successfully used in computer vision. More recently VDCNNs have been applied to speech recognition, using architectures adopted from computer vision. In this paper, we experimentally analyse the role of the components in VDCNN architectures for robust speech recognition. We have proposed a number of simplified VDCNN architectures, taking into account the use of fully-connected layers and down-sampling approaches. We have investigated three ways to down-sample feature maps: max-pooling, average-pooling, and convolution with increased stride. Our proposed model consisting solely of convolutional (conv) layers, and without any fully-connected layers, achieves a lower word error rate on Aurora 4 compared to other VDCNN architectures typically used in speech recognition. We have also extended our experiments to the MGB-3 task of multi-genre broadcast recognition using BBC TV recordings. The MGB-3 results indicate that the same architecture achieves the best result among our VDCNNs on this task as well. Joanna Rownicka, Steve Renals, Peter Bell 0001 |
ASRU | 2 |
| 2017 | Hierarchical recurrent neural network for story segmentation using fusion of lexical and acoustic featuresabstractA broadcast news stream consists of a number of stories and it is an important task to find the boundaries of stories automatically in news analysis. We capture the topic structure using a hierarchical model based on a Recurrent Neural Network (RNN) sentence modeling layer and a bidirectional Long Short-Term Memory (LSTM) topic modeling layer, with a fusion of acoustic and lexical features. Both features are accumulated with RNNs and trained jointly within the model to be fused at the sentence level. We conduct experiments on the topic detection and tracking (TDT4) task comparing combinations of two modalities trained with limited amount of parallel data. Further we utilize additional sufficient text data for training to polish our model. Experimental results indicate that the hierarchical RNN topic modeling takes advantage of the fusion scheme, especially with additional text training data, with a higher F1-measure compared to conventional state-of-the-art methods. Emiru Tsunoo, Ondrej Klejch, Peter Bell 0001, Steve Renals |
ASRU | 4 |
| 2017 | Sequence-to-sequence models for punctuated transcription combining lexical and acoustic featuresabstractIn this paper we present an extension of our previously described neural machine translation based system for punctuated transcription. This extension allows the system to map from per frame acoustic features to word level representations by replacing the traditional encoder in the encoder-decoder architecture with a hierarchical encoder. Furthermore, we show that a system combining lexical and acoustic features significantly outperforms systems using only a single source of features on all measured punctuation marks. The combination of lexical and acoustic features achieves a significant improvement in F-Measure of 1.5 absolute over the purely lexical neural machine translation based system. Ondrej Klejch, Peter Bell 0001, Steve Renals |
ICASSP | 3 |
| 2017 | Knowledge distillation for small-footprint highway networksabstractDeep learning has significantly advanced state-of-the-art of speech recognition in the past few years. However, compared to conventional Gaussian mixture acoustic models, neural network models are usually much larger, and are therefore not very deployable in embedded devices. Previously, we investigated a compact highway deep neural network (HDNN) for acoustic modelling, which is a type of depth-gated feedforward neural network. We have shown that HDNN-based acoustic models can achieve comparable recognition accuracy with much smaller number of model parameters compared to plain deep neural network (DNN) acoustic models. In this paper, we push the boundary further by leveraging on the knowledge distillation technique that is also known as teacher-student training, i.e., we train the compact HDNN model with the supervision of a high accuracy cumbersome model. Furthermore, we also investigate sequence training and adaptation in the context of teacher-student training. Our experiments were performed on the AMI meeting speech recognition corpus. With this technique, we significantly improved the recognition accuracy of the HDNN acoustic model with less than 0.8 million parameters, and narrowed the gap between this model and the plain DNN with 30 million parameters. Liang Lu 0001, Michelle Guo, Steve Renals |
ICASSP | 3 |
| 2017 | Factorised Representations for Neural Network Adaptation to Diverse Acoustic EnvironmentsabstractAdapting acoustic models jointly to both speaker and environment has been shown to be effective. In many realistic scenarios, however, either the speaker or environment at test time might be unknown, or there may be insufficient data to learn a joint transform. Generating independent speaker and environment transforms improves the match of an acoustic model to unseen combinations. Using i-vectors, we demonstrate that it is possible to factorise speaker or environment information using multi-condition training with neural networks. Specifically, we extract bottleneck features from networks trained to classify either speakers or environments. We perform experiments on the Wall Street Journal corpus combined with environment noise from the Diverse Environments Multichannel Acoustic Noise Database. Using the factorised i-vectors we show improvements in word error rates on perturbed versions of the eval92 and dev93 test sets, both when one factor is missing and when the factors are seen but not in the desired combination. Joachim Fainberg, Steve Renals, Peter Bell 0001 |
INTERSPEECH | 2 |
| 2017 | Hierarchical Recurrent Neural Network for Story SegmentationabstractA broadcast news stream consists of a number of stories and each story consists of several sentences. We capture this structure using a hierarchical model based on a word-level Recurrent Neural Network (RNN) sentence modeling layer and a sentence-level bidirectional Long Short-Term Memory (LSTM) topic modeling layer. First, the word-level RNN layer extracts a vector embedding the sentence information from the given transcribed lexical tokens of each sentence. These sentence embedding vectors are fed into a bidirectional LSTM that models the sentence and topic transitions. A topic posterior for each sentence is estimated discriminatively and a Hidden Markov model (HMM) follows to decode the story sequence and identify story boundaries. Experiments on the topic detection and tracking (TDT2) task indicate that the hierarchical RNN topic modeling achieves the best story segmentation performance with a higher F1-measure compared to conventional state-of-the-art methods. We also compare variations of our model to infer the optimal structure for the story segmentation task. Emiru Tsunoo, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2017 | Multitask Learning of Context-Dependent Targets in Deep Neural Network Acoustic ModelsabstractThis paper investigates the use of multitask learning to improve context-dependent deep neural network (DNN) acoustic models. The use of hybrid DNN systems with clustered triphone targets is now standard in automatic speech recognition. However, we suggest that using a single set of DNN targets in this manner may not be the most effective choice, since the targets are the result of a somewhat arbitrary clustering process that may not be optimal for discrimination. We propose to remedy this problem through the addition of secondary tasks predicting alternative content-dependent or context-independent targets. We present a comprehensive set of experiments on a lecture recognition task showing that DNNs trained through multitask learning in this manner give consistently improved performance compared to standard hybrid DNNs. The technique is evaluated across a range of data and output sizes. Improvements are seen when training uses the cross entropy criterion and also when sequence training is applied. Peter Bell 0001, Pawel Swietojanski, Steve Renals |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Small-Footprint Highway Deep Neural Networks for Speech RecognitionabstractState-of-the-art speech recognition systems typically employ neural network acoustic models. However, compared to Gaussian mixture models, deep neural network (DNN) based acoustic models often have many more model parameters, making it challenging for them to be deployed on resource-constrained platforms, such as mobile devices. In this paper, we study the application of the recently proposed highway deep neural network (HDNN) for training small-footprint acoustic models. HDNNs are a depth-gated feedforward neural network, which include two types of gate functions to facilitate the information flow through different layers. Our study demonstrates that HDNNs are more compact than regular DNNs for acoustic modeling, i.e., they can achieve comparable recognition accuracy with many fewer model parameters. Furthermore, HDNNs are more controllable than DNNs: The gate functions of an HDNN can control the behavior of the whole network using a very small number of model parameters. Finally, we show that HDNNs are more adaptable than DNNs. For example, simply updating the gate functions using adaptation data can result in considerable gains in accuracy. We demonstrate these aspects by experiments using the publicly available AMI corpus, which has around 80 h of training data. Liang Lu 0001, Steve Renals |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognitionabstractRecently, there has been an increasing interest in end-to-end speech recognition using neural networks, with no reliance on hidden Markov models (HMMs) for sequence modelling as in the standard hybrid framework. The recurrent neural network (RNN) encoderdecoder is such a model, performing sequence to sequence mapping without any predefined alignment. This model first transforms the input sequence into a fixed length vector representation, from which the decoder recovers the output sequence. In this paper, we extend our previous work on this model for large vocabulary end-to-end speech recognition. We first present a more effective stochastic gradient decent (SGD) learning rate schedule that can significantly improve the recognition accuracy. We then extend the decoder with long memory by introducing another recurrent layer that performs implicit language modelling. Finally, we demonstrate that using multiple recurrent layers in the encoder can reduce the word error rate. Our experiments were carried out on the Switchboard corpus using a training set of around 300 hours of transcribed audio data, and we have achieved significantly higher recognition accuracy, thereby reduced the gap compared to the hybrid baseline. Liang Lu 0001, Xingxing Zhang 0002, Steve Renals |
ICASSP | 3 |
| 2016 | SAT-LHUC: Speaker adaptive training for learning hidden unit contributionsabstractThis paper extends learning hidden unit contributions (LHUC) unsupervised speaker adaptation with speaker adaptive training (SAT). Contrary to other SAT approaches, the proposed technique does not require speaker-dependent features, the generation of auxiliary generative models to estimate or extract speaker-dependent information, or any changes to the speaker-independent model structure. SAT-LHUC is directly integrated into the objective and jointly learns speaker-independent and speaker-dependent representations. We demonstrate that the SAT-LHUC technique can match feature-space regression transforms for matched narrow-band data and outperform it on wide-band data when the runtime distribution differs significantly from training one. We have obtained 6.5%, 10% and 18.5% relative word error rate reductions compared to speaker-independent models on Switchboard, AMI meetings and TED lectures, respectively. This corresponds to relative gains of 2%, 4% and 6% compared with non-SAT LHUC adaptation. SAT-LHUC was also found to be complementary to SAT with feature-space maximum likelihood linear regression transforms. Pawel Swietojanski, Steve Renals |
ICASSP | 2 |
| 2016 | Automatic Dialect Detection in Arabic Broadcast SpeechabstractWe investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification. Ahmed Ali 0002, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James R. Glass, Peter Bell 0001, Steve Renals |
INTERSPEECH | 8 |
| 2016 | Improving Children's Speech Recognition Through Out-of-Domain Data AugmentationabstractChildren's speech poses challenges to speech recognition due to strong age-dependent anatomical variations and a lack of large, publicly-available corpora. In this paper we explore data augmentation for children's speech recognition using stochastic feature mapping (SFM) to transform out-of-domain adult data for both GMM-based and DNN-based acoustic models. We performed experiments on the English PF-STAR corpus, augmenting using WSJCAM0 and ABI. Our experimental results indicate that a DNN acoustic model for childrens speech can make use of adult data, and that out-of-domain SFM is more accurate than in-domain SFM. Joachim Fainberg, Peter Bell 0001, Mike Lincoln, Steve Renals |
INTERSPEECH | 4 |
| 2016 | Unsupervised Adaptation of Recurrent Neural Network Language ModelsabstractRecurrent neural network language models (RNNLMs) have been shown to consistently improve Word Error Rates (WERs) of large vocabulary speech recognition systems employing n-gram LMs. In this paper we investigate supervised and unsupervised discriminative adaptation of RNNLMs in a broadcast transcription task to target domains defined by either genre or show. We have explored two approaches based on (1) scaling forward-propagated hidden activations (Learning Hidden Unit Contributions (LHUC) technique) and (2) direct fine-tuning of the parameters of the whole RNNLM. To investigate the effectiveness of the proposed methods we carry out experiments on multi-genre broadcast (MGB) data following the MGB-2015 challenge protocol. We observe small but significant improvements in WER compared to a strong unadapted RNNLM model. Siva Reddy Gangireddy, Pawel Swietojanski, Peter Bell 0001, Steve Renals |
INTERSPEECH | 4 |
| 2016 | Segmental Recurrent Neural Networks for End-to-End Speech RecognitionabstractWe study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based acoustic models, it does not rely on an external system to provide features or segmentation boundaries. Instead, this model marginalises out all the possible segmentations, and features are extracted from the RNN trained together with the segmental CRF. In essence, this model is self-contained and can be trained end-to-end. In this paper, we discuss practical training and decoding issues as well as the method to speed up the training in the context of speech recognition. We performed experiments on the TIMIT dataset. We achieved 17.3 phone error rate (PER) from the first-pass decoding --- the best reported result using CRFs, despite the fact that we only used a zeroth-order CRF and without using any language model. Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith, Steve Renals |
INTERSPEECH | 5 |
| 2016 | Small-Footprint Deep Neural Networks with Highway Connections for Speech RecognitionabstractFor speech recognition, deep neural networks (DNNs) have significantly improved the recognition accuracy in most of benchmark datasets and application domains. However, compared to the conventional Gaussian mixture models, DNN-based acoustic models usually have much larger number of model parameters, making it challenging for their applications in resource constrained platforms, e.g., mobile devices. In this paper, we study the application of the recently proposed highway network to train small-footprint DNNs, which are {\it thinner} and {\it deeper}, and have significantly smaller number of model parameters compared to conventional DNNs. We investigated this approach on the AMI meeting speech transcription corpus which has around 70 hours of audio data. The highway neural networks constantly outperformed their plain DNN counterparts, and the number of model parameters can be reduced significantly without sacrificing the recognition accuracy. Liang Lu 0001, Steve Renals |
INTERSPEECH | 2 |
| 2016 | Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project
Guntis Barzdins, Steve Renals, Didzis Gosko |
LREC | 2 |
| 2016 | The MGB-2 challenge: Arabic multi-dialect broadcast media recognitionabstractThis paper describes the Arabic Multi-Genre Broadcast (MGB-2) Challenge for SLT-2016. Unlike last year's English MGB Challenge, which focused on recognition of diverse TV genres, this year, the challenge has an emphasis on handling the diversity in dialect in Arabic speech. Audio data comes from 19 distinct programmes from the Aljazeera Arabic TV channel between March 2005 and December 2015. Programmes are split into three groups: conversations, interviews, and reports. A total of 1,200 hours have been released with lightly supervised transcriptions for the acoustic modelling. For language modelling, we made available over 110M words crawled from Aljazeera Arabic website Aljazeera.net for a 10 year duration 2000-2011. Two lexicons have been provided, one phoneme based and one grapheme based. Finally, two tasks were proposed for this year's challenge: standard speech transcription, and word alignment. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained. Ahmed Ali 0002, Peter Bell 0001, James R. Glass, Yacine Messaoui, Hamdy Mubarak, Steve Renals |
SLT | 6 |
| 2016 | Punctuated transcription of multi-genre broadcasts using acoustic and lexical approachesabstractIn this paper we investigate the punctuated transcription of multi-genre broadcast media. We examine four systems, three of which are based on lexical features, the fourth of which uses acoustic features by integrating punctuation into the speech recognition acoustic models. We also explore the combination of these component systems using voting and log-linear interpolation. We performed experiments on the English language MGB Challenge data, which comprises about 1,600h of BBC television recordings. Our results indicate that a lexical system, based on a neural machine translation approach is significantly better than other systems achieving an F-Measure of 62.6% on reference text, with a relative degradation of 19% on ASR output. Our analysis of the results in terms of specific punctuation indicated that using longer context improves the prediction of question marks and acoustic information improves prediction of exclamation marks. Finally, we show that even though the systems are complementary, their straightforward combination does not yield better F-measures than a single system using neural machine translation. Ondrej Klejch, Peter Bell 0001, Steve Renals |
SLT | 3 |
| 2016 | Learning Hidden Unit Contributions for Unsupervised Acoustic Model AdaptationabstractThis work presents a broad study on the adaptation of neural network acoustic models by means of learning hidden unit contributions (LHUC) - a method that linearly re-combines hidden units in a speaker- or environment-dependent manner using small amounts of unsupervised adaptation data. We also extend LHUC to a speaker adaptive training (SAT) framework that leads to a more adaptable DNN acoustic model, working both in a speaker-dependent and a speaker-independent manner, without the requirements to maintain auxiliary speaker-dependent feature extractors or to introduce significant speaker-dependent changes to the DNN structure. Through a series of experiments on four different speech recognition benchmarks (TED talks, Switchboard, AMI meetings, and Aurora4) comprising 270 test speakers, we show that LHUC in both its test-only and SAT variants results in consistent word error rate reductions ranging from 5% to 23% relative depending on the task and the degree of mismatch between training and test data. In addition, we have investigated the effect of the amount of adaptation data per speaker, the quality of unsupervised adaptation targets, the complementarity to other adaptation techniques, one-shot adaptation, and an extension to adapting DNNs trained in a sequence discriminative manner. Pawel Swietojanski, Jinyu Li 0001, Steve Renals |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Differentiable Pooling for Unsupervised Acoustic Model AdaptationabstractWe present a deep neural network (DNN) acoustic model that includes parametrised and differentiable pooling operators. Unsupervised acoustic model adaptation is cast as the problem of updating the decision boundaries implemented by each pooling operator. In particular, we experiment with two types of pooling parametrisations: learned Lp-norm pooling and weighted Gaussian pooling, in which the weights of both operators are treated as speaker-dependent. We perform investigations using three different large vocabulary speech recognition corpora: AMI meetings, TED talks, and Switchboard conversational telephone speech. We demonstrate that differentiable pooling operators provide a robust and relatively low-dimensional way to adapt acoustic models, with relative word error rates reductions ranging from 5-20% with respect to unadapted systems, which themselves are better than the baseline fully-connected DNN-based acoustic models. We also investigate how the proposed techniques work under various adaptation conditions including the quality of adaptation data and complementarity to other feature- and model-space adaptation methods, as well as providing an analysis of the characteristics of each of the proposed approaches. Pawel Swietojanski, Steve Renals |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A system for automatic alignment of broadcast media captions using weighted finite-state transducersabstractWe describe our system for alignment of broadcast media captions in the 2015 MGB Challenge. A precise time alignment of previously-generated subtitles to media data is important in the process of caption generation by broadcasters. However, this task is challenging due to the highly diverse, often noisy content of the audio, and because the subtitles are frequently not a verbatim representation of the actual words spoken. Our system employs a two-pass approach with appropriately constrained weighted finite state transducers (WFSTs) to enable good alignment even when the audio quality would be challenging for conventional ASR. The system achieves an f-score of 0.8965 on the MGB Challenge development set. Peter Bell 0001, Steve Renals |
ASRU | 2 |
| 2015 | Multi-reference WER for evaluating ASR for languages with no orthographic rulesabstractLanguages with no standard orthographic representation faces a challenge to evaluate the output from Automatic Speech Recognition (ASR). Since the reference transcription text can vary widely from one user to another. We propose an innovative approach for evaluating speech recognition using Multi-References. For each recognized speech segments, we ask five different users to transcribe the speech. We combine the alignment for the multiple references, and use the combined alignment to report a modified version of Word Error Rate (WER). This approach is in favor of accepting a recognized word if any of the references typed it in the same form. Results are reported using two Dialectal Arabic (DA) as a language with no standard orthographic; Egyptian, and North African speech. The average WER for the five references individually is 71.4%, and 80.1% respectively. When considering all references combined, the Multi-References MR-WER was found to be 39.7%, and 45.9% respectively. Ahmed Ali 0002, Walid Magdy, Peter Bell 0001, Steve Renals |
ASRU | 4 |
| 2015 | The MGB challenge: Evaluating multi-genre broadcast media recognitionabstractThis paper describes the Multi-Genre Broadcast (MGB) Challenge at ASRU 2015, an evaluation focused on speech recognition, speaker diarization, and "lightly supervised" alignment of BBC TV recordings. The challenge training data covered the whole range of seven weeks BBC TV output across four channels, resulting in about 1,600 hours of broadcast audio. In addition several hundred million words of BBC subtitle text was provided for language modelling. A novel aspect of the evaluation was the exploration of speech recognition and speaker diarization in a longitudinal setting — i.e. recognition of several episodes of the same show, and speaker diarization across these episodes, linking speakers. The longitudinal tasks also offered the opportunity for systems to make use of supplied metadata including show title, genre tag, and date/time of transmission. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained. Peter Bell 0001, Mark J. F. Gales, Thomas Hain, Jonathan Kilgour, Pierre Lanchantin, Xunying Liu, Andrew McParland, Steve Renals, Oscar Saz-Torralba, Mirjam Wester, Philip C. Woodland |
ASRU | 8 |
| 2015 | Regularization of context-dependent deep neural networks with context-independent multi-task trainingabstractThe use of context-dependent targets has become standard in hybrid DNN systems for automatic speech recognition. However, we argue that despite the use of state-tying, optimising to context-dependent targets can lead to over-fitting, and that discriminating between arbitrary tied context-dependent targets may not be optimal. We propose a multitask learning method where the network jointly predicts context-dependent and monophone targets. We evaluate the method on a large-vocabulary lecture recognition task and show that it yields relative improvements of 3-10% over baseline systems. Peter Bell 0001, Steve Renals |
ICASSP | 2 |
| 2015 | Multi-frame factorisation for long-span acoustic modellingabstractAcoustic models based on Gaussian mixture models (GMMs) typically use short span acoustic feature inputs. This does not capture long-term temporal information from speech owing to the conditional independence assumption of hidden Markov models. In this paper, we present an implicit approach that approximates the joint distribution of long span features by product of factorized models, in contrast to deep neural networks (DNNs) that model feature correlations directly. The approach is applicable to a broad range of acoustic models. We present experiments using GMM and probabilistic linear discriminant analysis (PLDA) based models on Switchboard, observing consistent word error rate reductions. Liang Lu 0001, Steve Renals |
ICASSP | 2 |
| 2015 | Differentiable pooling for unsupervised speaker adaptationabstractThis paper proposes a differentiable pooling mechanism to perform model-based neural network speaker adaptation. The proposed technique learns a speaker-dependent combination of activations within pools of hidden units, was shown to work well unsupervised, and does not require speaker-adaptive training. We have conducted a set of experiments on the TED talks data, as used in the IWSLT evaluations. Our results indicate that the approach can reduce word error rates (WERs) on standard IWSLT test sets by about 5-11% relative compared to speaker-independent systems and was found complementary to the recently proposed learning hidden units contribution (LHUC) approach, reducing WER by 6-13% relative. Both methods were also found to work well when adapting with small amounts of unsupervised data - 10 seconds is able to decrease the WER by 5% relative compared to the baseline speaker independent system. Pawel Swietojanski, Steve Renals |
ICASSP | 2 |
| 2015 | Modelling acoustic feature dependencies with artificial neural networks: Trajectory-RNADEabstractGiven a transcription, sampling from a good model of acoustic feature trajectories should result in plausible realizations of an utterance. However, samples from current probabilistic speech synthesis systems result in low quality synthetic speech. Henter et al. have demonstrated the need to capture the dependencies between acoustic features conditioned on the phonetic labels in order to obtain high quality synthetic speech. These dependencies are often ignored in neural network based acoustic models. We tackle this deficiency by introducing a probabilistic neural network model of acoustic trajectories, trajectory RNADE, able to capture these dependencies. Benigno Uria, Iain Murray 0001, Steve Renals, Cassia Valentini-Botinhao, John Bridle |
ICASSP | 3 |
| 2015 | Complementary tasks for context-dependent deep neural network acoustic modelsabstractWe have previously found that context-dependent DNN models for automatic speech recognition can be improved with the use of monophone targets as a secondary task for the network. This paper asks whether the improvements derive from the regularising effect of having a much small number of monophone outputs – compared to the typical number of tied states – or from the use of targets that are not tied to an arbitrary stateclustering. We investigate the use of factorised targets for left and right context, and targets motivated by articulatory properties of the phonemes. We present results on a large-vocabulary lecture recognition task. Although the regularising effect of monophones seems to be important, all schemes give substantial improvements over the baseline single task system, even though the cardinality of the outputs is relatively high. Peter Bell 0001, Steve Renals |
INTERSPEECH | 2 |
| 2015 | Prosodically-enhanced recurrent neural network language modelsabstractRecurrent neural network language models have been shown to consistently reduce the word error rates (WERs) of large vocabulary speech recognition tasks. In this work we propose to enhance the RNNLMs with prosodic features computed using the context of the current word. Since it is plausible to compute the prosody features at the word and syllable level we have trained the models on prosody features computed at both these levels. To investigate the effectiveness of proposed models we report perplexity and WER for two speech recognition tasks, Switchboard and TED. We observed substantial improvements in perplexity and small improvements in WER. Index Terms: RNNLMs, 3-gram, prosody features, pause duration, duration of the word, syllable duration, syllable F0, GMMHMM, DNN-HMM, Switchboard conversations and TED lectures Siva Reddy Gangireddy, Steve Renals, Yoshihiko Nankaku, Akinobu Lee |
INTERSPEECH | 2 |
| 2015 | Feature-space speaker adaptation for probabilistic linear discriminant analysis acoustic modelsabstractProbabilistic linear discriminant analysis (PLDA) acoustic models extend Gaussian mixture models by factorizing the acoustic variability using state-dependent and observation-dependent variables. This enables the use of higher dimensional acoustic features, and the capture of intra-frame feature corre-lations. In this paper, we investigate the estimation of speaker adaptive feature-space (constrained) maximum likelihood lin-ear regression transforms from PLDA-based acoustic models. This feature-space speaker transformation estimation approach is potentially very useful due to the ability of PLDA acoustic models to use different types of acoustic features, for example applying these transforms to deep neural network (DNN) acous-tic models for cross adaptation. We evaluated the approach on the Switchboard corpus, and observe significant word error re-duction by using both the mel-frequency cepstral coefficients and DNN bottleneck features. Index Terms: speech recognition, probabilistic linear discrim-inant analysis, speaker adaptation, fMLLR, PLDA Liang Lu 0001, Steve Renals |
INTERSPEECH | 2 |
| 2015 | A study of the recurrent neural network encoder-decoder for large vocabulary speech recognitionabstractDeep neural networks have advanced the state-of-the-art in automatic speech recognition, when combined with hidden Markov models (HMMs). Recently there has been interest in using systems based on recurrent neural networks (RNNs) to perform sequence modelling directly, without the requirement of an HMM superstructure. In this paper, we study the RNN encoder-decoder approach for large vocabulary end-toend speech recognition, whereby an encoder transforms a sequence of acoustic vectors into a sequence of feature representations, from which a decoder recovers a sequence of words. We investigated this approach on the Switchboard corpus using a training set of around 300 hours of transcribed audio data. Without the use of an explicit language model or pronunciation lexicon, we achieved promising recognition accuracy, demonstrating that this approach warrants further investigation. Index Terms: end-to-end speech recognition, deep neural networks, recurrent neural networks, encoder-decoder. Liang Lu 0001, Xingxing Zhang 0002, Kyunghyun Cho, Steve Renals |
INTERSPEECH | 4 |
| 2015 | Structured output layer with auxiliary targets for context-dependent acoustic modellingabstractIn previous work we have introduced a multi-task training tech-nique for neural network acoustic modelling, in which context-dependent and context-independent targets are jointly learned. In this paper, we extend the approach by structuring the out-put layer such that the context-dependent outputs are depen-dent on the context-independent outputs, thus using the context-independent predictions at run-time. We have also investigated the applicability of this idea to unsupervised speaker adapta-tion as an approach to overcome the data sparsity issues that comes to the fore when estimating systems with a large num-ber of context-dependent states, when data is limited. We have experimented with various amounts of training material (from 10 to 300 hours) and find the proposed techniques are particu-larly well suited to data-constrained conditions allowing to bet-ter utilise large context-dependent state-clustered trees. Exper-imental results are reported for large vocabulary speech recog-nition using the Switchboard and TED corpora. Index Terms: multitask learning, structured output layer, adap-tation, deep neural networks Pawel Swietojanski, Peter Bell 0001, Steve Renals |
INTERSPEECH | 3 |
| 2015 | A study of speaker adaptation for DNN-based speech synthesisabstractA major advantage of statistical parametric speech synthesis (SPSS) over unit-selection speech synthesis is its adaptability and controllability in changing speaker characteristics and speaking style.Recently, several studies using deep neural networks (DNNs) as acoustic models for SPSS have shown promising results.However, the adaptability of DNNs in SPSS has not been systematically studied.In this paper, we conduct an experimental analysis of speaker adaptation for DNN-based speech synthesis at different levels.In particular, we augment a low-dimensional speaker-specific vector with linguistic features as input to represent speaker identity, perform model adaptation to scale the hidden activation weights, and perform a feature space transformation at the output layer to modify generated acoustic features.We systematically analyse the performance of each individual adaptation technique and that of their combinations.Experimental results confirm the adaptability of the DNN, and listening tests demonstrate that the DNN can achieve significantly better adaptation performance than the hidden Markov model (HMM) baseline in terms of naturalness and speaker similarity. Zhizheng Wu 0001, Pawel Swietojanski, Christophe Veaux, Steve Renals, Simon King 0001 |
INTERSPEECH | 4 |
| 2014 | Neural net word representations for phrase-break prediction without a part of speech taggerabstractThe use of shared projection neural nets of the sort used in language modelling is proposed as a way of sharing parameters between multiple text-to-speech system components. We experiment with pretraining the weights of such a shared projection on an auxiliary language modelling task and then apply the resulting word representations to the task of phrase-break prediction. Doing so allows us to build phrase-break predictors that rival conventional systems without any reliance on conventional knowledge-based resources such as part of speech taggers. Oliver Watts, Siva Reddy Gangireddy, Junichi Yamagishi, Simon King 0001, Steve Renals, Adriana Cornelia Stan, Mircea Giurgiu |
ICASSP | 5 |
| 2014 | Cross-lingual adaptation with multi-task adaptive networksabstractPosterior-based or bottleneck features derived from neural net-works trained on out-of-domain data may be successfully ap-plied to improve speech recognition performance when data is scarce for the target domain or language. In this paper we com-bine this approach with the use of a hierarchical deep neural net-work (DNN) network structure – which we term a multi-level adaptive network (MLAN) – and the use of multitask learning. We have applied the technique to cross-lingual speech recog-nition experiments on recordings of TED talks and European Parliament sessions in English (source language) and German (target language). We demonstrate that the proposed method can lead to improvements over standard methods, even when the quantity of training data for the target language is relatively high. When the complete method is applied, we achieve rela-tiveWER reductions of around 13 % compared to a monolingual hybrid DNN baseline. Index Terms: deep neural network, multilevel adaptive net- Peter Bell 0001, Joris Driesen, Steve Renals |
INTERSPEECH | 3 |
| 2014 | Automated production of true-cased punctuated subtitles for weather and news broadcasts
Joris Driesen, Alexandra Birch, Simon Grimsey, Saeid Safarfashandi, Juliet Gauthier, Matt Simpson, Steve Renals |
INTERSPEECH | 7 |
| 2014 | Feed forward pre-training for recurrent neural network language modelsabstractThe recurrent neural network language model (RNNLM) has been demonstrated to consistently reduce perplexities and au-tomatic speech recognition (ASR) word error rates across a variety of domains. In this paper we propose a pre-training method for the RNNLM, by sharing the output weights of the feed forward neural network language model (NNLM) with the RNNLM. This is accomplished by first fine-tuning the weights of the NNLM, which are then used to initialise the output weights of an RNNLM with the same number of hidden units. We have carried out text-based experiments on the Penn Tree-bank Wall Street Journal data, and ASR experiments on the TED talks data used in the International Workshop on Spoken Language Translation (IWSLT) evaluation campaigns. Across the experiments, we observe small improvements in perplexity and ASR word error rate. Siva Reddy Gangireddy, Fergus R. McInnes, Steve Renals |
INTERSPEECH | 3 |
| 2014 | Incorporating lexical and prosodic information at different levels for meeting summarizationabstractThis paper investigates how prosodic features can be used to augment lexical features for meeting summarization. Auto-matic detection of summary-worthy content using non-lexical features, like prosody, has generally focused on features cal-culated over dialogue acts. However, a salient role of prosody is to distinguish important words within utterances. To exam-ine whether including more fine grained prosodic information can help extractive summarization, we perform experiments incorporating lexical and prosodic features at different levels. For ICSI and AMI meeting corpora, we find that combining prosodic and lexical features at a lower level has better AUROC performance than adding in prosodic features derived over di-alogue acts. ROUGE F-scores also show the same pattern for the ICSI data. However, the differences are less clear for the AMI data where the range of scores is much more compressed. In order to understand the relationship between the generated summaries and differences in standard measures, we look at the distribution of extracted content over meeting as well as sum-mary redundancy. We find that summaries based on dialogue act level prosody better reflect the amount of human annotated summary content in meeting segments, while summaries de-rived from prosodically augmented lexical features exhibit less redundancy. Index Terms: meeting summarization, prosody, dialogue. 1. Catherine Lai, Steve Renals |
INTERSPEECH | 2 |
| 2014 | Probabilistic linear discriminant analysis with bottleneck features for speech recognitionabstractWe have recently proposed a new acoustic model based on probabilistic linear discriminant analysis (PLDA) which enjoys the flexibility of using higher dimensional acoustic features, and is more capable to capture the intra-frame feature correlations. In this paper, we investigate the use of bottleneck features obtained from a deep neural network (DNN) for the PLDA-based acoustic model. Experiments were performed on the Switchboard dataset — a large vocabulary conversational telephone speech corpus. We observe significant word error reduction by using the bottleneck features. In addition, we have also compared the PLDA-based acoustic model to three others using Gaussian mixture models (GMMs), subspace GMMs and hybrid deep neural networks (DNNs), and PLDA can achieve comparable or slightly higher recognition accuracy from our experiments. Liang Lu 0001, Steve Renals |
INTERSPEECH | 2 |
| 2014 | Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic modelsabstractThis paper proposes a simple yet effective model-based neural network speaker adaptation technique that learns speaker-specific hidden unit contributions given adaptation data, without requiring any form of speaker-adaptive training, or labelled adaptation data. An additional amplitude parameter is defined for each hidden unit; the amplitude parameters are tied for each speaker, and are learned using unsupervised adaptation. We conducted experiments on the TED talks data, as used in the International Workshop on Spoken Language Translation (IWSLT) evaluations. Our results indicate that the approach can reduce word error rates on standard IWSLT test sets by about 8-15% relative compared to unadapted systems, with a further reduction of 4-6% relative when combined with feature-space maximum likelihood linear regression (fMLLR). The approach can be employed in most existing feed-forward neural network architectures, and we report results using various hidden unit activation functions: sigmoid, maxout, and rectifying linear units (ReLU). Pawel Swietojanski, Steve Renals |
SLT | 2 |
| 2014 | Probabilistic Linear Discriminant Analysis for Acoustic ModelingabstractIn this letter, we propose a new acoustic modeling approach for automatic speech recognition based on probabilistic linear discriminant analysis (PLDA), which is used to model the state density function for the standard hidden Markov models (HMMs). Unlike the conventional Gaussian mixture models (GMMs) where the correlations are weakly modelled by using the diagonal covariance matrices, PLDA captures the correlations of feature vector in subspaces without vastly expanding the model. It also allows the usage of high dimensional feature input, and therefore is more flexible to make use of different type of acoustic features. We performed the preliminary experiments on the Switchboard corpus, and demonstrated the feasibility of this acoustic model. Liang Lu 0001, Steve Renals |
IEEE Signal Process. Lett. | 2 |
| 2014 | Convolutional Neural Networks for Distant Speech RecognitionabstractWe investigate convolutional neural networks (CNNs) for large vocabulary distant speech recognition, trained using speech recorded from a single distant microphone (SDM) and multiple distant microphones (MDM). In the MDM case we explore a beamformed signal input representation compared with the direct use of multiple acoustic channels as a parallel input to the CNN. We have explored different weight sharing approaches, and propose a channel-wise convolution with two-way pooling. Our experiments, using the AMI meeting corpus, found that CNNs improve the word error rate (WER) by 6.5% relative compared to conventional deep neural network (DNN) models and 15.7% over a discriminatively trained Gaussian mixture model (GMM) baseline. For cross-channel CNN training, the WER improves by 3.5% relative over the comparable DNN structure. Compared with the best beamformed GMM system, cross-channel convolution reduces the WER by 9.7% relative, and matches the accuracy of a beamformed DNN. Pawel Swietojanski, Arnab Ghoshal, Steve Renals |
IEEE Signal Process. Lett. | 3 |
| 2014 | Editorial: Expanding the Technical Reach of our TransactionsabstractWe thank the strong support of both IEEE Signal Processing Society and the ACM Publication Boards for this successful merger.IEEE's TASLP is a very well-established publication, strong in both quality and quantity.It is closely linked to ICASSP and to a number of workshops, such as ASRU and SLT.The language area is a relatively recent addition to TASLP, incorporated in 2006.ACM's TSLP is a more recent publication.The quality has been very high, but the quantity has only sustained a quarterly publication.There is no ACM Special Interest Group (SIG) or conference connection in this area, so it has been more difficult to maintain a direct link to the research community.One of the main motivations for the merger is that IEEE's TASLP does not yet have a strong profile in the language processing community, and this is reflected in the submissions received and the composition of the Editorial Board.ACM's TSLP has a stronger profile and Editorial Board membership in this area.Thus, it is clear that a joint transactions will be stronger than either publication on its own.For several years, the IEEE Signal Processing Society has recognized the importance of information processing in the work of a wide range of researchers within the Society beyond the traditional scope of signal processing.This led to the technical scope of the IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING being expanded to include Language Processing in 2006.While serving as editor of the IEEE SIGNAL PROCESSING MAGAZINE, one of the authors wrote in 2008 and 2010 two editorials [1], [2] that elaborated on the need for expanding the technical reach of signal processing by including the new "understanding" or "interpretation" component of signals consisting of language/text and bio-sequence data.They are both symbolic in nature, which were outside of the traditional definition of "signal" with numerical values in nature.This editorial on the formation of the joint IEEE/ACM TASLP, which highlights the importance of text or written language processing, is a concrete embodiment of the goal of expanding the technical reach of signal processing.As part of the merger of the two transactions, we have taken the opportunity to restructure the editorial board.In addition to the role of Editor-in-Chief, there will be six Senior Area Editors to advise the Editor-in-Chief across the full range of topics covered by the merged journal. Li Deng 0001, Steve Renals, Marcello Federico, Mari Ostendorf |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Cross-Lingual Subspace Gaussian Mixture Models for Low-Resource Speech RecognitionabstractThis paper studies cross-lingual acoustic modeling in the context of subspace Gaussian mixture models (SGMMs). SGMMs factorize the acoustic model parameters into a set that is globally shared between all the states of a hidden Markov model (HMM) and another that is specific to the HMM states. We demonstrate that the SGMM global parameters are transferable between languages, particularly when the parameters are trained multilingually. As a result, acoustic models may be trained using limited amounts of transcribed audio by borrowing the SGMM global parameters from one or more source languages, and only training the state-specific parameters on the target language audio. Model regularization using ℓ1-norm penalty is shown to be particularly effective at avoiding overtraining and leading to lower word error rates. We investigate maximum a posteriori (MAP) adaptation of subspace parameters in order to reduce the mismatch between the SGMM global parameters of the source and target languages. In addition, monolingual and cross-lingual speaker adaptive training is used to reduce the model variance introduced by speakers. We have systematically evaluated these techniques by experiments on the GlobalPhone corpus. Liang Lu 0001, Arnab Ghoshal, Steve Renals |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Lightly supervised automatic subtitling of weather forecastsabstractSince subtitling television content is a costly process, there are large potential advantages to automating it, using automatic speech recognition (ASR). However, training the necessary acoustic models can be a challenge, since the available training data usually lacks verbatim orthographic transcriptions. If there are approximate transcriptions, this problem can be overcome using light supervision methods. In this paper, we perform speech recognition on broadcasts of Weatherview, BBC's daily weather report, as a first step towards automatic subtitling. For training, we use a large set of past broadcasts, using their manually created subtitles as approximate transcriptions. We discuss and and compare two different light supervision methods, applying them to this data. The best training set finally obtained with these methods is used to create a hybrid deep neural network-based recognition system, which yields high recognition accuracies on three separate Weatherview evaluation sets. Joris Driesen, Steve Renals |
ASRU | 2 |
| 2013 | Acoustic data-driven pronunciation lexicon for large vocabulary speech recognitionabstractSpeech recognition systems normally use handcrafted pronunciation lexicons designed by linguistic experts. Building and maintaining such a lexicon is expensive and time consuming. This paper concerns automatically learning a pronunciation lexicon for speech recognition. We assume the availability of a small seed lexicon and then learn the pronunciations of new words directly from speech that is transcribed at word-level. We present two implementations for refining the putative pronunciations of new words based on acoustic evidence. The first one is an expectation maximization (EM) algorithm based on weighted finite state transducers (WFSTs) and the other is its Viterbi approximation. We carried out experiments on the Switchboard corpus of conversational telephone speech. The expert lexicon has a size of more than 30,000 words, from which we randomly selected 5,000 words to form the seed lexicon. By using the proposed lexicon learning method, we have significantly improved the accuracy compared with a lexicon learned using a grapheme-to-phoneme transformation, and have obtained a word error rate that approaches that achieved using a fully handcrafted lexicon. Liang Lu 0001, Arnab Ghoshal, Steve Renals |
ASRU | 3 |
| 2013 | Hybrid acoustic models for distant and multichannel large vocabulary speech recognitionabstractWe investigate the application of deep neural network (DNN)-hidden Markov model (HMM) hybrid acoustic models for far-field speech recognition of meetings recorded using microphone arrays. We show that the hybrid models achieve significantly better accuracy than conventional systems based on Gaussian mixture models (GMMs). We observe up to 8% absolute word error rate (WER) reduction from a discriminatively trained GMM baseline when using a single distant microphone, and between 4-6% absolute WER reduction when using beamforming on various combinations of array channels. By training the networks on audio from multiple channels, we find the networks can recover significant part of accuracy difference between the single distant microphone and beamformed configurations. Finally, we show that the accuracy of a network recognising speech from a single distant microphone can approach that of a multi-microphone setup by training with data from other microphones. Pawel Swietojanski, Arnab Ghoshal, Steve Renals |
ASRU | 3 |
| 2013 | Multi-level adaptive networks in tandem and hybrid ASR systemsabstractIn this paper we investigate the use of Multi-level adaptive networks (MLAN) to incorporate out-of-domain data when training large vocabulary speech recognition systems. In a set of experiments on multi-genre broadcast data and on TED lecture recordings we present results using of out-of-domain features in a hybrid DNN system and explore tandem systems using a variety of input acoustic features. Our experiments indicate using the MLAN approach in both hybrid and tandem systems results in consistent reductions in word error rate of 5-10% relative. Peter Bell 0001, Pawel Swietojanski, Steve Renals |
ICASSP | 3 |
| 2013 | Multilingual training of deep neural networksabstractWe investigate multilingual modeling in the context of a deep neural network (DNN) - hidden Markov model (HMM) hybrid, where the DNN outputs are used as the HMM state likelihoods. By viewing neural networks as a cascade of feature extractors followed by a logistic regression classifier, we hypothesise that the hidden layers, which act as feature extractors, will be transferable between languages. As a corollary, we propose that training the hidden layers on multiple languages makes them more suitable for such cross-lingual transfer. We experimentally confirm these hypotheses on the GlobalPhone corpus using seven languages from three different language families: Germanic, Romance, and Slavic. The experiments demonstrate substantial improvements over a monolingual DNN-HMM hybrid baseline, and hint at avenues of further exploration. Arnab Ghoshal, Pawel Swietojanski, Steve Renals |
ICASSP | 3 |
| 2013 | Revisiting hybrid and GMM-HMM system combination techniquesabstractIn this paper we investigate techniques to combine hybrid HMM-DNN (hidden Markov model - deep neural network) and tandem HMM-GMM (hidden Markov model - Gaussian mixture model) acoustic models using: (1) model averaging, and (2) lattice combination with Minimum Bayes Risk decoding. We have performed experiments on the “TED Talks” task following the protocol of the IWSLT-2012 evaluation. Our experimental results suggest that DNN-based and GMM-based acoustic models are complementary, with error rates being reduced by up to 8% relative when the DNN and GMM systems are combined at model-level in a multi-pass automatic speech recognition (ASR) system. Additionally, further gains were obtained by combining model-averaged lattices with the one obtained from baseline systems. Pawel Swietojanski, Arnab Ghoshal, Steve Renals |
ICASSP | 3 |
| 2013 | Recognition of overlapping speech using digital MEMS microphone arraysabstractThis paper presents a new corpus comprising single and overlapping speech recorded using digital MEMS and analogue microphone arrays. In addition to this, the paper presents results from speech separation and recognition experiments on this data. The corpus is a reproduction of the multi-channel Wall Street Journal audio-visual corpus (MC-WSJAV), containing recorded speech in both a meeting room and an anechoic chamber using two different microphone types as well as two different array geometries. The speech separation and speech recognition experiments were performed using SRP-PHAT-based speaker localisation, superdirective beamforming and multiple post-processing schemes, such as residual echo suppression and binary masking. Our simple, cMLLR-based recognition system matches the performance of state-of-the-art ASR systems on the single speaker task and outperforms them on overlapping speech. The corpus will be made publicly available via the LDC in spring 2013. Erich Zwyssig, Friedrich Faubel, Steve Renals, Mike Lincoln |
ICASSP | 3 |
| 2013 | A lecture transcription system combining neural network acoustic and language modelsabstractThis paper presents a new system for automatic transcription of lectures. The system combines a number of novel features, including deep neural network acoustic models using multi-level adaptive networks to incorporate out-of-domain information, and factored recurrent neural network language models. We demonstrate that the system achieves large improvements on the TED lecture transcription task from the 2012 IWSLT evaluation - our results are currently the best reported on this task, showing an relative WER reduction of more than 16% compared to the closest competing system from the evaluation. Peter Bell 0001, Hitoshi Yamamoto, Pawel Swietojanski, Youzheng Wu, Fergus R. McInnes, Chiori Hori, Steve Renals |
INTERSPEECH | 7 |
| 2013 | Detecting summarization hot spots in meetings using group level involvement and turn-taking featuresabstractIn this paper we investigate how participant involvement and turn-taking features relate to extractive summarization of meeting dialogues. In particular, we examine whether automatically derived measures of group level involvement, like participation equality and turn-taking freedom, can help detect where summarization relevant meeting segments will be. Results show that classification using turn-taking features performed better than the majority class baseline for data from both AMI and ICSI meeting corpora in identifying whether meeting segments contain extractive summary dialogue acts. The feature based approach also provided better recall than using manual ICSI involvement hot spot annotations. Turn-taking features were additionally found to be predictive of the amount of extractive summary content in a segment. In general, we find that summary content decreases with higher participation equality and overlap, while it increases with the number of very short utterances. Differences in results between the AMI and ICSI data sets suggest how group participatory structure can be used to understand what makes meetings easy or difficult to summarize. Index Terms: Turn-taking, involvement, hot spots, summarization, meetings, dialogue Catherine Lai, Jean Carletta, Steve Renals |
INTERSPEECH | 3 |
| 2013 | Noise adaptive training for subspace Gaussian mixture modelsabstractNoise adaptive training (NAT) is an effective approach to normalise the environmental distortions in the training data. This paper investigates the model-based NAT scheme using joint uncertainty decoding (JUD) for subspace Gaussian mixture models (SGMMs). A typical SGMM acoustic model has much larger number of surface Gaussian components, which makes it computationally infeasible to compensate each Gaussian explicitly. JUD tackles the problem by sharing the compensation parameters among the Gaussians and hence reduces the computational and memory demands. For noise adaptive training, JUD is reformulated into a generative model, which leads to an efficient expectation-maximisation (EM) based algorithm to update the SGMM acoustic model parameters. We evaluated the SGMMs with NAT on the Aurora 4 database, and obtained higher recognition accuracy compared to systems without adaptive training. Index Terms: adaptive training, noise robustness, joint uncertainty decoding, subspace Gaussian mixture models Liang Lu 0001, Arnab Ghoshal, Steve Renals |
INTERSPEECH | 3 |
| 2013 | Joint Uncertainty Decoding for Noise Robust Subspace Gaussian Mixture ModelsabstractJoint uncertainty decoding (JUD) is a model-based noise compensation technique for conventional Gaussian Mixture Model (GMM) based speech recognition systems. Unlike vector Taylor series (VTS) compensation which operates on the individual Gaussian components in an acoustic model, JUD clusters the Gaussian components into a smaller number of classes, sharing the compensation parameters for the set of Gaussians in a given class. This significantly reduces the computational cost. In this paper, we investigate noise compensation for subspace Gaussian mixture model (SGMM) based speech recognition systems using JUD. The total number of Gaussian components in an SGMM is typically very large. Therefore direct compensation of the individual Gaussian components, as performed by VTS, is computationally expensive. In this paper we show that JUD-based noise compensation can be successfully applied to SGMMs in a computationally efficient way. We evaluate the JUD/SGMM technique on the standard Aurora 4 corpus. Our experimental results indicate that the JUD/SGMM system results in lower word error rates compared with a conventional GMM system with either VTS-based or JUD-based noise compensation. Liang Lu 0001, K. K. Chin, Arnab Ghoshal, Steve Renals |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Maximum a posteriori adaptation of subspace Gaussian mixture models for cross-lingual speech recognitionabstractThis paper concerns cross-lingual acoustic modeling in the case when there are limited target language resources. We build on an approach in which a subspace Gaussian mixture model (SGMM) is adapted to the target language by reusing the globally shared parameters estimated from out-of-language training data. In current cross-lingual systems, these parameters are fixed when training the target system, which can give rise to a mismatch between the source and target systems. We investigate a maximum a posteriori (MAP) adaptation approach to alleviate the potential mismatch. In particular, we focus on the adaptation of phonetic subspace parameters using a matrix variate Gaussian prior distribution. Experiments on the GlobalPhone corpus using the MAP adaptation approach results in word error rate reductions, compared with the cross-lingual base-line systems and systems updated using maximum likelihood, for training conditions with 1 hour and 5 hours of target language data. Liang Lu 0001, Arnab Ghoshal, Steve Renals |
ICASSP | 3 |
| 2012 | On the effect of snr and superdirective beamforming in speaker diarisation in meetingsabstractThis paper examines the effect of sensor performance on speaker diarisation in meetings and investigates the use of more advanced beamforming techniques, beyond the typically employed delay-sum beamformer, for mitigating the effects of poorer sensor performance. We present superdirective beamforming and investigate how different time difference of arrival (TDOA) smoothing and beamforming techniques influence the performance of state-of-the-art diarisation systems. We produced and transcribed a new corpus of meetings recorded in the instrumented meeting room using a high SNR analogue and a newly developed low SNR digital MEMS microphone array (DMMA.2). This research demonstrates that TDOA smoothing has a significant effect on the diarisation error rate and that simple noise reduction and beamforming schemes suffice to overcome audio signal degradation due to the lower SNR of modern MEMS microphones. Erich Zwyssig, Steve Renals, Mike Lincoln |
ICASSP | 2 |
| 2012 | Determining the number of speakers in a meeting using microphone array featuresabstractThe accuracy of speaker diarisation in meetings relies heavily on determining the correct number of speakers. In this paper we present a novel algorithm based on time difference of arrival (TDOA) features that aims to find the correct number of active speakers in a meeting and thus aid the speaker segmentation and clustering process. With our proposed method the microphone array TDOA values and known geometry of the array are used to calculate a speaker matrix from which we determine the correct number of active speakers with the aid of the Bayesian information criterion (BIC). In addition, we analyse several well-known voice activity detection (VAD) algorithms and verified their fitness for meeting recordings. Experiments were performed using the NIST RT06, RT07 and RT09 data sets, and resulted in reduced error rates compared with BIC-based approaches. Erich Zwyssig, Steve Renals, Mike Lincoln |
ICASSP | 2 |
| 2012 | Noise Compensation for Subspace Gaussian Mixture ModelsabstractJoint uncertainty decoding (JUD) is an effective model-based noise compensation technique for conventional Gaussian mix-ture model (GMM) based speech recognition systems. In this paper, we apply JUD to subspace Gaussian mixture model (SGMM) based acoustic models. The total number of Gaus-sians in the SGMM acoustic model is usually much larger than for conventional GMMs, which limits the application of approaches which explicitly compensate each Gaussian, such as vector Taylor series (VTS). However, by clustering the Gaussian components into a number of regression classes, JUD-based noise compensation can be successfully applied to SGMM systems. We evaluate the JUD/SGMM technique us-ing the Aurora 4 corpus, and the experimental results indicated that it is more accurate than conventional GMM-based systems using either VTS or JUD noise compensation. 1. Liang Lu 0001, K. K. Chin, Arnab Ghoshal, Steve Renals |
INTERSPEECH | 4 |
| 2012 | Ultrax: An Animated Midsagittal Vocal Tract Display for Speech TherapyabstractSpeech sound disorders (SSD) are the most common communication impairment in childhood, and can hamper social development and learning. Current speech therapy interventions rely predominantly on the auditory skills of the child, as little technology is available to assist in diagnosis and therapy of SSDs. Realtime visualisation of tongue movements has the potential to bring enormous benefit to speech therapy. Ultrasound scanning offers this possibility, although its display may be hard to interpret. Our ultimate goal is to exploit ultrasound to track tongue movement, while displaying a simplified, diagrammatic vocal tract that is easier for the user to interpret. In this paper, we outline a general approach to this problem, combining a latent space model with a dimensionality reducing model of vocal tract shapes. We assess the feasibility of this approach using magnetic resonance imaging (MRI) scans to train a model of vocal tract shapes, which is animated using electromagnetic articulography (EMA) data from the same speaker. Index Terms: Ultrasound, speech therapy, vocal tract visualisation 1. Korin Richmond, Steve Renals |
INTERSPEECH | 2 |
| 2012 | Deep Architectures for Articulatory InversionabstractWe implement two deep architectures for the acoustic-articulatory inversion mapping problem: a deep neural network and a deep trajectory mixture density network. We find that in both cases, deep architectures produce more accurate predic-tions than shallow architectures and that this is due to the higher expressive capability of a deep model and not a consequence of adding more adjustable parameters. We also find that a deep trajectory mixture density network is able to obtain better in-version accuracies than smoothing the results of a deep neural network. Our best model obtained an average root mean square error of 0.885 mm on the MNGU0 test dataset. Index Terms: Articulatory inversion, deep neural network, deep belief network, deep regression network, pretraining Benigno Uria, Iain Murray 0001, Steve Renals, Korin Richmond |
INTERSPEECH | 3 |
| 2012 | Transcription of multi-genre media archives using out-of-domain dataabstractWe describe our work on developing a speech recognition system for multi-genre media archives. The high diversity of the data makes this a challenging recognition task, which may benefit from systems trained on a combination of in-domain and out-of-domain data. Working with tandem HMMs, we present Multi-level Adaptive Networks (MLAN), a novel technique for incorporating information from out-of-domain posterior features using deep neural networks. We show that it provides a substantial reduction in WER over other systems, with relative WER reductions of 15% over a PLP baseline, 9% over in-domain tandem features and 8% over the best out-of-domain tandem features. Peter Bell 0001, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanhua Long, Steve Renals, Pawel Swietojanski, Philip C. Woodland |
SLT | 6 |
| 2012 | Unsupervised cross-lingual knowledge transfer in DNN-based LVCSRabstractWe investigate the use of cross-lingual acoustic data to initialise deep neural network (DNN) acoustic models by means of unsupervised restricted Boltzmann machine (RBM) pre-training. DNNs for German are pretrained using one or all of German, Portuguese, Spanish and Swedish. The DNNs are used in a tandem configuration, where the network outputs are used as features for a hidden Markov model (HMM) whose emission densities are modeled by Gaussian mixture models (GMMs), as well as in a hybrid configuration, where the network outputs are used as the HMM state likelihoods. The experiments show that unsupervised pretraining is more crucial for the hybrid setups, particularly with limited amounts of transcribed training data. More importantly, unsupervised pretraining is shown to be language-independent. Pawel Swietojanski, Arnab Ghoshal, Steve Renals |
SLT | 3 |
| 2012 | Special issue on searching speechabstractNo abstract available. Martha A. Larson, Franciska de Jong, Wessel Kraaij, Steve Renals |
ACM Trans. Inf. Syst. | 4 |
| 2011 | Regularized subspace Gaussian mixture models for cross-lingual speech recognitionabstractWe investigate cross-lingual acoustic modelling for low resource languages using the subspace Gaussian mixture model (SGMM). We assume the presence of acoustic models trained on multiple source languages, and use the global subspace parameters from those models for improved modelling in a target language with limited amounts of transcribed speech. Experiments on the GlobalPhone corpus using Spanish, Portuguese, and Swedish as source languages and German as target language (with 1 hour and 5 hours of transcribed audio) show that multilingually trained SGMM shared parameters result in lower word error rates (WERs) than using those from a single source language. We also show that regularizing the estimation of the SGMM state vectors by penalizing their ℓ1-norm help to overcome numerical instabilities and lead to lower WER. Liang Lu 0001, Arnab Ghoshal, Steve Renals |
ASRU | 3 |
| 2011 | HMM-based speech synthesiser using the LF-model of the glottal sourceabstractA major factor which causes a deterioration in speech quality in HMM-based speech synthesis is the use of a simple delta pulse signal to generate the excitation of voiced speech. This paper sets out a new approach to using an acoustic glottal source model in HMM-based synthesisers instead of the traditional pulse signal. The goal is to improve speech quality and to better model and transform voice characteristics. We have found the new method decreases buzziness and also improves prosodic modelling. A perceptual evaluation has supported this finding by showing a 55.6% preference for the new system, as against the baseline. This improvement, while not being as significant as we had initially expected, does encourage us to work on developing the proposed speech synthesiser further. João P. Cabral, Steve Renals, Junichi Yamagishi, Korin Richmond |
ICASSP | 2 |
| 2011 | Regularized Subspace Gaussian Mixture Models for Speech RecognitionabstractSubspace Gaussian mixture models (SGMMs) provide a compact representation of the Gaussian parameters in an acoustic model, but may still suffer from over-fitting with insufficient training data. In this letter, the SGMM state parameters are estimated using a penalized maximum-likelihood objective, based on l1and l2regularization, as well as their combination, referred to as the elastic net, for robust model estimation. Experiments on the 5000-word Wall Street Journal transcription task show word error rate reduction and improved model robustness with regularization. Liang Lu 0001, Arnab Ghoshal, Steve Renals |
IEEE Signal Process. Lett. | 3 |
| 2010 | Power law discounting for n-gram language modelsabstractWe present an approximation to the Bayesian hierarchical Pitman-Yor process language model which maintains the power law distribution over word tokens, while not requiring a computationally expensive approximate inference process. This approximation, which we term power law discounting, has a similar computational complexity to interpolated and modified Kneser-Ney smoothing. We performed experiments on meeting transcription using the NIST RT06s evaluation data and the AMI corpus, with a vocabulary of 50,000 words and a language model training set of up to 211 million words. Our results indicate that power law discounting results in statistically significant reductions in perplexity and word error rate compared to both interpolated and modified Kneser-Ney smoothing, while producing similar results to the hierarchical Pitman-Yor process language model. Songfang Huang, Steve Renals |
ICASSP | 2 |
| 2010 | A digital microphone array for distant speech recognitionabstractIn this paper, the design, implementation and testing of a digital microphone array is presented. The array uses digital MEMS microphones which integrate the microphone, amplifier and analogue to digital converter on a single chip in place of the analogue microphones and external audio interfaces currently used. The device has the potential to be smaller, cheaper and more flexible than typical analogue arrays, however the effect on speech recognition performance of using digital microphones is as yet unknown. In order to evaluate the effect, an analogue array and the new digital array are used to simultaneously record test data for a speech recognition experiment. Initial results employing no adaptation show that performance using the digital array is significantly worse (14% absolute WER) than the analogue device. Subsequent experiments using MLLR and CMLLR channel adaptation reduce this gap, and employing MLLR for both channel and speaker adaptation reduces the difference between the arrays to 4.5% absolute WER. Erich Zwyssig, Mike Lincoln, Steve Renals |
ICASSP | 3 |
| 2010 | Augmentation of adaptation dataabstractLinear regression based speaker adaptation approaches can improve Automatic Speech Recognition (ASR) accuracy significantly for a target speaker. However, when the available adaptation data is limited to a few seconds, the accuracy of the speaker adapted models is often worse compared with speaker independent models. In this paper, we propose an approach to select a set of reference speakers acoustically close to the target speaker whose data can be used to augment the adaptation data. To determine the acoustic similarity of two speakers, we propose a distance metric based on transforming sample points in the acoustic space with the regression matrices of the two speakers. We show the validity of this approach through a speaker identification task. ASR results on SCOTUS and AMI corpora with limited adaptation data of 10 to 15 seconds augmented by data from selected reference speakers show a significant improvement in Word Error Rate over speaker independent and speaker adapted models. Ravichander Vipperla, Steve Renals, Joe Frankel |
INTERSPEECH | 2 |
| 2010 | Invited Talk: Recognition and Understanding of Meetings
Steve Renals |
HLT-NAACL | 1 |
| 2010 | Evaluation of a hierarchical reinforcement learning spoken dialogue system
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira |
Comput. Speech Lang. | 2 |
| 2010 | Hierarchical Bayesian Language Models for Conversational Speech RecognitionabstractTraditional$n$-gram language models are widely used in state-of-the-art large vocabulary speech recognition systems. This simple model suffers from some limitations, such as overfitting of maximum-likelihood estimation and the lack of rich contextual knowledge sources. In this paper, we exploit a hierarchical Bayesian interpretation for language modeling, based on a nonparametric prior called Pitman–Yor process. This offers a principled approach to language model smoothing, embedding the power-law distribution for natural language. Experiments on the recognition of conversational speech in multiparty meetings demonstrate that by using hierarchical Bayesian language models, we are able to achieve significant reductions in perplexity and word error rate. Songfang Huang, Steve Renals |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | A parallel training algorithm for hierarchical pitman-yor process language modelsabstractThe Hierarchical Pitman Yor Process Language Model (HPYLM) is a Bayesian language model based on a nonparametric prior, the Pitman-Yor Process. It has been demonstrated, both theoretically and practically, that the HPYLM can provide better smoothing for language modeling, compared with state-of-the-art approaches such as interpolated Kneser-Ney and modified Kneser-Ney smoothing. However, estimation of Bayesian language models is expensive in terms of both computation time and memory; the inference is approximate and requires a number of iterations to converge. In this paper, we present a parallel training algorithm for the HPYLM, which enables the approach to be applied in the context of automatic speech recognition, using large training corpora with large vocabularies. We demonstrate the effectiveness of the proposed algorithm by estimating language models from corpora for meeting transcription containing over 200 million words, and observe significant reductions in perplexity and word error rate. Index Terms: language model, Pitman-Yor processes, hierarchical Bayesian models, parallel training, meetings Songfang Huang, Steve Renals |
INTERSPEECH | 2 |
| 2009 | Age recognition for spoken dialogue systems: do we need it?abstractWhen deciding whether to adapt relevant aspects of the system to the particular needs of older users, spoken dialogue systems often rely on automatic detection of chronological age. In this paper, we show that vocal ageing as measured by acoustic features is an unreliable indicator of the need for adaptation. Simple lexical features greatly improve the prediction of both relevant aspects of cognition and interactions style. Lexical features also boost age group prediction. We suggest that adaptation should be based on observed behaviour, not on chronological age, unless it is not feasible to build classifiers for relevant adaptation decisions. Maria Klara Wolters, Ravichander Vipperla, Steve Renals |
INTERSPEECH | 3 |
| 2009 | Speech Recognition Using Augmented Conditional Random FieldsabstractAcoustic modeling based on hidden Markov models (HMMs) is employed by state-of-the-art stochastic speech recognition systems. Although HMMs are a natural choice to warp the time axis and model the temporal phenomena in the speech signal, their conditional independence properties limit their ability to model spectral phenomena well. In this paper, a new acoustic modeling paradigm based on augmented conditional random fields (ACRFs) is investigated and developed. This paradigm addresses some limitations of HMMs while maintaining many of the aspects which have made them successful. In particular, the acoustic modeling problem is reformulated in a data driven, sparse, augmented space to increase discrimination. Acoustic context modeling is explicitly integrated to handle the sequential phenomena of the speech signal. We present an efficient framework for estimating these models that ensures scalability and generality. In the TIMIT phone recognition task, a phone error rate of 23.0% was recorded on the full test set, a significant improvement over comparable HMM-based systems. Yasser Hifny, Steve Renals |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Robust Speaker-Adaptive HMM-Based Text-to-Speech SynthesisabstractThis paper describes a speaker-adaptive HMM-based speech synthesis system. The new system, called ldquoHTS-2007,rdquo employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available. In addition, a comparison study with several speech synthesis techniques shows the new system is very robust: It is able to build voices from less-than-ideal speech data and synthesize good-quality speech even for out-of-domain sentences. Junichi Yamagishi, Takashi Nose, Heiga Zen, Zhen-Hua Ling, Tomoki Toda, Keiichi Tokuda, Simon King 0001, Steve Renals |
IEEE Trans. Speech Audio Process. | 8 |
| 2008 | Glottal spectral separation for parametric speech synthesisabstractThe great advantage of using a glottal source model in parametric speech synthesis is the degree of parametric flexibility it gives to transform and model aspects of voice quality and speaker identity. However, few studies have addressed how the glottal source affects the quality of synthetic speech. Here, we have developed the Glottal Spectral Separation (GSS) method which consists of separating the glottal source effects from the spectral envelope of the speech. It enables us to compare the LF-model with the simple impulse excitation, using the same spectral envelope to synthesize speech. The results of a perceptual evaluation showed that the LF-model clearly outperformed the impulse. The GSS method was also used to successfully transform a modal voice into a breathy or tense voice, by modifying the LF-parameters. The proposed technique could be used to improve the speech quality and source parametrization of HMM-based speech synthesizers, which use an impulse excitation. João P. Cabral, Steve Renals, Korin Richmond, Junichi Yamagishi |
INTERSPEECH | 2 |
| 2008 | Pitch adaptive features for LVCSRabstractWe have investigated the use of a pitch adaptive spectral representation on large vocabulary speech recognition, in conjunction with speaker normalisation techniques. We have compared the effect of a smoothed spectrogram to the pitch adaptive spectral analysis by decoupling these two components of STRAIGHT. Experiments performed on a large vocabulary meeting speech recognition task highlight the importance of combining a pitch adaptive spectral representation with a conventional fixed window spectral analysis. We found evidence that STRAIGHT pitch adaptive features are more speaker independent than conventional MFCCs without pitch adaptation, thus they also provide better performances when combined using feature combination techniques such as Heteroscedastic Linear Discriminant Analysis. Giulia Garau, Steve Renals |
INTERSPEECH | 2 |
| 2008 | Unsupervised language model adaptation based on topic and role information in multiparty meetingsabstractWe continue our previous work on the modeling of topic and role information from multiparty meetings using a hierarchical Dirichlet process (HDP), in the context of language model adaptation. In this paper we focus on three problems: 1) an empirical analysis of the HDP as a nonparametric topic model; 2) the mismatch problem of vocabularies of the baseline n-gram model and the HDP; and 3) an automatic speech recognition experiment to further verify the effectiveness of our adaptation framework. Experiments on a large meeting corpus of more than 70 hours speech data show consistent and significant improvements in terms of word error rate for language model adaptation based on the topic and role information. Index Terms: language model, adaptation, topic model, hierarchical Dirichlet process, participant role Songfang Huang, Steve Renals |
INTERSPEECH | 2 |
| 2008 | Predicting tongue shapes from a few landmark locationsabstractWe present a method for predicting the midsagittal tongue contour from the locations of a few landmarks (metal pellets) on the tongue surface, as used in articulatory databases such as MOCHA and the Wisconsin XRDB. Our method learns a mapping using ground-truth tongue contours derived from ultrasound data and drastically improves over spline interpolation. We also determine the optimal locations of the landmarks, and the number of landmarks required to achieve a desired prediction error: 3–4 landmarks are enough to achieve 0.3–0.2 mm error per point on the tongue. Index Terms: ultrasound, midsagittal tongue contour, tongue tracking, articulatory database Miguel Á. Carreira-Perpiñán, Korin Richmond, Alan Wrench, Steve Renals |
INTERSPEECH | 5 |
| 2008 | Longitudinal study of ASR performance on ageing voicesabstractThis paper presents the results of a longitudinal study of ASR performance on ageing voices. Experiments were conducted on the audio recordings of the proceedings of the Supreme Court Of The United States (SCOTUS). Results show that the Automatic Speech Recognition (ASR) Word Error Rates (WERs) for elderly voices are significantly higher than those of adult voices. The word error rate increases gradually as the age of the elderly speakers increase. Use of maximum likelihood linear regression (MLLR) based speaker adaptation on ageing voices improves the WER though the performance is still considerably lower compared to adult voices. Speaker adaptation however reduces the increase in WER with age during old age. Ravichander Vipperla, Steve Renals, Joe Frankel |
INTERSPEECH | 2 |
| 2008 | Acoustic-Articulatory Modeling With the Trajectory HMMabstractIn this letter, we introduce an hidden Markov model (HMM)-based inversion system to recovery articulatory movements from speech acoustics. Trajectory HMMs are used as generative models for modelling articulatory data. Experiments on the MOCHA-TIMIT corpus indicate that the jointly trained acoustic-articulatory models are more accurate (lower RMS error) than the separately trained ones, and that trajectory HMM training results in greater accuracy compared with conventional maximum likelihood HMM training. Moreover, the system has the ability to synthesize articulatory movements directly from a textual representation. Steve Renals |
IEEE Signal Process. Lett. | 2 |
| 2008 | A Cascaded Broadcast News HighlighterabstractThis paper presents a fully automatic news skimming system which takes a broadcast news audio stream and provides the user with the segmented, structured, and highlighted transcript. This constitutes a system with three different, cascading stages: converting the audio stream to text using an automatic speech recognizer, segmenting into utterances and stories, and finally determining which utterance should be highlighted using a saliency score. Each stage must operate on the erroneous output from the previous stage in the system, an effect which is naturally amplified as the data progresses through the processing stages. We present a large corpus of transcribed broadcast news data enabling us to investigate to which degree information worth highlighting survives this cascading of processes. Both extrinsic and intrinsic experimental results indicate that mistakes in the story boundary detection has a strong impact on the quality of highlights, whereas erroneous utterance boundaries cause only minor problems. Further, the difference in transcription quality does not affect the overall performance greatly. Heidi Christensen, Yoshihiko Gotoh, Steve Renals |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Recognition of Dialogue Acts in Multiparty Meetings Using a Switching DBNabstractThis paper is concerned with the automatic recognition of dialogue acts (DAs) in multiparty conversational speech. We present a joint generative model for DA recognition in which segmentation and classification of DAs are carried out in parallel. Our approach to DA recognition is based on a switching dynamic Bayesian network (DBN) architecture. This generative approach models a set of features, related to lexical content and prosody, and incorporates a weighted interpolated factored language model. The switching DBN coordinates the recognition process by integrating the component models. The factored language model, which is estimated from multiple conversational data corpora, is used in conjunction with additional task-specific language models. In conjunction with this joint generative model, we have also investigated the use of a discriminative approach, based on conditional random fields, to perform a reclassification of the segmented DAs. We have carried out experiments on the AMI corpus of multimodal meeting recordings, using both manually transcribed speech, and the output of an automatic speech recognizer, and using different configurations of the generative model. Our results indicate that the system performs well both on reference and fully automatic transcriptions. A further significant improvement in recognition accuracy is obtained by the application of the discriminative reranking approach based on conditional random fields. Alfred Dielmann, Steve Renals |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Combining Spectral Representations for Large-Vocabulary Continuous Speech RecognitionabstractIn this paper, we investigate the combination of complementary acoustic feature streams in large-vocabulary continuous speech recognition (LVCSR). We have explored the use of acoustic features obtained using a pitch-synchronous analysis, Straight, in combination with conventional features such as Mel frequency cepstral coefficients. Pitch-synchronous acoustic features are of particular interest when used with vocal tract length normalization (VTLN) which is known to be affected by the fundamental frequency. We have combined these spectral representations directly at the acoustic feature level using heteroscedastic linear discriminant analysis (HLDA) and at the system level using ROVER. We evaluated this approach on three LVCSR tasks: dictated newspaper text (WSJCAM0), conversational telephone speech (CTS), and multiparty meeting transcription. The CTS and meeting transcription experiments were both evaluated using standard NIST test sets and evaluation protocols. Our results indicate that combining conventional and pitch-synchronous acoustic feature sets using HLDA results in a consistent, significant decrease in word error rate across all three tasks. Combining at the system level using ROVER resulted in a further significant decrease in word error rate. Giulia Garau, Steve Renals |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Hierarchical Pitman-Yor language models for ASR in meetingsabstractIn this paper we investigate the application of a hierarchical Bayesian language model (LM) based on the Pitman-Yor process for automatic speech recognition (ASR) of multiparty meetings. The hierarchical Pitman-Yor language model (HPYLM) provides a Bayesian interpretation of LM smoothing. An approximation to the HPYLM recovers the exact formulation of the interpolated Kneser-Ney smoothing method in n-gram models. This paper focuses on the application and scalability of HPYLM on a practical large vocabulary ASR system. Experimental results on NIST RT06s evaluation meeting data verify that HPYLM is a competitive and promising language modeling technique, which consistently performs better than interpolated Kneser-Ney and modified Kneser-Ney n-gram LMs in terms of both perplexity and word error rate. Songfang Huang, Steve Renals |
ASRU | 2 |
| 2007 | Recognition and understanding of meetings the AMI and AMIDA projectsabstractThe AMI and AMIDA projects are concerned with the recognition and interpretation of multiparty meetings. Within these projects we have: developed an infrastructure for recording meetings using multiple microphones and cameras; released a 100 hour annotated corpus of meetings; developed techniques for the recognition and interpretation of meetings based primarily on speech recognition and computer vision; and developed an evaluation framework at both component and system levels. In this paper we present an overview of these projects, with an emphasis on speech recognition and content extraction. Steve Renals, Thomas Hain, Hervé Bourlard |
ASRU | 1 |
| 2007 | DBN Based Joint Dialogue Act Recognition of Multiparty MeetingsabstractJoint dialogue act segmentation and classification of the new AMI meeting corpus has been performed through an integrated framework based on a switching dynamic Bayesian network and a set of continuous features and language models. The recognition process is based on a dictionary of 15 DA classes tailored for group decision-making. Experimental results show that a novel interpolated factored language model results in a low error rate on the automatic segmentation task, and thus good recognition results can be achieved on AMI multiparty conversational speech. Alfred Dielmann, Steve Renals |
ICASSP (4) | 2 |
| 2007 | Hierarchical dialogue optimization using semi-Markov decision processesabstractThis paper addresses the problem of dialogue optimization on large search spaces. For such a purpose, in this paper we propose to learn dialogue strategies using multiple Semi-Markov Decision Processes and hierarchical reinforcement learning. This approach factorizes state variables and actions in order to learn a hierarchy of policies. Our experiments are based on a simulated flight booking dialogue system and compare flat versus hierarchical reinforcement learning. Experimental results show that the proposed approach produced a dramatic search space reduction (99.36 than flat reinforcement learning with a very small loss in optimality (on average 0.3 system turns). Results also report that the learnt policies outperformed a hand-crafted one under three different conditions of ASR confidence levels. This approach is appealing to dialogue optimization due to faster learning, reusable subsolutions, and scalability to larger problems. Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira |
INTERSPEECH | 2 |
| 2007 | Towards online speech summarizationabstractThe majority of speech summarization research has focused on extracting the most informative dialogue acts from recorded, archived data. However, a potential use case for speech summarization in the meetings domain is to facilitate a meeting in progress by providing the participants- whether they are attending in-person or remotely- with an indication of the most important parts of the discussion so far. This requires being able to determine whether a dialogue act is extract-worthy before the global meeting context is available. This paper introduces a novel method for weighting dialogue acts using only very limited local context, and shows that high summary precision is possible even when information about the meeting as a whole is lacking. A new evaluation framework consisting of weighted precision, recall and f-score is detailed, and the novel online summarization method is shown to significantly increase recall and f-score compared with a method using no contextual information. Index Terms: speech summarization, online summarization, multiparty dialogues, meeting assistant, remote monitoring Gabriel Murray, Steve Renals |
INTERSPEECH | 2 |
| 2007 | Automatic Meeting Segmentation Using Dynamic Bayesian NetworksabstractMultiparty meetings are a ubiquitous feature of organizations, and there are considerable economic benefits that would arise from their automatic analysis and structuring. In this paper, we are concerned with the segmentation and structuring of meetings (recorded using multiple cameras and microphones) into sequences of group meeting actions such as monologue, discussion and presentation. We outline four families of multimodal features based on speaker turns, lexical transcription, prosody, and visual motion that are extracted from the raw audio and video recordings. We relate these low-level features to more complex group behaviors using a multistream modelling framework based on multi-stream dynamic Bayesian networks (DBNs). This results in an effective approach to the segmentation problem, resulting in an action error rate of 12.2%, compared with 43% using an approach based on hidden Markov models. Moreover, the multistream DBN developed here leaves scope for many further improvements and extensions Alfred Dielmann, Steve Renals |
IEEE Trans. Multim. | 2 |
| 2006 | Automatic Segmentation of Multiparty Dialogue
Pei-Yun Hsueh, Johanna D. Moore, Steve Renals |
EACL | 3 |
| 2006 | Learning multi-goal dialogue strategies using reinforcement learning with reduced state-action spacesabstractLearning dialogue strategies using the reinforcement learning framework is problematic due to its expensive computational cost. In this paper we propose an algorithm that reduces a state-action space to one which includes only valid state-actions. We performed experiments on full and reduced spaces using three systems (with 5, 9 and 20 slots) in the travel domain using a simulated environment. The task was to learn multi-goal dialogue strategies optimizing single and multiple confirmations. Average results using strategies learnt on reduced spaces reveal the following benefits against full spaces: 1) less computer memory (94 % reduction), 2) faster learning (93 % faster convergence) and better performance (8.4 % less time steps and 7.7 % higher reward). Index Terms: reinforcement learning, spoken dialogue systems. 1. Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira |
INTERSPEECH | 2 |
| 2006 | Dialogue act compression via pitch contour preservationabstractThis paper explores the usefulness of prosody in automatically compressing dialogue acts from meeting speech. Specifically, this work attempts to compress utterances by preserving the pitch contour of the original whole utterance. Two methods of doing this are described in detail and are evaluated subjectively using human annotators and objectively using edit distance with a human-authored gold-standard. Both metrics show that such a prosodic approach is much better than the random baseline approach and significantly better than a simple text compression method. Gabriel Murray, Steve Renals |
INTERSPEECH | 2 |
| 2006 | Phone recognition analysis for trajectory HMMabstractThe trajectory HMM has been shown to be useful for model-based speech synthesis where a smoothed trajectory is generated using temporal constraints imposed by dynamic features. To evaluate the performance of such model on an ASR task, we present a trajectory decoder based on tree search with delayed path merging. Experiment on a speaker-dependent phone recognition task using the MOCHA-TIMIT database shows that the MLE-trained trajectory model, while retaining attractive properties of being a proper generative model, tends to favour over-smoothed trajectory among competing hypothesises, and does not perform better than a conventional HMM. We use this to build an argument that models giving better fit on training data may suffer a reduction of discrimination by being too faithful to training data. This partially explains why alternative acoustic models that try to explicitly model temporal constraints do not achieve significant improvements in ASR. Steve Renals |
INTERSPEECH | 2 |
| 2006 | Incorporating Speaker and Discourse Features into Speech Summarization
Gabriel Murray, Steve Renals, Jean Carletta, Johanna D. Moore |
HLT-NAACL | 2 |
| 2006 | Reinforcement Learning of Dialogue Strategies with Hierarchical Abstract MachinesabstractIn this paper we propose partially specified dialogue strategies for dialogue strategy optimization, where part of the strategy is specified deterministically and the rest optimized with reinforcement learning (RL). To do this we apply RL with hierarchical abstract machines (HAMs). We also propose to build simulated users using HAMs, incorporating a combination of hierarchical deterministic and probabilistic behaviour. We performed experiments using a single-goal flight booking dialogue system, and compare two dialogue strategies (deterministic and optimized) using three types of simulated user (novice, experienced and expert). Our results show that HAMs are promising for both dialogue optimization and simulation, and provide evidence that indeed partially specified dialogue strategies can outperform deterministic ones (on average 4.7 fewer system turns) with faster learning than the traditional RL framework. Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira |
SLT | 2 |
| 2005 | Maximum entropy segmentation of broadcast newsabstractThe paper presents an automatic system for structuring and preparing a news broadcast for applications such as speech summarization, browsing, archiving and information retrieval. This process comprises transcribing the audio using an automatic speech recognizer and subsequently segmenting the text into utterances and topics. A maximum entropy approach is used to build statistical models for both utterance and topic segmentation. The experimental work addresses the effect on performance of the topic boundary detector of three factors - the types of feature used, the quality of the ASR transcripts, and the quality of the utterance boundary detector. The results show that the topic segmentation is not affected severely by transcript errors, whereas errors in utterance segmentation are more devastating. Heidi Christensen, BalaKrishna Kolluru, Yoshihiko Gotoh, Steve Renals |
ICASSP (1) | 4 |
| 2005 | Applying vocal tract length normalization to meeting recordingsabstractVocal Tract Length Normalisation (VTLN) is a commonly used technique to normalise for inter-speaker variability. It is based on the speaker-specific warping of the frequency axis, parameterised by a scalar warp factor. This factor is typically estimated using maximum likelihood. We discuss how VTLN may be applied to multiparty conversations, reporting a substantial decrease in word error rate in experiments using the ICSI meetings corpus. We investigate the behaviour of the VTLN warping factor and show that a stable estimate is not obtained. Instead it appears to be influenced by the context of the meeting, in particular the current conversational partner. These results are consistent with predictions made by the psycholinguistic interactive alignment account of dialogue, when applied at the acoustic and phonological levels. Giulia Garau, Steve Renals, Thomas Hain |
INTERSPEECH | 2 |
| 2005 | Transcription of conference room meetings: an investigationabstractThe automatic processing of speech collected in conference style meetings has attracted considerable interest with several large scale projects devoted to this area. In this paper we explore the use of various meeting corpora for the purpose of automatic speech recognition. In particular we investigate the similarity of these resources and how to efficiently use them in the construction of a meeting transcription system. The analysis shows distinctive features for each resource. However the benefit in pooling data and hence the similarity seems sufficient to speak of a generic conference meeting domain . In this context this paper also presents work on development for the AMI meeting transcription system, a joint effort by seven sites working on the AMI (augmented multi-party interaction) project. Thomas Hain, John Dines, Giulia Garau, Martin Karafiát, Darren Moore, Vincent Wan, Roeland Ordelman, Steve Renals |
INTERSPEECH | 8 |
| 2005 | A hybrid Maxent/HMM based ASR systemabstractThe aim of this work is to develop a practical framework, which extends the classical Hidden Markov Models (HMM) for continuous speech recognition based on the Maximum Entropy (MaxEnt) principle. The MaxEnt models can estimate the posterior probabilities directly as with Hybrid NN/HMM connectionist speech recognition systems. In particular, a new acoustic modelling based on discriminative MaxEnt models is formulated and is being developed to replace the generative Gaussian Mixture Models (GMM) commonly used to model acoustic variability. Initial experimental results using the TIMIT phone task are reported. Yasser Hifny, Steve Renals, Neil D. Lawrence |
INTERSPEECH | 2 |
| 2005 | Extractive summarization of meeting recordingsabstractSeveral approaches to automatic speech summarization are discussed below, using the ICSI Meetings corpus. We contrast feature-based approaches using prosodic and lexical features with maximal marginal relevance and latent semantic analysis approaches to summarization. While the latter two techniques are borrowed directly from the field of text summarization, feature-based approaches using prosodic information are able to utilize characteristics unique to speech data. We also investigate how the summarization results might deteriorate when carried out on ASR output as opposed to manual transcripts. All of the summaries are of an extractive variety, and are compared using the software ROUGE. Gabriel Murray, Steve Renals, Jean Carletta |
INTERSPEECH | 2 |
| 2005 | Speaker verification using sequence discriminant support vector machinesabstractThis paper presents a text-independent speaker verification system using support vector machines (SVMs) with score-space kernels. Score-space kernels generalize Fisher kernels and are based on underlying generative models such as Gaussian mixture models (GMMs). This approach provides direct discrimination between whole sequences, in contrast with the frame-level approaches at the heart of most current systems. The resultant SVMs have a very high dimensionality since it is related to the number of parameters in the underlying generative model. To address problems that arise in the resultant optimization we introduce a technique called spherical normalization that preconditions the Hessian matrix. We have performed speaker verification experiments using the PolyVar database. The SVM system presented here reduces the relative error rates by 34% compared to a GMM likelihood ratio system. Vincent Wan, Steve Renals |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Speech and crosstalk detection in multichannel audioabstractThe analysis of scenarios in which a number of microphones record the activity of speakers, such as in a round-table meeting, presents a number of computational challenges. For example, if each participant wears a microphone, speech from both the microphone's wearer (local speech) and from other participants (crosstalk) is received. The recorded audio can be broadly classified in four ways: local speech, crosstalk plus local speech, crosstalk alone and silence. We describe two experiments related to the automatic classification of audio into these four classes. The first experiment attempted to optimize a set of acoustic features for use with a Gaussian mixture model (GMM) classifier. A large set of potential acoustic features were considered, some of which have been employed in previous studies. The best-performing features were found to be kurtosis, "fundamentalness," and cross-correlation metrics. The second experiment used these features to train an ergodic hidden Markov model classifier. Tests performed on a large corpus of recorded meetings show classification accuracies of up to 96%, and automatic speech recognition performance close to that obtained using ground truth segmentation. Stuart N. Wrigley, Guy J. Brown, Vincent Wan, Steve Renals |
IEEE Trans. Speech Audio Process. | 4 |
| 2004 | From Text Summarisation to Style-Specific Summarisation for Broadcast News
Heidi Christensen, BalaKrishna Kolluru, Yoshihiko Gotoh, Steve Renals |
ECIR | 4 |
| 2004 | Acoustic space dimensionality selection and combination using the maximum entropy principleabstractWe propose a discriminative approach to acoustic space dimensionality selection based on maximum entropy modelling. We form a set of constraints by composing the acoustic space with the space of phone classes, and use a continuous feature formulation of maximum entropy modelling to select an optimal feature set. The suggested approach has two steps: (1) the selection of the best acoustic space that efficiently and economically represents the acoustic data and its variability; (2) the combination of selected acoustic features in the maximum entropy framework to estimate the posterior probabilities over the phonetic labels given the acoustic input. Specific contributions of the paper include a parameter estimation algorithm (generalized improved iterative scaling) that enables the use of negative features, the parameterization of constraint functions using Gaussian mixture models, and experimental results using the TIMIT database. Yasser H. Abdel-Haleem, Steve Renals, Neil D. Lawrence |
ICASSP (5) | 2 |
| 2004 | Dynamic Bayesian networks for meeting structuringabstractThe paper is about the automatic structuring of multiparty meetings using audio information. We have used a corpus of 53 meetings, recorded using a microphone array and lapel microphones for each participant. The task was to segment meetings into a sequence of meeting actions, or phases. We have adopted a statistical approach using dynamic Bayesian networks (DBNs). Two DBN architectures were investigated: a two-level hidden Markov model (HMM) in which the acoustic observations were concatenated; and a multistream DBN in which two separate observation sequences were modelled. We have also explored the use of counter variables to constrain the number of action transitions. Experimental results indicate that the DBN architectures are an improvement over a simple baseline HMM, with the multistream DBN with counter constraints producing an action error rate of 6%. Alfred Dielmann, Steve Renals |
ICASSP (5) | 2 |
| 2004 | Multi-stream segmentation of meetingsabstractThis paper investigates the automatic segmentation of meetings into a sequence of group actions or phases. Our work is based on a corpus of multiparty meetings collected in a meeting room instrumented with video cameras, lapel microphones and a microphone array. We have extracted a set of feature streams, in this case extracted from the audio data, based on speaker turns, prosody and a transcript of what was spoken. We have related these signals to the higher level semantic categories via a multistream statistical model based on dynamic Bayesian networks (DBNs). We report on a set of experiments in which different DBN architectures are compared, together with the different feature streams. The resultant system has an action error rate of 9%. Alfred Dielmann, Steve Renals |
MMSP | 2 |
| 2003 | Audio information access from meeting roomsabstractWe investigate approaches to accessing information from the streams of audio data that result from multi-channel recordings of meetings. The methods investigated use word-level transcriptions, and information derived from models of speaker activity and speaker turn patterns. Our experiments include spoken document retrieval for meetings, automatic structuring of meetings based on self-similarity matrices of speaker turn patterns and a simple model of speaker activity. Meeting recordings are rich in both lexical and non-lexical information; our results illustrate some novel kinds of analysis made possible by a transcribed corpus of natural meetings. Steve Renals, Daniel P. W. Ellis |
ICASSP (4) | 1 |
| 2003 | SVMSVM: support vector machine speaker verification methodologyabstractSupport vector machines with the Fisher and score-space kernels are used for text independent speaker verification to provide direct discrimination between complete utterances. This is unlike approaches such as discriminatively trained Gaussian mixture models or other discriminative classifiers that discriminate at the frame-level only. Using the sequence-level discrimination approach we are able to achieve error-rates that are significantly better than the current state-of-the-art on the PolyVar database. Vincent Wan, Steve Renals |
ICASSP (2) | 2 |
| 2003 | Multi-class extractive voicemail summarizationabstractThis paper is about a system that extracts principal content words from speech-recognized transcripts of voicemail messages and classifies them into proper names, telephone numbers, dates/times and `other'. The short text summaries generated are suitable for mobile messaging applications. The system uses a set of classifiers to identify the summary words, with each word being identified by a vector of lexical and prosodic features. The features are selected using Parcel, an ROC-based algorithm. We visually compare the role of a large number of individual features and discuss effective ways to combine them. We finally evaluate their performance on manual and automatic transcriptions derived from two different speech recognition systems. Konstantinos Koumpis, Steve Renals |
INTERSPEECH | 2 |
| 2003 | Feature selection for the classification of crosstalk in multi-channel audioabstractAn extension to the conventional speech / nonspeech classification framework is presented for a scenario in which a number of microphones record the activity of speakers present at a meeting (one microphone per speaker). Since each microphone can receive speech from both the participant wearing the microphone (local speech) and other participants (crosstalk), the recorded audio can be broadly classified in four ways: local speech, crosstalk plus local speech, crosstalk alone and silence. We describe a classifier in which a Gaussian mixture model (GMM) is used to model each class. A large set of potential acoustic features are considered, some of which have been employed in previous speech / nonspeech classifiers. A combination of two feature selection algorithms is used to identify the optimal feature set for each class. Results from the GMM classifier using the selected features are superior to those of a previously published approach. 1. Stuart N. Wrigley, Guy J. Brown, Vincent Wan, Steve Renals |
INTERSPEECH | 4 |
| 2002 | ASR system modeling for automatic evaluation and optimization of dialogue systemsabstractThough the field of spoken dialogue systems has developed quickly in the last decade, rapid design of dialogue strategies remains uneasy. Several approaches to the problem of automatic strategy learning have been proposed and aie use of Reinforcement Learning introduced by Levin and Pieraccini is becoming part of the state of the art in this area. However, the quality of the strategy learned by the system depends on the definition of the optimization criterion and on the accuracy of aie environment model. In this paper, we propose to bring a model of an ASR system in the simulated environment in order to enhance the learned strategy. To do so, we introduced recognition error rates and confidence levels produced by ASR systems in the optimization criterion. Olivier Pietquin, Steve Renals |
ICASSP | 2 |
| 2002 | Evaluation of kernel methods for speaker verification and identificationabstractSupport vector machines are evaluated on speaker verification and speaker identification tasks. We compare the polynomial kernel, the Fisher kernel, a likelihood ratio kernel and the pair hidden Markov model kernel with baseline systems based on a discriminative polynomial classifier and generative Gaussian mixture model classifiers. Simulations were carried out on the YOHO database and some promising results were obtained. Vincent Wan, Steve Renals |
ICASSP | 2 |
| 2002 | Connectionist speech recognition of Broadcast News
Anthony J. Robinson, Gary D. Cook, Daniel P. W. Ellis, Eric Fosler-Lussier, Steve Renals, D. A. G. Williams |
Speech Commun. | 5 |
| 2001 | Extractive summarization of voicemail using lexical and prosodic feature subset selectionabstractThis paper presents a novel data-driven approach to summarizing spoken audio transcripts utilizing lexical and prosodic features. The former are obtained from a speech recognizer and the latter are extracted automatically from speech waveforms. We employ a feature subset selection algorithm, based on ROC curves, which examines different combinations of features at different target operating conditions. The approach is evaluated on the IBM Voicemail corpus, demonstrating that it is possible and desirable to avoid complete commitment to a single best classifier or feature set. 1. Konstantinos Koumpis, Steve Renals, Mahesan Niranjan |
INTERSPEECH | 2 |
| 2000 | Variable word rate N-gramsabstractThe rate of occurrence of words is not uniform but varies from document to document. Despite this observation, parameters for conventional N-gram language models are usually derived using the assumption of a constant word rate. In this paper we investigate the use of variable word rate assumption, modelled by a Poisson distribution or a continuous mixture of Poissons. We present an approach to estimating the relative frequencies of words or N-grams taking prior information of their occurrences into account. Discounting and smoothing schemes are also considered. Using the Broadcast News task, the approach demonstrates a reduction of perplexity up to 10%. Yoshihiko Gotoh, Steve Renals |
ICASSP | 2 |
| 2000 | Transcription and summarization of voicemail speechabstractThis paper describes the development of a system to transcribe and summarize voicemail messages. The results of the research presented in this paper are two-fold. First, a hybrid connectionist approach to the Voicemail transcription task shows that competitive performance can be achieved using a context-independent system with fewer parameters than those based on mixtures of Gaussian likelihoods. Second, an effective and robust combination of statistical with prior knowledge sources for term weighting is used to extract information from the decoder's output in order to deliver summaries to the message recipients via a GSM Short Message Service (SMS) gateway. 1. INTRODUCTION As the emphasis in cellular networks changes from voice-only communication to a rich combination of content based applications and services, speech recognition can provide access to several types of information through a number of portable solutions, including mobile phones and personal digital assistants. This pa... Konstantinos Koumpis, Steve Renals |
INTERSPEECH | 2 |
| 2000 | Practical Identifiability of Finite Mixtures of Multivariate Bernoulli DistributionsabstractThe class of finite mixtures of multivariate Bernoulli distributions is known to be nonidentifiable; that is, different values of the mixture parameters can correspond to exactly the same probability distribution. In principle, this would mean that sample estimates using this model would give rise to different interpretations. We give empirical support to the fact that estimation of this class of mixtures can still produce meaningful results in practice, thus lessening the importance of the identifiability problem. We also show that the expectation-maximization algorithm is guaranteed to converge to a proper maximum likelihood estimate, owing to a property of the log-likelihood surface. Experiments with synthetic data sets show that an original generating distribution can be estimated from a sample. Experiments with an electropalatography data set show important structure in the data. Miguel Á. Carreira-Perpiñán, Steve Renals |
Neural Comput. | 2 |
| 2000 | Indexing and retrieval of broadcast news
Steve Renals, David C. Abberley, David Kirby, Tony Robinson |
Speech Commun. | 1 |
| 2000 | Accessing information in spoken audio
Steve Renals, Tony Robinson |
Speech Commun. | 1 |
| 1999 | Named entity tagged language modelsabstractWe introduce named entity (NE) language modelling, a stochastic finite state machine approach to identifying both words and NE categories from a stream of spoken data. We provide an overview of our approach to NE tagged language model (LM) generation together with results of the application of such a LM to the task of out-of-vocabulary (OOV) word reduction in large vocabulary speech recognition. Using the Wall Street Journal and Broadcast News corpora, it is shown that the tagged LM was able to reduce the overall word error rate by 14%, detecting up to 70% of previously OOV words. We also describe an example of the direct tagging of spoken data with NE categories. Yoshihiko Gotoh, Steve Renals, Gethin Williams |
ICASSP | 2 |
| 1999 | Integrated transcription and identification of named entities in broadcast speechabstract17 63($.(5 9(5,),&$7,21 21 /$%25$725< $1' ),(/' 7(67 '$7$%$6(6 ,1 7+( 0976 352-(&7(1) IMT, Neuchâtel (CH) [email protected](2) IDIAP, Martigny (CH) [email protected](3) now at EIV, Sion (CH) -Gilbert. Steve Renals, Yoshihiko Gotoh |
EUROSPEECH | 1 |
| 1999 | Recognition, indexing and retrieval of british broadcast news with the THISL systemabstractEfficient reduction of storage and complexity demands in VQ and MQ systems is a key issue when developing new, more powerful compression algorithms. Secondary Storage Quantisation (SSQ) is capable of drastically reducing VQ storage through an efficient representation of codebook elements. Rather than the conventional fixed or floating point representation, codebook elements are quantised using a set of “secondary codebooks” and represented as a set of quantisation indices. The number of bits required for these indices is relatively small and hence the amount of storage required for codebook representation is reduced. The potential of SSQ for codebook compression is demonstrated in a Split Matrix Quantisation (SMQ) application. A reduction of 65 75 % in the amount of memory required for SMQ codebooks is achieved. Tony Robinson, David C. Abberley, David Kirby, Steve Renals |
EUROSPEECH | 4 |
| 1999 | The THISL system for indexing and retrieval of broadcast newsabstractThis paper describes the THISL news retrieval system which maintains an archive of BBC radio and television news recordings. The system uses the ABBOT large vocabulary continuous speech recognition system to transcribe news broadcasts, and the thisIIR text retrieval system to index and access the transcripts. Decoding and indexing is performed automatically, and the archive is updated with three hours of new material every day. A Web-based interface to the retrieval system has been devised to facilitate access to the archive. Steve Renals, David C. Abberley, David Kirby, Tony Robinson |
MMSP | 1 |
| 1999 | Confidence measures from local posterior probability estimates
Gethin Williams, Steve Renals |
Comput. Speech Lang. | 2 |
| 1999 | Topic-based mixture language modellingabstractThis paper describes an approach for constructing a mixture of language models based on simple statistical notions of semantics using probabilistic models developed for information retrieval. The approach encapsulates corpus-derived semantic information and is able to model varying styles of text. Using such information, the corpus texts are clustered in an unsupervised manner and a mixture of topic-specific language models is automatically created. The principal contribution of this work is to characterise the document space resulting from information retrieval techniques and to demonstrate the approach for mixture language modelling. A comparison is made between manual and automatic clustering in order to elucidate how the global content information is expressed in the space. We also compare (in terms of association with manual clustering and language modelling accuracy) alternative term-weighting schemes and the effect of singular value decomposition dimension reduction (latent semantic analysis). Test set perplexity results using the British National Corpus indicate that the approach can improve the potential of statistical language modelling. Using an adaptive procedure, the conventional model may be tuned to track text data with a slight increase in computational cost. Yoshihiko Gotoh, Steve Renals |
Nat. Lang. Eng. | 2 |
| 1999 | Start-synchronous search for large vocabulary continuous speech recognitionabstractIn this paper, we present a novel, efficient search strategy for large vocabulary continuous speech recognition. The search algorithm, based on a stack decoder framework, utilizes phone-level posterior probability estimates (produced by a connectionist/hidden Markov model acoustic model) as a basis for phone deactivation pruning-a highly efficient method of reducing the required computation. The single-pass algorithm is naturally factored into the time-asynchronous processing of the word sequence and the time-synchronous processing of the hidden Markov model state sequence. This enables the search to be decoupled from the language model while still maintaining the computational benefits of time-synchronous processing. The incorporation of the language model in the search is discussed and computationally cheap approximations to the full language model are introduced. Experiments were performed on the North American Business News task using a 60000 word vocabulary and a trigram language model. Results indicate that the computational cost of the search may be reduced by more than a factor of 40 with a relative search error of less than 2% using the techniques discussed in the paper. Steve Renals, Mike Hochberg |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | Retrieval of broadcast news documents with the THISL systemabstractThis paper describes a spoken document retrieval system, combining the ABBOT large vocabulary continuous speech recognition (LVCSR) system developed by Cambridge University, Sheffield University and SoftSound, and the PRISE information retrieval engine developed by NIST. The system was constructed to enable us to participate in the TREC 6 Spoken Document Retrieval experimental evaluation. Our key aims in this work were to produce a complete system for the SDR task, to investigate the effect of a word error rate of 30-50% on retrieval performance and to investigate the integration of LVCSR and word spotting in a retrieval task. David C. Abberley, Steve Renals, Gary D. Cook |
ICASSP | 2 |
| 1998 | Acoustic confidence measures for segmenting broadcast newsabstractIn this paper we define an acoustic confidence measure based on the estimates of local posterior probabilities produced by a HMM/ANN large vocabulary continuous speech recognition system. We use this measure to segment continuous audio into regions where it is and is not appropriate to expend recognition effort. The segmentation is computationally inexpensive and provides reductions in both overall word error rate and decoding time. The technique is evaluated using material from the Broadcast News corpus. 1. Jon Barker, Gethin Williams, Steve Renals |
ICSLP | 3 |
| 1998 | Confidence measures derived from an acceptor HMMabstractIn this paper we define a number of confidence measures derived from an acceptor HMM and evaluate their performance for the task of utterance verification using the North American Business News (NAB) and Broadcast News (BN) corpora. Results are presented for decodings made at both the word and phone level which show the relative profitability of rejection provided by the diverse set of confidence measures. The results indicate that language model dependent confidence measures have reduced performance on BN data relative to that for the more grammatically constrained NAB data. An explanation linking the observations that rejection is more profitable for noisy acoustics, for a reduced vocabulary and at the phone level is also given. 1. INTRODUCTION We define a confidence measure as a function which quantifies how well a model matches some spoken utterance, where the values of the function must be comparable across utterances. More specifically, an acoustic confidence measure is one whic... Gethin Williams, Steve Renals |
ICSLP | 2 |
| 1998 | Dimensionality reduction of electropalatographic data using latent variable models
Miguel Á. Carreira-Perpiñán, Steve Renals |
Speech Commun. | 2 |
| 1997 | Document space models using latent semantic analysisabstractIn this paper, an approach for constructing mixture language models (LMs) based on some notion of semantics is discussed. To this end, a technique known as latent semantic analysis (LSA) is used. The approach encapsulates corpus-derived semantic information and is able to model the varying style of the text. Using such information, the corpus texts are clustered in an unsupervised manner and mixture LMs are automatically created. This work builds on previous work in the field of information retrieval which was recently applied by Bellegarda et. al. to the problem of clustering words by semantic categories. The principal contribution of this work is to characterize the document space resulting from the LSA modeling and to demonstrate the approach for mixture LM application. Comparison is made between manual and automatic clustering in order to elucidate how the semantic information is expressed in the space. It is shown that, using semantic information, mixture LMs performs better than a conventional single LM with slight increase of computational cost. Yoshihiko Gotoh, Steve Renals |
EUROSPEECH | 2 |
| 1997 | Estimation of global posteriors and forward-backward training of hybrid HMM/ANN systemsabstractThe results of our research presented in this paper is two-fold.First, an estimation of global posteriors is formalized in the framework of hybrid HMM/ANN systems.It is shown that hybrid HMM/ANN systems, in which the ANN part estimates local posteriors, can be used to modelize global model posteriors.This formalization provides us with a clear theory in which both REMAP and \classical" Viterbi trained hybrid systems are unied.Second, a new forward-backward training of hybrid HMM/ANN systems is derived from the previous formulation.Comparisons of performance between Viterbi and forward-backward hybrid systems are presented and discussed. Jean Hennebert, Christophe Ris, Hervé Bourlard, Steve Renals, Nelson Morgan |
EUROSPEECH | 4 |
| 1997 | Confidence measures for hybrid HMM/ANN speech recognitionabstractIn this paper we introduce four acoustic confidence measures which are derived from the output of a hybrid HMM/ANN large vocabulary continuous speech recognition system. These confidence measures, based on local posterior probability estimates computed by an ANN, are evaluated at both phone and word levels, using the North American Business News corpus. 1. INTRODUCTION A reliable measure of the confidence of a speech recogniser's output is useful in many circumstances. A word may be hypothesised with low confidence when an out-of-vocabulary (OOV) word is encountered or when the word model is matched against unclear acoustics caused by disfluencies or noise. Both OOV words and unclear acoustics are a major source of recogniser error. A confidence measure based on can be used to reject those hypotheses which are likely to be erroneous (i.e., have a low confidence) in a hypothesis test. Additionally, a reliable confidence measure may be of practical use in recognition search (confidence ... Gethin Williams, Steve Renals |
EUROSPEECH | 2 |
| 1996 | Efficient evaluation of the LVCSR search space using the NOWAY decoderabstractThis article further develops and analyses the large vocabulary continuous speech recognition (LVCSR) search strategy reported by Renals and Hochberg (see Proc. ICASSP '95, p.596-9, 1995). In particular, the posterior-based phone deactivation pruning approach has been extended to include phone-dependent thresholds and an improved estimate of the least upper bound on the utterance log-probability has been developed. Analysis of the pruning procedures and of the search's interaction with the language model has also been performed. Experiments were carried out using the ARPA North American Business News task with a 20,000 word vocabulary and a trigram language model. As a result of these improvements and analyses, the computational cost of the recognition process performed by the NOWAY decoder has been substantially reduced. Steve Renals, Mike Hochberg |
ICASSP | 1 |
| 1996 | The 1995 abbot LVCSR system for multiple unknown microphonesabstractABBOT is the hybrid connectionist-hidden Markov model largevocabulary speech recognition system developed at Cambridge University.In this system, a recurrent network maps each acoustic vector io an estimate of the posterior probabilities of the phone classes, which are used as observation probabilities within an HMM.This paper describes the system which participated in the November 1995 ARPA Hub-3 Multiple Unknown Microphones (MUM) evaluation of continuous speech recognition systems, under the guise of the CU-CON system.The emphasis of the paper is on the changes made to the 1994 ABBOT system, specifically to accomodate the H3 task.This includes improved acoustic modelling using limited word-intemal context-dependentmodels, training on the W a l l Street Joumal Secondary channel database, and using the linear input network for speaker and environmental adaptation.Experimental results are reported lor various test and development sets from the November 1994 and 1995 ARPA benchmark tests.Recent improvements to the ABBOT system include training of the recumnt networks for effective use of the SI284 training corpus [21, and local speaker-adaptation approaches 1121, while application of state-based contextdependent phone modelling is planned for the nwfunm.6. Dan J. Kershaw, Tony Robinson, Steve Renals |
ICSLP | 3 |
| 1996 | Phone deactivation pruning in large vocabulary continuous speech recognitionabstractIntroduces a new pruning strategy for large vocabulary continuous speech recognition based on direct estimates of local posterior phone probabilities. This approach is well suited to hybrid connectionist/hidden Markov model systems. Experiments on the Wall Street Journal task using a 20000 word vocabulary and a trigram language model have demonstrated that phone deactivation pruning can increase the speed of recognition-time search by up to a factor of 10, with a relative increase in error rate of less than 2%. Steve Renals |
IEEE Signal Process. Lett. | 1 |
| 1995 | Recent improvements to the ABBOT large vocabulary CSR systemabstractABBOT is the hybrid connectionist-hidden Markov model (HMM) large-vocabulary continuous speech recognition (CSR) system developed at Cambridge University. This system uses a recurrent network to estimate the acoustic observation probabilities within an HMM framework. A major advantage of this approach is that good performance is achieved using context-independent acoustic models and requiring many fewer parameters than comparable HMM systems. This paper presents substantial performance improvements gained from new approaches to connectionist model combination and phone-duration modeling. Additional capability has also been achieved by extending the decoder to handle larger vocabulary tasks (20000 words and greater) with a trigram language model. This paper describes the modifications to the system and experimental results are reported for various test and development sets from the November 1992, 1993, and 1994 ARPA evaluations of spoken language systems. Mike Hochberg, Steve Renals, Anthony J. Robinson, Gary D. Cook |
ICASSP | 2 |
| 1995 | Efficient search using posterior phone probability estimatesabstractWe present a novel, efficient search strategy for large vocabulary continuous speech recognition (LVCSR). The search algorithm, based on stack decoding, uses posterior phone probability estimates to substantially increase its efficiency with minimal effect on accuracy. In particular, the search space is dramatically reduced by phone deactivation pruning where phones with a small local posterior probability are deactivated. This approach is particularly well-suited to hybrid connectionist/hidden Markov model systems because posterior phone probabilities are directly computed by the acoustic model. On large vocabulary tasks, using a trigram language model, this increased the search speed by an order of magnitude, with 2% or less relative search error. Results from a hybrid system are presented using the Wall Street Journal LVCSR database for a 20,000 word task using a backed-off trigram language model. For this task, our single-pass decoder took around 15x realtime on an HP735 workstation. At a cost of 7% relative search error, the decoding time can be speeded up to approximately realtime. Steve Renals, Mike Hochberg |
ICASSP | 1 |
| 1995 | WSJCAMO: a British English speech corpus for large vocabulary continuous speech recognitionabstractA significant new speech corpus of British English has been recorded at Cambridge University. Derived from the Wall Street Journal text corpus, WSJCAMO constitutes one of the largest corpora of spoken British English currently in existence. It has been specifically designed for the construction and evaluation of speaker-independent speech recognition systems. The database consists of 140 speakers each speaking about 110 utterances. This paper describes the motivation for the corpus, the processes undertaken in its construction and the utilities needed as support tools. All utterance transcriptions have been verified and a phonetic dictionary has been developed to cover the training data and evaluation tasks. Two evaluation tasks have been defined using standard 5000 word bigram and 20000 word trigram language models. The paper concludes with comparative results on these tasks for British and American English. Tony Robinson, Jeroen Fransen, David Pye, Jonathan Foote, Steve Renals |
ICASSP | 5 |
| 1995 | Speaker-adaptation for hybrid HMM-ANN continuous speech recognition systemabstractIt is well known that recognition performance degrades significantly when moving from a speakerdependent to a speaker-independent system. Traditional hidden Markov model (HMM) systems have successfully applied speaker-adaptation approaches to reduce this degradation. In this paper we present and evaluate some techniques for speaker-adaptation of a hybrid HMM-artificial neural network (ANN) continuous speech recognition system. These techniques are applied to a well trained, speaker-independent, hybrid HMM-ANN system and the recognizer parameters are adapted to a new speaker through off-line procedures. The techniques are evaluated on the DARPA RM corpus using varying amounts of adaptation material and different ANN architectures. The results show that speaker-adaptation within the hybrid framework can substantially improve system performance. 1. INTRODUCTION Automatic speech recognition has been a major goal for a large research community in the last few years. The predominant approa... João Paulo da Silva Neto, Luís B. Almeida, Mike Hochberg, Ciro Martins, Luís Nunes, Steve Renals, Tony Robinson |
EUROSPEECH | 6 |
| 1994 | IPA: improved phone modelling with recurrent neural networksabstractThis paper describes phone modelling improvements to the hybrid connectionist-hidden Markov model speech recognition system developed at Cambridge University. These improvements are applied to phone recognition from the TIMIT task and word recognition from the Wall Street Journal (WSJ) task. A recurrent net is used to map acoustic vectors to posterior probabilities of phone classes. The maximum likelihood phone or word string is then extracted using Markov models. The paper describes three improvements: connectionist model merging; explicit presentation of acoustic context; and improved duration modelling. The first is shown to provide a significant improvement in the TIMIT phone recognition rate and all three provide an improvement in the WSJ word recognition rate.> Tony Robinson, Mike Hochberg, Steve Renals |
ICASSP (1) | 3 |
| 1994 | Large vocabulary continuous speech recognition using a hybrid connectionist-HMM system
Mike Hochberg, Steve Renals, Anthony J. Robinson, Dan J. Kershaw |
ICSLP | 2 |
| 1994 | Using gamma filters to model temporal dependencies in speechabstractHybrid systems which use connectionist networks to estimate the output probabilities of a hidden Markov model represent time both at the network level and the Markov chain level. In this paper we discuss modelling time in connectionist networks, and introduce local recurrences in a feed-forward network in the form of an adaptive gamma filter. Using the Resource Management (RM) database, we have performed continuous speech recognition experiments comparing a gamma filtered input representation to a delay line. We have also performed speaker adaptation experiments using the speaker-dependent RM database. Our results have not indicated that gamma filters offer an appreciable modelling advantage on this task. However, the baseline speaker adaptation experiments have indicated that supervised adaptation over 100 sentences reduced the word error by an average of 40%. 1. INTRODUCTION Hybrid connectionist/hidden Markov model (HMM) systems model time at two levels, although these levels are n... Steve Renals, Mike Hochberg |
ICSLP | 1 |
| 1994 | Connectionist probability estimators in HMM speech recognitionabstractThe authors are concerned with integrating connectionist networks into a hidden Markov model (HMM) speech recognition system. This is achieved through a statistical interpretation of connectionist networks as probability estimators. They review the basis of HMM speech recognition and point out the possible benefits of incorporating connectionist networks. Issues necessary to the construction of a connectionist HMM recognition system are discussed, including choice of connectionist probability estimator. They describe the performance of such a system using a multilayer perceptron probability estimator evaluated on the speaker-independent DARPA Resource Management database. In conclusion, they show that a connectionist component improves a state-of-the-art HMM system. Steve Renals, Nelson Morgan, Hervé Bourlard, Horacio Franco |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | Bayesian regularisation methods in a hybrid MLP-HMM systemabstractWe have applied Bayesian regularisation methods to multi-layer percepuon (MLP) training in the context of a hybrid MLP-HMM (hidden Markov model) continuous speech recognition system. The Bayesian framework adopted here allows an objective setting of the regularisation parameters, according to the training data. Experiments have been carried out on the ARPA Resource Management database. Steve Renals, David J. C. MacKay |
EUROSPEECH | 1 |
| 1993 | A neural network based, speaker independent, large vocabulary, continuous speech recognition system: the WERNICKE projectabstractInternational Computer Science Institute (ICSI), USA(Author list is alphabetical with the exception of the typist.)ABSTRACTThis paper describes the research underway for the ESPRITWERNICKE project. The project brings together a num-ber of different groups from Europe and the US and focuseson extending the state-of-the-art for hybrid hidden Markovmodel/connectionist approaches to large vocabulary, continu-ous speech recognition. Thispaper describes the specific goalsoftheresearchandpresentstheworkperformedtodate. Resultsare reported for the resource management talker-independentrecognition task. The paper concludes with a discussion of theprojected future work.Keywords: Recognition, Neural Nets, HMM.1. BACKGROUNDW Tony Robinson, Luís B. Almeida, Jean-Marc Boite, Hervé Bourlard, Frank Fallside, Mike Hochberg, Dan J. Kershaw, Phil Kohn, Yochai Konig, Nelson Morgan, João Paulo da Silva Neto, Steve Renals, Marco Saerens, Chuck Wooters |
EUROSPEECH | 12 |
| 1993 | Learning Temporal Dependencies in Connectionist Speech Recognition
Steve Renals, Mike Hochberg, Anthony J. Robinson |
NIPS | 1 |
| 1993 | Hybrid Neural Network/Hidden Markov Model Systems for Continuous Speech RecognitionabstractMultiLayer Perceptrons (MLP) are an effective family of algorithms for the smooth estimation of highly-dimensioned probability density functions that are useful in continuous speech recognition. Hidden Markov Models (HMM) provide a structure for the mapping of a temporal sequence of acoustic vectors to a generating sequence of states. For HMMs that are independent of phonetic context, the MLP approaches have consistently provided significant improvements (once we learned how to use them). Recently, these results have been extended to context-dependent models. In this paper, after having reviewed the basic principles of our hybrid HMM/MLP approach, we describe a series of experiments with continuous speech recognition. The hybrid methods directly trade off computational complexity for reduced requirements of memory and memory bandwidth. Results are presented on the widely used Resource Management speech database that is distributed by the National Institute of Standards and Technology. These results demonstrate performance that is at least as good as any other reported continuous speech recognition system (for this task). Nelson Morgan, Hervé Bourlard, Steve Renals, Horacio Franco |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1992 | CDNN: a context dependent neural network for continuous speech recognitionabstractA series of theoretical and experimental results have suggested that multilayer perceptrons (MLPs) are an effective family of algorithms for the smooth estimate of highly dimensioned probability density functions that are useful in continuous speech recognition. All of these systems have exclusively used context-independent phonetic models, in the sense that the probabilities or costs are estimated for simple speech units such as phonemes or words, rather than biphones or triphones. Numerous conventional systems based on hidden Markov models (HMMs) have been reported that use triphone or triphone like context-dependent models. In one case the outputs of many context-dependent MLPs (one per context class) were used to help choose the best sentence from the N best sentences as determined by a context-dependent HMM system. It is shown how, without any simplifying assumptions, one can estimate likelihoods for context-dependent phonetic models with nets that are not substantially larger than context-independent MLPs.> Hervé Bourlard, Nelson Morgan, Chuck Wooters, Steve Renals |
ICASSP | 4 |
| 1992 | Connectionist probability estimation in the DECIPHER speech recognition systemabstractThe authors have previously demonstrated that feedforward networks can be used to estimate local output probabilities in hidden Markov model (HMM) speech recognition systems (Renals et al., 1991). These connectionist techniques are integrated into the DECIPHER system, with experiments being performed using the speaker-independent DARPA RM database. The results indicate that: connectionist probability estimation can improve performance of a context-independent maximum-likelihood-trained HMM system; performance of the connectionist system is close to what can be achieved using (context-dependent) HMM systems of much higher complexity; and mixing connectionist and maximum-likelihood estimates can improve the performance of the state-of-the-art context-independent HMM system.> Steve Renals, Nelson Morgan, Horacio Franco |
ICASSP | 1 |
| 1992 | Neural nets and hidden Markov models: Review and generalizations
Hervé Bourlard, Nelson Morgan, Steve Renals |
Speech Commun. | 3 |
| 1991 | A comparative study of continuous speech recognition using neural networks and hidden Markov modelsabstractThe recognition performances of two front ends are compared for two continuous speech recognition tasks. First, a neural network model (NNM) front end was used, with frame labeling performed by a radial basis function network and segmentation by a Viterbi algorithm. The second front end was a discrete hidden Markov model (HMM), featuring explicit state duration probability distributions. Two experiments were performed. The first used a speaker-dependent database, with a lexicon of 571 words. Using a low-perplexity grammar, the NNM front end produced a word accuracy of 94% and a sentence accuracy of 86%. This was slightly inferior to the HMM front end, which produced word accuracies of 96% and sentence accuracies of 88%. Without a grammar, word accuracies of 58% (NNM) and 49% (HMM) were recorded. The second set of experiments used the MIT portion of the TIMIT database (415 speakers and 2072 sentences in total). Results were poor for both front ends, with the NNM producing marginally better results.> Steve Renals, David McKelvie, Fergus R. McInnes |
ICASSP | 1 |
| 1991 | Connectionist Optimisation of Tied Mixture Hidden Markov Models
Steve Renals, Nelson Morgan, Hervé Bourlard, Horacio Franco |
NIPS | 1 |
| 1989 | Learning phoneme recognition using neural networksabstractThe authors have applied two neural-network models (back-propagation network and radial-basis-functions network) to a static speech recognition problem. The radial-basis-functions network offers training times of over two orders of magnitude faster than back-propagation, when training networks to similar power and generality. The authors have computed recognition statistics of the two models with varying numbers of hidden units on this recognition problem. The back-propagation network may offer increased generalization and robustness. Both models compare favorably with a vector-quantized hidden Markov model on the same problem.> Steve Renals, Richard Rohwer |
ICASSP | 1 |
| 1989 | Analysis of a neural network model for speech recognition
Steve Renals, Jonathan Dalby |
EUROSPEECH | 1 |
| 1988 | Unstable connectionist networks in speech recognitionabstractConnectionist networks evolve in time according to a prescribed rule. Typically, they are designed to be stable so that their temporal activity ceases after a short transient period. However, meaningful patterns in speech have a temporal component: therefore it seems natural to attempt to map the temporality of speech patterns onto the temporality of an unstable network. The authors have begun some exploratory experiments to train networks to recognise temporal patterns. They have designed fully connected networks that are trained to emulate and classify sequences by regarding each temporal state of a network as a layer in a feedforward network. Training is then performed by a variant of the back-propagation algorithm. They have conducted initial experiments using the output of a peripheral auditory model.> Richard Rohwer, Steve Renals, Mark Terry |
ICASSP | 2 |
| 1988 | A connectionist approach to speech recognition using peripheral auditory modellingabstractA prototype isolated word recogniser was constructed, with an auditory-based analysis component and a pattern classification module based on a parallel distributed processing paradigm. The auditory model used was a band-pass non-linear (BPNL) configuration which incorporates the effects of lateral suppression. Pattern classification was performed by a layered, feed-forward neural network, consisting of an array of input nodes representing the binary features output by the auditory model, a set of hidden nodes and an array of output nodes representing the word to be recognised. A suitable internal representation was learned by the method of back-propagation of errors by gradient descent using the generalised delta rule. This prototype recogniser was trained to recognise English digits spoken by male and female speakers. Recognition rates for the digit set, (zero to ten) were better than 80%.> Mark Terry, Steve Renals, Richard Rohwer, Jonathan Harrington |
ICASSP | 2 |