EDBT 2026 Demo / reviewers in the wild / expert
Simon King 0001
dblp:68/2005
· DBLP profile ↗
190ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0002-2694-2843ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 162 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 135 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can We Reconstruct a Dysarthric Voice with the Large Speech Model Parler TTS?abstractSpeech disorders can make communication hard or even impossible for those who develop them. Personalised Text-to-Speech is an attractive option as a communication aid. We attempt voice reconstruction using a large speech model, with which we generate an approximation of a dysarthric speaker's voice prior to the onset of their condition. In particular, we investigate whether a state-of-the-art large speech model, Parler TTS, can generate intelligible speech while maintaining speaker identity. We curate a dataset and annotate it with relevant speaker and intelligibility information, and use this to fine-tune the model. Our results show that the model can indeed learn to generate from the distribution of this challenging data, but struggles to control intelligibility and to maintain consistent speaker identity. We propose future directions to improve controllability of this class of model, for the voice reconstruction task. Ariadna Sanchez, Simon King 0001 |
INTERSPEECH | 2 |
| 2025 | Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
Nicholas Sanders, Yuanchao Li, Korin Richmond, Simon King 0001 |
INTERSPEECH | 4 |
| 2025 | Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 2023abstractInternational audience Olivier Perrotin, Brooke Stephenson, Silvain Gerber, Gérard Bailly, Simon King 0001 |
Comput. Speech Lang. | 5 |
| 2024 | Controllable Speaking Styles Using A Large Language ModelabstractReference-based Text-to-Speech (TTS) models can generate multiple, prosodically-different renditions of the same target text. Such models jointly learn a latent acoustic space during training, which can be sampled from during inference. Controlling these models during inference typically requires finding an appropriate reference utterance, which is non-trivial.Large generative language models (LLMs) have shown excellent performance in various language-related tasks. Given only a natural language query text (the ‘prompt’), such models can be used to solve specific, context-dependent tasks. Recent work in TTS has attempted similar prompt-based control of novel speaking style generation. Those methods do not require a reference utterance and can, under ideal conditions, be controlled with only a prompt. But existing methods typically require a prompt-labelled speech corpus for jointly training a prompt-conditioned encoder.In contrast, we instead employ an LLM to directly suggest prosodic modifications for a controllable TTS model, using contextual information provided in the prompt. The prompt can be designed for a multitude of tasks. Here, we give two demonstrations: control of speaking style; prosody appropriate for a given dialogue context. The proposed method is rated most appropriate in 50% of cases vs. 31% for a baseline model. Atli Sigurgeirsson, Simon King 0001 |
ICASSP | 2 |
| 2024 | The limits of the Mean Opinion Score for speech synthesis evaluation
Sébastien Le Maguer, Simon King 0001, Naomi Harte |
Comput. Speech Lang. | 2 |
| 2023 | Do Prosody Transfer Models Transfer ProsodyƒabstractSome recent models for Text-to-Speech synthesis aim to transfer the prosody of a reference utterance to the generated target synthetic speech. This is done by using a learned embedding of the reference utterance, which is used to condition speech generation. During training, the reference utterance is identical to the target utterance. Yet, during synthesis, these models are often used to transfer prosody from a reference that differs from the text or speaker being synthesized.To address this inconsistency, we propose to use a different, but prosodically-related, utterance during training too. We believe this should encourage the model to learn to transfer only those characteristics that the reference and target have in common. If prosody transfer methods do indeed transfer prosody they should be able to be trained in the way we propose. However, results show that a model trained under these conditions performs significantly worse than one trained using the target utterance as a reference. To explain this, we hypothesize that prosody transfer models do not learn a transferable representation of prosody, but rather a utterance-level representation which is highly dependent on both the reference speaker and reference text. Atli Sigurgeirsson, Simon King 0001 |
ICASSP | 2 |
| 2023 | Ensemble Prosody Prediction For Expressive Speech SynthesisabstractGenerating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Yet, in evaluations it is generally found that no single model is preferred for all input texts. This suggests an approach that has rarely been used before for Text-to-Speech: an ensemble of models.We apply ensemble learning to prosody prediction. We construct simple ensembles of prosody predictors by varying either model architecture or model parameter values.To automatically select amongst the models in the ensemble when performing Text-to-Speech, we propose a novel, and computationally trivial, variance-based criterion. We demonstrate that even a small ensemble of prosody predictors yields useful diversity, which, combined with the proposed selection criterion, outperforms any individual model from the ensemble. Tian Huey Teh, Vivian Hu, Devang S. Ram Mohan, Zack Hodari, Christopher G. R. Wallis, Tomás Gómez Ibarrondo, Alexandra Torresquintero, James Leoni, Mark J. F. Gales, Simon King 0001 |
ICASSP | 10 |
| 2023 | Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal ProcessingabstractMost state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation.A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from a mel-spectrogram requires computationally expensive machine learning: a neural vocoder. Our proposed "autovocoder" reverses this arrangement. We use machine learning to obtain a representation that replaces the mel-spectrogram, and that can be inverted back to a waveform using simple, fast operations including a differentiable implementation of the inverse STFT.The autovocoder generates a waveform 5 times faster than the DSP-based Griffin-Lim algorithm, and 14 times faster than the neural vocoder HiFi-GAN. We provide perceptual listening test results to confirm that the speech is of comparable quality to HiFi-GAN in the copy synthesis task. Jacob J. Webber, Cassia Valentini-Botinhao, Evelyn Williams, Gustav Eje Henter, Simon King 0001 |
ICASSP | 5 |
| 2023 | Intonation Control for Neural Text-to-Speech Synthesis with Polynomial Models of F0
Niamh Corkey, Johannah O'Mahony, Simon King 0001 |
INTERSPEECH | 3 |
| 2022 | Speech Audio Corrector: using speech from non-target speakers for one-off correction of mispronunciations in grapheme-input text-to-speechabstractCorrect pronunciation is essential for text-to-speech (TTS) systems in production. Most production systems rely on pronouncing dictionaries to perform grapheme-to-phoneme conversion. Unlike end-to-end TTS, this enables pronunciation correction by manually altering the phoneme sequence, but the necessary dictionaries are labour-intensive to create and only exist in a few high-resourced languages. This work demonstrates that accurate TTS pronunciation control can be achieved without a dictionary. Moreover, we show that such control can be performed without requiring any model retraining or fine-tuning, merely by supplying a single correctly-pronounced reading of a word in a different voice and accent at synthesis time. Experimental results show that our proposed system successfully enables one-off correction of mispronunciations in grapheme-based TTS with maintained synthesis quality. This opens the door to production-level TTS in languages and applications where pronunciation dictionaries are unavailable. Jason Fong, Daniel Lyth, Gustav Eje Henter, Simon King 0001 |
INTERSPEECH | 5 |
| 2022 | Back to the Future: Extending the Blizzard Challenge 2013
Sébastien Le Maguer, Simon King 0001, Naomi Harte |
INTERSPEECH | 2 |
| 2022 | Combining conversational speech with read speech to improve prosody in Text-to-Speech synthesisabstractFor isolated utterances, speech synthesis quality has improved immensely thanks to the use of sequence-to-sequence models. However, these models are generally trained on read speech and fail to generalise to unseen speaking styles. Recently, more re-search is focused on the synthesis of expressive and conversa-tional speech. Conversational speech contains many prosodic phenomena that are not present in read speech. We would like to learn these prosodic patterns from data, but unfortunately, many large conversational corpora are unsuitable for speech synthesis due to low audio quality. We investigate whether a data mixing strategy can improve conversational prosody for a target voice based on monologue data from audiobooks by adding real con-versational data from podcasts. We filter the podcast data to create a set of 26k question and answer pairs. We evaluate two FastPitch models: one trained on 20 hours of monologue speech from a single speaker, and another trained on 5 hours of monologue speech from that speaker plus 15 hours of ques-tions and answers spoken by nearly 15k speakers. Results from three listening tests show that the second model generates more preferred question prosody. Johannah O'Mahony, Catherine Lai, Simon King 0001 |
INTERSPEECH | 3 |
| 2021 | Ctrl-P: Temporal Control of Prosodic Variation for Speech SynthesisabstractText does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data is to provide acoustic information as an additional learning signal. When generating speech, modifying this acoustic information enables multiple distinct renditions of a text to be produced. Since much of the unexplained variation is in the prosody, we propose a model that generates speech explicitly conditioned on the three primary acoustic correlates of prosody: $F_{0}$, energy and duration. The model is flexible about how the values of these features are specified: they can be externally provided, or predicted from text, or predicted then subsequently modified. Compared to a model that employs a variational auto-encoder to learn unsupervised latent features, our model provides more interpretable, temporally-precise, and disentangled control. When automatically predicting the acoustic features from text, it generates speech that is more natural than that from a Tacotron 2 model with reference encoder. Subsequent human-in-the-loop modification of the predicted acoustic features can significantly further increase naturalness. Devang S. Ram Mohan, Qinmin Hu, Tian Huey Teh, Alexandra Torresquintero, Christopher G. R. Wallis, Marlene Staib, Lorenzo Foglianti, Jiameng Gao, Simon King 0001 |
Interspeech | 9 |
| 2021 | ADEPT: A Dataset for Evaluating Prosody TransferabstractText-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity.One popular method is to transfer the prosody from a reference speech sample.There have been considerable advances in using prosody transfer to generate more expressive speech, but the field lacks a clear definition of what successful prosody transfer means and a method for measuring it.We introduce a dataset of prosodically-varied reference natural speech samples for evaluating prosody transfer.The samples include global variations reflecting emotion and interpersonal attitude, and local variations reflecting topical emphasis, propositional attitude, syntactic phrasing and marked tonicity.The corpus only includes prosodic variations that listeners are able to distinguish with reasonable accuracy, and we report these figures as a benchmark against which text-to-speech prosody transfer can be compared.We conclude the paper with a demonstration of our proposed evaluation methodology, using the corpus to evaluate two textto-speech models that perform prosody transfer. Alexandra Torresquintero, Tian Huey Teh, Christopher G. R. Wallis, Marlene Staib, Devang S. Ram Mohan, Vivian Hu, Lorenzo Foglianti, Jiameng Gao, Simon King 0001 |
Interspeech | 9 |
| 2021 | Detection and Analysis of Attention Errors in Sequence-to-Sequence Text-to-SpeechabstractSequence-to-sequence speech synthesis models are notorious for gross errors such as skipping and repetition, commonly associated with failures in the attention mechanism. While a lot has been done to improve attention and decrease errors, this paper focuses instead on automatic error detection and analysis. We evaluated three objective metrics against error detection scores collected by human listening. All metrics were derived from the synthesised attention matrix alone and do not require a reference signal, relying on the expectation that errors occur when attention is dispersed or insufficient. Using one of this metrics as an analysis tool, we observed that gross errors are more likely to occur in longer sentences and in sentences with punctuation marks that indicate pause or break. We also found that mechanisms such as forcibly incremented attention have the potential for decreasing gross errors but to the detriment of naturalness. The results of the error detection evaluation revealed that two of the evaluated metrics were able to detect errors with a relatively high success rate, obtaining F-scores of up to 0.89 and 0.96. Cassia Valentini-Botinhao, Simon King 0001 |
Interspeech | 2 |
| 2021 | An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep LearningabstractSpeaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech processing techniques, such as speech analysis, spectral conversion, prosody conversion, speaker characterization, and vocoding. With the recent advances in theory and practice, we are now able to produce human-like voice quality with high speaker similarity. In this article, we provide a comprehensive overview of the state-of-the-art of voice conversion techniques and their performance evaluation methods from the statistical approaches to deep learning, and discuss their promise and limitations. We will also report the recent Voice Conversion Challenges (VCC), the performance of the current state of technology, and provide a summary of the available resources for voice conversion research. Berrak Sisman, Junichi Yamagishi, Simon King 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Speaker Adaptation of a Multilingual Acoustic Model for Cross-Language SynthesisabstractSeveral studies have shown promising results in adapting DNN-based acoustic models as a mechanism to transfer characteristics from pre-trained models. One such example is speaker adaptation using a small amount of data, where fine-tuning has helped train models that extrapolate well to diverse linguistic contexts that are not present in the adaptation data. In the current work, our objective is to synthesize speech in different languages using the target speaker's voice, regardless of the language of their data. To achieve this goal, we create a multilingual model using a corpus that consists of recordings from a large number of monolingual and a few bilingual speakers in multiple languages. The model is then adapted using the target speaker's recordings in a language other than the target language. We also explore if additional adaptation data from a native speaker of the target language improves the performance. The subjective evaluation shows that the proposed approach of cross-language speaker adaptation is able to synthesize speech in the target language, in the target speaker's voice, without data spoken by the target speaker in that language. Also, extra data from a native speaker of the target language can improve model performance. Ivan Himawan, Sandesh Aryal, Iris Ouyang, Sam Kang, Pierre Lanchantin, Simon King 0001 |
ICASSP | 6 |
| 2020 | A Sound Engineering Approach to Near End Listening Enhancement
Carol Chermaz, Simon King 0001 |
INTERSPEECH | 2 |
| 2020 | Testing the Limits of Representation Mixing for Pronunciation Correction in End-to-End Speech SynthesisabstractAccurate pronunciation is an essential requirement for text-to-speech (TTS) systems. Systems trained on raw text exhibit pronunciation errors in output speech due to ambiguous letter-to-sound relations. Without an intermediate phonemic representation, it is difficult to intervene and correct these errors. Retaining explicit control over pronunciation runs counter to the current drive toward end-to-end (E2E) TTS using sequence-to-sequence models. On the one hand, E2E TTS aims to eliminate manual intervention, especially expert skill such as phonemic transcription of words in a lexicon. On the other, a system making difficult-to-correct pronunciation errors is of little practical use. Some intervention is necessary. We explore the minimal amount of linguistic features required to correct pronunciation errors in an otherwise E2E TTS system that accepts graphemic input. We use representation-mixing: within each sequence the system accepts either graphemic and/or phonemic input. We quantify how little training data needs to be phonemically labelled - that is, how small a lexicon must be written - to ensure control over pronunciation. We find modest correction is possible with 500 phonemised word types from the LJ speech dataset but correction works best when the majority of word types are phonemised with syllable boundaries. Jason Fong, Simon King 0001 |
INTERSPEECH | 3 |
| 2020 | An Unsupervised Method to Select a Speaker Subset from Large Multi-Speaker Speech Synthesis DatasetsabstractLarge multi-speaker datasets for TTS typically contain diverse speakers, recording conditions, styles and quality of data. Although one might generally presume that more data is better, in this paper we show that a model trained on a carefully-chosen subset of speakers from LibriTTS provides significantly better quality synthetic speech than a model trained on a larger set. We propose an unsupervised methodology to find this subset by clustering per-speaker acoustic representations. Pilar Oplustil, Jennifer Williams 0001, Joanna Rownicka, Simon King 0001 |
INTERSPEECH | 4 |
| 2020 | Hider-Finder-Combiner: An Adversarial Architecture for General Speech Signal ModificationabstractInternational audience Jacob J. Webber, Olivier Perrotin, Simon King 0001 |
INTERSPEECH | 3 |
| 2020 | A Vector Quantized Variational Autoencoder (VQ-VAE) Autoregressive Neural F0 Model for Statistical Parametric Speech SynthesisabstractRecurrent neural networks (RNNs) can predict fundamental frequency (F0) for statistical parametric speech synthesis systems, given linguistic features as input. However, these models assume conditional independence between consecutive F0values, given the RNN state. In a previous study, we proposed autoregressive (AR) neural F0models to capture the causal dependency of successive F0values. In subjective evaluations, a deep AR model (DAR) outperformed an RNN. Here, we propose a Vector Quantized Variational Autoencoder (VQ-VAE) neural F0model that is both more efficient and more interpretable than the DAR. This model has two stages: one uses the VQ-VAE framework to learn a latent code for the F0contour of each linguistic unit, and other learns to map from linguistic features to latent codes. In contrast to the DAR and RNN, which process the input linguistic features frame-by-frame, the new model converts one linguistic feature vector into one latent code for each linguistic unit. The new model achieves better objective scores than the DAR, has a smaller memory footprint and is computationally faster. Visualization of the latent codes for phones and moras reveals that each latent code represents an F0shape for a linguistic unit. Xin Wang 0037, Shinji Takaki, Junichi Yamagishi, Simon King 0001, Keiichi Tokuda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Attentive Filtering Networks for Audio Replay Attack DetectionabstractAn attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intention of measuring the limits of replay attack detection as well as developing countermeasures against them. In this work, we propose our replay attacks detection system - Attentive Filtering Network, which is composed of an attention-based filtering mechanism that enhances feature representations in both the frequency and time domains, and a ResNet-based classifier. We show that the network enables us to visualize the automatically acquired feature representations that are helpful for spoofing detection. Attentive Filtering Network attains an evaluation EER of 8.99% on the ASVspoof 2017 Version 2.0 dataset. With system fusion, our best system further obtains a 30% relative improvement over the ASVspoof 2017 enhanced baseline system. Cheng-I Lai, Alberto Abad, Korin Richmond, Junichi Yamagishi, Najim Dehak, Simon King 0001 |
ICASSP | 6 |
| 2019 | Speech Waveform Reconstruction Using Convolutional Neural Networks with Noise and Periodic InputsabstractThis paper presents a method for upsampling and transforming a compact representation of acoustics into a corresponding speech waveform. Similar to a conventional vocoder, the proposed system takes a pulse train derived from fundamental frequency and a noise sequence as inputs and shapes them to be consistent with the acoustic features. However, the filters that are used to shape the waveform in the proposed system are learned from data, and take the form of layers in a convolutional neural network. Because the network performs the transformation simultaneously for all waveform samples in a sentence, its synthesis speed is comparable with that of conventional vocoders on CPU, and many times faster on GPU. It is trained directly in a fast and straightforward manner, using a combined time- and frequency-domain objective function. We use publicly available data and provide code to allow our results to be reproduced. Oliver Watts, Cassia Valentini-Botinhao, Simon King 0001 |
ICASSP | 3 |
| 2019 | Improving Speech Synthesis with Discourse RelationsabstractThis paper explores whether adding Discourse Relation (DR) features improves the naturalness of neural statistical parametric speech synthesis (SPSS) in English. We hypothesize first - in the light of several previous studies - that DRs have a dedicated prosodic encoding. Secondly, we hypothesize that encoding DRs in a speech synthesizer's input will improve the naturalness of its output. In order to test our hypotheses, we prepare a dataset of DR-annotated transcriptions of audiobooks in English. We then perform an acoustic analysis of the corpus which supports our first hypothesis that DRs are acoustically encoded in speech prosody. The analysis reveals significant correlation between specific DR categories and acoustic features, such as F0 and intensity. Then, we use the corpus to train a neural SPSS system in two configurations: a baseline configuration making use only of conventional linguistic features, and an experimental one where these are supplemented with DRs. Augmenting the inputs with DR features improves objective acoustic scores on a test set and leads to significant preference by listeners in a forced choice AB test for naturalness. Adèle Aubin, Alessandra Cervone, Oliver Watts, Simon King 0001 |
INTERSPEECH | 4 |
| 2019 | Evaluating Near End Listening Enhancement Algorithms in Realistic EnvironmentsabstractSpeech playback (e.g., TV, radio, public address) becomes harder to understand in the presence of noise and reverberation.\nNELE (Near End Listening Enhancement) algorithms can improve intelligibility by modifying the signal before it is played back. Substantial intelligibility improvements have been achieved in the lab for both natural and synthetic speech. However, evidence is still scarce on how these algorithms work under conditions of realistic noise and reverberation.\n\nWe present a realistic test platform, featuring two representative everyday scenarios in which speech playback may occur (in the presence of both noise and reverberation): a domestic space (living room) and a public space (cafeteria). The generated stimuli are evaluated by measuring keyword accuracy rates in a listening test with normal hearing subjects.\n\nWe use the new platform to compare three state-of-theart NELE algorithms, employing either noise-adaptive or nonadaptive\nstrategies, and with or without compensation for reverberation. Carol Chermaz, Cassia Valentini-Botinhao, Henning F. Schepker, Simon King 0001 |
INTERSPEECH | 4 |
| 2019 | Investigating the Robustness of Sequence-to-Sequence Text-to-Speech Models to Imperfectly-Transcribed Training Data
Jason Fong, Pilar Oplustil, Zack Hodari, Simon King 0001 |
INTERSPEECH | 4 |
| 2019 | Using Pupil Dilation to Measure Cognitive Load When Listening to Text-to-Speech in Quiet and in NoiseabstractWith increased use of text-to-speech (TTS) systems in real-world applications, evaluating how such systems influence the human cognitive processing system becomes important. Particularly in situations where cognitive load is high, there may be negative implications such as fatigue. For example, noisy situations generally require the listener to exert increased mental effort. A better understanding of this could eventually suggest new ways of generating synthetic speech that demands low cognitive load. In our previous study, pupil dilation was used as an index of cognitive effort. Pupil dilation was shown to be sensitive to the quality of synthetic speech, but there were some uncertainties regarding exactly what was being measured. The current study resolves some of those uncertainties. Additionally, we investigate how the pupil dilates when listening to synthetic speech in the presence of speech-shaped noise. Our results show that, in quiet listening conditions, pupil dilation does not reflect listening effort but rather attention and engagement. In noisy conditions, increased pupil dilation indicates that listening effort increases as signal-to-noise ratio decreases, under all conditions tested. Avashna Govender, Anita E. Wagner, Simon King 0001 |
INTERSPEECH | 3 |
| 2019 | Disentangling Style Factors from Speaker RepresentationsabstractOur goal is to separate out speaking style from speaker identity in utterance-level representations of speech such as i-vectors and x-vectors. We first show that both i-vectors and x-vectors contain information not only about speaker but also about speaking style (for one data set) or emotion (for another data set), even when projected into a low-dimensional space. To disentangle these factors, we use an autoencoder in which the latent space is split into two subspaces. The entangled information about speaker and style/emotion is pushed apart by the use of auxiliary classifiers that take one of the two latent subspaces as input and that are jointly learned with the autoencoder. We evaluate how well the latent subspaces separate the factors by using them as input to separate style/emotion classification tasks. In traditional speaker identification tasks, speaker-invariant characteristics are factorized from channel and then the channel information is ignored. Our results suggest that this so-called channel may contain exploitable information, which we refer to as style factors. Finally, we propose future work to use information theory to formalize style factors in the context of speaker identity. Jennifer Williams 0001, Simon King 0001 |
INTERSPEECH | 2 |
| 2018 | Using Pupillometry to Measure the Cognitive Load of Synthetic SpeechabstractIt is common to evaluate synthetic speech using listening tests in which intelligibility is measured by asking listeners to transcribe the words heard, and naturalness is measured using Mean Opinion Scores. But, for real-world applications of synthetic speech, the effort (cognitive load) required to understand the synthetic speech may be a more appropriate measure. Cognitive load has been investigated in the past, when rule-based speech synthesizers were popular, but there is little or no recent work using state-of-the-art text-to-speech. Studies on the understanding of natural speech have shown that the pupil dilates when increased mental effort is exerted to perform a task. We use pupillometry to measure the cognitive load of synthetic speech submitted to two of the Blizzard Challenge evaluations. Our results show that pupil dilation is sensitive to the quality of synthetic speech. In all cases, synthetic speech imposes a higher cognitive load than natural speech. Pupillometry is therefore proposed as a sensitive measure that can be used to evaluate\nsynthetic speech. Avashna Govender, Simon King 0001 |
INTERSPEECH | 2 |
| 2018 | Measuring the Cognitive Load of Synthetic Speech Using a Dual Task ParadigmabstractWe present a methodology for measuring the cognitive load (listening effort) of synthetic speech using a dual task paradigm.Cognitive load is calculated from changes in a listener's performance on a secondary task (e.g., reaction time to decide if a visually-displayed digit is odd or even).Previous related studies have only found significant differences between the best and worst quality systems but failed to separate the systems that lie in between.A paradigm that is sensitive enough to detect differences between state-of-the-art, high quality speech synthesizers would be very useful for advancing the state of the art.In our work, four speech synthesis systems from a previous Blizzard Challenge, and the corresponding natural speech, were compared.Our results show that reaction times slow down as speech quality reduces, as we expected: lower quality speech imposes a greater cognitive load, taking resources away from the secondary task.However, natural speech did not have the fastest reaction times.This intriguing result might indicate that, as speech synthesizers attain near-perfect intelligibility, this paradigm is measuring something like the listener's level of sustained attention and not listening effort. Avashna Govender, Simon King 0001 |
INTERSPEECH | 2 |
| 2018 | Learning Interpretable Control Dimensions for Speech Synthesis by Using External DataabstractThere are many aspects of speech that we might want to control when creating text-to-speech (TTS) systems. We present a general method that enables control of arbitrary aspects of speech, which we demonstrate on the task of emotion control. Current TTS systems use supervised machine learning and are therefore heavily reliant on labelled data. If no labels are available for a desired control dimension, then creating interpretable control becomes challenging. We introduce a method that uses external, labelled data (i.e. not the original data used to train the acoustic model) to enable the control of dimensions that are not labelled in the original data. Adding interpretable control allows the voice to be manually controlled to produce more engaging speech, for applications such as audiobooks. We evaluate our method using a listening test. Zack Hodari, Oliver Watts, Srikanth Ronanki, Simon King 0001 |
INTERSPEECH | 4 |
| 2018 | Impact of Different Speech Types on Listening EffortabstractListeners are exposed to different types of speech in everyday life, from natural speech to speech that has undergone modifications or has been generated synthetically. While many studies have focused on measuring the intelligibility of these distinct speech types, their impact on listening effort is not known. The current study combined an objective measure of intelligibility, a physiological measure of listening effort (pupil size) and listeners’ subjective judgements, to examine the impact of four speech types: plain (natural) speech, speech produced in noise (Lombard speech), speech enhanced to promote intelligibility, and synthetic speech. For each speech type, listeners responded to sentences presented in one of three levels of speech-shaped noise. Subjective effort ratings and intelligibility scores showed an inverse ranking across speech types, with synthetic speech being the most demanding and enhanced speech the least. Pupil size measures indicated an increase in listening effort with decreasing signal-to-noise ratio for all speech types apart from\nsynthetic speech, which required significantly more effort at the most favourable noise level. Naturally and artificially modified speech were less effortful than plain speech at the more adverse noise levels. These outcomes indicate a clear impact of speech type on the cognitive demands required for comprehension. Olympia Simantiraki, Martin Cooke, Simon King 0001 |
INTERSPEECH | 3 |
| 2018 | Exemplar-based Speech Waveform GenerationabstractThis paper presents a simple but effective method for generating speech waveforms by selecting small units of stored speech to match a low-dimensional target representation. The method is designed as a drop-in replacement for the vocoder in a deep neural network-based text-to-speech system. Most previous work on hybrid unit selection waveform generation relies on phonetic annotation for determining unit boundaries, or for specifying target cost, or for candidate preselection. In contrast, our waveform generator requires no phonetic information, annotation, or alignment. Unit boundaries are determined by epochs, and spectral analysis provides representations which are compared directly with target features at runtime. As in unit selection, we minimise a combination of target cost and join cost, but find that greedy left-to-right nearest-neighbour search gives similar results to dynamic programming. The method is fast and can generate the waveform incrementally. We use publicly available data and provide a permissively-licensed open source toolkit for reproducing our results. Oliver Watts, Cassia Valentini-Botinhao, Felipe Espic, Simon King 0001 |
INTERSPEECH | 4 |
| 2018 | Examplar-Based Speechwaveform Generation for Text-To-SpeechabstractThis paper presents a hybrid text-to-speech framework that uses a waveform generation method based on examplars of natural speech waveform. These examplars are selected at synthesis time given a sequence of acoustic features generated from text by a statistical parametric speech synthesis model. In order to match the expected degradation of these target synthesis features, the database of units is constructed such that the units' target representations are generated from the same parametric model. We evaluate two variants of this framework by modifying the size of the examplar: a small unit variant (where unit boundaries are determined by pitch mark location) and a halfphone variant (where unit boundaries are determined by subphone state forced alignment). We found that for a larger dataset (around four hours of training data) the examplar-based waveform generation variants are rated higher than the vocoder-based system. Cassia Valentini-Botinhao, Oliver Watts, Felipe Espic, Simon King 0001 |
SLT | 4 |
| 2017 | The blizzard machine learning challenge 2017abstractThis paper describes the Blizzard Machine Learning Challenge (BMLC) 2017, which is a spin-off of the Blizzard Challenge. The annual Blizzard Challenges 2005-2017 were held to better understand and compare research techniques in building corpus-based text-to-speech (TTS) systems on the same data. The series of Blizzard Challenges has helped us measure progress in TTS technology. However, to get competitive performance, a lot time has to be spent on skilled tasks. This may make the Blizzard Challenge unattractive to machine learning researchers from other fields. Therefore, we recommend that the BMLC not involve these speech-specific tasks and that it allow participants to concentrate on the acoustic modeling task, framed as a straightforward machine learning problem, with a fixed dataset. In the BMLC 2017, two types of datasets consisting of four hours of speech data suitable for machine learning problems were distributed. This paper summarizes the purpose, design, and whole process of the challenge and its results. Kei Sawada, Keiichi Tokuda, Simon King 0001, Alan W. Black |
ASRU | 3 |
| 2017 | Direct Modelling of Magnitude and Phase Spectra for Statistical Parametric Speech SynthesisabstractWe propose a simple new representation for the FFT spectrum tailored to statistical parametric speech synthesis. It consists of four feature streams that describe magnitude, phase and fundamental frequency using real numbers. The proposed feature extraction method does not attempt to decompose the speech structure (e.g., into source+filter or harmonics+noise). By avoiding the simplifications inherent in decomposition, we can dramatically reduce the “phasiness” and “buzziness” typical of most vocoders. The method uses simple and computationally cheap operations and can operate at a lower frame rate than the 200 frames-per-second typical in many systems. It avoids heuristics and methods requiring approximate or iterative solutions, including phase unwrapping. Two DNN-based acoustic models were built - from male and female speech data - using the Merlin toolkit. Subjective comparisons were made with a state-of-the-art baseline, using the STRAIGHT vocoder. In all variants tested, and for both male and female voices, the proposed method substantially outperformed the baseline. We provide source code to enable our complete system to be replicated. Felipe Espic, Cassia Valentini-Botinhao, Simon King 0001 |
INTERSPEECH | 3 |
| 2017 | Nativization of Foreign Names in TTS for Automatic Reading of World News in SwahiliabstractWhen a text-to-speech (TTS) system is required to speak world news, a large fraction of the words to be spoken will be proper names originating in a wide variety of languages. Phonetization of these names based on target language letter-to-sound rules will typically be inadequate. This is detrimental not only during synthesis, when inappropriate phone sequences are produced, but also during training, if the system is trained on data from the same domain. This is because poor phonetization during forced alignment based on hidden Markov models can pollute the whole model set, resulting in degraded alignment even of normal target-language words. This paper presents four techniques designed to address this issue in the context of a Swahili TTS system: automatic transcription of proper names based on a lexicon from a better-resourced language; the addition of a parallel phone set and special part-of-speech tag exclusively dedicated to proper names; a manually-crafted phone mapping which allows substitutions for potentially more accurate phones in proper names during forced alignment; the addition in proper names of a grapheme-derived frame-level feature, supplementing the standard phonetic inputs to the acoustic model. We present results from objective and subjective evaluations of systems built using these four techniques. Joseph Mendelson, Pilar Oplustil, Oliver Watts, Simon King 0001 |
INTERSPEECH | 4 |
| 2017 | A Hierarchical Encoder-Decoder Model for Statistical Parametric Speech SynthesisabstractCurrent approaches to statistical parametric speech synthesis using Neural Networks generally require input at the same temporal resolution as the output, typically a frame every 5ms, or in some cases at waveform sampling rate. It is therefore necessary to fabricate highly-redundant frame-level (or sample level) linguistic features at the input. This paper proposes the use of a hierarchical encoder-decoder model to perform the sequence-to-sequence regression in a way that takes the input linguistic features at their original timescales, and preserves the relationships between words, syllables and phones. The proposed model is designed to make more effective use of suprasegmental features than conventional architectures, as well as being computationally efficient. Experiments were conducted on prosodically-varied audiobook material because the use of supra-segmental features is thought to be particularly important in this case. Both objective measures and results from subjective listening tests, which asked listeners to focus on prosody, show that the proposed method performs significantly better than a conventional architecture that requires the linguistic input to be at the acoustic frame rate. We provide code and a recipe to enable our system to be reproduced using the Merlin toolkit. Srikanth Ronanki, Oliver Watts, Simon King 0001 |
INTERSPEECH | 3 |
| 2017 | Locally Normalized Filter Banks Applied to Deep Neural-Network-Based Robust Speech RecognitionabstractThis letter describes modifications to locally normalized filter banks (LNFB), which substantially improve their performance on the Aurora-4 robust speech recognition task using a Deep Neural Network-Hidden Markov Model (DNN-HMM)-based speech recognition system. The modified coefficients, referred to as LNFB features, are a filter-bank version of locally normalized cepstral coefficients (LNCC), which have been described previously. The ability of the LNFB features is enhanced through the use of newly proposed dynamic versions of them, which are developed using an approach that differs somewhat from the traditional development of delta and delta-delta features. Further enhancements are obtained through the use of mean normalization and mean-variance normalization, which is evaluated both on a per-speaker and a per-utterance basis. The best performing feature combination (typically LNFB combined with LNFB delta and delta-delta features and mean-variance normalization) provides an average relative reduction in word error rate of 11.4% and 9.4%, respectively, compared to comparable features derived from Mel filter banks when clean and multinoise training are used for the Aurora-4 evaluation. The results presented here suggest that the proposed technique is more robust to channel mismatches between training and testing data than MFCC-derived features and is more effective in dealing with channel diversity. Josué Fredes, José Novoa, Simon King 0001, Richard M. Stern, Néstor Becerra Yoma |
IEEE Signal Process. Lett. | 3 |
| 2017 | Using Eigenvoices and Nearest-Neighbors in HMM-Based Cross-Lingual Speaker Adaptation With Limited DataabstractCross-lingual speaker adaptation for speech synthesis has many applications, such as use in speech-to-speech translation systems. Here, we focus on cross-lingual adaptation for statistical speech synthesis systems using limited adaptation data. To that end, we propose two eigenvoice adaptation approaches exploiting a bilingual Turkish-English speech database that we collected. In one approach, eigenvoice weights extracted using Turkish adaptation data and Turkish voice models are transformed into the eigenvoice weights for the English voice models using linear regression. Weighting the samples depending on the distance of reference speakers to target speakers during linear regression was found to improve the performance. Moreover, importance weighting the elements of the eigenvectors during regression further improved the performance. The second approach proposed here is speaker-specific state-mapping, which performed significantly better than the baseline state-mapping algorithm both in objective and subjective tests. Performance of the proposed state mapping algorithm was further improved when it was used with the intralingual eigenvoice approach instead of the linear-regression based algorithms used in the baseline system. Seyyed Saeed Sarfjoo, Cenk Demiroglu, Simon King 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesisabstractForced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do not degrade models, as long as the mistakes are consistent across training and synthesis. We present evidence that in the alignment of both standard read prompts and spontaneous speech this phoneme sequence is often wrong, and that this is likely to have a negative impact on acoustic models. A lattice-based forced alignment system allowing for pronunciation variation is implemented, resulting in improved phoneme identity accuracy for both types of speech. A perceptual evaluation of HMM-based voices showed that spontaneous models trained on this improved alignment also improved standard synthesis, despite breaking the consistency assumption. Rasmus Dall, Sandrine Brognaux, Korin Richmond, Cassia Valentini-Botinhao, Gustav Eje Henter, Julia Hirschberg, Junichi Yamagishi, Simon King 0001 |
ICASSP | 8 |
| 2016 | Robust TTS duration modelling using DNNSabstractAccurate modelling and prediction of speech-sound durations is an important component in generating more natural synthetic speech. Deep neural networks (DNNs) offer a powerful modelling paradigm, and large, found corpora of natural and expressive speech are easy to acquire for training them. Unfortunately, found datasets are seldom subject to the quality-control that traditional synthesis methods expect. Common issues likely to affect duration modelling include transcription errors, reductions, filled pauses, and forced-alignment inaccuracies. To combat this, we propose to improve modelling and prediction of speech durations using methods from robust statistics, which are able to disregard ill-fitting points in the training material. We describe a robust fitting criterion based on the density power divergence (the ß-divergence) and a robust generation heuristic using mixture density networks (MDNs). Perceptual tests indicate that subjects prefer synthetic speech generated using robust models of duration over the baselines. Gustav Eje Henter, Srikanth Ronanki, Oliver Watts, Mirjam Wester, Zhizheng Wu 0001, Simon King 0001 |
ICASSP | 6 |
| 2016 | Deep neural network-guided unit selection synthesisabstractVocoding of speech is a standard part of statistical parametric speech synthesis systems. It imposes an upper bound of the naturalness that can possibly be achieved. Hybrid systems using parametric models to guide the selection of natural speech units can combine the benefits of robust statistical models with the high level of naturalness of waveform concatenation. Existing hybrid systems use Hidden Markov Models (HMMs) as the statistical model. This paper demonstrates that the superiority of Deep Neural Network (DNN) acoustic models over HMMs in conventional statistical parametric speech synthesis also carries over to hybrid synthesis. We compare various DNN and HMM hybrid configurations, guiding the selection of waveform units in either the vocoder parameter domain, or in the domain of embeddings (bottleneck features). Thomas Merritt, Robert A. J. Clark, Zhizheng Wu 0001, Junichi Yamagishi, Simon King 0001 |
ICASSP | 5 |
| 2016 | Smooth talking: Articulatory join costs for unit selectionabstractJoin cost calculation has so far dealt exclusively with acoustic speech parameters, and a large number of distance metrics have previously been tested in conjunction with a wide variety of acoustic parameterisations. In contrast, we propose here to calculate distance in articulatory space. The motivation for this is simple: physical constraints mean a human talker's mouth cannot "jump" from one configuration to a different one, so smooth evolution of articulator positions would also seem desirable for a good candidate unit sequence. To test this, we built Festival Multisyn voices using a large articulatory-acoustic dataset. We first synthesised 460 TIMIT sentences and confirmed our articulatory join cost gives appreciably different unit sequences compared to the standard Multisyn acoustic join cost. A listening test (3 sets of 25 sentence pairs, 30 listeners) then showed our articulatory cost is preferred at a rate of 58% compared to the standard Multisyn acoustic join cost. Korin Richmond, Simon King 0001 |
ICASSP | 2 |
| 2016 | From HMMS to DNNS: Where do the improvements come from?abstractDeep neural networks (DNNs) have recently been the focus of much text-to-speech research as a replacement for decision trees and hidden Markov models (HMMs) in statistical parametric synthesis systems. Performance improvements have been reported; however, the configuration of systems evaluated makes it impossible to judge how much of the improvement is due to the new machine learning methods, and how much is due to other novel aspects of the systems. Specifically, whereas the decision trees in HMM-based systems typically operate at the state-level, and separate trees are used to handle separate acoustic streams, most DNN-based systems are trained to make predictions simultaneously for all streams at the level of the acoustic frame. This paper isolates the influence of three factors (machine learning method; state vs. frame predictions; separate vs. combined stream predictions) by building a continuum of systems along which only a single factor is varied at a time. We find that replacing decision trees with DNNs and moving from state-level to frame-level predictions both significantly improve listeners' naturalness ratings of synthetic speech produced by the systems. No improvement is found to result from switching from separate-stream to combined-stream predictions. Oliver Watts, Gustav Eje Henter, Thomas Merritt, Zhizheng Wu 0001, Simon King 0001 |
ICASSP | 5 |
| 2016 | Investigating gated recurrent networks for speech synthesisabstractRecently, recurrent neural networks (RNNs) as powerful sequence models have re-emerged as a potential acoustic model for statistical parametric speech synthesis (SPSS). The long short-term memory (LSTM) architecture is particularly attractive because it addresses the vanishing gradient problem in standard RNNs, making them easier to train. Although recent studies have demonstrated that LSTMs can achieve significantly better performance on SPSS than deep feedforward neural networks, little is known about why. Here we attempt to answer two questions: a) why do LSTMs work well as a sequence model for SPSS; b) which component (e.g., input gate, output gate, forget gate) is most important. We present a visual analysis alongside a series of experiments, resulting in a proposal for a simplified architecture. The simplified architecture has significantly fewer parameters than an LSTM, thus reducing generation complexity considerably without degrading quality. Zhizheng Wu 0001, Simon King 0001 |
ICASSP | 2 |
| 2016 | GlottDNN - A Full-Band Glottal Vocoder for Statistical Parametric Speech SynthesisabstractGlottHMM is a previously developed vocoder that has been successfully used in HMM-based synthesis by parameterizing speech into two parts (glottal flow, vocal tract) according to the functioning of the real human voice production mechanism. In this study, a new glottal vocoding method, GlottDNN, is proposed. The GlottDNN vocoder is built on the principles of its predecessor, GlottHMM, but the new vocoder introduces three main improvements: GlottDNN (1) takes advantage of a new, more accurate glottal inverse filtering method, (2) uses a new method of deep neural network (DNN) -based glottal excitation generation, and (3) proposes a new approach of band-wise processing of full-band speech. The proposed GlottDNN vocoder was evaluated as part of a full-band state-of-the-art DNN-based text-to-speech (TTS) synthesis system, and compared against the release version of the original GlottHMM vocoder, and the well-known STRAIGHT vocoder. The results of the subjective listening test indicate that GlottDNN improves the TTS quality over the compared methods. Manu Airaksinen, Bajibabu Bollepalli, Lauri Juvela, Zhizheng Wu 0001, Simon King 0001, Paavo Alku |
INTERSPEECH | 5 |
| 2016 | Waveform Generation Based on Signal Reshaping for Statistical Parametric Speech SynthesisabstractWe propose a new paradigm of waveform generation for Statistical Parametric Speech Synthesis that is based on neither source-filter separation nor sinusoidal modelling. We suggest that one of the main problems of current vocoding techniques is that they perform an extreme decomposition of the speech signal into source and filter, which is an underlying cause of “buzziness”, “musical artifacts”, or “ muffled sound” in the synthetic speech. The proposed method avoids making unnecessary assumptions and decompositions as far as possible, and uses only the spectral envelope and F0 as parameters. Prerecorded speech is used as a base signal, which is “reshaped” to match the acoustic specification predicted by the statistical model, without any source-filter decomposition. A detailed description of the method is presented, including implementation details and adjustments. Subjective listening test evaluations of complete DNN-based text-to-speech systems were conducted for two voices: one female and one male. The results show that the proposed method tends to outperform the state-of-theart standard vocoder STRAIGHT, whilst using fewer acoustic parameters. Felipe Espic, Cassia Valentini-Botinhao, Zhizheng Wu 0001, Simon King 0001 |
INTERSPEECH | 4 |
| 2016 | The Use of Locally Normalized Cepstral Coefficients (LNCC) to Improve Speaker Recognition Accuracy in Highly Reverberant Rooms
Víctor Poblete, Juan Pablo Escudero, Josué Fredes, José Novoa, Richard M. Stern, Simon King 0001, Néstor Becerra Yoma |
INTERSPEECH | 6 |
| 2016 | A Template-Based Approach for Speech Synthesis Intonation Generation Using LSTMs
Srikanth Ronanki, Gustav Eje Henter, Zhizheng Wu 0001, Simon King 0001 |
INTERSPEECH | 4 |
| 2016 | Median-based generation of synthetic speech durations using a non-parametric approachabstractThis paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike conventional approaches to duration modelling - which assume that duration distributions have a particular form (e.g., a Gaussian) and use the mean of that distribution for synthesis - our approach can in principle model any distribution supported on the non-negative integers. Generation from this model can be performed in many ways; here we consider output generation based on the median predicted duration. The median is more typical (more probable) than the conventional mean duration, is robust to training-data irregularities, and enables incremental generation. Furthermore, a frame-level approach to duration prediction is consistent with a longer-term goal of modelling durations and acoustic features together. Results indicate that the proposed method is competitive with baseline approaches in approximating the median duration of held-out natural speech. Srikanth Ronanki, Oliver Watts, Simon King 0001, Gustav Eje Henter |
SLT | 3 |
| 2016 | ALISA: An automatic lightly supervised speech segmentation and alignment tool
Adriana Cornelia Stan, Yoshitaka Mamiya, Junichi Yamagishi, Peter Bell 0001, Oliver Watts, Robert A. J. Clark, Simon King 0001 |
Comput. Speech Lang. | 7 |
| 2016 | Improving Trajectory Modelling for DNN-Based Speech Synthesis by Using Stacked Bottleneck Features and Minimum Generation Error TrainingabstractWe propose two novel techniques-stacking bottleneck features and minimum generation error (MGE) training criterion-to improve the performance of deep neural network (DNN)-based speech synthesis. The techniques address the related issues of frame-by-frame independence and ignorance of the relationship between static and dynamic features, within current typical DNN-based synthesis frameworks. Stacking bottleneck features, which are an acoustically informed linguistic representation, provides an efficient way to include more detailed linguistic context at the input. The MGE training criterion minimises overall output trajectory error across an utterance, rather than minimising the error per frame independently, and thus takes into account the interaction between static and dynamic features. The two techniques can be easily combined to further improve performance. We present both objective and subjective results that demonstrate the effectiveness of the proposed techniques. The subjective results show that combining the two techniques leads to significantly more natural synthetic speech than from conventional DNN or long short-term memory recurrent neural network systems. Zhizheng Wu 0001, Simon King 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Anti-Spoofing for Text-Independent Speaker Verification: An Initial Database, Comparison of Countermeasures, and Human PerformanceabstractIn this paper, we present a systematic study of the vulnerability of automatic speaker verification to a diverse range of spoofing attacks. We start with a thorough analysis of the spoofing effects of five speech synthesis and eight voice conversion systems, and the vulnerability of three speaker verification systems under those attacks. We then introduce a number of countermeasures to prevent spoofing attacks from both known and unknown attackers. Known attackers are spoofing systems whose output was used to train the countermeasures, while an unknown attacker is a spoofing system whose output was not available to the countermeasures during training. Finally, we benchmark automatic systems against human performance on both speaker verification and spoofing detection tasks. Zhizheng Wu 0001, Phillip L. De Leon, Cenk Demiroglu, Ali Khodabakhsh 0001, Simon King 0001, Zhen-Hua Ling, Daisuke Saito, Bryan Stewart, Tomoki Toda, Mirjam Wester, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2015 | Attributing modelling errors in HMM synthesis by stepping gradually from natural to modelled speechabstractEven the best statistical parametric speech synthesis systems do not achieve the naturalness of good unit selection. We investigated possible causes of this. By constructing speech signals that lie in between natural speech and the output from a complete HMM synthesis system, we investigated various effects of modelling. We manipulated the temporal smoothness and the variance of the spectral parameters to create stimuli, then presented these to listeners alongside natural and vocoded speech, as well as output from a full HMM-based text-to-speech system and from an idealised `pseudo-HMM'. All speech signals, except the natural waveform, were created using vocoders employing one of two popular spectral parameterisations: Mel-Cepstra or Mel-Line Spectral Pairs. Listeners made `same or different' pairwise judgements, from which we generated a perceptual map using Multidimensional Scaling. We draw conclusions about which aspects of HMM synthesis are limiting the naturalness of the synthetic speech. Thomas Merritt, Javier Latorre, Simon King 0001 |
ICASSP | 3 |
| 2015 | SAS: A speaker verification spoofing database containing diverse attacksabstractThis paper presents the first version of a speaker verification spoofing and anti-spoofing database, named SAS corpus. The corpus includes nine spoofing techniques, two of which are speech synthesis, and seven are voice conversion. We design two protocols, one for standard speaker verification evaluation, and the other for producing spoofing materials. Hence, they allow the speech synthesis community to produce spoofing materials incrementally without knowledge of speaker verification spoofing and anti-spoofing. To provide a set of preliminary results, we conducted speaker verification experiments using two state-of-the-art systems. Without any anti-spoofing techniques, the two systems are extremely vulnerable to the spoofing attacks implemented in our SAS corpus. Zhizheng Wu 0001, Ali Khodabakhsh 0001, Cenk Demiroglu, Junichi Yamagishi, Daisuke Saito, Tomoki Toda, Simon King 0001 |
ICASSP | 7 |
| 2015 | Deep neural networks employing Multi-Task Learning and stacked bottleneck features for speech synthesisabstractDeep neural networks (DNNs) use a cascade of hidden representations to enable the learning of complex mappings from input to output features. They are able to learn the complex mapping from text-based linguistic features to speech acoustic features, and so perform text-to-speech synthesis. Recent results suggest that DNNs can produce more natural synthetic speech than conventional HMM-based statistical parametric systems. In this paper, we show that the hidden representation used within a DNN can be improved through the use of Multi-Task Learning, and that stacking multiple frames of hidden layer activations (stacked bottleneck features) also leads to improvements. Experimental results confirmed the effectiveness of the proposed methods, and in listening tests we find that stacked bottleneck features in particular offer a significant improvement over both a baseline DNN and a benchmark HMM system. Zhizheng Wu 0001, Cassia Valentini-Botinhao, Oliver Watts, Simon King 0001 |
ICASSP | 4 |
| 2015 | Robustness to additive noise of locally-normalized cepstral coefficients in speaker verification
Josué Fredes, José Novoa, Víctor Poblete, Simon King 0001, Richard M. Stern, Néstor Becerra Yoma |
INTERSPEECH | 4 |
| 2015 | Reconstructing voices within the multiple-average-voice-model frameworkabstractPersonalisation of voice output communication aids (VOCAs) allows to preserve the vocal identity of people suffering from speech disorders. This can be achieved by the adaptation of HMM-based speech synthesis systems using a small amount of adaptation data. When the voice has begun to deteriorate, reconstruction is still possible in the statistical domain by correcting the parameters of the models associated with the speech disorder. This can be done by substituting those with parameters from a donor’s voice, at risk of losing part of the identity of the patient. Recently, the Multiple-Average-Voice-Model (Multiple AVM) framework has been proposed for speaker adaptation. Adaptation is performed via interpolation into a speaker eigenspace spanned by the mean vectors of speaker-adapted AVMs which can be tuned to the individual speaker. In this paper, we present the benefits of this framework for voice reconstruction: it requires only a very small amount of adaptation data, interpolation can be performed in a clean speech eigenspace and the resulting voice can be easily fine-tuned by acting on the interpolation weights. We illustrate our points with a subjective assessment of the reconstructed voice. Index Terms: HMM-Based speech synthesis, speaker adaptation, multiple average voice model, cluster adaptive training, voice reconstruction, voice output communication aids. Pierre Lanchantin, Christophe Veaux, Mark J. F. Gales, Simon King 0001, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2015 | Deep neural network context embeddings for model selection in rich-context HMM synthesisabstractThis paper introduces a novel form of parametric synthesis that uses context embeddings produced by the bottleneck layer of a deep neural network to guide the selection of models in a rich-context HMM-based synthesiser. Rich-context synthesis – in which Gaussian distributions estimated from single lin-guistic contexts seen in the training data are used for synthesis, rather than more conventional decision tree-tied models – was originally proposed to address over-smoothing due to averag-ing across contexts. Our previous investigations have confirmed experimentally that averaging across different contexts is in-deed one of the largest factors contributing to the limited quality of statistical parametric speech synthesis. However, a possible weakness of the rich context approach as previously formulated is that a conventional tied model is still used to guide selection of Gaussians at synthesis time. Our proposed approach replaces this with context embeddings derived from a neural network. Index Terms: speech synthesis, hidden Markov model, deep neural networks, rich context, embedding Thomas Merritt, Junichi Yamagishi, Zhizheng Wu 0001, Oliver Watts, Simon King 0001 |
INTERSPEECH | 5 |
| 2015 | Towards minimum perceptual error training for DNN-based speech synthesisabstractWe propose to use a perceptually-oriented domain to improve the quality of text-to-speech generated by deep neural networks (DNNs). We train a DNN that predicts the parameters required for speech reconstruction but whose cost function is calculated in another domain. In this paper, to represent this perceptual domain we extract an approximated version of the Spectro-Temporal Excitation Pattern that was originally proposed as part of a model of hearing speech in noise. We train DNNs that pre-dict band aperiodicity, fundamental frequency and Mel cepstral coefficients and compare generated speech when the spectral cost function is defined in the Mel cepstral, warped log spec-trum or perceptual domains. Objective results indicate that the perceptual domain system achieves the highest quality. Index Terms: speech synthesis, DNN, objective measures, spectral analysis Cassia Valentini-Botinhao, Zhizheng Wu 0001, Simon King 0001 |
INTERSPEECH | 3 |
| 2015 | Sentence-level control vectors for deep neural network speech synthesisabstractThis paper describes the use of a low-dimensional vector representation of sentence acoustics to control the output of a feed-forward deep neural network text-to-speech system on a sentence-by-sentence basis. Vector representations for sentences in the training corpus are learned during network training along with other parameters of the model. Although the network is trained on a frame-by-frame basis, the standard frame-level inputs representing linguistic features are supplemented by features from a projection layer which outputs a learned representation of sentence-level acoustic characteristics. The projection layer contains dedicated parameters for each sentence in the training data which are optimised jointly with the standard network weights. Sentence-specific parameters are optimised on all frames of the relevant sentence -- these parameters therefore allow the network to account for sentence-level variation in the data which is not predictable from the standard linguistic inputs. Results show that the global prosodic characteristics of synthetic speech can be controlled simply and robustly at run time by supplementing basic linguistic features with sentence-level control vectors which are novel but designed to be consistent with those observed in the training corpus. Oliver Watts, Zhizheng Wu 0001, Simon King 0001 |
INTERSPEECH | 3 |
| 2015 | Minimum trajectory error training for deep neural networks, combined with stacked bottleneck featuresabstractRecently, Deep Neural Networks (DNNs) have shown promise as an acoustic model for statistical parametric speech synthesis. Their ability to learn complex mappings from linguistic features to acoustic features has advanced the naturalness of synthesis speech significantly. However, because DNN parameter estimation methods typically attempt to minimise the mean squared error of each individual frame in the training data, the dynamic and continuous nature of speech parameters is neglected. In this paper, we propose a training criterion that minimises speech parameter trajectory errors, and so takes dynamic constraints from a wide acoustic context into account during training. We combine this novel training criterion with our previously proposed stacked bottleneck features, which provide wide linguistic context. Both objective and subjective evaluation results confirm the effectiveness of the proposed training criterion for improving model accuracy and naturalness of synthesised speech. Zhizheng Wu 0001, Simon King 0001 |
INTERSPEECH | 2 |
| 2015 | A study of speaker adaptation for DNN-based speech synthesisabstractA major advantage of statistical parametric speech synthesis (SPSS) over unit-selection speech synthesis is its adaptability and controllability in changing speaker characteristics and speaking style.Recently, several studies using deep neural networks (DNNs) as acoustic models for SPSS have shown promising results.However, the adaptability of DNNs in SPSS has not been systematically studied.In this paper, we conduct an experimental analysis of speaker adaptation for DNN-based speech synthesis at different levels.In particular, we augment a low-dimensional speaker-specific vector with linguistic features as input to represent speaker identity, perform model adaptation to scale the hidden activation weights, and perform a feature space transformation at the output layer to modify generated acoustic features.We systematically analyse the performance of each individual adaptation technique and that of their combinations.Experimental results confirm the adaptability of the DNN, and listening tests demonstrate that the DNN can achieve significantly better adaptation performance than the hidden Markov model (HMM) baseline in terms of naturalness and speaker similarity. Zhizheng Wu 0001, Pawel Swietojanski, Christophe Veaux, Steve Renals, Simon King 0001 |
INTERSPEECH | 5 |
| 2015 | A perceptually-motivated low-complexity instantaneous linear channel normalization technique applied to speaker verification
Víctor Poblete, Felipe Espic, Simon King 0001, Richard M. Stern, Fernando Huenupán, Josué Fredes, Néstor Becerra Yoma |
Comput. Speech Lang. | 3 |
| 2014 | Multiple-average-voice-based speech synthesisabstractThis paper describes a novel approach for the speaker adaptation of statistical parametric speech synthesis systems based on the interpolation of a set of average voice models (AVM). Recent results have shown that the quality/naturalness of adapted voices depends on the distance from the average voice model used for speaker adaptation. This suggests the use of several AVMs trained on carefully chosen speaker clusters from which a more suitable AVM can be selected/interpolated during the adaptation. In the proposed approach a set of AVMs, a multiple-AVM, is trained on distinct clusters of speakers which are iteratively re-assigned during the estimation process initialised according to metadata. During adaptation, each AVM from the multiple-AVM is first adapted towards the target speaker. The adapted means from the AVMs are then interpolated to yield the final speaker adapted mean for synthesis. It is shown, performing speaker adaptation on a corpus of British speakers with various regional accents, that the quality/naturalness of synthetic speech of adapted voices is significantly higher than when considering a single factor-independent AVM selected according to the target speaker characteristics. Pierre Lanchantin, Mark J. F. Gales, Simon King 0001, Junichi Yamagishi |
ICASSP | 3 |
| 2014 | Neural net word representations for phrase-break prediction without a part of speech taggerabstractThe use of shared projection neural nets of the sort used in language modelling is proposed as a way of sharing parameters between multiple text-to-speech system components. We experiment with pretraining the weights of such a shared projection on an auxiliary language modelling task and then apply the resulting word representations to the task of phrase-break prediction. Doing so allows us to build phrase-break predictors that rival conventional systems without any reliance on conventional knowledge-based resources such as part of speech taggers. Oliver Watts, Siva Reddy Gangireddy, Junichi Yamagishi, Simon King 0001, Steve Renals, Adriana Cornelia Stan, Mircea Giurgiu |
ICASSP | 4 |
| 2014 | Investigating automatic & human filled pause insertion for speech synthesisabstractFilled pauses are pervasive in conversational speech and have been shown to serve several psychological and structural purposes.Despite this, they are seldom modelled overtly by stateof-the-art speech synthesis systems.This paper seeks to motivate the incorporation of filled pauses into speech synthesis systems by exploring their use in conversational speech, and by comparing the performance of several automatic systems inserting filled pauses into fluent text.Two initial experiments are described which seek to determine whether people's predicted insertion points are consistent with actual practice and/or with each other.The experiments also investigate whether there are 'right' and 'wrong' places to insert filled pauses.The results show good consistency between people's predictions of usage and their actual practice, as well as a perceptual preference for the 'right' placement.The third experiment contrasts the performance of several automatic systems that insert filled pauses into fluent sentences.The best performance (determined by F-score) was achieved through the by-word interpolation of probabilities predicted by Recurrent Neural Network and 4gram Language Models.The results offer insights into the use and perception of filled pauses by humans, and how automatic systems can be used to predict their locations. Rasmus Dall, Marcus Tomalin, Mirjam Wester, William J. Byrne, Simon King 0001 |
INTERSPEECH | 5 |
| 2014 | A comparison of open-source segmentation architectures for dealing with imperfect data from the media in speech synthesisabstractProceedings of: 15th Annual Conference of the International Speech Communication Association. Singapore, September 14-18, 2014. Ascensión Gallardo-Antolín, Juan Manuel Montero-Martínez, Simon King 0001 |
INTERSPEECH | 3 |
| 2014 | Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speechabstractAcoustic models used for statistical parametric speech synthe-sis typically incorporate many modelling assumptions. It is an open question to what extent these assumptions limit the natu-ralness of synthesised speech. To investigate this question, we recorded a speech corpus where each prompt was read aloud multiple times. By combining speech parameter trajectories ex-tracted from different repetitions, we were able to quantify the perceptual effects of certain commonly used modelling assump-tions. Subjective listening tests show that taking the source and filter parameters to be conditionally independent, or using di-agonal covariance matrices, significantly limits the naturalness that can be achieved. Our experimental results also demonstrate the shortcomings of mean-based parameter generation. Index terms: speech synthesis, acoustic modelling, stream in-dependence, diagonal covariance matrices, repeated speech 1. Gustav Eje Henter, Thomas Merritt, Matt Shannon, Catherine Mayo, Simon King 0001 |
INTERSPEECH | 5 |
| 2014 | Investigating source and filter contributions, and their interaction, to statistical parametric speech synthesisabstractThis paper presents an investigation of the separate perceptual degradations introduced by the modelling of source and fil-ter features in statistical parametric speech synthesis. This is achieved using stimuli in which various permutations of natu-ral, vocoded and modelled source and filter are combined, op-tionally with the addition of filter modifications (e.g. global variance or modulation spectrum scaling). We also examine the assumption of independence between source and filter pa-rameters. Two complementary perceptual testing paradigms are adopted. In the first, we ask listeners to perform “same or differ-ent quality ” judgements between pairs of stimuli from different configurations. In the second, we ask listeners to give an opin-ion score for individual stimuli. Combining the findings from these tests, we draw some conclusions regarding the relative contributions of source and filter to the currently rather limited naturalness of statistical parametric synthetic speech, and test whether current independence assumptions are justified. Index Terms: speech synthesis, hidden Markov modelling, GlottHMM, source filter model, source filter interaction Thomas Merritt, Tuomo Raitio, Simon King 0001 |
INTERSPEECH | 3 |
| 2014 | Unsupervised lexical clustering of speech segments using fixed-dimensional acoustic embeddingsabstractUnsupervised speech processing methods are essential for applications ranging from zero-resource speech technology to modelling child language acquisition. One challenging problem is discovering the word inventory of the language: the lexicon. Lexical clustering is the task of grouping unlabelled acoustic word tokens according to type. We propose a novel lexical clustering model: variable-length word segments are embedded in a fixed-dimensional acoustic space in which clustering is then performed. We evaluate several clustering algorithms and find that the best methods produce clusters with wide variation in sizes, as observed in natural language. The best probabilistic approach is an infinite Gaussian mixture model (IGMM), which automatically chooses the number of clusters. Performance is comparable to that of non-probabilistic Chinese Whispers and average-linkage hierarchical clustering. We conclude that IGMM clustering of fixed-dimensional embeddings holds promise as the lexical clustering component in unsupervised speech processing systems. Herman Kamper, Aren Jansen, Simon King 0001, Sharon Goldwater |
SLT | 3 |
| 2014 | The listening talker: A review of human and algorithmic context-induced modifications of speech
Martin Cooke, Simon King 0001, Maëva Garnier, Vincent Aubanel |
Comput. Speech Lang. | 2 |
| 2014 | Introduction to the Special Issue on The listening talker: context-dependent speech production and perception
Martin Cooke, Simon King 0001, W. Bastiaan Kleijn, Yannis Stylianou |
Comput. Speech Lang. | 2 |
| 2014 | Feature analysis for discriminative confidence estimation in spoken term detection
Javier Tejedor, Doroteo T. Toledano, Dong Wang 0013, Simon King 0001, José Colás Pasamontes |
Comput. Speech Lang. | 4 |
| 2014 | Intelligibility enhancement of HMM-generated speech in additive noise by modifying Mel cepstral coefficients to increase the glimpse proportion
Cassia Valentini-Botinhao, Junichi Yamagishi, Simon King 0001, Ranniery Maia |
Comput. Speech Lang. | 3 |
| 2014 | Statistical parametric speech synthesis for Ibibio
Moses Ekpenyong, Eno-Abasi Urua, Oliver Watts, Simon King 0001, Junichi Yamagishi |
Speech Commun. | 4 |
| 2013 | Factorized context modelling for Text-to-Speech synthesisabstractBecause speech units are so context-dependent, a large number of linguistic context features are generally used by HMM-based Text-to-Speech (TTS) speech synthesis systems, via context-dependent models. Since it is impossible to train separate models for every context, decision trees are used to discover the most important combinations of features that should be modelled. The task of the decision tree is very hard - to generalize from a very small observed part of the context feature space to the rest - and they have a major weakness: they cannot directly take advantage of factorial properties: they subdivide the model space based on one feature at a time. We propose a Dynamic Bayesian Network (DBN) based Mixed Memory Markov Model (MMMM) to provide factorization of the context space. The results of a listening test are provided as evidence that the model successfully learns the factorial nature of this space. Simon King 0001 |
ICASSP | 2 |
| 2013 | Lightly supervised GMM VAD to use audiobook for speech synthesiserabstractAudiobooks have been focused on as promising data for training Text-to-Speech (TTS) systems. However, they usually do not have a correspondence between audio and text data. Moreover, they are usually divided only into chapter units. In practice, we have to make a correspondence of audio and text data before we use them for building TTS synthesisers. However aligning audio and text data is time-consuming and involves manual labor. It also requires persons skilled in speech processing. Previously, we have proposed to use graphemes for automatically aligning speech and text data. This paper further integrates a lightly supervised voice activity detection (VAD) technique to detect sentence boundaries as a pre-processing step before the grapheme approach. This lightly supervised technique requires time stamps of speech and silence only for the first fifty sentences. Combining those, we can semi-automatically build TTS systems from audiobooks with minimum manual intervention. From subjective evaluations we analyse how the grapheme-based aligner and/or the proposed VAD technique impact the quality of HMM-based speech synthesisers trained on audiobooks. Yoshitaka Mamiya, Junichi Yamagishi, Oliver Watts, Robert A. J. Clark, Simon King 0001, Adriana Cornelia Stan |
ICASSP | 5 |
| 2013 | Where are the challenges in speaker diarization?abstractWe present a study on the contributions to Diarization Error Rate by the various components of speaker diarization system. Following on from an earlier study by Huijbregts and Wooters, we extend into more areas and draw somewhat different conclusions. From a series of experiments combining real, oracle and ideal system components, we are able to conclude that the primary cause of error in diarization is the training of speaker models on impure data, something that is in fact done in every current system. We conclude by suggesting ways to improve future systems, including a focus on training the speaker models from smaller quantities of pure data instead of all the data, as is currently done. Mark Sinclair, Simon King 0001 |
ICASSP | 2 |
| 2013 | Improving intelligibility in noise of HMM-generated speech via noise-dependent and -independent methodsabstractIn order to improve the intelligibility of HMM-generated Text-to-Speech (TTS) in noise, this work evaluates several speech enhancement methods, exploring combinations of noise-independent and - dependent approaches as well as algorithms previously developed for natural speech. We evaluate one noise-dependent method proposed for TTS, based on the glimpse proportion measure, and three approaches originally proposed for natural speech - one that estimates the noise and is based on the speech intelligibility index, and two noise-independent methods based on different spectral shaping techniques followed by dynamic range compression. We demonstrate how these methods influence the average spectra for different phone classes. We then present results of a listening experiment with speech-shaped noise and a competing speaker. A few methods made the TTS voice even more intelligible than the natural one. Although noise-dependent methods did not improve gains, the intelligibility differences found in distinct noises motivates such dependency. Cassia Valentini-Botinhao, Elizabeth Godoy, Yannis Stylianou, Bastian Sauert, Simon King 0001, Junichi Yamagishi |
ICASSP | 5 |
| 2013 | Reactive accent interpolation through an interactive map application
Maria Astrinaki, Junichi Yamagishi, Simon King 0001, Nicolas D'Alessandro, Thierry Dutoit |
INTERSPEECH | 3 |
| 2013 | Combining in-domain and out-of-domain speech data for automatic recognition of disordered speechabstractRecently there has been increasing interest in ways of using out-of-domain (OOD) data to improve automatic speech recognition performance in domains where only limited data is available. This paper focuses on one such domain, namely that of disordered speech for which only very small databases exist, but where normal speech can be considered OOD. Standard approaches for handling small data domains use adaptation from OOD models into the target domain, but here we investigate an alternative approach with its focus on the feature extraction stage: OOD data is used to train feature-generating deep belief neural networks. Using AMI meeting and TED talk datasets, we investigate various tandem-based speaker independent systems as well as maximum a posteriori adapted speaker dependent systems. Results on the UAspeech isolated word task of disordered speech are very promising with our overall best system (using a combination of AMI and TED data) giving a correctness of 62.5 an increase of 15% on previously best published results based on conventional model adaptation. We show that the relative benefit of using OOD data varies considerably from speaker to speaker and is only loosely correlated with the severity of a speaker's impairments. Heidi Christensen, Magda B. Aniol, Peter Bell 0001, Phil D. Green, Thomas Hain, Simon King 0001, Pawel Swietojanski |
INTERSPEECH | 6 |
| 2013 | The edinburgh speech production facility doubletalk corpus
James M. Scobbie, Alice Turk, Christian Geng, Simon King 0001, Robin J. Lickley, Korin Richmond |
INTERSPEECH | 4 |
| 2013 | Lightly supervised discriminative training of grapheme models for improved sentence-level alignment of speech and text dataabstractThis paper introduces a method for lightly supervised discriminative training using MMI to improve the alignment of speech and text data for use in training HMM-based TTS systems for low-resource languages. In TTS applications, due to the use of long-span contexts, it is important to select training utterances which have wholly correct transcriptions. In a low-resource setting, when using poorly trained grapheme models, we show that the use of MMI discriminative training at the grapheme-level enables us to increase the amount of correctly aligned data by 40 while maintaining a 7% sentence error rate and 0.8% word error rate. We present the procedure for lightly supervised discriminative training with regard to the objective of minimising sentence error rate. Adriana Cornelia Stan, Peter Bell 0001, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 4 |
| 2013 | TUNDRA: a multilingual corpus of found data for TTS research created with light supervisionabstractSimple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech re-search with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Future versions of the corpus will include finer-grained alignment and prosodic annotation, all of which will be made freely available. This paper gives a gen-eral outline of the data collected so far, as well as a detailed description of how this has been done, emphasizing the mini-mal language-specific knowledge and manual intervention used to compile the corpus. To demonstrate its potential use, text-to-speech systems have been built for all languages using unsu-pervised or lightly supervised methods, also briefly presented in the paper. Index Terms: multilingual corpus, light supervision, imperfect data, found data, text-to-speech, audiobook data Adriana Cornelia Stan, Oliver Watts, Yoshitaka Mamiya, Mircea Giurgiu, Robert A. J. Clark, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 7 |
| 2013 | Combining perceptually-motivated spectral shaping with loudness and duration modification for intelligibility enhancement of HMM-based synthetic speech in noiseabstractThis paper presents our entry to a speech-in-noise intelligibility enhancement evaluation: the Hurricane Challenge. The system consists of a Text-To-Speech voice manipulated through a combination of enhancement strategies, each of which is known to be individually successful: a perceptually-motivated spectral shaper based on the Glimpse Proportion measure, dynamic range compression, and adaptation to Lombard excitation and duration patterns. We achieved substantial intelligibility improvements relative to unmodified synthetic speech: 4.9 dB in competing speaker and 4.1 dB in speech-shaped noise. An analysis conducted across this and other two similar evaluations shows that the spectral shaper and the compressor (both of which are loudness boosters) contribute most under higher SNR conditions, particularly for speech-shaped noise. Duration and excitation Lombard-adapted changes are more beneficial in lower SNR conditions, and for competing speaker noise. Cassia Valentini-Botinhao, Junichi Yamagishi, Simon King 0001, Yannis Stylianou |
INTERSPEECH | 3 |
| 2013 | Personalising speech-to-speech translation: Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis
John Dines, Lakshmi Babu Saheer, Matthew Gibson, William J. Byrne, Keiichiro Oura, Keiichi Tokuda, Junichi Yamagishi, Simon King 0001, Mirjam Wester, Teemu Hirsimäki, Reima Karhila, Mikko Kurimo |
Comput. Speech Lang. | 9 |
| 2013 | Cross-Lingual Automatic Speech Recognition Using Tandem FeaturesabstractAutomatic speech recognition depends on large amounts of transcribed speech recordings in order to estimate the parameters of the acoustic model. Recording such large speech corpora is time-consuming and expensive; as a result, sufficient quantities of data exist only for a handful of languages-there are many more languages for which little or no data exist. Given that there are acoustic similarities between speech in different languages, it may be fruitful to use data from a well-resourced source language to estimate the acoustic models for a recognizer in a poorly-resourced target language. Previous approaches to this task have often involved making assumptions about shared phonetic inventories between the languages. Unfortunately pairs of languages do not generally share a common phonetic inventory. We propose an indirect way of transferring information from a source language acoustic model to a target language acoustic model without having to make any assumptions about the phonetic inventory overlap. To do this, we employ tandem features, in which class-posteriors from a separate classifier are decorrelated and appended to conventional acoustic features. Tandem features have the advantage that the language of the speech data used to train the classifier need not be the same as the target language to be recognized. This is because the class-posteriors are not used directly, so do not have to be over any particular set of classes. We demonstrate the use of tandem features in cross-lingual settings, including training on one or several source languages. We also examine factors which may predict a priori how much relative improvement will be brought about by using such tandem features, for a given source and target pair. In addition to conventional phoneme class-posteriors, we also investigate whether articulatory features (AFs)-a multi-stream, discrete, multi-valued labeling of speech-can be used instead. This is motivated by an assumption that AFs are less language-specific than a phoneme set. Partha Lal, Simon King 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2012 | Cepstral analysis based on the glimpse proportion measure for improving the intelligibility of HMM-based synthetic speech in noiseabstractIn this paper we introduce a new cepstral coefficient extraction method based on an intelligibility measure for speech in noise, the Glimpse Proportion measure. This new method aims to increase the intelligibility of speech in noise by modifying the clean speech, and has applications in scenarios such as public announcement and car navigation systems. We first explain how the Glimpse Proportion measure operates and further show how we approximated it to integrate it into an existing spectral envelope parameter extraction method commonly used in the HMM-based speech synthesis framework. We then demonstrate how this new method changes the modelled spectrum according to the characteristics of the noise and show results for a listening test with vocoded and HMM-based synthetic speech. The test indicates that the proposed method can significantly improve intelligibility of synthetic speech in speech shaped noise. Cassia Valentini-Botinhao, Ranniery Maia, Junichi Yamagishi, Simon King 0001, Heiga Zen |
ICASSP | 4 |
| 2012 | Analysis of speaker clustering strategies for HMM-based speech synthesisabstractThis paper describes a method for speaker clustering, with the application of building average voice models for speakeradaptive HMM-based speech synthesis that are a good basis for adapting to specific target speakers. Our main hypothesis is that using perceptually similar speakers to build the average voice model will be better than use unselected speakers, even if the amount of data available from perceptually similar speakers is smaller. We measure the perceived similarities among a group of 30 female speakers in a listening test and then apply multiple linear regression to automatically predict these listener judgements of speaker similarity and thus to identify similar speakers automatically. We then compare a variety of average voice models trained on either speakers who were perceptually judged to be similar to the target speaker, or speakers selected by the multiple linear regression, or a large global set of unselected speakers. We find that the average voice model trained on perceptually similar speakers provides better performance than the global model, even though the latter is trained on more data, confirming our main hypothesis. However, the average voice model using speakers selected automatically by the multiple linear regression does not reach the same level of performance. Index Terms: Statistical parametric speech synthesis, hidden Markov models, speaker adaptation Rasmus Dall, Christophe Veaux, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 4 |
| 2012 | Using Bayesian Networks to find relevant context features for HMM-based speech synthesisabstractSpeech units are highly context-dependent, so taking contextual features into account is essential for speech modelling. Context is employed in HMM-based Text-to-Speech speech synthesis systems via context-dependent phone models. A very wide context is taken into account, represented by a large set of contextual factors. However, most of these factors probably have no significant influence on the speech, most of the time. To discover which combinations of features should be taken into account, decision tree-based context clustering is used. But the space of context-dependent models is vast, and the number of contexts seen in the training data is only a tiny fraction of this space, so the task of the decision tree is very hard: to generalise from observations of a tiny fraction of the space to the rest of the space, whilst ignoring uninformative or redundant context features. The structure of the context feature space has not been systematically studied for speech synthesis. In this paper we discover a dependency structure by learning a Bayesian Network over the joint distribution of the features and the speech. We demonstrate that it is possible to discard the majority of context features with minimal impact on quality, measured by a perceptual test. Index Terms: HMM-based speech synthesis, Bayesian Networks, context information Simon King 0001 |
INTERSPEECH | 2 |
| 2012 | Detecting Acronyms from Capital Letter Sequences in SpanishabstractThis paper presents an automatic strategy to decide how to pronounce a Capital Letter Sequence (CLS) in a Text to Speech system (TTS). If CLS is well known by the TTS, it can be expanded in several words. But when the CLS is unknown, the system has two alternatives: spelling it (abbreviation) or pronouncing it as a new word (acronym). In Spanish, there is a high relationship between letters and phonemes. Because of this, when a CLS is similar to other words in Spanish, there is a high tendency to pronounce it as a standard word. This paper proposes an automatic method for detecting acronyms. Additionaly, this paper analyses the discrimination capability of some features, and several strategies for combining them in order to obtain the best classifier. For the best classifier, the classification error is 8.45%. About the feature analysis, the best features have been the Letter Sequence Perplexity and the Average N-gram order. Index Terms: Capital letter sequence pronunciation, Rubén San-Segundo-Hernández, Juan Manuel Montero-Martínez, Verónica López-Ludeña, Simon King 0001 |
INTERSPEECH | 4 |
| 2012 | Mel cepstral coefficient modification based on the Glimpse Proportion measure for improving the intelligibility of HMM-generated synthetic speech in noiseabstractWe propose a method that modifies the Mel cepstral coefficients of HMM-generated synthetic speech in order to increase the intelligibility of the generated speech when heard by a listener in the presence of a known noise. This method is based on an approximation we previously proposed for the Glimpse Proportion measure. Here we show how to update the Mel cepstral coefficients using this measure as an optimization criterion and how to control the amount of distortion by limiting the frequency resolution of the modifications. To evaluate the method we built eight different voices from normal read-text speech data from a male speaker. Some voices were also built from Lombard speech data produced by the same speaker. Listening experiments with speech-shaped noise and with a single competing talker indicate that our method significantly improves intelligibility when compared to unmodified synthetic speech. The voices built from Lombard speech outperformed the proposed method particularly for the competing talker case. However, compared to a voice using only the spectral parameters from Lombard speech, the proposed method obtains similar or higher performance. Cassia Valentini-Botinhao, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 3 |
| 2012 | Using HMM-based Speech Synthesis to Reconstruct the Voice of Individuals with Degenerative Speech Disorders
Christophe Veaux, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 3 |
| 2012 | A grapheme-based method for automatic alignment of speech and text dataabstractThis paper introduces a method for automatic alignment of speech data with unsynchronised, imperfect transcripts, for a domain where no initial acoustic models are available. Using grapheme-based acoustic models, word skip networks and orthographic speech transcripts, we are able to harvest 55% of the speech with a 93% utterance-level accuracy and 99% word accuracy for the produced transcriptions. The work is based on the assumption that there is a high degree of correspondence between the speech and text, and that a full transcription of all of the speech is not required. The method is language independent and the only prior knowledge and resources required are the speech and text transcripts, and a few minor user interventions. Adriana Cornelia Stan, Peter Bell 0001, Simon King 0001 |
SLT | 3 |
| 2012 | Term-Dependent Confidence Normalisation for Out-of-Vocabulary Spoken Term Detection
Dong Wang 0013, Javier Tejedor, Simon King 0001, Joe Frankel |
J. Comput. Sci. Technol. | 3 |
| 2012 | Impacts of machine translation and speech synthesis on speech-to-speech translation
Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda |
Speech Commun. | 4 |
| 2012 | Analysis of unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis using KLD-based transform mapping
Keiichiro Oura, Junichi Yamagishi, Mirjam Wester, Simon King 0001, Keiichi Tokuda |
Speech Commun. | 4 |
| 2012 | Direct posterior confidence for out-of-vocabulary spoken term detectionabstractSpoken term detection (STD) is a key technology for spoken information retrieval. As compared to the conventional speech transcription and keyword spotting, STD is an open-vocabulary task and has to address out-of-vocabulary (OOV) terms. Approaches based on subword units, for example phones, are widely used to solve the OOV issue; however, performance on OOV terms is still substantially inferior to that of in-vocabulary (INV) terms. The performance degradation on OOV terms can be attributed to a multitude of factors. One particular factor we address in this article is the unreliable confidence estimation caused by weak acoustic and language modeling due to the absence of OOV terms in the training corpora. We propose a direct posterior confidence derived from a discriminative model, such as multilayer perceptron (MLP). The new confidence considers a wide-range acoustic context which is usually important for speech recognition and retrieval; moreover, it localizes on detected speech segments and therefore avoids the impact of long-span word context which is usually unreliable for OOV term detection. In this article, we first develop an extensive discussion about the modeling weakness problem associated with OOV terms, and then propose our approach to address this problem based on direct poster confidence. Our experiments carried out on spontaneous and conversational multiparty meeting speech, demonstrate that the proposed technique provides a significant improvement in STD performance as compared to conventional lattice-based confidence, in particular for OOV terms. Furthermore, the new confidence estimation approach is fused with other advanced techniques for OOV treatment, such as stochastic pronunciation modeling and discriminative confidence normalization. This leads to an integrated solution for OOV term detection that results in a large performance improvement. Dong Wang 0013, Simon King 0001, Joe Frankel, Ravichander Vipperla, Nicholas W. D. Evans, Raphaël Troncy |
ACM Trans. Inf. Syst. | 2 |
| 2011 | Voice banking and voice reconstruction for MND patientsabstractWhen the speech of an individual becomes unintelligible due to a degenerative disease such as motor neuron disease (MND), a voice output communication aid (VOCA) can be used. To fully replace all functions of speech communication: communication of information, maintenance of social relationships and displaying identity, the voice must be intelligible, natural-sounding and retain the vocal identity of the speaker. Attempts have been made to capture the voice before it is lost, using a process known as voice banking. But, for patients with MND, the speech deterioration frequently coincides or quickly follows diagnosis. Using model-based speech synthesis, it is now possible to retain the vocal identity of the patient with minimal data recordings and even deteriorating speech. The power of this approach is that it is possible to use the patient's recordings to adapt existing voice models pre-trained on many speakers. When the speech has begun to deteriorate, the adapted voice model can be further modified in order to compensate for the disordered characteristics found in the patient's speech. We present here an on-going project for voice banking and voice reconstruction based on this technology. Christophe Veaux, Junichi Yamagishi, Simon King 0001 |
ASSETS | 3 |
| 2011 | Vocal attractiveness of statistical speech synthesisersabstractOur previous analysis of speaker-adaptive HMM-based speech synthesis methods suggested that there are two possible reasons why average voices can obtain higher subjective scores than any individual adapted voice: 1) model adaptation degrades speech quality proportionally to the distance 'moved' by the transforms, and 2) psychoacoustic effects relating to the attractiveness of the voice. This paper is a follow-on from that analysis and aims to separate these effects out. Our latest perceptual experiments focus on attractiveness, using average voices and speaker-dependent voices without model trans formation, and show that using several speakers to create a voice improves smoothness (measured by Harmonics-to-Noise Ratio), reduces distance from the the average voice in the log F0-F1 space of the final voice and hence makes it more attractive at the segmental level. However, this is weakened or overridden at supra-segmental or sentence levels. Sandra Andraszewicz, Junichi Yamagishi, Simon King 0001 |
ICASSP | 3 |
| 2011 | An analysis of machine translation and speech synthesis in speech-to-speech translation systemabstractThis paper provides an analysis of the impacts of machine translation and speech synthesis on speech-to-speech translation systems. The speech-to-speech translation system consists of three components: speech recognition, machine translation and speech synthesis. Many techniques for integration of speech recognition and machine translation have been proposed. However, speech synthesis has not yet been considered. Therefore, in this paper, we focus on machine translation and speech synthesis, and report a subjective evaluation to analyze the impact of each component. The results of these analyses show that the naturalness and intelligibility of synthesized speech are strongly affected by the fluency of the translated sentences. Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda |
ICASSP | 4 |
| 2011 | Evaluation of objective measures for intelligibility prediction of HMM-based synthetic speech in noiseabstractIn this paper we evaluate four objective measures of speech with regards to intelligibility prediction of synthesized speech in diverse noisy situations. We evaluated three intelligibility measures, the Dau measure, the glimpse proportion and the Speech Intelligibility Index (SII) and a quality measure, the Perceptual Evaluation of Speech Quality (PESQ). For the generation of synthesized speech we used a state of the art HMM-based speech synthesis system. The noisy conditions comprised four additive noises. The measures were compared with subjective intelligibility scores obtained in listening tests. The results show the Dau and the glimpse measures to be the best predictors of intelligibility, with correlations of around 0.83 to subjective scores. All measures gave less accurate predictions of intelligibility for synthetic speech than have previously been found for natural speech; in particular the SII measure. In additional experiments, we processed the synthesized speech by an ideal binary mask before adding noise. The Glimpse measure gave the most accurate intelligibility predictions in this situation. Cassia Valentini-Botinhao, Junichi Yamagishi, Simon King 0001 |
ICASSP | 3 |
| 2011 | Handling overlaps in spoken term detectionabstractSpoken term detection (STD) systems usually arrive at many overlapping detections which are often addressed with some pragmatic approaches, e.g. choosing the best detection to represent all the overlaps. In this paper we present a theoretical study based on a concept of acceptance space. In particular, we present two confidence estimation approaches based on Bayesian and evidence perspectives respectively. Analysis shows that both approaches possess respective ad vantages and shortcomings, and that their combination has the potential to provide an improved confidence estimation. Experiments conducted on meeting data confirm our analysis and show considerable performance improvement with the combined approach, in particular for out-of-vocabulary spoken term detection with stochastic pronunciation modeling. Dong Wang 0013, Nicholas W. D. Evans, Raphaël Troncy, Simon King 0001 |
ICASSP | 4 |
| 2011 | Formant-Controlled HMM-Based Speech SynthesisabstractThis paper proposes a novel framework that enables us to manipulate and control formants in HMM-based speech synthesis. In this framework, the dependency between formants and spectral features is modelled by piecewise linear transforms; formant parameters are effectively mapped by these to the means of Gaussian distributions over the spectral synthesis parameters. The spectral envelope features generated under the influence of formants in this way may then be passed to high-quality vocoders to generate the speech waveform. This provides two major advantages over conventional frameworks. First, we can achieve spectral modification by changing formants only in those parts where we want control, whereas the user must specify all formants manually in conventional formant synthesisers (e.g. Klatt). Second, this can produce high-quality speech. Our results show the proposed method can control vowels in the synthesized speech by manipulating F 1 and F 2 without any degradation in synthesis quality. Junichi Yamagishi, Korin Richmond, Zhen-Hua Ling, Simon King 0001, Li-Rong Dai 0001 |
INTERSPEECH | 5 |
| 2011 | Announcing the Electromagnetic Articulography (Day 1) Subset of the mngu0 Articulatory CorpusabstractThis paper serves as an initial announcement of the availability of a corpus of articulatory data called mngu0. This corpus will ultimately consist of a collection of multiple sources of articulatory data acquired from a single speaker: electromagnetic articulography (EMA), audio, video, volumetric MRI scans, and 3D scans of dental impressions. This data will be provided free for research use. In this first stage of the release, we are making available one subset of EMA data, consisting of more than 1,300 phonetically diverse utterances recorded with a Carstens AG500 electromagnetic articulograph. Distribution of mngu0 will be managed by a dedicated ``forum-style'' web site. This paper both outlines the general goals motivating the distribution of the data and the creation of the mngu0 web forum, and also provides a description of the EMA data contained in this initial release. Korin Richmond, Phil Hoole, Simon King 0001 |
INTERSPEECH | 3 |
| 2011 | Can Objective Measures Predict the Intelligibility of Modified HMM-Based Synthetic Speech in Noise?abstractSynthetic speech can be modified to improve intelligibility in noise. In order to perform modifications automatically, it would be useful to have an objective measure that could predict the intelligibility of modified synthetic speech for human listeners. We analysed the impact on intelligibility – and on how well objective measures predict it – when we separately modify speaking rate, fundamental frequency, line spectral pairs and spectral peaks. Shifting LSPs can increase intelligibility for human listeners; other modifications had weaker effects. Among the objective measures we evaluated, the Dau model and the Glimpse proportion were the best predictors of human performance. Index Terms: objective measures for speech intelligibility, HMM-based speech synthesis, Lombard speech 1. Cassia Valentini-Botinhao, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 3 |
| 2011 | Unsupervised Continuous-Valued Word Features for Phrase-Break Prediction without a Part-of-Speech TaggerabstractPart of speech (POS) tags are foremost among the features conventionally used to predict intonational phrase-breaks for text to speech (TTS) conversion. The construction of such systems therefore presupposes the availability of a POS tagger for the relevant language, or of a corpus manually tagged with POS. However, such tools and resources are not available in the majority of the world’s languages, and manually labelling text with POS tags is an expensive and time-consuming process. We therefore propose the use of continuous-valued features that summarise the distributional characteristics of word types as surrogates for POS features. Importantly, such features are obtained in an unsupervised manner from an untagged text corpus. We present results on the phrase-break prediction task, where use of the features closes the gap in performance between a baseline system (using only basic punctuation-related features) and a topline system (incorporating a state-of-the-art POS tagger). Oliver Watts, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 3 |
| 2011 | Listeners' weighting of acoustic cues to synthetic speech naturalness: A multidimensional scaling analysis
Catherine Mayo, Robert A. J. Clark, Simon King 0001 |
Speech Commun. | 3 |
| 2011 | The Romanian speech synthesis (RSS) corpus: Building a high quality HMM-based speech synthesis system using a high sampling rate
Adriana Cornelia Stan, Junichi Yamagishi, Simon King 0001, Matthew P. Aylett |
Speech Commun. | 3 |
| 2011 | Letter-to-Sound Pronunciation Prediction Using Conditional Random FieldsabstractPronunciation prediction, or letter-to-sound (LTS) conversion, is an essential task for speech synthesis, open vocabulary spoken term detection and other applications dealing with novel words. Most current approaches (at least for English) employ data-driven methods to learn and represent pronunciation “rules” using statistical models such as decision trees, hidden Markov models (HMMs) or joint-multigram models (JMMs). The LTS task remains challenging, particularly for languages with a complex relationship between spelling and pronunciation such as English. In this paper, we propose to use a conditional random field (CRF) to perform LTS because it avoids having to model a distribution over observations and can perform global inference, suggesting that it may be more suitable for LTS than decision trees, HMMs or JMMs. One challenge in applying CRFs to LTS is that the phoneme and grapheme sequences of a word are generally of different lengths, which makes CRF training difficult. To solve this problem, we employed a joint-multigram model to generate aligned training exemplars. Experiments conducted with the AMI05 dictionary demonstrate that a CRF significantly outperforms other models, especially if n-best lists of predictions are generated. Dong Wang 0013, Simon King 0001 |
IEEE Signal Process. Lett. | 2 |
| 2011 | Stochastic Pronunciation Modeling for Out-of-Vocabulary Spoken Term DetectionabstractSpoken term detection (STD) is the name given to the task of searching large amounts of audio for occurrences of spoken terms, which are typically single words or short phrases. One reason that STD is a hard task is that search terms tend to contain a disproportionate number of out-of-vocabulary (OOV) words. The most common approach to STD uses subword units. This, in conjunction with some method for predicting pronunciations of OOVs from their written form, enables the detection of OOV terms but performance is considerably worse than for in-vocabulary terms. This performance differential can be largely attributed to the special properties of OOVs. One such property is the high degree of uncertainty in the pronunciation of OOVs. We present a stochastic pronunciation model (SPM) which explicitly deals with this uncertainty. The key insight is to search for all possible pronunciations when detecting an OOV term, explicitly capturing the uncertainty in pronunciation. This requires a probabilistic model of pronunciation, able to estimate a distribution over all possible pronunciations. We use a joint-multigram model (JMM) for this and compare the JMM-based SPM with the conventional soft match approach. Experiments using speech from the meetings domain demonstrate that the SPM performs better than soft match in most operating regions, especially at low false alarm probabilities. Furthermore, SPM and soft match are found to be complementary: their combination provides further performance gains. Dong Wang 0013, Simon King 0001, Joe Frankel |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesisabstractIn the EMIME project, we are developing a mobile device that performs personalized speech-to-speech translation such that a user's spoken input in one language is used to produce spoken output in another language, while continuing to sound like the user's voice. We integrate two techniques, unsupervised adaptation for HMM-based TTS using a word-based large-vocabulary continuous speech recognizer and cross-lingual speaker adaptation for HMM-based TTS, into a single architecture. Thus, an unsupervised cross-lingual speaker adaptation system can be developed. Listening tests show very promising results, demonstrating that adapted voices sound similar to the target speaker and that differences between supervised and unsupervised cross-lingual speaker adaptation are small. Keiichiro Oura, Keiichi Tokuda, Junichi Yamagishi, Simon King 0001, Mirjam Wester |
ICASSP | 4 |
| 2010 | Stochastic pronunciation modelling and soft match for out-of-vocabulary spoken term detectionabstractA major challenge faced by a spoken term detection (STD) system is the detection of out-of-vocabulary (OOV) terms. Although a subword-based STD system is able to detect OOV terms, performance reduction is always observed compared to in-vocabulary terms. One challenge that OOV terms bring to STD is the pronunciation uncertainty. A commonly used approach to address this problem is a soft matching procedure, and the other is the stochastic pronunciation modelling (SPM) proposed by the authors. In this paper we compare these two approaches, and combine them using a discriminative decision strategy. Experimental results demonstrated that SPM and soft match are highly complementary, and their combination gives significant performance improvement to OOV term detection. Dong Wang 0013, Simon King 0001, Joe Frankel, Peter Bell 0001 |
ICASSP | 2 |
| 2010 | Simple methods for improving speaker-similarity of HMM-based speech synthesisabstractIn this paper we revisit some basic configuration choices of HMM-based speech synthesis, such as waveform sampling rate, auditory frequency warping scale and the logarithmic scaling of F0, with the aim of improving speaker similarity which is an acknowledged weakness of current HMM-based speech synthesisers. All of the techniques investigated are simple but, as we demonstrate using perceptual tests, can make substantial differences to the quality of the synthetic speech. Contrary to common practice in automatic speech recognition, higher waveform sampling rates can offer enhanced feature extraction and improved speaker similarity for speech synthesis. In addition, a generalized logarithmic transform of F0results in larger intra-utterance variance of F0trajectories and hence more dynamic and natural-sounding prosody. Junichi Yamagishi, Simon King 0001 |
ICASSP | 2 |
| 2010 | A classifier-based target cost for unit selection speech synthesis trained on perceptual dataabstractOur goal is to automatically learn a perceptually-optimal target cost function for a unit selection speech synthesiser.The approach we take here is to train a classifier on human perceptual judgements of synthetic speech.The output of the classifier is used to make a simple three-way distinction rather than to estimate a continuously-valued cost.In order to collect the necessary perceptual data, we synthesised 145,137 short sentences with the usual target cost switched off, so that the search was driven by the join cost only.We then selected the 7200 sentences with the best joins and asked 60 listeners to judge them, providing their ratings for each syllable.From this, we derived a rating for each demiphone.Using as input the same context features employed in our conventional target cost function, we trained a classifier on these human perceptual ratings.We synthesised two sets of test sentences with both our standard target cost and the new target cost based on the classifier.A/B preference tests showed that the classifier-based target cost, which was learned completely automatically from modest amounts of perceptual data, is almost as good as our carefullyand expertly-tuned standard target cost. Volker Strom, Simon King 0001 |
INTERSPEECH | 2 |
| 2010 | Augmented set of features for confidence estimation in spoken term detectionabstractDiscriminative confidence estimation along with confidence normalisation have been shown to construct robust decision maker modules in spoken term detection (STD) systems. Discriminative confidence estimation, making use of termdependent features, has been shown to improve the widely used lattice-based confidence estimation in STD. In this work, we augment the set of these term-dependent features and show a significant improvement in the STD performance both in terms of ATWV and DET curves in experiments conducted on a Spanish geographical corpus. This work also proposes a multiple linear regression analysis to carry out the feature selection. Next, the most informative features derived from it are used within the discriminative confidence on the STD system. Javier Tejedor, Doroteo T. Toledano, Miguel Bautista, Simon King 0001, Dong Wang 0013, José Colás Pasamontes |
INTERSPEECH | 4 |
| 2010 | CRF-based stochastic pronunciation modeling for out-of-vocabulary spoken term detectionabstractOut-of-vocabulary (OOV) terms present a significant challenge to spoken term detection (STD). This challenge, to a large ex-tent, lies in the high degree of uncertainty in pronunciations of OOV terms. In previous work, we presented a stochastic pro-nunciation modeling (SPM) approach to compensate for this uncertainty. A shortcoming of our original work, however, is that the SPM was based on a joint-multigram model (JMM), which is suboptimal. In this paper, we propose to use con-ditional random fields (CRFs) for letter-to-sound conversion, which significantly improves quality of the predicted pronun-ciations. When applied to OOV STD, we achieve consider-able performance improvement with both a 1-best system and an SPM-based system. Index Terms: speech recognition, spoken term detection, con-ditional random field, joint multigram model Dong Wang 0013, Simon King 0001, Nicholas W. D. Evans, Raphaël Troncy |
INTERSPEECH | 2 |
| 2010 | The role of higher-level linguistic features in HMM-based speech synthesisabstractWe analyse the contribution of higher-level elements of the linguistic specification of a data-driven speech synthesiser to the naturalness of the synthetic speech which it generates. The system is trained using various subsets of the full feature-set, in which features relating to syntactic category, intonational phrase boundary, pitch accent and boundary tones are selectively removed. Utterances synthesised by the different configurations of the system are then compared in a subjective evaluation of their naturalness. The work presented forms background analysis for an ongoing set of experiments in performing text-to-speech (TTS) conversion based on shallow features: features that can be trivially extracted from text. By building a range of systems, each assuming the availability of a different level of linguistic annotation, we obtain benchmarks for our on-going work. Oliver Watts, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 3 |
| 2010 | Roles of the average voice in speaker-adaptive HMM-based speech synthesisabstractIn speaker-adaptive HMM-based speech synthesis, there are a few speakers whose synthetic speech sounds worse than that of other speakers, despite having the same amount of adapta-tion data from within the same corpus. This paper investigates these fluctuations in quality and found that as mel-cepstral dis-tance from the average voice becomes larger, the MOS scores generally become worse. Although the negative correlation ob-tained is not strong enough, this helps us improve the training and adaptation strategies for average voice models. Further-more we remark that this correlation is strongly linked to “vocal attractiveness.” Index Terms: speech synthesis, HMM, average voice, speaker adaptation Junichi Yamagishi, Oliver Watts, Simon King 0001, Bela Usabaev |
INTERSPEECH | 3 |
| 2010 | Analysis of statistical parametric and unit selection speech synthesis systems applied to emotional speech
Roberto Barra-Chicote, Junichi Yamagishi, Simon King 0001, Juan Manuel Montero-Martínez, Javier Macías Guarasa |
Speech Commun. | 3 |
| 2010 | Synthesis of Child Speech With HMM Adaptation and Voice ConversionabstractThe synthesis of child speech presents challenges both in the collection of data and in the building of a synthesizer from that data. We chose to build a statistical parametric synthesizer using the hidden Markov model (HMM)-based system HTS, as this technique has previously been shown to perform well for limited amounts of data, and for data collected under imperfect conditions. Six different configurations of the synthesizer were compared, using both speaker-dependent and speaker-adaptive modeling techniques, and using varying amounts of data. For comparison with HMM adaptation, techniques from voice conversion were used to transform existing synthesizers to the characteristics of the target speaker. Speaker-adaptive voices generally outperformed child speaker-dependent voices in the evaluation. HMM adaptation outperformed voice conversion style techniques when using the full target speaker corpus; with fewer adaptation data, however, no significant listener preference for either HMM adaptation or voice conversion methods was found. Oliver Watts, Junichi Yamagishi, Simon King 0001, Kay M. Berkling |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Thousands of Voices for HMM-Based Speech Synthesis-Analysis and Application of TTS Systems Built on Various ASR CorporaabstractIn conventional speech synthesis, large amounts of phonetically balanced speech data recorded in highly controlled recording studio environments are typically required to build a voice. Although using such data is a straightforward solution for high quality synthesis, the number of voices available will always be limited, because recording costs are high. On the other hand, our recent experiments with HMM-based speech synthesis systems have demonstrated that speaker-adaptive HMM-based speech synthesis (which uses an “average voice model” plus model adaptation) is robust to non-ideal speech data that are recorded under various conditions and with varying microphones, that are not perfectly clean, and/or that lack phonetic balance. This enables us to consider building high-quality voices on “non-TTS” corpora such as ASR corpora. Since ASR corpora generally include a large number of speakers, this leads to the possibility of producing an enormous number of voices automatically. In this paper, we demonstrate the thousands of voices for HMM-based speech synthesis that we have made from several popular ASR corpora such as the Wall Street Journal (WSJ0, WSJ1, and WSJCAM0), Resource Management, Globalphone, and SPEECON databases. We also present the results of associated analysis based on perceptual evaluation, and discuss remaining issues. Junichi Yamagishi, Bela Usabaev, Simon King 0001, Oliver Watts, John Dines, Jilei Tian, Rile Hu, Keiichiro Oura, Yi-Jian Wu, Keiichi Tokuda, Reima Karhila, Mikko Kurimo |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Diagonal priors for full covariance speech recognitionabstractWe investigate the use of full covariance Gaussians for large-vocabulary speech recognition. The large number of parameters gives high modelling power, but when training data is limited, the standard sample covariance matrix is often poorly conditioned, and has high variance. We explain how these problems may be solved by the use of a diagonal covariance smoothing prior, and relate this to the shrinkage estimator, for which the optimal shrinkage parameter may itself be estimated from the training data. We also compare the use of generatively and discriminatively trained priors. Results are presented on a large vocabulary conversational telephone speech recognition task. Peter Bell 0001, Simon King 0001 |
ASRU | 2 |
| 2009 | Posterior-based confidence measures for spoken term detectionabstractConfidence measures play a key role in spoken term detection (STD) tasks. The confidence measure expresses the posterior probability of the search term appearing in the detection period, given the speech. Traditional approaches are based on the acoustic and language model scores for candidate detections found using automatic speech recognition, with Bayes' rule being used to compute the desired posterior probability. In this paper, we present a novel direct posterior-based confidence measure which, instead of resorting to the Bayesian formula, calculates posterior probabilities from a multi-layer perceptron (MLP) directly. Compared with traditional Bayesian-based methods, the direct-posterior approach is conceptually and mathematically simpler. Moreover, the MLP-based model does not require assumptions to be made about the acoustic features such as their statistical distribution and the independence of static and dynamic co-efficients. Our experimental results in both English and Spanish demonstrate that the proposed direct posterior-based confidence improves STD performance. Dong Wang 0013, Javier Tejedor, Joe Frankel, Simon King 0001, José Colás Pasamontes |
ICASSP | 4 |
| 2009 | Speech synthesis without a phone inventoryabstractIn speech synthesis the unit inventory is decided using phonological and phonetic expertise. This process is resource intensive and potentially sub-optimal. In this paper we investigate how acoustic clustering, together with lexicon constraints, can be used to build a self-organised inventory. Six English speech synthesis systems were built using two frameworks, unit selection and parametric HTS for three inventory conditions: 1) a traditional phone set, 2) a system using orthographic units, and 3) a self-organised inventory. A listening test showed a strong preference for the classic system, and for the orthographic system over the self-organised system. Results also varied by letter to sound complexity and database coverage. This suggests the self-organised approach failed to generalise pronunciation as well as introducing noise above and beyond that caused by orthographic sound mismatch. Index Terms: speech synthesis, unit selection, parametric synthesis, phone inventory, orthographic synthesis Matthew P. Aylett, Simon King 0001, Junichi Yamagishi |
INTERSPEECH | 2 |
| 2009 | Measuring the gap between HMM-based ASR and TTSabstractThe EMIME European project is conducting research in the development of technologies for mobile, personalised speech-tospeech translation systems. The hidden Markov model is being used as the underlying technology in both automatic speech recognition (ASR) and text-to-speech synthesis (TTS) components, thus, the investigation of unified statistical modelling approaches has become an implicit goal of our research. As one of the first steps towards this goal, we have been investigating commonalities and differences between HMM-based ASR and TTS. In this paper we present results and analysis of a series of experiments that have been conducted on English ASR and TTS systems measuring their performance with respect to phone set and lexicon, acoustic feature type and dimensionality andHMM topology. Our results show that, although the fundamental statistical model may be essentially the same, optimal ASR and TTS performance often demands diametrically opposed system designs. This represents a major challenge to be addressed in the investigation of such unified modelling approaches. John Dines, Junichi Yamagishi, Simon King 0001 |
INTERSPEECH | 3 |
| 2009 | A posterior probability-based system hybridisation and combination for spoken term detectionabstractSpoken term detection (STD) is a fundamental task for multimedia information retrieval. To improve the detection performance, we have presented a direct posterior-based confidence measure generated from a neural network. In this paper, we propose a detection-independent confidence estimation based on the direct posterior confidence measure, in which the decision making is totally separated from the term detection. Based on this idea, we first present a hybrid system which conducts the term detection and confidence estimation based on different sub-word units and then propose a combination method which merges detections from heterogeneous term detectors based on the direct posterior-based confidence. Experimental results demonstrated that the proposed methods improved system performance considerably for both English and Spanish. Index Terms: speech recognition, spoken term detection, confidence estimation, grapheme Javier Tejedor, Dong Wang 0013, Simon King 0001, Joe Frankel, José Colás Pasamontes |
INTERSPEECH | 3 |
| 2009 | Stochastic pronunciation modelling for spoken term detectionabstractA major challenge faced by a spoken term detection (STD) system is the detection of out-of-vocabulary (OOV) terms. Although a subword-based STD system is able to detect OOV terms, performance reduction is always observed compared to in-vocabulary terms. Current approaches to STD do not acknowledge the particular properties of OOV terms, such as pronunciation uncertainty. In this paper, we use a stochastic pronunciation model to deal with the uncertain pronunciations of OOV terms. By considering all possible term pronunciations, predicted by a joint-multigram model, we observe a significant performance improvement. Index Terms: joint-multigram, pronunciation model, spoken term detection, speech recognition Dong Wang 0013, Simon King 0001, Joe Frankel |
INTERSPEECH | 2 |
| 2009 | Term-dependent confidence for out-of-vocabulary term detectionabstractWithin a spoken term detection (STD) system, the decision maker plays an important role in retrieving reliable detections. Most of the state-of-the-art STD systems make decisions based on a confidence measure that is term-independent, which poses a serious problem for out-of-vocabulary (OOV) term detection. In this paper, we study a term-dependent confidence measure based on confidence normalisation and discriminative modelling, particularly focusing on its remarkable effectiveness for detecting OOV terms. Experimental results indicate that the term-dependent confidence provides much more significant improvement for OOV terms than terms in-vocabulary. Index Terms: confidence estimation, spoken term detection, speech recognition Dong Wang 0013, Simon King 0001, Joe Frankel, Peter Bell 0001 |
INTERSPEECH | 2 |
| 2009 | HMM adaptation and voice conversion for the synthesis of child speech: a comparisonabstractThis study compares two different methodologies for producing data-driven synthesis of child speech from existing systems that have been trained on the speech of adults. On one hand, an existing statistical parametric synthesiser is transformed using model adaptation techniques, informed by linguistic and prosodic knowledge, to the speaker characteristics of a child speaker. This is compared with the application of voice conversion techniques to convert the output of an existing waveform concatenation synthesiser with no explicit linguistic or prosodic knowledge. In a subjective evaluation of the similarity of synthetic speech to natural speech from the target speaker, the HMM-based systems evaluated are generally preferred, although this is at least in part due to the higher dimensional acoustic features supported by these techniques. Index Terms: child speech, statistical parametric speech synthesis, HMM-based speech synthesis, voice conversion, HTS, Average Voice Models, Festival Oliver Watts, Junichi Yamagishi, Simon King 0001, Kay M. Berkling |
INTERSPEECH | 3 |
| 2009 | Thousands of voices for HMM-based speech synthesisabstractOur recent experiments with HMM-based speech synthesis systems have demonstrated that speaker-adaptive HMM-based speech synthesis (which uses an 'average voice model' plus model adaptation) is robust to non-ideal speech data that are recorded under various conditions and with varying microphones, that are not perfectly clean, and/or that lack of phonetic balance. This enables us consider building high-quality voices on 'non-TTS' corpora such as ASR corpora. Since ASR corpora generally include a large number of speakers, this leads to the possibility of producing an enormous number of voices automatically. In this paper we show thousands of voices for HMM-based speech synthesis that we have made from several popular ASR corpora such as the Wall Street Journal databases (WSJ0/WSJ1/WSJCAM0), Resource Management, Globalphone and Speecon. We report some perceptual evaluation results and outline the outstanding issues. Junichi Yamagishi, Bela Usabaev, Simon King 0001, Oliver Watts, John Dines, Jilei Tian, Rile Hu, Keiichiro Oura, Keiichi Tokuda, Reima Karhila, Mikko Kurimo |
INTERSPEECH | 3 |
| 2009 | Robust Speaker-Adaptive HMM-Based Text-to-Speech SynthesisabstractThis paper describes a speaker-adaptive HMM-based speech synthesis system. The new system, called ldquoHTS-2007,rdquo employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available. In addition, a comparison study with several speech synthesis techniques shows the new system is very robust: It is able to build voices from less-than-ideal speech data and synthesize good-quality speech even for out-of-domain sentences. Junichi Yamagishi, Takashi Nose, Heiga Zen, Zhen-Hua Ling, Tomoki Toda, Keiichi Tokuda, Simon King 0001, Steve Renals |
IEEE Trans. Speech Audio Process. | 7 |
| 2008 | A comparison of phone and grapheme-based spoken term detectionabstractWe propose grapheme-based sub-word units for spoken term detection (STD). Compared to phones, graphemes have a number of potential advantages. For out-of-vocabulary search terms, phone- based approaches must generate a pronunciation using letter-to-sound rules. Using graphemes obviates this potentially error-prone hard decision, shifting pronunciation modelling into the statistical models describing the observation space. In addition, long-span grapheme language models can be trained directly from large text corpora. We present experiments on Spanish and English data, comparing phone and grapheme-based STD. For Spanish, where phone and grapheme-based systems give similar transcription word error rates (WERs), grapheme-based STD significantly outperforms a phone- based approach. The converse is found for English, where the phone- based system outperforms a grapheme approach. However, we present additional analysis which suggests that phone-based STD performance levels may be achieved by a grapheme-based approach despite lower transcription accuracy, and that the two approaches may usefully be combined. We propose a number of directions for future development of these ideas, and suggest that if grapheme-based STD can match phone-based performance, the inherent flexibility in dealing with out-of-vocabulary terms makes this a desirable approach. Dong Wang 0013, Joe Frankel, Javier Tejedor, Simon King 0001 |
ICASSP | 4 |
| 2008 | A shrinkage estimator for speech recognition with full covariance HMMsabstractWe consider the problem of parameter estimation in full-covariance Gaussian mixture systems for automatic speech recognition. Due to the high dimensionality of the acoustic feature vector, the standard sample covariance matrix has a high variance and is often poorly-conditioned when the amount of training data is limited. We explain how the use of a shrinkage estimator can solve these problems, and derive a formula for the optimal shrinkage intensity. We present results of experiments on a phone recognition task, showing that the estimator gives a performance improvement over a standard full-covariance system Peter Bell 0001, Simon King 0001 |
INTERSPEECH | 2 |
| 2008 | Covariance updates for discriminative training by constrained line search
Peter Bell 0001, Simon King 0001 |
INTERSPEECH | 2 |
| 2008 | Growing bottleneck features for tandem ASR
Joe Frankel, Dong Wang 0013, Simon King 0001 |
INTERSPEECH | 3 |
| 2008 | Unsupervised adaptation for HMM-based speech synthesisabstractIt is now possible to synthesise speech using HMMs with a comparable quality to unit-selection techniques. Generating speech from a model has many potential advantages over concatenating waveforms. The most exciting is model adaptation. It has been shown that supervised speaker adaptation can yield highquality synthetic voices with an order of magnitude less data than required to train a speaker-dependent model or to build a basic unit-selection system. Such supervised methods require labelled adaptation data for the target speaker. In this paper, we introduce a method capable of unsupervised adaptation, using only speech from the target speaker without any labelling. Index Terms: speech synthesis, HMM-based speech synthesis, HTS, trajectory HMMs, speaker adaptation, MLLR Simon King 0001, Keiichi Tokuda, Heiga Zen, Junichi Yamagishi |
INTERSPEECH | 1 |
| 2008 | Investigating festival's target cost function using perceptual experimentsabstractWe describe an investigation of the target cost used in the Festival unit selection speech synthesis system [1].Our ultimate goal is to automatically learn a perceptually optimal target cost function.In this study, we investigated the behaviour of the target cost for one segment type.The target cost is based on counting the mismatches in several context features.A carrier sentence ("My name is Roger") was synthesised using all 147,820 possible combinations of the diphones /n ei/ and /ei m/.92 representative versions were selected and presented to listeners as 460 pairwise comparisons.The listeners' preference votes were used to analyse the behaviour of the target cost, with respect to the values of its component linguistic context features. Volker Strom, Simon King 0001 |
INTERSPEECH | 2 |
| 2008 | Cross-lingual portability of MLP-based tandem features - a case study for English and HungarianabstractOne promising approach for building ASR systems for lessresourced languages is cross-lingual adaptation.Tandem ASR is particularly well suited to such adaptation, as it includes two cascaded modelling steps: feature extraction using multi-layer perceptrons (MLPs), followed by modelling using a standard HMM.The language-specific tuning can be performed by adjusting the HMM only, leaving the MLP untouched.Here we examine the portability of feature extractor MLPs between an Indo-European (English) and a Finno-Ugric (Hungarian) language.We present experiments which use both conventional phone-posterior and articulatory feature (AF) detector MLPs, both trained on a much larger quantity of (English) data than the monolingual (Hungarian) system.We find that the cross-lingual configurations achieve similar performance to the monolingual system, and that, interestingly, the AF detectors lead to slightly worse performance, despite the expectation that they should be more language-independent than phone-based MLPs.However, the cross-lingual system outperforms all other configurations when the English phone MLP is adapted on the Hungarian data. László Tóth 0001, Joe Frankel, Gábor Gosztolya, Simon King 0001 |
INTERSPEECH | 4 |
| 2008 | A posterior approach for microphone array based speech recognitionabstractAutomatic speech recognition (ASR) becomes rather difficult in meetings domains because of the adverse acoustic conditions, including more background noise, more echo and reverberation and frequent cross-talking. Microphone arrays have been demonstrated able to boost ASR performance dramatically in such noisy and reverberant environments, with various beamforming algorithms. However, almost all existing beamforming measures work in the acoustic domain, resorting to signal processing theories and geometric explanation. This limits their application, and induces significant performance degradation when the geometric property is unavailable or hard to estimate, or if heterogenous channels exist in the audio system. In this paper, we preset a new posterior-based approach for array-based speech recognition. The main idea is, instead of enhancing speech signals, we try to enhance the posterior probabilities that frames belonging to recognition units, e.g., phones. These enhanced posteriors are then transferred to posterior probability based features and are modeled by HMMs, leading to a tandem ANN-HMM hybrid system presented by Hermansky et al.. Experimental results demonstrated the validity of this posterior approach. With the posterior accumulation or enhancement, significant improvement was achieved over the single channel baseline. Moreover, we can combine the acoustic enhancement and posterior enhancement together, leading to a hybrid acoustic-posterior beamforming approach, which works significantly better than just the acoustic beamforming, especially in the scenario with moving-speakers. Dong Wang 0013, Ivan Himawan, Joe Frankel, Simon King 0001 |
INTERSPEECH | 4 |
| 2008 | Robustness of HMM-based speech synthesisabstractAs speech synthesis techniques become more advanced, we are able to consider building high-quality voices from data collected outside the usual highly-controlled recording studio environment. This presents new challenges that are not present in conventional text-to-speech synthesis: the available speech data are not perfectly clean, the recording conditions are not consistent, and/or the phonetic balance of the material is not ideal. Although a clear picture of the performance of various speech synthesis techniques (e.g., concatenative, HMM-based or hybrid) under good conditions is provided by the Blizzard Challenge, it is not well understood how robust these algorithms are to less favourable conditions. In this paper, we analyse the performance of several speech synthesis methods under such conditions. This is, as far as we know, a new research topic: ``Robust speech synthesis.'' As a consequence of our investigations, we propose a new robust training method for the HMM-based speech synthesis in for use with speech data collected in unfavourable conditions. Junichi Yamagishi, Zhen-Hua Ling, Simon King 0001 |
INTERSPEECH | 3 |
| 2008 | Bayesian networks for phone duration prediction
Olga Goubanova, Simon King 0001 |
Speech Commun. | 2 |
| 2008 | A comparison of grapheme and phoneme-based units for Spanish spoken term detection
Javier Tejedor, Dong Wang 0013, Joe Frankel, Simon King 0001, José Colás Pasamontes |
Speech Commun. | 4 |
| 2007 | Monolingual and crosslingual comparison of tandem features derived from articulatory and phone MLPSabstractThe features derived from posteriors of a multilayer perceptron (MLP), known as tandem features, have proven to be very effective for automatic speech recognition. Most tandem features to date have relied on MLPs trained for phone classification. We recently showed on a relatively small data set that MLPs trained for articulatory feature classification can be equally effective. In this paper, we provide a similar comparison using MLPs trained on a much larger data set -2000 hours of English conversational telephone speech. We also explore how portable phone-and articulatory feature-based tandem features are in an entirely different language - Mandarin - without any retraining. We find that while the phone-based features perform slightly better than AF-based features in the matched-language condition, they perform significantly better in the cross-language condition. However, in the cross-language condition, neither approach is as effective as the tandem features extracted from an MLP trained on a relatively small amount of in-domain data. Beyond feature concatenation, we also explore novel factored observation modeling schemes that allow for greater flexibility in combining the tandem and standard features. Özgür Çetin, Mathew Magimai-Doss, Karen Livescu, Arthur Kantor, Simon King 0001, Chris D. Bartels, Joe Frankel |
ASRU | 5 |
| 2007 | An Articulatory Feature-Based Tandem Approach and Factored Observation ModelingabstractThe so-called tandem approach, where the posteriors of a multilayer perceptron (MLP) classifier are used as features in an automatic speech recognition (ASR) system has proven to be a very effective method. Most tandem approaches up to date have relied on MLPs trained for phone classification, and appended the posterior features to some standard feature hidden Markov model (HMM). In this paper, we develop an alternative tandem approach based on MLPs trained for articulatory feature (AF) classification. We also develop a factored observation model for characterizing the posterior and standard features at the HMM outputs, allowing for separate hidden mixture and state-tying structures for each factor. In experiments on a subset of Switchboard, we show that the AF-based tandem approach is as effective as the phone-based approach, and that the factored observation model significantly outperforms the simple feature concatenation approach while using fewer parameters. Özgür Çetin, Arthur Kantor, Simon King 0001, Chris D. Bartels, Mathew Magimai-Doss, Joe Frankel, Karen Livescu |
ICASSP (4) | 3 |
| 2007 | Manual Transcription of Conversational Speech at the Articulatory Feature LevelabstractWe present an approach for the manual labeling of speech at the articulatory feature level, and a new set of labeled conversational speech collected using this approach. A detailed transcription, including overlapping or reduced gestures, is useful for studying the great pronunciation variability in conversational speech. It also facilitates the testing of feature classifiers, such as those used in articulatory approaches to automatic speech recognition. We describe an effort to transcribe a small set of utterances drawn from the Switchboard database using eight articulatory tiers. Two transcribers have labeled these utterances in a multi-pass strategy, allowing for correction of errors. We describe the data collection methods and analyze the data to determine how quickly and reliably this type of transcription can be done. Finally, we demonstrate one use of the new data set by testing a set of multilayer perceptron feature classifiers against both the manual labels and forced alignments. Karen Livescu, Ari Bezman, Nash M. Borges, Lisa Yung, Özgür Çetin, Joe Frankel, Simon King 0001, Mathew Magimai-Doss, Xuemin Chi, Lisa Lavoie |
ICASSP (4) | 7 |
| 2007 | Articulatory Feature-Based Methods for Acoustic and Audio-Visual Speech Recognition: Summary from the 2006 JHU Summer workshopabstractWe report on investigations, conducted at the 2006 Johns Hopkins Workshop, into the use of articulatory features (AFs) for observation and pronunciation models in speech recognition. In the area of observation modeling, we use the outputs of AF classifiers both directly, in an extension of hybrid HMM/neural network models, and as part of the observation vector, an extension of the "tandem" approach. In the area of pronunciation modeling, we investigate a model having multiple streams of AF states with soft synchrony constraints, for both audio-only and audio-visual recognition. The models are implemented as dynamic Bayesian networks, and tested on tasks from the small-vocabulary switchboard (SVitchboard) corpus and the CUAVE audio-visual digits corpus. Finally, we analyze AF classification and forced alignment using a newly collected set of feature-level manual transcriptions. Karen Livescu, Özgür Çetin, Mark Hasegawa-Johnson, Simon King 0001, Chris D. Bartels, Nash M. Borges, Arthur Kantor, Partha Lal, Lisa Yung, Ari Bezman, Stephen Dawson-Haggerty, Bronwyn Woods, Joe Frankel, Mathew Magimai-Doss, Kate Saenko |
ICASSP (4) | 4 |
| 2007 | Sparse Gaussian graphical models for speech recognitionabstractWe address the problem of learning the structure of Gaussian graphical models for use in automatic speech recognition, a means of controlling the form of the inverse covariance matrices of such systems. With particular focus on data sparsity issues, we implement a method for imposing graphical model structure on a Gaussian mixture system, using a convex optimisation technique to maximise a penalised likelihood expression. The results of initial experiments on a phone recognition task show a performance improvement over an equivalent full-covariance system. Index Terms: speech recognition, acoustic models, graphical models, precision matrix models Peter Bell 0001, Simon King 0001 |
INTERSPEECH | 2 |
| 2007 | Articulatory feature classifiers trained on 2000 hours of telephone speechabstractThe so-called tandem approach, where the posteriors of a multilayer perceptron (MLP) classifier are used as features in an automatic speech recognition (ASR) system has proven to be a very effective method. Most tandem approaches up to date have relied on MLPs trained for phone classification, and appended the posterior features to some standard feature hidden Markov model (HMM). In this paper, we develop an alternative tandem approach based on MLPs trained for articulatory feature (AF) classification. We also develop a factored observation model for characterizing the posterior and standard features at the HMM outputs, allowing for separate hidden mixture and state-tying structures for each factor. In experiments on a subset of Switchboard, we show that the AFbased tandem approach is as effective as the phone-based approach, and that the factored observation model significantly outperforms the simple feature concatenation approach while using fewer parameters. Joe Frankel, Mathew Magimai-Doss, Simon King 0001, Karen Livescu, Özgür Çetin |
INTERSPEECH | 3 |
| 2007 | Modelling prominence and emphasis improves unit-selection synthesisabstractWe describe the results of large scale perception experiments showing improvements in synthesising two distinct kinds of prominence: standard pitch-accent and strong emphatic accents. Previously prominence assignment has been mainly evaluated by computing accuracy on a prominence-labelled test set. By contrast we integrated an automatic pitch-accent classifier into the unit selection target cost and showed that listeners preferred these synthesised sentences. We also describe an improved recording script for collecting emphatic accents, and show that generating emphatic accents leads to further improvements in the fiction genre over incorporating pitch accent only. Finally, we show differences in the effects of prominence between child-directed speech and news and fiction genres. Index Terms: speech synthesis, prosody, prominence, pitch accent, unit selection Volker Strom, Ani Nenkova, Robert A. J. Clark, Yolanda Vazquez-Alvarez, Jason M. Brenier, Simon King 0001, Daniel Jurafsky |
INTERSPEECH | 6 |
| 2007 | Articulatory feature recognition using dynamic Bayesian networks
Joe Frankel, Mirjam Wester, Simon King 0001 |
Comput. Speech Lang. | 3 |
| 2007 | Factoring Gaussian precision matrices for linear dynamic models
Joe Frankel, Simon King 0001 |
Pattern Recognit. Lett. | 2 |
| 2007 | Multisyn: Open-domain unit selection for the Festival speech synthesis system
Robert A. J. Clark, Korin Richmond, Simon King 0001 |
Speech Commun. | 3 |
| 2007 | Speech Recognition Using Linear Dynamic ModelsabstractThe majority of automatic speech recognition systems rely on hidden Markov models, in which Gaussian mixtures model the output distributions associated with sub-phone states. This approach, whilst successful, models consecutive feature vectors (augmented to include derivative information) as statistically independent. Furthermore, spatial correlations present in speech parameters are frequently ignored through the use of diagonal covariance matrices. This paper continues the work of Digalakis and others who proposed instead a first-order linear state-space model which has the capacity to model underlying dynamics, and furthermore give a model of spatial correlations. This paper examines the assumptions made in applying such a model and shows that the addition of a hidden dynamic state leads to increases in accuracy over otherwise equivalent static models. We also propose a time-asynchronous decoding strategy suited to recognition with segment models. We describe implementation of decoding for linear dynamic models and present TIMIT phone recognition results Joe Frankel, Simon King 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Joint prosodic and segmental unit selection speech synthesisabstractWe describe a unit selection technique for text-to-speech synthesis which jointly searches the space of possible diphone sequences and the space of possible prosodic unit sequences in order to produce synthetic speech with more natural prosody. We demonstrates that this search, although currently computationally expensive, can achieve improved intonation compared to a baseline in which only the space of possible diphone sequences is searched. We discuss ways in which the search could be made sufficiently efficient for use in a real-time system. Robert A. J. Clark, Simon King 0001 |
INTERSPEECH | 2 |
| 2006 | Expressive prosody for unit-selection speech synthesisabstractCurrent unit selection speech synthesis voices cannot produce emphasis or interrogative contours because of a lack of the necessary prosodic variation in the recorded speech database. A method of recording script design is proposed which addresses this shortcoming. Appropriate components were added to the target cost function of the Festival Multisyn engine, and a perceptual evaluation showed a clear preference over the baseline system. Volker Strom, Robert A. J. Clark, Simon King 0001 |
INTERSPEECH | 3 |
| 2006 | Observation process adaptation for linear dynamic models
Joe Frankel, Simon King 0001 |
Speech Commun. | 2 |
| 2006 | Subjective evaluation of join cost and smoothing methods for unit selection speech synthesisabstractIn unit selection-based concatenative speech synthesis, join cost (also known as concatenation cost), which measures how well two units can be joined together, is one of the main criteria for selecting appropriate units from the inventory. Usually, some form of local parameter smoothing is also needed to disguise the remaining discontinuities. This paper presents a subjective evaluation of three join cost functions and three smoothing methods. We also describe the design and performance of a listening test. The three join cost functions were taken from our previous study, where we proposed join cost functions derived from spectral distances, which have good correlations with perceptual scores obtained for a range of concatenation discontinuities. This evaluation allows us to further validate their ability to predict concatenation discontinuities. The units for synthesis stimuli are obtained from a state-of-the-art unit selection text-to-speech system: rVoice from Rhetorical Systems Ltd. In this paper, we report listeners' preferences for each join cost in combination with each smoothing method Jithendra Vepa, Simon King 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Detection of Symbolic Gestural Events in Articulatory Data for Use in Structural Representations of Continuous SpeechabstractOne of the crucial issues which often needs to be addressed in structural approaches to speech representation is the choice of fundamental symbolic units of representation. In this paper, a physiologically inspired methodology for defining these symbolic atomic units in terms of primitive articulatory events is proposed. It is shown how the atomic articulatory events (gestures) can be detected directly in the articulatory data. An algorithm for evaluating the reliability of the articulatory events is described and promising results of the experiments conducted on the MOCHA articulatory database are presented. Alexander Gutkin, Simon King 0001 |
ICASSP (1) | 2 |
| 2005 | Genetic triangulation of graphical models for speech and language processingabstractGraphical models are an increasingly popular approach for speech and language processing. As researchers design ever more complex models it becomes crucial to find triangulations that make inference problems tractable. This paper presents a genetic algorithm for triangulation search that is well-suited for speech and language graphical models. It is unique in two ways: First, it can find triangulations appropriate for graphs with a mix of stochastic and deterministic dependencies. Second, the search is guided by optimizing the inference speed (CPU runtime) on real data. We show results on 10 real-world speech and language graphs and demonstrate inference speed-ups over standard triangulation methods. 1. Chris D. Bartels, Kevin Duh, Jeff A. Bilmes, Katrin Kirchhoff, Simon King 0001 |
INTERSPEECH | 5 |
| 2005 | Multisyn voices from ARCTIC data for the blizzard challengeabstractThis paper describes the process of building unit selection voices for the Festival Multisyn engine using four ARCTIC datasets, as part of the Blizzard evaluation challenge. The build process is almost entirely automatic, with very little need for human intervention. We discuss the difference in the evaluation results for each voice and evaluate the suitability of the ARCTIC datasets for building this type of voice. Robert A. J. Clark, Korin Richmond, Simon King 0001 |
INTERSPEECH | 3 |
| 2005 | A hybrid ANN/DBN approach to articulatory feature recognitionabstractArtificial neural networks (ANN) have proven to be well suited to the task of articulatory feature (AF) recognition. Previous studies have taken a cascaded approach where separate ANNs are trained for each feature group, making the assumption that features are statistically independent. We address this by using ANNs to provide virtual evidence to a dynamic Bayesian network (DBN). This gives a hybrid ANN/DBN model and allows modelling of inter-feature dependencies. We demonstrate significant increases in AF recognition accuracy from modelling dependencies between features, and present the results of embedded training experiments in which a set of asynchronous feature changes are learned. Furthermore, we report on the application of a Viterbi training scheme in which we alternate between realigning the AF training labels and retraining the ANNs. Joe Frankel, Simon King 0001 |
INTERSPEECH | 2 |
| 2005 | Predicting consonant duration with Bayesian belief networksabstractConsonant duration is influenced by a number of linguistic factors such as the consonant’s identity, within-word position, stress level of the previous and following vowels, phrasal position of the word containing the target consonant, its syllabic position, identity of the previous and following segments. In our work, consonant duration is predicted from a Bayesian belief network (BN) consisting of discrete nodes for the linguistic factors and a single continuous node for the consonant’s duration. Interactions between factors are represented as conditional dependency arcs in this graphical model. Given the parameters of the belief network, the duration of each consonant in the test set is then predicted as the value with the maximum probability. We compare the results of the belief network model with those of sums-of-products (SoP) and classification and regression tree (CART) models using the same data. In terms of RMS error, our BN model performs better than both CART and SoP models. In terms of the correlation coefficient, our BN model performs better than SoP model, and no worse than CART model. In addition, the Bayesian model reliably predicts consonant duration in cases of missing or hidden linguistic factors. 1. Olga Goubanova, Simon King 0001 |
INTERSPEECH | 2 |
| 2005 | SVitchboard 1: small vocabulary tasks from SwitchboardabstractWe present a conversational telephone speech data set designed to support research on novel acoustic models. Small vocabulary tasks from 10 words up to 500 words are defined using subsets of the Switchboard-1 corpus; each task has a completely closed vocabulary (an OOV rate of 0%). We justify the need for these tasks, describe the algorithm for selecting them from a large corpus, give a statistical analysis of the data and present baseline whole-word hidden Markov model recognition results. The goal of the paper is to define a common data set and to encourage other researchers to use it. Simon King 0001, Chris D. Bartels, Jeff A. Bilmes |
INTERSPEECH | 1 |
| 2005 | Multidimensional scaling of listener responses to synthetic speechabstractThe move to unit-selection in speech synthesis has resulted in system improvements being made at subtle sub- and suprasegmental levels. Human perceptual evaluation of such subtle improvements requires a highly sophisticated level of perceptual attention to specific acoustic characteristics or cues. However, it is not well understood what acoustic cues listeners attend to by default when asked to evaluate synthetic speech. It may, therefore, be potentially quite difficult to design an evaluation method that allows listeners to concentrate on only one dimension of the signal, while ignoring others that are perceptually more important to them. This paper describes a pilot study which aims to evaluate multidimensional scaling (MDS) as a possible method of determining what acoustic characteristics of synthetic speech influence listeners’ judgements of the naturalness of the speech. Using distance measures (either real or perceived distances), MDS techniques represent stimuli as points in n-dimensional space. The space is configured so that similar stimuli are close together, while different stimuli are farther apart. Additionally, the dimensions of the space correspond to characteristics of the stimuli which influenced the perceived distances. Our results indicate that MDS techniques should be a useful tool in understanding the complex psychoacoustic processes that listeners undergo when evaluating synthetic speech. This method has allowed us to identify a number of cues that appear to be particularly perceptually salient to listeners evaluating synthetic speech naturalness, namely prosodic cues (in terms of duration and/or intonation) and segmental or unit level cues (in terms of appropriateness of units, or number of units). Catherine Mayo, Robert A. J. Clark, Simon King 0001 |
INTERSPEECH | 3 |
| 2004 | Articulatory feature recognition using dynamic Bayesian networksabstractWe describe a dynamic Bayesian network for articulatory feature recognition. The model is intended to be a component of a speech recognizer that avoids the problems of conventional ''beads-on-a-string'' phoneme-based models. We demonstrate that the model gives superior recognition of articulatory features from the speech signal compared with a state-of-the-art neural network system. We also introduce a training algorithm that offers two major advances: it does not require time-aligned feature labels and it allows the model to learn a set of asynchronous feature changes in a data-driven manner. Joe Frankel, Mirjam Wester, Simon King 0001 |
INTERSPEECH | 3 |
| 2004 | Phone classification in pseudo-euclidean vector spacesabstractRecently we have proposed a structural framework for modelling speech, which is based on patterns of phonological distinctive features, a linguistically well-motivated alternative to standard vector-space acoustic models like HMMs. This framework gives considerable representational freedom by working with features that have explicit linguistic interpretation, but at the expense of the ability to apply the wide range of analytical decision algorithms available in vector spaces, restricting oneself to more computationally expensive and less-developed symbolic metric tools. In this paper we show that a dissimilarity-based distance-preserving transition from the original structural representation to a corresponding pseudo-Euclidean vector space is possible. Promising results of phone classification experiments conducted on the TIMIT database are reported. Alexander Gutkin, Simon King 0001 |
INTERSPEECH | 2 |
| 2004 | Estimating detailed spectral envelopes using articulatory clusteringabstractThis paper presents an articulatory-acoustic mapping where detailed spectral envelopes are estimated. During the estimation, the harmonics of a range of F0 values are derived from the spectra of multiple voiced speech signals vocalized with similar articulator settings. The envelope formed by these harmonics is represented by a cepstrum, which is computed by fitting the peaks of all the harmonics based on the weighted least square method in the frequency domain. The experimental result shows that the spectral envelopes are estimated with the highest accuracy when the cepstral order is 48-64 for a female speaker, which suggests that representing the real response of the vocal tract requires high-quefrency elements that conventional speech synthesis methods are forced to discard in order to eliminate the pitch component of speech. 1. Yoshinori Shiga, Simon King 0001 |
INTERSPEECH | 2 |
| 2004 | Source-filter separation for articulation-to-speech synthesisabstractIn this paper we examine a method for separating out the vocal-tract filter response from the voice source characteristic using a large articulatory database. The method realises such separation for voiced speech using an iterative approximation procedure under the assumption that the speech production process is a linear system composed of a voice source and a vocal-tract filter, and that each of the components is controlled independently by different sets of factors. Experimental results show that the spectral variation is evidently influenced by the fundamental frequency or the power of speech, and that the tendency of the variation may be related closely to speaker identity. The method enables independent control over the voice source characteristic in our articulation-to-speech synthesis. 1. Yoshinori Shiga, Simon King 0001 |
INTERSPEECH | 2 |
| 2004 | Subjective evaluation of join cost functions used in unit selection speech synthesisabstractIn our previous papers, we have proposed join cost functions derived from spectral distances, which have good correlations with perceptual scores obtained for a range of concatenation discontinuities. To further validate their ability to predict concatenation discontinuities, we have chosen the best three spectral distances and evaluated them subjectively in a listening test. The unit sequences for synthesis stimuli are obtained from a state-of-the-art unit selection text-to-speech system: rVoice from Rhetorical Systems Ltd. In this paper, we report listeners' preferences for each of the three join cost functions. Jithendra Vepa, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Transforming voice qualityabstractVoice transformation is the process of transforming the characteristics of speech uttered by a source speaker, such that a listener would believe the speech was uttered by a target speaker. In this paper we address the problem of transforming voice quality. We do not attempt to transform prosody. Our system has two main parts corresponding to the two components of the source-filter model of speech production. The first component transforms the spectral envelope as represented by a linear prediction model. The transformation is achieved using a Gaussian mixture model, which is trained on aligned speech from source and target speakers. The second part of the system predicts the spectral detail from the transformed linear prediction coefficients. A novel approach is proposed, which is based on a classifier and residual codebooks. On the basis of a number of performance metrics it outperforms existing systems. Ben Gillett, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Transforming F0 contoursabstractVoice transformation is the process of transforming the characteristics of speech uttered by a source speaker, such that a listener would believe the speech was uttered by a target speaker. Training F0 contour generation models for speech synthesis requires a large corpus of speech. If it were possible to adapt the F0 contour of one speaker to sound like that of another speaker, using a small, easily obtainable parameter set, this would be extremely valuable. We present a new method for the transformation of F0 contours from one speaker to another based on a small linguistically motivated parameter set. The system performs a piecewise linear mapping using these parameters. A perceptual experiment clearly demonstrates that the presented system is at least as good as an existing technique for all speaker pairs, and that in many cases it is much better and almost as good as using the target F0 contour Ben Gillett, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Named entity extraction from word latticesabstractWe present a method for named entity extraction from word lattices produced by a speech recogniser. Previous work by others on named entity extraction from speech has used either a manual transcript or 1-best recogniser output. We describe how a single Viterbi search can recover both the named entity sequence and the corresponding word sequence from a word lattice, and further that it is possible to trade off an increase in word error rate for improved named entity extraction. James Horlock, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Discriminative methods for improving named entity extraction on speech dataabstractIn this paper we present a method of discriminatively training language models for spoken language understanding; we show improvements in named entity F-scores on speech data using these improved language models. A comparison between theoretical probabilities associated with manual markup and the actual probabilities of output markup is used to identify probabilities requiring adjustment. We present results which support our hypothesis that improvements in F-scores are possible by using either previously used training data or held out development data to improve discrimination amongst a set of N-gram language models. James Horlock, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Estimating the spectral envelope of voiced speech using multi-frame analysisabstractThis paper proposes a novel approach for estimating the spectral envelope of voiced speech independently of its harmonic structure. Because of the quasi-periodicity of voiced speech, its spectrum indicates harmonic structure and only has energy at frequencies corresponding to integral multiples of F0. It is hence impossible to identify transfer characteristics between the adjacent harmonics. In order to resolve this problem, Multi-frame Analysis (MFA) is introduced. The MFA estimates a spectral envelope using many portions of speech which are vocalised using the same vocal-tract shape. Since each of the portions usually has a different F0 and ensuing different harmonic structure, a number of harmonics can be obtained at various frequencies to form a spectral envelope. The method thereby gives a closer approximation to the vocal-tract transfer function. Yoshinori Shiga, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Estimation of voice source and vocal tract characteristics based on multi-frame analysisabstractThis paper presents a new approach for estimating voice source and vocal tract filter characteristics of voiced speech. When it is required to know the transfer function of a system in signal processing, the input and output of the system are experimentally observed and used to calculate the function. However, in the case of source-filter separation we deal with in this paper, only the output (speech) is observed and the characteristics of the system (vocal tract) and the input (voice source) must simultaneously be estimated. Hence the estimate becomes extremely difficult, and it is usually solved approximately using oversimplified models. We demonstrate that these characteristics are separable under the assumption that they are independently controlled by different factors. The separation is realised using an iterative approximation along with the Multi-frame Analysis method, which we have proposed to find spectral envelopes of voiced speech with minimum interference of the harmonic structure. Yoshinori Shiga, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Kalman-filter based join cost for unit-selection speech synthesisabstractWe introduce a new method for computing join cost in unit-selection speech synthesis which uses a linear dynamical model (also known as a Kalman filter) to model line spectral frequency trajectories. The model uses an underlying subspace in which it makes smooth, continuous trajectories. This subspace can be seen as an analogy for underlying articulator movement. Once trained, the model can be used to measure how well concatenated speech segments join together. The objective join cost is based on the error between model predictions and actual observations. We report correlations between this measure and mean listener scores obtained from a perceptual listening experiment. Our experiments use a state-of-the art unit-selection text-to-speech system: `rVoice' from Rhetorical Systems Ltd. Jithendra Vepa, Simon King 0001 |
INTERSPEECH | 2 |
| 2003 | Modelling the uncertainty in recovering articulation from acoustics
Korin Richmond, Simon King 0001, Paul Taylor 0001 |
Comput. Speech Lang. | 2 |
| 2002 | Framewise phone classification using support vector machinesabstractWe describe the use of Support Vector Machines for phonetic classification on the TIMIT corpus. Unlike previous work, in which entire phonemes are classified, our system operates in a framewise manner and is intended for use as the front-end of a hybrid system similar to ABBOT. We therefore avoid the problems of classifying variable-length vectors. Our frame-level phone classification accuracy on the complete TIMIT test set is competitive with other results from the literature. In addition, we address the serious problem of scaling Support Vector Machines by using the Kernel Fisher Discriminant. 1. Jesper Salomon, Simon King 0001, Miles Osborne |
INTERSPEECH | 2 |
| 2002 | Objective distance measures for spectral discontinuities in concatenative speech synthesisabstractIn unit selection based concatenative speech systems, join cost, which measures how well two units can be joined together, is one of the main criteria for selecting appropriate units from the inventory. The ideal join cost will measure perceived discontinuity, based on easily measurable spectral properties of the units being joined, in order to ensure smooth and natural-sounding synthetic speech. In this paper we report a perceptual experiment conducted to measure the correlation between subjective human perception and various objective spectrally-based measures proposed in the literature. Our experiments used a state-of-the art unit-selection text-to-speech system: rVoice from Rhetorical Systems Ltd. 1. Jithendra Vepa, Simon King 0001, Paul Taylor 0001 |
INTERSPEECH | 2 |
| 2001 | ASR - articulatory speech recognitionabstractWe propose that using a continuous trajectory model to describe an articulatory-based feature set will address some of the shortcomings inherent in the hidden Markov model (HMM) as a model for speech recognition. The articulatory parameters allow us to explicitly model effects such as co-articulation and assimilation. A linear dynamic model (LDM) is used to capture the characteristics of each segment type. These models are well suited to describing smoothly varying, continuous, yet noisy trajectories, such as we find present in speech data. Experimentation has been based on data for a single speaker from the MOCHA corpus. This consists of parallel acoustic and recorded articulatory parameters for 460 TIMIT sentences. We report the results of classification and recognition tasks using both real and recovered articulatory parameters, on their own and in conjunction with acoustic features. Joe Frankel, Simon King 0001 |
INTERSPEECH | 2 |
| 2000 | An automatic speech recognition system using neural networks and linear dynamic models to recover and model articulatory tracesabstractWe describe a speech recognition system which uses articulatory parameters as basic features and phone-dependent linear dynamic models. The system first estimates articulatory trajectories from the speech signal. Estimations of x and y coordinates of 7 actual articulator positions in the midsagittal plane are produced every 2 milliseconds by a recurrent neural network, trained on real articulatory data. The output of this network is then passed to a set of linear dynamic models, which perform phone recognition Joe Frankel, Korin Richmond, Simon King 0001, Paul Taylor 0001 |
INTERSPEECH | 3 |
| 2000 | Detection of phonological features in continuous speech using neural networks
Simon King 0001, Paul Taylor 0001 |
Comput. Speech Lang. | 1 |
| 1998 | Speech recognition via phonetically featured syllablesabstractSpeech can be naturally described by phonetic features, such as a set of acoustic phonetic features or a set of articulatory features. This thesis establi shes the effectiveness of using phonetic features in phoneme recognition by comparing a recogniser based on them to a recogniser using an established parametrisation as a baseline. The usefulness of phonetic features serves as the foundation for the subsequent modelling of syllables. Syllables are subject to fewer of the context-sensitivity effects that hamper phone-based speech recognition. I investigate the different questions involved in creating syllable models. After training a feature-based syllable recogniser, I compare the feature based syllables against a baseline. To conclude, the feature based syllable models are compared against the baseline phoneme models in word recognition. With the resultant feature-syllable models performing well in word recognition, the featuresyllables show their future potential for large vocabulary automatic speech recognition. Simon King 0001, Todd A. Stephenson, Stephen Isard, Paul Taylor 0001, Alex Strachan |
ICSLP | 1 |
| 1997 | Speech synthesis using non-uniform units in the Verbmobil project
Simon King 0001, Thomas Portele, Florian Höfer |
EUROSPEECH | 1 |
| 1997 | Using intonation to constrain language models in speech recognitionabstractThis paper describes a method for using intonation to reduce word error rate in a speech recognition system designed to recognise spontaneous dialogue speech. We use a form of dialogue analysis based on the theory of conversational games. Different move types under this analysis conform to different language models. Different move types are also characterised by different intonational tunes. Our overall recognition strategy is first to predict from intonation the type of game move that a test utterance represents, and then to use a bigram language model for that type of move during recognition. 1 INTRODUCTION This paper describes a method for using intonation to reduce word error rate in a speech recognition system designed to recognise spontaneous dialogue speech. Our experiments are on the DCIEM Maptask corpus [2], a corpus of spontaneous task-oriented dialogue speech. Our dialogue analysis is based on the theory of conversational games first introduced by Power [9] and adapted for ... Paul Taylor 0001, Simon King 0001, Stephen Isard, Helen Wright, Jacqueline C. Kowtko |
EUROSPEECH | 2 |
| 1996 | Using prosodic information to constrain language models for spoken dialogueabstractWe present w ork intended to improve speech recognition performance for computer dialogue by taking into account the way that dialogue context and intonational tune interact to limit the possibilities for what an utterance might be.We report here on the extra constraint a c hieved in a bigram language model, expressed in terms of entropy, b y using separate submodels for dierent sorts of dialogue acts, and trying to predict which submodel to apply by analysis of the intonation of the sentence being recognised. Paul Taylor 0001, Hiroshi Shimodaira, Stephen Isard, Simon King 0001, Jacqueline C. Kowtko |
ICSLP | 4 |