Tuomo Raitio

dblp:19/8809 · DBLP profile ↗
← Back
42ranked-venue papers
16as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 28 · 11 first-author · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2024 Dialog Modeling in Audiobook Synthesis
abstract
In audiobook synthesis, it is important to have the ability to differentiate between dialog and narration or different characters. In this work, we propose dialog modeling methods for audiobook synthesis. The proposed approach consists of two stages. First, a text-based dialog style classifier is employed to predict narration vs. dialog from text, and further predict the corresponding characters into soprano and baritone. Then, a dialog style adaptor is added to the text-to-speech (TTS) model to allow synthesizing speech with the corresponding styles. With a speaker verification (SV) based style adaptor, we can even control the strength of a given style. We evaluated the proposed approach in audiobook synthesis with a mean opinion score (MOS) listening test using 9 carefully designed questions. The results show an improvement of 0.35 MOS on dialog distinction without degradation in other aspects. Also a comparative MOS (CMOS) test is conducted to verify the effectiveness of the proposed method.
Cheng-chieh Yeh, Amirreza Shirani, Weicheng Zhang, Tuomo Raitio, Ramya Rasipuram, Ladan Golipour, David Winarsky
ICASSP4
2022 Hierarchical Prosody Modeling and Control in Non-Autoregressive Parallel Neural TTS
abstract
Neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the synthetic speech often represents the average prosodic style of the database instead of having more versatile prosodic variation. Moreover, many models lack the ability to control the output prosody, which does not allow for different styles for the same text input. In this work, we train a non-autoregressive parallel neural TTS front-end model hierarchically conditioned on both coarse and fine-grained acoustic speech features to learn a latent prosody space with intuitive and meaningful dimensions. Experiments show that a non-autoregressive TTS model hierarchically conditioned on utterance-wise pitch, pitch range, duration, energy, and spectral tilt can effectively control each prosodic dimension, generate a wide variety of speaking styles, and provide word-wise emphasis control, while maintaining equal or better quality to the baseline model.
Tuomo Raitio, Jiangchuan Li, Shreyas Seshadri
ICASSP1
2022 Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise
abstract
We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise.The method consists of first measuring the spectral tilt of unlabeled conventional speech data, and then conditioning a neural TTS model with normalized spectral tilt among other prosodic factors.Changing the spectral tilt parameter and keeping other prosodic factors unchanged enables effective vocal effort control at synthesis time independent of other prosodic factors.By extrapolation of the spectral tilt values beyond what has been seen in the original data, we can generate speech with high vocal effort levels, thus improving the intelligibility of speech in the presence of masking noise.We evaluate the intelligibility and quality of normal speech and speech with increased vocal effort in the presence of various masking noise conditions, and compare these to well-known speech intelligibility-enhancing algorithms.The evaluations show that the proposed method can improve the intelligibility of synthetic speech with little loss in speech quality.
Tuomo Raitio, Petko Petkov, Jiangchuan Li, P. V. Muhammed Shifas, Andrea Davis, Yannis Stylianou
INTERSPEECH1
2022 Emphasis Control for Parallel Neural TTS
abstract
Recent parallel neural text-to-speech (TTS) synthesis methods are able to generate speech with high fidelity while maintaining high performance.However, these systems often lack control over the output prosody, thus restricting the semantic information conveyable for a given text.This paper proposes a hierarchical parallel neural TTS system for prosodic emphasis control by learning a latent space that directly corresponds to a change in emphasis.Three candidate features for the latent space are compared: 1) Variance of pitch and duration within words in a sentence, 2) Wavelet-based feature computed from pitch, energy, and duration, and 3) Learned combination of the two aforementioned approaches.At inference time, word-level prosodic emphasis is achieved by increasing the feature values of the latent space for the given words.Experiments show that all the proposed methods are able to achieve the perception of increased emphasis with little loss in overall quality.Moreover, emphasized utterances were preferred in a pairwise comparison test over the non-emphasized utterances, indicating promise for real-world applications.
Shreyas Seshadri, Tuomo Raitio, Dan Castellani, Jiangchuan Li
INTERSPEECH2
2021 On-Device Neural Speech Synthesis
abstract
Recent advances in text-to-speech (TTS) synthesis, such as Tacotron and WaveRNN, have made it possible to construct a fully neural network based TTS system, by coupling the two components together. Such a system is conceptually sim-ple as it only takes grapheme or phoneme input, uses Mel-spectrogram as an intermediate feature, and directly generates speech samples. The system achieves quality equal or close to natural speech. However, the high computational cost of the system and issues with robustness have limited their usage in real-world speech synthesis applications and products. In this paper, we present key modeling improvements and optimization strategies that enable deploying these models, not only on GPU servers, but also on mobile devices. The proposed system can generate high-quality 24 kHz speech at 5x faster than real time on server and 3x faster than real time on mobile devices.
Sivanand Achanta, Albert Antony, Ladan Golipour, Jiangchuan Li, Tuomo Raitio, Ramya Rasipuram, Jennifer Shi, Jaimin Upadhyay, David Winarsky, Hepeng Zhang
ASRU5
2021 Whispered and Lombard Neural Speech Synthesis
abstract
It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user. In this paper, we present and compare various approaches for generating different speaking styles, namely, normal, Lombard, and whisper speech, using only limited data. The following systems are proposed and assessed: 1) Pre-training and fine-tuning a model for each style. 2) Lombard and whisper speech conversion through a signal processing based approach. 3) Multi-style generation using a single model based on a speaker verification model. Our mean opinion score and AB preference listening tests show that 1) we can generate high quality speech through the pre-training/fine-tuning approach for all speaking styles. 2) Although our speaker verification (SV) model is not explicitly trained to discriminate different speaking styles, and no Lombard and whisper voice is used for pretrain this system, SV model can be used as style encoder for generating different style embeddings as input for Tacotron system. We also show that the resulting synthetic Lombard speech has a significant positive impact on intelligibility gain.
Qiong Hu 0003, Tobias Bleisch, Petko Petkov, Tuomo Raitio, Erik Marchi, Varun Lakshminarasimhan
SLT4
2020 Controllable Neural Text-to-Speech Synthesis Using Intuitive Prosodic Features
abstract
Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech.However, the prosody of generated utterances often represents the average prosodic style of the database instead of having wide prosodic variation.Moreover, the generated prosody is solely defined by the input text, which does not allow for different styles for the same sentence.In this work, we train a sequence-to-sequence neural network conditioned on acoustic speech features to learn a latent prosody space with intuitive and meaningful dimensions.Experiments show that a model conditioned on sentencewise pitch, pitch range, phone duration, energy, and spectral tilt can effectively control each prosodic dimension and generate a wide variety of speaking styles, while maintaining similar mean opinion score (4.23) to our Tacotron baseline (4.26).
Tuomo Raitio, Ramya Rasipuram, Dan Castellani
INTERSPEECH1
2017 Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System
Tim Capes, Paul Coles, Alistair Conkie, Ladan Golipour, Abie Hadjitarkhani, Qiong Hu 0003, Nancy Huddleston, Melvyn Hunt, Jiangchuan Li, Matthias Neeracher, Kishore Prahallad, Tuomo Raitio, Ramya Rasipuram, Greg Townsend, Becci Williamson, David Winarsky, Zhizheng Wu 0001, Hepeng Zhang
INTERSPEECH12
2016 Phase perception of the glottal excitation and its relevance in statistical parametric speech synthesis
Tuomo Raitio, Lauri Juvela, Antti Suni, Martti Vainio, Paavo Alku
Speech Commun.1
2015 Noise robust estimation of the voice source using a deep neural network
abstract
In the analysis of speech production, information about the voice source can be obtained non-invasively with glottal inverse filtering (GIF) methods. Current state-of-the-art GIF methods are capable of producing high-quality estimates in suitable conditions (e.g. low noise and reverberation), but their performance deteriorates in nonideal conditions because they require noise-sensitive parameter estimation. This study proposes a method for noise robust estimation of the voice source by creating a mapping using a deep neural network (DNN) between robust low-level speech features and the desired reference, a time-domain glottal flow computed by a GIF method. The method was evaluated with two GIF methods, of which one (quasi closed phase analysis, QCP) requires additional parameter estimation and the other (iterative adaptive inverse filtering, IAIF) does not. The results show that the proposed method outperforms the QCP method with SNRs less than 50-20 dB, but the simple IAIF method only with very low SNRs.
Manu Airaksinen, Tuomo Raitio, Paavo Alku
ICASSP2
2015 Phase perception of the glottal excitation of vocoded speech
Tuomo Raitio, Lauri Juvela, Antti Suni, Martti Vainio, Paavo Alku
INTERSPEECH1
2015 A Deep Generative Architecture for Postfiltering in Statistical Parametric Speech Synthesis
abstract
The generated speech of hidden Markov model (HMM)-based statistical parametric speech synthesis still sounds “muffled.” One cause of this degradation in speech quality may be the loss of fine spectral structures. In this paper, we propose to use a deep generative architecture, a deep neural network (DNN) generatively trained, as a postfilter. The network models the conditional probability of the spectrum of natural speech given that of synthetic speech to compensate for such gap between synthetic and natural speech. The proposed probabilistic postfilter is generatively trained by cascading two restricted Boltzmann machines (RBMs) or deep belief networks (DBNs) with one bidirectional associative memory (BAM). We devised two types of DNN postfilters: one operating in the mel-cepstral domain and the other in the higher dimensional spectral domain. We compare these two new data-driven postfilters with other types of postfilters that are currently used in speech synthesis: a fixed mel-cepstral based postfilter, the global variance based parameter generation, and the modulation spectrum-based enhancement. Subjective evaluations using the synthetic voices of a male and female speaker confirmed that the proposed DNN-based postfilter in the spectral domain significantly improved the segmental quality of synthetic speech compared to that with conventional methods.
Linghui Chen, Tuomo Raitio, Cassia Valentini-Botinhao, Zhen-Hua Ling, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Toward a Universal Synthetic Speech Spoofing Detection Using Phase Information
abstract
In the field of speaker verification (SV) it is nowadays feasible and relatively easy to create a synthetic voice to deceive a speech driven biometric access system. This paper presents a synthetic speech detector that can be connected at the front-end or at the back-end of a standard SV system, and that will protect it from spoofing attacks coming from state-of-the-art statistical Text to Speech (TTS) systems. The system described is a Gaussian Mixture Model (GMM) based binary classifier that uses natural and copy-synthesized signals obtained from the Wall Street Journal database to train the system models. Three different state-of-the-art vocoders are chosen and modeled using two sets of acoustic parameters: 1) relative phase shift and 2) canonical Mel Frequency Cepstral Coefficients (MFCC) parameters, as baseline. The vocoder dependency of the system and multivocoder modeling features are thoroughly studied. Additional phase-aware vocoders are also tested. Several experiments are carried out, showing that the phase-based parameters perform better and are able to cope with new unknown attacks. The final evaluations, testing synthetic TTS signals obtained from the Blizzard challenge, validate our proposal.
Jon Sánchez, Ibon Saratxaga, Inma Hernáez Rioja, Eva Navas, Daniel Erro, Tuomo Raitio
IEEE Trans. Inf. Forensics Secur.6
2014 Parametric representation for singing voice synthesis: A comparative evaluation
abstract
Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be straightforward. The goal of this paper is twofold. First, a comparative subjective evaluation is performed across four existing techniques suitable for statistical parametric synthesis: traditional pulse vocoder, Deterministic plus Stochastic Model, Harmonic plus Noise Model and GlottHMM. The behavior of these techniques as a function of the singer type (baritone, counter-tenor and soprano) is studied. Secondly, the artifacts occurring in high-pitched voices are discussed and possible approaches to overcome them are suggested.
Onur Babacan, Thomas Drugman, Tuomo Raitio, Daniel Erro, Thierry Dutoit
ICASSP3
2014 A comparative evaluation of vocoding techniques for HMM-based laughter synthesis
abstract
This paper presents an experimental comparison of various leading vocoders for the application of HMM-based laughter synthesis. Four vocoders, commonly used in HMM-based speech synthesis, are used in copy-synthesis and HMM-based synthesis of both male and female laughter. Subjective evaluations are conducted to assess the performance of the vocoders. The results show that all vocoders perform relatively well in copy-synthesis. In HMM-based laughter synthesis using original phonetic transcriptions, all synthesized laughter voices were significantly lower in quality than copy-synthesis, indicating a challenging task and room for improvements. Interestingly, two vocoders using rather simple and robust excitation modeling performed the best, indicating that robustness in speech parameter extraction and simple parameter representation in statistical modeling are key factors in successful laughter synthesis.
Bajibabu Bollepalli, Jérôme Urbain, Tuomo Raitio, Joakim Gustafson, Hüseyin Çakmak
ICASSP3
2014 COVAREP - A collaborative voice analysis repository for speech technologies
abstract
Speech processing algorithms are often developed demonstrating improvements over the state-of-the-art, but sometimes at the cost of high complexity. This makes algorithm reimplementations based on literature difficult, and thus reliable comparisons between published results and current work are hard to achieve. This paper presents a new collaborative and freely available repository for speech processing algorithms called COVAREP, which aims at fast and easy access to new speech processing algorithms and thus facilitating research in the field. We envisage that COVAREP will allow more reproducible research by strengthening complex implementations through shared contributions and openly available code which can be discussed, commented on and corrected by the community. Presently COVAREP contains contributions from five distinct laboratories and we encourage contributions from across the speech processing research field. In this paper, we provide an overview of the current offerings of COVAREP and also include a demonstration of the algorithms through an emotion classification experiment.
Gilles Degottex, John Kane 0002, Thomas Drugman, Tuomo Raitio, Stefan Scherer
ICASSP4
2014 Excitation modeling for HMM-based speech synthesis: Breaking down the impact of periodic and aperiodic components
abstract
HMM-based speech synthesis generally suffers from typical buzzi-ness due to over-simplified excitation modeling of voiced speech. In order to alleviate this effect, several studies have proposed various new excitation models. No consensus has however been reached on what is the perceptual importance of the accurate modeling of the periodic and aperiodic components of voiced speech, and to what extent they separately contribute in improving naturalness. This paper considers a generalized mixed excitation modeling, common to various existing approaches, in which both periodic and aperiodic components coexist. At least three main factors may alter the quality of synthesis: periodic waveform, noise spectral weighting, and noise time envelope. Based on a large subjective evaluation, the goal of this paper is threefold: i) to evaluate the relative perceptual importance of each factor, ii) to investigate what is the most appropriate method to model the periodic and aperiodic components, and iii) to provide prospective clues for future work in excitation modeling.
Thomas Drugman, Tuomo Raitio
ICASSP2
2014 DNN-based stochastic postfilter for HMM-based speech synthesis
abstract
In this paper we propose a deep neural network to model the conditional probability of the spectral differences between nat-ural and synthetic speech. This allows us to reconstruct the spectral fine structures in speech generated by HMMs. We com-pared the new stochastic data-driven postfilter with global vari-ance based parameter generation and modulation spectrum en-hancement. Our results confirm that the proposed method sig-nificantly improves the segmental quality of synthetic speech compared to the conventional methods. Index Terms: HMM, speech synthesis, DNN, modulation spectrum, postfilter, segmental quality
Linghui Chen, Tuomo Raitio, Cassia Valentini-Botinhao, Junichi Yamagishi, Zhen-Hua Ling
INTERSPEECH2
2014 Investigating source and filter contributions, and their interaction, to statistical parametric speech synthesis
abstract
This paper presents an investigation of the separate perceptual degradations introduced by the modelling of source and fil-ter features in statistical parametric speech synthesis. This is achieved using stimuli in which various permutations of natu-ral, vocoded and modelled source and filter are combined, op-tionally with the addition of filter modifications (e.g. global variance or modulation spectrum scaling). We also examine the assumption of independence between source and filter pa-rameters. Two complementary perceptual testing paradigms are adopted. In the first, we ask listeners to perform “same or differ-ent quality ” judgements between pairs of stimuli from different configurations. In the second, we ask listeners to give an opin-ion score for individual stimuli. Combining the findings from these tests, we draw some conclusions regarding the relative contributions of source and filter to the currently rather limited naturalness of statistical parametric synthetic speech, and test whether current independence assumptions are justified. Index Terms: speech synthesis, hidden Markov modelling, GlottHMM, source filter model, source filter interaction
Thomas Merritt, Tuomo Raitio, Simon King 0001
INTERSPEECH2
2014 Deep neural network based trainable voice source model for synthesis of speech with varying vocal effort
abstract
This paper studies a deep neural network (DNN) based voice source modelling method in the synthesis of speech with varying vocal effort. The new trainable voice source model learns a mapping between the acoustic features and the time-domain pitch-synchronous glottal flow waveform using a DNN. The voice source model is trained with various speech material from breathy, normal, and Lombard speech. In synthesis, a normal voice is first adapted to a desired style, and using the flexible DNN-based voice source model, a style-specific excitation waveform is automatically generated based on the adapted acoustic features. The proposed voice source model is compared to a robust and high-quality excitation modelling method based on manually selected mean glottal flow pulses for each vocal effort level and using a spectral matching filter to correctly match the voice source spectrum to a desired style. Subjective evaluations show that the proposed DNN-based method is rated comparable to the baseline method, but avoids the manual selection of the pulses and is computationally faster than a system using a spectral matching filter.
Tuomo Raitio, Antti Suni, Lauri Juvela, Martti Vainio, Paavo Alku
INTERSPEECH1
2014 Automatic glottal inverse filtering with the Markov chain Monte Carlo method
Harri Auvinen, Tuomo Raitio, Manu Airaksinen, Samuli Siltanen, Brad H. Story, Paavo Alku
Comput. Speech Lang.2
2014 Synthesis and perception of breathy, normal, and Lombard speech in the presence of noise
Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku
Comput. Speech Lang.1
2014 Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction
abstract
This study presents a new glottal inverse filtering (GIF) technique based on closed phase analysis over multiple fundamental periods. The proposed quasi closed phase (QCP) analysis method utilizes weighted linear prediction (WLP) with a specific attenuated main excitation (AME) weight function that attenuates the contribution of the glottal source in the linear prediction model optimization. This enables the use of the autocorrelation criterion in linear prediction in contrast to the covariance criterion used in conventional closed phase analysis. The QCP method was compared to previously developed methods by using synthetic vowels produced with the conventional source-filter model as well as with a physical modeling approach. The obtained objective measures show that the QCP method improves the GIF performance in terms of errors in typical glottal source parametrizations for both low- and high-pitched vowels. Additionally, QCP was tested in a physiologically oriented vocoder, where the analysis/synthesis quality was evaluated with a subjective listening test indicating improved perceived quality for normal speaking style.
Manu Airaksinen, Tuomo Raitio, Brad H. Story, Paavo Alku
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Prediction of creaky voice from contextual factors
abstract
Creaky voice, also referred to as vocal fry, is a voice quality frequently produced in many languages, in both read and conversational speech. In order to enhance the naturalness of speech synthesisers, these latter should be able to generate speech in all its expressive diversity. This includes a proper use of creaky voice. The goal of this paper is two-fold. Firstly we analyse how contextual factors can be informative for the prediction of creaky use. It is observed that a few contextual factors related to speech production preceding a silence or a pause are of particular interest. This study validates that creaky voice plays a crucial syntactic role, allowing for a better structuring of phrases. In a second experiment, we investigate the prediction of creakiness from contextual factors based on HMMs. Four methods are compared on a US English and a Finnish speaker. It is shown that the best prediction technique achieves a promising performance comparable to what is carried out with the creaky detection algorithm on which HMMs were trained.
Thomas Drugman, John Kane 0002, Tuomo Raitio, Christer Gobl
ICASSP3
2013 Comparing glottal-flow-excited statistical parametric speech synthesis methods
abstract
This paper studies the performance of glottal flow signal based excitation methods in statistical parametric speech synthesis. The current state of the art in excitationmodeling is reviewed and three excitation methods are selected for experiments. Two of the methods are based on the principal component analysis (PCA) decomposition of estimated glottal flow pulses. While the first one uses only the mean of the pulses, the second method uses 12 principal components in addition to the mean signal for modeling the glottal flow waveform. The third method utilizes a glottal flow pulse library from which pulses are selected according to target and concatenation costs. Subjective listening tests are carried out to determine the quality and similarity of the synthetic speech of one male and one female speaker. The results show that the PCA-based methods are rated best both in quality and similarity, but adding more components does not yield any improvements.
Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku
ICASSP1
2013 Effect of MPEG audio compression on HMM-based speech synthesis
abstract
In this paper, the effect of MPEG audio compression on HMMbased speech synthesis is studied. Speech signals are encoded with various compression rates and analyzed using the GlottHMM vocoder. Objec ...
Bajibabu Bollepalli, Tuomo Raitio, Paavo Alku
INTERSPEECH2
2013 HMM-based synthesis of creaky voice
abstract
Creaky voice, also referred to as vocal fry, is a voice quality frequently produced in many languages, in both read and conversational speech. To enhance the naturalness of speech synthesis, these latter should be able to generate speech in all its expressive diversity, including creaky voice. The present study looks to exploit our recent developments, including creaky voice detection, prediction of creaky voice from context, and rendering of the creaky excitation, into a fully functioning and automatic HMM-based synthesis system. HMM-based synthetic creaky voices are built and evaluated in subjective listening tests, which show that the best synthetic creaky voices are rated more natural and more creaky compared to a conventional voice. A noncreaky voice is also successfully transformed to use creak by modifying the F0 contour and excitation of the predicted creaky parts. The transformed voice is rated equal in terms of naturalness and clearly more creaky compared to the original voice. Index Terms: speech synthesis, creaky voice, contextual factors, F0 estimation, excitation modeling
Tuomo Raitio, John Kane 0002, Thomas Drugman, Christer Gobl
INTERSPEECH1
2013 Analysis and synthesis of shouted speech
abstract
In this study, the acoustic properties of shouted speech are analyzed in relation to normal speech, and various synthesis techniques for shouting are investigated. The analysis shows large differences between the two styles, which induces difficulties in synthesis. Analysis-synthesis experiments show that the use of spectral estimation methods that are not biased by the sparse harmonics of shouted speech is beneficial. The synthesis of shouting is performed through adaptation and voice conversion. Subjective evaluation of synthesis reveals that, despite quality degradation, the impression of shouting and use of vocal effort is fairly well preserved. In addition, the use of specific spectral estimation methods is found to be beneficial also in adaptation.
Tuomo Raitio, Antti Suni, Jouni Pohjalainen, Manu Airaksinen, Martti Vainio, Paavo Alku
INTERSPEECH1
2013 Lombard modified text-to-speech synthesis for improved intelligibility: submission for the hurricane challenge 2013
abstract
This paper describes modification of a TTS system for im-proving the intelligibility of speech in various noise conditions. First, the GlottHMM vocoder is used for training a voice with modal speech data. The vocoder and voice parameters are then modified to mimic the properties of Lombard effect based on a small amount of Lombard speech from the same speaker. More specifically, the durations are increased, fundamental frequency is raised, spectral tilt is decreased, the harmonic-to-noise ratio is increased, and a pressed glottal flow pulses are used in cre-ating excitation. The formants of the speech are also enhanced and finally the speech is compressed in order to increase noise robustness of the voice. The evaluation results of the Hurricane Challenge 2013 indicate that the modified voice is mostly less intelligible than the unmodified natural speech, as expected, but more intelligible than the reference TTS voice, especially in the low SNR conditions.
Antti Suni, Reima Karhila, Tuomo Raitio, Mikko Kurimo, Martti Vainio, Paavo Alku
INTERSPEECH3
2012 On measuring the intelligibility of synthetic speech in noise - Do we need a realistic noise environment?
abstract
Assessing the intelligibility of synthetic speech is important in creating synthetic voices to be used in real life applications, especially for the ones involving interfering noise. This raises the question how to measure the intelligibility of synthetic speech to correctly simulate such conditions. Conventionally, this has been done using a simple listening test setup where diotic speech and noise are played to both ears with headphones. This is indeed very different from the real noise environment where speech and noise are spatially distributed. This paper addresses the question whether a realistic noise environment should be used to test the intelligibility of synthetic speech. Three different test conditions, one with multichannel reproduction of noise and speech, and two headphone setups are evaluated. Tests are performed with natural and synthetic speech, including speech especially intended for noisy conditions. The results indicate a general trend in all setups but also some interesting differences.
Tuomo Raitio, Marko Takanen, Olli Santala, Antti Suni, Martti Vainio, Paavo Alku
ICASSP1
2012 Utilizing Markov Chain Monte Carlo (MCMC) Method for Improved Glottal Inverse Filtering
abstract
This paper presents a new glottal inverse filtering (GIF) method that utilizes Markov chain Monte Carlo (MCMC) algorithm. First, initial estimates of the vocal tract and glottal flow are eval-uated by an existing GIF method, the iterative adaptive inverse filtering (IAIF). Simultaneously, the initially estimated glottal flow is synthesized using the Klatt model and filtered with the estimated vocal tract filter. In the MCMC estimation process, the first few poles of the initial vocal tract model and the Klatt parameter are refined in order to minimize the error between the original and the synthetic signals. MCMC converges to the optimal result, and the final estimate of the vocal tract is found by averaging the parameter values of the Markov chain. Ex-periments show that the MCMC-based GIF method gives more accurate results compared to the original IAIF method.
Harri Auvinen, Tuomo Raitio, Samuli Siltanen, Paavo Alku
INTERSPEECH2
2012 Towards Glottal Source Controllability in Expressive Speech Synthesis
abstract
In order to obtain more human like sounding humanmachine interfaces we must first be able to give them expressive capabilities in the way of emotional and stylistic features so as to closely adequate them to the intended task. If we want to replicate those features it is not enough to merely replicate the prosodic information of fundamental frequency and speaking rhythm. The proposed additional layer is the modification of the glottal model, for which we make use of the GlottHMM parameters. This paper analyzes the viability of such an approach by verifying that the expressive nuances are captured by the aforementioned features, obtaining 95% recognition rates on styled speaking and 82% on emotional speech. Then we evaluate the effect of speaker bias and recording environment on the source modeling in order to quantify possible problems when analyzing multi-speaker databases. Finally we propose a speaking styles separation for Spanish based on prosodic features and check its perceptual significance.
Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Tuomo Raitio, Nicolas Obin, Paavo Alku, Junichi Yamagishi, Juan Manuel Montero-Martínez
INTERSPEECH3
2012 Voice source analysis using biomechanical modeling and glottal inverse filtering
abstract
This paper studies the use of glottal inverse filtering together with a biomechanical model of the vocal folds to simulate the glottal flow waveform. The glottal flow waveform is first estimated by inverse filtering the acoustic speech pressure signal of natural speech. The estimated glottal flow is used as a template in an optimization process which searches for a set of parameters for a deterministic vocal fold model such that the model output reproduces the estimated glottal flow. The results indicate that the method can reproduce the main deterministic components of the glottal flow signal with good accuracy. Index Terms: vocal folds, glottal flow, biomechanical simulation, glottal inverse filtering.
Alan Pinheiro, Tuomo Raitio, Danyane Gomes, Paavo Alku
INTERSPEECH2
2012 Automatic Detection of High Vocal Effort in Telephone Speech
abstract
A system is proposed for the automatic detection of high vocal effort in speech. The system is evaluated using both PCMcoded speech and AMR-coded telephone speech. In addition, the effect of far-end noise in the telephone conditions is studied using both matched-condition training and cases with additive noise mismatch. The proposed system is based on Bayesian classification of mel-frequency cepstral feature vectors. Concerning the MFCC feature extraction process, the substitution of a spectrum analysis method emphasizing the fine structure improves the results in the noisy cases.
Jouni Pohjalainen, Tuomo Raitio, Hannu Pulakka, Paavo Alku
INTERSPEECH2
2012 Wideband Parametric Speech Synthesis Using Warped Linear Prediction
abstract
This paper studies the use of warped linear prediction (WLP) for wideband parametric speech synthesis. As the sampling fre-quency is increased from the usual 16 kHz, linear frequency res-olution of conventional linear prediction (LP) cannot efficiently model the speech spectrum. By using frequency warping that weights perceptually the most important formant information, spectral models with better accuracy and lower model orders can be utilized. In this work, WLP is embedded in a paramet-ric speech synthesizer to efficiently create wideband synthetic speech. Experiments show that WLP-based wideband synthetic speech is rated better compared to narrowband speech and wide-band LP-based speech. Index Terms: statistical parametric speech synthesis, wide-band, warped linear prediction, WLP
Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku
INTERSPEECH1
2012 Effect of noise type and level on focus related fundamental frequency changes
abstract
Speech in noise, or Lombard speech, is characterized by increased intensity and higher fundamental frequency as well as lengthened segmental durations as speakers try to maintain a beneficial signal-to-noise ratio to fill both communicative and self-monitoring requirements. The phenomenon has been studied with regard to different noise types and different noise levels, as well as with respect to different communicative tasks (e.g., reading out loud vs. speaking to a real listener). However, there are no studies where the effect has been measured with different noises keeping the loudness levels equal. Here we study the Lombard effect with three different noise types at three levels with equal loudness while varying focus structure to elicit different pitch contours. The results show that people adapt their intonation contours depending on both noise level and type even when the noises are similar with respect to their perceived loudness. This points to a special role for pitch in Lombard speech.
Martti Vainio, Daniel Aalto, Antti Suni, Anja Arnhold, Tuomo Raitio, Henri Seijo, Juhani Järvikivi, Paavo Alku
INTERSPEECH5
2011 Utilizing glottal source pulse library for generating improved excitation signal for HMM-based speech synthesis
abstract
This paper describes a source modeling method for hidden Markov model (HMM) based speech synthesis for improved naturalness. A speech corpus is first decomposed into the glottal source signal and the model of the vocal tract filter using glottal inverse filtering, and parametrized into excitation and spectral features. Additionally, a library of glottal source pulses is extracted from the estimated voice source signal. In the synthesis stage, the excitation signal is generated by selecting appropriate pulses from the library according to the target cost of the excitation features and a concatenation cost between adjacent glottal source pulses. Finally, speech is synthesized by filtering the excitation signal by the vocal tract filter. Experiments show that the naturalness of the synthetic speech is better or equal, and speaker similarity is better, compared to a system using only single glottal source pulse.
Tuomo Raitio, Antti Suni, Hannu Pulakka, Martti Vainio, Paavo Alku
ICASSP1
2011 Detection of Shouted Speech in the Presence of Ambient Noise
abstract
This study focuses on the detection of shouted speech in realistic noisy conditions. An automatic system based on modified mel frequency cepstral coefficient (MFCC) feature extraction and Gaussian mixture model (GMM) classification is developed. The performance of the automatic system is compared against human perception measured by a listening test. At moderate noise levels, the automatic system outperforms humans. In severe conditions, classification by humans is clearly better.
Jouni Pohjalainen, Tuomo Raitio, Paavo Alku
INTERSPEECH2
2011 Analysis of HMM-Based Lombard Speech Synthesis
abstract
Humans modify their voice in interfering noise in order to maintain the intelligibility of their speech – this is called the Lombard effect. This ability, however, has not been extensively modeled in speech synthesis. Here we compare several methods of synthesizing speech in noise using a physiologically based statistical speech synthesis system (GlottHMM). The results show that in a realistic street noise situation the synthetic Lombard speech is judged by listeners both as appropriate for the situation and as intelligible as natural Lombard speech. Of the different types of models, one using adaptation and extrapolation performed the best.
Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku
INTERSPEECH1
2011 HMM-Based Speech Synthesis Utilizing Glottal Inverse Filtering
abstract
This paper describes an hidden Markov model (HMM)-based speech synthesizer that utilizes glottal inverse filtering for generating natural sounding synthetic speech. In the proposed method, speech is first decomposed into the glottal source signal and the model of the vocal tract filter through glottal inverse filtering, and thus parametrized into excitation and spectral features. The source and filter features are modeled individually in the framework of HMM and generated in the synthesis stage according to the text input. The glottal excitation is synthesized through interpolating and concatenating natural glottal flow pulses, and the excitation signal is further modified according to the spectrum of the desired voice source characteristics. Speech is synthesized by filtering the reconstructed source signal with the vocal tract filter. Experiments show that the proposed system is capable of generating natural sounding speech, and the quality is clearly better compared to two HMM-based speech synthesis systems based on widely used vocoder techniques.
Tuomo Raitio, Antti Suni, Junichi Yamagishi, Hannu Pulakka, Jani Nurminen, Martti Vainio, Paavo Alku
IEEE Trans. Speech Audio Process.1
2009 New method for delexicalization and its application to prosodic tagging for text-to-speech synthesis
abstract
This paper describes a new flexible delexicalization method based on glottal excited parametric speech synthesis scheme. The system utilizes inverse filtered glottal flow and all-pole modelling of the vocal tract. The method provides a possibility to retain and manipulate all relevant prosodic features of any kind of speech. Most importantly, the features include voice quality, which has not been properly modeled in earlier delexicalization methods. The functionality of the new method was tested in a prosodic tagging experiment aimed at providing word prominence data for a text-to-speech synthesis system. The experiment confirmed the usefulness of the method and further corroborated earlier evidence that linguistic factors influence the perception of prosodic prominence.
Martti Vainio, Antti Suni, Tuomo Raitio, Jani Nurminen, Juhani Järvikivi, Paavo Alku
INTERSPEECH3
2008 HMM-based Finnish text-to-speech system utilizing glottal inverse filtering
abstract
Abstract This paper describes an HMM-based speech synthesis sys-tem that utilizes glottal inverse filtering for generating naturalsounding synthetic speech. In the proposed system, speech isfirst parametrized into spectral and excitation features using aglottal inverse filtering based method. The parameters are fedinto an HMM system for training and then generated from thetrained HMM according to text input. Glottal flow pulses ex-tracted from real speech are used as a voice source, and thevoice source is further modified according to the all-pole modelparameters generated by the HMM. Preliminary experimentsshow that the proposed system is capable of generating naturalsounding speech, and the quality is clearly better compared to asystem utilizing a conventional impulse train excitation model.Index Terms: speech synthesis, glottal inverse filtering, HMM 1. Introduction The ultimate goal of text-to-speech synthesis (TTS) is to enablecreating natural sounding speech from arbitrary text. More-over, the current trend in TTS research calls for systems thatenable producing speech in different speaking styles with dif-ferent speaker characteristics and even emotions. In order tofulfill these stringent general requirements, two major synthe-sis techniques have attracted increasing interest in the speechresearch community during the past decade. These two alter-natives are (1) the unit selection technique and (2) the hiddenMarkov model (HMM) based approach. The former has beenshown to yield synthetic speech of highly natural quality. How-ever, unit selection techniques do not allow for easy adaptationof the TTS system to different speaking styles and speaker char-acteristics. In addition, their implementation requires databasesof extensive sizes, which severely limit the use of this TTS tech-nique, for example, in mobile terminals. HMM-based tech-niques, in turn, benefit from better adaptability and a clearlysmaller memory requirement. However, the current HMM sys-tems often suffer from degraded naturalness in quality. It canbe argued that a potential reason for the reduced naturalness inthe current HMM-based TTS systems can be explained by theuse of signal generation techniques which are oversimplified toproperly mimic natural speech pressure waveforms.A large part of what can be characterized as naturalnessin speech emerges from different voice characteristics as wellas their context dependent changes. Therefore, it is justifiedin speech synthesis to search for methods aiming at accuratemodeling of different voice characteristics as well as prosodicfeatures of speech. Towards these goals, HMM-based synthe-sizers have been developed with special emphasis on voice char-acteristics such as speaker individualities, speaking styles, andemotions [1]. Moreover, some recent studies have introducedimprovements to the parametric HMM systems’ signal genera-tion techniques by utilizing, for example, mixed excitation [2]and residual modeling [3]. These techniques have been shownto improve the quality of synthetic speech compared to systemsutilizing a traditional impulse train excitation model. However,the quality of the systems using these techniques still remainsfar from the quality of natural speech.In the real human voice production mechanism, the excita-tion of (voiced) speech is represented by the glottal volume ve-locity waveform generated by the vibrating vocal folds. This ex-citation signal, the glottal source, has naturally attracted interestin speech synthesis and many techniques have been proposed tomimic the glottal source of natural speech. One such techniqueis the Liljencrants-Fant (LF) model of the differentiated glottalsource that has been used both in traditional rule-based synthe-sis [4, 5] as well as within an HMM-based speech synthesizer[6]. However, the use of artificial glottal flow pulses usuallyresults in a somewhat buzzy quality due to a strong harmonicstructure at higher frequencies. To overcome this problem, theidea of utilizing glottal flow pulses extracted from real speechwith the help of glottal inverse filtering has been proposed [7, 8].However, previous studies based on glottal flow pulses extractedfrom natural speech are limited to special purposes such as thegeneration of isolated vowels, and the benefits from combiningautomatic glottal inverse filtering with an HMM-based speechsynthesizer have not been utilized.In this paper, a novel HMM-based speech synthesis sys-tem that utilizes glottal inverse filtering for generating naturalsounding synthetic speech is presented. The rest of the paper isorganized as follows: Section 2 describes the proposed speechsynthesis system. The results of the experiments with the newsynthesizer are presented in Section 3. Discussion on the pro-posed speech synthesis system and future plans are presented inSection 4, and final conclusions are presented in Section 5.
Tuomo Raitio, Antti Suni, Hannu Pulakka, Martti Vainio, Paavo Alku
INTERSPEECH1