VLDB 2026 Research / reviewers in the wild / expert
Roberto Barra-Chicote
dblp:93/2127
· DBLP profile ↗
51ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0003-0844-7037ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 34 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Universal Semantic Disentangled Privacy-preserving Speech Representation LearningabstractThe use of human speech to train LLMs poses privacy concerns due to these models' ability to generate samples that closely resemble artifacts in the training data. We propose a speaker privacy-preserving representation learning method through the Universal Speech Codec (USC), a computationally efficient codec that disentangles speech into: (i) privacy-preserving semantically rich representations, capturing content and speech paralinguistics, and (ii) residual acoustic and speaker representations that enable high-fidelity reconstruction. Evaluations show that USC's semantic representation preserves content, prosody, and sentiment, while removing identifiable traits. Additionally, we present an evaluation methodology for measuring privacy-preserving properties. We compare USC against other speech codecs and demonstrate its effectiveness on privacy-preserving representation learning, showcasing the trade-offs between speaker anonymization and paralinguistics retention.1 Biel Tura Vecino, Subhadeep Maji, Aravind Varier, Antonio Bonafonte, Ivan Valles, Michael Owen, Constantinos Papayiannis, Leif Rädel, Grant P. Strimel, Oluwaseyi Feyisetan, Roberto Barra-Chicote, Ariya Rastrow, Volker Leutnant, Trevor Wood |
INTERSPEECH | 11 |
| 2023 | SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic SpacesabstractNumerous examples in the literature proved that deep learning models have the ability to work well with multimodal data. Recently, CLIP has enabled deep learning systems to learn shared latent spaces between images and text descriptions, with outstanding zero- or few-shot results in downstream tasks. In this paper we explore the same idea proposed by CLIP but applied to the speech domain, where the phonetic and acoustic spaces usually coexist. We train a CLIP-based model with the aim to learn shared representations of phonetic and acoustic spaces. The results show that the proposed model is sensible to phonetic changes, with a 91% of score drops when replacing 20% of the phonemes at random, while providing substantial robustness against different kinds of noise, with a 10% performance drop when mixing the audio with 75% of Gaussian noise. We also provide empirical evidence showing that the resulting embeddings are useful for a variety of downstream applications, such as intelligibility evaluation and the ability to leverage rich pre-trained phonetic embeddings in speech generation task. Finally, we discuss potential applications with interesting implications for the speech generation and recognition fields. Iván Vallés-Pérez, Grzegorz Beringer, Piotr Bilinski, Gary Cook, Roberto Barra-Chicote |
ECAI | 5 |
| 2023 | Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech
Guangyan Zhang, Thomas Merritt, Manuel Sam Ribeiro, Biel Tura Vecino, Kayoko Yanagisawa, Kamil Pokora, Abdelhamid Ezzerg, Sebastian Cygert, Ammar Abbas, Piotr Bilinski, Roberto Barra-Chicote, Daniel Korzekwa, Jaime Lorenzo-Trueba |
INTERSPEECH | 11 |
| 2022 | Duration Modeling of Neural TTS for Automatic DubbingabstractAutomatic dubbing (AD) addresses the problem of translating speech in a video with speech in another language while preserving the viewer experience. A most important requirement of AD is isochrony, i.e. dubbed speech has to closely match the timing of speech and pauses of the original audio. In our automatic dubbing system, isochrony is modeled by controlling the verbosity of machine translation; inserting pauses in the translations, a.k.a. prosodic alignment; and controlling the duration of text-to-speech (TTS) utterances. The latter two steps heavily rely on speech duration information, either to predict or control TTS duration. So far, duration prediction was based on a proxy method while duration control on linear warping of the TTS speech spectrogram. In this study, we propose novel duration models for neural TTS that can be leveraged both to predict and control TTS duration. Experimental results show that compared to previous work, the new models improve or match the performance of prosodic alignment and significantly enhance neural TTS speech quality for both slow and fast speaking rates. Johanes Effendi, Yogesh Virkar, Roberto Barra-Chicote, Marcello Federico |
ICASSP | 3 |
| 2022 | Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing ModuleabstractState-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and intelligibility degradations, making training low-resource TTS systems problematic. In this paper, we propose a novel extremely low-resource TTS method called Voice Filter that uses as little as one minute of speech from a target speaker. It uses voice conversion (VC) as a post-processing module appended to a pre-existing high-quality TTS system and marks a conceptual shift in the existing TTS paradigm, framing the few-shot TTS problem as a VC task. Furthermore, we propose to use a duration-controllable TTS system to create a parallel speech corpus to facilitate the VC task. Results show that the Voice Filter outperforms state-of-the-art few-shot speech synthesis techniques in terms of objective and subjective metrics on one minute of speech on a diverse set of voices, while being competitive against a TTS model built on 30 times more data.1 Adam Gabrys, Goeric Huybrechts, Manuel Sam Ribeiro, Chung-Ming Chien, Julian Roth, Giulia Comini, Roberto Barra-Chicote, Bartek Perz, Jaime Lorenzo-Trueba |
ICASSP | 7 |
| 2022 | Text-Free Non-Parallel Many-To-Many Voice Conversion Using Normalising FlowabstractNon-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all other information from the source speech is retained is a large challenge. This is particularly challenging in the scenario where at inference-time we have no knowledge of the text being read, i.e., text-free VC. To mitigate this, we investigate information-preserving VC approaches.Normalising flows have gained attention for text-to-speech synthesis, however have been under-explored for VC. Flows utilize invertible functions to learn the likelihood of the data, thus provide a lossless encoding of speech. We investigate normalising flows for VC in both text-conditioned and text-free scenarios. Furthermore, for text-free VC we compare pre-trained and jointly-learnt priors. Flow-based VC evaluations show no degradation between text-free and text-conditioned VC, resulting in improvements over the state-of-the-art. Also, joint-training of the prior is found to negatively impact text-free VC quality. Thomas Merritt, Abdelhamid Ezzerg, Piotr Bilinski, Magdalena Proszewska, Kamil Pokora, Roberto Barra-Chicote, Daniel Korzekwa |
ICASSP | 6 |
| 2022 | Creating New Voices using Normalizing FlowsabstractCreating realistic and natural-sounding synthetic speech remains a big challenge for voice identities unseen during training. As there is growing interest in synthesizing voices of new speakers, here we investigate the ability of normalizing flows in text-to-speech (TTS) and voice conversion (VC) modes to extrapolate from speakers observed during training to create unseen speaker identities. Firstly, we create an approach for TTS and VC, and then we comprehensively evaluate our methods and baselines in terms of intelligibility, naturalness, speaker similarity, and ability to create new voices. We use both objective and subjective metrics to benchmark our techniques on 2 evaluation tasks: zero-shot and new voice speech synthesis. The goal of the former task is to measure the precision of the conversion to an unseen voice. The goal of the latter is to measure the ability to create new voices. Extensive evaluations demonstrate that the proposed approach systematically allows to obtain state-of-the-art performance in zero-shot speech synthesis and creates various new voices, unobserved in the training set. We consider this work to be the first attempt to synthesize new voices based on mel-spectrograms and normalizing flows, along with a comprehensive analysis and comparison of the TTS and VC modes. Piotr Bilinski, Thomas Merritt, Abdelhamid Ezzerg, Kamil Pokora, Sebastian Cygert, Kayoko Yanagisawa, Roberto Barra-Chicote, Daniel Korzekwa |
INTERSPEECH | 7 |
| 2022 | GlowVC: Mel-spectrogram space disentangling model for language-independent text-free voice conversionabstractIn this paper, we propose GlowVC: a multilingual multispeaker flow-based model for language-independent text-free voice conversion.We build on Glow-TTS, which provides an architecture that enables use of linguistic features during training without the necessity of using them for VC inference.We consider two versions of our model: GlowVC-conditional and GlowVC-explicit.GlowVC-conditional models the distribution of mel-spectrograms with speaker-conditioned flow and disentangles the mel-spectrogram space into content-and pitchrelevant dimensions, while GlowVC-explicit models the explicit distribution with unconditioned flow and disentangles said space into content-, pitch-and speaker-relevant dimensions.We evaluate our models in terms of intelligibility, speaker similarity and naturalness for intra-and cross-lingual conversion in seen and unseen languages.GlowVC models greatly outperform Au-toVC baseline in terms of intelligibility, while achieving just as high speaker similarity in intra-lingual VC, and slightly worse in the cross-lingual setting.Moreover, we demonstrate that GlowVC-explicit surpasses both GlowVC-conditional and Au-toVC in terms of naturalness. Magdalena Proszewska, Grzegorz Beringer, Daniel Saez-Trigueros, Thomas Merritt, Abdelhamid Ezzerg, Roberto Barra-Chicote |
INTERSPEECH | 6 |
| 2022 | Prosodic alignment for off-screen automatic dubbingabstractThe goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence.This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into phrases and pauses, especially when the speaker's mouth is visible.In previous work, we introduced a prosodic alignment model to address isochrone or on-screen dubbing.In this work, we extend the prosodic alignment model to also address off-screen dubbing that requires less stringent synchronization constraints.We conduct experiments on four dubbing directions -English to French, Italian, German and Spanish -on a publicly available collection of TED Talks and on publicly available YouTube videos.Empirical results show that compared to our previous work the extended prosodic alignment model provides significantly better subjective viewing experience on videos in which on-screen and off-screen automatic dubbing is applied for sentences with speakers mouth visible and not visible, respectively. Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote |
INTERSPEECH | 4 |
| 2022 | Remap, Warp and Attend: Non-Parallel Many-to-Many Accent Conversion with Normalizing FlowsabstractRegional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intonation. This paper investigates a novel flow-based approach to accent conversion using normalizing flows. The proposed approach revolves around three steps: remapping the phonetic conditioning, to better match the target accent, warping the duration of the converted speech, to better suit the target phonemes, and an attention mechanism that implicitly aligns source and target speech sequences. The proposed remap-warp-attend system enables adaptation of both phonetic and prosodic aspects of speech while allowing for source and converted speech signals to be of different lengths. Objective and subjective evaluations show that the proposed approach significantly outperforms a competitive CopyCat baseline model in terms of similarity to the target accent, naturalness and intelligibility. Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa, Piotr Bilinski, Magdalena Proszewska, Kamil Pokora, Renard Korzeniowski, Roberto Barra-Chicote, Daniel Korzekwa |
SLT | 8 |
| 2021 | Machine Translation Verbosity Control for Automatic DubbingabstractAutomatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterances. In this paper, we focus on the problem of controlling the verbosity of machine translation out-put, so that subsequent steps of our automatic dubbing pipeline can generate dubs of better quality. We propose new methods to control the verbosity of MT output and compare them against the state of the art with both intrinsic and extrinsic evaluations. For our experiments we use a public data set to dub English speeches into French, Italian, German and Spanish. Finally, we report extensive subjective tests that measure the impact of MT verbosity control on the final quality of dubbed video clips. Surafel Melaku Lakew, Marcello Federico, Yue Wang 0034, Cuong Hoang, Yogesh Virkar, Roberto Barra-Chicote, Robert Enyedi |
ICASSP | 6 |
| 2021 | Improvements to Prosodic Alignment for Automatic DubbingabstractAutomatic dubbing is an extension of speech-to-speech translation such that the resulting target speech is carefully aligned in terms of duration, lip movements, timbre, emotion, prosody, etc. of the speaker in order to achieve audiovisual coherence. Dubbing quality strongly depends on isochrony, i.e., arranging the translation of the original speech to optimally match its sequence of phrases and pauses. To this end, we present improvements to the prosodic alignment component of our recently introduced dubbing architecture. We present empirical results for four dubbing directions – English to French, Italian, German and Spanish – on a publicly available collection of TED Talks. Compared to previous work, our enhanced prosodic alignment model significantly improves prosodic alignment accuracy and provides segmentation perceptibly better or on par with manually annotated reference segmentation. Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote |
ICASSP | 4 |
| 2021 | Exploring the application of synthetic audio in training keyword spottersabstractThe study of keyword spotting, a subfield within the broader field of speech recognition that centers around identifying individual keywords in speech audio, has gained particular importance in recent years with the rise of personal voice assistants such as Alexa. As voice assistants aim to rapidly expand to support new languages, keywords, and use cases, stakeholders face the issue of limited training data for these unseen scenarios. This paper details some initial exploration into the application of Text-To-Speech (TTS) audio as a "helper" tool for training keyword spotters in these low-resource scenarios. In the experiments studied in this paper, the careful mixing of TTS audio with human speech audio during training led to a reduction of over 11% in the detection-error-tradeoff (DET) area under the curve (AUC) metric. Andrew Werchniak, Roberto Barra-Chicote, Yuriy Mishchenko, Jasha Droppo, Jeff Condal, Anish Shah |
ICASSP | 2 |
| 2021 | SynthASR: Unlocking Synthetic Data for Speech RecognitionabstractEnd-to-end (E2E) automatic speech recognition (ASR) models have recently demonstrated superior performance over the traditional hybrid ASR models. Training an E2E ASR model requires a large amount of data which is not only expensive but may also raise dependency on production data. At the same time, synthetic speech generated by the state-of-the-art text-to-speech (TTS) engines has advanced to near-human naturalness. In this work, we propose to utilize synthetic speech for ASR training (SynthASR) in applications where data is sparse or hard to get for ASR model training. In addition, we apply continual learning with a novel multi-stage training strategy to address catastrophic forgetting, achieved by a mix of weighted multi-style training, data augmentation, encoder freezing, and parameter regularization. In our experiments conducted on in-house datasets for a new application of recognizing medication names, training ASR RNN-T models with synthetic audio via the proposed multi-stage training improved the recognition performance on new application by more than 65% relative, without degradation on existing general applications. Our observations show that SynthASR holds great promise in training the state-of-the-art large-scale E2E ASR models for new applications while reducing the costs and dependency on production data. Amin Fazel, Yulan Liu, Roberto Barra-Chicote, Yixiong Meng, Roland Maas, Jasha Droppo |
Interspeech | 4 |
| 2021 | Improving the Expressiveness of Neural Vocoding with Non-Affine Normalizing FlowsabstractThis paper proposes a general enhancement to the Normalizing Flows (NF) used in neural vocoding. As a case study, we improve expressive speech vocoding with a revamped Parallel Wavenet (PW). Specifically, we propose to extend the affine transformation of PW to the more expressive invertible non-affine function. The greater expressiveness of the improved PW leads to better-perceived signal quality and naturalness in the waveform reconstruction and text-to-speech (TTS) tasks. We evaluate the model across different speaking styles on a multi-speaker, multi-lingual dataset. In the waveform reconstruction task, the proposed model closes the naturalness and signal quality gap from the original PW to recordings by $10\%$, and from other state-of-the-art neural vocoding systems by more than $60\%$. We also demonstrate improvements in objective metrics on the evaluation test set with L2 Spectral Distance and Cross-Entropy reduced by $3\%$ and $6\unicode{x2030}$ comparing to the affine PW. Furthermore, we extend the probability density distillation procedure proposed by the original PW paper, so that it works with any non-affine invertible and differentiable function. Adam Gabrys, Yunlong Jiao, Viacheslav Klimkov, Daniel Korzekwa, Roberto Barra-Chicote |
Interspeech | 5 |
| 2021 | Detection of Lexical Stress Errors in Non-Native (L2) English with Data Augmentation and AttentionabstractThis paper describes two novel complementary techniques that improve the detection of lexical stress errors in non-native (L2) English speech: attention-based feature extraction and data augmentation based on Neural Text-To-Speech (TTS). In a classical approach, audio features are usually extracted from fixed regions of speech such as the syllable nucleus. We propose an attention-based deep learning model that automatically derives optimal syllable-level representation from frame-level and phoneme-level audio features. Training this model is challenging because of the limited amount of incorrect stress patterns. To solve this problem, we propose to augment the training set with incorrectly stressed words generated with Neural TTS. Combining both techniques achieves 94.8% precision and 49.2% recall for the detection of incorrectly stressed words in L2 English speech of Slavic and Baltic speakers. Daniel Korzekwa, Roberto Barra-Chicote, Szymon Zaporowski, Grzegorz Beringer, Jaime Lorenzo-Trueba, Alicja Serafinowicz, Jasha Droppo, Thomas Drugman, Bozena Kostek |
Interspeech | 2 |
| 2021 | Intra-Sentential Speaking Rate Control in Neural Text-To-Speech for Automatic Dubbing
Yogesh Virkar, Marcello Federico, Roberto Barra-Chicote, Robert Enyedi |
Interspeech | 4 |
| 2021 | Improving Multi-Speaker TTS Prosody Variance with a Residual Encoder and Normalizing FlowsabstractText-to-speech systems recently achieved almost indistinguishable quality from human speech. However, the prosody of those systems is generally flatter than natural speech, producing samples with low expressiveness. Disentanglement of speaker id and prosody is crucial in text-to-speech systems to improve on naturalness and produce more variable syntheses. This paper proposes a new neural text-to-speech model that approaches the disentanglement problem by conditioning a Tacotron2-like architecture on flow-normalized speaker embeddings, and by substituting the reference encoder with a new learned latent distribution responsible for modeling the intra-sentence variability due to the prosody. By removing the reference encoder dependency, the speaker-leakage problem typically happening in this kind of systems disappears, producing more distinctive syntheses at inference time. The new model achieves significantly higher prosody variance than the baseline in a set of quantitative prosody features, as well as higher speaker distinctiveness, without decreasing the speaker intelligibility. Finally, we observe that the normalized speaker embeddings enable much richer speaker interpolations, substantially improving the distinctiveness of the new interpolated speakers. Iván Vallés-Pérez, Julian Roth, Grzegorz Beringer, Roberto Barra-Chicote, Jasha Droppo |
Interspeech | 4 |
| 2020 | Using Vaes and Normalizing Flows for One-Shot Text-To-Speech Synthesis of Expressive SpeechabstractWe propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence based system with a Variational AutoEncoder (VAE) and a Householder Flow. The proposed system provides a 22% KL-divergence reduction while jointly improving perceptual metrics over state-of-the-art. At synthesis time we use one example of expressive style as a reference input to the encoder for generating any text in the desired style. Perceptual MUSHRA evaluations show that we can create a voice with a 9% relative naturalness improvement over standard Neural Text-to-Speech, while also improving the perceived emotional intensity (59 compared to the 55 of neutral speech). Vatsal Aggarwal, Marius Cotescu, Nishant Prateek, Jaime Lorenzo-Trueba, Roberto Barra-Chicote |
ICASSP | 5 |
| 2020 | BOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian OptimizationabstractWe present BOFFIN TTS (Bayesian Optimization For FIne-tuning Neural Text To Speech), a novel approach for few-shot speaker adaptation. Here, the task is to fine-tune a pre-trained TTS model to mimic a new speaker using a small corpus of target utterances. We demonstrate that there does not exist a one-size-fits-all adaptation strategy, with convincing synthesis requiring a corpus-specific configuration of the hyper-parameters that control fine-tuning. By using Bayesian optimization to efficiently optimize these hyper-parameter values for a target speaker, we are able to perform adaptation with an average 30% improvement in speaker similarity over standard techniques. Results indicate, across multiple corpora, that BOFFIN TTS can learn to synthesize new speakers using less than ten minutes of audio, achieving the same naturalness as produced for the speakers used to train the base model. Henry B. Moss, Vatsal Aggarwal, Nishant Prateek, Javier González 0002, Roberto Barra-Chicote |
ICASSP | 5 |
| 2020 | Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing
Marcello Federico, Yogesh Virkar, Robert Enyedi, Roberto Barra-Chicote |
INTERSPEECH | 4 |
| 2019 | Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric SpeechabstractThis paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and encoding of the input text. We showed that the latent space conveys interpretable characteristics of dysarthria, such as intelligibility and fluency of speech. MUSHRA perceptual test demonstrated that the adaptation of the latent space let the model generate speech of improved fluency. The multi-task supervised approach for predicting both the probability of dysarthric speech and the mel-spectrogram helps improve the detection of dysarthria with higher accuracy. This is thanks to a low-dimensional latent space of the auto-encoder as opposed to directly predicting dysarthria from a highly dimensional mel-spectrogram. Daniel Korzekwa, Roberto Barra-Chicote, Bozena Kostek, Thomas Drugman, Mateusz Lajszczak |
INTERSPEECH | 2 |
| 2019 | Towards Achieving Robust Universal Neural VocodingabstractThis paper explores the potential universality of neural vocoders.We train a WaveRNN-based vocoder on 74 speakers coming from 17 languages.This vocoder is shown to be capable of generating speech of consistently good quality (98% relative mean MUSHRA when compared to natural speech) regardless of whether the input spectrogram comes from a speaker or style seen during training or from an out-of-domain scenario when the recording conditions are studio-quality.When the recordings show significant changes in quality, or when moving towards non-speech vocalizations or singing, the vocoder still significantly outperforms speaker-dependent vocoders, but operates at a lower average relative MUSHRA of 75%.These results are shown to be consistent across languages, regardless of them being seen during training (e.g.English or Japanese) or unseen (e.g. Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, Vatsal Aggarwal |
INTERSPEECH | 6 |
| 2018 | Comprehensive Evaluation of Statistical Speech Waveform SynthesisabstractStatistical TTS systems that directly predict the speech waveform have recently reported improvements in synthesis quality. This investigation evaluates Amazon's statistical speech waveform synthesis (SSWS) system. An in-depth evaluation of SSWS is conducted across a number of domains to better understand the consistency in quality. The results of this evaluation are validated by repeating the procedure on a separate group of testers. Finally, an analysis of the nature of speech errors of SSWS compared to hybrid unit selection synthesis is conducted to identify the strengths and weaknesses of SSWS. Having a deeper insight into SSWS allows us to better define the focus of future work to improve this new technology. Thomas Merritt, Bartosz Putrycz, Adam Nadolski, Tianjun Ye, Daniel Korzekwa, Wiktor Dolecki, Thomas Drugman, Viacheslav Klimkov, Alexis Moinet, Andrew P. Breen, Rafal Kuklinski, Nikko Strom, Roberto Barra-Chicote |
SLT | 13 |
| 2017 | Phrase Break Prediction for Long-Form Reading TTS: Exploiting Text Structure Information
Viacheslav Klimkov, Adam Nadolski, Alexis Moinet, Bartosz Putrycz, Roberto Barra-Chicote, Thomas Merritt, Thomas Drugman |
INTERSPEECH | 5 |
| 2016 | Continuous Expressive Speaking Styles Synthesis based on CVSM and MR-HMMabstractThis paper introduces a continuous system capable of automatically producing the most adequate speaking style to synthesize a desired target text. This is done thanks to a joint modeling of the acoustic and lexical parameters of the speaker models by adapting the CVSM projection of the training texts using MR-HMM techniques. As such, we consider that as long as sufficient variety in the training data is available, we should be able to model a continuous lexical space into a continuous acoustic space. The proposed continuous automatic text to speech system was evaluated by means of a perceptual evaluation in order to compare them with traditional approaches to the task. The system proved to be capable of conveying the correct expressiveness (average adequacy of 3.6) with an expressive strength comparable to oracle traditional expressive speech synthesis (average of 3.6) although with a drop in speech quality mainly due to the semi-continuous nature of the data (average quality of 2.9). This means that the proposed system is capable of improving traditional neutral systems without requiring any additional user interaction. Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Ascensión Gallardo-Antolín, Junichi Yamagishi, Juan Manuel Montero-Martínez |
COLING | 2 |
| 2016 | Feature extraction from smartphone inertial signals for human activity segmentation
Rubén San-Segundo-Hernández, Juan Manuel Montero-Martínez, Roberto Barra-Chicote, Fernando Fernández Martínez, José Manuel Pardo |
Signal Process. | 3 |
| 2015 | Knowledge versus data in TTS: evaluation of a continuum of synthesis systemsabstractGrapheme-based models have been proposed for both ASR and TTS as a way of circumventing the lack of expert-compiled pronunciation lexicons in under-resourced languages. It is a common observation that this should work well in languages employing orthographies with a transparent letter-to-phoneme relationship,such as Spanish. Our experience has shown, however,that there is still a significant difference in intelligibility between grapheme-based systems and conventional ones for this language. This paper explores the contribution of different levels of linguistic annotation to system intelligibility, and the trade-off between those levels and the quantity of data used for training. Ten systems spaced across these two continua of knowledge and data were subjectively evaluated for intelligibility. Rosie Kay, Oliver Watts, Roberto Barra-Chicote, Cassie Mayo |
INTERSPEECH | 3 |
| 2015 | Emotion transplantation through adaptation in HMM-based speech synthesis
Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Rubén San-Segundo-Hernández, Javier Ferreiros, Junichi Yamagishi, Juan Manuel Montero-Martínez |
Comput. Speech Lang. | 2 |
| 2014 | Generating segmental foreign accentabstractFor most of us, speaking in a non-native language involves deviating to some extent from native pronunciation norms. However, the detailed basis for foreign accent (FA) remains elusive, in part due to methodological challenges in isolating segmental from suprasegmental factors. The current study examines the role of segmental features in conveying FA through the use of a generative approach in which accent is localised to single consonantal segments. Three techniques are evaluated: the first requires a highly-proficiency bilingual to produce words with isolated accented segments; the second uses cross-splicing of context-dependent consonants from the non-native language into native words; the third employs hidden Markov model synthesis to blend voice models for both languages. Using English and Spanish as the native/non-native languages respectively, listener cohorts from both languages identified words and rated their degree of FA. All techniques were capable of generating accented words, but to differing degrees. Naturally-produced speech led to the strongest FA ratings and synthetic speech the weakest, which we interpret as the outcome of over-smoothing. Nevertheless, the flexibility offered by synthesising localised accent encourages further development of the method. María Luisa García Lecumberri, Roberto Barra-Chicote, Rubén Pérez Ramón, Junichi Yamagishi, Martin Cooke |
INTERSPEECH | 2 |
| 2014 | Translating bus information into sign language for deaf people
Verónica López-Ludeña, Carlos González-Morcillo, Juan Carlos López 0001, Roberto Barra-Chicote, Ricardo de Córdoba, Rubén San-Segundo-Hernández |
Eng. Appl. Artif. Intell. | 4 |
| 2013 | NEMOHIFI: an affective HiFi agentabstractThis demo concerns a recently developed prototype of an emotionally-sensitive autonomous HiFi Spoken Conversational Agent, called NEMOHIFI. The baseline agent was developed by the Speech Technology Group (GTH) and has recently been integrated with an emotional engine called NEMO (Need-inspired Emotional Model) to enable it to adapt to users' emotion and respond to the users using appropriate expressive speech. NEMOHIFI controls and manages the HiFi audio system, and for end users, its functions equate a remote control, except that instead of clicking, the user interacts with the agent using voice. A pairwise comparison between the baseline (non-adaptive) and NEMO-HIFI showed that the latter was not only statistically substantially preferred by users to the former, but they are also significantly more satisfied with it than the former. Syaheerah L. Lutfi, Fernando Fernández Martínez, Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Juan Manuel Montero-Martínez |
ICMI | 4 |
| 2013 | LSESpeak: A spoken language generator for Deaf people
Verónica López-Ludeña, Roberto Barra-Chicote, Syaheerah L. Lutfi, Juan Manuel Montero-Martínez, Rubén San-Segundo-Hernández |
Expert Syst. Appl. | 2 |
| 2012 | Towards Glottal Source Controllability in Expressive Speech SynthesisabstractIn order to obtain more human like sounding humanmachine interfaces we must first be able to give them expressive capabilities in the way of emotional and stylistic features so as to closely adequate them to the intended task. If we want to replicate those features it is not enough to merely replicate the prosodic information of fundamental frequency and speaking rhythm. The proposed additional layer is the modification of the glottal model, for which we make use of the GlottHMM parameters. This paper analyzes the viability of such an approach by verifying that the expressive nuances are captured by the aforementioned features, obtaining 95% recognition rates on styled speaking and 82% on emotional speech. Then we evaluate the effect of speaker bias and recording environment on the source modeling in order to quantify possible problems when analyzing multi-speaker databases. Finally we propose a speaking styles separation for Spanish based on prosodic features and check its perceptual significance. Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Tuomo Raitio, Nicolas Obin, Paavo Alku, Junichi Yamagishi, Juan Manuel Montero-Martínez |
INTERSPEECH | 2 |
| 2012 | Towards an Unsupervised Speaking Style Voice Building Framework: Multi-Style Speaker DiarizationabstractCurrent text-to-speech systems are developed using studio-recorded speech in a neutral style or based on acted emotions. However, the proliferation of media sharing sites would allow developing a new generation of speech-based systems which could cope with spontaneous and styled speech. This paper proposes an architecture to deal with realistic recordings and carries out some experiments on unsupervised speaker diarization. In order to maximize the speaker purity of the clusters while keeping a high speaker coverage, the paper evaluates the F-measure of a diarization module, achieving high scores (>85%) especially when the clusters are longer than 30 seconds, even for the more spontaneous and expressive styles (such as talk shows or sports). Jaime Lorenzo-Trueba, Beatriz Martínez-González, Roberto Barra-Chicote, Verónica López-Ludeña, Javier Ferreiros, Junichi Yamagishi, Juan Manuel Montero-Martínez |
INTERSPEECH | 3 |
| 2012 | Selection of TDOA Parameters for MDM Speaker DiarizationabstractSeveral methods to improve multiple distant microphone (MDM) speaker diarization based on Time Delay of Arrival (TDOA) features are evaluated in this paper. All of them avoid the use of a single reference channel to calculate the TDOA values and, based on different criteria, select among all possible pairs of microphones a set of pairs that will be used to estimate the TDOA's. The evaluated methods have been named the "Dynamic Margin" (DM), the "Extreme Regions" (ER), the "Most Common" (MC), the "Cross Correlation" (XCorr) and the "Principle Component Analysis" (PCA). It is shown that all methods improve the baseline results for the development set and four of them improve also the results for the evaluation set. Improvements of 3.49% and 10.77% DER relative are obtained for DM and ER respectively for the test set. The XCorr and PCA methods achieve an improvement of 36.72% and 30.82% DER relative for the test set. Moreover, the computational cost for the XCorr method is 20% less than the baseline. Beatriz Martínez-González, José Manuel Pardo, Julián D. Echeverry-Correa, José A. Vallejo, Roberto Barra-Chicote |
INTERSPEECH | 5 |
| 2012 | Speaker Diarization Features: The UPM Contribution to the RT09 EvaluationabstractTwo new features have been proposed and used in the Rich Transcription Evaluation 2009 by the Universidad Politécnica de Madrid, which outperform the results of the baseline system. One of the features is the intensity channel contribution, a feature related to the location of the speaker. The second feature is the logarithm of the interpolated fundamental frequency. It is the first time that both features are applied to the clustering stage of multiple distant microphone meetings diarization. It is shown that the inclusion of both features improves the baseline results by 15.36% and 16.71% relative to the development set and the RT 09 set, respectively. If we consider speaker errors only, the relative improvement is 23% and 32.83% on the development set and the RT09 set, respectively. José Manuel Pardo, Roberto Barra-Chicote, Rubén San-Segundo-Hernández, Ricardo de Córdoba, Beatriz Martínez-González |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Speaker Diarization Based on Intensity Channel ContributionabstractThe time delay of arrival (TDOA) between multiple microphones has been used since 2006 as a source of information (localization) to complement the spectral features for speaker diarization. In this paper, we propose a new localization feature, the intensity channel contribution (ICC) based on the relative energy of the signal arriving at each channel compared to the sum of the energy of all the channels. We have demonstrated that by joining the ICC features and the TDOA features, the robustness of the localization features is improved and that the diarization error rate (DER) of the complete system (using localization and spectral features) has been reduced. By using this new localization feature, we have been able to achieve a 5.2% DER relative improvement in our development data, a 3.6% DER relative improvement in the RT07 evaluation data and a 7.9% DER relative improvement in the last year's RT09 evaluation data. Roberto Barra-Chicote, José Manuel Pardo, Javier Ferreiros, Juan Manuel Montero-Martínez |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | HIFI-AV: An Audio-visual Corpus for Spoken Language Human-Machine Dialogue Research in Spanish
Fernando Fernández Martínez, Juan Manuel Lucas, Roberto Barra-Chicote, Javier Ferreiros, Javier Macías Guarasa |
LREC | 3 |
| 2010 | Spoken Spanish generation from sign languageabstractThis paper describes the development of a Spoken Spanish generator from sign-writing. The sign language considered was the Spanish sign language (LSE: Lengua de Signos Española). This system consists of an advanced visual interface (where a deaf person can specify a sequence of signs in sign-writing), a language translator (for generating the sequence of words in Spanish), and finally, a text to speech converter. The visual interface allows a sign sequence to be defined using several sign-writing alternatives. The paper details the process for designing the visual interface proposing solutions for HCI-specific challenges when working with the Deaf (i.e. important difficulties in writing Spanish or limited sign coverage for describing abstract or conceptual ideas). Three strategies were developed and combined for language translation to implement the final version of the language translator module. The summative evaluation, carried out with Deaf from Madrid and Toledo, includes objective measurements from the system and subjective information from questionnaires. The paper also describes the first Spanish-LSE parallel corpus for language processing research focused on specific domains. This corpus includes more than 4000 Spanish sentences translated into LSE. These sentences focused on two restricted domains: the renewal of the identity document and driver’s license. This corpus also contains all sign descriptions in several sign-writing specifications generated with a new version of the eSign Editor. This new version includes a grapheme to phoneme system for Spanish and a SEA-HamNoSys converter. Rubén San-Segundo-Hernández, José Manuel Pardo, Javier Ferreiros, Valentín Sama Rojo, Roberto Barra-Chicote, Juan Manuel Lucas, D. Sánchez |
Interact. Comput. | 5 |
| 2010 | Analysis of statistical parametric and unit selection speech synthesis systems applied to emotional speech
Roberto Barra-Chicote, Junichi Yamagishi, Simon King 0001, Juan Manuel Montero-Martínez, Javier Macías Guarasa |
Speech Commun. | 1 |
| 2009 | Acoustic emotion recognition using dynamic Bayesian networks and multi-space distributionsabstractIn this paper we describe the acoustic emotion recognition system built at the Speech Technology Group of the Universidad Politecnica de Madrid (Spain) to participate in the INTERSPEECH 2009 Emotion Challenge. Our proposal is based on the use of a Dynamic Bayesian Network (DBN) to deal with the temporal modelling of the emotional speech information. The selected features (MFCC, F0, Energy and their variants) are modelled as different streams, and the F0 related ones are integrated under a Multi Space Distribution (MSD) framework, to properly model its dual nature (voiced/unvoiced). Experimental evaluation on the challenge test set, show a 67.06%and 38.24% of unweighted recall for the 2 and 5-classes tasks respectively. In the 2-class case, we achieve similar results compared with the baseline, with a considerable less number of features. In the 5-class case, we achieve a statistically significant 6.5% relative improvement Roberto Barra-Chicote, Fernando Fernández Martínez, Syaheerah L. Lutfi, Juan Manuel Lucas, Javier Macías Guarasa, Juan Manuel Montero-Martínez, Rubén San-Segundo-Hernández, José Manuel Pardo |
INTERSPEECH | 1 |
| 2009 | Speeding Up the Design of Dialogue Applications by Using Database Contents and Structure Information
Luis Fernando D'Haro, Ricardo de Córdoba, Juan Manuel Lucas, Roberto Barra-Chicote, Rubén San-Segundo-Hernández |
SIGDIAL Conference | 4 |
| 2008 | Evaluation of a spoken dialogue system for controlling a Hifi audio systemabstractIn this paper a Bayesian Networks, BNs, approach to dialogue modelling [1] is evaluated in terms of a battery of both subjective and objective metrics. A significant effort in improving the contextual information handling capabilities of the system has been done. Consequently, besides typical dialogue measurement rates for usability like task or dialogue completion rates, dialogue time, etc. we have included a new figure measuring the contextuality of the dialogue as the number of turns where contextual information is helpful for dialogue resolution. The evaluation is developed through a set of predefined scenarios according to different initiative styles and focusing on the impact of the user's level of experience. Fernando Fernández Martínez, Juan Blázquez, Javier Ferreiros, Roberto Barra-Chicote, Javier Macías Guarasa, Juan Manuel Lucas |
SLT | 4 |
| 2008 | Speech to sign language translation system for Spanish
Rubén San-Segundo-Hernández, Roberto Barra-Chicote, Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Javier Ferreiros, Juan Manuel Lucas, Javier Macías Guarasa, Juan Manuel Montero-Martínez, José Manuel Pardo |
Speech Commun. | 2 |
| 2007 | On the limitations of voice conversion techniques in emotion identification tasksabstractThe growing interest in emotional speech synthesis urges effective emotion conversion techniques to be explored. This paper estimates the relevance of three speech components (spectral envelope, residual excitation and prosody) for synthesizing identifiable emotional speech, in order to be able to customize voice conversion techniques to the specific characteristics of each emotion. The analysis has been based on a listening test with a set of synthetic mixed-emotion utterances that draw their speech components from emotional and neutral recordings. Results prove the importance of transforming residual excitation for the identification of emotions that are not fully conveyed through prosodic means (such as cold anger or sadness in our Spanish corpus). Roberto Barra-Chicote, Juan Manuel Montero-Martínez, Javier Macías Guarasa, Juana M. Gutiérrez-Arriola, Javier Ferreiros, José Manuel Pardo |
INTERSPEECH | 1 |
| 2007 | Language identification using several sources of information with a multiple-Gaussian classifierabstractWe present several innovative techniques that can be applied in a PPRLM system for language identification (LID). To normalize the scores, eliminate the bias in the scores and improve the classifier, we compared the bias removal technique (up to 19 % relative improvement (RI)) and a Gaussian classifier (up to 37 % RI). Then, we include additional sources of information in different feature vectors of the Gaussian classifier: the sentence acoustic score (11% RI), the average acoustic score for each phoneme (11 % RI), and the average duration for each phoneme (7.8 % RI). The use of a multiple-Gaussian classifier with 4 feature vectors meant an additional 15.1 % RI. Using 4 feature vectors instead of just PPRLM provides a 26.1 % RI. Finally, we include additional acoustic HMMs of the same language with success (10 % relative improvement). We will show how all these improvements have been mostly additive. Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Juan Manuel Montero-Martínez, Roberto Barra-Chicote |
INTERSPEECH | 5 |
| 2007 | Automatic phonetic segmentation of Spanish emotional speechabstractTo achieve high quality synthetic emotional speech, unitselection is the state-of-the-art technique. Nevertheless, a large expensive phonetically-segmented corpus is needed, and cost-effective automatic techniques should be studied. According to the HMM experiments in this paper: segmentation performance can depend heavily on the segmental or prosodic nature of the intended emotion (segmental emotions are more difficult to segment than prosodic ones), several emotions should be combined to obtain a larger training set (especially when prosodic emotions are involved; this is especially true for small training sets) and a combination of emphatic and nonemphatic emotional recordings (short sentences vs. long paragraphs) can degrade overall performance. Index Terms: expressive speech, automatic phonetic segmentation, emotional speech synthesis Ascensión Gallardo-Antolín, Roberto Barra-Chicote, Marc Schröder 0001, Sacha Krstulovic, Juan Manuel Montero-Martínez |
INTERSPEECH | 2 |
| 2006 | Prosodic and Segmental Rubrics in Emotion IdentificationabstractIt is well known that the emotional state of a speaker usually alters the way she/he speaks. Although all the components of the voice can be affected by emotion in some statistically-significant way, not all these deviations from a neutral voice are identified by human listeners as conveying emotional information. In this paper we have carried out several perceptual and objective experiments that show the relevance of prosody and segmental spectrum in the characterization and identification of four emotions in Spanish. A Bayes classifier has been used in the objective emotion identification task. Emotion models were generated as the contribution of every emotion to the build-up of a universal background emotion codebook. According to our experiments, surprise is primarily identified by humans through its prosodic rubric (in spite of some automatically-identifiable segmental characteristics); while for anger the situation is just the opposite. Sadness and happiness need a combination of prosodic and segmental rubrics to be reliably identified Roberto Barra-Chicote, Juan Manuel Montero-Martínez, Javier Macías Guarasa, Luis Fernando D'Haro, Rubén San-Segundo-Hernández, Ricardo de Córdoba |
ICASSP (1) | 1 |
| 2006 | A Spanish speech to sign language translation system for assisting deaf-mute peopleabstractThis paper describes the first experiments of a speech to sign language translation system in a real domain. The developed system is focused on the sentences spoken by an officer when assisting people in applying for, or renewing the National Identification Document (NID) and the Passport. This system translates officer explanations into sign language for deafmute people. The translation system is composed by a speech recognizer (for decoding the spoken utterance into a word sequence), a natural language translator (for converting a word sequence into a sequence of gestures belonging to the sign language), and a 3D avatar animation module (for playing the gestures). The field experiments have reported a 27.2 % GER (Gesture Error Rate) and a 0.62 BLEU Rubén San-Segundo-Hernández, Roberto Barra-Chicote, Luis Fernando D'Haro, Juan Manuel Montero-Martínez, Ricardo de Córdoba, Javier Ferreiros |
INTERSPEECH | 2 |
| 2005 | New word-level and sentence-level confidence scoring using graph theory calculus and its evaluation on speech understandingabstractA lot of work has been devoted to the estimation of confidence measures for speech recognizers. In the quite extended case where a word-graph speech recognizer is in use, we will present new confidence measures employing the graph theory that shows us how to estimate some interesting characteristics about the different paths through the graph that constitute the recognition solutions, without the need of expanding them all. We will take advantage of some of these features to generate confidence scores both at the word and sentence level. We will also compare this new confidence scoring to more traditional ones and will find similar behavior with less computational load and with an increase in the simplicity of the approach that will lead to more generalization power of the confidence estimation to different applications of the recognizer. Javier Ferreiros, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Luis Fernando D'Haro, Valentín Sama Rojo, Roberto Barra-Chicote, Pedro Mellén |
INTERSPEECH | 6 |