Giulia Comini

dblp:207/5179 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-9391-6565ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Lightweight neural front-ends for low-resource on-device Text-to-Speech
abstract
We propose a lightweight neural front-end framework for on-device speech generation and highlight its benefits towards low-resource language scaling. While data-driven models have shown potential in front-end literature, especially since they can enable fast language expansion, they are often extremely large and of high latency. There is limited work focusing on their usability in real-time settings, and none for on-device TTS applications. At the cost of small performances trade-offs, we build lightweight neural Grapheme-to-Phoneme and verbalization models which achieve, on average across three languages, a p90 latency reduction of 95.98% per token on single-threaded [email protected], with respect to a traditional transformer-based baseline, while having 99.26% less parameters. Additionally, leveraging pre-trained teacher models to bootstrap lightweight students, we enable low-resource language scaling on both Grapheme-to-Phoneme conversion and verbalization.
Giulia Comini, Heereen Shim, Manuel Sam Ribeiro
ICASSP1
2023 Multilingual context-based pronunciation learning for Text-to-Speech
Giulia Comini, Manuel Sam Ribeiro, Heereen Shim, Jaime Lorenzo-Trueba
INTERSPEECH1
2023 Improving grapheme-to-phoneme conversion by learning pronunciations from speech recordings
Manuel Sam Ribeiro, Giulia Comini, Jaime Lorenzo-Trueba
INTERSPEECH2
2022 Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module
abstract
State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and intelligibility degradations, making training low-resource TTS systems problematic. In this paper, we propose a novel extremely low-resource TTS method called Voice Filter that uses as little as one minute of speech from a target speaker. It uses voice conversion (VC) as a post-processing module appended to a pre-existing high-quality TTS system and marks a conceptual shift in the existing TTS paradigm, framing the few-shot TTS problem as a VC task. Furthermore, we propose to use a duration-controllable TTS system to create a parallel speech corpus to facilitate the VC task. Results show that the Voice Filter outperforms state-of-the-art few-shot speech synthesis techniques in terms of objective and subjective metrics on one minute of speech on a diverse set of voices, while being competitive against a TTS model built on 30 times more data.1
Adam Gabrys, Goeric Huybrechts, Manuel Sam Ribeiro, Chung-Ming Chien, Julian Roth, Giulia Comini, Roberto Barra-Chicote, Bartek Perz, Jaime Lorenzo-Trueba
ICASSP6
2022 Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation
abstract
We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting conversational expressive data from different speakers. Our goal is to build a TTS system that is expressive, while retaining the target speaker’s identity. The proposed approach relies on voice conversion to first generate high-quality data from the set of supporting expressive speakers. The voice converted data is then pooled with natural data from the target speaker and used to train a single-speaker multi-style TTS system. We provide evidence that this approach is efficient, flexible, and scalable. The method is evaluated using one or more supporting speakers, as well as a variable amount of supporting data. We further provide evidence that this approach allows some controllability of speaking style, when using multiple supporting speakers. We conclude by scaling our proposed technology to a set of 14 speakers across 7 languages. Results indicate that our technology consistently improves synthetic samples in terms of style similarity, while retaining the target speaker’s identity.
Manuel Sam Ribeiro, Julian Roth, Giulia Comini, Goeric Huybrechts, Adam Gabrys, Jaime Lorenzo-Trueba
ICASSP3
2022 Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation
abstract
The availability of data in expressive styles across languages is limited, and recording sessions are costly and time consuming.To overcome these issues, we demonstrate how to build low-resource, neural text-to-speech (TTS) voices with only 1 hour of conversational speech, when no other conversational data are available in the same language.Assuming the availability of non-expressive speech data in that language, we propose a 3-step technology: 1) we train an F0-conditioned voice conversion (VC) model as data augmentation technique; 2) we train an F0 predictor to control the conversational flavour of the voice-converted synthetic data; 3) we train a TTS system that consumes the augmented data.We prove that our technology enables F0 controllability, is scalable across speakers and languages and is competitive in terms of naturalness over a state-of-the-art baseline model, another augmented method which does not make use of F0 information.
Giulia Comini, Goeric Huybrechts, Manuel Sam Ribeiro, Adam Gabrys, Jaime Lorenzo-Trueba
INTERSPEECH1
2021 Low-Resource Expressive Text-To-Speech Using Data Augmentation
abstract
While recent neural text-to-speech (TTS) systems perform remarkably well, they typically require a substantial amount of recordings from the target speaker reading in the desired speaking style. In this work, we present a novel 3-step methodology to circumvent the costly operation of recording large amounts of target data in order to build expressive style voices with as little as 15 minutes of such recordings. First, we augment data via voice conversion by leveraging recordings in the desired speaking style from other speakers. Next, we use that synthetic data on top of the available recordings to train a TTS model. Finally, we fine-tune that model to further increase quality. Our evaluations show that the proposed changes bring significant improvements over non-augmented models across many perceived aspects of synthesised speech. We demonstrate the proposed approach on 2 styles (news-caster and conversational), on various speakers, and on both single and multi-speaker models, illustrating the robustness of our approach.1
Goeric Huybrechts, Thomas Merritt, Giulia Comini, Bartek Perz, Raahil Shah, Jaime Lorenzo-Trueba
ICASSP3
2017 Abstract Games of Argumentation Strategy and Game-Theoretical Argument Strength
Pietro Baroni, Giulia Comini, Antonio Rago 0001, Francesca Toni
PRIMA2