Vincent Colotte

dblp:15/9230 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0000-5040-6971ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Towards realtime co-speech gestures synthesis using STARGATE
abstract
International audience
Louis Abel, Vincent Colotte, Slim Ouni
INTERSPEECH2
2023 Stochastic Pitch Prediction Improves the Diversity and Naturalness of Speech in Glow-TTS
Sewade Ogun, Vincent Colotte, Emmanuel Vincent 0001
INTERSPEECH2
2022 Analysis of expressivity transfer in non-autoregressive end-to-end multispeaker TTS systems
abstract
International audience
Ajinkya Kulkarni, Vincent Colotte, Denis Jouvet
INTERSPEECH2
2022 Can We Use Common Voice to Train a Multi-Speaker TTS System?
abstract
Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies have leveraged the availability of large, crowdsourced automatic speech recognition (ASR) datasets. A major problem with such datasets is the presence of noisy and/or distorted samples, which degrade TTS quality. In this paper, we propose to automatically select high-quality training samples using a non-intrusive mean opinion score (MOS) estimator, WV-MOS. We show the viability of this approach for training a multi-speaker GlowTTS model on the Common Voice English dataset. Our approach improves the overall quality of generated utterances by 1.26 MOS point with respect to training on all the samples and by 0.35 MOS point with respect to training on the LibriTTS dataset. This opens the door to au-tomatic TTS dataset curation for a wider range of languages.
Sewade Ogun, Vincent Colotte, Emmanuel Vincent 0001
SLT2
2021 Duration modelling and evaluation for Arabic statistical parametric speech synthesis
Imene Zangar, Zied Mnasri, Vincent Colotte, Denis Jouvet
Multim. Tools Appl.3
2021 Learning emotions latent representation with CVAE for text-driven expressive audiovisual speech synthesis
Sara Dahmani, Vincent Colotte, Valérian Girard, Slim Ouni
Neural Networks2
2020 Transfer Learning of the Expressivity Using FLOW Metric Learning in Multispeaker Text-to-Speech Synthesis
abstract
International audience
Ajinkya Kulkarni, Vincent Colotte, Denis Jouvet
INTERSPEECH2
2019 Conditional Variational Auto-Encoder for Text-Driven Expressive AudioVisual Speech Synthesis
abstract
In recent years, the performance of speech synthesis systems has been improved thanks to deep learning-based models, but generating expressive audiovisual speech is still an open issue. The variational auto-encoders (VAE)s are recently proposed to learn latent representations of data. In this paper, we present a system for expressive text-to-audiovisual speech synthesis that learns a latent embedding space of emotions using a conditional generative model based on the variational auto-encoder framework. When conditioned on textual input, the VAE is able to learn an embedded representation that captures emotion characteristics from the signal, while being invariant to the phonetic content of the utterances. We applied this method in an unsuper-vised manner to generate duration, acoustic and visual features of speech. This conditional variational auto-encoder (CVAE) has been used to blend emotions together. This model was able to generate nuances of a given emotion or to generate new emotions that do not exist in our database. We conducted three perceptive experiments to evaluate our findings.
Sara Dahmani, Vincent Colotte, Valérian Girard, Slim Ouni
INTERSPEECH2
2016 Acoustic and Visual Analysis of Expressive Speech: A Case Study of French Acted Speech
abstract
International audience
Slim Ouni, Vincent Colotte, Sara Dahmani, Soumaya Azzi
INTERSPEECH2
2016 The IFCASL Corpus of French and German Non-native and Native Read Speech
Jürgen Trouvain, Anne Bonneau, Vincent Colotte, Camille Fauth, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius, Frank Zimmerer
LREC3
2014 Designing a Bilingual Speech Corpus for French and German Language Learners: a Two-Step Process
Camille Fauth, Anne Bonneau, Frank Zimmerer, Jürgen Trouvain, Bistra Andreeva, Vincent Colotte, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius
LREC6
2011 Weight Optimization for Bimodal Unit-Selection Talking Head Synthesis
abstract
(accepted)
Asterios Toutios, Utpala Musti, Slim Ouni, Vincent Colotte
INTERSPEECH4
2010 HMM-based automatic visual speech segmentation using facial data
abstract
International audience
Utpala Musti, Asterios Toutios, Slim Ouni, Vincent Colotte, Brigitte Wrobel-Dautcourt, Marie-Odile Berger
INTERSPEECH4
2010 Setup for acoustic-visual speech synthesis by concatenating bimodal units
abstract
International audience
Asterios Toutios, Utpala Musti, Slim Ouni, Vincent Colotte, Brigitte Wrobel-Dautcourt, Marie-Odile Berger
INTERSPEECH4
2005 Linguistic features weighting for a text-to-speech system without prosody model
abstract
This paper presents a Non-Uniform Units selection-based Text-To-Speech synthesizer. Nowadays, systems use prosodic models that do not allow the prosody to vary as far as we should hope, involving a listening comfort degradation. Our system has the advantage to avoid the using of prosodic model. Speech units selection builds its features set exclusively from the linguistic information generated by the natural language analysis. We also present an original method to automatically weight these features. Therefore, selected units are not restricted by a predetermined prosody. With only using linguistic features, we obtain a various prosody and the units concatenation is performed without resort to heavy signal processing
Vincent Colotte, Richard Beaufort
INTERSPEECH1
2001 Perceptual experiments on enhanced and slowed down speech sentences for second language acquisition
abstract
Colloque avec actes et comité de lecture. internationale.
Vincent Colotte, Yves Laprie, Anne Bonneau
INTERSPEECH1
2000 Automatic enhancement of speech intelligibility
abstract
This paper presents a speech signal transformation which slows down speech signals selectively and enhances some important acoustic cues. This transformation can be used not only for hearing aids but also for second language acquisition by facilitating oral comprehension. Selective slowing down relies on the use of the TD-PSOLA synthesis method. An automatic pitch marking algorithm was designed to apply this method automatically. The strategy used to control slowing down exploits a spectral variation function which locates rapid spectral changes. The enhancement simply consists of amplifying stop bursts and unvoiced fricatives. These acoustic cues are detected automatically through the examination of energy criteria. This approach was evaluated in the context of second language acquisition, more precisely by evaluating improvements in oral comprehension. Transformations triggered properly, i.e. the signal regions modified are those which were expected to be modified. Experiments show that the oral comprehension is improved.
Vincent Colotte, Yves Laprie
ICASSP1