Olivier Perrotin

dblp:159/4902 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-9909-6078ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Is self-supervised learning enough to fill in the gap? A study on speech inpainting
abstract
Speech inpainting consists in reconstructing corrupted or missing speech segments using surrounding context, a process that closely resembles the pretext tasks in Self-Supervised Learning (SSL) for speech encoders. This study investigates using SSL-trained speech encoders for inpainting without any additional training beyond the initial pretext task, and simply adding a decoder to generate a waveform. We compare this approach to supervised fine-tuning of speech encoders for a downstream task—here, inpainting. Practically, we integrate HuBERT as the SSL encoder and HiFi-GAN as the decoder in two configurations: (1) fine-tuning the decoder to align with the frozen pre-trained encoder’s output and (2) fine-tuning the encoder for an inpainting task based on a frozen decoder’s input. Evaluations are conducted under single- and multi-speaker conditions using in-domain datasets and out-of-domain datasets (including unseen speakers, diverse speaking styles, and noise). Both informed and blind inpainting scenarios are considered, where the position of the corrupted segment is either known or unknown. The proposed SSL-based methods are benchmarked against several baselines, including a text-informed method combining automatic speech recognition with zero-shot text-to-speech synthesis. Performance is assessed using objective metrics and perceptual evaluations. The results demonstrate that both approaches outperform baselines, successfully reconstructing speech segments up to 200 ms, and sometimes up to 400 ms. Notably, fine-tuning the SSL encoder achieves more accurate speech reconstruction in single-speaker settings, while a pre-trained encoder proves more effective for multi-speaker scenarios. This demonstrates that an SSL pretext task can transfer to speech inpainting, enabling successful speech reconstruction with a pre-trained encoder.
Ihab Asaad, Maxime Jacquelin, Olivier Perrotin, Laurent Girin, Thomas Hueber
Comput. Speech Lang.3
2026 A closer look at internal representations of end-to-end Text-to-Speech models: How is phonetic and acoustic information encoded?
abstract
In recent years, deep neural architectures have demonstrated groundbreaking performances in various speech processing areas, including Text-To-Speech (TTS). Models have grown larger, including more layers and millions of trainable parameters to achieve near-natural synthesis, at the expense of interpretability of computed intermediate representations. However, the statistical learning performed by these neural models offers a valuable source of information about language and speech production. The present study aims to develop statistical tools to narrow the gap between these advanced processing techniques and speech sciences. By linearly probing phonetic and acoustic features in model representations, the proposed methods help to understand how neural TTS are able to organize speech information in an unsupervised manner and provide novel insights on phonetic regularities captured through statistical learning on massive datasets that extend beyond human expertise. This study takes a step further by leveraging these insights to design emerging control mechanisms for speech synthesis models, without requiring additional data or training processes. The proposed control is evaluated across a variety of acoustic and prosodic parameters relevant to the perception of speech expressivity. Experiments on the two foundational TTS models Tacotron2 and FastSpeech2 on a multi-speaker French dataset demonstrate promising performance of these control mechanisms, and underscore the value of employing explainability methods in a broader range of domains, enabling neural models to be viewed not merely as modeling tools, but as analysis frameworks that invite a deeper exploration of their underlying mechanisms and structures. Such an approach fosters more comprehensive insights that can improve both the technology and its applications.
Martin Lenglet, Olivier Perrotin, Gérard Bailly
Comput. Speech Lang.2
2026 Hand gesture realisation of contrastive focus in real-time whisper-to-speech synthesis: Investigating the transfer from implicit to explicit control of intonation
abstract
The ability of speakers to externalise the control of their intonation in the context of voice substitution communication is evaluated in terms of the realisation of a contrastive focus in French. A whisper-to-speech synthesiser is used with gestural interfaces for intonation control, enabling two types of gesture: an isometric finger pressure and an isotonic wrist movement. An original experimental paradigm is designed to elicit a contrastive focus on the syllables of nine-syllable sentences by means of a read-question-answer scenario. For all 16 participants, focus was successfully achieved in speech and in both modality transfer situations by increasing the fundamental frequency and duration of the target syllable. Coordination of the articulation of the whispered syllables and the manual intonational control was acquired quickly and easily. Focus realisation by finger pressure or wrist movement showed very similar dynamics in intonation and duration. Overall, although wrist movement was preferred in terms of ease of control, both interfaces were judged to be equal in terms of learning, performance, emotional experience, and cognitive load.
Delphine Charuau, Nathalie Henrich Bernardoni, Silvain Gerber, Olivier Perrotin
Speech Commun.4
2025 LombardTokenizer: Disentanglement and Control of Vocal Effort in a Neural Speech Codec
abstract
International audience
Maxime Jacquelin, Maëva Garnier, Laurent Girin, Rémy Vincent, Olivier Perrotin
INTERSPEECH5
2025 Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 2023
abstract
International audience
Olivier Perrotin, Brooke Stephenson, Silvain Gerber, Gérard Bailly, Simon King 0001
Comput. Speech Lang.1
2024 Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS
abstract
We developped a web app for ascribing verbal descriptions to expressive audiovisual utterances. These descriptions are limited to lists of adjectives that are either suggested via a navigation in emotional latent spaces built using discriminant analysis of BERT embeddings or entered freely by subjects. We show that such verbal descriptions collected on-line via Prolific on massive data (310 participants, 12620 labelled utterances up-to-now) provide Expressive Multimodal Text-to-Speech Synthesis with precise verbal control over desired emotional content
Gérard Bailly, Romain Legrand, Martin Lenglet, Frédéric Elisei, Maëva Hueber, Olivier Perrotin
LREC/COLING6
2024 FastLips: an End-to-End Audiovisual Text-to-Speech System with Lip Features Prediction for Virtual Avatars
abstract
International audience
Martin Lenglet, Olivier Perrotin, Gérard Bailly
INTERSPEECH2
2023 Investigating the dynamics of hand and lips in French Cued Speech using attention mechanisms and CTC-based decoding
abstract
Hard of hearing or profoundly deaf people make use of cued speech (CS) as a communication tool to understand spoken language. By delivering cues that are relevant to the phonetic information, CS offers a way to enhance lipreading. In literature, there have been several studies on the dynamics between the hand and the lips in the context of human production. This article proposes a way to investigate how a neural network learns this relation for a single speaker while performing a recognition task using attention mechanisms. Further, an analysis of the learnt dynamics is utilized to establish the relationship between the two modalities and extract automatic segments. For the purpose of this study, a new dataset has been recorded for French CS. Along with the release of this dataset, a benchmark will be reported for word-level recognition, a novelty in the automatic recognition of French CS.
Sanjana Sankar, Denis Beautemps, Frédéric Elisei, Olivier Perrotin, Thomas Hueber
INTERSPEECH4
2022 Voicing decision based on phonemes classification and spectral moments for whisper-to-speech conversion
abstract
International audience
Luc Ardaillon, Nathalie Henrich Bernardoni, Olivier Perrotin
INTERSPEECH3
2022 Speaking Rate Control of end-to-end TTS Models by Direct Manipulation of the Encoder's Output Embeddings
abstract
International audience
Martin Lenglet, Olivier Perrotin, Gérard Bailly
INTERSPEECH2
2021 Evaluating the Extrapolation Capabilities of Neural Vocoders to Extreme Pitch Values
abstract
International audience
Olivier Perrotin, Hussein El Amouri, Gérard Bailly, Thomas Hueber
Interspeech1
2020 Hider-Finder-Combiner: An Adversarial Architecture for General Speech Signal Modification
abstract
International audience
Jacob J. Webber, Olivier Perrotin, Simon King 0001
INTERSPEECH2
2020 Glottal Flow Synthesis for Whisper-to-Speech Conversion
abstract
Whisper-to-speech conversion is motivated by laryngeal disorders, in which malfunction of the vocal folds leads to loss of voicing. Many patients with laryngeal disorders can still produce functional whispers, since these are characterised by the absence of vocal fold vibration. Whispers therefore constitute a common ground for speech rehabilitation across many kinds of laryngeal disorder. Whisper-to-speech conversion involves recreating natural-sounding speech from recorded whispers, and is a non-invasive and non-surgical rehabilitation that can maintain a natural method of speaking, unlike the existing methods of rehabilitation. This article proposes a new rule-based method for whisper-to-speech conversion that replaces the noisy whisper sound source with a synthesised speech-like harmonic source, while maintaining the vocal tract component unaltered. In particular, a novel glottal source generator is developed in which whisper information is used to parameterise the excitation through a high-quality glottis model. Evaluation of the system against the standard pulse train excitation method reveals significantly improved performance. Since our method is glottis-based, it is potentially compatible with the many existing vocal tract component adaptation systems.
Olivier Perrotin, Ian McLoughlin 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 A Spectral Glottal Flow Model for Source-filter Separation of Speech
abstract
The estimation of glottal flow from a speech waveform is an essential technique used in speech analysis and parameterisation. Significant research effort has been addressed at separating the first vocal tract resonance from the glottal formant (the low-frequency resonance that describes the open-phase of the vocal fold vibration), but few methods are capable of estimating the high-frequency spectral tilt, characteristic of the closing phase of the vocal fold vibration (which is crucial to the perception of vocal effort). This paper proposes an improved Iterative Adaptive Inverse Filtering (IAIF) method based on a Glottal Flow Model, which we call GFM-IAIF. The proposed method models the wide-band glottis response, incorporating both glottal formant and spectral tilt characteristics. Evaluation against IAIF and recently proposed IOP-IAIF shows that, while GFM-IAIF maintains good performance on vocal tract modelling, it significantly improves the glottis model. This ensures that timbral variations associated to voice quality can be correctly attributed and described.
Olivier Perrotin, Ian McLoughlin 0001
ICASSP1
2019 GFM-Voc: A Real-Time Voice Quality Modification System
Olivier Perrotin, Ian McLoughlin 0001
INTERSPEECH1
2017 Seeing, Listening, Drawing: Interferences between Sensorimotor Modalities in the Use of a Tablet Musical Interface
abstract
Audio, visual, and proprioceptive actions are involved when manipulating a graphic tablet musical interface. Previous works suggested a possible dominance of the visual over the auditory modality in this situation. The main goal of the present study is to examine the interferences between these modalities in visual, audio, and audio-visual target acquisition tasks. Experiments are based on a movement replication paradigm, where a subject controls a cursor on a screen or the pitch of a synthesized sound by changing the stylus position on a covered graphic tablet. The experiments consisted of the following tasks: (1) a target acquisition task that was aimed at a visual target (reaching a cue with the cursor displayed on a screen), an audio target (reaching a reference note by changing the pitch of the sound played in headsets), or an audio-visual target, and (2) the replication of the target acquisition movement in the opposite direction. In the return phase, visual and audio feedback were suppressed. Different gain factors perturbed the relationships among the stylus movements, visual cursor movements, and audio pitch movements. The deviations between acquisition and return movements were analyzed. The results showed that hand amplitudes varied in accordance with visual, audio, and audio-visual perturbed gains, showing a larger effect for the visual modality. This indicates that visual, audio, and audio-visual actions interfered with the motor modality and confirms the spatial representation of pitch reported in previous studies. In the audio-visual situation, vision dominated over audition, as the latter had no significant influence on motor movement. Consequently, visual feedback is helpful for musical targeting of pitch on a graphic tablet, at least during the learning phase of the instrument. This result is linked to the underlying spatial organization of pitch perception. Finally, this work brings a complementary approach to previous studies showing that audition may dominate over vision for other aspects of musical sound (e.g., timing, rhythm, and timbre).
Olivier Perrotin, Christophe d'Alessandro
ACM Trans. Appl. Percept.1
2016 Vocal Effort Modification for Singing Synthesis
abstract
International audience
Olivier Perrotin, Christophe d'Alessandro
INTERSPEECH1
2016 Target Acquisition vs. Expressive Motion: Dynamic Pitch Warping for Intonation Correction
abstract
The purpose of pitch correction is to assist a musician in playing notes with accuracy and precision, without preventing expressive pitch variations. This study presents and examines a new method for automatic pitch correction: Dynamic Pitch Warping (DPW). The analytic formulation of the warping function is derived. In the context of live playing of continuous pitch trajectories, the dynamics of pitch correction must be considered. Methods for triggering and releasing the correction are discussed, and a performance test is conducted. DPW is evaluated in the context of digital musical instruments that are controlled by a stylus on a graphic tablet. The results show significant improvement in note accuracy and precision with the addition of the correction method. Analyses of various types of modulations (including vibrato, portamento, and glissando) demonstrate that expressive pitch variations are preserved by the DPW correction. Perceptual tests show that the effects of DPW correction are well perceived and positively assessed by listeners. The proposed method allows for accurate pitch target acquisition together with preservation of expressive motion, a result that could be extended to other situations that require dynamic trajectory correction.
Olivier Perrotin, Christophe d'Alessandro
ACM Trans. Comput. Hum. Interact.1