Gérard Bailly

dblp:48/1036 · DBLP profile ↗
← Back
102ranked-venue papers
28as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 21 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 76 · 25 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A closer look at internal representations of end-to-end Text-to-Speech models: How is phonetic and acoustic information encoded?
abstract
In recent years, deep neural architectures have demonstrated groundbreaking performances in various speech processing areas, including Text-To-Speech (TTS). Models have grown larger, including more layers and millions of trainable parameters to achieve near-natural synthesis, at the expense of interpretability of computed intermediate representations. However, the statistical learning performed by these neural models offers a valuable source of information about language and speech production. The present study aims to develop statistical tools to narrow the gap between these advanced processing techniques and speech sciences. By linearly probing phonetic and acoustic features in model representations, the proposed methods help to understand how neural TTS are able to organize speech information in an unsupervised manner and provide novel insights on phonetic regularities captured through statistical learning on massive datasets that extend beyond human expertise. This study takes a step further by leveraging these insights to design emerging control mechanisms for speech synthesis models, without requiring additional data or training processes. The proposed control is evaluated across a variety of acoustic and prosodic parameters relevant to the perception of speech expressivity. Experiments on the two foundational TTS models Tacotron2 and FastSpeech2 on a multi-speaker French dataset demonstrate promising performance of these control mechanisms, and underscore the value of employing explainability methods in a broader range of domains, enabling neural models to be viewed not merely as modeling tools, but as analysis frameworks that invite a deeper exploration of their underlying mechanisms and structures. Such an approach fosters more comprehensive insights that can improve both the technology and its applications.
Martin Lenglet, Olivier Perrotin, Gérard Bailly
Comput. Speech Lang.3
2025 Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model
abstract
This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by leveraging a pre-trained audiovisual autoregressive text-to-speech model (AVTacotron2). This model is reprogrammed to infer Cued Speech (CS) hand and lip movements from text input. Experiments are conducted on two publicly available datasets, including one recorded specifically for this study. Performance is assessed using an automatic CS recognition system. With a decoding accuracy at the phonetic level reaching approximately 77%, the results demonstrate the effectiveness of our approach.
Sanjana Sankar, Martin Lenglet, Gérard Bailly, Denis Beautemps, Thomas Hueber
ICASSP3
2025 Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 2023
abstract
International audience
Olivier Perrotin, Brooke Stephenson, Silvain Gerber, Gérard Bailly, Simon King 0001
Comput. Speech Lang.4
2025 THERADIA WoZ: An Ecological Corpus for Appraisal-Based Affect Research in Healthcare
abstract
We present THERADIA WoZ, an ecological corpus designed for audiovisual research on affect in healthcare. Two groups of senior individuals, consisting of 52 healthy participants and 9 individuals with Mild Cognitive Impairment (MCI), performed Computerised Cognitive Training (CCT) exercises while receiving support from a virtual assistant, tele-operated by a human in the role of a Wizard-of-Oz (WoZ). The audiovisual expressions produced by the participants were fully transcribed, and partially annotated based on dimensions derived from recent appraisal theory models, including novelty, intrinsic pleasantness, goal conduciveness, and coping. Additionally, the annotations included 23 affective labels from the literature of achievement affects. We present the data collection, transcription, and annotation protocols, alongside a detailed analysis of the annotated dimensions and labels. Baseline methods and results for their automatic prediction are also presented. Results reveal that the dimensions of appraisal theory can be predicted, with the performance varying across different modalities. The corpus aims to serve as a valuable resource for researchers in affective computing, and is made available to both industry and academia.
Hippolyte Fournier, Sina Alisamir, Safaa Azzakhnini, Isabella Zsoldos, Eléonore Trân, Gérard Bailly, Frédéric Elisei, Béatrice Bouchot, Brice Varini, Patrick Constant, Joan Fruitet, Franck Tarpin-Bernard, Solange Rossato, François Portet, Olivier Koenig, Hanna Chainay, Fabien Ringeval
IEEE Trans. Affect. Comput.6
2024 Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS
abstract
We developped a web app for ascribing verbal descriptions to expressive audiovisual utterances. These descriptions are limited to lists of adjectives that are either suggested via a navigation in emotional latent spaces built using discriminant analysis of BERT embeddings or entered freely by subjects. We show that such verbal descriptions collected on-line via Prolific on massive data (310 participants, 12620 labelled utterances up-to-now) provide Expressive Multimodal Text-to-Speech Synthesis with precise verbal control over desired emotional content
Gérard Bailly, Romain Legrand, Martin Lenglet, Frédéric Elisei, Maëva Hueber, Olivier Perrotin
LREC/COLING1
2024 EVAC 2024 - Empathic Virtual Agent Challenge: Appraisal-based Recognition of Affective States
abstract
As autonomous interactive agents become increasingly prevalent, it is crucial for these virtual agents to understand and respond to both our verbal content and emotions, enabling deeper interactions. Despite significant advancements in the automatic recognition and understanding of human speech, challenges remain in accurately identifying and addressing the nuances of human emotions, hindering the development of more empathic virtual agents. We believe that empathic virtual agents should excel in three key tasks: (i) recognising spontaneous emotional expressions alongside understanding verbal content, (ii) generating timely and appropriate responses, and (iii) providing insightful feedback while comprehending user responses. To advance the development of empathic agents, we introduce the first Empathic Virtual Agent Challenge (EVAC). The inaugural edition focuses on robustly recognising spontaneous human expressions during interactions with a virtual agent, using the newly introduced THERADIA WoZ dataset. This paper provides an overview of the baseline systems operated on the pseudonymised version of the corpus on the two following modeling tasks: core affect presence and intensity, and appraisal based dimensions.
Fabien Ringeval, Björn W. Schuller, Gérard Bailly, Safaa Azzakhnini, Hippolyte Fournier
ICMI3
2024 Training speech-breathing coordination in computer-assisted reading
abstract
International audience
Delphine Charuau, Andrea Briglia, Erika Godde, Gérard Bailly
INTERSPEECH4
2024 FastLips: an End-to-End Audiovisual Text-to-Speech System with Lip Features Prediction for Virtual Avatars
abstract
International audience
Martin Lenglet, Olivier Perrotin, Gérard Bailly
INTERSPEECH3
2022 Speaking Rate Control of end-to-end TTS Models by Direct Manipulation of the Encoder's Output Embeddings
abstract
International audience
Martin Lenglet, Olivier Perrotin, Gérard Bailly
INTERSPEECH3
2022 Automatic Verbal Depiction of a Brick Assembly for a Robot Instructing Humans
abstract
Verbal and nonverbal communication skills are essential for human-robot interaction, in particular when the agents are involved in a shared task.We address the specific situation where the robot is the only agent knowing about both the plan and the goal of the task, and has to instruct the human partners.The case study is a brick assembly.We here describe a multilayered verbal depictor whose semantic, syntactic, and lexical settings have been collected and evaluated via crowdsourcing.One crowdsourced experiment involves a robot-instructed pick-and-place task.We show that implicitly referring to achieved subgoals (stairs, pillars, etc) increases the performance of human partners.
Rami Younes, Gérard Bailly, Frédéric Elisei, Damien Pellier
SIGDIAL2
2022 Automatic assessment of oral readings of young pupils
abstract
We propose a computational framework for estimating multidimensional subjective ratings of the reading performance of young readers from speech-based objective measures. We combine linguistic features (number of correct words, repetitions, deletions, insertions uttered per minute, etc.) with prosodic features. Expressivity is particularly difficult to predict since there is no unique gold standard. We propose a novel framework for performing such an estimation that exploits multiple references performed by adults and we demonstrate its effectiveness using recordings from a large data set of 1063 oral readings from 442 children (more than 30 h of speech), 84 oral readings from 42 adults and 6853 subjective scores delivered by 29 different human raters. We show that robust and accurate estimations of reading fluency can be achieved using combined features. This automatic assessment tool provides teachers and speech therapists with reliable estimates of the maturation of several reading skills.
Gérard Bailly, Erika Godde, Anne-Laure Piat-Marchand, Marie-Line Bosse
Speech Commun.1
2021 Evaluating the Extrapolation Capabilities of Neural Vocoders to Extreme Pitch Values
abstract
International audience
Olivier Perrotin, Hussein El Amouri, Gérard Bailly, Thomas Hueber
Interspeech3
2020 Predicting Multidimensional Subjective Ratings of Children' Readings from the Speech Signals for the Automatic Assessment of Fluency
abstract
The objective of this research is to estimate multidimensional subjective ratings of the reading performance of young readers from signal-based objective measures. We here combine linguistic features (number of correct words, repetitions, deletions, insertions uttered per minute . . . ) with phonetic features. Expressivity is particularly difficult to predict since there is no unique golden standard. We here propose a novel framework for performing such an estimation that exploits multiple references performed by adults and demonstrate its efficiency using recordings of 273 pupils.
Gérard Bailly, Erika Godde, Anne-Laure Piat-Marchand, Marie-Line Bosse
LREC1
2019 Transfer and Extraction of the Style of Handwritten Letters using Deep Learning
abstract
How can we learn, transfer and extract handwriting styles using deep neural networks? This paper explores these questions using a deep conditioned autoencoder on the IRON-OFF handwriting data-set. We perform three experiments that systematically explore the quality of our style extraction procedure. First, We compare our model to handwriting benchmarks using multidimensional performance metrics. Second, we explore the quality of style transfer, i.e. how the model performs on new, unseen writers. In both experiments, we improve the metrics of state of the art methods by a large margin. Lastly, we analyze the latent space of our model, and we show that it separates consistently writing styles.
Omar Mohammed, Gérard Bailly, Damien Pellier
ICAART (2)2
2018 A Weighted Superposition of Functional Contours Model for Modelling Contextual Prominence of Elementary Prosodic Contours
abstract
International audience
Branislav Gerazov, Gérard Bailly, Yi Xu 0007
INTERSPEECH2
2018 Audio-visual synchronization in reading while listening to texts: Effects on visual behavior and verbal learning
Emilie Gerbier, Gérard Bailly, Marie-Line Bosse
Comput. Speech Lang.2
2018 Introduction to the special issue on auditory-visual expressive speech and gesture in humans and machines
Jeesun Kim, Gérard Bailly, Chris Davis 0001
Speech Commun.2
2017 Learning off-line vs. on-line models of interactive multimodal behaviors with recurrent neural networks
Duc Canh Nguyen, Gérard Bailly, Frédéric Elisei
Pattern Recognit. Lett.2
2017 Which prosodic features contribute to the recognition of dramatic attitudes?
Adela Barbulescu, Rémi Ronfard, Gérard Bailly
Speech Commun.3
2016 Quantitative Analysis of Backchannels Uttered by an Interviewer During Neuropsychological Tests
abstract
International audience
Gérard Bailly, Frédéric Elisei, Alexandra Juphard, Olivier Moreaud
INTERSPEECH1
2016 Characterization of Audiovisual Dramatic Attitudes
abstract
International audience
Adela Barbulescu, Rémi Ronfard, Gérard Bailly
INTERSPEECH3
2016 Introduction to Poster Presentation of Part II
Jeesun Kim, Gérard Bailly
INTERSPEECH2
2016 Adaptive Latency for Part-of-Speech Tagging in Incremental Text-to-Speech Synthesis
abstract
International audience
Maël Pouget, Olha Nahorna, Thomas Hueber, Gérard Bailly
INTERSPEECH4
2016 Statistical conversion of silent articulation into audible speech using full-covariance HMM
Thomas Hueber, Gérard Bailly
Comput. Speech Lang.2
2016 Graphical models for social behavior modeling in face-to face interaction
Alaeddine Mihoub, Gérard Bailly, Christian Wolf 0001, Frédéric Elisei
Pattern Recognit. Lett.2
2015 HMM training strategy for incremental speech synthesis
abstract
Incremental speech synthesis aims at delivering the synthetic voice while the sentence is still being typed. One of the main challenges is the online estimation of the target prosody from a partial knowledge of the sentence's syntactic structure. In the context of HMM-based speech synthesis, this typically results in missing segmental and suprasegmental features, which describe the linguistic context of each phoneme. This study describes a voice training procedure which integrates explicitly a potential uncertainty on some contextual features. The proposed technique is compared to a baseline approach (previously published), which consists in substituting a missing contextual feature by a default value calculated on the training set. Both techniques were implemented in a HMM-based Text-To-Speech system for French, and compared using objective and perceptual measurements. Experimental results show that the proposed strategy outperforms the baseline technique for this language.
Maël Pouget, Thomas Hueber, Gérard Bailly, Timo Baumann
INTERSPEECH3
2015 Speaker-Adaptive Acoustic-Articulatory Inversion Using Cascaded Gaussian Mixture Regression
abstract
This paper addresses the adaptation of an acoustic-articulatory model of a reference speaker to the voice of another speaker, using a limited amount of audio-only data. In the context of pronunciation training, a virtual talking head displaying the internal speech articulators (e.g., the tongue) could be automatically animated by means of such a model using only the speaker's voice. In this study, the articulatory-acoustic relationship of the reference speaker is modeled by a gaussian mixture model (GMM). To address the speaker adaptation problem, we propose a new framework called cascaded Gaussian mixture regression (C-GMR), and derive two implementations. The first one, referred to as Split-C-GMR, is a straightforward chaining of two distinct GMRs: one mapping the acoustic features of the source speaker into the acoustic space of the reference speaker, and the other estimating the articulatory trajectories with the reference model. In the second implementation, referred to as Integrated-C-GMR, the two mapping steps are tied together in a single probabilistic model. For this latter model, we present the full derivation of the exact EM training algorithm, that explicitly exploits the missing data methodology of machine learning. Other adaptation schemes based on maximum-a posteriori (MAP), maximum likelihood linear regression (MLLR) and direct cross-speaker acoustic-to-articulatory GMR are also investigated. Experiments conducted on two speakers for different amount of adaptation data show the interest of the proposed C-GMR techniques.
Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, Gérard Bailly
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Modeling perception-action loops: comparing sequential models with frame-based classifiers
abstract
Modeling multimodal perception-action loops in face-to-face interactions is a crucial step in the process of building sensory-motor behaviors for social robots or users-aware Embodied Conversational Agents (ECA). In this paper, we compare trainable behavioral models based on sequential models (HMMs) and classifiers (SVMs and Decision Trees) inherently inappropriate to model sequential aspects. These models aim at giving pertinent perception/action skills for robots in order to generate optimal actions given the perceived actions of others and joint goals. We applied these models to parallel speech and gaze data collected from interacting dyads. The challenge was to predict the gaze of one subject given the gaze of the interlocutor and the voice activity of both. We show that Incremental Discrete HMM (IDHMM) generally outperforms classifiers and that injecting input context in the modeling process significantly improves the performances of all algorithms.
Alaeddine Mihoub, Gérard Bailly, Christian Wolf 0001
HAI2
2014 Assessing objective characterizations of phonetic convergence
abstract
This paper focuses on the study of the convergence between characteristics of speech segments- i.e. spectral characteristics of speech sounds - during live interactions between speaking dyads. The interaction data has been collected using an original verbal game called 'verbal dominoes' that provides a dense sampling of the acoustic spaces of the interlocutors. Two methods for characterizing phonetic convergence are here compared. The first one is based on a fine-grained analysis of the spectra of central frames of vowels (LDA) while the second one uses a more global speaker recognition technique (LLR). We show that convergence rates calculated by the two techniques correlate as the number of dominoes increases and that the LDA method well resists to the decrease of training and test material. We finally comment the impact of several factors on the computed convergence rates, i.e. interlocutors' familiarity and sex pairs.
Gérard Bailly, Amélie Martin
INTERSPEECH1
2014 Beyond basic emotions: expressive virtual actors with social attitudes
abstract
The purpose of this work is to evaluate the contribution of audio-visual prosody to the perception of complex mental states of virtual actors. We propose that global audio-visual prosodic contours - i.e. melody, rhythm and head movements over the utterance - constitute discriminant features for both the generation and recognition of social attitudes. The hypothesis is tested on an acted corpus of social attitudes in virtual actors and evaluation is done using objective measures and perceptual tests.
Adela Barbulescu, Rémi Ronfard, Gérard Bailly, Georges Gagneré, Hüseyin Çakmak
MIG3
2013 Adaptation of respiratory patterns in collaborative reading
abstract
Speech and variation of respiratory chest circumferences of eight French dyads were monitored while reading texts with increasing constraints on mutual synchrony. In line with previous research, we find that speakers mutually adapt their respiratory patterns. However a significant alignment is observed only when speakers need to perform together, i.e. when reading in alternation or synchronously. From quiet breathing to listening, to speech reading, we didn't find the gradual asymmetric shaping of respiratory cycles generally assumed in literature (e.g. from symmetric inhalation and exhalation phases towards short inhalation and long exhalation). In contrast, the control of breathing seems to switch abruptly between two systems: vital vs. speech production. We also find that the syllabic and the respiratory cycles are strongly phased at speech onsets. This phenomenon is in agreement with the quantal nature of speech rhythm beyond the utterance, previously observed via pause durations.
Gérard Bailly, Amélie Rochet-Capellan, Coriandre Vilain
INTERSPEECH1
2013 Speaker adaptation of an acoustic-articulatory inversion model using cascaded Gaussian mixture regressions
abstract
The article presents a method for adapting a GMM-based acoustic-articulatory inversion model trained on a reference speaker to another speaker. The goal is to estimate the articulatory trajectories in the geometrical space of a reference speaker from the speech audio signal of another speaker. This method is developed in the context of a system of visual biofeedback, aimed at pronunciation training. This system provides a speaker with visual information about his/her own articulation, via a 3D orofacial clone. In previous work, we proposed to use GMM-based voice conversion for speaker adaptation. Acoustic-articulatory mapping was achieved in 2 consecutive steps: 1) converting the spectral trajectories of the target speaker (i.e. the system user) into spectral trajectories of the reference speaker (voice conversion), and 2) estimating the most likely articulatory trajectories of the reference speaker from the converted spectral features (acoustic-articulatory inversion). In this work, we propose to combine these two steps into the same statistical mapping framework, by fusing multiple regressions based on trajectory GMM and maximum likelihood criterion (MLE). The proposed technique is compared to two standard speaker adaptation techniques based respectively on MAP and MLLR.
Thomas Hueber, Gérard Bailly, Pierre Badin, Frédéric Elisei
INTERSPEECH2
2012 Pauses and respiratory markers of the structure of book reading
abstract
The automatic reading of books by text-to-speech synthesizers requires not only the adequate encoding of the many levels of information and discourse structures in the acoustic signals but also the proper patterns of breathing, so that to pace information and organize discourse at an ecological rhythm. We analyze here the locations and durations of near 4,000 pauses produced by voice donor who has read several audiobooks, freely available on the web. Since the voice was recorded by a close microphone, we also characterized the acoustic markers of inhalation and show that the delay between end of phonation and air intake can be considered as an additional marker of thematic continuity between the two adjacent speech chunks that complements well-documented prosodic cues such as the preboundary tone and lengthening or the pause duration.
Gérard Bailly, Cécilia Gouvernayre
INTERSPEECH1
2012 Continuous Articulatory-to-Acoustic Mapping using Phone-based Trajectory HMM for a Silent Speech Interface
abstract
The article presents an HMM-based mapping approach for converting ultrasound and video images of the vocal tract into an audible speech signal, for a silent speech interface application. The proposed technique is based on the joint modeling of articulatory and spectral features, for each phonetic class, using Hidden Markov Models (HMM) and multivariate Gaussian distributions with full covariance matrices. The articulatory-toacoustic mapping is achieved in 2 steps: 1) finding the most likely HMM state sequence from the articulatory observations; 2) inferring the spectral trajectories from both the decoded state sequence and the articulatory observations. The proposed technique is compared to our previous approach, in which only the decoded state sequence was used for the inference of the spectral trajectories, independently from the articulatory observations. Both objective and perceptual evaluations show that this new approach leads to a better estimation of the spectral trajectories.
Thomas Hueber, Gérard Bailly, Bruce Denby
INTERSPEECH2
2012 Cross-speaker Acoustic-to-Articulatory Inversion using Phone-based Trajectory HMM for Pronunciation Training
abstract
The article presents a statistical mapping approach for crossspeaker acoustic-to-articulatory inversion. The goal is to estimate the most likely articulatory trajectories for a reference speaker from the speech audio signal of another speaker. This approach is developed in the framework of our system of visual articulatory feedback developed for computer-assisted pronunciation training applications (CAPT). The proposed technique is based on the joint modeling of articulatory and acoustic features, for each phonetic class, using full-covariance trajectory HMM. The acoustic-to-articulatory inversion is achieved in 2 steps: 1) finding the most likely HMM state sequence from the acoustic observations; 2) inferring the articulatory trajectories from both the decoded state sequence and the acoustic observations. The problem of speaker adaptation is addressed using a voice conversion approach, based on trajectory GMM. Index Terms: acoustic-to-articulatory inversion, intelligent tutoring systems, pronunciation training, trajectory HMM, voice
Thomas Hueber, Atef Ben Youssef, Gérard Bailly, Pierre Badin, Frédéric Elisei
INTERSPEECH3
2011 Synchronous Reading: Learning French Orthography by Audiovisual Training
abstract
International audience
Gérard Bailly, Will Barbour
INTERSPEECH1
2011 Toward a Multi-Speaker Visual Articulatory Feedback System
abstract
de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Atef Ben Youssef, Thomas Hueber, Pierre Badin, Gérard Bailly
INTERSPEECH4
2011 A pilot study on augmented speech communication based on Electro-Magnetic Articulography
Panikos Heracleous, Pierre Badin, Gérard Bailly, Norihiro Hagita
Pattern Recognit. Lett.3
2010 Exploiting multimodal data fusion in robust speech recognition
abstract
This article introduces automatic speech recognition based on Electro-Magnetic Articulography (EMA). Movements of the tongue, lips, and jaw are tracked by an EMA device, which are used as features to create Hidden Markov Models (HMM) and recognize speech only from articulation, that is, without any audio information. Also, automatic phoneme recognition experiments are conducted to examine the contribution of the EMA parameters to robust speech recognition. Using feature fusion, multistream HMM fusion, and late fusion methods, noisy audio speech has been integrated with EMA speech and recognition experiments have been conducted. The achieved results show that the integration of the EMA parameters significantly increases an audio speech recognizer's accuracy, in noisy environments.
Panikos Heracleous, Pierre Badin, Gérard Bailly, Norihiro Hagita
ICME3
2010 Speech dominoes and phonetic convergence
abstract
come from teaching and research institutions in France or abroad, or from public or private research centers.L'archive ouverte pluridisciplinaire
Gérard Bailly, Amélie Lelong
INTERSPEECH1
2010 Can tongue be recovered from face? the answer of data-driven statistical models
abstract
International audience
Atef Ben Youssef, Pierre Badin, Gérard Bailly
INTERSPEECH3
2010 Can you 'read' tongue movements? Evaluation of the contribution of tongue display to speech understanding
Pierre Badin, Yuliya Tarabalka, Frédéric Elisei, Gérard Bailly
Speech Commun.4
2010 Gaze, conversational agents and face-to-face communication
Gérard Bailly, Stephan Raidt, Frédéric Elisei
Speech Commun.1
2010 Speech and face-to-face communication - An introduction
Marion Dohen, Jean-Luc Schwartz, Gérard Bailly
Speech Commun.3
2010 Improvement to a NAM-captured whisper-to-speech system
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Tomoki Toda
Speech Commun.2
2009 Multimodal HMM-based NAM-to-speech conversion
abstract
Although the segmental intelligibility of converted speech from silent speech using direct signal-to-signal mapping proposed by Toda et al. [1] is quite acceptable, listeners have sometimes difficulty in chunking the speech continuum into meaningful words due to incomplete phonetic cues provided by output signals. This paper studies another approach consisting in combining HMM-based statistical speech recognition and synthesis techniques, as well as training on aligned corpora, to convert silent speech to audible voice. By introducing phonological constraints, such systems are expected to improve the phonetic consistency of output signals. Facial movements are used in order to improve the performance of both recognition and synthesis procedures. The results show that including these movements improves the recognition rate by 6.2% and a final improvement of the spectral distortion by 2.7% is observed. The comparison between direct signal-to-signal and phonetic-based mappings is finally commented in this paper.
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Tomoki Toda
INTERSPEECH2
2009 Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models
abstract
In order to recover the movements of usually hidden articulators such as tongue or velum, we have developed a data-based speech inversion method. HMMs are trained, in a multistream framework, from two synchronous streams: articulatory movements measured by EMA, and MFCC + energy from the speech signal. A speech recognition procedure based on the acoustic part of the HMMs delivers the chain of phonemes and together with their durations, information that is subsequently used by a trajectory formation procedure based on the articulatory part of the HMMs to synthesise the articulatory movements. The RMS reconstruction error ranged between 1.1 and 2. mm. Index Terms: Speech inversion, augmented speech, automatic speech recognition, HTK, Electro-Magnetic Articulography (EMA), hidden Markov model (HMM), trajectory formation, HTS.
Atef Ben Youssef, Pierre Badin, Gérard Bailly, Panikos Heracleous
INTERSPEECH3
2008 Can you "read tongue movements"?
abstract
Lip reading relies on visible articulators to ease audiovisual speech understanding.However, lips and face alone provide very incomplete phonetic information: the tongue, that is generally not entirely seen, carries an important part of the articulatory information not accessible through lip reading.The question was thus whether the direct and full vision of the tongue allows tongue reading.We have therefore generated a set of audiovisual VCV stimuli by controlling an audiovisual talking head that can display all speech articulators, including tongue, in an augmented speech mode, from articulators movements tracked on a speaker.These stimuli have been played to subjects in a series of audiovisual perception tests in various presentation conditions (audio signal alone, audiovisual signal with profile cutaway display with or without tongue, complete face), at various Signal-to-Noise Ratios.The results show a given implicit effect of tongue reading learning, a preference for the more ecological rendering of the complete face in comparison with the cutaway presentation, a predominance of lip reading over tongue reading, but the capability of tongue reading to take over when the audio signal is strongly degraded or absent.We conclude that these tongue reading capabilities could be used for applications in the domain of speech therapy for speech retarded children, perception and production rehabilitation of hearing impaired children, and pronunciation training for second language learners.
Pierre Badin, Yuliya Tarabalka, Frédéric Elisei, Gérard Bailly
INTERSPEECH4
2008 A trainable trajectory formation model TD-HMM parameterized for the LIPS 2008 challenge
abstract
We describe here the trainable trajectory formation model that will be used for the LIPS'2008 challenge organized at InterSpeech'2008.It predicts articulatory trajectories of a talking face from phonetic input.It basically uses HMMbased synthesis but asynchrony between acoustic and gestural boundaries -taking for example into account non audible anticipatory gestures -is handled by a phasing model that predicts the delays between the acoustic boundaries of allophones to be synthesized and the gestural boundaries of HMM triphones.The HMM triphones and the phasing model are trained simultaneously using an iterative analysissynthesis loop.Convergence is obtained within a few iterations.Using different motion capture data, we demonstrate here that the phasing model improves significantly the prediction error and captures subtle contextdependent anticipatory phenomena.
Gérard Bailly, Oxana Govokhina, Gaspard Breton, Frédéric Elisei, Christophe Savariaux
INTERSPEECH1
2008 From 3-d speaker cloning to text-to-audiovisual-speech
Sascha Fagel, Frédéric Elisei, Gérard Bailly
INTERSPEECH3
2008 LIPS2008: visual speech synthesis challenge
abstract
In this paper we present an overview of LIPS2008: Visual Speech Synthesis Challenge.The aim of this challenge is to bring together researchers in the field of visual speech synthesis to firstly evaluate their systems within a common framework, and secondly to identify the needs of the wider community in terms of evaluation.In doing so we hope to better understand the differences between the various approaches and to identify the strengths/weaknesses of the competing approaches.In this paper we firstly motivate the need for the challenge, before describing the capture and preparation of the training data, the evaluation framework, and conclude with an outline of possible directions for standardising the evaluation of talking heads.
Barry-John Theobald, Sascha Fagel, Gérard Bailly, Frédéric Elisei
INTERSPEECH3
2008 Improvement to a NAM captured whisper-to-speech system
abstract
In this paper, new techniques to improve whisper-to-speech conversion are investigated, in the framework of silent speech telephone communication. A preliminary conversion method from Non-Audible Murmur (NAM) to modal speech, based on statistical mapping trained using aligned corpora has been proposed. Although it is a very promising technique, its performance is still insufficient due to the difficulties in estimating F0 from unvoiced speech. In this paper, two distinct modifications are proposed, in order to improve the naturalness of the synthesized speech. In the first modification, LDA (Linear Discriminant Analysis) is used instead of PCA (Principal Component Analysis) to reduce the dimensionality of the input spectral vectors. In addition, the influence of long-term variation of spectral information on pitch estimation is examined. The second modification is an attempt to integrate visual information as a complementary input to improve spectral estimation, F0 estimation and voicing decision.
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Christian Jutten
INTERSPEECH2
2007 Scrutinizing Natural Scenes: Controlling the Gaze of an Embodied Conversational Agent
Antoine Picot, Gérard Bailly, Frédéric Elisei, Stephan Raidt
IVA2
2007 Analyzing Gaze During Face-to-Face Interaction
Stephan Raidt, Gérard Bailly, Frédéric Elisei
IVA2
2006 Generating German intonation with a trainable prosodic model
abstract
Abstract A trainable prosodic model called SFC (Superposition of Functional Contours), proposed by Holm and Bailly, is here confronted to German intonation. Training material is the publicly available Siemens Synthesis Corpus that provides spoken utterances for high-quality speech synthesis. We describe the labeling framework and first evaluation results that compares the original prosody of test sentences of this corpus with their prosodic rendering by the proposed model and state-of-the-art systems available on-line on the web. Index Terms : speech synthesis, prosody, evaluation Introduction The trainable prosodic model SFC (Superposition of Functional Contours) has been developed by Holm and Bailly [1-3]. It implements a theoretical model of intonation initially sketched by Auberge [4, 5] that promotes an intimate link between phonetic forms and linguistic functions: metalinguistic functions acting on different discourse units (thus at different scopes) are directly implemented as global multiparametric contours. These metalinguistic functions refer to the general ability of intonation to demarcate phonological units and convey information about the propositional and interactional func tions of these units within the discourse. This trainable prosodic model has been confronted to speech styles (from read speech to spoken maths) and different languages including French, Galician or more recently Chinese [6]. German is of most interest because of its rich morphology and its potentially deep recursive syntactic embedding. Analysis of German prosody notably induces Schreuder and Gilbers [7] to question the Strict Layer Hypothesis [8] and claim for the existence of recursive prosodic phrases. While most quantitative models of German intonation that have been so far applied to speech synthesis use a phonological representation with few levels when not limited to prosodic phrases [9, 10]. We describe here our first efforts in confronting the SFC - that may potentially capture rich embedded performance structures [11] – to German intonation. Our first parameterization of the SFC usi ng limited training material is evaluated against state-of-the-art text-to-speech systems available on the web.
Gérard Bailly, Jan Gorisch
INTERSPEECH1
2006 Evaluating a virtual speech cuer
abstract
This paper presents the virtual speech cuer built in the context of the ARTUS project aiming at watermarking hand and face gestures of a virtual animated agent in a broadcasted audiovisual sequence.For deaf televiewers that master cued speech, the animated agent can be then superimposed -on demand and at the reception -on the original broadcast as an alternative to subtitling.The paper presents the multimodal text-to-speech synthesis system and the first evaluation performed by deaf users.
Guillaume Gibert, Gérard Bailly, Frédéric Elisei
INTERSPEECH2
2006 TDA: a new trainable trajectory formation system for facial animation
Oxana Govokhina, Gérard Bailly, Gaspard Breton, Paul C. Bagshaw
INTERSPEECH2
2006 A joint prosody evaluation of French text-to-speech synthesis systems
Marie-Neige Garcia, Christophe d'Alessandro, Gérard Bailly, Philippe Boula de Mareüil, Michel Morel
LREC3
2006 A joint intelligibility evaluation of French text-to-speech synthesis systems: the EvaSy SUS/ACR campaign
Philippe Boula de Mareüil, Christophe d'Alessandro, Alexander Raake, Gérard Bailly, Marie-Neige Garcia, Michel Morel
LREC4
2006 Does a Virtual Talking Face Generate Proper Multimodal Cues to Draw User's Attention to Points of Interest?
Stephan Raidt, Gérard Bailly, Frédéric Elisei
LREC2
2006 Rackham: An Interactive Robot-Guide
abstract
Rackham is an interactive robot-guide that has been used in several places and exhibitions. This paper presents its design and reports on results that have been obtained after its deployment in a permanent exhibition. The project is conducted so as to incrementally enhance the robot functional and decisional capabilities based on the observation of the interaction between the public and the robot. Besides robustness and efficiency in the robot navigation abilities in a dynamic environment, our focus was to develop and test a methodology to integrate human-robot interaction abilities in a systematic way. We first present the robot and some of its key design issues. Then, we discuss a number of lessons that we have drawn from its use in interaction with the public and how that will serve to refine our design choices and to enhance robot efficiency and acceptability
Aurélie Clodic, Sara Fleury, Rachid Alami 0001, Raja Chatila 0001, Gérard Bailly, Ludovic Brethes, Maxime Cottret, Patrick Danès, Xavier Dollat, Frédéric Elisei, Isabelle Ferrané, Matthieu Herrb, Guillaume Infantes, Christian Lemaire, Frédéric Lerasle, Jérôme Manhes, Patrick Marcoul, Paulo Menezes 0001, Vincent Montreuil
RO-MAN5
2005 Statistical active model for mouth components segmentation
abstract
Mouth segmentation is an important issue which applies in many multimedia applications as speech reading, face synthesis, recognition or audiovisual communication. In this paper, we propose a method based on a statistical model of shape and appearance to detect the lips. To create the model, the outline of the lips and teeth has to be manually annotated with 30 key-points on a few visemes (450). Once the model has been trained on this set, it is used for segmentation. After a step to situate mouth corners, the goal is to find the parameters to fit the model to an unknown image. The originalities of this work are (a) an initialization step which broadly classify lip and skin pixels, (b) the mouth corners local model, and (c) the automatically extracted dynamic and static sampled-appearance which are well adapted to describe the mouth area and its components.
Pierre Gacon, Pierre-Yves Coulon, Gérard Bailly
ICASSP (2)3
2005 Evaluating the pronunciation of proper names by four French grapheme-to-phoneme converters
abstract
International Speech Communication Association (Isca) - International Astronautical Federation. ISBN : 13 9781604234480.
Philippe Boula de Mareüil, Christophe d'Alessandro, Gérard Bailly, Frédéric Béchet, Marie-Neige Garcia, Michel Morel, Romain Prudon, Jean Véronis
INTERSPEECH3
2005 SFC: A trainable prosodic model
Gérard Bailly, Bleicke Holm
Speech Commun.1
2004 A trainable prosodic model: learning the contours implementing communicative functions within a superpositional model of intonation
abstract
This paper introduces a new model-constrained, datadriven method to generate prosody from metalinguistic information. We refer here to the general ability of intonation to demarcate speech units and convey information about the propositional and interactional functions of these units within the discourse. Our strong hypotheses are that (1) these functions are directly implemented as prototypical prosodic contours that are coextensive to the unit(s) they apply to, (2) the prosody of the message is obtained by superposing and adding all the contributing contours. We describe here an analysis-by-synthesis scheme that consists in identifying these prototypical contours and separating out their contributions in the prosodic contours of the training data. We will show that such a trainable prosodic model generates faithful prosodic contours with very few prototypical movements.
Gérard Bailly, Bleicke Holm, Véronique Aubergé
INTERSPEECH1
2004 Audiovisual perceptual evaluation of resynthesised speech movements
Matthias Odisio, Gérard Bailly
INTERSPEECH2
2004 Evaluation of a Speech Cuer: From Motion Capture to a Concatenative Text-to-cued Speech System
Guillaume Gibert, Gérard Bailly, Frédéric Elisei, Denis Beautemps, Rémi Brun
LREC2
2004 Tracking talking faces with shape and appearance models
Matthias Odisio, Gérard Bailly, Frédéric Elisei
Speech Commun.2
2003 ISCA special session: hot topics in speech synthesis
Gérard Bailly, Nick Campbell 0001, Bernd Möbius
INTERSPEECH1
2002 Audiovisual speech synthesis. from ground truth to models
abstract
We present here the main approaches used to synthesize and drive talking faces. Illustrative systems are described. We distinguish between facial synthesis itself (i.e the manner in which facial movements are rendered on a computer screen), and the way these movements may be controlled and predicted using phonetic input. We then focus on the necessity to capture, model and render with maximum fidelity the intimate coherence of the facial deformations observed on a human face.
Gérard Bailly
INTERSPEECH1
2002 Seeing tongue movements from outside
Gérard Bailly, Pierre Badin
INTERSPEECH1
2001 Generating prosodic attitudes in French: Data, model and evaluation
Yann Morlec, Gérard Bailly, Véronique Aubergé
Speech Commun.2
2000 Generating prosody by superposing multi-parametric overlapping contours
abstract
We present here a model for generating prosody by superposing overlapping multi-parametric contours. These contours are associated with high-level communication tasks such as segmentation, hierarchisation or emphasis of discourse units. We propose a analysis-by-synthesis scheme for automatically learning these contours and apply this new paradigm to the enunciation of mathematical formulae and utterances carrying various attitudes. 1. INTRODUCTION It is a commonly accepted that prosody takes part in the transmission of linguistic information during the speech act. One of the most studied linguistic functions that prosody assumes is hierarchisation and segmentation of units. If there is a large consensus in literature concerning the major role of prosody in language acquisition and structuring of speech, authors diverge on the way prosody actually encodes information. In the framework of automatic prosody generation the aim is to establish an acceptable evolution of prosodic parameter...
Bleicke Holm, Gérard Bailly
INTERSPEECH2
2000 MOTHER: a new generation of talking heads providing a flexible articulatory control for video-realistic speech animation
abstract
International audience
Lionel Revéret, Gérard Bailly, Pierre Badin
INTERSPEECH2
2000 The Cost258 Signal Generation Test Array
Gérard Bailly, Eduardo Rodríguez Banga, Alex I. C. Monaghan, Erhard Rank
LREC1
1999 Accurate estimation of sinusoidal parameters in an harmonic+noise model for speech synthesis
abstract
A spoken dialog interface of a mobile office robot is described. To realize robust speech recognition in noisy office environments, a microphone array system and a technique of switching multiple speech recognition processes with different dictionaries are introduced. To realize flexible and natural dialog, task dependent semantic frames and keeping track of attentional state of dialog are used. The system is implemented on a real mobile robot and evaluated with sample dialogs.
Gérard Bailly
EUROSPEECH1
1999 Training an application-dependent prosodic model corpus, model and evaluation
Yann Morlec, Gérard Bailly, Véronique Aubergé
EUROSPEECH2
1998 A three-dimensional linear articulatory model based on MRI data
abstract
Based on a set of 3D vocal tract images obtained by MRI, a 3D statistical articulatory model has been built using guided Principal Component Analysis. It constitutes an extension to the lateral dimension of the mid-sagittal model previously developed from a radiofilm recorded on the same subject. The parameters of the 2D model have been found to be good predictors of the 3D shapes, for most configurations. A first evaluation of the model in terms of area functions and formants is presented.
Pierre Badin, Gérard Bailly, Monica Raybaudi, Christoph Segebarth
ICSLP2
1998 Synergy between jaw and lips/tongue movements : consequences in articulatory modelling
Gérard Bailly, Pierre Badin, Anne Vilain
ICSLP1
1998 Cooperation and competition of burst and formant transitions for the perception and identification of French stops
abstract
In this paper, we study the influence of the vocalic context on the perception and automatic recognition of stops. In a previous perception experiment [1] using conflicting cues stimuli, we have shown that place of articulation cued by formant transitions may be overwritten by the place cued by the burst. This effect is inversely proportional to the vowel aperture. Here we give special attention to /i/ context where nor burst, nor formant transitions seem to carry rich information on place of articulation. We present here automatic recognition experiments that confirm perception results. Taking into account both segments increase identification rates, early fusion of segmental cues performs best and most errors come from the front unrounded vocalic context. We introduce the "burst characteristic frequency" (BF) that palliates for the poor discriminative power of the traditional cues in the front context. Moreover we present perception results showing the perceptual relevance of BF. ...
Adrian Neagu, Gérard Bailly
ICSLP2
1998 Evaluation of grapheme-to phoneme conversion for text-to-speech synthesis in French
Philippe Boula de Mareüil, François Yvon, Christophe d'Alessandro, V. Auberg, Michel Bagein, Gérard Bailly, Frédéric Béchet, S. Fonkia, Jean-Philippe Goldman, Eric Keller, Douglas D. O'Shaughnessy, Steve Pagel, F. Sannier, Jean Véronis, Brigitte Zellner Keller
LREC6
1998 Evaluating the adeqnacy of synthetic prosody in signaling syntactic boundaries: methodology and first results
Yann Morlec, Albert Rilliard, Gérard Bailly, Véronique Aubergé
LREC3
1998 Objective evaluation of grapheme to phoneme conversion for text-to-speech synthesis in French
François Yvon, Philippe Boula de Mareüil, Christophe d'Alessandro, Véronique Aubergé, Michel Bagein, Gérard Bailly, Frédéric Béchet, S. Foukia, J.-F. Goldman, Eric Keller, Douglas D. O'Shaughnessy, Vincent Pagel, Fred Sannier, Jean Véronis, Brigitte Zellner
Comput. Speech Lang.6
1997 Synthesis of fricative consonants by audiovisual-to-articulatory inversion
abstract
We present here results of audio-visual to articulatory inversion for French fricatives embedded into VCVs. The inversion technique is evaluated using both experimental and synthetic data. The final synthesis is assessed by a perceptual categorisation test. Synthetic stimuli have similar scores as natural ones.
Khaled Mawass, Pierre Badin, Gérard Bailly
EUROSPEECH3
1997 Synthesising attitudes with global rhythmic and intonation contours
abstract
We present here a trainable generative model of French prosody. We focus on the sentence level and design SNNs able to generate both rhythmic and intonation contours for diverse attitudes. First results of a perceptual test show that listeners are able to retrieve the right definition of attitudes by listening to synthetic PSOLA stimuli.
Yann Morlec, Gérard Bailly, Véronique Aubergé
EUROSPEECH2
1997 Relative contributions of noise burst and vocalic transitions to the perceptual identification of stop consonants
abstract
A set of three perceptual experiments is described. These experiments were designed to provide identification scores on CV sequences for French. Original stimuli were augmented with acoustic "monsters" where burst were excised or replaced. The first identification task shows that information carried by vocalic transitions can be overwritten by burst information. The importance of this phenomenon is inversely proportional to vowel aperture. The second experiment shows that these results are almost insensitive to relative amplitudes between the burst and the vowel. In the third experiment we manipulated the voice onset time (VOT) of the monsters using high quality analysis-resynthesis. Stimuli with a very short VOT were perceived as bilabials but VOT manipulation did not affect the /t/-/k/ confusions. These experiments claim for a dynamic model of stop identification where burst and vocalic transitions both contribute and compete to the phonetic decision. 1. INTRODUCTION The quest for ...
Adrian Neagu, Gérard Bailly
EUROSPEECH2
1997 Learning to speak. Sensori-motor control of speech movements
Gérard Bailly
Speech Commun.1
1996 Building sensori-motor prototypes from audiovisual exemplars
Gérard Bailly
ICSLP1
1996 Generating intonation by superposing gestures
Yann Morlec, Gérard Bailly, Véronique Aubergé
ICSLP2
1995 Generation of intonation: a global approach
Véronique Aubergé, Gérard Bailly
EUROSPEECH2
1995 Articulatori-acoustic vowel prototypes for speech production
Gérard Bailly, Louis-Jean Boë, Nathalie Vallée, Pierre Badin
EUROSPEECH1
1995 Synthesis and evaluation of intonation with a superposition model
abstract
A data-driven method based on a new paradigm is introduced in this paper. We assume that cognitive representations of the discourse are prosodically encoded by means of global multiparametric prototypes. The generation of adequate prosodic contours is then obtained by retrieving and combining these elementary prototypic contours accessed by linguistic or paralinguistic keys. We examine here F0 generation. Two procedures have been applied: the first consists of a superposition model via a structured lexicon, the second uses a recurrent neural network. Preliminary experiments shows that both methods can lead to good quality F0 contours. 1. INTRODUCTION Two main approaches are currently used to describe intonation: . The classical method lies in building phonological representations from phonetic constructs with a bottom-up analysis. These phonetic constructs include both global and local elements. Global elements consist of templates such as declination lines [1] or tonal grids [2]. L...
Yann Morlec, Gérard Bailly, Véronique Aubergé
EUROSPEECH2
1994 Characterisation of rhythmic patterns for text-to-speech synthesis
Plínio Barbosa, Gérard Bailly
Speech Communication2
1993 COMPOST: a client-server model for applications using text-to-speech systems
abstract
This article presents a Client-Server Model for multilingual text-to-speech synthesis. The server maintains a collection of TTS systems together with related recongurable descriptions, called scenarios. Applications of an authorized client can access to this collection via an Ethernet network on a simple request to the server. This server allows the client to customize the TTS processing (language, speaker, speech rate, intonation. . . ) to its requirements by switching between different systems and/or reconfiguring the one it is currently using. The working environment, called COMPOST, has a three layered architecture: the development layer including a powerfull rule-compiler [3] and language-independent processing facilities (linguistic analyzers, PSOLA and Klatt synthesizers . . . ), the system construction layer including the Scenario Definition Language, and the server layer which has two main components: the process manager and the ressource manager.
Mamoun Alissali, Gérard Bailly
EUROSPEECH2
1993 Resonances as possible representation of speech in the auditory-to-articulatory transform
abstract
This article presents a characterisation of formant trajectories based on an active tracking of each resonance of the vocal tract. This tracking is necessary due to the presence of focal points i.e. changes of affiliations between F2 and F3 in most of transitions between back and front sounds. We show that this tracking can be easely implemented without any analysis-by-synthesis process by assuming a minimum elasticity principle (smoothness) and a monotonous relation between resonances associated with the same cavity. Our proposal is supported by intensive analysis of natural speech, articulatory simulations and perceptual experiments. Keywords: Perception, Formants, Vocal Tract, Resonances. 1 INTRODUCTION Formants are widely used for the acoustic characterization of sounds in phonetic science. Their effective tracking by the human ear faces however many obstacles. If the presence of the spectral pics in the auditory representations have been attested by the neurophysiology (cf. [11]...
Gérard Bailly
EUROSPEECH1
1991 Synthesis-by-rule using compost: modelling resonance trajectories
M. Guerti, Gérard Bailly
EUROSPEECH2
1990 Automatic segmentation and alignment of continuous speech based on temporal decomposition model
Hai-Dong Wang, Gérard Bailly, Denis Tuffelli
ICSLP2
1989 A new algorithm for temporal decomposition of speech-application to a numerical model of coarticulation
abstract
The authors propose an algorithm based on a constrained iterative optimization process using a gradient method. The constraints can be applied on the time as well as the spectral dimension. Without any temporal constraints the algorithm produces relatively compact functions which exhibit a secondary lobe structure. Applying B.S. Atal's (1983) method each important secondary lobe is modeled by a different parametric target and associated compact function. After presenting the new decomposition technique, the authors focus on an experiment where they have been able automatically to infer directly from the speech signal articulatory gestures which intervene in a stimulus like /iyiy/ by means of an articulation with two degrees of freedom (lip protrusion versus lip opening). This study reports the first step toward a numerical model of articulatory inversion.>
Gérard Bailly, Pierre-François Marteau, Christian Abry
ICASSP1
1989 Compost: a rule-compiler for speech synthesis
Gérard Bailly, A. Tran
EUROSPEECH1
1989 Integration of rhythmic and syntactic constraints in a model of generation of French prosody
Gérard Bailly
Speech Commun.1
1988 Stochastic model of diphone-like segments based on trajectory concepts
abstract
A new global approach to coarse classification of speech segments is presented. Markov modeling is applied on an analytic approach to coarticulation. Speech signal evolution of diphone-like segments is modelized by a point moving frame by frame in a factorial space. Kinematic segmentation applied to the trajectory covered by this point enables the authors to build stochastic models of these segments. Input parameters of a Markov model are extracted from a skeleton of this trajectory considered as a functional model of overlapping segments. The evaluation of such representations in a recognition task gives some elements of discussion about the relative information contained in steady states versus transient segments and acoustical trajectories in general.>
Pierre-François Marteau, Gérard Bailly, M. T. Janot-Giorgetti
ICASSP2
1986 Multiparametric generation of French prosody from unrestricted text
abstract
Automatic assignment of prosodic parameters for a text-to-speech synthesis system for French is presented. A structural linguistic preprocessing using a lexical processor and a pre-syntactic analyser parses the text into sense groups which are assigned prosodic markers. A multiparametric qualitative model of French prosody generates the temporal. pausal, intonative and intensity structure of each utterance, using Fujisaki's source command model. Tested with the INRS text-to-speech system, this prosodic generator increases naturalness for naive listeners.
Gérard Bailly
ICASSP1