EDBT 2026 Demo / reviewers in the wild / expert
Thomas Hueber
dblp:08/8014
· DBLP profile ↗
45ranked-venue papers
14as first author
14since 2021 · last 2026
0000-0002-8296-5177ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 12 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 11 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units DiscoveryabstractThis paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models. Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber, Emmanuel Dupoux |
ACL (1) | 4 |
| 2026 | Is self-supervised learning enough to fill in the gap? A study on speech inpaintingabstractSpeech inpainting consists in reconstructing corrupted or missing speech segments using surrounding context, a process that closely resembles the pretext tasks in Self-Supervised Learning (SSL) for speech encoders. This study investigates using SSL-trained speech encoders for inpainting without any additional training beyond the initial pretext task, and simply adding a decoder to generate a waveform. We compare this approach to supervised fine-tuning of speech encoders for a downstream task—here, inpainting. Practically, we integrate HuBERT as the SSL encoder and HiFi-GAN as the decoder in two configurations: (1) fine-tuning the decoder to align with the frozen pre-trained encoder’s output and (2) fine-tuning the encoder for an inpainting task based on a frozen decoder’s input. Evaluations are conducted under single- and multi-speaker conditions using in-domain datasets and out-of-domain datasets (including unseen speakers, diverse speaking styles, and noise). Both informed and blind inpainting scenarios are considered, where the position of the corrupted segment is either known or unknown. The proposed SSL-based methods are benchmarked against several baselines, including a text-informed method combining automatic speech recognition with zero-shot text-to-speech synthesis. Performance is assessed using objective metrics and perceptual evaluations. The results demonstrate that both approaches outperform baselines, successfully reconstructing speech segments up to 200 ms, and sometimes up to 400 ms. Notably, fine-tuning the SSL encoder achieves more accurate speech reconstruction in single-speaker settings, while a pre-trained encoder proves more effective for multi-speaker scenarios. This demonstrates that an SSL pretext task can transfer to speech inpainting, enabling successful speech reconstruction with a pre-trained encoder. Ihab Asaad, Maxime Jacquelin, Olivier Perrotin, Laurent Girin, Thomas Hueber |
Comput. Speech Lang. | 5 |
| 2025 | From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelabstractHuman infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction.We present a computational model that addresses the acoustic-to-articulatory mapping problem through self-supervised learning.Our model comprises a feature extractor that transforms speech into latent representations, an inverse model that maps these representations to articulatory parameters, and a synthesizer that generates speech outputs.Experiments conducted in both single-and multi-speaker settings reveal that intermediate layers of a pretrained wav2vec 2.0 model provide optimal representations for articulatory learning, significantly outperforming MFCC features.These representations enable our model to learn articulatory trajectories that correlate with human patterns, discriminate between places of articulation, and produce intelligible speech.Critical to successful articulatory learning are representations that balance phonetic discriminability with speaker invariance -precisely the characteristics of self-supervised representation learning models.Our findings provide computational evidence consistent with developmental theories proposing that perceptual learning of phonetic categories guides articulatory development, offering insights into how infants might acquire speech production capabilities despite the complex mapping problem they face. Marvin Lavechin, Thomas Hueber |
EMNLP | 2 |
| 2025 | Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech ModelabstractThis paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by leveraging a pre-trained audiovisual autoregressive text-to-speech model (AVTacotron2). This model is reprogrammed to infer Cued Speech (CS) hand and lip movements from text input. Experiments are conducted on two publicly available datasets, including one recorded specifically for this study. Performance is assessed using an automatic CS recognition system. With a decoding accuracy at the phonetic level reaching approximately 77%, the results demonstrate the effectiveness of our approach. Sanjana Sankar, Martin Lenglet, Gérard Bailly, Denis Beautemps, Thomas Hueber |
ICASSP | 5 |
| 2024 | Simulating articulatory trajectories with phonological feature interpolationabstractAs a first step towards a complete computational model of speech learning involving perception-production loops, we investigate the forward mapping between pseudo-motor commands and articulatory trajectories. Two phonological feature sets, based respectively on generative and articulatory phonology, are used to encode a phonetic target sequence. Different interpolation techniques are compared to generate smooth trajectories in these feature spaces, with a potential optimisation of the target value and timing to capture co-articulation effects. We report the Pearson correlation between a linear projection of the generated trajectories and articulatory data derived from a multi-speaker dataset of electromagnetic articulography (EMA) recordings. A correlation of 0.67 is obtained with an extended feature set based on generative phonology and a linear interpolation technique. We discuss the implications of our results for our understanding of the dynamics of biological motion. Angelo Ortiz Tandazo, Thomas Schatz, Thomas Hueber, Emmanuel Dupoux |
INTERSPEECH | 3 |
| 2023 | Investigating the dynamics of hand and lips in French Cued Speech using attention mechanisms and CTC-based decodingabstractHard of hearing or profoundly deaf people make use of cued speech (CS) as a communication tool to understand spoken language. By delivering cues that are relevant to the phonetic information, CS offers a way to enhance lipreading. In literature, there have been several studies on the dynamics between the hand and the lips in the context of human production. This article proposes a way to investigate how a neural network learns this relation for a single speaker while performing a recognition task using attention mechanisms. Further, an analysis of the learnt dynamics is utilized to establish the relationship between the two modalities and extract automatic segments. For the purpose of this study, a new dataset has been recorded for French CS. Along with the release of this dataset, a benchmark will be reported for word-level recognition, a novelty in the automatic recognition of French CS. Sanjana Sankar, Denis Beautemps, Frédéric Elisei, Olivier Perrotin, Thomas Hueber |
INTERSPEECH | 5 |
| 2022 | Repeat after Me: Self-Supervised Learning of Acoustic-to-Articulatory Mapping by Vocal ImitationabstractWe propose a computational model of speech production combining a pre-trained neural articulatory synthesizer able to reproduce complex speech stimuli from a limited set of interpretable articulatory parameters, a DNN-based internal forward model predicting the sensory consequences of articulatory commands, and an internal inverse model based on a recurrent neural network recovering articulatory commands from the acoustic speech input. Both forward and inverse models are jointly trained in a self-supervised way from raw acoustic-only speech data from different speakers. The imitation simulations are evaluated objectively and subjectively and display quite encouraging performances. Marc-Antoine Georges, Julien Diard, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber |
ICASSP | 5 |
| 2022 | Multistream Neural Architectures for Cued Speech Recognition Using a Pre-Trained Visual Feature Extractor and Constrained CTC DecodingabstractThis paper proposes a simple and effective approach for automatic recognition of Cued Speech (CS), a visual communication tool that helps people with hearing impairment to understand spoken language with the help of hand gestures that can uniquely identify the uttered phonemes in complement to lip-reading. The proposed approach is based on a pre-trained hand and lips tracker used for visual feature extraction and a phonetic decoder based on a multistream recurrent neural network trained with connectionist temporal classification loss and combined with a pronunciation lexicon. The proposed system is evaluated on an updated version of the French CS dataset CSF18 for which the phonetic transcription has been manually checked and corrected. With a decoding accuracy at the phonetic level of 70.88%, the proposed system outperforms our previous CNN-HMM decoder and competes with more complex baselines. Sanjana Sankar, Denis Beautemps, Thomas Hueber |
ICASSP | 3 |
| 2022 | Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAEabstractInternational audience Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber |
INTERSPEECH | 3 |
| 2022 | BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language modelabstractInternational audience Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber |
INTERSPEECH | 4 |
| 2021 | A Benchmark of Dynamical Variational Autoencoders Applied to Speech Spectrogram ModelingabstractAccepted to Interspeech 2021. arXiv admin note: text overlap with arXiv:2008.12595 Xiaoyu Bie, Laurent Girin, Simon Leglaive, Thomas Hueber, Xavier Alameda-Pineda |
Interspeech | 4 |
| 2021 | Learning Robust Speech Representation with an Articulatory-Regularized Variational AutoencoderabstractIt is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task. Marc-Antoine Georges, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber |
Interspeech | 4 |
| 2021 | Evaluating the Extrapolation Capabilities of Neural Vocoders to Extreme Pitch ValuesabstractInternational audience Olivier Perrotin, Hussein El Amouri, Gérard Bailly, Thomas Hueber |
Interspeech | 4 |
| 2021 | Alternate Endings: Improving Prosody for Incremental Neural TTS with Predicted Future Text InputabstractThe prosody of a spoken word is determined by its surrounding context. In incremental text-to-speech synthesis, where the synthesizer produces an output before it has access to the complete input, the full context is often unknown which can result in a loss of naturalness in the synthesized speech. In this paper, we investigate whether the use of predicted future text can attenuate this loss. We compare several test conditions of next future word: (a) unknown (zero-word), (b) language model predicted, (c) randomly predicted and (d) ground-truth. We measure the prosodic features (pitch, energy and duration) and find that predicted text provides significant improvements over a zero-word lookahead, but only slight gains over random-word lookahead. We confirm these results with a perceptive test. Brooke Stephenson, Thomas Hueber, Laurent Girin, Laurent Besacier |
Interspeech | 2 |
| 2020 | What the Future Brings: Investigating the Impact of Lookahead for Incremental Neural TTSabstractInternational audience Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber |
INTERSPEECH | 4 |
| 2020 | Evaluating the Potential Gain of Auditory and Audiovisual Speech-Predictive Coding Using Deep LearningabstractSensory processing is increasingly conceived in a predictive framework in which neurons would constantly process the error signal resulting from the comparison of expected and observed stimuli. Surprisingly, few data exist on the accuracy of predictions that can be computed in real sensory scenes. Here, we focus on the sensory processing of auditory and audiovisual speech. We propose a set of computational models based on artificial neural networks (mixing deep feedforward and convolutional networks), which are trained to predict future audio observations from present and past audio or audiovisual observations (i.e., including lip movements). Those predictions exploit purely local phonetic regularities with no explicit call to higher linguistic levels. Experiments are conducted on the multispeaker LibriSpeech audio speech database (around 100 hours) and on the NTCD-TIMIT audiovisual speech database (around 7 hours). They appear to be efficient in a short temporal range (25-50 ms), predicting 50% to 75% of the variance of the incoming stimulus, which could result in potentially saving up to three-quarters of the processing power. Then they quickly decrease and almost vanish after 250 ms. Adding information on the lips slightly improves predictions, with a 5% to 10% increase in explained variance. Interestingly the visual gain vanishes more slowly, and the gain is maximum for a delay of 75 ms between image and predicted sound. Thomas Hueber, Eric Tatulli, Laurent Girin, Jean-Luc Schwartz |
Neural Comput. | 1 |
| 2018 | Visual Recognition of Continuous Cued Speech Using a Tandem CNN-HMM ApproachabstractInternational audience Li Liu 0036, Thomas Hueber, Gang Feng 0002, Denis Beautemps |
INTERSPEECH | 2 |
| 2017 | Feature extraction using multimodal convolutional neural networks for visual speech recognitionabstractThis article addresses the problem of continuous speech recognition from visual information only, without exploiting any audio signal. Our approach combines a video camera and an ultrasound imaging system for monitoring simultaneously the speaker's lips and the movement of the tongue. We investigate the use of convolutional neural networks (CNN) to extract visual features directly from the raw ultrasound and video images. We propose different architectures among which a multimodal CNN processing jointly the two visual modalities. Combined with an HMM-GMM decoder, the CNN-based approach outperforms our previous baseline based on Principal Component Analysis. Importantly, the recognition accuracy is only 4% lower than the one obtained when decoding the audio signal, which makes it a good candidate for a practical visual speech recognition system. Eric Tatulli, Thomas Hueber |
ICASSP | 2 |
| 2017 | Automatic animation of an articulatory tongue model from ultrasound images of the vocal tract
Diandra Fabre, Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, Pierre Badin |
Speech Commun. | 2 |
| 2017 | Extending the Cascaded Gaussian Mixture Regression Framework for Cross-Speaker Acoustic-Articulatory MappingabstractThis paper addresses the adaptation of an acoustic-articulatory inversion model of a reference speaker to the voice of another source speaker, using a limited amount of audio-only data. In this study, the articulatory-acoustic relationship of the reference speaker is modeled by a Gaussian mixture model and inference of articulatory data from acoustic data is made by the associated Gaussian mixture regression (GMR). To address speaker adaptation, we previously proposed a general framework called Cascaded-GMR (C-GMR) which decomposes the adaptation process into two consecutive steps: spectral conversion between source and reference speaker and acoustic-articulatory inversion of converted spectral trajectories. In particular, we proposed the integrated C-GMR technique (IC-GMR) in which both steps are tied together in the same probabilistic model. In this paper, we extend the C-GMR framework with another model called Joint-GMR (J-GMR). Contrary to the IC-GMR, this model aims at exploiting all potential acoustic-articulatory relationships, including those between the source speaker's acoustics and the reference speaker's articulation. We present the full derivation of the exact expectation-maximization (EM) training algorithm for the J-GMR. It exploits the missing data methodology of machine learning to deal with limited adaptation data. We provide an extensive evaluation of the J-GMR on both synthetic acoustic-articulatory data and on the multispeaker MOCHA EMA database. We compare the J-GMR performance to other models of the C-GMR framework, notably the IC-GMR, and discuss their respective merits. Laurent Girin, Thomas Hueber, Xavier Alameda-Pineda |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Introduction to the Special Issue on Biosignal-Based Spoken CommunicationabstractThe papers in this special section focus on biosignal-based spoken communication. Speech production is a complex process resulting from human activities initiated in the brain, eventually leading to muscle activities that produce respiratory, laryngeal, and articulatory gestures which finally create acoustic signals. Traditional speech processing systems capture and interpret the acoustic signal of speech. However, speech is not only limited to acoustics – speech-related activities can be measured at each level of speech processing, including the central and peripheral nervous systems, muscular action potentials, and speech kinematics. Their measurement, obtained through recordings from variety of sensor technologies, results in speech-related “biosignals” that have been studied for decades to better understand the underlying mechanisms of human speech processing. However, there is more: speech-related biosignals have the potential to overcome limitations of traditional acoustic-based systems for spoken communication. Biosignals can be captured before the airborne acoustic signal and are thus less prone to environmental noise. Also, they do not rely on the production of audible speech - both features open up newtracks for “Biosignalbased Spoken Communication”. Examples of these tracks include Brain-Computer Interfaces allowing for communication by directly decoding cortical brain activity into speech representations, and Silent-Speech Interfaces, which offer a way to communicate privately without disturbing bystanders and to restore spoken communication for people who lost their voice due to severe speech impairments. Furthermore, biosignals could provide valuable articulatory biofeedback to speakers about their own voice production for increasing articulatory awareness in speech therapy or language learning. Tanja Schultz, Thomas Hueber, Dean J. Krusienski, Jonathan S. Brumberg |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Biosignal-Based Spoken Communication: A SurveyabstractSpeech is a complex process involving a wide range of biosignals, including but not limited to acoustics. These biosignals-stemming from the articulators, the articulator muscle activities, the neural pathways, and the brain itself-can be used to circumvent limitations of conventional speech processing in particular, and to gain insights into the process of speech production in general. Research on biosignal-based speech processing is a wide and very active field at the intersection of various disciplines, ranging from engineering, computer science, electronics and machine learning to medicine, neuroscience, physiology, and psychology. Consequently, a variety of methods and approaches have been used to investigate the common goal of creating biosignal-based speech processing devices for communication applications in everyday situations and for speech rehabilitation, as well as gaining a deeper understanding of spoken communication. This paper gives an overview of the various modalities, research approaches, and objectives for biosignal-based spoken communication. Tanja Schultz, Michael Wand 0002, Thomas Hueber, Dean J. Krusienski, Christian Herff, Jonathan S. Brumberg |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Adaptive Latency for Part-of-Speech Tagging in Incremental Text-to-Speech SynthesisabstractInternational audience Maël Pouget, Olha Nahorna, Thomas Hueber, Gérard Bailly |
INTERSPEECH | 3 |
| 2016 | Statistical conversion of silent articulation into audible speech using full-covariance HMM
Thomas Hueber, Gérard Bailly |
Comput. Speech Lang. | 1 |
| 2016 | Real-Time Control of an Articulatory-Based Speech Synthesizer for Brain Computer InterfacesabstractRestoring natural speech in paralyzed and aphasic people could be achieved using a Brain-Computer Interface (BCI) controlling a speech synthesizer in real-time. To reach this goal, a prerequisite is to develop a speech synthesizer producing intelligible speech in real-time with a reasonable number of control parameters. We present here an articulatory-based speech synthesizer that can be controlled in real-time for future BCI applications. This synthesizer converts movements of the main speech articulators (tongue, jaw, velum, and lips) into intelligible speech. The articulatory-to-acoustic mapping is performed using a deep neural network (DNN) trained on electromagnetic articulography (EMA) data recorded on a reference speaker synchronously with the produced speech signal. This DNN is then used in both offline and online modes to map the position of sensors glued on different speech articulators into acoustic parameters that are further converted into an audio signal using a vocoder. In offline mode, highly intelligible speech could be obtained as assessed by perceptual evaluation performed by 12 listeners. Then, to anticipate future BCI applications, we further assessed the real-time control of the synthesizer by both the reference speaker and new speakers, in a closed-loop paradigm using EMA data recorded in real time. A short calibration period was used to compensate for differences in sensor positions and articulatory differences between new speakers and the reference speaker. We found that real-time synthesis of vowels and consonants was possible with good intelligibility. In conclusion, these results open to future speech BCI applications using such articulatory-based speech synthesizer. Florent Bocquelet, Thomas Hueber, Laurent Girin, Christophe Savariaux, Blaise Yvert |
PLoS Comput. Biol. | 2 |
| 2015 | Real-time control of a DNN-based articulatory synthesizer for silent speech conversion: a pilot studyabstractInternational audience Florent Bocquelet, Thomas Hueber, Laurent Girin, Christophe Savariaux, Blaise Yvert |
INTERSPEECH | 2 |
| 2015 | Tongue tracking in ultrasound images using eigentongue decomposition and artificial neural networksabstractInternational audience Diandra Fabre, Thomas Hueber, Florent Bocquelet, Pierre Badin |
INTERSPEECH | 2 |
| 2015 | HMM training strategy for incremental speech synthesisabstractIncremental speech synthesis aims at delivering the synthetic voice while the sentence is still being typed. One of the main challenges is the online estimation of the target prosody from a partial knowledge of the sentence's syntactic structure. In the context of HMM-based speech synthesis, this typically results in missing segmental and suprasegmental features, which describe the linguistic context of each phoneme. This study describes a voice training procedure which integrates explicitly a potential uncertainty on some contextual features. The proposed technique is compared to a baseline approach (previously published), which consists in substituting a missing contextual feature by a default value calculated on the training set. Both techniques were implemented in a HMM-based Text-To-Speech system for French, and compared using objective and perceptual measurements. Experimental results show that the proposed strategy outperforms the baseline technique for this language. Maël Pouget, Thomas Hueber, Gérard Bailly, Timo Baumann |
INTERSPEECH | 2 |
| 2015 | Speaker-Adaptive Acoustic-Articulatory Inversion Using Cascaded Gaussian Mixture RegressionabstractThis paper addresses the adaptation of an acoustic-articulatory model of a reference speaker to the voice of another speaker, using a limited amount of audio-only data. In the context of pronunciation training, a virtual talking head displaying the internal speech articulators (e.g., the tongue) could be automatically animated by means of such a model using only the speaker's voice. In this study, the articulatory-acoustic relationship of the reference speaker is modeled by a gaussian mixture model (GMM). To address the speaker adaptation problem, we propose a new framework called cascaded Gaussian mixture regression (C-GMR), and derive two implementations. The first one, referred to as Split-C-GMR, is a straightforward chaining of two distinct GMRs: one mapping the acoustic features of the source speaker into the acoustic space of the reference speaker, and the other estimating the articulatory trajectories with the reference model. In the second implementation, referred to as Integrated-C-GMR, the two mapping steps are tied together in a single probabilistic model. For this latter model, we present the full derivation of the exact EM training algorithm, that explicitly exploits the missing data methodology of machine learning. Other adaptation schemes based on maximum-a posteriori (MAP), maximum likelihood linear regression (MLLR) and direct cross-speaker acoustic-to-articulatory GMR are also investigated. Experiments conducted on two speakers for different amount of adaptation data show the interest of the proposed C-GMR techniques. Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, Gérard Bailly |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Robust articulatory speech synthesis using deep neural networks for BCI applicationsabstractBrain-Computer Interfaces (BCIs) usually propose typing strategies to restore communication for paralyzed and aphasic people. A more natural way would be to use speech BCI directly controlling a speech synthesizer. Toward this goal, a prerequisite is the development a synthesizer that should i) produce intelligible speech, ii) run in real time, iii) depend on as few parameters as possible, and iv) be robust to error fluctuations on the control parameters. In this context, we describe here an articulatory-to-acoustic mapping approach based on deep neural network (DNN) trained on electromagnetic articulography (EMA) data recorded synchronously with produced speech sounds. On this corpus, the DNN-based model provided a speech synthesis quality (as assessed by automatic speech recognition and behavioral testing) comparable to a state-of-the-art Gaussian mixture model (GMM), yet showing higher robustness when noise was added to the EMA coordinates. Moreover, to envision BCI applications, this robustness was also assessed when the space covered by the 12 original articulatory parameters was reduced to 7 parameters using deep auto-encoders (DAE). Given that this method can be implemented in real time, DNN-based articulatory speech synthesis seems a good candidate for speech BCI applications. Index Terms: articulatory speech synthesis, brain computer interface (BCI), deep neural networks, deep auto-encoder, EMA, noise robustness, dimensionality reduction Florent Bocquelet, Thomas Hueber, Laurent Girin, Pierre Badin, Blaise Yvert |
INTERSPEECH | 2 |
| 2014 | Automatic animation of an articulatory tongue model from ultrasound images using Gaussian mixture regressionabstractThis paper presents a method for automatically animating the articulatory tongue model of a reference speaker from ultrasound images of the tongue of another speaker. This work is developed in the context of speech therapy based on visual biofeedback, where a speaker is provided with visual information about his/her own articulation. In our approach, the feedback is delivered via an articulatory talking head, which displays the tongue during speech production using augmented reality (e.g. transparent skin). The user’s tongue movements are captured using ultrasound imaging and parameterized using the PCA-based EigenTongue technique. Extracted features are then converted into control parameters of the articulatory tongue model using Gaussian Mixture Regression. This procedure was evaluated by decoding the converted tongue movements at the phonetic level using an HMM-based decoder trained on the reference speaker's articulatory data. Decoding errors were then manually reassessed in order to take into account possible phonetic idiosyncrasies (i.e. speaker / phoneme specific articulatory strategies). With a system trained on a limited set of 88 VCV sequences, the recognition accuracy at the phonetic level was found to be approximately 70%. Index Terms: articulatory tongue model, articulatory talking head, ultrasound imaging, GMM, speech therapy Diandra Fabre, Thomas Hueber, Pierre Badin |
INTERSPEECH | 2 |
| 2013 | Ultraspeech-player: intuitive visualization of ultrasound articulatory data for speech therapy and pronunciation training
Thomas Hueber |
INTERSPEECH | 1 |
| 2013 | Speaker adaptation of an acoustic-articulatory inversion model using cascaded Gaussian mixture regressionsabstractThe article presents a method for adapting a GMM-based acoustic-articulatory inversion model trained on a reference speaker to another speaker. The goal is to estimate the articulatory trajectories in the geometrical space of a reference speaker from the speech audio signal of another speaker. This method is developed in the context of a system of visual biofeedback, aimed at pronunciation training. This system provides a speaker with visual information about his/her own articulation, via a 3D orofacial clone. In previous work, we proposed to use GMM-based voice conversion for speaker adaptation. Acoustic-articulatory mapping was achieved in 2 consecutive steps: 1) converting the spectral trajectories of the target speaker (i.e. the system user) into spectral trajectories of the reference speaker (voice conversion), and 2) estimating the most likely articulatory trajectories of the reference speaker from the converted spectral features (acoustic-articulatory inversion). In this work, we propose to combine these two steps into the same statistical mapping framework, by fusing multiple regressions based on trajectory GMM and maximum likelihood criterion (MLE). The proposed technique is compared to two standard speaker adaptation techniques based respectively on MAP and MLLR. Thomas Hueber, Gérard Bailly, Pierre Badin, Frédéric Elisei |
INTERSPEECH | 1 |
| 2012 | Continuous Articulatory-to-Acoustic Mapping using Phone-based Trajectory HMM for a Silent Speech InterfaceabstractThe article presents an HMM-based mapping approach for converting ultrasound and video images of the vocal tract into an audible speech signal, for a silent speech interface application. The proposed technique is based on the joint modeling of articulatory and spectral features, for each phonetic class, using Hidden Markov Models (HMM) and multivariate Gaussian distributions with full covariance matrices. The articulatory-toacoustic mapping is achieved in 2 steps: 1) finding the most likely HMM state sequence from the articulatory observations; 2) inferring the spectral trajectories from both the decoded state sequence and the articulatory observations. The proposed technique is compared to our previous approach, in which only the decoded state sequence was used for the inference of the spectral trajectories, independently from the articulatory observations. Both objective and perceptual evaluations show that this new approach leads to a better estimation of the spectral trajectories. Thomas Hueber, Gérard Bailly, Bruce Denby |
INTERSPEECH | 1 |
| 2012 | Cross-speaker Acoustic-to-Articulatory Inversion using Phone-based Trajectory HMM for Pronunciation TrainingabstractThe article presents a statistical mapping approach for crossspeaker acoustic-to-articulatory inversion. The goal is to estimate the most likely articulatory trajectories for a reference speaker from the speech audio signal of another speaker. This approach is developed in the framework of our system of visual articulatory feedback developed for computer-assisted pronunciation training applications (CAPT). The proposed technique is based on the joint modeling of articulatory and acoustic features, for each phonetic class, using full-covariance trajectory HMM. The acoustic-to-articulatory inversion is achieved in 2 steps: 1) finding the most likely HMM state sequence from the acoustic observations; 2) inferring the articulatory trajectories from both the decoded state sequence and the acoustic observations. The problem of speaker adaptation is addressed using a voice conversion approach, based on trajectory GMM. Index Terms: acoustic-to-articulatory inversion, intelligent tutoring systems, pronunciation training, trajectory HMM, voice Thomas Hueber, Atef Ben Youssef, Gérard Bailly, Pierre Badin, Frédéric Elisei |
INTERSPEECH | 1 |
| 2011 | Statistical Mapping Between Articulatory and Acoustic Data for an Ultrasound-Based Silent Speech InterfaceabstractThis paper presents recent developments on our “silent speech interface ” that converts tongue and lip motions, captured by ultrasound and video imaging, into audible speech. In our previous studies, the mapping between the observed articulatory movements and the resulting speech sound was achieved using a unit selection approach. We investigate here the use of statistical mapping techniques, based on the joint modeling of visual and spectral features, using respectively Gaussian Mixture Models (GMM) and Hidden Markov Models (HMM). The prediction of the voiced/unvoiced parameter from visual articulatory data is also investigated using an artificial neural network (ANN). A continuous speech database consisting of one-hour of high-speed ultrasound and video sequences was specifically recorded to evaluate the proposed mapping techniques. Index Terms: silent speech interface, GMM, HMM, ultrasound, video, multimodal, statistical mapping Thomas Hueber, Elie-Laurent Benaroya, Bruce Denby, Gérard Chollet |
INTERSPEECH | 1 |
| 2011 | Toward a Multi-Speaker Visual Articulatory Feedback Systemabstractde niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Atef Ben Youssef, Thomas Hueber, Pierre Badin, Gérard Bailly |
INTERSPEECH | 2 |
| 2010 | Silent vs vocalized articulation for a portable ultrasound-based silent speech interfaceabstractInternational audience Victoria M. Florescu, Lise Crevier-Buchman, Bruce Denby, Thomas Hueber, Antonia Colazo-Simon, Claire Pillot-Loiseau, Pierre Roussel-Ragot, Cédric Gendrot, Sophie Quattrocchi |
INTERSPEECH | 4 |
| 2010 | Silent speech interfaces
Bruce Denby, Tanja Schultz, Kiyoshi Honda, Thomas Hueber, J. M. Gilbert, Jonathan S. Brumberg |
Speech Commun. | 4 |
| 2010 | Development of a silent speech interface driven by ultrasound and optical images of the tongue and lips
Thomas Hueber, Elie-Laurent Benaroya, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001 |
Speech Commun. | 1 |
| 2009 | Visuo-phonetic decoding using multi-stream and context-dependent models for an ultrasound-based silent speech interfaceabstractRecent improvements are presented for phonetic decoding of continuous-speech from ultrasound and optical observations of the tongue and lips in a silent speech interface application. In a new approach to this critical step, the visual streams are modeled by context-dependent multi-stream Hidden Markov Models (CD-MSHMM). Results are compared to a baseline system using context-independent modeling and a visual feature fusion strategy, with both systems evaluated on a one-hour, phonetically balanced English speech database. Tongue and lip images are coded using PCA-based feature extraction techniques. The uttered speech signal, also recorded, is used to initialize the training of the visual HMMs. Visual phonetic decoding performance is evaluated successively with and without the help of linguistic constraints introduced via a 2.5k-word decoding dictionary. Index Terms: silent speech interface, visual speech recognition, multi-stream modeling 1. Thomas Hueber, Elie-Laurent Benaroya, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001 |
INTERSPEECH | 1 |
| 2008 | Towards a segmental vocoder driven by ultrasound and optical images of the tongue and lipsabstractThis article presents a framework for a phonetic vocoder driven by ultrasound and optical images of the tongue and lips for a “silent speech interface” application. The system is built around an HMM-based visual phone recognition step which provides target phonetic sequences from a continuous visual observation stream. The phonetic target constrains the search for the optimal sequence of diphones that maximizes similarity to the input test data in visual space subject to a unit concatenation cost in the acoustic domain. The final speech waveform is generated using “Harmonic plus Noise Model” synthesis techniques. Experimental results are based on a onehour continuous speech audiovisual database comprising ultrasound images of the tongue and both frontal and lateral view of the speaker’s lips. Thomas Hueber, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001 |
INTERSPEECH | 1 |
| 2008 | Phone recognition from ultrasound and optical video sequences for a silent speech interfaceabstractLatest results on continuous speech phone recognition from video observations of the tongue and lips are described in the context of an ultrasound-based silent speech interface. The study is based on a new 61-minute audiovisual database containing ultrasound sequences of the tongue as well as both frontal and lateral view of the speaker’s lips. Phonetically balanced and exhibiting good diphone coverage, this database is designed both for recognition and corpus-based synthesis purposes. Acoustic waveforms are phonetically labeled, and visual sequences coded using PCA-based robust feature extraction techniques. Visual and acoustic observations of each phonetic class are modeled by continuous HMMs, allowing the performance of the visual phone recognizer to be compared to a traditional acoustic-based phone recognition experiment. The phone recognition confusion matrix is also discussed in detail. Thomas Hueber, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001 |
INTERSPEECH | 1 |
| 2007 | Eigentongue Feature Extraction for an Ultrasound-Based Silent Speech InterfaceabstractThe article compares two approaches to the description of ultrasound vocal tract images for application in a "silent speech interface," one based on tongue contour modeling, and a second, global coding approach in which images are projected onto a feature space of Eigentongues. A curvature-based lip profile feature extraction method is also presented. Extracted visual features are input to a neural network which learns the relation between the vocal tract configuration and line spectrum frequencies (LSF) contained in a one-hour speech corpus. An examination of the quality of LSFs derived from the two approaches demonstrates that the Eigemongues approach has a more efficient implementation and provides superior results based on a normalized mean squared error criterion. Thomas Hueber, Guido Aversano, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Yacine Oussar, Pierre Roussel-Ragot, Maureen Stone 0001 |
ICASSP (1) | 1 |
| 2007 | Continuous-speech phone recognition from ultrasound and optical images of the tongue and lipsabstractThe article describes a video-only speech recognition system for a “silent speech interface” application, using ultrasound and optical images of the voice organ. A one-hour audiovisual speech corpus was phonetically labeled using an automatic speech alignment procedure and robust visual feature extraction techniques. HMM-based stochastic models were estimated separately on the visual and acoustic corpus. The performance of the visual speech recognition system is compared to a traditional acoustic-based recognizer. Thomas Hueber, Gérard Chollet, Bruce Denby, Gérard Dreyfus, Maureen Stone 0001 |
INTERSPEECH | 1 |