Thierry Dutoit

dblp:79/2569 · DBLP profile ↗
← Back
111ranked-venue papers
11as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 86 · 10 first-author · 11 since 2021Artificial intelligence and machine learning · 58 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Robustness of Self-Avatar Animation Beyond Sparse Tracking: Effects of Pose Estimator Discrepancies and Inaccuracies
abstract
In immersive applications, reconstructing full-body self-avatar motion remains a challenging task, particularly in the absence of lowerbody tracking devices. A promising approach involves supplementing sparse upper-body motion signals with 3D Cartesian joint positions estimated from external RGB cameras. While effective, such pose estimators are prone to inaccuracies, jitter, and occlusions, which can negatively impact reconstruction quality. In this work, we investigate the robustness of self-avatar animation models to such artifacts and to variations across pose estimators. We evaluate two state-of-the-art models, HMD-Poser and AvatarJLM, and show that integrating RGB-based pose data improves lower-body accuracy over VR-only baselines, even for inaccurate 3D pose. However, model performance degrades when training and testing rely on different pose estimators. This highlights the sensitivity of current approaches to estimator variability and underscores the need for estimator-aware training to ensure robustness in real-world deployment.
Antoine Maiorca, Thierry Ravet, George Fletcher 0002, Thierry Dutoit
ISMAR4
2024 Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
abstract
Human social behaviors are inherently multi-modal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended version of Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE), which is pre-trained on audiovisual social data. Specifically, we modify CAV-MAE to receive a larger number of frames as input and pre-train it on a large dataset of human social interaction (VoxCeleb2) in a self-supervised manner. We demonstrate the effectiveness of this model by fine-tuning and evaluating the model on different social and affective downstream tasks, namely, emotion recognition, laughter detection and apparent personality estimation. The model achieves state-of-the-art results on multimodal emotion recognition and laughter recognition and competitive results for apparent personality estimation, demonstrating the effectiveness of in-domain self-supervised pre-training. Code and model weight are available here https://github.com/HuBohy/SocialMAE.
Hugo Bohy, Minh Tran 0004, Kevin El Haddad, Thierry Dutoit, Mohammad Soleymani 0001
FG4
2024 Latent Space Interpolation of Synthesizer Parameters Using Timbre-Regularized Auto-Encoders
abstract
Sound synthesizers are ubiquitous in modern music production but manipulating their presets, i.e. the sets of synthesis parameters, demands expert skills. This study presents a novel variational auto-encoder model tailored for black-box synthesizer preset interpolation, which enables the intuitive generation of new presets from pre-existing ones. Leveraging multi-head self-attention networks, the model efficiently learns latent representations of synthesis parameters, aligning these with perceived timbre dimensions through attribute-based regularization. It is able to gradually transition between diverse presets, surpassing traditional linear parametric interpolation methods. Furthermore, we introduce an objective and reproducible evaluation method, based on linearity and smoothness metrics computed on a broad set of audio features. The model's efficacy is demonstrated through subjective experiments, whose results also highlight significant correlations with the proposed objective metrics. The model is validated using a widespread frequency modulation synthesizer with a large set of interdependent parameters. It can be adapted to various commercial synthesizers, and can perform other tasks such as modulations and extrapolations.
Gwendal Le Vaillant, Thierry Dutoit
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Cardiotocography Signal Abnormality Detection Based on Deep Semi-Unsupervised Learning
abstract
Cardiotocography (CTG) plays a vital role in fetal well-being monitoring by tracking fetal heart rate (FHR) and uterine contractions (UC). However, CTG interpretation suffers from subjectivity, resulting in low agreement among observers and potentially unnecessary medical interventions. Existing AI-based diagnostic models struggle with generalization and are only effective on distinct CTG samples. This study introduces a novel approach employing deep semi-supervised learning for anomaly detection in CTG signals, marking the first attempt in this direction. A modified GANomaly model is proposed, trained on normal CTG data, and evaluated for abnormality detection. This model employs an encoder-decoder structure with a discriminator, minimizing reconstruction and latent space errors while learning the normal CTG signal distribution. Leveraging the CTU-UHB dataset, our model demonstrates superior performance compared to existing methods during inference.
Julien Bertieaux, Mohammadhadi Shateri, Fabrice Labeau, Thierry Dutoit
BDCAT4
2023 Deep Learning-Based Stereo Camera Multi-Video Synchronization
abstract
Stereo vision is essential for many applications. Currently, the synchronization of the streams coming from two cameras is done using mostly hardware. A software-based synchronization method would reduce the cost, weight and size of the entire system and allow for more flexibility when building such systems. With this goal in mind, we present here a comparison of different deep learning-based systems and prove that some are efficient and generalizable enough for such a task. This study paves the way to a production ready software-based video synchronization system.
Nicolas Boizard, Kevin El Haddad, Thierry Ravet, François Cresson, Thierry Dutoit
ICASSP5
2023 Synthesizer Preset Interpolation Using Transformer Auto-Encoders
abstract
Sound synthesizers are widespread in modern music production but they increasingly require expert skills to be mastered. This work focuses on interpolation between presets, i.e., sets of values of all sound synthesis parameters, to enable the intuitive creation of new sounds from existing ones.We introduce a bimodal auto-encoder neural network, which simultaneously processes presets using multi-head attention blocks, and audio using convolutions. This model has been tested on a popular frequency modulation synthesizer with more than one hundred parameters. Experiments have compared the model to related architectures and methods, and have demonstrated that it performs smoother interpolations. After training, the proposed model can be integrated into commercial synthesizers for live interpolation or sound design tasks.
Gwendal Le Vaillant, Thierry Dutoit
ICASSP2
2023 Objective Evaluation Metric for Motion Generative Models: Validating Fréchet Motion Distance on Foot Skating and Over-smoothing Artifacts
abstract
Nowadays, Deep Learning-powered generative models are able to generate new synthetic samples nearly indistinguishable from natural data. The development of such systems necessarily involves the design of evaluation protocols to assess their performance. Quantitative objective metrics, such as Fréchet distance, in addition to human-centered subjective surveys, have become a standard for evaluating generative algorithms. Although motion generation is a popular research field, only a few works addressed the problem of the design and validation of a robust objective evaluation metric for motion-generative models. These previous works proposed to degrade ground truth motion samples with synthetic noises (e.g., Gaussian, Salt& Pepper) and studied the behavior of the proposed metric. However, this degradation does not mimic common motion artifacts produced by generative models. In this work, we propose (1) to validate Fréchet distance-based objective metrics on motion datasets degraded by two realistic motion artifacts, foot skating and over-smoothing, often found in motion synthesis results, and (2) a Fréchet Motion Distance (FMD), using Transformer-based feature extractor, able to capture the motion artifacts and also robust towards the variation of motion length.
Antoine Maiorca, Hugo Bohy, Youngwoo Yoon, Thierry Dutoit
MIG4
2023 Validating Objective Evaluation Metric: Is Fréchet Motion Distance able to Capture Foot Skating Artifacts ?
abstract
Automatically generating character motion is one of the technologies required for virtual reality, graphics, and robotics. Motion synthesis with deep learning is an emerging research topic. A key component of the development of such an algorithm involves the design of a proper objective metric to evaluate the quality and diversity of the synthesized motion dataset, two key factors of the performance of generative models. The Fréchet distance is nowadays a common method to assess this performance. In the motion generation field, the validation of such evaluation methods relies on the computation of the Fréchet distance between embeddings of the ground truth dataset and motion samples polluted by synthetic noise to mimic the artifacts produced by generative algorithms. However, the synthetic noise degradation does not fully represent motion perturbations that are commonly perceived. One of these artifacts is foot skating: the unnatural foot slides on the ground during locomotion. In this work-in-progress paper, we tested how well the Fréchet Motion Distance (FMD), which was proposed in previous works, is able to measure foot skating artifacts, and we found that FMD is not able to measure efficiently the intensity of the skating degradation.
Antoine Maiorca, Youngwoo Yoon, Thierry Dutoit
IMX3
2023 Developing an Interactive Agent for Blind and Visually Impaired People
abstract
The aim of this project is to create an interactive assistant that incorporates different assistive features for blind and visually impaired people. The assistant might incorporate screen readers, magnifiers, voice synthesis, OCR, GPS, face recognition, and object recognition among other tools. Recently, the work done by OpenAI and Be My Eyes with the implementation of GPT-4 is comparable to the aim of this project. It shows the development of an interactive assistant has become simpler due to recent developments in large language models. However, older methods like named entity recognition and intent classification are still valuable to build lightweight assistants. A hybrid solution combining both methods seems possible, would help to reduce the computational cost of the assistant, and would facilitate the data collection process. Despite being more complex to implement in a multilingual and multimodal context, a hybrid solution has the potential to be used offline and to consume less resources.
Vincent Stragier, Omar Seddati, Thierry Dutoit
IMX3
2022 A New Perspective on Smiling and Laughter Detection: Intensity Levels Matter
abstract
Smiles and laughs detection systems have attracted a lot of attention in the past decade contributing to the improvement of human-agent interaction systems. But very few considered these expressions as distinct, although no prior work clearly proves them to belong to the same category or not. In this work, we present a deep learning-based multimodal smile and laugh classification system, considering them as two different entities. We compare the use of audio and vision-based models as well as a fusion approach. We show that, as expected, the fusion leads to a better generalization on unseen data. We also present an in-depth analysis of the behavior of these models on the smiles and laughs intensity levels. The analyses on the intensity levels show that the relationship between smiles and laughs might not be as simple as a binary one or even grouping them in a single category, and so, a more complex approach should be taken when dealing with them. We also tackle the problem of limited resources by showing that transfer learning allows the models to improve the detection of confusing intensity levels.
Hugo Bohy, Kevin El Haddad, Thierry Dutoit
ACII3
2022 Towards Human Performance on Sketch-Based Image Retrieval
abstract
Sketch-based image retrieval (SBIR) solutions are attracting increased interest in the field of computer vision. These solutions provide an intuitive and powerful tool to retrieve images in large-scale image databases. In this paper, we conduct a comprehensive study of classic triplet CNN training pipelines within the SBIR context. We study the impact of embeddings normalization, model sharing, margin selection, batch size, hard mining selection and the evolution of the number of hard triplets during training to propose several avenues for improvement. We also propose dropout column, an adaptation of dropout for triplet network and similar pipelines. In addition, we also introduce a novel approach to build state-of-the-art SBIR solutions that can be used with low power systems. The whole study is conducted using The Sketchy Database, a large-scale SBIR database. We carry out a series of experiments and show that adopting a few simple modifications enhances significantly existing SBIR pipelines (faster training & higher accuracy). Our study enables us to propose an enhanced pipeline that outperforms previous state-of-the-art on the Sketchy Database by a significant margin (a recall of 53.92% compared to 46.2% at k = 1) and reaches almost human performance (54.27%) on a large-scale benchmark.
Omar Seddati, Stéphane Dupont, Saïd Mahmoudi, Thierry Dutoit
CBMI4
2022 Spatio-Temporal Analysis of Transformer based Architecture for Attention Estimation from EEG
abstract
For many years now, understanding the brain mechanism has been a great research subject in many different fields. Brain signal processing and especially electroencephalogram (EEG) has recently known a growing interest both in academia and industry. One of the main examples is the increasing number of Brain-Computer Interfaces (BCI) aiming to link brains and computers. In this paper, we present a novel framework allowing us to retrieve the attention state, i.e degree of attention given to a specific task, from EEG signals. While previous methods often consider the spatial relationship in EEG through electrodes and process them in recurrent or convolutional based architecture, we propose here to also exploit the time and frequency information with a transformer-based network that has already shown its supremacy in many machine-learning (ML) related studies, e.g. machine translation. In addition to this novel architecture, an extensive study on the feature extraction methods, frequential bands and temporal windows length has also been carried out. The proposed network has been trained and validated on two public datasets and achieves higher results compared to state-of-the-art models. As well as proposing better results, the framework could be used in real applications, e.g. Attention Deficit Hyperactivity Disorder (ADHD) symptoms or vigilance during a driving assessment.
Victor Delvigne, Hazem Wannous, Jean-Philippe Vandeborre, Laurence Ris, Thierry Dutoit
ICPR5
2022 PhyDAA: Physiological Dataset Assessing Attention
abstract
Attention Deficit Hyperactivity Disorder (ADHD) is the most prevalent neurodevelopmental disorder among children. It affects patients’ lives in many ways: inattention, difficulty with stimuli inhibition or motor function regulation. Different treatments exist today, but these can present side effects or are not effective for all subgroups. Neurofeedback (NF) is an innovative treatment consisting of brain activity display. NF training could consist of a virtual reality (VR) video-game in which the participant’s attention affects the game. Attention being assessed through physiological signals, one of the main steps is to design an estimator for the attention state. We present a novel framework able to record physiological signals in specific attention states and able to estimate the corresponding attention state. We propose a database composed of electroencephalography signals (EEG), and an eye-tracker labelled with a score representing the attention span for 32 healthy participants. Different features are extracted from the signals and machine learning (ML) algorithms are proposed. Our approach exhibits high accuracy for attention estimation, which corroborates a correlation between attention state and physiological signals (i.e. EEG, eye-tracking signals). The dataset has been made publicly available to promote research in the domain and we encourage other scientists to use their own approach for attention estimation.
Victor Delvigne, Hazem Wannous, Thierry Dutoit, Laurence Ris, Jean-Philippe Vandeborre
IEEE Trans. Circuits Syst. Video Technol.3
2021 Improving Synthesizer Programming From Variational Autoencoders Latent Space
abstract
Deep neural networks have been recently applied to the task of automatic synthesizer programming, i.e., finding optimal values of sound synthesis parameters in order to reproduce a given input sound. This paper focuses on generative models, which can infer parameters as well as generate new sets of parameters or perform smooth morphing effects between sounds. We introduce new models to ensure scalability and to increase performance by using heterogeneous representations of parameters as numerical and categorical random variables. Moreover, a spectral variational autoencoder architecture with multi-channel input is proposed in order to improve inference of parameters related to the pitch and intensity of input sounds. Model performance was evaluated according to several criteria such as parameters estimation error and audio reconstruction accuracy. Training and evaluation were performed using a 30k presets dataset which is published with this paper. They demonstrate significant improvements in terms of parameter inference and audio accuracy and show that presented models can be used with subsets or full sets of synthesizer parameters.
Gwendal Le Vaillant, Thierry Dutoit, Sebastien Dekeyser
DAFx2
2020 ICE-Talk: An Interface for a Controllable Expressive Talking Machine
Noé Tits, Kevin El Haddad, Thierry Dutoit
INTERSPEECH3
2020 Laughter Synthesis: Combining Seq2seq Modeling with Transfer Learning
abstract
Despite the growing interest for expressive speech synthesis, synthesis of nonverbal expressions is an under-explored area. In this paper we propose an audio laughter synthesis system based on a sequence-to-sequence TTS synthesis system. We leverage transfer learning by training a deep learning model to learn to generate both speech and laughs from annotations. We evaluate our model with a listening test, comparing its performance to an HMM-based laughter synthesis one and assess that it reaches higher perceived naturalness. Our solution is a first step towards a TTS system that would be able to synthesize speech with a control on amusement level with laughter integration.
Noé Tits, Kevin El Haddad, Thierry Dutoit
INTERSPEECH3
2020 Depth prediction from 2D images: A taxonomy and an evaluation study
Ambroise Moreau, Matei Mancas, Thierry Dutoit
Image Vis. Comput.3
2019 An End-to-end Network to Synthesize Intonation Using a Generalized Command Response Model
abstract
The generalized command response (GCR) model represents intonation as a superposition of muscle responses to spike command signals. We have previously shown that the spikes can be predicted by a two-stage system, consisting of a recurrent neural network and a post-processing procedure, but the responses themselves were fixed dictionary atoms. We propose an end-to-end neural architecture that replaces the dictionary atoms with trainable second-order recurrent elements analogous to recursive filters. We demonstrate gradient stability under modest conditions, and show that the system can be trained by imposing temporal sparsity constraints. Subjective listening tests demonstrate that the system can synthesize intonation with high naturalness, comparable to state-of-the-art acoustic models, and retains the physiological plausibility of the GCR model.
François Marelli, Bastian Schnell, Hervé Bourlard, Thierry Dutoit, Philip N. Garner
ICASSP4
2019 Leveraging Pre-trained CNN Models for Skeleton-Based Action Recognition
Sohaib Laraba, Joëlle Tilmanne, Thierry Dutoit
ICVS3
2019 Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis
abstract
The field of Text-to-Speech has experienced huge improvements last years benefiting from deep learning techniques. Producing realistic speech becomes possible now. As a consequence, the research on the control of the expressiveness, allowing to generate speech in different styles or manners, has attracted increasing attention lately. Systems able to control style have been developed and show impressive results. However the control parameters often consist of latent variables and remain complex to interpret. In this paper, we analyze and compare different latent spaces and obtain an interpretation of their influence on expressive speech. This will enable the possibility to build controllable speech synthesis systems with an understandable behaviour.
Noé Tits, Fengna Wang, Kevin El Haddad, Vincent Pagel, Thierry Dutoit
INTERSPEECH5
2018 A Dyadic Conversation Dataset on Moral Emotions
abstract
In this paper, we present a dyadic conversation dataset involving topics related to moral emotions which are ethically relevant. To the best of our knowledge, it is the first dataset where the main focus is moral emotions. This dataset also focuses on speaker-listener reactions during a dyadic conversation. Although some of the currently available datasets contain dyadic conversations, they were not conceived with the idea of focusing on the speaker-listener setup. Thus making it difficult to use them to study reactions related to speakers and listeners. Some preliminary analyses of the data are presented as well as our thoughts on future work related to this dataset.
Louise Heron, Jaebok Kim, Minha Lee, Kevin El Haddad, Stéphane Dupont, Thierry Dutoit, Khiet P. Truong
FG6
2018 HMM-based generation of laughter facial expression
Hüseyin Çakmak, Thierry Dutoit
Speech Commun.2
2017 3D skeleton-based action recognition by representing motion capture sequences as 2D-RGB images
abstract
Abstract In recent years, 3D skeleton‐based action recognition has become a popular technique of action classification, thanks to development and availability of cheaper depth sensors. State‐of‐the‐art methods generally represent motion sequences as high dimensional trajectories followed by a time‐warping technique. These trajectories are used to train a classification model to predict the classes of new sequences. Despite the success of these techniques in some fields, particularly when the data used are captured by a high‐precision motion capture system, action classification is still less successful than the field of image classification, especially with the advance of deep learning. In this paper, we present a new representation of motion sequences (Seq2Im—for sequence to image), which projects motion sequences onto the RGB domain. The 3D coordinates of joints are mapped to red, green, and blue values, and therefore, action classification becomes an image classification problem and algorithms for this field can be applied. This representation was tested with basic image classification algorithms (namely, support vector machine,k‐nearest neighbor, and random forests) in addition to convolutional neural networks. Evaluation of the proposed method on standard 3D human action recognition datasets shows its potential for action recognition and outperforms most of the state‐of‐the‐art results.
Sohaib Laraba, Mohammed Brahimi 0001, Joëlle Tilmanne, Thierry Dutoit
Comput. Animat. Virtual Worlds4
2016 Multi-task learning for speech recognition: an overview
Gueorgui Pironkov, Stéphane Dupont, Thierry Dutoit
ESANN3
2016 Towards a listening agent: a system generating audiovisual laughs and smiles to show interest
abstract
In this work, we experiment with the use of smiling and laughter in order to help create more natural and efficient listening agents. We present preliminary results on a system which predicts smile and laughter sequences in one dialogue participant based on observations of the other participant's behavior. This system also predicts the level of intensity or arousal for these sequences. We also describe an audiovisual concatenative synthesis process used to generate laughter and smiling sequences, producing multilevel amusement expressions from a dataset of audiovisual laughs. We thus present two contributions: one in the generation of smiling and laughter responses, the other in the prediction of what laughter and smiles to use in response to an interlocutor's behaviour. Both the synthesis system and the prediction system have been evaluated via Mean Opinion Score tests and have proved to give satisfying and promising results which open the door to interesting perspectives.
Kevin El Haddad, Hüseyin Çakmak, Emer Gilmartin, Stéphane Dupont, Thierry Dutoit
ICMI5
2016 Speaker-aware Multi-Task Learning for automatic speech recognition
abstract
Overfitting is a commonly met issue in automatic speech recognition and is especially impacting when the amount of training data is limited. In order to address this problem, this article investigates acoustic modeling through Multi-Task Learning, with two speaker-related auxiliary tasks. Multi-Task Learning is a regularization method which aims at improving the network's generalization ability, by training a unique model to solve several different, but related tasks. In this article, two auxiliary tasks are jointly examined. On the one hand, we consider speaker classification as an auxiliary task by training the acoustic model to recognize the speaker, or find the closest one inside the training set. On the other hand, the acoustic model is also trained to extract i-vectors from the standard acoustic features. I-Vectors are efficiently applied in the speaker identification community in order to characterize a speaker and its acoustic environment. The core idea of using these auxiliary tasks is to give the network an additional inter-speaker awareness, and thus, reduce overfitting.We investigate this Multi-Task Learning setup on the TIMIT database, while the acoustic modeling is performed using a Recurrent Neural Network with Long Short-Term Memory cells.
Gueorgui Pironkov, Stéphane Dupont, Thierry Dutoit
ICPR3
2016 AVAB-DBS: an Audio-Visual Affect Bursts Database for Synthesis
Kevin El Haddad, Hüseyin Çakmak, Stéphane Dupont, Thierry Dutoit
LREC4
2016 I-Vector estimation as auxiliary task for Multi-Task Learning based acoustic modeling for automatic speech recognition
abstract
I-Vectors have been successfully applied in the speaker identification community in order to characterize the speaker and its acoustic environment. Recently, i-vectors have also shown their usefulness in automatic speech recognition, when concatenated to standard acoustic features. Instead of directly feeding the acoustic model with i-vectors, we here investigate a Multi-Task Learning approach, where a neural network is trained to simultaneously recognize the phone-state posterior probabilities and extract i-vectors, using the standard acoustic features. Multi-Task Learning is a regularization method which aims at improving the network's generalization ability, by training a unique network to solve several different, but related tasks. The core idea of using i-vector extraction as an auxiliary task is to give the network an additional inter-speaker awareness, and thus, reduce overfitting. Overfitting is a commonly met issue in speech recognition and is especially impacting when the amount of training data is limited. The proposed setup is trained and tested on the TIMIT database, while the acoustic modeling is performed using a Recurrent Neural Network with Long Short-Term Memory cells.
Gueorgui Pironkov, Stéphane Dupont, Thierry Dutoit
SLT3
2015 GMM-based synchronization rules for HMM-based audio-visual laughter synthesis
abstract
In this paper we propose synchronization rules between acoustic and visual laughter synthesis systems. Previous works have addressed separately the acoustic and visual laughter synthesis following an HMM-based approach. The need of synchronization rules comes from the constraint that in laughter, HMM-based synthesis cannot be performed using a unified system where common transcriptions may be used as it has been shown to be the case for audio-visual speech synthesis. Therefore acoustic and visual models are trained independently without any synchronization constraints. In this work, we propose rules derived from the analysis of audio and visual laughter transcriptions in order to be able to generate a visual laughter transcriptions corresponding to an audio laughter data.
Hüseyin Çakmak, Kevin El Haddad, Thierry Dutoit
ACII3
2015 Investigating sparse deep neural networks for speech recognition
abstract
We propose an organized sparse deep neural network architecture for automatic speech recognition. The proposed method is inspired by the tonotopic organization in the auditory nerve/cortex. The approach consists of limiting the neurons connections between the hidden layers, in a manner that preserves frequency proximity, resulting in a diffuse integration of the spectral information inside the neural network. This method is put in perspective with related work on sparser neural network architectures for speech recognition (tonotopy, convolutional nets, dropout). The model is trained and tested on the TIMIT database, showing encouraging results compared to the traditional fully connected architecture.
Gueorgui Pironkov, Stéphane Dupont, Thierry Dutoit
ASRU3
2015 Synchronization rules for HMM-based audio-visual laughter synthesis
abstract
In this paper we propose synchronization rules between acoustic and visual laughter synthesis systems. This work follows up our previous studies on acoustics laughter synthesis and visual laughter synthesis. The need of synchronization rules comes from the constraint that in laughter, HMM-based synthesis of laughter cannot be performed using a unified system where common transcriptions may be used. Therefore acoustic and visual models are trained independently without any synchronization constraints. In this work, we propose simple rules derived from the analysis of audio and visual laughter transcriptions in order to generate visual laughter transcriptions starting from acoustic transcriptions. A perceptive Mean Opinion Score (MOS) test is conducted to evaluate the method.
Hüseyin Çakmak, Jérôme Urbain, Thierry Dutoit
ICASSP3
2015 Speech-laughs: An HMM-based approach for amused speech synthesis
abstract
This paper presents an HMM-based synthesis approach for speechlaughs. The building stone of this project was the idea of the co-occurrence of smile and laughter bursts in varying proportions within amused speech utterances. A corpus with three complementary speaking styles was used to train the underlying HMM models: neutral speech, speech-smile, and finally laughter in different articulatory configurations. Two types of speech-laughs were then synthesized: one made by combining neutral speech and laughter bursts, and the other made by combining speech-smile and laughter bursts. Synthesized stimuli were then rated in terms of perceived amusement and naturalness levels. Results show the compound effect of both laughter bursts and smile on both amusement and naturalness and inspire interesting perspectives.
Kevin El Haddad, Stéphane Dupont, Jérôme Urbain, Thierry Dutoit
ICASSP4
2015 Adaptation procedure for HMM-based sensor-dependent gesture recognition
abstract
In this paper, we address the problem of sensor-dependent gesture recognition thanks to adaptation procedure. Capturing human movements by a motion capture (MoCap) system provides very accurate data. Unfortunately, such systems are very expensive, unlike recent depth sensors, like Microsoft Kinect, which are much cheaper, but provide lower data quality. Hidden Markov Models (HMMs) are widely used in gesture recognition to learn the dynamics of each gesture class. However, models trained on one type of data can only be used on data of the same type. For this reason, we propose to adapt HMMs trained on Mocap data to a small set of Kinect data using Maximum Likelihood Linear Regression (MLLR) to recognize gestures captured by a Kinect. Results show that using this method, we can achieve a recognition average accuracy of 84.48% using a small set of adaptation data while, using the same set to create new models, we obtain only 72.41% of accuracy.
Sohaib Laraba, Joëlle Tilmanne, Thierry Dutoit
MIG3
2015 An evaluation criterion of saliency models for video seam carving
abstract
Modeling human attention has been arousing a lot of interest due to its numerous applications. The process that allows us to focus on some more important stimuli is defined as the “attention”. Seam carving is an approach to resize images or video sequences while preserving the semantic content. To define what is important, gradient was first used but due to limitations of this approach, saliency models, which are able to predict whether regions in the images attract human attention, are now used. Most of them are optimized to conform a ground truth like eye tracking but there is no way to know the efficiency of the saliency models applied in a specific application like in seam carving. In this paper, we propose a criterion, based on the quantity of geometric deformation and the image's reduction, which evaluate the quality of a resizing by seam carving. This criterion is applied on evaluation of image (SCES) or video (SCED) resizing. We validate our criterion with subjective evaluation and used it to rank state of the art saliency models for seam carving. Evaluation of the image by SCES gives a Spearman correlation of −0.92196 and a Pearson correlation of −0.8812. For the video, the final SCED gives a Spearman correlation of −0.81351 and a Pearson correlation of −0.80581.
Marc Décombas, Pierre Marighetto, Matei Mancas, Ioannis Cassagne, Nicolas Riche, Bernard Gosselin, Thierry Dutoit, Robert Laganière
MMSP7
2014 Parametric representation for singing voice synthesis: A comparative evaluation
abstract
Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be straightforward. The goal of this paper is twofold. First, a comparative subjective evaluation is performed across four existing techniques suitable for statistical parametric synthesis: traditional pulse vocoder, Deterministic plus Stochastic Model, Harmonic plus Noise Model and GlottHMM. The behavior of these techniques as a function of the singer type (baritone, counter-tenor and soprano) is studied. Secondly, the artifacts occurring in high-pitched voices are discussed and possible approaches to overcome them are suggested.
Onur Babacan, Thomas Drugman, Tuomo Raitio, Daniel Erro, Thierry Dutoit
ICASSP5
2014 Evaluation of HMM-based visual laughter synthesis
abstract
In this paper we apply speaker-dependent training of Hidden Markov Models (HMMs) to audio and visual laughter synthesis separately. The two modalities are synthesized with a forced durations approach and are then combined together to render audio-visual laughter on a 3D avatar. This paper focuses on visual synthesis of laughter and its perceptive evaluation when combined with synthesized audio laughter. Previous work on audio and visual synthesis has been successfully applied to speech. The extrapolation to audio laughter synthesis has already been done. This paper shows that it is possible to extrapolate to visual laughter synthesis as well.
Hüseyin Çakmak, Jérôme Urbain, Joëlle Tilmanne, Thierry Dutoit
ICASSP4
2014 elite-HTS: a NLP tool for French HMM-based speech synthesis
Sophie Roekhaut, Sandrine Brognaux, Richard Beaufort, Thierry Dutoit
INTERSPEECH4
2014 The AV-LASYN Database : A synchronous corpus of audio and 3D facial marker data for audio-visual laughter synthesis
Hüseyin Çakmak, Jérôme Urbain, Thierry Dutoit, Joëlle Tilmanne
LREC3
2014 Scenarizing Metropolitan Views: FlanoGraphing the Urban Spaces
abstract
The recent decade has seen a rapid evolution in the field of digital media. Mobile devices are now being integrated into every aspect of urban life. GPS, sensor technologies and augmented reality have transformed the new generation of mobile devices from a communication and information platform into a navigational tool, fostering new ways of perceiving reality and image building. Touch sensor technology has changed the screen into a joint input and display device. In this paper we present the FlanoGraph , an application for smartphones and tablets designed to take benefit of the changes induced by mobile devices. We first briefly outline the conceptual background, evoking the work of some researchers in the fields of ‘Non Representational Theory’, mobile media, and computational data processing. We then present and describe the FlanoGraph through a set of use cases. Finally, we conclude discussing some techniques necessary for the development of the application.
Bénédicte Jacobs, Laure-Anne Jacobs, Christian Frisson, Willy Yvart, Thierry Dutoit, Sylvie Leleu-Merviel
MMM (2)5
2014 Scenarizing CADastre Exquisse: A Crossover between Snoezeling in Hospitals/Domes, and Authoring/Experiencing Soundful Comic Strips
Cédric Sabato, Aurélien Giraudet, Virginie Delattre, Yves Desnos, Christian Frisson, Rudi Giot, Willy Yvart, François Rocca, Stéphane Dupont, Guy Vandem Bemden, Sylvie Leleu-Merviel, Thierry Dutoit
MMM (2)12
2014 Tangible needle, digital haystack: tangible interfaces for reusing media content organized by similarity
abstract
This paper presents the design process of a desk-set tangible user interface for the navigation and manipulation of media content organized by content-based similarity with off-the-shelf/flea market devices. For intra-media navigation, a refurbished portable vinyl player has its inside mechanics replaced by a webcam monitoring circular gray code analyzed through computer vision for position/speed tracking. For inter-media navigation, a 3D force-feedback controller is mounted in upright position on a truss with cell clamps, repurposed as trackpad. For media recomposition, motorized faders recall the effect presets of the closest/last selected media item.
Christian Frisson, François Rocca, Stéphane Dupont, Thierry Dutoit, Damien Grobet, Rudi Giot, Mohammed El Brouzi, Samir Bouaziz, Willy Yvart, Sylvie Leleu-Merviel
TEI4
2014 Analysis and HMM-based synthesis of hypo and hyperarticulated speech
Benjamin Picart, Thomas Drugman, Thierry Dutoit
Comput. Speech Lang.3
2014 Speech polarity determination: A comparative evaluation
Thomas Drugman, Thierry Dutoit
Neurocomputing2
2014 HMM-based speech synthesis with various degrees of articulation: A perceptual study
Benjamin Picart, Thomas Drugman, Thierry Dutoit
Neurocomputing3
2013 Automatic Phonetic Transcription of Laughter and Its Application to Laughter Synthesis
abstract
In this paper, automatic phonetic transcription of laughter is achieved with the help of Hidden Markov Models (HMMs). The models are evaluated in a speaker-independent way. Several measures to evaluate the quality of the transcriptions are discussed, some focusing on the recognized sequences (without paying attention to the segmentation of the phones), other only taking into account the segmentation boundaries (without involving the phonetic labels). Although the results are far from perfect recognition, it is shown that using this kind of automatic transcriptions does not impair too much the naturalness of laughter synthesis. The paper opens interesting perspectives in automatic laughter analysis as well as in laughter synthesis, as it will enable faster developments of laughter synthesis on large sets of laughter data.
Jérôme Urbain, Hüseyin Çakmak, Thierry Dutoit
ACII3
2013 A comparative study of pitch extraction algorithms on a large variety of singing sounds
abstract
The problem of pitch tracking has been extensively studied in the speech research community. The goal of this paper is to investigate how these techniques should be adapted to singing voice analysis, and to provide a comparative evaluation of the most representative state-of-the-art approaches. This study is carried out on a large database of annotated singing sounds with aligned EGG recordings, comprising a variety of singer categories and singing exercises. The algorithmic performance is assessed according to the ability to detect voicing boundaries and to accurately estimate pitch contour. First, we evaluate the usefulness of adapting existing methods to singing voice analysis. Then we compare the accuracy of several pitch-extraction algorithms, depending on singer category and laryngeal mechanism. Finally, we analyze their robustness to reverberation.
Onur Babacan, Thomas Drugman, Nicolas D'Alessandro, Nathalie Henrich Bernardoni, Thierry Dutoit
ICASSP5
2013 Evaluation of HMM-based laughter synthesis
abstract
In this paper we explore the potential of Hidden Markov Models (HMMs) for laughter synthesis. Several versions of HMMs are developed, with varying contextual information and algorithms for estimating the parameters of the source-filter synthesis model. These methods are compared, in a perceptive tests, to the naturalness of actual human laughs and copy-synthesis laughs. The evaluation shows that 1) the addition of contextual information did not increase the naturalness, 2) the proposed method is significantly less natural than human and copy-synthesized laughs, but 3) significantly improves laughter synthesis naturalness compared to the state of the art. The evaluation also demonstrates that the duration of the laughter units can be efficiently learnt by the HMM-based parametric synthesis methods.
Jérôme Urbain, Hüseyin Çakmak, Thierry Dutoit
ICASSP3
2013 Saliency and Human Fixations: State-of-the-Art and Study of Comparison Metrics
abstract
Visual saliency has been an increasingly active research area in the last ten years with dozens of saliency models recently published. Nowadays, one of the big challenges in the field is to find a way to fairly evaluate all of these models. In this paper, on human eye fixations, we compare the ranking of 12 state-of-the art saliency models using 12 similarity metrics. The comparison is done on Jian Li's database containing several hundreds of natural images. Based on Kendall concordance coefficient, it is shown that some of the metrics are strongly correlated leading to a redundancy in the performance metrics reported in the available benchmarks. On the other hand, other metrics provide a more diverse picture of models' overall performance. As a recommendation, three similarity metrics should be used to obtain a complete point of view of saliency model performance.
Nicolas Riche, Matthieu Duvinage, Matei Mancas, Bernard Gosselin, Thierry Dutoit
ICCV5
2013 Spatio-temporal saliency based on rare model
abstract
In this paper, a new spatio-temporal saliency model is presented. Based on the idea that both spatial and temporal features are needed to determine the saliency of a video, this model builds upon the fact that locally contrasted and globally rare features are salient. The features used in the model are both spatial (color and orientations) and temporal (motion amplitude and direction) at several scales. To be more robust to moving camera a module computes the global motion and to be more consistent in time, the saliency maps are combined together after a temporal filtering. The model is evaluated on a dataset of 24 videos split into 5 categories (Abnormal, Surveillance, Crowds, Moving camera, and Noisy). This model achieves better performance when compared to several state-of-the-art saliency models.
Marc Décombas, Nicolas Riche, Frédéric Dufaux, Béatrice Pesquet-Popescu, Matei Mancas, Bernard Gosselin, Thierry Dutoit
ICIP7
2013 Biologically plausible context recognition algorithms
abstract
In this paper, four new approaches of global context recognition algorithms (gist) are introduced. They are able to automatically distinguish context differences like buildings, coast, home (indoor), mountain or streets. All proposed models are biologically plausible and are able to deal with both color and gray-level images. They use Gabor or Log-Gabor filters to extract features that better mimic human visual perception. Those features are then classified using a Mahalanobis space (when a subset of features is extracted) or in a high-dimensional Gaussian space (when all features are taken into account) with Support Vector Machines (SVM). The proposed models are compared to a standard state of the art gist model to proof their efficiency.
Makiese Mibulumukini, Nicolas Riche, Matei Mancas, Bernard Gosselin, Thierry Dutoit
ICIP5
2013 Reactive accent interpolation through an interactive map application
Maria Astrinaki, Junichi Yamagishi, Simon King 0001, Nicolas D'Alessandro, Thierry Dutoit
INTERSPEECH5
2013 A quantitative comparison of glottal closure instant estimation algorithms on a large variety of singing sounds
abstract
Glottal closure instant (GCI) estimation is a well-studied topic that plays a critical role in several speech processing applications. Many GCI estimation algorithms have been proposed in the literature and shown to provide excellent results on the speech signal. Nonetheless the efficiency of these algorithms for the analysis of the singing voice is still unknown. The goal of this paper is to assess the performance of existing GCI estimation methods on the singing voice with a quantitative comparison. A second goal is to provide a starting point for the adaptation of these algorithms to the singing voice by identifying weaknesses and strengths under different conditions. This study is carried out on a large database of singing sounds with synchronous electroglottography (EGG) recordings, containing a variety of singer categories and singing techniques. The evaluated algorithms are Dynamic Programming Phase Slope
Onur Babacan, Thomas Drugman, Nicolas D'Alessandro, Nathalie Henrich Bernardoni, Thierry Dutoit
INTERSPEECH5
2013 VideoCycle: User-Friendly Navigation by Similarity in Video Databases
Christian Frisson, Stéphane Dupont, Alexis Moinet, Cécile Picard-Limpens, Thierry Ravet, Xavier Siebert, Thierry Dutoit
MMM (2)7
2013 RARE2012: A multi-scale rarity-based saliency detection with its comparative statistical analysis
Nicolas Riche, Matei Mancas, Matthieu Duvinage, Makiese Mibulumukini, Bernard Gosselin, Thierry Dutoit
Signal Process. Image Commun.6
2013 Objective Study of Sensor Relevance for Automatic Cough Detection
abstract
The development of a system for the automatic, objective, and reliable detection of cough events is a need underlined by the medical literature for years. The benefit of such a tool is clear as it would allow the assessment of pathology severity in chronic cough diseases. Even though some approaches have recently reported solutions achieving this task with a relative success, there is still no standardization about the method to adopt or the sensors to use. The goal of this paper is to study objectively the performance of several sensors for cough detection: ECG, thermistor, chest belt, accelerometer, contact, and audio microphones. Experiments are carried out on a database of 32 healthy subjects producing, in a confined room and in three situations, voluntary cough at various volumes as well as other event categories which can possibly lead to some detection errors: background noise, forced expiration, throat clearing, speech, and laugh. The relevance of each sensor is evaluated at three stages: mutual information conveyed by the features, ability to discriminate at the frame level cough from these latter other sources of ambiguity, and ability to detect cough events. In this latter experiment, with both an averaged sensitivity and specificity of about 94.5%, the proposed approach is shown to clearly outperform the commercial Karmelsonix system which achieved a specificity of 95.3% and a sensitivity of 64.9%.
Thomas Drugman, Jérôme Urbain, Nathalie Bauwens, Ricardo Chessini, Carlos Valderrama 0001, Patrick Lebecque, Thierry Dutoit
IEEE J. Biomed. Health Informatics7
2012 Dynamic Saliency Models and Human Attention: A Comparative Study on Videos
Nicolas Riche, Matei Mancas, Dubravko Culibrk, Vladimir S. Crnojevic, Bernard Gosselin, Thierry Dutoit
ACCV (3)6
2012 Rare: A new bottom-up saliency model
abstract
In this paper, a new bottom-up visual saliency model is proposed. Based on the idea that locally contrasted and globally rare features are salient, this model will be called “RARE” in the following sections. It uses a sequential bottom-up features extraction where first low-level features as luminance and chrominance are computed and from those results medium-level features as image orientations are extracted. A qualitative and a quantitative comparison are achieved on a 120 images dataset. The RARE algorithm powerfully predicts human fixations compared with most of the freely available saliency models.
Nicolas Riche, Matei Mancas, Bernard Gosselin, Thierry Dutoit
ICIP4
2012 Are current gait-related artifact removal techniques useful for low-complexity BCIs?
abstract
A recent study has shown that several gait-related artifact removal techniques are not helpful to improve the performances of an ultra-compact P300-based BCI system integrating only three electrodes and designed for out of the lab experiments. Moreover, the authors advise to use seven electrodes with a standard xDawn spatial filter - known to magnify the P300 response - in order to obtain the best performance/compactedness ratio. However, the positive impact of the xDawn filter and the gaitrelated artifact removal techniques on the BCI performance with seven EEG electrodes had not been precisely evaluated at that time. This is precisely what we do in this study using seven healthy subjects walking at three different speeds. Astonishingly, none of the methods, even the so-called xDawn spatial filter, does significantly outperform the raw data. Thereby, the recommendation considering a low-complexity four-state P300 BCI under ambulatory conditions would be to use filtered raw data without specific gait-related artifact removal techniques.
Matthieu Duvinage, Thierry Castermans, Mathieu Petieau, Guy Cheron, Thierry Dutoit
IJCNN5
2012 Audio and Contact Microphones for Cough Detection
abstract
In the framework of assessing the pathology severity in chronic cough diseases, medical literature underlines the lack of tools for allowing the automatic, objective and reliable detection of cough events. This paper describes a system based on two microphones which we developed for this purpose. The proposed approach relies on a large variety of audio descriptors, an efficient algorithm of feature selection based on their mutual information and the use of artificial neural networks. First, the possible use of a contact microphone (placed on the patient's thorax or trachea) in complement to the audio signal is investigated. This study underlines that this contact microphone suffers from reliability issues, and conveys little new relevant information compared to the audio modality. Secondly, the proposed audio-only approach is compared to a commercially available system using four sensors on a database with different sound categories often misdetected as coughs, and produced in various conditions. With average sensitivity and specificity of 94.7% and 95% respectively, the proposed method achieves better cough detection performance than the commercial system.
Thomas Drugman, Jérôme Urbain, Nathalie Bauwens, Ricardo Chessini, Anne-Sophie Aubriot, Patrick Lebecque, Thierry Dutoit
INTERSPEECH7
2012 Walker Speed Adaptation in Gait Synthesis
Joëlle Tilmanne, Thierry Dutoit
MIG2
2012 Reactive and continuous control of HMM-based speech synthesis
abstract
In this paper, we present a modified version of HTS, called performative HTS or pHTS. The objective of pHTS is to enhance the control ability and reactivity of HTS. pHTS reduces the phonetic context used for training the models and generates the speech parameters within a 2-label window. Speech waveforms are generated on-the-fly and the models can be re-actively modified, impacting the synthesized speech with a delay of only one phoneme. It is shown that HTS and pHTS have comparable output quality. We use this new system to achieve reactive model interpolation and conduct a new test where articulation degree is modified within the sentence.
Maria Astrinaki, Nicolas D'Alessandro, Benjamin Picart, Thomas Drugman, Thierry Dutoit
SLT5
2012 Statistical methods for varying the degree of articulation in new HMM-based voices
abstract
This paper focuses on the automatic modification of the degree of articulation (hypo/hyperarticulation) of an existing standard neutral voice in the framework of HMM-based speech synthesis. Starting from a source speaker for which neutral, hypo and hyperarticulated speech data are available, two sets of transformations are computed during the adaptation of the neutral speech synthesizer. These transformations are then applied to a new target speaker for which no hypo/hyperarticulated recordings are available. Four statistical methods are investigated, differing in the speaking style adaptation technique (MLLR vs. CMLLR) and in the speaking style transposition approach (phonetic vs. acoustic correspondence) they use. This study focuses on the prosody model although such techniques can be applied to any stream of parameters exhibiting suited interpolability properties. Two subjective evaluations are performed in order to determine which statistical transformation method achieves the better segmental quality and reproduction of the articulation degree.
Benjamin Picart, Thomas Drugman, Thierry Dutoit
SLT3
2012 A comparative study of glottal source estimation techniques
Thomas Drugman, Baris Bozkurt, Thierry Dutoit
Comput. Speech Lang.3
2012 The Deterministic Plus Stochastic Model of the Residual Signal and Its Applications
abstract
The modeling of speech production often relies on a source-filter approach. Although methods parameterizing the filter have nowadays reached a certain maturity, there is still a lot to be gained for several speech processing applications in finding an appropriate excitation model. This manuscript presents a Deterministic plus Stochastic Model (DSM) of the residual signal. The DSM consists of two contributions acting in two distinct spectral bands delimited by a maximum voiced frequency. Both components are extracted from an analysis performed on a speaker-dependent dataset of pitch-synchronous residual frames. The deterministic part models the low-frequency contents and arises from an orthonormal decomposition of these frames. As for the stochastic component, it is a high-frequency noise modulated both in time and frequency. Some interesting phonetic and computational properties of the DSM are also highlighted. The applicability of the DSM in two fields of speech processing is then studied. First, it is shown that incorporating the DSM vocoder in HMM-based speech synthesis enhances the delivered quality. The proposed approach turns out to significantly outperform the traditional pulse excitation and provides a quality equivalent to STRAIGHT. In a second application, the potential of glottal signatures derived from the proposed DSM is investigated for speaker identification purpose. Interestingly, these signatures are shown to lead to better recognition rates than other glottal-based methods.
Thomas Drugman, Thierry Dutoit
IEEE Trans. Speech Audio Process.2
2012 Detection of Glottal Closure Instants From Speech Signals: A Quantitative Review
abstract
The pseudo-periodicity of voiced speech can be exploited in several speech processing applications. This requires however that the precise locations of the glottal closure instants (GCIs) are available. The focus of this paper is the evaluation of automatic methods for the detection of GCIs directly from the speech waveform. Five state-of-the-art GCI detection algorithms are compared using six different databases with contemporaneous electroglottographic recordings as ground truth, and containing many hours of speech by multiple speakers. The five techniques compared are the Hilbert Envelope-based detection (HE), the Zero Frequency Resonator-based method (ZFR), the Dynamic Programming Phase Slope Algorithm (DYPSA), the Speech Event Detection using the Residual Excitation And a Mean-based Signal (SEDREAMS) and the Yet Another GCI Algorithm (YAGA). The efficacy of these methods is first evaluated on clean speech, both in terms of reliabililty and accuracy. Their robustness to additive noise and to reverberation is also assessed. A further contribution of the paper is the evaluation of their performance on a concrete application of speech processing: the causal-anticausal decomposition of speech. It is shown that for clean speech, SEDREAMS and YAGA are the best performing techniques, both in terms of identification rate and accuracy. ZFR and SEDREAMS also show a superior robustness to additive noise and reverberation.
Thomas Drugman, Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor, Thierry Dutoit
IEEE Trans. Speech Audio Process.5
2011 A Phonetic Analysis of Natural Laughter, for Use in Automatic Laughter Processing Systems
Jérôme Urbain, Thierry Dutoit
ACII (1)2
2011 Continuous Control of Style through Linear Interpolation in Hidden Markov Model Based Stylistic Walk Synthesis
abstract
In this work, we present a Hidden Markov Model (HMM) based stylistic walk synthesizer, where the synthesized styles are combinations or exaggerations of the walk styles present in the training database. In a first stage, Hidden Markov Models of eleven different styles of gait are trained, using a database of motion capture walk sequences. In a second stage, the probability density functions inside the stylistic models are interpolated or extrapolated in order to synthesize walks with styles or style intensities that were not present in the training database. A continuous model of the style parameter space is thus constructed around the eleven original walk styles. An informal user evaluation of the synthesized sequences showed that the naturalness of motions is preserved after linear interpolation.
Joëlle Tilmanne, Thierry Dutoit
CW2
2011 Phase-based information for voice pathology detection
abstract
In most current approaches of speech processing, information is extracted from the magnitude spectrum. However re cent perceptual studies have underlined the importance of the phase component. The goal of this paper is to investigate the potential of using phase-based features for automatically detecting voice disorders. It is shown that group delay functions are appropriate for characterizing irregularities in the phonation. Besides the respect of the mixed-phase model of speech is discussed. The proposed phase-based features are evaluated and compared to other parameters derived from the magnitude spectrum. Both streams are shown to be interestingly complementary. Furthermore phase-based features turn out to convey a great amount of relevant information, leading to high discrimination performance.
Thomas Drugman, Thomas Dubuisson, Thierry Dutoit
ICASSP3
2011 3D Saliency for Abnormal Motion Selection: The Role of the Depth Map
Nicolas Riche, Matei Mancas, Bernard Gosselin, Thierry Dutoit
ICVS4
2011 Continuous Control of the Degree of Articulation in HMM-Based Speech Synthesis
abstract
This paper focuses on the implementation of a continuous control of the degree of articulation (hypo/hyperarticulation) in the framework of HMM-based speech synthesis. The adaptation of a neutral speech synthesizer to generate hypo and hyperarticulated speech using a limited amount of speech data is first studied. This is done using inter-speaker voice adaptation techniques, applied here to intra-speaker voice adaptation. The implementation of a continuous control of the degree of articulation is then proposed in a second step. Finally, a subjective evaluation shows that good quality neutral/hypo/hyperarticulated speech, and also any intermediate, interpolated or extrapolated articulation degrees, can be obtained from an HMM-based speech synthesizer.
Benjamin Picart, Thomas Drugman, Thierry Dutoit
INTERSPEECH3
2011 Causal-anticausal decomposition of speech using complex cepstrum for glottal source estimation
Thomas Drugman, Baris Bozkurt, Thierry Dutoit
Speech Commun.3
2010 On the potential of glottal signatures for speaker recognition
abstract
Most of current speaker recognition systems are based on features extracted from the magnitude spectrum of speech. However the excitation signal produced by the glottis is expected to convey complementary relevant information about the speaker identity. This paper explores the use of two proposed glottal signatures, derived from the residual signal, for speaker identification. Experiments using these signatures are performed on both TIMIT and YOHO databases. Promising results are shown to outperform other approaches based on glottal features. Besides it is highlighted that the signatures can be used for text-independent speaker recognition and that only several seconds of voiced speech are sufficient for estimating them reliably.
Thomas Drugman, Thierry Dutoit
INTERSPEECH2
2010 Chirp complex cepstrum-based decomposition for asynchronous glottal analysis
abstract
It was recently shown that complex cepstrum can be effectively used for glottal flow estimation by separating the causal and anticausal components of speech. In order to guarantee a correct estimation, some constraints on the window have been derived. Among these, the window has to be synchronized on a Glottal Closure Instant. This paper proposes an extension of the complex cepstrum-based decomposition by incorporating a chirp analysis. The resulting method is shown to give a reliable estimation of the glottal flow wherever the window is located. This technique is then suited for its integration in usual speech processing systems, which generally operate in an asynchronous way. Besides its potential for automatic voice quality analysis is highlighted.
Thomas Drugman, Thierry Dutoit
INTERSPEECH2
2010 Glottal-based analysis of the lombard effect
abstract
The Lombard effect refers to the speech changes due to the immersion of the speaker in a noisy environment. Among these changes, studies have already reported acoustic modifications mainly related to the vocal tract behaviour. In a complementary way, this paper investigates the variation of the glottal flow in Lombard speech. For this, the glottal flow is estimated by a closed-phase analysis and parametrized by a set of time and spectral features. Through a study on a database containing 25 speakers uttering in clean and noisy environments (with 4 noise types at 2 levels), it is highlighted that the glottal source is significantly modified due to the increased vocal effort. Such changes are of interest in several applications of speech processing, such as speech or speaker recognition, or speech synthesis.
Thomas Drugman, Thierry Dutoit
INTERSPEECH2
2010 The AVLaughterCycle Database
Jérôme Urbain, Elisabetta Bevacqua, Thierry Dutoit, Alexis Moinet, Radoslaw Niewiadomski, Catherine Pelachaud, Benjamin Picart, Joëlle Tilmanne, Johannes Wagner 0001
LREC3
2010 Expressive Gait Synthesis Using PCA and Gaussian Modeling
Joëlle Tilmanne, Thierry Dutoit
MIG2
2009 Using a pitch-synchronous residual codebook for hybrid HMM/frame selection speech synthesis
abstract
This paper proposes a method to improve the quality delivered by statistical parametric speech synthesizers. For this, we use a codebook of pitch-synchronous residual frames, so as to construct a more realistic source signal. First a limited codebook of typical excitations is built from some training database. During the synthesis part, HMMs are used to generate filter and source coefficients. The latter coefficients contain both the pitch and a compact representation of target residual frames. The source signal is obtained by concatenating excitation frames picked up from the codebook, based on a selection criterion and taking target residual coefficients as input. Subjective results show a relevant improvement compared to the basic technique.
Thomas Drugman, Alexis Moinet, Thierry Dutoit, Geoffrey Wilfart
ICASSP3
2009 Generating Robot/Agent backchannels during a storytelling experiment
abstract
This work presents the development of a real-time framework for the research of multimodal feedback of robots/talking agents in the context of Human Robot Interaction (HRI) and Human Computer Interaction (HCI). For evaluating the framework, a Multimodal corpus is built (ENTERFACE_STEAD), and a study on the important multimodal features was done for building an active Robot/Agent listener of a storytelling experience with Humans. The experiments show that even when building the same reactive behavior models for Robot and Talking Agents, the interpretation and the realization of the behavior communicated is different due to the different communicative channels Robots/Agents offer be it physical but less human-like in Robots, and virtual but more expressive and human-like in Talking agents.
Sames Al Moubayed, Malek Baklouti, Mohamed Chetouani, Thierry Dutoit, Ammar Mahdhaoui, Jean-Claude Martin, Stanislav Ondás, Catherine Pelachaud, Jérôme Urbain
ICRA4
2009 Cross-language voice conversion based on eigenvoices
abstract
INTERSPEECH2009: 10th Annual Conference of the International Speech Communication Association, September 6-10, 2009, Brighton, UK.
Malorie Charlier, Yamato Ohtani, Tomoki Toda, Alexis Moinet, Thierry Dutoit
INTERSPEECH5
2009 Complex cepstrum-based decomposition of speech for glottal source estimation
abstract
Homomorphic analysis is a well-known method for the separation of non-linearly combined signals. More particularly, the use of complex cepstrum for source-tract deconvolution has been discussed in various articles. However there exists no study which proposes a glottal flow estimation methodology based on cepstrum and reports effective results. In this paper, we show that complex cepstrum can be effectively used for glottal flow estimation by separating the causal and anticausal components of a windowed speech signal as done by the Zeros of the Z-Transform (ZZT) decomposition. Based on exactly the same principles presented for ZZT decomposition, windowing should be applied such that the windowed speech signals exhibit mixed-phase characteristics which conform the speech production model that the anticausal component is mainly due to the glottal flow open phase. The advantage of the complex cepstrum-based approach compared to the ZZT decomposition is its much higher speed.
Thomas Drugman, Baris Bozkurt, Thierry Dutoit
INTERSPEECH3
2009 Glottal closure and opening instant detection from speech signals
abstract
This paper proposes a new procedure to detect Glottal Closure and Opening Instants (GCIs and GOIs) directly from speech waveforms.The procedure is divided into two successive steps.First a mean-based signal is computed, and intervals where speech events are expected to occur are extracted from it.Secondly, at each interval a precise position of the speech event is assigned by locating a discontinuity in the Linear Prediction residual.The proposed method is compared to the DYPSA algorithm on the CMU ARCTIC database.A significant improvement as well as a better noise robustness are reported.Besides, results of GOI identification accuracy are promising for the glottal source characterization.
Thomas Drugman, Thierry Dutoit
INTERSPEECH2
2009 On the mutual information between source and filter contributions for voice pathology detection
abstract
This paper addresses the problem of automatic detection of voice pathologies directly from the speech signal. For this, we investigate the use of the glottal source estimation as a means to detect voice disorders. Three sets of features are proposed, depending on whether they are related to the speech or the glottal signal, or to prosody. The relevancy of these features is assessed through mutual information-based measures. This allows an intuitive interpretation in terms of discrimation power and redundancy between the features, independently of any subsequent classifier. It is discussed which characteristics are interestingly informative or complementary for detecting voice pathologies.
Thomas Drugman, Thomas Dubuisson, Thierry Dutoit
INTERSPEECH3
2009 A deterministic plus stochastic model of the residual signal for improved parametric speech synthesis
abstract
Speech generated by parametric synthesizers generally suffers from a typical buzziness, similar to what was encountered in old LPC-like vocoders.In order to alleviate this problem, a more suited modeling of the excitation should be adopted.For this, we hereby propose an adaptation of the Deterministic plus Stochastic Model (DSM) for the residual.In this model, the excitation is divided into two distinct spectral bands delimited by the maximum voiced frequency.The deterministic part concerns the low-frequency contents and consists of a decomposition of pitch-synchronous residual frames on an orthonormal basis obtained by Principal Component Analysis.The stochastic component is a high-pass filtered noise whose time structure is modulated by an energy-envelope, similarly to what is done in the Harmonic plus Noise Model (HNM).The proposed residual model is integrated within a HMM-based speech synthesizer and is compared to the traditional excitation through a subjective test.Results show a significative improvement for both male and female voices.In addition the proposed model requires few computational load and memory, which is essential for its integration in commercial applications.
Thomas Drugman, Geoffrey Wilfart, Thierry Dutoit
INTERSPEECH3
2008 Dynamic modality weighting for multi-stream hmms inaudio-visual speech recognition
abstract
Merging decisions from different modalities is a crucial problem in Audio-Visual Speech Recognition. To solve this, state synchronous multi-stream HMMs have been proposed for their important advantage of incorporating stream reliability in their fusion scheme. This paper focuses on stream weight adaptation based on modality confidence estimators. We assume different and time-varying environment noise, as can be encountered in realistic applications, and, for this, adaptive methods are best suited. Stream reliability is assessed directly through classifier outputs since they are not specific to either noise type or level. The influence of constraining the weights to sum to one is also discussed.
Mihai Gurban, Jean-Philippe Thiran, Thomas Drugman, Thierry Dutoit
ICMI4
2007 Towards a Voice Conversion System Based on Frame Selection
abstract
The subject of this paper is the conversion of a given speaker's voice (the source speaker) into another identified voice (the target one). We assume we have at our disposal a large amount of speech samples from source and target voice with at least a part of them being parallel. The proposed system is built on a mapping function between source and target spectral envelopes followed by a frame selection algorithm to produce final spectral envelopes. Converted speech is produced by a basic LP analysis of the source and LP synthesis using the converted spectral envelopes. We compared three types of conversion: without mapping, with mapping and using the excitation of the source speaker and finally with mapping using the excitation of the target. Results show that the combination of mapping and frame selection provide the best results, and underline the interest to work on methods to convert the LP excitation.
Thierry Dutoit, Andre Holzapfel, Matthieu Jottrand, Alexis Moinet, Javier Pérez, Yannis Stylianou
ICASSP (4)1
2007 RAMCESS/handsketch: a multi-representation framework for realtime and expressive singing synthesis
Nicolas D'Alessandro, Thierry Dutoit
INTERSPEECH2
2007 Chirp group delay analysis of speech signals
Baris Bozkurt, Laurent Couvreur, Thierry Dutoit
Speech Commun.3
2006 Dynamic Bayesian Networks for NLU Simulation with Applications to Dialog Optimal Strategy Learning
abstract
In this paper, we propose to add a model for NLU-related error generation in a modular environment for computer-based simulation of man-machine spoken dialogs. This model is jointly designed with a user model. Both of them are based on the same underlying Bayesian network used with different parameters in such a way that it can generate a consistent user behavior, according to a goal and the interaction history, and been used as a concept classifier. The proposed simulation environment was used to train a reinforcement-learning algorithm on a simple form-filling task and the results of this experiment show that the addition of the NLU model helps pointing out problematic situations that may occur because of misunderstandings and modifying the dialog strategy accordingly
Olivier Pietquin, Thierry Dutoit
ICASSP (1)2
2006 Multimodal human-computer interfaces
Thierry Dutoit, Laurence Nigay, Michael Schnaider
Signal Process.1
2006 A probabilistic framework for dialog simulation and optimal strategy learning
abstract
The design of Spoken Dialog Systems cannot be considered as the simple combination of speech processing technologies. Indeed, speech-based interface design has been an expert job for a long time. It necessitates good skills in speech technologies and low-level programming. Moreover, rapid development and reusability of previously designed systems remains uneasy. This makes optimality and objective evaluation of design very difficult. The design process is therefore a cyclic process composed of prototype releases, user satisfaction surveys, bug reports and refinements. It is well known that human intervention for testing is time-consuming and above all very expensive. This is one of the reasons for the recent interest in dialog simulation for evaluation as well as for design automation and optimization. In this paper we expose a probabilistic framework for a realistic simulation of spoken dialogs in which the major components of a dialog system are modeled and parameterized thanks to independent data or expert knowledge. Especially, an Automatic Speech Recognition (ASR) system model and a User Model (UM) have been developed. The ASR model, based on articulatory similarities in language models, provides task-adaptive performance prediction and Confidence Level (CL) distribution estimation. The user model relies on the Bayesian Networks (BN) paradigm and is used both for user behavior modeling and Natural Language Understanding (NLU) modeling. The complete simulation framework has been used to train a reinforcement-learning agent on two different tasks. These experiments helped to point out several potentially problematic dialog scenarios.
Olivier Pietquin, Thierry Dutoit
IEEE Trans. Speech Audio Process.2
2005 TTSBOX: a MATLAB toolbox for teaching text-to-speech synthesis
abstract
This paper presents a new toolbox for teaching TTS synthesis. TTSBOX performs the synthesis of Genglish (for "generic English"), an imaginary language obtained by replacing English words by generic words. Genglish therefore has a rather limited lexicon, but its pronunciation maintains most of the problems encountered in natural languages. TTSBOX uses simple data-driven techniques (bigrams, CARTs, NUUs) while trying to keep the code minimal, so as to keep it readable for students with reasonable MATLAB practice. TTSBOX was designed with the hope that it can help to increase the personal involvement of undergraduate and graduate students in their TTS courses.
Thierry Dutoit, Milos Cernak
ICASSP (5)1
2005 Zeros of Z-transform representation with application to source-filter separation in speech
abstract
We propose a new spectral representation called the zeros of z-transform (ZZT), which is an all-zero representation of the z-transform of the signal. We show that separate patterns exist in ZZT representations of speech signals for the glottal flow and the vocal tract contributions. A decomposition method for source-tract separation is presented based on ZZT. The ZZT-decomposition consists in grouping the zeros into two sets, according to their location in the z-plane. This type of decomposition leads to separating glottal flow contribution (without a return phase) from vocal tract contribution in the z domain.
Baris Bozkurt, Boris Doval, Christophe d'Alessandro, Thierry Dutoit
IEEE Signal Process. Lett.4
2004 Unusual teaching short-cuts to the Levinson and lattice algorithms
abstract
The Levinson and lattice algorithms are taught in many signal processing curricula. If only for one reason, the fact that every cell phone solves Yule-Walker equations every 10 ms justifies it all. These algorithms, however, tend to be hard to conceptualize in a few mental images. The paper proposes two such short-cut views, geometric in the wide sense, which have proved to help students "see" the essence of these tools.
Thierry Dutoit
ICASSP (5)1
2004 A method for glottal formant frequency estimation
abstract
This study presents a method for estimation of glottal formant frequency (Fg) from speech signals. Our method is based on zeros of z-transform decomposition of speech spectra into two spectra: glottal flow dominated spectrum and vocal tract dominated spectrum. Peak picking is performed on the amplitude spectrum of the glottal flow dominated part. The algorithm is tested on synthetic speech. It is shown to be effective especially when glottal formant and first formant of vocal tract are not too close. In addition, tests on a real speech example are also presented where open quotient estimates from EGG signals are used as reference and correlated with the glottal formant frequency estimates. 1.
Baris Bozkurt, Thierry Dutoit, Boris Doval, Christophe d'Alessandro
INTERSPEECH2
2004 Improved differential phase spectrum processing for formant tracking
abstract
This study presents an improved version of our previously introduced formant tracking algorithm. The algorithm is based on processing the negative derivative of the argument of the chirp-z transform (termed as the differential phase spectrum) of a given speech signal. No modeling is included in the procedure but only peak picking on differential phase spectrum. We discuss the effects of roots of z-transform to differential phase spectrum and the need to ensure that all zeros are at some distance from the circle where chirp-z transform is computed. For that, we include an additional zero-decomposition step in our previously presented algorithm to improve its robustness. The final version of the algorithm is tested for analysis of synthetic speech and real speech signals and compared to two other formant tracking systems. 1.
Baris Bozkurt, Thierry Dutoit, Boris Doval, Christophe d'Alessandro
INTERSPEECH2
2004 Zeros of z-transform (ZZT) decomposition of speech for source-tract separation
abstract
This study proposes a new spectral decomposition method for source-tract separation. It is based on a new spectral representation called the Zeros of Z-Transform (ZZT), which is an all-zero representation of the z-transform of the signal. We show that separate patterns exist in ZZT representations of speech signals for the glottal flow and the vocal tract contributions. The ZZT-decomposition is simply composed of grouping the zeros into two sets, according to their location in the z-plane. This type of decomposition leads to separating glottal flow contribution (without a return phase) from vocal tract contributions in z domain.
Boris Doval, Baris Bozkurt, Christophe d'Alessandro, Thierry Dutoit
INTERSPEECH4
2003 Aided design of finite-state dialogue management systems
abstract
Due to recent progresses in the field of speech and natural language processing, spoken dialogue systems are becoming more and more common. Nevertheless, the design of complete dialogue systems remains uneasy. On the one hand, developing such a system involves defining a dialogue strategy. Though automatic learning of dialogue strategies has been introduced in several researches, it stays hard to use in practice. On the other hand, system design implies some coding skills and is still a job for specialists despite the emergence of the voiceXML language. In this paper, we describe a graphical interface dedicated to ease the development of dialogue systems. The user of the interface may be helped along his design thanks to automatically learned strategies.
Olivier Pietquin, Thierry Dutoit
ICME2
2003 Text design for TTS speech corpus building using a modified greedy selection
abstract
Speech corpora design is one of the key issues in building high quality text to speech synthesis systems. Often read speech is used since it seems to be the easiest way to obtain a recorded speech corpus with highest control of the content. The main topic of this study is designing text for recording read speech corpora for concatenative text to speech systems. We will discuss application of the greedy algorithm for text selection by proposing a new way of implementing it and comparing with the standard implementation. Additionally, a text corpus design for Turkish TTS is presented. 1.
Baris Bozkurt, Özlem Öztürk, Thierry Dutoit
INTERSPEECH3
2003 Phonetic alignment: speech synthesis-based vs. Viterbi-based
Fabrice Malfrère, Olivier Deroo, Thierry Dutoit, Christophe Ris
Speech Commun.3
2000 EULER: an Open, Generic, Multilingual and Multi-platform Text-to-Speech System
Thierry Dutoit, Michel Bagein, Fabrice Malfrère, Vincent Pagel, Alain Ruelle, Nawfal Tounsi, Dominique Wynsberghe
LREC1
1998 Plug and play software for designing high-level speech processing systems
abstract
Software engineering for research and development in the area of signal processing is by no means unimportant. For speech processing, in particular, it should be a priority: given the intrinsic complexity of text-to-speech or recognition systems, there is little hope to do state-of-the-art research without solid and extensible code. This paper describes a simple and efficient methodology for the design of maximally reusable and extensible software components for speech and signal processing. The resulting programming paradigm allows software components to be advantageously combined with each other in a way that recalls the concept of hardware plug-andplay, without the need for incorporating complex schedulers to control data flows. It has been successfully used for the design of a software library for high-level speech processing systems at AT&T Labs, as well as for several other large-scale software projects.
Thierry Dutoit, Juergen Schroeter
ICSLP1
1998 Phonetic alignment: speech synthesis based vs. hybrid HMM/ANN
abstract
In this paper we compare two different methods for phonetically labeling a speech database. The first approach is based on the alignment of the speech signal on a high quality synthetic speech pattern, and the second one uses a hybrid HMM/ANN system. Both systems have been evaluated on French read utterances from a speaker never seen in the training stage of the HMM/ANN system and manually segmented. This study outlines the advantages and drawbacks of both methods. The high quality speech synthetic system has the great advantage that no training stage is needed, while the classical HMM/ANN system easily allows multiple phonetic transcriptions. We deduce a method for the automatic constitution of phonetically labeled speech databases based on using the synthetic speech segmentation tool to bootstrap the training process of our hybrid HMM/ANN system. The importance of such segmentation tools will be a key point for the development of improved speech synthesis and recognition systems.
Fabrice Malfrère, Olivier Deroo, Thierry Dutoit
ICSLP3
1998 Fully automatic prosody generator for text-to-speech
abstract
Text-to-Prosody systems based on the use of prosodic databases extracted from natural speech will be a key point for further development of new Text-to-Speech systems. This paper describes a system using such speech databases to generate the rhythm and the intonation of a French written text. The system is based on a very crude chinks ’n chunks prosodic phrasing algorithm and on a prosodic analysis of a natural speech database. The rhythm of the synthetic speech is generated with a CART tree trained on a large mono-speaker speech corpus. The acoustic aspect of the intonation is derived from a set of prosodic patterns automatically derived from the same speech corpus. The system has been tested on single sentences and news paragraphs. Informal listening tests have shown that the resulting prosody is convincing most of the time.
Fabrice Malfrère, Thierry Dutoit, Piet Mertens
ICSLP2
1997 High-quality speech synthesis for phonetic speech segmentation
abstract
This paper presents an original technique for solving the phonetic segmentation problem. It is based on the use of a speech synthesizer for the alignment of a text on its corresponding speech signal. A high-quality digital speech synthesizer is used to create a synthetic reference speech pattern used in the alignment process. This approach has the great advantage on other approaches that no training stage (hence no labeled database) is needed. The system has been mainly evaluated on French read utterances. Other evaluations have been made on other languages like English, German, Romanian and Spanish. Following these experiments, the system seems to be a powerful tool for the automatic constitution of large phonetically and prosodically labeled speech databases. The availability of such corpora will be a key point for the development of improved speech synthesis and recognition systems.
Fabrice Malfrère, Thierry Dutoit
EUROSPEECH2
1997 Diphone concatenation using a harmonic plus noise model of speech
abstract
In this paper we present a high-quality text-to-speech system using diphones. The system is based on a Harmonic plus Noise (HNM) representation of the speech signal. HNM is a pitch-synchronous analysis-synthesis system but does not require pitch marks to be determined as necessary in PSOLA-based methods. HNM assumes the speech signal to be composed of a periodic part and a stochastic part. As a result, different prosody and spectral envelope modification methods can be applied to each part, yielding more natural-sounding synthetic speech. The fully parametric representation of speech using HNM also provides a straightforward way of smoothing diphone boundaries. Informal listening tests, using natural prosody, have shown that the synthetic speech quality is close to the quality of the original sentences, without smoothing problems and without buzziness or other oddities observed with other speech representations used for TTS. 1. INTRODUCTION Many current Text-To-Speech (TTS) systems a...
Yannis Stylianou, Thierry Dutoit, Juergen Schroeter
EUROSPEECH2
1997 A simple and efficient algorithm for the compression of MBROLA segment databases
abstract
Most state-of-the-art TTS synthesizers are based on a technique known as synthesis by concatenation, in which speech is produced by concatenating elementary speech units. The design of a high-quality TTS system implies the storage of a large number of segments. To facilitate the storage of these segments, this paper proposes a very low complexity coder to compress unit databases with a toll quality. A particular interest has been taken in the databases used by the MBROLA synthesizer, composed of fixed-length pitch periods with constrained harmonic phases. The coder developed here uses this special characteristic to reach compression rates from 7 to 9 without degrading the speech quality produced by the synthesizer, and with very limited computational cost.
Olivier van der Vrecken, Nicolas Pierret, Thierry Dutoit, Vincent Pagel, Fabrice Malfrère
EUROSPEECH3
1996 The MBROLA project: towards a set of high quality speech synthesizers free of use for non commercial purposes
Thierry Dutoit, Vincent Pagel, Nicolas Pierret, F. Bataille, Olivier van der Vrecken
ICSLP1
1996 On the use of a hybrid harmonic/stochastic model for TTS synthesis-by-concatenation
Thierry Dutoit, Bernard Gosselin
Speech Commun.1
1994 High quality text-to-speech synthesis: a comparison of four candidate algorithms
abstract
We investigate the use of four candidate speech models in the context of high quality text-to-speech systems (HQ-TTS), address problems typically encountered by their prosody matching and segment concatenation modules, and compare their performances regarding: the segment database compression ratio they allow, the computational load of the related synthesis algorithms, as well as their intelligibility and subjective segmental quality. The models addressed are: the classical auto-regressive (LPC) one, the hybrid harmonic/stochastic (H/S) model proposed by Griffin and Lim (1988) and by Abrantes, Marques and Transcoso (1991), the 'null' model, as implemented by the time-domain pitch-synchronous overlap-add (TD-PSOLA) synthesis algorithm, and the multi-band re-synthesis pitch-synchronous overlap-add (MBR-PSOLA) model.>
Thierry Dutoit
ICASSP (1)1
1993 An analysis of the performances of the MBE model when used in the context of a text-to-speech system
Thierry Dutoit, Henri Leich
EUROSPEECH1
1993 MBR-PSOLA: Text-To-Speech synthesis based on an MBE re-synthesis of the segments database
Thierry Dutoit, Henri Leich
Speech Commun.1