EDBT 2026 Demo / reviewers in the wild / expert
Laurent Girin
dblp:64/516
· DBLP profile ↗
82ranked-venue papers
13as first author
14since 2021 · last 2026
0000-0002-9214-8760ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 45 · 7 first-author · 10 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is self-supervised learning enough to fill in the gap? A study on speech inpaintingabstractSpeech inpainting consists in reconstructing corrupted or missing speech segments using surrounding context, a process that closely resembles the pretext tasks in Self-Supervised Learning (SSL) for speech encoders. This study investigates using SSL-trained speech encoders for inpainting without any additional training beyond the initial pretext task, and simply adding a decoder to generate a waveform. We compare this approach to supervised fine-tuning of speech encoders for a downstream task—here, inpainting. Practically, we integrate HuBERT as the SSL encoder and HiFi-GAN as the decoder in two configurations: (1) fine-tuning the decoder to align with the frozen pre-trained encoder’s output and (2) fine-tuning the encoder for an inpainting task based on a frozen decoder’s input. Evaluations are conducted under single- and multi-speaker conditions using in-domain datasets and out-of-domain datasets (including unseen speakers, diverse speaking styles, and noise). Both informed and blind inpainting scenarios are considered, where the position of the corrupted segment is either known or unknown. The proposed SSL-based methods are benchmarked against several baselines, including a text-informed method combining automatic speech recognition with zero-shot text-to-speech synthesis. Performance is assessed using objective metrics and perceptual evaluations. The results demonstrate that both approaches outperform baselines, successfully reconstructing speech segments up to 200 ms, and sometimes up to 400 ms. Notably, fine-tuning the SSL encoder achieves more accurate speech reconstruction in single-speaker settings, while a pre-trained encoder proves more effective for multi-speaker scenarios. This demonstrates that an SSL pretext task can transfer to speech inpainting, enabling successful speech reconstruction with a pre-trained encoder. Ihab Asaad, Maxime Jacquelin, Olivier Perrotin, Laurent Girin, Thomas Hueber |
Comput. Speech Lang. | 4 |
| 2025 | AnCoGen: Analysis, Control and Generation of Speech with a Masked AutoencoderabstractThis article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise ratio, and clarity index. In addition, it can generate speech from these attributes and allow precise control of the synthesized speech by modifying them. Extensive experiments demonstrated the effectiveness of AnCoGen across speech analysis-resynthesis, pitch estimation, pitch modification, and speech enhancement. Code and audio examples are available online1. Samir Sadok, Simon Leglaive, Laurent Girin, Gaël Richard, Xavier Alameda-Pineda |
ICASSP | 3 |
| 2025 | LombardTokenizer: Disentanglement and Control of Vocal Effort in a Neural Speech CodecabstractInternational audience Maxime Jacquelin, Maëva Garnier, Laurent Girin, Rémy Vincent, Olivier Perrotin |
INTERSPEECH | 3 |
| 2024 | A multimodal dynamical variational autoencoder for audiovisual speech representation learningabstractHigh-dimensional data such as natural images or speech signals exhibit some form of regularity, preventing their dimensions from varying independently. This suggests that there exists a lower dimensional latent representation from which the high-dimensional observed data were generated. Uncovering the hidden explanatory features of complex data is the goal of representation learning, and deep latent variable generative models have emerged as promising unsupervised approaches. In particular, the variational autoencoder (VAE) which is equipped with both a generative and an inference model allows for the analysis, transformation, and generation of various types of data. Over the past few years, the VAE has been extended to deal with data that are either multimodal or dynamical (i.e., sequential). In this paper, we present a multimodal and dynamical VAE (MDVAE) applied to unsupervised audiovisual speech representation learning. The latent space is structured to dissociate the latent dynamical factors that are shared between the modalities from those that are specific to each modality. A static latent variable is also introduced to encode the information that is constant over time within an audiovisual speech sequence. The model is trained in an unsupervised manner on an audiovisual emotional speech dataset, in two stages. In the first stage, a vector quantized VAE (VQ-VAE) is learned independently for each modality, without temporal modeling. The second stage consists in learning the MDVAE model on the intermediate representation of the VQ-VAEs before quantization. The disentanglement between static versus dynamical and modality-specific versus modality-common information occurs during this second training stage. Extensive experiments are conducted to investigate how audiovisual speech latent factors are encoded in the latent space of MDVAE. These experiments include manipulating audiovisual speech, audiovisual facial image denoising, and audiovisual speech emotion recognition. The results show that MDVAE effectively combines the audio and visual information in its latent space. They also show that the learned static representation of audiovisual speech can be used for emotion recognition with few labeled data, and with better accuracy compared with unimodal baselines and a state-of-the-art supervised model based on an audiovisual transformer architecture. Samir Sadok, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda, Renaud Séguier |
Neural Networks | 3 |
| 2023 | Speech Modeling with a Hierarchical Transformer Dynamical VAEabstractThe dynamical variational autoencoders (DVAEs) are a family of latent-variable deep generative models that extends the VAE to model a sequence of observed data and a corresponding sequence of latent vectors. In almost all the DVAEs of the literature, the temporal dependencies within each sequence and across the two sequences are modeled with recurrent neural networks. In this paper, we propose to model speech signals with the Hierarchical Transformer DVAE (HiT-DVAE), which is a DVAE with two levels of latent variable (sequence-wise and frame-wise) and in which the temporal dependencies are implemented with the Transformer architecture. We show that HiT-DVAE outperforms several other DVAEs for speech spectrogram modeling, while enabling a simpler training procedure, revealing its high potential for downstream low-level speech processing tasks such as speech enhancement. Xiaoyu Bie, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda |
ICASSP | 4 |
| 2023 | Unsupervised speech enhancement with deep dynamical generative speech and noise modelsabstractThis work builds on a previous work on unsupervised speech enhancement using a dynamical variational autoencoder (DVAE) as the clean speech model and non-negative matrix factorization (NMF) as the noise model. We propose to replace the NMF noise model with a deep dynamical generative model (DDGM) depending either on the DVAE latent variables, or on the noisy observations, or on both. This DDGM can be trained in three configurations: noise-agnostic, noise-dependent and noise adaptation after noise-dependent training. Experimental results show that the proposed method achieves competitive performance compared to state-of-the-art unsupervised speech enhancement methods, while the noise-dependent training configuration yields a much more time-efficient inference process. Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda |
INTERSPEECH | 3 |
| 2023 | Learning and controlling the source-filter representation of speech with a variational autoencoderabstractNational audience Samir Sadok, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda, Renaud Séguier |
Speech Commun. | 3 |
| 2022 | Repeat after Me: Self-Supervised Learning of Acoustic-to-Articulatory Mapping by Vocal ImitationabstractWe propose a computational model of speech production combining a pre-trained neural articulatory synthesizer able to reproduce complex speech stimuli from a limited set of interpretable articulatory parameters, a DNN-based internal forward model predicting the sensory consequences of articulatory commands, and an internal inverse model based on a recurrent neural network recovering articulatory commands from the acoustic speech input. Both forward and inverse models are jointly trained in a self-supervised way from raw acoustic-only speech data from different speakers. The imitation simulations are evaluated objectively and subjectively and display quite encouraging performances. Marc-Antoine Georges, Julien Diard, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber |
ICASSP | 3 |
| 2022 | BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language modelabstractInternational audience Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber |
INTERSPEECH | 3 |
| 2022 | Unsupervised Speech Enhancement Using Dynamical Variational AutoencodersabstractDynamical variational autoencoders (DVAEs) are a class of deep generative models with latent variables, dedicated to model time series of high-dimensional data. DVAEs can be considered as extensions of the variational autoencoder (VAE) that include temporal dependencies between successive observed and/or latent vectors. Previous work has shown the interest of using DVAEs over the VAE for speech spectrograms modeling. Independently, the VAE has been successfully applied to speech enhancement in noise, in an unsupervised noise-agnostic set-up that requires neither noise samples nor noisy speech samples at training time, but only requires clean speech signals. In this paper, we extend these works to DVAE-based single-channel unsupervised speech enhancement, hence exploiting both speech signals unsupervised representation learning and dynamics modeling. We propose an unsupervised speech enhancement algorithm that combines a DVAE speech prior pre-trained on clean speech signals with a noise model based on nonnegative matrix factorization, and we derive a variational expectation-maximization (VEM) algorithm to perform speech enhancement. The algorithm is presented with the most general DVAE formulation and is then applied with three specific DVAE models to illustrate the versatility of the framework. Experimental results show that the proposed DVAE-based approach outperforms its VAE-based counterpart, as well as several supervised and unsupervised noise-dependent baselines, especially when the noise type is unseen during training. Xiaoyu Bie, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | A Benchmark of Dynamical Variational Autoencoders Applied to Speech Spectrogram ModelingabstractAccepted to Interspeech 2021. arXiv admin note: text overlap with arXiv:2008.12595 Xiaoyu Bie, Laurent Girin, Simon Leglaive, Thomas Hueber, Xavier Alameda-Pineda |
Interspeech | 2 |
| 2021 | Learning Robust Speech Representation with an Articulatory-Regularized Variational AutoencoderabstractIt is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task. Marc-Antoine Georges, Laurent Girin, Jean-Luc Schwartz, Thomas Hueber |
Interspeech | 2 |
| 2021 | Alternate Endings: Improving Prosody for Incremental Neural TTS with Predicted Future Text InputabstractThe prosody of a spoken word is determined by its surrounding context. In incremental text-to-speech synthesis, where the synthesizer produces an output before it has access to the complete input, the full context is often unknown which can result in a loss of naturalness in the synthesized speech. In this paper, we investigate whether the use of predicted future text can attenuate this loss. We compare several test conditions of next future word: (a) unknown (zero-word), (b) language model predicted, (c) randomly predicted and (d) ground-truth. We measure the prosodic features (pitch, energy and duration) and find that predicted text provides significant improvements over a zero-word lookahead, but only slight gains over random-word lookahead. We confirm these results with a perceptive test. Brooke Stephenson, Thomas Hueber, Laurent Girin, Laurent Besacier |
Interspeech | 3 |
| 2021 | Variational Bayesian Inference for Audio-Visual Tracking of Multiple SpeakersabstractIn this article, we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature and roles of these two modalities in order to accurately estimate smooth trajectories of the tracked persons, to deal with the partial or total absence of one of the modalities over short periods of time, and to estimate the acoustic status-either speaking or silent-of each tracked person over time. We propose to cast the problem at hand into a generative audio-visual fusion (or association) model formulated as a latent-variable temporal graphical model. This may well be viewed as the problem of maximizing the posterior joint distribution of a set of continuous and discrete latent variables given the past and current observations, which is intractable. We propose a variational inference model which amounts to approximate the joint distribution with a factorized distribution. The solution takes the form of a closed-form expectation maximization procedure. We describe in detail the inference algorithm, we evaluate its performance and we compare it with several baseline methods. These experiments show that the proposed audio-visual tracker performs well in informal meetings involving a time-varying number of people. Yutong Ban, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | A Recurrent Variational Autoencoder for Speech EnhancementabstractThis paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix factorization noise model for speech enhancement. We propose a variational expectation-maximization algorithm where the encoder of the RVAE is finetuned at test time, to approximate the distribution of the latent variables given the noisy speech observations. Compared with previous approaches based on feed-forward fully-connected architectures, the proposed recurrent deep generative speech model induces a posterior temporal dynamic over the latent variables, which is shown to improve the speech enhancement results. Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
ICASSP | 3 |
| 2020 | What the Future Brings: Investigating the Impact of Lookahead for Incremental Neural TTSabstractInternational audience Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber |
INTERSPEECH | 3 |
| 2020 | Evaluating the Potential Gain of Auditory and Audiovisual Speech-Predictive Coding Using Deep LearningabstractSensory processing is increasingly conceived in a predictive framework in which neurons would constantly process the error signal resulting from the comparison of expected and observed stimuli. Surprisingly, few data exist on the accuracy of predictions that can be computed in real sensory scenes. Here, we focus on the sensory processing of auditory and audiovisual speech. We propose a set of computational models based on artificial neural networks (mixing deep feedforward and convolutional networks), which are trained to predict future audio observations from present and past audio or audiovisual observations (i.e., including lip movements). Those predictions exploit purely local phonetic regularities with no explicit call to higher linguistic levels. Experiments are conducted on the multispeaker LibriSpeech audio speech database (around 100 hours) and on the NTCD-TIMIT audiovisual speech database (around 7 hours). They appear to be efficient in a short temporal range (25-50 ms), predicting 50% to 75% of the variance of the incoming stimulus, which could result in potentially saving up to three-quarters of the processing power. Then they quickly decrease and almost vanish after 250 ms. Adding information on the lips slightly improves predictions, with a 5% to 10% increase in explained variance. Interestingly the visual gain vanishes more slowly, and the gain is maximum for a delay of 75 ms between image and predicted sound. Thomas Hueber, Eric Tatulli, Laurent Girin, Jean-Luc Schwartz |
Neural Comput. | 3 |
| 2020 | Audio-Visual Speech Enhancement Using Conditional Variational Auto-EncodersabstractVariational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. One advantage of this generative approach is that it does not require pairs of clean and noisy speech signals at training. In this article, we propose audio-visual variants of VAEs for single-channel and speaker-independent speech enhancement. We develop a conditional VAE (CVAE) where the audio speech generative process is conditioned on visual information of the lip region. At test time, the audio-visual speech generative model is combined with a noise model based on nonnegative matrix factorization, and speech enhancement relies on a Monte Carlo expectation-maximization algorithm. Experiments are conducted with the recently published NTCD-TIMIT dataset as well as the GRID corpus. The results confirm that the proposed audio-visual CVAE effectively fuses audio and visual information, and it improves the speech enhancement performance compared with the audio-only VAE model, especially when the speech signal is highly corrupted by noise. We also show that the proposed unsupervised audio-visual speech enhancement approach outperforms a state-of-the-art supervised deep learning method. Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Semi-supervised Multichannel Speech Enhancement with Variational Autoencoders and Non-negative Matrix FactorizationabstractIn this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neural network for modeling the speech spectro-temporal content. The parameters of this supervised model are learned using the framework of variational autoencoders. The noisy recording environment is supposed to be unknown, so the noise spectro-temporal modeling remains unsupervised and is based on non-negative matrix factorization (NMF). We develop a Monte Carlo expectation-maximization algorithm and we experimentally show that the proposed approach outperforms its NMF-based counterpart, where speech is modeled using supervised NMF. Simon Leglaive, Laurent Girin, Radu Horaud |
ICASSP | 2 |
| 2019 | Speech Enhancement with Variational Autoencoders and Alpha-stable DistributionsabstractThis paper focuses on single-channel semi-supervised speech enhancement. We learn a speaker-independent deep generative speech model using the framework of variational autoencoders. The noise model remains unsupervised because we do not assume prior knowledge of the noisy recording environment. In this context, our contribution is to propose a noise model based on alpha-stable distributions, instead of the more conventional Gaussian non-negative matrix factorization approach found in previous studies. We develop a Monte Carlo expectation-maximization algorithm for estimating the model parameters at test time. Experimental results show the superiority of the proposed approach both in terms of perceptual quality and intelligibility of the enhanced speech signal. Simon Leglaive, Umut Simsekli, Antoine Liutkus, Laurent Girin, Radu Horaud |
ICASSP | 4 |
| 2019 | Audio-Visual Variational Fusion for Multi-Person Tracking with RobotsabstractRobust multi-person tracking with robots opens the door to analysing engagement and social signals in real-world environments. Multi-person scenarios are charaterised by (i) a time-varying number of people, (ii) intermittent auditory (\eg speech turns) and visual cues (\eg person appearing/disappearing) and (iii) impact of the robot actions in perception. The various sensors (cameras and microphones) available for perception, provide a rich flow of information of intermittent and complementary nature. How to jointly exploit these cues to tackle the multi-person tracking problem with an autonomous system has been an intense research line of the Perception Team in the past few years. In this demo we want to present our, now mature, achievements in the field, and demonstrate two robotic systems able to track multiple persons using auditory and visual cues, when they are available. We will bring the two robots and the necessary computing resources with us, as well as the required presentation materials to discuss the models, methods and tools supporting this technology with the attendants. Xavier Alameda-Pineda, Soraya Arias, Yutong Ban, Guillaume Delorme 0002, Laurent Girin, Radu Horaud, Xiaofei Li 0001, Bastien Mourgue, Guillaume Sarrazin |
ACM Multimedia | 5 |
| 2019 | Assessing the performances of different neural network architectures for the detection of screams and shouts in public transportation
Pierre Laffitte, David Sodoyer, Laurent Girin |
Expert Syst. Appl. | 4 |
| 2019 | Audio-Noise Power Spectral Density Estimation Using Long Short-Term MemoryabstractWe propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short-time Fourier transform (STFT) domain. An LSTM network common to all frequency bands is trained, which processes each frequency band individually by mapping the noisy STFT magnitude sequence to its corresponding noise PSD sequence. Unlike deep-learning-based speech-enhancement methods, which learn the full-band spectral structure of speech segments, the proposed method exploits the sub-band STFT magnitude evolution of noise with long time dependence, in the spirit of the unsupervised noise estimators described in the literature. Speaker- and speech-independent experiments with different types of noise show that the proposed method outperforms the unsupervised estimators, and it generalizes well to noise types that are not present in the training set. Xiaofei Li 0001, Simon Leglaive, Laurent Girin, Radu Horaud |
IEEE Signal Process. Lett. | 3 |
| 2019 | Multichannel Speech Separation and Enhancement Using the Convolutive Transfer FunctionabstractThis paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, assuming known mixing filters. We propose to perform speech separation and enhancement in the short-time Fourier transform domain using the convolutive transfer function (CTF) approximation. Compared to time-domain filters, the CTF has much less taps. Consequently, it requires less computational cost and sometimes is more robust against the filter perturbations. We propose three methods: 1) for the multisource case, the multichannel inverse filtering method, i.e., the multiple input/output inverse theorem (MINT), is exploited in the CTF domain; 2) a beamforming-like multichannel inverse filtering method applying the single-source MINT and using power minimization, which is suitable whenever the source CTFs are not all known; and 3) a basis pursuit method, where the sources are recovered by minimizing their ℓ1-norm to impose spectral sparsity, while the ℓ2-norm fitting cost between microphone signals and mixing model is constrained to be lower than a tolerance. The noise can be reduced by setting this tolerance at the noise power level. Experiments under various acoustic conditions are carried out to evaluate and compare the three proposed methods. Comparison with four baseline methods-beamforming-based, two time-domain inverse filters, and time-domain Lasso-shows the applicability of the proposed methods. Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Multichannel Online Dereverberation Based on Spectral Magnitude Inverse FilteringabstractThis paper addresses the problem of multichannel online dereverberation. The proposed method is carried out in the short-time Fourier transform (STFT) domain, and for each frequency band independently. In the STFT domain, the time-domain room impulse response is approximately represented by the convolutive transfer function (CTF). The multichannel CTFs are adaptively identified based on the cross-relation method, and using the recursive least square criterion. Instead of the complex-valued CTF convolution model, we use a nonnegative convolution model between the STFT magnitude of the source signal and the CTF magnitude, which is just a coarse approximation of the former model, but is shown to be more robust against the CTF perturbations. Based on this nonnegative model, we propose an online STFT magnitude inverse filtering method. The inverse filters of the CTF magnitude are formulated based on the multiple-input/output inverse theorem, and adaptively estimated based on the gradient descent criterion. Finally, the inverse filtering is applied to the STFT magnitude of the microphone signals, obtaining an estimate of the STFT magnitude of the source signal. Experiments regarding both speech enhancement and automatic speech recognition are conducted, which demonstrate that the proposed method can effectively suppress reverberation, even for the difficult case of a moving speaker. Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Accounting for Room Acoustics in Audio-Visual Multi-Speaker TrackingabstractMultiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reverberation. We investigate how to combine state-of-the-art audio-source localization techniques with Bayesian multi-person tracking. Our experiments demonstrate that the performance of the proposed system is not affected by changes in the acoustic environment. Yutong Ban, Xiaofei Li 0001, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
ICASSP | 4 |
| 2018 | Multisource Mint Using Convolutive Transfer FunctionabstractThe multichannel inverse filtering method, i.e. multiple input/output inverse theorem (MINT), is widely used. However, it is usually performed in the time domain, and based on the long room impulse responses, thus it has a high computational complexity and a large number of near-common zeros. In this paper, we propose to perform MINT in the short-time Fourier transform (STFT) domain, in which the time-domain filter is approximated by the convolutive transfer function. The oversampled STFT is used to avoid frequency aliasing, which however leads to a common zero region in the subband frequency response due to the frequency response of the STFT window. A new inverse filtering target function concerning the STFT window is proposed to overcome this problem. In addition, unlike most studies using MINT for single source dereverberation, the multisource MINT is proposed for both source separation and dereverberation. Xiaofei Li 0001, Sharon Gannot, Laurent Girin, Radu Horaud |
ICASSP | 3 |
| 2018 | Multichannel Identification and Nonnegative Equalization for Dereverberation and Noise Reduction Based on Convolutive Transfer FunctionabstractThis paper addresses the problems of blind multichannel identification and equalization for joint speech dereverberation and noise reduction. The time-domain cross-relation method is hardly applicable for blind room impulse response identification due to the near-common zeros of the long impulse responses. We extend the cross-relation method to the short-time Fourier transform (STFT) domain, in which the time-domain impulse response is approximately represented by the convolutive transfer function (CTF) with much less coefficients. For the oversampled STFT, CTFs suffer from the common zeros caused by the nonflat frequency response of the STFT window. To overcome this, we propose to identify CTFs using the STFT framework with oversampled signals and critically sampled CTFs, which is a good tradeoff between the frequency aliasing of the signals and the common zeros problem of CTFs. The identified complex-valued CTFs are not accurate enough for multichannel equalization due to the frequency aliasing of the CTFs. Hence, we only use the CTF magnitudes, which leads to a nonnegative multichannel equalization method based on a nonnegative convolution model between the STFT magnitude of the source signal and the CTF magnitude. Compared with the complex-valued convolution model, this nonnegative convolution model is shown to be more robust against the CTF perturbations. To recover the STFT magnitude of the source signal and to reduce the additive noise, the l2-norm fitting error between the STFT magnitude of the microphone signals and the nonnegative convolution is constrained to be less than a noise power related tolerance. Meanwhile, the l1-norm of the STFT magnitude of the source signal is minimized to impose the sparsity. Xiaofei Li 0001, Sharon Gannot, Laurent Girin, Radu Horaud |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixturesabstractWe present a probabilistic model for joint source separation and diarisation of multichannel convolutive speech mixtures. We build upon the framework of local Gaussian model (LGM) with non-negative matrix factorization (NMF). The diarisation is introduced as a temporal labeling of each source in the mix as active or inactive at the short-term frame level. We devise an EM algorithm in which the source separation process is aided by the diarisation state, since the latter indicates the sources actually present in the mixture. The diarisation state is tracked with a Hidden Markov Model (HMM) with emission probabilities calculated from the estimated source signals. The proposed EM has separation performance comparable with a state-of-the-art LGM NMF method, while outperforming a state-of-the-art speaker diarisation pipeline. Dionyssos Kounades-Bastian, Laurent Girin, Xavier Alameda-Pineda, Sharon Gannot, Radu Horaud |
ICASSP | 2 |
| 2017 | Audio source separation based on convolutive transfer function and frequency-domain lasso optimizationabstractThis paper addresses the problem of under-determined convolutive audio source separation in a semi-oracle configuration where the mixing filters are assumed to be known. We propose a separation procedure based on the convolutive transfer function (CTF), which is a more appropriate model for strongly reverberant signals than the widely-used multiplicative transfer function approximation. In the short-time Fourier transform domain, source signals are estimated by minimizing the mixture fitting cost using Lasso optimization, with a ℓ1-norm regularization to exploit the spectral sparsity of source signals. Experiments show that the proposed method achieves satisfactory performance on highly reverberant speech mixtures, with a much lower computational cost compared to time-domain dual techniques. Xiaofei Li 0001, Laurent Girin, Radu Horaud |
ICASSP | 2 |
| 2017 | Automatic animation of an articulatory tongue model from ultrasound images of the vocal tract
Diandra Fabre, Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, Pierre Badin |
Speech Commun. | 3 |
| 2017 | Extending the Cascaded Gaussian Mixture Regression Framework for Cross-Speaker Acoustic-Articulatory MappingabstractThis paper addresses the adaptation of an acoustic-articulatory inversion model of a reference speaker to the voice of another source speaker, using a limited amount of audio-only data. In this study, the articulatory-acoustic relationship of the reference speaker is modeled by a Gaussian mixture model and inference of articulatory data from acoustic data is made by the associated Gaussian mixture regression (GMR). To address speaker adaptation, we previously proposed a general framework called Cascaded-GMR (C-GMR) which decomposes the adaptation process into two consecutive steps: spectral conversion between source and reference speaker and acoustic-articulatory inversion of converted spectral trajectories. In particular, we proposed the integrated C-GMR technique (IC-GMR) in which both steps are tied together in the same probabilistic model. In this paper, we extend the C-GMR framework with another model called Joint-GMR (J-GMR). Contrary to the IC-GMR, this model aims at exploiting all potential acoustic-articulatory relationships, including those between the source speaker's acoustics and the reference speaker's articulation. We present the full derivation of the exact expectation-maximization (EM) training algorithm for the J-GMR. It exploits the missing data methodology of machine learning to deal with limited adaptation data. We provide an extensive evaluation of the J-GMR on both synthetic acoustic-articulatory data and on the multispeaker MOCHA EMA database. We compare the J-GMR performance to other models of the C-GMR framework, notably the IC-GMR, and discuss their respective merits. Laurent Girin, Thomas Hueber, Xavier Alameda-Pineda |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Multiple-Speaker Localization Based on Direct-Path Features and Likelihood Maximization With Spatial Sparsity RegularizationabstractThis paper addresses the problem of multiple-speaker localization in noisy and reverberant environments, using binaural recordings of an acoustic scene. A complex-valued Gaussian mixture model (CGMM) is adopted, whose components correspond to all the possible candidate source locations defined on a grid. After optimizing the CGMM-based objective function, given an observed set of complex-valued binaural features, both the number of sources and their locations are estimated by selecting the CGMM components with the largest weights. An entropy-based penalty term is added to the likelihood to impose sparsity over the set of CGMM component weights. This favors a small number of detected speakers with respect to the large number of initial candidate source locations. In addition, the direct-path relative transfer function (DP-RTF) is used to build robust binaural features. The DP-RTF, recently proposed for single-source localization, encodes interchannel information corresponding to the direct path of sound propagation and is thus robust to reverberations. In this paper, we extend the DP-RTF estimation to the case of multiple sources. In the short-time Fourier transform domain, a consistency test is proposed to check whether a set of consecutive frames is associated with the same source or not. Reliable DP-RTF features are selected from the frames that pass the consistency test to be used for source localization. Experiments carried out using both simulation data and real data recorded with a robotic head confirm the efficiency of the proposed multisource localization method. Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | An inverse-gamma source variance prior with factorized parameterization for audio source separationabstractIn this paper we present a new statistical model for the power spectral density (PSD) of an audio signal and its application to multichannel audio source separation (MASS). The source signal is modeled with the local Gaussian model (LGM) and we propose to model its variance with an inverse-Gamma distribution, whose scale parameter is factorized as a rank-1 model. We discuss the interest of this approach and evaluate it in a MASS task with underdetermined convolutive mixtures. For this aim, we derive a variational EM algorithm for parameter estimation and source inference. The proposed model shows a benefit in source separation performance compared to a state-of-the-art LGM NMF-based technique. Dionyssos Kounades-Bastian, Laurent Girin, Xavier Alameda-Pineda, Sharon Gannot, Radu Horaud |
ICASSP | 2 |
| 2016 | Deep neural networks for automatic detection of screams and shouted speech in subway trainsabstractDeep Neural Networks (DNNs) have recently become a popular technique for regression and classification problems. Their capacity to learn high-order correlations between input and output data proves to be very powerful for automatic speech recognition. In this paper we investigate the use of DNNs for automatic scream and shouted speech detection, within the framework of surveillance systems in public transportation. We recorded a database of sounds occurring in subway trains in real conditions of exploitation and used DNNs to classify the sounds into screams, shouts and other categories. We report encouraging results, given the difficulty of the task, especially when a high level of surrounding noise is present. Pierre Laffitte, David Sodoyer, Charles Tatkeu, Laurent Girin |
ICASSP | 4 |
| 2016 | Non-stationary noise power spectral density estimation based on regional statisticsabstractEstimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past and present periodograms in a short-time period. We show that these features are efficient in characterizing the statistical difference between noise PSD and noisy speech PSD. We therefore propose to use these features for estimating the speech presence probability (SPP). The noise PSD is recursively estimated by averaging past spectral power values with a time-varying smoothing parameter controlled by the SPP. The proposed method exhibits good tracking capability for non-stationary noise, even for abruptly increasing noise level. Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud |
ICASSP | 2 |
| 2016 | Reverberant sound localization with a robot head based on direct-path relative transfer functionabstractThis paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for human-robot interaction. The microphone-pair response corresponding to the direct-path sound propagation is a function of the source direction. In practice, this response is contaminated by noise and reverberations. The direct-path relative transfer function (DP-RTF) is defined as the ratio between the direct-path acoustic transfer function (ATF) of the two microphones, and it is an important feature for SSL. We propose a method to estimate the DP-RTF from noisy and reverberant signals in the short-time Fourier transform (STFT) domain. First, the convolutive transfer function (CTF) approximation is adopted to accurately represent the impulse response of the microphone array, and the first coefficient of the CTF is mainly composed of the direct-path ATF. At each frequency, the frame-wise speech auto- and cross-power spectral density (PSD) are obtained by spectral subtraction. Then a set of linear equations is constructed by the speech auto- and cross-PSD of multiple frames, in which the DP-RTF is an unknown variable, and is estimated by solving the equations. Finally, the estimated DP-RTFs are concatenated across frequencies and used as a feature vector for SSL. Experiments with a robot, placed in various reverberant environments, show that the proposed method outperforms two state-of-the-art methods. Xiaofei Li 0001, Laurent Girin, Fabien Badeig, Radu Horaud |
IROS | 2 |
| 2016 | Real-Time Control of an Articulatory-Based Speech Synthesizer for Brain Computer InterfacesabstractRestoring natural speech in paralyzed and aphasic people could be achieved using a Brain-Computer Interface (BCI) controlling a speech synthesizer in real-time. To reach this goal, a prerequisite is to develop a speech synthesizer producing intelligible speech in real-time with a reasonable number of control parameters. We present here an articulatory-based speech synthesizer that can be controlled in real-time for future BCI applications. This synthesizer converts movements of the main speech articulators (tongue, jaw, velum, and lips) into intelligible speech. The articulatory-to-acoustic mapping is performed using a deep neural network (DNN) trained on electromagnetic articulography (EMA) data recorded on a reference speaker synchronously with the produced speech signal. This DNN is then used in both offline and online modes to map the position of sensors glued on different speech articulators into acoustic parameters that are further converted into an audio signal using a vocoder. In offline mode, highly intelligible speech could be obtained as assessed by perceptual evaluation performed by 12 listeners. Then, to anticipate future BCI applications, we further assessed the real-time control of the synthesizer by both the reference speaker and new speakers, in a closed-loop paradigm using EMA data recorded in real time. A short calibration period was used to compensate for differences in sensor positions and articulatory differences between new speakers and the reference speaker. We found that real-time synthesis of vowels and consonants was possible with good intelligibility. In conclusion, these results open to future speech BCI applications using such articulatory-based speech synthesizer. Florent Bocquelet, Thomas Hueber, Laurent Girin, Christophe Savariaux, Blaise Yvert |
PLoS Comput. Biol. | 3 |
| 2016 | A Variational EM Algorithm for the Separation of Time-Varying Convolutive Audio MixturesabstractThis paper addresses the problem of separating audio sources from time-varying convolutive mixtures. We propose a probabilistic framework based on the local complex-Gaussian model combined with non-negative matrix factorization. The time-varying mixing filters are modeled by a continuous temporal stochastic process. We present a variational expectation-maximization (VEM) algorithm that employs a Kalman smoother to estimate the time-varying mixing matrix, and that jointly estimate the source parameters. The sound sources are then separated by Wiener filters constructed with the estimators provided by the VEM algorithm. Extensive experiments on simulated data show that the proposed method outperforms a blockwise version of a state-of-the-art baseline method. Dionyssos Kounades-Bastian, Laurent Girin, Xavier Alameda-Pineda, Sharon Gannot, Radu Horaud |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Estimation of the Direct-Path Relative Transfer Function for Supervised Sound-Source LocalizationabstractThis paper addresses the problem of sound-source localization of a single speech source in noisy and reverberant environments. For a given binaural microphone setup, the binaural response corresponding to the direct-path propagation of a single source is a function of the source direction. In practice, this response is contaminated by noise and reverberations. The direct-path relative transfer function (DP-RTF) is defined as the ratio between the direct-path acoustic transfer function of the two channels. We propose a method to estimate the DP-RTF from the noisy and reverberant microphone signals in the short-time Fourier transform (STFT) domain. First, the convolutive transfer function approximation is adopted to accurately represent the impulse response of the sensors in the STFT domain. Second, the DP-RTF is estimated by using the auto- and cross-power spectral densities at each frequency and over multiple frames. In the presence of stationary noise, an interframe spectral subtraction algorithm is proposed, which enables to achieve the estimation of noise-free auto- and cross-power spectral densities. Finally, the estimated DP-RTFs are concatenated across frequencies and used as a feature vector for the localization of speech source. Experiments with both simulated and real data show that the proposed localization method performs well, even under severe adverse acoustic conditions, and outperforms state-of-the-art localization methods under most of the acoustic conditions. Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtractionabstractThis paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresponding to speech-plus-noise activity and noise-only. Then, the subtraction of two segmental PSD matrices leads to an almost noise-free PSD matrix by reducing the stationary noise component and preserving non-stationary speech component. This noise-free PSD matrix is used for single speaker RTF identification by eigenvalue decomposition. Experiments are performed in the context of sound source localization to evaluate the efficiency of the proposed method. Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot |
ICASSP | 2 |
| 2015 | Real-time control of a DNN-based articulatory synthesizer for silent speech conversion: a pilot studyabstractInternational audience Florent Bocquelet, Thomas Hueber, Laurent Girin, Christophe Savariaux, Blaise Yvert |
INTERSPEECH | 3 |
| 2015 | Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear RegressionabstractThis paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on monaural segregation. The method starts with a training stage that establishes a locally linear Gaussian regression model between the directional coordinates of all the sources and the auditory features extracted from binaural measurements. While fixed-length wide-spectrum sounds (white noise) are used for training to reliably estimate the model parameters, we show that the testing (localization) can be extended to variable-length sparse-spectrum sounds (such as speech), thus enabling a wide range of realistic applications. Indeed, we demonstrate that the method can be used for audio-visual fusion, namely to map speech signals onto images and hence to spatially align the audio and visual modalities, thus enabling to discriminate between speaking and non-speaking faces. We release a novel corpus of real-room recordings that allow quantitative evaluation of the co-localization method in the presence of one or two sound sources. Experiments demonstrate increased accuracy and speed relative to several state-of-the-art methods. Antoine Deleforge, Radu Horaud, Yoav Y. Schechner, Laurent Girin |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Speaker-Adaptive Acoustic-Articulatory Inversion Using Cascaded Gaussian Mixture RegressionabstractThis paper addresses the adaptation of an acoustic-articulatory model of a reference speaker to the voice of another speaker, using a limited amount of audio-only data. In the context of pronunciation training, a virtual talking head displaying the internal speech articulators (e.g., the tongue) could be automatically animated by means of such a model using only the speaker's voice. In this study, the articulatory-acoustic relationship of the reference speaker is modeled by a gaussian mixture model (GMM). To address the speaker adaptation problem, we propose a new framework called cascaded Gaussian mixture regression (C-GMR), and derive two implementations. The first one, referred to as Split-C-GMR, is a straightforward chaining of two distinct GMRs: one mapping the acoustic features of the source speaker into the acoustic space of the reference speaker, and the other estimating the articulatory trajectories with the reference model. In the second implementation, referred to as Integrated-C-GMR, the two mapping steps are tied together in a single probabilistic model. For this latter model, we present the full derivation of the exact EM training algorithm, that explicitly exploits the missing data methodology of machine learning. Other adaptation schemes based on maximum-a posteriori (MAP), maximum likelihood linear regression (MLLR) and direct cross-speaker acoustic-to-articulatory GMR are also investigated. Experiments conducted on two speakers for different amount of adaptation data show the interest of the proposed C-GMR techniques. Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, Gérard Bailly |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Sound representation and classification benchmark for domestic robotsabstractWe address the problem of sound representation and classification and present results of a comparative study in the context of a domestic robotic scenario. A dataset of sounds was recorded in realistic conditions (background noise, presence of several sound sources, reverberations, etc.) using the humanoid robot NAO. An extended benchmark is carried out to test a variety of representations combined with several classifiers. We provide results obtained with the annotated dataset and we assess the methods quantitatively on the basis of their classification scores, computation times and memory requirements. The annotated dataset is publicly available at https://team.inria.fr/perception/nard/. Maxime Janvier, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud |
ICRA | 3 |
| 2014 | Robust articulatory speech synthesis using deep neural networks for BCI applicationsabstractBrain-Computer Interfaces (BCIs) usually propose typing strategies to restore communication for paralyzed and aphasic people. A more natural way would be to use speech BCI directly controlling a speech synthesizer. Toward this goal, a prerequisite is the development a synthesizer that should i) produce intelligible speech, ii) run in real time, iii) depend on as few parameters as possible, and iv) be robust to error fluctuations on the control parameters. In this context, we describe here an articulatory-to-acoustic mapping approach based on deep neural network (DNN) trained on electromagnetic articulography (EMA) data recorded synchronously with produced speech sounds. On this corpus, the DNN-based model provided a speech synthesis quality (as assessed by automatic speech recognition and behavioral testing) comparable to a state-of-the-art Gaussian mixture model (GMM), yet showing higher robustness when noise was added to the EMA coordinates. Moreover, to envision BCI applications, this robustness was also assessed when the space covered by the 12 original articulatory parameters was reduced to 7 parameters using deep auto-encoders (DAE). Given that this method can be implemented in real time, DNN-based articulatory speech synthesis seems a good candidate for speech BCI applications. Index Terms: articulatory speech synthesis, brain computer interface (BCI), deep neural networks, deep auto-encoder, EMA, noise robustness, dimensionality reduction Florent Bocquelet, Thomas Hueber, Laurent Girin, Pierre Badin, Blaise Yvert |
INTERSPEECH | 3 |
| 2013 | Informed Source Separation from compressed mixtures using spatial wiener filter and quantization noise estimationabstractIn a previous work, we proposed an Informed Source Separation system based on Wiener filtering for active listening of music from uncompressed (16-bit PCM) multichannel mix signals. In the present work, the system is improved to work with (MPEG-2 AAC) compressed mix signals: quantization noise is estimated from the AAC bitstream at the decoder and explicitly taken into account in the source separation process. Also a direct MDCT-to-STFT transform is used to optimize the computational efficiency of the process in the STFT domain from AAC-decoded MDCT coefficients. Laurent Girin, Antoine Liutkus |
ICASSP | 2 |
| 2013 | Fast and Accurate Direct MDCT to DFT Conversion With Arbitrary Window FunctionsabstractIn this paper, we propose a method for direct conversion of MDCT coefficients to DFT coefficients, without passing through time signal reconstruction. In contrast to previous works, this method is valid for any pair of MDCT and DFT window functions. It is based on the decomposition of the MDCT-to-DFT conversion matrices into a Toeplitz part plus a Hankel part. The latter is split, then mirrored and combined with the former to construct a global Toeplitz matrix. This leads to a fast FIR filtering implementation of the conversion process. The filter taps are DFT coefficients of window functions products, and concentrate most of their energy in a few low-frequency taps. The conversion can thus be efficiently approximated by keeping only a few most significant taps, as confirmed by numerical experiments: For example, for frame size of 2048, Hanning-windowed DFT is obtained from KBD-windowed MDCT with SNR over 60 dB when keeping only 20 taps. Laurent Girin |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A Simple Hybrid Acoustic / Morphologically-Constrained Technique for the Synthesis of Stop Consonants in Various Vocalic ContextsabstractThe predominant way to synthesize stop consonants is currently to use an articulatory model controlled by vocal tract parameters.We propose a new method to make this synthesis in various vocalic contexts.To generate the formant transitions, the basic principle is to apply an opening function on the (equal-length section) area function derived from the linear predictive (LP) model of speech signals.The definition of this opening function is empirically based on morphological considerations, and the main parameter is the place of articulation.Syllabic sounds with /b d g/ in /a i u/ vowel contexts are generated using LP synthesis with reflections coefficients corresponding to the interpolated area function.We show that the general structure of the formant transitions can be well represented using this model, and provide intelligible sound examples. Frédéric Berthommier, Laurent Girin, Louis-Jean Boë |
INTERSPEECH | 2 |
| 2012 | Informed source separation through spectrogram coding and data embedding
Antoine Liutkus, Jonathan Pinel, Roland Badeau, Laurent Girin, Gaël Richard |
Signal Process. | 4 |
| 2011 | A Long-Term Harmonic Plus Noise Model for Speech SignalsabstractThe harmonic plus noise model (HNM) is widely used for spectral modeling of mixed harmonic/noise speech sounds. In this paper, we present an analysis/synthesis system based on a long-term two-band HNM. “Long-term ” means that the time-trajectories of the HNM parameters are modeled using “smooth ” (discrete cosine) functions depending on a small set of parameters. The goal is to capture and exploit the longterm correlation of spectral components on time segments of up to several hundreds of ms. The proposed long-term HNM enables joint compact representation of signals (thus a potential for low bit-rate coding) and easy signal transformation (e.g. time stretching) directly from the long-term parameters. Experiments show that it can be compared favourably with the shortterm version in terms of parameter rates and signal quality. Index Terms: speech analysis/synthesis, harmonic + noise model, long-term processing. Faten Ben Ali, Laurent Girin, Sonia Djaziri Larbi |
INTERSPEECH | 2 |
| 2011 | An Informed Source Separation System for Speech SignalsabstractIn two previous papers, we proposed an audio Informed Source Separation (ISS) system which can achieve the separation of I> 2 musical sources from linear instantaneous stationary stereo (2-channel) mixtures, based on audio signal’s natural sparsity, pre-mix source signals analysis, and side-information embedding (within the mix signal). In the present paper and for the first time, we apply this system to mixtures of (up to seven) simultaneous speech signals. Compared to the reference MPEG-4 Spatial Audio Object Coding system, our system provides much cleaner separated speech signals (consistently 10– 20 dB higher Signal to Interference Ratios), revealing strong potential for audio conference applications. Index Terms: underdetermined source separation, speech mixture, speech signals sparsity, signal compression 1. Laurent Girin |
INTERSPEECH | 2 |
| 2011 | Informed Source Separation of Linear Instantaneous Under-Determined Audio Mixtures by Source Index EmbeddingabstractIn this paper, we address the issue of underdetermined source separation ofInonstationary audio sources from aJ-channel linear instantaneous mixture (JI). This problem is addressed with a specific coder-decoder configuration. At the coder, source signals are assumed to be available before the mixing is processed. A time-frequency (TF) joint analysis of each source signal and mixture signal enables to select the subset of sources (amongI) leading to the best separation results in each TF region. A corresponding source(s) index code is imperceptibly embedded into the mix signal using a watermarking technique. At the decoder, where the original source signals are unknown, the extraction of the watermark enables to invert the mixture in each TF region to recover the source signals. With such an informed approach, it is shown that five instruments and singing voice signals can be efficiently separated from two-channel stereo mixtures, with a quality that significantly overcomes the quality obtained by a semi-blind reference method and enables separate manipulation of the source signals during stereo music restitution (i.e., remixing). Mathieu Parvaix, Laurent Girin |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Informed source separation of underdetermined instantaneous stereo mixtures using source index embeddingabstractIn this paper, we address the issue of underdetermined source separation of non-stationary audio sources from a stereo (i.e. 2-channel) linear instantaneous mixture. This problem is addressed with a specific coder-decoder configuration. At the coder, source signals are assumed to be available before the mixing is processed. A time-frequency (TF) analysis of each source enables to select the one or two predominant sources (among I>2) in each TF region, and a corresponding source(s) index code is imperceptibly embedded into the mix signals using a watermarking technique. At the decoder level, where the original sources signals are unknown, the extraction of the watermark enables to locally reduce the underdetermined configuration to an (over)determined configuration. Sources signals can then be estimated using a classical (over)determined separation technique. Thereby several instruments or voice signals can be separated from stereo mixtures, enabling separate manipulation of the source signals during restitution (i.e. remastering). Mathieu Parvaix, Laurent Girin |
ICASSP | 2 |
| 2010 | A Watermarking-Based Method for Informed Source Separation of Audio Signals With a Single SensorabstractIn this paper, the issue of audio source separation from a single channel is addressed, i.e., the estimation of several source signals from a single observation of their mixture. This challenging problem is tackled with a specific two levels coder-decoder configuration. At the coder, source signals are assumed to be available before the mix is processed. Each source signal is characterized by a set of parameters that provide additional information useful for separation. We propose an original method using a watermarking technique to imperceptibly embed this information about the source signals into the mix signal. At the decoder, the watermark is extracted from the mix signal to enable an end-user who has no access to the original sources to separate these signals from their mixture. Hence, we call this separation process informed source separation (ISS). Thereby, several instruments or voice signals can be segregated from a single piece of music to enable post-mixing processing such as volume control, echo addition, spatialization, or timbre transformation. Good performances are obtained for the separation of up to four source signals, from mixtures of speech or music signals. Promising results open up new perspectives in both under-determined source separation and audio watermarking domains. Mathieu Parvaix, Laurent Girin, Jean-Marc Brossier |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | A watermarking-based method for single-channel audio source separationabstractIn this paper, we address the issue of audio source separation with a single channel, i.e. the estimation of source signals from a single mixture of these signals. This problem is addressed with a specific configuration: source signals are assumed to be available before the mix is processed. We propose an original method that uses a watermarking technique to embed information about the source signals into the mix signal. Extracting this watermark enables an end-user who has no access to the original sources to separate these signals from their mixture. Thereby several instruments or voice signals can be segregated from a single piece of music to enable post-mixing processing such as volume control. Mathieu Parvaix, Laurent Girin, Jean-Marc Brossier |
ICASSP | 2 |
| 2008 | Long-term flexible 2D cepstral modeling of speech spectral amplitudesabstractThis paper presents a method for modeling the envelope of spectral amplitude parameters of speech signals in "two dimensions" (2D). It consists of two cascaded modelings: the first one along the frequency axis is the usual cepstrum technique, which consists of modeling the log-scaled spectral envelope with a discrete cosine model (DCM). The second one, along the time axis, consists of modeling the trajectory of the envelope DCM coefficients by another similar DCM model. An iterative algorithm is proposed to optimally fit this 2D-model to the data according to a perceptual criterion based on frequency masking. This approach is shown to provide an efficient and flexible representation of spectral amplitude parameters in terms of coefficient rates, while providing good signal quality, opening new perspectives in very-low bit-rate sinusoidal speech coding. Mohammad Firouzmand, Laurent Girin |
ICASSP | 2 |
| 2008 | Estimation of the voicing cut-off frequency contour of natural speech based on harmonic and aperiodic energiesabstractWe present a new algorithm for the automatic estimation of the voicing cut-off frequency (VCO), i.e., the frequency that separates the periodic low-frequency part from the aperiodic high-frequency part in voiced segments of natural speech. Starting from the power spectrum of a two pitch period speech frame, we define the VCO to be located at the frequency for which the sum of the periodic and aperiodic energy in the spectral band below and above that frequency respectively, is maximised. By formulating the problem in terms of a score function we are able to apply a dynamic programming based smoothing technique. Remarkably smooth and accurate VCO contours were obtained, despite the simplicity of the proposed algorithm. In a formal evaluation the algorithm compares favourably to two existing VCO estimation techniques. Kris Hermus, Laurent Girin, Hugo Van hamme, Sufian Irhimeh |
ICASSP | 2 |
| 2007 | Long-Term Quantization of Speech LSF ParametersabstractThis paper addresses the problem of coding the LSF parameters of LPC speech coders on a "long-term" basis, i.e. beyond the usual #20 ms frame duration. The objective is to provide efficient LSF quantization for a speech coder with very large delay but very- to ultra-low bit-rate and good quality. To do this, a long-term model of the time-trajectory of the LSF vectors is applied on long segments of speech to capture the inter-frame correlation of the vectors over each whole segment. Using this model, it is shown that only a reduced set of LSF vectors need to be quantized to derive quantized LSF vectors at every original location. Experiments show that large gains in bit-rate over usual frame-by-frame quantization can be achieved (up to more than 50%) while preserving signal quality. Laurent Girin |
ICASSP (4) | 1 |
| 2007 | Visual voice activity detection as a help for speech source separation from convolutive mixtures
Bertrand Rivet, Laurent Girin, Christian Jutten |
Speech Commun. | 2 |
| 2007 | Perceptual Long-Term Variable-Rate Sinusoidal Modeling of SpeechabstractIn this paper, the problem of modeling the time-trajectory of the sinusoidal components of voiced speech signals is addressed. A new global approach is presented: a single so-called long-term (LT) model, based on discrete cosine functions, is used to model the overall trajectories of amplitude and phase parameters, for each entire voiced section of speech, differing from usual (short-term) models defined on a frame-by-frame basis. The complete analysis-modeling-synthesis process is presented, including an iterative algorithm for optimal fitting between LT model and measures. A major issue of this paper concerns the use of perceptual criteria in the LT model fitting process (both for amplitude and phase modeling). The adaptation of perceptual criteria usually defined in the short-term and/or stationary cases to the long-term processing is proposed. Experiments dealing with the ten first harmonics of voiced signals show that the proposed approach provides an efficient variable-rate representation of voiced speech signals. Promising results are given in terms of modeling accuracy, synthesis quality, and data compression. The interest of the presented approach for speech coding and speech watermarking is discussed Laurent Girin, Mohammad Firouzmand, Sylvain Marchand |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Mixing Audiovisual Speech Processing and Blind Source Separation for the Extraction of Speech Signals From Convolutive MixturesabstractLooking at the speaker's face can be useful to better hear a speech signal in noisy environment and extract it from competing sources before identification. This suggests that the visual signals of speech (movements of visible articulators) could be used in speech enhancement or extraction systems. In this paper, we present a novel algorithm plugging audiovisual coherence of speech signals, estimated by statistical tools, on audio blind source separation (BSS) techniques. This algorithm is applied to the difficult and realistic case of convolutive mixtures. The algorithm mainly works in the frequency (transform) domain, where the convolutive mixture becomes an additive mixture for each frequency channel. Frequency by frequency separation is made by an audio BSS algorithm. The audio and visual informations are modeled by a newly proposed statistical model. This model is then used to solve the standard source permutation and scale factor ambiguities encountered for each frequency after the audio blind separation stage. The proposed method is shown to be efficient in the case of 2 times 2 convolutive mixtures and offers promising perspectives for extracting a particular speech source of interest from complex mixtures Bertrand Rivet, Laurent Girin, Christian Jutten |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Log-Rayleigh Distribution: A Simple and Efficient Statistical Representation of Log-Spectral CoefficientsabstractIn this paper, we study the distribution of the log-modulus of a Gaussian complex random variable. In the circular case, it is a Log-Rayleigh (LR) variable, whose probability distribution function (pdf) depends on only one parameter. In the noncircular case, the pdf is more complicated, although we show that it can be adequately modeled by an LR pdf, for which the optimal fitting parameter is derived. These results can be used in any application using the log-modulus of discrete Fourier transform coefficients, e.g., for speech/audio signals, and suggest that a mixture of LR pdf kernels is preferable to more classical models such as mixtures of Gaussian kernels, which are more costly and less efficient Bertrand Rivet, Laurent Girin, Christian Jutten |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | An Analysis of Visual Speech Information Applied to Voice Activity DetectionabstractWe present a new approach to the voice activity detection (VAD) problem for speech signals embedded in non-stationary noise. The method is based on automatic lipreading: the objective is to detect voice activity or non-activity by exploiting the coherence between the speech acoustic signal and the speaker's lip movements. From a comprehensive analysis of lip shape parameters during speech and non-speech events, we show that a single appropriate visual parameter, defined to characterize the lip movements, can be used for the detection of sections of voice activity or more precisely, for the detection of silence sections. Detection scores obtained on spontaneous speech confirm the efficiency of the visual voice activity detector (VVAD) David Sodoyer, Bertrand Rivet, Laurent Girin, Jean-Luc Schwartz, Christian Jutten |
ICASSP (1) | 3 |
| 2005 | Perceptually Weighted Long Term Modeling of Sinusoidal Speech Amplitude TrajectoriesabstractIn this paper, the problem of modeling the trajectory of the amplitudes of speech signals is addressed within the context of the sinusoidal model of speech. A long-term model of the trajectory of the amplitude of the partials is proposed for each entire voiced section of speech, contrary to standard models, which are defined on a frame-by-frame basis. The complete analysis-modeling-synthesis process is presented. We compare a DCT-based long-term model with classical (frame-by-frame) interpolation schemes, given that the analysis process is identical in both cases. Perceptual constraints are taken into account since the distortion criterion in this approach is the level of modeling noise above the masking threshold. Promising results are given and the interest of the presented models for speech coding and watermarking applications is discussed. Mohammad Firouzmand, Laurent Girin |
ICASSP (1) | 2 |
| 2005 | Solving the indeterminations of blind source separation of convolutive speech mixturesabstractLooking at the speaker's face seems useful for hearing a speech signal better and extracting it from competing sources before identification. We present a novel algorithm plugging the audiovisual coherence of speech signals, estimated by statistical tools, on audio blind source separation (BSS) algorithms in the difficult case of convolutive mixtures. The algorithm mainly works in the frequency (transform) domain, where the convolutive mixture becomes an additive mixture for each frequency channel. Frequency by frequency separation is made by an audio BSS algorithm, and the audiovisual information is used to solve the standard source permutation and scale factor problems at the output of the separation stage, for each frequency. The proposed method is shown to be efficient in the case of 2/spl times/2 convolutive mixtures. Bertrand Rivet, Laurent Girin, Christian Jutten |
ICASSP (5) | 2 |
| 2005 | Comparing several models for perceptual long-term modeling of amplitude and phase trajectories of sinusoidal speech
Mohammad Firouzmand, Laurent Girin, Sylvain Marchand |
INTERSPEECH | 2 |
| 2004 | Watermarking of speech signals using the sinusoidal model and frequency modulation of the partialsabstractIn this paper, the application of the sinusoidal model for audio/speech signals to the watermarking task is proposed. The basic idea is that adequate modulation of medium rank partials (frequency) trajectories is not perceptible and thus this modulation may contain the data to be embedded in the signal. The modulation (encoding) and estimation (decoding) of the message are described and preliminary promising results are given in the case of speech signals. Laurent Girin, Sylvain Marchand |
ICASSP (1) | 1 |
| 2004 | Characterizing and classifying cued speech vowels from labial parametersabstractAs part of the THIMP project (Telephony for Hearing- IMpaired People), we aim at automatically analyzing Cued Speech [1] and translating it into oral spoken language. This work focuses on vowel classification and will be part of this transcoding process as a preprocessing step of the input data analysis. Its objective is to identify vowels produced by a speaker pronouncing and coding in Cued Speech a set of French sentences, knowing: - The Cued Speech Hand Placement, - The analysis of defined Labial Parameters. Here, we will show that the crossing of these two sources of information allows to automatically identify vowels. These results have to be compared to performances of hearingimpaired people in perception of Cued Speech. Denis Beautemps, Thomas Burger, Laurent Girin |
INTERSPEECH | 3 |
| 2004 | Long term modeling of phase trajectories within the speech sinusoidal model frameworkabstractAbstract In this paper, the problem of modeling the trajectory of the phase of speech signal is addressed within the context of the sinusoidal model of speech. A global or long-term model of the trajectory of the phase of the partials is proposed for each entire voiced section of speech, contrary to standard models, which are defined on a frame-by-frame basis. The complete analysis-modeling-synthesis process is presented. We compare two basic long-term models, namely a polynomial and a DCT-based model, with classical (frame-by-frame) interpolation schemes, given that the analysis process is the same in all cases. Promising results are given and the interest of the presented models for speech coding and speech watermarking applications is discussed. 1. Introduction Sinusoidal modeling of audio signals has been extensively studied since the eighties and successfully applied to a wide range of applications, such as coding or time- and frequency-stretching [1-5]. The signal is modeled as the sum of a small number Laurent Girin, Mohammad Firouzmand, Sylvain Marchand |
INTERSPEECH | 1 |
| 2004 | Using audiovisual speech processing to improve the robustness of the separation of convolutive speech mixturesabstractLooking at the speaker's face seems useful in hearing better a speech signal and extract it from the competing sources before identification. In this paper, we present a novel algorithm plugging audiovisual coherence of speech signals, estimated by statistical tools, on audio blind source separation (BSS) algorithms in the difficult case of convolutive mixtures. The algorithm mainly works in the frequency (transform) domain, where the convolutive mixture becomes an additive mixture for each frequency channel. Frequency by frequency separation is made by an audio BSS algorithm, and the audiovisual information is used to solve the standard source permutation problem at the output of the separation stage, for each frequency. The proposed method is shown to be efficient in the case of 2 /spl times/ 2 convolutive mixtures. Bertrand Rivet, Laurent Girin, Christian Jutten, Jean-Luc Schwartz |
MMSP | 2 |
| 2004 | Developing an audio-visual speech source separation algorithm
David Sodoyer, Laurent Girin, Christian Jutten, Jean-Luc Schwartz |
Speech Commun. | 2 |
| 2004 | Joint matrix quantization of face parameters and LPC coefficients for low bit rate audiovisual speech codingabstractA key problem for videophony, that is telephony including the processing of images of the speaker's face in addition to acoustic speech, concerns signal compression for transmission. In such systems, audio and video compression are separately achieved by using both audio and video coders. In this paper, an audio-visual approach to this problem is considered, since we claim that the fundamental property of coherence (redundancy) between the two modalities of speech should be exploited by coding systems. We consider the framework of parametric analysis, modeling and synthesis of talking faces, which allows efficient representation of video information. Thus, we propose to jointly encode several face parameters, namely lip shape geometric descriptors, together with sets of audio coefficients, namely quite usual LPC parameters. The definition of an audiovisual distance between vectors of concatenated audio and video parameters allows to generate audiovisual single stage vector and matrix quantizers by using the generalized Lloyd algorithm. Calculation of video and audio mean distortion measures shows a significant gain in quantization accuracy and/or resolution compared to separate video and audio quantization. An alternative sub-optimal tree-like structure for audiovisual joint coding is also tested and yields interesting results while decreasing the computational complexity of the quantization process. Laurent Girin |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Extracting an AV speech source from a mixture of signals
David Sodoyer, Laurent Girin, Christian Jutten, Jean-Luc Schwartz |
INTERSPEECH | 2 |
| 2002 | Audio-visual speech sources separation: a new approach exploiting the audio-visual coherence of speech stimuli
David Sodoyer, Laurent Girin, Christian Jutten, Jean-Luc Schwartz |
INTERSPEECH | 2 |
| 2001 | Speech signals separation: a new approach exploiting the coherence of audio and visual speechabstractWe present a new approach to the source separation problem in the case of multiple speech signals. The method is based on the use of automatic lip reading: the objective is to extract an acoustic speech signal from other acoustic signals by exploiting its coherence with the speaker's lip movements. For this aim, a statistical model is used to quantify this coherence. The results, while very preliminary, are encouraging. They show that this method can achieve a good separation of a speech source in the case of simple 2/spl times/2 additive mixtures. Moreover, it presents some interesting complementarity with traditional pure audio techniques. Laurent Girin, A. Allard, Jean-Luc Schwartz |
MMSP | 1 |
| 1998 | Fusion of auditory and visual information for noisy speech enhancement: a preliminary study of vowel transitionsabstractThis paper deals with a noisy speech enhancement technique based on the fusion of auditory and visual information. We first present the global structure of the system, and then we focus on the tool we used to melt both sources of information. The whole noise reduction system is implemented in the context of vowel transitions corrupted with white noise. A complete evaluation of the system in this context is presented, including distance measures, Gaussian classification scores, and a perceptive test. The results are very promising. Laurent Girin, Gang Feng 0002, Jean-Luc Schwartz |
ICASSP | 1 |
| 1998 | A signal processing system for having the sound "pop-out" in noise thanks to the image of the speaker's lips: new advances using multi-layer perceptrons
Laurent Girin, Laurent Varin, Gang Feng 0002, Jean-Luc Schwartz |
ICSLP | 1 |
| 1998 | An audio-visual distance for audio-visual speech vector quantizationabstractSpeech is both an acoustic and a visual signal, and there exists some complementarity and redundancy between the two modalities. In the speech coding domain, it is of great interest to use this redundancy to improve speech coder performance. In this paper, we consider some audio and video joint coding process based on an audio-visual vector quantization. The method is shown to exploit quite well the audio-visual redundancy as it can reduce the bit rate while decreasing the quantization error. A notion of audio-visual distance has to be introduced and adapted to the different nature of the data. It is defined from an existing audio distance and a new visual distance, which is particularly focussed. Laurent Girin, Elodie Foucher, Gang Feng 0002 |
MMSP | 1 |
| 1998 | Audiovisual speech enhancement: new advances using multi-layer perceptronsabstractThis paper deals with the improvement of a noisy speech enhancement system based on the fusion of auditory and visual information. The system was presented in previous papers and implemented with a simple stimuli corrupted with white noise. Its principle consists of an analysis-enhancement-synthesis process based on a linear prediction (LP) model of the signal: the LP filter is enhanced thanks to associative tools that estimate the LP cleaned parameters from both noisy audio and lip shape information. The structure of the system is reviewed and we focus on the improvement that concerns the associators: multi-layers perceptrons are used instead of linear regression. It is shown that in the context of VCV transitions corrupted with white noise, the performances of the system are improved in terms of the intelligibility gain, distance measures and classification tests. Laurent Girin, Laurent Varin, Gang Feng 0002, Jean-Luc Schwartz |
MMSP | 1 |
| 1997 | Noisy speech enhancement by fusion of auditory and visual information: a study of vowel transitionsabstractThis paper deals with a noisy speech enhancement technique based on the fusion of auditory and visual information. We first present the global structure of the system, and then we focus on the tool we used to melt both sources of information. The whole noise reduction system is implemented in the context of vowel transitions corrupted with white noise. A complete evaluation of the system in this context is presented, including distance measures, gaussian classification scores, and a perceptive test. The results are very promising. Laurent Girin, Gang Feng 0002, Jean-Luc Schwartz |
EUROSPEECH | 1 |
| 1995 | Noisy speech enhancement with filters estimated from the speaker's lips
Laurent Girin, Gang Feng 0002, Jean-Luc Schwartz |
EUROSPEECH | 1 |