Antoine Deleforge

dblp:47/10875 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-0339-7472ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 100%
Artificial intelligence
3 papers
Generative modeling · 57% Robot navigation and mapping · 19% Face, body and person analysis · 12%
Computer networks
1 paper
Physical-layer communications · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
audio generation
0.712023
From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023
Machine learning › Generative modeling
diffusion model
0.712023
From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023
Audio and music processing
sound synthesis
0.712023
From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023
Audio and music processing › speech coding
vocoder
0.712023
From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023
Audio and music processing
acoustic signal processing
0.312018
MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval · NeurIPS 2018
Audio and music processing
beamforming
0.312018
Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Audio and music processing › speech enhancement
multichannel wiener filter
0.312018
Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Audio and music processing
speech enhancement
0.312018
Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Physical-layer communications › channel estimation
blind channel estimation
0.312018
MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval · NeurIPS 2018
Physical-layer communications
channel estimation
0.312018
MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval · NeurIPS 2018
Computer vision › Face, body and person analysis
head pose estimation
0.312017
Robust Head-Pose Estimation Based on Partially-Latent Mixture of Linear Regressions · IEEE Trans. Image Process. 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture of linear regressions
0.312017
Robust Head-Pose Estimation Based on Partially-Latent Mixture of Linear Regressions · IEEE Trans. Image Process. 2017
Robotics › Robot navigation and mapping › sound source localization
binaural localization
0.212015
Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Robotics › Robot navigation and mapping
sound source localization
0.212015
Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression · IEEE ACM Trans. Audio Speech Lang. Process. 2015

Methods — techniques the papers use, named apart from their topics

multi-band diffusion · 1.3parameter-space estimation · 0.7finite-rate-of-innovation sampling · 0.7compressed sensing · 0.7sample covariance matrix modeling · 0.3bivariate normal distribution · 0.3partially-latent output · 0.3mixture of regressions · 0.3manifold learning · 0.3locally linear gaussian regression · 0.2binaural features · 0.2generative probabilistic model · 0.1constrained mixture model · 0.1EM algorithm · 0.1
YearPublicationVenuePosition
2025 Latent Watermarking of Audio Generative Models
abstract
The advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. To address this, watermarking techniques have been recently developed, enabling the detection of content generated by a deployed model. For such techniques to be useful, the watermark must resist typical modifications applied to the model or its outputs. The use case of an open-source model trained on proprietary data is challenging, as post-hoc watermarks can then be trivially removed. In response, we introduce a method that watermarks latent audio generative models by directly watermarking their training data. We show the method to be robust against a broad range of audio edits including filtering, compression or even to changing the model’s decoder, maintaining high detection rates with very few false positives. Interestingly, we show that even fine-tuning the model on another dataset can only significantly lower the detection rate at the cost of degrading the generation performance near the level of re-training the model without the protected training data.
Robin San-Roman, Pierre Fernandez, Antoine Deleforge, Yossi Adi, Romain Serizel
ICASSP3
2023 How to (Virtually) Train Your Speaker Localizer
abstract
Learning-based methods have become ubiquitous in speaker localization. Existing systems rely on simulated training sets for the lack of sufficiently large, diverse and annotated real datasets. Most room acoustics simulators used for this purpose rely on the image source method (ISM) because of its computational efficiency. This paper argues that carefully extending the ISM to incorporate more realistic surface, source and microphone responses into training sets can significantly boost the real-world performance of speaker localization systems. It is shown that increasing the training-set realism of a state-of-the-art direction-of-arrival estimator yields consistent improvements across three different real test sets featuring human speakers in a variety of rooms and various microphone arrays. An ablation study further reveals that every added layer of realism contributes positively to these improvements.
Prerak Srivastava, Antoine Deleforge, Archontis Politis, Emmanuel Vincent 0001
INTERSPEECH2
2023 From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion
abstract
Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms conditioned on highly compressed representations. Although such methods produce impressive results, they are prone to generate audible artifacts when the conditioning is flawed or imperfect. An alternative modeling approach is to use diffusion models. However, these have mainly been used as speech vocoders (i.e., conditioned on mel-spectrograms) or generating relatively low sampling rate signals. In this work, we propose a high-fidelity multi-band diffusion-based framework that generates any type of audio modality (e.g., speech, music, environmental sounds) from low-bitrate discrete representations. At equal bit rate, the proposed approach outperforms state-of-the-art generative techniques in terms of perceptual quality. Training and evaluation code are available on the facebookresearch/ audiocraft github project. Samples are available on the following link (https://ai.honu.io/papers/mbd/).
Robin San-Roman, Yossi Adi, Antoine Deleforge, Romain Serizel, Gabriel Synnaeve, Alexandre Défossez
NeurIPS3
2022 Gridless 3D Recovery of Image Sources From Room Impulse Responses
abstract
Given a sound field generated by a sparse distribution of impulse image sources, can the continuous 3D positions and amplitudes of these sources be recovered from discrete, band-limited measurements of the field at a finite set of locations,e.g., a multichannel room impulse response? Borrowing from recent advances in super-resolution imaging, it is shown that this non-linear, non-convex inverse problem can be efficiently relaxed into a convex linear inverse problem over the space of Radon measures in$\mathbb {R}^{3}$. The new linear operator introduced here stems from the fundamental solution of the wave equation combined with the receivers' responses. An adaptation of the Sliding Frank-Wolfe algorithm is proposed to numerically solve the problemoff-the-grid,i.e., in continuous 3D space. Idealized simulated experiments show that the approach can recover hundreds of image sources at a rate and accuracy that are not achievable by previous methods, using a compact microphone array and source placed at random in random-sized shoe-box rooms. The impact of noise, sampling rate and array diameter on these results is also examined.
Tom Sprunck, Antoine Deleforge, Yannick Privat, Cédric Foy
IEEE Signal Process. Lett.2
2021 Detecting Acoustic Reflectors Using A Robot's Ego-Noise
abstract
In this paper, we propose a method to estimate the proximity of an acoustic reflector, e.g., a wall, using ego-noise, i.e., the noise produced by the moving parts of a listening robot. This is achieved by estimating the times of arrival of acoustic echoes reflected from the surface. Simulated experiments show that the proposed non-intrusive approach is capable of accurately estimating the distance of a reflector up to 1 meter and outperforms a previously proposed intrusive approach under loud ego-noise conditions. The proposed method is helped by a probabilistic echo detector that estimates whether or not an acoustic reflector is within a short range of the robotic platform. This preliminary investigation paves the way towards a new kind of collision avoidance system that would purely rely on audio sensors rather than conventional proximity sensors.
Usama Saqib, Antoine Deleforge, Jesper Rindom Jensen
ICASSP2
2020 Blaster: An Off-Grid Method for Blind and Regularized Acoustic Echoes Retrieval
abstract
Acoustic echoes retrieval is a research topic that is gaining importance in many speech and audio signal processing applications such as speech enhancement, source separation, dereverberation and room geometry estimation. This work proposes a novel approach to blindly retrieve the off-grid timing of early acoustic echoes from a stereophonic recording of an unknown sound source such as speech. It builds on the recent framework of continuous dictionaries. In contrast with existing methods, the proposed approach does not rely on parameter tuning nor peak picking techniques by working directly in the parameter space of interest. The accuracy and robustness of the method are assessed on challenging simulated setups with varying noise and reverberation levels and are compared to two state-of-the-art methods.
Diego Di Carlo, Clement Elvira, Antoine Deleforge, Nancy Bertin, Rémi Gribonval
ICASSP3
2020 Filterbank Design for End-to-end Speech Separation
abstract
Single-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet. In parallel, parameterized filterbanks have been proposed for speaker recognition where only center frequencies and bandwidths are learned. In this work, we extend real-valued learned and parameterized filterbanks into complex-valued analytic filterbanks and define a set of corresponding representations and masking strategies. We evaluate these filterbanks on a newly released noisy speech separation dataset (WHAM). The results show that the proposed analytic learned filterbank consistently outperforms the real-valued filterbank of ConvTasNet. Also, we validate the use of parameterized filterbanks and show that complex-valued representations and masks are beneficial in all conditions. Finally, we show that the STFT achieves its best performance for 2 ms windows.
Manuel Pariente, Samuele Cornell, Antoine Deleforge, Emmanuel Vincent 0001
ICASSP3
2020 Asteroid: The PyTorch-Based Audio Source Separation Toolkit for Researchers
abstract
This paper describes Asteroid, the PyTorch-based audio source separation toolkit for researchers. Inspired by the most successful neural source separation systems, it provides all neural building blocks required to build such a system. To improve reproducibility, Kaldi-style recipes on common audio source separation datasets are also provided. This paper describes the software architecture of Asteroid and its most important features. By showing experimental results obtained with Asteroid's recipes, we show that our implementations are at least on par with most results reported in reference papers. The toolkit is publicly available at https://github.com/mpariente/asteroid .
Manuel Pariente, Samuele Cornell, Joris Cosentino, Sunit Sivasankaran, Efthymios Tzinis, Jens Heitkaemper, Michel Olvera, Fabian-Robert Stöter, Mathieu Hu, Juan M. Martín-Doñas, David Ditter, Ariel Frank, Antoine Deleforge, Emmanuel Vincent 0001
INTERSPEECH13
2019 Mirage: 2D Source Localization Using Microphone Pair Augmentation with Echoes
abstract
It is commonly observed that acoustic echoes hurt per mance of sound source localization (SSL) methods. We troduce the concept of microphone array augmentation echoes (MIRAGE) and show how estimation of early-e characteristics can in fact benefit SSL. We propose a learn based scheme for echo estimation combined with a phys based scheme for echo aggregation. In a simple scenario volving 2 microphones close to a reflective surface and source, we show using simulated data that the proposed proach performs similarly to a correlation-based metho azimuth estimation while retrieving elevation as well from 2 microphones only, an impossible task in anechoic settings.
Diego Di Carlo, Antoine Deleforge, Nancy Bertin
ICASSP2
2019 A Statistically Principled and Computationally Efficient Approach to Speech Enhancement Using Variational Autoencoders
abstract
Recent studies have explored the use of deep generative models of speech spectra based of variational autoencoders (VAEs), combined with unsupervised noise models, to perform speech enhancement. These studies developed iterative algorithms involving either Gibbs sampling or gradient descent at each step, making them computationally expensive. This paper proposes a variational inference method to iteratively estimate the power spectrogram of the clean speech. Our main contribution is the analytical derivation of the variational steps in which the en-coder of the pre-learned VAE can be used to estimate the varia-tional approximation of the true posterior distribution, using the very same assumption made to train VAEs. Experiments show that the proposed method produces results on par with the afore-mentioned iterative methods using sampling, while decreasing the computational cost by a factor 36 to reach a given performance .
Manuel Pariente, Antoine Deleforge, Emmanuel Vincent 0001
INTERSPEECH2
2018 Blind Source Separation Using Mixtures of Alpha-Stable Distributions
abstract
We propose a new blind source separation algorithm based on mixtures of α-stable distributions. Complex symmetric α-stable distributions have been recently showed to better model audio signals in the time-frequency domain than classical Gaussian distributions thanks to their larger dynamic range. However, inference with these models is notoriously hard to perform because their probability density functions do not have a closed-form expression in general. Here, we introduce a novel method for estimating mixtures of α-stable distributions based on characteristic function matching. We apply this to the blind estimation of binary masks in individual frequency bands from multichannel convolutive audio mixtures. We show that the proposed method yields better separation performance than Gaussian-based binary-masking methods.
Nicolas Keriven, Antoine Deleforge, Antoine Liutkus
ICASSP2
2018 Audio Source Separation with Magnitude Priors: The Beads Model
abstract
Audio source separation comes with the need to devise multichannel filters that can exploit priors about the target signals. In that context, experience shows that modeling magnitude spectra is effective. However, devising a probabilistic model on complex spectral data with a prior on magnitudes is non trivial, because it should both reflect the prior but also be tractable for easy inference. In this paper, we approximate the ideal donut-shaped distribution of a complex variable with approximately known magnitude as a Gaussian mixture model called BEADS (Bayesian Expansion Approximating the Donut Shape) and show that it permits straightforward inference and filtering while effectively constraining the magnitudes of the signals to comply with the prior. As a result, we demonstrate large improvements over the Gaussian baseline for multichannel audio coding when exploiting the BEADS model.
Antoine Liutkus, Christian Rohlfing, Antoine Deleforge
ICASSP3
2018 Separake: Source Separation with a Little Help from Echoes
abstract
It is commonly believed that multipath hurts various audio processing algorithms. At odds with this belief, we show that multipath in fact helps sound source separation, even with very simple propagation models. Unlike most existing methods, we neither ignore the room impulse responses, nor we attempt to estimate them fully. We rather assume to know the positions of a few virtual microphones generated by echoes and we show how this gives us enough spatial diversity to get a performance boost over the anechoic case. We show improvements for two standard algorithms-one that uses only magnitudes of the transfer functions, and one that also uses the phases. Concretely, we show that multi-channel non-negative matrix factorization aided with a small number of echoes beats the vanilla variant of the same algorithm, and that with magnitude information only, echoes enable separation where it was previously impossible.
Robin Scheibler, Diego Di Carlo, Antoine Deleforge, Ivan Dokmanic
ICASSP3
2018 DREGON: Dataset and Methods for UAV-Embedded Sound Source Localization
abstract
This paper introduces DREGON, a novel publicly-available dataset that aims at pushing research in sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). The dataset contains both clean and noisy in-flight audio recordings continuously annotated with the 3D position of the target sound source using an accurate motion capture system. In addition, various signals of interests are available such as the rotational speed of individual rotors and inertial measurements at all time. Besides introducing the dataset, this paper sheds light on the specific properties, challenges and opportunities brought by the emerging task of UAV-embedded sound source localization. Several baseline methods are evaluated and compared on the dataset, with real-time applicability in mind. Very promising results are obtained for the localization of a broad-band source in loud noise conditions, while speech localization remains a challenge under extreme noise levels.
Martin Strauss 0003, Pol Mordel, Victor Miguet, Antoine Deleforge
IROS4
2018 MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval
abstract
This paper addresses the general problem of blind echo retrieval, i.e., given M sensors measuring in the discrete-time domain M mixtures of K delayed and attenuated copies of an unknown source signal, can the echo location and weights be recovered? This problem has broad applications in fields such as sonars, seismology, ultrasounds or room acoustics. It belongs to the broader class of blind channel identification problems, which have been intensively studied in signal processing. All existing methods proceed in two steps: (i) blind estimation of sparse discrete-time filters and (ii) echo information retrieval by peak picking. The precision of these methods is fundamentally limited by the rate at which the signals are sampled: estimated echo locations are necessary on-grid, and since true locations never match the sampling grid, the weight estimation precision is also strongly limited. This is the so-called basis-mismatch problem in compressed sensing. We propose a radically different approach to the problem, building on top of the framework of finite-rate-of-innovation sampling. The approach operates directly in the parameter-space of echo locations and weights, and enables near-exact blind and off-grid echo retrieval from discrete-time measurements. It is shown to outperform conventional methods by several orders of magnitudes in precision.
Helena Peic Tukuljac, Antoine Deleforge, Rémi Gribonval
NeurIPS2
2018 Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices
abstract
This paper studies the statistical performance of the multichannel Wiener filter (MWF) when the weights are computed using estimates of the sample covariance matrices of the noisy and the noise signals. It is well known that the optimal weights of the minimum variance distortionless response beamformer are only determined by the noisy sample covariance matrix or the noise sample covariance matrix, while those of the MWF are determined by both of them. Therefore, the difficulty increases dramatically in statistically analyzing the MWF when compared to analyzing the MVDR, where the main reason is that expressing the general joint probability density function (p.d.f.) of the two sample covariance matrices presented a Hitherto unsolved problem, to the best of our knowledge. For a deeper insight into the statistical performance of the MWF, this paper first introduces a bivariate normal distribution to approximately model the joint p.d.f. of the noisy and the noise sample covariance matrices. Each sample covariance matrix is approximately modeled by a random scalar multiplied by its true covariance matrix. This approximation is designed to preserve both the bias and the mean squared error of the matrix with respect to a natural distance on covariance matrices. The correlation of the bivariate normal distribution, referred to as the sample covariance matrices intrinsic correlation coefficient, captures all second-order dependencies of the noisy and the noise sample covariance matrices. By using the proposed bivariate normal distribution, the performance of the MWF can be predicted from the derived analytical expressions and many interesting results are revealed. As an example, the theoretical analysis demonstrates that the MWF performance may degrade in terms of noise reduction and signal-to-noise-ratio improvement when using more sensors in some noise scenarios.
Chengshi Zheng, Antoine Deleforge, Xiaodong Li 0002, Walter Kellermann
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Phase unmixing: Multichannel source separation with magnitude constraints
abstract
We consider the problem of estimating the phases of K mixed complex signals from a multichannel observation, when the mixing matrix and signal magnitudes are known. This problem can be cast as a non-convex quadratically constrained quadratic program which is known to be NP-hard in general. We propose three approaches to tackle it: a heuristic method, an alternate minimization method, and a convex relaxation into a semi-definite program. The last two approaches are showed to outperform the oracle multichannel Wiener filter in under-determined informed source separation tasks, using simulated and speech signals. The convex relaxation approach yields best results, including the potential for exact source separation in under-determined settings.
Antoine Deleforge, Yann Traonmilin
ICASSP1
2017 Phase retrieval with a multivariate Von Mises prior: From a Bayesian formulation to a lifting solution
abstract
In this paper, we investigate a new method for phase recovery when prior information on the missing phases is available. In particular, we propose to take into account this information in a generic fashion by means of a multivariate Von Mises distribution. Building on a Bayesian formulation (a Maximum A Posteriori estimation), we show that the problem can be expressed using a Mahalanobis distance and be solved by a lifting optimization procedure.
Angélique Dremeau, Antoine Deleforge
ICASSP2
2017 Hearing in a shoe-box: Binaural source position and wall absorption estimation using virtually supervised learning
abstract
This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes. These scenes are used to build a dataset of spatial binaural features annotated with acoustic properties such as the 3D source position and the walls' absorption coefficients. A probabilistic high- to low-dimensional regression framework is used to learn a mapping from these features to the acoustic properties. Results indicate that this mapping successfully estimates the azimuth and elevation of new sources, but also their range and even the walls' absorption coefficients solely based on binaural signals. Results also reveal that incorporating random-diffusion effects in the data significantly improves the estimation of all parameters.
Saurabh Kataria 0001, Clément Gaultier, Antoine Deleforge
ICASSP3
2017 Robust Head-Pose Estimation Based on Partially-Latent Mixture of Linear Regressions
abstract
Head-pose estimation has many applications, such as social event analysis, human-robot and human-computer interaction, driving assistance, and so forth. Head-pose estimation is challenging, because it must cope with changing illumination conditions, variabilities in face orientation and in appearance, partial occlusions of facial landmarks, as well as bounding-box-to-face alignment errors. We propose to use a mixture of linear regressions with partially-latent output. This regression method learns to map high-dimensional feature vectors (extracted from bounding boxes of faces) onto the joint space of head-pose angles and bounding-box shifts, such that they are robustly predicted in the presence of unobservable phenomena. We describe in detail the mapping method that combines the merits of unsupervised manifold learning techniques and of mixtures of regressions. We validate our method with three publicly available data sets and we thoroughly benchmark four variants of the proposed algorithm with several state-of-the-art head-pose estimation methods.
Vincent Drouard, Radu Horaud, Antoine Deleforge, Sileye O. Ba, Georgios Evangelidis 0002
IEEE Trans. Image Process.3
2016 Ego-noise reduction using a motor data-guided multichannel dictionary
abstract
We address the problem of ego-noise reduction, i.e., suppressing the noise a robot causes by its own motions. Such noise degrades the recorded microphone signal massively such that the robot's auditory capabilities suffer. To suppress it, it is intuitive to use also motor data, since it provides additional information about the robot's joints and thereby the noise sources. We propose to fuse motor data to a recently proposed multichannel dictionary algorithm for ego-noise reduction. At training, a dictionary is learned that captures spatial and spectral characteristics of ego-noise. At testing, nonlinear classifiers are used to efficiently associate the current robot's motor state to relevant sets of entries in the learned dictionary. By this, computational load is reduced by one third in typical scenarios while achieving at least the same noise reduction performance. Moreover, we propose to train dictionaries on different microphone array geometries and use them for ego-noise reduction while the head to which the microphones are mounted is moving. In such scenarios, the motor guided approach results in significantly better performance values.
Alexander Schmidt 0004, Antoine Deleforge, Walter Kellermann
IROS2
2015 Phase-optimized K-SVD for signal extraction from underdetermined multichannel sparse mixtures
abstract
We propose a novel sparse representation for heavily underdetermined multichannel sound mixtures, i.e., with much more sources than microphones. The proposed approach operates in the complex Fourier domain, thus preserving spatial characteristics carried by phase differences. We derive a generalization of K-SVD which jointly estimates a dictionary capturing both spectral and spatial features, a sparse activation matrix, and all instantaneous source phases from a set of signal examples. This dictionary can be used to extract the learned signal from a new input mixture. The method is applied to the challenging problem of ego-noise reduction for robot audition. We demonstrate its superiority relative to conventional dictionary-based techniques using real-room recordings.
Antoine Deleforge, Walter Kellermann
ICASSP1
2015 Head pose estimation via probabilistic high-dimensional regression
abstract
This paper addresses the problem of head pose estimation with three degrees of freedom (pitch, yaw, roll) from a single image. Pose estimation is formulated as a high-dimensional to low-dimensional mixture of linear regression problem. We propose a method that maps HOG-based descriptors, extracted from face bounding boxes, to corresponding head poses. To account for errors in the observed bounding-box position, we learn regression parameters such that a HOG descriptor is mapped onto the union of a head pose and an offset, such that the latter optimally shifts the bounding box towards the actual position of the face in the image. The performance of the proposed method is assessed on publicly available datasets. The experiments that we carried out show that a relatively small number of locally-linear regression functions is sufficient to deal with the non-linear mapping problem at hand. Comparisons with state-of-the-art methods show that our method outperforms several other techniques.
Vincent Drouard, Sileye O. Ba, Georgios Evangelidis 0002, Antoine Deleforge, Radu Horaud
ICIP4
2015 Acoustic Space Learning for Sound-Source Separation and Localization on Binaural Manifolds
abstract
In this paper, we address the problems of modeling the acoustic space generated by a full-spectrum sound source and using the learned model for the localization and separation of multiple sources that simultaneously emit sparse-spectrum sounds. We lay theoretical and methodological grounds in order to introduce the binaural manifold paradigm. We perform an in-depth study of the latent low-dimensional structure of the high-dimensional interaural spectral data, based on a corpus recorded with a human-like audiomotor robot head. A nonlinear dimensionality reduction technique is used to show that these data lie on a two-dimensional (2D) smooth manifold parameterized by the motor states of the listener, or equivalently, the sound-source directions. We propose a probabilistic piecewise affine mapping model (PPAM) specifically designed to deal with high-dimensional data exhibiting an intrinsic piecewise linear structure. We derive a closed-form expectation-maximization (EM) procedure for estimating the model parameters, followed by Bayes inversion for obtaining the full posterior density function of a sound-source direction. We extend this solution to deal with missing data and redundancy in real-world spectrograms, and hence for 2D localization of natural sound sources such as speech. We further generalize the model to the challenging case of multiple sound sources and we propose a variational EM framework. The associated algorithm, referred to as variational EM for source separation and localization (VESSL) yields a Bayesian estimation of the 2D locations and time-frequency masks of all the sources. Comparisons of the proposed approach with several existing methods reveal that the combination of acoustic-space learning with Bayesian inference enables our method to outperform state-of-the-art methods.
Antoine Deleforge, Florence Forbes, Radu Horaud
Int. J. Neural Syst.1
2015 Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression
abstract
This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on monaural segregation. The method starts with a training stage that establishes a locally linear Gaussian regression model between the directional coordinates of all the sources and the auditory features extracted from binaural measurements. While fixed-length wide-spectrum sounds (white noise) are used for training to reliably estimate the model parameters, we show that the testing (localization) can be extended to variable-length sparse-spectrum sounds (such as speech), thus enabling a wide range of realistic applications. Indeed, we demonstrate that the method can be used for audio-visual fusion, namely to map speech signals onto images and hence to spatially align the audio and visual modalities, thus enabling to discriminate between speaking and non-speaking faces. We release a novel corpus of real-room recordings that allow quantitative evaluation of the co-localization method in the presence of one or two sound sources. Experiments demonstrate increased accuracy and speed relative to several state-of-the-art methods.
Antoine Deleforge, Radu Horaud, Yoav Y. Schechner, Laurent Girin
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Variational EM for binaural sound-source separation and localization
abstract
The sound-source separation and localization (SSL) problems are addressed within a unified formulation. Firstly, a mapping between white-noise source locations and binaural cues is estimated. Secondly, SSL is solved via Bayesian inversion of this mapping in the presence of multiple sparse-spectrum emitters (such as speech), noise and reverberations. We propose a variational EM algorithm which is described in detail together with initialization and convergence issues. Extensive real-data experiments show that the method outperforms the state-of-the-art both in separation and localization (azimuth and elevation).
Antoine Deleforge, Florence Forbes, Radu Horaud
ICASSP1
2012 The cocktail party robot: sound source separation and localisation with an active binaural head
abstract
Human-robot communication is often faced with the difficult problem of interpreting ambiguous auditory data. For example, the acoustic signals perceived by a humanoid with its on-board microphones contain a mix of sounds such as speech, music, electronic devices, all in the presence of attenuation and reverberations. In this paper we propose a novel method, based on a generative probabilistic model and on active binaural hearing, allowing a robot to robustly perform sound-source separation and localization. We show how interaural spectral cues can be used within a constrained mixture model specifically designed to capture the richness of the data gathered with two microphones mounted onto a human-like artificial head. We describe in detail a novel EM algorithm, we analyse its initialization, speed of convergence and complexity, and we assess its performance with both simulated and real data.
Antoine Deleforge, Radu Horaud
HRI1