EDBT 2026 Demo / reviewers in the wild / expert
Antoine Deleforge
dblp:47/10875
· DBLP profile ↗
27ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-0339-7472ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 100% | |
| Artificial intelligence
3 papers |
Generative modeling · 57% Robot navigation and mapping · 19% Face, body and person analysis · 12% | |
| Computer networks
1 paper |
Physical-layer communications · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
audio generation |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Audio and music processing
sound synthesis |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Audio and music processing › speech coding
vocoder |
0.7 | 1 | 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion · NeurIPS 2023 |
Audio and music processing
acoustic signal processing |
0.3 | 1 | 2018 | MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval · NeurIPS 2018 |
Audio and music processing
beamforming |
0.3 | 1 | 2018 | Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Audio and music processing › speech enhancement
multichannel wiener filter |
0.3 | 1 | 2018 | Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Audio and music processing
speech enhancement |
0.3 | 1 | 2018 | Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance Matrices · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Physical-layer communications › channel estimation
blind channel estimation |
0.3 | 1 | 2018 | MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval · NeurIPS 2018 |
Physical-layer communications
channel estimation |
0.3 | 1 | 2018 | MULAN: A Blind and Off-Grid Method for Multichannel Echo Retrieval · NeurIPS 2018 |
Computer vision › Face, body and person analysis
head pose estimation |
0.3 | 1 | 2017 | Robust Head-Pose Estimation Based on Partially-Latent Mixture of Linear Regressions · IEEE Trans. Image Process. 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture of linear regressions |
0.3 | 1 | 2017 | Robust Head-Pose Estimation Based on Partially-Latent Mixture of Linear Regressions · IEEE Trans. Image Process. 2017 |
Robotics › Robot navigation and mapping › sound source localization
binaural localization |
0.2 | 1 | 2015 | Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Robotics › Robot navigation and mapping
sound source localization |
0.2 | 1 | 2015 | Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Methods — techniques the papers use, named apart from their topics
multi-band diffusion · 1.3parameter-space estimation · 0.7finite-rate-of-innovation sampling · 0.7compressed sensing · 0.7sample covariance matrix modeling · 0.3bivariate normal distribution · 0.3partially-latent output · 0.3mixture of regressions · 0.3manifold learning · 0.3locally linear gaussian regression · 0.2binaural features · 0.2generative probabilistic model · 0.1constrained mixture model · 0.1EM algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Watermarking of Audio Generative ModelsabstractThe advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. To address this, watermarking techniques have been recently developed, enabling the detection of content generated by a deployed model. For such techniques to be useful, the watermark must resist typical modifications applied to the model or its outputs. The use case of an open-source model trained on proprietary data is challenging, as post-hoc watermarks can then be trivially removed. In response, we introduce a method that watermarks latent audio generative models by directly watermarking their training data. We show the method to be robust against a broad range of audio edits including filtering, compression or even to changing the model’s decoder, maintaining high detection rates with very few false positives. Interestingly, we show that even fine-tuning the model on another dataset can only significantly lower the detection rate at the cost of degrading the generation performance near the level of re-training the model without the protected training data. Robin San-Roman, Pierre Fernandez, Antoine Deleforge, Yossi Adi, Romain Serizel |
ICASSP | 3 |
| 2023 | How to (Virtually) Train Your Speaker LocalizerabstractLearning-based methods have become ubiquitous in speaker localization. Existing systems rely on simulated training sets for the lack of sufficiently large, diverse and annotated real datasets. Most room acoustics simulators used for this purpose rely on the image source method (ISM) because of its computational efficiency. This paper argues that carefully extending the ISM to incorporate more realistic surface, source and microphone responses into training sets can significantly boost the real-world performance of speaker localization systems. It is shown that increasing the training-set realism of a state-of-the-art direction-of-arrival estimator yields consistent improvements across three different real test sets featuring human speakers in a variety of rooms and various microphone arrays. An ablation study further reveals that every added layer of realism contributes positively to these improvements. Prerak Srivastava, Antoine Deleforge, Archontis Politis, Emmanuel Vincent 0001 |
INTERSPEECH | 2 |
| 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionabstractDeep generative models can generate high-fidelity audio conditioned on various
types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients
(MFCC)). Recently, such models have been used to synthesize audio
waveforms conditioned on highly compressed representations. Although such
methods produce impressive results, they are prone to generate audible artifacts
when the conditioning is flawed or imperfect. An alternative modeling approach is
to use diffusion models. However, these have mainly been used as speech vocoders
(i.e., conditioned on mel-spectrograms) or generating relatively low sampling
rate signals. In this work, we propose a high-fidelity multi-band diffusion-based
framework that generates any type of audio modality (e.g., speech, music, environmental
sounds) from low-bitrate discrete representations. At equal bit rate,
the proposed approach outperforms state-of-the-art generative techniques in terms
of perceptual quality. Training and evaluation code are available on the facebookresearch/
audiocraft github project. Samples are available on the following
link (https://ai.honu.io/papers/mbd/). Robin San-Roman, Yossi Adi, Antoine Deleforge, Romain Serizel, Gabriel Synnaeve, Alexandre Défossez |
NeurIPS | 3 |
| 2022 | Gridless 3D Recovery of Image Sources From Room Impulse ResponsesabstractGiven a sound field generated by a sparse distribution of impulse image sources, can the continuous 3D positions and amplitudes of these sources be recovered from discrete, band-limited measurements of the field at a finite set of locations,e.g., a multichannel room impulse response? Borrowing from recent advances in super-resolution imaging, it is shown that this non-linear, non-convex inverse problem can be efficiently relaxed into a convex linear inverse problem over the space of Radon measures in$\mathbb {R}^{3}$. The new linear operator introduced here stems from the fundamental solution of the wave equation combined with the receivers' responses. An adaptation of the Sliding Frank-Wolfe algorithm is proposed to numerically solve the problemoff-the-grid,i.e., in continuous 3D space. Idealized simulated experiments show that the approach can recover hundreds of image sources at a rate and accuracy that are not achievable by previous methods, using a compact microphone array and source placed at random in random-sized shoe-box rooms. The impact of noise, sampling rate and array diameter on these results is also examined. Tom Sprunck, Antoine Deleforge, Yannick Privat, Cédric Foy |
IEEE Signal Process. Lett. | 2 |
| 2021 | Detecting Acoustic Reflectors Using A Robot's Ego-NoiseabstractIn this paper, we propose a method to estimate the proximity of an acoustic reflector, e.g., a wall, using ego-noise, i.e., the noise produced by the moving parts of a listening robot. This is achieved by estimating the times of arrival of acoustic echoes reflected from the surface. Simulated experiments show that the proposed non-intrusive approach is capable of accurately estimating the distance of a reflector up to 1 meter and outperforms a previously proposed intrusive approach under loud ego-noise conditions. The proposed method is helped by a probabilistic echo detector that estimates whether or not an acoustic reflector is within a short range of the robotic platform. This preliminary investigation paves the way towards a new kind of collision avoidance system that would purely rely on audio sensors rather than conventional proximity sensors. Usama Saqib, Antoine Deleforge, Jesper Rindom Jensen |
ICASSP | 2 |
| 2020 | Blaster: An Off-Grid Method for Blind and Regularized Acoustic Echoes RetrievalabstractAcoustic echoes retrieval is a research topic that is gaining importance in many speech and audio signal processing applications such as speech enhancement, source separation, dereverberation and room geometry estimation. This work proposes a novel approach to blindly retrieve the off-grid timing of early acoustic echoes from a stereophonic recording of an unknown sound source such as speech. It builds on the recent framework of continuous dictionaries. In contrast with existing methods, the proposed approach does not rely on parameter tuning nor peak picking techniques by working directly in the parameter space of interest. The accuracy and robustness of the method are assessed on challenging simulated setups with varying noise and reverberation levels and are compared to two state-of-the-art methods. Diego Di Carlo, Clement Elvira, Antoine Deleforge, Nancy Bertin, Rémi Gribonval |
ICASSP | 3 |
| 2020 | Filterbank Design for End-to-end Speech SeparationabstractSingle-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet. In parallel, parameterized filterbanks have been proposed for speaker recognition where only center frequencies and bandwidths are learned. In this work, we extend real-valued learned and parameterized filterbanks into complex-valued analytic filterbanks and define a set of corresponding representations and masking strategies. We evaluate these filterbanks on a newly released noisy speech separation dataset (WHAM). The results show that the proposed analytic learned filterbank consistently outperforms the real-valued filterbank of ConvTasNet. Also, we validate the use of parameterized filterbanks and show that complex-valued representations and masks are beneficial in all conditions. Finally, we show that the STFT achieves its best performance for 2 ms windows. Manuel Pariente, Samuele Cornell, Antoine Deleforge, Emmanuel Vincent 0001 |
ICASSP | 3 |
| 2020 | Asteroid: The PyTorch-Based Audio Source Separation Toolkit for ResearchersabstractThis paper describes Asteroid, the PyTorch-based audio source separation toolkit for researchers. Inspired by the most successful neural source separation systems, it provides all neural building blocks required to build such a system. To improve reproducibility, Kaldi-style recipes on common audio source separation datasets are also provided. This paper describes the software architecture of Asteroid and its most important features. By showing experimental results obtained with Asteroid's recipes, we show that our implementations are at least on par with most results reported in reference papers. The toolkit is publicly available at https://github.com/mpariente/asteroid . Manuel Pariente, Samuele Cornell, Joris Cosentino, Sunit Sivasankaran, Efthymios Tzinis, Jens Heitkaemper, Michel Olvera, Fabian-Robert Stöter, Mathieu Hu, Juan M. Martín-Doñas, David Ditter, Ariel Frank, Antoine Deleforge, Emmanuel Vincent 0001 |
INTERSPEECH | 13 |
| 2019 | Mirage: 2D Source Localization Using Microphone Pair Augmentation with EchoesabstractIt is commonly observed that acoustic echoes hurt per mance of sound source localization (SSL) methods. We troduce the concept of microphone array augmentation echoes (MIRAGE) and show how estimation of early-e characteristics can in fact benefit SSL. We propose a learn based scheme for echo estimation combined with a phys based scheme for echo aggregation. In a simple scenario volving 2 microphones close to a reflective surface and source, we show using simulated data that the proposed proach performs similarly to a correlation-based metho azimuth estimation while retrieving elevation as well from 2 microphones only, an impossible task in anechoic settings. Diego Di Carlo, Antoine Deleforge, Nancy Bertin |
ICASSP | 2 |
| 2019 | A Statistically Principled and Computationally Efficient Approach to Speech Enhancement Using Variational AutoencodersabstractRecent studies have explored the use of deep generative models of speech spectra based of variational autoencoders (VAEs), combined with unsupervised noise models, to perform speech enhancement. These studies developed iterative algorithms involving either Gibbs sampling or gradient descent at each step, making them computationally expensive. This paper proposes a variational inference method to iteratively estimate the power spectrogram of the clean speech. Our main contribution is the analytical derivation of the variational steps in which the en-coder of the pre-learned VAE can be used to estimate the varia-tional approximation of the true posterior distribution, using the very same assumption made to train VAEs. Experiments show that the proposed method produces results on par with the afore-mentioned iterative methods using sampling, while decreasing the computational cost by a factor 36 to reach a given performance . Manuel Pariente, Antoine Deleforge, Emmanuel Vincent 0001 |
INTERSPEECH | 2 |
| 2018 | Blind Source Separation Using Mixtures of Alpha-Stable DistributionsabstractWe propose a new blind source separation algorithm based on mixtures of α-stable distributions. Complex symmetric α-stable distributions have been recently showed to better model audio signals in the time-frequency domain than classical Gaussian distributions thanks to their larger dynamic range. However, inference with these models is notoriously hard to perform because their probability density functions do not have a closed-form expression in general. Here, we introduce a novel method for estimating mixtures of α-stable distributions based on characteristic function matching. We apply this to the blind estimation of binary masks in individual frequency bands from multichannel convolutive audio mixtures. We show that the proposed method yields better separation performance than Gaussian-based binary-masking methods. Nicolas Keriven, Antoine Deleforge, Antoine Liutkus |
ICASSP | 2 |
| 2018 | Audio Source Separation with Magnitude Priors: The Beads ModelabstractAudio source separation comes with the need to devise multichannel filters that can exploit priors about the target signals. In that context, experience shows that modeling magnitude spectra is effective. However, devising a probabilistic model on complex spectral data with a prior on magnitudes is non trivial, because it should both reflect the prior but also be tractable for easy inference. In this paper, we approximate the ideal donut-shaped distribution of a complex variable with approximately known magnitude as a Gaussian mixture model called BEADS (Bayesian Expansion Approximating the Donut Shape) and show that it permits straightforward inference and filtering while effectively constraining the magnitudes of the signals to comply with the prior. As a result, we demonstrate large improvements over the Gaussian baseline for multichannel audio coding when exploiting the BEADS model. Antoine Liutkus, Christian Rohlfing, Antoine Deleforge |
ICASSP | 3 |
| 2018 | Separake: Source Separation with a Little Help from EchoesabstractIt is commonly believed that multipath hurts various audio processing algorithms. At odds with this belief, we show that multipath in fact helps sound source separation, even with very simple propagation models. Unlike most existing methods, we neither ignore the room impulse responses, nor we attempt to estimate them fully. We rather assume to know the positions of a few virtual microphones generated by echoes and we show how this gives us enough spatial diversity to get a performance boost over the anechoic case. We show improvements for two standard algorithms-one that uses only magnitudes of the transfer functions, and one that also uses the phases. Concretely, we show that multi-channel non-negative matrix factorization aided with a small number of echoes beats the vanilla variant of the same algorithm, and that with magnitude information only, echoes enable separation where it was previously impossible. Robin Scheibler, Diego Di Carlo, Antoine Deleforge, Ivan Dokmanic |
ICASSP | 3 |
| 2018 | DREGON: Dataset and Methods for UAV-Embedded Sound Source LocalizationabstractThis paper introduces DREGON, a novel publicly-available dataset that aims at pushing research in sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). The dataset contains both clean and noisy in-flight audio recordings continuously annotated with the 3D position of the target sound source using an accurate motion capture system. In addition, various signals of interests are available such as the rotational speed of individual rotors and inertial measurements at all time. Besides introducing the dataset, this paper sheds light on the specific properties, challenges and opportunities brought by the emerging task of UAV-embedded sound source localization. Several baseline methods are evaluated and compared on the dataset, with real-time applicability in mind. Very promising results are obtained for the localization of a broad-band source in loud noise conditions, while speech localization remains a challenge under extreme noise levels. Martin Strauss 0003, Pol Mordel, Victor Miguet, Antoine Deleforge |
IROS | 4 |
| 2018 | MULAN: A Blind and Off-Grid Method for Multichannel Echo RetrievalabstractThis paper addresses the general problem of blind echo retrieval, i.e., given M sensors measuring in the discrete-time domain M mixtures of K delayed and attenuated copies of an unknown source signal, can the echo location and weights be recovered? This problem has broad applications in fields such as sonars, seismology, ultrasounds or room acoustics. It belongs to the broader class of blind channel identification problems, which have been intensively studied in signal processing. All existing methods proceed in two steps: (i) blind estimation of sparse discrete-time filters and (ii) echo information retrieval by peak picking. The precision of these methods is fundamentally limited by the rate at which the signals are sampled: estimated echo locations are necessary on-grid, and since true locations never match the sampling grid, the weight estimation precision is also strongly limited. This is the so-called basis-mismatch problem in compressed sensing. We propose a radically different approach to the problem, building on top of the framework of finite-rate-of-innovation sampling. The approach operates directly in the parameter-space of echo locations and weights, and enables near-exact blind and off-grid echo retrieval from discrete-time measurements. It is shown to outperform conventional methods by several orders of magnitudes in precision. Helena Peic Tukuljac, Antoine Deleforge, Rémi Gribonval |
NeurIPS | 2 |
| 2018 | Statistical Analysis of the Multichannel Wiener Filter Using a Bivariate Normal Distribution for Sample Covariance MatricesabstractThis paper studies the statistical performance of the multichannel Wiener filter (MWF) when the weights are computed using estimates of the sample covariance matrices of the noisy and the noise signals. It is well known that the optimal weights of the minimum variance distortionless response beamformer are only determined by the noisy sample covariance matrix or the noise sample covariance matrix, while those of the MWF are determined by both of them. Therefore, the difficulty increases dramatically in statistically analyzing the MWF when compared to analyzing the MVDR, where the main reason is that expressing the general joint probability density function (p.d.f.) of the two sample covariance matrices presented a Hitherto unsolved problem, to the best of our knowledge. For a deeper insight into the statistical performance of the MWF, this paper first introduces a bivariate normal distribution to approximately model the joint p.d.f. of the noisy and the noise sample covariance matrices. Each sample covariance matrix is approximately modeled by a random scalar multiplied by its true covariance matrix. This approximation is designed to preserve both the bias and the mean squared error of the matrix with respect to a natural distance on covariance matrices. The correlation of the bivariate normal distribution, referred to as the sample covariance matrices intrinsic correlation coefficient, captures all second-order dependencies of the noisy and the noise sample covariance matrices. By using the proposed bivariate normal distribution, the performance of the MWF can be predicted from the derived analytical expressions and many interesting results are revealed. As an example, the theoretical analysis demonstrates that the MWF performance may degrade in terms of noise reduction and signal-to-noise-ratio improvement when using more sensors in some noise scenarios. Chengshi Zheng, Antoine Deleforge, Xiaodong Li 0002, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Phase unmixing: Multichannel source separation with magnitude constraintsabstractWe consider the problem of estimating the phases of K mixed complex signals from a multichannel observation, when the mixing matrix and signal magnitudes are known. This problem can be cast as a non-convex quadratically constrained quadratic program which is known to be NP-hard in general. We propose three approaches to tackle it: a heuristic method, an alternate minimization method, and a convex relaxation into a semi-definite program. The last two approaches are showed to outperform the oracle multichannel Wiener filter in under-determined informed source separation tasks, using simulated and speech signals. The convex relaxation approach yields best results, including the potential for exact source separation in under-determined settings. Antoine Deleforge, Yann Traonmilin |
ICASSP | 1 |
| 2017 | Phase retrieval with a multivariate Von Mises prior: From a Bayesian formulation to a lifting solutionabstractIn this paper, we investigate a new method for phase recovery when prior information on the missing phases is available. In particular, we propose to take into account this information in a generic fashion by means of a multivariate Von Mises distribution. Building on a Bayesian formulation (a Maximum A Posteriori estimation), we show that the problem can be expressed using a Mahalanobis distance and be solved by a lifting optimization procedure. Angélique Dremeau, Antoine Deleforge |
ICASSP | 2 |
| 2017 | Hearing in a shoe-box: Binaural source position and wall absorption estimation using virtually supervised learningabstractThis paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes. These scenes are used to build a dataset of spatial binaural features annotated with acoustic properties such as the 3D source position and the walls' absorption coefficients. A probabilistic high- to low-dimensional regression framework is used to learn a mapping from these features to the acoustic properties. Results indicate that this mapping successfully estimates the azimuth and elevation of new sources, but also their range and even the walls' absorption coefficients solely based on binaural signals. Results also reveal that incorporating random-diffusion effects in the data significantly improves the estimation of all parameters. Saurabh Kataria 0001, Clément Gaultier, Antoine Deleforge |
ICASSP | 3 |
| 2017 | Robust Head-Pose Estimation Based on Partially-Latent Mixture of Linear RegressionsabstractHead-pose estimation has many applications, such as social event analysis, human-robot and human-computer interaction, driving assistance, and so forth. Head-pose estimation is challenging, because it must cope with changing illumination conditions, variabilities in face orientation and in appearance, partial occlusions of facial landmarks, as well as bounding-box-to-face alignment errors. We propose to use a mixture of linear regressions with partially-latent output. This regression method learns to map high-dimensional feature vectors (extracted from bounding boxes of faces) onto the joint space of head-pose angles and bounding-box shifts, such that they are robustly predicted in the presence of unobservable phenomena. We describe in detail the mapping method that combines the merits of unsupervised manifold learning techniques and of mixtures of regressions. We validate our method with three publicly available data sets and we thoroughly benchmark four variants of the proposed algorithm with several state-of-the-art head-pose estimation methods. Vincent Drouard, Radu Horaud, Antoine Deleforge, Sileye O. Ba, Georgios Evangelidis 0002 |
IEEE Trans. Image Process. | 3 |
| 2016 | Ego-noise reduction using a motor data-guided multichannel dictionaryabstractWe address the problem of ego-noise reduction, i.e., suppressing the noise a robot causes by its own motions. Such noise degrades the recorded microphone signal massively such that the robot's auditory capabilities suffer. To suppress it, it is intuitive to use also motor data, since it provides additional information about the robot's joints and thereby the noise sources. We propose to fuse motor data to a recently proposed multichannel dictionary algorithm for ego-noise reduction. At training, a dictionary is learned that captures spatial and spectral characteristics of ego-noise. At testing, nonlinear classifiers are used to efficiently associate the current robot's motor state to relevant sets of entries in the learned dictionary. By this, computational load is reduced by one third in typical scenarios while achieving at least the same noise reduction performance. Moreover, we propose to train dictionaries on different microphone array geometries and use them for ego-noise reduction while the head to which the microphones are mounted is moving. In such scenarios, the motor guided approach results in significantly better performance values. Alexander Schmidt 0004, Antoine Deleforge, Walter Kellermann |
IROS | 2 |
| 2015 | Phase-optimized K-SVD for signal extraction from underdetermined multichannel sparse mixturesabstractWe propose a novel sparse representation for heavily underdetermined multichannel sound mixtures, i.e., with much more sources than microphones. The proposed approach operates in the complex Fourier domain, thus preserving spatial characteristics carried by phase differences. We derive a generalization of K-SVD which jointly estimates a dictionary capturing both spectral and spatial features, a sparse activation matrix, and all instantaneous source phases from a set of signal examples. This dictionary can be used to extract the learned signal from a new input mixture. The method is applied to the challenging problem of ego-noise reduction for robot audition. We demonstrate its superiority relative to conventional dictionary-based techniques using real-room recordings. Antoine Deleforge, Walter Kellermann |
ICASSP | 1 |
| 2015 | Head pose estimation via probabilistic high-dimensional regressionabstractThis paper addresses the problem of head pose estimation with three degrees of freedom (pitch, yaw, roll) from a single image. Pose estimation is formulated as a high-dimensional to low-dimensional mixture of linear regression problem. We propose a method that maps HOG-based descriptors, extracted from face bounding boxes, to corresponding head poses. To account for errors in the observed bounding-box position, we learn regression parameters such that a HOG descriptor is mapped onto the union of a head pose and an offset, such that the latter optimally shifts the bounding box towards the actual position of the face in the image. The performance of the proposed method is assessed on publicly available datasets. The experiments that we carried out show that a relatively small number of locally-linear regression functions is sufficient to deal with the non-linear mapping problem at hand. Comparisons with state-of-the-art methods show that our method outperforms several other techniques. Vincent Drouard, Sileye O. Ba, Georgios Evangelidis 0002, Antoine Deleforge, Radu Horaud |
ICIP | 4 |
| 2015 | Acoustic Space Learning for Sound-Source Separation and Localization on Binaural ManifoldsabstractIn this paper, we address the problems of modeling the acoustic space generated by a full-spectrum sound source and using the learned model for the localization and separation of multiple sources that simultaneously emit sparse-spectrum sounds. We lay theoretical and methodological grounds in order to introduce the binaural manifold paradigm. We perform an in-depth study of the latent low-dimensional structure of the high-dimensional interaural spectral data, based on a corpus recorded with a human-like audiomotor robot head. A nonlinear dimensionality reduction technique is used to show that these data lie on a two-dimensional (2D) smooth manifold parameterized by the motor states of the listener, or equivalently, the sound-source directions. We propose a probabilistic piecewise affine mapping model (PPAM) specifically designed to deal with high-dimensional data exhibiting an intrinsic piecewise linear structure. We derive a closed-form expectation-maximization (EM) procedure for estimating the model parameters, followed by Bayes inversion for obtaining the full posterior density function of a sound-source direction. We extend this solution to deal with missing data and redundancy in real-world spectrograms, and hence for 2D localization of natural sound sources such as speech. We further generalize the model to the challenging case of multiple sound sources and we propose a variational EM framework. The associated algorithm, referred to as variational EM for source separation and localization (VESSL) yields a Bayesian estimation of the 2D locations and time-frequency masks of all the sources. Comparisons of the proposed approach with several existing methods reveal that the combination of acoustic-space learning with Bayesian inference enables our method to outperform state-of-the-art methods. Antoine Deleforge, Florence Forbes, Radu Horaud |
Int. J. Neural Syst. | 1 |
| 2015 | Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear RegressionabstractThis paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on monaural segregation. The method starts with a training stage that establishes a locally linear Gaussian regression model between the directional coordinates of all the sources and the auditory features extracted from binaural measurements. While fixed-length wide-spectrum sounds (white noise) are used for training to reliably estimate the model parameters, we show that the testing (localization) can be extended to variable-length sparse-spectrum sounds (such as speech), thus enabling a wide range of realistic applications. Indeed, we demonstrate that the method can be used for audio-visual fusion, namely to map speech signals onto images and hence to spatially align the audio and visual modalities, thus enabling to discriminate between speaking and non-speaking faces. We release a novel corpus of real-room recordings that allow quantitative evaluation of the co-localization method in the presence of one or two sound sources. Experiments demonstrate increased accuracy and speed relative to several state-of-the-art methods. Antoine Deleforge, Radu Horaud, Yoav Y. Schechner, Laurent Girin |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Variational EM for binaural sound-source separation and localizationabstractThe sound-source separation and localization (SSL) problems are addressed within a unified formulation. Firstly, a mapping between white-noise source locations and binaural cues is estimated. Secondly, SSL is solved via Bayesian inversion of this mapping in the presence of multiple sparse-spectrum emitters (such as speech), noise and reverberations. We propose a variational EM algorithm which is described in detail together with initialization and convergence issues. Extensive real-data experiments show that the method outperforms the state-of-the-art both in separation and localization (azimuth and elevation). Antoine Deleforge, Florence Forbes, Radu Horaud |
ICASSP | 1 |
| 2012 | The cocktail party robot: sound source separation and localisation with an active binaural headabstractHuman-robot communication is often faced with the difficult problem of interpreting ambiguous auditory data. For example, the acoustic signals perceived by a humanoid with its on-board microphones contain a mix of sounds such as speech, music, electronic devices, all in the presence of attenuation and reverberations. In this paper we propose a novel method, based on a generative probabilistic model and on active binaural hearing, allowing a robot to robustly perform sound-source separation and localization. We show how interaural spectral cues can be used within a constrained mixture model specifically designed to capture the richness of the data gathered with two microphones mounted onto a human-like artificial head. We describe in detail a novel EM algorithm, we analyse its initialization, speed of convergence and complexity, and we assess its performance with both simulated and real data. Antoine Deleforge, Radu Horaud |
HRI | 1 |