François Grondin

dblp:14/2074 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0002-0563-4251ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 6 since 2021Systems, architecture and hardware · 10 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Sound Source Localization for Human-Robot Interaction in Outdoor Environments
abstract
This paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment strategy is combined with a time-domain acoustic echo cancellation algorithm to estimate a time-frequency ideal ratio mask to isolate the target speech from interferences and environmental noise. This allows selective sound source localization, and provides the robot with the direction of arrival of sound from the active operator, which enables rich interaction in noisy scenarios. Results demonstrate an average angle error of 4 degrees and an accuracy within 5 degrees of 95% at a signal-to-noise ratio of 1dB, which is significantly superior to the state-of-the-art localization methods.
Victor Liu, Tim D. Barfoot, Jordy Sehn, Jack Collier, François Grondin
IROS5
2024 Resource-Efficient Separation Transformer
abstract
Transformers have recently achieved state-of-the-art performance in speech separation. These models, however, are computationally demanding and require a lot of learnable parameters. This paper explores Transformer-based speech separation with a reduced computational cost. Our main contribution is the development of the Resource-Efficient Separation Transformer (RE-SepFormer), a self-attention-based architecture that reduces the computational burden in two ways. First, it uses non-overlapping blocks in the latent space. Second, it operates on compact latent summaries calculated from each chunk. The RE-SepFormer reaches a competitive performance on the popular WSJ0-2Mix and WHAM! datasets in both causal and non-causal settings. Remarkably, it scales significantly better than the previous Transformer-based architectures in terms of memory and inference time, making it more suitable for processing long mixtures.
Luca Della Libera, Cem Subakan, Mirco Ravanelli, Samuele Cornell, Frédéric Lepoutre, François Grondin
ICASSP6
2024 Unsupervised Improved MVDR Beamforming for Sound Enhancement
Jacob Kealey, John R. Hershey, François Grondin
INTERSPEECH3
2023 Fast Cross-Correlation for TDoA Estimation on Small Aperture Microphone Arrays
abstract
This paper introduces the Fast Cross-Correlation (FCC) method for Time Difference of Arrival (TDoA) Estimation for pairs of microphones on a small aperture microphone array. FCC relies on low-rank decomposition and exploits symmetry in even and odd bases to speed up computation while preserving TDoA accuracy. FCC reduces the number of flops by a factor of 4.5 and the execution speed by factors between 3.5 and 8.3 on embedded hardware, compared to the state-of-the-art Generalized Cross-Correlation (GCC) method that relies on the Fast Fourier Transform (FFT). This improvement can provide portable microphone arrays with extended battery life and allow real-time processing on low-cost hardware.
François Grondin, Marc-Antoine Maheux, Jean-Samuel Lauzon, Jonathan Vincent, François Michaud
ICASSP1
2023 Ego-Noise Reduction of a Mobile Robot Using Noise Spatial Covariance Matrix Learning and Minimum Variance Distortionless Response
abstract
The performance of speech and events recognition systems significantly improved recently thanks to deep learning methods. However, some of these tasks remain challenging when algorithms are deployed on robots due to the unseen mechanical noise and electrical interference generated by their actuators while training the neural networks. Ego-noise reduction as a preprocessing step therefore can help solve this issue when using pre-trained speech and event recognition algorithms on robots. In this paper, we propose a new method to reduce ego-noise using only a microphone array and less than two minute of noise recordings. Using Principal Component Analysis (PCA), the best covariance matrix candidate is selected from a dictionary created online during calibration and used with the Minimum Variance Distortionless Response (MVDR) beamformer. Results show that the proposed method runs in real-time, improves the signal-to-distortion ratio (SDR) by up to 10 dB, decreases the word error rate (WER) by 55% in some cases and increases the Average Precision (AP) of event detection by up to 0.2.
Pierre-Olivier Lagacé, François Ferland, François Grondin
IROS3
2023 SmartBelt: A Wearable Microphone Array for Sound Source Localization with Haptic Feedback
abstract
This paper introduces SmartBelt, a wearable microphone array on a belt that performs sound source localization and returns the direction of arrival with respect to the user waist. One of the haptic motors on the belt then vibrates in the corresponding direction to provide useful feedback to the user. We also introduce a simple calibration step to adapt the belt to different waist sizes. Experiments are performed to confirm the accuracy of this wearable sound source localization system, and results show a Mean Average Error (MAE) of 2.90°, and a correct haptic motor selection with a rate of 92.3%. Results suggest the device can provide useful haptic feedback, and will be evaluated in a study with people having hearing impairments.
Simon Michaud, Benjamin Moffett, Ana Tapia Rousiouk, Victoria Duda, François Grondin
RO-MAN5
2023 Exploring Self-Attention Mechanisms for Speech Separation
abstract
Transformers have enabled impressive improvements in deep learning. They often outperform recurrent and convolutional models in many tasks while taking advantage of parallel processing. Recently, we proposed the SepFormer, which obtains state-of-the-art performance in speech separation with the WSJ0-2/3 Mix datasets. This paper studies in-depth Transformers for speech separation. In particular, we extend our previous findings on the SepFormer by providing results on more challenging noisy and noisy-reverberant datasets, such as LibriMix, WHAM!, and WHAMR!. Moreover, we extend our model to perform speech enhancement and provide experimental evidence on denoising and dereverberation tasks. Finally, we investigate, for the first time in speech separation, the use of efficient self-attention mechanisms such as Linformers, Lonformers, and ReFormers. We found that they reduce memory requirements significantly. For example, we show that the Reformer-based attention outperforms the popular Conv-TasNet model on the WSJ0-2Mix dataset while being faster at inference and comparable in terms of memory consumption.
Cem Subakan, Mirco Ravanelli, Samuele Cornell, François Grondin, Mirko Bronzi
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Learning Filterbanks for End-to-End Acoustic Beamforming
abstract
Recent work on monaural source separation has shown that performance can be increased by using fully learned filterbanks with short windows. On the other hand it is widely known that, for conventional beamforming techniques, performance increases with long analysis windows. This applies also to most hybrid neural beamforming methods which rely on a deep neural network (DNN) to estimate the spatial covariance matrices. In this work we try to bridge the gap between these two worlds and explore fully end-to-end hybrid neural beamforming in which, instead of using the Short-Time-Fourier Transform, also the analysis and synthesis filterbanks are learnt jointly with the DNN. In detail, we explore two different types of learned filterbanks: fully learned and analytic. We perform a detailed analysis using the recent Clarity Challenge data and show that by using learnt filterbanks it is possible to surpass oracle-mask based beamforming for short windows.
Samuele Cornell, Manuel Pariente, François Grondin, Stefano Squartini
ICASSP3
2022 Real-M: Towards Speech Separation on Real Mixtures
abstract
In recent years, deep learning based source separation has achieved impressive results. Most studies, however, still evaluate separation models on synthetic datasets, while the performance of state-of-the-art techniques on in-the-wild speech data remains an open question. This paper contributes to fill this gap in two ways. First, we release the REAL-M dataset, a crowd-sourced corpus of real-life mixtures. Secondly, we address the problem of performance evaluation of real-life mixtures, where the ground truth is not available. We bypass this issue by carefully designing a blind Scale-Invariant Signal-to-Noise Ratio (SI-SNR) neural estimator. Through a user study, we show that our estimator reliably evaluates the separation performance on real mixtures, i.e. we observe that the performance predictions of the SI-SNR estimator correlate well with human opinions. Moreover, when evaluating popular speech separation models, we observe that the performance trends predicted by our estimator on the REAL-M dataset closely follow the performance trends achieved on synthetic benchmarks.
Cem Subakan, Mirco Ravanelli, Samuele Cornell, François Grondin
ICASSP4
2022 Audio Scene Monitoring Using Redundant Ad Hoc Microphone Array Networks
abstract
We present a system for localizing sound sources in a room with severalad hocmicrophone arrays. Each circular array performs direction of arrival (DOA) estimation independently using commercial software. The DOAs are fed to a fusion center, concatenated, and used to perform the localization based on two proposed methods, which require only a few labeled source locations (anchor points) for training. The first proposed method is based on principal component analysis (PCA) of the observed DOA and does not require any knowledge of anchor points. The array cluster can then perform localization on a manifold defined by the PCA of concatenated DOAs over time. The second proposed method performs localization using an affine transformation between the DOA vectors and the room manifold. The PCA has fewer requirements on the training sequence, but is less robust to missing DOAs from one of the arrays. The methods are demonstrated with five IoT 8-microphone circular arrays, placed at unspecified fixed locations in an office. Both the PCA and the affine method can easily map out a rectangle based on a few anchor points with similar accuracy. The proposed methods provide a step toward monitoring activities in a smart home and require little installation effort as the array locations are not needed.
Peter Gerstoft, Yihan Hu 0002, Michael Bianco, Chaitanya Patil, Ardel Alegre, Yoav Freund, François Grondin
IEEE Internet Things J.7
2021 ECAPA-TDNN Embeddings for Speaker Diarization
abstract
Learning robust speaker embeddings is a crucial step in speaker diarization. Deep neural networks can accurately capture speaker discriminative characteristics and popular deep embeddings such as x-vectors are nowadays a fundamental component of modern diarization systems. Recently, some improvements over the standard TDNN architecture used for x-vectors have been proposed. The ECAPA-TDNN model, for instance, has shown impressive performance in the speaker verification domain, thanks to a carefully designed neural model. In this work, we extend, for the first time, the use of the ECAPA-TDNN model to speaker diarization. Moreover, we improved its robustness with a powerful augmentation scheme that concatenates several contaminated versions of the same signal within the same training batch. The ECAPA-TDNN model turned out to provide robust speaker embeddings under both close-talking and distant-talking conditions. Our results on the popular AMI meeting corpus show that our system significantly outperforms recently proposed approaches.
Nauman Dawalatabad, Mirco Ravanelli, François Grondin, Jenthe Thienpondt, Brecht Desplanques, Hwidong Na
Interspeech3
2020 Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT
abstract
This paper proposes a straightforward 2-D method to spatially calibrate the visual field of a camera with the auditory field of an array microphone by generating and overlaying an acoustic image over an optical image. Using a low-cost microphone array and an off-the-shelf camera, we show that polynomial regression can deal efficiently with non-linear camera distortion, and that a recently proposed sound source localization method for real-time processing, SVD-PHAT, can be adapted for this task.
François Grondin, Hao Tang 0002, James R. Glass
ICASSP1
2020 GEV Beamforming Supported by DOA-Based Masks Generated on Pairs of Microphones
abstract
Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion ratio (SDR). Recently, a combination of beamforming and speech separation networks have been proposed to improve the target source quality in the direction of arrival of interest. However, with this type of approach, the neural network needs to be trained in advance for a specific microphone array geometry, which limits versatility when adding/removing microphones, or changing the shape of the array. The solution presented in this paper is to train a neural network on pairs of microphones with different spacing and acoustic environmental conditions, and then use this network to estimate a time-frequency mask from all the pairs of microphones forming the array with an arbitrary shape. Using this mask, the target and noise covariance matrices can be estimated, and then used to perform generalized eigenvalue (GEV) beamforming. Results show that the proposed approach improves the SDR from 4.78 dB to 7.69 dB on average, for various microphone array geometries that correspond to commercially available hardware.
François Grondin, Jean-Samuel Lauzon, Jonathan Vincent, François Michaud
INTERSPEECH1
2020 3D Localization of a Sound Source Using Mobile Microphone Arrays Referenced by SLAM
abstract
A microphone array can provide a mobile robot with the capability of localizing, tracking and separating distant sound sources in 2D, i.e., estimating their relative elevation and azimuth. To combine acoustic data with visual information in real world settings, spatial correlation must be established. The approach explored in this paper consists of having two robots, each equipped with a microphone array, localizing themselves in a shared reference map using SLAM. Based on their locations, data from the microphone arrays are used to triangulate in 3D the location of a sound source in relation to the same map. This strategy results in a novel cooperative sound mapping approach using mobile microphone arrays. Trials are conducted using two mobile robots localizing a static or a moving sound source to examine in which conditions this is possible. Results suggest that errors under 0.3 m are observed when the relative angle between the two robots are above 30° for a static sound source, while errors under 0.3 m for angles between 40° and 140° are observed with a moving sound source.
Simon Michaud, Samuel Faucher, François Grondin, Jean-Samuel Lauzon, Mathieu Labbé, Dominic Létourneau, François Ferland, François Michaud
IROS3
2020 Dynamic Object Tracking and Masking for Visual SLAM
abstract
In dynamic environments, performance of visual SLAM techniques can be impaired by visual features taken from moving objects. One solution is to identify those objects so that their visual features can be removed for localization and mapping. This paper presents a simple and fast pipeline that uses deep neural networks, extended Kalman filters and visual SLAM to improve both localization and mapping in dynamic environments (around 14 fps on a GTX 1080). Results on the dynamic sequences from the TUM dataset using RTAB-Map as visual SLAM suggest that the approach achieves similar localization performance compared to other state-of-the-art methods, while also providing the position of the tracked dynamic objects, a 3D map free of those dynamic objects, better loop closure detection with the whole pipeline able to run on a robot moving at moderate speed.
Jonathan Vincent, Mathieu Labbé, Jean-Samuel Lauzon, François Grondin, Pier-Marc Comtois-Rivet, François Michaud
IROS4
2019 SVD-PHAT: A Fast Sound Source Localization Method
abstract
This paper introduces a new localization method called SVD-PHAT. The SVD-PHAT method relies on Singular Value De-composition of the SRP-PHAT projection matrix. A k-d tree is also proposed to speed up the search for the most likely direction of arrival of sound. We show that this method per-forms as accurately as SRP-PHAT, while reducing significantly the amount of computation required.
François Grondin, James R. Glass
ICASSP1
2019 A Deep Residual Network for Large-Scale Acoustic Scene Analysis
Logan Ford, Hao Tang 0002, François Grondin, James R. Glass
INTERSPEECH3
2019 Multiple Sound Source Localization with SVD-PHAT
abstract
This paper introduces a modification of phase transform on singular value decomposition (SVD-PHAT) to localize multiple sound sources. This work aims to improve localization accuracy and keeps the algorithm complexity low for real-time applications. This method relies on multiple scans of the search space, with projection of each low-dimensional observation onto orthogonal subspaces. We show that this method localizes multiple sound sources more accurately than discrete SRP-PHAT, with a reduction in the Root Mean Square Error up to 0.0395 radians.
François Grondin, James R. Glass
INTERSPEECH1
2019 State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH10
2019 Fast and Robust 3-D Sound Source Localization with DSVD-PHAT
abstract
This paper introduces a variant of the Singular Value Decomposition with Phase Transform (SVD-PHAT), named Difference SVD-PHAT (DSVD-PHAT), to achieve robust Sound Source Localization (SSL) in noisy conditions. Experiments are performed on a Baxter robot with a four-microphone planar array mounted on its head. Results show that this method offers similar robustness to noise as the state-of-the-art Multiple Signal Classification based on Generalized Singular Value Decomposition (GSVD-MUSIC) method, and considerably reduces the computational load by a factor of 250. This performance gain thus makes DSVD-PHAT appealing for real-time application on robots with limited on-board computing power.
François Grondin, James R. Glass
IROS1
2018 A Study of Enhancement, Augmentation and Autoencoder Methods for Domain Adaptation in Distant Speech Recognition
abstract
Speech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute.Most studies focus on tackling distant speech recognition as a separate problem, leaving little effort to adapting close-talking speech recognizers to distant speech.In this work, we review several approaches from a domain adaptation perspective.These approaches, including speech enhancement, multi-condition training, data augmentation, and autoencoders, all involve a transformation of the data between domains.We conduct experiments on the AMI data set, where these approaches can be realized under the same controlled setting.These approaches lead to different amounts of improvement under their respective assumptions.The purpose of this paper is to quantify and characterize the performance gap between the two domains, setting up the basis for studying adaptation of speech recognizers from close-talking speech to distant speech.Our results also have implications for improving distant speech recognition.
Hao Tang 0002, Wei-Ning Hsu, François Grondin, James R. Glass
INTERSPEECH3
2017 Localization of RW-UAVs using particle filtering over distributed microphone arrays
abstract
Rotary-Wing Air Vehicles (RW-UAVs), also referred to as drones, have gained in popularity over the last few years. Intrusions over secured areas have become common and authorities are actively looking for solutions to detect and localize undesired drones. The sound generated by the propellers of the RW-UAVs is powerful enough to be perceived by a human observer nearby. In this paper, we examine the use of particle filtering to detect and localize in 3D the position of a RW-UAV based on sound source localization (SSL) over distributed microphone arrays (MAs). Results show that the proposed method is able to detect and track a drone with precision, as long as the noise emitted by the RW-UAVs dominates the background noise.
Jean-Samuel Lauzon, François Grondin, Dominic Létourneau, Alexis Lussier Desbiens, François Michaud
IROS2
2016 Robust speech/non-speech discrimination based on pitch estimation for mobile robots
abstract
To be used on a mobile robot, speech/non-speech discrimination must be robust to environmental noise and to the position of the interlocutor, without necessarily having to satisfy low-latency requirements. To address these conditions, this paper presents a speech/non-speech discrimination approach based on pitch estimation. Pitch features are robust to noise and reverberation, and can be estimated over a few seconds. Results suggest that our approach is more robust compared to the use of Mel-Frequency Cepstrum Coefficients with Gaussian Mixture Models (MFCC-GMM) under high reverberation levels and additive noise (with an accuracy above 98% with a latency of 2.21 sec), which makes it ideal for mobile robot applications. The approach is also validated on a mobile robot equipped with a 8-microphone array, using speech/non-speech discrimination based on pitch estimation as a post-processing module of a localization, tracking and separation system.
François Grondin, François Michaud
ICRA1
2016 Noise mask for TDOA sound source localization of speech on mobile robots in noisy environments
abstract
Sound source localization is an important challenge for mobile robots operating in real life settings. Sound sources of interest, such as speech, are often corrupted by broadband coherent noise sound source(s) that are non-stationary during transitions between steady-state segments. The interfering noise introduces localization ambiguities leading to the localization of invalid sound sources. Masks to reduce such interferences perform well under stationary noise, but the performance degrades as localization of invalid sound sources generated by noise appear and disappear suddenly during transitions between steady-state. This paper presents a new mask based on speech non-stationarity to discriminate between the time difference of arrival (TDOA) of speech source and noise transition. Simulations and experiments on a mobile robot suggest that the proposed technique improve TDOA discrimination and reduces significantly localization of invalid sound sources caused by noise.
François Grondin, François Michaud
ICRA1
2016 Integration framework for speech processing with live visualization interfaces
abstract
Audition is a rich source of spatial, identity, linguistic and paralinguistic information. Processing all this information requires acquisition, processing and interpretation of sound sources, which are instantaneous, invisible and noisy signals. This can lead to different responses by the system in relation to the information perceived. This paper presents our first implementation of an integration framework for speech processing. Acquisition includes sound capture, sound source localization, tracking, separation and enhancement, and voice activity detection. Processing involves speech and emotion recognition. Interpretation consists of translating speech utterances into commands that can influence interaction through dialogue management and speech synthesis. The paper also describes two visualization interfaces, inspired by comic strips, to represent live vocal interactions in real life environments. These interfaces are used to demonstrate how the framework performs in live interactions and its use in a usability study.
David Brodeur, François Grondin, Yazid Attabi, Pierre Dumouchel, François Michaud
RO-MAN2
2015 Time difference of arrival estimation based on binary frequency mask for sound source localization on mobile robots
abstract
Localization of sound sources in adverse environments is an important challenge in robot audition. The target sound source is often corrupted by coherent broadband noise, which introduces localization ambiguities as noise is often mistaken as the target source. To discriminate the time difference of arrival (TDOA) parameters of the target source and noise, this paper presents a binary mask for weighted generalized cross-correlation with phase transform (GCC-PHAT). Simulation and experiments on a mobile robot suggest that the proposed technique improves TDOA discrimination. It also brings the additional benefit of modulating the computing load requirement according to voice activity.
François Grondin, François Michaud
IROS1
2014 Multimodal biometric identification system for mobile robots combining human metrology to face recognition and speaker identification
abstract
Recognizing a person from a distance is important to establish meaningful social interaction and to provide additional cues regarding the situations experienced by a robot. To do so, face recognition and speaker identification are biometrics commonly used, with identification performance that are influenced by the distance between the person and the robot. This paper presents a system that combines these biometrics with human metrology (HM) to increase identification performance and range. HM measures are derived from 2D silhouettes extracted online using a dynamic background subtraction approach, processing in parallel 45 front features and 24 side features in 400 ms compared to 38 front and 22 side features extracted in sequence in 30 sec by using the approach presented by Lin and Wang [1]. By having each modality identify a set of up to five possible candidates, results suggest that combining modalities provide better performance compared to what each individual modality provides, from a wider range of distances.
Simon Ouellet, François Grondin, Francis Leconte, François Michaud
RO-MAN2
2012 WISS, a speaker identification system for mobile robots
abstract
This paper presents WISS, a speaker identification system for mobile robots integrated to ManyEars, a sound source localization, tracking and separation system. Speaker identification consists in recognizing an individual among a group of known speakers. For mobile robots, performing speaker identification in presence of noise that changes over time is one important challenge. To deal with this issue, WISS uses Parallel Model Combination (PMC) and masks to update in real-time the speaker models (obtained in clean conditions) to both additive and convolutive noises. The results show that the weighted rate of good speaker identifications is 96% on average for a Signal-to-Noise Ratio (SNR) of 16 dB, whereas it only decreases to 84% when the SNR drops to 2 dB.
François Grondin, François Michaud
ICRA1