VLDB 2026 Research / reviewers in the wild / expert
François G. Germain
dblp:07/10306
· DBLP profile ↗
25ranked-venue papers
6as first author
20since 2021 · last 2025
0000-0002-8973-5315ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingabstractAudio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost of annotating training data with the class labels of all these events, along with their start and end times, imposes constraints on the scalability of AVVP techniques unless they can be trained in a weakly-supervised setting, where only modality-agnostic, video-level labels are available in the training data. To this end, recently proposed approaches seek to generate segment-level pseudo-labels to better guide model training. However, the absence of inter-segment dependencies when generating these pseudo-labels and the general bias towards predicting labels that are absent in a segment limit their performance. This work proposes a novel approach towards overcoming these weaknesses called Uncertainty-Weighted Weakly-Supervised Audio-Visual Video Parsing (UWAV). Additionally, our innovative approach factors in the uncertainty associated with these estimated pseudo-labels and incorporates a feature mixup based training regularization for improved training. Empirical results show that UWAV outperforms state-of-the-art methods for the AVVP task on multiple metrics, across two different datasets, attesting to its effectiveness and generalizability.1 Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain, Michael J. Jones 0001, Moitreya Chatterjee |
CVPR | 4 |
| 2025 | No Class Left Behind: A Closer Look at Class Balancing for Audio TaggingabstractLarge-scale audio tagging datasets like AudioSet usually suffer from severe class imbalance comprising many audio examples for common sound classes but only few examples of rare sound classes. The latter, however, may yet be equally or even more important to recognize. Therefore, it is common practice to sample examples from rare classes more frequently during training. At the same time, the effects of such balancing on a model’s training and tagging performance are still little understood. In this work, we investigate how it affects training convergence and tagging performance. We consider varying degrees of balancing and investigate whether classes converge simultaneously or if there is a benefit from selecting different balancing rates for each class. Furthermore, we investigate data efficient oversampling, which keeps audio files from rare classes in memory, and repeats them in close succession over multiple batches, minimizing data loading from disk. Finally, we show that for AudioSet, the optimal amount of class balancing is different when fine-tuning a model pre-trained via self-supervised learning, versus training a supervised model from scratch. Janek Ebbers, François G. Germain, Kevin Wilkinghoff, Gordon Wichern, Jonathan Le Roux |
ICASSP | 2 |
| 2025 | Retrieval-Augmented Neural Field for HRTF Upsampling and PersonalizationabstractHead-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields, spatial upsampling only from a few measured directions, e.g., 3 or 5 measurements, is still challenging. To tackle this problem, we propose a retrieval-augmented neural field (RANF). RANF retrieves a subject whose HRTFs are close to those of the target subject from a dataset. The HRTF of the retrieved subject at the desired direction is fed into the neural field in addition to the sound source direction itself. Furthermore, we present a neural network that can efficiently handle multiple retrieved subjects, inspired by a multi-channel processing technique called transform-average-concatenate. Our experiments confirm the benefits of RANF on the SONICOM dataset, and it is a key component in the winning solution of Task 2 of the listener acoustic personalization challenge 2024. Yoshiki Masuyama, Gordon Wichern, François G. Germain, Christopher Ick, Jonathan Le Roux |
ICASSP | 3 |
| 2025 | Leveraging Audio-Only Data for Text-Queried Target Sound ExtractionabstractThe goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text queries, the limited number of available high-quality text-audio pairs hinders the data scaling. To this end, this work explores how to leverage audio-only data without any captions for the text-queried TSE task to potentially scale up the data amount. A straightforward way to do so is to use a joint audio-text embedding model, such as the contrastive language-audio pre-training (CLAP) model, as a query encoder and train a TSE model using audio embeddings obtained from the ground-truth audio. The TSE model can then accept text queries at inference time by switching to the text encoder. While this approach should work if the audio and text embedding spaces in CLAP were well aligned, in practice, the embeddings have domain-specific information that causes the TSE model to overfit to audio queries. We investigate several methods to avoid overfitting and show that simple embedding-manipulation methods such as dropout can effectively alleviate this issue. Extensive experiments demonstrate that using audio-only data with embedding dropout is as effective as using text captions during training, and audio-only data can be effectively leveraged to improve text-queried TSE models. Kohei Saijo, Janek Ebbers, François G. Germain, Sameer Khurana, Gordon Wichern, Jonathan Le Roux |
ICASSP | 3 |
| 2025 | Task-Aware Unified Source SeparationabstractSeveral attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or cinematic audio source separation (CASS) with a single model. These models are trained on large-scale data including speech, instruments, or sound events and can often successfully separate a wide range of sources. However, it is still challenging for such models to cover all separation tasks because some of them are contradictory (e.g., musical instruments are separated in MSS while they have to be grouped in CASS). To overcome this issue and support all the major separation tasks, we propose a task-aware unified source separation (TUSS) model. The model uses a variable number of learnable prompts to specify which source to separate, and changes its behavior depending on the given prompts, enabling it to handle all the major separation tasks including contradictory ones. Experimental results demonstrate that the proposed TUSS model successfully handles the five major separation tasks mentioned earlier. We also provide some audio examples, including both synthetic mixtures and real recordings, to demonstrate how flexibly the TUSS model changes its behavior at inference depending on the prompts. Kohei Saijo, Janek Ebbers, François G. Germain, Gordon Wichern, Jonathan Le Roux |
ICASSP | 3 |
| 2025 | Keeping the Balance: Anomaly Score Calculation for Domain GeneralizationabstractEmitted sounds may drastically change when using different microphones, when properties of the sound sources change, or when recording in different acoustic environments. Ideally, anomalous sound detection (ASD) systems should be able to generalize well to unseen target domains by only providing a few target domain samples to define how normal data samples sound like, without needing to re-train or modify the system. In contrast with the source domain, for which many normal training samples are available, accurately estimating the underlying distribution of normal data after a domain shift based on very few samples is challenging. This usually leads to a mismatch between the corresponding anomaly scores of source and target domains and significantly reduces performance. In this work, we propose a framework for re-scaling anomaly scores based on the ratio between the cosine distance of a test sample to a normal reference sample and the distances to this sample’s next-closest neighbors in the reference set. In experimental evaluations, it is shown that the re-scaled anomaly scores reduce the domain mismatch for multiple domains. As a result, we obtain new state-of-the-art performances on the DCASE2020 and DCASE2023 ASD datasets. Kevin Wilkinghoff, Haici Yang, Janek Ebbers, François G. Germain, Gordon Wichern, Jonathan Le Roux |
ICASSP | 4 |
| 2025 | HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
Amir Hussein, Sameer Khurana, Gordon Wichern, François G. Germain, Jonathan Le Roux |
INTERSPEECH | 4 |
| 2025 | Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses
Christopher Ick, Gordon Wichern, Yoshiki Masuyama, François G. Germain, Jonathan Le Roux |
INTERSPEECH | 4 |
| 2025 | Factorized RVQ-GAN For Disentangled Speech TokenizationabstractInternational audience Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos, Juraj Novosad, Peter Gazdik, Ellen Zhang, Zili Huang, Amir Hussein, Ricard Marxer, Yoshiki Masuyama, Ryo Aihara, Chiori Hori, François G. Germain, Gordon Wichern, Jonathan Le Roux |
INTERSPEECH | 14 |
| 2025 | Investigating continuous autoregressive generative speech enhancement
Haici Yang, Gordon Wichern, Ryo Aihara, Yoshiki Masuyama, Sameer Khurana, François G. Germain, Jonathan Le Roux |
INTERSPEECH | 6 |
| 2024 | Generation or Replication: Auscultating Audio Latent Diffusion ModelsabstractThe introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusion models by investigating how their audio outputs compare with the training data, similar to how a doctor auscultates a patient by listening to the sounds of their organs. Using text-to-audio latent diffusion models trained on the AudioCaps dataset, we systematically analyze memorization behavior as a function of training set size. We also evaluate different retrieval metrics for evidence of training data memorization, finding the similarity between mel spectrograms to be more robust in detecting matches than learned embedding vectors. In the process of analyzing memorization in audio latent diffusion models, we also discover a large amount of duplicated audio clips within the AudioCaps database. Dimitrios Bralios, Gordon Wichern, François G. Germain, Zexu Pan, Sameer Khurana, Chiori Hori, Jonathan Le Roux |
ICASSP | 3 |
| 2024 | NIIRF: Neural IIR Filter Field for HRTF Upsampling and PersonalizationabstractHead-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimating the magnitude of the HRTF from a given sound source direction, and the magnitude is converted to a finite impulse response (FIR) filter. We propose the neural infinite impulse response filter field (NIIRF) method that instead estimates the coefficients of cascaded IIR filters. IIR filters mimic the modal nature of HRTFs, thus needing fewer coefficients to approximate them well compared to FIR filters. We find that our method can match the performance of existing NF-based methods on multiple datasets, even outperforming them when measurements are sparse. We also explore approaches to personalize the NF to a subject and experimentally find low-rank adaptation to be effective. Yoshiki Masuyama, Gordon Wichern, François G. Germain, Zexu Pan, Sameer Khurana, Chiori Hori, Jonathan Le Roux |
ICASSP | 3 |
| 2024 | NeuroHeed+: Improving Neuro-Steered Speaker Extraction with Joint Auditory Attention DetectionabstractNeuro-steered speaker extraction aims to extract the listener’s brainattended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using electroencephalography (EEG) devices. Though promising, current methods often have a high speaker confusion error, where the interfering speaker is extracted instead of the attended speaker, degrading the listening experience. In this work, we aim to reduce the speaker confusion error in the neuro-steered speaker extraction model through a jointly fine-tuned auxiliary auditory attention detection model. The latter reinforces the consistency between the extracted target speech signal and the EEG representation, and also improves the EEG representation. Experimental results show that the proposed network significantly outperforms the baseline in terms of speaker confusion and overall signal quality in two-talker scenarios. Zexu Pan, Gordon Wichern, François G. Germain, Sameer Khurana, Jonathan Le Roux |
ICASSP | 3 |
| 2024 | Improving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up AugmentationabstractAutomated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a sequence-to-sequence (seq2seq) backbone powered by strong models such as Transformers. Following the macro-trend of applied machine learning research, in this work, we strive to improve the performance of seq2seq AAC models by extensively leveraging pretrained models and large language models (LLMs). Specifically, we utilize BEATS to extract fine-grained audio features. Then, we employ Instructor LLM to fetch text embeddings of captions, and infuse their language-modality knowledge into BEATs audio features via an auxiliary InfoNCE loss function. Moreover, we propose a novel data augmentation method that uses ChatGPT to produce caption mix-ups (i.e., grammatical and compact combinations of two captions) which, together with the corresponding audio mixtures, increase not only the amount but also the complexity and diversity of training data. During inference, we propose to employ nucleus sampling and a hybrid reranking algorithm, which has not been explored in AAC research. Combining our efforts, our model achieves a new state-of-the-art 32.6 SPIDEr-FL score on the Clotho evaluation split, and wins the 2023 DCASE AAC challenge. Shih-Lun Wu, Xuankai Chang, Gordon Wichern, Jee-Weon Jung, François G. Germain, Jonathan Le Roux, Shinji Watanabe 0001 |
ICASSP | 5 |
| 2024 | Sound Event Bounding Boxes
Janek Ebbers, François G. Germain, Gordon Wichern, Jonathan Le Roux |
INTERSPEECH | 2 |
| 2024 | PARIS: Pseudo-AutoRegressIve Siamese Training for Online Speech Separation
Zexu Pan, Gordon Wichern, François G. Germain, Kohei Saijo, Jonathan Le Roux |
INTERSPEECH | 3 |
| 2024 | Enhanced Reverberation as Supervision for Unsupervised Speech Separation
Kohei Saijo, Gordon Wichern, François G. Germain, Zexu Pan, Jonathan Le Roux |
INTERSPEECH | 3 |
| 2023 | Scenario-Aware Audio-Visual TF-Gridnet for Target Speech ExtractionabstractTarget speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. Building upon the achievements of the state-of-the-art (SOTA) time-frequency speaker separation model TF-GridNet, we propose AV-GridNet, a visual-grounded variant that incorporates the face recording of a target speaker as a conditioning factor during the extraction process. Recognizing the inherent dissimilarities between speech and noise signals as interfering sources, we also propose SAV-GridNet, a scenario-aware model that identifies the type of interfering scenario first and then applies a dedicated expert model trained specifically for that scenario. Our proposed model achieves SOTA results on the second COG-MHEAR Audio-Visual Speech Enhancement Challenge, outperforming other models by a significant margin, objectively and in a listening test. We also perform an extensive analysis of the results under the two scenarios. Zexu Pan, Gordon Wichern, Yoshiki Masuyama, François G. Germain, Sameer Khurana, Chiori Hori, Jonathan Le Roux |
ASRU | 4 |
| 2023 | Cold Diffusion for Speech EnhancementabstractDiffusion models have recently shown promising results for difficult enhancement tasks such as the conditional and unconditional restoration of natural images and audio signals. In this work, we explore the possibility of leveraging a recently proposed advanced iterative diffusion model, namely cold diffusion, to recover clean speech signals from noisy signals. The unique mathematical properties of the sampling process from cold diffusion could be utilized to restore high-quality samples from arbitrary degradations. Based on these properties, we propose an improved training algorithm and objective to help the model generalize better during the sampling process. We verify our proposed framework by investigating two model architectures. Experimental results on benchmark speech enhancement dataset VoiceBank-DEMAND demonstrate the strong performance of the proposed approach compared to representative discriminative models and diffusion-based enhancement models. Hao Yen, François G. Germain, Gordon Wichern, Jonathan Le Roux |
ICASSP | 2 |
| 2021 | Practical Virtual Analog Modeling Using MÖbius TransformsabstractMÖbius transforms provide for the definition of a family of one-step discretization methods offering a framework for alleviating well-known limitations of common one-step methods, such as the trapezoidal method, at no cost in model compactness or complexity. In this paper, we extend the existing theory around these methods. Here, we show how it can be applied to common frameworks used to structure virtual analog models. Then, we propose practical strategies to tune the transform parameters for best simulation results. Finally, we show how such strategies enable us to formulate much improved non-oversampled virtual analog models for several historical audio circuits. François G. Germain |
DAFx | 1 |
| 2019 | Speech Denoising with Deep Feature LossesabstractWe present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed signal that contains only the speech content. Recent approaches have shown promising results using various deep network architectures. In this paper, we propose to train a fully-convolutional context aggregation network using a deep feature loss. That loss is based on comparing the internal feature activations in a different network, trained for acoustic environment detection and domestic audio tagging. Our approach outperforms the state-of-the-art in objective speech quality metrics and in large-scale perceptual experiments with human listeners. It also outperforms an identical network trained using traditional regression losses. The advantage of the new approach is particularly pronounced for the hardest data with the most intrusive background noise, for which denoising is most needed and most challenging. François G. Germain, Qifeng Chen 0001, Vladlen Koltun |
INTERSPEECH | 1 |
| 2016 | Equalization matching of speech recordings in real-world environmentsabstractWhen different parts of speech content such as voice-overs and narration are recorded in real-world environments with different acoustic properties and background noise, the difference in sound quality between the recordings is typically quite audible and therefore undesirable. We propose an algorithm to equalize multiple such speech recordings so that they sound like they were recorded in the same environment. As the timbral content of the speech and background noise typically differ considerably, a simple equalization matching results in a noticeable mismatch in the output signals. A single equalization filter affects both timbres equally and thus cannot disambiguate the competing matching equations of each source. We propose leveraging speech enhancement methods in order to separate speech and background noise, independently apply equalization filtering to each source, and recombine the outputs. By independently equalizing the separated sources, our method is able to better disambiguate the matching equations associated with each source. Therefore the resulting matched signals are perceptually very similar. Additionally, by retaining the background noise in the final output signals, most artifacts from speech enhancement methods are considerably reduced and in general perceptually masked. Subjective listening tests show that our approach significantly outperforms simple equalization matching. François G. Germain, Gautham J. Mysore, Takako Fujioka |
ICASSP | 1 |
| 2015 | Speaker and noise independent online single-channel speech enhancementabstractDesirable properties of real-world speech enhancement methods include online operation, single-channel operation, operation in the presence of a variety of noise types including non-stationary noise, and no requirement for isolated training examples of the specific speaker and noise type at hand. Methods in the literature typically possess only a subset of these properties. Source separation methods particularly rarely simultaneously possess the first and last properties. We extend universal speech model-based speech enhancement to adaptively learn a noise model in an online fashion. We learn a model from a general corpus of speech in place of speaker-dependent training examples before deployment. This setup provides all of these desirable properties, making it easy to deploy in real-world systems without the need to provide additional training examples, while explicitly modeling speech. Our experimental results show that our method achieves the same performance as in the case in which speaker-dependent training data is available. François G. Germain, Gautham J. Mysore |
ICASSP | 1 |
| 2014 | Stopping Criteria for Non-Negative Matrix Factorization Based Supervised and Semi-Supervised Source SeparationabstractNumerous audio signal processing and analysis techniques using non-negative matrix factorization (NMF) have been developed in the past decade, particularly for the task of source separation. NMF-based algorithms iteratively optimize a cost function. However, the correlation between cost functions and application-dependent performance metrics is less known. Furthermore, to the best of our knowledge, no formal heuristic to compute a stopping criterion tailored to a given application exists in the literature. In this paper, we examine this problem for the case of supervised and semi-supervised NMF-based source separation and show that iterating these algorithms to convergence is not optimal for this application. We propose several heuristic stopping criteria that we empirically found to be well correlated with source separation performance. Moreover, our results suggest that simply integrating the learning of an appropriate stopping criterion in a sweep for model size selection could lead to substantial performance improvements with minimal additional effort. François G. Germain, Gautham J. Mysore |
IEEE Signal Process. Lett. | 1 |
| 2013 | Speaker and noise independent voice activity detectionabstractVoice activity detection (VAD) in the presence of heavy, nonstationary noise is a challenging problem that has attracted attention in recent years. Most modern VAD systems require training on highly specialized data: either labeled mixtures of speech and noise that are matched to the application, or, at the very least, noise data similar to that encountered in the application. Because obtaining labeled data can be a laborious task in practical applications, it is desirable for a voice activity detector to be able to perform well in the presence of any type of noise without the need for matched training data. In this paper, we propose a VAD method based on non-negative matrix factorization. We train a universal speech model from a corpus of clean speech but do not train a noise model. Rather, the universal speech model is sufficient to detect the presence of speech in noisy signals. Our experimental results show that our technique is robust to a variety of non-stationary noises mixed at a wide range of signal-to-noise ratios and significantly outperforms baseline algorithms. François G. Germain, Dennis L. Sun, Gautham J. Mysore |
INTERSPEECH | 1 |