VLDB 2026 Research / reviewers in the wild / expert
Romain Serizel
dblp:95/9192
· DBLP profile ↗
53ranked-venue papers
14as first author
31since 2021 · last 2025
0000-0002-6848-0114ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 9 first-author · 28 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion-based Unsupervised Audio-visual Speech EnhancementabstractThis paper proposes a new unsupervised audiovisual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on corresponding video data to simulate the speech generative distribution. This pre-trained model is then paired with the NMF-based noise model to estimate clean speech iteratively. Specifically, a diffusion-based posterior sampling approach is implemented within the reverse diffusion process, where after each iteration, a speech estimate is obtained and used to update the noise parameters. Experimental results confirm that the proposed AVSE approach not only outperforms its audio-only counterpart but also generalizes better than a recent supervised-generative AVSE method. Additionally, the new inference algorithm offers a better balance between inference speed and performance compared to the previous diffusion-based method. Code and demo available at: https://jeaneudesayilo.github.io/fast_UdiffSE Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda |
ICASSP | 3 |
| 2025 | Energy Consumption Trends in Sound Event Detection SystemsabstractDeep learning systems have become increasingly energy-and computation-intensive, raising concerns about their environmental impact. As organizers of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, we recognize the importance of addressing this issue. For the past three years, we have integrated energy consumption metrics into the evaluation of sound event detection (SED) systems. In this paper, we analyze the impact of this energy criterion on the challenge results and explore the evolution of system complexity and energy consumption over the years. We highlight a shift towards more energy-efficient approaches during training without compromising performance, while the number of operations and the system complexity continue to grow. Through this analysis, we hope to promote more environmentally friendly practices within the SED community. Constance Douwes, Romain Serizel |
ICASSP | 2 |
| 2025 | A decade of DCASE: Achievements, practices, evaluations and future challengesabstractThis paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic Signal Processing area. Its success comes from a combination of factors: the challenge offers a large variety of tasks that are renewed each year; and the workshop offers a channel for dissemination of related work, engaging a young and dynamic community. At the same time, DCASE faces its own challenges, growing and expanding to different areas. One of the core principles of DCASE is open science and reproducibility: publicly available datasets, baseline systems, technical reports and workshop publications. While the DCASE challenge and workshop are independent of IEEE SPS, the challenge receives annual endorsement from the AASP TC, and the DCASE community contributes significantly to the ICASSP flagship conference and the success of SPS in many of its activities. Annamaria Mesaros, Romain Serizel, Toni Heittola, Tuomas Virtanen, Mark D. Plumbley |
ICASSP | 2 |
| 2025 | Latent Watermarking of Audio Generative ModelsabstractThe advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. To address this, watermarking techniques have been recently developed, enabling the detection of content generated by a deployed model. For such techniques to be useful, the watermark must resist typical modifications applied to the model or its outputs. The use case of an open-source model trained on proprietary data is challenging, as post-hoc watermarks can then be trivially removed. In response, we introduce a method that watermarks latent audio generative models by directly watermarking their training data. We show the method to be robust against a broad range of audio edits including filtering, compression or even to changing the model’s decoder, maintaining high detection rates with very few false positives. Interestingly, we show that even fine-tuning the model on another dataset can only significantly lower the detection rate at the cost of degrading the generation performance near the level of re-training the model without the protected training data. Robin San-Roman, Pierre Fernandez, Antoine Deleforge, Yossi Adi, Romain Serizel |
ICASSP | 5 |
| 2025 | Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker EmbeddingsabstractSpeaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding extractors usually struggle with short temporal contexts and overlapping speech, which imposes long-term identity reassignment to exploit longer temporal contexts. However, this increases the probability of tracking system errors, which in turn impacts negatively on identity reassignment. To address this, we propose a Knowledge Distillation (KD) based training approach for short context speaker embedding extraction from two speaker mixtures. We leverage the spatial information of the speaker of interest using beamforming to reduce overlap. We study the feasibility of performing identity reassignment over blocks of fixed size, i.e., blockwise identity reassignment, to go towards a low-latency speaker embedding based tracking system. Results demonstrate that our distilled models are effective at short-context embedding extraction and more robust to overlap. Although, blockwise reassignment results indicate that further work is needed to handle simultaneous speech more effectively. Taous Iatariene, Alexandre Guérin, Romain Serizel |
MMSP | 3 |
| 2025 | Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech EnhancementabstractRecent advances in deep learning have significantly improved multichannel speech enhancement algorithms, yet conventional training loss functions such as the scale-invariant signal-to-distortion ratio (SDR) may fail to preserve fine-grained spectral cues essential for phoneme intelligibility. In this work, we propose perceptually-informed variants of the SDR loss, formulated in the time-frequency domain and modulated by frequency-dependent weighting schemes. These weights are designed to emphasize time-frequency regions where speech is prominent or where the interfering noise is particularly strong. We investigate both fixed and adaptive strategies, including ANSI band-importance weights, spectral magnitude-based weighting, and dynamic weighting based on the relative amount of speech and noise. We train the FaSNet multichannel speech enhancement model using these various losses. Experimental results show that while standard metrics such as the SDR are only marginally improved, their perceptual frequency-weighted counterparts exhibit a more substantial improvement. Besides, spectral and phoneme-level analysis indicates better consonant reconstruction, which points to a better preservation of certain acoustic cues. Nasser-Eddine Monir, Paul Magron, Romain Serizel |
MMSP | 3 |
| 2025 | Posterior Transition Modeling for Unsupervised Diffusion-Based Speech EnhancementabstractWe explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise-perturbed likelihood score, combined with the unconditional score via a trade-off hyperparameter. In this work, we propose two alternative algorithms that directly model the conditional reverse transition distribution of diffusion states. The first method integrates the diffusion prior with the observation model in a principled way, removing the need for hyperparameter tuning. The second defines a diffusion process over the noisy speech itself, yielding a fully tractable and exact likelihood score. Experiments on the WSJ0-QUT and VoiceBank-DEMAND datasets demonstrate improved enhancement metrics and greater robustness to domain shifts compared to both supervised and unsupervised baselines. Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel, Xavier Alameda-Pineda |
IEEE Signal Process. Lett. | 3 |
| 2024 | RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile RobotabstractIn this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. RoboVox is a French corpus recorded by a mobile robot. The files are recorded from different distances under severe acoustical conditions with the presence of several types of noise and reverberation. In addition to noise and reverberation, the robot’s internal noise acts as an extra additive noise. RoboVox can be used for both single-channel and multi-channel speaker recognition. In the evaluation protocols, we are considering both cases. The obtained results demonstrate a significant decline in performance in far-filed speaker recognition and urge the community to further research in this domain Mohammad MohammadAmini, Driss Matrouf, Mickael Rouvier, Jean-François Bonastre, Romain Serizel, Théophile Gonos |
LREC/COLING | 5 |
| 2024 | Diffusion-Based Speech Enhancement with a Weighted Generative-Supervised Learning LossabstractDiffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian noise, usually centered on noisy speech, and subsequently learn a parameterized model to reverse this process, conditionally on noisy speech. Unlike supervised methods, generative-based SE approaches often rely solely on an unsupervised loss, which may result in less efficient incorporation of conditioned noisy speech. To address this issue, we propose augmenting the original diffusion training objective with an ℓ2loss, measuring the discrepancy between ground-truth clean speech and its estimation at each diffusion time-step. Experimental results demonstrate the effectiveness of our proposed methodology. Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel |
ICASSP | 3 |
| 2024 | A Weighted-Variance Variational Autoencoder Model for Speech EnhancementabstractWe address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed for the generative model, where the speech information is encoded in the variance as a function of a latent variable. In contrast to this commonly used approach, we propose a weighted variance generative model, where the contribution of each spectrogram time-frame in parameter learning is weighted. We impose a Gamma prior distribution on the weights, which would effectively lead to a Student’s t-distribution instead of Gaussian for speech generative modeling. We develop efficient training and speech enhancement algorithms based on the proposed generative model. Our experimental results on spectrogram auto-encoding and speech enhancement demonstrate the effectiveness and robustness of the proposed approach compared to the standard unweighted variance model. Ali Golmakani, Mostafa Sadeghi, Xavier Alameda-Pineda, Romain Serizel |
ICASSP | 4 |
| 2024 | Unsupervised Speech Enhancement with Diffusion-Based Generative ModelsabstractRecently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when generalising to unseen conditions. To address this issue, we introduce an alternative approach that operates in an unsupervised manner, leveraging the generative power of diffusion models. Specifically, in a training phase, a clean speech prior distribution is learnt in the short-time Fourier transform (STFT) domain using score-based diffusion models, allowing it to unconditionally generate clean speech from Gaussian noise. Then, we develop a posterior sampling methodology for speech enhancement by combining the learnt clean speech prior with a noise model for speech signal inference. The noise parameters are simultaneously learnt along with clean speech estimation through an iterative expectation-maximisation (EM) approach. To the best of our knowledge, this is the first work exploring diffusion-based generative models for unsupervised speech enhancement, demonstrating promising results compared to a recent variational auto-encoder (VAE)-based unsupervised approach and a state-of-the-art diffusion-based supervised method. It thus opens a new direction for future research in unsupervised speech enhancement. Berné Nortier, Mostafa Sadeghi, Romain Serizel |
ICASSP | 3 |
| 2024 | Performance and Energy Balance: A Comprehensive Study of State-of-the-Art Sound Event Detection SystemsabstractIn recent years, deep learning systems have shown a concerning trend toward increased complexity and higher energy consumption. As researchers in this domain and organizers of one of the Detection and Classification of Acoustic Scenes and Events challenges task, we recognize the importance of addressing the environmental impact of data-driven SED systems. In this paper, we propose an analysis focused on SED systems based on the challenge submissions. This includes a comparison across the past two years and a detailed analysis of this year’s SED systems. Through this research, we aim to explore how the SED systems are evolving every year in relation to their energy efficiency implications1. Francesca Ronchini, Romain Serizel |
ICASSP | 2 |
| 2024 | Posterior Sampling Algorithms for Unsupervised Speech Enhancement with Recurrent Variational AutoencoderabstractIn this paper, we address the unsupervised speech enhancement problem based on recurrent variational autoencoder (RVAE). This approach offers promising generalization performance over the supervised counterpart. Nevertheless, the involved iterative variational expectation-maximization (VEM) process at test time, which relies on a variational inference method, results in high computational complexity. To tackle this issue, we present efficient sampling techniques based on Langevin dynamics and Metropolis-Hasting algorithms, adapted to the EM-based speech enhancement with RVAE. By directly sampling from the intractable posterior distribution within the EM process, we circumvent the intricacies of variational inference. We conduct a series of experiments, comparing the proposed methods with VEM and a state-of-the-art supervised speech enhancement approach based on diffusion models. The results reveal that our sampling-based algorithms significantly outperform VEM, not only in terms of computational efficiency but also in overall performance. Furthermore, when compared to the supervised baseline, our methods showcase robust generalization performance in mismatched test conditions. Mostafa Sadeghi, Romain Serizel |
ICASSP | 2 |
| 2024 | Multi-Channel Extension of Pre-trained Models for Speaker VerificationabstractInternational audience Ladislav Mosner, Romain Serizel, Lukás Burget, Oldrich Plchot, Emmanuel Vincent 0001, Junyi Peng, Jan Cernocký |
INTERSPEECH | 2 |
| 2023 | Lightweight Annotation and Class Weight Training for Automatic Estimation of Alarm Audibility in NoiseabstractIn an effort to improve occupational health and safety, we recently proposed an approach to assess the audibility of acoustic danger signals. It is based on the use of a binary classifier trained on perceptual data to predict the audibility of acoustic alarms in audio clips. In the present article, we first investigate the impact of label noise in the training data induced by a flexible annotation procedure on the model performance. We show that a lighter annotation procedure at training still allows for reaching close to human performance at test time. Besides, threshold selection is a crucial aspect in our application as it can have a direct impact on user safety. We thus explore class weight to train a model that allows for a more robust decision threshold selection, ensuring a low false positive rate. François Effa, Romain Serizel, Jean-Pierre Arz, Nicolas Grimault |
ICASSP | 2 |
| 2023 | Audio-Visual Speech Enhancement with a Deep Kalman Filter Generative ModelabstractDeep latent variable generative models based on variational autoencoder (VAE) have shown promising performance for audio-visual speech enhancement (AVSE). The underlying idea is to learn a VAE-based audio-visual prior distribution for clean speech data, and then combine it with a statistical noise model to recover a speech signal from a noisy audio recording and video (lip images) of the target speaker. Existing generative models developed for AVSE do not take into account the sequential nature of speech data, which prevents them from fully incorporating the power of visual data. In this paper, we present an audio-visual deep Kalman filter (AV-DKF) generative model which assumes a first-order Markov chain model for the latent variables and effectively fuses audio-visual data. Moreover, we develop an efficient inference methodology to estimate speech signals at test time. We conduct a set of experiments to compare different variants of generative models for speech enhancement. The results demonstrate the superiority of the AV-DKF model compared with both its audio-only version and the non-sequential audio-only and audio-visual VAE-based models. Ali Golmakani, Mostafa Sadeghi, Romain Serizel |
ICASSP | 3 |
| 2023 | Spice+: Evaluation of Automatic Audio Captioning Systems with Pre-Trained Language ModelsabstractAudio captioning aims at describing acoustic scenes with natural language. Systems are currently evaluated by image captioning metrics CIDEr and SPICE. However, recent studies have highlighted a poor correlation of these metrics with human assessments. In this paper, we propose SPICE+, a modification of SPICE that improves caption annotation and comparison with pre-trained language models. The metric parses captions to semantic graphs with a deep dependency annotation model and a refined set of linguistic rules, then compares sentence embeddings of candidate and reference semantic elements. We formulate a score for general-purpose captioning evaluation, that can be tailored to more specific applications. Combined with fluency error detection, the metric achieves competitive performance on the FENSE benchmark, with 84.0% accuracy on AudioCaps and 74.1% on Clotho. Further experiments show that the metric behaves similarly to the full sentence embedding similarity, while the decomposition into semantic elements allows better interpretability of scores and can provide additional information on the properties of captioning systems. Félix Gontier, Romain Serizel, Christophe Cerisara |
ICASSP | 2 |
| 2023 | Fast and Efficient Speech Enhancement with Variational AutoencodersabstractUnsupervised speech enhancement based on variational autoencoders has shown promising performance compared with the commonly used supervised methods. This approach involves the use of a pre-trained deep speech prior along with a parametric noise model, where the noise parameters are learned from the noisy speech signal with an expectation-maximization (EM)-based method. The E-step involves an intractable latent posterior distribution. Existing algorithms to solve this step are either based on computationally heavy Monte Carlo Markov Chain sampling methods and variational inference, or inefficient optimization-based methods. In this paper, we propose a new approach based on Langevin dynamics that generates multiple sequences of samples and comes with a total variation-based regularization to incorporate temporal correlations of latent vectors. Our experiments demonstrate that the developed framework makes an effective compromise between computational efficiency and enhancement quality, and outperforms existing methods. Mostafa Sadeghi, Romain Serizel |
ICASSP | 2 |
| 2023 | Performance Above All? Energy Consumption vs. Performance, a Study on Sound Event Detection with Heterogeneous DataabstractInternational audience Romain Serizel, Samuele Cornell, Nicolas Turpault |
ICASSP | 1 |
| 2023 | Self-supervised learning with Diffusion-based multichannel speech enhancement for speaker verification under noisy conditionsabstractProceedings of Interspeech 2023 Sandipana Dowerah, Ajinkya Kulkarni, Romain Serizel, Denis Jouvet |
INTERSPEECH | 3 |
| 2023 | From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionabstractDeep generative models can generate high-fidelity audio conditioned on various
types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients
(MFCC)). Recently, such models have been used to synthesize audio
waveforms conditioned on highly compressed representations. Although such
methods produce impressive results, they are prone to generate audible artifacts
when the conditioning is flawed or imperfect. An alternative modeling approach is
to use diffusion models. However, these have mainly been used as speech vocoders
(i.e., conditioned on mel-spectrograms) or generating relatively low sampling
rate signals. In this work, we propose a high-fidelity multi-band diffusion-based
framework that generates any type of audio modality (e.g., speech, music, environmental
sounds) from low-bitrate discrete representations. At equal bit rate,
the proposed approach outperforms state-of-the-art generative techniques in terms
of perceptual quality. Training and evaluation code are available on the facebookresearch/
audiocraft github project. Samples are available on the following
link (https://ai.honu.io/papers/mbd/). Robin San-Roman, Yossi Adi, Antoine Deleforge, Romain Serizel, Gabriel Synnaeve, Alexandre Défossez |
NeurIPS | 4 |
| 2022 | Threshold Independent Evaluation of Sound Event Detection ScoresabstractPerforming an adequate evaluation of sound event detection (SED) systems is far from trivial and is still subject to ongoing research. The recently proposed polyphonic sound detection (PSD)-receiver operating characteristic (ROC) and PSD score (PSDS) make an important step into the direction of an evaluation of SED systems which is independent from a certain decision threshold. This allows to obtain a more complete picture of the overall system behavior which is less biased by threshold tuning. Yet, the PSD-ROC is currently only approximated using a finite set of thresholds. The choice of the thresholds used in approximation, however, can have a severe impact on the resulting PSDS. In this paper we propose a method which allows for computing system performance on an evaluation set for all possible thresholds jointly, enabling accurate computation not only of the PSD-ROC and PSDS but also of other collar-based and intersection-based performance curves. It further allows to select the threshold which best fulfills the requirements of a given application. Source code is publicly available in our SED evaluation package sed_scores_eval1. Janek Ebbers, Reinhold Häb-Umbach, Romain Serizel |
ICASSP | 3 |
| 2022 | A Benchmark of State-of-the-Art Sound Event Detection Systems Evaluated on Synthetic SoundscapesabstractThis paper proposes a benchmark of submissions to Detection and Classification Acoustic Scene and Events 2021 Challenge (DCASE) Task 4 representing a sampling of the state-of-the-art in Sound Event Detection task. The submissions are evaluated according to the two polyphonic sound detection score scenarios proposed for the DCASE 2021 Challenge Task 4, which allow to make an analysis on whether submissions are designed to perform fine-grained temporal segmentation, coarse-grained temporal segmentation, or have been designed to be polyvalent on the scenarios proposed.We study the solutions proposed by participants to analyze their robustness to varying level target to non-target signal-to-noise ratio and to temporal localization of target sound events. A last experiment is proposed in order to study the impact of non-target events on systems outputs. Results show that systems adapted to provide coarse segmentation outputs are more robust to different target to non-target signal-to-noise ratio and, with the help of specific data augmentation methods, they are more robust to time localization of the original event. Results of the last experiment display that systems tend to spuriously predict short events when non-target events are present. This is particularly true for systems that are tailored to have a fine segmentation. Francesca Ronchini, Romain Serizel |
ICASSP | 2 |
| 2022 | Barlow Twins self-supervised learning for robust speaker recognitionabstractInternational audience Mohammad MohammadAmini, Driss Matrouf, Jean-François Bonastre, Sandipana Dowerah, Romain Serizel, Denis Jouvet |
INTERSPEECH | 5 |
| 2022 | Joint Optimization of Diffusion Probabilistic-Based Multichannel Speech Enhancement with Far-Field Speaker VerificationabstractSmart devices using speaker verification are getting equipped with multiple microphones, improving spatial ambiguity and directivity. However, unlike other speech-based applications, the performance of speaker verification degrades in far-field scenarios due to the adverse effects of a noisy environment and room reverberation. This paper presents a novel diffusion probabilistic models-based multichannel speech enhancement as a front-end for the ECAPA-TDNN speaker verification system in a far-field noisy-reverberant scenario. The proposed approach incorporates a two-stage training approach. In the first stage, we individually train the speech enhancement and speaker verification modules. In the second stage, we combined both modules and trained them jointly. We use similarity-preserving knowledge distillation loss that guides the network to produce similar activation for enhanced signals like clean signals. Joint optimization achieved the best results on synthetic and VOiCES datasets. Sandipana Dowerah, Romain Serizel, Denis Jouvet, Mohammad MohammadAmini, Driss Matrouf |
SLT | 2 |
| 2021 | Improving Sound Event Detection Metrics: Insights from DCASE 2020abstractThe ranking of sound event detection (SED) systems may be biased by assumptions inherent to evaluation criteria and to the choice of an operating point. This paper compares conventional event-based and segment-based criteria against the Polyphonic Sound Detection Score (PSDS)'s intersection-based criterion, over a selection of systems from DCASE 2020 Challenge Task 4. It shows that, by relying on collars, the conventional event-based criterion introduces different strictness levels depending on the length of the sound events, and that the segment-based criterion may lack precision and be application dependent. Alternatively, PSDS's intersection-based criterion overcomes the dependency of the evaluation on sound event duration and provides robustness to labelling subjectivity, by allowing valid detections of interrupted events. Furthermore, PSDS enhances the comparison of SED systems by measuring sound event modelling performance independently from the systems' operating points. Giacomo Ferroni, Nicolas Turpault, Juan Azcarreta, Francesco Tuveri, Romain Serizel, Cagdas Bilen, Sacha Krstulovic |
ICASSP | 5 |
| 2021 | Distributed Speech Separation in Spatially Unconstrained Microphone ArraysabstractSpeech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different sources using sophisticated deep neural networks which are very tedious to train. When several microphones are available, spatial information can be exploited to design much simpler algorithms to discriminate speakers. We propose a distributed algorithm that can process spatial information in a spatially unconstrained microphone array. The algorithm relies on a convolutional recurrent neural network that can exploit the signal diversity from the distributed nodes. In a typical case of a meeting room, this algorithm can capture an estimate of each source in a first step and propagate it over the microphone array in order to increase the separation performance in a second step. We show that this approach performs even better when the number of sources and nodes increases. We also study the influence of a mismatch in the number of sources between the training and testing conditions. Nicolas Furnon, Romain Serizel, Irina Illina, Slim Essid |
ICASSP | 2 |
| 2021 | Sound Event Detection and Separation: A Benchmark on Desed Synthetic SoundscapesabstractWe propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of an event and length of clips) and study the impact of non-target sound events and reverberation. We show that temporal localization of sound events remains a challenge for SED systems. We also show that reverberation and non-target sound events severely degrade system performance. In the latter case, sound separation seems like a promising solution. Nicolas Turpault, Romain Serizel, Scott Wisdom, Hakan Erdogan, John R. Hershey, Eduardo Fonseca, Prem Seetharaman, Justin Salamon |
ICASSP | 2 |
| 2021 | What's all the Fuss about Free Universal Sound Separation Data?abstractWe introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtures of one to four sources. To simulate reverberation, an acoustic room simulator is used to generate impulse responses of box-shaped rooms with frequency-dependent reflective walls. Additional open-source data augmentation tools are also provided to produce new mixtures with different combinations of sources and room simulations. Finally, we introduce an open-source baseline separation model, based on an improved time-domain convolutional network (TDCN++), that can separate a variable number of sources in a mixture. This model achieves 9.8 dB of scale-invariant signal-to-noise ratio improvement (SI-SNRi) on mixtures with two to four sources, while reconstructing single-source inputs with 35.8 dB absolute SI-SNR. We hope this dataset will lower the barrier to new research and allow for fast iteration and application of novel techniques from other machine learning domains to the sound separation challenge. Scott Wisdom, Hakan Erdogan, Daniel P. W. Ellis, Romain Serizel, Nicolas Turpault, Eduardo Fonseca, Justin Salamon, Prem Seetharaman, John R. Hershey |
ICASSP | 4 |
| 2021 | UIAI System for Short-Duration Speaker Verification Challenge 2020abstractIn this work, we present the system description of the UIAI entry for the short-duration speaker verification (SdSV) challenge 2020. Our focus is on Task 1 dedicated to text-dependent speaker verification. We investigate different feature extraction and modeling approaches for automatic speaker verification (ASV) and utterance verification (UV). We have also studied different fusion strategies for combining UV and ASV modules. Our primary submission to the challenge is the fusion of seven subsystems which yields a normalized minimum detection cost function (minDCF) of 0.072 and an equal error rate (EER) of 2.14% on the evaluation set. The single system consisting of a pass-phrase identification based model with phone-discriminative bottleneck features gives a normalized minDCF of 0.118 and achieves 19% relative improvement over the state-of-the-art challenge baseline. Md. Sahidullah, Achintya Kumar Sarkar, Ville Vestman, Xuechen Liu 0001, Romain Serizel, Tomi Kinnunen, Zheng-Hua Tan, Emmanuel Vincent 0001 |
SLT | 5 |
| 2021 | DNN-Based Mask Estimation for Distributed Speech Enhancement in Spatially Unconstrained Microphone ArraysabstractDeep neural network (DNN)-based speech enhancement algorithms in microphone arrays have now proven to be efficient solutions to speech understanding and speech recognition in noisy environments. However, in the context of ad-hoc microphone arrays, many challenges remain and raise the need for distributed processing. In this paper, we propose to extend a previously introduced distributed DNN-based time-frequency mask estimation scheme that can efficiently use spatial information in form of so-called compressed signals which are pre-filtered target estimations. We study the performance of this algorithm named Tango under realistic acoustic conditions and investigate practical aspects of its optimal application. We show that the nodes in the microphone array cooperate by taking profit of their spatial coverage in the room. We also propose to use the compressed signals not only to convey the target estimation but also the noise estimation in order to exploit the acoustic diversity recorded throughout the microphone array. Nicolas Furnon, Romain Serizel, Slim Essid, Irina Illina |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | DNN-based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone ArraysabstractMultichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions in the real world. Distributed sensor arrays that consider several devices with a few microphones is a viable solution which allows for exploiting the multiple devices equipped with microphones that we are using in our everyday life. In this context, we propose to extend the distributed adaptive node-specific signal estimation approach to a neural network framework. At each node, a local filtering is performed to send one signal to the other nodes where a mask is estimated by a neural network in order to compute a global multichannel Wiener filter. In an array of two nodes, we show that this additional signal can be leveraged to predict the masks and leads to better speech enhancement performance than when the mask estimation relies only on the local signals. Nicolas Furnon, Romain Serizel, Irina Illina, Slim Essid |
ICASSP | 2 |
| 2020 | Sound Event Detection in Synthetic Domestic EnvironmentsabstractWe present a comparative analysis of the performance of state-of-the-art sound event detection systems. In particular, we study the robustness of the systems to noise and signal degradation, which is known to impact model generalization. Our analysis is based on the results of task 4 of the DCASE 2019 challenge, where submitted systems were evaluated on, in addition to real-world recordings, a series of synthetic soundscapes that allow us to carefully control for different soundscape characteristics. Our results show that while overall systems exhibit significant improvements compared to previous work, they still suffer from biases that could prevent them from generalizing to real-world scenarios. Romain Serizel, Nicolas Turpault, Ankit Parag Shah, Justin Salamon |
ICASSP | 1 |
| 2020 | Limitations of Weak Labels for Embedding and TaggingabstractMany datasets and approaches in ambient sound analysis use weakly labeled data. Weak labels are employed because annotating every data sample with a strong label is too expensive. Yet, their impact on the performance in comparison to strong labels remains unclear. Indeed, weak labels must often be dealt with at the same time as other challenges, namely multiple labels per sample, unbalanced classes and/or overlapping events. In this paper, we formulate a supervised learning problem which involves weak labels. We create a dataset that focuses on the difference between strong and weak labels as opposed to other challenges. We investigate the impact of weak labels when training an embedding or an end-to-end classifier. Different experimental scenarios are discussed to provide insights into which applications are most sensitive to weakly labeled data. Nicolas Turpault, Romain Serizel, Emmanuel Vincent 0001 |
ICASSP | 2 |
| 2020 | Joint NN-Supported Multichannel Reduction of Acoustic Echo, Reverberation and NoiseabstractWe consider the problem of simultaneous reduction of acoustic echo, reverberation and noise. In real scenarios, these distortion sources may occur simultaneously and reducing them implies combining the corresponding distortion-specific filters. As these filters interact with each other, they must be jointly optimized. We propose to model the target and residual signals after linear echo cancellation and dereverberation using a multichannel Gaussian modeling framework and to jointly represent their spectra by means of a neural network. We develop an iterative block-coordinate ascent algorithm to update all the filters. We evaluate our system on real recordings of acoustic echo, reverberation and noise acquired with a smart speaker in various situations. The proposed approach outperforms in terms of overall distortion a cascade of the individual approaches and a joint reduction approach which does not rely on a spectral model of the target and residual signals. Guillaume Carbajal, Romain Serizel, Emmanuel Vincent 0001, Eric Humbert |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Semi-supervised Triplet Loss Based Learning of Ambient Audio EmbeddingsabstractDeep neural networks are particularly useful to learn relevant representations from data. Recent studies have demonstrated the potential of unsupervised representation learning for ambient sound analysis using various flavors of the triplet loss. They have compared this approach to supervised learning. However, in real situations, it is common to have a small labeled dataset and a large unlabeled one. In this paper, we combine unsupervised and supervised triplet loss based learning into a semi-supervised representation learning approach. We propose two flavors of this approach, whereby the positive samples for those triplets whose anchors are unlabeled are obtained either by applying a transformation to the anchor, or by selecting the nearest sample in the training set. We compare our approach to supervised and unsupervised representation learning as well as the ratio between the amount of labeled and unlabeled data. We evaluate all the above approaches on an audio tagging task using the DCASE 2018 Task 4 dataset, and we show the impact of this ratio on the tagging performance. Nicolas Turpault, Romain Serizel, Emmanuel Vincent 0001 |
ICASSP | 2 |
| 2018 | Multiple-Input Neural Network-Based Residual Echo SuppressionabstractA residual echo suppressor (RES) aims to suppress the residual echo in the output of an acoustic echo canceler (AEC). Spectral-based RES approaches typically estimate the magnitude spectra of the near-end speech and the residual echo from a single input, that is either the far-end speech or the echo computed by the AEC, and derive the RES filter coefficients accordingly. These single inputs do not always suffice to discriminate the near-end speech from the remaining echo. In this paper, we propose a neural network-based approach that directly estimates the RES filter coefficients from multiple inputs, including the AEC output, the far-end speech, and/or the echo computed by the AEC. We evaluate our system on real recordings of acoustic echo and near-end speech acquired in various situations with a smart speaker. We compare it to two single-input spectral-based approaches in terms of echo reduction and near-end speech distortion. Guillaume Carbajal, Romain Serizel, Emmanuel Vincent 0001, Eric Humbert |
ICASSP | 2 |
| 2018 | Multichannel Speech Separation with Recurrent Neural Networks from High-Order Ambisonics RecordingsabstractWe present a source separation system for high-order ambisonics (HOA) contents. We derive a multichannel spatial filter from a mask estimated by a long short-term memory (LSTM) recurrent neural network. We combine one channel of the mixture with the outputs of basic HOA beamformers as inputs to the LSTM, assuming that we know the directions of arrival of the directional sources. In our experiments, the speech of interest can be corrupted either by diffuse noise or by an equally loud competing speaker. We show that adding as input the output of the beamformer steered toward the competing speech in addition to that of the beamformer steered toward the target speech brings significant improvements in terms of word error rate. Lauréline Perotin, Romain Serizel, Emmanuel Vincent 0001, Alexandre Guérin |
ICASSP | 2 |
| 2018 | Rank-1 constrained Multichannel Wiener Filter for speech recognition in noisy environments
Emmanuel Vincent 0001, Romain Serizel, Yonghong Yan 0002 |
Comput. Speech Lang. | 3 |
| 2017 | Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identificationabstractThis paper presents supervised feature learning approaches for speaker identification that rely on nonnegative matrix factorisation. Recent studies have shown that group nonnegative matrix factorisation and task-driven supervised dictionary learning can help performing effective feature learning for audio classification problems. This paper proposes to integrate a recent method that relies on group nonnegative matrix factorisation into a task-driven supervised framework for speaker identification. The goal is to capture both the speaker variability and the session variability while exploiting the discriminative learning aspect of the task-driven approach. Results on a subset of the ESTER corpus prove that the proposed approach can be competitive with I-vectors. Romain Serizel, Victor Bisot, Slim Essid, Gaël Richard |
ICASSP | 1 |
| 2017 | Deep-neural network approaches for speech recognition with heterogeneous groups of speakers including childrenabstractAbstract This paper introduces deep neural network (DNN)–hidden Markov model (HMM)-based methods to tackle speech recognition in heterogeneous groups of speakers including children. We target three speaker groups consisting of children, adult males and adult females. Two different kind of approaches are introduced here: approaches based on DNN adaptation and approaches relying on vocal-tract length normalisation (VTLN). First, the recent approach that consists in adapting a general DNN to domain/language specific data is extended to target age/gender groups in the context of DNN–HMM. Then, VTLN is investigated by training a DNN–HMM system by using either mel frequency cepstral coefficients normalised with standard VTLN or mel frequency cepstral coefficients derived acoustic features combined with the posterior probabilities of the VTLN warping factors. In this later, novel, approach the posterior probabilities of the warping factors are obtained with a separate DNN and the decoding can be operated in a single pass when the VTLN approach requires two decoding passes. Finally, the different approaches presented here are combined to take advantage of their complementarity. The combination of several approaches is shown to improve the baseline phone error rate performance by thirty per cent to thirty-five per cent relative and the baseline word error rate performance by about ten per cent relative. Romain Serizel, Diego Giuliani |
Nat. Lang. Eng. | 1 |
| 2017 | Feature Learning With Matrix Factorization Applied to Acoustic Scene ClassificationabstractIn this paper, we study the usefulness of various matrix factorization methods for learning features to be used for the specific acoustic scene classification (ASC) problem. A common way of addressing ASC has been to engineer features capable of capturing the specificities of acoustic environments. Instead, we show that better representations of the scenes can be automatically learned from time–frequency representations using matrix factorization techniques. We mainly focus on extensions including sparse, kernel-based, convolutive and a novel supervised dictionary learning variant of principal component analysis and nonnegative matrix factorization. An experimental evaluation is performed on two of the largest ASC datasets available in order to compare and discuss the usefulness of these methods for the task. We show that the unsupervised learning methods provide better representations of acoustic scenes than the best conventional hand-crafted features on both datasets. Furthermore, the introduction of a novel nonnegative supervised matrix factorization model and deep neural networks trained on spectrograms, allow us to reach further improvements. Victor Bisot, Romain Serizel, Slim Essid, Gaël Richard |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Acoustic scene classification with matrix factorization for unsupervised feature learningabstractIn this paper we study the use of unsupervised feature learning for acoustic scene classification (ASC). The acoustic environment recordings are represented by time-frequency images from which we learn features in an unsupervised manner. After a set of preprocessing and pooling steps, the images are decomposed using matrix factorization methods. By decomposing the data on a learned dictionary, we use the projection coefficients as features for classification. An experimental evaluation is done on a large ASC dataset to study popular matrix factorization methods such as Principal Component Analysis (PCA) and Non-negative Matrix Factorization (NMF) as well as some of their extensions including sparse, kernel based and convolutive variants. The results show the compared variants lead to significant improvement compared to the state-of-the-art results in ASC. Victor Bisot, Romain Serizel, Slim Essid, Gaël Richard |
ICASSP | 2 |
| 2016 | Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identificationabstractThis paper presents a feature learning approach for speaker identification that is based on nonnegative matrix factorisation. Recent studies have shown that with such models, the dictionary atoms can represent well the speaker identity. The approaches proposed so far focused only on speaker variability and not on session variability. However, this later point is a crucial aspect in the success of the I-vector approach that is now the state-of-the-art in speaker identification. This paper proposes a method that relies on group nonnegative matrix factorisation and that is inspired by the I-vector training procedure. By doing so the proposed approach intends to capture both the speaker variability and the session variability. Results on a small corpus prove that the proposed approach can be competitive with I-vectors. Romain Serizel, Slim Essid, Gaël Richard |
ICASSP | 1 |
| 2016 | Machine listening techniques as a complement to video image analysis in forensicsabstractVideo is now one of the major sources of information for forensics. However, video documents can be originating from various recording devices (CCTV, mobile devices, etc.) with inconsistent quality and can sometimes be recorded in challenging light or motion conditions. Therefore, the amount of information that can be extracted relying solely on video image can vary to a great extent. Most of the videos however generally include audio recording as well. Machine listening can then become a valuable complement to video image analysis in challenging scenarios. In this paper, the authors present a brief overview of some machine listening techniques and their application to the analysis of video documents for forensics. The applicability of these techniques to forensics problems is then discussed in the light of machine listening system performances. Romain Serizel, Victor Bisot, Slim Essid, Gaël Richard |
ICIP | 1 |
| 2014 | Vocal tract length normalisation approaches to DNN-based children's and adults' speech recognitionabstractThis paper introduces approaches based on vocal tract length normalisation (VTLN) techniques for hybrid deep neural network (DNN) - hidden Markov model (HMM) automatic speech recognition when targeting children's and adults' speech. VTLN is investigated by training a DNN-HMM system by using first mel frequency cepstral coefficients (MFCCs) normalised with standard VTLN. Then, MFCCs derived acoustic features are combined with the VTLN warping factors to obtain an augmented set of features as input to a DNN. In this later, novel, approach the warping factors are obtained with a separate DNN and the decoding can be operated in a single pass when standard VTLN approach requires two decoding passes. Both VTLN-based approaches are shown to improve phone error rate performance, up to 20% relative improvement, compared to a baseline trained on a mixture of children's and adults' speech. Romain Serizel, Diego Giuliani |
SLT | 1 |
| 2014 | Low-rank Approximation Based Multichannel Wiener Filter Algorithms for Noise Reduction with Application in Cochlear ImplantsabstractThis paper presents low-rank approximation based multichannel Wiener filter algorithms for noise reduction in speech plus noise scenarios, with application in cochlear implants. In a single speech source scenario, the frequency-domain autocorrelation matrix of the speech signal is often assumed to be a rank-1 matrix, which then allows to derive different rank-1 approximation based noise reduction filters. In practice, however, the rank of the autocorrelation matrix of the speech signal is usually greater than one. Firstly, the link between the different rank-1 approximation based noise reduction filters and the original speech distortion weighted multichannel Wiener filter is investigated when the rank of the autocorrelation matrix of the speech signal is indeed greater than one. Secondly, in low input signal-to-noise-ratio scenarios, due to noise non-stationarity, the estimation of the autocorrelation matrix of the speech signal can be problematic and the noise reduction filters can deliver unpredictable noise reduction performance. An eigenvalue decomposition based filter and a generalized eigenvalue decomposition based filter are introduced that include a more robust rank-1, or more generally rank-R, approximation of the autocorrelation matrix of the speech signal. These noise reduction filters are demonstrated to deliver a better noise reduction performance especially in low input signal-to-noise-ratio scenarios. The filters are especially useful in cochlear implants, where more speech distortion and hence a more aggressive noise reduction can be tolerated. Romain Serizel, Marc Moonen, Bas van Dijk, Jan Wouters |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Rank-1 approximation based multichannel wiener filtering algorithms for noise reduction in cochlear implantsabstractThis paper presents multichannel Wiener filtering-based algorithms for noise reduction in cochlear implants. In a single speech scenario, the autocorrelation matrix of the speech signal can be approximated by a rank-1 matrix. It is then possible to derive noise reduction filters that deliver improved signal-to-noise ratio performance. The link between these different filters is investigated here and an eigenvalue decomposition based algorithm is demonstrated to be more stable at low input signal-to-noise ratio compared to previous algorithms. Romain Serizel, Marc Moonen, Bas van Dijk, Jan Wouters |
ICASSP | 1 |
| 2013 | A speech distortion weighting based approach to integrated active noise control and noise reduction in hearing aids
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen |
Signal Process. | 1 |
| 2013 | Binaural Integrated Active Noise Control and Noise Reduction in Hearing AidsabstractThis paper presents a binaural approach to integrated active noise control and noise reduction in hearing aids and aims at demonstrating that a binaural setup indeed provides significant advantages in terms of the number of noise sources that can be compensated for and in terms of the causality margins. Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | A Zone-of-Quiet Based Approach to Integrated Active Noise Control and Noise Reduction for Speech Enhancement in Hearing AidsabstractThis paper focuses on speech enhancement in hearing aids and presents an integrated approach to active noise control and noise reduction which is based on an optimization over a zone-of-quiet generated by the active noise control. A basic integrated active noise control and noise reduction scheme has been introduced previously to tackle secondary path effects and effects of noise leakage through an open fitting. This scheme however, only takes the sound pressure at the ear canal microphone into account. For an integrated active noise control and noise reduction scheme to be efficient, it is desired to achieve active noise control at the eardrum which in practice is away from the ear canal microphone. In some cases, it can also be desired to achieve noise control over a zone not limited to a single point. Two different schemes are presented. The first scheme is based on a mean squared error criterion expressed at a remote point (RP) away from the ear canal microphone and the second scheme is based on an average mean squared error criterion over a desired zone-of-quiet. They are both compared experimentally with the original scheme for both active noise control and integrated active noise control and noise reduction, respectively. The remote-point approach then allows to restore the performance of the original scheme at the desired remote point while the zone-of-quiet approach allows to increase performance up to 3 dB on the desired zone-of-quiet. Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Output SNR analysis of integrated active noise control and noise reduction in hearing aids under a single speech source scenario
Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen |
Signal Process. | 1 |
| 2010 | Integrated Active Noise Control and Noise Reduction in Hearing AidsabstractThis paper presents combined active noise control and noise reduction schemes for hearing aids to tackle secondary path effects and effects of noise leakage through an open fitting. While such leakage contributions and the secondary acoustic path from the loudspeaker to the tympanic membrane are usually not taken into account in standard noise reduction systems, they appear to have a non-negligible impact on the final signal-to-noise ratio. Using a noise-reduction algorithm and an active noise control system in cascade may be efficient as long as the causality margin of the system is large enough. Putting the two functional blocks in parallel and then integrating them is found to lead to a more robust algorithm. A Filtered-x Multichannel Wiener Filter is presented and applied to integrate noise reduction and active noise control. The cascaded scheme and the integrated scheme are compared experimentally with a Multichannel Wiener Filter in a classic noise reduction framework without active noise control, where the integrated scheme is found to provide the best performance. Romain Serizel, Marc Moonen, Jan Wouters, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 1 |