Théo Mariotte

dblp:330/9578 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-2108-101XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Multiple Choice Learning for Efficient Speech Separation with Many Speakers
abstract
Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily solved using Permutation Invariant Training (PIT). In this article, we instead consider using the Multiple Choice Learning (MCL) framework, which was originally introduced to tackle ambiguous tasks. We demonstrate experimentally on the popular WSJ0-mix and LibriMix benchmarks that MCL matches the performances of PIT, while being computationally advantageous. This opens the door to a promising research direction, as MCL can be naturally extended to handle a variable number of speakers, or to tackle speech separation in the unsupervised setting.
David Perera, François Derrida, Théo Mariotte, Gaël Richard, Slim Essid
ICASSP3
2024 Unsupervised multiple domain translation through controlled Disentanglement in variational autoencoder
abstract
Unsupervised Multiple Domain Translation is the task of transforming data from one domain to other domains without having paired data to train the systems. Typically, methods based on Generative Adversarial Networks (GANs) are used to address this task. However, our proposal exclusively relies on a modified version of a Variational Autoencoder. This modification consists of the use of two latent variables disentangled in a controlled way by design. One of this latent variables is imposed to depend exclusively on the domain, while the other one must depend on the rest of the variability factors of the data. Additionally, the conditions imposed over the domain latent variable allow for better control and understanding of the latent space. We empirically demonstrate that our approach works on different vision datasets improving the performance of other well known methods. Finally, we prove that, indeed, one of the latent variables stores all the information related to the domain and the other one hardly contains any domain information.
Antonio Almudévar, Théo Mariotte, Alfonso Ortega Giménez, Marie Tahon
ICASSP2
2024 An Explainable Proxy Model for Multilabel Audio Segmentation
abstract
Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for transparency of decision-making with machine learning. In this paper, we propose an explainable multilabel segmentation model that solves speech activity (SAD), music (MD), noise (ND), and overlapped speech detection (OSD) simultaneously. This proxy uses the non-negative matrix factorization (NMF) to map the embeddings used for the segmentation to the frequency domain. Experiments conducted on two datasets show similar performances as the pre-trained black box model while strong explainable features arise. Specifically, the frequency bins used for the decision can be easily identified at both the segment level (local explanations) and global level (class prototypes).
Théo Mariotte, Antonio Almudévar, Marie Tahon, Alfonso Ortega Giménez
ICASSP1
2024 Predefined Prototypes for Intra-Class Separation and Disentanglement
abstract
International audience
Antonio Almudévar, Théo Mariotte, Alfonso Ortega Giménez, Marie Tahon, Luis Vicente, Antonio Miguel, Eduardo Lleida
INTERSPEECH2
2024 Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
Martin Lebourdais, Théo Mariotte, Antonio Almudévar, Marie Tahon, Alfonso Ortega Giménez
INTERSPEECH2
2024 ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
abstract
Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually capture the audio signal. Beamforming, i.e., spatial filtering, is a common practice to process multi-microphone audio data. However, it often requires an explicit localization of the active source to steer the filter. This paper proposes a self-attention-based algorithm to select the output of a bank of fixed spatial filters. This method serves as a feature extractor for joint Voice Activity (VAD) and Overlapped Speech Detection (OSD). The speaker diarization is then inferred from the detected segments. The approach shows convincing distant VAD, OSD, and SD performance, e.g. 14.5% DER on the AISHELL-4 dataset. The analysis of the self-attention weights demonstrates their explainability, as they correlate with the speaker's angular locations.
Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas
INTERSPEECH1
2024 Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing
abstract
We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversity of the predictions. However, this scheme may converge toward an arbitrarily suboptimal local minimum, due to the greedy nature of WTA. We overcome this limitation using annealing, which enhances the exploration of the hypothesis space during training. We leverage insights from statistical physics and information theory to provide a detailed description of the model training trajectory. Additionally, we validate our algorithm by extensive experiments on synthetic datasets, on the standard UCI benchmark, and on speech separation.
David Perera, Victor Letzelter, Théo Mariotte, Adrien Cortés, Mickaël Chen, Slim Essid, Gaël Richard
NeurIPS3
2024 Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
abstract
Voice Activity Detection (VAD) and Overlapped Speech Detection (OSD) are key pre-processing tasks for speaker diarization. In the meeting context, it is often easier to capture speech with a distant device. This consideration however leads to severe performance degradation. We study a unified supervised learning framework to solve distant multi-microphone joint VAD and OSD (VAD+OSD). This paper investigates various multi-channel VAD+OSD front-ends that weight and combine incoming channels. We propose three algorithms based on the Self-Attention Channel Combinator (SACC), previously proposed in the literature. Experiments conducted on the AMI meeting corpus exhibit that channel combination approaches bring significant VAD+OSD improvements in the distant speech scenario. Specifically, we explore the use of learned complex combination weights and demonstrate the benefits of such an approach in terms of explainability. Channel combination-based VAD+OSD systems are evaluated on the final back-end task, i.e. speaker diarization, and show significant improvements. Finally, since multi-channel systems are trained given a fixed array configuration, they may fail in generalizing to other array set-ups, e.g. mismatched number of microphones. A channel-number invariant loss is proposed to learn a unique feature representation regardless of the number of available microphones. The evaluation conducted on mismatched array configurations highlights the robustness of this training strategy.
Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Multi-microphone Automatic Speech Segmentation in Meetings Based on Circular Harmonics Features
abstract
Speaker diarization is the task of answering Who spoke and when? in an audio stream.Pipeline systems rely on speech segmentation to extract speakers' segments and achieve robust speaker diarization.This paper proposes a common framework to solve three segmentation tasks in the distant speech scenario: Voice Activity Detection (VAD), Overlapped Speech Detection (OSD), and Speaker Change Detection (SCD).In the literature, a few studies investigate the multi-microphone distant speech scenario.In this work, we propose a new set of spatial features based on direction-of-arrival estimations in the circular harmonic domain (CH-DOA).These spatial features are extracted from multi-microphone audio data and combined with standard acoustic features.Experiments on the AMI meeting corpus show that CH-DOA can improve the segmentation while being robust in case of deactivated microphones.
Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas
INTERSPEECH1
2022 Microphone Array Channel Combination Algorithms for Overlapped Speech Detection
abstract
International audience
Théo Mariotte, Anthony Larcher, Silvio Montrésor, Jean-Hugh Thomas
INTERSPEECH1