VLDB 2026 Research / reviewers in the wild / expert
Marvin Lavechin
dblp:252/5332
· DBLP profile ↗
11ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-6005-9368ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelabstractHuman infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction.We present a computational model that addresses the acoustic-to-articulatory mapping problem through self-supervised learning.Our model comprises a feature extractor that transforms speech into latent representations, an inverse model that maps these representations to articulatory parameters, and a synthesizer that generates speech outputs.Experiments conducted in both single-and multi-speaker settings reveal that intermediate layers of a pretrained wav2vec 2.0 model provide optimal representations for articulatory learning, significantly outperforming MFCC features.These representations enable our model to learn articulatory trajectories that correlate with human patterns, discriminate between places of articulation, and produce intelligible speech.Critical to successful articulatory learning are representations that balance phonetic discriminability with speaker invariance -precisely the characteristics of self-supervised representation learning models.Our findings provide computational evidence consistent with developmental theories proposing that perceptual learning of phonetic categories guides articulatory development, offering insights into how infants might acquire speech production capabilities despite the complex mapping problem they face. Marvin Lavechin, Thomas Hueber |
EMNLP | 1 |
| 2025 | Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type ClassifierabstractInternational audience Tarek Kunze, Marianne Métais, Hadrien Titeux, Lucas Elbert, Joseph Coffey, Emmanuel Dupoux, Alejandrina Cristià, Marvin Lavechin |
INTERSPEECH | 8 |
| 2025 | Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to ValidityabstractAudio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and therefore high validity. The sheer volume of resulting data necessitates automated analysis to extract relevant metrics for researchers and clinicians. This paper summarizes collective knowledge on this technique, providing entry points to existing resources. We also highlight various sources of error that threaten the accuracy of automated annotations and the interpretation of resulting metrics. To address this, we propose potential troubleshooting metrics to help users assess data quality. While a fully automated quality control system is not feasible, we outline practical strategies for researchers to improve data collection and contextualize their analyses. Loann Peurey, Marvin Lavechin, Tarek Kunze, Manel Khentout, Lucas Gautheron, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 2 |
| 2023 | Brouhaha: Multi-Task Training for Voice Activity Detection, Speech-to-Noise Ratio, and C50 Room Acoustics EstimationabstractMost automatic speech processing systems register degraded performance when applied to noisy or reverberant speech. But how can one tell whether speech is noisy or reverberant? We propose Brouhaha, a neural network jointly trained to extract speech/non-speech segments, speech-to-noise ratios, and C50 room acoustics from single-channel recordings. Brouhaha is trained using a data-driven approach in which noisy and reverberant audio segments are synthesized. We first evaluate its performance and demonstrate that the proposed multi-task regime is beneficial. We then present two scenarios illustrating how Brouhaha can be used on naturally noisy and reverberant data: 1) to investigate the errors made by a speaker diarization model (pyannote.audio); and 2) to assess the reliability of an automatic speech recognition model (Whisper from OpenAI). Both our pipeline and a pretrained model are open source and shared with the speech community. Marvin Lavechin, Marianne Métais, Hadrien Titeux, Alodie Boissonnet, Jade Copet, Morgane Rivière, Elika Bergelson, Alejandrina Cristià, Emmanuel Dupoux, Hervé Bredin |
ASRU | 1 |
| 2023 | BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language modelsabstractInternational audience Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón, Okko Johannes Räsänen, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 1 |
| 2023 | ProsAudit, a prosodic benchmark for self-supervised speech modelsabstractISSN: 2958-1796 Maureen de Seyssel, Marvin Lavechin, Hadrien Titeux, Arthur Thomas, Gwendal Virlet, Andrea Santos Revilla, Guillaume Wisniewski, Bogdan Ludusan, Emmanuel Dupoux |
INTERSPEECH | 2 |
| 2023 | Measuring Language Development From Child-centered RecordingsabstractStandard ways to measure child language development from spontaneous corpora rely on detailed linguistic descriptions of a language as well as exhaustive transcriptions of the child’s speech, which today can only be done through costly human labor. We tackle both issues by proposing (1) a new language development metric (based on entropy) that does not require linguistic knowledge other than having a corpus of text in the language in question to train a language model, (2) a method to derive this metric directly from speech based on a smaller text-speech parallel corpus. Here, we present descriptive results on an open archive including data from six Englishlearning children as a proof of concept. We document that our entropy metric documents a gradual convergence of children’s speech towards adults’ speech as a function of age, and it also correlates moderately with lexical and morphosyntactic measures derived from morphologically parsed transcriptions.The source code of the experiments is released at https://github.com/yaya-sy/EntropyBasedCLDMetricsIndex Terms: L1 acquisition, child speech, morphosyntax, phonetics, speech technology application Yaya Sy, William Havard, Marvin Lavechin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 3 |
| 2022 | Probing phoneme, language and speaker information in unsupervised speech representationsabstractInternational audience Maureen de Seyssel, Marvin Lavechin, Yossi Adi, Emmanuel Dupoux, Guillaume Wisniewski |
INTERSPEECH | 2 |
| 2020 | Pyannote.Audio: Neural Building Blocks for Speaker DiarizationabstractWe introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines. pyannote.audio also comes with pre-trained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding - reaching state-of-the-art performance for most of them. Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, Marie-Philippe Gill |
ICASSP | 6 |
| 2020 | An Open-Source Voice Type Classifier for Child-Centered Daylong RecordingsabstractInternational audience Marvin Lavechin, Ruben Bousbib, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 1 |
| 2020 | End-to-End Domain-Adversarial Voice Activity DetectionabstractInternational audience Marvin Lavechin, Marie-Philippe Gill, Ruben Bousbib, Hervé Bredin, L. Paola García-Perera |
INTERSPEECH | 1 |