VLDB 2026 Research / reviewers in the wild / expert
Emmanuel Dupoux
dblp:41/8160
· DBLP profile ↗
94ranked-venue papers
2as first author
37since 2021 · last 2026
0000-0002-7814-2952ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 2 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 1 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpidR-Adapt: A Universal Speech Representation Model for Few-Shot AdaptationabstractMahi Luthra, Jiayi Shen, Maxime Poli, Angelo Ortiz Tandazo, Yosuke Higuchi, Youssef Benchekroun, Martin Gleize, Charles-Éric Saint-James, Dongyan Lin, Phillip Rust, Angel Villar-Corrales, Surya, Vanessa Stark, Rashel Moritz, Juan Pino, Yann LeCun, Emmanuel Dupoux. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mahi Luthra, Maxime Poli, Angelo Ortiz Tandazo, Yosuke Higuchi, Youssef Benchekroun, Martin Gleize, Charles-Éric Saint-James, Dongyan Lin, Phillip Rust, Angel Villar-Corrales, Surya Parimi, Vanessa Stark, Rashel Moritz, Juan Pino 0001, Yann LeCun, Emmanuel Dupoux |
ACL (1) | 17 |
| 2026 | MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units DiscoveryabstractThis paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models. Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber, Emmanuel Dupoux |
ACL (1) | 5 |
| 2025 | CASPER: A Large Scale Spontaneous Speech DatasetabstractThe success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community. Cihan Xiao, Ruixing Liang, Xiangyu Zhang 0005, Mehmet Emre Tiryaki, Veronica Bae, Lavanya Shankar, Ethan Poon, Emmanuel Dupoux, Sanjeev Khudanpur, L. Paola García-Perera |
ASRU | 9 |
| 2025 | Frequency & Compositionality in Emergent CommunicationabstractIn natural languages, frequency and compositionality exhibit an inverse relationship: the most frequent words often resist regular patterns, developing idiosyncratic forms.This phenomenon, exemplified by irregular verbs where the most frequent verbs resist regular patterns, raises a compelling question: do artificial communication systems follow similar principles?Through systematic experiments with neural network agents in a referential game setting, and by manipulating input frequency through Zipfian distributions, we investigate if these systems mirror the irregular verbs phenomenon, where messages referring to frequent objects develop less compositional structure than messages referring to rare ones.We establish that compositionality is not an inherent property of the frequency itself and provide compelling evidence that limited data exposure, which frequency distributions naturally create, serves as a fundamental driver for the emergence of compositional structure in communication systems, offering insights into the cognitive and computational pressures that shape linguistic systems. Jean-Baptiste Sevestre, Emmanuel Dupoux |
EMNLP | 2 |
| 2025 | Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type ClassifierabstractInternational audience Tarek Kunze, Marianne Métais, Hadrien Titeux, Lucas Elbert, Joseph Coffey, Emmanuel Dupoux, Alejandrina Cristià, Marvin Lavechin |
INTERSPEECH | 6 |
| 2025 | Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to ValidityabstractAudio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and therefore high validity. The sheer volume of resulting data necessitates automated analysis to extract relevant metrics for researchers and clinicians. This paper summarizes collective knowledge on this technique, providing entry points to existing resources. We also highlight various sources of error that threaten the accuracy of automated annotations and the interpretation of resulting metrics. To address this, we propose potential troubleshooting metrics to help users assess data quality. While a fully automated quality control system is not feasible, we outline practical strategies for researchers to improve data collection and contextualize their analyses. Loann Peurey, Marvin Lavechin, Tarek Kunze, Manel Khentout, Lucas Gautheron, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 6 |
| 2025 | A simple method for predicting Clinical Scores in Huntington's Disease by leveraging ASR's uncertainty on spontaneous speech
Hadrien Titeux, Quang Tuan Rémy Nguyen, Andres Gil-Salcedo, Anne-Catherine Bachoud-Lévi, Emmanuel Dupoux |
INTERSPEECH | 5 |
| 2024 | Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning ApproachabstractRecent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible.Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and nonverbal vocalizations.Modeling directly from speech opens up the path to more natural and expressive systems.On the other hand, speechonly systems require up to three orders of magnitude more data to catch up to their text-based counterparts in terms of their semantic abilities.We show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data. Maxime Poli, Emmanuel Chemla, Emmanuel Dupoux |
EMNLP | 3 |
| 2024 | EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech ModelsabstractWe introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis.We apply this to two tasks: speech resynthesis and speech-to-speech translation.In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language.As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level. Maureen de Seyssel, Antony D'Avirro, Adina Williams, Emmanuel Dupoux |
EMNLP | 4 |
| 2024 | Simulating articulatory trajectories with phonological feature interpolationabstractAs a first step towards a complete computational model of speech learning involving perception-production loops, we investigate the forward mapping between pseudo-motor commands and articulatory trajectories. Two phonological feature sets, based respectively on generative and articulatory phonology, are used to encode a phonetic target sequence. Different interpolation techniques are compared to generate smooth trajectories in these feature spaces, with a potential optimisation of the target value and timing to capture co-articulation effects. We report the Pearson correlation between a linear projection of the generated trajectories and articulatory data derived from a multi-speaker dataset of electromagnetic articulography (EMA) recordings. A correlation of 0.67 is obtained with an extended feature set based on generative phonology and a linear interpolation technique. We discuss the implications of our results for our understanding of the dynamics of biological motion. Angelo Ortiz Tandazo, Thomas Schatz, Thomas Hueber, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2023 | Brouhaha: Multi-Task Training for Voice Activity Detection, Speech-to-Noise Ratio, and C50 Room Acoustics EstimationabstractMost automatic speech processing systems register degraded performance when applied to noisy or reverberant speech. But how can one tell whether speech is noisy or reverberant? We propose Brouhaha, a neural network jointly trained to extract speech/non-speech segments, speech-to-noise ratios, and C50 room acoustics from single-channel recordings. Brouhaha is trained using a data-driven approach in which noisy and reverberant audio segments are synthesized. We first evaluate its performance and demonstrate that the proposed multi-task regime is beneficial. We then present two scenarios illustrating how Brouhaha can be used on naturally noisy and reverberant data: 1) to investigate the errors made by a speaker diarization model (pyannote.audio); and 2) to assess the reliability of an automatic speech recognition model (Whisper from OpenAI). Both our pipeline and a pretrained model are open source and shared with the speech community. Marvin Lavechin, Marianne Métais, Hadrien Titeux, Alodie Boissonnet, Jade Copet, Morgane Rivière, Elika Bergelson, Alejandrina Cristià, Emmanuel Dupoux, Hervé Bredin |
ASRU | 9 |
| 2023 | Generative Spoken Language Model based on continuous word-sized audio tokensabstractIn NLP, text language models based on words or subwords are known to outperform their character-based counterparts.Yet, in the speech community, the standard input of spoken LMs are 20ms or 40ms-long discrete units (shorter than a phoneme).Taking inspiration from wordbased LM, we introduce a Generative Spoken Language Model (GSLM) based on word-size continuous-valued audio embeddings that can generate diverse and expressive language output.This is obtained by replacing lookup table for lexical types with a Lexical Embedding function, the cross entropy loss by a contrastive loss, and multinomial sampling by k-NN sampling.The resulting model is the first generative language model based on word-size continuous embeddings.Its performance is on par with discrete unit GSLMs regarding generation quality as measured by automatic metrics and subjective human judgements.Moreover, it is five times more memory efficient thanks to its large 200ms units.In addition, the embeddings before and after the Lexical Embedder are phonetically and semantically interpretable. 1 Robin Algayres, Yossi Adi, Tu Anh Nguyen, Jade Copet, Gabriel Synnaeve, Benoît Sagot, Emmanuel Dupoux |
EMNLP | 7 |
| 2023 | Do Coarser Units Benefit Cluster Prediction-Based Speech Pre-Training?abstractThe research community has produced many successful self-supervised speech representation learning methods over the past few years. Discrete units have been utilized in various self-supervised learning frameworks, such as VQ-VAE [1], wav2vec 2.0 [2], Hu-BERT [3], and Wav2Seq [4]. This paper studies the impact of altering the granularity and improving the quality of these discrete acoustic units for pre-training encoder-only and encoder-decoder models. We systematically study the current proposals of using Byte-Pair Encoding (BPE) and new extensions that use cluster smoothing and Brown clustering. The quality of learned units is studied intrinsically using zero speech metrics and on the down-stream speech recognition (ASR) task. Our results suggest that longer-range units are helpful for encoder-decoder pre-training; however, encoder-only masked-prediction models cannot yet benefit from self-supervised word-like targets. Ali Elkahky, Wei-Ning Hsu, Paden Tomasello, Tu Anh Nguyen, Robin Algayres, Yossi Adi, Jade Copet, Emmanuel Dupoux, Abdel-rahman Mohamed |
ICASSP | 8 |
| 2023 | Introducing Topography in Convolutional Neural NetworksabstractParts of the brain that carry sensory tasks are organized topographically: nearby neurons are responsive to the same properties of input signals. Thus, in this work, inspired by the neuroscience literature, we proposed a new topographic inductive bias in Convolutional Neural Networks (CNNs). To achieve this, we introduced a new topographic loss and an efficient implementation to topographically organize each convolutional layer of any CNN. We benchmarked our new method on 4 datasets and 3 models in vision and audio tasks and showed equivalent performance to all benchmarks. Besides, we also showcased the generalizability of our topographic loss with how it can be used with different topographic organizations in CNNs. Finally, we demonstrated that adding the topographic inductive bias made CNNs more resistant to pruning. Our approach provides a new avenue to obtain models that are more memory efficient while maintaining better accuracy. Maxime Poli, Emmanuel Dupoux, Rachid Riad |
ICASSP | 2 |
| 2023 | Neural Agents Struggle to Take Turns in Bidirectional Emergent Communication
Valentin Taillandier, Dieuwke Hupkes, Benoît Sagot, Emmanuel Dupoux, Paul Michel |
ICLR | 4 |
| 2023 | Evaluating context-invariance in unsupervised speech representations
Mark Hallap, Emmanuel Dupoux, Ewan Dunbar |
INTERSPEECH | 2 |
| 2023 | BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language modelsabstractInternational audience Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón, Okko Johannes Räsänen, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 7 |
| 2023 | Expresso: A Benchmark and Analysis of Discrete Expressive Speech ResynthesisabstractInternational audience Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro, Bowen Shi 0002, Itai Gat, Maryam Fazel-Zarandi, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, Felix Kreuk, Yossi Adi, Emmanuel Dupoux |
INTERSPEECH | 13 |
| 2023 | ProsAudit, a prosodic benchmark for self-supervised speech modelsabstractISSN: 2958-1796 Maureen de Seyssel, Marvin Lavechin, Hadrien Titeux, Arthur Thomas, Gwendal Virlet, Andrea Santos Revilla, Guillaume Wisniewski, Bogdan Ludusan, Emmanuel Dupoux |
INTERSPEECH | 9 |
| 2023 | Measuring Language Development From Child-centered RecordingsabstractStandard ways to measure child language development from spontaneous corpora rely on detailed linguistic descriptions of a language as well as exhaustive transcriptions of the child’s speech, which today can only be done through costly human labor. We tackle both issues by proposing (1) a new language development metric (based on entropy) that does not require linguistic knowledge other than having a corpus of text in the language in question to train a language model, (2) a method to derive this metric directly from speech based on a smaller text-speech parallel corpus. Here, we present descriptive results on an open archive including data from six Englishlearning children as a proof of concept. We document that our entropy metric documents a gradual convergence of children’s speech towards adults’ speech as a function of age, and it also correlates moderately with lexical and morphosyntactic measures derived from morphologically parsed transcriptions.The source code of the experiments is released at https://github.com/yaya-sy/EntropyBasedCLDMetricsIndex Terms: L1 acquisition, child speech, morphosyntax, phonetics, speech technology application Yaya Sy, William Havard, Marvin Lavechin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 4 |
| 2023 | Textually Pretrained Speech Language ModelsabstractSpeech language models (SpeechLMs) process and generate acoustic data only, without textual supervision. In this work, we propose TWIST, a method for training SpeechLMs using a warm-start from a pretrained textual language models. We show using both automatic and human evaluations that TWIST outperforms a cold-start SpeechLM across the board. We empirically analyze the effect of different model design choices such as the speech tokenizer, the pretrained textual model, and the dataset size. We find that model and dataset scale both play an important role in constructing better-performing SpeechLMs. Based on our observations, we present the largest (to the best of our knowledge) SpeechLM both in terms of number of parameters and training data. We additionally introduce two spoken versions of the StoryCloze textual benchmark to further improve model evaluation and advance future research in the field. We make speech samples, code and models publicly available. Michael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat, Alexis Conneau, Felix Kreuk, Jade Copet, Alexandre Défossez, Gabriel Synnaeve, Emmanuel Dupoux, Roy Schwartz 0001, Yossi Adi |
NeurIPS | 10 |
| 2023 | Generative Spoken Dialogue Language ModelingabstractAbstract We introduce dGSLM, the first “textless” model able to generate audio samples of naturalistic spoken dialogues. It uses recent work on unsupervised spoken unit discovery coupled with a dual-tower transformer architecture with cross-attention trained on 2000 hours of two-channel raw conversational audio (Fisher dataset) without any text or labels. We show that our model is able to generate speech, laughter, and other paralinguistic signals in the two channels simultaneously and reproduces more naturalistic and fluid turn taking compared to a text-based cascaded model.1,2 Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoît Sagot, Abdel-rahman Mohamed, Emmanuel Dupoux |
Trans. Assoc. Comput. Linguistics | 11 |
| 2022 | Text-Free Prosody-Aware Generative Spoken Language ModelingabstractEugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Morgane Riviere, Abdelrahman Mohamed, Emmanuel Dupoux, Wei-Ning Hsu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Eugene Kharitonov, Ann Lee 0001, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Morgane Rivière, Abdel-rahman Mohamed, Emmanuel Dupoux, Wei-Ning Hsu |
ACL (1) | 10 |
| 2022 | Is the Language Familiarity Effect gradual ? A computational modelling approach
Maureen de Seyssel, Guillaume Wisniewski, Emmanuel Dupoux |
CogSci | 3 |
| 2022 | Textless Speech Emotion Conversion using Discrete & Decomposed RepresentationsabstractFelix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu Anh Nguyen, Morgan Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdel-rahman Mohamed, Emmanuel Dupoux, Yossi Adi |
EMNLP | 9 |
| 2022 | On the role of population heterogeneity in emergent communication
Mathieu Rita, Florian Strub, Jean-Bastien Grill, Olivier Pietquin, Emmanuel Dupoux |
ICLR | 5 |
| 2022 | Speech Sequence Embeddings using Nearest Neighbors Contrastive LearningabstractInternational audience Robin Algayres, Adel Nabli, Benoît Sagot, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2022 | Probing phoneme, language and speaker information in unsupervised speech representationsabstractInternational audience Maureen de Seyssel, Marvin Lavechin, Yossi Adi, Emmanuel Dupoux, Guillaume Wisniewski |
INTERSPEECH | 4 |
| 2022 | Emergent Communication: Generalization and Overfitting in Lewis GamesabstractLewis signaling games are a class of simple communication games for simulating the emergence of language. In these games, two agents must agree on a communication protocol in order to solve a cooperative task. Previous work has shown that agents trained to play this game with reinforcement learning tend to develop languages that display undesirable properties from a linguistic point of view (lack of generalization, lack of compositionality, etc). In this paper, we aim to provide better understanding of this phenomenon by analytically studying the learning problem in Lewis games. As a core contribution, we demonstrate that the standard objective in Lewis games can be decomposed in two components: a co-adaptation loss and an information loss. This decomposition enables us to surface two potential sources of overfitting, which we show may undermine the emergence of a structured communication protocol. In particular, when we control for overfitting on the co-adaptation loss, we recover desired properties in the emergent languages: they are more compositional and generalize better. Mathieu Rita, Corentin Tallec, Paul Michel, Jean-Bastien Grill, Olivier Pietquin, Emmanuel Dupoux, Florian Strub |
NeurIPS | 6 |
| 2022 | Stop: A Dataset for Spoken Task Oriented Semantic ParsingabstractEnd-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging acoustic information lost in the intermediate textual representation and preventing cascading errors from Automatic Speech Recognition (ASR). Further, having one unified model has efficiency advantages when deploying assistant systems on-device. However, the limited number of public audio datasets with semantic parse labels hinders the research progress in this area. In this paper, we release the Spoken Task-Oriented semantic Parsing (STOP) dataset1, the largest and most complex SLU dataset publicly available. Additionally, we define low-resource splits to establish a benchmark for improving SLU when limited labeled data is available. Furthermore, in addition to the human-recorded audio, we are releasing a TTS-generated versions to benchmark the performance for low-resource and domain adaptation of end-to-end SLU systems. Paden Tomasello, Akshat Shrivastava, Daniel Lazar, Po-Chun Hsu, Adithya Sagar, Ali Elkahky, Jade Copet, Wei-Ning Hsu, Yossi Adi, Robin Algayres, Tu Anh Nguyen, Emmanuel Dupoux, Luke Zettlemoyer, Abdel-rahman Mohamed |
SLT | 13 |
| 2022 | IntPhys 2019: A Benchmark for Visual Intuitive Physics UnderstandingabstractIn order to reach human performance on complex visual tasks, artificial systems need to incorporate a significant amount of understanding of the world in terms of macroscopic objects, movements, forces, etc. Inspired by work on intuitive physics in infants, we propose an evaluation benchmark which diagnoses how much a given system understands about physics by testing whether it can tell apart well matched videos of possible versus impossible events constructed with a game engine. The test requires systems to compute a physical plausibility score over an entire video. To prevent perceptual biases, the dataset is made of pixel matched quadruplets of videos, enforcing systems to focus on high level temporal dependencies between frames rather than pixel-level details. We then describe two Deep Neural Networks systems aimed at learning intuitive physics in an unsupervised way, using only physically possible videos. The systems are trained with a future semantic mask prediction objective and tested on the possible versus impossible discrimination task. The analysis of their results compared to human data gives novel insights in the potentials and limitations of next frame prediction architectures. Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, Emmanuel Dupoux |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | DP-Parse: Finding Word Boundaries from Raw Speech with an Instance LexiconabstractAbstract Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a ‘space’ delimiter between words. Popular Bayesian non-parametric models for text segmentation (Goldwater et al., 2006, 2009) use a Dirichlet process to jointly segment sentences and build a lexicon of word types. We introduce DP-Parse, which uses similar principles but only relies on an instance lexicon of word tokens, avoiding the clustering errors that arise with a lexicon of word types. On the Zero Resource Speech Benchmark 2017, our model sets a new speech segmentation state-of-the-art in 5 languages. The algorithm monotonically improves with better input representations, achieving yet higher scores when fed with weakly supervised inputs. Despite lacking a type lexicon, DP-Parse can be pipelined to a language model and learn semantic and syntactic representations as assessed by a new spoken word embedding benchmark. 1 Robin Algayres, Tristan Ricoul, Julien Karadayi, Hugo Laurençon, Mohamed Salah Zaïem, Abdel-rahman Mohamed, Benoît Sagot, Emmanuel Dupoux |
Trans. Assoc. Comput. Linguistics | 8 |
| 2021 | VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationabstractChanghan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, Emmanuel Dupoux. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Changhan Wang, Morgane Rivière, Ann Lee 0001, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino 0001, Emmanuel Dupoux |
ACL/IJCNLP (1) | 9 |
| 2021 | The Zero Resource Speech Challenge 2021: Spoken Language ModellingabstractWe present the Zero Resource Speech Challenge 2021, which asks participants to learn a language model directly from audio, without any text or labels. The challenge is based on the Libri-light dataset, which provides up to 60k hours of audio from English audio books without any associated text. We provide a pipeline baseline system consisting on an encoder based on contrastive predictive coding (CPC), a quantizer ($k$-means) and a standard language model (BERT or LSTM). The metrics evaluate the learned representations at the acoustic (ABX discrimination), lexical (spot-the-word), syntactic (acceptability judgment) and semantic levels (similarity judgment). We present an overview of the eight submitted systems from four groups and discuss the main results. Ewan Dunbar, Mathieu Bernard, Nicolas Hamilakis, Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Eugene Kharitonov, Emmanuel Dupoux |
Interspeech | 9 |
| 2021 | Speech Resynthesis from Discrete Disentangled Self-Supervised RepresentationsabstractWe propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker identity. This allows to synthesize speech in a controllable manner. We analyze various state-of-the-art, self-supervised representation learning methods and shed light on the advantages of each method while considering reconstruction quality and disentanglement properties. Specifically, we evaluate the F0 reconstruction, speaker identification performance (for both resynthesis and voice conversion), recordings' intelligibility, and overall quality using subjective human evaluation. Lastly, we demonstrate how these representations can be used for an ultra-lightweight speech codec. Using the obtained representations, we can get to a rate of 365 bits per second while providing better speech quality than the baseline methods. Audio samples can be found under the following link: speechbot.github.io/resynthesis. Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdel-rahman Mohamed, Emmanuel Dupoux |
Interspeech | 8 |
| 2021 | Data Augmenting Contrastive Learning of Speech Representations in the Time DomainabstractContrastive Predictive Coding (CPC), based on predicting future segments of speech from past segments is emerging as a powerful algorithm for representation learning of speech signal. However, it still under-performs compared to other methods on unsupervised evaluation benchmarks. Here, we intro-duce WavAugment, a time-domain data augmentation library which we adapt and optimize for the specificities of CPC (raw waveform input, contrastive loss, past versus future structure). We find that applying augmentation only to the segments from which the CPC prediction is performed yields better results than applying it also to future segments from which the samples (both positive and negative) of the contrastive loss are drawn. After selecting the best combination of pitch modification, additive noise and reverberation on unsupervised metrics on LibriSpeech (with a gain of 18-22% relative on the ABX score), we apply this combination without any change to three new datasets in the Zero Resource Speech Benchmark 2017 and beat the state-of-the-art using out-of-domain training data. Finally, we show that the data-augmented pretrained features improve a downstream phone recognition task in the Libri-light semi-supervised setting (10 min, 1 h or 10 h of labelled data) reducing the PER by 15% relative. Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf, Pierre-Emmanuel Mazaré, Matthijs Douze, Emmanuel Dupoux |
SLT | 7 |
| 2021 | Towards Unsupervised Learning of Speech Features in the WildabstractRecent work on unsupervised contrastive learning of speech representation has shown promising results, but so far has mostly been applied to clean, curated speech datasets. Can it also be used with unprepared audio data "in the wild"? Here, we explore three potential problems in this setting: (i) presence of non-speech data, (ii) noisy or low quality speech data, and (iii) imbalance in speaker distribution. We show that on the Libri-light train set, which is itself a relatively clean speech-only dataset, these problems combined can already have a performance cost of up to 30% relative for the ABX score. We show that the first two problems can be alleviated by data filtering, with voice activity detection selecting speech segments, while perplexity of a model trained with clean data helping to discard entire files. We show that the third problem can be alleviated by learning a speaker embedding in the predictive branch of the model. We show that these techniques build more robust speech features that can be transferred to an ASR task in the low resource setting. Morgane Rivière, Emmanuel Dupoux |
SLT | 2 |
| 2020 | Compositionality and Generalization In Emergent LanguagesabstractNatural language allows us to refer to novel composite concepts by combining expressions denoting their parts according to systematic rules, a property known as \emph{compositionality}. In this paper, we study whether the language emerging in deep multi-agent simulations possesses a similar ability to refer to novel primitive combinations, and whether it accomplishes this feat by strategies akin to human-language compositionality. Equipped with new ways to measure compositionality in emergent languages inspired by disentanglement in representation learning, we establish three main results. First, given sufficiently large input spaces, the emergent language will naturally develop the ability to refer to novel composite concepts. Second, there is no correlation between the degree of compositionality of an emergent language and its ability to generalize. Third, while compositionality is not necessary for generalization, it provides an advantage in terms of language transmission: The more compositional a language is, the more easily it will be picked up by new learners, even when the latter differ in architecture from the original agents. We conclude that compositionality does not arise from simple generalization pressure, but if an emergent language does chance upon it, it will be more likely to survive and thrive. Rahma Chaabouni 0001, Eugene Kharitonov, Diane Bouchacourt, Emmanuel Dupoux, Marco Baroni |
ACL | 4 |
| 2020 | Modelling Perceptual Effects of Phonology with ASR Systems
Bing'er Jiang, Ewan Dunbar, Morgan Sonderegger, Meghan Clayards, Emmanuel Dupoux |
CogSci | 5 |
| 2020 | Does bilingual input hurt? A simulation of language discrimination and clustering using i-vectors
Maureen de Seyssel, Emmanuel Dupoux |
CogSci | 2 |
| 2020 | Analogies minus analogy test: measuring regularities in word embeddingsabstractVector space models of words have long been claimed to capture linguistic regularities as simple vector translations, but problems have been raised with this claim. We decompose and empirically analyze the classic arithmetic word analogy test, to motivate two new metrics that address the issues with the standard test, and which distinguish between class-wise offset concentration (similar directions between pairs of words drawn from different broad classes, such as France--London, China--Ottawa, ...) and pairing consistency (the existence of a regular transformation between correctly-matched pairs such as France:Paris::China:Beijing). We show that, while the standard analogy test is flawed, several popular word embeddings do nevertheless encode linguistic regularities. Louis Fournier, Emmanuel Dupoux, Ewan Dunbar |
CoNLL | 2 |
| 2020 | "LazImpa": Lazy and Impatient neural agents learn to communicate efficientlyabstractPrevious work has shown that artificial neural agents naturally develop surprisingly nonefficient codes.This is illustrated by the fact that in a referential game involving a speaker and a listener neural networks optimizing accurate transmission over a discrete channel, the emergent messages fail to achieve an optimal length.Furthermore, frequent messages tend to be longer than infrequent ones, a pattern contrary to the Zipf Law of Abbreviation (ZLA) observed in all natural languages.Here, we show that near-optimal and ZLA-compatible messages can emerge, but only if both the speaker and the listener are modified.We hence introduce a new communication system, "Laz-Impa", where the speaker is made increasingly lazy, i.e., avoids long messages, and the listener impatient, i.e., seeks to guess the intended content as soon as possible. Mathieu Rita, Rahma Chaabouni 0001, Emmanuel Dupoux |
CoNLL | 3 |
| 2020 | Libri-Light: A Benchmark for ASR with Limited or No SupervisionabstractWe introduce a new collection of spoken English audio suitable for training speech recognition systems under limited or no supervision. It is derived from open-source audio books from the LibriVox project. It contains over 60K hours of audio, which is, to our knowledge, the largest freely-available corpus of speech. The audio has been segmented using voice activity detection and is tagged with SNR, speaker ID and genre descriptions. Additionally, we provide baseline systems and evaluation metrics working under three settings: (1) the zero resource/unsupervised setting (ABX), (2) the semi- supervised setting (PER, CER) and (3) the distant supervision setting (WER). Settings (2) and (3) use limited textual resources (10 minutes to 10 hours) aligned with the speech. Setting (3) uses large amounts of unaligned text. They are evaluated on the standard LibriSpeech dev and test sets for comparison with the supervised state-of-the-art. Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fügen, Tatiana Likhomanenko, Gabriel Synnaeve, Armand Joulin, Abdel-rahman Mohamed, Emmanuel Dupoux |
ICASSP | 15 |
| 2020 | Unsupervised Pretraining Transfers Well Across LanguagesabstractCross-lingual and multi-lingual training of Automatic Speech Recognition (ASR) has been extensively investigated in the supervised setting. This assumes the existence of a parallel corpus of speech and orthographic transcriptions. Recently, contrastive predictive coding (CPC) algorithms have been proposed to pretrain ASR systems with unlabelled data. In this work, we investigate whether unsupervised pretraining transfers well across languages. We show that a slight modification of the CPC pretraining extracts features that transfer well to other languages, being on par or even outperforming supervised pretraining. This shows the potential of unsupervised methods for languages with few linguistic resources. Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, Emmanuel Dupoux |
ICASSP | 4 |
| 2020 | Evaluating the Reliability of Acoustic Speech EmbeddingsabstractInternational audience Robin Algayres, Mohamed Salah Zaïem, Benoît Sagot, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2020 | The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word UnitsabstractInternational audience Ewan Dunbar, Julien Karadayi, Mathieu Bernard, Xuan-Nga Cao, Robin Algayres, Lucas Ondel Yang, Laurent Besacier, Sakriani Sakti, Emmanuel Dupoux |
INTERSPEECH | 9 |
| 2020 | An Open-Source Voice Type Classifier for Child-Centered Daylong RecordingsabstractInternational audience Marvin Lavechin, Ruben Bousbib, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 4 |
| 2020 | Vocal Markers from Sustained Phonation in Huntington's DiseaseabstractDisease-modifying treatments are currently assessed in neurodegenerative diseases. Huntington's Disease represents a unique opportunity to design automatic sub-clinical markers, even in premanifest gene carriers. We investigated phonatory impairments as potential clinical markers and propose them for both diagnosis and gene carriers follow-up. We used two sets of features: Phonatory features and Modulation Power Spectrum Features. We found that phonation is not sufficient for the identification of sub-clinical disorders of premanifest gene carriers. According to our regression results, Phonatory features are suitable for the predictions of clinical performance in Huntington's Disease. Rachid Riad, Hadrien Titeux, Laurie Lemoine, Justine Montillot, Jennifer Hamet Bagnou, Xuan-Nga Cao, Emmanuel Dupoux, Anne-Catherine Bachoud-Lévi |
INTERSPEECH | 7 |
| 2020 | Identification of Primary and Collateral Tracks in Stuttered SpeechabstractDisfluent speech has been previously addressed from two main perspectives: the clinical perspective focusing on diagnostic, and the Natural Language Processing (NLP) perspective aiming at modeling these events and detect them for downstream tasks. In addition, previous works often used different metrics depending on whether the input features are text or speech, making it difficult to compare the different contributions. Here, we introduce a new evaluation framework for disfluency detection inspired by the clinical and NLP perspective together with the theory of performance from (Clark, 1996) which distinguishes between primary and collateral tracks. We introduce a novel forced-aligned disfluency dataset from a corpus of semi-directed interviews, and present baseline results directly comparing the performance of text-based features (word and span information) and speech-based (acoustic-prosodic information). Finally, we introduce new audio features inspired by the word-based span features. We show experimentally that using these features outperformed the baselines for speech-based predictions on the present dataset. Rachid Riad, Anne-Catherine Bachoud-Lévi, Frank Rudzicz, Emmanuel Dupoux |
LREC | 4 |
| 2020 | Seshat: a Tool for Managing and Verifying Annotation Campaigns of Audio DataabstractWe introduce Seshat, a new, simple and open-source software to efficiently manage annotations of speech corpora. The Seshat software allows users to easily customise and manage annotations of large audio corpora while ensuring compliance with the formatting and naming conventions of the annotated output files. In addition, it includes procedures for checking the content of annotations following specific rules that can be implemented in personalised parsers. Finally, we propose a double-annotation mode, for which Seshat computes automatically an associated inter-annotator agreement with the gamma measure taking into account the categorisation and segmentation discrepancies. Hadrien Titeux, Rachid Riad, Xuan-Nga Cao, Nicolas Hamilakis, Kris Madden, Alejandrina Cristià, Anne-Catherine Bachoud-Lévi, Emmanuel Dupoux |
LREC | 8 |
| 2020 | Speech Technology for Unwritten LanguagesabstractSpeech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible. Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 11 |
| 2019 | Word-order Biases in Deep-agent Emergent CommunicationabstractSequence-processing neural networks led to remarkable progress on many NLP tasks.As a consequence, there has been increasing interest in understanding to what extent they process language as humans do.We aim here to uncover which biases such models display with respect to "natural" word-order constraints.We train models to communicate about paths in a simple gridworld, using miniature languages that reflect or violate various natural language trends, such as the tendency to avoid redundancy or to minimize long-distance dependencies.We study how the controlled characteristics of our miniature languages affect individual learning and their stability across multiple network generations.The results draw a mixed picture.On the one hand, neural networks show a strong tendency to avoid long-distance dependencies.On the other hand, there is no clear preference for the efficient, non-redundant encoding of information that is widely attested in natural language.We thus suggest inoculating a notion of "effort" into neural networks, as a possible way to make their linguistic behavior more humanlike. Rahma Chaabouni 0001, Eugene Kharitonov, Alessandro Lazaric, Emmanuel Dupoux, Marco Baroni |
ACL (1) | 4 |
| 2019 | Phoneme learning is influenced by the taxonomic similarity of the semantic referents
Abdellah Fourtassi, Emmanuel Dupoux |
CogSci | 2 |
| 2019 | The Zero Resource Speech Challenge 2019: TTS Without TabstractWe present the Zero Resource Speech Challenge 2019, which proposes to build a speech synthesizer without any text or phonetic labels: hence, TTS without T (text-to-speech without text). We provide raw audio for a target voice in an unknown language (the Voice dataset), but no alignment, text or labels. Participants must discover subword units in an unsupervised way (using the Unit Discovery dataset) and align them to the voice recordings in a way that works best for the purpose of synthesizing novel utterances from novel speakers, similar to the target speaker's voice. We describe the metrics used for evaluation, a baseline system consisting of unsupervised subword unit discovery plus a standard TTS system, and a topline TTS using gold phoneme transcriptions. We present an overview of the 19 submitted systems from 10 teams and discuss the main results. Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel Yang, Alan W. Black, Laurent Besacier, Sakriani Sakti, Emmanuel Dupoux |
INTERSPEECH | 13 |
| 2019 | Anti-efficient encoding in emergent communicationabstractDespite renewed interest in emergent language simulations with neural networks, little is known about the basic properties of the induced code, and how they compare to human language. One fundamental characteristic of the latter, known as Zipf's Law of Abbreviation (ZLA), is that more frequent words are efficiently associated to shorter strings. We study whether the same pattern emerges when two neural networks, a speaker'' and alistener'', are trained to play a signaling game. Surprisingly, we find that networks develop an \emph{anti-efficient} encoding scheme, in which the most frequent inputs are associated to the longest messages, and messages in general are skewed towards the maximum length threshold. This anti-efficient code appears easier to discriminate for the listener, and, unlike in human communication, the speaker does not impose a contrasting least-effort pressure towards brevity. Indeed, when the cost function includes a penalty for longer messages, the resulting message distribution starts respecting ZLA. Our analysis stresses the importance of studying the basic features of emergent communication in a highly controlled setup, to ensure the latter will not strand too far from human language. Moreover, we present a concrete illustration of how different functional pressures can lead to successful communication codes that lack basic properties of human language, thus highlighting the role such pressures play in the latter. Rahma Chaabouni 0001, Eugene Kharitonov, Emmanuel Dupoux, Marco Baroni |
NeurIPS | 3 |
| 2018 | Bayesian Models for Unit Discovery on a Very Low Resource LanguageabstractDeveloping speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the I-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus. Lucas Ondel Yang, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Lukás Burget, François Yvon, Sanjeev Khudanpur |
ICASSP | 7 |
| 2018 | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 WorkshopabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech. Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux |
ICASSP | 19 |
| 2018 | Learning Filterbanks from Raw Speech for Phone RecognitionabstractWe train a bank of complex filters that operates on the raw waveform and is fed into a convolutional neural network for end-to-end phone recognition. These time-domain filterbanks (TD-filterbanks) are initialized as an approximation of mel-filterbanks, and then fine-tuned jointly with the remaining convolutional architecture. We perform phone recognition experiments on TIMIT and show that for several architectures, models trained on TD- filterbanks consistently outperform their counterparts trained on comparable mel-filterbanks. We get our best performance by learning all front-end steps, from pre-emphasis up to averaging. Finally, we observe that the filters at convergence have an asymmetric impulse response, and that some of them remain almost analytic. Neil Zeghidour, Nicolas Usunier, Iasonas Kokkinos, Thomas Schatz, Gabriel Synnaeve, Emmanuel Dupoux |
ICASSP | 6 |
| 2018 | Learning Word Embeddings: Unsupervised Methods for Fixed-size Representations of Variable-length Speech SegmentsabstractInternational audience Nils Holzenberger, Mingxing Du, Julien Karadayi, Rachid Riad, Emmanuel Dupoux |
INTERSPEECH | 5 |
| 2018 | Sampling Strategies in Siamese Networks for Unsupervised Speech Representation LearningabstractRecent studies have investigated siamese network architectures for learning invariant speech representations using same-different side information at the word level. Here we investigate systematically an often ignored component of siamese networks: the sampling procedure (how pairs of same vs. different tokens are selected). We show that sampling strategies taking into account Zipf's Law, the distribution of speakers and the proportions of same and different pairs of words significantly impact the performance of the network. In particular, we show that word frequency compression improves learning across a large range of variations in number of training pairs. This effect does not apply to the same extent to the fully unsupervised setting, where the pairs of same-different words are obtained by spoken term discovery. We apply these results to pairs of words discovered using an unsupervised algorithm and show an improvement on state-of-the-art in unsupervised representation learning using siamese networks. Rachid Riad, Corentin Dancette, Julien Karadayi, Neil Zeghidour, Thomas Schatz, Emmanuel Dupoux |
INTERSPEECH | 6 |
| 2018 | End-to-End Speech Recognition from the Raw WaveformabstractState-of-the-art speech recognition systems rely on fixed, hand-crafted features such as mel-filterbanks to preprocess the waveform before the training pipeline. In this paper, we study end-to-end systems trained directly from the raw waveform, building on two alternatives for trainable replacements of mel-filterbanks that use a convolutional architecture. The first one is inspired by gammatone filterbanks (Hoshen et al., 2015; Sainath et al, 2015), and the second one by the scattering transform (Zeghidour et al., 2017). We propose two modifications to these architectures and systematically compare them to mel-filterbanks, on the Wall Street Journal dataset. The first modification is the addition of an instance normalization layer, which greatly improves on the gammatone-based trainable filterbanks and speeds up the training of the scattering-based filterbanks. The second one relates to the low-pass filter used in these approaches. These modifications consistently improve performances for both approaches, and remove the need for a careful initialization in scattering-based trainable filterbanks. In particular, we show a consistent improvement in word error rate of the trainable filterbanks relatively to comparable mel-filterbanks. It is the first time end-to-end models trained from the raw signal significantly outperform mel-filterbanks on a large vocabulary task under clean recording conditions. Neil Zeghidour, Nicolas Usunier, Gabriel Synnaeve, Ronan Collobert, Emmanuel Dupoux |
INTERSPEECH | 5 |
| 2018 | BabyCloud, a Technological Platform for Parents and Researchers
Xuan-Nga Cao, Cyrille Dakhlia, Patricia Del Carmen, Mohamed-Amine Jaouani, Malik Ould-Arbi, Emmanuel Dupoux |
LREC | 6 |
| 2018 | A K-Nearest Neighbours Approach To Unsupervised Spoken Term DiscoveryabstractThe following topics are dealt with: speech recognition; neural nets; speech processing; learning (artificial intelligence); natural language processing; recurrent neural nets; speaker recognition; speech synthesis; feature extraction; and text analysis. Alexis Thual, Corentin Dancette, Julien Karadayi, Juan Benjumea, Emmanuel Dupoux |
SLT | 5 |
| 2017 | The zero resource speech challenge 2017abstractWe describe a new challenge aimed at discovering subword and word units from raw speech. This challenge is the followup to the Zero Resource Speech Challenge 2015. It aims at constructing systems that generalize across languages and adapt to new speakers. The design features and evaluation metrics of the challenge are presented and the results of seventeen models are discussed. Ewan Dunbar, Xuan-Nga Cao, Juan Benjumea, Julien Karadayi, Mathieu Bernard, Laurent Besacier, Xavier Anguera Miró, Emmanuel Dupoux |
ASRU | 8 |
| 2017 | ASR Systems as Models of Phonetic Category Perception in Adults
Thomas Schatz, Francis R. Bach, Emmanuel Dupoux |
CogSci | 3 |
| 2017 | Learning Weakly Supervised Multimodal Phoneme EmbeddingsabstractRecent works have explored deep architectures for learning multimodal speech representation (e.g. audio and images, articulation and audio) in a supervised way. Here we investigate the role of combining different speech modalities, i.e. audio and visual information representing the lips movements, in a weakly supervised way using Siamese networks and lexical same-different side information. In particular, we ask whether one modality can benefit from the other to provide a richer representation for phone recognition in a weakly supervised setting. We introduce mono-task and multi-task methods for merging speech and visual modalities for phone recognition. The mono-task learning consists in applying a Siamese network on the concatenation of the two modalities, while the multi-task learning receives several different combinations of modalities at train time. We show that multi-task learning enhances discriminability for visual and multimodal inputs while minimally impacting auditory inputs. Furthermore, we present a qualitative analysis of the obtained phone embeddings, and show that cross-modal visual input can improve the discriminability of phonological features which are visually discernable (rounding, open/close, labial place of articulation), resulting in representations that are closer to abstract linguistic features than those based on audio only. Rahma Chaabouni 0001, Ewan Dunbar, Neil Zeghidour, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2017 | Predicting Epenthetic Vowel Quality from AcousticsabstractInternational audience Adriana Guevara-Rukoz, Erika Parlato-Oliveira, Yuki Hirose, Sharon Peperkamp, Emmanuel Dupoux |
INTERSPEECH | 6 |
| 2017 | Relating Unsupervised Word Segmentation to Reported Vocabulary AcquisitionabstractInternational audience Elin Larsen, Alejandrina Cristià, Emmanuel Dupoux |
INTERSPEECH | 3 |
| 2017 | A Quantitative Measure of the Impact of Coarticulation on Phone DiscriminabilityabstractInternational audience Thomas Schatz, Rory Turnbull, Francis R. Bach, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2016 | Discriminability of sound contrasts in the face of speaker variation quantified
Christina Bergmann, Alejandrina Cristià, Emmanuel Dupoux |
CogSci | 3 |
| 2016 | Modeling language discrimination in infants using i-vector representations
M. Julia Carbajal, Radek Fér, Emmanuel Dupoux |
CogSci | 3 |
| 2016 | The role of word-word co-occurrence in word learning
Abdellah Fourtassi, Emmanuel Dupoux |
CogSci | 2 |
| 2016 | A deep scattering spectrum - Deep Siamese network pipeline for unsupervised acoustic modelingabstractRecent work has explored deep architectures for learning acoustic features in an unsupervised or weakly-supervised way for phone recognition. Here we investigate the role of the input features, and in particular we test whether standard mel-scaled filterbanks could be replaced by inherently richer representations, such as derived from an analytic scattering spectrum. We use a Siamese network using lexical side information similar to a well-performing architecture used in the Zero Resource Speech Challenge (2015), and show a substantial improvement when the filterbanks are replaced by scattering features, even though these features yield similar performance when tested without training. This shows that unsupervised and weakly-supervised architectures can benefit from richer features than the traditional ones. Neil Zeghidour, Gabriel Synnaeve, Maarten Versteegh, Emmanuel Dupoux |
ICASSP | 4 |
| 2016 | A new efficient measure for accuracy prediction and its application to multistream-based unsupervised adaptationabstractA new efficient measure for predicting estimation accuracy is proposed and successfully applied to multistream-based unsupervised adaptation of ASR systems to address data uncertainty when the ground-truth is unknown. The proposed measure is an extension of the M-measure, which predicts confidence in the output of a probability estimator by measuring the divergences of probability estimates spaced at specific time intervals. In this study, the M-measure was extended by considering the latent phoneme information, resulting in an improved reliability. Experimental comparisons carried out in a multistream-based ASR paradigm demonstrated that the extended M-measure yields a significant improvement over the original M-measure, especially under narrow-band noise conditions. Tetsuji Ogawa, Sri Harish Reddy Mallidi, Emmanuel Dupoux, Jordan Cohen, Naomi Feldman, Hynek Hermansky |
ICPR | 3 |
| 2016 | Joint Learning of Speaker and Phonetic Similarities with Siamese Networks
Neil Zeghidour, Gabriel Synnaeve, Nicolas Usunier, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2016 | Assessing the Ability of LSTMs to Learn Syntax-Sensitive DependenciesabstractThe success of long short-term memory (LSTM) neural networks in language processing is typically attributed to their ability to capture long-distance statistical regularities. Linguistic regularities are often sensitive to syntactic structure; can such dependencies be captured by LSTMs, which do not have explicit structural representations? We begin addressing this question using number agreement in English subject-verb dependencies. We probe the architecture’s grammatical competence both using training objectives with an explicit grammatical target (number prediction, grammaticality judgments) and using language models. In the strongly supervised settings, the LSTM achieved very high overall accuracy (less than 1% errors), but errors increased when sequential and structural information conflicted. The frequency of such errors rose sharply in the language-modeling setting. We conclude that LSTMs can capture a non-trivial amount of grammatical structure given targeted supervision, but stronger architectures may be required to further reduce errors; furthermore, the language modeling signal is insufficient for capturing syntax-sensitive dependencies, and should be supplemented with more direct supervision if such dependencies need to be captured. Tal Linzen, Emmanuel Dupoux, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 2 |
| 2015 | Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial WorkshopabstractA group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which the estimator was trained. The paper describes the problem and summarizes approaches that were taken by the group1. Hynek Hermansky, Lukás Burget, Jordan Cohen, Emmanuel Dupoux, Naomi Feldman, John Godfrey, Sanjeev Khudanpur, Matthew Maciejewski, Sri Harish Reddy Mallidi, Anjali Menon, Tetsuji Ogawa, Vijayaditya Peddinti, Richard C. Rose, Richard M. Stern, Matthew Wiesner, Karel Veselý |
ICASSP | 4 |
| 2015 | Salient dimensions in implicit phonotactic learningabstractAdults are able to learn sound co-occurrences without conscious knowledge after brief exposures. But which dimensions of sounds are most salient in this process? Using an artificial phonology paradigm, we explored potential learnability differences involving consonant-, speaker-, and tone-vowel cooccurrences. Results revealed that participants, whose native language was not tonal, implicitly encoded consonant-vowel patterns with a high level of accuracy; were above chance for tone-vowel co-occurrences; and were at chance for speakervowel co-occurrences. This pattern of results is exactly what would be expected if both language-specific experience and innate biases to encode potentially contrastive linguistic dimensions affect the salience of different dimensions during implicit learning of sound patterns. Elise Michon, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 2 |
| 2015 | A hybrid dynamic time warping-deep neural network architecture for unsupervised acoustic modelingabstractWe report on an architecture for the unsupervised discovery of talker-invariant subword embeddings. It is made out of two components: a dynamic-time warping based spoken term discovery (STD) system and a Siamese deep neural network (DNN). The STD system clusters word-sized repeated fragments in the acoustic streams while the DNN is trained to minimize the distance between time aligned frames of tokens of the same cluster, and maximize the distance between tokens of different clusters. We use additional side information regarding the average duration of phonemic units, as well as talker identity tags. For evaluation we use the datasets and metrics of the Zero Resource Speech Challenge. The model shows improvement over the baseline in subword unit modeling. Roland Thiollière, Ewan Dunbar, Gabriel Synnaeve, Maarten Versteegh, Emmanuel Dupoux |
INTERSPEECH | 5 |
| 2015 | The zero resource speech challenge 2015abstractétablissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Maarten Versteegh, Roland Thiollière, Thomas Schatz, Xuan-Nga Cao, Xavier Anguera Miró, Aren Jansen, Emmanuel Dupoux |
INTERSPEECH | 7 |
| 2015 | Sign constraints on feature weights improve a joint model of word segmentation and phonologyabstractMark Johnson, Joe Pater, Robert Staubs, Emmanuel Dupoux. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Mark Johnson 0001, Joe Pater, Robert Staubs, Emmanuel Dupoux |
HLT-NAACL | 4 |
| 2015 | Prosodic boundary information helps unsupervised word segmentationabstractIt is well known that prosodic information is used by infants in early language acquisition. In particular, prosodic boundaries have been shown to help infants with sentence and wordlevel segmentation. In this study, we extend an unsupervised method for word segmentation to include information about prosodic boundaries. The boundary information used was either derived from oracle data (handannotated), or extracted automatically with a system that employs only acoustic cues for boundary detection. The approach was tested on two different languages, English and Japanese, and the results show that boundary information helps word segmentation in both cases. The performance gain obtained for two typologically distinct languages shows the robustness of prosodic information for word segmentation. Furthermore, the improvements are not limited to the use of oracle information, similar performances being obtained also with automatically extracted boundaries. Bogdan Ludusan, Gabriel Synnaeve, Emmanuel Dupoux |
HLT-NAACL | 3 |
| 2014 | Modelling function words improves unsupervised word segmentationabstractInspired by experimental psychological findings suggesting that function words play a special role in word learning, we make a simple modification to an Adaptor Grammar based Bayesian word segmentation model to allow it to learn sequences of monosyllabic "function words" at the beginnings and endings of collocations of (possibly multi-syllabic) words.This modification improves unsupervised word segmentation on the standard Bernstein-Ratner (1987) corpus of child-directed English by more than 4% token f-score compared to a model identical except that it does not special-case "function words", setting a new state-of-the-art of 92.4% token f-score.Our function word model assumes that function words appear at the left periphery, and while this is true of languages such as English, it is not true universally.We show that a learner can use Bayesian model selection to determine the location of function words in their language, even though the input to the model only consists of unsegmented sequences of phones.Thus our computational models support the hypothesis that function words play a special role in word learning. Mark Johnson 0001, Anne Christophe, Emmanuel Dupoux, Katherine Demuth |
ACL (1) | 3 |
| 2014 | Self-Consistency as an Inductive Bias in Early Language Acquisition
Abdellah Fourtassi, Ewan Dunbar, Emmanuel Dupoux |
CogSci | 3 |
| 2014 | Unsupervised Word Segmentation in Context
Gabriel Synnaeve, Isabelle Dautriche, Benjamin Börschinger, Mark Johnson 0001, Emmanuel Dupoux |
COLING | 5 |
| 2014 | A Rudimentary Lexicon and Semantics Help Bootstrap Phoneme AcquisitionabstractInfants spontaneously discover the relevant phonemes of their language without any direct supervision.This acquisition is puzzling because it seems to require the availability of high levels of linguistic structures (lexicon, semantics), that logically suppose the infants having a set of phonemes already.We show how this circularity can be broken by testing, in realsize language corpora, a scenario whereby infants would learn approximate representations at all levels, and then refine them in a mutually constraining way.We start with corpora of spontaneous speech that have been encoded in a varying number of detailed context-dependent allophones.We derive, in an unsupervised way, an approximate lexicon and a rudimentary semantic representation.Despite the fact that all these representations are poor approximations of the ground truth, they help reorganize the fine grained categories into phoneme-like categories with a high degree of accuracy. Abdellah Fourtassi, Emmanuel Dupoux |
CoNLL | 2 |
| 2014 | Evaluating speech features with the minimal-pair ABX task (II): resistance to noiseabstractThe Minimal-Pair ABX (MP-ABX) paradigm has been pro-posed as a method for evaluating speech features for zero-resource/unsupervised speech technologies. We apply it in a phoneme discrimination task on the Articulation Index corpus to evaluate the resistance to noise of various speech features. In Experiment 1, we evaluate the robustness to additive noise at different signal-to-noise ratios, using car and babble noise from the Aurora-4 database and white noise. In Experiment 2, we ex-amine the robustness to different kinds of convolutional noise. In both experiments we consider two classes of techniques to induce noise resistance: smoothing of the time-frequency rep-resentation and short-term adaptation in the time-domain. We consider smoothing along the spectral axis (as in PLP) and along the time axis (as in FDLP). For short-term adaptation in the time-domain, we compare the use of a static compressive non-linearity followed by RASTA filtering to an adaptive com-pression scheme. Index Terms: noise resistance, zero-resource, speech features, evaluation framework, minimal-pair ABX task Thomas Schatz, Vijayaditya Peddinti, Xuan-Nga Cao, Francis R. Bach, Hynek Hermansky, Emmanuel Dupoux |
INTERSPEECH | 6 |
| 2014 | Bridging the gap between speech technology and natural language processing: an evaluation toolbox for term discovery systems
Bogdan Ludusan, Maarten Versteegh, Aren Jansen, Guillaume Gravier, Xuan-Nga Cao, Mark Johnson 0001, Emmanuel Dupoux |
LREC | 7 |
| 2014 | Phonetics embedding learning with side informationabstractWe show that it is possible to learn an efficient acoustic model using only a small amount of easily available word-level similarity annotations. In contrast to the detailed phonetic labeling required by classical speech recognition technologies, the only information our method requires are pairs of speech excerpts which are known to be similar (same word) and pairs of speech excerpts which are known to be different (different words). An acoustic model is obtained by training shallow and deep neural networks, using an architecture and a cost function well-adapted to the nature of the provided information. The resulting model is evaluated in an ABX minimal-pair discrimination task and is shown to perform much better (11.8% ABX error rate) than raw speech features (19.6%), not far from a fully supervised baseline (best neural network: 9.2%, HMM-GMM: 11%). Gabriel Synnaeve, Thomas Schatz, Emmanuel Dupoux |
SLT | 3 |
| 2013 | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisitionabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies. Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001 |
ICASSP | 2 |
| 2013 | Evaluating speech features with the minimal-pair ABX task: analysis of the classical MFC/PLP pipelineabstractWe present a new framework for the evaluation of speech representations in zero-resource settings, that extends and complements previous work by Carlin, Jansen and Hermansky [1].In particular, we replace their Same/Different discrimination task by several Minimal-Pair ABX (MP-ABX) tasks.We explain the analytical advantages of this new framework and apply it to decompose the standard signal processing pipelines for computing PLP and MFC coefficients.This method enables us to confirm and quantify a variety of well-known and not-so-well-known results in a single framework. Thomas Schatz, Vijayaditya Peddinti, Francis R. Bach, Aren Jansen, Hynek Hermansky, Emmanuel Dupoux |
INTERSPEECH | 6 |
| 2011 | Templatic features for modeling phoneme acquisition
Emmanuel Dupoux, Guillaume Beraud-Sudreau, Shigeki Sagayama |
CogSci | 1 |
| 1999 | Prelexical locus of an illusory vowel effect in JapaneseabstractInformation extraction from speech is a crucial step on the way from speech recognition to speech understanding. A preliminary step toward speech understanding is the detection of topic boundaries, sentence boundaries, and proper names in speech recognizer output. This is important since speech recognizer output lacks the usual textual cues to these entities (such as headers, paragraphs, sentence punctuation, and capitalization). Numerous word-based approaches to these tasks have been developed in the past; in this work we demonstrate the use of prosodic cues, alone and in combination with words, for segmentation and name finding. In experiments on the Broadcast News corpus, we find that prosodic cues alone allow sentence and topic segmentation that is at least as good as word-based methods alone, and that combining both types of cues gives significant wins. Named entity recognition, on the other hand, currently does not seem to benefit from prosodic cues, for several interesting reas... Emmanuel Dupoux, Takao Fushimi, Jacques Mehler |
EUROSPEECH | 1 |
| 1999 | Perception of stress by French, Spanish, and bilingual subjects
Sharon Peperkamp, Emmanuel Dupoux, Núria Sebastián-Gallés |
EUROSPEECH | 2 |