Emmanuel Vincent 0001

dblp:55/3279 · DBLP profile ↗
← Back
125ranked-venue papers
13as first author
27since 2021 · last 2027
0000-0002-0183-7289ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 80 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 78 · 8 first-author · 20 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Privacy attacks on voice anonymization systems: Overview and key findings from the First VoicePrivacy Attacker Challenge
Natalia A. Tomashenko, Xiaoxiao Miao, Emmanuel Vincent 0001, Junichi Yamagishi
Comput. Speech Lang.3
2026 The third VoicePrivacy challenge: Preserving emotional expressiveness and linguistic content in voice anonymization
Natalia A. Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer, Michele Panariello, Xin Wang 0037, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi, Massimiliano Todisco
Comput. Speech Lang.8
2025 Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
abstract
In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker’s speech rate and, more importantly, the speaker’s phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity.
Natalia A. Tomashenko, Emmanuel Vincent 0001, Marc Tommasi
ICASSP2
2025 The First VoicePrivacy Attacker Challenge
abstract
The First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge which focuses on evaluating attacker systems against a set of voice anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation datasets were provided along with a baseline attacker. Participants developed their attacker systems in the form of automatic speaker verification systems and submitted their scores on the development and evaluation data. The best attacker systems reduced the equal error rate (EER) by 25–44% relative w.r.t. the baseline.
Natalia A. Tomashenko, Xiaoxiao Miao, Emmanuel Vincent 0001, Junichi Yamagishi
ICASSP3
2025 Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition
Raphaël Bagat, Irina Illina, Emmanuel Vincent 0001
INTERSPEECH3
2025 Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
abstract
International audience
Natalia A. Tomashenko, Emmanuel Vincent 0001, Marc Tommasi
INTERSPEECH2
2025 Legally validated evaluation framework for voice anonymization
abstract
International audience
Nathalie Vauquier, Brij Mohan Lal Srivastava, Seyed Ahmad Hosseini, Emmanuel Vincent 0001
INTERSPEECH4
2025 Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
Can Cui 0007, Paul Magron, Mostafa Sadeghi, Emmanuel Vincent 0001
MMSP4
2024 Multi-Channel Extension of Pre-trained Models for Speaker Verification
abstract
International audience
Ladislav Mosner, Romain Serizel, Lukás Burget, Oldrich Plchot, Emmanuel Vincent 0001, Junyi Peng, Jan Cernocký
INTERSPEECH5
2024 Training RNN language models on uncertain ASR hypotheses in limited data scenarios
Imran A. Sheikh, Emmanuel Vincent 0001, Irina Illina
Comput. Speech Lang.2
2024 The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
abstract
The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datasets used for system development and evaluation, present the different attack models used for evaluation, and the associated objective and subjective metrics. We describe three anonymisation baselines, provide a summary description of the anonymisation systems developed by challenge participants, and report objective and subjective evaluation results for all. In addition, we describe post-evaluation analyses and a summary of related work reported in the open literature. Results show that solutions based on voice conversion better preserve utility, that an alternative which combines automatic speech recognition with synthesis achieves greater privacy, and that a privacy-utility trade-off remains inherent to current anonymisation solutions. Finally, we present our ideas and priorities for future VoicePrivacy Challenge editions.
Michele Panariello, Natalia A. Tomashenko, Xin Wang 0037, Xiaoxiao Miao, Pierre Champion, Hubert Nourtel, Massimiliano Todisco, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.9
2023 End-to-End Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis
abstract
We present an end-to-end multichannel speaker-attributed automatic speech recognition (MC-SA-ASR) system that combines a Conformer-based encoder with multi-frame cross-channel attention and a speaker-attributed Transformer-based decoder. To the best of our knowledge, this is the first model that efficiently integrates ASR and speaker identification modules in a multichannel setting. On simulated mixtures of LibriSpeech data, our system reduces the word error rate (WER) by up to 12% and 16% relative compared to previously proposed single-channel and multichannel approaches, respectively. Furthermore, we investigate the impact of different input features, including multichannel magnitude and phase information, on the ASR performance. Finally, our experiments on the AMI corpus confirm the effectiveness of our system for real-world multichannel meeting transcription.
Can Cui 0007, Imran A. Sheikh, Mostafa Sadeghi, Emmanuel Vincent 0001
ASRU4
2023 Find-2-Find: Multitask Learning for Anaphora Resolution and Object Localization
abstract
In multimodal understanding tasks, visual and linguistic ambiguities can arise.Visual ambiguity can occur when visual objects require a model to ground a referring expression in a video without strong supervision, while linguistic ambiguity can occur from changes in entities in action flows.As an example from the cooking domain, "oil" mixed with "salt" and "pepper" could later be referred to as a "mixture".Without a clear visual-linguistic alignment, we cannot know which among several objects shown is referred to by the language expression "mixture", and without resolved antecedents, we cannot pinpoint what the mixture is.We define this chicken-and-egg problem as visual-linguistic ambiguity.In this paper, we present Find2Find, a joint anaphora resolution and object localization dataset targeting the problem of visual-linguistic ambiguity, consisting of 500 anaphora-annotated recipes with corresponding videos.We present experimental results of a novel end-to-end joint multitask learning framework for Find2Find that fuses visual and textual information and shows improvements both for anaphora resolution and object localization as compared to a strong single-task baseline.
Cennet Oguz, Pascal Denis, Emmanuel Vincent 0001, Simon Ostermann 0002, Josef van Genabith
EMNLP3
2023 Stochastic Pitch Prediction Improves the Diversity and Naturalness of Speech in Glow-TTS
Sewade Ogun, Vincent Colotte, Emmanuel Vincent 0001
INTERSPEECH3
2023 How to (Virtually) Train Your Speaker Localizer
abstract
Learning-based methods have become ubiquitous in speaker localization. Existing systems rely on simulated training sets for the lack of sufficiently large, diverse and annotated real datasets. Most room acoustics simulators used for this purpose rely on the image source method (ISM) because of its computational efficiency. This paper argues that carefully extending the ISM to incorporate more realistic surface, source and microphone responses into training sets can significantly boost the real-world performance of speaker localization systems. It is shown that increasing the training-set realism of a state-of-the-art direction-of-arrival estimator yields consistent improvements across three different real test sets featuring human speakers in a variety of rooms and various microphone arrays. An ablation study further reveals that every added layer of realism contributes positively to these improvements.
Prerak Srivastava, Antoine Deleforge, Archontis Politis, Emmanuel Vincent 0001
INTERSPEECH4
2023 Guest editorial: Special issue on advances in deep learning based speech processing
Eric Fosler-Lussier, Emmanuel Vincent 0001
Neural Networks4
2023 Differentially Private Speaker Anonymization
abstract
Sharing real-world speech utterances is key to the training and deployment of voice-based services. However, it also raises privacy risks as speech contains a wealth of personal data. Speaker anonymization aims to remove speaker information from a speech utterance while leaving its linguistic and prosodic attributes intact. State-of-the-art techniques operate by disentangling the speaker information (represented via a speaker embedding) from these attributes and re-synthesizing speech based on the speaker embedding of another speaker. Prior research in the privacy community has shown that anonymization often provides brittle privacy protection, even less so any provable guarantee. In this work, we show that disentanglement is indeed not perfect: linguistic and prosodic attributes still contain speaker information. We remove speaker information from these attributes by introducing differentially private feature extractors based on an autoencoder and an automatic speech recognizer, respectively, trained using noise layers. We plug these extractors in the state-of-the-art anonymization pipeline and generate, for the first time, private speech utterances with a provable upper bound on the speaker information they contain. We evaluate empirically the privacy and utility resulting from our differentially private speaker anonymization approach on the LibriSpeech data set. Experimental results show that the generated utterances retain very high utility for automatic speech recognition training and inference, while being much better protected against strong adversaries who leverage the full knowledge of the anonymization process to try to infer the speaker identity.
Ali Shahin Shamsabadi, Brij Mohan Lal Srivastava, Aurélien Bellet, Nathalie Vauquier, Emmanuel Vincent 0001, Mohamed Maouche, Marc Tommasi, Nicolas Papernot
Proc. Priv. Enhancing Technol.5
2022 On The Impact of Normalization Strategies in Unsupervised Adversarial Domain Adaptation for Acoustic Scene Classification
abstract
Acoustic scene classification systems face performance degradation when trained and tested on data recorded by different devices. Unsupervised domain adaptation methods have been studied to reduce the impact of this mismatch. While they do not assume the availability of labels at test time, they often exploit parallel data recorded by both devices, and thus are not fully blind to the target domain. In this paper, we address a more practical scenario where parallel data are not available. We thoroughly analyze the impact of normalization and moment matching strategies to compensate for the linear distortion introduced by the recording device and propose their integration with adversarial domain adaptation to handle the remaining non-linear distortion. Experiments on the DCASE Challenge 2018 Task 1B dataset show that the proposed integrated approach considerably reduces domain mismatch, reaching an accuracy in the target domain close to that obtained in the source domain.
Michel Olvera, Emmanuel Vincent 0001, Gilles Gasso
ICASSP2
2022 Enhancing Speech Privacy with Slicing
abstract
Privacy preservation calls for speech anonymization methods which hide the speaker's identity while minimizing the impact on downstream tasks such as automatic speech recognition (ASR) training or decoding. In the recent VoicePrivacy 2020 Challenge, several anonymization methods have been proposed to transform speech utterances in a way that preserves their verbal and prosodic contents while reducing the accuracy of a speaker verification system. In this paper, we propose to further increase the privacy achieved by such methods by segmenting the utterances into shorter slices. We show that our approach has two major impacts on privacy. First, it reduces the accuracy of speaker verification with respect to unsegmented utterances. Second, it also reduces the amount of personal information that can be extracted from the verbal content, in a way that cannot easily be reversed by an attacker. We also show that it is possible to train an ASR system from anonymized speech slices with negligible impact on the word error rate.
Mohamed Maouche, Brij Mohan Lal Srivastava, Nathalie Vauquier, Aurélien Bellet, Marc Tommasi, Emmanuel Vincent 0001
INTERSPEECH6
2022 Transformer versus LSTM Language Models trained on Uncertain ASR Hypotheses in Limited Data Scenarios
abstract
In several ASR use cases, training and adaptation of domain-specific LMs can only rely on a small amount of manually verified text transcriptions and sometimes a limited amount of in-domain speech. Training of LSTM LMs in such limited data scenarios can benefit from alternate uncertain ASR hypotheses, as observed in our recent work. In this paper, we propose a method to train Transformer LMs on ASR confusion networks. We evaluate whether these self-attention based LMs are better at exploiting alternate ASR hypotheses as compared to LSTM LMs. Evaluation results show that Transformer LMs achieve 3-6% relative reduction in perplexity on the AMI scenario meetings but perform similar to LSTM LMs on the smaller Verbmobil conversational corpus. Evaluation on ASR N-best rescoring shows that LSTM and Transformer LMs trained on ASR confusion networks do not bring significant WER reductions. However, a qualitative analysis reveals that they are better at predicting less frequent words.
Imran A. Sheikh, Emmanuel Vincent 0001, Irina Illina
LREC2
2022 Adapting Language Models When Training on Privacy-Transformed Data
abstract
In recent years, voice-controlled personal assistants have revolutionized the interaction with smart devices and mobile applications. The collected data are then used by system providers to train language models (LMs). Each spoken message reveals personal information, hence removing private information from the input sentences is necessary. Our data sanitization process relies on recognizing and replacing named entities by other words from the same class. However, this may harm LM training because privacy-transformed data is unlikely to match the test distribution. This paper aims to fill the gap by focusing on the adaptation of LMs initially trained on privacy-transformed sentences using a small amount of original untransformed data. To do so, we combine class-based LMs, which provide an effective approach to overcome data sparsity in the context of n-gram LMs, and neural LMs, which handle longer contexts and can yield better predictions. Our experiments show that training an LM on privacy-transformed data result in a relative 11% word error rate (WER) increase compared to training on the original untransformed data, and adapting that model on a limited amount of original untransformed data leads to a relative 8% WER improvement over the model trained solely on privacy-transformed data.
M. A. Tugtekin Turan, Dietrich Klakow, Emmanuel Vincent 0001, Denis Jouvet
LREC3
2022 Can We Use Common Voice to Train a Multi-Speaker TTS System?
abstract
Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies have leveraged the availability of large, crowdsourced automatic speech recognition (ASR) datasets. A major problem with such datasets is the presence of noisy and/or distorted samples, which degrade TTS quality. In this paper, we propose to automatically select high-quality training samples using a non-intrusive mean opinion score (MOS) estimator, WV-MOS. We show the viability of this approach for training a multi-speaker GlowTTS model on the Common Voice English dataset. Our approach improves the overall quality of generated utterances by 1.26 MOS point with respect to training on all the samples and by 0.35 MOS point with respect to training on the LibriTTS dataset. This opens the door to au-tomatic TTS dataset curation for a wider range of languages.
Sewade Ogun, Vincent Colotte, Emmanuel Vincent 0001
SLT3
2022 Overlapped Speech Detection and speaker counting using distant microphone arrays
Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001
Comput. Speech Lang.4
2022 The VoicePrivacy 2020 Challenge: Results and findings
Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Jose Patino 0001, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
Comput. Speech Lang.3
2022 Privacy and Utility of X-Vector Based Speaker Anonymization
abstract
We study the scenario where individuals (speakers) contribute to the publication of an anonymized speech corpus. Datausersleverage this public corpus for downstream tasks, e.g., training an automatic speech recognition (ASR) system, whileattackersmay attempt to de-anonymize it using auxiliary knowledge. Motivated by this scenario, speaker anonymization aims to conceal speaker identity while preserving the quality and usefulness of speech data. In this article, we study x-vector based speaker anonymization, the leading approach in the VoicePrivacy Challenge, which converts the speaker’s voice into that of a random pseudo-speaker. We show that the strength of anonymization varies significantly depending on how the pseudo-speaker is chosen. We explore four design choices for this step: the distance metric between speakers, the region of speaker space where the pseudo-speaker is picked, its gender, and whether to assign it to one or all utterances of the original speaker. We assess the quality of anonymization from the perspective of the three actors involved in our threat model, namely the speaker, the user and the attacker. To measure privacy and utility, we use respectively the linkability score achieved by the attackers and the decoding word error rate achieved by an ASR model trained on the anonymized data. Experiments on LibriSpeech show that the best combination of design choices yields state-of-the-art performance in terms of both privacy and utility. Experiments on Mozilla Common Voice further show that it guarantees the same anonymization level against re-identification attacks among 50 speakers as original speech among 20,000 speakers.
Brij Mohan Lal Srivastava, Mohamed Maouche, Md. Sahidullah, Emmanuel Vincent 0001, Aurélien Bellet, Marc Tommasi, Natalia A. Tomashenko, Xin Wang 0037, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Explaining Deep Learning Models for Speech Enhancement
abstract
International audience
Sunit Sivasankaran, Emmanuel Vincent 0001, Dominique Fohr
Interspeech2
2021 UIAI System for Short-Duration Speaker Verification Challenge 2020
abstract
In this work, we present the system description of the UIAI entry for the short-duration speaker verification (SdSV) challenge 2020. Our focus is on Task 1 dedicated to text-dependent speaker verification. We investigate different feature extraction and modeling approaches for automatic speaker verification (ASV) and utterance verification (UV). We have also studied different fusion strategies for combining UV and ASV modules. Our primary submission to the challenge is the fusion of seven subsystems which yields a normalized minimum detection cost function (minDCF) of 0.072 and an equal error rate (EER) of 2.14% on the evaluation set. The single system consisting of a pass-phrase identification based model with phone-discriminative bottleneck features gives a normalized minDCF of 0.118 and achieves 19% relative improvement over the state-of-the-art challenge baseline.
Md. Sahidullah, Achintya Kumar Sarkar, Ville Vestman, Xuechen Liu 0001, Romain Serizel, Tomi Kinnunen, Zheng-Hua Tan, Emmanuel Vincent 0001
SLT8
2020 Filterbank Design for End-to-end Speech Separation
abstract
Single-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet. In parallel, parameterized filterbanks have been proposed for speaker recognition where only center frequencies and bandwidths are learned. In this work, we extend real-valued learned and parameterized filterbanks into complex-valued analytic filterbanks and define a set of corresponding representations and masking strategies. We evaluate these filterbanks on a newly released noisy speech separation dataset (WHAM). The results show that the proposed analytic learned filterbank consistently outperforms the real-valued filterbank of ConvTasNet. Also, we validate the use of parameterized filterbanks and show that complex-valued representations and masks are beneficial in all conditions. Finally, we show that the STFT achieves its best performance for 2 ms windows.
Manuel Pariente, Samuele Cornell, Antoine Deleforge, Emmanuel Vincent 0001
ICASSP4
2020 SLOGD: Speaker Location Guided Deflation Approach to Speech Separation
abstract
Speech separation is the process of separating multiple speakers from an audio recording. In this work we propose to separate the sources using a Speaker LOcalization Guided Deflation (SLOGD) approach wherein we estimate the sources iteratively. In each iteration we first estimate the location of the speaker and use it to estimate a mask corresponding to the localized speaker. The estimated source is removed from the mixture before estimating the location and mask of the next source. Experiments are conducted on a reverberated, noisy multichannel version of the well-studied WSJ-2MIX dataset using word error rate (WER) as a metric. The proposed method achieves a WER of 44.2 %, a 34% relative improvement over the system without separation and 17% relative improvement over Conv-TasNet.
Sunit Sivasankaran, Emmanuel Vincent 0001, Dominique Fohr
ICASSP2
2020 Evaluating Voice Conversion-Based Privacy Protection against Informed Attackers
abstract
Speech data conveys sensitive speaker attributes like identity or accent. With a small amount of found data, such attributes can be inferred and exploited for malicious purposes: voice cloning, spoofing, etc. Anonymization aims to make the data unlinkable, i.e., ensure that no utterance can be linked to its original speaker. In this paper, we investigate anonymization methods based on voice conversion. In contrast to prior work, we argue that various linkage attacks can be designed depending on the attackers' knowledge about the anonymization scheme. We compare two frequency warping-based conversion methods and a deep learning based method in three attack scenarios. The utility of converted speech is measured via the word error rate achieved by automatic speech recognition, while privacy protection is assessed by the increase in equal error rate achieved by state-of-the-art i-vector or x-vector based speaker verification. Our results show that voice conversion schemes are unable to effectively protect against an attacker that has extensive knowledge of the type of conversion and how it has been applied, but may provide some protection against less knowledgeable attackers.
Brij Mohan Lal Srivastava, Nathalie Vauquier, Md. Sahidullah, Aurélien Bellet, Marc Tommasi, Emmanuel Vincent 0001
ICASSP6
2020 Limitations of Weak Labels for Embedding and Tagging
abstract
Many datasets and approaches in ambient sound analysis use weakly labeled data. Weak labels are employed because annotating every data sample with a strong label is too expensive. Yet, their impact on the performance in comparison to strong labels remains unclear. Indeed, weak labels must often be dealt with at the same time as other challenges, namely multiple labels per sample, unbalanced classes and/or overlapping events. In this paper, we formulate a supervised learning problem which involves weak labels. We create a dataset that focuses on the difference between strong and weak labels as opposed to other challenges. We investigate the impact of weak labels when training an embedding or an end-to-end classifier. Different experimental scenarios are discussed to provide insights into which applications are most sensitive to weakly labeled data.
Nicolas Turpault, Romain Serizel, Emmanuel Vincent 0001
ICASSP3
2020 Detecting and Counting Overlapping Speakers in Distant Speech Scenarios
abstract
International audience
Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001
INTERSPEECH4
2020 Kaldi-Web: An Installation-Free, On-Device Speech Recognition System
Mathieu Hu, Laurent Pierron, Emmanuel Vincent 0001, Denis Jouvet
INTERSPEECH3
2020 A Comparative Study of Speech Anonymization Metrics
abstract
Speech anonymization techniques have recently been proposed for preserving speakers' privacy. They aim at concealing speak-ers' identities while preserving the spoken content. In this study, we compare three metrics proposed in the literature to assess the level of privacy achieved. We exhibit through simulation the differences and blindspots of some metrics. In addition, we conduct experiments on real data and state-of-the-art anonymiza-tion techniques to study how they behave in a practical scenario. We show that the application-independent log-likelihood-ratio cost function C min llr provides a more robust evaluation of privacy than the equal error rate (EER), and that detection-based metrics provide different information from linkability metrics. Interestingly , the results on real data indicate that current anonymiza-tion design choices do not induce a regime where the differences between those metrics become apparent.
Mohamed Maouche, Brij Mohan Lal Srivastava, Nathalie Vauquier, Aurélien Bellet, Marc Tommasi, Emmanuel Vincent 0001
INTERSPEECH6
2020 Asteroid: The PyTorch-Based Audio Source Separation Toolkit for Researchers
abstract
This paper describes Asteroid, the PyTorch-based audio source separation toolkit for researchers. Inspired by the most successful neural source separation systems, it provides all neural building blocks required to build such a system. To improve reproducibility, Kaldi-style recipes on common audio source separation datasets are also provided. This paper describes the software architecture of Asteroid and its most important features. By showing experimental results obtained with Asteroid's recipes, we show that our implementations are at least on par with most results reported in reference papers. The toolkit is publicly available at https://github.com/mpariente/asteroid .
Manuel Pariente, Samuele Cornell, Joris Cosentino, Sunit Sivasankaran, Efthymios Tzinis, Jens Heitkaemper, Michel Olvera, Fabian-Robert Stöter, Mathieu Hu, Juan M. Martín-Doñas, David Ditter, Ariel Frank, Antoine Deleforge, Emmanuel Vincent 0001
INTERSPEECH14
2020 On Semi-Supervised LF-MMI Training of Acoustic Models with Limited Data
abstract
This work investigates semi-supervised training of acoustic models (AM) with the lattice-free maximum mutual information (LF-MMI) objective in practically relevant scenarios with a limited amount of labeled in-domain data. An error detection driven semi-supervised AM training approach is proposed, in which an error detector controls the hypothesized transcriptions or lattices used as LF-MMI training targets on additional unlabeled data. Under this approach, our first method uses a single error-tagged hypothesis whereas our second method uses a modified supervision lattice. These methods are evaluated and compared with existing semi-supervised AM training methods in three different matched or mismatched, limited data setups. Word error recovery rates of 28 to 89% are reported.
Imran A. Sheikh, Emmanuel Vincent 0001, Irina Illina
INTERSPEECH2
2020 Design Choices for X-Vector Based Speaker Anonymization
abstract
International audience
Brij Mohan Lal Srivastava, Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Junichi Yamagishi, Mohamed Maouche, Aurélien Bellet, Marc Tommasi
INTERSPEECH4
2020 Introducing the VoicePrivacy Initiative
abstract
The VoicePrivacy initiative aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a series of challenges. In this paper, we formulate the voice anonymization task selected for the VoicePrivacy 2020 Challenge and describe the datasets used for system development and evaluation. We also present the attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and report objective evaluation results.
Natalia A. Tomashenko, Brij Mohan Lal Srivastava, Xin Wang 0037, Emmanuel Vincent 0001, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Paul-Gauthier Noé, Massimiliano Todisco
INTERSPEECH4
2020 Achieving Multi-Accent ASR via Unsupervised Acoustic Model Adaptation
abstract
Current automatic speech recognition (ASR) systems trained on native speech often perform poorly when applied to non-native or accented speech. In this work, we propose to compute x-vector-like accent embeddings and use them as auxiliary inputs to an acoustic model trained on native data only in order to improve the recognition of multi-accent data comprising native, non-native, and accented speech. In addition, we leverage untranscribed accented training data by means of semi-supervised learning. Our experiments show that acoustic models trained with the proposed accent embeddings outperform those trained with conventional i-vector or x-vector speaker embeddings, and achieve a 15% relative word error rate (WER) reduction on non-native and accented speech w.r.t. acoustic models trained with regular spectral features only. Semi-supervised training using just 1 hour of untranscribed speech per accent yields an additional 15% relative WER reduction w.r.t. models trained on native data only.
M. A. Tugtekin Turan, Emmanuel Vincent 0001, Denis Jouvet
INTERSPEECH2
2020 Joint NN-Supported Multichannel Reduction of Acoustic Echo, Reverberation and Noise
abstract
We consider the problem of simultaneous reduction of acoustic echo, reverberation and noise. In real scenarios, these distortion sources may occur simultaneously and reducing them implies combining the corresponding distortion-specific filters. As these filters interact with each other, they must be jointly optimized. We propose to model the target and residual signals after linear echo cancellation and dereverberation using a multichannel Gaussian modeling framework and to jointly represent their spectra by means of a neural network. We develop an iterative block-coordinate ascent algorithm to update all the filters. We evaluate our system on real recordings of acoustic echo, reverberation and noise acquired with a smart speaker in various situations. The proposed approach outperforms in terms of overall distortion a cascade of the individual approaches and a joint reduction approach which does not rely on a spectral model of the target and residual signals.
Guillaume Carbajal, Romain Serizel, Emmanuel Vincent 0001, Eric Humbert
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Lead2Gold: Towards Exploiting the Full Potential of Noisy Transcriptions for Speech Recognition
abstract
The transcriptions used to train an Automatic Speech Recognition (ASR) system may contain errors. Usually, either a quality control stage discards transcriptions with too many errors, or the noisy transcriptions are used as is. We introduce Lead2Gold, a method to train an ASR system that exploits the full potential of noisy transcriptions. Based on a noise model of transcription errors, Lead2Gold searches for better transcriptions of the training data with a beam search that takes this noise model into account. The beam search is differentiable and does not require a forced alignment step, thus the whole system is trained end-to-end. Lead2Gold can be viewed as a new loss function that can be used on top of any sequence-to-sequence deep neural network. We conduct proof-of-concept experiments on noisy transcriptions generated from letter corruptions with different noise levels. We show that Lead2Gold obtains a better ASR accuracy than a competitive baseline which does not account for the (artificially-introduced) transcription noise.
Adrien Dufraux, Emmanuel Vincent 0001, Awni Y. Hannun, Armelle Brun, Matthijs Douze
ASRU2
2019 An Improved Uncertainty Propagation Method for Robust I-vector Based Speaker Recognition
abstract
The performance of automatic speaker recognition systems degrades when facing distorted speech data containing additive noise and/or reverberation. Statistical uncertainty propagation has been introduced as a promising paradigm to address this challenge. So far, different uncertainty propagation methods have been proposed to compensate noise and reverberation in i-vectors in the context of speaker recognition. They have achieved promising results on small datasets such as YOHO and Wall Street Journal, but little or no improvement on the larger, highly variable NIST Speaker Recognition Evaluation (SRE) corpus. In this paper, we propose a complete uncertainty propagation method, whereby we model the effect of uncertainty both in the computation of unbiased Baum-Welch statistics and in the derivation of the posterior expectation of the i-vector. We conduct experiments on the NIST-SRE corpus mixed with real domestic noise and reverberation from the CHiME-2 corpus and preprocessed by multichannel speech enhancement. The proposed method improves the equal error rate (EER) by 4% relative compared to a conventional i-vector based speaker verification baseline. This is to be compared with previous methods which degrade performance.
Dayana Ribas González, Emmanuel Vincent 0001
ICASSP2
2019 Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings
abstract
Deep neural networks are particularly useful to learn relevant representations from data. Recent studies have demonstrated the potential of unsupervised representation learning for ambient sound analysis using various flavors of the triplet loss. They have compared this approach to supervised learning. However, in real situations, it is common to have a small labeled dataset and a large unlabeled one. In this paper, we combine unsupervised and supervised triplet loss based learning into a semi-supervised representation learning approach. We propose two flavors of this approach, whereby the positive samples for those triplets whose anchors are unlabeled are obtained either by applying a transformation to the anchor, or by selecting the nearest sample in the training set. We compare our approach to supervised and unsupervised representation learning as well as the ratio between the amount of labeled and unlabeled data. We evaluate all the above approaches on an audio tagging task using the DCASE 2018 Task 4 dataset, and we show the impact of this ratio on the tagging performance.
Nicolas Turpault, Romain Serizel, Emmanuel Vincent 0001
ICASSP3
2019 A Statistically Principled and Computationally Efficient Approach to Speech Enhancement Using Variational Autoencoders
abstract
Recent studies have explored the use of deep generative models of speech spectra based of variational autoencoders (VAEs), combined with unsupervised noise models, to perform speech enhancement. These studies developed iterative algorithms involving either Gibbs sampling or gradient descent at each step, making them computationally expensive. This paper proposes a variational inference method to iteratively estimate the power spectrogram of the clean speech. Our main contribution is the analytical derivation of the variational steps in which the en-coder of the pre-learned VAE can be used to estimate the varia-tional approximation of the true posterior distribution, using the very same assumption made to train VAEs. Experiments show that the proposed method produces results on par with the afore-mentioned iterative methods using sampling, while decreasing the computational cost by a factor 36 to reach a given performance .
Manuel Pariente, Antoine Deleforge, Emmanuel Vincent 0001
INTERSPEECH3
2019 Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?
abstract
Automatic speech recognition (ASR) is a key technology in many services and applications. This typically requires user devices to send their speech data to the cloud for ASR decoding. As the speech signal carries a lot of information about the speaker, this raises serious privacy concerns. As a solution, an encoder may reside on each user device which performs local computations to anonymize the representation. In this paper, we focus on the protection of speaker identity and study the extent to which users can be recognized based on the encoded representation of their speech as obtained by a deep encoder-decoder architecture trained for ASR. Through speaker identification and verification experiments on the Librispeech corpus with open and closed sets of speakers, we show that the representations obtained from a standard architecture still carry a lot of information about speaker identity. We then propose to use adversarial training to learn representations that perform well in ASR while hiding speaker identity. Our results demonstrate that adversarial training dramatically reduces the closed-set classification accuracy, but this does not translate into increased open-set verification error hence into increased protection of the speaker identity in practice. We suggest several possible reasons behind this negative result.
Brij Mohan Lal Srivastava, Aurélien Bellet, Marc Tommasi, Emmanuel Vincent 0001
INTERSPEECH4
2019 VoiceHome-2, an extended corpus for multichannel speech processing in real homes
Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent 0001, Sunit Sivasankaran, Irina Illina, Frédéric Bimbot
Speech Commun.4
2019 Sound Event Detection in the DCASE 2017 Challenge
abstract
Each edition of the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) contained several tasks involving sound event detection in different setups. DCASE 2017 presented participants with three such tasks, each having specific datasets and detection requirements: Task 2, in which target sound events were very rare in both training and testing data, Task 3 having overlapping events annotated in real-life audio, and Task 4, in which only weakly labeled data were available for training. In this paper, we present three tasks, including the datasets and baseline systems, and analyze the challenge entries for each task. We observe the popularity of methods using deep neural networks, and the still widely used mel frequency-based representations, with only few approaches standing out as radically different. Analysis of the systems behavior reveals that task-specific optimization has a big role in producing good performance; however, often this optimization closely follows the ranking metric, and its maximization/minimization does not result in universally good performance. We also introduce the calculation of confidence intervals based on a jackknife resampling procedure to perform statistical analysis of the challenge results. The analysis indicates that while the 95% confidence intervals for many systems overlap, there are significant differences in performance between the top systems and the baseline for all tasks.
Annamaria Mesaros, Aleksandr Diment, Benjamin Elizalde, Toni Heittola, Emmanuel Vincent 0001, Bhiksha Raj, Tuomas Virtanen
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 Multiple-Input Neural Network-Based Residual Echo Suppression
abstract
A residual echo suppressor (RES) aims to suppress the residual echo in the output of an acoustic echo canceler (AEC). Spectral-based RES approaches typically estimate the magnitude spectra of the near-end speech and the residual echo from a single input, that is either the far-end speech or the echo computed by the AEC, and derive the RES filter coefficients accordingly. These single inputs do not always suffice to discriminate the near-end speech from the remaining echo. In this paper, we propose a neural network-based approach that directly estimates the RES filter coefficients from multiple inputs, including the AEC output, the far-end speech, and/or the echo computed by the AEC. We evaluate our system on real recordings of acoustic echo and near-end speech acquired in various situations with a smart speaker. We compare it to two single-input spectral-based approaches in terms of echo reduction and near-end speech distortion.
Guillaume Carbajal, Romain Serizel, Emmanuel Vincent 0001, Eric Humbert
ICASSP3
2018 Multichannel Speech Separation with Recurrent Neural Networks from High-Order Ambisonics Recordings
abstract
We present a source separation system for high-order ambisonics (HOA) contents. We derive a multichannel spatial filter from a mask estimated by a long short-term memory (LSTM) recurrent neural network. We combine one channel of the mixture with the outputs of basic HOA beamformers as inputs to the LSTM, assuming that we know the directions of arrival of the directional sources. In our experiments, the speech of interest can be corrupted either by diffuse noise or by an equally loud competing speaker. We show that adding as input the output of the beamformer steered toward the competing speech in addition to that of the beamformer steered toward the target speech brings significant improvements in terms of word error rate.
Lauréline Perotin, Romain Serizel, Emmanuel Vincent 0001, Alexandre Guérin
ICASSP3
2018 Semi-Supervised Learning with Deep Neural Networks for Relative Transfer Function Inverse Regression
abstract
Prior knowledge of the relative transfer function (RTF) is useful in many applications but remains little studied. In this paper, we propose a semi-supervised learning algorithm based on deep neural networks (DNNs) for RTF inverse regression, that is to generate the full-band RTF vector directly from the source-receiver pose (position and orientation). Two typical scenarios are discussed: training on labeled RTFs only, or on additional unlabeled RTFs. Both setups utilize the low-dimensional manifold property of RTF in stationary environments. With this property as an additional regularization term, a smooth mapping solution with respect to the manifold is obtained. Experimental simulations show that the proposed method achieves a lower mean prediction error than the free field model with few labeled RTFs, and the unlabeled RTFs are essential in improving the inverse regression performance.
Yonghong Yan 0002, Emmanuel Vincent 0001
ICASSP4
2018 The Fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, Task and Baselines
abstract
The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning. This paper introduces the 5th CHiME Challenge, which considers the task of distant multi-microphone conversational ASR in real home environments. Speech material was elicited using a dinner party scenario with efforts taken to capture data that is representative of natural conversational speech and recorded by 6 Kinect microphone arrays and 4 binaural microphone pairs. The challenge features a single-array track and a multiple-array track and, for each track, distinct rankings will be produced for systems focusing on robustness with respect to distant-microphone capture vs. systems attempting to address all aspects of the task including conversational language modeling. We discuss the rationale for the challenge and provide a detailed description of the data collection procedure, the task, and the baseline systems for array synchronization, speech enhancement, and conventional and end-to-end ASR.
Jon Barker, Shinji Watanabe 0001, Emmanuel Vincent 0001, Jan Trmal
INTERSPEECH3
2018 Keyword Based Speaker Localization: Localizing a Target Speaker in a Multi-speaker Environment
abstract
Speaker localization is a hard task, especially in adverse environmental conditions involving reverberation and noise. In this work we introduce the new task of localizing the speaker who uttered a given keyword, e.g., the wake-up word of a distant-microphone voice command system, in the presence of overlapping speech. We employ a convolutional neural network based localization system and investigate multiple identifiers as additional inputs to the system in order to characterize this speaker. We conduct experiments using ground truth identifiers which are obtained assuming the availability of clean speech and also in realistic conditions where the identifiers are computed from the corrupted speech. We find that the identifier consisting of the ground truth time-frequency mask corresponding to the target speaker provides the best localization performance and we propose methods to estimate such a mask in adverse reverberant and noisy conditions using the considered keyword.
Sunit Sivasankaran, Emmanuel Vincent 0001, Dominique Fohr
INTERSPEECH2
2018 Rank-1 constrained Multichannel Wiener Filter for speech recognition in noisy environments
Emmanuel Vincent 0001, Romain Serizel, Yonghong Yan 0002
Comput. Speech Lang.2
2018 DNN Uncertainty Propagation Using GMM-Derived Uncertainty Features for Noise Robust ASR
abstract
The uncertainty decoding framework is known to improve the deep neural network (DNN)-based automatic speech recognition (ASR) performance in noisy environments. It operates by estimating the statistical uncertainty about the input features and propagating it to the output senone posteriors by sampling. Unfortunately, this approximate propagation scheme limits the performance improvement. In this letter, we exploit the fact that uncertainty propagation can be achieved in closed form for Gaussian mixture acoustic models (GMMs). We introduce new GMM-derived (GMMD) uncertainty features for the robust DNN-based acoustic model training and decoding. The GMMD features are computed as the difference between the GMM log-likelihoods obtained with versus without uncertainty. They are concatenated with conventional acoustic features and used as inputs to the DNN. We evaluate the resulting ASR performance on the CHiME-2 and CHiME-3 datasets. The proposed features are shown to improve the performance on both datasets, both for the conventional decoding and for the uncertainty decoding with different uncertainty estimation/propagation techniques.
Karan Nathwani, Emmanuel Vincent 0001, Irina Illina
IEEE Signal Process. Lett.2
2017 Consistent DNN uncertainty training and decoding for robust ASR
abstract
We consider the problem of robust automatic speech recognition (ASR) in noisy conditions. The performance improvement brought by speech enhancement is often limited by residual distortions of the enhanced features, which can be seen as a form of statistical uncertainty. Uncertainty estimation and propagation methods have recently been proposed to improve the ASR performance with deep neural network (DNN) acoustic models. However, the performance is still limited due to the use of uncertainty only during decoding. In this paper, we propose a consistent approach to account for uncertainty in the enhanced features during both training and decoding. We estimate the variance of the distortions using a DNN uncertainty estimator that operates directly in the feature maximum likelihood linear regression (fMLLR) domain and we then sample the uncertain features using the unscented transform (UT). We report the resulting ASR performance on the CHiME-2 and CHiME-3 datasets for different uncertainty estimation/propagation techniques. The proposed DNN uncertainty training method brings 4% and 8% relative improvement on these two datasets, respectively, compared to a competitive fMLLR-domain DNN acoustic modeling baseline.
Karan Nathwani, Emmanuel Vincent 0001, Irina Illina
ASRU2
2017 Recursive Bayesian estimation of the acoustic noise emitted by wind farms
abstract
Wind turbine noise is often annoying for humans living in close proximity to a wind farm. Reliably estimating the intensity of wind turbine noise is a necessary step towards quantifying and reducing annoyance, but it is challenging because of the overlap with background noise sources. Current approaches involve measurements with on/off turbine cycles and acoustic simulations, which are expensive and unreliable. This raises the problem of separating the noise of wind turbines from that of background noise sources and coping with the uncertainties associated with the source separation output. In this paper we propose to assist a black-box source separation system with a model of wind turbine noise emission and propagation in a recursive Bayesian estimation framework. We validate our approach on real data with simulated uncertainties using different nonlinear Kalman filters.
Baldwin Dumortier, Emmanuel Vincent 0001, Madalina Deaconu
ICASSP2
2017 Discriminative importance weighting of augmented training data for acoustic model training
abstract
DNN based acoustic models require a large amount of training data. Parametric data augmentation techniques such as adding noise, reverberation, or changing the speech rate, are often employed to boost the dataset size and the ASR performance. The choice of augmentation techniques and the associated parameters has been handled heuristically so far. In this work we propose an algorithm to automatically weight data perturbed using a variety of augmentation techniques and/or parameters. The weights are learned in a discriminative fashion so as to minimize the frame error rate using the standard gradient descent algorithm in an iterative manner. Experiments were performed using the CHiME-3 dataset. Data augmentation was done by adding noise at different SNRs. A relative WER improvement of 15% was obtained with the proposed data weighting algorithm compared to the unweighted augmented dataset. Interestingly, the resulting distribution of SNRs in the weighted training set differs significantly from that of the test set.
Sunit Sivasankaran, Emmanuel Vincent 0001, Irina Illina
ICASSP2
2017 Multi-microphone speech recognition in everyday environments
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001
Comput. Speech Lang.3
2017 The third 'CHiME' speech separation and recognition challenge: Analysis and outcomes
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001
Comput. Speech Lang.3
2017 A combined evaluation of established and new approaches for speech recognition in varied reverberation conditions
Sunit Sivasankaran, Emmanuel Vincent 0001, Irina Illina
Comput. Speech Lang.2
2017 An analysis of environment, microphone and data simulation mismatches in robust speech recognition
Emmanuel Vincent 0001, Shinji Watanabe 0001, Aditya Arie Nugraha, Jon Barker, Ricard Marxer
Comput. Speech Lang.1
2017 A Consolidated Perspective on Multimicrophone Speech Enhancement and Source Separation
abstract
Speech enhancement and separation are core problems in audio signal processing, with commercial applications in devices as diverse as mobile phones, conference call systems, hands-free systems, or hearing aids. In addition, they are crucial preprocessing steps for noise-robust automatic speech and speaker recognition. Many devices now have two to eight microphones. The enhancement and separation capabilities offered by these multichannel interfaces are usually greater than those of single-channel interfaces. Research in speech enhancement and separation has followed two convergent paths, starting with microphone array processing and blind source separation, respectively. These communities are now strongly interrelated and routinely borrow ideas from each other. Yet, a comprehensive overview of the common foundations and the differences between these approaches is lacking at present. In this paper, we propose to fill this gap by analyzing a large number of established and recent techniques according to four transverse axes: 1) the acoustic impulse response model, 2) the spatial filter design criterion, 3) the parameter estimation algorithm, and 4) optional postfiltering. We conclude this overview paper by providing a list of software and data resources and by discussing perspectives and future trends in the field.
Sharon Gannot, Emmanuel Vincent 0001, Shmulik Markovich-Golan, Alexey Ozerov
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Estimating the Structural Segmentation of Popular Music Pieces Under Regularity Constraints
abstract
Music structure estimation has recently emerged as a central topic within the field of music information retrieval. Indeed, as music is a highly structured information stream, knowledge of how a music piece is organized represents a key challenge to enhance the management and exploitation of large music collections. This paper focuses on the benefits that can be expected from a regularity constraint on the structural segmentation of popular music pieces. Specifically, here, we study how a constraint that favors structural segments of comparable size provides a better conditioning of the boundary estimation process. First, we propose a formulation of the structural segmentation task as an optimization process, which separates the contribution from the audio features and the one from the constraint. We illustrate how the corresponding cost function can be minimized using a Viterbi algorithm. We present briefly its implementation and results in three systems designed for and submitted to the MIREX 2010, 2011, and 2012 evaluation campaigns. Then, we explore the benefits of the regularity constraint as an efficient mean for combining the outputs of a selection of systems presented at MIREX between 2010 and 2015, yielding a level of performance competitive to that of the state-of-the-art on the “MIREX10” dataset (100 J-Pop songs from the RWC database).
Gabriel Sargent, Frédéric Bimbot, Emmanuel Vincent 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 A French Corpus for Distant-Microphone Speech Processing in Real Homes
abstract
International audience
Nancy Bertin, Ewen Camberlein, Emmanuel Vincent 0001, Romain Lebarbenchon, Stéphane Peillon, Éric Lamande, Sunit Sivasankaran, Frédéric Bimbot, Irina Illina, Ariane Tom, Sylvain Fleury, Eric Jamet
INTERSPEECH3
2016 Discussion
Dayana Ribas González, Emmanuel Vincent 0001, John H. L. Hansen, Emma Jokinen, Mirco Ravanelli, Hannes Gamper, Fred Richardson
INTERSPEECH2
2016 Localizing an intermittent and moving sound source using a mobile robot
abstract
This paper addresses the problem of localizing and tracking one intermittent, moving sound source using a microphone array on a mobile robot. Robot motion provides a solution for estimating the distance to the source and avoiding front-back ambiguity. We propose a mixture Kalman filter (MKF) framework in order to fuse the robot motion information and the measurements taken at different poses of the robot. Experiments and statistical results demonstrate the ability of the proposed method to track one intermittent sound source in a reverberant environment where false measurements of the source angle of arrival (AoA) and the source activity often occur compared to a method that does not consider tracking source activity into account.
Quan V. Nguyen, Francis Colas, Emmanuel Vincent 0001, François Charpillet
IROS3
2016 A study of speech distortion conditions in real scenarios for speech processing applications
abstract
The growing demand for robust speech processing applications able to operate in adverse scenarios calls for new evaluation protocols and datasets beyond artificial laboratory conditions. The characteristics of real data for a given scenario are rarely discussed in the literature. As a result, methods are often tested based on the author expertise and not always in scenarios with actual practical value. This paper aims to open this discussion by identifying some of the main problems with data simulation or collection procedures used so far and summarizing the important characteristics of real scenarios to be taken into account, including the properties of reverberation, noise and Lombard effect. At last, we provide some preliminary guidelines towards designing experimental setup and speech recognition results for proposal validation.
Dayana Ribas González, Emmanuel Vincent 0001, José Ramón Calvo de Lara
SLT2
2016 Variational Bayesian Inference for Source Separation and Robust Feature Extraction
abstract
We consider the task of separating and classifying individual sound sources mixed together. The main challenge is to achieve robust classification despite residual distortion of the separated source signals. A promising paradigm is to estimate the uncertainty about the separated source signals and to propagate it through the subsequent feature extraction and classification stages. We argue that variational Bayesian (VB) inference offers a mathematically rigorous way of deriving uncertainty estimators, which contrasts with state-of-the-art estimators based on heuristics or on maximum likelihood (ML) estimation. We propose a general VB source separation algorithm, which makes it possible to jointly exploit spatial and spectral models of the sources. This algorithm achieves 6% and 5% relative error reduction compared to ML uncertainty estimation on the CHiME noise-robust speaker identification and speech recognition benchmarks, respectively, and it opens the way for more complex VB approximations of uncertainty.
Kamil Adiloglu, Emmanuel Vincent 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Fusion Methods for Speech Enhancement and Audio Source Separation
abstract
A wide variety of audio source separation techniques exist and can already tackle many challenging industrial issues. However, in contrast with other application domains, fusion principles were rarely investigated in audio source separation despite their demonstrated potential in classification tasks. In this paper, we propose a general fusion framework which takes advantage of the diversity of existing separation techniques in order to improve separation quality. We obtain new source estimates by summing the individual estimates given by different separation techniques weighted by a set of fusion coefficients. We investigate three alternative fusion methods which are based on standard nonlinear optimization, Bayesian model averaging, or deep neural networks. Experiments conducted for both speech enhancement and singing voice extraction demonstrate that all the proposed methods outperform traditional model selection. The use of deep neural networks for the estimation of time-varying coefficients notably leads to large quality improvements, up to 3 dB in terms of signal-to-distortion ratio compared to model selection.
Xabier Jaureguiberry, Emmanuel Vincent 0001, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Multichannel Audio Source Separation With Deep Neural Networks
abstract
This article addresses the problem of multichannel audio source separation. We propose a framework where deep neural networks (DNNs) are used to model the source spectra and combined with the classical multichannel Gaussian model to exploit the spatial information. The parameters are estimated in an iterative expectation-maximization (EM) fashion and used to derive a multichannel Wiener filter. We present an extensive experimental study to show the impact of different design choices on the performance of the proposed technique. We consider different cost functions for the training of DNNs, namely the probabilistically motivated Itakura-Saito divergence, and also Kullback-Leibler, Cauchy, mean squared error, and phase-sensitive cost functions. We also study the number of EM iterations and the use of multiple DNNs, where each DNN aims to improve the spectra estimated by the preceding EM iteration. Finally, we present its application to a speech enhancement problem. The experimental results show the benefit of the proposed multichannel approach over a single-channel DNN-based approach and the conventional multichannel nonnegative matrix factorization-based iterative EM algorithm.
Aditya Arie Nugraha, Antoine Liutkus, Emmanuel Vincent 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 The third 'CHiME' speech separation and recognition challenge: Dataset, task and baselines
abstract
The CHiME challenge series aims to advance far field speech recognition technology by promoting research at the interface of signal processing and automatic speech recognition. This paper presents the design and outcomes of the 3rd CHiME Challenge, which targets the performance of automatic speech recognition in a real-world, commercially-motivated scenario: a person talking to a tablet device that has been fitted with a six-channel microphone array. The paper describes the data collection, the task definition and the baseline systems for data simulation, enhancement and recognition. The paper then presents an overview of the 26 systems that were submitted to the challenge focusing on the strategies that proved to be most successful relative to the MVDR array processing and DNN acoustic modeling reference system. Challenge findings related to the role of simulated data in system training and evaluation are discussed.
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001
ASRU3
2015 Robust ASR using neural network based speech enhancement and feature simulation
abstract
We consider the problem of robust automatic speech recognition (ASR) in the context of the CHiME-3 Challenge. The proposed system combines three contributions. First, we propose a deep neural network (DNN) based multichannel speech enhancement technique, where the speech and noise spectra are estimated using a DNN based regressor and the spatial parameters are derived in an expectation-maximization (EM) like fashion. Second, a conditional restricted Boltzmann machine (CRBM) model is trained using the obtained enhanced speech and used to generate simulated training and development datasets. The goal is to increase the similarity between simulated and real data, so as to increase the benefit of multicondition training. Finally, we make some changes to the ASR backend. Our system ranked 4th among 25 entries.
Sunit Sivasankaran, Aditya Arie Nugraha, Emmanuel Vincent 0001, Juan Andres Morales-Cordovilla, Siddharth Dalmia, Irina Illina, Antoine Liutkus
ASRU3
2015 Micbots: Collecting large realistic datasets for speech and audio research using mobile robots
abstract
Speech and audio signal processing research is a tale of data collection efforts and evaluation campaigns. Large benchmark datasets for automatic speech recognition (ASR) have been instrumental in the advancement of speech recognition technologies. However, when it comes to robust ASR, source separation, and localization, especially using microphone arrays, the perfect dataset is out of reach, and many different data collection efforts have each made different compromises between the conflicting factors in terms of realism, ground truth, and costs. Our goal here is to escape some of the most difficult trade-offs by proposing MICbots, a low-cost method of collecting large amounts of realistic data where annotations and ground truth are readily available. Our key idea is to use freely moving robots equiped with microphones and loudspeakers, playing recorded utterances from existing (already annotated) speech datasets. We give an overview of previous data collection efforts and the trade-offs they make, and describe the benefits of using our robot-based approach. We finally explain the use of this method to collect room impulse response measurement.
Jonathan Le Roux, Emmanuel Vincent 0001, John R. Hershey, Daniel P. W. Ellis
ICASSP2
2015 Music separation guided by cover tracks: Designing the joint NMF model
abstract
In audio source separation, reference guided approaches are a class of methods that use reference signals to guide the separation. In prior work, we proposed a general framework to model the deformation between the sources and the references. In this paper, we investigate a specific scenario within this framework: music separation guided by the multitrack recording of a cover interpretation of the song to be processed. We report a series of experiments highlighting the relevance of joint Non-negative Matrix Factorization (NMF), dictionary transformation, and specific transformation models for different types of sources. A signal-to-distortion ratio improvement (SDRI) of almost 11 decibels (dB) is achieved, improving by 2 dB compared to previous study on the same data set. These observations contribute to validate the relevance of the theoretical general framework and can be useful in practice for designing models for other reference guided source separation problems.
Nathan Souviraà-Labastie, Emmanuel Vincent 0001, Frédéric Bimbot
ICASSP2
2015 Fast DNN training based on auxiliary function technique
abstract
Deep neural networks (DNN) are typically optimized with stochastic gradient descent (SGD) using a fixed learning rate or an adaptive learning rate approach (ADAGRAD). In this paper, we introduce a new learning rule for neural networks that is based on an auxiliary function technique without parameter tuning. Instead of minimizing the objective function, a quadratic auxiliary function is recursively introduced layer by layer which has a closed-form optimum. We prove the monotonic decrease of the new learning rule. Our experiments show that the proposed algorithm converges faster and to a better local minimum than SGD. In addition, we propose a combination of the proposed learning rule and ADAGRAD which further accelerates convergence. Experimental evaluation on the MNIST database shows the benefit of the proposed approach in terms of digit recognition accuracy.
Dung T. Tran, Nobutaka Ono, Emmanuel Vincent 0001
ICASSP3
2015 Discriminative uncertainty estimation for noise robust ASR
abstract
We consider the problem of uncertainty estimation for noiserobust ASR. Existing uncertainty estimation techniques improve ASR accuracy but they still exhibit a gap compared to the use of oracle uncertainty. This comes partly from the highly non-linear feature transformation and from additional assumptions such as Gaussian distribution and independence between frequency bins in the spectral domain. In this paper, we propose a method to rescale the estimated feature-domain full uncertainty covariance matrix in a statedependent fashion according to a discriminative criterion. The state-dependent and feature index-dependent scaling factors are learned from development data. Experimental evaluation on Track 1 of the 2nd CHiME challenge data shows that discriminative rescaling leads to better results than generative rescaling. Moreover, discriminative rescaling of the Wiener uncertainty estimator leads to 12% relative word error rate reduction compared to discriminative rescaling of the alternative estimator in [1].
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
ICASSP2
2015 Audio source localization by optimal control of a mobile robot
abstract
We consider the task of audio source localization using a microphone array on a mobile robot. Active localization algorithms have been proposed in the literature that can estimate the 3D position of a source by fusing the measurements taken for different poses of the robot. The robot movements are typically fixed, however, or they obey heuristic strategies, such as turning the head and moving towards the source, which may be suboptimal. In this paper, we propose to control the robot movements so as to locate the source as quickly as possible. We represent the belief about the source position by a discrete grid and we introduce a dynamic programming algorithm to find the optimal robot motion minimizing the entropy of the grid. We report initial results in a real environment.
Emmanuel Vincent 0001, Aghilas Sini, François Charpillet
ICASSP1
2015 Uncertainty propagation through deep neural networks
abstract
In order to improve the ASR performance in noisy environments, distorted speech is typically pre-processed by a speech enhancement algorithm, which usually results in a speech estimate containing residual noise and distortion.We may also have some measures of uncertainty or variance of the estimate.Uncertainty decoding is a framework that utilizes this knowledge of uncertainty in the input features during acoustic model scoring.Such frameworks have been well explored for traditional probabilistic models, but their optimal use for deep neural network (DNN)-based ASR systems is not yet clear.In this paper, we study the propagation of observation uncertainties through the layers of a DNN-based acoustic model.Since this is intractable due to the nonlinearities of the DNN, we employ approximate propagation methods, including Monte Carlo sampling, the unscented transform, and the piecewise exponential approximation of the activation function, to estimate the distribution of acoustic scores.Finally, the expected value of the acoustic score distribution is used for decoding, which is shown to further improve the ASR accuracy on the CHiME database, relative to a highly optimized DNN baseline.
Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, John R. Hershey, Emmanuel Vincent 0001, Dorothea Kolossa
INTERSPEECH4
2015 Uncertainty propagation for noise robust speaker recognition: the case of NIST-SRE
abstract
Uncertainty propagation is an established approach to handle noisy and reverberant conditions in automatic speech recognition (ASR), but it has little been studied for speaker recognition so far. Yu et al. recently proposed to propagate uncertainty to the Baum-Welch (BW) statistics without changing the posterior probability of each mixture component. They obtained good results on a small dataset (YOHO) but little improvement on the NIST-SRE dataset, despite the use of oracle uncertainty estimates. In this paper, we propose to modify the computation of the posterior probability of each mixture component in order to obtain unbiased BW statistics. We show that our approach improves the accuracy of BW statistics on the Wall Street Journal (WSJ) corpus, but yields little or no improvement on NIST-SRE again. We provide a theoretical explanation for this that opens the way for more efficient exploitation of uncertainty on NIST-SRE and other large datasets in the future.
Dayana Ribas González, Emmanuel Vincent 0001, José Ramón Calvo de Lara
INTERSPEECH2
2015 Full multicondition training for robust i-vector based speaker recognition
abstract
Multicondition training (MCT) is an established technique to handle noisy and reverberant conditions. Previous works in the field of i-vector based speaker recognition have applied MCT to linear discriminant analysis (LDA) and probabilistic LDA (PLDA), but not to the universal background model (UBM) and the total variability (T) matrix, arguing that this would be too much time consuming due to the increase of the size of the training set by the number of noise and reverberation conditions. In this paper, we propose a full MCT approach which consists of applying MCT in all stages of training, including the UBM and the T matrix, while keeping the size of the training set fixed. Experiments in highly nonstationary noise conditions show a decrease of the equal error rate (EER) to 14.16% compared to 17.90% for clean training and 18.08% for MCT of LDA and PLDA only. We also evaluate the impact of state-of-the-art multichannel speech enhancement and show further reduction of the EER down to 10.47%.
Dayana Ribas González, Emmanuel Vincent 0001, José Ramón Calvo de Lara
INTERSPEECH2
2015 Multi-Channel Audio Source Separation Using Multiple Deformed References
abstract
We present a general multi-channel source separation framework where additional audio references are available for one (or more) source(s) of a given mixture. Each audio reference is another mixture which is supposed to contain at least one source similar to one of the target sources. Deformations between the sources of interest and their references are modeled in a linear manner using a generic formulation. This is done by adding transformation matrices to an excitation-filter model, hence affecting different axes, namely frequency, dictionary component or time. A nonnegative matrix co-factorization algorithm and a generalized expectation-maximization algorithm are used to estimate the parameters of the model. Different model parameterizations and different combinations of algorithms are tested on music plus voice mixtures guided by music and/or voice references and on professionally-produced music recordings guided by cover references. Our algorithms improve the signal-to-distortion ratio (SDR) of the sources with the lowest intensity by 9 to 15 decibels (dB) with respect to original mixtures.
Nathan Souviraà-Labastie, Anaïk Olivero, Emmanuel Vincent 0001, Frédéric Bimbot
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Nonparametric Uncertainty Estimation and Propagation for Noise Robust ASR
abstract
We consider the framework of uncertainty propagation for automatic speech recognition (ASR) in highly nonstationary noise environments. Uncertainty is considered as the variance of speech distortion. Yet, its accurate estimation in the spectral domain and its propagation to the feature domain remain difficult. Existing methods typically rely on a single uncertainty estimator and propagator fixed by mathematical approximation. In this paper, we propose a new paradigm where we seek to learn more powerful mappings to predict uncertainty from data. We investigate two such possible mappings: linear fusion of multiple uncertainty estimators/propagators and nonparametric uncertainty estimation/propagation. In addition, a procedure to propagate the estimated spectral-domain uncertainty to the static Mel frequency cepstral coefficients (MFCCs), to the log-energy, and to their first- and second-order time derivatives is proposed. This results in a full uncertainty covariance matrix over both static and dynamic MFCCs. Experimental evaluation on Tracks 1 and 2 of the 2nd CHiME Challenge resulted in up to 29% and 28% relative keyword error rate reduction with respect to speech enhancement alone.
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Blind RT60 estimation robust across room sizes and source distances
abstract
The reverberation time or RT60 is an essential acoustic parameter of a room. In many situations, the room impulse response (RIR) is not available and the RT60 must be blindly estimated from a speech or music signal. Current methods often implicitly assume that reverberation dominates direct sound, which restricts their applicability to relatively small rooms or distant sound sources. This paper features two contributions. Firstly, we propose a blind RT60 estimation method that is independent of the room size and the source distance by preprocessing the input signal using a beamformer to cancel direct sound and early echoes. Secondly, we perform the largest experimental evaluation to our knowledge using a set of 342 RIRs. We show that the estimation error is significantly reduced even in the case when reverberation dominates.
Baldwin Dumortier, Emmanuel Vincent 0001
ICASSP2
2014 Extension of uncertainty propagation to dynamic MFCCS for noise robust ASR
abstract
Uncertainty propagation has been successfully employed for speech recognition in nonstationary noise environments. The uncertainty about the features is typically represented as a diagonal covariance matrix for static features only. We present a framework for estimating the uncertainty over both static and dynamic features as a full covariance matrix. The estimated covariance matrix is then multiplied by scaling coefficients optimized on development data. We achieve 21% relative error rate reduction on the 2nd CHiME Challenge with respect to conventional decoding without uncertainty, that is five times more than the reduction achieved with diagonal uncertainty covariance for static features only.
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
ICASSP2
2014 Fusion of multiple uncertainty estimators and propagators for noise robust ASR
abstract
Uncertainty decoding has been successfully used for speech recognition in highly nonstationary noise environments. Yet, accurate estimation of the uncertainty on the denoised signals and propagation to the features remain difficult. In this work, we propose to fuse the uncertainty estimates obtained from different uncertainty estimators and propagators by linear combination. The fusion coefficients are optimized by minimizing a measure of divergence with oracle estimates on development data. Using the Kullback-Leibler divergence, we obtain 18% relative error rate reduction on the 2nd CHiME Challenge with respect to conventional decoding, that is about twice as much as the reduction achieved by the best single uncertainty estimator and propagator.
Dung T. Tran, Emmanuel Vincent 0001, Denis Jouvet
ICASSP2
2014 Multiple-order non-negative matrix factorization for speech enhancement
abstract
Amongst the speech enhancement techniques, statistical models based on Non-negative Matrix Factorization (NMF) have received great attention. In a single channel configuration, NMF is used to describe the spectral content of both the speech and noise sources. As the number of components can have a crucial influence on separation quality, we here propose to investigate model order selection based on the variational Bayesian approximation to the marginal likelihood of models of different orders. To go further, we propose to use model averaging to combine several single-order NMFs and we show that a straightforward application of model averaging principles is inefficient as it turned out to be equivalent to model selection. We thus introduce a parameter to control the entropy of the model order distribution which makes the averaging effective. We also show that our probabilistic model nicely extends to a multiple-order NMF model where several NMFs are jointly estimated and averaged. Experiments are conducted on real data from the CHiME challenge and give an interesting insight on the entropic parameter and model order priors. Separation results are also promising as model averaging outperforms single-order model selection. Finally, our multiple-order NMF shows an interesting gain in computation time.
Xabier Jaureguiberry, Emmanuel Vincent 0001, Gaël Richard
INTERSPEECH2
2014 An investigation of likelihood normalization for robust ASR
abstract
International audience
Emmanuel Vincent 0001, Aggelos Gkiokas, Dominik Schnitzer, Arthur Flexer
INTERSPEECH1
2014 Genre-Based Music Language Modeling with Latent Hierarchical Pitman-Yor Process Allocation
abstract
In this work we present a new Bayesian topic model: latent hierarchical Pitman-Yor process allocation (LHPYA), which uses hierarchical Pitman-Yor process priors for both word and topic distributions, and generalizes a few of the existing topic models, including the latent Dirichlet allocation (LDA), the bigram topic model and the hierarchical Pitman-Yor topic model. Using such priors allows for integration of n-grams with a topic model, while smoothing them with the state-of-the-art method. Our model is evaluated by measuring its perplexity on a dataset of musical genre and harmony annotations 3 Genre Database (3GDB) and by measuring its ability to predict musical genre from chord sequences. In terms of perplexity, for a 262-chord dictionary we achieve a value of 2.74, compared to 18.05 for trigrams and 7.73 for a unigram topic model. In terms of genre prediction accuracy with 9 genres, the proposed approach performs about 33% better in relative terms than genre-dependent n-grams, achieving 60.4% of accuracy.
Stanislaw Andrzej Raczynski, Emmanuel Vincent 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 The second 'CHiME' speech separation and recognition challenge: An overview of challenge systems and outcomes
abstract
Distant-microphone automatic speech recognition (ASR) remains a challenging goal in everyday environments involving multiple background sources and reverberation. This paper reports on the results of the 2nd ‘CHiME’ Challenge, an initiative designed to analyse and evaluate the performance of ASR systems in a real-world domestic environment. We discuss the rationale for the challenge and provide a summary of the datasets, tasks and baseline systems. The paper overviews the systems that were entered for the two challenge tracks: small-vocabulary with moving talker and medium-vocabulary with stationary talker. We present a summary of the challenge findings including novel results produced by challenge system combination. Possible directions for future challenges are discussed.
Emmanuel Vincent 0001, Jon Barker, Shinji Watanabe 0001, Jonathan Le Roux, Francesco Nesta, Marco Matassoni
ASRU1
2013 A fundamental pitfall in blind deconvolution with sparse and shift-invariant priors
abstract
We consider the problem of blind sparse deconvolution, which is common in both image and signal processing. To counter-balance the ill-posedness of the problem, many approaches are based on the minimization of a cost function. A well-known issue is a tendency to converge to an undesirable trivial solution. Besides domain specific explanations (such as the nature of the spectrum of the blurring filter in image processing) a widespread intuition behind this phenomenon is related to scaling issues and the nonconvexity of the optimized cost function. We prove that a fundamental issue lies in fact in the intrinsic properties of the cost function itself: for a large family of shift-invariant cost functions promoting the sparsity of either the filter or the source, the only global minima are trivial. We complete the analysis with an empirical method to verify the existence of more useful local minima.
Alexis Benichoux, Emmanuel Vincent 0001, Rémi Gribonval
ICASSP2
2013 The second 'chime' speech separation and recognition challenge: Datasets, tasks and baselines
abstract
Distant-microphone automatic speech recognition (ASR) remains a challenging goal in everyday environments involving multiple background sources and reverberation. This paper is intended to be a reference on the 2nd `CHiME' Challenge, an initiative designed to analyze and evaluate the performance of ASR systems in a real-world domestic environment. Two separate tracks have been proposed: a small-vocabulary task with small speaker movements and a medium-vocabulary task without speaker movements. We discuss the rationale for the challenge and provide a detailed description of the datasets, tasks and baseline performance results for each track.
Emmanuel Vincent 0001, Jon Barker, Shinji Watanabe 0001, Jonathan Le Roux, Francesco Nesta, Marco Matassoni
ICASSP1
2013 MODIS: an audio motif discovery software
Laurence Catanese, Nathan Souviraà-Labastie, Bingqing Qu, Sébastien Campion, Guillaume Gravier, Emmanuel Vincent 0001, Frédéric Bimbot
INTERSPEECH6
2013 Special issue on speech separation and recognition in multisource environments
Jon Barker, Emmanuel Vincent 0001
Comput. Speech Lang.2
2013 The PASCAL CHiME speech separation and recognition challenge
Jon Barker, Emmanuel Vincent 0001, Ning Ma 0002, Heidi Christensen, Phil D. Green
Comput. Speech Lang.2
2013 Uncertainty-based learning of acoustic models from noisy data
Alexey Ozerov, Mathieu Lagrange, Emmanuel Vincent 0001
Comput. Speech Lang.3
2013 Consistent Wiener Filtering for Audio Source Separation
abstract
Wiener filtering is one of the most ubiquitous tools in signal processing, in particular for signal denoising and source separation. In the context of audio, it is typically applied in the time-frequency domain by means of the short-time Fourier transform (STFT). Such processing does generally not take into account the relationship between STFT coefficients in different time-frequency bins due to the redundancy of the STFT, which we refer to as consistency. We propose to enforce this relationship in the design of the Wiener filter, either as a hard constraint or as a soft penalty. We derive two conjugate gradient algorithms for the computation of the filter coefficients and show improved audio source separation performance compared to the classical Wiener filter both in oracle and in blind conditions.
Jonathan Le Roux, Emmanuel Vincent 0001
IEEE Signal Process. Lett.2
2013 Dynamic Bayesian Networks for Symbolic Polyphonic Pitch Modeling
abstract
Symbolic pitch modeling is a way of incorporating knowledge about relations between pitches into the process of analyzing musical information or signals. In this paper, we propose a family of probabilistic symbolic polyphonic pitch models, which account for both the “horizontal” and the “vertical” pitch structure. These models are formulated as linear or log-linear interpolations of up to five sub-models, each of which is responsible for modeling a different type of relation. The ability of the models to predict symbolic pitch data is evaluated in terms of their cross-entropy, and of a newly proposed “contextual cross-entropy” measure. Their performance is then measured on synthesized polyphonic audio signals in terms of the accuracy of multiple pitch estimation in combination with a Nonnegative Matrix Factorization-based acoustic model. In both experiments, the log-linear combination of at least one “vertical” (e.g., harmony) and one “horizontal” (e.g., note duration) sub-model outperformed a pitch-dependent Bernoulli prior by more than 60% in relative cross-entropy and 3% in absolute multiple pitch estimation accuracy. This work provides a proof of concept of the usefulness of model interpolation, which may be used for improved symbolic modeling of other aspects of music in the future.
Stanislaw Andrzej Raczynski, Emmanuel Vincent 0001, Shigeki Sagayama
IEEE Trans. Speech Audio Process.2
2012 A general variational Bayesian framework for robust feature extraction in multisource recordings
abstract
We consider the problem of extracting features from individual sources in a multisource audio recording using a general source separation algorithm. The main issue is to estimate and propagate the uncertainty over the separated source signals, so as to robustly estimate the features despite source separation errors. While state-of-the-art techniques estimate the uncertainty in a heuristic manner, we propose to integrate over the parameter space of the source separation algorithm. We apply variational Bayes to estimate the posterior probability of the sources and subsequently derive the expectation of the features by moment matching. Experiments over stereo mixtures of three or four sources show that the proposed method provides the best results in terms of the root mean square (RMS) error on the estimated features.
Kamil Adiloglu, Emmanuel Vincent 0001
ICASSP2
2012 Multi-source TDOA estimation in reverberant audio using angular spectra and clustering
Charles Blandin, Alexey Ozerov, Emmanuel Vincent 0001
Signal Process.3
2012 Latent variable analysis and signal separation
Vincent Vigneron, Vicente Zarzoso, Rémi Gribonval, Emmanuel Vincent 0001
Signal Process.4
2012 The signal separation evaluation campaign (2007-2010): Achievements and remaining challenges
Emmanuel Vincent 0001, Shoko Araki, Fabian J. Theis, Guido Nolte, Pau Bofill, Hiroshi Sawada, Alexey Ozerov, Vikrham Gowreesunker, Dominik Lutter, Ngoc Q. K. Duong
Signal Process.1
2012 A General Flexible Framework for the Handling of Prior Information in Audio Source Separation
abstract
Most audio source separation methods are developed for a particular scenario characterized by the number of sources and channels and the characteristics of the sources and the mixing process. In this paper, we introduce a general audio source separation framework based on a library of structured source models that enable the incorporation of prior knowledge about each source via user-specifiable constraints. While this framework generalizes several existing audio source separation methods, it also allows to imagine and implement new efficient methods that were not yet reported in the literature. We first introduce the framework by describing the model structure and constraints, explaining its generality, and summarizing its algorithmic implementation using a generalized expectation-maximization algorithm. Finally, we illustrate the above-mentioned capabilities of the framework by applying it in several new and existing configurations to different source separation problems. We have released a software tool named Flexible Audio Source Separation Toolbox (FASST) implementing a baseline version of the framework in Matlab.
Alexey Ozerov, Emmanuel Vincent 0001, Frédéric Bimbot
IEEE Trans. Speech Audio Process.2
2011 Stability analysis of multiplicative update algorithms for non-negative matrix factorization
abstract
Multiplicative update algorithms have encountered a great success to solve optimization problems with non-negativity constraints, such as the famous non-negative matrix factorization (NMF) and its many variants. However, despite several years of research on the topic, the understanding of their convergence properties is still to be improved. In this paper, we show that Lyapunov's stability theory provides a very enlightening viewpoint on the problem. We prove the stability of supervised NMF and study the more difficult case of unsupervised NMF. Numerical simulations illustrate those theoretical results, and the convergence speed of NMF multiplicative updates is analyzed.
Roland Badeau, Nancy Bertin, Emmanuel Vincent 0001
ICASSP3
2011 Multi-source TDOA estimation using SNR-based angular spectra
abstract
This paper deals with the localization of multiple sources from two-channel mixtures recorded in a reverberant environment. We introduce new angular spectrum-based methods relying on the signal-to-noise ratio (SNR) to estimate the time difference of arrival (TDOA) of each source. We propose and compare five ways of estimating the SNR in each time-frequency point and in each direction, using beamforming techniques and statistical models. Large-scale evaluation considering a high number of situations shows the effectiveness of the proposed approach compared to state-of-the-art angular spectrum-based techniques.
Charles Blandin, Emmanuel Vincent 0001, Alexey Ozerov
ICASSP2
2011 Multichannel harmonic and percussive component separation by joint modeling of spatial and spectral continuity
abstract
This paper considers the blind separation of the harmonic and percussive components of multichannel music signals. We model the contribution of each source to all mixture channels in the time-frequency domain via a spatial covariance matrix, which encodes its spatial characteristics, and a scalar spectral variance, which represents its spectral structure. We then exploit the spatial continuity and the different spectral continuity structures of harmonic and percussive components as prior information to derive maximum a posteriori (MAP) estimates of the parameters using the expectation-maximization (EM) algorithm. Experimental results over professional musical mixtures show the effectiveness of the proposed approach.
Ngoc Q. K. Duong, Hideyuki Tachibana, Emmanuel Vincent 0001, Nobutaka Ono, Rémi Gribonval, Shigeki Sagayama
ICASSP3
2011 An acoustically-motivated spatial prior for under-determined reverberant source separation
abstract
We consider the task of under-determined reverberant audio source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random vector with full-rank spatial co variance matrix. We introduce an inverse Wishart prior over the covariance matrices, whose mean is given by the theory of statistical room acoustics and whose variance is learned from training data. We then derive an Expectation-Maximization (EM) algorithm to estimate the model parameters in the Maximum A Posteriori (MAP) sense given prior knowledge about the microphone spacing and the source positions. This algorithm provides a principled solution to the well-known per mutation problem and achieves better separation performance than other algorithms exploiting the same prior knowledge.
Ngoc Q. K. Duong, Emmanuel Vincent 0001, Rémi Gribonval
ICASSP2
2011 Multipitch estimation by joint modeling of harmonic and transient sounds
abstract
Multipitch estimation techniques are widely used for music transcription and acquisition of musical data from digital signals. In this paper, we propose a flexible harmonic temporal timbre model to decompose the spectral energy of the signal in the time-frequency domain into individual pitched notes. Each note is modeled with a 2-dimensional Gaussian mixture. Unlike previous approaches, the proposed model is able to represent not only the harmonic partials but also the inharmonic attack of each note. We derive an Expectation-Maximization (EM) algorithm to estimate the parameters of this model and illustrate the higher performance of the proposed algorithm than NMF algorithm and HTC algorithm for the task of multipitch estimation over synthetic and real-world data.
Emmanuel Vincent 0001, Stanislaw Andrzej Raczynski, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama
ICASSP2
2011 Subjective and Objective Quality Assessment of Audio Source Separation
abstract
We aim to assess the perceived quality of estimated source signals in the context of audio source separation. These signals may involve one or more kinds of distortions, including distortion of the target source, interference from the other sources or musical noise artifacts. We propose a subjective test protocol to assess the perceived quality with respect to each kind of distortion and collect the scores of 20 subjects over 80 sounds. We then propose a family of objective measures aiming to predict these subjective scores based on the decomposition of the estimation error into several distortion components and on the use of the PEMO-Q perceptual salience measure to provide multiple features that are then combined. These measures increase correlation with subjective scores up to 0.5 compared to nonlinear mapping of individual state-of-the-art source separation measures. Finally, we released the data and code presented in this paper in a freely available toolkit called PEASS.
Valentin Emiya, Emmanuel Vincent 0001, Niklas Harlander, Volker Hohmann
IEEE Trans. Speech Audio Process.2
2010 Under-determined convolutive blind source separation using spatial covariance models
abstract
This paper deals with the problem of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial properties of the source. We consider two covariance models and address the estimation of their parameters from the recorded mixture by a suitable initialization scheme followed by an iterative expectation-maximization (EM) procedure in each frequency bin. We then align the order of the estimated sources across all frequency bins based on their estimated directions of arrival (DOA). Experimental results over a stereo reverberant speech mixture show the effectiveness of the proposed approach.
Ngoc Q. K. Duong, Emmanuel Vincent 0001, Rémi Gribonval
ICASSP2
2010 Designing the Wiener post-filter for diffuse noise suppression using imaginary parts of inter-channel cross-spectra
abstract
This paper describes a new design of the Wiener post-filter for diffuse noise suppression. The Wiener post-filter is well-known as an effective post-processing of the minimum variance distortionless response beamformer, and its output is the optimal estimate of the target signal in the sense of the minimum mean square error. It is essential to accurately estimate the target power spectrum from the observed signals contaminated by noise when designing the Wiener post-filter. In our method, it is estimated from the imaginary parts of the inter-channel observation cross-spectra, under the assumption that the inter-channel noise cross-spectra are real-valued. The post-filter is designed using the estimate and this design is shown to be effective even for a small-sized array through experiments using simulated and real environmental noise.
Nobutaka Ito, Nobutaka Ono, Emmanuel Vincent 0001, Shigeki Sagayama
ICASSP3
2010 Enforcing Harmonicity and Smoothness in Bayesian Non-Negative Matrix Factorization Applied to Polyphonic Music Transcription
abstract
This paper presents theoretical and experimental results about constrained non-negative matrix factorization (NMF) in a Bayesian framework. A model of superimposed Gaussian components including harmonicity is proposed, while temporal continuity is enforced through an inverse-Gamma Markov chain prior. We then exhibit a space-alternating generalized expectation-maximization (SAGE) algorithm to estimate the parameters. Computational time is reduced by initializing the system with an original variant of multiplicative harmonic NMF, which is described as well. The algorithm is then applied to perform polyphonic piano music transcription. It is compared to other state-of-the-art algorithms, especially NMF-based. Convergence issues are also discussed on a theoretical and experimental point of view. Bayesian NMF with harmonicity and temporal continuity constraints is shown to outperform other standard NMF-based transcription systems, providing a meaningful mid-level representation of the data. However, temporal smoothness has its drawbacks, as far as transients are concerned in particular, and can be detrimental to transcription performance when it is the only constraint used. Possible improvements of the temporal prior are discussed.
Nancy Bertin, Roland Badeau, Emmanuel Vincent 0001
IEEE Trans. Speech Audio Process.3
2010 Under-Determined Reverberant Audio Source Separation Using a Full-Rank Spatial Covariance Model
abstract
This paper addresses the modeling of reverberant recording environments in the context of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial characteristics of the source. We then consider four specific covariance models, including a full-rank unconstrained model. We derive a family of iterative expectation-maximization (EM) algorithms to estimate the parameters of each model and propose suitable procedures adapted from the state-of-the-art to initialize the parameters and to align the order of the estimated sources across all frequency bins. Experimental results over reverberant synthetic mixtures and live recordings of speech data show the effectiveness of the proposed approach.
Ngoc Q. K. Duong, Emmanuel Vincent 0001, Rémi Gribonval
IEEE Trans. Speech Audio Process.2
2010 Beyond the Narrowband Approximation: Wideband Convex Methods for Under-Determined Reverberant Audio Source Separation
abstract
We consider the problem of extracting the source signals from an under-determined convolutive mixture assuming known mixing filters. State-of-the-art methods operate in the time-frequency domain and rely on narrowband approximation of the convolutive mixing process by complex-valued multiplication in each frequency bin. The source signals are then estimated by minimizing either a mixture fitting cost or aℓ1source sparsity cost, under possible constraints on the number of active sources. In this paper, we define a wideband ℓ2mixture fitting cost circumventing the above approximation and investigate the use of a ℓ1,2mixed-norm cost promoting disjointness of the source time-frequency representations. We design a family of convex functionals combining these costs and derive suitable optimization algorithms. Experiments indicate that the proposed wideband methods result in a signal-to-distortion ratio improvement of 2 to 5 dB compared to the state-of-the-art on reverberant speech mixtures.
Matthieu Kowalski, Emmanuel Vincent 0001, Rémi Gribonval
IEEE Trans. Speech Audio Process.2
2010 Adaptive Harmonic Spectral Decomposition for Multiple Pitch Estimation
abstract
Multiple pitch estimation consists of estimating the fundamental frequencies and saliences of pitched sounds over short time frames of an audio signal. This task forms the basis of several applications in the particular context of musical audio. One approach is to decompose the short-term magnitude spectrum of the signal into a sum of basis spectra representing individual pitches scaled by time-varying amplitudes, using algorithms such as nonnegative matrix factorization (NMF). Prior training of the basis spectra is often infeasible due to the wide range of possible musical instruments. Appropriate spectra must then be adaptively estimated from the data, which may result in limited performance due to overfitting issues. In this paper, we model each basis spectrum as a weighted sum of narrowband spectra representing a few adjacent harmonic partials, thus enforcing harmonicity and spectral smoothness while adapting the spectral envelope to each instrument. We derive a NMF-like algorithm to estimate the model parameters and evaluate it on a database of piano recordings, considering several choices for the narrowband spectra. The proposed algorithm performs similarly to supervised NMF using pre-trained piano spectra but improves pitch estimation performance by 6% to 10% compared to alternative unsupervised NMF algorithms.
Emmanuel Vincent 0001, Nancy Bertin, Roland Badeau
IEEE Trans. Speech Audio Process.1
2010 Stability Analysis of Multiplicative Update Algorithms and Application to Nonnegative Matrix Factorization
abstract
Multiplicative update algorithms have proved to be a great success in solving optimization problems with nonnegativity constraints, such as the famous nonnegative matrix factorization (NMF) and its many variants. However, despite several years of research on the topic, the understanding of their convergence properties is still to be improved. In this paper, we show that Lyapunov's stability theory provides a very enlightening viewpoint on the problem. We prove the exponential or asymptotic stability of the solutions to general optimization problems with nonnegative constraints, including the particular case of supervised NMF, and finally study the more difficult case of unsupervised NMF. The theoretical results presented in this paper are confirmed by numerical simulations involving both supervised and unsupervised NMF, and the convergence speed of NMF multiplicative updates is investigated.
Roland Badeau, Nancy Bertin, Emmanuel Vincent 0001
IEEE Trans. Neural Networks3
2009 Benchmarking flexible adaptive time-frequency transforms for underdetermined audio source separation
abstract
We have implemented several fast and flexible adaptive lapped orthogonal transform (LOT) schemes for underdetermined audio source separation. This is generally addressed by time-frequency masking, requiring the sources to be disjoint in the time-frequency domain. We have already shown that disjointness can be increased via adaptive dyadic LOTs. By taking inspiration from the windowing schemes used in many audio coding frameworks, we improve on earlier results in two ways. Firstly, we consider non-dyadic LOTs which match the time-varying signal structures better. Secondly, we allow for a greater range of overlapping window profiles to decrease window boundary artifacts. This new scheme is benchmarked through oracle evaluations, and is shown to decrease computation time by over an order of magnitude compared to using very general schemes, whilst maintaining high separation performance and flexible signal adaptivity. As the results demonstrate, this work may find practical applications in high fidelity audio source separation.
Andrew Nesbit, Emmanuel Vincent 0001, Mark D. Plumbley
ICASSP2
2009 Robust modeling of musical chord sequences using probabilistic N-grams
abstract
The modeling of music as a language is a core issue for a wide range of applications such as polyphonic music retrieval, automatic style identification, audio to symbolic music transcription and computer-assisted composition. In this paper, we focus on the modeling of chord sequences by probabilistic N-grams. Previous studies using these models have achieved limited success, due to overfitting and to the use of a single chord labeling scheme. We investigate these issues using model smoothing and selection techniques initially designed for spoken language modeling. This approach is evaluated over a set of songs by The Beatles, considering several chord labeling schemes. Initial results show that the accuracy of N-grams is increased but that additional improvements may still be achieved in the future using more advanced, possibly music-specific, smoothing techniques.
Ricardo Scholz, Emmanuel Vincent 0001, Frédéric Bimbot
ICASSP2
2008 Harmonic and inharmonic Nonnegative Matrix Factorization for Polyphonic Pitch transcription
abstract
Polyphonic pitch transcription consists of estimating the onset time, duration and pitch of each note in a music signal. This task is difficult in general, due to the wide range of possible instruments. This issue has been studied using adaptive models such as Nonnegative Matrix Factorization (NMF), which describe the signal as a weighted sum of basis spectra. However basis spectra representing multiple pitches result in inaccurate transcription. To avoid this, we propose a family of constrained NMF models, where each basis spectrum is expressed as a weighted sum of narrowband spectra consisting of a few adjacent partials at harmonic or inharmonic frequencies. The model parameters are adapted via combined multiplicative and Newton updates. The proposed method is shown to outperform standard NMF on a database of piano excerpts.
Emmanuel Vincent 0001, Nancy Bertin, Roland Badeau
ICASSP1
2008 An adaptive stereo basis method for convolutive blind audio source separation
Maria G. Jafari, Emmanuel Vincent 0001, Samer A. Abdallah, Mark D. Plumbley, Mike E. Davies 0001
Neurocomputing2
2008 Efficient Bayesian inference for harmonic models via adaptive posterior factorization
Emmanuel Vincent 0001, Mark D. Plumbley
Neurocomputing1
2008 Instrument-Specific Harmonic Atoms for Mid-Level Music Representation
abstract
Several studies have pointed out the need for accurate mid-level representations of music signals for information retrieval and signal processing purposes. In this paper, we propose a new mid-level representation based on the decomposition of a signal into a small number of sound atoms or molecules bearing explicit musical instrument labels. Each atom is a sum of windowed harmonic sinusoidal partials whose relative amplitudes are specific to one instrument, and each molecule consists of several atoms from the same instrument spanning successive time windows. We design efficient algorithms to extract the most prominent atoms or molecules and investigate several applications of this representation, including polyphonic instrument recognition and music visualization.
Pierre Leveau, Emmanuel Vincent 0001, Gaël Richard, Laurent Daudet
IEEE Trans. Speech Audio Process.2
2007 Oracle estimators for the benchmarking of source separation algorithms
Emmanuel Vincent 0001, Rémi Gribonval, Mark D. Plumbley
Signal Process.1
2007 Low Bit-Rate Object Coding of Musical Audio Using Bayesian Harmonic Models
abstract
This paper deals with the decomposition of music signals into pitched sound objects made of harmonic sinusoidal partials for very low bit-rate coding purposes. After a brief review of existing methods, we recast this problem in the Bayesian framework. We propose a family of probabilistic signal models combining learned object priors and various perceptually motivated distortion measures. We design efficient algorithms to infer object parameters and build a coder based on the interpolation of frequency and amplitude parameters. Listening tests suggest that the loudness-based distortion measure outperforms other distortion measures and that our coder results in a better sound quality than baseline transform and parametric coders at 8 and 2 kbit/s. This work constitutes a new step towards a fully object-based coding system, which would represent audio signals as collections of meaningful note-like sound objects
Emmanuel Vincent 0001, Mark D. Plumbley
IEEE Trans. Speech Audio Process.1
2006 Musical source separation using time-frequency source priors
abstract
This article deals with the source separation problem for stereo musical mixtures using prior information about the sources (instrument names and localization). After a brief review of existing methods, we design a family of probabilistic mixture generative models combining modified positive independent subspace analysis (ISA), localization models, and segmental models (SM). We express source separation as a Bayesian estimation problem and we propose efficient resolution algorithms. The resulting separation methods rely on a variable number of cues including harmonicity, spectral envelope, azimuth, note duration, and monophony. We compare these methods on two synthetic mixtures with long reverberation. We show that they outperform methods exploiting spatial diversity only and that they are robust against approximate localization of the sources.
Emmanuel Vincent 0001
IEEE Trans. Speech Audio Process.1
2006 Performance measurement in blind audio source separation
abstract
In this paper, we discuss the evaluation of blind audio source separation (BASS) algorithms. Depending on the exact application, different distortions can be allowed between an estimated source and the wanted true source. We consider four different sets of such allowed distortions, from time-invariant gains to time-varying filters. In each case, we decompose the estimated source into a true source part plus error terms corresponding to interferences, additive noise, and algorithmic artifacts. Then, we derive a global performance measure using an energy ratio, plus a separate performance measure for each error term. These measures are computed and discussed on the results of several BASS problems with various difficulty levels
Emmanuel Vincent 0001, Rémi Gribonval, Cédric Févotte
IEEE Trans. Speech Audio Process.1