VLDB 2026 Research / reviewers in the wild / expert
Natalia A. Tomashenko
dblp:133/5680
· DBLP profile ↗
41ranked-venue papers
16as first author
22since 2021 · last 2027
0000-0002-7125-2382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 11 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Privacy attacks on voice anonymization systems: Overview and key findings from the First VoicePrivacy Attacker Challenge
Natalia A. Tomashenko, Xiaoxiao Miao, Emmanuel Vincent 0001, Junichi Yamagishi |
Comput. Speech Lang. | 1 |
| 2026 | The third VoicePrivacy challenge: Preserving emotional expressiveness and linguistic content in voice anonymization
Natalia A. Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer, Michele Panariello, Xin Wang 0037, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi, Massimiliano Todisco |
Comput. Speech Lang. | 1 |
| 2025 | Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice AnonymizationabstractIn this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker’s speech rate and, more importantly, the speaker’s phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity. Natalia A. Tomashenko, Emmanuel Vincent 0001, Marc Tommasi |
ICASSP | 1 |
| 2025 | The First VoicePrivacy Attacker ChallengeabstractThe First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge which focuses on evaluating attacker systems against a set of voice anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation datasets were provided along with a baseline attacker. Participants developed their attacker systems in the form of automatic speaker verification systems and submitted their scores on the development and evaluation data. The best attacker systems reduced the equal error rate (EER) by 25–44% relative w.r.t. the baseline. Natalia A. Tomashenko, Xiaoxiao Miao, Emmanuel Vincent 0001, Junichi Yamagishi |
ICASSP | 1 |
| 2025 | Exploiting Context-dependent Duration Features for Voice Anonymization Attack SystemsabstractInternational audience Natalia A. Tomashenko, Emmanuel Vincent 0001, Marc Tommasi |
INTERSPEECH | 1 |
| 2025 | Adapting general disentanglement-based speaker anonymization for enhanced emotion preservationabstractA general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, needs to be preserved to a greater extent. Two strategies for this are examined. First, we show that integrating emotion embeddings from a pre-trained emotion encoder can help preserve emotional cues, even though this approach slightly compromises privacy protection. Alternatively, we propose an emotion compensation strategy as a post-processing step applied to anonymized speaker embeddings. This conceals the original speaker’s identity and reintroduces the emotional traits lost during speaker embedding anonymization. Specifically, we model the emotion attribute using support vector machines to learn separate boundaries for each emotion. During inference, the original speaker embedding is processed in two ways: one, by an emotion indicator to predict emotion and select the emotion-matched SVM accurately; and two, by a speaker anonymizer to conceal speaker characteristics. The anonymized speaker embedding is then modified along the corresponding SVM boundary towards an enhanced emotional direction to save the emotional cues. The proposed strategies are also expected to be useful for adapting a general disentanglement-based speaker anonymization system to preserve other target paralinguistic attributes, with potential for a range of downstream tasks 2 . Xiaoxiao Miao, Xin Wang 0037, Natalia A. Tomashenko, Cheng Lock Donny Soh, Ian McLoughlin 0001 |
Comput. Speech Lang. | 4 |
| 2024 | Anonymizing Speaker Voices: Easy to Imitate, Difficult to Recognize?abstractA vastly under-explored area in speech anonymization involves characterizing how different speakers perform in voice privacy tasks. In this paper, we present a deeper analysis by creating and analyzing groups of challenging speakers categorized based on their performance in two related facets of voice anonymization evaluation: (1) speaker similarity using automatic speaker verification (ASV) and (2) human perception using a large-scale A/B listening test. We group speakers into four categories (sheep, goats, lambs, and wolves) based on their anonymization properties. We present an extension of voice anonymization evaluation by identifying speakers who are easy to imitate or difficult to recognize. This knowledge is important for trustworthy anonymization evaluation, and it has the potential to influence how evaluation datasets are created from a pool of speakers. We provide further insights on speaker influence on anonymized speech between human perception and automatic speaker similarity scoring. Jennifer Williams 0001, Karla Pizzi, Natalia A. Tomashenko, Sneha Das |
ICASSP | 3 |
| 2024 | LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Comput. Speech Lang. | 9 |
| 2024 | Evaluating the effects of task design on unfamiliar Francophone listener and automatic speaker identification performance
Benjamin O'Brien, Christine Meunier, Natalia A. Tomashenko, Alain Ghio, Jean-François Bonastre |
Multim. Tools Appl. | 3 |
| 2024 | The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice AnonymisationabstractThe VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datasets used for system development and evaluation, present the different attack models used for evaluation, and the associated objective and subjective metrics. We describe three anonymisation baselines, provide a summary description of the anonymisation systems developed by challenge participants, and report objective and subjective evaluation results for all. In addition, we describe post-evaluation analyses and a summary of related work reported in the open literature. Results show that solutions based on voice conversion better preserve utility, that an alternative which combines automatic speech recognition with synthesis achieves greater privacy, and that a privacy-utility trade-off remains inherent to current anonymisation solutions. Finally, we present our ideas and priorities for future VoicePrivacy Challenge editions. Michele Panariello, Natalia A. Tomashenko, Xin Wang 0037, Xiaoxiao Miao, Pierre Champion, Hubert Nourtel, Massimiliano Todisco, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Federated Learning for ASR Based on wav2vec 2.0abstractThis paper presents a study on the use of federated learning to train an ASR model based on a wav2vec 2.0 model pre-trained by self supervision. Carried out on the well-known TED-LIUM 3 dataset, our experiments show that such a model can obtain, with no use of a language model, a word error rate of 10.92% on the official TEDLIUM 3 test set, without sharing any data from the different users. We also analyse the ASR performance for speakers depending to their participation to the federated learning. Since federated learning was first introduced for privacy purposes, we also measure its ability to protect speaker identity. To do that, we exploit an approach to analyze information contained in exchanged models based on a neural network footprint on an indicator dataset. This analysis is made layer-wise and shows which layers in an exchanged wav2vec 2.0based model bring the speaker identity information. Salima Mdhaffar, Natalia A. Tomashenko, Jean-François Bonastre, Yannick Estève |
ICASSP | 3 |
| 2023 | Speaker Anonymization Using Orthogonal Householder Neural NetworkabstractSpeaker anonymization aims to conceal a speaker's identity while preserving content information in speech. Current mainstream neural-network speaker anonymization systems disentangle speech into prosody-related, content, and speaker representations. The speaker representation is then anonymized by a selection-based speaker anonymizer that uses a mean vector over a set of randomly selected speaker vectors from an external pool of English speakers. However, the resulting anonymized vectors are subject to severe privacy leakage against powerful attackers, reduction in speaker diversity, and language mismatch problems for unseen-language speaker anonymization. To generate diverse, language-neutral speaker vectors, this paper proposes an anonymizer based on an orthogonal Householder neural network (OHNN). Specifically, the OHNN acts like a rotation to transform the original speaker vectors into anonymized speaker vectors, which are constrained to follow the distribution over the original speaker vector space. A basic classification loss is introduced to ensure that anonymized speaker vectors from different speakers have unique speaker identities. To further protect speaker identities, an improved classification loss and similarity loss are used to push original-anonymized sample pairs away from each other. Experiments on VoicePrivacy Challenge datasets in English and theAISHELL-3dataset in Mandarin demonstrate the proposed anonymizer's effectiveness. Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Natalia A. Tomashenko |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Retrieving Speaker Information from Personalized Acoustic Models for Speech RecognitionabstractThe widespread of powerful personal devices capable of collecting voice of their users has opened the opportunity to build speaker adapted speech recognition system (ASR) or to participate to collaborative learning of ASR. In both cases, personalized acoustic models (AM), i.e. fine-tuned AM with specific speaker data, can be built. A question that naturally arises is whether the dissemination of personalized acoustic models can leak personal information. In this paper, we show that it is possible to retrieve the gender of the speaker, but also his identity, by just exploiting the weight matrix changes of a neural acoustic model locally adapted to this speaker. Incidentally we observe phenomena that may be useful towards explainability of deep neural networks in the context of speech processing. Gender can be identified almost surely using only the first layers and speaker verification performs well when using middle-up layers. Our experimental study on the TED-LIUM 3 dataset with HMM/TDNN models shows a purity of 95% for gender detection, and an Equal Error Rate of 9.07% for a speaker verification task by only exploiting the weights from personalized models that could be exchanged instead of user data. Salima Mdhaffar, Jean-François Bonastre, Marc Tommasi, Natalia A. Tomashenko, Yannick Estève |
ICASSP | 4 |
| 2022 | Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning FrameworkabstractThis paper investigates methods to effectively retrieve speaker information from the personalized speaker adapted neural network acoustic models (AMs) in automatic speech recognition (ASR). This problem is especially important in the context of federated learning of ASR acoustic models where a global model is learnt on the server based on the updates received from multiple clients. We propose an approach to analyze information in neural network AMs based on a neural network footprint on the so-called Indicator dataset. Using this method, we develop two attack models that aim to infer speaker identity from the updated personalized models without access to the actual users’ speech data. Experiments on the TED-LIUM 3 corpus demonstrate that the proposed approaches are very effective and can provide equal error rate (EER) of 1–2%. Natalia A. Tomashenko, Salima Mdhaffar, Marc Tommasi, Yannick Estève, Jean-François Bonastre |
ICASSP | 1 |
| 2022 | A Study of Gender Impact in Self-supervised Models for Speech-to-Text SystemsabstractSelf-supervised models for speech processing emerged recently as popular foundation blocks in speech processing pipelines.These models are pre-trained on unlabeled audio data and then used in speech processing downstream tasks such as automatic speech recognition (ASR) or speech translation (ST).Since these models are now used in research and industrial systems alike, it becomes necessary to understand the impact caused by some features such as gender distribution within pre-training data.Using French as our investigation language, we train and compare gender-specific wav2vec 2.0 models against models containing different degrees of gender balance in their pretraining data.The comparison is performed by applying these models to two speech-to-text downstream tasks: ASR and ST.Results show the type of downstream integration matters.We observe lower overall performance using gender-specific pretraining before fine-tuning an end-to-end ASR system.However, when self-supervised models are used as feature extractors, the overall ASR and ST results follow more complex patterns in which the balanced pre-trained model does not necessarily lead to the best results.Lastly, our crude 'fairness' metric, the relative performance difference measured between female and male test sets, does not display a strong variation from balanced to gender-specific pre-trained wav2vec 2.0 models. Marcely Zanon Boito, Laurent Besacier, Natalia A. Tomashenko, Yannick Estève |
INTERSPEECH | 3 |
| 2022 | Analyzing Language-Independent Speaker Anonymization Framework under Unseen ConditionsabstractIn our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models.Although the system can anonymize speech data of any language, the anonymization was imperfect, and the speech content of the anonymized speech was distorted.This limitation is more severe when the input speech is from a domain unseen in the training data.This study analyzed the bottleneck of the anonymization system under unseen conditions.It was found that the domain (e.g., language and channel) mismatch between the training and test data affected the neural waveform vocoder and anonymized speaker vectors, which limited the performance of the whole system.Increasing the training data diversity for the vocoder was found to be helpful to reduce its implicit language and channel dependency.Furthermore, a simple correlation-alignment-based domain adaption strategy was found to be significantly effective to alleviate the mismatch on the anonymized speaker vectors.Audio samples 1 and source code 2 are available online. Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Natalia A. Tomashenko |
INTERSPEECH | 5 |
| 2022 | Towards a unified assessment framework of speech pseudonymisationabstractAnonymisation and pseudonymisation are two similar concepts used in privacy preservation for speech data. With no established definitions for these tasks, nor standard approaches to assessment, this paper provides definitions and presents two complementary assessment frameworks. The first is based on voice similarity matrices which provide both an immediate visualisation of privacy protection performance at the speaker level and two objective measures in the form of de-identification and voice distinctiveness preservation. The approach readily highlights imbalances in system performance at the speaker level. The second, referred to as the zero evidence biometric recognition assessment (ZEBRA) framework, is based on information theory and measures the amount of private information disclosed in speech data. The paper presents also an extension to the original ZEBRA framework. It aims to reflect the robustness of the privacy safeguard when a privacy adversary adapts to the protected speech. We demonstrate the application of both frameworks to assess pseudonymisation performance on the two VoicePrivacy 2020 challenge baseline solutions plus a third one. The two frameworks were designed independently of each other. The ZEBRA framework is fully consistent with the Bayesian decision theory and the other framework focuses instead on speaker-wise visualisations of a system performance. Thus, while metrics derived from them bear similarities, they expose differences in safeguard behavior. The assessment of pseudonymisation remains challenging and merits greater attention in the future. Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Natalia A. Tomashenko, Driss Matrouf |
Comput. Speech Lang. | 6 |
| 2022 | The VoicePrivacy 2020 Challenge: Results and findings
Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Jose Patino 0001, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche |
Comput. Speech Lang. | 1 |
| 2022 | Privacy and Utility of X-Vector Based Speaker AnonymizationabstractWe study the scenario where individuals (speakers) contribute to the publication of an anonymized speech corpus. Datausersleverage this public corpus for downstream tasks, e.g., training an automatic speech recognition (ASR) system, whileattackersmay attempt to de-anonymize it using auxiliary knowledge. Motivated by this scenario, speaker anonymization aims to conceal speaker identity while preserving the quality and usefulness of speech data. In this article, we study x-vector based speaker anonymization, the leading approach in the VoicePrivacy Challenge, which converts the speaker’s voice into that of a random pseudo-speaker. We show that the strength of anonymization varies significantly depending on how the pseudo-speaker is chosen. We explore four design choices for this step: the distance metric between speakers, the region of speaker space where the pseudo-speaker is picked, its gender, and whether to assign it to one or all utterances of the original speaker. We assess the quality of anonymization from the perspective of the three actors involved in our threat model, namely the speaker, the user and the attacker. To measure privacy and utility, we use respectively the linkability score achieved by the attackers and the decoding word error rate achieved by an ASR model trained on the anonymized data. Experiments on LibriSpeech show that the best combination of design choices yields state-of-the-art performance in terms of both privacy and utility. Experiments on Mozilla Common Voice further show that it guarantees the same anonymization level against re-identification attacks among 50 speakers as original speech among 20,000 speakers. Brij Mohan Lal Srivastava, Mohamed Maouche, Md. Sahidullah, Emmanuel Vincent 0001, Aurélien Bellet, Marc Tommasi, Natalia A. Tomashenko, Xin Wang 0037, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2021 | Speaker Anonymisation Using the McAdams CoefficientabstractAnonymisation has the goal of manipulating speech signals in order to degrade the reliability of automatic approaches to speaker recognition, while preserving other aspects of speech, such as those relating to intelligibility and naturalness. This paper reports an approach to anonymisation that, unlike other current approaches, requires no training data, is based upon well-known signal processing techniques and is both efficient and effective. The proposed solution uses the McAdams coefficient to transform the spectral envelope of speech signals. Results derived using common VoicePrivacy 2020 databases and protocols show that random, optimised transformations can outperform competing solutions in terms of anonymisation while causing only modest, additional degradations to intelligibility, even in the case of a semi-informed privacy adversary. Jose Patino 0001, Natalia A. Tomashenko, Massimiliano Todisco, Andreas Nautsch, Nicholas W. D. Evans |
Interspeech | 2 |
| 2021 | LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from SpeechabstractSelf-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve performance on downstream tasks such as automatic speech recognition (ASR). While these works suggest it is possible to reduce dependence on labeled data for building efficient speech systems, their evaluation was mostly made on ASR and using multiple and heterogeneous experimental settings (most of them for English). This questions the objective comparison of SSL approaches and the evaluation of their impact on building speech systems. In this paper, we propose LeBenchmark: a reproducible framework for assessing SSL from speech. It not only includes ASR (high and low resource) tasks but also spoken language understanding, speech translation and emotion recognition. We also focus on speech technologies in a language different than English: French. SSL models of different sizes are trained from carefully sourced and documented datasets. Experiments show that SSL is beneficial for most but not all tasks which confirms the need for exhaustive and reliable benchmarks to evaluate its real impact. LeBenchmark is shared with the scientific community for reproducible research in SSL from speech. Solène Evain, Hang Le 0001, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia A. Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Estève, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Interspeech | 8 |
| 2021 | Anonymous Speaker Clusters: Making Distinctions Between Anonymised Speech Recordings with Clustering InterfaceabstractOur study examined the performance of evaluators tasked to group natural and anonymised speech recordings into clusters based on their perceived similarities. Speech stimuli were selected from the VCTK corpus; two systems developed for the VoicePrivacy 2020 Challenge were used for anonymisation. The Baseline-1 (B1) system was developed by using x-vectors and neural waveform models, while the Baseline-2 (B2) system relied on digital-signal-processing techniques. 74 evaluators completed three trials composed of 16 recordings with either natural or anonymised speech generated from a single system. F-measure and cluster purity metrics were used to assess evaluator accuracy. Probabilistic linear discriminant analysis (PLDA) scores from an automatic speaker verification system were generated to quantify similarity between recordings and used to correlate subjective results. Our findings showed that non-native English speaking evaluators significantly lowered their F-measure means when presented anonymised recordings. We observed no significance for cluster purity. Pearson correlation procedures revealed that PLDA scores generated from natural and B2-anonymised speech recordings correlated positively to F-measure and cluster purity metrics. These findings show evaluators were able to use the interface to cluster natural and anonymised speech recordings and suggest anonymisation systems modelled like B1 are more effective at suppressing identifiable speech characteristics. Benjamin O'Brien, Natalia A. Tomashenko, Anaïs Chanclu, Jean-François Bonastre |
Interspeech | 2 |
| 2020 | Error Analysis Applied to End-to-End Spoken Language UnderstandingabstractThis paper presents a qualitative study of errors produced by an end-to-end spoken language understanding (SLU) system (speech signal to concepts) that reaches state of the art performance. Different studies are proposed to better understand the weaknesses of such systems: comparison to a classical pipeline SLU system, a study on the cause of concept deletions (the most frequent error), observation of a problem in the capability of the end-to-end SLU system to segment correctly concepts, analysis of the system behavior to process unseen concept/value pairs, analysis of the benefit of the curriculum-based transfer learning approach. Last, we proposed a way to compute embeddings of sub-sequences that seem to contain relevant information for future work. Antoine Caubrière, Sahar Ghannay, Natalia A. Tomashenko, Renato De Mori, Antoine Laurent, Emmanuel Morin, Yannick Estève |
ICASSP | 3 |
| 2020 | Dialogue History Integration into End-to-End Signal-to-Concept Spoken Language Understanding SystemsabstractThis work investigates the embeddings for representing dialog history in spoken language understanding (SLU) systems. We focus on the scenario when the semantic information is extracted directly from the speech signal by means of a single end-to-end neural network model. We proposed to integrate dialogue history into an end-to-end signal-to-concept SLU system. The dialog history is represented in the form of dialog history embedding vectors (so-called h-vectors) and is provided as an additional information to end-to-end SLU models in order to improve the system performance. Three following types of h-vectors are proposed and experimentally evaluated in this paper: (1) supervised-all embeddings predicting bag-of-concepts expected in the answer of the user from the last dialog system response; (2) supervised-freq embeddings focusing on predicting only a selected set of semantic concept (corresponding to the most frequent errors in our experiments); and (3) unsupervised embeddings. Experiments on the MEDIA corpus for the semantic slot filling task demonstrate that the proposed h-vectors improve the model performance. Natalia A. Tomashenko, Christian Raymond, Antoine Caubrière, Renato De Mori, Yannick Estève |
ICASSP | 1 |
| 2020 | The Privacy ZEBRA: Zero Evidence Biometric Recognition AssessmentabstractInternational audience Andreas Nautsch, Jose Patino 0001, Natalia A. Tomashenko, Junichi Yamagishi, Paul-Gauthier Noé, Jean-François Bonastre, Massimiliano Todisco, Nicholas W. D. Evans |
INTERSPEECH | 3 |
| 2020 | Investigating Self-Supervised Pre-Training for End-to-End Speech TranslationabstractInternational audience Fethi Bougares, Natalia A. Tomashenko, Yannick Estève, Laurent Besacier |
INTERSPEECH | 3 |
| 2020 | Speech Pseudonymisation Assessment Using Voice Similarity MatricesabstractThe proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and potentially sensitive information. Anonymisation, the focus of the recent VoicePrivacy initiative, is one strategy to protect speaker identity information. Pseudonymisation solutions aim not only to mask the speaker identity and preserve the linguistic content, quality and naturalness, as is the goal of anonymisation, but also to preserve voice distinctiveness. Existing metrics for the assessment of anonymisation are ill-suited and those for the assessment of pseudonymisation are completely lacking. Based upon voice similarity matrices, this paper proposes the first intuitive visualisation of pseudonymisation performance for speech signals and two novel metrics for objective assessment. They reflect the two, key pseudonymisation requirements of de-identification and voice distinctiveness. Paul-Gauthier Noé, Jean-François Bonastre, Driss Matrouf, Natalia A. Tomashenko, Andreas Nautsch, Nicholas W. D. Evans |
INTERSPEECH | 4 |
| 2020 | Design Choices for X-Vector Based Speaker AnonymizationabstractInternational audience Brij Mohan Lal Srivastava, Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Junichi Yamagishi, Mohamed Maouche, Aurélien Bellet, Marc Tommasi |
INTERSPEECH | 2 |
| 2020 | Introducing the VoicePrivacy InitiativeabstractThe VoicePrivacy initiative aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a series of challenges. In this paper, we formulate the voice anonymization task selected for the VoicePrivacy 2020 Challenge and describe the datasets used for system development and evaluation. We also present the attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and report objective evaluation results. Natalia A. Tomashenko, Brij Mohan Lal Srivastava, Xin Wang 0037, Emmanuel Vincent 0001, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Paul-Gauthier Noé, Massimiliano Todisco |
INTERSPEECH | 1 |
| 2019 | Curriculum-Based Transfer Learning for an Effective End-to-End Spoken Language Understanding and Domain PortabilityabstractWe present an end-to-end approach to extract semantic concepts directly from the speech audio signal. To overcome the lack of data available for this spoken language understanding approach, we investigate the use of a transfer learning strategy based on the principles of curriculum learning. This approach allows us to exploit out-of-domain data that can help to prepare a fully neural architecture. Experiments are carried out on the French MEDIA and PORTMEDIA corpora and show that this end-to-end SLU approach reaches the best results ever published on this task. We compare our approach to a classical pipeline approach that uses ASR, POS tagging, lemmatizer, chunker... and other NLP tools that aim to enrich ASR outputs that feed an SLU text to concepts system. Last, we explore the promising capacity of our end-to-end SLU approach to address the problem of domain portability. Antoine Caubrière, Natalia A. Tomashenko, Antoine Laurent, Emmanuel Morin, Nathalie Camelin, Yannick Estève |
INTERSPEECH | 2 |
| 2019 | Investigating Adaptation and Transfer Learning for End-to-End Spoken Language Understanding from SpeechabstractInternational audience Natalia A. Tomashenko, Antoine Caubrière, Yannick Estève |
INTERSPEECH | 1 |
| 2018 | An Investigation of Mixup Training Strategies for Acoustic Models in ASRabstractInternational audience Ivan Medennikov, Yuri Y. Khokhlov, Aleksei Romanenko, Natalia A. Tomashenko, Ivan Sorokin, Alexander Zatvornitsky |
INTERSPEECH | 5 |
| 2018 | Speaker Adaptive Training and Mixup Regularization for Neural Network Acoustic Models in Automatic Speech RecognitionabstractInternational audience Natalia A. Tomashenko, Yuri Y. Khokhlov, Yannick Estève |
INTERSPEECH | 1 |
| 2018 | Evaluation of Feature-Space Speaker Adaptation for End-to-End Acoustic Models
Natalia A. Tomashenko, Yannick Estève |
LREC | 1 |
| 2017 | The STC Keyword Search System for OpenKWS 2016 Evaluation
Yuri Y. Khokhlov, Ivan Medennikov, Aleksei Romanenko, Valentin Mendelev, Maxim Korenevsky, Alexey Prudnikov, Natalia A. Tomashenko, Alexander Zatvornitsky |
INTERSPEECH | 7 |
| 2017 | Fast and Accurate OOV Decoder on High-Level FeaturesabstractInternational audience Yuri Y. Khokhlov, Natalia A. Tomashenko, Ivan Medennikov, Aleksei Romanenko |
INTERSPEECH | 2 |
| 2016 | On the Use of Gaussian Mixture Model Framework to Improve Speaker Adaptation of Deep Neural Network Acoustic ModelsabstractInternational audience Natalia A. Tomashenko, Yuri Y. Khokhlov, Yannick Estève |
INTERSPEECH | 1 |
| 2016 | LIUM ASR systems for the 2016 Multi-Genre Broadcast Arabic challengeabstractThis paper describes the automatic speech recognition (ASR) systems developed by LIUM in the framework of the 2016 Multi-Genre Broadcast (MGB-2) Challenge in the Arabic language. LIUM participated in the first of the two proposed tasks, namely the speech-to-text transcription of Aljazeera recordings. We present the approaches and details found in our systems, as well as our results in the evaluation campaign: the primary LIUM ASR system attained the second position. The main aspects come from the use of GMM-derived features for training a DNN, combined with the use of time-delay neural networks for acoustic models, the use of two different approaches in order to automatically phonetize Arabic words, and finally, the training data selection strategy for acoustic and language models. Natalia A. Tomashenko, Kévin Vythelingum, Anthony Rousseau, Yannick Estève |
SLT | 1 |
| 2015 | GMM-derived features for effective unsupervised adaptation of deep neural network acoustic models
Natalia A. Tomashenko, Yuri Y. Khokhlov |
INTERSPEECH | 1 |
| 2014 | Automated closed captioning for Russian live broadcasting
Kirill Levin, Irina Ponomareva, Anna Bulusheva, German A. Chernykh, Ivan Medennikov, Nickolay Merkin, Alexey Prudnikov, Natalia A. Tomashenko |
INTERSPEECH | 8 |
| 2014 | Speaker adaptation of context dependent deep neural networks based on MAP-adaptation and GMM-derived feature processing
Natalia A. Tomashenko, Yuri Y. Khokhlov |
INTERSPEECH | 1 |