Jean-François Bonastre

dblp:72/3130 · DBLP profile ↗
← Back
134ranked-venue papers
15as first author
22since 2021 · last 2025
0000-0001-7741-3346ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 115 · 14 first-author · 18 since 2021Artificial intelligence and machine learning · 99 · 11 first-author · 15 since 2021
YearPublicationVenuePosition
2025 Unified Text and Speaker Verification using SSL model for Text-Dependent Speaker Verification
Nathan Griot, Driss Matrouf, Raphaël Blouet, Jean-François Bonastre, Ana Mantecon
INTERSPEECH4
2025 Using gender, phonation and age to interpret automatically discovered speech attributes for explainable speaker recognition
Carole Millot, Clara Ponchard, Cédric Gendrot, Jean-François Bonastre, Orane Dufour
INTERSPEECH4
2024 RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile Robot
abstract
In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. RoboVox is a French corpus recorded by a mobile robot. The files are recorded from different distances under severe acoustical conditions with the presence of several types of noise and reverberation. In addition to noise and reverberation, the robot’s internal noise acts as an extra additive noise. RoboVox can be used for both single-channel and multi-channel speaker recognition. In the evaluation protocols, we are considering both cases. The obtained results demonstrate a significant decline in performance in far-filed speaker recognition and urge the community to further research in this domain
Mohammad MohammadAmini, Driss Matrouf, Mickael Rouvier, Jean-François Bonastre, Romain Serizel, Théophile Gonos
LREC/COLING4
2024 Explaining a Probabilistic Prediction on the Simplex with Shapley Compositions
abstract
Originating in game theory, Shapley values are widely used for explaining a machine learning model’s prediction by quantifying the contribution of each feature’s value to the prediction. This requires a scalar prediction as in binary classification, whereas a multiclass probabilistic prediction is a discrete probability distribution, living on a multidimensional simplex. In such a multiclass setting the Shapley values are typically computed separately on each class in a one-vs-rest manner, ignoring the compositional nature of the output distribution. In this paper, we introduce Shapley compositions as a well-founded way to properly explain a multiclass probabilistic prediction, using the Aitchison geometry from compositional data analysis. We prove that the Shapley composition is the unique quantity satisfying linearity, symmetry and efficiency on the Aitchison simplex, extending the corresponding axiomatic properties of the standard Shapley value. We demonstrate this proper multiclass treatment in a range of scenarios.
Paul-Gauthier Noé, Miquel Perelló-Nieto, Jean-François Bonastre, Peter A. Flach
ECAI3
2024 Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset
abstract
The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech collected from real human speakers. For example, the widely-used VoxCeleb2 dataset for speaker recognition is no longer accessible from the official website. To mitigate these concerns, this work presents an initiative to generate a privacyfriendly synthetic VoxCeleb2 dataset that ensures the quality of the generated speech in terms of privacy, utility, and fairness. We also discuss the challenges of using synthetic data for the downstream task of speaker verification.
Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Nicholas W. D. Evans, Massimiliano Todisco, Jean-François Bonastre, Mickael Rouvier
ICASSP7
2024 Extraction of interpretable and shared speaker-specific speech attributes through binary auto-encoder
abstract
International audience
Imen Ben Amor, Jean-François Bonastre, Salima Mdhaffar
INTERSPEECH2
2024 Evaluating the effects of task design on unfamiliar Francophone listener and automatic speaker identification performance
Benjamin O'Brien, Christine Meunier, Natalia A. Tomashenko, Alain Ghio, Jean-François Bonastre
Multim. Tools Appl.5
2024 Improving Tone Recognition Performance using Wav2vec 2.0-Based Learned Representation in Yoruba, a Low-Resourced Language
abstract
Many sub-Saharan African languages are categorized as tone languages, and for the most part, they are classified as low-resource languages due to the limited resources and tools available to process these languages. Identifying the tone associated with a syllable is therefore a key challenge for speech recognition in these languages. We propose models that automate the recognition of tones in continuous speech that can easily be incorporated into a speech recognition pipeline for these languages. We have investigated different neural architectures as well as several feature extraction algorithms in speech (FBs (Filter Banks), LEAF (Learnable Frontend), CS (Cestrogram), MFCC (Mel-Frequency Cepstral Coefficients)). In the context of low-resource languages, we also evaluated W2V (Wav2vec 2.0) models for this task. In this work, we use a public speech recognition dataset on Yoruba. As for the results, using the combination of features obtained from CS and FBs, we obtain a minimum TER (Tone Error Rate) of 19.54%, whereas for the evaluations of the models using W2V, we have a TER of 17.72%, demonstrating that the use of W2V provides better performance than the models used in the literature for tone identification on low-resource languages.
Saint Germes Bienvenu Bengono Obiang, Norbert Tsopzé, Paulin Melatagia Yonta, Jean-François Bonastre, Tania Jiménez
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2023 Federated Learning for ASR Based on wav2vec 2.0
abstract
This paper presents a study on the use of federated learning to train an ASR model based on a wav2vec 2.0 model pre-trained by self supervision. Carried out on the well-known TED-LIUM 3 dataset, our experiments show that such a model can obtain, with no use of a language model, a word error rate of 10.92% on the official TEDLIUM 3 test set, without sharing any data from the different users. We also analyse the ASR performance for speakers depending to their participation to the federated learning. Since federated learning was first introduced for privacy purposes, we also measure its ability to protect speaker identity. To do that, we exploit an approach to analyze information contained in exchanged models based on a neural network footprint on an indicator dataset. This analysis is made layer-wise and shows which layers in an exchanged wav2vec 2.0based model bring the speaker identity information.
Salima Mdhaffar, Natalia A. Tomashenko, Jean-François Bonastre, Yannick Estève
ICASSP4
2023 Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline
abstract
The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representation fed into a HiFiGAN vocoder is protected using a neural-discriminant analysis approach, which is consistent with the zero-evidence concept of privacy. This approach significantly reduces the information in speech related to the speaker’s sex while preserving speech content and some consistency in the resulting protected voices.
Paul-Gauthier Noé, Xiaoxiao Miao, Xin Wang 0037, Junichi Yamagishi, Jean-François Bonastre, Driss Matrouf
ICASSP5
2023 Describing the phonetics in the underlying speech attributes for deep and interpretable speaker recognition
abstract
International audience
Imen Ben Amor, Jean-François Bonastre, Benjamin O'Brien, Pierre-Michel Bousquet
INTERSPEECH2
2023 Differentiating acoustic and physiological features in speech for hypoxia detection
abstract
International audience
Benjamin O'Brien, Adrien Gresse, Jean-Baptise Billaud, Guilhem Belda, Jean-François Bonastre
INTERSPEECH5
2022 Retrieving Speaker Information from Personalized Acoustic Models for Speech Recognition
abstract
The widespread of powerful personal devices capable of collecting voice of their users has opened the opportunity to build speaker adapted speech recognition system (ASR) or to participate to collaborative learning of ASR. In both cases, personalized acoustic models (AM), i.e. fine-tuned AM with specific speaker data, can be built. A question that naturally arises is whether the dissemination of personalized acoustic models can leak personal information. In this paper, we show that it is possible to retrieve the gender of the speaker, but also his identity, by just exploiting the weight matrix changes of a neural acoustic model locally adapted to this speaker. Incidentally we observe phenomena that may be useful towards explainability of deep neural networks in the context of speech processing. Gender can be identified almost surely using only the first layers and speaker verification performs well when using middle-up layers. Our experimental study on the TED-LIUM 3 dataset with HMM/TDNN models shows a purity of 95% for gender detection, and an Equal Error Rate of 9.07% for a speaker verification task by only exploiting the weights from personalized models that could be exchanged instead of user data.
Salima Mdhaffar, Jean-François Bonastre, Marc Tommasi, Natalia A. Tomashenko, Yannick Estève
ICASSP2
2022 A Bridge between Features and Evidence for Binary Attribute-Driven Perfect Privacy
abstract
Attribute-driven privacy aims to conceal a single user’s attribute, contrary to anonymisation that tries to hide the full identity of the user in some data. When the attribute to protect from malicious inferences is binary, perfect privacy requires the log-likelihood-ratio to be zero resulting in no strength-of-evidence. This work presents an approach based on normalizing flow that maps a feature vector into a latent space where the evidence, related to the binary attribute, and an independent residual are disentangled. It can be seen as a non-linear discriminant analysis where the mapping is invertible al-lowing generation by mapping the latent variable back to the original space. This framework allows to manipulate the log-likelihood-ratio of the data and therefore allows to set it to zero for privacy. We show the applicability of the approach on an attribute-driven privacy task where the sex information is removed from speaker embeddings. Results on VoxCeleb2 dataset show the efficiency of the method that outperforms in terms of privacy and utility our previous experiments based on adversarial disentanglement.
Paul-Gauthier Noé, Andreas Nautsch, Driss Matrouf, Pierre-Michel Bousquet, Jean-François Bonastre
ICASSP5
2022 Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning Framework
abstract
This paper investigates methods to effectively retrieve speaker information from the personalized speaker adapted neural network acoustic models (AMs) in automatic speech recognition (ASR). This problem is especially important in the context of federated learning of ASR acoustic models where a global model is learnt on the server based on the updates received from multiple clients. We propose an approach to analyze information in neural network AMs based on a neural network footprint on the so-called Indicator dataset. Using this method, we develop two attack models that aim to infer speaker identity from the updated personalized models without access to the actual users’ speech data. Experiments on the TED-LIUM 3 corpus demonstrate that the proposed approaches are very effective and can provide equal error rate (EER) of 1–2%.
Natalia A. Tomashenko, Salima Mdhaffar, Marc Tommasi, Yannick Estève, Jean-François Bonastre
ICASSP5
2022 Reliability criterion based on learning-phase entropy for speaker recognition with neural network
abstract
International audience
Pierre-Michel Bousquet, Mickael Rouvier, Jean-François Bonastre
INTERSPEECH3
2022 Barlow Twins self-supervised learning for robust speaker recognition
abstract
International audience
Mohammad MohammadAmini, Driss Matrouf, Jean-François Bonastre, Sandipana Dowerah, Romain Serizel, Denis Jouvet
INTERSPEECH3
2022 Towards a unified assessment framework of speech pseudonymisation
abstract
Anonymisation and pseudonymisation are two similar concepts used in privacy preservation for speech data. With no established definitions for these tasks, nor standard approaches to assessment, this paper provides definitions and presents two complementary assessment frameworks. The first is based on voice similarity matrices which provide both an immediate visualisation of privacy protection performance at the speaker level and two objective measures in the form of de-identification and voice distinctiveness preservation. The approach readily highlights imbalances in system performance at the speaker level. The second, referred to as the zero evidence biometric recognition assessment (ZEBRA) framework, is based on information theory and measures the amount of private information disclosed in speech data. The paper presents also an extension to the original ZEBRA framework. It aims to reflect the robustness of the privacy safeguard when a privacy adversary adapts to the protected speech. We demonstrate the application of both frameworks to assess pseudonymisation performance on the two VoicePrivacy 2020 challenge baseline solutions plus a third one. The two frameworks were designed independently of each other. The ZEBRA framework is fully consistent with the Bayesian decision theory and the other framework focuses instead on speaker-wise visualisations of a system performance. Thus, while metrics derived from them bear similarities, they expose differences in safeguard behavior. The assessment of pseudonymisation remains challenging and merits greater attention in the future.
Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Natalia A. Tomashenko, Driss Matrouf
Comput. Speech Lang.5
2022 The VoicePrivacy 2020 Challenge: Results and findings
Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Jose Patino 0001, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
Comput. Speech Lang.12
2021 Automatic Classification of Phonation Types in Spontaneous Speech: Towards a New Workflow for the Characterization of Speakers' Voice Quality
abstract
International audience
Anaïs Chanclu, Imen Ben Amor, Cédric Gendrot, Emmanuel Ferragne, Jean-François Bonastre
Interspeech5
2021 Adversarial Disentanglement of Speaker Representation for Attribute-Driven Privacy Preservation
abstract
In speech technologies, speaker's voice representation is used in many applications such as speech recognition, voice conversion, speech synthesis and, obviously, user authentication. Modern vocal representations of the speaker are based on neural embeddings. In addition to the targeted information, these representations usually contain sensitive information about the speaker, like the age, sex, physical state, education level or ethnicity. In order to allow the user to choose which information to protect, we introduce in this paper the concept of attribute-driven privacy preservation in speaker voice representation. It allows a person to hide one or more personal aspects to a potential malicious interceptor and to the application provider. As a first solution to this concept, we propose to use an adversarial autoencoding method that disentangles in the voice representation a given speaker attribute thus allowing its concealment. We focus here on the sex attribute for an Automatic Speaker Verification (ASV) task. Experiments carried out using the VoxCeleb datasets have shown that the proposed method enables the concealment of this attribute while preserving ASV ability.
Paul-Gauthier Noé, Mohammad MohammadAmini, Driss Matrouf, Titouan Parcollet, Andreas Nautsch, Jean-François Bonastre
Interspeech6
2021 Anonymous Speaker Clusters: Making Distinctions Between Anonymised Speech Recordings with Clustering Interface
abstract
Our study examined the performance of evaluators tasked to group natural and anonymised speech recordings into clusters based on their perceived similarities. Speech stimuli were selected from the VCTK corpus; two systems developed for the VoicePrivacy 2020 Challenge were used for anonymisation. The Baseline-1 (B1) system was developed by using x-vectors and neural waveform models, while the Baseline-2 (B2) system relied on digital-signal-processing techniques. 74 evaluators completed three trials composed of 16 recordings with either natural or anonymised speech generated from a single system. F-measure and cluster purity metrics were used to assess evaluator accuracy. Probabilistic linear discriminant analysis (PLDA) scores from an automatic speaker verification system were generated to quantify similarity between recordings and used to correlate subjective results. Our findings showed that non-native English speaking evaluators significantly lowered their F-measure means when presented anonymised recordings. We observed no significance for cluster purity. Pearson correlation procedures revealed that PLDA scores generated from natural and B2-anonymised speech recordings correlated positively to F-measure and cluster purity metrics. These findings show evaluators were able to use the interface to cluster natural and anonymised speech recordings and suggest anonymisation systems modelled like B1 are more effective at suppressing identifiable speech characteristics.
Benjamin O'Brien, Natalia A. Tomashenko, Anaïs Chanclu, Jean-François Bonastre
Interspeech4
2020 Learning Voice Representation Using Knowledge Distillation for Automatic Voice Casting
abstract
The search for professional voice-actors for audiovisual productions is a sensitive task, performed by the artistic directors (ADs). The ADs have a strong appetite for new talents/voices but cannot perform large scale auditions. Automatic tools able to suggest the most suited voices are of a great interest for audiovisual industry. In previous works, we showed the existence of acoustic information allowing to mimic the AD's choices. However, the only available information is the ADs' choices from the already dubbed multimedia productions. In this paper, we propose a representation-learning based strategy to build a character/role representation, called p-vector. In addition, the large variability between audiovisual productions makes difficult to have homogeneous training datasets. We overcome this difficulty by using knowledge distillation methods to take advantage of external datasets. Experiments are conducted on video-game voice excerpts. Results show a significant improvement using the p-vector, compared to the speaker-based x-vectors representation.
Adrien Gresse, Mathias Quillot, Richard Dufour, Jean-François Bonastre
INTERSPEECH4
2020 Multi-Task Learning for Voice Related Recognition Tasks
Ana Montalvo, José Ramón Calvo de Lara, Jean-François Bonastre
INTERSPEECH3
2020 The Privacy ZEBRA: Zero Evidence Biometric Recognition Assessment
abstract
International audience
Andreas Nautsch, Jose Patino 0001, Natalia A. Tomashenko, Junichi Yamagishi, Paul-Gauthier Noé, Jean-François Bonastre, Massimiliano Todisco, Nicholas W. D. Evans
INTERSPEECH6
2020 Speech Pseudonymisation Assessment Using Voice Similarity Matrices
abstract
The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and potentially sensitive information. Anonymisation, the focus of the recent VoicePrivacy initiative, is one strategy to protect speaker identity information. Pseudonymisation solutions aim not only to mask the speaker identity and preserve the linguistic content, quality and naturalness, as is the goal of anonymisation, but also to preserve voice distinctiveness. Existing metrics for the assessment of anonymisation are ill-suited and those for the assessment of pseudonymisation are completely lacking. Based upon voice similarity matrices, this paper proposes the first intuitive visualisation of pseudonymisation performance for speech signals and two novel metrics for objective assessment. They reflect the two, key pseudonymisation requirements of de-identification and voice distinctiveness.
Paul-Gauthier Noé, Jean-François Bonastre, Driss Matrouf, Natalia A. Tomashenko, Andreas Nautsch, Nicholas W. D. Evans
INTERSPEECH2
2020 Introducing the VoicePrivacy Initiative
abstract
The VoicePrivacy initiative aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a series of challenges. In this paper, we formulate the voice anonymization task selected for the VoicePrivacy 2020 Challenge and describe the datasets used for system development and evaluation. We also present the attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and report objective evaluation results.
Natalia A. Tomashenko, Brij Mohan Lal Srivastava, Xin Wang 0037, Emmanuel Vincent 0001, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Paul-Gauthier Noé, Massimiliano Todisco
INTERSPEECH9
2020 Introduction to the special issue "Speaker and language characterization and recognition: Voice modeling, conversion, synthesis and ethical aspects"
Jean-François Bonastre, Tomi Kinnunen, Anthony Larcher, Junichi Yamagishi
Comput. Speech Lang.1
2019 Representation Learning for Underdefined Tasks
Jean-François Bonastre
CIARP1
2019 Similarity Metric Based on Siamese Neural Networks for Voice Casting
abstract
Dubbing contributes to a larger international distribution of multimedia documents. It aims to replace the original voice in a source language by a new one in a target language. For now, the target voice selection procedure, called voice casting, is manually performed by human experts. This selection is not exclusively based on acoustic similarity between the two voices. Actually, it is also supported by more subjective criteria such as the "color" of the voice, sociocultural choices... The objective of this work is to model a voice similarity metric able to embed all the concerned voice characteristics, including the observers' receptive interests. In this paper, we propose a Siamese Neural Networks-based approach, measuring proximity between the original and dubbed voices. We propose an adapted jackknifing cross-validation method to evaluate our similarity model on unseen voices. The results show that we successfully capture information allowing two voices to be associated, with respect to the character's or role's abstract dimension.
Adrien Gresse, Mathias Quillot, Richard Dufour, Vincent Labatut, Jean-François Bonastre
ICASSP5
2019 Effects of Waveform PMF on Anti-Spoofing Detection
abstract
International audience
Itshak Lapidot, Jean-François Bonastre
INTERSPEECH2
2018 Voice Comparison and Rhythm: Behavioral Differences between Target and Non-target Comparisons
abstract
International audience
Moez Ajili, Jean-François Bonastre, Solange Rossato
INTERSPEECH2
2018 Speech Database and Protocol Validation Using Waveform Entropy
Itshak Lapidot, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Jean-François Bonastre
INTERSPEECH5
2018 A Unified Joint Model to Deal With Nuisance Variabilities in the i-Vector Space
abstract
The past decade has witnessed a significant improvement in speaker recognition (SR) technology in terms of performance with the introduction of the i-vectors framework. Despite these advances, the performance of SR systems considerably suffers in the presence of acoustic nuisances and variabilities. In this paper, we develop a data-driven nuisance compensation technique in the i-vector space without referring to the effects of the targeted nuisances in the temporal domain. This approach is nonparametric as it does not suppose a specific relationship between a “good” version of an i-vector and its corrupted version. Instead, our algorithm models directly the joint distribution of both representations (the good i-vector and its corrupted version) and takes advantage of the reproducibility of acoustic corruptions to generate the corrupted i-vectors. We then build an MMSE estimator that computes an improved version of a corrupted test i-vector, given this joint distribution. Experiments are carried out on NIST SRE 2010 and speakers in the wild databases where the proposed algorithm is used to deal with additive noise and short utterances. Our technique is shown to be efficient, improving the baseline system performance in terms of equal-error rate by up to 70% when used on known test noises and up to 65% in the context of unseen noises using a generic model. It was also proven efficient in the context of duration mismatch reaching up to 40% of relative improvement when used on short utterances using multiple models corresponding to different durations and up to 36% when used on arbitrary duration test segments.
Waad Ben Kheder, Driss Matrouf, Moez Ajili, Jean-François Bonastre
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Phonological content impact on wrongful convictions in Forensic Voice Comparison context
abstract
Forensic Voice Comparison (FVC) is increasingly using the likelihood ratio (LR) in order to indicate whether the evidence supports the prosecution (same-speaker) or defender (different-speakers) hypotheses. Nevertheless, the LR accepts some practical limitations due both to its estimation process itself and to a lack of knowledge about the reliability of this (practical) estimation process. It is particularly true when FVC is considered using Automatic Speaker Recognition (ASR) systems. Indeed, in the LR estimation performed by ASR systems, different factors are not considered such as speaker intrinsic characteristics, denoted “speaker factor”, the amount of information involved in the comparison as well as the phonological content and so on. This article focuses on the impact of phonological content on FVC involving two different speakers and more precisely the potential implication of a specific phonemic category on wrongful conviction cases (innocents are send behind bars). We show that even though the vast majority of speaker pairs (more than 90%) are well discriminated, few pairs are difficult to distinguish. For the “best” discriminated pairs, all the phonemic content play a positive role in speaker discrimination while for the “worst” pairs, it appears that nasals have a negative effect and lead to a confusion between speakers.
Moez Ajili, Jean-François Bonastre, Waad Ben Kheder, Solange Rossato, Juliette Kahn
ICASSP2
2017 Homogeneity Measure Impact on Target and Non-Target Trials in Forensic Voice Comparison
Moez Ajili, Jean-François Bonastre, Waad Ben Kheder, Solange Rossato, Juliette Kahn
INTERSPEECH2
2017 Acoustic Pairing of Original and Dubbed Voices in the Context of Video Game Localization
abstract
International audience
Adrien Gresse, Mickael Rouvier, Richard Dufour, Vincent Labatut, Jean-François Bonastre
INTERSPEECH5
2017 The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah
INTERSPEECH61
2017 Fast i-vector denoising using MAP estimation and a noise distributions database for robust speaker recognition
Waad Ben Kheder, Driss Matrouf, Pierre-Michel Bousquet, Jean-François Bonastre, Moez Ajili
Comput. Speech Lang.4
2017 Generalized Viterbi-based models for time-series segmentation and clustering applied to speaker diarization
Itshak Lapidot, Alon Shoa, Tal Furmanov, Lidiya Aminov, Ami Moyal, Jean-François Bonastre
Comput. Speech Lang.6
2016 Inter-speaker variability in forensic voice comparison: A preliminary evaluation
abstract
In forensic voice comparison, it is strongly recommended to follow Bayesian paradigm. In this paradigm, the strength of the forensic evidence is summarized by a likelihood ratio (LR). The LR magnitude quantifies the strength of the evidence: far from unity for a meaningful LR (a LR which supports strongly one of the hypothesis); close to unity when the evidence is next to useless. Despite this nice theoretical aspect, the LR does not embed the reliability of its estimation process itself. And, in various cases, a lack in reliability inside the estimation process is able to destroy the reliability of the resulting LR. It is particularly true when voice comparison is considered, as Speaker Recognition (SR) systems are outputting a score in all situations regardless of the case specific conditions. Furthermore, SR systems use different normalization steps to see their scores as LR and these normalization steps are clearly a potential source of bias. Consequently, a complete view of reliability should be taken into account for forensic voice comparison. This article focuses on one part of this question, the "speaker factor", the characteristics and the behaviors of the two speakers involved in a voice comparison trial.
Moez Ajili, Jean-François Bonastre, Solange Rossato, Juliette Kahn
ICASSP2
2016 Speaker Comparison for Forensic and Investigative Applications II
Jean-François Bonastre, Joseph P. Campbell, Anders P. Eriksson, Hirotaka Nakasone, Reva Schwartz
INTERSPEECH1
2016 LIA System for the SITW Speaker Recognition Challenge
Waad Ben Kheder, Moez Ajili, Pierre-Michel Bousquet, Driss Matrouf, Jean-François Bonastre
INTERSPEECH5
2016 Probabilistic Approach Using Joint Long and Short Session i-Vectors Modeling to Deal with Short Utterances for Speaker Recognition
Waad Ben Kheder, Driss Matrouf, Moez Ajili, Jean-François Bonastre
INTERSPEECH4
2016 Probabilistic Approach Using Joint Clean and Noisy i-Vectors Modeling for Speaker Recognition
Waad Ben Kheder, Driss Matrouf, Moez Ajili, Jean-François Bonastre
INTERSPEECH4
2016 On the Importance of Efficient Transition Modeling for Speaker Diarization
Itshak Lapidot, Jean-François Bonastre
INTERSPEECH2
2016 FABIOLE, a Speech Database for Forensic Speaker Comparison
Moez Ajili, Jean-François Bonastre, Juliette Kahn, Solange Rossato, Guillaume Bernard 0002
LREC2
2016 Phonetic content impact on Forensic Voice Comparison
abstract
Forensic Voice Comparison (FVC) is increasingly using the likelihood ratio (LR) in order to indicate whether the evidence supports the prosecution (same-speaker) or defender (different-speakers) hypotheses. In addition to support one hypothesis, the LR provides a theoretically founded estimate of the relative strength of its support. Despite this nice theoretical aspect, the LR accepts some practical limitations due both to its estimation process itself and to a lack of knowledge about the reliability of this (practical) estimation process. In a large set of situations, a lack in reliability at the estimation process level potentially destroys the reliability of the resulting LR. It is particularly true when automatic FVC is considered, as Automatic Speaker Recognition (ASpR) systems are outputting a score in all situations regardless of the case specific conditions. Furthermore, ASpR systems use different normalization steps to see their scores as LR and these normalization steps are potential sources of bias. In the LR estimation done by ASpR systems, different factors are not taken into account such as the amount of information involved in the comparison, the phonemic content and finally the speaker intrinsic characteristics, denoted here “speaker factor”. Consequently, a more complete view of reliability seems to be a mandatory point for FVC, even if a LR-like approach is used. This article focuses on the impact of phonemic content on FVC performance and variability. The experimental part is using FABIOLE database. This database is dedicated to this kind of studies and allows to examine both inter-speaker variability and intra-speaker variability. The results demonstrate the importance of the phonemic content and highlight interesting differences between inter-speakers effects and intra-speaker's ones.
Moez Ajili, Jean-François Bonastre, Waad Ben Kheder, Solange Rossato, Juliette Kahn
SLT2
2016 Speaker recognition using temporal information and session variability compensation in a binary framework
abstract
In recent years a simple representation of a speech excerpt has been proposed, as a binary matrix allowing easy access to the speaker discriminant information. In addition to the time-related abilities of this representation, it also allows the system to work with a temporal information representat ion based on sequential changes present in the binary representation. A new temporal information is proposed in order to add it to speaker recognition systems. A new specificity selection approach using a mask in the cumulative vector space is also proposed. Furthermore in this space, temporal information can be exploited to compensate for the effects of session variability. A new variability compensation method in the temporal space is proposed in order to remove the unwanted attributes of session variability and the common attributes among speakers. This aims to increase effectiveness in the speaker binary key paradigm. The experimental validation, done on the NIST-SRE framework, demonstrates the efficiency of the proposed solutions, which shows an EER improvement of 9%. The combination of i-vector and binary approaches, using the proposed methods, showed the complementarity of the discriminatory information exploited by each of them.
Gabriel Hernández Sierra, José Ramón Calvo de Lara, Jean-François Bonastre
Intell. Data Anal.3
2015 Homogeneity Measure for Forensic Voice Comparison: A Step Forward Reliability
Moez Ajili, Jean-François Bonastre, Solange Rossato, Juliette Kahn, Itshak Lapidot
CIARP2
2015 Additive noise compensation in the i-vector space for speaker recognition
abstract
State-of-the-art speaker recognition systems performance degrades considerably in noisy environments even though they achieve very good results in clean conditions. In order to deal with this strong limitation, we aim in this work to remove the noisy part of an i-vector directly in the i-vector space. Our approach offers the advantage to operate only at the i-vector extraction level, letting the other steps of the system unchanged. A maximum a posteriori (MAP) procedure is applied in order to obtain clean version of the noisy i-vectors taking advantage of prior knowledge about clean i-vectors distribution. To perform this MAP estimation, Gaussian assumptions over clean and noise i-vectors distributions are made. Operating on NIST 2008 data, we show a relative improvement up to 60% compared with baseline system. Our approach also outperforms the “multi-style” backend training technique. The efficiency of the proposed method is obtained at the price of relative high computational cost. We present at the end some ideas to improve this aspect.
Waad Ben Kheder, Driss Matrouf, Jean-François Bonastre, Moez Ajili, Pierre-Michel Bousquet
ICASSP3
2015 An information theory based data-homogeneity measure for voice comparison
Moez Ajili, Jean-François Bonastre, Solange Rossato, Juliette Kahn, Itshak Lapidot
INTERSPEECH2
2014 Temporal Information in a Binary Framework for Speaker Recognition
Gabriel Hernández Sierra, José Ramón Calvo de Lara, Jean-François Bonastre
CIARP3
2014 Session compensation using binary speech representation for speaker recognition
Gabriel Hernández Sierra, José Ramón Calvo de Lara, Jean-François Bonastre, Pierre-Michel Bousquet
Pattern Recognit. Lett.3
2013 Identify the Benefits of the Different Steps in an i-Vector Based Speaker Verification System
Pierre-Michel Bousquet, Jean-François Bonastre, Driss Matrouf
CIARP (2)2
2013 ALIZE 3.0 - open source toolkit for state-of-the-art speaker recognition
abstract
International audience
Anthony Larcher, Jean-François Bonastre, Benoit G. B. Fauve, Kong-Aik Lee, Christophe Lévy, Haizhou Li 0001, John S. D. Mason, Jean-Yves Parfait
INTERSPEECH2
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH31
2012 Speaker Recognition Using a Binary Representation and Specificities Models
Gabriel Hernández Sierra, Jean-François Bonastre, José Ramón Calvo de Lara
CIARP2
2012 Typicality extraction in a Speaker Binary Keys model
abstract
In the field of speaker recognition, the recently proposed notion of “Speaker Binary Key” provides a representation of each acoustic frame in a discriminant binary space. This approach relies on an unique acoustic model composed by a large set of speaker specific local likelihood peaks (called specificities). The model proposes a spatial coverage where each frame is characterized in terms of neighborhood. The most frequent specificities, picked up to represent the whole utterance, generate a binary key vector. The flexibility of this modeling allows to capture non-parametric behaviors. In this paper, we introduce a concept of “typicality” between binary keys, with a discriminant goal. We describe an algorithm able to extract such typicalities, which involves a singular value decomposition in a binary space. The theoretical aspects of this decomposition as well as its potential in terms of future developments are presented. All the propositions are also experimentally validated using NIST SRE 2008 framework.
Pierre-Michel Bousquet, Jean-François Bonastre
ICASSP2
2012 I-vectors in the context of phonetically-constrained short utterances for speaker verification
abstract
Short speech duration remains a critical factor of performance degradation when deploying a speaker verification system. To overcome this difficulty, a large number of commercial applications impose the use of fixed pass-phrases. In this context, we show that the performance of the popular i-vector approach can be greatly improved by taking advantage of the phonetic information that they convey. Moreover, as i-vectors require a conditioning process to reach high accuracy, we show that further improvements are possible by taking advantage of this phonetic information within the normalisation process. We compare two methods, Within Class Covariance Normalization (WCCN) and Eigen Factor Radial (EFR), both relying on parameters estimated on the same development data. Our study suggests that WCCN is more robust to data mismatch but less efficient than EFR when the development data has a better match with the test data.
Anthony Larcher, Pierre-Michel Bousquet, Kong-Aik Lee, Driss Matrouf, Haizhou Li 0001, Jean-François Bonastre
ICASSP6
2012 Computationally efficient speaker identification using fast-MLLR based anchor modeling
abstract
In this paper, we propose a computationally efficient method to identify a speaker from a large population of speakers. The proposed method is based on our earlier [1] Fast Maximum Likelihood Linear Linear Regression (MLLR) anchor modeling technique which provides performance comparable to the conventional anchor modeling system and yet reduces computation time significantly by computing likelihood efficiently using sufficient statistics of data and anchor specific MLLR matrix. However, both these systems still require a Gaussian Mixture Model-Universal Background Model (GMM-UBM) based back-end system to choose the optimal speaker, which is computationally heavy. In our proposed method, we show that applying Linear-Discriminant Analysis (LDA) and Within-Class-Covariance Normalization (WCCN) on the Speaker characterization Vector (SCV) of our recently proposed Fast-MLLR method, we can combine the computational efficiency and the discriminant capability to have a system that uses simple cosine-distance measure to identify speakers and yet has significantly superior performance compared to both full-blown GMM-UBM system and the anchor-model system. More importantly, there is no need of the “back-end” system. Experimental result on NIST 2004 SRE shows that the proposed method reduces identification error rate by an absolute 2% and takes only 2/3 of the time taken by efficient Fast-MLLR system and only 20% of the time taken by the stand-alone GMM-UBM system.
Achintya Kumar Sarkar, Srinivasan Umesh, Jean-François Bonastre
ICASSP3
2012 Study of the Effect of I-vector Modeling on Short and Mismatch Utterance Duration for Speaker Verification
abstract
International audience
Achintya Kumar Sarkar, Driss Matrouf, Pierre-Michel Bousquet, Jean-François Bonastre
INTERSPEECH4
2011 Discriminant binary data representation for speaker recognition
abstract
In supervector UBM/GMM paradigm, each acoustic file is represented by the mean parameters of a GMM model. This supervector space is used as a data representation space, which has a high dimensionality. Moreover, this space is not intrinsically discriminant and a complete speech segment is represented by only one vector, withdrawing mainly the possibility to take into account temporal or sequential information. This work proposes a new approach where each acoustic frame is represented in a discriminant binary space. The proposed approach relies on a UBM to structure the acoustic space in regions. Each region is then populated with a set of Gaussian models, denoted as “specificities”, able to emphasize speaker specific information. Each acoustic frame is mapped in the discriminant binary space, turning “on” or “off” all the specificities to create a large binary vector. All the following steps, speaker reference extraction, likelihood estimation or decision take place in this binary space. Even if this work is a first step in this avenue, the experiments based on NIST SRE 2008 framework demonstrate the potential of the proposed approach. Moreover, this approach opens the opportunity to rethink all the classical processes using a discrete, binary view.
Jean-François Bonastre, Pierre-Michel Bousquet, Driss Matrouf, Xavier Anguera Miró
ICASSP1
2011 Speaker verification by inexperienced and experienced listeners vs. speaker verification system
abstract
This paper describes the participation of the LIA in the Human Assisted Speaker Recognition (HASR) task of the NIST-SRE 2010 evaluation campaign and its extension to a larger number of listeners.The human performance in such unfavorable conditions is analyzed in relation to the decision of a speaker recognition automatic system. Results of the perception test showed an important inter-trial variability (from 3% to 90% of correct answers for non-target trials) whereas there was no significant difference between the experienced and inexperienced listeners. Some complementarity between speaker verification system and human decisions was also found.
Juliette Kahn, Nicolas Audibert, Solange Rossato, Jean-François Bonastre
ICASSP4
2011 Fast speaker diarization based on binary keys
abstract
Splitting a speech signal into speakers is the main goal of a speaker diarization system, which has become an important building block in many speech processing algorithms. Current state of the art systems are able to obtain good diarization error rates, but most of them are rather slow, which is a strong handicap in applications that require overall faster than real-time processing. In this paper we present a novel speaker diarization system which is built following a bottom-up agglomerative clustering approach and based on speaker binary keys, recently proposed for speaker modeling. After initialization, processing is entirely done over binary vectors and using exclusively binary metrics, which makes the system very fast. On tests performed using all conference meetings datasets released for the NIST RT evaluation campaigns we achieve diarization error rates just slightly worse than a classic acoustic-based system while running over 10 times faster.
Xavier Anguera Miró, Jean-François Bonastre
ICASSP2
2011 Speaker Modeling Using Local Binary Decisions
abstract
International audience
Jean-François Bonastre, Xavier Anguera Miró, Gabriel Hernández Sierra, Pierre-Michel Bousquet
INTERSPEECH1
2011 Intersession Compensation and Scoring Methods in the i-vectors Space for Speaker Recognition
abstract
International audience
Pierre-Michel Bousquet, Driss Matrouf, Jean-François Bonastre
INTERSPEECH3
2011 Modeling nuisance variabilities with factor analysis for GMM-based audio pattern classification
Driss Matrouf, Florian Verdet, Mickael Rouvier, Jean-François Bonastre, Georges Linarès
Comput. Speech Lang.4
2011 Applying SVMs and weight-based factor analysis to unsupervised adaptation for speaker verification
Mitchell McLaren, Driss Matrouf, Robbie Vogt, Jean-François Bonastre
Comput. Speech Lang.4
2010 Beyond Doddington menagerie, a first step towards
abstract
During the last decade, speaker verification systems have shown significant progress and have reached a level of performance and accuracy that support their utilization in practical applications, including the forensic ones. This context emphasizes the importance of a deeper analysis of the system's performance over basic error rate. In this paper, the influence of the speaker (his/her `voice') on the performance is studied and the effect of the model (the training excerpt) is investigated. The experimental setup is based on an open source system and the experimental context of NIST-SRE 2008. The results confirm that the lower performances are obtained from a reduced number of speakers. Even more than speaker factor, speaker verification system performances are shown to be highly dependant on the voice samples used to train speaker models.
Juliette Kahn, Solange Rossato, Jean-François Bonastre
ICASSP3
2010 Model and Score Adaptation for Biometric Systems: Coping With Device Interoperability and Changing Acquisition Conditions
abstract
The performance of biometric systems can be significantly affected by changes in signal quality. In this paper, two types of changes are considered: change in acquisition environment and in sensing devices. We investigated three solutions: (i) model-level adaptation, (ii) score-level adaptation (normalisation), and (iii) the combination of the two, called “compound” adaptation. In order to cope with the above changing conditions, the model-level adaptation attempts to update the parameters of the expert systems (classifiers). This approach requires the authenticity of the candidate samples used for adaptation be known (corresponding to supervised adaptation), or can be estimated (unsupervised adaptation). In comparison, the score-level adaptation merely involves post processing the expert output, with the objective of rendering the associated decision threshold to be dependent only on the class priors despite the changing acquisition conditions. Since the above adaptation strategies treat the underlying biometric experts/classifiers as a black-box, they can be applied to any unimodal or multimodal biometric system, thus facilitating system-level integration and performance optimisation. Our contributions are: (i) proposal of compound adaptation; (ii) investigation and comparison of two different quality-dependent score normalisation strategies; and, (iii) empirical comparison of the merit of the above three solutions on the BANCA face (video) and speech database.
Norman Poh, Josef Kittler, Sébastien Marcel, Driss Matrouf, Jean-François Bonastre
ICPR5
2010 A novel speaker binary key derived from anchor models
abstract
The approach presented in this paper represents voice recordings by a novel acoustic key composed only of binary values. Except for the process being used to extract such keys, there is no need for acoustic modeling and processing in the approach proposed, as all the other elements in the system are based on the binary vectors. We show that this binary key is able to effectively model a speaker’s voice and to distinguish it from other speakers. Its main properties are its small size compared to current speaker modeling techniques and its low computational cost when comparing different speakers as it is limited to obtaining a similarity metric between two binary vectors. Furthermore, the binary key vector extraction process does not need any threshold and offers the opportunity to set the decision steps in a well defined binary domain where scores and decisions are easy to interpret and implement. Index Terms: binary key, speaker modeling, biometrics
Xavier Anguera Miró, Jean-François Bonastre
INTERSPEECH2
2010 Decoupling session variability modelling and speaker characterisation
abstract
The Factor Analysis framework demonstrated its high power to model session variability during the past years. However, training the FA parameters implies to have a large amount of training data. When the size of the available database is limited, the number of components of the core statistical model, the UBM, is also limited as the UBM drives the dimension of the FA main matrix. As the size of the UBM gives directly the size of the speaker supervector (concatenation of the GMM mean parameters), it limits also the intrinsic capacity of the recognition system , reducing the performance expectation. This paper aims to withdraw this limitation by breaking the intrinsic link between the FA dimensionality and the UBM dimensionality. The session variability modelling is done on a smaller dimension compared to the UBM, which drives the discriminative power of the system. The first experimental results proposed in this paper, done using the NIST-SRE 2008 framework, are encouraging with a relative EER improvement of about 18% when a 512 components UBM is associated to a 32 components session variability modelling compared with a 32 components UBM associated with the same variability modelling.
Anthony Larcher, Christophe Lévy, Driss Matrouf, Jean-François Bonastre
INTERSPEECH4
2010 Topological representation of speech for speaker recognition
abstract
International audience
Gabriel Hernández Sierra, Jean-François Bonastre, Driss Matrouf, José Ramón Calvo de Lara
INTERSPEECH2
2010 Channel detectors for system fusion in the context of NIST LRE 2009
abstract
One of the difficulties in Language Recognition is the variability of the speech signal due to speakers and channels. If channel mismatch is too big and when different categories of channels can be identified, one possibility is to build a separate language recognition system for each category and then to fuse them together. This article uses a system selector that takes, for each utterance, the scores of one of the channel-category dependent systems. This selection is guided by a channel detector. We analyze different ways to design such channel detectors: based on cepstral features or on the Factor Analysis channel variability term. The systems are evaluated in the context of NIST’s LRE 2009 and run at 1:65% minCavg for a subset of 8 languages and at 3:85% minCavg for the 23 language setup. Index Terms: language recognition, channel, channel category, fusion, factor analysis, channel detector.
Florian Verdet, Driss Matrouf, Jean-François Bonastre, Jean Hennebert
INTERSPEECH3
2010 The DesPho-APaDy Project: Developing an Acoustic-phonetic Characterization of Dysarthric Speech in French
Cécile Fougeron, Lise Crevier-Buchman, Corinne Fredouille, Alain Ghio, Christine Meunier, Claude Chevrie-Muller, Jean-François Bonastre, Antonia Colazo-Simon, Céline De Looze, Danielle Duez, Cédric Gendrot, Thierry Legou, Nathalie Lévêque, Claire Pillot-Loiseau, Serge Pinto, Gilles Pouchoulin, Danièle Robert, Jacqueline Vaissière, François Viallet, Coralie Vincent
LREC7
2009 Feature Selection Based on Information Theory for Speaker Verification
Rafael Fernández, Jean-François Bonastre, Driss Matrouf, José Ramón Calvo de Lara
CIARP2
2009 Speaker diarization using unsupervised discriminant analysis of inter-channel delay features
abstract
When multiple microphones are available estimates of inter-channel delay, which characterise a speaker's location, can be used as features for speaker diarization. Background noise and reverberation can, however, lead to noisy features and poor performance. To ameliorate these problems, this paper presents a new approach to the discriminant analysis of delay features for speaker diarization. This novel and nonetheless unsupervised approach aims to increase speaker separability in delay-space. We assess the approach on subsets of four standard NIST RT datasets and demonstrate a relative improvement in diarization error rate of 25% on a separate evaluation set using delay features alone.
Nicholas W. D. Evans, Corinne Fredouille, Jean-François Bonastre
ICASSP3
2009 Factor analysis and SVM for language recognition
abstract
International audience
Florian Verdet, Driss Matrouf, Jean-François Bonastre, Jean Hennebert
INTERSPEECH3
2008 Reinforced temporal structure information for embedded utterance-based speaker recognition
abstract
Embedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of com-puting resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a rel-atively large quantity of speech data, but with negligible har-nessing of linguistic content. The proposed approach addresses these limitations and takes advantage from the linguistic nature of the speech material into the GMM/UBM framework by us-ing client-customised utterances. The GMM/UBM is then rein-forced with new temporal information. Experiments on the MyIdea database are performed when im-postors know the client-utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 45 % in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations. 1.
Anthony Larcher, Jean-François Bonastre, John S. D. Mason
INTERSPEECH2
2008 Factor analysis multi-session training constraint in session compensation for speaker verification
abstract
For a few years now, the problem of session variability in text-independent automatic speaker verification is being tackled actively. A new paradigm based on a Latent Factor Analysis (LFA) model has been applied successfully for this task. However, using this approach, a large training corpus with several sessions per speaker is required. This constraint is hard to satisfy in many real applications. In this paper, we try to analyze if the LFA paradigm still holds even when the constraint of multiple sessions per speaker isn’t satisfied. We propose to study two approaches. The first one consists in using the basic paradigm of the LFA model and the second one is founded on a new interpretation of the interaction between the session and the speaker. The experiments were carried out with NIST SRE 2005 and 2006 protocols. We show that even with only one session per speaker the gain obtained by LFA session compensation (with the two strategies) is still very important. 1
Driss Matrouf, Jean-François Bonastre, Salah Eddine Mezaache
INTERSPEECH2
2008 Combining continuous progressive model adaptation and factor analysis for speaker verification
abstract
International audience
Mitchell McLaren, Driss Matrouf, Robbie Vogt, Jean-François Bonastre
INTERSPEECH4
2008 Analysis of impostor tests with high scores in NIST-SRE context
abstract
International audience
Salah Eddine Mezaache, Jean-François Bonastre, Driss Matrouf
INTERSPEECH2
2008 Dysphonic voices and the 0-3000 hz frequency band
abstract
Concerned with pathological voice assessment, this paper aims at characterizing dysphonia in the frequency domain for a better understanding of related phenomena while most of the studies have focused only on improving classification systems for diag-nosis help purposes. Based on a first study which demonstrates that the low frequencies ([0-3000]Hz) are more relevant for dys-phonia discrimination compared with higher frequencies, the authors propose in this paper to pursue by analyzing the impact of the restricted frequency band ([0-3000]Hz) on the dysphonic voice discrimination from a phonetical and perceptual point of views. A discussion around the frequency band limitation of telephone channel is also proposed. Index Terms: Voice disorder, dysphonia characterization, au-tomatic dysphonic voice classification, frequency analysis
Gilles Pouchoulin, Corinne Fredouille, Jean-François Bonastre, Alain Ghio, Antoine Giovanni
INTERSPEECH3
2008 Short utterance-based video aided speaker recognition
abstract
Embedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of computing resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a relatively large quantity of speech data, but with negligible harnessing of linguistic content. The proposed approach addresses these limitations and takes advantage of the linguistic nature of the speech material into the GMM/UBM framework by using clientcustomised utterances. Furthermore, the acoustic structure is then reinforced with video information. Experiments on the MyIdea database are performed when impostors know the client utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 47% in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations.
Anthony Larcher, Jean-François Bonastre, John S. D. Mason
MMSP2
2007 Complementary approaches for voice disorder assessment
abstract
This paper describes two comparative studies of voice quality assessment based on complementary approaches. The first study was undertaken on 449 speakers (including 391 dysphonic patients) whose voice quality was evaluated in parallel by a perceptual judgment and objective measurements on acoustic and aerodynamic data. Results showed that a nonlinear combination of 7 parameters allowed the classification of 82% voice samples in the same grade as the jury. The second study relates to the adaptation of Automatic Speaker Recognition (ASR) techniques to pathological voice assessment. The system designed for this particular task relies on a GMM based approach, which is the state-of-the-art for ASR. Experiments conducted on 80 female voices provide promising results, underlining the interest of such an approach. We benefit from the multiplicity of theses techniques to evaluate the methodological situation which points fundamental differences between these complementary approaches (bottom-up vs. top-down, global vs. analytic). We also discuss some theoretical aspects about relationship between acoustic measurement and perceptual mechanisms which are often forgotten in the performance race.
Jean-François Bonastre, Corinne Fredouille, Alain Ghio, Antoine Giovanni, Gilles Pouchoulin, Joana Revis, Bernard Teston
INTERSPEECH1
2007 Artificial impostor voice transformation effects on false acceptance rates
abstract
This paper investigates the effect of a transfer function-based voice transformation on automatic speaker recognition system performance. We focus on increasing the impostor acceptance rate, by modifying the voice of an impostor in order to target a specific speaker. This paper follows previous works where we demonstrate that, if someone has a knowledge on the speaker recognition method used, it is possible to impersonate a given speaker, in the view of this speaker recognition method. In this paper we extend the previous work by relaxing the needed knowledge on the targeted speaker recognition system. The results show that the voice transformation allows a drastic increase of the false acceptance rate, without damaging the natural perception of the voice, and without needing a large knowledge on the targeted speaker recognition system.
Jean-François Bonastre, Driss Matrouf, Corinne Fredouille
INTERSPEECH1
2007 Influence of task duration in text-independent speaker verification
abstract
International audience
Benoit G. B. Fauve, Nicholas W. D. Evans, Neil Pearson, Jean-François Bonastre, John S. D. Mason
INTERSPEECH4
2007 An interactive timeline for speech database browsing
abstract
Speech databases lack efficient interfaces to explore information along time. We introduce an interactive timeline that helps the user in browsing an audio stream on a large time scale and recontextualize targeted information. Time can be explored at different granularities using synchronized scales. We try to take advantage of automatic transcription to generate a conceptual structure of the database. The timeline is annotated with two elements to reflect the information distribution relevant to a user need. Information density is computed using an information retrieval model and displayed as a continuous shade on the timeline whereas anchorage points are expected to provide a stronger structure and to guide the user through his exploration. These points are generated using an extractive summarization algorithm. We present a prototype implementing the interactive timeline to browse broadcast news recordings.
Benoît Favre, Jean-François Bonastre, Patrice Bellot
INTERSPEECH2
2007 Fast adaptation of GMM-based compact models
abstract
In this paper, a new strategy for a fast adaptation of acoustic models is proposed for embedded speech recognition. It relies on a general GMM, which represents the whole acoustic space, associated with a set of HMM state-dependent probability functions modeled as transformations of this GMM. The work presented here takes advantage of this architecture to propose a fast and efficient way to adapt the acoustic models. The adaptation is performed only on the general GMM model, using techniques gathered from the speaker recognition domain. It does not require state-dependent adaptation data and it is very efficient in terms of computational cost. Weevaluate our approach in the voice-command task, using a car-based corpus. This adaptation method achieved a relative error-rate decrease of about 10% even if few adaptation data are available. The complete system allows a total relative gain of more than 20% compared to a basic HMM-based system. Index Terms: speech recognition, compact acoustic models, adaptation
Christophe Lévy, Georges Linarès, Jean-François Bonastre
INTERSPEECH3
2007 A straightforward and efficient implementation of the factor analysis model for speaker verification
abstract
For a few years, the problem of session variability in textindependent automatic speaker verification is being tackled actively. A new paradigm based on a factor analysis model have successfully been applied for this task. While very efficient, its implementation is demanding. In this paper, the algorithms involved in the eigenchannel MAP model are written down for a straightforward implementation, without referring to previous work or complex mathematics. In addition, a different compensation scheme is proposed where the standard GMM likelihood can be used without any modification to obtain good performance (even without the need of score normalization). The use of the compensated supervectors within a SVM classifier through a distance based kernel is also investigated. Experiments results shows an overall 50 % relative gain over the standard
Driss Matrouf, Nicolas Scheffer, Benoit G. B. Fauve, Jean-François Bonastre
INTERSPEECH4
2007 Information retrieval strategies for accessing african audio corpora
abstract
In this paper we present a first approach to access African oral corpora, combining automatic speech recognition and information retrieval. Firstly, we present the principal characteristics of our Somali speech recognizer [8] and the results obtained on real audio archives gathered from Djibouti Radio. Secondly, we present a Hybrid Language Model (HLM) including words and sub-words to improve the robustness against OOV words. We proceed to Information Retrieval experiments with various strategies. We search on the different outputs of the ASR system (words, sub-words and hybrid). We finally present a new strategy combining sub-words and words to enhance the information retrieval results. Index Terms: speech recognition, information retrieval, hybrid language model, Somali language.
Abdillahi Nimaan, Pascal Nocera, Frédéric Béchet, Jean-François Bonastre
INTERSPEECH4
2007 Frequency study for the characterization of the dysphonic voices
abstract
Concerned with pathological voice assessment, this paper aims at characterizing dysphonia in the frequency domain for a better understanding of relating phenomena while most of the studies have focused only on improving classification systems for diagnosis help purposes.In this context, a GMM-based automatic classification system is applied on different frequency ranges in order to investigate which ones are relevant for dysphonia characterization.Experiment results demonstrate that the low frequencies [0-3000]Hz are more relevant for dysphonia discrimination compared with higher frequencies.
Gilles Pouchoulin, Corinne Fredouille, Jean-François Bonastre, Alain Ghio, Antoine Giovanni
INTERSPEECH3
2007 Confidence measure based unsupervised target model adaptation for speaker verification
abstract
International audience
Alexandre Preti, Jean-François Bonastre, Driss Matrouf, François Capman, Bertrand Ravera
INTERSPEECH2
2007 State-of-the-Art Performance in Text-Independent Speaker Verification Through Open-Source Software
abstract
This paper illustrates an evolution in state-of-the-art speaker verification by highlighting the contribution from newly developed techniques. Starting from a baseline system based on Gaussian mixture models that reached state-of-the-art performances during the NIST'04 SRE, final systems with new intersession compensation techniques show a relative gain of around 50%. This work highlights that a key element in recent improvements is still the classical maximum a posteriori (MAP) adaptation, while the latest compensation methods have a crucial impact on overall performances. Nuisance attribute projection (NAP) and factor analysis (FA) are examined and shown to provide significant improvements. For FA, a new symmetrical scoring (SFA) approach is proposed. We also show further improvement with an original combination between a support vector machine and SFA. This work is undertaken through the open-source ALIZE toolkit.
Benoit G. B. Fauve, Driss Matrouf, Nicolas Scheffer, Jean-François Bonastre, John S. D. Mason
IEEE Trans. Speech Audio Process.4
2006 On the Use of Linguistic Information for Broadcast News Speaker Tracking
abstract
In this paper, we have explored a speaker characterization at two different linguistic levels, the lexical content and the syntactical form, in the context of Broadcast News (BN) speaker tracking task. The modeling of the information is done classically, by a n-gram approach applied on the word sequence for the linguistic content and on the syntactical tags issued from this sequence for the syntactical level (n-class modeling). The experiments were done on a subset of the BN rich transcription French evaluation campaign, ESTER. We have observed that lexical information is not useful for BN data while the syntactical information seems promising as it allows significant speaker identification and verification performance (until 40% of correct identification rate and 35% of EER).
William Antoni, Corinne Fredouille, Jean-François Bonastre
ICASSP (1)3
2006 Effect of Speech Transformation on Impostor Acceptance
abstract
This paper investigates the effect of voice transformation on automatic speaker recognition system performance. We focus on increasing the impostor acceptance rate, by modifying the voice of an impostor in order to target a specific speaker. This paper is based on the following idea: in several applications and particularly in forensic situations, it is reasonable to think that some organizations have a knowledge on the speaker recognition method used and could impersonate a given, well known speaker. This paper presents some experiments based on NIST SRE 2005 protocol and a simple impostor voice transformation method. The results show that this simple voice transformation allows a drastic increase of the false acceptance rate, without a degradation of the natural aspect of the voice
Driss Matrouf, Jean-François Bonastre, Corinne Fredouille
ICASSP (1)2
2006 Imperfect transcript driven speech recognition
abstract
In many cases, textual information can be associated with speech signals such as movie subtitles, theater scenarios, broadcast news summaries etc. This information could be considered as approximated transcripts and corresponds rarely to the exact word utterances. The goal of this work is to use this kind of information to improve the performance of an automatic speech recognition (ASR) system. Multiple applications are possible: to follow a play with closed caption aligned to the voice signal (while respecting to performer variations) to help deaf people, to watch a movie in another language using aligned and corrected closed captions, etc. We propose in this paper a method combining a linguistic analysis of the imperfect transcripts and a dynamic synchronization of these transcripts inside the search algorithm. The proposed technique is based on language model adaptation and on-line synchronization of the search algorithm. Experiments are carried out on an extract of the ESTER evaluation campaign [4] database, using the LIA Broadcast News system. The results show that the transcript-driven system outperforms significantly both the original recognizer and the imperfect transcript itself.
Benjamin Lecouteux, Georges Linarès, Pascal Nocera, Jean-François Bonastre
INTERSPEECH4
2006 GMM-based acoustic modeling for embedded speech recognition
abstract
International audience
Christophe Lévy, Georges Linarès, Jean-François Bonastre
INTERSPEECH3
2006 Automatic transcription of Somali language
Abdillahi Nimaan, Pascal Nocera, Jean-François Bonastre
INTERSPEECH3
2006 Unsupervised model adaptation for speaker verification
abstract
International audience
Alexandre Preti, Jean-François Bonastre
INTERSPEECH2
2006 A multiclass framework for speaker verification within an acoustic event sequence system
abstract
International audience
Nicolas Scheffer, Jean-François Bonastre
INTERSPEECH2
2006 Corpus description of the ESTER Evaluation Campaign for the Rich Transcription of French Broadcast News
Sylvain Galliano, Edouard Geoffrois, Guillaume Gravier, Jean-François Bonastre, Djamel Mostefa, Khalid Choukri
LREC4
2006 Towards automatic transcription of Somali language
Abdillahi Nimaan, Pascal Nocera, Jean-François Bonastre
LREC3
2006 Step-by-step and integrated approaches in broadcast news speaker diarization
Sylvain Meignier, Daniel Moraru, Corinne Fredouille, Jean-François Bonastre, Laurent Besacier
Comput. Speech Lang.4
2005 ALIZE, a free toolkit for speaker recognition
abstract
This paper presents the ALIZE free speaker recognition toolkit. ALIZE is designed and developed within the framework of the ALIZE project, a part of the French Research Ministry Technolangue program. The paper focuses on the innovative aspects of ALIZE and illustrates them by some examples. An experimental validation of the toolkit during the NIST 2004 speaker recognition evaluation campaign is also proposed.
Jean-François Bonastre, Frédéric Wils, Sylvain Meignier
ICASSP (1)1
2005 Application of automatic speaker recognition techniques to pathological voice assessment (dysphonia)
abstract
HAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Corinne Fredouille, Gilles Pouchoulin, Jean-François Bonastre, M. Azzarello, Antoine Giovanni, Alain Ghio
INTERSPEECH3
2005 The ESTER phase II evaluation campaign for the rich transcription of French broadcast news
abstract
This paper gives the final results of the ESTER evaluation campaign which started in 2003 and ended in January 2005. The aim of this campaign was to evaluate automatic broadcast news rich transcription systems for the French language. The evaluation tasks were divided into three main categories: orthographic transcription, event detection and tracking (e.g. speech vs. music, speaker tracking), and information extraction. The last one, limited to named entity detection in this evaluation, was a preliminary test. The paper reports on protocols and gives the results obtained in the campaign. 1.
Sylvain Galliano, Edouard Geoffrois, Djamel Mostefa, Khalid Choukri, Jean-François Bonastre, Guillaume Gravier
INTERSPEECH5
2005 Broadcast news speaker tracking for ESTER 2005 campaign
abstract
This paper presents the speaker tracking system of the LIA laboratory, validated during ESTER 2005 campaign on a radio broadcast news corpus of about 90 h. The LIA speaker tracking system firstly uses an acoustic class segmentation in order to suppress non speech frames and to detect the speech conditions. Secondly, a speaker diarization process is applied in order to provide speaker detection system (the last step) with speaker homogeneous segments (boundaries and clustering). The speaker detection system uses UBM/GMM likelihood ratios in order to decide if a segment belongs to one tracked speaker. The speaker tracking system is presented and some results obtained during ESTER 2005 campaign are proposed. The presented systems are based on the ALIZE platform (Automatic speaker recognition C++ library).
Dan Istrate, Nicolas Scheffer, Corinne Fredouille, Jean-François Bonastre
INTERSPEECH4
2005 Speaker detection using acoustic event sequences
abstract
Novel approaches using high level features have recently shown up in the speaker recognition field. They basically consist in modeling speakers using linguistic features such as words, phonemes, idiolects. The benefit of these features was demonstrated in NIST campaigns. Their main disadvantage is their need of a huge amount of data to be efficient. The purpose of this study is to generalize this approach by using acoustic events, generated by a GMM, as input features. A methodology to build a dictionary and to model speakers using symbol sequences from this dictionary is derived. Different experiments on NIST SRE 2004 database show that the information produced is speaker specific and that a fusion experiment with a GMM verification system improves performance.
Nicolas Scheffer, Jean-François Bonastre
INTERSPEECH2
2004 Reducing computational and memory cost for cellular phone embedded speech recognition system
abstract
We present several methods able to fit speech recognition system requirements to cellular phone resources. The proposed techniques are evaluated on a digit recognition task using both French and English corpora. We investigate particularly three aspects of speech processing: acoustic parameterization, recognition algorithms; acoustic modeling. Several parameterization algorithms (LPCC, MFCC and PLP) are compared to the linear predictive coding (LPC) included in the GSM norm. The MFCC and PLP parameterization algorithms perform significantly better than the others. Moreover, feature vector size can be reduced to 6 PLP coefficients, allowing memory and computation resources to be decreased without a significant loss of performance. In order to achieve good performance with reasonable resource needs, we develop several methods to embed a classical HMM-based speech recognition system in a cellular phone. We first propose an automatic on-line building of a phonetic lexicon which allows a minimal but unlimited lexicon. Then we reduce the HMM complexity by decreasing the number of (Gaussian) components per state. Finally, we evaluate our propositions by comparing dynamic time warping (DTW) with our HMM system - in the cellular phone context - for clean conditions. The experiments show that our HMM system outperforms DTW for speaker independent tasks and allows more practical applications for the cellular-phone user interface.
Christophe Lévy, Georges Linarès, Pascal Nocera, Jean-François Bonastre
ICASSP (5)4
2004 Benefits of prior acoustic segmentation for automatic speaker segmentation
abstract
The paper investigates the interest of segmentation in acoustic macro classes (like gender or bandwidth) as front-end processing for the segmentation/diarization task. The impact of this prior acoustic segmentation is evaluated in terms of speaker diarization performance in the particular context of NIST RT'03 evaluation (done on the HUB4 broadcast news corpora). It is rarely discussed in the literature, but our work shows that the application of prior acoustic segmentation, in a similar way to the automatic speech recognition task, may be very useful to the speaker segmentation task. Experiments were conducted using two different kinds of speaker segmentation systems developed individually by the LIA and CLIPS laboratories in the framework of the ELISA consortium. For both systems, improvement was observed when combined with prior acoustic segmentation. However, a larger impact, in terms of performance, is observed on the LIA system based on an ascending/HMM approach compared to the CLIPS system based on speaker turn detection.
Sylvain Meignier, Daniel Moraru, Corinne Fredouille, Laurent Besacier, Jean-François Bonastre
ICASSP (1)5
2004 The ELISA consortium approaches in broadcast news speaker segmentation during the NIST 2003 rich transcription evaluation
abstract
The paper presents the ELISA consortium activities in automatic speaker segmentation, also known as speaker diarization, during the NIST rich transcription (RT), 2003, evaluation. The experiments were conducted on real broadcast news data (HUB4). Two different approaches from the CLIPS and LIA laboratories are presented and different possibilities of combining them are investigated, in the framework of the ELISA consortium. The system submitted as an ELISA primary system obtained the second lowest segmentation error rate compared to the other RT03-participant primary systems. Another ELISA system submitted as a secondary system outperformed the best primary system and obtained the lowest speaker segmentation error rate.
Daniel Moraru, Sylvain Meignier, Corinne Fredouille, Laurent Besacier, Jean-François Bonastre
ICASSP (1)5
2004 The ESTER Evaluation Campaign for the Rich Transcription of French Broadcast News
Guillaume Gravier, Jean-François Bonastre, Edouard Geoffrois, Sylvain Galliano, Kevin McTait, Khalid Choukri
LREC2
2003 Structural speaker adaptation using maximum a posteriori approach and a Gaussian distributions merging technique
abstract
The aim of speaker adaptation techniques is to enhance speaker-independent acoustic models to bring their recognition accuracy as close as possible to the one obtained with speaker-dependent models. Recently, a technique based on a hierarchical structure and the maximum a posteriori criterion was proposed (SMAP) (Shinoda, K. and Lee, C.-H., Proc IEEE ICASSP, 1998). As in SMAP, we assume that the acoustic model parameters are organized in a tree containing all the Gaussian distributions. Each node in that tree represents a cluster of Gaussian distributions sharing a common affine transformation representing the mismatch between training and test conditions. To estimate this affine transformation, we propose a new technique based on merging Gaussians and the standard MAP adaptation. This new technique is very fast and allows a good unsupervised adaptation for both means and variances even with a small amount of adaptation data. This adaptation strategy has shown a significant performance improvement in a large vocabulary speech recognition task, alone and combined with the MLLR (maximum likelihood linear regression) adaptation.
Olivier Bellot, Driss Matrouf, Pascal Nocera, Georges Linarès, Jean-François Bonastre
ICASSP (2)5
2003 Speaker detection using multi-speaker audio files for both enrollment and test
abstract
This paper focuses on speaker detection using multispeaker files both for the enrollment phase and for the test phase. This task was introduced during the 2002 NIST speaker recognition evaluation campaign. Enrollment data is composed of three two-speaker files. Test files are also two-speaker records. The system presented here uses a speaker segmentation process based on an HMM conversation model followed by a speaker matching technique to produce one-speaker segments. Speaker detection is then achieved using AMIRAL, LIA's GMM-based speaker verification system. Validation of the proposed strategy is done using extracts from the NIST 2002 results.
Jean-François Bonastre, Sylvain Meignier, Téva Merlin
ICASSP (2)1
2003 The ELISA consortium approaches in speaker segmentation during the NIST 2002 speaker recognition evaluation
abstract
This paper presents the ELISA consortium activities in automatic speaker segmentation during last NIST 2002 evaluation: two different approaches from CLIPS and LIA laboratories are presented and the possibility of combining them either by applying them consecutively, or by fusing the decisions made by each of them, is investigated. Various types of data were available for NIST 2002. The ELISA systems obtained the lower error rates for two corpora: the CLIPS system obtained the best performance on the Meeting data, the LIA system obtained the best performance on the Switchboard data. The combining strategies proposed in this paper allowed us to improve the performance of the best single system on both data types (up to 30 % of error rate reduction).
Daniel Moraru, Sylvain Meignier, Laurent Besacier, Jean-François Bonastre, Ivan Magrin-Chagnolleau
ICASSP (2)4
2003 Person authentication by voice: a need for caution
abstract
Because of recent events and as members of the scientific community working in the field of speech processing, we feel compelled to publicize our views concerning the possibility of identifying or authenticating a person from his or her voice. The need for a clear and common message was indeed shown by the diversity of information that has been circulating on this matter in the media and general public over the past year. In a press release initiated by the AFCP and further elaborated in collaboration with the SpLC ISCA-SIG, the two groups herein discuss and present a summary of the current state of scientific knowledge and technological development in the field of speaker recognition, in accessible wording for nonspecialists. Our main conclusion is that, despite the existence of technological solutions to some constrained applications, at the present time, there is no scientific process that enables one to uniquely characterize a person’s voice or to identify with absolute certainty an individual from his or her voice. 1.
Jean-François Bonastre, Frédéric Bimbot, Louis-Jean Boë, Joseph P. Campbell, Douglas A. Reynolds, Ivan Magrin-Chagnolleau
INTERSPEECH1
2003 Gaussian dynamic warping (GDW) method applied to text-dependent speaker detection and verification
abstract
International audience
Jean-François Bonastre, Philippe Morin, Jean-Claude Junqua
INTERSPEECH1
2003 Structural linear model-space transformations for speaker adaptation
abstract
Within the framework of speaker-adaptation, a technique based on tree structure and the maximum a posteriori criterion was proposed (SMAP). In SMAP, the parameters estimation, at each node in the tree is based on the assumption that the mismatch between the training and adaptation data is a Gaussian PDF which parameters are estimated by using the Maximum Likelihood criterion. To avoid poor transformation parameters estimation accuracy due to an insufcienc y of adaptation data in a node, we propose a new technique based on the maximum a posteriori approach and PDF Gaussians Merging. The basic idea behind this new technique is to estimate an afne transformations which bring the training acoustic models as close as possible to the test acoustic models rather than transformation maximizing the likelihood of the adaptation data. In this manner, even with very small amount of adaptation data, the parameters transformations are accurately estimated for means and variances. This adaptation strategy has shown a signicant performance improvement in a large vocabulary speech recognition task, alone and combined with the MLLR adaptation.
Driss Matrouf, Olivier Bellot, Pascal Nocera, Georges Linarès, Jean-François Bonastre
INTERSPEECH5
2002 Speaker utterances tying among speaker segmented audio documents using hierarchical classification: towards speaker indexing of audio databases
abstract
International audience
Sylvain Meignier, Jean-François Bonastre, Ivan Magrin-Chagnolleau
INTERSPEECH2
2001 SPeaker and language characterization (spLC): a special interest group (SIG) of ISCA
abstract
Last year, SpLC - an ISCA Special Interest Group (SIG) centered around Speaker and Language Characterization was born. The aims of this paper are to present the SpLC SIG, its objectives, and the work done during the first year.
Jean-François Bonastre, Ivan Magrin-Chagnolleau, Stephan Euler, François Pellegrino, Régine André-Obrecht, John S. D. Mason, Frédéric Bimbot
INTERSPEECH1
2001 A posteriori and a priori transformations for speaker adaptation in large vocabulary speech recognition systems
abstract
International audience
Driss Matrouf, Olivier Bellot, Pascal Nocera, Georges Linarès, Jean-François Bonastre
INTERSPEECH5
2000 A speaker tracking system based on speaker turn detection for NIST evaluation
abstract
A speaker tracking system (STS) is built by using successively a speaker change detector and a speaker verification system. The aim of the STS is to find in a conversation between several persons (some of them having already enrolled and other being totally unknown) target speakers chosen in a set of enrolled users. In a first step, speech is segmented into homogeneous segments containing only one speaker, without any use of a priori knowledge about speakers. Then, the resulting segments are checked to belong to one of the target speakers. The system has been used in a NIST evaluation test with satisfactory results.
Jean-François Bonastre, Perrine Delacourt, Corinne Fredouille, Téva Merlin, Christian Wellekens
ICASSP1
2000 Evolutive HMM for multi-speaker tracking system
abstract
Seeking within a speech sequence the speaker utterances is one of the main tasks of indexing. In this paper, the proposed speaker tracking system is defined in the case where all speaker identities are known beforehand. The conversation is modeled as an evolutive HMM-like model, in which speaker models computed are added one by one. A temporary indexing is proposed after each speaker adding and then challenged at the next step. This process is iterated until all the speakers are detected. The system has been assessed using multi-speaker messages generated by concatenation of Switchboard mono-speaker segments. The obtained results show the potentiality of the proposed solution.
Sylvain Meignier, Jean-François Bonastre, Corinne Fredouille, Téva Merlin
ICASSP2
2000 Additive and convolutional noises compensation for speaker recognition
Olivier Bellot, Driss Matrouf, Téva Merlin, Jean-François Bonastre
INTERSPEECH4
2000 Subband architecture for automatic speaker recognition
Laurent Besacier, Jean-François Bonastre
Signal Process.2
2000 Localization and selection of speaker-specific information with statistical modeling
Laurent Besacier, Jean-François Bonastre, Corinne Fredouille
Speech Commun.2
1999 Similarity normalization method based on world model and a posteriori probability for speaker verification
Corinne Fredouille, Jean-François Bonastre, Téva Merlin
EUROSPEECH2
1998 Frame pruning for speaker recognition
abstract
In this paper, we propose a frame selection procedure for text-independent speaker identification. Instead of averaging the frame likelihoods along the whole test utterance, some of these are rejected (pruning) and the final score is computed with a limited number of frames. This pruning stage requires a prior frame level likelihood normalization in order to make comparison between frames meaningful. This normalization procedure alone leads to a significant performance enhancement. As far as pruning is concerned, the optimal number of frames pruned is learned on a tuning data set for normal and telephone speech. Validation of the pruning procedure on 567 speakers leads to a 27% identification rate improvement on TIMIT, and to 17% on NTIMIT.
Laurent Besacier, Jean-François Bonastre
ICASSP2
1998 Time and frequency pruning for speaker identification
abstract
This work is an attempt to refine decisions in speaker identification. A test utterance is divided into multiple time-frequency blocks on which a normalized likelihood score is calculated. Instead of averaging the block-likelihoods along the whole test utterance, some of them are rejected (pruning) and the final score is computed with a limited number of time-frequency blocks. The results obtained in the special case of time pruning lead the authors to experiment a joint time and frequency pruning approach. The optimal percentage of blocks pruned is learned on a tuning data set with the minimum identification error criterion. Validation of the time-frequency pruning process on 567 speakers leads to a significant error rate reduction for short training and test duration.
Laurent Besacier, Jean-François Bonastre
ICPR2
1995 Effect of utterance duration and phonetic content on speaker identification using second-order statistical methods
abstract
International audience
Ivan Magrin-Chagnolleau, Jean-François Bonastre, Frédéric Bimbot
EUROSPEECH2
1993 Automatic speaker recognition and analytic process
Jean-François Bonastre, Henri Meloni
EUROSPEECH1
1991 Analytical strategy for speaker identification
Jean-François Bonastre, Henri Meloni, Philippe Langlais
EUROSPEECH1