Driss Matrouf

dblp:53/1360 · DBLP profile ↗
← Back
66ranked-venue papers
10as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 50 · 7 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Unified Text and Speaker Verification using SSL model for Text-Dependent Speaker Verification
Nathan Griot, Driss Matrouf, Raphaël Blouet, Jean-François Bonastre, Ana Mantecon
INTERSPEECH2
2024 RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile Robot
abstract
In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. RoboVox is a French corpus recorded by a mobile robot. The files are recorded from different distances under severe acoustical conditions with the presence of several types of noise and reverberation. In addition to noise and reverberation, the robot’s internal noise acts as an extra additive noise. RoboVox can be used for both single-channel and multi-channel speaker recognition. In the evaluation protocols, we are considering both cases. The obtained results demonstrate a significant decline in performance in far-filed speaker recognition and urge the community to further research in this domain
Mohammad MohammadAmini, Driss Matrouf, Mickael Rouvier, Jean-François Bonastre, Romain Serizel, Théophile Gonos
LREC/COLING2
2023 Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline
abstract
The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representation fed into a HiFiGAN vocoder is protected using a neural-discriminant analysis approach, which is consistent with the zero-evidence concept of privacy. This approach significantly reduces the information in speech related to the speaker’s sex while preserving speech content and some consistency in the resulting protected voices.
Paul-Gauthier Noé, Xiaoxiao Miao, Xin Wang 0037, Junichi Yamagishi, Jean-François Bonastre, Driss Matrouf
ICASSP6
2022 A Bridge between Features and Evidence for Binary Attribute-Driven Perfect Privacy
abstract
Attribute-driven privacy aims to conceal a single user’s attribute, contrary to anonymisation that tries to hide the full identity of the user in some data. When the attribute to protect from malicious inferences is binary, perfect privacy requires the log-likelihood-ratio to be zero resulting in no strength-of-evidence. This work presents an approach based on normalizing flow that maps a feature vector into a latent space where the evidence, related to the binary attribute, and an independent residual are disentangled. It can be seen as a non-linear discriminant analysis where the mapping is invertible al-lowing generation by mapping the latent variable back to the original space. This framework allows to manipulate the log-likelihood-ratio of the data and therefore allows to set it to zero for privacy. We show the applicability of the approach on an attribute-driven privacy task where the sex information is removed from speaker embeddings. Results on VoxCeleb2 dataset show the efficiency of the method that outperforms in terms of privacy and utility our previous experiments based on adversarial disentanglement.
Paul-Gauthier Noé, Andreas Nautsch, Driss Matrouf, Pierre-Michel Bousquet, Jean-François Bonastre
ICASSP3
2022 Barlow Twins self-supervised learning for robust speaker recognition
abstract
International audience
Mohammad MohammadAmini, Driss Matrouf, Jean-François Bonastre, Sandipana Dowerah, Romain Serizel, Denis Jouvet
INTERSPEECH2
2022 Joint Optimization of Diffusion Probabilistic-Based Multichannel Speech Enhancement with Far-Field Speaker Verification
abstract
Smart devices using speaker verification are getting equipped with multiple microphones, improving spatial ambiguity and directivity. However, unlike other speech-based applications, the performance of speaker verification degrades in far-field scenarios due to the adverse effects of a noisy environment and room reverberation. This paper presents a novel diffusion probabilistic models-based multichannel speech enhancement as a front-end for the ECAPA-TDNN speaker verification system in a far-field noisy-reverberant scenario. The proposed approach incorporates a two-stage training approach. In the first stage, we individually train the speech enhancement and speaker verification modules. In the second stage, we combined both modules and trained them jointly. We use similarity-preserving knowledge distillation loss that guides the network to produce similar activation for enhanced signals like clean signals. Joint optimization achieved the best results on synthetic and VOiCES datasets.
Sandipana Dowerah, Romain Serizel, Denis Jouvet, Mohammad MohammadAmini, Driss Matrouf
SLT5
2022 Towards a unified assessment framework of speech pseudonymisation
abstract
Anonymisation and pseudonymisation are two similar concepts used in privacy preservation for speech data. With no established definitions for these tasks, nor standard approaches to assessment, this paper provides definitions and presents two complementary assessment frameworks. The first is based on voice similarity matrices which provide both an immediate visualisation of privacy protection performance at the speaker level and two objective measures in the form of de-identification and voice distinctiveness preservation. The approach readily highlights imbalances in system performance at the speaker level. The second, referred to as the zero evidence biometric recognition assessment (ZEBRA) framework, is based on information theory and measures the amount of private information disclosed in speech data. The paper presents also an extension to the original ZEBRA framework. It aims to reflect the robustness of the privacy safeguard when a privacy adversary adapts to the protected speech. We demonstrate the application of both frameworks to assess pseudonymisation performance on the two VoicePrivacy 2020 challenge baseline solutions plus a third one. The two frameworks were designed independently of each other. The ZEBRA framework is fully consistent with the Bayesian decision theory and the other framework focuses instead on speaker-wise visualisations of a system performance. Thus, while metrics derived from them bear similarities, they expose differences in safeguard behavior. The assessment of pseudonymisation remains challenging and merits greater attention in the future.
Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Natalia A. Tomashenko, Driss Matrouf
Comput. Speech Lang.7
2021 Adversarial Disentanglement of Speaker Representation for Attribute-Driven Privacy Preservation
abstract
In speech technologies, speaker's voice representation is used in many applications such as speech recognition, voice conversion, speech synthesis and, obviously, user authentication. Modern vocal representations of the speaker are based on neural embeddings. In addition to the targeted information, these representations usually contain sensitive information about the speaker, like the age, sex, physical state, education level or ethnicity. In order to allow the user to choose which information to protect, we introduce in this paper the concept of attribute-driven privacy preservation in speaker voice representation. It allows a person to hide one or more personal aspects to a potential malicious interceptor and to the application provider. As a first solution to this concept, we propose to use an adversarial autoencoding method that disentangles in the voice representation a given speaker attribute thus allowing its concealment. We focus here on the sex attribute for an Automatic Speaker Verification (ASV) task. Experiments carried out using the VoxCeleb datasets have shown that the proposed method enables the concealment of this attribute while preserving ASV ability.
Paul-Gauthier Noé, Mohammad MohammadAmini, Driss Matrouf, Titouan Parcollet, Andreas Nautsch, Jean-François Bonastre
Interspeech3
2020 Speech Pseudonymisation Assessment Using Voice Similarity Matrices
abstract
The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of rich, personal and potentially sensitive information. Anonymisation, the focus of the recent VoicePrivacy initiative, is one strategy to protect speaker identity information. Pseudonymisation solutions aim not only to mask the speaker identity and preserve the linguistic content, quality and naturalness, as is the goal of anonymisation, but also to preserve voice distinctiveness. Existing metrics for the assessment of anonymisation are ill-suited and those for the assessment of pseudonymisation are completely lacking. Based upon voice similarity matrices, this paper proposes the first intuitive visualisation of pseudonymisation performance for speech signals and two novel metrics for objective assessment. They reflect the two, key pseudonymisation requirements of de-identification and voice distinctiveness.
Paul-Gauthier Noé, Jean-François Bonastre, Driss Matrouf, Natalia A. Tomashenko, Andreas Nautsch, Nicholas W. D. Evans
INTERSPEECH3
2019 Preserving privacy in speaker and speech characterisation
abstract
Speech recordings are a rich source of personal, sensitive data that can be used to support a plethora of diverse applications, from health profiling to biometric recognition. It is therefore essential that speech recordings are adequately protected so that they cannot be misused. Such protection, in the form of privacy-preserving technologies, is required to ensure that: (i) the biometric profiles of a given individual (e.g., across different biometric service operators) are unlinkable; (ii) leaked, encrypted biometric information is irreversible, and that (iii) biometric references are renewable. Whereas many privacy-preserving technologies have been developed for other biometric characteristics, very few solutions have been proposed to protect privacy in the case of speech signals. Despite privacy preservation this is now being mandated by recent European and international data protection regulations. With the aim of fostering progress and collaboration between researchers in the speech, biometrics and applied cryptography communities, this survey article provides an introduction to the field, starting with a legal perspective on privacy preservation in the case of speech data. It then establishes the requirements for effective privacy preservation, reviews generic cryptography-based solutions, followed by specific techniques that are applicable to speaker characterisation (biometric applications) and speech characterisation (non-biometric applications). Glancing at non-biometrics, methods are presented to avoid function creep, preventing the exploitation of biometric information, e.g., to single out an identity in speech-assisted health care via speaker characterisation. In promoting harmonised research, the article also outlines common, empirical evaluation metrics for the assessment of privacy-preserving technologies for speech data.
Andreas Nautsch, Abelino Jiménez, Amos Treiber, Jascha Kolberg, Catherine Jasserand, Els Kindt, Héctor Delgado, Massimiliano Todisco, Mohamed Amine Hmani, Aymen Mtibaa, Mohammed Ahmed Abdelraheem, Alberto Abad, Francisco Teixeira, Driss Matrouf, Marta Gomez-Barrero, Dijana Petrovska-Delacrétaz, Gérard Chollet, Nicholas W. D. Evans, Christoph Busch 0001
Comput. Speech Lang.14
2018 A Unified Joint Model to Deal With Nuisance Variabilities in the i-Vector Space
abstract
The past decade has witnessed a significant improvement in speaker recognition (SR) technology in terms of performance with the introduction of the i-vectors framework. Despite these advances, the performance of SR systems considerably suffers in the presence of acoustic nuisances and variabilities. In this paper, we develop a data-driven nuisance compensation technique in the i-vector space without referring to the effects of the targeted nuisances in the temporal domain. This approach is nonparametric as it does not suppose a specific relationship between a “good” version of an i-vector and its corrupted version. Instead, our algorithm models directly the joint distribution of both representations (the good i-vector and its corrupted version) and takes advantage of the reproducibility of acoustic corruptions to generate the corrupted i-vectors. We then build an MMSE estimator that computes an improved version of a corrupted test i-vector, given this joint distribution. Experiments are carried out on NIST SRE 2010 and speakers in the wild databases where the proposed algorithm is used to deal with additive noise and short utterances. Our technique is shown to be efficient, improving the baseline system performance in terms of equal-error rate by up to 70% when used on known test noises and up to 65% in the context of unseen noises using a generic model. It was also proven efficient in the context of duration mismatch reaching up to 40% of relative improvement when used on short utterances using multiple models corresponding to different durations and up to 36% when used on arbitrary duration test segments.
Waad Ben Kheder, Driss Matrouf, Moez Ajili, Jean-François Bonastre
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah
INTERSPEECH45
2017 Fast i-vector denoising using MAP estimation and a noise distributions database for robust speaker recognition
Waad Ben Kheder, Driss Matrouf, Pierre-Michel Bousquet, Jean-François Bonastre, Moez Ajili
Comput. Speech Lang.2
2016 LIA System for the SITW Speaker Recognition Challenge
Waad Ben Kheder, Moez Ajili, Pierre-Michel Bousquet, Driss Matrouf, Jean-François Bonastre
INTERSPEECH4
2016 Probabilistic Approach Using Joint Long and Short Session i-Vectors Modeling to Deal with Short Utterances for Speaker Recognition
Waad Ben Kheder, Driss Matrouf, Moez Ajili, Jean-François Bonastre
INTERSPEECH2
2016 Probabilistic Approach Using Joint Clean and Noisy i-Vectors Modeling for Speaker Recognition
Waad Ben Kheder, Driss Matrouf, Moez Ajili, Jean-François Bonastre
INTERSPEECH2
2015 Additive noise compensation in the i-vector space for speaker recognition
abstract
State-of-the-art speaker recognition systems performance degrades considerably in noisy environments even though they achieve very good results in clean conditions. In order to deal with this strong limitation, we aim in this work to remove the noisy part of an i-vector directly in the i-vector space. Our approach offers the advantage to operate only at the i-vector extraction level, letting the other steps of the system unchanged. A maximum a posteriori (MAP) procedure is applied in order to obtain clean version of the noisy i-vectors taking advantage of prior knowledge about clean i-vectors distribution. To perform this MAP estimation, Gaussian assumptions over clean and noise i-vectors distributions are made. Operating on NIST 2008 data, we show a relative improvement up to 60% compared with baseline system. Our approach also outperforms the “multi-style” backend training technique. The efficiency of the proposed method is obtained at the price of relative high computational cost. We present at the end some ideas to improve this aspect.
Waad Ben Kheder, Driss Matrouf, Jean-François Bonastre, Moez Ajili, Pierre-Michel Bousquet
ICASSP2
2015 A comparison of normalization techniques applied to latent space representations for speech analytics
abstract
In the context of noisy environments, Automatic Speech Recognition (ASR) systems usually produce poor transcription quality which also negatively impact performance of speech analyt-ics. Various methods have then been proposed to compensate the bad effect of ASR errors, mainly by projecting transcribed words in an abstract space. In this paper, we seek to identify themes from dialogues of telephone conversation services using latent topic-spaces estimated from a latent Dirichlet allocation (LDA). As an outcome, a document can be represented with a vector containing probabilities to be associated to each topic estimated with LDA. This vector should nonetheless be normalized to condition document representations. We propose to compare the original LDA vector representation (without normalization) with two normalization approaches, the Eigen Factor Radial (EFR) and the Feature Warping (FW) methods, already successfully applied in speaker recognition field, but never compared and evaluated in the context of a speech analytic task. Results show the interest of these normalization techniques for theme identification tasks using automatic transcriptions The EFR normalization approach allows a gain of 3.67 and 3.06 points respectively in comparison to the absence of normalization and to the FW normalization technique.
Mohamed Morchid, Richard Dufour, Driss Matrouf
INTERSPEECH3
2015 Compact Multiview Representation of Documents Based on the Total Variability Space
abstract
Mapping text documents in an LDA-based topic-space is a classical way to extract high-level representation of text documents. Unfortunately, LDA is highly sensitive to hyper-parameters related to the number of classes, or word and topic distribution, and there is no systematic way to pre-estimate optimal configurations. Moreover, various hyper-parameter configurations offer complementary views on the document. In this paper, we propose a method based on a two-step process that, first, expands the representation space by using a set of topic spaces and, second, compacts the representation space by removing poorly relevant dimensions. These two steps are based respectively on multi-view LDA-based representation spaces and factor-analysis models. This model provides a view-independent representation of documents while extracting complementary information from a massive multi-view representation. Experiments are conducted on the DECODA conversation corpus and the Reuters-21578 textual dataset. Results show the efficiency of the proposed multiview compact representation paradigm. The proposed categorization system reaches an accuracy of 86.5% with automatic transcriptions of conversations from DECODA corpus and a Macro-F1 of 80% during a classification task of the well-known Reuters-21578 corpus, with a significant gain compared to the baseline (best single topic space configuration), as well as methods and document representations previously studied.
Mohamed Morchid, Mohamed Bouallegue, Richard Dufour, Georges Linarès, Driss Matrouf, Renato De Mori
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Audio-Based Video Genre Identification
abstract
This paper presents investigations about the automatic identification of video genre by audio channel analysis. Genre refers to editorial styles such commercials, movies, sports... We propose and evaluate some methods based on both low and high level descriptors, in cepstral or time domains, but also by analyzing the global structure of the document and the linguistic contents. Then, the proposed features are combined and their complementarity is evaluated. On a database composed of single-stories web-videos, the best audio-only based system performs 9% of Classification Error Rate (CER). Finally, we evaluate the complementarity of the proposed audio features and video features that are classically used for Video Genre Identification (VGI). Results demonstrate the complementarity of the modalities for genre recognition, the final audio-video system reaching 6% CER.
Mickael Rouvier, Stanislas Oger, Georges Linarès, Driss Matrouf, Bernard Mérialdo, Yingbo Li
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 An I-vector Based Approach to Compact Multi-Granularity Topic Spaces Representation of Textual Documents
abstract
Mohamed Morchid, Mohamed Bouallegue, Richard Dufour, Georges Linarès, Driss Matrouf, Renato de Mori. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Mohamed Morchid, Mohamed Bouallegue, Richard Dufour, Georges Linarès, Driss Matrouf, Renato De Mori
EMNLP5
2014 Subspace Gaussian mixture models for dialogues classification
abstract
International audience
Mohamed Bouallegue, Mohamed Morchid, Richard Dufour, Driss Matrouf, Georges Linarès, Renato De Mori
INTERSPEECH4
2014 Factor analysis based semantic variability compensation for automatic conversation representation
abstract
International audience
Mohamed Bouallegue, Mohamed Morchid, Richard Dufour, Driss Matrouf, Georges Linarès, Renato De Mori
INTERSPEECH4
2014 I-vector based representation of highly imperfect automatic transcriptions
abstract
The performance of Automatic Speech Recognition (ASR) systems drops dramatically when used in noisy environments. Speech analytics suffer from this poor quality of automatic transcriptions. In this paper, we seek to identify themes from dialogues of telephone conversation services using multiple topic-spaces estimated with a Latent Dirichlet Allocation (LDA) approach. This technique consists in estimating several topic models that offer different views of the document. Unfortunately, such a multi-model approach also introduces additional vari-abilities due to the model diversity. We propose to extract the useful information from the full model-set by using an i-vector based approach, previously developed in the context of speaker recognition. Experiments are conducted on the DECODA corpus , that contains records from the call center of the Paris Transportation Company. Results show the effectiveness of the proposed representation paradigm, our identification system reaching an accuracy of 84.7%, with a gain of 3.3 points compared to the baseline.
Mohamed Morchid, Mohamed Bouallegue, Richard Dufour, Georges Linarès, Driss Matrouf, Renato De Mori
INTERSPEECH5
2013 Identify the Benefits of the Different Steps in an i-Vector Based Speaker Verification System
Pierre-Michel Bousquet, Jean-François Bonastre, Driss Matrouf
CIARP (2)3
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH22
2012 I-vectors in the context of phonetically-constrained short utterances for speaker verification
abstract
Short speech duration remains a critical factor of performance degradation when deploying a speaker verification system. To overcome this difficulty, a large number of commercial applications impose the use of fixed pass-phrases. In this context, we show that the performance of the popular i-vector approach can be greatly improved by taking advantage of the phonetic information that they convey. Moreover, as i-vectors require a conditioning process to reach high accuracy, we show that further improvements are possible by taking advantage of this phonetic information within the normalisation process. We compare two methods, Within Class Covariance Normalization (WCCN) and Eigen Factor Radial (EFR), both relying on parameters estimated on the same development data. Our study suggests that WCCN is more robust to data mismatch but less efficient than EFR when the development data has a better match with the test data.
Anthony Larcher, Pierre-Michel Bousquet, Kong-Aik Lee, Driss Matrouf, Haizhou Li 0001, Jean-François Bonastre
ICASSP4
2012 Subspace Gaussian Mixture Models Based on Noise Compensation for Speech Recognition
Mohamed Bouallegue, Driss Matrouf, Georges Linarès, Mickael Rouvier
INTERSPEECH2
2012 Study of the Effect of I-vector Modeling on Short and Mismatch Utterance Duration for Speaker Verification
abstract
International audience
Achintya Kumar Sarkar, Driss Matrouf, Pierre-Michel Bousquet, Jean-François Bonastre
INTERSPEECH2
2012 Acoustic modeling for under-resourced languages based on vectorial HMM-states representation using Subspace Gaussian Mixture Models
abstract
This paper explores a novel method for context-dependent models in automatic speech recognition (ASR), in the context of under-resourced languages. We present a simple way to realize a tying states approach, based on a new vectorial representation of the HMM states. This vectorial representation is considered as a vector of a low number of parameters obtained by the Subspace Gaussian Mixture Models paradigm (SGMM). The proposed method does not require phonetic knowledge or a large amount of data, which represent the major problems of acoustic modeling for under-resourced languages. This paper shows how this representation can be obtained and used for tying states. Our experiments, applied on Vietnamese, show that this approach achieves a stable gain compared to the classical approach which is based on decision trees. Furthermore, this method appears to be portable to other languages, as shown in the preliminary study conducted on Berber.
Mohamed Bouallegue, Emmanuel Ferreira, Driss Matrouf, Georges Linarès, Maria Goudi, Pascal Nocera
SLT3
2011 Subspace Gaussian Mixture Models for vectorial HMM-states representation
abstract
In this paper we present a vectorial representation of the HMM states that is inspired by the Subspace Gaussian Mixture Models paradigm (SGMM). This vectorial representation of states will make possible a large number of applications, such as HMM-states clustering and graphical visualization. Thanks to this representation, the Hidden Markov Model (HMM) states can be seen as sets of points in multi-dimensional space and then can be studied using statistical data analysis techniques. In this paper, we show how this representation can be obtained and used for tying states of an HHM-based automatic speech recognition system without any use of linguistic or phonetic knowledge. In experiments, this approach achieves significant and stable gain, while conserving the classical approach based on decision trees. We also show how it can be used for graphical visualization, which can be useful in other domains like phonetics or clinical phonetics.
Mohamed Bouallegue, Driss Matrouf, Mickael Rouvier, Georges Linarès
ASRU2
2011 Factor analysis based session variability compensation for Automatic Speech Recognition
abstract
In this paper we propose a new feature normalization based on Factor Analysis (FA) for the problem of acoustic variability in Automatic Speech Recognition (ASR). The FA paradigm was previously used in the field of ASR, in order to model the usefull information: the HMM state dependent acoustic information. In this paper, we propose to use the FA paradigm to model the useless information (speaker- or channel-variability) in order to remove it from acoustic data frames. The transformed training data frames are then used to train new HMM models using the standard training algorithm. The transformation is also applied to the test data before the decoding process. With this approach we obtain, on french broadcast news, an absolute WER reduction of 1.3%.
Mickael Rouvier, Mohamed Bouallegue, Driss Matrouf, Georges Linarès
ASRU3
2011 Discriminant binary data representation for speaker recognition
abstract
In supervector UBM/GMM paradigm, each acoustic file is represented by the mean parameters of a GMM model. This supervector space is used as a data representation space, which has a high dimensionality. Moreover, this space is not intrinsically discriminant and a complete speech segment is represented by only one vector, withdrawing mainly the possibility to take into account temporal or sequential information. This work proposes a new approach where each acoustic frame is represented in a discriminant binary space. The proposed approach relies on a UBM to structure the acoustic space in regions. Each region is then populated with a set of Gaussian models, denoted as “specificities”, able to emphasize speaker specific information. Each acoustic frame is mapped in the discriminant binary space, turning “on” or “off” all the specificities to create a large binary vector. All the following steps, speaker reference extraction, likelihood estimation or decision take place in this binary space. Even if this work is a first step in this avenue, the experiments based on NIST SRE 2008 framework demonstrate the potential of the proposed approach. Moreover, this approach opens the opportunity to rethink all the classical processes using a discrete, binary view.
Jean-François Bonastre, Pierre-Michel Bousquet, Driss Matrouf, Xavier Anguera Miró
ICASSP3
2011 A simplified Subspace Gaussian Mixture to compact acoustic models for speech recognition
abstract
Speech recognition applications are known to require a significant amount of resources (memory, computing power). However, embedded speech recognition systems, such as in mobile phones, only authorizes few KB of memory and few MIPS. In the context of HMM-based speech recognizers, each HMM-state distribution is modeled independently from to the other and has a large amount of parameters. In spite of using state-tying techniques, the size of the acoustic models stays large and certain redundancy remains between states. In this paper, we investigate the capacity of the Subspace Gaussian Mixture approach to reduce the acoustic models size while keeping good performances. We introduce a simplification concerning state specific Gaussians weights estimation, which is a very complex and time consuming procedure in the original approach. With this approach, we show that the acoustic model size can be reduced by 92% with almost the same performance as the standard acoustic modeling.
Mohamed Bouallegue, Driss Matrouf, Georges Linarès
ICASSP2
2011 Intersession Compensation and Scoring Methods in the i-vectors Space for Speaker Recognition
abstract
International audience
Pierre-Michel Bousquet, Driss Matrouf, Jean-François Bonastre
INTERSPEECH2
2011 Modeling nuisance variabilities with factor analysis for GMM-based audio pattern classification
Driss Matrouf, Florian Verdet, Mickael Rouvier, Jean-François Bonastre, Georges Linarès
Comput. Speech Lang.1
2011 Applying SVMs and weight-based factor analysis to unsupervised adaptation for speaker verification
Mitchell McLaren, Driss Matrouf, Robbie Vogt, Jean-François Bonastre
Comput. Speech Lang.2
2010 On-the-fly video genre classification by combination of audio features
abstract
Video genre identification methods are frequently based on image or motion analysis, which are relatively time-consuming processes. Since such approaches are tractable by batch processing, as-soon-as-possible identification requires faster methods. In this paper, we investigate the use of audio-only methods for on-the-fly video classification. We propose to use several acoustic feature streams and we evaluate various combination schemes at the frame or at the score level. Results are compared to those obtained by humans, according to the listening duration. Although the system based on model combination slightly outperforms the humans on very soon detection. The latter remain significantly more accurate on long sessions.
Mickael Rouvier, Georges Linarès, Driss Matrouf
ICASSP3
2010 Model and Score Adaptation for Biometric Systems: Coping With Device Interoperability and Changing Acquisition Conditions
abstract
The performance of biometric systems can be significantly affected by changes in signal quality. In this paper, two types of changes are considered: change in acquisition environment and in sensing devices. We investigated three solutions: (i) model-level adaptation, (ii) score-level adaptation (normalisation), and (iii) the combination of the two, called “compound” adaptation. In order to cope with the above changing conditions, the model-level adaptation attempts to update the parameters of the expert systems (classifiers). This approach requires the authenticity of the candidate samples used for adaptation be known (corresponding to supervised adaptation), or can be estimated (unsupervised adaptation). In comparison, the score-level adaptation merely involves post processing the expert output, with the objective of rendering the associated decision threshold to be dependent only on the class priors despite the changing acquisition conditions. Since the above adaptation strategies treat the underlying biometric experts/classifiers as a black-box, they can be applied to any unimodal or multimodal biometric system, thus facilitating system-level integration and performance optimisation. Our contributions are: (i) proposal of compound adaptation; (ii) investigation and comparison of two different quality-dependent score normalisation strategies; and, (iii) empirical comparison of the merit of the above three solutions on the BANCA face (video) and speech database.
Norman Poh, Josef Kittler, Sébastien Marcel, Driss Matrouf, Jean-François Bonastre
ICPR4
2010 Decoupling session variability modelling and speaker characterisation
abstract
The Factor Analysis framework demonstrated its high power to model session variability during the past years. However, training the FA parameters implies to have a large amount of training data. When the size of the available database is limited, the number of components of the core statistical model, the UBM, is also limited as the UBM drives the dimension of the FA main matrix. As the size of the UBM gives directly the size of the speaker supervector (concatenation of the GMM mean parameters), it limits also the intrinsic capacity of the recognition system , reducing the performance expectation. This paper aims to withdraw this limitation by breaking the intrinsic link between the FA dimensionality and the UBM dimensionality. The session variability modelling is done on a smaller dimension compared to the UBM, which drives the discriminative power of the system. The first experimental results proposed in this paper, done using the NIST-SRE 2008 framework, are encouraging with a relative EER improvement of about 18% when a 512 components UBM is associated to a 32 components session variability modelling compared with a 32 components UBM associated with the same variability modelling.
Anthony Larcher, Christophe Lévy, Driss Matrouf, Jean-François Bonastre
INTERSPEECH3
2010 Topological representation of speech for speaker recognition
abstract
International audience
Gabriel Hernández Sierra, Jean-François Bonastre, Driss Matrouf, José Ramón Calvo de Lara
INTERSPEECH3
2010 Channel detectors for system fusion in the context of NIST LRE 2009
abstract
One of the difficulties in Language Recognition is the variability of the speech signal due to speakers and channels. If channel mismatch is too big and when different categories of channels can be identified, one possibility is to build a separate language recognition system for each category and then to fuse them together. This article uses a system selector that takes, for each utterance, the scores of one of the channel-category dependent systems. This selection is guided by a channel detector. We analyze different ways to design such channel detectors: based on cepstral features or on the Factor Analysis channel variability term. The systems are evaluated in the context of NIST’s LRE 2009 and run at 1:65% minCavg for a subset of 8 languages and at 3:85% minCavg for the 23 language setup. Index Terms: language recognition, channel, channel category, fusion, factor analysis, channel detector.
Florian Verdet, Driss Matrouf, Jean-François Bonastre, Jean Hennebert
INTERSPEECH2
2009 Feature Selection Based on Information Theory for Speaker Verification
Rafael Fernández, Jean-François Bonastre, Driss Matrouf, José Ramón Calvo de Lara
CIARP3
2009 Robust audio-based classification of video genre
abstract
International audience
Mickael Rouvier, Georges Linarès, Driss Matrouf
INTERSPEECH3
2009 Factor analysis for audio-based video genre classification
abstract
Statistical classifiers operate on features that generally include both useful and useless information. These two types of information are difficult to separate in the feature domain. Recently, a new paradigm based on a Latent Factor Analysis (LFA) proposed a model decomposition into usefull and useless components. This method was successfully applied to speaker and language recognition tasks. In this paper, we study the use of LFA for video genre classification by using only the audio channel. We propose a classification method based on short-term cep-stral features and Gaussian Mixture Models (GMM) or Support Vector Machine (SVM) classifiers, that are combined with Factor Analysis (FA). Experiments are conducted on a corpus composed of 5 types of video (musics, commercials, cartoons, movies and news). The relative classification error reduction obtained by using the best factor analysis configuration with respect to the baseline system, Gaussian Mixture Model Universal Background Model (GMM-UBM), is about 56%, corresponding to a correct identification rate of about 90%.
Mickael Rouvier, Driss Matrouf, Georges Linarès
INTERSPEECH2
2009 Factor analysis and SVM for language recognition
abstract
International audience
Florian Verdet, Driss Matrouf, Jean-François Bonastre, Jean Hennebert
INTERSPEECH2
2008 Frame-based acoustic feature integration for speech understanding
abstract
With the purpose of improving spoken language understanding (SLU) performance, a combination of different acoustic speech recognition (ASR) systems is proposed. State a posteriori probabilities obtained with systems using different acoustic feature sets are combined with log-linear interpolation. In order to perform a coherent combination of these probabilities, acoustic models must have the same topology (i.e. same set of states). For this purpose, a fast and efficient twin model training protocol is proposed. By a wise choice of acoustic feature sets and log-linear interpolation of their likelihood ratios, a substantial concept error rate (CER) reduction has been observed on the test part of the French MEDIA corpus.
Loïc Barrault, Christophe Servan, Driss Matrouf, Georges Linarès, Renato De Mori
ICASSP3
2008 Factor analysis multi-session training constraint in session compensation for speaker verification
abstract
For a few years now, the problem of session variability in text-independent automatic speaker verification is being tackled actively. A new paradigm based on a Latent Factor Analysis (LFA) model has been applied successfully for this task. However, using this approach, a large training corpus with several sessions per speaker is required. This constraint is hard to satisfy in many real applications. In this paper, we try to analyze if the LFA paradigm still holds even when the constraint of multiple sessions per speaker isn’t satisfied. We propose to study two approaches. The first one consists in using the basic paradigm of the LFA model and the second one is founded on a new interpretation of the interaction between the session and the speaker. The experiments were carried out with NIST SRE 2005 and 2006 protocols. We show that even with only one session per speaker the gain obtained by LFA session compensation (with the two strategies) is still very important. 1
Driss Matrouf, Jean-François Bonastre, Salah Eddine Mezaache
INTERSPEECH1
2008 Combining continuous progressive model adaptation and factor analysis for speaker verification
abstract
International audience
Mitchell McLaren, Driss Matrouf, Robbie Vogt, Jean-François Bonastre
INTERSPEECH2
2008 Analysis of impostor tests with high scores in NIST-SRE context
abstract
International audience
Salah Eddine Mezaache, Jean-François Bonastre, Driss Matrouf
INTERSPEECH3
2007 Artificial impostor voice transformation effects on false acceptance rates
abstract
This paper investigates the effect of a transfer function-based voice transformation on automatic speaker recognition system performance. We focus on increasing the impostor acceptance rate, by modifying the voice of an impostor in order to target a specific speaker. This paper follows previous works where we demonstrate that, if someone has a knowledge on the speaker recognition method used, it is possible to impersonate a given speaker, in the view of this speaker recognition method. In this paper we extend the previous work by relaxing the needed knowledge on the targeted speaker recognition system. The results show that the voice transformation allows a drastic increase of the false acceptance rate, without damaging the natural perception of the voice, and without needing a large knowledge on the targeted speaker recognition system.
Jean-François Bonastre, Driss Matrouf, Corinne Fredouille
INTERSPEECH2
2007 A straightforward and efficient implementation of the factor analysis model for speaker verification
abstract
For a few years, the problem of session variability in textindependent automatic speaker verification is being tackled actively. A new paradigm based on a factor analysis model have successfully been applied for this task. While very efficient, its implementation is demanding. In this paper, the algorithms involved in the eigenchannel MAP model are written down for a straightforward implementation, without referring to previous work or complex mathematics. In addition, a different compensation scheme is proposed where the standard GMM likelihood can be used without any modification to obtain good performance (even without the need of score normalization). The use of the compensated supervectors within a SVM classifier through a distance based kernel is also investigated. Experiments results shows an overall 50 % relative gain over the standard
Driss Matrouf, Nicolas Scheffer, Benoit G. B. Fauve, Jean-François Bonastre
INTERSPEECH1
2007 Confidence measure based unsupervised target model adaptation for speaker verification
abstract
International audience
Alexandre Preti, Jean-François Bonastre, Driss Matrouf, François Capman, Bertrand Ravera
INTERSPEECH3
2007 State-of-the-Art Performance in Text-Independent Speaker Verification Through Open-Source Software
abstract
This paper illustrates an evolution in state-of-the-art speaker verification by highlighting the contribution from newly developed techniques. Starting from a baseline system based on Gaussian mixture models that reached state-of-the-art performances during the NIST'04 SRE, final systems with new intersession compensation techniques show a relative gain of around 50%. This work highlights that a key element in recent improvements is still the classical maximum a posteriori (MAP) adaptation, while the latest compensation methods have a crucial impact on overall performances. Nuisance attribute projection (NAP) and factor analysis (FA) are examined and shown to provide significant improvements. For FA, a new symmetrical scoring (SFA) approach is proposed. We also show further improvement with an original combination between a support vector machine and SFA. This work is undertaken through the open-source ALIZE toolkit.
Benoit G. B. Fauve, Driss Matrouf, Nicolas Scheffer, Jean-François Bonastre, John S. D. Mason
IEEE Trans. Speech Audio Process.2
2006 Characterizing Feature Variability in Automatic Speech Recognition Systems
abstract
A method is described for predicting acoustic feature variability by analyzing the consensus and relative entropy of phoneme posterior probability distributions obtained with different acoustic models having the same type of observations. Variability prediction is used for diagnosis of automatic speech recognition (ASR) systems. When errors are likely to occur, different feature sets are considered for correcting recognition results. Experimental results are provided on the CH1 Italian portion of AURORA3
Loïc Barrault, Driss Matrouf, Renato De Mori, Roberto Gemello, Franco Mana
ICASSP (5)2
2006 Effect of Speech Transformation on Impostor Acceptance
abstract
This paper investigates the effect of voice transformation on automatic speaker recognition system performance. We focus on increasing the impostor acceptance rate, by modifying the voice of an impostor in order to target a specific speaker. This paper is based on the following idea: in several applications and particularly in forensic situations, it is reasonable to think that some organizations have a knowledge on the speaker recognition method used and could impersonate a given, well known speaker. This paper presents some experiments based on NIST SRE 2005 protocol and a simple impostor voice transformation method. The results show that this simple voice transformation allows a drastic increase of the false acceptance rate, without a degradation of the natural aspect of the voice
Driss Matrouf, Jean-François Bonastre, Corinne Fredouille
ICASSP (1)1
2005 Variability of automatic speech recognition systems using different features
abstract
International audience
Loïc Barrault, Renato De Mori, Roberto Gemello, Franco Mana, Driss Matrouf
INTERSPEECH5
2003 Structural speaker adaptation using maximum a posteriori approach and a Gaussian distributions merging technique
abstract
The aim of speaker adaptation techniques is to enhance speaker-independent acoustic models to bring their recognition accuracy as close as possible to the one obtained with speaker-dependent models. Recently, a technique based on a hierarchical structure and the maximum a posteriori criterion was proposed (SMAP) (Shinoda, K. and Lee, C.-H., Proc IEEE ICASSP, 1998). As in SMAP, we assume that the acoustic model parameters are organized in a tree containing all the Gaussian distributions. Each node in that tree represents a cluster of Gaussian distributions sharing a common affine transformation representing the mismatch between training and test conditions. To estimate this affine transformation, we propose a new technique based on merging Gaussians and the standard MAP adaptation. This new technique is very fast and allows a good unsupervised adaptation for both means and variances even with a small amount of adaptation data. This adaptation strategy has shown a significant performance improvement in a large vocabulary speech recognition task, alone and combined with the MLLR (maximum likelihood linear regression) adaptation.
Olivier Bellot, Driss Matrouf, Pascal Nocera, Georges Linarès, Jean-François Bonastre
ICASSP (2)2
2003 Structural linear model-space transformations for speaker adaptation
abstract
Within the framework of speaker-adaptation, a technique based on tree structure and the maximum a posteriori criterion was proposed (SMAP). In SMAP, the parameters estimation, at each node in the tree is based on the assumption that the mismatch between the training and adaptation data is a Gaussian PDF which parameters are estimated by using the Maximum Likelihood criterion. To avoid poor transformation parameters estimation accuracy due to an insufcienc y of adaptation data in a node, we propose a new technique based on the maximum a posteriori approach and PDF Gaussians Merging. The basic idea behind this new technique is to estimate an afne transformations which bring the training acoustic models as close as possible to the test acoustic models rather than transformation maximizing the likelihood of the adaptation data. In this manner, even with very small amount of adaptation data, the parameters transformations are accurately estimated for means and variances. This adaptation strategy has shown a signicant performance improvement in a large vocabulary speech recognition task, alone and combined with the MLLR adaptation.
Driss Matrouf, Olivier Bellot, Pascal Nocera, Georges Linarès, Jean-François Bonastre
INTERSPEECH1
2001 A posteriori and a priori transformations for speaker adaptation in large vocabulary speech recognition systems
abstract
International audience
Driss Matrouf, Olivier Bellot, Pascal Nocera, Georges Linarès, Jean-François Bonastre
INTERSPEECH1
2000 Additive and convolutional noises compensation for speaker recognition
Olivier Bellot, Driss Matrouf, Téva Merlin, Jean-François Bonastre
INTERSPEECH2
1999 Using AR HMM state-dependent filtering for speech enhancement
abstract
In this paper we address the problem of enhancing speech which has been degraded by additive noise. As proposed by Ephraim et al. (1989), autoregressive hidden Markov models (AR-HMM) for the clean speech and an autoregressive Gaussian for the noise are used. The filter applied to a given frame of noisy speech is estimated using the noise model and the autoregressive Gaussian having the highest a posteriori probability given the decoded state sequence. The success of this technique is highly dependent on accurate estimation of the best state sequence. A new strategy combining the use of cepstral-based HMMs, autoregressive HMMs, and a model combination technique, is proposed. The intelligibility of the enhanced speech is indirectly assessed via speech recognition, by comparing performance on noisy speech with compensated models to performance on the enhanced speech with clean-speech models. The results on enhanced speech are as good as our best results obtained with noise compensated models.
Driss Matrouf, Jean-Luc Gauvain
ICASSP1
1999 Comparing different model configurations for language identification using a phonotactic approach
abstract
In this paper different model configurations for language identification using a phonotactic approach are explored. Identification experiments were carried out on the 11-language telephone speech corpus OGI-TS, containing calls in French, English, German, Spanish, Japanese, Korean, Mandarin, Tamil, Farsi, Hindi, and Vietnamese. Phone sequences output by one or multiple phone recognizers are rescored with language-dependent phonotactic models approximated by phone bigrams. The parameters of different sets of acoustic phone models were estimated using the 4-language IDEAL corpus. Sets of language-specific phonotactic models were trained using the training portion of the OGITS CORPUS. Error rates are significantly reduced by combining language-dependent and language-independent acoustic decoders, especially for short segments. A 9.9% LID error rate was obtained on the 11-language task using phonotactic models trained on spontaneous speech data. These results show that the phonotactic approach is relative insensitive to an acoustic mismatch between training and test conditions.
Driss Matrouf, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel
EUROSPEECH1
1998 Language identification incorporating lexical information
abstract
In this paper we explore the use of lexical information for language identification (LID). Our reference LID system uses language-dependent acoustic phone models and phone-based bigram language models. For each language, lexical information is introduced by augmenting the phone vocabulary with the N most frequent words in the training data. Combined phone and word bigram models are used to provide linguistic constraints during acoustic decoding. Experiments were carried out on a 4-language telephone speech corpus. Using lexical information achieves a relative error reduction of about 20% on spontaneous and read speech compared to the reference phone-based system. Identification rates of 92%, 96% and 99% are achieved for spontaneous, read and task-specific speech segments respectively, with prior speech detection.
Driss Matrouf, Martine Adda-Decker, Lori Lamel, Jean-Luc Gauvain
ICSLP1
1997 Model compensation for noises in training and test data
abstract
It is well known that the performance of speech recognition systems degrade rapidly as the mismatch between the training and test conditions increases. Approaches to compensate for this mismatch generally assume that the training data is noise-free, and the test data is noisy. In practice, this assumption is seldom correct. We propose an iterative technique to compensate for noise in both the training and test data. The adopted approach compensates the speech model parameters using the noise present in the test data, and compensates the test data frames using the noise present in the training data. The training and test data are assumed to come from different and unknown microphones and acoustic environments. The interest of such a compensation scheme has been assessed on the MASK task using a continuous density HMM-based speech recognizer. Experimental results show the advantage of compensating for both test and training noise.
Driss Matrouf, Jean-Luc Gauvain
ICASSP1
1996 Developments in continuous speech dictation using the 1995 ARPA NAB news task
abstract
We report on the LIMSI recognizer evaluated in the ARPA 1995 North American Business (NAB) news benchmark test. In contrast to previous evaluations, the new Hub 3 test aims at improving basic SI, CSR performance on unlimited-vocabulary read speech recorded under more varied acoustical conditions (background environmental noise and unknown microphones). The LIMSI recognizer is an HMM-based system with a Gaussian mixture. Decoding is carried out in multiple forward acoustic passes, where more refined acoustic and language models are used in successive passes and information is transmitted via word graphs. In order to deal with the varied acoustic conditions, channel compensation is performed iteratively, refining the noise estimates before the first three decoding passes. The final decoding pass is carried out with speaker-adapted models obtained via unsupervised adaptation using the MLLR method. On the Sennheiser microphone (average SNR 29 dB) a word error of 9.1% was obtained, which can be compared to 17.5% on the secondary microphone data (average SNR 15 dB) using the same recognition system.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Driss Matrouf
ICASSP4