EDBT 2026 Demo / reviewers in the wild / expert
Corinne Fredouille
dblp:42/2411
· DBLP profile ↗
57ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-0413-8950ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 39 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speech Reduction in French: The Relationship Between Vowel Space and Articulation DynamicsabstractInternational audience Kübra Bodur, Corinne Fredouille, Christine Meunier |
INTERSPEECH | 2 |
| 2025 | Exploring the nuances of reduction in conversational speech: lexicalized and non-lexicalized reductionsabstractIn spoken language, a significant proportion of words are produced with missing or underspecified segments, a phenomenon known as reduction. In this study, we distinguish two types of reductions in spontaneous speech: lexicalized reductions, which are well-documented, regularly occurring forms driven primarily by lexical processes, and non-lexicalized reductions, which occur irregularly and lack consistent patterns or representations. The latter are inherently more difficult to detect, and existing methods struggle to capture their full range. We introduce a novel bottom-up approach for detecting potential reductions in French conversational speech, complemented by a top-down method focused on detecting previously known reduced forms. Our bottom-up method targets sequences consisting of at least six phonemes produced within a 230ms window, identifying temporally condensed segments, indicative of reduction. Our findings reveal significant variability in reduction patterns across the corpus. Lexicalized reductions displayed relatively stable and consistent ratios, whereas non-lexicalized reductions varied substantially and were strongly influenced by speaker characteristics. Notably, gender had a significant effect on non-lexicalized reductions, with male speakers showing higher reduction ratios, while no such effect was observed for lexicalized reductions. The two reduction types were influenced differently by speaking time and articulation rate. A positive correlation between lexicalized and non-lexicalized reduction ratios suggested speaker-specific tendencies. Non-lexicalized reductions showed a higher prevalence of certain phonemes and word categories, whereas lexicalized reductions were more closely linked to morpho-syntactic roles. In a focused investigation of selected lexicalized items, we found that “tu sais” was more frequently reduced when functioning as a discourse marker than when used as a pronoun + verb construction. These results support the interpretation that lexicalized reductions are integrated into the mental lexicon, while non-lexicalized reductions are more context-dependent, further supporting the distinction between the two types of reductions. Kübra Bodur, Corinne Fredouille, Stéphane Rauzy, Christine Meunier |
Speech Commun. | 2 |
| 2024 | Interpretable Assessment of Speech Intelligibility Using Deep Learning: A Case Study on Speech Disorders Due to Head and Neck CancersabstractThis paper sheds light on a relatively unexplored area which is deep learning interpretability for speech disorder assessment and characterization. Building upon a state-of-the-art methodology for the explainability and interpretability of hidden representation inside a deep-learning speech model, we provide a deeper understanding and interpretation of the final intelligibility assessment of patients experiencing speech disorders due to Head and Neck Cancers (HNC). Promising results have been obtained regarding the prediction of speech intelligibility and severity of HNC patients while giving relevant interpretations of the final assessment both at the phonemes and phonetic feature levels. The potential of this approach becomes evident as clinicians can acquire more valuable insights for speech therapy. Indeed, this can help identify the specific linguistic units that affect intelligibility from an acoustic point of view and enable the development of tailored rehabilitation protocols to improve the patient’s ability to communicate effectively, and thus, the patient’s quality of life. Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Mathieu Balaguer, Virginie Woisard |
LREC/COLING | 2 |
| 2024 | Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce ContextabstractAutomatic speech quality assessment has raised more attention as an alternative or support to traditional perceptual clinical evaluation. However, most research so far only gains good results on simple tasks such as binary classification, largely due to data scarcity. To deal with this challenge, current works tend to segment patients’ audio files into many samples to augment the datasets. Nevertheless, this approach has limitations, as it indirectly relates overall audio scores to individual segments. This paper introduces a novel approach where the system learns at the audio level instead of segments despite data scarcity. This paper proposes to use the pre-trained Wav2Vec2 architecture for both SSL, and ASR as feature extractor in speech assessment. Carried out on the HNC dataset, our ASR-driven approach established a new baseline compared with other approaches, obtaining average MSE = 0.73 and MSE = 1.15 for the prediction of intelligibility and severity scores respectively, using only 95 training samples. It shows that the ASR based Wav2Vec2 model brings the best results and may indicate a strong correlation between ASR and speech quality assessment. We also measure its ability on variable segment durations and speech content, exploring factors influencing its decision. Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard |
LREC/COLING | 2 |
| 2024 | Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based modelsabstractHead and Neck Cancers (HNC) significantly impact patients' ability to speak, affecting their quality of life. Commonly used metrics for assessing pathological speech are subjective, prompting the need for automated and unbiased evaluation methods. This study proposes a self-supervised Wav2Vec2-based model for phone classification with HNC patients, to enhance accuracy and improve the discrimination of phonetic features for subsequent interpretability purpose. The impact of pre-training datasets, model size, and fine-tuning datasets and parameters are explored. Evaluation on diverse corpora reveals the effectiveness of the Wav2Vec2 architecture, outperforming a CNN-based approach, used in previous work. Correlation with perceptual measures also affirms the model relevance for impaired speech analysis. This work paves the way for better understanding of pathological speech with interpretable approaches for clinicians, by leveraging complex self-learnt speech representations. Malo Maisonneuve, Corinne Fredouille, Muriel Lalain, Alain Ghio, Virginie Woisard |
INTERSPEECH | 2 |
| 2024 | Exploring ASR-Based WAV2VEC2 for Automated Speech Disorder Assessment: Insights and AnalysisabstractWith the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pretrained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability. Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard |
SLT | 2 |
| 2023 | Speech reduction: position within French prosodic structureabstractInternational audience Kübra Bodur, Roxane Bertrand, James Sneed German, Stéphane Rauzy, Corinne Fredouille, Christine Meunier |
INTERSPEECH | 5 |
| 2023 | Interpreting Deep Representations of Phonetic Features via Neuro-Based Concept Detector: Application to Speech Disorders Due to Head and Neck CancerabstractThe popularity of Deep Neural Networks (DNNs) is growing significantly, and so is the interest in gaining a better understanding of their functioning. In this work, it is even more interesting to reveal the behavior of these black-boxes since we are involved in a clinical context. To this end, we propose a general analytic framework, namedNeuro-based Concept Detector (NCD), for interpreting deep representations of a DNN. Based on the activation patterns of the hidden neurons, this framework highlights the capacity of neurons to detect a specific concept related to the final task. The key strength of our framework is that it provides an interpretability tool for any type of DNN performing a classification task regardless of the application field. In this paper, we evaluate this framework on a Convolutional Neural Network (CNN) trained for the task of French phone classification. This choice was guided by the final objective of a long-term research project, which aims to identify the linguistic units best contributing to the maintenance or loss of intelligibility in the context of speech disorders. ThroughNCD, we demonstrate the emergence of phonetic features in the classification layers of the CNN-based model, while applied on healthy speech, a concept with a great interest in the field of clinical phonetics. Indeed, we further show that these interesting findings shed light on the characteristics of speech disorders in terms of altered phonetic features and provide relevant information for clinical practice, notably, patients' rehabilitation and follow-up. Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders Step 2: Contribution of the Emergence of Phonetic TraitsabstractApart from the impressive performance it has achieved in several tasks, one of the most important factors remaining for the continuous progress of deep learning is the increased work related to interpretability, especially in a medical context. In a recent work, we presented competitive performance achieved with a CNN-based model trained on normal speech for the French phone classification and how it correlates well with different perceptual measures when exposed to disordered speech. This paper extends that work by focusing on interpretability. Here, the goal is to get insights into the way in which neural representations shape the final task of phone classification so that it can be used further to explain the loss of intelligibility in disordered speech. In this way, an original framework is proposed, relying firstly on the neural activity and a novel representation per neuron, here considering the phone classification, and, secondly, permitting to identify a set of neurons devoted to the detection of specific phonetic traits on normal speech. Faced to disordered speech, a degradation of that set of neurons is observed, demonstrating a loss of specific phonetic traits in some patients involved, and the potentiality of the proposed approaches to inform about speech alteration. Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
ICASSP | 2 |
| 2022 | Validation of the Neuro-Concept Detector framework for the characterization of speech disorders: A comparative study including Dysarthria and DysphoniaabstractInternational audience Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
INTERSPEECH | 2 |
| 2020 | Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders - Step 1: CNN Model-Based Phone ClassificationabstractInternational audience Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
INTERSPEECH | 2 |
| 2020 | How to Compare Automatically Two Phonological Strings: Application to Intelligibility Measurement in the Case of Atypical SpeechabstractAtypical speech productions, regardless of their origins (accents, learning, pathology), need to be assessed with regard to “typical” or “expected” productions. Evaluation is necessarily based on comparisons between linguistic forms produced and linguistic forms expected. In the field of speech disorders, the intelligibility of a patient is evaluated in order to measure the functional impact of his/her pathology on his/her oral communication. The usual method is to transcribe orthographic linguistic forms perceived and to assign a global and imprecise rating based on their correctness or incorrect. To obtain a more precise evaluation of the production deviations, we propose a measurement method based on phonological transcriptions. An algorithm computes automatically and finely the distances between the phonological forms produced and expected from cost matrices based on the differences of features between phonemes. A first test of this method among a large population of healthy speakers and patients treated for cancer of the oral and pharyngeal cavities has proved its validity. Alain Ghio, Muriel Lalain, Laurence Giusti, Corinne Fredouille, Virginie Woisard |
LREC | 4 |
| 2020 | Have a Cake and Eat it Too: Assessing Discriminating Performance of an Intelligibility Index Obtained from a Reduced Sample SizeabstractThis paper investigates random vs. phonetically motivated reduction of linguistic material used in an intelligibility task in speech disordered populations and the subsequent impact on the discrimination classifier quantified by the area under the receiver operating characteristics curve (AUC of ROC). The comparison of obtained accuracy indexes shows that when the sample size is reduced based on a phonetic criterium—here, related to phonotactic complexity—, the classifier has a higher ranking ability than when the linguistic material is arbitrarily reduced. Crucially, downsizing the linguistic sample to about 30% of the original dataset does not diminish the discriminatory performance of the classifier. This result is of significant interest to both clinicians and patients as it validates a tool that is both reliable and efficient. Anna K. Marczyk, Alain Ghio, Muriel Lalain, Marie Rebourg, Corinne Fredouille, Virginie Woisard |
LREC | 5 |
| 2018 | Automatic Evaluation of Speech Intelligibility Based on I-vectors in the Context of Head and Neck CancersabstractInternational audience Imed Laaridh, Corinne Fredouille, Alain Ghio, Muriel Lalain, Virginie Woisard |
INTERSPEECH | 2 |
| 2018 | Carcinologic Speech Severity Index Project: A Database of Speech Disorder Productions to Assess Quality of Life Related to Speech After Cancer
Corine Astésano, Mathieu Balaguer, Jérôme Farinas, Corinne Fredouille, Pascal Gaillard, Alain Ghio, Imed Laaridh, Muriel Lalain, Benoît Lepage, Julie Mauclair, Olivier Nocaudie, Julien Pinquier, Oriol Pont, Gilles Pouchoulin, Michèle Puech, Danièle Robert, Etienne Sicard, Virginie Woisard |
LREC | 4 |
| 2018 | Dysarthric speech evaluation: automatic and perceptual approaches
Imed Laaridh, Christine Meunier, Corinne Fredouille |
LREC | 3 |
| 2018 | Perceptual evaluation for automatic anomaly detection in disordered speech: Focus on ambiguous cases
Imed Laaridh, Christine Meunier, Corinne Fredouille |
Speech Commun. | 3 |
| 2017 | Automatic Prediction of Speech Evaluation Metrics for Dysarthric SpeechabstractInternational audience Imed Laaridh, Waad Ben Kheder, Corinne Fredouille, Christine Meunier |
INTERSPEECH | 3 |
| 2016 | Evaluation of a Phone-Based Anomaly Detection Approach for Dysarthric SpeechabstractInternational audience Imed Laaridh, Corinne Fredouille, Christine Meunier |
INTERSPEECH | 2 |
| 2016 | Automatic Anomaly Detection for Dysarthria across Two Speech Styles: Read vs Spontaneous Speech
Imed Laaridh, Corinne Fredouille, Christine Meunier |
LREC | 2 |
| 2016 | The TYPALOC Corpus: A Collection of Various Dysarthric Speech Recordings in Read and Spontaneous Styles
Christine Meunier, Cécile Fougeron, Corinne Fredouille, Brigitte Bigi, Lise Crevier-Buchman, Elisabeth Delais-Roussarie, Laurianne Georgeton, Alain Ghio, Imed Laaridh, Thierry Legou, Claire Pillot-Loiseau, Gilles Pouchoulin |
LREC | 3 |
| 2015 | Novel clustering selection criterion for fast binary key speaker diarizationabstractSpeaker diarization has become an important building block in many speech-related systems. Given the great increase of audiovisual media, fast systems are required in order to process large amounts of data in a reasonable time. In this regard, the recently proposed speaker diarization system based on binary key speaker modeling provides a very fast alternative to state-of-the-art systems at the cost of a slight decrease in performance. This decrease is mainly due to drawbacks in the final clustering selection algorithm, which is far from returning the optimum clustering the system is actually able to generate. At the same time, we have identified potential points of our system which can be further sped up. This paper aims to face these two issues by first lightening the processing at the main identified bottleneck, and second by proposing an alternative clustering selection technique capable of providing near-optimum clustering outputs. Experimental results on the REPERE test database validate the effectiveness of the proposed improvements, obtaining a relative performance gain of 20% and execution times of 0.037 xRT (being xRT the Real-Time factor). Héctor Delgado, Xavier Anguera Miró, Corinne Fredouille, Javier Serrano 0001 |
INTERSPEECH | 3 |
| 2015 | Fast Single- and Cross-Show Speaker Diarization Using Binary Key Speaker ModelingabstractInternational audience Héctor Delgado, Xavier Anguera Miró, Corinne Fredouille, Javier Serrano 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Multimodal understanding for person recognition in video broadcastsabstractInternational audience Frédéric Béchet, Meriem Bendris, Delphine Charlet, Géraldine Damnati, Benoît Favre, Mickael Rouvier, Rémi Auguste, Benjamin Bigot, Richard Dufour, Corinne Fredouille, Georges Linarès, Jean Martinet, Grégory Senay, Pierre Tirilly |
INTERSPEECH | 10 |
| 2014 | Towards a complete binary key system for the speaker diarization taskabstractInternational audience Héctor Delgado, Corinne Fredouille, Javier Serrano 0001 |
INTERSPEECH | 2 |
| 2014 | Analysis of i-vector framework for speaker identification in TV-showsabstractInspired from the Joint Factor Analysis, the I-vector-based analysis has become the most popular and state-of-the-art framework for the speaker verification task. Mainly applied within the NIST/SRE evaluation campaigns, many studies have been proposed to improve more and more performance of speaker verification systems. Nevertheless, while the i-vector framework has been used in other speech processing fields like language recognition, a very few studies have been reported for the speaker identification task on TV shows. This work was done in the REPERE challenge context, focused on the people recognition task in multimodal conditions (audio, video, text) from TV show corpora. Moreover, the challenge participants are invited for providing systems for monomodal tasks, like speaker identification. The application of the i-vector framework is investi-gatedthrough different points of views: (1) some of the i-vector based approaches are compared, (2) a specific i-vector extraction protocol is proposed in order to deal with widely varying amounts of training data among speaker population, (3) the joint use of both speaker diarization and identification is finally analyzed. Based on a 533 speaker dictionary, this joint system wins the monomodal speaker identification task of the 2014 REPERE challenge. Corinne Fredouille, Delphine Charlet |
INTERSPEECH | 1 |
| 2013 | Person name recognition in ASR outputs using continuous context modelsabstractThe detection and characterization, in audiovisual documents, of speech utterances where person names are pronounced, is an important cue for spoken content analysis. This paper tackles the problematic of retrieving spoken person names in the 1-Best ASR outputs of broadcast TV shows. Our assumption is that a person name is a latent variable produced by the lexical context it appears in. Thereby, a spoken name could be derived from ASR outputs even if it has not been proposed by the speech recognition system. A new context modelling is proposed in order to capture lexical and structural information surrounding a spoken name. The fundamental hypothesis of this study has been validated on broadcast TV documents available in the context of the REPERE challenge. Benjamin Bigot, Grégory Senay, Georges Linarès, Corinne Fredouille, Richard Dufour |
ICASSP | 4 |
| 2013 | Combining acoustic name spotting and continuous context models to improve spoken person name recognition in speechabstractRetrieving pronounced person names in spoken documents is a critical problematic in the context of audiovisual content indexing. In this paper, we present a cascading strategy for two methods dedicated to spoken name recognition in speech. The first method is an acoustic name spotting in phoneme confusion networks. It is based on a phonetic edition distance criterion based on phoneme probabilities held in confusion networks. The second method is a continuous context modelling approach applied on the 1-best transcription output. It relies on a probabilistic modelling of name-to-context dependencies. We assume that the combination of these methods, based on different types of information, may improve spoken name recognition performance. This assumption is studied through experiments done on a set of audiovisual documents from the development set of the REPERE challenge. Results report that combining acoustic and linguistic methods produces an absolute gain of 3% in terms of F-measure compared to the best system taken alone. Benjamin Bigot, Grégory Senay, Georges Linarès, Corinne Fredouille, Richard Dufour |
INTERSPEECH | 4 |
| 2013 | Improving speaker identification in TV-shows using person name detection in overlaid text and speechabstractInternational audience Delphine Charlet, Corinne Fredouille, Géraldine Damnati, Grégory Senay |
INTERSPEECH | 2 |
| 2013 | Person name spotting by combining acoustic matching and LDA topic modelsabstractIn this article, we are interested in spoken term detection task, with a particular focus on Person Name (PN) spotting in automatic speech recognition (ASR) system outputs. We propose a two-step method that combines an acoustic matching based on a Phoneme Confusion Network (PCN) with a semantic rescoring based on the Latent Dirichlet Allocation (LDA) models. The first module allows to find, in the PCN, potential PN candidates in speech segments, while the second is in charge of ranking the competing PN, according to a LDA topic model. The proposed LDA-based approach outperforms significantly the baseline system based on a search in the ASR phoneme lattice, obtaining a F-measure score of 77.04% on PN detection. Grégory Senay, Benjamin Bigot, Richard Dufour, Georges Linarès, Corinne Fredouille |
INTERSPEECH | 5 |
| 2012 | How to manage sound, physiological and clinical data of 2500 dysphonic and dysarthric speakers?
Alain Ghio, Gilles Pouchoulin, Bernard Teston, Serge Pinto, Corinne Fredouille, Céline De Looze, Danièle Robert, François Viallet, Antoine Giovanni |
Speech Commun. | 5 |
| 2012 | A Comparative Study of Bottom-Up and Top-Down Approaches to Speaker DiarizationabstractThis paper presents a theoretical framework to analyze the relative merits of the two most general, dominant approaches to speaker diarization involving bottom-up and top-down hierarchical clustering. We present an original qualitative comparison which argues how the two approaches are likely to exhibit different behavior in speaker inventory optimization and model training: bottom-up approaches will capture comparatively purer models and will thus be more sensitive to nuisance variation such as that related to the speech content; top-down approaches, in contrast, will produce less discriminative speaker models but, importantly, models which are potentially better normalized against nuisance variation. We report experiments conducted on two standard, single-channel NIST RT evaluation datasets which validate our hypotheses. Results show that competitive performance can be achieved with both bottom-up and top-down approaches (average DERs of 21% and 22%), and that neither approach is superior. Speaker purification, which aims to improve speaker discrimination, gives more consistent improvements with the top-down system than with the bottom-up system (average DERs of 19% and 25%), thereby confirming that the top-down system is less discriminative and that the bottom-up system is less stable. Finally, we report a new combination strategy that exploits the merits of the two approaches. Combination delivers an average DER of 17% and confirms the intrinsic complementary of the two approaches. Nicholas W. D. Evans, Simon Bozonnet, Dong Wang 0013, Corinne Fredouille, Raphaël Troncy |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Speaker Diarization: A Review of Recent ResearchabstractSpeaker diarization is the task of determining “who spoke when?” in an audio or video recording that contains an unknown amount of speech and also an unknown number of speakers. Initially, it was proposed as a research topic related to automatic speech recognition, where speaker diarization serves as an upstream processing step. Over recent years, however, speaker diarization has become an important key technology for many tasks, such as navigation, retrieval, or higher level inference on audio data. Accordingly, many important improvements in accuracy and robustness have been reported in journals and conferences in the area. The application domains, from broadcast news, to lectures and meetings, vary greatly and pose different problems, such as having access to multiple microphones and multimodal information or overlapping speech. The most recent review of existing technology dates back to 2006 and focuses on the broadcast news domain. In this paper, we review the current state-of-the-art, focusing on research developed since 2006 that relates predominantly to speaker diarization for conference meetings. Finally, we present an analysis of speaker diarization performance as reported through the NIST Rich Transcription evaluations on meeting data and identify important areas for future research. Xavier Anguera Miró, Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Gerald Friedland, Oriol Vinyals |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Speaker diarization of heterogeneous web video files: A preliminary studyabstractIn the last ten years, internet as well as its applications changed significantly, mainly thanks to the raising of available personal resources. Concerning multimedia, the most impressive evolution is the continuous growing success of the video sharing websites. But with this success come the difficulties to efficiently search, index and access relevant information about these documents. Speaker diarization is an important task in the overall information retrieval process. This paper describes an audio/video database, especially built for the speaker diarization task, based on different video genres. Through some preliminary experiments, it highlights the difficulties encountered in this context, mainly linked to the database heterogeneity. Pierre Clément, Thierry Bazillon, Corinne Fredouille |
ICASSP | 3 |
| 2010 | The lia-eurecom RT'09 speaker diarization system: Enhancements in speaker modelling and cluster purificationabstractThere are two approaches to speaker diarization. They are bottom-up and top-down. Our work on top-down systems show that they can deliver competitive results compared to bottom-up systems and that they are extremely computationally efficient, but also that they are particularly prone to poor model initialisation and cluster impurities. In this paper we present enhancements to our state-of-the-art, top-down approach to speaker diarization that deliver improved stability across three different datasets composed of conference meetings from five standard NIST RT evaluations. We report an improved approach to speaker modelling which, despite having greater chances for cluster impurities, delivers a 35% relative improvement in DER for the MDM condition. We also describe new work to incorporate cluster purification into a top-down system which delivers relative improvements of 44% over the baseline system without compromising computational efficiency. Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille |
ICASSP | 3 |
| 2010 | System output combination for improved speaker diarizationabstractInternational audience Simon Bozonnet, Nicholas W. D. Evans, Xavier Anguera Miró, Oriol Vinyals, Gerald Friedland, Corinne Fredouille |
INTERSPEECH | 6 |
| 2010 | An integrated top-down/bottom-up approach to speaker diarizationabstractInternational audience Simon Bozonnet, Nicholas W. D. Evans, Corinne Fredouille, Dong Wang 0013, Raphaël Troncy |
INTERSPEECH | 3 |
| 2010 | The DesPho-APaDy Project: Developing an Acoustic-phonetic Characterization of Dysarthric Speech in French
Cécile Fougeron, Lise Crevier-Buchman, Corinne Fredouille, Alain Ghio, Christine Meunier, Claude Chevrie-Muller, Jean-François Bonastre, Antonia Colazo-Simon, Céline De Looze, Danielle Duez, Cédric Gendrot, Thierry Legou, Nathalie Lévêque, Claire Pillot-Loiseau, Serge Pinto, Gilles Pouchoulin, Danièle Robert, Jacqueline Vaissière, François Viallet, Coralie Vincent |
LREC | 3 |
| 2009 | Speaker diarization using unsupervised discriminant analysis of inter-channel delay featuresabstractWhen multiple microphones are available estimates of inter-channel delay, which characterise a speaker's location, can be used as features for speaker diarization. Background noise and reverberation can, however, lead to noisy features and poor performance. To ameliorate these problems, this paper presents a new approach to the discriminant analysis of delay features for speaker diarization. This novel and nonetheless unsupervised approach aims to increase speaker separability in delay-space. We assess the approach on subsets of four standard NIST RT datasets and demonstrate a relative improvement in diarization error rate of 25% on a separate evaluation set using delay features alone. Nicholas W. D. Evans, Corinne Fredouille, Jean-François Bonastre |
ICASSP | 2 |
| 2008 | New implementations of the E-HMM-based system for speaker diarization in meeting roomsabstractThis paper addresses the problem of speaker diarization in the specific context of meeting room recordings. Some new enhancements to the E-HMM-based speaker diarization system are reported. These involve a different approach to speaker modelling utilising EM/ML-based training rather than MAP adaptation as in our previous work. Using the new system we investigate the effects of speech activity detection through speaker diarization experiments conducted on 23 meetings extracted from the NIST/RT evaluation campaign datasets. We propose a new approach, which assigns confidence values according to the type of information carried by the signal and incorporates these values directly into the speaker diarization system. Experimental results show that, perhaps surprisingly, the non-speech segments do not systematically affect the robustness of the speaker diarization system, and more precisely the speaker model training process. Corinne Fredouille, Nick Evans |
ICASSP | 1 |
| 2008 | Dysphonic voices and the 0-3000 hz frequency bandabstractConcerned with pathological voice assessment, this paper aims at characterizing dysphonia in the frequency domain for a better understanding of related phenomena while most of the studies have focused only on improving classification systems for diag-nosis help purposes. Based on a first study which demonstrates that the low frequencies ([0-3000]Hz) are more relevant for dys-phonia discrimination compared with higher frequencies, the authors propose in this paper to pursue by analyzing the impact of the restricted frequency band ([0-3000]Hz) on the dysphonic voice discrimination from a phonetical and perceptual point of views. A discussion around the frequency band limitation of telephone channel is also proposed. Index Terms: Voice disorder, dysphonia characterization, au-tomatic dysphonic voice classification, frequency analysis Gilles Pouchoulin, Corinne Fredouille, Jean-François Bonastre, Alain Ghio, Antoine Giovanni |
INTERSPEECH | 2 |
| 2007 | Complementary approaches for voice disorder assessmentabstractThis paper describes two comparative studies of voice quality assessment based on complementary approaches. The first study was undertaken on 449 speakers (including 391 dysphonic patients) whose voice quality was evaluated in parallel by a perceptual judgment and objective measurements on acoustic and aerodynamic data. Results showed that a nonlinear combination of 7 parameters allowed the classification of 82% voice samples in the same grade as the jury. The second study relates to the adaptation of Automatic Speaker Recognition (ASR) techniques to pathological voice assessment. The system designed for this particular task relies on a GMM based approach, which is the state-of-the-art for ASR. Experiments conducted on 80 female voices provide promising results, underlining the interest of such an approach. We benefit from the multiplicity of theses techniques to evaluate the methodological situation which points fundamental differences between these complementary approaches (bottom-up vs. top-down, global vs. analytic). We also discuss some theoretical aspects about relationship between acoustic measurement and perceptual mechanisms which are often forgotten in the performance race. Jean-François Bonastre, Corinne Fredouille, Alain Ghio, Antoine Giovanni, Gilles Pouchoulin, Joana Revis, Bernard Teston |
INTERSPEECH | 2 |
| 2007 | Artificial impostor voice transformation effects on false acceptance ratesabstractThis paper investigates the effect of a transfer function-based voice transformation on automatic speaker recognition system performance. We focus on increasing the impostor acceptance rate, by modifying the voice of an impostor in order to target a specific speaker. This paper follows previous works where we demonstrate that, if someone has a knowledge on the speaker recognition method used, it is possible to impersonate a given speaker, in the view of this speaker recognition method. In this paper we extend the previous work by relaxing the needed knowledge on the targeted speaker recognition system. The results show that the voice transformation allows a drastic increase of the false acceptance rate, without damaging the natural perception of the voice, and without needing a large knowledge on the targeted speaker recognition system. Jean-François Bonastre, Driss Matrouf, Corinne Fredouille |
INTERSPEECH | 3 |
| 2007 | The influence of speech activity detection and overlap on speaker diarization for meeting room recordingsabstractAbstract This paper addresses the problem of speaker diarization inthe specific context of meeting room recordings which of-ten involve a high degree of spontaneous speech with largeoverlapped speech segments, speaker noise (laughs, whispers,coughs, etc.) and very short speaker turns. A large variabilityin signal quality has brought an additional level of complexity.This paper investigates the effects of speech activity detectionand overlapped speech through speaker diarization experimentsconducted on the NIST RT’05 and RT’06 data sets. Resultsindicate that our system is highly sensitive to the shape of theinitial segmentation and that, perhaps surprisingly, perfect ref-erences can even degrade performance. Finally we propose adirection for future research to incorporate confidence valuesaccording to acoustic attributes in order to unify what is cur-rently a somewhat disjointed approach to speaker diarization. Index Terms : speaker diarization, meeting room, speech activ-ity detection, overlapped speech. Corinne Fredouille, Nicholas W. D. Evans |
INTERSPEECH | 1 |
| 2007 | Frequency study for the characterization of the dysphonic voicesabstractConcerned with pathological voice assessment, this paper aims at characterizing dysphonia in the frequency domain for a better understanding of relating phenomena while most of the studies have focused only on improving classification systems for diagnosis help purposes.In this context, a GMM-based automatic classification system is applied on different frequency ranges in order to investigate which ones are relevant for dysphonia characterization.Experiment results demonstrate that the low frequencies [0-3000]Hz are more relevant for dysphonia discrimination compared with higher frequencies. Gilles Pouchoulin, Corinne Fredouille, Jean-François Bonastre, Alain Ghio, Antoine Giovanni |
INTERSPEECH | 2 |
| 2006 | On the Use of Linguistic Information for Broadcast News Speaker TrackingabstractIn this paper, we have explored a speaker characterization at two different linguistic levels, the lexical content and the syntactical form, in the context of Broadcast News (BN) speaker tracking task. The modeling of the information is done classically, by a n-gram approach applied on the word sequence for the linguistic content and on the syntactical tags issued from this sequence for the syntactical level (n-class modeling). The experiments were done on a subset of the BN rich transcription French evaluation campaign, ESTER. We have observed that lexical information is not useful for BN data while the syntactical information seems promising as it allows significant speaker identification and verification performance (until 40% of correct identification rate and 35% of EER). William Antoni, Corinne Fredouille, Jean-François Bonastre |
ICASSP (1) | 2 |
| 2006 | Effect of Speech Transformation on Impostor AcceptanceabstractThis paper investigates the effect of voice transformation on automatic speaker recognition system performance. We focus on increasing the impostor acceptance rate, by modifying the voice of an impostor in order to target a specific speaker. This paper is based on the following idea: in several applications and particularly in forensic situations, it is reasonable to think that some organizations have a knowledge on the speaker recognition method used and could impersonate a given, well known speaker. This paper presents some experiments based on NIST SRE 2005 protocol and a simple impostor voice transformation method. The results show that this simple voice transformation allows a drastic increase of the false acceptance rate, without a degradation of the natural aspect of the voice Driss Matrouf, Jean-François Bonastre, Corinne Fredouille |
ICASSP (1) | 3 |
| 2006 | Step-by-step and integrated approaches in broadcast news speaker diarization
Sylvain Meignier, Daniel Moraru, Corinne Fredouille, Jean-François Bonastre, Laurent Besacier |
Comput. Speech Lang. | 3 |
| 2005 | Application of automatic speaker recognition techniques to pathological voice assessment (dysphonia)abstractHAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Corinne Fredouille, Gilles Pouchoulin, Jean-François Bonastre, M. Azzarello, Antoine Giovanni, Alain Ghio |
INTERSPEECH | 1 |
| 2005 | Broadcast news speaker tracking for ESTER 2005 campaignabstractThis paper presents the speaker tracking system of the LIA laboratory, validated during ESTER 2005 campaign on a radio broadcast news corpus of about 90 h. The LIA speaker tracking system firstly uses an acoustic class segmentation in order to suppress non speech frames and to detect the speech conditions. Secondly, a speaker diarization process is applied in order to provide speaker detection system (the last step) with speaker homogeneous segments (boundaries and clustering). The speaker detection system uses UBM/GMM likelihood ratios in order to decide if a segment belongs to one tracked speaker. The speaker tracking system is presented and some results obtained during ESTER 2005 campaign are proposed. The presented systems are based on the ALIZE platform (Automatic speaker recognition C++ library). Dan Istrate, Nicolas Scheffer, Corinne Fredouille, Jean-François Bonastre |
INTERSPEECH | 3 |
| 2004 | Benefits of prior acoustic segmentation for automatic speaker segmentationabstractThe paper investigates the interest of segmentation in acoustic macro classes (like gender or bandwidth) as front-end processing for the segmentation/diarization task. The impact of this prior acoustic segmentation is evaluated in terms of speaker diarization performance in the particular context of NIST RT'03 evaluation (done on the HUB4 broadcast news corpora). It is rarely discussed in the literature, but our work shows that the application of prior acoustic segmentation, in a similar way to the automatic speech recognition task, may be very useful to the speaker segmentation task. Experiments were conducted using two different kinds of speaker segmentation systems developed individually by the LIA and CLIPS laboratories in the framework of the ELISA consortium. For both systems, improvement was observed when combined with prior acoustic segmentation. However, a larger impact, in terms of performance, is observed on the LIA system based on an ascending/HMM approach compared to the CLIPS system based on speaker turn detection. Sylvain Meignier, Daniel Moraru, Corinne Fredouille, Laurent Besacier, Jean-François Bonastre |
ICASSP (1) | 3 |
| 2004 | The ELISA consortium approaches in broadcast news speaker segmentation during the NIST 2003 rich transcription evaluationabstractThe paper presents the ELISA consortium activities in automatic speaker segmentation, also known as speaker diarization, during the NIST rich transcription (RT), 2003, evaluation. The experiments were conducted on real broadcast news data (HUB4). Two different approaches from the CLIPS and LIA laboratories are presented and different possibilities of combining them are investigated, in the framework of the ELISA consortium. The system submitted as an ELISA primary system obtained the second lowest segmentation error rate compared to the other RT03-participant primary systems. Another ELISA system submitted as a secondary system outperformed the best primary system and obtained the lowest speaker segmentation error rate. Daniel Moraru, Sylvain Meignier, Corinne Fredouille, Laurent Besacier, Jean-François Bonastre |
ICASSP (1) | 3 |
| 2000 | A speaker tracking system based on speaker turn detection for NIST evaluationabstractA speaker tracking system (STS) is built by using successively a speaker change detector and a speaker verification system. The aim of the STS is to find in a conversation between several persons (some of them having already enrolled and other being totally unknown) target speakers chosen in a set of enrolled users. In a first step, speech is segmented into homogeneous segments containing only one speaker, without any use of a priori knowledge about speakers. Then, the resulting segments are checked to belong to one of the target speakers. The system has been used in a NIST evaluation test with satisfactory results. Jean-François Bonastre, Perrine Delacourt, Corinne Fredouille, Téva Merlin, Christian Wellekens |
ICASSP | 3 |
| 2000 | Behavior of a Bayesian adaptation method for incremental enrollment in speaker verificationabstractClassical adaptation approaches are generally used for speaker or environment adaptation of speech recognition systems. In this paper, we use such techniques for the incremental training of client models in a speaker verification system. The initial model is trained on a very limited amount of data and then progressively updated with access data, using a segmental-EM procedure. In supervised mode (i.e. when access utterances are certified), the incremental approach yields equivalent performance to the batch one. We also investigate on the impact of various scenarios of impostor attacks during the incremental enrollment phase. All results are obtained with the Picassoft platform-the state-of-the-art speaker verification system developed in the PICASSO project. Corinne Fredouille, Johnny Mariéthoz, Cédric Jaboulet, Jean Hennebert, Chafic Mokbel, Frédéric Bimbot |
ICASSP | 1 |
| 2000 | Evolutive HMM for multi-speaker tracking systemabstractSeeking within a speech sequence the speaker utterances is one of the main tasks of indexing. In this paper, the proposed speaker tracking system is defined in the case where all speaker identities are known beforehand. The conversation is modeled as an evolutive HMM-like model, in which speaker models computed are added one by one. A temporary indexing is proposed after each speaker adding and then challenged at the next step. This process is iterated until all the speakers are detected. The system has been assessed using multi-speaker messages generated by concatenation of Switchboard mono-speaker segments. The obtained results show the potentiality of the proposed solution. Sylvain Meignier, Jean-François Bonastre, Corinne Fredouille, Téva Merlin |
ICASSP | 3 |
| 2000 | Localization and selection of speaker-specific information with statistical modeling
Laurent Besacier, Jean-François Bonastre, Corinne Fredouille |
Speech Commun. | 3 |
| 1999 | Similarity normalization method based on world model and a posteriori probability for speaker verification
Corinne Fredouille, Jean-François Bonastre, Téva Merlin |
EUROSPEECH | 1 |