EDBT 2026 Demo / reviewers in the wild / expert
Virginie Woisard
dblp:38/10649
· DBLP profile ↗
20ranked-venue papers
0as first author
13since 2021 · last 2024
0000-0003-3895-2827ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Interpretable Assessment of Speech Intelligibility Using Deep Learning: A Case Study on Speech Disorders Due to Head and Neck CancersabstractThis paper sheds light on a relatively unexplored area which is deep learning interpretability for speech disorder assessment and characterization. Building upon a state-of-the-art methodology for the explainability and interpretability of hidden representation inside a deep-learning speech model, we provide a deeper understanding and interpretation of the final intelligibility assessment of patients experiencing speech disorders due to Head and Neck Cancers (HNC). Promising results have been obtained regarding the prediction of speech intelligibility and severity of HNC patients while giving relevant interpretations of the final assessment both at the phonemes and phonetic feature levels. The potential of this approach becomes evident as clinicians can acquire more valuable insights for speech therapy. Indeed, this can help identify the specific linguistic units that affect intelligibility from an acoustic point of view and enable the development of tailored rehabilitation protocols to improve the patient’s ability to communicate effectively, and thus, the patient’s quality of life. Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Mathieu Balaguer, Virginie Woisard |
LREC/COLING | 7 |
| 2024 | Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce ContextabstractAutomatic speech quality assessment has raised more attention as an alternative or support to traditional perceptual clinical evaluation. However, most research so far only gains good results on simple tasks such as binary classification, largely due to data scarcity. To deal with this challenge, current works tend to segment patients’ audio files into many samples to augment the datasets. Nevertheless, this approach has limitations, as it indirectly relates overall audio scores to individual segments. This paper introduces a novel approach where the system learns at the audio level instead of segments despite data scarcity. This paper proposes to use the pre-trained Wav2Vec2 architecture for both SSL, and ASR as feature extractor in speech assessment. Carried out on the HNC dataset, our ASR-driven approach established a new baseline compared with other approaches, obtaining average MSE = 0.73 and MSE = 1.15 for the prediction of intelligibility and severity scores respectively, using only 95 training samples. It shows that the ASR based Wav2Vec2 model brings the best results and may indicate a strong correlation between ASR and speech quality assessment. We also measure its ability on variable segment durations and speech content, exploring factors influencing its decision. Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard |
LREC/COLING | 5 |
| 2024 | Electroglottography for the assessment of dysphonia in Parkinson's disease and multiple system atrophyabstractElectroglottography (EGG) is a noninvasive technique which allows accurate measurements of vocal folds dynamics and perturbations during speech. It has been widely used in medical diagnosis and monitoring of several laryngeal pathologies, but its use in neurological disorders has been very limited. This paper presents the first study on EGG in early stages of Parkinson's disease (PD) and an atypical parkinsonian disorder, multiple system atrophy (MSA). Our first objective was to investigate whether EGG can reveal distinctive dysphonia features that could help in the challenging early differential diagnosis between PD and MSA-P, the parkinsonian variant of MSA. The second objective was to verify some hypothesis on early markers of PD drawn from speech-alone acoustic analysis. For MSA-P patients, our analysis revealed a reduced open quotient and confirmed excessive pitch variation. The analysis also mitigated some hypothesis on dysphonia in early stages of PD. © 2024 International Speech Communication Association. All rights reserved. Khalid Daoudi, Solange Milhé de Saint Victor, Alexandra Foubert-Samier, Margherita Fabbri, Anne Pavy-Le Traon, Olivier Rascol, Virginie Woisard, Wassilios Meissner |
INTERSPEECH | 7 |
| 2024 | Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based modelsabstractHead and Neck Cancers (HNC) significantly impact patients' ability to speak, affecting their quality of life. Commonly used metrics for assessing pathological speech are subjective, prompting the need for automated and unbiased evaluation methods. This study proposes a self-supervised Wav2Vec2-based model for phone classification with HNC patients, to enhance accuracy and improve the discrimination of phonetic features for subsequent interpretability purpose. The impact of pre-training datasets, model size, and fine-tuning datasets and parameters are explored. Evaluation on diverse corpora reveals the effectiveness of the Wav2Vec2 architecture, outperforming a CNN-based approach, used in previous work. Correlation with perceptual measures also affirms the model relevance for impaired speech analysis. This work paves the way for better understanding of pathological speech with interpretable approaches for clinicians, by leveraging complex self-learnt speech representations. Malo Maisonneuve, Corinne Fredouille, Muriel Lalain, Alain Ghio, Virginie Woisard |
INTERSPEECH | 5 |
| 2024 | Exploring ASR-Based WAV2VEC2 for Automated Speech Disorder Assessment: Insights and AnalysisabstractWith the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pretrained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability. Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard |
SLT | 5 |
| 2023 | Can We Use Speaker Embeddings On Spontaneous Speech Obtained From Medical Conversations To Predict Intelligibility?abstractThe automatic prediction of speech intelligibility is a recurrent problem in the context of pathological speech. Despite recent developments, these systems are normally applied to specific speech tasks recorded in clean conditions that do not necessarily reflect a healthcare environment. In the present paper, we intend to test the reliability of an intelligibility predictor on data obtained in clinical conditions, in the specific case of head and neck cancer. In order to do so, we present a system based on speaker embeddings trained on a multi-task methodology to simultaneous predict speech intelligibility and speech disorder severity. The results obtained on the different evaluation tasks display correlations as high as 0.891 on a hospital patient set, showing robustness to the type of speech material used in these automatic assessments. Moreover, the usage of spontaneous speech during the evaluation shed light on an understudied, but with more ecological validity, type of speech material which displayed promising results. The reliability displayed across the different tasks suggests a direct deployment of the developed systems in a hospital setting. Sebastião Quintas, Mathieu Balaguer, Julie Mauclair, Virginie Woisard, Julien Pinquier |
ASRU | 4 |
| 2023 | Towards Reducing Patient Effort for the Automatic Prediction of Speech Intelligibility in Head and Neck CancersabstractThe automatic prediction of speech intelligibility can be seen as a growing and relevant alternative to the perceptual evaluations used clinically, which are known to be biased, variant and subjective. We propose an automatic way to regress an intelligibility score based on a recurrent model with a self-attention mechanism. This approach not only presented a high correlation of 0.87 when applied to a pseudo-word task designed for head and neck cancers, but also a significant decrease in error of more than 50%, when compared to previous approaches. Moreover, we have also studied the reliability of the same system when operating with smaller amounts of data at inference time. The results suggest that we can reduce the linguistic sample size to only 30% of the full sample, without losing performance. This aspect validates the reliability of using a smaller subset of data when predicting intelligibility, which can be extremely useful to prevent patient’s fatigue by creating smaller batteries of clinical exams. Sebastião Quintas, Alberto Abad, Julie Mauclair, Virginie Woisard, Julien Pinquier |
ICASSP | 4 |
| 2023 | Interpreting Deep Representations of Phonetic Features via Neuro-Based Concept Detector: Application to Speech Disorders Due to Head and Neck CancerabstractThe popularity of Deep Neural Networks (DNNs) is growing significantly, and so is the interest in gaining a better understanding of their functioning. In this work, it is even more interesting to reveal the behavior of these black-boxes since we are involved in a clinical context. To this end, we propose a general analytic framework, namedNeuro-based Concept Detector (NCD), for interpreting deep representations of a DNN. Based on the activation patterns of the hidden neurons, this framework highlights the capacity of neurons to detect a specific concept related to the final task. The key strength of our framework is that it provides an interpretability tool for any type of DNN performing a classification task regardless of the application field. In this paper, we evaluate this framework on a Convolutional Neural Network (CNN) trained for the task of French phone classification. This choice was guided by the final objective of a long-term research project, which aims to identify the linguistic units best contributing to the maintenance or loss of intelligibility in the context of speech disorders. ThroughNCD, we demonstrate the emergence of phonetic features in the classification layers of the CNN-based model, while applied on healthy speech, a concept with a great interest in the field of clinical phonetics. Indeed, we further show that these interesting findings shed light on the characteristics of speech disorders in terms of altered phonetic features and provide relevant information for clinical practice, notably, patients' rehabilitation and follow-up. Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2022 | Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders Step 2: Contribution of the Emergence of Phonetic TraitsabstractApart from the impressive performance it has achieved in several tasks, one of the most important factors remaining for the continuous progress of deep learning is the increased work related to interpretability, especially in a medical context. In a recent work, we presented competitive performance achieved with a CNN-based model trained on normal speech for the French phone classification and how it correlates well with different perceptual measures when exposed to disordered speech. This paper extends that work by focusing on interpretability. Here, the goal is to get insights into the way in which neural representations shape the final task of phone classification so that it can be used further to explain the loss of intelligibility in disordered speech. In this way, an original framework is proposed, relying firstly on the neural activity and a novel representation per neuron, here considering the phone classification, and, secondly, permitting to identify a set of neurons devoted to the detection of specific phonetic traits on normal speech. Faced to disordered speech, a degradation of that set of neurons is observed, demonstrating a loss of specific phonetic traits in some patients involved, and the potentiality of the proposed approaches to inform about speech alteration. Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
ICASSP | 6 |
| 2022 | Validation of the Neuro-Concept Detector framework for the characterization of speech disorders: A comparative study including Dysarthria and DysphoniaabstractInternational audience Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
INTERSPEECH | 6 |
| 2022 | A comparative study on vowel articulation in Parkinson's disease and multiple system atrophyabstractAcoustic realisation of the working vowel space has been widely studied in Parkinson's disease (PD).However, it has never been studied in atypical parkinsonian disorders (APD).The latter are neurodegenerative diseases which share similar clinical features with PD, rendering the differential diagnosis very challenging in early disease stages.This paper presents the first contribution in vowel space analysis in APD, by comparing corner vowel realisation in PD and the parkinsonian variant of Multiple System Atrophy (MSA-P).Our study has the particularity of focusing exclusively on early stage PD and MSA-P patients, as our main purpose was early differential diagnosis between these two diseases.We analysed the corner vowels, extracted from a spoken sentence, using traditional vowel space metrics.We found no statistical difference between the PD group and healthy controls (HC) while MSA-P exhibited significant differences with the PD and HC groups.We also found that some metrics conveyed complementary discriminative information.Consequently, we argue that restriction in the acoustic realisation of corner vowels cannot be a viable early marker of PD, as hypothesised by some studies, but it might be a candidate as an early hypokinetic marker of MSA-P (when the clinical target is discrimination between PD and MSA-P). Khalid Daoudi, Biswajit Das, Solange Milhé de Saint Victor, Alexandra Foubert-Samier, Margherita Fabbri, Anne Pavy-Le Traon, Olivier Rascol, Virginie Woisard, Wassilios Meissner |
INTERSPEECH | 8 |
| 2022 | Automatic Assessment of Speech Intelligibility using Consonant Similarity for Head and Neck CancerabstractInternational audience Sebastião Quintas, Julie Mauclair, Virginie Woisard, Julien Pinquier |
INTERSPEECH | 3 |
| 2021 | Distortion of Voiced Obstruents for Differential Diagnosis Between Parkinson's Disease and Multiple System AtrophyabstractParkinson's disease (PD) and the parkinsonian variant of Multiple System Atrophy (MSA-P) are two neurodegenerative diseases which share similar clinical features, particularly in early disease stages. The differential diagnosis can be thus very challenging. Dysarthria is known to be a frequent and early clinical feature of PD and MSA. It can be thus used as a vehicle to provide a vocal biomarker which could help in the differential diagnosis. In particular, distortion of consonants is known to be a frequent impairment in these diseases. The aim of this study is to investigate distinctive patterns in the distortion of voiced obstruents (plosives and fricatives). It is the first study which attempts to examine such distortions in the French language for the purpose of the differential diagnosis between PD and MSAP (and among the very few studies if we consider all languages). We carry out a perceptual and objective analysis of voiced obstruents extracted from isolated pseudo-words initials. We first show that devoicing is a significant impairment which predominates in MSA-P. We then show that voice onset time (VOT) of voiced plosives (prevoicing duration) can be a complementary feature to improve the accuracy in discrimination between PD and MSA-P. Khalid Daoudi, Biswajit Das, Solange Milhé de Saint Victor, Alexandra Foubert-Samier, Anne Pavy-Le Traon, Olivier Rascol, Wassilios Meissner, Virginie Woisard |
Interspeech | 8 |
| 2020 | Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders - Step 1: CNN Model-Based Phone ClassificationabstractInternational audience Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard |
INTERSPEECH | 6 |
| 2020 | Automatic Prediction of Speech Intelligibility Based on X-Vectors in the Context of Head and Neck CancerabstractInternational audience Sebastião Quintas, Julie Mauclair, Virginie Woisard, Julien Pinquier |
INTERSPEECH | 3 |
| 2020 | How to Compare Automatically Two Phonological Strings: Application to Intelligibility Measurement in the Case of Atypical SpeechabstractAtypical speech productions, regardless of their origins (accents, learning, pathology), need to be assessed with regard to “typical” or “expected” productions. Evaluation is necessarily based on comparisons between linguistic forms produced and linguistic forms expected. In the field of speech disorders, the intelligibility of a patient is evaluated in order to measure the functional impact of his/her pathology on his/her oral communication. The usual method is to transcribe orthographic linguistic forms perceived and to assign a global and imprecise rating based on their correctness or incorrect. To obtain a more precise evaluation of the production deviations, we propose a measurement method based on phonological transcriptions. An algorithm computes automatically and finely the distances between the phonological forms produced and expected from cost matrices based on the differences of features between phonemes. A first test of this method among a large population of healthy speakers and patients treated for cancer of the oral and pharyngeal cavities has proved its validity. Alain Ghio, Muriel Lalain, Laurence Giusti, Corinne Fredouille, Virginie Woisard |
LREC | 5 |
| 2020 | Have a Cake and Eat it Too: Assessing Discriminating Performance of an Intelligibility Index Obtained from a Reduced Sample SizeabstractThis paper investigates random vs. phonetically motivated reduction of linguistic material used in an intelligibility task in speech disordered populations and the subsequent impact on the discrimination classifier quantified by the area under the receiver operating characteristics curve (AUC of ROC). The comparison of obtained accuracy indexes shows that when the sample size is reduced based on a phonetic criterium—here, related to phonotactic complexity—, the classifier has a higher ranking ability than when the linguistic material is arbitrarily reduced. Crucially, downsizing the linguistic sample to about 30% of the original dataset does not diminish the discriminatory performance of the classifier. This result is of significant interest to both clinicians and patients as it validates a tool that is both reliable and efficient. Anna K. Marczyk, Alain Ghio, Muriel Lalain, Marie Rebourg, Corinne Fredouille, Virginie Woisard |
LREC | 6 |
| 2018 | Automatic Evaluation of Speech Intelligibility Based on I-vectors in the Context of Head and Neck CancersabstractInternational audience Imed Laaridh, Corinne Fredouille, Alain Ghio, Muriel Lalain, Virginie Woisard |
INTERSPEECH | 5 |
| 2018 | Carcinologic Speech Severity Index Project: A Database of Speech Disorder Productions to Assess Quality of Life Related to Speech After Cancer
Corine Astésano, Mathieu Balaguer, Jérôme Farinas, Corinne Fredouille, Pascal Gaillard, Alain Ghio, Imed Laaridh, Muriel Lalain, Benoît Lepage, Julie Mauclair, Olivier Nocaudie, Julien Pinquier, Oriol Pont, Gilles Pouchoulin, Michèle Puech, Danièle Robert, Etienne Sicard, Virginie Woisard |
LREC | 18 |
| 2011 | Is the Perception of Voice Quality Language-Dependant? A Comparison of French and Italian Listeners and Dysphonic Speakers
Alain Ghio, Frédérique Weisz, Giovanna Baracca, Giovanna Cantarella, Danièle Robert, Virginie Woisard, Franco Fussi, Antoine Giovanni |
INTERSPEECH | 6 |