Alain Ghio

dblp:26/4212 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0001-7302-0799ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Interpretable Assessment of Speech Intelligibility Using Deep Learning: A Case Study on Speech Disorders Due to Head and Neck Cancers
abstract
This paper sheds light on a relatively unexplored area which is deep learning interpretability for speech disorder assessment and characterization. Building upon a state-of-the-art methodology for the explainability and interpretability of hidden representation inside a deep-learning speech model, we provide a deeper understanding and interpretation of the final intelligibility assessment of patients experiencing speech disorders due to Head and Neck Cancers (HNC). Promising results have been obtained regarding the prediction of speech intelligibility and severity of HNC patients while giving relevant interpretations of the final assessment both at the phonemes and phonetic feature levels. The potential of this approach becomes evident as clinicians can acquire more valuable insights for speech therapy. Indeed, this can help identify the specific linguistic units that affect intelligibility from an acoustic point of view and enable the development of tailored rehabilitation protocols to improve the patient’s ability to communicate effectively, and thus, the patient’s quality of life.
Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Mathieu Balaguer, Virginie Woisard
LREC/COLING3
2024 Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
abstract
Automatic speech quality assessment has raised more attention as an alternative or support to traditional perceptual clinical evaluation. However, most research so far only gains good results on simple tasks such as binary classification, largely due to data scarcity. To deal with this challenge, current works tend to segment patients’ audio files into many samples to augment the datasets. Nevertheless, this approach has limitations, as it indirectly relates overall audio scores to individual segments. This paper introduces a novel approach where the system learns at the audio level instead of segments despite data scarcity. This paper proposes to use the pre-trained Wav2Vec2 architecture for both SSL, and ASR as feature extractor in speech assessment. Carried out on the HNC dataset, our ASR-driven approach established a new baseline compared with other approaches, obtaining average MSE = 0.73 and MSE = 1.15 for the prediction of intelligibility and severity scores respectively, using only 95 training samples. It shows that the ASR based Wav2Vec2 model brings the best results and may indicate a strong correlation between ASR and speech quality assessment. We also measure its ability on variable segment durations and speech content, exploring factors influencing its decision.
Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard
LREC/COLING3
2024 Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
abstract
Head and Neck Cancers (HNC) significantly impact patients' ability to speak, affecting their quality of life. Commonly used metrics for assessing pathological speech are subjective, prompting the need for automated and unbiased evaluation methods. This study proposes a self-supervised Wav2Vec2-based model for phone classification with HNC patients, to enhance accuracy and improve the discrimination of phonetic features for subsequent interpretability purpose. The impact of pre-training datasets, model size, and fine-tuning datasets and parameters are explored. Evaluation on diverse corpora reveals the effectiveness of the Wav2Vec2 architecture, outperforming a CNN-based approach, used in previous work. Correlation with perceptual measures also affirms the model relevance for impaired speech analysis. This work paves the way for better understanding of pathological speech with interpretable approaches for clinicians, by leveraging complex self-learnt speech representations.
Malo Maisonneuve, Corinne Fredouille, Muriel Lalain, Alain Ghio, Virginie Woisard
INTERSPEECH4
2024 Exploring ASR-Based WAV2VEC2 for Automated Speech Disorder Assessment: Insights and Analysis
abstract
With the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pretrained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability.
Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard
SLT3
2024 Evaluating the effects of task design on unfamiliar Francophone listener and automatic speaker identification performance
Benjamin O'Brien, Christine Meunier, Natalia A. Tomashenko, Alain Ghio, Jean-François Bonastre
Multim. Tools Appl.4
2024 Evaluating the effects of continuous pitch and speech tempo modifications on perceptual speaker verification performance by familiar and unfamiliar listeners
Benjamin O'Brien, Christine Meunier, Alain Ghio
Speech Commun.3
2023 Interpreting Deep Representations of Phonetic Features via Neuro-Based Concept Detector: Application to Speech Disorders Due to Head and Neck Cancer
abstract
The popularity of Deep Neural Networks (DNNs) is growing significantly, and so is the interest in gaining a better understanding of their functioning. In this work, it is even more interesting to reveal the behavior of these black-boxes since we are involved in a clinical context. To this end, we propose a general analytic framework, namedNeuro-based Concept Detector (NCD), for interpreting deep representations of a DNN. Based on the activation patterns of the hidden neurons, this framework highlights the capacity of neurons to detect a specific concept related to the final task. The key strength of our framework is that it provides an interpretability tool for any type of DNN performing a classification task regardless of the application field. In this paper, we evaluate this framework on a Convolutional Neural Network (CNN) trained for the task of French phone classification. This choice was guided by the final objective of a long-term research project, which aims to identify the linguistic units best contributing to the maintenance or loss of intelligibility in the context of speech disorders. ThroughNCD, we demonstrate the emergence of phonetic features in the classification layers of the CNN-based model, while applied on healthy speech, a concept with a great interest in the field of clinical phonetics. Indeed, we further show that these interesting findings shed light on the characteristics of speech disorders in terms of altered phonetic features and provide relevant information for clinical practice, notably, patients' rehabilitation and follow-up.
Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders Step 2: Contribution of the Emergence of Phonetic Traits
abstract
Apart from the impressive performance it has achieved in several tasks, one of the most important factors remaining for the continuous progress of deep learning is the increased work related to interpretability, especially in a medical context. In a recent work, we presented competitive performance achieved with a CNN-based model trained on normal speech for the French phone classification and how it correlates well with different perceptual measures when exposed to disordered speech. This paper extends that work by focusing on interpretability. Here, the goal is to get insights into the way in which neural representations shape the final task of phone classification so that it can be used further to explain the loss of intelligibility in disordered speech. In this way, an original framework is proposed, relying firstly on the neural activity and a novel representation per neuron, here considering the phone classification, and, secondly, permitting to identify a set of neurons devoted to the detection of specific phonetic traits on normal speech. Faced to disordered speech, a degradation of that set of neurons is observed, demonstrating a loss of specific phonetic traits in some patients involved, and the potentiality of the proposed approaches to inform about speech alteration.
Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard
ICASSP3
2022 Validation of the Neuro-Concept Detector framework for the characterization of speech disorders: A comparative study including Dysarthria and Dysphonia
abstract
International audience
Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard
INTERSPEECH3
2022 Evaluating the effects of modified speech on perceptual speaker identification performance
abstract
International audience
Benjamin O'Brien, Christine Meunier, Alain Ghio
INTERSPEECH3
2022 The Speed-Vel Project: a Corpus of Acoustic and Aerodynamic Data to Measure Droplets Emission During Speech Interaction
abstract
Conversations (normal speech) or professional interactions (e.g., projected speech in the classroom) have been identified as situations with increased risk of exposure to SARS-CoV-2 due to the high production of droplets in the exhaled air. However, it is still unclear to what extent speech properties influence droplets emission during everyday life conversations. Here, we report the experimental protocol of three experiments aiming at measuring the velocity and the direction of the airflow, the number and size of droplets spread during speech interactions in French. We consider different phonetic conditions, potentially leading to a modulation of speech droplets production, such as voice intensity (normal vs. loud voice), articulation manner of phonemes (type of consonants and vowels) and prosody (i.e., the melody of the speech). Findings from these experiments will allow future simulation studies to predict the transport, dispersion and evaporation of droplets emitted under different speech conditions.
Francesca Carbone, Gilles Bouchet, Alain Ghio, Thierry Legou, Carine André, Muriel Lalain, Sabrina Kadri, Caterina Petrone, Federica Procino, Antoine Giovanni
LREC3
2021 Presentation Matters: Evaluating Speaker Identification Tasks
Benjamin O'Brien, Christine Meunier, Alain Ghio
Interspeech3
2020 Towards Interpreting Deep Learning Models to Understand Loss of Speech Intelligibility in Speech Disorders - Step 1: CNN Model-Based Phone Classification
abstract
International audience
Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Virginie Woisard
INTERSPEECH3
2020 How to Compare Automatically Two Phonological Strings: Application to Intelligibility Measurement in the Case of Atypical Speech
abstract
Atypical speech productions, regardless of their origins (accents, learning, pathology), need to be assessed with regard to “typical” or “expected” productions. Evaluation is necessarily based on comparisons between linguistic forms produced and linguistic forms expected. In the field of speech disorders, the intelligibility of a patient is evaluated in order to measure the functional impact of his/her pathology on his/her oral communication. The usual method is to transcribe orthographic linguistic forms perceived and to assign a global and imprecise rating based on their correctness or incorrect. To obtain a more precise evaluation of the production deviations, we propose a measurement method based on phonological transcriptions. An algorithm computes automatically and finely the distances between the phonological forms produced and expected from cost matrices based on the differences of features between phonemes. A first test of this method among a large population of healthy speakers and patients treated for cancer of the oral and pharyngeal cavities has proved its validity.
Alain Ghio, Muriel Lalain, Laurence Giusti, Corinne Fredouille, Virginie Woisard
LREC1
2020 Have a Cake and Eat it Too: Assessing Discriminating Performance of an Intelligibility Index Obtained from a Reduced Sample Size
abstract
This paper investigates random vs. phonetically motivated reduction of linguistic material used in an intelligibility task in speech disordered populations and the subsequent impact on the discrimination classifier quantified by the area under the receiver operating characteristics curve (AUC of ROC). The comparison of obtained accuracy indexes shows that when the sample size is reduced based on a phonetic criterium—here, related to phonotactic complexity—, the classifier has a higher ranking ability than when the linguistic material is arbitrarily reduced. Crucially, downsizing the linguistic sample to about 30% of the original dataset does not diminish the discriminatory performance of the classifier. This result is of significant interest to both clinicians and patients as it validates a tool that is both reliable and efficient.
Anna K. Marczyk, Alain Ghio, Muriel Lalain, Marie Rebourg, Corinne Fredouille, Virginie Woisard
LREC2
2018 Automatic Evaluation of Speech Intelligibility Based on I-vectors in the Context of Head and Neck Cancers
abstract
International audience
Imed Laaridh, Corinne Fredouille, Alain Ghio, Muriel Lalain, Virginie Woisard
INTERSPEECH3
2018 Carcinologic Speech Severity Index Project: A Database of Speech Disorder Productions to Assess Quality of Life Related to Speech After Cancer
Corine Astésano, Mathieu Balaguer, Jérôme Farinas, Corinne Fredouille, Pascal Gaillard, Alain Ghio, Imed Laaridh, Muriel Lalain, Benoît Lepage, Julie Mauclair, Olivier Nocaudie, Julien Pinquier, Oriol Pont, Gilles Pouchoulin, Michèle Puech, Danièle Robert, Etienne Sicard, Virginie Woisard
LREC6
2017 The Phonological Status of the French Initial Accent and its Role in Semantic Processing: An Event-Related Potentials Study
abstract
International audience
Noémie te Rietmolen, Radouane El Yagoubi, Alain Ghio, Corine Astésano
INTERSPEECH3
2016 The TYPALOC Corpus: A Collection of Various Dysarthric Speech Recordings in Read and Spontaneous Styles
Christine Meunier, Cécile Fougeron, Corinne Fredouille, Brigitte Bigi, Lise Crevier-Buchman, Elisabeth Delais-Roussarie, Laurianne Georgeton, Alain Ghio, Imed Laaridh, Thierry Legou, Claire Pillot-Loiseau, Gilles Pouchoulin
LREC8
2015 Vocal tremor analysis via AM-FM decomposition of empirical modes of the glottal cycle length time series
abstract
The presentation concerns a method that obtains the size and frequency of vocal tremor in speech sounds sustained by normal speakers and patients suffering from neurological disorders.The glottal cycle lengths are tracked in the temporal domain via salience analysis and dynamic programming.The cycle length time series is then decomposed into a sum of oscillating components by empirical mode decomposition the instantaneous envelopes and frequencies of which are obtained via an AM-FM decomposition.Based on their average instantaneous frequencies, the empirical modes are then assigned to four categories (intonation, physiological tremor, neurological tremor as well as jitter) and added within each.The within-category size of the cycle length perturbations is estimated via the standard deviation of the empirical mode sum divided by the average cycle length.The tremor frequency within the neurological tremor category is obtained via a weighted instantaneous average of the mode frequencies followed by a weighted temporal average.The method is applied to two corpora of vowels sustained by 123 and 74 control and 456 and 205 Parkinson speakers respectively.
Christophe Mertens, Francis Grenez, François Viallet, Alain Ghio, Sabine Skodda, Jean Schoentgen
INTERSPEECH4
2013 Perceptual interference between regional accent and voice/speech disorders
abstract
We present a study where we examined the influence of a regional accent in the perception of voice and/or speech disorders.These aspects are most of the time overshadowed in clinical context.This protocol, involving multiple sources of speech variations, is also interesting for perception theories.For the experiment, speakers with or without a Southern French accent and with or without speech/voice disorders were recorded on reading a text.The samples were then randomly played back to two groups of listeners (familiar vs unfamiliar with the regional accent), specialists in speech therapy.The task was the perceptual evaluation of voice quality, articulation disorders and dysprosody.We focused in this paper on the voice dimension.The main results on this part concern the weak influence of regional accent on the perception of moderate or severe dysphonia, where the speech signal is strongly disturbed by the disorder.By contrast, the effect of regional accent is important on normal voices perception: listeners unfamiliar with the regional accent judge speakers with accent without voice disorder as slightly dysphonic.This last result can be interpreted as a form of perceptual interference between different dimensions of speech variations around a central position.
Alain Ghio, Médéric Gasquet-Cyrus, Juliette Roquel, Antoine Giovanni
INTERSPEECH1
2012 How to manage sound, physiological and clinical data of 2500 dysphonic and dysarthric speakers?
Alain Ghio, Gilles Pouchoulin, Bernard Teston, Serge Pinto, Corinne Fredouille, Céline De Looze, Danièle Robert, François Viallet, Antoine Giovanni
Speech Commun.1
2011 Is the Perception of Voice Quality Language-Dependant? A Comparison of French and Italian Listeners and Dysphonic Speakers
Alain Ghio, Frédérique Weisz, Giovanna Baracca, Giovanna Cantarella, Danièle Robert, Virginie Woisard, Franco Fussi, Antoine Giovanni
INTERSPEECH1
2010 The DesPho-APaDy Project: Developing an Acoustic-phonetic Characterization of Dysarthric Speech in French
Cécile Fougeron, Lise Crevier-Buchman, Corinne Fredouille, Alain Ghio, Christine Meunier, Claude Chevrie-Muller, Jean-François Bonastre, Antonia Colazo-Simon, Céline De Looze, Danielle Duez, Cédric Gendrot, Thierry Legou, Nathalie Lévêque, Claire Pillot-Loiseau, Serge Pinto, Gilles Pouchoulin, Danièle Robert, Jacqueline Vaissière, François Viallet, Coralie Vincent
LREC4
2008 Dysphonic voices and the 0-3000 hz frequency band
abstract
Concerned with pathological voice assessment, this paper aims at characterizing dysphonia in the frequency domain for a better understanding of related phenomena while most of the studies have focused only on improving classification systems for diag-nosis help purposes. Based on a first study which demonstrates that the low frequencies ([0-3000]Hz) are more relevant for dys-phonia discrimination compared with higher frequencies, the authors propose in this paper to pursue by analyzing the impact of the restricted frequency band ([0-3000]Hz) on the dysphonic voice discrimination from a phonetical and perceptual point of views. A discussion around the frequency band limitation of telephone channel is also proposed. Index Terms: Voice disorder, dysphonia characterization, au-tomatic dysphonic voice classification, frequency analysis
Gilles Pouchoulin, Corinne Fredouille, Jean-François Bonastre, Alain Ghio, Antoine Giovanni
INTERSPEECH4
2007 Complementary approaches for voice disorder assessment
abstract
This paper describes two comparative studies of voice quality assessment based on complementary approaches. The first study was undertaken on 449 speakers (including 391 dysphonic patients) whose voice quality was evaluated in parallel by a perceptual judgment and objective measurements on acoustic and aerodynamic data. Results showed that a nonlinear combination of 7 parameters allowed the classification of 82% voice samples in the same grade as the jury. The second study relates to the adaptation of Automatic Speaker Recognition (ASR) techniques to pathological voice assessment. The system designed for this particular task relies on a GMM based approach, which is the state-of-the-art for ASR. Experiments conducted on 80 female voices provide promising results, underlining the interest of such an approach. We benefit from the multiplicity of theses techniques to evaluate the methodological situation which points fundamental differences between these complementary approaches (bottom-up vs. top-down, global vs. analytic). We also discuss some theoretical aspects about relationship between acoustic measurement and perceptual mechanisms which are often forgotten in the performance race.
Jean-François Bonastre, Corinne Fredouille, Alain Ghio, Antoine Giovanni, Gilles Pouchoulin, Joana Revis, Bernard Teston
INTERSPEECH3
2007 Frequency study for the characterization of the dysphonic voices
abstract
Concerned with pathological voice assessment, this paper aims at characterizing dysphonia in the frequency domain for a better understanding of relating phenomena while most of the studies have focused only on improving classification systems for diagnosis help purposes.In this context, a GMM-based automatic classification system is applied on different frequency ranges in order to investigate which ones are relevant for dysphonia characterization.Experiment results demonstrate that the low frequencies [0-3000]Hz are more relevant for dysphonia discrimination compared with higher frequencies.
Gilles Pouchoulin, Corinne Fredouille, Jean-François Bonastre, Alain Ghio, Antoine Giovanni
INTERSPEECH4
2005 Application of automatic speaker recognition techniques to pathological voice assessment (dysphonia)
abstract
HAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Corinne Fredouille, Gilles Pouchoulin, Jean-François Bonastre, M. Azzarello, Antoine Giovanni, Alain Ghio
INTERSPEECH6