Julie M. Liss

dblp:136/5109 · also Julie Liss · DBLP profile ↗
← Back
37ranked-venue papers
0as first author
15since 2021 · last 2027
0000-0001-8782-2901ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 12 since 2021Artificial intelligence and machine learning · 18 · 10 since 2021
YearPublicationVenuePosition
2027 Evaluating the interpretability of clinical speech AI models: Lessons from two user studies
Lingfeng Xu, Visar Berisha, Julie M. Liss
Comput. Speech Lang.3
2025 Cross-lingual Evaluation Of Hypernasality Using Wav2Vec2 Features
abstract
Hypernasality, a speech resonance disorder characterized by excessive nasal airflow, presents challenges in accurate detection across languages. Traditional assessments of hypernasality include perceptual evaluation and nasometry. More recently, objective acoustic measures based on formant analysis and other acoustic features have been proposed as proxies for hypernasality; however, these acoustic measures exhibit considerable variability, especially in multilingual contexts. In this study, we utilize the wav2vec2-large-xlsr-53 model, a cross-lingual speech representation framework, and evaluate hypernasality-related features within its transformer layers. We extracted features from each layer of the model’s transformer and trained a machine learning model to predict hypernasality ratings across three datasets: Americleft (English), New Mexico Cleft Palate Center (English), and All India Institute of Speech and Hearing (Kannada). Our analysis reveals that the 11th and 12th layer contextualized embeddings effectively model hypernasality cross-lingually, demonstrating significant within-language average correlations (0.78 and 0.78) and cross-lingual average correlations (0.68 and 0.67) between predicted and perceptual ratings. These findings suggest that the wav2vec2-large-xlsr-53 model’s intermediate layers effectively capture hypernasality cross-lingually.
Krupaben Kothadia, Vikram C. M., Ajish K. Abraham, Pushpavathi M, S. R. Mahadeva Prasanna, Nancy Scherer, Kathy Chapman, Julie M. Liss, Visar Berisha
ICASSP8
2025 Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
abstract
Existing methods for analyzing linguistic content from picture descriptions for assessment of cognitive-linguistic impairment often overlook the participant’s visual narrative path, which typically requires eye tracking to assess. Spatio-semantic graphs are a useful tool for analyzing this narrative path from transcripts alone, however they are limited by the need for manual tagging of content information units (CIUs). In this paper, we propose an automated approach for estimation of spatio-semantic graphs (via automated extraction of CIUs) from the Cookie Theft picture commonly used in cognitive-linguistic analyses. The method enables the automatic characterization of the visual semantic path during picture description. Experiments demonstrate that the automatic spatio-semantic graphs effectively differentiate between cognitively impaired and unimpaired speakers. Statistical analyses reveal that the features derived by the automated method produce comparable results to the manual method, with even greater group differences between clinical groups of interest. These results highlight the potential of the automated approach for extracting spatio-semantic features in developing clinical speech models for cognitive impairment assessment.
Si-Ioi Ng, Pranav S. Ambadi, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
ICASSP4
2025 The Impact of Decorrelation on Transformer Interpretation Methods: Applications to Clinical Speech AI
abstract
Recent applications of decorrelation methods to the multi-head attention layers and output embeddings of transformer-based models have resulted in improvements in efficiency and accuracy. Despite these advancements, there is a lack of research focused on the influence of decorrelation on transformer interpretation techniques. This study investigated the impact of two decorrelation methods on interpreting the decision-making logic of a Bidirectional Encoder Representations from Transformers (BERT) model. Two metrics, namely Comprehensiveness and Sufficiency, were used to quantify the interpretation quality, while the changes in correlation within each multi-head self-attention layer was statistically analyzed. Results indicate that decorrelating BERT embeddings leads to a sparser distribution of weights in the middle attention layers and a significantly improved interpretation quality. Conversely, decorrelating the attention maps of specific attention layers increases the correlation in the corresponding attention weight matrices, yielding a less marked improvement in interpretation quality and, in some instances, degraded model performance.
Lingfeng Xu, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
ICASSP3
2025 Mitigating Overfitting During Speech Foundation Model Fine-tuning: Applications to Dysarthric Speech Detection
Yan Xiong 0002, Visar Berisha, Julie M. Liss, Chaitali Chakrabarti
INTERSPEECH3
2024 How Does Alignment Error Affect Automated Pronunciation Scoring in Children's Speech?
abstract
following cross utterance averaging. Thus, practical comparisons between child speakers should be very comparable across the two methods.
Prad Kadambi, Tristan J. Mahr, Lucas Annear, Henry Nomeland, Julie M. Liss, Katherine C. Hustad, Visar Berisha
INTERSPEECH5
2024 Segmental and Suprasegmental Speech Foundation Models for Classifying Cognitive Risk Factors: Evaluating Out-of-the-Box Performance
abstract
Speech foundation models are remarkably successful in various consumer applications, prompting their extension to clinical use-cases. This is challenged by small clinical datasets, which precludes effective fine-tuning. We tested the efficacy of two models to classify participants by segmental (Wav2Vec2.0) and suprasegmental (Trillsson) speech analysis windows. Analysis at both time scales has shown differences in the context of cognitive decline. Speakers were classified as healthy controls (HC), Amyloid-β+ (Aβ+), mild cognitive impairment (MCI), or dementia. A subset of W2V2 and Trillsson representations showed large effect size between HC and each risk factor. Cross-validation showed W2V2 consistently outperforms Trillsson. Mean macro-F1 of 54.1%, 63.5%, and 72.0% in were found for classifying Aβ+, MCI, and dementia from HC. Repeatability of Trillsson and W2V2 showed intraclass correlations of 0.30 and 0.41. Reliability of such models must be enhanced for clinical speech analysis and longitudinal tracking.
Si-Ioi Ng, Lingfeng Xu, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
INTERSPEECH4
2024 Improving Speech-Based Dysarthria Detection using Multi-task Learning with Gradient Projection
Yan Xiong 0002, Visar Berisha, Julie M. Liss, Chaitali Chakrabarti
INTERSPEECH3
2023 Decorrelating Language Model Embeddings for Speech-Based Prediction of Cognitive Impairment
abstract
Training robust clinical speech-based models that generalize requires large sample sizes because speech is variable and high-dimensional. Researchers have turned to foundational models, such as the Bidirectional Encoder Representations from Transformers (BERT), to generate lower-dimensional embeddings, and then finetuned the models for a specific down-stream clinical task. While there is empirical evidence that this approach is helpful, a recent study reveals that the embeddings generated by BERT models tend to be highly correlated, which makes the downstream models difficult to fine-tune, particularly in the small sample size regime. In this work, we propose a new regularization scheme to penalize correlated embeddings during fine tuning of BERT and apply the approach to speech-based assessment of cognitive impairment. Compared to existing methods, the proposed method yields lower estimation errors and smaller false alarm rates in a Mini-Mental State Examination (MMSE) score regression task.
Lingfeng Xu, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
ICASSP3
2023 Consonant-Vowel Transition Models Based on Deep Learning for Objective Evaluation of Articulation
abstract
Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the objective articulation measure (OAM), by analyzing the CV transitions segmented around vowel onsets. The OAM is derived based on the posteriors of a convolutional neural network pre-trained to classify between different consonants using CV regions as input. We demonstrate that the OAM is correlated with perceptual measures in a variety of contexts including (a) adult dysarthric speech, (b) the speech of children with cleft lip/palate, and (c) a database of accented English speech from native Mandarin and Spanish speakers.
Vikram C. M., Julie M. Liss, Kathy Chapman, Nancy Scherer, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Robust Vocal Quality Feature Embeddings for Dysphonic Voice Detection
abstract
Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment often fail to generalize outside the training conditions or to other related applications. In this paper, we propose a deep learning framework for generating acoustic feature embeddings sensitive to vocal quality and robust across different corpora. A contrastive loss is combined with a classification loss to train our deep learning model jointly. Data warping methods are used on input voice samples to improve the robustness of our method. Empirical results demonstrate that our method not only achieves high in-corpus and cross-corpus classification accuracy but also generates good embeddings sensitive to voice quality and robust across different corpora. We also compare our results against three baseline methods on clean and three variations of deteriorated in-corpus and cross-corpus datasets and demonstrate that the proposed model consistently outperforms the baseline methods.
Julie M. Liss, Suren Jayasuriya, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Are reported accuracies in the clinical speech machine learning literature overoptimistic?
Visar Berisha, Chelsea Krantsevich, Gabriela Stegmann, Shira Hahn, Julie M. Liss
INTERSPEECH5
2022 Investigating the Impact of Speech Compression on the Acoustics of Dysarthric Speech
abstract
Acoustic analysis plays an important role in the assessment of dysarthria. Out of a public health necessity, telepractice has become increasingly adopted as the modality in which clinical care is given. While there are differences in software among telepractice platforms, they all use some form of speech compression to preserve bandwidth, with the most common algorithm being the Opus codec. Opus has been optimized for compression of speech from the general (mostly healthy) population. As a result, for speech-language pathologists, this begs the question: is the remotely transmitted speech signal a faithful representation of dysarthric speech? Existing high-fidelity audio recordings from 20 speakers of various dysarthria types were encoded at three different bit rates defined within Opus to simulate different internet bandwidth conditions. Acoustic measures of articulation, voice, and prosody were extracted, and mixed-effect models were used to evaluate the impact of bandwidth conditions on the measures. Significant differences in cepstral peak prominence, degree of voice breaks, jitter, vowel space area, pitch, and vowel space area were observed after Opus processing, providing insight into the types of acoustic measures that are susceptible to speech compression algorithms.
Kelvin Tran, Lingfeng Xu, Gabriela Stegmann, Julie M. Liss, Visar Berisha, Rene Utianski
INTERSPEECH4
2021 An Attention Model for Hypernasality Prediction in Children with Cleft Palate
abstract
Hypernasality refers to the perception of abnormal nasal resonances in vowels and voiced consonants. Estimation of hypernasality severity from connected speech samples involves learning a mapping between the frame-level features and utterance-level clinical ratings of hypernasality. However, not all speech frames contribute equally to the perception of hypernasality. In this work, we propose an attention-based bidirectional long-short memory (BLSTM) model that directly maps the frame-level features to utterance-level ratings by focusing only on specific speech frames carrying hyper-nasal cues. The models performance is evaluated on the Americleft database containing speech samples of children with cleft palate and clinical ratings of hypernasality. We analyzed the attention weights over broad phonetic categories and found that the model yields results consistent with what is known in the speech science literature. Further, the correlation between the predicted and perceptual rating is found to be significant (r = 0.684, p < 0.001) and better than conventional BLSTMs trained using frame-wise and last-frame approaches.
Vikram C. M., Nancy Scherer, Kathy Chapman, Julie M. Liss, Visar Berisha
ICASSP4
2021 The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation
Vikram C. M., Tristan J. Mahr, Nancy Scherer, Kathy Chapman, Katherine C. Hustad, Julie M. Liss, Visar Berisha
Interspeech6
2020 Deep Learning Based Prediction of Hypernasality for Clinical Applications
abstract
Hypernasality refers to the perception of excessive nasal resonance during the production of oral sounds. Existing methods for automatic assessment of hypernasality from speech are based on machine learning models trained on disordered speech databases rated by speech-language pathologists. However, the performance of such systems critically depends on the availability of hypernasal speech samples and the reliability of clinical ratings. In this paper, we propose a new approach that uses the speech samples from healthy controls to model the acoustic characteristics of nasalized speech. Using healthy speech samples, we develop a 4-class deep neural network classifier for the classification of nasal consonants, oral consonants, nasalized vowels, and oral vowels. We use the classifier to compute nasalization scores for clinical speech samples and show that the resulting scores correlate with clinical perception of hypernasality. The proposed approach is evaluated on the speech samples of speakers with dysarthria and cleft lip and palate speakers.
Vikram C. M., Kathy Chapman, Julie M. Liss, Nancy Scherer, Visar Berisha
ICASSP3
2020 Robust Estimation of Hypernasality in Dysarthria With Acoustic Model Likelihood Features
abstract
Hypernasality is a common characteristic symptom across many motor-speech disorders. For voiced sounds, hypernasality introduces an additional resonance in the lower frequencies and, for unvoiced sounds, there is reduced articulatory precision due to air escaping through the nasal cavity. However, the acoustic manifestation of these symptoms is highly variable, making hypernasality estimation very challenging, both for human specialists and automated systems. Previous work in this area relies on either engineered features based on statistical signal processing or machine learning models trained on clinical ratings. Engineered features often fail to capture the complex acoustic patterns associated with hypernasality, whereas metrics based on machine learning are prone to overfitting to the small disease-specific speech datasets on which they are trained. Here we propose a new set of acoustic features that capture these complementary dimensions. The features are based on two acoustic models trained on a large corpus of healthy speech. The first acoustic model aims to measure nasal resonance from voiced sounds, whereas the second acoustic model aims to measure articulatory imprecision from unvoiced sounds. To demonstrate that the features derived from these acoustic models are specific to hypernasal speech, we evaluate them across different dysarthria corpora. Our results show that the features generalize even when training on hypernasal speech from one disease and evaluating on hypernasal speech from another disease (e.g., training on Parkinson's disease, evaluation on Huntington's disease), and when training on neurologically disordered speech but evaluating on cleft palate speech.
Michael Saxon, Ayush Tripathi, Yishan Jiao, Julie M. Liss, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Objective Assessment of Vocal Tremor
abstract
Detecting early signs of neurodegeneration is vital for planning treatments for neurological diseases. Speech plays an important role in this context because it has been shown to be a promising early indicator of neurological decline, and because it can be acquired remotely without the need for specialized hardware. Typically, symptoms are characterized by clinicians using subjective and discrete scales. The poor resolution and subjectivity of these scales can make the earliest speech changes hard to detect. In this paper, we propose an algorithm for the objective assessment of vocal tremor, a phenomenon associated with many neurological disorders. The algorithm extracts and aggregates a feature set from the average spectra of the energy and fundamental frequency profiles of a sustained phonation. We show that the resultant low-dimensional feature set reliably classifies healthy controls and patients with amyotrophic lateral sclerosis perceptually rated for tremor by speech language pathologists.
Jacob Peplinski, Visar Berisha, Julie M. Liss, Shira Hahn, Jeremy Shefner, Seward B. Rutkove, Kristin Qi, Kerisa Shelton
ICASSP3
2019 Objective Measures of Plosive Nasalization in Hypernasal Speech
abstract
Hypernasal speech is a common symptom across several neurological disorders; however it has a variable acoustic signature, making it difficult to quantify acoustically or perceptually. In this paper, we propose the nasal cognate distinctiveness features as an objective proxy for hypernasal speech. Our method is motivated by the observation that incomplete velopharyngeal closure changes the acoustics of the resultant speech such that alveolar stops /t/ and /d/ map to the alveolar nasal /n/ and bilabial stops /b/ and /p/ map to bilabial nasal /m/. We propose a new family of features based on likelihood ratios between the plosives and their respective nasal cognates. These features are based on an acoustic model that is trained only on healthy speech, and evaluated on a set of 75 speakers diagnosed with different dysarthria subtypes and exhibiting varying levels of hypernasality. Our results show that the family of features compares favorably with the clinical perception of speech-language pathologists subjectively evaluating hypernasality.
Michael Saxon, Julie M. Liss, Visar Berisha
ICASSP2
2019 Investigating the Effects of Word Substitution Errors on Sentence Embeddings
abstract
A key initial step in several natural language processing (NLP) tasks involves embedding phrases of text to vectors of real numbers that preserve semantic meaning. To that end, several methods have been recently proposed with impressive results on semantic similarity tasks. However, all of these approaches assume that perfect transcripts are available when generating the embeddings. While this is a reasonable assumption for analysis of written text, it is limiting for analysis of transcribed text. In this paper we investigate the effects of word substitution errors, such as those coming from automatic speech recognition errors (ASR), on several state-of-the-art sentence embedding methods. To do this, we propose a new simulator that allows the experimenter to induce ASR-plausible word substitution errors in a corpus at a desired word error rate. We use this simulator to evaluate the robustness of several sentence embedding methods. Our results show that pre-trained neural sentence encoders are both robust to ASR errors and perform well on textual similarity tasks after errors are introduced. Meanwhile, unweighted averages of word vectors perform well with perfect transcriptions, but their performance degrades rapidly on textual similarity tasks for text with word substitution errors.
Rohit Voleti, Julie M. Liss, Visar Berisha
ICASSP2
2019 Objective Assessment of Social Skills Using Automated Language Analysis for Identification of Schizophrenia and Bipolar Disorder
abstract
Several studies have shown that speech and language features, automatically extracted from clinical interviews or spontaneous discourse, have diagnostic value for mental disorders such as schizophrenia and bipolar disorder. They typically make use of a large feature set to train a classifier for distinguishing between two groups of interest, i.e. a clinical and control group. However, a purely data-driven approach runs the risk of overfitting to a particular data set, especially when sample sizes are limited. Here, we first down-select the set of language features to a small subset that is related to a well-validated test of functional ability, the Social Skills Performance Assessment (SSPA). This helps establish the concurrent validity of the selected features. We use only these features to train a simple classifier to distinguish between groups of interest. Linear regression reveals that a subset of language features can effectively model the SSPA, with a correlation coefficient of 0.75. Furthermore, the same feature set can be used to build a strong binary classifier to distinguish between healthy controls and a clinical group (AUC = 0.96) and also between patients within the clinical group with schizophrenia and bipolar I disorder (AUC = 0.83).
Rohit Voleti, Stephanie Woolridge, Julie M. Liss, Melissa Milanovic, Christopher R. Bowie, Visar Berisha
INTERSPEECH3
2018 Simulating Dysarthric Speech for Training Data Augmentation in Clinical Speech Applications
abstract
Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is prohibitively expensive because of privacy concerns or lack of access. As a result, clinical speech applications typically rely on small data sets with only tens of speakers. In this paper, we propose a method for simulating training data for clinical applications by transforming healthy speech to dysarthric speech using adversarial training. We evaluate the efficacy of our approach using both objective and subjective criteria. We present the transformed samples to five experienced speech-language pathologists (SLPs) and ask them to identify the samples as healthy or dysarthric. The results reveal that the SLPs identify the transformed speech as dysarthric 65% of the time. In a pilot classification experiment, we show that by using the simulated speech samples to balance an existing dataset, the classification accuracy improves by ~10% after data augmentation.
Yishan Jiao, Visar Berisha, Julie M. Liss
ICASSP4
2018 Investigating the Role of L1 in Automatic Pronunciation Evaluation of L2 Speech
abstract
Automatic pronunciation evaluation plays an important role in pronunciation training and second language education. This field draws heavily on concepts from automatic speech recognition (ASR) to quantify how close the pronunciation of non-native speech is to native-like pronunciation. However, it is known that the formation of accent is related to pronunciation patterns of both the target language (L2) and the speaker's first language (L1). In this paper, we propose to use two native speech acoustic models, one trained on L2 speech and the other trained on L1 speech. We develop two sets of measurements that can be extracted from two acoustic models given accented speech. A new utterance-level feature extraction scheme is used to convert these measurements into a fixed-dimension vector which is used as an input to a statistical model to predict the accentedness of a speaker. On a data set consisting of speakers from 4 different L1 backgrounds, we show that the proposed system yields improved correlation with human evaluators compared to systems only using the L2 acoustic model.
Anna Grabek, Julie M. Liss, Visar Berisha
INTERSPEECH3
2017 Interpretable phonological features for clinical applications
abstract
Instrumental analysis of speech sometimes complements subjective evaluations in speech and language therapy; however, apart from elemental speech features such as pitch and formant statistics, higher dimensional spectral features are rarely used in practice because they are clinically uninterpretable. While these features are likely to somehow be related to clinical intervention, this relationship remains to be determined. This paper uses artificial recurrent neural networks to map high-dimensional spectral features into phonological features that are easily interpretable and provide fine-resolution information regarding articulation quality. The evaluation on a dysarthric speech data set shows strong correlation between the phonological feature measures and perceptual ratings. To increase clinical utility, we provide a new way to visualize phonological disturbances that provides clinicians with actionable information about intervention strategies.
Yishan Jiao, Visar Berisha, Julie M. Liss
ICASSP3
2017 Objective assessment of pathological speech using distribution regression
abstract
Objective assessment of pathological speech is an important part of existing systems for automatic diagnosis and treatment of various speech disorders. In this paper, we propose a new regression method for this application. Rather than treating speech samples from each speaker as individual data instances, we treat each speaker's data as a probability distribution. We propose a simple non-parametric learning method to make predictions for out-of-sample speakers based on a probability distance measure to the speakers in the training set. This is in contrast to traditional learning methods that rely on Euclidean distances between individual instances. We evaluate the method on two pathological speech data sets with promising results.
Visar Berisha, Julie M. Liss
ICASSP3
2017 Float Like a Butterfly Sting Like a Bee: Changes in Speech Preceded Parkinsonism Diagnosis for Muhammad Ali
Visar Berisha, Julie M. Liss, Timothy Huston, Alan Wisler, Yishan Jiao, Jonathan Eig
INTERSPEECH2
2017 Interpretable Objective Assessment of Dysarthric Speech Based on Deep Neural Networks
Visar Berisha, Julie M. Liss
INTERSPEECH3
2017 Articulation Entropy: An Unsupervised Measure of Articulatory Precision
abstract
Articulatory precision is a critical factor that influences speaker intelligibility. In this letter, we propose a new measure we call “articulation entropy” that serves as a proxy for the number of distinct phonemes a person produces when he or she speaks. The method is based on the observation that the ability of a speaker to achieve an articulatory target, and hence clearly produce distinct phonemes, is related to the variation of the distribution of speech features that capture articulation-the larger the variation, the larger the number of distinct phonemes produced. In contrast to previous work, the proposed method is completely unsupervised, does not require phonetic segmentation or formant estimation, and can be estimated directly from continuous speech. We evaluate the performance of this measure with several experiments on two data sets: a database of English speakers with various neurological disorders and a database of Mandarin speakers with Parkinson's disease. The results reveal that our measure correlates with subjective evaluation of articulatory precision and reveals differences between healthy individuals and individuals with neurological impairment.
Yishan Jiao, Visar Berisha, Julie M. Liss, Sih-Chiao Hsu, Erika Levy, Megan McAuliffe
IEEE Signal Process. Lett.3
2016 Online speaking rate estimation using recurrent neural networks
abstract
A reliable online speaking rate estimation tool is useful in many domains, including speech recognition, speech therapy intervention, speaker identification, etc. This paper proposes an online speaking rate estimation model based on recurrent neural networks (RNNs). Speaking rate is a long-term feature of speech, which depends on how many syllables were spoken over an extended time window (seconds). We posit that since RNNs can capture long-term dependencies through the memory of previous hidden states, they are a good match for the speaking rate estimation task. Here we train a long short-term memory (LSTM) RNN on a set of speech features that are known to correlate with speech rhythm. An evaluation on spontaneous speech shows that the method yields a higher correlation between the estimated rate and the ground-truth rate when compared to the state-of-the-art alternatives. The evaluation on longitudinal pathological speech shows that the proposed method can capture long-term and short-term changes in speaking rate.
Yishan Jiao, Visar Berisha, Julie M. Liss
ICASSP4
2016 Accent Identification by Combining Deep Neural Networks and Recurrent Neural Networks Trained on Long and Short Term Features
Yishan Jiao, Visar Berisha, Julie M. Liss
INTERSPEECH4
2015 Hilbert spectral analysis of vowels using intrinsic mode functions
abstract
In recent work, we presented mathematical theory and algorithms for time-frequency analysis of non-stationary signals. In that work, we generalized the definition of the Hilbert spectrum by using a superposition of complex AM-FM components parameterized by the Instantaneous Amplitude (IA) and Instantaneous Frequency (IF). Using our Hilbert Spectral Analysis (HSA) approach, the IA and IF estimates can be far more accurate at revealing underlying signal structure than prior approaches to time-frequency analysis. In this paper, we have applied HSA to speech and compared to both narrowband and wideband spectrograms. We demonstrate how the AM-FM components, assumed to be intrinsic mode functions, align well with the energy concentrations of the spectrograms and highlight fine structure present in the Hilbert spectrum. As an example, we show never before seen intra-glottal pulse phenomena that are not readily apparent in other analyses. Such fine-scale analyses may have application in speech-based medical diagnosis and automatic speech recognition (ASR) for pathological speakers.
Steven Sandoval, Phillip L. De Leon, Julie M. Liss
ASRU3
2015 Removing data with noisy responses in regression analysis
abstract
In regression analysis, outliers in the data can induce a bias in the learned function, resulting in larger errors. In this paper we derive an empirically estimable bound on the regression error based on a Euclidean minimum spanning tree generated from the data. Using this bound as motivation, we propose an iterative approach to remove data with noisy responses from the training set. We evaluate the performance of the algorithm on experiments with real-world pathological speech (speech from individuals with neurogenic disorders). Comparative results show that removing noisy examples during training using the proposed approach yields better predictive performance on out-of- sample data.
Alan Wisler, Visar Berisha, Karthikeyan Natesan Ramamurthy, Andreas Spanias, Julie M. Liss
ICASSP5
2015 Convex Weighting Criteria for Speaking Rate Estimation
abstract
Speaking rate estimation directly from the speech waveform is a long-standing problem in speech signal processing. In this paper, we pose the speaking rate estimation problem as that of estimating a temporal density function whose integral over a given interval yields the speaking rate within that interval. In contrast to many existing methods, we avoid the more difficult task of detecting individual phonemes within the speech signal and we avoid heuristics such as thresholding the temporal envelope to estimate the number of vowels. Rather, the proposed method aims to learn an optimal weighting function that can be directly applied to time-frequency features in a speech signal to yield a temporal density function. We propose two convex cost functions for learning the weighting functions and an adaptation strategy to customize the approach to a particular speaker using minimal training. The algorithms are evaluated on the TIMIT corpus, on a dysarthric speech corpus, and on the ICSI Switchboard spontaneous speech corpus. Results show that the proposed methods outperform three competing methods on both healthy and dysarthric speech. In addition, for spontaneous speech rate estimation, the result show a high correlation between the estimated speaking rate and ground truth values.
Yishan Jiao, Visar Berisha, Julie M. Liss
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Modeling pathological speech perception from data with similarity labels
abstract
The current state of the art in judging pathological speech intelligibility is subjective assessment performed by trained speech pathologists (SLP). These tests, however, are inconsistent, costly and, oftentimes suffer from poor intra- and inter-judge reliability. As such, consistent, reliable, and perceptually-relevant objective evaluations of pathological speech are critical. Here, we propose a data-driven approach to this problem. We propose new cost functions for examining data from a series of experiments, whereby we ask certified SLPs to rate pathological speech along the perceptual dimensions that contribute to decreased intelligibility. We consider qualitative feedback from SLPs in the form of comparisons similar to statements "Is Speaker A's rhythm more similar to Speaker B or Speaker C?" Data of this form is common in behavioral research, but is different from the traditional data structures expected in supervised (data matrix + class labels) or unsupervised (data matrix) machine learning. The proposed method identifies relevant acoustic features that correlate with the ordinal data collected during the experiment. Using these features, we show that we are able to develop objective measures of the speech signal degradation that correlate well with SLP responses.
Visar Berisha, Julie M. Liss, Steven Sandoval, Rene Utianski, Andreas Spanias
ICASSP2
2014 Domain invariant speech features using a new divergence measure
abstract
Existing speech classification algorithms often perform well when evaluated on training and test data drawn from the same distribution. In practice, however, these distributions are not always the same. In these circumstances, the performance of trained models will likely decrease. In this paper, we discuss an underutilized divergence measure and derive an estimable upper bound on the test error rate that depends on the error rate on the training data and the distance between training and test distributions. Using this bound as motivation, we develop a feature learning algorithm that aims to identify invariant speech features that generalize well to data similar to, but different from, the training set. Comparative results confirm the efficacy of the algorithm on a set of cross-domain speech classification tasks.
Alan Wisler, Visar Berisha, Julie M. Liss, Andreas Spanias
SLT3
2013 Selecting disorder-specific features for speech pathology fingerprinting
abstract
The general aim of this work is to learn a unique statistical signature for the state of a particular speech pathology. We pose this as a speaker identification problem for dysarthric individuals. To that end, we propose a novel algorithm for feature selection that aims to minimize the effects of speaker-specific features (e.g., fundamental frequency) and maximize the effects of pathology-specific features (e.g., vocal tract distortions and speech rhythm). We derive a cost function for optimizing feature selection that simultaneously trades off between these two competing criteria. Furthermore, we develop an efficient algorithm that optimizes this cost function and test the algorithm on a set of 34 dysarthric and 13 healthy speakers. Results show that the proposed method yields a set of features related to the speech disorder and not an individual's speaking style. When compared to other feature-selection algorithms, the proposed approach results in an improvement in a disorder fingerprinting task by selecting features that are specific to the disorder.
Visar Berisha, Steven Sandoval, Rene Utianski, Julie M. Liss, Andreas Spanias
ICASSP4
2013 Towards a clinical tool for automatic intelligibility assessment
abstract
An important, yet under-explored, problem in speech processing is the automatic assessment of intelligibility for pathological speech. In practice, intelligibility assessment is often done through subjective tests administered by speech pathologists; however research has shown that these tests are inconsistent, costly, and exhibit poor reliability. Although some automatic methods for intelligibility assessment for telecommunications exist, research specific to pathological speech has been limited. Here, we propose an algorithm that captures important multi-scale perceptual cues shown to correlate well with intelligibility. Nonlinear classifiers are trained at each time scale and a final intelligibility decision is made using ensemble learning methods from machine learning. Preliminary results indicate a marked improvement in intelligibility assessment over published baseline results.
Visar Berisha, Rene Utianski, Julie M. Liss
ICASSP3