Visar Berisha

dblp:12/2088 · DBLP profile ↗
← Back
77ranked-venue papers
14as first author
22since 2021 · last 2027
0000-0001-8804-8874ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 13 first-author · 15 since 2021Artificial intelligence and machine learning · 34 · 2 first-author · 16 since 2021Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2027 Evaluating the interpretability of clinical speech AI models: Lessons from two user studies
Lingfeng Xu, Visar Berisha, Julie M. Liss
Comput. Speech Lang.2
2025 Cross-lingual Evaluation Of Hypernasality Using Wav2Vec2 Features
abstract
Hypernasality, a speech resonance disorder characterized by excessive nasal airflow, presents challenges in accurate detection across languages. Traditional assessments of hypernasality include perceptual evaluation and nasometry. More recently, objective acoustic measures based on formant analysis and other acoustic features have been proposed as proxies for hypernasality; however, these acoustic measures exhibit considerable variability, especially in multilingual contexts. In this study, we utilize the wav2vec2-large-xlsr-53 model, a cross-lingual speech representation framework, and evaluate hypernasality-related features within its transformer layers. We extracted features from each layer of the model’s transformer and trained a machine learning model to predict hypernasality ratings across three datasets: Americleft (English), New Mexico Cleft Palate Center (English), and All India Institute of Speech and Hearing (Kannada). Our analysis reveals that the 11th and 12th layer contextualized embeddings effectively model hypernasality cross-lingually, demonstrating significant within-language average correlations (0.78 and 0.78) and cross-lingual average correlations (0.68 and 0.67) between predicted and perceptual ratings. These findings suggest that the wav2vec2-large-xlsr-53 model’s intermediate layers effectively capture hypernasality cross-lingually.
Krupaben Kothadia, Vikram C. M., Ajish K. Abraham, Pushpavathi M, S. R. Mahadeva Prasanna, Nancy Scherer, Kathy Chapman, Julie M. Liss, Visar Berisha
ICASSP9
2025 Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
abstract
Existing methods for analyzing linguistic content from picture descriptions for assessment of cognitive-linguistic impairment often overlook the participant’s visual narrative path, which typically requires eye tracking to assess. Spatio-semantic graphs are a useful tool for analyzing this narrative path from transcripts alone, however they are limited by the need for manual tagging of content information units (CIUs). In this paper, we propose an automated approach for estimation of spatio-semantic graphs (via automated extraction of CIUs) from the Cookie Theft picture commonly used in cognitive-linguistic analyses. The method enables the automatic characterization of the visual semantic path during picture description. Experiments demonstrate that the automatic spatio-semantic graphs effectively differentiate between cognitively impaired and unimpaired speakers. Statistical analyses reveal that the features derived by the automated method produce comparable results to the manual method, with even greater group differences between clinical groups of interest. These results highlight the potential of the automated approach for extracting spatio-semantic features in developing clinical speech models for cognitive impairment assessment.
Si-Ioi Ng, Pranav S. Ambadi, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
ICASSP5
2025 The Impact of Decorrelation on Transformer Interpretation Methods: Applications to Clinical Speech AI
abstract
Recent applications of decorrelation methods to the multi-head attention layers and output embeddings of transformer-based models have resulted in improvements in efficiency and accuracy. Despite these advancements, there is a lack of research focused on the influence of decorrelation on transformer interpretation techniques. This study investigated the impact of two decorrelation methods on interpreting the decision-making logic of a Bidirectional Encoder Representations from Transformers (BERT) model. Two metrics, namely Comprehensiveness and Sufficiency, were used to quantify the interpretation quality, while the changes in correlation within each multi-head self-attention layer was statistically analyzed. Results indicate that decorrelating BERT embeddings leads to a sparser distribution of weights in the middle attention layers and a significantly improved interpretation quality. Conversely, decorrelating the attention maps of specific attention layers increases the correlation in the corresponding attention weight matrices, yielding a less marked improvement in interpretation quality and, in some instances, degraded model performance.
Lingfeng Xu, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
ICASSP4
2025 Mitigating Overfitting During Speech Foundation Model Fine-tuning: Applications to Dysarthric Speech Detection
Yan Xiong 0002, Visar Berisha, Julie M. Liss, Chaitali Chakrabarti
INTERSPEECH2
2025 Statistically Valid Post-Deployment Monitoring Should Be Standard for AI-Based Digital Health
abstract
This position paper argues that post-deployment monitoring in clinical AI is underdeveloped and proposes statistically valid and label-efficient testing frameworks as a principled foundation for ensuring reliability and safety in real-world deployment. A recent review found that only 9\% of FDA-registered AI-based healthcare tools include a post-deployment surveillance plan. Existing monitoring approaches are often manual, sporadic, and reactive, making them ill-suited for the dynamic environments in which clinical models operate. We contend that post-deployment monitoring should be grounded in label-efficient and statistically valid testing frameworks, offering a principled alternative to current practices. We use the term "statistically valid" to refer to methods that provide explicit guarantees on error rates (e.g., Type I/II error), enable formal inference under pre-defined assumptions, and support reproducibility—features that align with regulatory requirements. Specifically, we propose that the detection of changes in the data and model performance degradation should be framed as distinct statistical hypothesis testing problems. Grounding monitoring in statistical rigor ensures a reproducible and scientifically sound basis for maintaining the reliability of clinical AI systems. Importantly, it also opens new research directions for the technical community---spanning theory, methods, and tools for statistically principled detection, attribution, and mitigation of post-deployment model failures in real-world settings.
Pavel Dolin, Weizhi Li, Gautam Dasarathy, Visar Berisha
NeurIPS4
2024 How Does Alignment Error Affect Automated Pronunciation Scoring in Children's Speech?
abstract
following cross utterance averaging. Thus, practical comparisons between child speakers should be very comparable across the two methods.
Prad Kadambi, Tristan J. Mahr, Lucas Annear, Henry Nomeland, Julie M. Liss, Katherine C. Hustad, Visar Berisha
INTERSPEECH7
2024 Segmental and Suprasegmental Speech Foundation Models for Classifying Cognitive Risk Factors: Evaluating Out-of-the-Box Performance
abstract
Speech foundation models are remarkably successful in various consumer applications, prompting their extension to clinical use-cases. This is challenged by small clinical datasets, which precludes effective fine-tuning. We tested the efficacy of two models to classify participants by segmental (Wav2Vec2.0) and suprasegmental (Trillsson) speech analysis windows. Analysis at both time scales has shown differences in the context of cognitive decline. Speakers were classified as healthy controls (HC), Amyloid-β+ (Aβ+), mild cognitive impairment (MCI), or dementia. A subset of W2V2 and Trillsson representations showed large effect size between HC and each risk factor. Cross-validation showed W2V2 consistently outperforms Trillsson. Mean macro-F1 of 54.1%, 63.5%, and 72.0% in were found for classifying Aβ+, MCI, and dementia from HC. Repeatability of Trillsson and W2V2 showed intraclass correlations of 0.30 and 0.41. Reliability of such models must be enhanced for clinical speech analysis and longitudinal tracking.
Si-Ioi Ng, Lingfeng Xu, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
INTERSPEECH5
2024 Improving Speech-Based Dysarthria Detection using Multi-task Learning with Gradient Projection
Yan Xiong 0002, Visar Berisha, Julie M. Liss, Chaitali Chakrabarti
INTERSPEECH2
2023 Smoothly Giving up: Robustness for Simple Models
abstract
There is a growing need for models that are interpretable and have reduced energy/computational cost (e.g., in health care analytics and federated learning). Examples of algorithms to train such models include logistic regression and boosting. However, one challenge facing these algorithms is that they provably suffer from label noise; this has been attributed to the joint interaction between oft-used convex loss functions and simpler hypothesis classes, resulting in too much emphasis being placed on outliers. In this work, we use the margin-based $\alpha$-loss, which continuously tunes between canonical convex and quasi-convex losses, to robustly train simple models. We show that the $\alpha$ hyperparameter smoothly introduces non-convexity and offers the benefit of “giving up” on noisy training examples. We also provide results on the Long-Servedio dataset for boosting and a COVID-19 survey dataset for logistic regression, highlighting the efficacy of our approach across multiple relevant domains.
Tyler Sypherd, Nathaniel Stromberg 0001, Richard Nock, Visar Berisha, Lalitha Sankar
AISTATS4
2023 Does Human Speech Follow Benford's Law?
abstract
Researchers have observed that the frequencies of leading digits in many man-made and naturally occurring datasets follow a logarithmic curve, with digits that start with the number 1 accounting for ~ 30% of all numbers in the dataset and digits that start with the number 9 accounting for ~ 5% of all numbers in the dataset. This phenomenon, known as Benford’s Law, is highly repeatable and appears in lists of numbers from electricity bills, stock prices, tax returns, house prices, death rates, lengths of rivers, and naturally occurring images. In this paper we demonstrate that human speech spectra also follow Benford’s Law, on average. That is, when averaged over many speakers, the frequencies of leading digits in speech magnitude spectra follow this distribution, although with some variability at the individual sample level. We use this observation to motivate a new set of features that can be efficiently extracted from speech and demonstrate that these features can be used to classify between human speech and synthetic speech.
Leo Hsu, Visar Berisha
ICASSP2
2023 Decorrelating Language Model Embeddings for Speech-Based Prediction of Cognitive Impairment
abstract
Training robust clinical speech-based models that generalize requires large sample sizes because speech is variable and high-dimensional. Researchers have turned to foundational models, such as the Bidirectional Encoder Representations from Transformers (BERT), to generate lower-dimensional embeddings, and then finetuned the models for a specific down-stream clinical task. While there is empirical evidence that this approach is helpful, a recent study reveals that the embeddings generated by BERT models tend to be highly correlated, which makes the downstream models difficult to fine-tune, particularly in the small sample size regime. In this work, we propose a new regularization scheme to penalize correlated embeddings during fine tuning of BERT and apply the approach to speech-based assessment of cognitive impairment. Compared to existing methods, the proposed method yields lower estimation errors and smaller false alarm rates in a Mini-Mental State Examination (MMSE) score regression task.
Lingfeng Xu, Kimberly D. Mueller, Julie M. Liss, Visar Berisha
ICASSP4
2023 Aligning Speech Enhancement for Improving Downstream Classification Performance
Yan Xiong 0002, Visar Berisha, Chaitali Chakrabarti
INTERSPEECH2
2023 Learning Repeatable Speech Embeddings Using An Intra-class Correlation Regularizer
abstract
A good supervised embedding for a specific machine learning task is only sensitive to changes in the label of interest and is invariant to other confounding factors. We leverage the concept of repeatability from measurement theory to describe this property and propose to use the intra-class correlation coefficient (ICC) to evaluate the repeatability of embeddings. We then propose a novel regularizer, the ICC regularizer, as a complementary component for contrastive losses to guide deep neural networks to produce embeddings with higher repeatability. We use simulated data to explain why the ICC regularizer works better on minimizing the intra-class variance than the contrastive loss alone. We implement the ICC regularizer and apply it to three speech tasks: speaker verification, voice style conversion, and a clinical application for detecting dysphonic voice. The experimental results demonstrate that adding an ICC regularizer can improve the repeatability of learned embeddings compared to only using the contrastive loss; further, these embeddings lead to improved performance in these downstream tasks.
Suren Jayasuriya, Visar Berisha
NeurIPS3
2023 Consonant-Vowel Transition Models Based on Deep Learning for Objective Evaluation of Articulation
abstract
Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the objective articulation measure (OAM), by analyzing the CV transitions segmented around vowel onsets. The OAM is derived based on the posteriors of a convolutional neural network pre-trained to classify between different consonants using CV regions as input. We demonstrate that the OAM is correlated with perceptual measures in a variety of contexts including (a) adult dysarthric speech, (b) the speech of children with cleft lip/palate, and (c) a database of accented English speech from native Mandarin and Spanish speakers.
Vikram C. M., Julie M. Liss, Kathy Chapman, Nancy Scherer, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Robust Vocal Quality Feature Embeddings for Dysphonic Voice Detection
abstract
Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment often fail to generalize outside the training conditions or to other related applications. In this paper, we propose a deep learning framework for generating acoustic feature embeddings sensitive to vocal quality and robust across different corpora. A contrastive loss is combined with a classification loss to train our deep learning model jointly. Data warping methods are used on input voice samples to improve the robustness of our method. Empirical results demonstrate that our method not only achieves high in-corpus and cross-corpus classification accuracy but also generates good embeddings sensitive to voice quality and robust across different corpora. We also compare our results against three baseline methods on clean and three variations of deteriorated in-corpus and cross-corpus datasets and demonstrate that the proposed model consistently outperforms the baseline methods.
Julie M. Liss, Suren Jayasuriya, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Are reported accuracies in the clinical speech machine learning literature overoptimistic?
Visar Berisha, Chelsea Krantsevich, Gabriela Stegmann, Shira Hahn, Julie M. Liss
INTERSPEECH1
2022 Investigating the Impact of Speech Compression on the Acoustics of Dysarthric Speech
abstract
Acoustic analysis plays an important role in the assessment of dysarthria. Out of a public health necessity, telepractice has become increasingly adopted as the modality in which clinical care is given. While there are differences in software among telepractice platforms, they all use some form of speech compression to preserve bandwidth, with the most common algorithm being the Opus codec. Opus has been optimized for compression of speech from the general (mostly healthy) population. As a result, for speech-language pathologists, this begs the question: is the remotely transmitted speech signal a faithful representation of dysarthric speech? Existing high-fidelity audio recordings from 20 speakers of various dysarthria types were encoded at three different bit rates defined within Opus to simulate different internet bandwidth conditions. Acoustic measures of articulation, voice, and prosody were extracted, and mixed-effect models were used to evaluate the impact of bandwidth conditions on the measures. Significant differences in cepstral peak prominence, degree of voice breaks, jitter, vowel space area, pitch, and vowel space area were observed after Opus processing, providing insight into the types of acoustic measures that are susceptible to speech compression algorithms.
Kelvin Tran, Lingfeng Xu, Gabriela Stegmann, Julie M. Liss, Visar Berisha, Rene Utianski
INTERSPEECH5
2022 A label efficient two-sample test
abstract
Two-sample tests evaluate whether two samples are realizations of the same distribution (the null hypothesis) or two different distributions (the alternative hypothesis). We consider a new setting for this problem where sample features are easily measured whereas sample labels are unknown and costly to obtain. Accordingly, we devise a three-stage framework in service of performing an effective two-sample test with only a small number of sample label queries: first, a classifier is trained with samples uniformly labeled to model the posterior probabilities of the labels; second, a novel query scheme dubbed bimodal query is used to query labels of samples from both classes, and last, the classical Friedman-Rafsky (FR) two-sample test is performed on the queried samples. Theoretical analysis and extensive experiments performed on several datasets demonstrate that the proposed test controls the Type I error and has decreased Type II error relative to uniform querying and certainty-based querying. Source code for our algorithms and experimental results is available at https://github.com/wayne0908/Label-Efficient-Two-Sample.
Weizhi Li, Gautam Dasarathy, Karthikeyan Natesan Ramamurthy, Visar Berisha
UAI4
2021 An Attention Model for Hypernasality Prediction in Children with Cleft Palate
abstract
Hypernasality refers to the perception of abnormal nasal resonances in vowels and voiced consonants. Estimation of hypernasality severity from connected speech samples involves learning a mapping between the frame-level features and utterance-level clinical ratings of hypernasality. However, not all speech frames contribute equally to the perception of hypernasality. In this work, we propose an attention-based bidirectional long-short memory (BLSTM) model that directly maps the frame-level features to utterance-level ratings by focusing only on specific speech frames carrying hyper-nasal cues. The models performance is evaluated on the Americleft database containing speech samples of children with cleft palate and clinical ratings of hypernasality. We analyzed the attention weights over broad phonetic categories and found that the model yields results consistent with what is known in the speech science literature. Further, the correlation between the predicted and perceptual rating is found to be significant (r = 0.684, p < 0.001) and better than conventional BLSTMs trained using frame-wise and last-frame approaches.
Vikram C. M., Nancy Scherer, Kathy Chapman, Julie M. Liss, Visar Berisha
ICASSP5
2021 The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation
Vikram C. M., Tristan J. Mahr, Nancy Scherer, Kathy Chapman, Katherine C. Hustad, Julie M. Liss, Visar Berisha
Interspeech7
2021 Restoring Degraded Speech via a Modified Diffusion Model
abstract
There are many deterministic mathematical operations (e.g. compression, clipping, downsampling) that degrade speech quality considerably. In this paper we introduce a neural network architecture, based on a modification of the DiffWave model, that aims to restore the original speech signal. DiffWave, a recently published diffusion-based vocoder, has shown state-of-the-art synthesized speech quality and relatively shorter waveform generation times, with only a small set of parameters. We replace the mel-spectrum upsampler in DiffWave with a deep CNN upsampler, which is trained to alter the degraded speech mel-spectrum to match that of the original speech. The model is trained using the original speech waveform, but conditioned on the degraded speech mel-spectrum. Post-training, only the degraded mel-spectrum is used as input and the model generates an estimate of the original speech. Our model results in improved speech quality (original DiffWave model as baseline) on several different experiments. These include improving the quality of speech degraded by LPC-10 compression, AMR-NB compression, and signal clipping. Compared to the original DiffWave architecture, our scheme achieves better performance on several objective perceptual metrics and in subjective comparisons. Improvements over baseline are further amplified in a out-of-corpus evaluation setting.
Suren Jayasuriya, Visar Berisha
Interspeech3
2020 Regularization via Structural Label Smoothing
abstract
Regularization is an effective way to promote the generalization performance of machine learning models. In this paper, we focus on label smoothing, a form of output distribution regularization that prevents overfitting of a neural network by softening the ground-truth labels in the training data in an attempt to penalize overconfident outputs. Existing approaches typically use cross-validation to impose this smoothing, which is uniform across all training data. In this paper, we show that such label smoothing imposes a quantifiable bias in the Bayes error rate of the training data, with regions of the feature space with high overlap and low marginal likelihood having a lower bias and regions of low overlap and high marginal likelihood having a higher bias. These theoretical results motivate a simple objective function for data-dependent smoothing to mitigate the potential negative consequences of the operation while maintaining its desirable properties as a regularizer. We call this approach Structural Label Smoothing (SLS). We implement SLS and empirically validate on synthetic, Higgs, SVHN, CIFAR-10, and CIFAR-100 datasets. The results confirm our theoretical insights and demonstrate the effectiveness of the proposed method in comparison to traditional label smoothing.
Weizhi Li, Gautam Dasarathy, Visar Berisha
AISTATS3
2020 Deep Learning Based Prediction of Hypernasality for Clinical Applications
abstract
Hypernasality refers to the perception of excessive nasal resonance during the production of oral sounds. Existing methods for automatic assessment of hypernasality from speech are based on machine learning models trained on disordered speech databases rated by speech-language pathologists. However, the performance of such systems critically depends on the availability of hypernasal speech samples and the reliability of clinical ratings. In this paper, we propose a new approach that uses the speech samples from healthy controls to model the acoustic characteristics of nasalized speech. Using healthy speech samples, we develop a 4-class deep neural network classifier for the classification of nasal consonants, oral consonants, nasalized vowels, and oral vowels. We use the classifier to compute nasalization scores for clinical speech samples and show that the resulting scores correlate with clinical perception of hypernasality. The proposed approach is evaluated on the speech samples of speakers with dysarthria and cleft lip and palate speakers.
Vikram C. M., Kathy Chapman, Julie M. Liss, Nancy Scherer, Visar Berisha
ICASSP5
2020 Compressing LSTM Networks with Hierarchical Coarse-Grain Sparsity
Deepak Kadetotad, Jian Meng, Visar Berisha, Chaitali Chakrabarti, Jae-sun Seo
INTERSPEECH3
2020 UncommonVoice: A Crowdsourced Dataset of Dysphonic Speech
Meredith Moore 0001, Piyush Papreja, Michael Saxon, Visar Berisha, Sethuraman Panchanathan
INTERSPEECH4
2020 Finding the Homology of Decision Boundaries with Active Learning
abstract
Accurately and efficiently characterizing the decision boundary of classifiers is important for problems related to model selection and meta-learning. Inspired by topological data analysis, the characterization of decision boundaries using their homology has recently emerged as a general and powerful tool. In this paper, we propose an active learning algorithm to recover the homology of decision boundaries. Our algorithm sequentially and adaptively selects which samples it requires the labels of. We theoretically analyze the proposed framework and show that the query complexity of our active learning algorithm depends naturally on the intrinsic complexity of the underlying manifold. We demonstrate the effectiveness of our framework in selecting best-performing machine learning models for datasets just using their respective homological summaries. Experiments on several standard datasets show the sample complexity improvement in recovering the homology and demonstrate the practical utility of the framework for model selection.
Weizhi Li, Gautam Dasarathy, Karthikeyan Natesan Ramamurthy, Visar Berisha
NeurIPS4
2020 Robust Estimation of Hypernasality in Dysarthria With Acoustic Model Likelihood Features
abstract
Hypernasality is a common characteristic symptom across many motor-speech disorders. For voiced sounds, hypernasality introduces an additional resonance in the lower frequencies and, for unvoiced sounds, there is reduced articulatory precision due to air escaping through the nasal cavity. However, the acoustic manifestation of these symptoms is highly variable, making hypernasality estimation very challenging, both for human specialists and automated systems. Previous work in this area relies on either engineered features based on statistical signal processing or machine learning models trained on clinical ratings. Engineered features often fail to capture the complex acoustic patterns associated with hypernasality, whereas metrics based on machine learning are prone to overfitting to the small disease-specific speech datasets on which they are trained. Here we propose a new set of acoustic features that capture these complementary dimensions. The features are based on two acoustic models trained on a large corpus of healthy speech. The first acoustic model aims to measure nasal resonance from voiced sounds, whereas the second acoustic model aims to measure articulatory imprecision from unvoiced sounds. To demonstrate that the features derived from these acoustic models are specific to hypernasal speech, we evaluate them across different dysarthria corpora. Our results show that the features generalize even when training on hypernasal speech from one disease and evaluating on hypernasal speech from another disease (e.g., training on Parkinson's disease, evaluation on Huntington's disease), and when training on neurologically disordered speech but evaluating on cleft palate speech.
Michael Saxon, Ayush Tripathi, Yishan Jiao, Julie M. Liss, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Objective Assessment of Vocal Tremor
abstract
Detecting early signs of neurodegeneration is vital for planning treatments for neurological diseases. Speech plays an important role in this context because it has been shown to be a promising early indicator of neurological decline, and because it can be acquired remotely without the need for specialized hardware. Typically, symptoms are characterized by clinicians using subjective and discrete scales. The poor resolution and subjectivity of these scales can make the earliest speech changes hard to detect. In this paper, we propose an algorithm for the objective assessment of vocal tremor, a phenomenon associated with many neurological disorders. The algorithm extracts and aggregates a feature set from the average spectra of the energy and fundamental frequency profiles of a sustained phonation. We show that the resultant low-dimensional feature set reliably classifies healthy controls and patients with amyotrophic lateral sclerosis perceptually rated for tremor by speech language pathologists.
Jacob Peplinski, Visar Berisha, Julie M. Liss, Shira Hahn, Jeremy Shefner, Seward B. Rutkove, Kristin Qi, Kerisa Shelton
ICASSP2
2019 Objective Measures of Plosive Nasalization in Hypernasal Speech
abstract
Hypernasal speech is a common symptom across several neurological disorders; however it has a variable acoustic signature, making it difficult to quantify acoustically or perceptually. In this paper, we propose the nasal cognate distinctiveness features as an objective proxy for hypernasal speech. Our method is motivated by the observation that incomplete velopharyngeal closure changes the acoustics of the resultant speech such that alveolar stops /t/ and /d/ map to the alveolar nasal /n/ and bilabial stops /b/ and /p/ map to bilabial nasal /m/. We propose a new family of features based on likelihood ratios between the plosives and their respective nasal cognates. These features are based on an acoustic model that is trained only on healthy speech, and evaluated on a set of 75 speakers diagnosed with different dysarthria subtypes and exhibiting varying levels of hypernasality. Our results show that the family of features compares favorably with the clinical perception of speech-language pathologists subjectively evaluating hypernasality.
Michael Saxon, Julie M. Liss, Visar Berisha
ICASSP3
2019 Joint Optimization of Quantization and Structured Sparsity for Compressed Deep Neural Networks
abstract
The usage of Deep Neural Networks (DNN) on resource-constrained edge devices has been limited due to their high computation and large memory requirement. In this work, we propose an algorithm to compress DNNs by jointly optimizing structured sparsity and quantization constraints in a single DNN training framework. The proposed algorithm has been extensively validated on high/low capacity DNNs and wide/deep sparse DNNs. Further, we perform Pareto-optimal analysis to extract optimal DNN models from a large set of trained DNN models. The optimal structurally-compressed DNN model achieves ~50X weight memory reduction without test accuracy degradation, compared to floating-point uncompressed DNN.
Gaurav Srivastava 0001, Deepak Kadetotad, Shihui Yin, Visar Berisha, Chaitali Chakrabarti, Jae-sun Seo
ICASSP4
2019 Investigating the Effects of Word Substitution Errors on Sentence Embeddings
abstract
A key initial step in several natural language processing (NLP) tasks involves embedding phrases of text to vectors of real numbers that preserve semantic meaning. To that end, several methods have been recently proposed with impressive results on semantic similarity tasks. However, all of these approaches assume that perfect transcripts are available when generating the embeddings. While this is a reasonable assumption for analysis of written text, it is limiting for analysis of transcribed text. In this paper we investigate the effects of word substitution errors, such as those coming from automatic speech recognition errors (ASR), on several state-of-the-art sentence embedding methods. To do this, we propose a new simulator that allows the experimenter to induce ASR-plausible word substitution errors in a corpus at a desired word error rate. We use this simulator to evaluate the robustness of several sentence embedding methods. Our results show that pre-trained neural sentence encoders are both robust to ASR errors and perform well on textual similarity tasks after errors are introduced. Meanwhile, unweighted averages of word vectors perform well with perfect transcriptions, but their performance degrades rapidly on textual similarity tasks for text with word substitution errors.
Rohit Voleti, Julie M. Liss, Visar Berisha
ICASSP3
2019 Do Conversational Partners Entrain on Articulatory Precision?
abstract
The communication phenomenon known as conversational entrainment occurs when dialogue partners align or adapt their behavior to one another while conversing. Associated with rapport, trust, and communicative efficiency, entrainment appears to facilitate conversational success. In this work, we explore how conversational partners entrain or align on articulatory precision or the clarity with which speakers articulate their spoken productions. Articulatory precision also has implications for conversational success as precise articulation can enhance speech understanding and intelligibility. However, in conversational speech, speakers tend to reduce their articulatory precision, preferring low-cost, imprecise speech. Speakers may adapt their articulation and become more precise depending on feedback from their listeners. Given the potential of entrainment, we are interested in how conversational partners adapt or entrain their articulatory precision to one another. We explore this phenomenon in 57 task-based dialogues. Controlling for the influence of speaking rate, we find that speakers entrain on articulatory precision, with significant alignment on articulation of consonants. We discuss the potential applications that speaker alignment on precision might have for modeling conversation and implementing strategies for enhancing communicative success in human-human and human-computer interactions.
Nichola Lubold, Stephanie A. Borrie, Tyson S. Barrett, Megan M. Willi, Visar Berisha
INTERSPEECH5
2019 Say What? A Dataset for Exploring the Error Patterns That Two ASR Engines Make
Meredith Moore 0001, Michael Saxon, Hemanth Venkateswara, Visar Berisha, Sethuraman Panchanathan
INTERSPEECH4
2019 Objective Assessment of Social Skills Using Automated Language Analysis for Identification of Schizophrenia and Bipolar Disorder
abstract
Several studies have shown that speech and language features, automatically extracted from clinical interviews or spontaneous discourse, have diagnostic value for mental disorders such as schizophrenia and bipolar disorder. They typically make use of a large feature set to train a classifier for distinguishing between two groups of interest, i.e. a clinical and control group. However, a purely data-driven approach runs the risk of overfitting to a particular data set, especially when sample sizes are limited. Here, we first down-select the set of language features to a small subset that is related to a well-validated test of functional ability, the Social Skills Performance Assessment (SSPA). This helps establish the concurrent validity of the selected features. We use only these features to train a simple classifier to distinguish between groups of interest. Linear regression reveals that a subset of language features can effectively model the SSPA, with a correlation coefficient of 0.75. Furthermore, the same feature set can be used to build a strong binary classifier to distinguish between healthy controls and a clinical group (AUC = 0.96) and also between patients within the clinical group with schizophrenia and bipolar I disorder (AUC = 0.83).
Rohit Voleti, Stephanie Woolridge, Julie M. Liss, Melissa Milanovic, Christopher R. Bowie, Visar Berisha
INTERSPEECH6
2019 Residual + Capsule Networks (ResCap) for Simultaneous Single-Channel Overlapped Keyword Recognition
Yan Xiong 0002, Visar Berisha, Chaitali Chakrabarti
INTERSPEECH2
2018 Online Machine Learning Experiments in HTML5
abstract
This work in progress paper describes software that enables online machine learning experiments in an undergraduate DSP course. This software operates in HTML5 and embeds several digital signal processing functions. The software can process natural signals such as speech and can extract various features, for machine learning applications. For example in the case of speech processing, LPC coefficients and formant frequencies can be computed. In this paper, we present speech processing, feature extraction and clustering of features using the K-means machine learning algorithm. The primary objective is to provide a machine learning experience to undergraduate students. The functions and simulations described provide a user-friendly visualization of phoneme recognition tasks. These tasks make use of the Levinson-Durbin linear prediction and the K-means machine learning algorithms. The exercise was assigned as a class project in our undergraduate DSP class. The description of the exercise along with assessment results is described.
Abhinav Dixit, Uday Shankar Shanthamallu, Andreas Spanias, Visar Berisha, Mahesh K. Banavar
FIE4
2018 Simulating Dysarthric Speech for Training Data Augmentation in Clinical Speech Applications
abstract
Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is prohibitively expensive because of privacy concerns or lack of access. As a result, clinical speech applications typically rely on small data sets with only tens of speakers. In this paper, we propose a method for simulating training data for clinical applications by transforming healthy speech to dysarthric speech using adversarial training. We evaluate the efficacy of our approach using both objective and subjective criteria. We present the transformed samples to five experienced speech-language pathologists (SLPs) and ask them to identify the samples as healthy or dysarthric. The results reveal that the SLPs identify the transformed speech as dysarthric 65% of the time. In a pilot classification experiment, we show that by using the simulated speech samples to balance an existing dataset, the classification accuracy improves by ~10% after data augmentation.
Yishan Jiao, Visar Berisha, Julie M. Liss
ICASSP3
2018 Towards a Wearable Cough Detector Based on Neural Networks
abstract
Persistent cough is a symptom common to a number of respiratory disorders; however, reliable monitoring of cough frequency and cough severity over an extended period of time can be a challenge. Traditional methods involve subjective evaluation by care providers or patient self-reports. As an alternative, we propose an objective method for monitoring cough using a wearable microphone. We collected 24-hour audio recordings from 9 patients suffering from chronic obstructive pulmonary disease, asthma, and lung cancer using the VitaloJAK wearable microphone. Trained professionals carefully listened to each audio stream and manually labeled each cough event. Using this data, we propose a new neural-network-based cough detection scheme. A pre-processing algorithm is used to estimate the start and end of each cough and the deep neural network is trained using each cough instance. Experiments demonstrate an average leave-one-participant-out cross-validation specificity and sensitivity of 93.7% and 97.6% respectively.
Prad Kadambi, Abinash Mohanty, Jaclyn Smith, Kevin McGuinnes, Kimberly Holt, Armin Furtwaengler, Roberto Slepetys, Jae-sun Seo, Junseok Chae, Yu Cao 0001, Visar Berisha
ICASSP13
2018 Direct Ensemble Estimation of Density Functionals
abstract
Estimating density functionals of analog sources is an important problem in statistical signal processing and information theory. Traditionally, estimating these quantities requires either making parametric assumptions about the underlying distributions or using non-parametric density estimation followed by integration. In this paper we introduce a direct nonparametric approach which bypasses the need for density estimation by using the error rates of k-NN classifiers as “data-driven” basis functions that can be combined to estimate a range of density functionals. However, this method is subject to a non-trivial bias that dramatically slows the rate of convergence in higher dimensions. To overcome this limitation, we develop an ensemble method for estimating the value of the basis function which, under some minor constraints on the smoothness of the underlying distributions, achieves the parametric rate of convergence regardless of data dimension.
Alan Wisler, Kevin R. Moon, Visar Berisha
ICASSP3
2018 Triplet Network with Attention for Speaker Diarization
abstract
In automatic speech processing systems, speaker diarization is a crucial front-end component to separate segments from different speakers.Inspired by the recent success of deep neural networks (DNNs) in semantic inferencing, triplet loss-based architectures have been successfully used for this problem.However, existing work utilizes conventional i-vectors as the input representation and builds simple fully connected networks for metric learning, thus not fully leveraging the modeling power of DNN architectures.This paper investigates the importance of learning effective representations from the sequences directly in metric learning pipelines for speaker diarization.More specifically, we propose to employ attention models to learn embeddings and the metric jointly in an end-to-end fashion.Experiments are conducted on the CALLHOME conversational speech corpus.The diarization results demonstrate that, besides providing a unified model, the proposed approach achieves improved performance when compared against existing approaches.
Huan Song, Megan M. Willi, Jayaraman J. Thiagarajan, Visar Berisha, Andreas Spanias
INTERSPEECH4
2018 Investigating the Role of L1 in Automatic Pronunciation Evaluation of L2 Speech
abstract
Automatic pronunciation evaluation plays an important role in pronunciation training and second language education. This field draws heavily on concepts from automatic speech recognition (ASR) to quantify how close the pronunciation of non-native speech is to native-like pronunciation. However, it is known that the formation of accent is related to pronunciation patterns of both the target language (L2) and the speaker's first language (L1). In this paper, we propose to use two native speech acoustic models, one trained on L2 speech and the other trained on L1 speech. We develop two sets of measurements that can be extracted from two acoustic models given accented speech. A new utterance-level feature extraction scheme is used to convert these measurements into a fixed-dimension vector which is used as an input to a statistical model to predict the accentedness of a speaker. On a data set consisting of speakers from 4 different L1 backgrounds, we show that the proposed system yields improved correlation with human evaluators compared to systems only using the L2 acoustic model.
Anna Grabek, Julie M. Liss, Visar Berisha
INTERSPEECH4
2018 A Discriminative Acoustic-Prosodic Approach for Measuring Local Entrainment
abstract
Acoustic-prosodic entrainment describes the tendency of humans to align or adapt their speech acoustics to each other in conversation. This alignment of spoken behavior has important implications for conversational success. However, modeling the subtle nature of entrainment in spoken dialogue continues to pose a challenge. In this paper, we propose a straightforward definition for local entrainment in the speech domain and operationalize an algorithm based on this: acoustic-prosodic features that capture entrainment should be maximally different between real conversations involving two partners and sham conversations generated by randomly mixing the speaking turns from the original two conversational partners. We propose an approach for measuring local entrainment that quantifies alignment of behavior on a turn-by-turn basis, projecting the differences between interlocutors' acoustic-prosodic features for a given turn onto a discriminative feature subspace that maximizes the difference between real and sham conversations. We evaluate the method using the derived features to drive a classifier aiming to predict an objective measure of conversational success (i.e., low versus high), on a corpus of task-oriented conversations. The proposed entrainment approach achieves 72% classification accuracy using a Naive Bayes classifier, outperforming three previously established approaches evaluated on the same conversational corpus.
Megan M. Willi, Stephanie A. Borrie, Tyson S. Barrett, Visar Berisha
INTERSPEECH5
2017 Interpretable phonological features for clinical applications
abstract
Instrumental analysis of speech sometimes complements subjective evaluations in speech and language therapy; however, apart from elemental speech features such as pitch and formant statistics, higher dimensional spectral features are rarely used in practice because they are clinically uninterpretable. While these features are likely to somehow be related to clinical intervention, this relationship remains to be determined. This paper uses artificial recurrent neural networks to map high-dimensional spectral features into phonological features that are easily interpretable and provide fine-resolution information regarding articulation quality. The evaluation on a dysarthric speech data set shows strong correlation between the phonological feature measures and perceptual ratings. To increase clinical utility, we provide a new way to visualize phonological disturbances that provides clinicians with actionable information about intervention strategies.
Yishan Jiao, Visar Berisha, Julie M. Liss
ICASSP2
2017 Objective assessment of pathological speech using distribution regression
abstract
Objective assessment of pathological speech is an important part of existing systems for automatic diagnosis and treatment of various speech disorders. In this paper, we propose a new regression method for this application. Rather than treating speech samples from each speaker as individual data instances, we treat each speaker's data as a probability distribution. We propose a simple non-parametric learning method to make predictions for out-of-sample speakers based on a probability distance measure to the speakers in the training set. This is in contrast to traditional learning methods that rely on Euclidean distances between individual instances. We evaluate the method on two pathological speech data sets with promising results.
Visar Berisha, Julie M. Liss
ICASSP2
2017 Float Like a Butterfly Sting Like a Bee: Changes in Speech Preceded Parkinsonism Diagnosis for Muhammad Ali
Visar Berisha, Julie M. Liss, Timothy Huston, Alan Wisler, Yishan Jiao, Jonathan Eig
INTERSPEECH1
2017 Interpretable Objective Assessment of Dysarthric Speech Based on Deep Neural Networks
Visar Berisha, Julie M. Liss
INTERSPEECH2
2017 Improving efficiency in sparse learning with the feedforward inhibitory motif
Steven Skorheim, Visar Berisha, Shimeng Yu, Jae-sun Seo, Maxim Bazhenov, Yu Cao 0001
Neurocomputing4
2017 Articulation Entropy: An Unsupervised Measure of Articulatory Precision
abstract
Articulatory precision is a critical factor that influences speaker intelligibility. In this letter, we propose a new measure we call “articulation entropy” that serves as a proxy for the number of distinct phonemes a person produces when he or she speaks. The method is based on the observation that the ability of a speaker to achieve an articulatory target, and hence clearly produce distinct phonemes, is related to the variation of the distribution of speech features that capture articulation-the larger the variation, the larger the number of distinct phonemes produced. In contrast to previous work, the proposed method is completely unsupervised, does not require phonetic segmentation or formant estimation, and can be estimated directly from continuous speech. We evaluate the performance of this measure with several experiments on two data sets: a database of English speakers with various neurological disorders and a database of Mandarin speakers with Parkinson's disease. The results reveal that our measure correlates with subjective evaluation of articulatory precision and reveals differences between healthy individuals and individuals with neurological impairment.
Yishan Jiao, Visar Berisha, Julie M. Liss, Sih-Chiao Hsu, Erika Levy, Megan McAuliffe
IEEE Signal Process. Lett.2
2016 Online speaking rate estimation using recurrent neural networks
abstract
A reliable online speaking rate estimation tool is useful in many domains, including speech recognition, speech therapy intervention, speaker identification, etc. This paper proposes an online speaking rate estimation model based on recurrent neural networks (RNNs). Speaking rate is a long-term feature of speech, which depends on how many syllables were spoken over an extended time window (seconds). We posit that since RNNs can capture long-term dependencies through the memory of previous hidden states, they are a good match for the speaking rate estimation task. Here we train a long short-term memory (LSTM) RNN on a set of speech features that are known to correlate with speech rhythm. An evaluation on spontaneous speech shows that the method yields a higher correlation between the estimated rate and the ground-truth rate when compared to the state-of-the-art alternatives. The evaluation on longitudinal pathological speech shows that the proposed method can capture long-term and short-term changes in speaking rate.
Yishan Jiao, Visar Berisha, Julie M. Liss
ICASSP3
2016 Ranking the parameters of deep neural networks using the fisher information
abstract
The large number of parameters in deep neural networks (DNNs) often makes them prohibitive for low-power devices, such as field-programmable gate arrays (FPGA). In this paper, we propose a method to determine the relative importance of all network parameters by measuring the amount of information that the network output carries about each of the parameters - the Fisher Information. Based on the importance ranking, we design a complexity reduction scheme that discards unimportant parameters and assigns more quantization bits to more important parameters. For evaluation, we construct a deep autoencoder and learn a non-linear dimensionality reduction scheme for accelerometer data measuring the gait of individuals with Parkinson's disease. Experimental results confirm that the proposed ranking method can help reduce the complexity of the network with minimal impact on performance.
Visar Berisha, Martin Woolf, Jae-sun Seo, Yu Cao 0001
ICASSP2
2016 Empirically-estimable multi-class classification bounds
abstract
In this paper, we extend previously developed non-parametric bounds on the Bayes risk in binary classification problems to multi-class problems. In comparison with the well-known Bhattacharyya bound which is typically calculated by employing parametric assumptions, the bounds proposed in this paper are directly estimable from data, provably tighter, and more robust to different types of data. We verify the tightness and validity of this bound using an illustrative synthetic example, and further demonstrate its value by incorporating it into a feature selection algorithm which we apply to the real-world problem of distinguishing between different neuro-motor disorders based on sentence-level speech data.
Alan Wisler, Visar Berisha, Dennis Wei, Karthikeyan Natesan Ramamurthy, Andreas Spanias
ICASSP2
2016 Accent Identification by Combining Deep Neural Networks and Recurrent Neural Networks Trained on Long and Short Term Features
Yishan Jiao, Visar Berisha, Julie M. Liss
INTERSPEECH3
2016 A Convex Model for Linguistic Influence in Group Conversations
Kan Kawabata, Visar Berisha, Anna Scaglione, Amy LaCross
INTERSPEECH2
2015 Removing data with noisy responses in regression analysis
abstract
In regression analysis, outliers in the data can induce a bias in the learned function, resulting in larger errors. In this paper we derive an empirically estimable bound on the regression error based on a Euclidean minimum spanning tree generated from the data. Using this bound as motivation, we propose an iterative approach to remove data with noisy responses from the training set. We evaluate the performance of the algorithm on experiments with real-world pathological speech (speech from individuals with neurogenic disorders). Comparative results show that removing noisy examples during training using the proposed approach yields better predictive performance on out-of- sample data.
Alan Wisler, Visar Berisha, Karthikeyan Natesan Ramamurthy, Andreas Spanias, Julie M. Liss
ICASSP2
2015 Active data labeling for improved classifier generalizability
Visar Berisha, Douglas Cochran
Signal Process.1
2015 Empirical Non-Parametric Estimation of the Fisher Information
abstract
The Fisher information matrix (FIM) is a foundational concept in statistical signal processing. The FIM depends on the probability distribution, assumed to belong to a smooth parametric family. Traditional approaches to estimating the FIM require estimating the probability distribution function (PDF), or its parameters, along with its gradient or Hessian. However, in many practical situations the PDF of the data is not known but the statistician has access to an observation sample for any parameter value. Here we propose a method of estimating the FIM directly from sampled data that does not require knowledge of the underlying PDF. The method is based on non-parametric estimation of an f-divergence over a local neighborhood of the parameter space and a relation between curvature of the f-divergence and the FIM. Thus we obtain an empirical estimator of the FIM that does not require density estimation and is asymptotically consistent. We empirically evaluate the validity of our approach using two experiments.
Visar Berisha, Alfred O. Hero III
IEEE Signal Process. Lett.1
2015 Convex Weighting Criteria for Speaking Rate Estimation
abstract
Speaking rate estimation directly from the speech waveform is a long-standing problem in speech signal processing. In this paper, we pose the speaking rate estimation problem as that of estimating a temporal density function whose integral over a given interval yields the speaking rate within that interval. In contrast to many existing methods, we avoid the more difficult task of detecting individual phonemes within the speech signal and we avoid heuristics such as thresholding the temporal envelope to estimate the number of vowels. Rather, the proposed method aims to learn an optimal weighting function that can be directly applied to time-frequency features in a speech signal to yield a temporal density function. We propose two convex cost functions for learning the weighting functions and an adaptation strategy to customize the approach to a particular speaker using minimal training. The algorithms are evaluated on the TIMIT corpus, on a dysarthric speech corpus, and on the ICSI Switchboard spontaneous speech corpus. Results show that the proposed methods outperform three competing methods on both healthy and dysarthric speech. In addition, for spontaneous speech rate estimation, the result show a high correlation between the estimated speaking rate and ground truth values.
Yishan Jiao, Visar Berisha, Julie M. Liss
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Modeling pathological speech perception from data with similarity labels
abstract
The current state of the art in judging pathological speech intelligibility is subjective assessment performed by trained speech pathologists (SLP). These tests, however, are inconsistent, costly and, oftentimes suffer from poor intra- and inter-judge reliability. As such, consistent, reliable, and perceptually-relevant objective evaluations of pathological speech are critical. Here, we propose a data-driven approach to this problem. We propose new cost functions for examining data from a series of experiments, whereby we ask certified SLPs to rate pathological speech along the perceptual dimensions that contribute to decreased intelligibility. We consider qualitative feedback from SLPs in the form of comparisons similar to statements "Is Speaker A's rhythm more similar to Speaker B or Speaker C?" Data of this form is common in behavioral research, but is different from the traditional data structures expected in supervised (data matrix + class labels) or unsupervised (data matrix) machine learning. The proposed method identifies relevant acoustic features that correlate with the ordinal data collected during the experiment. Using these features, we show that we are able to develop objective measures of the speech signal degradation that correlate well with SLP responses.
Visar Berisha, Julie M. Liss, Steven Sandoval, Rene Utianski, Andreas Spanias
ICASSP1
2014 Domain invariant speech features using a new divergence measure
abstract
Existing speech classification algorithms often perform well when evaluated on training and test data drawn from the same distribution. In practice, however, these distributions are not always the same. In these circumstances, the performance of trained models will likely decrease. In this paper, we discuss an underutilized divergence measure and derive an estimable upper bound on the test error rate that depends on the error rate on the training data and the distance between training and test distributions. Using this bound as motivation, we develop a feature learning algorithm that aims to identify invariant speech features that generalize well to data similar to, but different from, the training set. Comparative results confirm the efficacy of the algorithm on a set of cross-domain speech classification tasks.
Alan Wisler, Visar Berisha, Julie M. Liss, Andreas Spanias
SLT2
2013 Selecting disorder-specific features for speech pathology fingerprinting
abstract
The general aim of this work is to learn a unique statistical signature for the state of a particular speech pathology. We pose this as a speaker identification problem for dysarthric individuals. To that end, we propose a novel algorithm for feature selection that aims to minimize the effects of speaker-specific features (e.g., fundamental frequency) and maximize the effects of pathology-specific features (e.g., vocal tract distortions and speech rhythm). We derive a cost function for optimizing feature selection that simultaneously trades off between these two competing criteria. Furthermore, we develop an efficient algorithm that optimizes this cost function and test the algorithm on a set of 34 dysarthric and 13 healthy speakers. Results show that the proposed method yields a set of features related to the speech disorder and not an individual's speaking style. When compared to other feature-selection algorithms, the proposed approach results in an improvement in a disorder fingerprinting task by selecting features that are specific to the disorder.
Visar Berisha, Steven Sandoval, Rene Utianski, Julie M. Liss, Andreas Spanias
ICASSP1
2013 Towards a clinical tool for automatic intelligibility assessment
abstract
An important, yet under-explored, problem in speech processing is the automatic assessment of intelligibility for pathological speech. In practice, intelligibility assessment is often done through subjective tests administered by speech pathologists; however research has shown that these tests are inconsistent, costly, and exhibit poor reliability. Although some automatic methods for intelligibility assessment for telecommunications exist, research specific to pathological speech has been limited. Here, we propose an algorithm that captures important multi-scale perceptual cues shown to correlate well with intelligibility. Nonlinear classifiers are trained at each time scale and a final intelligibility decision is made using ensemble learning methods from machine learning. Preliminary results indicate a marked improvement in intelligibility assessment over published baseline results.
Visar Berisha, Rene Utianski, Julie M. Liss
ICASSP1
2010 An auditory-domain based speech enhancement algorithm
abstract
Typically, speech enhancement algorithms minimize a suitable error criterion in the spectral or time domain. Although the error criterions have included perceptual properties such as masking thresholds, non-uniform frequency resolution and sensitivity of the auditory system, these are only done heuristically and the error criterion does not explicitly include an auditory model in their formulation. In this paper, we propose an auditory-domain based speech enhancement algorithm that minimizes the distortion between the auditory representation of the estimated and desired signal. Simulation results indicate that the proposed algorithm performs effectively under different noise conditions and also results in a lower average loudness error.
Harish Krishnamoorthi, Andreas Spanias, Visar Berisha, Homin Kwon, Harvey D. Thornburg
ICASSP3
2009 Low-complexity sinusoidal component selection using loudness patterns
abstract
Sinusoidal modeling of audio at low-bit rates involves selecting a limited number of parameters according to a quantitative or perceptual criterion. Most perceptual sinusoidal component selection strategies are computationally intensive and not suitable for real-time applications. In this paper, a computationally efficient sinusoidal selection algorithm based on a novel hybrid loudness estimation scheme is presented. The hybrid scheme first estimates efficiently the loudness of a multi-tone signal from the loudness patterns of its constituent sinusoidal components. Then it refines this estimate by performing a full evaluation of loudness but only in select critical bands. Experimental results show that the proposed technique maintains a low perceptual sinusoidal synthesis error at a much lower computational complexity.
Harish Krishnamoorthi, Visar Berisha, Andreas Spanias, Homin Kwon
ICASSP2
2009 Energy-constrained discriminant analysis
abstract
Dimensionality reduction algorithms have become an indispensable tool for working with high-dimensional data in classification. Linear discriminant analysis (LDA) is a popular analysis technique used to project high-dimensional data into a lower-dimensional space while maximizing class separability. Although this technique is widely used in many applications, it suffers from overfitting when the number of training examples is on the same order as the dimension of the original data space. When overfitting occurs, the direction of the LDA solution can be dominated by low-energy noise and therefore the solution becomes non-robust to unseen data. In this paper, we propose a novel algorithm, energy-constrained discriminant analysis (ECDA), that overcomes the limitations of LDA by finding lower dimensional projections that maximize inter-class separability, while also preserving signal energy. Our results show that the proposed technique results in higher classification rates when compared to comparable methods. The results are given in terms of SAR image classification, however the algorithm is broadly applicable and can be generalized to any classification problem.
Scott Philips, Visar Berisha, Andreas Spanias
ICASSP2
2009 A Sensor Network for Real-time Acoustic Scene Analysis
abstract
Acoustic scene analysis can be used to extract relevant information in applications such as homeland security, surveillance and environmental monitoring. Wireless sensor networks have been of particular interest in monitoring acoustic scenes. Sensors embedded in such a network typically operate under several constraints such as low power and limited bandwidth. In this paper, we consider resource-efficient acoustic sensing tasks that extract and transmit relevant information to a central station where information assessment can be conducted. We propose a series of acoustic scene analysis tasks that are performed in a hierarchical manner. Hierarchical tasks include sound and speech discrimination, estimation of the number of speakers from the acquired sound, gender and emotional state, and ultimately voice monitoring and key word spotting. We apply support vector machine and Gaussian mixture model algorithms on sound features. A real-time implementation is accomplished using crossbow motes interfaced with a TI DSP board. A series of experiments are presented to characterize the performance of the algorithms under different conditions.
Homin Kwon, Harish Krishnamoorthi, Visar Berisha, Andreas Spanias
ISCAS3
2009 A Frequency/Detector Pruning Approach for Loudness Estimation
abstract
In this letter, we propose a frequency and detector pruning approach for reducing the computational complexity associated with loudness estimation. The frequency pruning approach exploits the principles of psychoacoustics such that the total neural activity is preserved. The detector pruning approach evaluates the excitation/loudness patterns at nonuniform sample locations and employs signal interpolation techniques to obtain their corresponding high resolution estimates. Comparative results with the Moore and Glasberg loudness estimation process reveal that the proposed pruning approach for loudness estimation performs consistently well for different types of audio signals with a significant reduction in the computational complexity.
Harish Krishnamoorthi, Andreas Spanias, Visar Berisha
IEEE Signal Process. Lett.3
2008 A low-complexity loudness estimation algorithm
abstract
Audio processing applications such as rate determination, bandwidth extension, compression, and noise reduction make use of loudness metrics. Most loudness estimation algorithms are computationally expensive and often not suitable for real time applications. In this paper, we present a low-complexity loudness estimation algorithm applicable to both steady and time-varying sounds. The model computes an estimate of the excitation pattern by simultaneously pruning the frequency components and detector locations. Comparative results indicate that the proposed algorithm performs consistently well for different types of audio signals at a reduced complexity.
Harish Krishnamoorthi, Visar Berisha, Andreas Spanias
ICASSP2
2008 Gradient projection-based channel equalization under sustained fading
Venkatraman Atti, Andreas Spanias, Kostas Tsakalis, Constantinos Panayiotou, Leonidas D. Iasemidis, Visar Berisha
Signal Process.6
2007 A Scalable Bandwidth Extension Algorithm
abstract
Most modern bandwidth extension techniques predict the high- frequency band based on features extracted from the lower band. While this works for some frames, problems arise when the correlation between the low and the high band is insufficient. In these situations, additional high-band information must be sent to the decoder. In this paper, we propose a scalable speech coding method based on the principles of bandwidth extension. The rate selection is based on explicit psychoacoustic criteria, while the bandwidth extension is performed using a constrained MMSE estimation technique. Objective and subjective evaluations indicate that the proposed system performs at a lower average bit rate when compared to other similar algorithms while improving speech quality.
Visar Berisha, Andreas Spanias
ICASSP (4)1
2007 Sparse Manifold Learning with Applications to SAR Image Classification
abstract
Nonlinear data-driven dimensionality reduction techniques have recently gained popularity due to the emergence of high dimensional data sets. The algorithmic complexity and storage requirements of these techniques, however, can make them prohibitive in resource-limited applications. It is therefore beneficial to reduce the number of exemplar samples required for performing an out-of-sample extension to a test point. In this paper, we propose a novel method for selecting a minimal set of exemplars and performing the out-of-sample extension. In the case of two-class target recognition with synthetic aperture radar (SAR) data, we compare the efficacy of the proposed approach with other approaches for selecting a subset of the available training samples. We show that the proposed algorithm outperforms the existing methods by providing low-dimensional embeddings that maintain interclass separability using fewer retained exemplars.
Visar Berisha, Nitesh Shah, Donald E. Waagen, Harry Schmitt, Salvatore Bellofiore, Andreas Spanias, Douglas Cochran
ICASSP (3)1
2007 Dual-Mode Wideband Speech Compression
abstract
Many bandwidth extension techniques attempt to predict the high-band frequencies based on features extracted from the lower band. Recent work suggests that such methods are limiting because the correlation between the low band and the high band is insufficient for adequate representation. As a result, additional high-band information must be sent to the decoder. In this paper, we propose a dual mode wideband speech coding algorithm based on the principles of bandwidth extension. The principal contributions include a mode selection algorithm based on greedy algorithm that maximizes the loudness criteria, and a bandwidth extension algorithm based on a constrained MMSE estimator. Results reveal that the proposed system improves the quality of narrowband speech while performing at a lower bit rate.
Visar Berisha, Andreas Spanias
MMSP1
2006 Real-Time Collaborative Monitoring in Wireless Sensor Networks
abstract
In recent years, wireless sensor networks (WSN) have shown success in distributed real-time signal processing systems. In collaborative signal processing environments, each sensor is responsible for extracting pertinent information from the surrounding environment and transmitting it to other sensors and/or to the main processing station. Often times, the sensors operate under a number of constraints, such as limited processing power and low bandwidth. In this paper we propose a collaborative signal processing framework that is implemented in an acoustic monitoring scenario. A low-complexity voice activity detector and a gender classifier are implemented on the Crossbow sensor motes. A series of experiments are presented that characterize the performance of the algorithms under varying SNR conditions and in different environments.
Visar Berisha, Homin Kwon, Andreas Spanias
ICASSP (3)1
2006 Real-time acoustic monitoring using wireless sensor motes
abstract
Wireless sensor networks (WSN) have recently gained popularity in distributed monitoring and surveillance applications. The objective of these devices is to extract pertinent information under several constrains such as low computational capabilities, limited arithmetic precision, and the need to conserve power. One of the most revealing environmental cues is audio. In this paper, we propose a voice activity detector and a simple gender classifier for use in a distributed acoustic sensing system. This algorithm makes use of low-complexity audio features and a pre-trained regression tree to classify incoming speech by gender. The algorithm is implemented real-time on the Crossbow sensor motes and a series of results are given that characterize the algorithm performance and complexity. Challenges in this real-time implementation include designing the algorithm and software architecture such that the signal processing is appropriately distributed between the sensor mote and the base station. At the base station, a data fusion algorithm considers a linear combination of individual mote decisions to form a final decision.
Visar Berisha, Homin Kwon, Andreas Spanias
ISCAS1
2006 Bandwidth Extension of Audio Based on Partial Loudness Criteria
abstract
Most modern speech coders operate on a limited bandwidth. This tends to decrease the naturalness of the synthesized audio and often also affects the intelligibility of certain sounds. While a few wideband speech coders have been standardized, implementing them in existing systems would require significant changes to the infrastructure. One solution is to use bandwidth extension techniques that predict the high-frequency band based on low-band features. Problems arise however when the correlation between the low and the high band is insufficient for an adequate representation of the wideband signal. In this paper, we propose a novel source-filter bandwidth extension algorithm that makes use of psychoacoustic concepts to determine the perceptual benefits that a particular audio frame gains from a more exact representation of the high band. Preliminary results indicate that the proposed system performs at a lower average bit rate when compared to other similar algorithms without compromising the audio quality
Visar Berisha, Andreas Spanias
MMSP1
2005 Interactive Java modules for the MPEG-1 psychoacoustic model [audio coding teaching applications]
abstract
This paper presents a collection of interactive Java modules for the purpose of introducing undergraduate DSP students to perceptual audio coding principles. This effort is part of a combined research and curriculum program funded by NSF that aims towards exposing undergraduate students to advanced concepts and research in signal processing. A computer laboratory with several supporting exercises and Java functions has been developed for use in our undergraduate DSP course. This exercise along with the accompanying Java software was assigned and assessed in the Summer of 2004 and will be reassessed in the Fall of 2004. Results of this assessment along with student comments are presented at the end of the paper.
Andreas Spanias, Venkatraman Atti, Visar Berisha
ICASSP (5)4
2005 Enhancing the Quality of Coded Audio Using Perceptual Criteria
abstract
Code excited linear predictive (CELP) coding standards often fail to properly represent non-speech signals because they are inherently optimized for speech. Most modern CELP coders include provisions for the inclusion of indirect perceptual criteria to counteract this problem; however no direct psychoacoustic models are employed. In this paper, we present a pre- and postprocessor for the vocoder that makes use of the MPEG-1 psychoacoustic model 1 in order to enhance the quality of the coded audio. A novel frequency-domain technique is proposed that attempts to shape the residual of a vocoder such that it falls below psychoacoustic thresholds
Visar Berisha, Andreas Spanias
MMSP1