Shri Narayanan

dblp:19/3899 · also Shrikanth Narayanan, Shrikanth S. Narayanan, Shrikanth Shri Narayanan · DBLP profile ↗
← Back
729ranked-venue papers
20as first author
107since 2021 · last 2027
0000-0002-1052-6204ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 564 · 14 first-author · 68 since 2021Artificial intelligence and machine learning · 415 · 10 first-author · 53 since 2021Human-computer interaction and ubiquitous computing · 37 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2027 Exploring contrastive alignment across conversational turns for modeling vocal entrainment in interactions involving children with autism
Rimita Lahiri, So Hyun Kim, Somer Bishop, Catherine Lord, Helen Tager-Flusberg, Shri Narayanan
Comput. Speech Lang.6
2026 The Subjectivity of Respect in Police Traffic Stops: Modeling Community Perspectives in Body-Worn Camera Footage
abstract
Preni Golazizian, Elnaz Rahmati, Jackson Trager, Zhivar Sourati, Nona Ghazizadeh, Georgios Chochlakis, Jose J. Alcocer, Kerby Bennett, Aarya Vijay Devnani, Parsa Hejabi, Harry G. Muttram, Akshay Kiran Padte, Mehrshad Saadatinia, Chenhao Wu, Alireza Salkhordeh Ziabari, Michael Sierra-Arévalo, Nicholas Weller, Shrikanth Narayanan, Benjamin A.t. Graham, Morteza Dehghani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Preni Golazizian, Elnaz Rahmati, Jackson Trager, Zhivar Sourati, Nona Ghazizadeh, Georgios Chochlakis, Jose Alcocer, Kerby Bennett, Aarya Vijay Devnani, Parsa Hejabi, Harry G. Muttram, Akshay Kiran Padte, Mehrshad Saadatinia, Alireza S. Ziabari, Michael Sierra-Arévalo, Nicholas Weller, Shri Narayanan, Benjamin A. T. Graham, Morteza Dehghani
ACL (1)18
2026 RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification
abstract
Respiratory diseases remain a leading cause of global mortality, where timely and accurate diagnosis is critical to improving patient outcomes and reducing healthcare burdens.While prior work has explored audio-based models for respiratory disease detection, such unimodal approaches often suffer from limited generalizability and diagnostic precision.In this paper, we propose RespiraMFM, a Multimodal Foundation Model that integrates respiratory sounds with patient medical history and symptoms to enhance diagnostic accuracy and disease detection capabilities.We introduce an effective contrastive alignment strategy for audio-text multimodal integration, allowing the model to learn better cross-modal representations between respiratory sounds and corresponding textual clinical information.We evaluate RespiraMFM across five major respiratory diseases using seven real-world datasets in both supervised fine-tuning and zero-shot settings, achieving a 9.15% improvement in AU-ROC on supervised tasks and a 20.98% gain on zero-shot tasks over existing baselines.These findings underscore the potential of our framework to advance early diagnosis and improve clinical decision-making in respiratory disease management.
Shakhrul Iman Siam, Tiantian Feng, Shri Narayanan, Mi Zhang 0002
ACL (1)4
2026 Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
abstract
We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the RAIL license at https://github.com/tiantiaf0627/voxlect.
Tiantian Feng, Anfeng Xu, Xuan Shi, Thanathai Lertpetchpun, Yoonjeong Lee, Dani Byrd, Shri Narayanan
KDD (1)9
2026 Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model
Xuan Shi, Tiantian Feng, Jay Park, Christina Hagedorn, Louis Goldstein, Shri Narayanan
Comput. Speech Lang.6
2025 Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
abstract
Modeling complex subjective tasks in Natural Language Processing, such as recognizing emotion and morality, is considerably challenging due to significant variation in human annotations.This variation often reflects reasonable differences in semantic interpretations rather than mere noise, necessitating methods to distinguish between legitimate subjectivity and error.We address this challenge by exploring label verification in these contexts using Large Language Models (LLMs).First, we propose a simple In-Context Learning binary filtering baseline that estimates the reasonableness of a document-label pair.We then introduce the Label-in-a-Haystack setting: the query and its label(s) are included in the demonstrations shown to LLMs, which are prompted to predict the label(s) again, while receiving task-specific instructions (e.g., emotion recognition) rather than label copying.We show how the failure to copy the label(s) to the output of the LLM are task-relevant and informative.Building on this, we propose the Label-in-a-Haystack Rectification (LiaHR) framework for subjective label correction: when the model outputs diverge from the reference gold labels, we assign the generated labels to the example instead of discarding it.This approach can be integrated into annotation pipelines to enhance signal-to-noise ratios.Comprehensive analyses, human evaluations, and ecological validity studies verify the utility of LiaHR for label correction.
Georgios Chochlakis, Peter Wu, Arjun Bedi, Marcus Ma, Kristina Lerman, Shri Narayanan
EMNLP6
2025 Large Language Models Do Multi-Label Classification Differently
abstract
Multi-label classification is prevalent in realworld settings, but the behavior of Large Language Models (LLMs) in this setting is understudied.We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the models at each label generation step.We find that the initial probability distribution for the first label often does not reflect the eventual final output, even in terms of relative order and find LLMs tend to suppress all but one label at each generation step.We further observe that as model scale increases, their token distributions exhibit lower entropy and higher single-label confidence, but the internal relative ranking of the labels improves.Finetuning methods such as supervised finetuning and reinforcement learning amplify this phenomenon.We introduce the task of distribution alignment for multi-label settings: aligning LLM-derived label distributions with empirical distributions estimated from annotator responses in subjective tasks.We propose both zero-shot and supervised methods which improve both alignment and predictive performance over existing approaches.We find one method -taking the max probability over all label generation distributions instead of just using the initial probability distribution -improves both distribution alignment and overall F1 classification without adding any additional computation.
Marcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason, Shri Narayanan
EMNLP5
2025 Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
abstract
In-Context Learning (ICL) in Large Language Models (LLM) has emerged as the dominant technique for performing natural language tasks, as it does not require updating the model parameters with gradient-based methods. ICL promises to "adapt" the LLM to perform the present task at a competitive or state-of-the-art level at a fraction of the computational cost. ICL can be augmented by incorporating the reasoning process to arrive at the final label explicitly in the prompt, a technique called Chain-of-Thought (CoT) prompting. However, recent work has found that ICL relies mostly on the retrieval of task priors and less so on "learning" to perform tasks, especially for complex subjective domains like emotion and morality, where priors ossify posterior predictions. In this work, we examine whether "enabling" reasoning also creates the same behavior in LLMs, wherein the format of CoT retrieves reasoning priors that remain relatively unchanged despite the evidence in the prompt. We find that, surprisingly, CoT indeed suffers from the same posterior collapse as ICL for larger language models. Code is available at https://github.com/gchochla/cot-priors.
Georgios Chochlakis, Niyantha Maruthu Pandiyan, Kristina Lerman, Shri Narayanan
ICASSP4
2025 Wavelet Scattering Network Features for Intensity Category Classification and Prediction of SPL from Speech
abstract
Speakers change vocal intensity in daily life to communicate over long distances and to express vocal emotions. Humans produce speech using different intensity categories (e.g. soft, normal and loud voice) and they can regulate intensity across a wide sound pressure level (SPL) range. Knowing the intensity category or the SPL of speech is beneficial in speech-based biomarking of health. Recent studies have explored the vocal intensity category classification and prediction of SPL from speech, which has been recorded without SPL calibration information and is presented on an arbitrary amplitude scale. Using speech signals in such scenario, this study investigates the wavelet scattering network (WSN) features in two tasks: (1) classification of speech into four intensity categories (soft, normal, loud, very loud) (multi-class classification task) and (2) prediction of SPL (regression task). In the former task, the WSN features showed absolute accuracy improvements of 4-14% compared to reference features. For the latter task, the WSN features improved the prediction of SPL by an average of 1-2 dB compared to the reference features.
Manila Kodali, Sudarsana Reddy Kadiri, Shri Narayanan, Paavo Alku
ICASSP3
2025 Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
abstract
Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We propose a novel approach to enhance listened speech decoding from electroencephalography (EEG) signals by utilizing an auxiliary phoneme predictor that simultaneously decodes textual phoneme sequences. The proposed model architecture consists of three main parts: EEG module, speech module, and phoneme predictor. The EEG module learns to properly represent EEG signals into EEG embeddings. The speech module generates speech waveforms from the EEG embeddings. The phoneme predictor outputs the decoded phoneme sequences in text modality. Our proposed approach allows users to obtain decoded listened speech from EEG signals in both modalities (speech waveforms and textual phoneme sequences) simultaneously, eliminating the need for a concatenated sequential pipeline for each modality. The proposed approach also outperforms previous methods in both modalities. The source code and speech samples are publicly available1.
Tiantian Feng, Aditya Kommineni, Sudarsana Reddy Kadiri, Shri Narayanan
ICASSP5
2025 Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
abstract
Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using an speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate. The source code is available for the public at: https://github.com/Hong7Cong/SPAN-rtmri.git
Hong Nguyen, Sean Foley, Xuan Shi, Tiantian Feng, Shri Narayanan
ICASSP6
2025 Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
abstract
Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies "who spoke when", is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available.
Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Catherine Lord, Shri Narayanan
ICASSP5
2025 Scaling Wearable Foundation Models
abstract
Wearable sensors have become ubiquitous thanks to a variety of health tracking features. The resulting continuous and longitudinal measurements from everyday life generate large volumes of data. However, making sense of these observations for scientific and actionable insights is non-trivial. Inspired by the empirical success of generative modeling, where large neural networks learn powerful representations from vast amounts of text, image, video, or audio data, we investigate the scaling properties of wearable sensor foundation models across compute, data, and model size. Using a dataset of up to 40 million hours of in-situ heart rate, heart rate variability, accelerometer, electrodermal activity, skin temperature, and altimeter per-minute data from over 165,000 people, we create LSM, a multimodal foundation model built on the largest wearable-signals dataset with the most extensive range of sensor modalities to date. Our results establish the scaling laws of LSM for tasks such as imputation, interpolation and extrapolation across both time and sensor modalities. Moreover, we highlight how LSM enables sample-efficient downstream learning for tasks including exercise and activity recognition.
Girish Narayanswamy, Xin Liu 0034, Kumar Ayush, Yuzhe Yang 0003, Xuhai Xu, Shun Liao, Jake Garrison, Shyam A. Tailor, Jacob E. Sunshine, Yun Liu 0013, Tim Althoff, Shri Narayanan, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Samy Abdel-Ghaffar, Daniel McDuff
ICLR12
2025 Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
Tiantian Feng, Thanathai Lertpetchpun, Dani Byrd, Shri Narayanan
INTERSPEECH4
2025 Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
Tiantian Feng, Anfeng Xu, Xuan Shi, Somer Bishop, Shri Narayanan
INTERSPEECH5
2025 On the Relationship between Accent Strength and Articulatory Features
Sean Foley, Yoonjeong Lee, Dani Byrd, Shri Narayanan
INTERSPEECH6
2025 Can Multimodal Foundation Models Help Analyze Child-Inclusive Autism Diagnostic Videos?
Aditya Kommineni, Digbalay Bose, Tiantian Feng, So Hyun Kim, Helen Tager-Flusberg, Somer Bishop, Catherine Lord, Sudarsana Reddy Kadiri, Shri Narayanan
INTERSPEECH9
2025 Articulatory Feature Prediction from Surface EMG during Speech Production
Kleanthis Avramidis, Simon Pistrosch, Monica González Machorro, Yoonjeong Lee, Björn W. Schuller, Louis Goldstein, Shri Narayanan
INTERSPEECH9
2025 Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shri Narayanan
INTERSPEECH4
2025 Examining Test-Time Adaptation for Personalized Child Speech Recognition
Zhonghao Shi, Xuan Shi, Anfeng Xu, Tiantian Feng, Harshvardhan Srivastava, Shri Narayanan, Maja J. Mataric
INTERSPEECH6
2025 75-Speaker Annot-16: A benchmark dataset for speech articulatory rt-MRI annotation with articulator contours and phonetic alignment
Xuan Shi, Yubin Zhang, Yijing Lu, Marcus Ma, Tiantian Feng, Asterios Toutios, Haley Hsu, Louis Goldstein, Shri Narayanan
INTERSPEECH9
2025 Large Language Models based ASR Error Correction for Child Conversations
Anfeng Xu, Tiantian Feng, So Hyun Kim, Somer Bishop, Catherine Lord, Shri Narayanan
INTERSPEECH6
2025 Co-registration of real-time MRI and respiration for speech research
Yubin Zhang, Prakash Kumar, Xuan Shi, Haley Hsu, Shri Narayanan, Krishna S. Nayak, Louis Goldstein
INTERSPEECH9
2025 Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
abstract
Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shrikanth Narayanan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shri Narayanan
NAACL (Long Papers)4
2025 ModalityMirror: Enhancing Audio Classification in Modality Heterogeneity Federated Learning via Multimodal Distillation
abstract
Multimodal Federated Learning frequently encounters challenges of client modality heterogeneity, leading to undesired performances for secondary modality in multimodal learning. It is particularly prevalent in audiovisual learning, with audio is often assumed to be the weaker modality in recognition tasks. To address this challenge, we introduce ModalityMirror to improve audio model performance by leveraging knowledge distillation from an audiovisual federated learning model. ModalityMirror involves two phases: a modality-wise FL stage to aggregate unimodal encoders; and a federated knowledge distillation stage on multimodality clients to train a unimodal student model. Our results demonstrate that ModalityMirror significantly improves the audio classification compared to the state-of-the-art FL methods such as Harmony, particularly in audiovisual FL facing video missing. Our approach unlocks the potential for exploiting the diverse modality spectrum inherent in multimodal FL.
Tiantian Feng, Amir Salman Avestimehr, Shri Narayanan
NOSSDAV4
2025 Automatic classification of vocal intensity categories from amplitude-normalized speech signals by comparing acoustic features and classifier models
abstract
Regulation of vocal intensity is a fundamental phenomenon in speech communication. Speakers use different intensity categories (e.g., soft, normal, and loud voice) to generate different vocal emotions or to communicate in noisy conditions or over varying distances. Vocal intensity categories have been studied in fundamental research of speech, but much less is known about their automatic classification. This study investigates the classification of vocal intensity categories from speech signals in a scenario, where the original level information of speech is absent and the signal is presented on a normalized amplitude scale. Different acoustic features were studied together with machine learning (ML) and deep learning (DL) classifiers using two different labeling approaches. Speech signals recorded from 50 speakers reciting sentences in four intensity categories (soft, normal, loud, and very loud) were analyzed. Altogether 15 feature sets including different cepstral, spectral and handcrafted (eGeMAPS) features were compared. Three ML classifiers (support vector machine, random forest and AdaBoost), and four DL classifiers (deep neural network, convolutional neural network, recurrent neural network and bidirectional long short-term memory network) were compared. The best classification accuracy of 86.0% was obtained by combining the best performing cepstral and spectral features and using the bidirectional long short-term memory classifier. • Multi-class classification of vocal intensity categories is studied. • The classification is studied using amplitude-normalized speech signals. • Various labelling approaches, features and classifiers are compared. • DL models outperformed ML models, with BiLSTM achieving the best performance.
Manila Kodali, Luna Ansari, Sudarsana Reddy Kadiri, Shri Narayanan, Paavo Alku
Speech Commun.4
2025 Do all features matter? Layer-wise feature probing of self-supervised speech models for dysarthria severity classification
Paban Sapkota, Harsh Srivastava, Hemant Kumar Kathania, Shri Narayanan, Sudarsana Reddy Kadiri
Speech Commun.4
2025 Can Layer-Wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
abstract
Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL) models have greatly enhanced the transcription of adult speech, accurately transcribing children's speech remains a significant challenge. This study investigates the effectiveness of layer-wise features extracted from state-of-the-art SSL pre-trained models - specifically, Wav2Vec2, HuBERT, Data2Vec, and WavLM in improving the performance of ASR for children's speech in zero-shot scenarios. A detailed analysis of features extracted from these models was conducted, integrating them into a simplified DNN-based ASR system using the Kaldi toolkit. The analysis identified the most effective layers for enhancing ASR performance on children's speech in a zero-shot scenario, where WSJCAM0 adult speech was used for training and PFSTAR children speech for testing. Experimental results indicated that Layer 22 of the Wav2Vec2 model achieved the lowest Word Error Rate (WER) of 5.15%, representing a 51.64% relative improvement over the direct zero-shot decoding using Wav2Vec2 (WER of 10.65%). Additionally, age group-wise analysis demonstrated consistent performance improvements with increasing age, along with significant gains observed even in younger age groups using the SSL features. Further experiments on the CMU Kids dataset confirmed similar trends, highlighting the generalizability of the proposed approach.
Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shri Narayanan
IEEE Signal Process. Lett.4
2024 The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
abstract
In-context Learning (ICL) has emerged as a power-ful paradigm for performing natural language tasks with Large Language Models (LLM) without updating the models' parameters, in contrast to the traditional gradient-based finetuning. The promise of ICL is that the LLM can adapt to perform the present task at a competitive or state-of-the-art level at a fraction of the cost. The ability of LLMs to perform tasks in this few-shot manner relies on their background knowledge of the task (or task priors). However, recent work has found that, unlike traditional learning, LLMs are unable to fully integrate information from demonstrations that contrast task priors. This can lead to performance saturation at suboptimal levels, especially for subjective tasks such as emotion recognition, where the mapping from text to emotions can differ widely due to variability in human annotations. In this work, we design experiments and propose measurements to explicitly quantify the consistency of proxies of LLM priors and their pull on the posteriors. We show that LLMs have strong yet inconsistent priors in emotion recognition that ossify their predictions. We also find that the larger the model, the stronger these effects become. Our results suggest that caution is needed when using ICL with larger LLMs for affect-centered tasks outside their pretraining domain and when interpreting ICL results.1
Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shri Narayanan
ACII4
2024 RobuSER: A Robustness Benchmark for Speech Emotion Recognition
abstract
The recent surge in deep learning has improved Speech Emotion Recognition (SER) model performance; however, ensuring robustness across diverse scenarios beyond the training dataset remains a problem. This challenge becomes pronounced in real-world situations characterized by noisy conditions, where model adaptability to unclean data is crucial. Despite ongoing efforts to develop noise-robust models, the lack of standardized evaluation protocols hampers fair comparisons among different models. This paper tackles this issue by introducing Robuser, a benchmarking procedure designed specifically for evaluating the robustness of SER models under noise. Robuser is a comprehensive open-source benchmark that can be applied to any speech dataset, focusing on diverse corruption types in two pivotal dimensions: additive background noise and various signal distortion corruptions, each in varying levels of severity. Furthermore, through the evaluation of a state-of-the-art SER model against this benchmark, we offer quantitative insights into the impact of the different corruption types and severity levels on performance. The baseline model reveals a notable performance degradation of up to 22.77% in Unweighted Accuracy (UA) and 20.32% in Weighted Accuracy (WA) on corrupted IEMOCAP, underscoring the substantial room for improvement in this domain. Our code is openly available at the following URL: https://github.com/BehavioralSignalTechnologies/ser_robustness.git
Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Sofia Eleftheriou, Petros Mitseas, Dimitris Sgouropoulos, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan
ACII9
2024 Emotion-Aware Speech Popularity Prediction: A Use-Case on TED Talks
abstract
In the context of the ever-growing influence of social media, understanding and predicting the popularity of content has become crucial for creators and marketers alike. Our research addresses this need by introducing a method to forecast the success of oral presentations, focusing on the nuanced use of paralinguistic features and insights derived from speech emotion recognition models. This innovative approach is designed to enhance verbal communication skills by providing public speakers with targeted feedback. We leverage a dataset of 2,462 TED talk videos, complete with metadata such as user comments, tags, and views, to establish a set of four objective metrics for determining presentation popularity. These metrics form the foundation of our analysis, enabling us to evaluate the efficacy of our predictive methodology. By integrating audio-based emotional cues with text-based content analysis we showcase the capability of the proposed speech analytics system to capture user assessments of presentation quality. This research highlights the role of emotional expression in speech as a component of content's appeal, advocating for a broader analytical perspective beyond just text-only analysis. It suggests new directions for improving the impact of public speaking and calls for further investigation into multimodal content analysis, aiming to deepen our understanding of audience engagement on social media and content delivery platforms.
Dimitris Sgouropoulos, Petros Mitseas, Sofia Eleftheriou, Theodoros Giannakopoulos, Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Athanasios Katsamanis, Shri Narayanan
ACII9
2024 Character Attribute Extraction from Movie Scripts Using LLMs
abstract
Narrative understanding is an integrative task of studying characters, plots, events, and relations in a story. It involves natural language processing tasks such as named entity recognition and coreference resolution to identify the characters, semantic role labeling and argument mining to find character actions and events, information extraction and question answering to describe character attributes, causal analysis to relate different events, and summarization to find the main storyline. In this work, we aim to formally operationalize the task of character attribute extraction, motivated by analyzing inclusive character representations and portrayals. We focus on a mix of static and dynamic attribute types that require varying context sizes for their accurate retrieval. We use automated screenplay parsing, entity recognition, and external knowledge bases to collect character descriptions from movie scripts, and explore different prompting strategies (zero-shot, few-shot, and chain-of-thought) to leverage large language models for attribute extraction.1
Sabyasachee Baruah, Shri Narayanan
ICASSP2
2024 TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion Recognition
abstract
Recent studies have explored using pre-trained embeddings for speech emotion recognition, achieving comparable performance to conventional methods that rely on low-level knowledge-inspired acoustic features. These embeddings are often generated from models trained on large-scale speech datasets using self-supervised or weakly-supervised learning objectives. Despite the significant advancements made in SER through pre-trained embeddings, there is a limited understanding of the trustworthiness of these methods, including privacy breaches, unfair performance, vulnerability to adversarial attacks, and computational cost, all of which may hinder the real-world deployment of these systems. In response, we introduce TrustSER, a general framework designed to evaluate the trustworthiness of SER systems using deep learning methods, focusing on privacy, safety, fairness, and sustainability, offering unique insights into future research in the field of SER. Our code is publicly available under: https://github.com/usc-sail/trust-ser.
Tiantian Feng, Rajat Hebbar, Shri Narayanan
ICASSP3
2024 Foundation Model Assisted Automatic Speech Emotion Recognition: Transcribing, Annotating, and Augmenting
abstract
Significant advances are being made in speech emotion recognition (SER) using deep learning models. Nonetheless, training SER systems remains challenging, requiring both time and costly resources. Like many other machine learning tasks, acquiring datasets for SER requires substantial data annotation efforts, including transcription and labeling. These annotation processes present challenges when attempting to scale up conventional SER systems. Recent developments in foundational models have had a tremendous impact, giving rise to applications such as ChatGPT. These models have enhanced human-computer interactions including bringing unique possibilities for streamlining data collection in fields like SER. In this research, we explore the use of foundational models to assist in automating SER from transcription and annotation to augmentation. Our study demonstrates that these models can generate transcriptions to enhance the performance of SER systems that rely solely on speech data. Furthermore, we note that annotating emotions from transcribed speech remains a challenging task. However, combining outputs from multiple LLMs enhances the quality of annotations. Lastly, our findings suggest the feasibility of augmenting existing speech emotion datasets by annotating unlabeled speech samples.
Tiantian Feng, Shri Narayanan
ICASSP2
2024 Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization
abstract
Video summarization remains a huge challenge in computer vision due to the size of the input videos to be summarized. We propose an efficient, language-only video summarizer that achieves competitive accuracy with high data efficiency. Using only textual captions obtained via a zero-shot approach, we train a language transformer model and forego image representations. This method allows us to perform filtration amongst the representative text vectors and condense the sequence. With our approach, we gain explainability with natural language that comes easily for human interpretation and textual summaries of the videos. An ablation study that focuses on modality and data compression shows that leveraging text modality only effectively reduces input data processing while retaining comparable results.
Yoonsoo Nam, Adam Lehavi, Daniel Yang, Digbalay Bose, Swabha Swayamdipta, Shri Narayanan
ICASSP6
2024 Emotion-Aligned Contrastive Learning Between Images and Music
abstract
Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using queries of various modalities including text, video, and speech. While most approaches aim to match general music semantics to the input queries, only a few focus on affective qualities. In this work, we address the task of retrieving emotionally-relevant music from image queries by learning an affective alignment between images and music audio. Our approach focuses on learning an emotion-aligned joint embedding space between images and music. This embedding space is learned via emotion-supervised contrastive learning, using an adapted cross-modal version of the SupCon loss. We evaluate the joint embeddings through cross-modal retrieval tasks (image-to-music and music-to-image) based on emotion labels. Furthermore, we investigate the generalizability of the learned music embeddings via automatic music tagging. Our experiments show that the proposed approach successfully aligns images and music, and that the learned embedding space is effective for cross-modal retrieval applications.
Shanti Stewart, Kleanthis Avramidis, Tiantian Feng, Shri Narayanan
ICASSP4
2024 Audio-Visual Child-Adult Speaker Classification in Dyadic Interactions
abstract
Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer scale and robustness across diverse and wide-ranging conditions. Identifying the speech segments belonging to the child is a critical step in such modeling. Conventional child-adult speaker classification typically relies on audio modeling approaches, overlooking visual signals that convey speech articulation information, such as lip motion. Building on the foundation of an audio-only child-adult speaker classification pipeline, we propose incorporating visual cues through active speaker detection and visual processing models. Our framework involves video preprocessing, utterance-level child-adult speaker detection, and late fusion of modality-specific predictions. We demonstrate from extensive experiments that a visually aided classification pipeline enhances the accuracy and robustness of the classification. We show relative improvements of 2.38% and 3.97% in F1 macro score when one face and two faces are visible, respectively.
Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Shri Narayanan
ICASSP5
2024 Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?
abstract
Multi-modal learning has emerged as an increasingly promising avenue in vision recognition, driving innovations across diverse domains. Despite its success, the robustness of multi-modal learning for visual recognition is often challenged by the unavailability of a subset of modalities, especially the visual modality. Conventional approaches to mitigate missing modalities in multi-modal learning rely heavily on modality fusion schemes. In contrast, this paper explores the use of text-to-image models to assist multi-modal learning. Specifically, we propose and explore a simple but effective multi-modal learning framework GTI-MM to enhance the data efficiency and model robustness against missing visual modality by imputing the missing data with generative models. Using multiple multi-modal datasets with visual recognition tasks, we present a comprehensive analysis of diverse conditions involving missing visual data. Our findings show that synthetic images benefit training data efficiency with missing visual data during training and improve model robustness with visual data missing during both training and testing. Moreover, we demonstrate GTI-MM is effective with lower generation quantity and simple prompt techniques. Our code base and synthetic images are at https://github.com/usc-sail/GTI-MM.
Tiantian Feng, Daniel Yang, Digbalay Bose, Shri Narayanan
ICMI4
2024 Socio-Linguistic Characteristics of Coordinated Inauthentic Accounts
abstract
Online manipulation is a pressing concern for democracies, but the actions and strategies of coordinated inauthentic accounts, which have been used to interfere in elections, are not well understood. We analyze a five million-tweet multilingual dataset related to the 2017 French presidential election, when a major information campaign led by Russia called "#MacronLeaks" took place. We utilize heuristics to identify coordinated inauthentic accounts and detect attitudes, concerns and emotions within their tweets, collectively known as socio-linguistic characteristics. We find that coordinated accounts retweet other coordinated accounts far more than expected by chance, while being exceptionally active just before the second round of voting. Concurrently, socio-linguistic characteristics reveal that coordinated accounts share tweets promoting a candidate at three times the rate of non-coordinated accounts. Coordinated account tactics also varied in time to reflect news events and rounds of voting. Our analysis highlights the utility of socio-linguistic characteristics to inform researchers about tactics of coordinated accounts and how these may feed into online social manipulation.
Keith Burghardt, Ashwin Rao, Georgios Chochlakis, Sabyasachee Baruah, Siyi Guo, Andrew Rojecki, Shri Narayanan, Kristina Lerman
ICWSM8
2024 CVAT-BWV: A Web-Based Video Annotation Platform for Police Body-Worn Video
Parsa Hejabi, Akshay Kiran Padte, Preni Golazizian, Rajat Hebbar, Jackson Trager, Georgios Chochlakis, Aditya Kommineni, Ellie Graeden, Shri Narayanan, Benjamin A. T. Graham, Morteza Dehghani
IJCAI9
2024 Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
Tiantian Feng, Dimitrios Dimitriadis, Shri Narayanan
INTERSPEECH3
2024 Analysis of articulatory setting for L1 and L2 English speakers using MRI data
abstract
This paper investigates the extent to which the geographical region (country) where a speaker acquired their English language affects the articulatory setting in their speech.To obtain accurate measurements for evaluating articulatory setting, we utilized a large real-time MRI corpus of vocal tract articulation.The corpus was obtained from speakers from a variety of linguistic backgrounds producing continuous English speech.We use an automated pipeline to process and extract articulatory positional information from the MRI video data.This data is used to draw comparisons between English language speakers from the United States and speakers who acquired their English in India, Korea, and China.Analysis of the speaker groups reveals statistically significant articulatory setting posture differences in multiple places of articulation.
Jack Goldberg, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2024 State-of-the-art speech production MRI protocol for new 0.55 Tesla scanners
Prakash Kumar, Yongwan Lim, Sophia X. Cui, Christina Hagedorn, Dani Byrd, Uttam K. Sinha, Shri Narayanan, Krishna S. Nayak
INTERSPEECH8
2024 Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
abstract
Speech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli.We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals.Our approach aims to directly reconstruct listened speech waveforms given EEG signals, where no intermediate acoustic feature processing step is required.The proposed method consists of an EEG module and a speech module along with a connector.The EEG module learns to better represent EEG signals, while the speech module generates speech waveforms from model representations.The connector learns to bridge the distributions of the latent spaces of EEG and speech.The proposed framework is both simple and efficient, by allowing single-step inference, and outperforms prior works on objective metrics.A fine-grained phoneme analysis is conducted to unveil model characteristics of speech decoding.The source code is available here: github.com/lee-jhwn/fesde.
Aditya Kommineni, Tiantian Feng, Kleanthis Avramidis, Xuan Shi, Sudarsana Reddy Kadiri, Shri Narayanan
INTERSPEECH7
2024 Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
Anfeng Xu, Tiantian Feng, Lue Shen, Helen Tager-Flusberg, Shri Narayanan
INTERSPEECH6
2024 Scaling Representation Learning From Ubiquitous ECG With State-Space Models
abstract
Ubiquitous sensing from wearable devices in the wild holds promise for enhancing human well-being, from diagnosing clinical conditions and measuring stress to building adaptive health promoting scaffolds. But the large volumes of data therein across heterogeneous contexts pose challenges for conventional supervised learning approaches. Representation Learning from biological signals is an emerging realm catalyzed by the recent advances in computational modeling and the abundance of publicly shared databases. The electrocardiogram (ECG) is the primary researched modality in this context, with applications in health monitoring, stress and affect estimation. Yet, most studies are limited by small-scale controlled data collection and over-parameterized architecture choices. We introduce WildECG, a pre-trained state-space model for representation learning from ECG signals. We train this model in a self-supervised manner with 275 000 10 s ECG recordings collected in the wild and evaluate it on a range of downstream tasks. The proposed model is a robust backbone for ECG analysis, providing competitive performance on most of the tasks considered, while demonstrating efficacy in low-resource regimes.
Kleanthis Avramidis, Dominika Kunc, Bartosz Perz, Kranti Adsul, Tiantian Feng, Przemyslaw Kazienko, Stanislaw Saganowski, Shri Narayanan
IEEE J. Biomed. Health Informatics8
2023 PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
abstract
Many recent studies have focused on fine-tuning pretrained models for speech emotion recognition (SER), resulting in promising performance compared to traditional methods that rely largely on low-level, knowledge-inspired acoustic features. These pre-trained speech models learn general-purpose speech representations using self-supervised or weakly-supervised learning objectives from large-scale datasets. Despite the significant advances made in SER through the use of pre-trained architecture, fine-tuning these large pre-trained models for different datasets requires saving copies of entire weight parameters, rendering them impractical to deploy in real-world settings. As an alternative, this work explores parameter-efficient fine-tuning (PEFT) approaches for adapting pre-trained speech models for emotion recognition. Specifically, we evaluate the efficacy of adapter tuning, embedding prompt tuning, and LoRa (Low-rank approximation) on four popular SER testbeds. Our results reveal that LoRa achieves the best fine-tuning performance in emotion recognition while enhancing fairness and requiring only a minimal extra amount of weight parameters. Furthermore, our findings offer novel insights into future research directions in SER, distinct from existing approaches focusing on directly fine-tuning the model architecture. Our code is publicly available under: https://github.com/usc-sail/peft-ser.
Tiantian Feng, Shri Narayanan
ACII2
2023 Context Unlocks Emotions: Text-based Emotion Classification Dataset Auditing with Large Language Models
abstract
The lack of contextual information in text data can make the annotation process of text-based emotion classification datasets challenging. As a result, such datasets often contain labels that fail to consider all the relevant emotions in the vocabulary. This misalignment between text inputs and labels can degrade the performance of machine learning models trained on top of them. As re-annotating entire datasets is a costly and time-consuming task that cannot be done at scale, we propose to use the expressive capabilities of large language models to synthesize additional context for input text to increase its alignment with the annotated emotional labels. In this work, we propose a formal definition of textual context to motivate a prompting strategy to enhance such contextual information. We provide both human and empirical evaluation to demonstrate the efficacy of the enhanced context. Our method improves alignment between inputs and their human-annotated labels from both an empirical and human-evaluated standpoint.
Daniel Yang, Aditya Kommineni, Mohammad Alshehri, Nilamadhab Mohanty, Vedant Modi, Jonathan Gratch, Shri Narayanan
ACII7
2023 Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAP
abstract
There is an imminent need for guidelines and standard test sets to allow direct and fair comparisons of speech emotion recognition (SER). While resources, such as the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, have emerged as widely-adopted reference corpora for researchers to develop and test models for SER, published work reveals a wide range of assumptions and variety in its use that challenge reproducibility and generalization. Based on a critical review of the latest advances in SER using IEMOCAP as the use case, our work aims at two contributions: First, using an analysis of the recent literature, including assumptions made and metrics used therein, we provide a set of SER evaluation guidelines. Second, using recent publications with open-sourced implementations, we focus on reproducibility assessment in SER.
Nikolaos Antoniou, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan
ICASSP4
2023 Navigating and Reaching Therapeutic Goals with Dynamical Systems in Conversation-Based Interventions
abstract
Modern human behavioral signal processing and machine-learning methods have introduced novel ways for representing and estimating internal states of people in goal-based conversational interactions, such as psychotherapy. By combining these methods with systems theoretic approaches, we demonstrate how canonical approaches to control policy design can be utilized for improving the quality of goal-oriented talk-based interactions.
Victor Ardulov, Shri Narayanan
ICASSP2
2023 Signal Processing Grand Challenge 2023 - E-Prevention: Sleep Behavior as an Indicator of Relapses in Psychotic Patients
abstract
This paper presents the approach and results of USC SAIL’s submission to the Signal Processing Grand Challenge 2023 – e-Prevention (Task 2), on detecting relapses in psychotic patients. Relapse prediction has proven to be challenging, primarily due to the heterogeneity of symptoms and responses to treatment between individuals. We address these challenges by investigating the use of sleep behavior features to estimate relapse days as outliers in an unsupervised machine learning setting. We extract informative features from human activity and heart rate data collected in the wild, and evaluate various combinations of feature types and time resolutions. We found that short-time sleep behavior features outperformed their awake counterparts and larger time intervals. Our submission was ranked 3rd in the Task’s official leaderboard, demonstrating the potential of such features as an objective and non-invasive predictor of psychotic relapses.
Kleanthis Avramidis, Kranti Adsul, Digbalay Bose, Shri Narayanan
ICASSP4
2023 On the Role of Visual Context in Enriching Music Representations
abstract
Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval. Incorporating multimodal context from diverse sources provides a promising approach toward modeling this variability. Music presented in media such as movies and music videos provide rich multimodal context that modulates underlying human experiences. However, such context modeling is underexplored, as it requires large amounts of multimodal data along with relevant annotations. Self-supervised learning can help address these challenges by automatically extracting rich, high-level correspondences between different modalities, hence alleviating the need for fine-grained annotations at scale. In this study, we propose VCMR – Video-Conditioned Music Representations, a contrastive learning framework that learns music representations from audio and the accompanying music videos. The contextual visual information enhances representations of music audio, as evaluated on the downstream task of music tagging. Experimental results show that the proposed framework can contribute additive robustness to audio representations and indicates to what extent musical elements are affected or determined by visual context.1
Kleanthis Avramidis, Shanti Stewart, Shri Narayanan
ICASSP3
2023 Contextually-Rich Human Affect Perception Using Multimodal Scene Information
abstract
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived by humans rely on multiple contextual cues including social settings, foreground interactions, and ambient visual scenes. In this work, we leverage pretrained vision-language (VLN) models to extract descriptions of foreground context from images. Further, we propose a multimodal context fusion (MCF) module to combine foreground cues with the visual scene and person-based contextual information for emotion prediction. We show the effectiveness of our proposed modular design on two datasets associated with natural scenes and TV shows.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP4
2023 Leveraging Label Correlations in a Multi-Label Setting: a Case Study in Emotion
abstract
Detecting emotions expressed in text has become critical to a range of fields. In this work, we investigate ways to exploit label correlations in multi-label emotion recognition models to improve emotion detection. First, we develop two modeling approaches to the problem in order to capture word associations of the emotion words themselves, by either including the emotions in the input, or by leveraging Masked Language Modeling (MLM). Second, we inte-grate pairwise constraints of emotion representations as regularization terms alongside the classification loss of the models. We split these terms into two categories, local and global. The former dynamically change based on the gold labels, while the latter remain static during training. We demonstrate state-of-the-art performance across Spanish, English, and Arabic in SemEval 2018 Task 1 E-c using monolingual BERT-based models. On top of better performance, we also demonstrate improved robustness. Code is available at https://github.com/gchochla/Demux-MEmo.1
Georgios Chochlakis, Gireesh Mahajan, Sabyasachee Baruah, Keith Burghardt, Kristina Lerman, Shri Narayanan
ICASSP6
2023 Using Emotion Embeddings to Transfer Knowledge between Emotions, Languages, and Annotation Formats
abstract
The need for emotional inference from text continues to diversify as more and more disciplines integrate emotions into their theories and applications. These needs include inferring different emotion types, handling multiple languages, and different annotation formats. A shared model between different configurations would enable the sharing of knowledge and a decrease in training costs, and would simplify the process of deploying emotion recognition models in novel environments. In this work, we study how we can build a single model that can transition between these different configurations by leveraging multilingual models and Demux, a transformer-based model whose input includes the emotions of interest, enabling us to dynamically change the emotions predicted by the model. Demux also produces emotion embeddings, and performing operations on them allows us to transition to clusters of emotions by pooling the embeddings of each cluster. We show that Demux can simultaneously transfer knowledge in a zero-shot manner to a new language, to a novel annotation format and to unseen emotions. Code is available at https://github.com/gchochla/Demux-MEmo.1
Georgios Chochlakis, Gireesh Mahajan, Sabyasachee Baruah, Keith Burghardt, Kristina Lerman, Shri Narayanan
ICASSP6
2023 A Dataset for Audio-Visual Sound Event Detection in Movies
abstract
Audio event detection is a widely studied field, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically involve manual annotation and verification, which is expensive to perform at scale. Movies depict various real-life and fictional scenarios which makes them a rich resource for mining a wide range of audio events. In this work, we present a dataset of audio events called Subtitle-Aligned Movie Sounds (SAM-S). We use publicly available closed-caption transcripts to automatically mine over 110K audio events from 430 movies. We identify three dimensions to categorize audio events: sound, source, quality, and present the steps involved to produce a final taxonomy of 245 sounds. We discuss the choices involved in generating the taxonomy, and also highlight the human-centered nature of sounds in our dataset. We establish a baseline performance for audio-only sound classification of 34.76% mean average precision, and show that incorporating visual information can further improve the performance by 5%. Data and code are made available for research at https://github.com/usc-sail/mica-subtitle-aligned-movie-sounds
Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai, Shri Narayanan
ICASSP5
2023 A Context-Aware Computational Approach for Measuring Vocal Entrainment in Dyadic Conversations
abstract
Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual’s cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal entrainment in dyadic conversations. We use conformers (a combination of convolutional network and transformer) for capturing both short-term and long-term conversational context to model entrainment patterns in interactions across different domains. Specifically we use cross-subject attention layers to learn intra- as well as interpersonal signals from dyadic conversations. We first validate the proposed method based on classification experiments to distinguish between real (consistent) and fake (inconsistent/shuffled) conversations. Experimental results on interactions involving individuals with Autism Spectrum Disorder (ASD) also show evidence of a statistically-significant association between the introduced entrainment measure and clinical scores relevant to symptoms, including across gender and age groups.
Rimita Lahiri, Md. Nasir, Catherine Lord, So Hyun Kim, Shri Narayanan
ICASSP5
2023 Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal Data
abstract
The sensitivity of data collected via ambulatory monitoring, which regularly involve the recording of speech signals and sensor information, can cause strong privacy concerns. We investigate user re-identification risk in a corpus of such data collected to observe the interplay between behavior, physiology, and well-being of healthcare workers in their daily life. We then develop a user anonymization approach that preserves well-being information (i.e., anxiety), but eliminates user identify (ID) information. We formulate this via an auto-encoder that learns a transformed version of the original feature set in an adversarial manner so that it minimizes the anxiety estimation loss and maximizes the user classification loss. Results indicate that the original features bear a large user re-identification risk, while also having a good ability to classify a user’s anxiety. After removing the most prone features to user re-identification from the original feature set, the user classification accuracy decreases, while the anxiety classification performance is preserved. The final features transformed via the auto-encoder further reduce evidence of user ID and preserve anxiety classification ability. Findings from this study can contribute to the design privacy-aware bio-behavioral models that can be used for responsible ambulatory monitoring in healthcare and beyond.
Ravi Pranjal, Ranjana Seshadri, Rakesh Kumar Sanath Kumar Kadaba, Tiantian Feng, Shri Narayanan, Theodora Chaspari
ICASSP5
2023 Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?
abstract
With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-based MIDI-to-audio system and improve it in terms of feature computation, model selection, and training strategy, aiming to synthesize highly natural-sounding audio. Moreover, we conducted an extensive model evaluation through listening tests, pitch measurement, and spectrogram analysis. This work demonstrates not only synthesis of highly natural music but offers a thorough analytical approach and useful outcomes for the community. Our code, pre-trained models, supplementary materials, and audio samples are open sourced at https://github.com/nii-yamagishilab/midi-to-audio.
Xuan Shi, Erica Cooper, Xin Wang 0037, Junichi Yamagishi, Shri Narayanan
ICASSP5
2023 FedAudio: A Federated Learning Benchmark for Audio Tasks
abstract
Federated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio data and audio-related tasks. In this paper, we fill this critical gap by introducing a new FL benchmark for audio tasks which we refer to as FedAudio. FedAudio includes four representative and commonly used audio datasets from three important audio tasks that are well aligned with FL use cases. In particular, a unique contribution of FedAudio is the introduction of data noises and label errors to the datasets to emulate challenges when deploying FL systems in real-world settings. FedAudio also includes the benchmark results of the datasets and a PyTorch library with the objective of facilitating researchers to fairly compare their algorithms. We hope FedAudio could act as a catalyst to inspire new FL research for audio tasks and thus benefit the acoustic and speech research community. The datasets and benchmark results can be accessed at https://github.com/zhang-tuo-pdf/FedAudio.
Tiantian Feng, Samiul Alam, Sunwoo Lee 0001, Mi Zhang 0002, Shri Narayanan, Amir Salman Avestimehr
ICASSP6
2023 Beatboxing Kick Drum Kinematics
Reed Blaylock, Shri Narayanan
INTERSPEECH2
2023 Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism
Rimita Lahiri, Tiantian Feng, Rajat Hebbar, Catherine Lord, So Hyun Kim, Shri Narayanan
INTERSPEECH6
2023 Cross-Lingual Features for Alzheimer's Dementia Detection from Speech
Thomas Melistas, Lefteris Kapelonis, Nikolaos Antoniou, Petros Mitseas, Dimitris Sgouropoulos, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan
INTERSPEECH8
2023 Bridging Speech Science and Technology - Now and Into the Future
Shri Narayanan
INTERSPEECH1
2023 Understanding Spoken Language Development of Children with ASD Using Pre-trained Speech Embeddings
Anfeng Xu, Rajat Hebbar, Rimita Lahiri, Tiantian Feng, Lindsay Butler, Lue Shen, Helen Tager-Flusberg, Shri Narayanan
INTERSPEECH8
2023 FedMultimodal: A Benchmark for Multimodal Federated Learning
abstract
Over the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal.
Tiantian Feng, Digbalay Bose, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta 0001, Mi Zhang 0002, Amir Salman Avestimehr, Shri Narayanan
KDD9
2023 MM-AU: Towards Multimodal Understanding of Advertisement Videos
abstract
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce, as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise narrative structures. The narrative structures of advertisements involve several elements like reasoning about the broad content (topic and the underlying message) and examining fine-grained details involving the transition of perceived tone due to the sequence of events and interaction among characters. In this work, to facilitate the understanding of advertisements along the three dimensions of topic categorization, perceived tone transition, and social message detection, we introduce a multimodal multilingual benchmark called MM-AU comprised of 8.4 K videos (147hrs) curated from multiple web-based sources. We explore multiple zero-shot reasoning baselines through the application of large language models on the ads transcripts. Further, we demonstrate that leveraging signals from multiple modalities, including audio, video, and text, in multimodal transformer-based supervised models leads to improved performance compared to unimodal approaches.
Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shri Narayanan
ACM Multimedia6
2023 SEAR: Semantically-grounded Audio Representations
abstract
Audio supports visual story-telling in movies through the use of different sounds. These sounds are often tied to different visual elements, including foreground entities, the interactions between them as well as background context. Visual captions provide a condensed view of an image, providing a natural language description of entities and the relationships between them. In this work, we utilize visual captions to semantically ground audio representations in a self-supervised setup. We leverage state-of-the-art vision-language models to augment movie datasets with visual captions at scale to the order of 9.6M captions to learn audio representations from over 2500 hours of movie data. We evaluate the utility of the learned representations and show state-of-the art performance on two movie understanding tasks, genre and speaking-style classification, outperforming video based methods and audio baselines. Finally, we show that the learned model can be transferred in a zero-shot manner through application in both movie understanding tasks and general action recognition.
Rajat Hebbar, Digbalay Bose, Shri Narayanan
ACM Multimedia3
2023 MovieCLIP: Visual Scene Recognition in Movies
abstract
Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a wide array of real-life and fictional scenarios. Existing visual scene datasets in movies have limited taxonomies and don’t consider the visual scene transition within movie clips. In this work, we address the problem of visual scene recognition in movies by first automatically curating a new and extensive movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets. Instead of manual annotations which can be expensive, we use CLIP to weakly label 1.12 million shots from 32K movie clips based on our proposed taxonomy. We provide baseline visual models trained on the weakly labeled dataset called MovieCLIP and evaluate them on an independent dataset verified by human raters. We show that leveraging features from models pretrained on MovieCLIP benefits downstream tasks such as multi-label scene and genre classification of web videos and movie trailers.
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Yin Cui, Kree Cole-McLaughlin, Huisheng Wang, Shri Narayanan
WACV8
2023 A study of bias mitigation strategies for speaker recognition
Raghuveer Peri, Krishna Somandepalli, Shri Narayanan
Comput. Speech Lang.3
2023 An Engineering View on Emotions and Speech: From Analysis and Predictive Models to Responsible Human-Centered Applications
abstract
The substantial growth of Internet-of-Things technology and the ubiquity of smartphone devices has increased the public and industry focus on speech emotion recognition (SER) technologies. Yet, conceptual, technical, and societal challenges restrict the wide adoption of these technologies in various domains, including, healthcare, and education. These challenges are amplified when automated emotion recognition systems are called to function “in-the-wild” due to the inherent complexity and subjectivity of human emotion, the difficulty of obtaining reliable labels at high temporal resolution, and the diverse contextual and environmental factors that confound the expression of emotion in real life. In addition, societal and ethical challenges hamper the wide acceptance and adoption of these technologies, with the public raising questions about user privacy, fairness, and explainability. This article briefly reviews the history of affective speech processing, provides an overview of current state-of-the-art approaches to SER, and discusses algorithmic approaches to render these technologies accessible to all, maximizing their benefits and leading to responsible human-centered computing applications.
Chi-Chun Lee, Theodora Chaspari, Emily Mower Provost, Shri Narayanan
Proc. IEEE4
2023 Cross Modal Video Representations for Weakly Supervised Active Speaker Localization
abstract
An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is talking, and not. Speaker activity can be automatically discerned from the rich multimodal information present in the media content. This is however a challenging problem due to the vast variety and contextual variability in media content, and the lack of labeled data. In this work, we present a cross-modal neural network for learning visual representations, which have implicit information pertaining to the spatial location of a speaker in the visual frames. Avoiding the need for manual annotations for active speakers in visual frames, acquiring of which is very expensive, we present a weakly supervised system for the task of localizing active speakers in movie content. We use the learned cross-modal visual representations, and provide weak supervision from movie subtitles acting as a proxy for voice activity, thus requiring no manual annotations. Furthermore, we propose an audio-assisted post-processing formulation for the task of active speaker detection. We evaluate the performance of the proposed system on three benchmark datasets: i) AVA active speaker dataset, ii) Visual person clustering dataset, and iii) Columbia datset, and demonstrate the effectiveness of the cross-modal embeddings for localizing active speakers in comparison to fully supervised systems.
Krishna Somandepalli, Shri Narayanan
IEEE Trans. Multim.3
2022 Audio and ASR-based Filled Pause Detection
abstract
Filled pauses (or fillers) are the most common form of speech disfluencies and they can be recognized as hesitation markers (“um”, “uh” and “er”) made by speakers, usually to gain extra time while thinking their next words. Filled pauses are very frequent in spontaneous speech. Their detection is therefore rather important for two basic reasons: (a) their existence influences the performance of individual components, like Automatic Speech Recognition system (ASR), in human-machine interaction and (b) their frequency can characterize the overall speech quality of a particular speaker, as it can be strongly associated with the speaker's confidence. Despite that, only limited work has been published for the detection of filled pauses in speech, especially through audio. In this work, we propose a framework for filled pause detection using both audio and textual information. For the audio modality, we transfer knowledge from a plethora of supervised tasks, such as emotion or speaking rate, using Convolutional Neural Networks (CNNs). For the text modality, we develop a temporal Recurrent Neural Network (RNN) method that takes into account textual information derived from an ASR system. In addition, the proposed transfer learning approach for the audio classifier leads to better results when benchmarked on our internal dataset for which the text is not transcribed but estimated by an ASR system. In this case, a simple late fusion approach boosts the performance even further. This proves that the audio approach is suitable for real-world applications where the transcribed text is not available and has to leverage imperfect ASR results, or even the absence of textual information (to reduce computational cost).
Aggelina Chatziagapi, Dimitris Sgouropoulos, Constantinos Karouzos, Thomas Melistas, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan
ACII7
2022 Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition
abstract
Speech Emotion Recognition (SER) techniques have gained considerable interest in many applications including smart virtual assistants and health state tracking. SER systems often acquire and transmit speech data collected at the client-side to remote cloud platforms for inference and decision making. However, speech data carries rich information not only about emotions conveyed in vocal expressions, but also other sensitive demographic traits, such as gender, age, and language background. It is desirable to select only features that are necessary for the emotion classification while protecting sensitive features. However, there are some features that are necessary for emotion classification. These features may also reveal other demographic traits. In this work, we propose a method to improve inference privacy for sensitive features by injecting noise into the input speech data, but without degrading the SER system performance. The approach combines a noise representation learning architecture, called Cloak [1], with adversarial training to keep relevant information inside the data for emotion classification while removing information that would enable inferring sensitive demographic attributes. Experimental results show that our method can effectively prevent inference of sensitive demographic information, and that the improved privacy comes at a cost of only a minor utility loss for the emotion classification.
Tiantian Feng, Hanieh Hashemi, Murali Annavaram, Shri Narayanan
ICASSP4
2022 Automating Detection of Papilledema in Pediatric Fundus Images with Explainable Machine Learning
abstract
Papilledema is an ophthalmic neurologic disorder in which increased intracranial pressure leads to swelling of the optic nerves. Undiagnosed papilledema in children may lead to blindness and may be a sign of life-threatening conditions, such as brain tumors. Robust and accurate clinical diagnosis of this syndrome can be facilitated by automated analysis of fundus images using deep learning, especially in the presence of challenges posed by pseudopapilledema that has similar fundus appearance but distinct clinical implications. We present a deep learning-based algorithm for the automatic detection of pediatric papilledema. Our approach is based on optic disc localization and detection of explainable papilledema indicators through data augmentation. Experiments on real-world clinical data demonstrate that our proposed method is effective with a diagnostic accuracy comparable to expert ophthalmologists1.
Kleanthis Avramidis, Melinda Chang, Shri Narayanan
ICIP4
2022 Semi-FedSER: Semi-supervised Learning for Speech Emotion Recognition On Federated Learning using Multiview Pseudo-Labeling
abstract
Speech Emotion Recognition (SER) application is frequently associated with privacy concerns as it often acquires and transmits speech data at the client-side to remote cloud platforms for further processing. These speech data can reveal not only speech content and affective information but the speaker's identity, demographic traits, and health status. Federated learning (FL) is a distributed machine learning algorithm that coordinates clients to train a model collaboratively without sharing local data. This algorithm shows enormous potential for SER applications as sharing raw speech or speech features from a user's device is vulnerable to privacy attacks. However, a major challenge in FL is limited availability of high-quality labeled data samples. In this work, we propose a semi-supervised federated learning framework, Semi-FedSER, that utilizes both labeled and unlabeled data samples to address the challenge of limited labeled data samples in FL. We show that our Semi-FedSER can generate desired SER performance even when the local label rate l=20 using two SER benchmark datasets: IEMOCAP and MSP-Improv.
Tiantian Feng, Shri Narayanan
INTERSPEECH2
2022 User-Level Differential Privacy against Attribute Inference Attack of Speech Emotion Recognition on Federated Learning
abstract
Many existing privacy-enhanced speech emotion recognition (SER) frameworks focus on perturbing the original speech data through adversarial training within a centralized machine learning setup. However, this privacy protection scheme can fail since the adversary can still access the perturbed data. In recent years, distributed learning algorithms, especially federated learning (FL), have gained popularity to protect privacy in machine learning applications. While FL provides good intuition to safeguard privacy by keeping the data on local devices, prior work has shown that privacy attacks, such as attribute inference attacks, are achievable for SER systems trained using FL. In this work, we propose to evaluate the user-level differential privacy (UDP) in mitigating the privacy leaks of the SER system in FL. UDP provides theoretical privacy guarantees with privacy parameters $\epsilon$ and $\delta$. Our results show that the UDP can effectively decrease attribute information leakage while keeping the utility of the SER system with the adversary accessing one model update. However, the efficacy of the UDP suffers when the FL system leaks more model updates to the adversary. We make the code publicly available to reproduce the results in https://github.com/usc-sail/fed-ser-leakage.
Tiantian Feng, Raghuveer Peri, Shri Narayanan
INTERSPEECH3
2022 Multimodal Clustering with Role Induced Constraints for Speaker Diarization
abstract
Speaker clustering is an essential step in conventional speaker diarization systems and is typically addressed as an audio-only speech processing task.The language used by the participants in a conversation, however, carries additional information that can help improve the clustering performance.This is especially true in conversational interactions, such as business meetings, interviews, and lectures, where specific roles assumed by interlocutors (manager, client, teacher, etc.) are often associated with distinguishable linguistic patterns.In this paper we propose to employ a supervised text-based model to extract speaker roles and then use this information to guide an audio-based spectral clustering step by imposing must-link and cannot-link constraints between segments.The proposed method is applied on two different domains, namely on medical interactions and on podcast episodes, and is shown to yield improved results when compared to the audio-only approach.
Nikolaos Flemotomos, Shri Narayanan
INTERSPEECH2
2022 An automated quality evaluation framework of psychotherapy conversations with local quality estimates
Zhuohao Chen, Nikolaos Flemotomos, Karan Singla, Torrey A. Creed, David C. Atkins, Shri Narayanan
Comput. Speech Lang.6
2022 Causal indicators for assessing the truthfulness of child speech in forensic interviews
Zane Durante, Victor Ardulov, Manoj Kumar 0007, Jennifer Gongola, Thomas D. Lyon, Shri Narayanan
Comput. Speech Lang.6
2022 A review of speaker diarization: Recent advances with deep learning
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu Jeong Han, Shinji Watanabe 0001, Shri Narayanan
Comput. Speech Lang.6
2022 End-to-end neural systems for automatic children speech recognition: An empirical study
Prashanth Gurunath Shivakumar, Shri Narayanan
Comput. Speech Lang.2
2022 Multi-Label Multi-Task Deep Learning for Behavioral Coding
abstract
We propose a methodology for estimating human behaviors in psychotherapy sessions using multi-label and multi-task learning paradigms. We discuss the problem of behavioral coding in which data of human interactions are annotated with labels to describe relevant human behaviors of interest. We describe two related, yet distinct, corpora consisting of therapist-client interactions in psychotherapy sessions. We experimentally compare the proposed learning approaches for estimating behaviors of interest in these datasets. Specifically, we compare single and multiple label learning approaches, single and multiple task learning approaches, and evaluate the performance of these approaches when incorporating turn context. We demonstrate that the best multi-label, multi-task learning model with turn context achieves 18.9 and 19.5 percent absolute improvements with respect to a logistic regression classifier (for each behavioral coding task respectively) and 6.4 and 6.1 percent absolute improvements with respect to the best single-label, single-task deep neural network models. Lastly, we discuss the insights these modeling paradigms provide into these complex interactions including key commonalities and differences of behaviors within and between the two prevalent psychotherapy approaches–Motivational Interviewing and Cognitive Behavioral Therapy–considered.
James Gibson, David C. Atkins, Torrey A. Creed, Zac E. Imel, Panayiotis G. Georgiou, Shri Narayanan
IEEE Trans. Affect. Comput.6
2022 Modeling Vocal Entrainment in Conversational Speech Using Deep Unsupervised Learning
abstract
In interpersonal spoken interactions, individuals tend to adapt to their conversation partner's vocal characteristics to become similar, a phenomenon known as entrainment. A majority of the previous computational approaches are often knowledge driven and linear and fail to capture the inherent nonlinearity of entrainment. In this article, we present an unsupervised deep learning framework to derive a representation from speech features containing information relevant for vocal entrainment. We investigate both an encoding based approach and a more robust triplet network based approach within the proposed framework. We also propose a number of distance measures in the representation space and use them for quantification of entrainment. We first validate the proposed distances by using them to distinguish real conversations from fake ones. Then we also demonstrate their applications in relation to modeling several entrainment-relevant behaviors in observational psychotherapy, namely agreement, blame and emotional bond.
Md. Nasir, Brian R. Baucom, Craig J. Bryan, Shri Narayanan, Panayiotis G. Georgiou
IEEE Trans. Affect. Comput.4
2022 Joint Multi-Dimensional Model for Global and Time-Series Annotations
Anil Ramakrishna, Rahul Gupta 0001, Shri Narayanan
IEEE Trans. Affect. Comput.3
2022 Robust Character Labeling in Movie Videos: Data Resources and Self-Supervised Feature Adaptation
abstract
Robust face clustering is a vital step in enabling computational understanding of visual character portrayal in media. Face clustering for long-form content is challenging because of variations in appearance and lack of supporting large-scale labeled data. Our work in this paper focuses on two key aspects of this problem: the lack of domain-specific training or benchmark datasets, and adapting face embeddings learned on web images to long-form content, specifically movies. First, we present a dataset of over 169000 face tracks curated from 240 Hollywood movies with weak labels on whether a pair of face tracks belong to the same or a different character. We propose an offline algorithm based on nearest-neighbor search in the embedding space to mine hard-examples from these tracks. We then investigate triplet-loss and multiview correlation-based methods for adapting face embeddings to hard-examples. Our experimental results highlight the usefulness of weakly labeled data for domain-specific feature adaptation. Overall, we find that multiview correlation-based adaptation yields more discriminative and robust face embeddings. Its performance on downstream face verification and clustering tasks is comparable to that of the state-of-the-art results in this domain. We also present the SAIL-Movie Character Benchmark corpus developed to augment existing benchmarks. It consists of racially diverse actors and provides face-quality labels for subsequent error analysis. We hope that the large-scale datasets developed in this work can further advance automatic character labeling in videos. All resources are available freely athttps://sail.usc.edu/~ccmi/multiface.
Krishna Somandepalli, Rajat Hebbar, Shri Narayanan
IEEE Trans. Multim.3
2021 Privacy and Utility Preserving Data Transformation for Speech Emotion Recognition
abstract
Speech carries rich information not only about an individual’s intent but about demographic traits, physical and psychological state among other things. Notably, continuously worn wearable sensors enable researchers to collect egocentric speech data to study and assess real-life expressed emotions, offering unprecedented opportunities for applications in the field of assistive agents, medical diagnoses, and personalized education. Many existing systems collect and transmit these speech data, either processed or unprocessed, from users’ devices to a central server for post analysis. However, egocentric audio sensing for speech emotion recognition has created concerns and risks to privacy, where unintended/improper inferences of sensitive information and demographic information may occur without user consent. Toward addressing these concerns, in this work, we propose a privacy-preserving data transformation technique to mitigate potential threats associated with sensitive information and demographic inferences. The proposed mechanism combines an autoencoder architecture, called replacement autoencoder, with gradient reversal layer to remove sensitive information inside the data, such as sensitive labels and demographics. We empirically validate our approach for predicting emotions using three commonly used datasets for speech emotion recognition. We show that our method can effectively prevent inferences of sensitive emotions and demographic information. We further show that the improved privacy comes at a cost of a minor utility loss for the target application.
Tiantian Feng, Shri Narayanan
ACII2
2021 Mitigating the Bias of Heterogeneous Human Behavior in Affective Computing
abstract
Affective computing is broadly applied to decision making systems ranging from mental health assessment to employability evaluation. The heterogeneity of human behavioral data poses challenges for both model validity and fairness. The limited access to sensitive attributes (e,g., race, gender) in real-world settings makes it more difficult to mitigate the unfairness of the model outcomes. In this work, we focus on the heterogeneity of human behavioral signals and analyze its impact on model fairness. We design a novel method named multi-layer factor analysis to automatically identify the heterogeneity patterns in high-dimensional behavioral data and propose a framework to enhance fairness of behavioral modeling without accessing sensitive attributes.
Shen Yan 0007, Hsien-Te Kao, Kristina Lerman, Shri Narayanan, Emilio Ferrara
ACII4
2021 A Computational Tool to Study Vocal Participation of Women in UN-ITU Meetings
abstract
International organizations such as the United Nations drive policies that impact our everyday lives. Diverse representation of people and ideas in the decision making process of such bodies is critical to ensure that the policies work for everyone. One aspect of the representation is the partipants' expressed gender. In this work, we focus on analyzing meetings at the International Telecommunication Union (ITU). These meetings include a moderator who mediates the proceedings between delegates from across the world speaking in different languages. For the purpose of quantifying the participation of delegates, we propose a scalable, human-in-the-loop system to first identify the moderator's speech and estimate the speaking time with respect to gender for all the speakers. Our proposed system includes three main audio modules: speech activity detection, gender identification and moderator verification using a human-labelled speech probe. We then estimate percentage of speaking time controlled for the moderator's speech. We present detailed and multilingual performance evaluation of the component systems using state-of-the-art technologies for these tasks. Finally, we examine the vocal participation of female delegates in the 2018 ITU Plenipotentiary Conference spanning for 18 days and about 108 hours of audio recordings.
Rajat Hebbar, Krishna Somandepalli, Raghuveer Peri, Ruchir Travadi, Tracy Tuplin, Fernando Rivera, Shri Narayanan
CBMI7
2021 Loss Function Approaches for Multi-label Music Tagging
abstract
Given the ever-increasing volume of music created and released every day, it has never been more important to study automatic music tagging. In this paper, we present an ensemble-based convolutional neural network (CNN) model trained using various loss functions for tagging musical genres from audio. We investigate the effect of different loss functions and resampling strategies on prediction performance, finding that using focal loss improves overall performance on the the MTG-Jamendo dataset: an imbalanced, multi-label dataset with over 18,000 songs in the public domain, containing 57 labels. Additionally, we report results from varying the receptive field on our base classifier-a CNN-based architecture trained using Mel spectrograms-which also results in a model performance boost and state-of-the-art performance on the Jamendo dataset. We conclude that the choice of the loss function is paramount for improving on existing methods in music tagging, particularly in the presence of class imbalance.
Dillon Knox, Timothy Greer, Benjamin Ma, Emily Kuo, Krishna Somandepalli, Shri Narayanan
CBMI6
2021 Context-Aware Speech Stress Detection in Hospital Workers Using Bi-LSTM Classifiers
abstract
Hospital workers are known to work long hours in a highly stressful environment. The COVID-19 pandemic has increased this burden multi-fold. Pre-COVID statistics already showed that one in every three nurses reported burnout, thus affecting patient satisfaction and the quality of their provided service. Real-time monitoring of burnout, and other underlying factors, such as stress, could provide feedback not only to the clinical staff, but also to hospital administrators, thus allowing for supportive measures to be taken early. In this paper, we present a context-aware speech-based system for stress detection. We consider data from 144 hospital workers who were monitored during their daily shifts over a 10-week period; subjective stress readings were collected daily. Wearable devices measured speech features and physiological readings, such as heart rate. Environment sensors, in turn, were used to track staff movement within the hospital. Here, we show the importance of context-awareness for stress level detection based on a bidirectional LSTM deep neural network. In particular, we show the importance of hospital location and circadian rhythm based contextual cues for stress prediction. Overall, we show improvements as high as 14% in F1 scores once context is incorporated, relative to using the speech features alone.
Amr Gaballah, Abhishek Tiwari 0003, Shri Narayanan, Tiago H. Falk
ICASSP3
2021 Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training
abstract
Deep neural network based speaker recognition systems can easily be deceived by an adversary using minuscule imperceptible perturbations to the input speech samples. These adversarial attacks pose serious security threats to the speaker recognition systems that use speech biometric. To address this concern, in this work, we propose a new defense mechanism based on a hybrid adversarial training (HAT) setup. In contrast to existing works on countermeasures against adversarial attacks in deep speaker recognition that only use class-boundary information by supervised cross-entropy (CE) loss, we propose to exploit additional information from supervised and unsupervised cues to craft diverse and stronger perturbations for adversarial training. Specifically, we employ multi-task objectives using CE, feature-scattering (FS), and margin losses to create adversarial perturbations and include them for adversarial training to enhance the robustness of the model. We conduct speaker recognition experiments on the Librispeech dataset, and compare the performance with state-of-the-art projected gradient descent (PGD)-based adversarial training which employs only CE objective. The proposed HAT improves adversarial accuracy by absolute 3.29% and 3.18% for PGD and Carlini-Wagner (CW) attacks respectively, while retaining high accuracy on benign examples.
Monisankha Pal, Arindam Jati, Raghuveer Peri, Chin-Cheng Hsu, Wael Abd-Almageed, Shri Narayanan
ICASSP6
2021 Multi-Scale Speaker Diarization with Neural Affinity Score Fusion
abstract
Predicting the speaker’s identity of short speech segments in human dialogue has been considered one of the most challenging problems in speech signal processing. Speaker representations of short speech segments tend to be unreliable, resulting in poor fidelity of speaker representations in tasks requiring speaker recognition. In this paper, we propose an unconventional method that tackles the trade-off between temporal resolution and the quality of the speaker representations. To find a set of weights that balance the scores from multiple temporal scales of segments, a neural affinity score fusion model is presented. Using the CALLHOME dataset, we show that our proposed multi-scale segmentation and integration approach can achieve a state-of-the-art diarization performance.
Tae Jin Park, Manoj Kumar 0007, Shri Narayanan
ICASSP3
2021 Analyzing Short Term Dynamic Speech Features for Understanding Behavioral Traits of Children with Autism Spectrum Disorder
Young-Kyung Kim, Rimita Lahiri, Md. Nasir, So Hyun Kim, Somer Bishop, Catherine Lord, Shri Narayanan
Interspeech7
2021 Acted vs. Improvised: Domain Adaptation for Elicitation Approaches in Audio-Visual Emotion Recognition
abstract
Key challenges in developing generalized automatic emotion recognition systems include scarcity of labeled data and lack of gold-standard references. Even for the cues that are labeled as the same emotion category, the variability of associated expressions can be high depending on the elicitation context e.g., emotion elicited during improvised conversations vs. acted sessions with predefined scripts. In this work, we regard the emotion elicitation approach as domain knowledge, and explore domain transfer learning techniques on emotional utterances collected under different emotion elicitation approaches, particularly with limited labeled target samples. Our emotion recognition model combines the gradient reversal technique with an entropy loss function as well as the softlabel loss, and the experiment results show that domain transfer learning methods can be employed to alleviate the domain mismatch between different elicitation approaches. Our work provides new insights into emotion data collection, particularly the impact of its elicitation strategies, and the importance of domain adaptation in emotion recognition aiming for generalized systems.
Yelin Kim, Cheng-Hao Kuo, Shri Narayanan
Interspeech4
2021 Leveraging Real-Time MRI for Illuminating Linguistic Velum Action
Miran Oh, Dani Byrd, Shri Narayanan
Interspeech3
2021 Developing Neural Representations for Robust Child-Adult Diarization
abstract
Automated processing and analysis of child speech has been long acknowledged as a harder problem compared to understanding speech by adults. Specifically, conversations between a child and adult involve spontaneous speech which often compounds idiosyncrasies associated with child speech. In this work, we improve upon the task of speaker diarization (determining who spoke when) from audio of child-adult conversations in naturalistic settings. We select conversations from the autism diagnosis and intervention domains, wherein speaker diarization forms an important step towards computational behavioral analysis in support of clinical research and decision making. We train deep speaker embeddings using publicly available child speech and adult speech corpora, unlike predominant state-of-art models which typically utilize only adult speech for speaker embedding training. We demonstrate significant reductions in relative diarization error rate (DER) on DIHARD II (dev) sessions containing child speech (22.88%) and two internal corpora representing interactions involving children with Autism: excerpts from ADOS Mod3 sessions (33.7%) and combination of full-length ADOS and BOSCC sessions (44.99%). Further, we validate our improvements in identifying the child speaker (typically with short speaking time) using the recall measure. Finally, we analyze the effect of fundamental frequency augmentation and the effect of child age, gender on speaker diarization performance.
Suchitra Krishnamachari, Manoj Kumar 0007, So Hyun Kim, Catherine Lord, Shri Narayanan
SLT5
2021 RNN Based Incremental Online Spoken Language Understanding
abstract
Spoken Language Understanding (SLU) typically comprises of an automatic speech recognition (ASR) followed by a natural language understanding (NLU) module. The two modules process signals in a blocking sequential fashion, i.e., the NLU often has to wait for the ASR to finish processing on an utterance basis, potentially leading to high latencies that render the spoken interaction less natural. In this paper, we propose recurrent neural network (RNN) based incremental processing towards the SLU task of intent detection. The proposed methodology offers lower latencies than a typical SLU system, without any significant reduction in system accuracy. We introduce and analyze different recurrent neural network architectures for incremental and online processing of the ASR transcripts and compare it to the existing offline systems. A lexical End-of-Sentence (EOS) detector is proposed for segmenting the stream of transcript into sentences for intent classification. Intent detection experiments are conducted on benchmark ATIS, Snips and Facebook's multilingual task oriented dialog datasets modified to emulate a continuous incremental stream of words with no utterance demarcation. We also analyze the prospects of early intent detection, before EOS, with our proposed system.
Prashanth Gurunath Shivakumar, Naveen Kumar 0004, Panayiotis G. Georgiou, Shri Narayanan
SLT4
2021 An analysis of observation length requirements for machine understanding of human behaviors from spoken language
Sandeep Nallan Chakravarthula, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou
Comput. Speech Lang.3
2021 Adversarial attack and defense strategies for deep speaker recognition systems
Arindam Jati, Chin-Cheng Hsu, Monisankha Pal, Raghuveer Peri, Wael Abd-Almageed, Shri Narayanan
Comput. Speech Lang.6
2021 Unsupervised speech representation learning for behavior modeling using triplet enhanced contextualized networks
Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou
Comput. Speech Lang.3
2021 Computational Media Intelligence: Human-Centered Machine Analysis of Media
abstract
Media is created by humans for humans to tell stories. There exists a natural and imminent need for creating human-centered media analytics to illuminate the stories being told and to understand their impact on individuals and society at large. An objective understanding of media content has numerous applications for different stakeholders, from creators to decision-/policy-makers to consumers. Advances in multimodal signal processing and machine learning (ML) can enable detailed and nuanced characterization of media content (of who, what, how, where, and why) at scale. They can also aid our understanding of the impact of media on a range of issues, including individual experiences, behavioral, cultural, and societal trends, and commercial outcomes. Modern deep learning models combined with audiovisual signal processing can analyze entertainment media, such as Film & TV content to quantify gender, age, and race representations. This creates awareness in an objective way that was hitherto impossible. On the other hand, text mining and natural language processing allow nuanced understanding of language use and spoken interactions in media, such as News to track patterns and trends across different contexts. Moreover, advances in human sensing have enabled us to directly measure the influence of media on an individual’s physiology (and brain), while social media analysis enables tracking the societal impact of media content on different cross sections of the society. This article reviews representative methodologies and algorithms, tools, and systems advancing human-centered media understanding through ML in the pursuit of developing computational media intelligence.
Krishna Somandepalli, Tanaya Guha, Victor R. Martinez, Naveen Kumar 0004, Hartwig Adam, Shri Narayanan
Proc. IEEE6
2021 Extending the Beta divergence to complex values
Colin Vaz, Shri Narayanan
Pattern Recognit. Lett.2
2021 Multimodal Embeddings From Language Models for Emotion Recognition in the Wild
abstract
Word embeddings such as ELMo and BERT have been shown to model word usage in language with greater efficacy through contextualized learning on large-scale language corpora, resulting in significant performance improvement across many natural language processing tasks. In this work we integrate acoustic information into contextualized lexical embeddings through the addition of a parallel stream to the bidirectional language model. This multimodal language model is trained on spoken language data that includes both text and audio modalities. We show that embeddings extracted from this model integrate paralinguistic cues into word meanings and can provide vital affective information by applying these multimodal embeddings to the task of speaker emotion recognition.
Shao-Yen Tseng, Shri Narayanan, Panayiotis G. Georgiou
IEEE Signal Process. Lett.2
2021 Temporal Dynamics of Workplace Acoustic Scenes: Egocentric Analysis and Prediction
abstract
Identification of the acoustic environment from an audio recording, also known as acoustic scene classification, is an active area of research. In this paper, we study dynamically-changing background acoustic scenes from the egocentric perspective of an individual in a workplace. In a novel data collection setup, wearable sensors were deployed on individuals to collect audio signals within a built environment, while Bluetooth-based hubs continuously tracked the individual's location which represents the acoustic scene at a certain time. The data of this paper come from 170 hospital workers gathered continuously during work shifts for a 10 week period. In the first part of our study, we investigate temporal patterns in the egocentric sequence of acoustic scenes encountered by an employee, and the association of those patterns with factors such as job-role and daily routine of the individual. Motivated by evidence of multifaceted effects of ambient sounds on human psychology, we also analyze the association of the temporal dynamics of the perceived acoustic scenes with particular behavioral traits of the individual. Experiments reveal rich temporal patterns in the acoustic scenes experienced by the individuals during their work shifts, and a strong association of those patterns with various constructs related to job-roles and behavior of the employees. In the second part of our study, we employ deep learning models to predict the temporal sequence of acoustic scenes from the egocentric audio signal. We propose a two-stage framework where a recurrent neural network is trained on top of the latent acoustic representations learned by a segment-level neural network. The experimental results show the efficacy of the proposed system in predicting sequence of acoustic scenes, highlighting the existence of underlying temporal patterns in the acoustic scenes experienced in workplace.
Arindam Jati, Amrutha Nadarajan, Raghuveer Peri, Karel Mundnich, Tiantian Feng, Benjamin Girault, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.7
2021 Meta-Learning With Latent Space Clustering in Generative Adversarial Network for Speaker Diarization
abstract
The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness. Earlier work on speaker diarization using generative adversarial network (GAN) with an encoder network (ClusterGAN) to project input x-vectors into a latent space has shown promising performance on meeting data. In this paper, we extend the ClusterGAN network to improve diarization robustness and enable rapid generalization across various challenging domains. To this end, we fetch the pre-trained encoder from the ClusterGAN and fine tune it by using prototypical loss (meta-ClusterGAN or MCGAN) under the meta-learning paradigm. Experiments are conducted on CALLHOME telephonic conversations, AMI meeting data, DIHARD-II (dev set) which includes challenging multi-domain corpus, and two child-clinician interaction corpora (ADOS, BOSCC) related to the autism spectrum disorder domain. Extensive analyses of the experimental data are done to investigate the effectiveness of the proposed ClusterGAN and MCGAN embeddings over x-vectors. The results show that the proposed embeddings with normalized maximum eigengap spectral clustering (NME-SC) back-end consistently outperform the Kaldi state-of-the-art x-vector diarization system. Finally, we employ embedding fusion with x-vectors to provide further improvement in diarization performance. We achieve a relative diarization error rate (DER) improvement of 6.67% to 53.93% on the aforementioned datasets using the proposed fused embeddings over x-vectors. Besides, the MCGAN embeddings provide better performance in the number of speakers estimation and short speech segment diarization compared to x-vectors and ClusterGAN on telephonic conversations.
Monisankha Pal, Manoj Kumar 0007, Raghuveer Peri, Tae Jin Park, So Hyun Kim, Catherine Lord, Somer Bishop, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.8
2021 Evidence of Task-Independent Person-Specific Signatures in EEG Using Subspace Techniques
abstract
Electroencephalography (EEG) signals are promising as alternatives to other biometrics owing to their protection against spoofing. Previous studies have focused on capturing individual variability by analyzing task/condition-specific EEG. This work attempts to model biometric signatures independent of task/condition by normalizing the associated variance. Toward this goal, the paper extends ideas from subspace-based text-independent speaker recognition and proposes novel modifications for modeling multi-channel EEG data. The proposed techniques assume that biometric information is present in the entire EEG signal and accumulate statistics across time in a high dimensional space. These high dimensional statistics are then projected to a lower dimensional space where the biometric information is preserved. The lower dimensional embeddings obtained using the proposed approach are shown to be task-independent. The best subspace system identifies individuals with accuracies of 86.4% and 35.9% on datasets with 30 and 920 subjects, respectively, using just nine EEG channels. The paper also provides insights into the subspace model's scalability to unseen tasks and individuals during training and the number of channels needed for subspace modeling.
Mari Ganesh Kumar, Shri Narayanan, Mriganka Sur, Hema A. Murthy
IEEE Trans. Inf. Forensics Secur.2
2020 Towards end-2-end learning for predicting behavior codes from spoken utterances in psychotherapy conversations
abstract
Spoken language understanding tasks usually rely on pipelines involving complex processing blocks such as voice activity detection, speaker diarization and Automatic speech recognition (ASR). We propose a novel framework for predicting utterance level labels directly from speech features, thus removing the dependency on first generating transcripts, and transcription free behavioral coding. Our classifier uses a pretrained Speech-2-Vector encoder as bottleneck to generate word-level representations from speech features. This pre-trained encoder learns to encode speech features for a word using an objective similar to Word2Vec. Our proposed approach just uses speech features and word segmentation information for predicting spoken utterance-level target labels. We show that our model achieves competitive results to other state-of-the-art approaches which use transcribed text for the task of predicting psychotherapy-relevant behavior codes.
Karan Singla, Zhuohao Chen, David C. Atkins, Shri Narayanan
ACL4
2020 Joint Estimation and Analysis of Risk Behavior Ratings in Movie Scripts
abstract
Exposure to violent, sexual, or substanceabuse content in media increases the willingness of children and adolescents to imitate similar behaviors.Computational methods that identify portrayals of risk behaviors from audio-visual cues are limited in their applicability to films in post-production, where modifications might be prohibitively expensive.To address this limitation, we propose a model that estimates content ratings based on the language use in movie scripts, making our solution available at the earlier stages of creative production.Our model significantly improves the state-of-the-art by adapting novel techniques to learn better movie representations from the semantic and sentiment aspects of a character's language use, and by leveraging the co-occurrence of risk behaviors, following a multi-task approach.Additionally, we show how this approach can be useful to learn novel insights on the joint portrayal of these behaviors, and on the subtleties that filmmakers may otherwise not pick up on.
Victor R. Martinez, Krishna Somandepalli, Yalda T. Uhls, Shri Narayanan
EMNLP (1)4
2020 Identifying Truthful Language in Child Interviews
abstract
When a child is suspected to be the victim or sole witness of a crime, the manner in which information is gathered from the child becomes critical. A child forensic interview is the guided conversation that a legal expert conducts to elicit reliable information from a child. To help substantiate child testimony, it is important to discern characteristics of truthful and deceptive behavior in these interviews. The work presented uses various machine learning algorithms to identify differences in the speech of children when they are lying or being truthful, particularly when they have been asked by a confederate to deceive an interviewer. Results show that vocabulary and psycho-linguistic norms of a child's language use, in response to directed questions, provide substantial information to outperform human adults in detecting truthful statements.
Victor Ardulov, Zane Durante, Shanna Williams, Thomas D. Lyon, Shri Narayanan
ICASSP5
2020 Trapezoidal Segment Sequencing: A Novel Approach for Fusion of Human-Produced Continuous Annotations
abstract
Generating accurate ground truth representations of human subjective experiences and judgements is essential for advancing our understanding of human-centered constructs such as emotions. Often, this requires the collection and fusion of annotations from several people where each one is subject to valuation disagreements, distraction artifacts, and other error sources. This work proposes trapezoidal segment sequencing, a new method for fusing annotations into a single representation that, when used alongside a recently proposed signal warping pipeline for correcting annotation artifacts, produces accurate ground truths. We prove that annotations can be well approximated with trapezoidal signals and present results showing the proposed method performs competitively with state-of-the-art fusion methods on a data set where the true target signal being annotated is known. The main utility of the proposed approach is its ability to help segment individual annotations into interpretable regions where either changes or no perceived changes to the construct occur.
Brandon M. Booth, Shri Narayanan
ICASSP2
2020 Automatic Prediction of Suicidal Risk in Military Couples Using Multimodal Interaction Cues from Couples Conversations
abstract
Suicide is a major societal challenge globally, with a wide range of risk factors, from individual health, psychological and behavioral elements to socio-economic aspects. Military personnel, in particular, are at especially high risk. Crisis resources, while helpful, are often constrained by access to clinical visits or therapist availability, especially when needed in a timely manner. There have hence been efforts on identifying whether communication patterns between couples at home can provide preliminary information about potential suicidal behaviors, prior to intervention. In this work, we investigate whether acoustic, lexical, behavior and turn-taking cues from military couples’ conversations can provide meaningful markers of suicidal risk. We test their effectiveness in real-world noisy conditions by extracting these cues through an automatic diarization and speech recognition front-end. Evaluation is performed by classifying 3 degrees of suicidal risk: none, ideation, attempt. Our automatic system performs significantly better than chance in all classification scenarios and we find that behavior and turn-taking cues are the most informative ones. We also observe that conditioning on factors such as speaker gender and topic of discussion tends to improve classification performance.
Sandeep Nallan Chakravarthula, Md. Nasir, Shao-Yen Tseng, Tae Jin Park, Brian R. Baucom, Craig J. Bryan, Shri Narayanan, Panayiotis G. Georgiou
ICASSP8
2020 Modeling Behavior as Mutual Dependency between Physiological Signals and Indoor Location in Large-Scale Wearable Sensor Study
abstract
Wearable sensors today can unobtrusively collect rich time-series of physiological states and human movement patterns over a prolonged period. Gaining a better understanding of how an individual's physiological responses vary in different workplace environments can be valuable in understanding human behavior related to wellness and performance. In this work, we describe our exploration in discovering the correlation between one's physiological responses and movement patterns within different indoor locations using data collected from nurses in a hospital workplace for a ten week period. In this work, we use simple heuristics to empirically validate the idea that such a relationship may exist and then quantify it using mutual information analysis. We propose and demonstrate a data analysis approach that can also detect variations in the level of mutual dependency between different locations and physiological responses. The mutual dependency measures derived from our method are empirically shown to provide valuable information for improving modeling of self-reported work behavior patterns compared to using features derived from a single data stream.
Tiantian Feng, Brandon M. Booth, Shri Narayanan
ICASSP3
2020 Modeling Behavioral Consistency in Large-Scale Wearable Recordings of Human Bio-Behavioral Signals
abstract
Continuously-worn wearable sensors provide an unprecedented opportunity to unobtrusively measure rich bio-behavioral time-series recordings in natural settings such as the workplace. These time-series data can be helpful in inferring broad patterns of behavior such as common routines and daily stress. Many existing approaches either rely on rigid pre-defined notions of activities or use sensitive contextual measurements, such as GPS location or localization within the home, that present privacy concerns and measurement challenges. In this work, we introduce a novel data processing pipeline to model behavioral consistency in a large real-world wearable recording data-set collected in a hospital workplace setting from nurses and direct clinical providers for a period of ten weeks. We use a non-parametric clustering method to generate time series clusters and capture behavioral consistency via the activity curve model. We evaluate the behavioral consistency model under different work roles and conditions such as between different groups of nursing professions and day versus night shift individuals. We also demonstrate that the learned behavioral consistency feature can assist in predicting self-reported work behaviors and anxiety levels.
Tiantian Feng, Shri Narayanan
ICASSP2
2020 The Role of Annotation Fusion Methods in the Study of Human-Reported Emotion Experience During Music Listening
abstract
Music is a universally-enjoyed art form, but listeners often respond to it in tremendously different ways. The same song can bring one person great joy and another deep sorrow. This paper focuses on modeling human music experience at the group level. In this scenario, human annotations serve an important role in computational modeling, especially where the target constructs under study are hidden, such as dimensions of emotion or enjoyment to music listening. In this work, we investigate several ways to represent aggregate human annotations of the complex, subjective emotional experience of listening to music. We show the utility of several methods for fusing self-reported emotion and enjoyment ratings by predicting these responses with auditory features. Using traditional methods such as time alignment with simple averaging and Dynamic Time Warping, as well as state-of-the-art methods based on Expectation Maximization and Triplet Embeddings, we show that it is possible to accurately represent hidden constructs in time under noisy sampling conditions, evidenced by better performance on behavioral response predictions. That subjective responses to complex musical stimuli can be accurately captured using these methods suggests more general applications to research in areas such as affective computing and music perception.
Timothy Greer, Karel Mundnich, Matthew E. Sachs, Shri Narayanan
ICASSP4
2020 Vocal Tract Articulatory Contour Detection in Real-Time Magnetic Resonance Images Using Spatio-Temporal Context
abstract
Due to its ability to visualize and measure the dynamics of vocal tract shaping during speech production, real-time magnetic resonance imaging (rtMRI) has emerged as one of the prominent research tools. The ability to track different articulators such as the tongue, lips, velum, and the pharynx is a crucial step toward automating further scientific and clinical analysis. Recently, various researchers have addressed the problem of detecting articulatory boundaries, but those are primarily limited to static-image based methods. In this work, we propose to use information from temporal dynamics together with the spatial structure to detect the articulatory boundaries in rtMRI videos. We train a convolutional LSTM network to detect and label the articulatory contours. We compare the produced contours against reference labels generated by iteratively fitting a manually created subject-specific template. We observe that the proposed method outperforms solely image-based methods, especially for the difficult-to-track articulators involved in airway constriction formation during speech.
S. Ashwin Hebbar, Krishna Somandepalli, Asterios Toutios, Shri Narayanan
ICASSP5
2020 Meta-Learning for Robust Child-Adult Classification from Speech
abstract
Computational modeling of naturalistic conversations in clinical applications has seen growing interest in the past decade. An important use-case involves child-adult interactions within the autism diagnosis and intervention domain. In this paper, we address a specific sub-problem of speaker diarization, namely child-adult speaker classification in such dyadic conversations with specified roles. Training a speaker classification system robust to speaker and channel conditions is challenging due to inherent variability in the speech within children and the adult interlocutors. In this work, we propose the use of meta-learning, in particular prototypical networks which optimize a metric space across multiple tasks. By modeling every child-adult pair in the training set as a separate task during meta-training, we learn a representation with improved generalizability compared to conventional supervised learning. We demonstrate improvements over state-of-the-art speaker embeddings (x-vectors) under two evaluation settings: weakly supervised classification (upto 14.53% relative improvement in F1-scores) and clustering (upto relative 9.66% improvement in cluster purity). Our results show that protonets can potentially extract robust speaker embeddings for child-adult classification from speech.
Nithin Rao Koluguri, Manoj Kumar 0007, So Hyun Kim, Catherine Lord, Shri Narayanan
ICASSP5
2020 Learning Domain Invariant Representations for Child-Adult Classification from Speech
abstract
Diagnostic procedures for ASD (autism spectrum disorder) involve semi-naturalistic interactions between the child and a clinician. Computational methods to analyze these sessions require an end-to-end speech and language processing pipeline that goes from raw audio to clinically-meaningful behavioral features. An important component of this pipeline is the ability to automatically detect who is speaking when i.e., perform child-adult speaker classification. This binary classification task is often confounded due to variability associated with the participants' speech and background conditions. Further, scarcity of training data often restricts direct application of conventional deep learning methods. In this work, we address two major sources of variability-age of the child and data source collection location-using domain adversarial learning which does not require labeled target domain data. We use two methods, generative adversarial training with inverted label loss and gradient reversal layer to learn speaker embeddings invariant to the above sources of variability, and analyze different conditions under which the proposed techniques improve over conventional learning methods. Using a large corpus of ADOS-2 (autism diagnostic observation schedule, 2nd edition) sessions, we demonstrate up to 13.45% and 6.44% relative improvements over conventional learning methods.
Rimita Lahiri, Manoj Kumar 0007, Somer Bishop, Shri Narayanan
ICASSP4
2020 Speaker-Invariant Affective Representation Learning via Adversarial Training
abstract
Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold-standard references. In addition, there is much variability from input speech signals, human subjective perception of the signals and emotion label ambiguity. In this paper, we propose a machine learning framework to obtain speech emotion representations by limiting the effect of speaker variability in the speech signals. Specifically we propose to disentangle the speaker characteristics from emotion through an adversarial training network in order to better represent emotion. Our method combines the gradient reversal technique with an entropy loss function to remove such speaker information. Our approach is evaluated on both IEMOCAP and CMU-MOSEI datasets. We show that our method improves speech emotion classification and increases generalization to unseen speakers.
Jing Huang 0019, Shri Narayanan, Panayiotis G. Georgiou
ICASSP4
2020 Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network
abstract
In this work, we propose deep latent space clustering for speaker diarization using generative adversarial network (GAN) back-projection with the help of an encoder network. The proposed diarization system is trained jointly with GAN loss, latent variable recovery loss, and a clustering-specific loss. It uses x-vector speaker embeddings at the input, while the latent variables are sampled from a combination of continuous random variables and discrete one-hot encoded variables using the original speaker labels. We benchmark our proposed system on the AMI meeting corpus, and two child-clinician interaction corpora (ADOS and BOSCC) from the autism diagnosis domain. ADOS and BOSCC contain diagnostic and treatment outcome sessions respectively obtained in clinical settings for verbal children and adolescents with autism. Experimental results show that our proposed system significantly outperform the state-of-the-art x-vector based diarization system on these databases. Further, we perform embedding fusion with x-vectors to achieve a relative diarization error rate (DER) improvement of 31%, 36% and 49% on AMI eval, ADOS and BOSCC corpora respectively, when compared to the x-vector baseline using oracle speech segmentation.
Monisankha Pal, Manoj Kumar 0007, Raghuveer Peri, Tae Jin Park, So Hyun Kim, Catherine Lord, Somer Bishop, Shri Narayanan
ICASSP8
2020 Robust Speaker Recognition Using Unsupervised Adversarial Invariance
abstract
In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker embeddings extracted using a pretrained model onto two lower dimensional embedding spaces. The embedding spaces are learnt to disentangle speaker-discriminative information from all other information present in the audio recordings, without supervision about the acoustic conditions. We analyze the robustness of the proposed embeddings to various sources of variability present in the signal for speaker verification and unsupervised clustering tasks on a large-scale speaker recognition corpus. Our analyses show that the proposed system substantially outperforms the baseline in a variety of challenging acoustic scenarios. Furthermore, for the task of speaker diarization on a real-world meeting corpus, our system shows a relative improvement of 36% in the diarization error rate compared to the state-of-the-art baseline.
Raghuveer Peri, Monisankha Pal, Arindam Jati, Krishna Somandepalli, Shri Narayanan
ICASSP5
2020 Multitask Learning for Darpa Lorelei's Situation Frame Extraction Task
Karan Singla, Shri Narayanan
ICASSP2
2020 Bringing in the Outliers: A Sparse Subspace Clustering Approach to Learn a Dictionary of Mouse Ultrasonic Vocalizations
abstract
Mice vocalize in the ultrasonic range during social interactions. These vocalizations are used in neuroscience and clinical studies to tap into complex behaviors and states. The analysis of these ultrasonic vocalizations (USVs) has been traditionally a manual process, which is prone to errors and human bias, and is not scalable to large scale analysis. We propose a new method to automatically create a dictionary of USVs based on a two-step spectral clustering approach, where we split the set of USVs into inlier and outlier data sets. This approach is motivated by the known degrading performance of sparse subspace clustering with outliers. We apply spectral clustering to the inlier data set and later find the clusters for the outliers. We propose quantitative and qualitative performance measures to evaluate our method in this setting, where there is no ground truth. Our approach outperforms two baselines based on k-means and spectral clustering in all of the proposed performance measures, showing greater distances between clusters and more variability between clusters.
Karel Mundnich, Allison T. Knoll, Pat Levitt, Shri Narayanan
ICASSP5
2020 Fifty Shades of Green: Towards a Robust Measure of Inter-annotator Agreement for Continuous Signals
abstract
Continuous human annotations of complex human experiences are essential for enabling psychological and machine-learned inquiry into the human mind, but establishing a reliable set of annotations for analysis and ground truth generation is difficult. Measures of consensus or agreement are often used to establish the reliability of a collection of annotations and thereby purport their suitability for further research and analysis. This work examines many of the commonly used agreement metrics for continuous-scale and continuous-time human annotations and demonstrates their shortcomings, especially in measuring agreement in general annotation shape and structure. Annotation quality is carefully examined in a controlled study where the true target signal is known and evidence is presented suggesting that annotators' perceptual distortions can be modeled using monotonic functions. A novel measure of agreement is proposed which is agnostic to these perceptual differences between annotators and provides unique information when assessing agreement. We illustrate how this measure complements existing agreement metrics and can serve as a tool for curating a reliable collection of human annotations based on differential consensus.
Brandon M. Booth, Shri Narayanan
ICMI2
2020 Human-centered Multimodal Machine Intelligence
Shri Narayanan
ICMI1
2020 Exploiting Conic Affinity Measures to Design Speech Enhancement Systems Operating in Unseen Noise Conditions
Pavlos Papadopoulos, Shri Narayanan
INTERSPEECH2
2020 The INTERSPEECH 2020 Far-Field Speaker Verification Challenge
abstract
The INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field textindependent speaker verification from single microphone array, and far-field text-dependent speaker verification from distributed microphone arrays.All three tasks pose a cross-channel challenge to the participants.To simulate the real-life scenario, the enrollment utterances are recorded from close-talk cellphone, while the test utterances are recorded from the far-field microphone arrays.In this paper, we describe the database, the challenge, and the baseline system, which is based on a ResNetbased deep speaker network with cosine similarity scoring.For a given utterance, the speaker embeddings of different channels are equally averaged as the final embedding.The baseline system achieves minDCFs of 0.62, 0.66, and 0.64 and EERs of 6.27%, 6.55%, and 7.18% for task 1, task 2, and task 3, respectively.
Xiaoyi Qin, Ming Li 0026, Hui Bu, Wei Rao 0002, Rohan Kumar Das, Shri Narayanan, Haizhou Li 0001
INTERSPEECH6
2020 Sentence Level Estimation of Psycholinguistic Norms Using Joint Multidimensional Annotations
abstract
Psycholinguistic normatives represent various affective and mental constructs using numeric scores and are used in a variety of applications in natural language processing. They are commonly used at the sentence level, the scores of which are estimated by extrapolating word level scores using simple aggregation strategies, which may not always be optimal. In this work, we present a novel approach to estimate the psycholinguistic norms at sentence level. We apply a multidimensional annotation fusion model on annotations at the word level to estimate a parameter which captures relationships between different norms. We then use this parameter at sentence level to estimate the norms. We evaluate our approach by predicting sentence level scores for various normative dimensions and compare with standard word aggregation schemes.
Anil Ramakrishna, Shri Narayanan
INTERSPEECH2
2020 Affective Conditioning on Hierarchical Attention Networks Applied to Depression Detection from Transcribed Clinical Interviews
abstract
In this work we propose a machine learning model for depression detection from transcribed clinical interviews. Depression is a mental disorder that impacts not only the subject's mood but also the use of language. To this end we use a Hierarchical Attention Network to classify interviews of depressed subjects. We augment the attention layer of our model with a conditioning mechanism on linguistic features, extracted from affective lexica. Our analysis shows that individuals diagnosed with depression use affective language to a greater extent than not-depressed. Our experiments show that external affective information improves the performance of the proposed architecture in the General Psychotherapy Corpus and the DAIC-WoZ 2017 depression datasets, achieving state-of-the-art 71.6 and 68.6 F1 scores respectively.
Danai Xezonaki, Georgios Paraskevopoulos, Alexandros Potamianos, Shri Narayanan
INTERSPEECH4
2020 ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends
abstract
The Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process.
Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong
ACM Multimedia9
2020 Vocal tract shaping of emotional speech
Jangwon Kim, Asterios Toutios, Sungbok Lee, Shri Narayanan
Comput. Speech Lang.4
2020 Leveraging Linguistic Context in Dyadic Interactions to Improve Automatic Speech Recognition for Children
abstract
Automatic speech recognition for child speech has been long considered a more challenging problem than for adult speech. Various contributing factors have been identified such as larger acoustic speech variability including mispronunciations due to continuing biological changes in growth, developing vocabulary and linguistic skills, and scarcity of training corpora. A further challenge arises when dealing with spontaneous speech of children involved in a conversational interaction, and especially when the child may have limited or impaired communication ability. This includes health applications, one of the motivating domains of this paper, that involve goal-oriented dyadic interactions between a child and clinician/adult social partner as a part of behavioral assessment. In this work, we use linguistic context information from the interaction to adapt speech recognition models for children speech. Specifically, spoken language from the interacting adult speech provides the context for the child's speech. We propose two methods to exploit this context: lexical repetitions and semantic response generation. For the latter, we make use of sequence-to-sequence models that learn to predict the target child utterance given context adult utterances. Long-term context is incorporated in the model by propagating the cell-state across the duration of conversation. We use interpolation techniques to adapt language models at the utterance level, and analyze the effect of length and direction of context (forward and backward). Two different domains are used in our experiments to demonstrate the generalized nature of our methods - interactions between a child with ASD and an adult social partner in a play-based, naturalistic setting, and in forensic interviews between a child and a trained interviewer. In both cases, context-adapted models yield significant improvement (upto 10.71% in absolute word error rate) over the baseline and perform consistently across context windows and directions. Using statistical analysis, we investigate the effect of source-based (adult) and target-based (child) factors on adaptation methods. Our results demonstrate the applicability of our modeling approach in improving child speech recognition by employing information transfer from the adult interlocutor.
Manoj Kumar 0007, So Hyun Kim, Catherine Lord, Thomas D. Lyon, Shri Narayanan
Comput. Speech Lang.5
2020 Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap
abstract
In this study, we propose a new spectral clustering framework that can auto-tune the parameters of the clustering algorithm in the context of speaker diarization. The proposed framework uses normalized maximum eigengap (NME) values to estimate the number of clusters and the parameters for the threshold of the elements of each row in an affinity matrix during spectral clustering, without the use of parameter tuning on the development set. Even through this hands-off approach, we achieve a comparable or better performance across various evaluation sets than the results found using traditional clustering methods that apply careful parameter tuning and development data. A relative improvement of 17% in the speaker error rate on the well-known CALLHOME evaluation set shows the effectiveness of our proposed spectral clustering with auto-tuning.
Tae Jin Park, Kyu Jeong Han, Manoj Kumar 0007, Shri Narayanan
IEEE Signal Process. Lett.4
2019 Violence Rating Prediction from Movie Scripts
abstract
Violent content in movies can influence viewers’ perception of the society. For example, frequent depictions of certain demographics as perpetrators or victims of abuse can shape stereotyped attitudes. In this work, we propose to characterize aspects of violent content in movies solely from the language used in the scripts. This makes our method applicable to a movie in the earlier stages of content creation even before it is produced. This is complementary to previous works which rely on audio or video post production. Our approach is based on a broad range of features designed to capture lexical, semantic, sentiment and abusive language characteristics. We use these features to learn a vector representation for (1) complete movie, and (2) for an act in the movie. The former representation is used to train a movie-level classification model, and the latter, to train deep-learning sequence classifiers that make use of context. We tested our models on a dataset of 732 Hollywood scripts annotated by experts for violent content. Our performance evaluation suggests that linguistic features are a good indicator for violent content. Furthermore, our ablation studies show that semantic and sentiment features are the most important predictors of violence in this data. To date, we are the first to show the language used in movie scripts is a strong indicator of violent content. This offers novel computational tools to assist in creating awareness of storytelling.
Victor R. Martinez, Krishna Somandepalli, Karan Singla, Anil Ramakrishna, Yalda T. Uhls, Shri Narayanan
AAAI6
2019 Trapezoidal Segmented Regression: A Novel Continuous-scale Real-time Annotation Approximation Algorithm
abstract
Accurate ground truth representations of human behavior and experiences are essential for furthering our understanding of the complex relationships between everyday events and interactions and their effects on people. Producing accurate ground truth signals for subjective or latent experiences is difficult because it requires human annotation and is subject to annotator bias, distraction artifacts, valuation errors, among others. We build on previous work aiming to produce highly accurate continuous-scale ground truth labels for human experiences which advocates using supplemental human observations to warp the continuous-scale annotations to correct these errors. We propose a new method, trapezoidal segmented regression, for optimally approximating fused human-produced continuous-scale annotations to simplify its segmentation into intervals of low and high confidence in valuation. We evaluate this algorithm as an alternative to the total variation denoising method used in prior work by comparing the ground truths that both methods produce in experiments where the true annotation target signal is known a priori. Results show that the proposed signal approximation technique performs on par with the prior method, producing ground truth signals in close alignment with the true target, but with the added advantages of being more easily tuned and intuitive. We conclude that the proposed algorithm enables accurate and more robust ground truth generation.
Brandon M. Booth, Shri Narayanan
ACII2
2019 Predicting Human-Reported Enjoyment Responses in Happy and Sad Music
abstract
Whether in a happy mood or a sad mood, humans enjoy listening to music. In this paper, we introduce a novel method to identify auditory features that best predict listener-reported enjoyment ratings by splitting the features into qualitative feature groups, then training predictive models on these feature groups and comparing prediction performance. Using audio features that relate to dynamics, timbre, harmony, and rhythm, we predicted continuous enjoyment ratings for a set of happy and sad songs. We found that a distributed lag model with Ll regularization best predicted these responses and that timbre-related features were most relevant for predicting enjoyment ratings in happy music, while harmony-related features were most relevant to predicting enjoyment ratings in sad music. This work adds to our understanding of how music influences affective human experience.
Benjamin Ma, Timothy Greer, Matthew E. Sachs, Assal Habibi, Jonas T. Kaplan, Shri Narayanan
ACII6
2019 A system for the 2019 Sentiment, Emotion and Cognitive State Task of DARPA's LORELEI project
abstract
During the course of a Humanitarian Assistance-Disaster Relief (HADR) crisis, that can happen anywhere in the world, real-time information is often posted online by the people in need of help which, in turn, can be used by different stakeholders involved with management of the crisis. Automated processing of such posts can considerably improve the effectiveness of such efforts by, for example, understanding the aggregated emotion from affected populations in specific areas, which may help inform decision-makers on how to best allocate resources for an effective disaster response. However, these efforts may be severely limited by the availability of resources for a particular local language. The ongoing DARPA project Low Resource Languages for Emergent Incidents (LORELEI) aims to further language processing technologies for low resource languages in the context of such a humanitarian crisis. In this work, we describe our submission for the 2019 Sentiment, Emotion and Cognitive state (SEC) pilot task of the LORELEI project. We describe a collection of sentiment analysis systems included in our submission along with the features extracted. Our fielded systems obtained the best results in both English and Spanish language evaluations of the SEC pilot task.
Victor R. Martinez, Anil Ramakrishna, Ming-Chang Chiu, Karan Singla, Shri Narayanan
ACII5
2019 Using Oliver API for emotion-aware movie content characterization
abstract
This paper demonstrates the utilization of Oliver11https://behavioralsignals.com/oliver/, the speech emotion recognition (SER) API created by Behavioral Signals, in the context of a movie content visualization application. Oliver API provides an emotion recognition as-a-service solution that can be accessed via a Web API. In this work, we demonstrate how one can send sound recordings from famous movies, retrieve respective emotional descriptors and use simple aggregations on these descriptors to visualize movie content. We have compiled a dataset of 60 movies, categorized over 8 directors. The classification examples included in this paper indicate the ability of simple emotion aggregations to discriminate between movie directors. In order for others to also experiment with the output of both the API's Emotional and Automatic Speech Recognition, the responses are provided as JSON files in this link: https://tinyurl.com/yxeqvvy2.
Theodoros Giannakopoulos, Spiros Dimopoulos, Georgios Pantazopoulos, Aggelina Chatziagapi, Dimitris Sgouropoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan
CBMI8
2019 On Evaluating CNN Representations for Low Resource Medical Image Classification
abstract
Convolutional Neural Networks (CNNs) have revolutionized performances in several machine learning tasks such as image classification, object tracking, and keyword spotting. However, given that they contain a large number of parameters, their direct applicability into low resource tasks is not straightforward. In this work, we experiment with an application of CNN models to gastrointestinal landmark classification with only a few thousands of training samples through transfer learning. As in a standard transfer learning approach, we train CNNs on a large external corpus, followed by representation extraction for the medical images. Finally, a classifier is trained on these CNN representations. However, given that several variants of CNNs exist, the choice of CNN is not obvious. To address this, we develop a novel metric that can be used to predict test performances, given CNN representations on the training set. Not only we demonstrate the superiority of the CNN based transfer learning approach against an assembly of knowledge driven features, but the proposed metric also carries an 87% correlation with the test set performances as obtained using various CNN representations.
Taruna Agrawal, Rahul Gupta 0001, Shri Narayanan
ICASSP3
2019 Toward Robust Interpretable Human Movement Pattern Analysis in a Workplace Setting
abstract
Gaining a better understanding of how people move about and interact with their environment is an important piece of understanding human behavior. Careful analysis of individuals' deviations or variations in movement over time can provide an awareness about changes to their physical or mental state and may be helpful in tracking performance and well-being especially in workplace settings. We propose a technique for clustering and discovering patterns in human movement data by extracting motifs from the time series of durations where participants linger at different locations. Using a data set of over 200 participants moving around a hospital for ten weeks, we show this technique intuitively captures local temporal relationships between hospital rooms and also clusters them in a fashion consistent with the room type labels (e.g. lounge, break room, etc.) without using prior knowledge. Machine learning features derived from these clusters are empirically shown to provide information similar to features attained using domain knowledge of the room type labels directly when predicting mental wellness from self-reports.
Brandon M. Booth, Tiantian Feng, Abhishek Jangalwa, Shri Narayanan
ICASSP4
2019 Improving the Prediction of Therapist Behaviors in Addiction Counseling by Exploiting Class Confusions
abstract
In this work we address the problem of joint prosodic and lexical behavioral annotation for addiction counseling. We expand on past work that employed Recurrent Neural Networks (RNNs) on multimodal features by grouping and classifying subsets of classes. We propose two implementations: One is hierarchical classification, which uses the behavior confusion matrix to cluster similar classes and makes the prediction based on a tree structure. The second is a graph-based method which uses the result of the original classification just to find a certain subset of the most probable candidate classes, where the candidate sets of different predicted classes are determined by the class confusions. We make a second prediction with simpler classifier to discriminate the candidates. The evaluation shows that the strict hierarchical approach degrades performance, likely due to error propagation, while the graph-based hierarchy provides significant gains.
Zhuohao Chen, Karan Singla, James Gibson, Dogan Can, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan
ICASSP8
2019 Discovering Optimal Variable-length Time Series Motifs in Large-scale Wearable Recordings of Human Bio-behavioral Signals
abstract
Continuously-worn wearable sensors produce copious amounts of rich bio-behavioral time series recordings. Exploring recurring patterns, often known as motifs, in wearable time series offers critical insights into understanding the nature of human behavior. Challenges in discovering motifs from wearable recordings include noise removal, pattern generalization, and accounting for subtle variations between subsequences in one motif set. In this work, we introduce a time series processing pipeline to summarize an optimal set of variable-length motifs in a real-world wearable recording data-set collected in a hospital workplace setting. We propose the use of the Savitzky-Golay filter for noise removal without significant data distortion. We then combine the previously developed HierarchIcal based Motif Enumeration (HIME) algorithm with a principled optimization approach to obtain the most repetitive patterns in long-term wearable time-series. We also describe challenges in using just a single method to detect motifs in wearable time series in our experiments. We demonstrate our pipeline can effectively identify meaningful variable-length motifs in large-scale heart rate signals collected continuously from over 100 individuals both at and outside their workplace over 10 weeks through two machine learning experiments.
Tiantian Feng, Shri Narayanan
ICASSP2
2019 Role Specific Lattice Rescoring for Speaker Role Recognition from Speech Recognition Outputs
abstract
The language patterns followed by different speakers who play specific roles in conversational interactions provide valuable cues for the task of Speaker Role Recognition (SRR). Given the speech signal, existing algorithms typically try to find such patterns in the output of an Automatic Speech Recognition (ASR) system. In this work we propose an alternative way of revealing role-specific linguistic characteristics, by making use of role-specific ASR outputs, which are built by suitably rescoring the lattice produced after a first pass of ASR decoding. That way, we avoid pruning the lattice too early, eliminating the potential risk of information loss.
Nikolaos Flemotomos, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan
ICASSP4
2019 Learning Shared Vector Representations of Lyrics and Chords in Music
abstract
Music has a powerful influence on a listener's emotions. In this paper, we represent lyrics and chords in a shared vector space using a phrase-aligned chord-and-lyrics corpus. We show that models that use these shared representations predict a listener's emotion while hearing musical passages better than models that do not use these representations. Additionally, we conduct a visual analysis of these learnt shared vector representations and explain how they support existing theories in music. This work adds to our understanding of how lyrics and chords interact with one another in music and bears applications in music emotion recognition tasks and music information retrieval.
Timothy Greer, Karan Singla, Benjamin Ma, Shri Narayanan
ICASSP4
2019 Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation
abstract
Speech activity detection in highly variable acoustic conditions is a challenging task. Many approaches to detect speech activity in such conditions involve an inherent knowledge of the noise types involved. Movie audio can offer an excellent research test-bed for developing speech activity models. A robust speech detection in movie audio is also a crucial step for subsequent content analyses such as audio diarization. Obtaining labels for supervision of such data can be very expensive, and may not be scalable. In this paper, we employ a simple, yet effective approach to obtain speech labels for movie data by coarse aligning the subtitles with movie audio. We compiled a dataset, called Subtitle-aligned Movie Corpus (SAM) of nearly 23 hours of data labelled as speech from ninety-five Hollywood movies. We propose convolutional neural network architectures that use log-mel spectrograms as input features to predict speech at a segment-level, as opposed to frame-level. We show that our models trained on SAM outperform existing baselines on two independent, publicly released movie speech datasets. We have made the SAM corpus and pretrained models publicly available for further research.
Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
ICASSP3
2019 On Role and Location of Normalization before Model-based Data Augmentation in Residual Blocks for Classification Tasks
abstract
Regularization is crucial to the success of many practical deep learning models, in particular in frequent scenarios where there are only a few to a moderate number of accessible training samples. In addition to weight decay, noise injection and dropout, regularization based on multi-branch architectures, such as Shake-Shake regularization, has been proven successful in many applications and attracted more and more attention. However, beyond model-based representation augmentation, it is unclear how Shake-Shake regularization helps to provide further improvement on classification tasks, let alone the baffling interaction between batch normalization and shaking. In this work, we present our investigation on Shake-Shake regularization. One of our findings illustrates the phenomenon that batch normalization in residual blocks is indispensable when shaking is applied to model branches, along with which we also empirically demonstrate the most effective location to place a batch normalization layer in a shaking regularized residual block. Based on these findings, we believe our work is beneficial to future studies on the research topic of refining control for model-based representation augmentation.
Che-Wei Huang, Shri Narayanan
ICASSP2
2019 Bluetooth Based Indoor Localization Using Triplet Embeddings
abstract
We propose a novel algorithm for indoor localization using triplet embeddings through Bluetooth connectivity streams obtained in very noisy settings with irregular sampling schemes using environmental sensors distributed ad hoc inside buildings. We pose the problem as a matrix completion problem, where a single row and column is added to the (noisy) distances matrix of sensors. Since this is an underdetermined problem, we use information from connectivity between the sender and the trackers to find the missing distances with the help of triplet comparisons. We test our algorithm in a busy hospital setting, where locations such as patient rooms, intensive care units and nursing stations have been equipped with Bluetooth trackers, that are capable of sending Bluetooth packets as well. We assign the sender role to three different trackers, and estimate their locations. We achieve a mean error of 4.01m in indoor localization for three different senders.
Karel Mundnich, Benjamin Girault, Shri Narayanan
ICASSP3
2019 Speaker Agnostic Foreground Speech Detection from Audio Recordings in Workplace Settings from Wearable Recorders
abstract
Audio-signal acquisition as part of wearable sensing adds an important dimension for applications such as understanding human behaviors. As part of a large study on work place behaviours, we collected audio data from individual hospital staff using custom wearable recorders. The audio features collected were limited to preserve privacy of the interactions in the hospital. A first step towards audio processing is to identify the foreground speech of the person wearing the audio badge. This task is challenging because of the multi-party nature of possible ambulatory interactions, lack of access to speaker information and varying channel and ambient conditions. In this paper, we present a speaker-agnostic approach to foreground detection. We propose a convolutional neural network model to predict foreground regions using a limited set of audio features. We show that these models generalize across the proxy corpora we collected in house to approximately match the deployment environment. The proxy corpora contained full audio and was used as a test-bed to analyze our models in greater detail. We also evaluated the models in the workplace setting to measure speech activity. Our experimental results show promising direction for analyzing workplace behaviors with privacy protected sensing.
Amrutha Nadarajan, Krishna Somandepalli, Shri Narayanan
ICASSP3
2019 An Empirical Study of Speech Processing in the Brain by Analyzing the Temporal Syllable Structure in Speech-input Induced EEG
abstract
Clinical applicability of electroencephalography (EEG) is well established, however the use of EEG as a choice for constructing brain computer interfaces to develop communication platforms is relatively recent. To provide more natural means of communication, there is an increasing focus on bringing together speech and EEG signal processing. Quantifying the way our brain processes speech is one way of approaching the problem of speech recognition using brain waves. This paper analyses the feasibility of recognizing syllable level units by studying the temporal structure of speech reflected in the EEG signals. The slowly varying component of the delta band EEG(0.3-3Hz) is present in all other EEG frequency bands. Analysis shows that removing the delta trend in EEG signals results in signals that reveals syllable like structure. Using a 25 syllable framework, classification of EEG data obtained from 13 subjects yields promising results, underscoring the potential of revealing speech related temporal structure in EEG.
Rini A. Sharon, Shri Narayanan, Mriganka Sur, Hema A. Murthy
ICASSP2
2019 Reinforcing Self-expressive Representation with Constraint Propagation for Face Clustering in Movies
abstract
The ability to robustly cluster faces in movies is a necessary step in understanding media content representations of people along dimensions such as gender and age. Building upon the successes of sparse subspace clustering (SSC) in uncovering the underlying structure of the data, in this paper we propose an algorithm called Constraint Propagation Sparse Subspace Clustering (CP-SSC) for applications such as face clustering in videos where pairwise sample constraints (must-link and cannot-link sample pairs) are available in the processing pipeline since detected faces can be tracked locally in time. We learn the subspace structure while simultaneously incorporating the pairwise constraints to construct a similarity matrix needed for clustering. Our joint formulation uses low-rank matrix completion to propagate the initial pairwise constraints, that are used to reinforce the subspace representation during optimization. We evaluate CP-SSC for clustering faces in movies with pre-trained neural network embeddings as features. We first analyze CP-SSC with synthetic data and then show that it can be effectively used to cluster faces in movie videos. We evaluate our method for two movies annotated in-house and two benchmark movies released publicly. We also compare the performance of our algorithm with other clustering approaches that use pairwise constraint information.
Krishna Somandepalli, Shri Narayanan
ICASSP2
2019 Toward Visual Voice Activity Detection for Unconstrained Videos
abstract
The prevalent audio-based Voice Activity Detection (VAD) systems are challenged by the presence of ambient noise and are sensitive to variations in the type of the noise. The use of information from the visual modality, when available, can help overcome some of the problems of audio-based VAD. Existing visual-VAD systems however do not operate directly on the whole image but require intermediate face detection, face landmark detection and subsequent facial feature extraction from the lip region. In this work we present an end-to-end trainable Hierarchical Context Aware (HiCA) architecture for visual-VAD for videos obtained in unconstrained environments which can be trained with videos as input and audio speech labels as output. The network is designed to account for local and global temporal information in a video sequence. In contrast to existing visual-VAD systems our proposed approach does not rely on face detection and subsequent facial feature extraction. It can obtain a VAD accuracy of 66% on a dataset of Hollywood movie videos just with visual information. Further analysis of the representations learned from our visual-VAD system shows that the network learns to localize on human faces, and sometimes speaking human faces specifically. Our quantitative analysis of the effectiveness of face localization shows that our system performs better than sound-localization networks designed for unconstrained videos.
Krishna Somandepalli, Shri Narayanan
ICIP3
2019 Data Augmentation Using GANs for Speech Emotion Recognition
Aggelina Chatziagapi, Georgios Paraskevopoulos, Dimitris Sgouropoulos, Georgios Pantazopoulos, Malvina Nikandrou, Theodoros Giannakopoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan
INTERSPEECH9
2019 Multi-Task Discriminative Training of Hybrid DNN-TVM Model for Speaker Verification with Noisy and Far-Field Speech
Arindam Jati, Raghuveer Peri, Monisankha Pal, Tae Jin Park, Naveen Kumar 0004, Ruchir Travadi, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH8
2019 Identifying Therapist and Client Personae for Therapeutic Alliance Estimation
abstract
. We measure the strength of the relation between personae and alliance in two experiments. Our results show that (1) alliance can be explained by the interactions between the discovered character types, and (2) models trained on therapist and client personae achieve significant performance gains compared to competitive supervised baselines. Finally, exploratory analysis reveals important character traits that lead to an improved perception of alliance.
Victor R. Martinez, Nikolaos Flemotomos, Victor Ardulov, Krishna Somandepalli, Simon B. Goldberg, Zac E. Imel, David C. Atkins, Shri Narayanan
INTERSPEECH8
2019 Modeling Interpersonal Linguistic Coordination in Conversations Using Word Mover's Distance
abstract
embeddings and extend it to measure the dissimilarity in language used in multiple consecutive speaker turns. To validate our approach, we apply this measure for two case studies in the clinical psychology domain. We find that our proposed measure is correlated with the therapist's empathy towards their patient in Motivational Interviewing and with affective behaviors in Couples Therapy. In both case studies, our proposed metric exhibits higher correlation than previously proposed measures. When applied to the couples with relationship improvement, we also notice a significant decrease in the proposed measure over the course of therapy, indicating higher linguistic coordination.
Md. Nasir, Sandeep Nallan Chakravarthula, Brian R. Baucom, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH6
2019 The Second DIHARD Challenge: System Description for USC-SAIL Team
Tae Jin Park, Manoj Kumar 0007, Nikolaos Flemotomos, Monisankha Pal, Raghuveer Peri, Rimita Lahiri, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH8
2019 Speaker Diarization with Lexical Information
abstract
This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with speaker embeddings into a speaker clustering process to improve the overall diarization accuracy. To integrate lexical and acoustic information in a comprehensive way during clustering, we introduce an adjacency matrix integration for spectral clustering. Since words and word boundary information for word-level speaker turn probability estimation are provided by a speech recognition system, our proposed method works without any human intervention for manual transcriptions. We show that the proposed method improves diarization performance on various evaluation datasets compared to the baseline diarization system using acoustic information only in speaker embeddings.
Tae Jin Park, Kyu Jeong Han, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH7
2019 Multiview Shared Subspace Learning Across Speakers and Speech Commands
Krishna Somandepalli, Naveen Kumar 0004, Arindam Jati, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH5
2019 A Multimodal View into Music's Effect on Human Neural, Physiological, and Emotional Experience
abstract
Music has a powerful influence on human experience. In this paper, we investigate how music affects brain activity, physiological response, and human-reported behavior. Using auditory features related to dynamics, timbre, harmony, rhythm, and register, we predicted brain activity in the form of phase synchronizations in bilateral Heschl's gyri and superior temporal gyri; physiological response in the form of galvanic skin response and heart activity; and emotional experience in the form of continuous, subjective descriptions reported by music listeners. We found that using multivariate time series models with attention mechanisms are effective in predicting emotional ratings, while vector-autoregressive models are effective in predicting involuntary human responses. Musical features related to dynamics, register, rhythm, and harmony were found to be particularly helpful in predicting these human reactions. This work adds to our understanding of how music affects multimodal human experience and has applications in affective computing, music emotion recognition, neuroscience, and music information retrieval.
Timothy Greer, Benjamin Ma, Matthew E. Sachs, Assal Habibi, Shri Narayanan
ACM Multimedia5
2019 Efficient estimation and model generalization for the totalvariability model
Ruchir Travadi, Shri Narayanan
Comput. Speech Lang.2
2019 Generating labels for regression of subjective constructs using triplet embeddings
Karel Mundnich, Brandon M. Booth, Benjamin Girault, Shri Narayanan
Pattern Recognit. Lett.4
2019 Total Variability Layer in Deep Neural Network Embeddings for Speaker Verification
abstract
The total variability model (TVM) has been extensively used as a tool to obtain a vector representation of the sources of variability present in a signal. However, recent studies have shown that embeddings derived from a deep neural network (DNN) architecture can provide significant performance improvement over TVM for the speaker verification task. In this letter, we show that TVM can also be reformulated in a manner that enables the integration of a DNN within the model. In addition, we show that this TVM architecture can also be incorporated as one of the layers within a DNN embedding system. Through experiments on speakers in the wild (SITW) corpus, we show that the inclusion of total variability layer in a DNN embedding system provides around 20% relative improvement in equal error rate performance.
Ruchir Travadi, Shri Narayanan
IEEE Signal Process. Lett.2
2018 A Novel Method for Human Bias Correction of Continuous- Time Annotations
abstract
Human annotations are of integral value in human behavior studies and in particular for the generation of ground truth for behavior prediction using various machine learning methods. These often subjective human annotations are especially required for studies involving measuring and predicting hidden mental states (e.g. emotions) that cannot effectively be measured or assessed by other means. Human annotations are noisy and prone to the influence of several factors including personal bias, task ambiguity, environmental distractions, and health state. We propose a novel method for fusion of continuous real-time human annotations to generate accurate ground truth estimates. We introduce a signal warping method that uses additional comparative rank-based information about specific subsets of the annotations to correct for specific types of human annotation artifacts. This approach is validated using a mechanically simple but perceptually demanding psychophysical annotation experiment where objective truth labels are known. Our method yields ground truth estimates that are in better agreement with the objective truth than state-of-the-art approaches.
Brandon M. Booth, Karel Mundnich, Shri Narayanan
ICASSP3
2018 Pykaldi: A Python Wrapper for Kaldi
abstract
We present PyKaldi, a free and open-source Python wrapper for the widely-used Kaldi speech recognition toolkit. PyKaldi is more than a collection of Python bindings into Kaldi libraries. It is an extensible scripting layer that allows users to work with Kaldi and OpenFst types interactively in Python. It tightly integrates Kaldi vector and matrix types with NumPy arrays. We believe Py Kaldi will significantly improve the user experience and simplify the integration of Kaldi into Python workflows. PyKaldi comes with extensive documentation and tests. It is released under the Apache License v2.0 with support for both Python 2.7 and 3.5+.
Dogan Can, Victor R. Martinez, Pavlos Papadopoulos, Shri Narayanan
ICASSP4
2018 Semi-Supervised and Transfer Learning Approaches for Low Resource Sentiment Classification
abstract
Sentiment classification involves quantifying the affective reaction of a human to a document, media item or an event. Although researchers have investigated several methods to reliably infer sentiment from lexical, speech and body language cues, training a model with a small set of labeled datasets is still a challenge. For instance, in expanding sentiment analysis to new languages and cultures, it may not always be possible to obtain comprehensive labeled datasets. In this paper, we investigate the application of semi- supervised and transfer learning methods to improve performances on low resource sentiment classification tasks. We experiment with extracting dense feature representations, pre-training and manifold regularization in enhancing the performance of sentiment classification systems. Our goal is a coherent implementation of these methods and we evaluate the gains achieved by these methods in matched setting involving training and testing on a single corpus setting as well as two cross corpora settings. In both the cases, our experiments demonstrate that the proposed methods can significantly enhance the model performance against a purely supervised approach, particularly in cases involving a handful of training data.
Rahul Gupta 0001, Saurabh Sahu, Carol Y. Espy-Wilson, Shri Narayanan
ICASSP4
2018 Shaking Acoustic Spectral Sub-Bands can Letxer Regularize Learning in Affective Computing
abstract
In this work, we investigate a recently proposed regularization technique based on multi-branch architectures, called Shake-Shake regularization, for the task of speech emotion recognition. In addition, we also propose variants to incorporate domain knowledge into model configurations. The experimental results demonstrate: 1) independently shaking subbands delivers favorable models compared to shaking the entire spectral-temporal feature maps. 2) with proper patience in early stopping, the proposed models can simultaneously outperform the baseline and maintain a smaller performance gap between training and validation.
Che-Wei Huang, Shri Narayanan
ICASSP2
2018 Improving Semi-Supervised Classification for Low-Resource Speech Interaction Applications
abstract
We propose a semi-supervised learning method to improve classification performance in scenarios with limited labeled data. We employ adaptation strategies such as entropy-filtering and self-training, and show that our method achieves up to 17.2% relative improvement in UAR for a multi-class problem. We apply our method to two different tasks: speaker clustering for adult-child interactions during autism assessment sessions, and a variation of the language identification task (LID). We show that in both tasks our method improves classification accuracy while using lesser training data than the baseline and demonstrate the robustness of our setup to the degree of adaptation by controlling the threshold on uncertainty of classification.
Manoj Kumar 0007, Pavlos Papadopoulos, Ruchir Travadi, Daniel Bone, Shri Narayanan
ICASSP5
2018 Multimodal Interaction Modeling of Child Forensic Interviewing
abstract
Constructing computational models of interactions during Forensic Interviews (FI) with children presents a unique challenge in being able to maximize complete and accurate information disclosure, while minimizing emotional trauma experienced by the child. Leveraging multiple channels of observational signals, dynamical system modeling is employed to track and identify patterns in the influence interviewers' linguistic and paralinguistic behavior has on children's verbal recall productivity. Specifically, linear mixed effects modeling and dynamical mode decomposition allow for robust analysis of acoustic-prosodic features, aligned with lexical features at turn-level utterances. By varying the window length, the model parameters evaluate both interviewer and child behaviors at different temporal resolutions, thus capturing both rapport-building and disclosure phases of FI. Making use of a recently proposed definition of productivity, the dynamic systems modeling provides insight into the characteristics of interaction that are most relevant to effectively eliciting narrative and task-relevant information from a child.
Victor Ardulov, Madelyn Mendlen, Manoj Kumar 0007, Neha Anand, Shanna Williams, Thomas D. Lyon, Shri Narayanan
ICMI7
2018 A Multimodal Approach to Understanding Human Vocal Expressions and Beyond
abstract
Human verbal and nonverbal expressions carry crucial information not only about intent but also emotions, individual identity, and the state of health and wellbeing. From a basic science perspective, understanding how such rich information is encoded in these signals can illuminate underlying production mechanisms including the variability therein, within and across individuals. From a technology perspective, finding ways for automatically processing and decoding this complex information continues to be of interest across a variety of applications.
Shri Narayanan
ICMI1
2018 Multimodal Representation of Advertisements Using Segment-level Autoencoders
abstract
Automatic analysis of advertisements (ads) poses an interesting problem for learning multimodal representations. A promising direction of research is the development of deep neural network autoencoders to obtain inter-modal and intra-modal representations. In this work, we propose a system to obtain segment-level unimodal and joint representations. These features are concatenated, and then averaged across the duration of an ad to obtain a single multimodal representation. The autoencoders are trained using segments generated by time-aligning frames between the audio and video modalities with forward and backward context. In order to assess the multimodal representations, we consider the tasks of classifying an ad as funny or exciting in a publicly available dataset of 2,720 ads. For this purpose we train the segment-level autoencoders on a larger, unlabeled dataset of 9,740 ads, agnostic of the test set. Our experiments show that: 1) the multimodal representations outperform joint and unimodal representations, 2) the different representations we learn are complementary to each other, and 3) the segment-level multimodal representations perform better than classical autoencoders and cross-modal representations -- within the context of the two classification tasks. We obtain an improvement of about 5% in classification accuracy compared to a competitive baseline.
Krishna Somandepalli, Victor R. Martinez, Naveen Kumar 0004, Shri Narayanan
ICMI4
2018 Language Features for Automated Evaluation of Cognitive Behavior Psychotherapy Sessions
Nikolaos Flemotomos, Victor R. Martinez, James Gibson, David C. Atkins, Torrey A. Creed, Shri Narayanan
INTERSPEECH6
2018 Combined Speaker Clustering and Role Recognition in Conversational Speech
Nikolaos Flemotomos, Pavlos Papadopoulos, James Gibson, Shri Narayanan
INTERSPEECH4
2018 Improving Gender Identification in Movie Audio Using Cross-Domain Data
Rajat Hebbar, Krishna Somandepalli, Shri Narayanan
INTERSPEECH3
2018 Stochastic Shake-Shake Regularization for Affective Learning from Speech
Che-Wei Huang, Shri Narayanan
INTERSPEECH2
2018 A Knowledge Driven Structural Segmentation Approach for Play-Talk Classification During Autism Assessment
Manoj Kumar 0007, Pooja Chebolu, So Hyun Kim, Kassandra Martinez, Catherine Lord, Shri Narayanan
INTERSPEECH6
2018 Towards an Unsupervised Entrainment Distance in Conversational Speech Using Deep Neural Networks
abstract
Entrainment is a known adaptation mechanism that causes interaction participants to adapt or synchronize their acoustic characteristics. Understanding how interlocutors tend to adapt to each other's speaking style through entrainment involves measuring a range of acoustic features and comparing those via multiple signal comparison methods. In this work, we present a turn-level distance measure obtained in an unsupervised manner using a Deep Neural Network (DNN) model, which we call Neural Entrainment Distance (NED). This metric establishes a framework that learns an embedding from the population-wide entrainment in an unlabeled training corpus. We use the framework for a set of acoustic features and validate the measure experimentally by showing its efficacy in distinguishing real conversations from fake ones created by randomly shuffling speaker turns. Moreover, we show real world evidence of the validity of the proposed measure. We find that high value of NED is associated with high ratings of emotional bond in suicide assessment interviews, which is consistent with prior studies.
Md. Nasir, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou
INTERSPEECH3
2018 Exploring the Relationship between Conic Affinity of NMF Dictionaries and Speech Enhancement Metrics
Pavlos Papadopoulos, Colin Vaz, Shri Narayanan
INTERSPEECH3
2018 Computational Modeling of Conversational Humor in Psychotherapy
abstract
Humor is an important social construct that serves several roles in human communication. Though subjective, it is culturally ubiquitous and is often used to diffuse tension, specially in intense conversations such as those in psychotherapy sessions. Automatic recognition of humor has been of considerable interest in the natural language processing community thanks to its relevance in conversational agents. In this work, we present a model for humor recognition in Motivational Interviewing based psychotherapy sessions. We use a Long Short Term Memory (LSTM) based recurrent neural network sequence model trained on dyadic conversations from psychotherapy sessions and our model outperforms a standard baseline with linguistic humor features.
Anil Ramakrishna, Timothy Greer, David C. Atkins, Shri Narayanan
INTERSPEECH4
2018 Denoising and Raw-waveform Networks for Weakly-Supervised Gender Identification on Noisy Speech
abstract
This paper presents a raw-waveform neural network and uses it along with a denoising network for clustering in weakly supervised learning scenarios under extreme noise conditions. Specifically, we consider language independent Automatic Gender Recognition (AGR) on a set of varied noise conditions and Signal to Noise Ratios (SNRs). We formulate the denoising problem as a source separation task and train the system using a discriminative criterion in order to enhance output SNRs. A denoising Recurrent Neural Network (RNN) is first trained on a small subset (roughly one-fifth) of the data for learning a speech specific mask. The denoised speech signal is then directly fed as input to a raw-waveform convolutional neural network (CNN) trained with denoised speech. We evaluate the standalone performance of denoiser in terms of various signal-to-noise measures and discuss its contribution towards robust AGR. An absolute improvement of 11.06% and 13.33% is achieved by the combined pipeline over the i-vector SVM baseline system for 0 dB and -5 dB SNR conditions, respectively. We further analyse the information captured by the first CNN layer in both noisy and denoised speech.
Jilt Sebastian, Manoj Kumar 0007, Pavan Kumar D. S., Mathew Magimai-Doss, Hema A. Murthy, Shri Narayanan
INTERSPEECH6
2018 Using Prosodic and Lexical Information for Learning Utterance-level Behaviors in Psychotherapy
abstract
In this paper, we present an approach for predicting utterance level behaviors in psychotherapy sessions using both speech and lexical features. We train long short term memory (LSTM) networks with an attention mechanism using words, both manually and automatically transcribed, and prosodic features, at the word level, to predict the annotated behaviors. We demonstrate that prosodic features provide discriminative information relevant to the behavior task and show that they improve prediction when fused with automatically derived lexical features. Additionally, we investigate the weights of the attention mechanism to determine words and prosodic patterns which are of importance to the behavior prediction task.
Karan Singla, Zhuohao Chen, Nikolaos Flemotomos, James Gibson, Dogan Can, David C. Atkins, Shri Narayanan
INTERSPEECH7
2018 Role Annotated Speech Recognition for Conversational Interactions
abstract
Speaker Role Recognition (SRR) assigns a specific speaker role to each speaker-homogeneous speech segment in a conversation. Typically, those segments have to be identified first through a diarization step. Additionally, since SRR is usually based on the different linguistic patterns observed between the roles to be recognized, an Automatic Speech Recognition (ASR) system is also indispensable for the task in hand to convert speech to text. In this work we introduce a Role Annotated Speech Recognition (RASR) system which, given a speech signal, outputs a sequence of words annotated with the corresponding speaker roles. Thus, the need of different component modules which are connected in a way that may lead to error propagation is eliminated. We present, analyze, and test our system for the case of two speaker roles to show-case an end-to-end approach for automatic rich transcription with application to clinical dyadic interactions.
Nikolaos Flemotomos, Zhuohao Chen, David C. Atkins, Shri Narayanan
SLT4
2018 Analysis of speech production real-time MRI
Vikram Ramanarayanan, Sam Tilsen, Michael I. Proctor, Johannes Töger, Louis Goldstein, Krishna S. Nayak, Shri Narayanan
Comput. Speech Lang.7
2018 The ELISA Situation Frame extraction for low resource languages pipeline for LoReHLT'2016
Nikos Malandrakis, Anil Ramakrishna, Victor R. Martinez, Tanner Sorensen, Dogan Can, Shri Narayanan
Mach. Transl.6
2018 A Computational Study of Expressive Facial Dynamics in Children with Autism
abstract
Several studies have established that facial expressions of children with autism are often perceived as atypical, awkward or less engaging by typical adult observers. Despite this clear deficit in the quality of facial expression production, very little is understood about its underlying mechanisms and characteristics. This paper takes a computational approach to studying details of facial expressions of children with high functioning autism (HFA). The objective is to uncover those characteristics of facial expressions, notably distinct from those in typically developing children, and which are otherwise difficult to detect by visual inspection. We use motion capture data obtained from subjects with HFA and typically developing subjects while they produced various facial expressions. This data is analyzed to investigate how the overall and local facial dynamics of children with HFA differ from their typically developing peers. Our major observations include reduced complexity in the dynamic facial behavior of the HFA group arising primarily from the eye region.
Tanaya Guha, Ruth B. Grossman, Shri Narayanan
IEEE Trans. Affect. Comput.4
2018 Modeling Multiple Time Series Annotations as Noisy Distortions of the Ground Truth: An Expectation-Maximization Approach
abstract
Studies of time-continuous human behavioral phenomena often rely on ratings from multiple annotators. Since the ground truth of the target construct is often latent, the standard practice is to use ad-hoc metrics (such as averaging annotator ratings). Despite being easy to compute, such metrics may not provide accurate representations of the underlying construct. In this paper, we present a novel method for modeling multiple time series annotations over a continuous variable that computes the ground truth by modeling annotator specific distortions. We condition the ground truth on a set of features extracted from the data and further assume that the annotators provide their ratings as modification of the ground truth, with each annotator having specific distortion tendencies. We train the model using an Expectation-Maximization based algorithm and evaluate it on a study involving natural interaction between a child and a psychologist, to predict confidence ratings of the children's smiles. We compare and analyze the model against two baselines where: (i) the ground truth in considered to be framewise mean of ratings from various annotators and, (ii) each annotator is assumed to bear a distinct time delay in annotation and their annotations are aligned before computing the framewise mean.
Rahul Gupta 0001, Kartik Audhkhasi, Zach Jacokes, Agata Rozga, Shri Narayanan
IEEE Trans. Affect. Comput.5
2018 Acoustic Denoising Using Dictionary Learning With Spectral and Temporal Regularization
abstract
We present a method for speech enhancement of data collected in extremely noisy environments, such as those obtained during magnetic resonance imaging (MRI) scans. We propose an algorithm based on dictionary learning to perform this enhancement. We use complex nonnegative matrix factorization with intra-source additivity (CMF-WISA) to learn dictionaries of the noise and speech+noise portions of the data and use these to factor the noisy spectrum into estimated speech and noise components. We augment the CMF-WISA cost function with spectral and temporal regularization terms to improve the noise modeling. Based on both objective and subjective assessments, we find that our algorithm significantly outperforms traditional techniques such as Least Mean Squares (LMS) filtering, while not requiring prior knowledge or specific assumptions such as periodicity of the noise waveforms that current state-of-the-art algorithms require.
Colin Vaz, Vikram Ramanarayanan, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Unsupervised Discovery of Character Dictionaries in Animation Movies
abstract
Automatic content analysis of animation movies can enable an objective understanding of character (actor) representations and their portrayals. It can also help illuminate potential markers of unconscious biases and their impact. However, multimedia analysis of movie content has predominantly focused on live-action features. A dearth of multimedia research in this field is because of the complexity and heterogeneity in the design of animated characters-an extremely challenging problem to be generalized by a single method or model. In this paper, we address the problem of automatically discovering characters in animation movies as a first step toward automatic character labeling in these media. Movie-specific character dictionaries can act as a powerful first step for subsequent content analysis at scale. We propose an unsupervised approach which requires no prior information about the characters in a movie. We first use a deep neural network-based object detector that is trained on natural images to identify a set of initial character candidates. These candidates are further pruned using saliency constraints and visual object tracking. A character dictionary per movie is then generated from exemplars obtained by clustering these candidates. We are able to identify both anthropomorphic and nonanthropomorphic characters in a dataset of 46 animation movies with varying composition and character design. Our results indicate high precision and recall of the automatically detected characters compared to human-annotated ground truth, demonstrating the generalizability of our approach.
Krishna Somandepalli, Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
IEEE Trans. Multim.4
2017 Designing Contestability: Interaction Design, Machine Learning, and Mental Health
abstract
We describe the design of an automated assessment and training tool for psychotherapists to illustrate challenges with creating interactive machine learning (ML) systems, particularly in contexts where human life, livelihood, and wellbeing are at stake. We explore how existing theories of interaction design and machine learning apply to the psychotherapy context, and identify "contestability" as a new principle for designing systems that evaluate human behavior. Finally, we offer several strategies for making ML systems more accountable to human actors.
Tad Hirsch, Kritzia Merced, Shri Narayanan, Zac E. Imel, David C. Atkins
Conference on Designing Interactive Systems3
2017 Toward active and unobtrusive engagement assessment of distance learners
abstract
Student behavior and lecturer oversight in the classroom is known to modulate study behaviors and impact performance and learning outcomes, but cannot at present be managed for distance learning students. Quantifying and automatically measuring student engagement during lectures in a scalable and accessible manner for these students is essential for improving academic success, but has not been studied widely in natural distance learning environments. We collect video recordings from a screen-mounted camera of students studying online lectures in a mostly unstructured setting and gather annotations from a panel of humans for assessing student engagement levels. We present results on the prediction of different representations of engagement, both with subject-independent and individual-specific models, and quantify the performance gap between the generalized and personalized models for engagement prediction. While the subject-independent performance is challenged by data sparsity, results show that the individual-specific models can predict engagement well even with very few labeled examples.
Brandon M. Booth, Asem M. Ali, Shri Narayanan, Ian Bennett, Aly A. Farag
ACII3
2017 Exploring sparse representation measures of physiological synchrony for romantic couples
abstract
Quantifying the inherent coordination between interacting individuals can afford us new insights into their emotions, communicative intent, and relationship quality. We propose a novel framework to capture the physiological synchrony between romantic partners through sparse representation techniques and appropriately designed parametric dictionaries that take into account the characteristic structure of the considered signals. Physiological synchrony is operationalized as the similarity of co-occurring electrodermal activity (EDA) streams captured through the distance in the corresponding parametric representation space, as well as through the joint signal representation errors. Results indicate that the proposed sparse EDA synchrony measures (SESM)-evaluated on two datasets of couples' interactions-differ across tasks of various emotional intensity and are associated with the partners' attachment style. These results provide a foundation towards designing novel descriptors of interaction and physiological linkage between individuals for emerging affective computing applications.
Theodora Chaspari, Adela C. Timmons, Brian R. Baucom, Laura Perrone, Katherine J. W. Baucom, Panayiotis G. Georgiou, Gayla Margolin, Shri Narayanan
ACII8
2017 Weighted geodesic flow kernel for interpersonal mutual influence modeling and emotion recognition in dyadic interactions
abstract
Interpersonal mutual influence occurs naturally in social interactions through various behavioral aspects of spoken words, speech prosody, body gestures and so on. Such interpersonal behavior dynamic flow along an interaction is often modulated by the underlying emotional states. This work focuses on modeling how a participant in a dyadic interaction adapts his/her behavior to the multimodal behavior of the interlocutor, to express the emotions. We propose a weighted geodesic flow kernel (WGFK) to capture the complex interpersonal relationship in the expressive human interactions. In our framework, we parameterize the interaction between two partners using WGFK in a Grassmann manifold by fine-grained modeling of the varying contributions in the behavior subspaces of interaction partners. We verify the effectiveness of the WGFK-based interaction modeling in multimodal emotion recognition tasks drawn from dyadic interactions.
Boqing Gong, Shri Narayanan
ACII3
2017 Linguistic analysis of differences in portrayal of movie characters
abstract
Anil Ramakrishna, Victor R. Martínez, Nikolaos Malandrakis, Karan Singla, Shrikanth Narayanan. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Anil Ramakrishna, Victor R. Martinez, Nikos Malandrakis, Karan Singla, Shri Narayanan
ACL (1)5
2017 A knowledge-driven framework for ECG representation and interpretation for wearable applications
abstract
The increasing use of wearable technology creates the need for reliable signal representations with low storage and transmission cost, as well as interpretable models that can be used to translate signals into meaningful constructs. We propose a knowledge-driven sparse representation of the electrocardiogram (ECG) that takes into account the characteristic structure of the corresponding signal through the use of appropriately designed parametric dictionaries containing Hermite and amplitude-modulated sinusoidal atoms for the P, T waves and QRS complex, respectively. We further demonstrate how these atoms can be used to automatically interpret the ECG morphology through the QRS detection and beat classification. Our results indicate relative errors of the order of 10-2, compression rates 10 times smaller than the actual signal, as well as reliable QRS detection (93%) and beat classification (78%). These are discussed in terms of developing efficient and reliable wearable ECG applications.
Ramasubramanian Balasubramanian, Theodora Chaspari, Shri Narayanan
ICASSP3
2017 A knowledge transfer and boosting approach to the prediction of affect in movies
abstract
Affect prediction is a classical problem and has recently garnered special interest in multimedia applications. Affect prediction in movies is one such domain, potentially aiding the design as well as the impact analysis of movies. Given the large diversity in movies (such as different genres and languages), obtaining a comprehensive movie dataset for modeling affect is challenging while models trained on smaller datasets may not generalize. In this paper, we address the problem of continuous affect ratings with the availability of limited in-domain data resources. We initially setup several baseline models trained on in-domain data, followed by a proposal of a Knowledge Transfer (KT) + Gradient Boosting (GB) approach. KT learns models on a larger (mismatched) data which are then adapted to make predictions on the data of interest. GB further updates these predictions based on models learnt from the in-domain data. We observe that the KT + GB models provide Concordance Correlation Coefficient values of 0.13 and 0.27 for valence and affect prediction on the continuous LIRIS ACCEDE dataset against best baseline prediction values of 0.12 and 0.11. Not only the KT + GB models improve the overall performance metrics, we also observe a more consistent model performance across movies of various genres.
Sabyasachee Baruah, Rahul Gupta 0001, Shri Narayanan
ICASSP3
2017 Quantifying regulation mechanisms in dating couples through a dynamical systems model of acoustic and physiological arousal
abstract
Negative emotional arousal during conflict has been related to negative outcomes in romantic relationships and degraded quality of family life. Despite its extensive study in psychology, it is still challenging to quantify emotional arousal in a meaningful way with objective indices beyond traditionally-used self-reported scores. We examine the association of acoustic and physiological arousal between dating couples through speech prosodic patterns and Electrodermal Activity (EDA) features. We use a dynamical systems model (DSM) approach to capture the interplay of arousal indices within and between people. The DSM parameters reflect the amount of self-regulation with respect to the acoustic and physiological cues within a person, the degree of cross-regulation between the two modalities, as well as the within-couple co-regulation. Our results through statistical analysis and classification experiments indicate a significant association between the estimated system parameters and the participants' self-reported relationship satisfaction measures. This is consistent with previous findings and can help towards better understanding regulation mechanisms and escalation effects of emotional arousal during couples' discussions.
Theodora Chaspari, Sohyun C. Han, Daniel Bone, Adela C. Timmons, Laura Perrone, Gayla Margolin, Shri Narayanan
ICASSP7
2017 Towards a definition of local stationarity for graph signals
abstract
In this paper, we extend the recent definition of graph stationarity into a definition of local stationarity. Doing so, we present a metric to assess local stationarity using projections on localized atoms on the graph. Energy of these projections defines the local power spectrum of the signal. We use this local power spectrum to characterize local stationarity and identify sources of non-stationarity through differences of local power spectrum. Finally, we take advantage of the knowledge of the spectrum of the atoms to give a new power spectrum estimator.
Benjamin Girault, Shri Narayanan, Antonio Ortega
ICASSP2
2017 Grasp: A matlab toolbox for graph signal processing
abstract
The GraSP toolbox aims at processing and visualizing graphs and graphs signal with ease. In the demo, we show those capabilities using several examples from the literature and from our own experiments.
Benjamin Girault, Shri Narayanan, Antonio Ortega, Paulo Gonçalves 0001, Eric Fleury
ICASSP2
2017 Estimation of vocal tract area function from volumetric Magnetic Resonance Imaging
abstract
The acoustic properties of speech signals are largely determined by the shaping of the vocal tract. Thus, measurements of vocal-tract area functions and their relationship to various properties of the speech signal have been of interest to the speech research community. Recent advances in Magnetic Resonance Imaging (MRI) allow direct 3D volumetric imaging of the upper airway during production of sustained sounds in as little as seven seconds, therefore allowing direct measurements of vocal-tract area functions for a variety of speech sounds, including fricative and liquid consonants. In this work we present a tool for semi-automatic vocal-tract area function estimation from such data and demonstrate its utility for estimation of the area function for various sustained sounds. Such estimations can be used to address the problem of sagittal-to-area conversion in order to allow inference of 3D vocal-tract shaping dynamics from mid-sagittal real-time MRI data.
Z.-I. Skordilis, Asterios Toutios, Johannes Töger, Shri Narayanan
ICASSP4
2017 Deep convolutional recurrent neural network with attention mechanism for robust speech emotion recognition
abstract
We present a deep convolutional recurrent neural network for speech emotion recognition based on the log-Mel filterbank energies, where the convolutional layers are responsible for the discriminative feature learning. Based on the hypothesis that a better understanding of the internal configuration within an utterance would help reduce misclassification, we further propose a convolutional attention mechanism to learn the utterance structure relevant to the task. In addition, we quantitatively measure the performance gain contributed by each module in our model in order to characterize the nature of emotion expressed in speech. The experimental results on the eNTERFACE'05 emotion database validate our hypothesis and also demonstrate an absolute improvement by 4.62% compared to the state-of-the-art approach.
Che-Wei Huang, Shri Narayanan
ICME2
2017 VCV Synthesis Using Task Dynamics to Animate a Factor-Based Articulatory Model
Rachel Alexander, Tanner Sorensen, Asterios Toutios, Shri Narayanan
INTERSPEECH4
2017 Sounds of the Human Vocal Tract
Reed Blaylock, Nimisha Patil, Timothy Greer, Shri Narayanan
INTERSPEECH4
2017 Acoustic-Prosodic and Physiological Response to Stressful Interactions in Children with Autism Spectrum Disorder
abstract
Social anxiety is a prevalent condition affecting individuals to varying degrees. Research on autism spectrum disorder (ASD), a group of neurodevelopmental disorders marked by impairments in social communication, has found that social anxiety occurs more frequently in this population. Our study aims to further understand the multimodal manifestation of social stress for adolescents with ASD versus neurotypically developing (TD) peers. We investigate this through objective measures of speech behavior and physiology (mean heart rate) acquired during three tasks: a low-stress conversation, a medium-stress interview, and a high-stress presentation. Measurable differences are found to exist for speech behavior and heart rate in relation to task-induced stress. Additionally, we find the acoustic measures are particularly effective for distinguishing between diagnostic groups. Individuals with ASD produced higher prosodic variability, agreeing with previous reports. Moreover, the most informative features captured an individual's vocal changes between low and high social-stress, suggesting an interaction between vocal production and social stressors in ASD.
Daniel Bone, Julia Mertens, Emily Zane, Sungbok Lee, Shri Narayanan, Ruth B. Grossman
INTERSPEECH5
2017 Attention Networks for Modeling Behaviors in Addiction Counseling
James Gibson, Dogan Can, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan
INTERSPEECH5
2017 An Affect Prediction Approach Through Depression Severity Parameter Incorporation in Neural Networks
Rahul Gupta 0001, Saurabh Sahu, Carol Y. Espy-Wilson, Shri Narayanan
INTERSPEECH4
2017 Multi-Scale Context Adaptation for Improving Child Automatic Speech Recognition in Child-Adult Spoken Interactions
Manoj Kumar 0007, Daniel Bone, Kelly McWilliams, Shanna Williams, Thomas D. Lyon, Shri Narayanan
INTERSPEECH6
2017 Transfer Learning Between Concepts for Human Behavior Modeling: An Application to Sincerity and Deception Prediction
Qinyi Luo, Rahul Gupta 0001, Shri Narayanan
INTERSPEECH3
2017 Extracting Situation Frames from Non-English Speech: Evaluation Framework and Pilot Results
Nikos Malandrakis, Ondrej Glembek, Shri Narayanan
INTERSPEECH3
2017 Exploiting Intra-Annotator Rating Consistency Through Copeland's Method for Estimation of Ground Truth Labels in Couples' Therapy
Karel Mundnich, Md. Nasir, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH4
2017 Complexity in Speech and its Relation to Emotional Bond in Therapist-Patient Interactions During Suicide Risk Assessment Interviews
Md. Nasir, Brian R. Baucom, Craig J. Bryan, Shri Narayanan, Panayiotis G. Georgiou
INTERSPEECH4
2017 Global SNR Estimation of Speech Signals for Unknown Noise Conditions Using Noise Adapted Non-Linear Regression
Pavlos Papadopoulos, Ruchir Travadi, Shri Narayanan
INTERSPEECH3
2017 Team ELISA System for DARPA LORELEI Speech Evaluation 2016
Pavlos Papadopoulos, Ruchir Travadi, Colin Vaz, Nikos Malandrakis, Ulf Hermjakob, Nima Pourdamghani, Michael Pust, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Ondrej Glembek, Murali Karthick Baskar, Martin Karafiát, Lukás Burget, Mark Hasegawa-Johnson, Heng Ji 0001, Jonathan May, Kevin Knight, Shri Narayanan
INTERSPEECH20
2017 Comparison of Basic Beatboxing Articulations Between Expert and Novice Artists Using Real-Time Magnetic Resonance Imaging
Nimisha Patil, Timothy Greer, Reed Blaylock, Shri Narayanan
INTERSPEECH4
2017 Semantic Edge Detection for Tracking Vocal Tract Air-Tissue Boundaries in Real-Time Magnetic Resonance Images
Krishna Somandepalli, Asterios Toutios, Shri Narayanan
INTERSPEECH3
2017 Database of Volumetric and Real-Time Vocal Tract MRI for Speech Science
Tanner Sorensen, Z.-I. Skordilis, Asterios Toutios, Yoon-Chul Kim, Yinghua Zhu, Jangwon Kim, Adam C. Lammert, Vikram Ramanarayanan, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shri Narayanan
INTERSPEECH12
2017 Test-Retest Repeatability of Articulatory Strategies Using Real-Time Magnetic Resonance Imaging
Tanner Sorensen, Asterios Toutios, Johannes Töger, Louis Goldstein, Shri Narayanan
INTERSPEECH5
2017 A Distribution Free Formulation of the Total Variability Model
Ruchir Travadi, Shri Narayanan
INTERSPEECH2
2017 Multiple Instance Learning for Behavioral Coding
abstract
We propose a computational methodology for automatically estimating human behavioral patterns using the multiple instance learning (MIL) paradigm. We describe the incremental diverse density algorithm, a particular formulation of multiple instance learning, and discuss its suitability for behavioral coding. We use a rich multi-modal corpus comprised of chronically distressed married couples having problem-solving discussions as a case study to experimentally evaluate our approach. In the multiple instance learning framework, we treat each discussion as a collection of short-term behavioral expressions which are manifested in the acoustic, lexical, and visual channels. We experimentally demonstrate that this approach successfully learns representations that carry relevant information about the behavioral coding task. Furthermore, we employ this methodology to gain novel insights into human behavioral data, such as the local versus global nature of behavioral constructs as well as the level of ambiguity in the expression of behaviors through each respective modality. Finally, we assess the success of each modality for behavioral classification and compare schemes for multimodal fusion within the proposed framework.
James Gibson, Athanasios Katsamanis, Francisco Romero, Bo Xiao 0003, Panayiotis G. Georgiou, Shri Narayanan
IEEE Trans. Affect. Comput.6
2017 Modeling Dynamics of Expressive Body Gestures In Dyadic Interactions
abstract
Body gestures are an important non-verbal expression channel during affective communication. They convey human attitudes and emotions as they dynamically unfold during an interpersonal interaction. Hence, it is highly desirable to understand the dynamics of body gestures associated with emotion expression in human interactions. We present a statistical framework for robustly modeling the dynamics of body gestures in dyadic interactions. Our framework is based on high-level semantic gesture patterns and consists of three components. First, we construct a universal background model (UBM) using Gaussian mixture modeling (GMM) to represent subject-independent gesture variability. Next, we describe each gesture sequence as a concatenation of semantic gesture patterns which are derived from a parallel HMM structure. Then, we probabilistically compare the segments of each gesture sequence extracted from the second step with the UBM obtained from the first step, in order to select highly probabilistic gesture patterns for the sequence. The dynamics of each gesture sequence are represented by a statistical variation profile computed from the selected patterns, and are further described in a well-defined kernel space. This framework is compared with three baseline models and is evaluated in emotion recognition experiments, i.e., recognizing the overall emotional state of a participant in a dyadic interaction from the gesture dynamics. The recognition performance demonstrates the superiority of the proposed framework over the baseline models. The analysis of the relationship between the emotion recognition performance and the number of the selected segments also indicates that a few local salient events, rather than the whole gesture sequence, are sufficiently informative to trigger the human summarization of their overall global emotion perception.
Shri Narayanan
IEEE Trans. Affect. Comput.2
2016 Developing an Automated Report Card for Addiction Counseling: The Counselor Observer Ratings Expert for MI (CORE-MI)
Tad Hirsch, Geoff Gray, James Gibson, Shri Narayanan, Zac E. Imel, David S. Atkins
AMIA4
2016 A multimodal mixture-of-experts model for dynamic emotion prediction in movies
abstract
This paper addresses the problem of continuous emotion prediction in movies from multimodal cues. The rich emotion content in movies is inherently multimodal, where emotion is evoked through both audio (music, speech) and video modalities. To capture such affective information, we put forth a set of audio and video features that includes several novel features such as, Video Compressibility and Histogram of Facial Area (HFA). We propose a Mixture of Experts (MoE)-based fusion model that dynamically combines information from the audio and video modalities for predicting the emotion evoked in movies. A learning module, based on hard Expectation-Maximization (EM) algorithm, is presented for the MoE model. Experiments on a database of popular movies demonstrate that our MoE-based fusion method outperforms popular fusion strategies (e.g. early and late fusion) in the context of dynamic emotion prediction.
Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
ICASSP4
2016 Pathological speech processing: State-of-the-art, current challenges, and future directions
abstract
The study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP) scientists. This growth in interest is led by advances in novel data collection technology, data science, speech processing and computational modeling. These in turn have enabled scientists in better understanding both the causes and effects of pathological speech conditions. In this paper, we review the application of machine learning and signal processing techniques to speech pathology and specifically focus on three different aspects. First, we list challenges such as controlling subjectivity in pathological speech assessments and patient variability in the application of ML-SP tools to the domain. Second, we discuss feature design methods and machine learning algorithms using a combination of domain knowledge and data driven methods. Finally, we present some case studies related to analysis of pathological speech and discuss their design.
Rahul Gupta 0001, Theodora Chaspari, Jangwon Kim, Naveen Kumar 0004, Daniel Bone, Shri Narayanan
ICASSP6
2016 Opening big in box office? Trailer content can help
abstract
Computational prediction of a movie's financial success usually relies only on metadata such as - genre, budget, actors, Motion Picture Association of America (MPAA) rating and critics' reviews. We argue that movie trailers, created to invoke viewers' interest and curiosity about a movie, carry complementary information for predicting a movie's financial future. We created a database consisting of 474 American movie trailers along with various metadata and information about movie's financial success in the opening weekend. A number of features that capture the emotional information contained in the audiovisual stream of the trailers are designed and extracted. We observe that the content-based features have as much predictive information as the meta features. Through regression analysis on our database, we show that signal information from trailer content improves the prediction performance.
Adarsh Tadimari, Naveen Kumar 0004, Tanaya Guha, Shri Narayanan
ICASSP4
2016 CNMF-based acoustic features for noise-robust ASR
abstract
We present an algorithm using convolutive non-negative matrix factorization (CNMF) to create noise-robust features for automatic speech recognition (ASR). Typically in noise-robust ASR, CNMF is used to remove noise from noisy speech prior to feature extraction. However, we find that denoising introduces distortion and artifacts, which can degrade ASR performance. Instead, we propose using the time-activation matrices from CNMF as acoustic model features. In this paper, we describe how to create speech and noise dictionaries that generate noise-robust time-activation matrices from noisy speech. Using the time-activation matrices created by our proposed algorithm, we achieve a 11.8% relative improvement in the word error rate on the Aurora 4 corpus compared to using log-mel filterbank energies. Furthermore, we attain a 13.8% relative improvement over log-mel filterbank energies when we combine them with our proposed features, indicating that our features contain complementary information to log-mel features.
Colin Vaz, Dimitrios Dimitriadis, Samuel Thomas 0001, Shri Narayanan
ICASSP4
2016 Lightly-supervised utterance-level emotion identification using latent topic modeling of multimodal words
abstract
Research on multimodal emotion recognition has drawn much attention recently in diverse disciplines. With the increasing amount of multimodal data, unsupervised or semi-supervised learning has become highly desirable to automatically discover expression of emotion patterns in behavioral data. We present a novel approach for multimodal emotion learning using only a small amount of labels. Our approach is hinging on probabilistic latent semantic analysis (pLSA) that defines the latent variable as the emotion class, motivated by the conceptualization that human emotion acts as a latent control variable that regulates the external behavior manifestations, such as through speech and body gesture. In our approach, we represent the audio-visual information in an utterance as a bag of multimodal words. To exploit the interrelation between speech and gesture modalities, we propose a canonical correlation analysis (CCA) based vocabulary of multimodal words. Our approach has achieved promising experimental results. We have also demonstrated the superiority of the CCA-based multimodal words over those derived directly from the original cues.
Shri Narayanan
ICASSP2
2016 Velum Control for Oral Sounds
Reed Blaylock, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2016 Acoustic-Prosodic and Turn-Taking Features in Interactions with Children with Neurodevelopmental Disorders
Daniel Bone, Somer Bishop, Rahul Gupta 0001, Sungbok Lee, Shri Narayanan
INTERSPEECH5
2016 Automatic Estimation of Perceived Sincerity from Spoken Language
Brandon M. Booth, Rahul Gupta 0001, Pavlos Papadopoulos, Ruchir Travadi, Shri Narayanan
INTERSPEECH5
2016 A Deep Learning Approach to Modeling Empathy in Addiction Counseling
James Gibson, Dogan Can, Bo Xiao 0003, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH7
2016 Predicting Affective Dimensions Based on Self Assessed Depression Severity
Rahul Gupta 0001, Shri Narayanan
INTERSPEECH2
2016 Laughter Valence Prediction in Motivational Interviewing Based on Lexical and Acoustic Cues
Rahul Gupta 0001, Nishant Nath, Taruna Agrawal, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan
INTERSPEECH6
2016 L2 Acquisition and Production of the English Rhotic Pharyngeal Gesture
Sarah Harper, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2016 Attention Assisted Discovery of Sub-Utterance Structure in Speech Emotion Recognition
Che-Wei Huang, Shri Narayanan
INTERSPEECH2
2016 Objective Language Feature Analysis in Children with Neurodevelopmental Disorders During Autism Assessment
Manoj Kumar 0007, Rahul Gupta 0001, Daniel Bone, Nikos Malandrakis, Somer Bishop, Shri Narayanan
INTERSPEECH6
2016 Robust Multichannel Gender Classification from Speech in Movie Audio
Naveen Kumar 0004, Md. Nasir, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH4
2016 Investigation of Speed-Accuracy Tradeoffs in Speech Production Using Real-Time Magnetic Resonance Imaging
Adam C. Lammert, Christine H. Shadle, Shri Narayanan, Thomas F. Quatieri
INTERSPEECH3
2016 Improved Depiction of Tissue Boundaries in Vocal Tract Real-Time MRI Using Automatic Off-Resonance Correction
Yongwan Lim, Sajan Goud Lingala, Asterios Toutios, Shri Narayanan, Krishna S. Nayak
INTERSPEECH4
2016 State-of-the-Art MRI Protocol for Comprehensive Assessment of Vocal Tract Structure and Function
Sajan Goud Lingala, Asterios Toutios, Johannes Töger, Yongwan Lim, Yinghua Zhu, Yoon-Chul Kim, Colin Vaz, Shri Narayanan, Krishna S. Nayak
INTERSPEECH8
2016 Perceptual Lateralization of Coda Rhotic Production in Puerto Rican Spanish
Mairym Lloréns Monteserín, Shri Narayanan, Louis Goldstein
INTERSPEECH2
2016 Complexity in Prosody: A Nonlinear Dynamical Systems Approach for Dyadic Conversations; Behavior and Outcomes in Couples Therapy
Md. Nasir, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou
INTERSPEECH3
2016 Noise Aware and Combined Noise Models for Speech Denoising in Unknown Noise Conditions
Pavlos Papadopoulos, Colin Vaz, Shri Narayanan
INTERSPEECH3
2016 An Expectation Maximization Approach to Joint Modeling of Multidimensional Ratings Derived from Multiple Annotators
abstract
Ratings from multiple human annotators are often pooled in applications where the ground truth is hidden. Examples include annotating perceived emotions and assessing quality metrics for speech and image. These ratings are not restricted to a single dimension and can be multidimensional. In this paper, we propose an Expectation-Maximization based algorithm to model such ratings. Our model assumes that there exists a latent multidimensional ground truth that can be determined from the observation features and that the ratings provided by the annotators are noisy versions of the ground truth. We test our model on a study conducted on children with autism to predict a four dimensional rating of expressivity, naturalness, pronunciation goodness and engagement. Our goal in this application is to reliably predict the individual annotator ratings which can be used to address issues of cognitive load on the annotators as well as the rating cost. We initially train a baseline directly predicting annotator ratings from the features and compare it to our model under three different settings assuming: (i) each entry in the multidimensional rating is independent of others, (ii) a joint distribution among rating dimensions exists, (iii) a partial set of ratings to predict the remaining entries is available.
Anil Ramakrishna, Rahul Gupta 0001, Ruth B. Grossman, Shri Narayanan
INTERSPEECH4
2016 Characterizing Vocal Tract Dynamics Across Speakers Using Real-Time MRI
Tanner Sorensen, Asterios Toutios, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2016 Sensitivity of Quantitative RT-MRI Metrics of Vocal Tract Dynamics to Image Reconstruction Settings
Johannes Töger, Yongwan Lim, Sajan Goud Lingala, Shri Narayanan, Krishna S. Nayak
INTERSPEECH4
2016 Illustrating the Production of the International Phonetic Alphabet Sounds Using Fast Real-Time Magnetic Resonance Imaging
Asterios Toutios, Sajan Goud Lingala, Colin Vaz, Jangwon Kim, John H. Esling, Patricia A. Keating, Matthew Gordon, Dani Byrd, Louis Goldstein, Krishna S. Nayak, Shri Narayanan
INTERSPEECH11
2016 Articulatory Synthesis Based on Real-Time Magnetic Resonance Imaging Data
Asterios Toutios, Tanner Sorensen, Krishna Somandepalli, Rachel Alexander, Shri Narayanan
INTERSPEECH5
2016 Non-Iterative Parameter Estimation for Total Variability Model Using Randomized Singular Value Decomposition
Ruchir Travadi, Shri Narayanan
INTERSPEECH2
2016 Convex Hull Convolutive Non-Negative Matrix Factorization for Uncovering Temporal Patterns in Multivariate Time-Series Data
Colin Vaz, Asterios Toutios, Shri Narayanan
INTERSPEECH3
2016 Behavioral Coding of Therapist Language in Addiction Counseling Using Recurrent Neural Networks
Bo Xiao 0003, Dogan Can, James Gibson, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH7
2016 Analyzing Temporal Dynamics of Dyadic Synchrony in Affective Interactions
Shri Narayanan
INTERSPEECH2
2016 Comparison of feature-level and kernel-level data fusion methods in multi-sensory fall detection
abstract
In this work, we studied the problem of fall detection using signals from tri-axial wearable sensors. In particular, we focused on the comparison of methods to combine signals from multiple tri-axial accelerometers which were attached to different body parts in order to recognize human activities. To improve the detection rate while maintaining a low false alarm rate, previous studies developed detection algorithms by cascading base algorithms and experimented on each sensory data separately. Rather than combining base algorithms, we explored the combination of multiple data sources. Based on the hypothesis that these sensor signals should provide complementary information to the characterization of human's physical activities, we benchmarked a feature level and a kernel-level fusions to learn the kernel that incorporates multiple sensors in the support vector classifier. The results show that given the same false alarm rate constraint, the detection rate improves when using signals from multiple sensors, compared to the baseline where no fusion was employed.
Che-Wei Huang, Shri Narayanan
MMSP2
2016 Comparison of feature-level and kernel-level data fusion methods in multi-sensory fall detection
abstract
In this work, we studied the problem of fall detection using signals from tri-axial wearable sensors. In particular, we focused on the comparison of methods to combine signals from multiple tri-axial accelerometers which were attached to different body parts in order to recognize human activities. To improve the detection rate while maintaining a low false alarm rate, previous studies developed detection algorithms by cascading base algorithms and experimented on each sensory data separately. Rather than combining base algorithms, we explored the combination of multiple data sources. Based on the hypothesis that these sensor signals should provide complementary information to the characterization of human's physical activities, we benchmarked a feature level and a kernel-level fusions to learn the kernel that incorporates multiple sensors in the support vector classifier. The results show that given the same false alarm rate constraint, the detection rate improves when using signals from multiple sensors, compared to the baseline where no fusion was employed.
Che-Wei Huang, Shri Narayanan
MMSP2
2016 Novel affective features for multiscale prediction of emotion in music
abstract
The majority of computational work on emotion in music concentrates on developing machine learning methodologies to build new, more accurate prediction systems, and usually relies on generic acoustic features. Relatively less effort has been put to the development and analysis of features that are particularly suited for the task. The contribution of this paper is twofold. First, the paper proposes two features that can efficiently capture the emotion-related properties in music. These features are named compressibility and sparse spectral components. These features are designed to capture the overall affective characteristics of music (global features). We demonstrate that they can predict emotional dimensions (arousal and valence) with high accuracy as compared to generic audio features. Secondly, we investigate the relationship between the proposed features and the dynamic variation in the emotion ratings. To this end, we propose a novel Haar transform-based technique to predict dynamic emotion ratings using only global features.
Naveen Kumar 0004, Tanaya Guha, Che-Wei Huang, Colin Vaz, Shri Narayanan
MMSP5
2016 Detecting paralinguistic events in audio stream using context in features and probabilistic decisions
Rahul Gupta 0001, Kartik Audhkhasi, Sungbok Lee, Shri Narayanan
Comput. Speech Lang.4
2016 Analysis of engagement behavior in children during dyadic interactions using prosodic cues
Rahul Gupta 0001, Daniel Bone, Sungbok Lee, Shri Narayanan
Comput. Speech Lang.4
2016 Speaker verification based on the fusion of speech acoustics and inverted articulatory signals
Ming Li 0026, Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan
Comput. Speech Lang.6
2016 Directly data-derived articulatory gesture-like representations retain discriminatory information about phone categories
Vikram Ramanarayanan, Maarten Van Segbroeck, Shri Narayanan
Comput. Speech Lang.3
2016 The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing
abstract
Work on voice sciences over recent decades has led to a proliferation of acoustic parameters that are used quite selectively and are not always extracted in a similar fashion. With many independent teams working in different research areas, shared standards become an essential safeguard to ensure compliance with state-of-the-art methods allowing appropriate comparison of results across studies and potential integration and combination of extraction and recognition systems. In this paper we propose a basic standard acoustic parameter set for various areas of automatic voice analysis, such as paralinguistic or clinical speech analysis. In contrast to a large brute-force parameter set, we present a minimalistic set of voice parameters here. These were selected based on a) their potential to index affective physiological changes in voice production, b) their proven value in former studies as well as their automatic extractability, and c) their theoretical significance. The set is intended to provide a common baseline for evaluation of future research and eliminate differences caused by varying parameter sets or even different implementations of the same parameters. Our implementation is publicly available with the openSMILE toolkit. Comparative evaluations of the proposed feature set and large baseline feature sets of INTERSPEECH challenges show a high performance of the proposed set in relation to its size.
Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shri Narayanan, Khiet P. Truong
IEEE Trans. Affect. Comput.10
2016 Long-Term SNR Estimation of Speech Signals in Known and Unknown Channel Conditions
abstract
Many speech processing algorithms and applications rely on the explicit knowledge of signal-to-noise ratio (SNR) in their design and implementation. Estimating the SNR of a signal can enhance the performance of such technologies. We propose a novel method for estimating the long-term SNR of speech signals based on features, from which we can approximately detect regions of speech presence in a noisy signal. By measuring the energy in these regions, we create sets of energy ratios, from which we train regression models for different types of noise. If the type of noise that corrupts a signal is known, we use the corresponding regression model to estimate the SNR. When the noise is unknown, we use a deep neural network to find the “closest” regression model to estimate the SNR. Evaluations were done based on the TIMIT speech corpus, using noises from the NOISEX-92 noise database. Furthermore, we performed cross-corpora experiments by training on TIMIT and NOISEX-92 and testing on the Wall Street Journal speech corpus and DEMAND noise database. Our results show that our system provides accurate SNR estimations across different noise types, corpora, and that it outperforms other SNR estimation methods.
Pavlos Papadopoulos, Andreas Tsiartas, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Context-sensitive learning for enhanced audiovisual emotion classification (Extended abstract)
abstract
Human emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs) and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations.
Angeliki Metallinou, Athanasios Katsamanis, Martin Wöllmer, Florian Eyben, Björn W. Schuller, Shri Narayanan
ACII6
2015 Modeling head motion entrainment for prediction of couples' behavioral characteristics
abstract
Our work examines the link between head motion entrainment of interacting couples and human expert's judgment on certain overall behavioral characteristics (e.g., Blame patterns). We employ a data-driven model that clusters head motion in an unsupervised manner into elementary types called kinemes. We propose three groups of similarity measures based on Kullback-Leibler divergence to model entrainment. We find that the divergence of the (joint) distribution of kinemes yields consistent and significant correlation with target behavior characteristics. The divergence of the conditional distribution of kinemes is shown to predict the polarity of the behavioral characteristics. We partly explain the strong correlations via associating the conditional distributions with the prominent behavioral implications of their respective associated kinemes. These results show the possibility of inferring human behavioral characteristics through the modeling of dyadic head motion entrainment.
Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan
ACII4
2015 A Dynamic Programming Algorithm for Computing N-gram Posteriors from Lattices
abstract
Efficient computation of n-gram posterior probabilities from lattices has applications in lattice-based minimum Bayes-risk decoding in statistical machine translation and the estimation of expected document frequencies from spoken corpora.In this paper, we present an algorithm for computing the posterior probabilities of all ngrams in a lattice and constructing a minimal deterministic weighted finite-state automaton associating each n-gram with its posterior for efficient storage and retrieval.Our algorithm builds upon the best known algorithm in literature for computing ngram posteriors from lattices and leverages the following observations to significantly improve the time and space requirements: i) the n-grams for which the posteriors will be computed typically comprises all n-grams in the lattice up to a certain length, ii) posterior is equivalent to expected count for an n-gram that do not repeat on any path, iii) there are efficient algorithms for computing n-gram expected counts from lattices.We present experimental results comparing our algorithm with the best known algorithm in literature as well as a baseline algorithm based on weighted finite-state automata operations.
Dogan Can, Shri Narayanan
EMNLP2
2015 A quantitative analysis of gender differences in movies using psycholinguistic normatives
abstract
Direct content analysis reveals important details about movies including those of gender representations and potential biases.We investigate the differences between male and female character depictions in movies, based on patterns of language used.Specifically, we use an automatically generated lexicon of linguistic norms characterizing gender ladenness.We use multivariate analysis to investigate gender depictions and correlate them with elements of movie production.The proposed metric differentiates between male and female utterances and exhibits some interesting interactions with movie genres and the screenplay writer gender.
Anil Ramakrishna, Nikos Malandrakis, Elizabeth Staruk, Shri Narayanan
EMNLP4
2015 Quantifying EDA synchrony through joint sparse representation: A case-study of couples' interactions
abstract
The co-variation degree between individuals in their physiological signals can reveal insights about the quality of their interaction as well as their personal characteristics. In an effort to capture the amount of synchrony between Electrodermal Activity (EDA) streams occurring in parallel during dyadic interactions, we propose Sparse EDA Synchrony Measure (SESM), an index derived from the joint sparse representation of EDA ensembles. Sparse decomposition is performed using Simultaneous Orthogonal Matching Pursuit (SOMP) from a knowledge-driven dictionary of tonic and phasic atoms, capturing the slow-varying trends and high-frequency signal fluctuations, respectively. At each iteration the atom having the maximum average correlation with the residuals is selected. We compute SESM as the negative natural logarithm of the joint reconstruction error and evaluate it with data from interactions of married and young dating couples participating in tasks of varying emotional intensity. Through statistical analysis and multiple linear regression experiments, our results indicate that SESM depicts significant differences across tasks in both datasets considered and can be associated to individuals' attachment-related characteristics.
Theodora Chaspari, Brian R. Baucom, Adela C. Timmons, Andreas Tsiartas, Larissa Borofsky Del Piero, Katherine J. W. Baucom, Panayiotis G. Georgiou, Gayla Margolin, Shri Narayanan
ICASSP9
2015 Computationally deconstructing movie narratives: An informatics approach
abstract
In general, popular films and screenplays follow a well defined storytelling paradigm that comprises three essential segments or acts: exposition (act I), conflict (act II) and resolution (act III). Deconstructing a movie into its narrative units can enrich semantic understanding of movies, and help in movie summarization, navigation and detection of the key events. A multimodal framework for detecting such three act narrative structure is developed in this paper. Various low-level features are designed and extracted from video, audio and text channels of a movie so as to capture the pace and excitement of the movie's narrative. Information from the three modalities is combined to compute a continuous dynamic measure of the movie's narrative flow, referred to as the story intensity of the movie in this paper. Guided by the knowledge of film grammar, the act boundaries are detected, and compared against annotations collected from human experts. Promising results are demonstrated for nine full-length Hollywood feature films of various genres.
Tanaya Guha, Naveen Kumar 0004, Shri Narayanan, Stacy L. Smith
ICASSP3
2015 On quantifying facial expression-related atypicality of children with Autism Spectrum Disorder
abstract
by adult observers. This paper focuses on data driven ways to analyze and quantify atypicality in facial expressions of children with ASD. Our objective is to uncover those characteristics of facial gestures that induce the sense of perceived atypicality in observers. Using a carefully collected motion capture database, facial expressions of children with and without ASD are compared within six basic emotion categories employing methods from information theory, time-series modeling and statistical analysis. Our experiments show that children with ASD usually have less complex expression producing mechanisms; the differences in facial dynamics between children with and without ASD primarily come from the eye region. Our study also notes that children with ASD exhibit lower symmetry between left and right regions, and lower variation in motion intensity across facial regions.
Tanaya Guha, Anil Ramakrishna, Ruth B. Grossman, Darren Hedley, Sungbok Lee, Shri Narayanan
ICASSP7
2015 A mixture of experts approach towards intelligibility classification of pathological speech
abstract
Pathological speech involves atypical speech production which may result from several factors including oral diseases, physical disabilities in the voice production system and atypical anatomy. Automatic evaluation of intelligibility in patients with pathological speech can assist accurate diagnosis of pathological conditions. Loss of intelligibility may be associated with one of the several pathological conditions, making automatic evaluation a challenging computational problem. A Mixture of Experts (MoE) models class boundaries using a weighted combination of several experts and can characterize the complex class boundaries arising due to pathological variability. We train an MoE for intelligibility evaluation using a modified Expectation Maximization (EM) algorithm based on joint simulated annealing-gradient ascent procedure. Our algorithm optimizes the expert parameters and simultaneously obtains the feature subsets for each expert. We observe that the MoE trained using the new EM algorithm not only outperforms a single classifier baseline but also the vanilla MoE. We perform further data analysis and interpret the weights assigned to each expert during inference. Also, we obtain a different feature subset per expert in the mixture. This illustrates feature use based on location of the data point in the feature space.
Rahul Gupta 0001, Kartik Audhkhasi, Shri Narayanan
ICASSP3
2015 Redundancy analysis of behavioral coding for couples therapy and improved estimation of behavior from noisy annotations
abstract
Assessment and quantification of behavior is an important research objective in the recently developed field of behavioral signal processing. This paper focuses on the estimation of behavior from noisy human assessment. It aims to address the redundancy of behavioral descriptors for couples therapy by introducing a lower-dimensional representation of the behavioral space. We present an improved method for estimating the ground truth of behavioral ratings from assessment by multiple experts or annotators. The results show improved estimation performance using the proposed method and provide an insightful analysis of reconstruction error and decorrelation of annotator bias in the reduced behavioral space.
Md. Nasir, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan
ICASSP4
2015 Improvements to the IBM speech activity detection system for the DARPA RATS program
abstract
In this paper we describe improvements to the IBM speech activity detection (SAD) system for the third phase of the DARPA RATS program. The progress during this final phase comes from jointly training convolutional and regular deep neural networks with rich time-frequency representations of speech. With these additions, the phase 3 system reduces the equal error rate (EER) significantly on both of the program's development sets (relative improvements of 20% on dev1 and 7% on dev2) compared to an earlier phase 2 system. For the final program evaluation, the newly developed system also performs well past the program target of 3% Pmissat 1% Pfawith a performance of 1.2% Pmissat 1% Pfaand 0.3% Pfaat 3% Pmiss.
Samuel Thomas 0001, George Saon, Maarten Van Segbroeck, Shri Narayanan
ICASSP4
2015 Modeling mutual influence of multimodal behavior in affective dyadic interactions
abstract
To accomplish effective communication, interaction partners generally adapt their verbal and non-verbal behavior to that of their interlocutors. This behavior adaptation is often modulated by the underlying emotional states of partners. Modeling such mutual behavioral influence is critical for emotion characterization in an interaction. In this paper, we focus on explicitly modeling the mutual influence of multimodal behavior (speech and hand gesture) in affective dyadic interactions. In our framework, the behavior adaptation in each interaction is modeled by an interaction matrix which assembles all behavioral information on the path between the dyad's behavior. Experimental results show that our modeling approach can significantly improve the performance of emotion recognition. We further investigate the properties of the interaction model. Analysis results reveal that the entrainment effect of dyad's behavior can be better embodied by interaction modeling, and that the interaction patterns captured by interaction matrices are dependent on the emotional states of interaction partners. These observations corroborate the validity of our interaction model for capturing emotion-dependent mutual influence of dyad's interaction behavior.
Shri Narayanan
ICASSP2
2015 Gender Representation in Cinematic Content: A Multimodal Approach
abstract
The goal of this paper is to enable an objective understanding of gender portrayals in popular films and media through multimodal content analysis. An automated system for analyzing gender representation in terms of screen presence and speaking time is developed. First, we perform independent processing of the video and the audio content to estimate gender distribution of screen presence at shot level, and of speech at utterance level. A measure of the movie's excitement or intensity is computed using audiovisual features for every scene. This measure is used as a weighting function to combine the gender-based screen/speaking time information at shot/utterance level to compute gender representation for the entire movie. Detailed results and analyses are presented on seventeen full length Hollywood movies.
Tanaya Guha, Che-Wei Huang, Naveen Kumar 0004, Shri Narayanan
ICMI5
2015 A discriminative reliability-aware classification model with applications to intelligibility classification in pathological speech
Naveen Kumar 0004, Shri Narayanan
INTERSPEECH2
2015 Automated evaluation of non-native English pronunciation quality: combining knowledge- and data-driven features at multiple time scales
abstract
Automatically evaluating pronunciation quality of non-native speech has seen tremendous success in both research and com-mercial settings, with applications in L2 learning. In this paper, submitted for the INTERSPEECH 2015 Degree of Nativeness Sub-Challenge, this problem is posed under a challenging cross-corpora setting using speech data drawn from multiple speakers from a variety of language backgrounds (L1) reading different English sentences. Since the perception of non-nativeness is re-alized at the segmental and suprasegmental linguistic levels, we explore a number of acoustic cues at multiple time scales. We experiment with both data-driven and knowledge-inspired fea-tures that capture degree of nativeness from pauses in speech, speaking rate, rhythm/stress, and goodness of phone pronunci-ation. One promising finding is that highly accurate automated assessment can be attained using a small diverse set of intuitive and interpretable features. Performance is further boosted by smoothing scores across utterances from the same speaker; our best system significantly outperforms the challenge baseline.
Matthew Black, Daniel Bone, Z.-I. Skordilis, Rahul Gupta 0001, Pavlos Papadopoulos, Sandeep Nallan Chakravarthula, Bo Xiao 0003, Maarten Van Segbroeck, Jangwon Kim, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH12
2015 Acoustic-prosodic correlates of 'awkward' prosody in story retellings from adolescents with autism
abstract
. Acoustic-prosodic features are also able to significantly differentiate subjects with ASD from typically developing (TD) subjects in a classification task, emphasizing the potential of automated methods for diagnostic efficiency and clarity.
Daniel Bone, Matthew Black, Anil Ramakrishna, Ruth B. Grossman, Shri Narayanan
INTERSPEECH5
2015 A dialog act tagging approach to behavioral coding: a case study of addiction counseling conversations
Dogan Can, David C. Atkins, Shri Narayanan
INTERSPEECH3
2015 Predicting therapist empathy in motivational interviews using language features inspired by psycholinguistic norms
abstract
Therapist language plays a critical role in influencing the overall quality of psychotherapy. Notably, it is a major contributor to the perceived level of empathy expressed by therapists, a primary measure for judging their efficacy. We explore psycholinguistics inspired features for predicting therapist empathy. These features model language which conveys information about affective and cognitive processes, which is central to the therapist expressing understanding of the patient’s perspective. We describe the dimensional features obtained based on psycholinguisitic norms, and their application to predicting empathy expressed in motivational interviewing sessions for addiction counseling. We compare these to standard lexical features (n-grams) and demonstrate that these features contain complementary information for predicting therapist empathy. The highest empathy prediction results achieved are 75.28% UAR and 0.6112 Spearman’s correlation.
James Gibson, Nikos Malandrakis, Francisco Romero, David C. Atkins, Shri Narayanan
INTERSPEECH5
2015 Analysis and modeling of the role of laughter in motivational interviewing based psychotherapy conversations
Rahul Gupta 0001, Theodora Chaspari, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan
INTERSPEECH5
2015 Automatic estimation of parkinson's disease severity from diverse speech tasks
abstract
The need for reliable, scalable and efficient diagnosis of Parkin-son’s Disease (PD) is a major clinical need. Automating the diagnosis can lead to more accurate and objective predictions as well as provide insights regarding the nature of Parkinson’s condition. This paper proposes a fully automated system to rate the severity (UPDRS-III scale) of PD from patients ’ speech. Specifically, the system captures atypicalities in an individ-ual’s voice when performing multiple diverse speaking tasks and makes a unified prediction of the PD severity. The perfor-mance is tested in a cross-data setting, with different subjects and dissimilar recording conditions. Results indicate that (i) effective features vary depending on the nature of the specific speech task, (ii) additional novel feature sets to detect distor-tions in Parkinson’s speech significantly improve the prediction accuracy from the Interspeech15 Challenge baseline system and (iii) our fusion system based on an unsupervised clustering tech-nique also improves the accuracy. Our system incorporates i-vector and functionals for segmental features, non-linear time series features, speech rhythm and automatic speech recogni-tion decoding based features. By its application on the Inter-speech15 eating condition challenge, the system also shows its potential for detecting other sources of speech variability.
Jangwon Kim, Md. Nasir, Rahul Gupta 0001, Maarten Van Segbroeck, Daniel Bone, Matthew Black, Z.-I. Skordilis, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH10
2015 An analysis of the relationship between signal-derived vocal arousal score and human emotion production and perception
abstract
Bone et al. recently proposed an unsupervised signal-derived vocal arousal score (VC-AS) based on fusion of three intuitive acoustic features, i.e., pitch, intensity, and HF500, and have shown the effectiveness of quantifying humans perceptual rat-ings of arousal robustly across multiple corpora. Due to the readily-applicable nature of the system, this objective quantifi-cation scheme could foresee-ably be used in multiple fields of behavioral science as an objective measure of affect. In this work, we investigate in detail the relationship of this signal-derived measure to both intended arousal expression (i.e., pro-duction aspect) and perceived arousal rating (i.e., perception as-pect). On the perception side, our results in three databases (EMA, VAM, and IEMOCAP) indicate that VC-AS agrees with mean perception at least as well as an average individ-ual rater does. Regarding production, we demonstrate that in-tended arousal correlates more with VC-AS than mean percep-tion (EMA and IEMOCAP). We also show that VC-AS corre-lates more with intended arousal than perceived arousal (EMA); this finding is quite surprising given that the framework is sup-ported by extensive affective perception studies, although there is physiological motivation as well. Implication for utilizing VC-AS for novel scientific study in engineering (e.g., to miti-gate subjectivity) is further discussed. Index Terms: vocal arousal rating, affective perception, affec-tive production
Chi-Chun Lee, Daniel Bone, Shri Narayanan
INTERSPEECH3
2015 Therapy language analysis using automatically generated psycholinguistic norms
abstract
Lexical norms, normative and usually numeric, ratings of word meaning are popular tools in research domains relating to human expression and perception of language, especially with regards to emotion. In this paper we are proposing an algorithm of psycholinguistic norm expansion capable of generating high quality norms representing aspects of language beyond emotion, including language concreteness and indicators of age and gender association. Starting from small manually annotated norm lexica, continuous norms for new words are estimated using semantic similarity and a simple linear model along eleven expression-related dimensions. The model is shown to achieve state of the art level performance of word norm estimation. To investigate the potential of these norms as analysis tools of more complex phenomena we use them to investigate the differences in therapist speech in sessions conducted by practitioners adhering to the psychoanalytic and client-centered schools of therapy.
Nikos Malandrakis, Shri Narayanan
INTERSPEECH2
2015 Still together?: the role of acoustic features in predicting marital outcome
abstract
The assessment and prediction of marital outcome in couple therapy has intrigued many clinical psychologists. In this work, we analyze the significance of various acoustic features extracted from couples ’ spoken interaction in predicting the success or failure of their marriage. We also investigate whether speech acoustic features can provide complementary information to behavioral descriptions or codes provided by human experts (e.g., relationship satisfaction, blame patterns, global negativity). We formulate marital outcome prediction as both binary (improvement vs. no improvement) and multiclass (different levels of improvement) classification problem. Our experiments show that acoustic features can predict marital outcome more accurately than those based on behavioral descriptors provided by human experts. We also find that dialog turn-level acoustic features generally perform better than frame-level signal descriptors. This observation supports the notion that the impact of the behavior of one interlocutor on the other is more important than the behavior itself looked in isolation. Finally, acoustic features together with human-derived behavioral codes show the best performance in outcome prediction, suggesting some complementarity in the information captured by these behavioral representations.
Md. Nasir, Bo Xiao 0003, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou
INTERSPEECH5
2015 Experimental assessment of the tongue incompressibility hypothesis during speech production
abstract
The human tongue is an important organ for speech production. Its deformation and motion control the shape of the vocal tract significantly and thereby the acoustic properties of the speech signal produced. Thus, much effort in the speech research com-munity has been directed towards its biomechanical modeling. A common assumption incorporated into many models of the human tongue is the tissue incompressibility hypothesis: the tongue is considered a muscular hydrostat and therefore its vol-ume should remain constant regardless of its posture. To the best of our knowledge, experimental assessment of the constant volume hypothesis during actual speech production is limited. In this work, the aim is to experimentally assess the incom-pressibility hypothesis during actual speech production using a dataset of volumetric Magnetic Resonance (MR) images of 17 subjects sustaining contextualized continuants (27 continu-ants per subject). A seeded region growing based algorithm is used to segment the tongue and calculate its volume. Then the intra-subject variability of the tongue volume along the differ-ent tongue postures is examined. Within the accuracy of our tongue volume measurements, our empirical results seem con-sistent with the incompressibility hypothesis. Index Terms: speech production, tongue volume, muscular hy-drostat, tissue incompressibility, volumetric MRI
Z.-I. Skordilis, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2015 Ensemble of Gaussian mixture localized neural networks with application to phone recognition
Ruchir Travadi, Shri Narayanan
INTERSPEECH2
2015 Learning a speech manifold for signal subspace speech denoising
Colin Vaz, Shri Narayanan
INTERSPEECH2
2015 Analyzing speech rate entrainment and its relation to therapist empathy in drug addiction counseling
Bo Xiao 0003, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH5
2015 Automatic intelligibility classification of sentence-level pathological speech
Jangwon Kim, Naveen Kumar 0004, Andreas Tsiartas, Ming Li 0026, Shri Narayanan
Comput. Speech Lang.5
2015 Rapid Language Identification
abstract
A critical challenge to automatic language identification (LID) is achieving accurate performance with the shortest possible speech segment in a rapid fashion. The accuracy to correctly identify the spoken language is highly sensitive to the duration of speech and is bounded by the amount of information available. The proposed approach for rapid language identification transforms the utterances to a low dimensional i-vector representation upon which language classification methods are applied. In order to meet the challenges involved in rapidly making reliable decisions about the spoken language, a highly accurate and computationally efficient framework of i-vector extraction is proposed. The LID framework integrates the approach of universal background model (UBM) fused total variability modeling. UBM-fused modeling yields the estimation of a more discriminant, single i-vector space. This way, it is also a computationally more efficient alternative than system level fusion. A further reduction in equal error rate is achieved by training the i-vector model on long duration speech utterances and by the deployment of a robust feature extraction scheme that aims to capture the relevant language cues under various acoustic conditions. Evaluation results on the DARPA RATS data corpus suggest the potential of performing successful automated language identification at the level of one second of speech or even shorter duration.
Maarten Van Segbroeck, Ruchir Travadi, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Head Motion Modeling for Human Behavior Analysis in Dyadic Interaction
abstract
This paper presents a computational study of head motion in human interaction, notably of its role in conveying interlocutors' behavioral characteristics. Head motion is physically complex and carries rich information; current modeling approaches based on visual signals, however , are still limited in their ability to adequately capture these important properties. Guided by the methodology of kinesics , we propose a data-driven approach to identify typical head motion patterns. The approach follows the steps of first segmenting motion events, then parametrically representing the motion by linear predictive features, and finally generalizing the motion types using Gaussian mixture models. The proposed approach is experimentally validated using video recordings of communication sessions from real couples involved in a couples therapy study. In particular we use the head motion model to classify binarized expert judgments of the interactants' specific behavioral characteristics where entrainment in head motion is hypothesized to play a role: Acceptance, Blame, Positive, and Negative behavior. We achieve accuracies in the range of 60% to 70% for the various experimental settings and conditions. In addition, we describe a measure of motion similarity between the interaction partners based on the proposed model. We show that the relative change of head motion similarity during the interaction significantly correlates with the expert judgments of the interactants' behavioral characteristics. These findings demonstrate the effectiveness of the proposed head motion model, and underscore the promise of analyzing human behavioral characteristics through signal processing methods.
Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan
IEEE Trans. Multim.4
2014 Hull detection based on largest empty sector angle with application to analysis of realtime MR images
abstract
We present a novel view of the hull detection problem in two dimensions. Our proposed method is based on the principle of finding Pareto optimal boundaries and extends it to the general problem of finding a hull for a given set of points. We first compute the largest empty sector angle (LESA) score for each point. The desired hull can then be obtained as a super-level set of this score. We show how the proposed representation is related to a convex hull and demonstrate the flexibility it provides in choosing the geometry of the hull. As a target application we also present a head movement correction technique for real-time MR images of the dynamic vocal tract.
Naveen Kumar 0004, Shri Narayanan
ICASSP2
2014 Fusion of diverse denoising systems for robust automatic speech recognition
abstract
We present a framework for combining different denoising front-ends for robust speech enhancement for recognition in noisy conditions. This is contrasted against results of optimally fusing diverse parameter settings for a single denoising algorithm. All frontends in the latter case exploit the same denoising algorithm, which combines harmonic decomposition, with noise estimation and spectral subtraction. The set of associated parameters involved in these steps are dependent on the noise conditions. Rather than explicitly tuning them, we suggest a strategy that tries to account for the trade-off between average word error rate and diversity to find an optimal subset of these parameter settings. We present the results on Aurora4 database and also compare against traditional speech enhancement methods e.g. Wiener filtering and spectral subtraction.
Naveen Kumar 0004, Maarten Van Segbroeck, Kartik Audhkhasi, Peter Drotár, Shri Narayanan
ICASSP5
2014 Semi-supervised term-weighted value rescoring for keyword search
abstract
We present a semi-supervised algorithm for rescoring the output of a speech keyword search (KWS) system. Conventional loss functions such as squared-error and logistic loss are not suitable for optimizing the commonly-used KWS term-weighted value (TWV) performance metric. We derive a novel concave modified logistic log-likelihood function which lower-bounds TWV. We then use a manifold-regularized kernel classifier that maximizes this lower-bound. A manifold regularization term in our objective function uses available unlabeled speech data and makes our approach semi-supervised. This term is particularly useful for KWS in low-resource languages and ensures that the predicted keyword confidence scores are smooth on a low-dimensional manifold in the feature space. We conduct KWS experiments on the IARPA Babel Vietnamese task and show performance improvements in terms of the maximum TWV (MTWV). Our estimated confidence score is complementary with respect to the ASR posterior score and gives MTWV improvement upon interpolation with it.
Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shri Narayanan
ICASSP4
2014 Barista: A framework for concurrent speech processing by usc-sail
abstract
We present Barista, an open-source framework for concurrent speech processing based on the Kaldi speech recognition toolkit and the libcppa actor library. With Barista, we aim to provide an easy-to-use, extensible framework for constructing highly customizable concurrent (and/or distributed) networks for a variety of speech processing tasks. Each Barista network specifies a flow of data between simple actors, concurrent entities communicating by message passing, modeled after Kaldi tools. Leveraging the fast and reliable concurrency and distribution mechanisms provided by libcppa, Barista lets demanding speech processing tasks, such as real-time speech recognizers and complex training workflows, to be scheduled and executed on parallel (and/or distributed) hardware. Barista is released under the Apache License v2.0.
Dogan Can, James Gibson, Colin Vaz, Panayiotis G. Georgiou, Shri Narayanan
ICASSP5
2014 A non-homogeneous poisson process model of Skin Conductance Responses integrated with observed regulatory behaviors for Autism intervention
abstract
Early intervention in individuals with Autism Spectrum Disorder (ASD) can improve core and associated symptoms and facilitate skills that increase social opportunities. However, determining effective intervention success in this population, and the mechanisms that produce it, is currently restricted to observable behavior. The need of therapy assessment metrics beyond traditional behavioral criteria, led to the use of physiological signals for capturing child-therapist internal dynamics during an intervention session. Internal physiological states were measured through Electrodermal Activity (EDA) and modeled in relation to observed self- and co-regulatory behaviors. A common measure of EDA, Skin Conductance Response (SCR), was the primary signal of interest and assumed to form a non-homogeneous Poisson Process whose rate function is determined by observed regulatory behaviors. Through likelihood and residual goodness of fit analysis, statistical tests and classification tasks, our results indicate that SCR changes and observable behavior in child-therapist dyads are temporally associated and the estimated model parameters can be linked to the types of regulation stimuli.
Theodora Chaspari, Matthew S. Goodwin, Oliver Wilder-Smith, Amanda Gulsrud, Charlotte A. Mucchetti, Connie Kasari, Shri Narayanan
ICASSP7
2014 Learning multiple concepts with incremental diverse density
abstract
We present a novel method of learning multiple disjunct concepts with diverse density using an incremental approach. We demonstrate that by maximizing the diverse density over individual target concept points and minimizing the probability of their intersection, concepts can be learned incrementally. This method reduces the complexity of the algorithm from factorial, with respect to the number of targets, to exponential order. We demonstrate that this greedy approach successfully learns disjunctive target concepts with competitive classification accuracy on a benchmark multiple instance learning dataset in comparison to other common diverse density approaches. We also introduce a novel application of the multiple instance learning framework to an emotion recognition task using prosodic and spectral speech features.
James Gibson, Shri Narayanan
ICASSP2
2014 Training ensemble of diverse classifiers on feature subsets
abstract
Ensembles of diverse classifiers often out-perform single classifiers as has been well-demonstrated across several applications. Existing training algorithms either learn a classifier ensemble on pre-defined feature sets or independently perform classifier training and feature selection. Neither of these schemes is optimal. We pose feature subset selection and training of diverse classifiers on selected subsets as a joint optimization problem. We propose a novel greedy algorithm to solve this problem. We sequentially learn an ensemble of classifiers where each subsequent classifier is encouraged to learn data instances misclassified by previous classifiers on a concurrently selected feature set. Our experiments on synthetic and real-world data sets show the effectiveness of our algorithm. We observe that ensembles trained by our algorithm performs better than both a single classifier and an ensemble of classifiers learnt on pre-defined feature sets. We also test our algorithm as a feature selector on a synthetic dataset to filter out irrelevant features.
Rahul Gupta 0001, Kartik Audhkhasi, Shri Narayanan
ICASSP3
2014 Affective language model adaptation via corpus selection
abstract
Motivated by methods used in language modeling and grammar induction, we propose the use of pragmatic constraints and perplexity as criteria to filter the unlabeled data used to generate the semantic similarity model. We investigate unsupervised adaptation algorithms of the semantic-affective models proposed in [1, 2]. Affective ratings at the utterance level are generated based on an emotional lexicon, which in turn is created using a semantic (similarity) model estimated over raw, unlabeled text. The proposed adaptation method creates task-dependent semantic similarity models and task-dependent word/term affective ratings. The proposed adaptation algorithms are tested on anger/distress detection of transcribed speech data and sentiment analysis in tweets showing significant relative classification error reduction of up to 10%.
Nikos Malandrakis, Alexandros Potamianos, Kean J. Hsu, Kalina N. Babeva, Michelle C. Feng, Gerald C. Davison, Shri Narayanan
ICASSP7
2014 A supervised signal-to-noise ratio estimation of speech signals
abstract
This paper introduces a supervised statistical framework for estimating the signal-to-noise (SNR) ratio of speech signals. Information on how noise corrupts a signal can help us compensate for its effects, especially in real life applications where the usual assumption of white Gaussian noise does not hold and speech boundaries in the signal are not known. We use features from which we can detect speech regions in a signal, without using Voice Activity Detection, and estimate the energies of those regions. Then we use these features to train ordinary least squares regression models for various noise types. We compare this supervised method with state-of-the-art SNR estimation algorithms and show its superior performance with respect to the tested noise types.
Pavlos Papadopoulos, Andreas Tsiartas, James Gibson, Shri Narayanan
ICASSP4
2014 Simplified and supervised i-vector modeling for speaker age regression
abstract
We propose a simplified and supervised i-vector modeling scheme for the speaker age regression task. The supervised i-vector is obtained by concatenating the label vector and the linear regression matrix at the end of the mean super-vector and the i-vector factor loading matrix, respectively. Different label vector designs are proposed to increase the robustness of the supervised i-vector models. Finally, Support Vector Regression (SVR) is deployed to estimate the age of the speakers. The proposed method outperforms the conventional i-vector baseline for speaker age estimation. A relative 2.4% decrease in Mean Absolute Error and 3.33% increase in correlation coefficient is achieved using supervised i-vector modeling using different label designs on the NIST SRE 2008 dataset male part.
Prashanth Gurunath Shivakumar, Ming Li 0026, Vedant Dhandhania, Shri Narayanan
ICASSP4
2014 Classification of clean and noisy bilingual movie audio for speech-to-speech translation corpora design
abstract
Identifying suitable sources of bilingual audio and text data is a crucial part of statistical Speech to Speech (S2S) research and development. Movies, often dubbed in other languages, offer a good source for this purpose; but not all data are directly usable because of noise and other audio condition differences. Hence, automatically selecting the bilingual audio data that are suitable for analysis, and training S2S systems for specific environments becomes crucial. In this work, we extract bilingual speech segments from movies and aim at classifying segments as clean speech or speech with background noise (i.e. music, babble noise etc.). We examine various features in solving this problem and our best performing method delivers accuracy up to 87% in discriminating clean and noisy speech in bilingual data.
Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan
ICASSP4
2014 Energy-constrained minimum variance response filter for robust vowel spectral estimation
abstract
We propose the energy-constrained minimum-variance response (ECMVR) filter to perform robust spectral estimation of vowels. We modify the distortionless constraint of the minimum-variance distortionless response (MVDR) filter and add an energy constraint to its formulation to mitigate the influence of noise on the speech spectrum. We test our ECMVR filter on a vowel classification task with different background noises at various SNR levels. Results show that vowels are classified more accurately in certain noises using MFCC and PLP features extracted from the ECMVR spectrum compared to using features extracted from the FFT and MVDR spectra.
Colin Vaz, Andreas Tsiartas, Shri Narayanan
ICASSP3
2014 Power-spectral analysis of head motion signal for behavioral modeling in human interaction
abstract
We examine whether head motion can be used for predicting human expert's judgments of behavioral characteristics relevant to the couples therapy domain. Specifically we predict “high” or “low” presence of several behavioral characteristics such as “Blame” that are discerned by human experts, through data-driven clustering of the head motion signal based on power-spectral features. We employ the distribution of motion samples in each cluster for behavior judgment prediction. We find clustering horizontal and vertical motion separately is superior to combined clustering in predicting behavior. The performance of gender-specific and gender-independent clustering of head motion is comparable in average while different for each gender. The proposed power-spectral features outperform linear prediction features in average. Using data from a clinical study of distressed couples, we empirically show that the derived clusters quantize head motion into meaningful types that relate to interpretable behavior characteristics. These findings demonstrate the feasibility of inferring behavior characteristics from head motion signals.
Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan
ICASSP4
2014 Analysis of interaction attitudes using data-driven hand gesture phrases
abstract
Hand gesture is one of the most expressive, natural and common types of body language for conveying attitudes and emotions in human interactions. In this paper, we study the role of hand gesture in expressing attitudes of friendliness or conflict towards the interlocutors during interactions. We first employ an unsupervised clustering method using a parallel HMM structure to extract recurring patterns of hand gesture (hand gesture phrases or primitives). We further investigate the validity of the derived hand gesture phrases by examining the correlation of dyad's hand gesture for different interaction types defined by the attitudes of interlocutors. Finally, we model the interaction attitudes with SVM using the dynamics of the derived hand gesture phrases over an interaction. The classification results are promising, suggesting the expressiveness of the derived hand gesture phrases for conveying attitudes and emotions.
Angeliki Metallinou, Engin Erzin, Shri Narayanan
ICASSP4
2014 Graph-based approach for motion capture data representation and analysis
abstract
Providing better representation methods for motion capture data can lead to improved performance in terms of classification, recognition, synthesis and dimensionality reduction. In this paper, we propose a novel representation method inspired by algebraic and spectral graph theoretic concepts. Our proposed method represents motion data in a space constructed with bases for skeleton-like graphs. We introduce two criteria and as well as its influences on the generated bases. With experiments on CMU MoCap database, we will also discuss how this method may act as a good preprocessing tool in order to enhance the further analysis steps on MoCap data.
Jiun-Yu Kao, Antonio Ortega, Shri Narayanan
ICIP3
2014 Gesture dynamics modeling for attitude analysis using graph based transform
abstract
Gesture dynamic pattern is an essential indicator of emotions or attitudes during human communication. However, there might exist great variability of gesture dynamics among gesture sequences within the same emotion, which form a major obstacle to detect emotion from body motion in general interpersonal interactions. In this paper, we propose a graph-based framework for modeling gesture dynamics towards attitude recognition. We demonstrate that the dynamics derived from a weighted graph based method provide a better separation between distinct emotion classes and maintain less variability within the same emotion class. This helps capture salient dynamic patterns for specific emotions by removing interaction-dependent variations. In this framework, we represent each gesture sequence as an undirected graph of connected gesture units and use the graph-based transform to generate features to describe gesture dynamics. In our experiments, we apply the graph-based dynamics for attitude recognition, i.e., classifying the attitude of an individual as friendly or conflictive. Experimental results verify the effectiveness of our approach.
Antonio Ortega, Shri Narayanan
ICIP3
2014 A real-time MRI study of articulatory setting in second language speech
abstract
Previous work has shown that languages differ in their articulatory setting, the postural configuration that the vocal tract articulators tend to adopt when they are not engaged in any active speech gesture, and that this posture might be specified as part of the phonological knowledge speakers have of the language. This study tests whether the articulatory setting of a language can be acquired by non-native speakers. Three native speakers of German who had learned English as a second language were imaged using real-time MRI of the vocal tract while reading passages in German and English, and features that capture vocal tract posture were extracted from the inter-speech pauses in their native and non-native languages. Results show that the speakers exhibit distinct inter-speech postures in each language, with a lower and more retracted tongue in English, consistent with classic descriptions of the differences between the German and the English articulatory settings. This supports the view that non-native speakers may acquire relevant features of the articulatory setting of a second language, and also lends further support to the idea that articulatory setting is part of a speaker’s phonological competence in a language. Index Terms: articulatory setting, speech production, second language speech acquisition, real-time MRI.
Andrés Benítez, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2014 An investigation of vocal arousal dynamics in child-psychologist interactions using synchrony measures and a conversation-based model
abstract
Researchers from various disciplines are concerned with the study of affective phenomena, especially arousal. Expressed affective modulations, which reflect both an individual’s in-ternal state and external factors, are central to the commu-nicative process. Bone et al. developed a robust, unsuper-vised (rule-based) method which provides a scale-continuous, bounded arousal rating from the vocal signal. In this study, we investigate the joint-dynamics of child and psychologist vocal arousal in autism spectrum disorder (ASD) diagnostic interac-tions. Arousal synchrony is assessed with multiple methods. Results indicate that children with higher ASD severity tend to lead the arousal dynamics more, seemingly because the children aren’t as responsive to the psychologist’s affective modulations. A vocal arousal model is also proposed which incorporates so-cial and conversational constructs. The model captures conver-sational signal relations, and is able to distinguish between high and low ASD severity at accuracies well-above chance.
Daniel Bone, Chi-Chun Lee, Alexandros Potamianos, Shri Narayanan
INTERSPEECH4
2014 Robust language identification using convolutional neural network features
abstract
The language identification (LID) task in the Robust Automatic Transcription of Speech (RATS) program is challenging due to the noisy nature of the audio data collected over highly degraded radio communication channels as well as the use of short duration speech segments for testing. In this paper, we report the recent advances made in the RATS LID task by using bottleneck features from a convolutional neural network (CNN). The CNN, which is trained with labelled data from one of target languages, generates bottleneck features which are used in a Gaussian mixture model (GMM)-ivector LID system. The CNN bottleneck features provide substantial complimentary information to the conventional acoustic features even on languages not seen in its training. Using these bottleneck features in conjunction with acoustic features, we obtain significant improvements (average relative improvements of 25% in terms of equal error rate (EER) compared to the corresponding acoustic system) for the LID task. Furthermore, these improvements are consistent for various choices of acoustic features as well as speech segment durations.
Sriram Ganapathy, Kyu Jeong Han, Samuel Thomas 0001, Mohamed Kamal Omar, Maarten Van Segbroeck, Shri Narayanan
INTERSPEECH6
2014 Comparing time-frequency representations for directional derivative features
abstract
We compare the performance of Directional Derivatives features for automatic speech recognition when extracted from different time-frequency representations. Specifically, we use the short-time Fourier transform, Mel-frequency, and Gammatone spectrograms as a base from which we extract spectrotemporal modulations. We then assess the noise robustness of each representation with varied number of frequency bins and dynamic range compression schemes for both word and phone recognition. We find that the choice of dynamic range compression approach has the most significant impact on recognition performance. Whereas, the performance differences between perceptually motivated filter-banks are minimal in the proposed framework. Furthermore, this work presents significant gains in speech recognition accuracy for low SNRs over MFCCs, GFCCs, and Directional Derivatives extracted from the log-Mel spectrogram.
James Gibson, Maarten Van Segbroeck, Shri Narayanan
INTERSPEECH3
2014 Variable Span disfluency detection in ASR transcripts
Rahul Gupta 0001, Sankaranarayanan Ananthakrishnan, Shri Narayanan
INTERSPEECH4
2014 Predicting client's inclination towards target behavior change in motivational interviewing and investigating the role of laughter
Rahul Gupta 0001, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan
INTERSPEECH4
2014 Unsupervised speaker diarization using riemannian manifold clustering
abstract
We address the problem of speaker clustering for robust unsupervised speaker diarization. We model each speakerhomogeneous segment as one single full multivariate Gaussian probability density function (pdf) and take into consideration the Riemannian property of Gaussian pdfs. By assuming that segments from different speakers lie on different (possibly intersected) sub-manifolds of the manifold of Gaussian pdfs, we formulate the original problem as a Riemannian manifold clustering problem. To apply the computationally simple Riemannian locally linear embedding (LLE) algorithm, we impose a constraint on the length of each segment so as to ensure the fitness of single-Gaussian modeling and to increase the chance that all k-nearest neighbors of a pdf are from the same submanifold (speaker). Experiments on the microphone-recorded conversational interviews from NIST 2010 speaker recognition evaluation set demonstrate promising results of less than 1%
Che-Wei Huang, Bo Xiao 0003, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH4
2014 A study of invariant properties and variation patterns in the converter/distributor model for emotional speech
abstract
Invariant properties of vocal organ controls at an abstract level are crucial for better understanding and modeling of the speech production mechanism. Despite the large variability of articulatory movements at the execution level, the Converter/Distributor (C/D) model provides a systematic and comprehensive framework for the prosodic organization of speech production, based on the invariant properties of articulatory movements with the concept of “iceberg” region. The goal of this paper is two-fold: (i) to examine the invariant properties in the C/D model in emotional speech, and (ii) to understand emotion-dependent variation patterns of important parameters in the C/D model framework. Experimental results support the validity of strong linear relationship between the speed and excursion of critical articulatorsat the iceberg points for emotional speech. Also, emotion-dependent variation patterns of the C/D model parameters, (e.g., relatively smaller “shadow” angle and greater syllable magnitude for happiness) are reported. Finally, the emotion-dependent relationships between the abstract-level C/D model parameters and the surface-level parameters of the invariant articulatory behaviors are reported. Index Terms: emotional variation, temporal organization of speech, C/D model, invariant property
Jangwon Kim, Donna Erickson, Sungbok Lee, Shri Narayanan
INTERSPEECH4
2014 Estimation of the movement trajectories of non-crucial articulators based on the detection of crucial moments and physiological constraints
abstract
This study develops a mathematical model that estimates the movements of (linguistically) non-crucial articulators in speech production, which provides a systematic way to study the relationship between the behaviors of crucial and non-crucial articulators; crucial articulators are those essential for realizing a speech task. The underlying assumption of our model is that non-crucial articulatory movements are governed by the physiological constraints in relation to the corresponding crucial articulators as well as by the contextual constraint from the nearest crucial time of the non-crucial articulator. These constraints have been generally assumed in the speech production literature, but they have not been incorporated directly into articulatory models. The crucial articulatory moments in an utterance are automatically determined by a novel forced-alignment algorithm for articulatory trajectories, which uses the inherent physical properties of crucial articulatory movements. Experimental results suggest that the proposed algorithm is capable of estimating non-crucial articulatory positions well in both neutral and emotional speech, significantly better than the simple interpolation of crucial points. Index Terms: non-crucial articulators, articulatory modeling, emotional speech
Jangwon Kim, Sungbok Lee, Shri Narayanan
INTERSPEECH3
2014 Selection of optimal vocal tract regions using real-time magnetic resonance imaging for robust voice activity detection
abstract
Real time magnetic resonance imaging (rtMRI) enables direct video capture of the moving vocal tract concurrent with audio signal providing valuable data for speech research. We consider a multimodal approach to voice activity detection (VAD) in the rtMRI recording that uses audio signal as well as MRI image sequence. The degraded quality of the audio recorded in the scanner motivates this multimodal scheme for robust VAD. Optimal regions in the MRI image are selected for performing VAD with a novel algorithm. VAD experiments using rtMRI data of two male and two female subjects show that VAD performance using optimally selected regions from MRI images is comparable to that using only audio signal. The optimal regions turn out to be parts of jaw, velum, glottis and lips. VAD performance using audio signal and MRI image sequence together is found to be significantly better (∼14% absolute improvement in VAD accuracy) than that using the audio only when the audio is contaminated with additive noise at low SNR. Index Terms: voice activity detection, vocal tract imaging, optimal region selection.
Abhay Prasad, Prasanta Kumar Ghosh, Shri Narayanan
INTERSPEECH3
2014 Motor control primitives arising from a learned dynamical systems model of speech articulation
abstract
We present a method to derive a small number of speech motor control “primitives” that can produce linguisticallyinterpretable articulatory movements. We envision that such a dictionary of primitives can be useful for speech motor control, particularly in finding a low-dimensional subspace for such control. First, we use the iterative Linear Quadratic Gaussian with Learned Dynamics (iLQG-LD) algorithm to derive (for a set of utterances) a set of stochastically optimal control inputs to a learned dynamical systems model of the vocal tract that produces desired movement sequences. Second, we use a convolutive Nonnegative Matrix Factorization with sparseness constraints (cNMFsc) algorithm to find a small dictionary of control input primitives that can be used to reproduce the aforementioned optimal control inputs that produce the observed articulatory movements. The method performs favorably on both qualitative and quantitative evaluations conducted on synthetic data produced by an articulatory synthesizer. Such a primitivesbased framework could help inform theories of speech motor control and coordination. Index Terms: speech motor control, motor primitives, synergies, dynamical systems, iLQG, NMF.
Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH3
2014 UBM fused total variability modeling for language identification
abstract
This paper proposes Universal Background Model (UBM) fusion in the framework of total variability or i-vector modeling with the application to language identification (LID). The total variability subspace which is typically exploited to discriminate between the language classes of different speech recordings, is trained by combining the normalized Baum-Welch statistics of multiple UBMs. When the UBMs model a diverse set of feature representations, the method yields an i-vector representation which is more discriminant between the classes of interest. This approach is particularly useful when applied to shortduration utterances, and is a computationally less complex alternative to performance boosting as compared to system level fusion. We assess the performance of UBM fused total variability modeling on the task of robust language identification on short-duration utterances, as part of Phase-III of the DARPA RATS (Robust Automatic Transcription of Speech) program. Index Terms: language identification, i-vector representation, short-duration, noise robustness, RATS
Maarten Van Segbroeck, Ruchir Travadi, Shri Narayanan
INTERSPEECH3
2014 Classification of cognitive load from speech using an i-vector framework
abstract
The goal in this work is to automatically classify speakers ’ level of cognitive load (low, medium, high) from a standard battery of reading tasks requiring varying levels of working memory. This is a challenging machine learning problem because of the inherent difficulty in defining/measuring cognitive load and due to intra-/inter-speaker differences in how their effects are man-ifested in behavioral cues. We experimented with a number of static and dynamic features extracted directly from the audio signal (prosodic, spectral, voice quality) and from automatic speech recognition hypotheses (lexical information, speaking rate). Our approach to classification addressed the wide vari-ability and heterogeneity through speaker normalization and by adopting an i-vector framework that affords a systematic way to factorize the multiple sources of variability. Index Terms: computational paralinguistics, behavioral signal processing (BSP), prosody, ASR, i-vector, cognitive load
Maarten Van Segbroeck, Ruchir Travadi, Colin Vaz, Jangwon Kim, Matthew Black, Alexandros Potamianos, Shri Narayanan
INTERSPEECH7
2014 Modified-prior i-vector estimation for language identification of short duration utterances
abstract
In this paper, we address the problem of Language Identification (LID) on short duration segments. Current state-of-the-art LID systems typically employ total variability i-Vector modeling for obtaining fixed length representation of utterances. However, when the utterances are short, only a small amount of data is available, and the estimated i-Vector representation will con-sequently exhibit significant variability, making the identifica-tion problem challenging. In this paper, we propose novel tech-niques to modify the standard normal prior distribution of the i-Vectors, to obtain a more discriminative i-Vector extraction given the small amount of available utterance data. Improved performance was observed by using the proposed i-Vector esti-mation techniques on short segments of the DARPA RATS cor-pora, with lengths as small as 3 seconds.
Ruchir Travadi, Maarten Van Segbroeck, Shri Narayanan
INTERSPEECH3
2014 Enhancing audio source separability using spectro-temporal regularization with NMF
Colin Vaz, Dimitrios Dimitriadis, Shri Narayanan
INTERSPEECH3
2014 Joint filtering and factorization for recovering latent structure from noisy speech data
abstract
We propose a joint filtering and factorization algorithm to re-cover latent structure from noisy speech. We incorporate the minimum variance distortionless response (MVDR) formula-tion within the non-negative matrix factorization (NMF) frame-work to derive a single, unified cost function for both filtering and factorization. Minimizing this cost function jointly opti-mizes three quantities – a filter that removes noise, a basis ma-trix that captures latent structure in the data, and an activation matrix that captures how the elements in the basis matrix can be linearly combined to reconstruct input data. Results show that the proposed algorithm recovers the speech basis matrix from noisy speech significantly better than NMF alone or Wiener fil-tering followed by NMF. Furthermore, PESQ scores show that our algorithm is a viable choice for speech denoising. Index Terms: NMF, MVDR, denoising, filtering. 1.
Colin Vaz, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH3
2014 Modeling therapist empathy through prosody in drug addiction counseling
abstract
Empathy measures the capacity of the therapist to experience the same cognitive and emotional dispositions as the patient, and is a key quality factor in counseling. In this work we build computational models to infer the empathy of therapist using prosodic cues. We extract pitch, energy, jitter, shimmer and utterance duration from the speech signal, and normalize and quantize these features in order to estimate the distribution of certain prosodic patterns during each interaction. We find significant correlation between empathy and the distribution of prosodic patterns, and achieve 75% accuracy in classifying therapist empathy levels using this distribution. Experiment results suggest high pitch and energy of the therapist are negatively correlated with empathy. These observations agree with domain literature and human intuition.
Bo Xiao 0003, Daniel Bone, Maarten Van Segbroeck, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH7
2014 Analysis of emotional effect on speech-body gesture interplay
abstract
In interpersonal interactions, speech and body gesture channels are internally coordinated towards conveying communicative intentions. The speech-gesture relationship is influenced by the internal emotion state underlying the communication. In this paper, we focus on uncovering the emotional effect on the interrelation between speech and body gestures. We investigate acoustic features describing speech prosody (pitch and energy) and vocal tract configuration (MFCCs), as well as three types of body gestures, viz., head motion, lower and upper body motions. We employ mutual information to measure the coordination between the two communicative channels, and analyze the quantified speech-gesture link with respect to distinct levels of emotion attributes, i.e., activation and valence. The results reveal that the speech-gesture coupling is generally tighter for low-level activation and high-level valence, compared to high-level activation and low-level valence. We further propose a framework for modeling the dynamics of speech-gesture interaction. Experimental studies suggest that such quantified coupling representations can well discriminate different levels of activation and valence, reinforcing that emotions are encoded in the dynamics of the multimodal link. We also verify that the structures of the coupling representations are emotiondependent using subspace-based analysis. Index Terms: emotion attributes, body gesture, speech prosody, speech-gesture interplay, mutual information
Shri Narayanan
INTERSPEECH2
2014 Intoxicated speech detection: A fusion framework with speaker-normalized hierarchical functionals and GMM supervectors
Daniel Bone, Ming Li 0026, Matthew Black, Shri Narayanan
Comput. Speech Lang.4
2014 Computing vocal entrainment: A signal-derived PCA-based quantification scheme with application to affect analysis in married couple interactions
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan
Comput. Speech Lang.7
2014 Simplified supervised i-vector modeling with application to robust and efficient language identification and speaker verification
Ming Li 0026, Shri Narayanan
Comput. Speech Lang.2
2014 Robust Unsupervised Arousal Rating: A Rule-Based Framework withKnowledge-Inspired Vocal Features
abstract
Studies in classifying affect from vocal cues have produced exceptional within-corpus results, especially for arousal (activation or stress); yet cross-corpora affect recognition has only recently garnered attention. An essential requirement of many behavioral studies is affect scoring that generalizes across different social contexts and data conditions. We present a robust, unsupervised (rule-based) method for providing a scale-continuous, bounded arousal rating operating on the vocal signal. The method incorporates just three knowledge-inspired features chosen based on empirical and theoretical evidence. It constructs a speaker's baseline model for each feature separately, and then computes single-feature arousal scores. Lastly, it advantageously fuses the single-feature arousal scores into a final rating without knowledge of the true affect. The baseline data is preferably labeled as neutral, but some initial evidence is provided to suggest that no labeled data is required in certain cases. The proposed method is compared to a state-of-the-art supervised technique which employs a high-dimensional feature set. The proposed framework achieves highly-competitive performance with additional benefits. The measure is interpretable, scale-continuous as opposed to discrete, and can operate without any affective labeling. An accompanying Matlab tool is made available with the paper.
Daniel Bone, Chi-Chun Lee, Shri Narayanan
IEEE Trans. Affect. Comput.3
2014 Theoretical Analysis of Diversity in an Ensemble of Automatic Speech Recognition Systems
abstract
Diversity or complementarity of automatic speech recognition (ASR) systems is crucial for achieving a reduction in word error rate (WER) upon fusion using the ROVER algorithm. We present a theoretical proof explaining this often-observed link between ASR system diversity and ROVER performance. This is in contrast to many previous works that have only presented empirical evidence for this link or have focused on designing diverse ASR systems using intuitive algorithmic modifications. We prove that the WER of the ROVER output approximately decomposes into a difference of the average WER of the individual ASR systems and the average WER of the ASR systems with respect to the ROVER output. We refer to the latter quantity as the diversity of the ASR system ensemble because it measures the spread of the ASR hypotheses about the ROVER hypothesis. This result explains the trade-off between the WER of the individual systems and the diversity of the ensemble. We support this result through ROVER experiments using multiple ASR systems trained on standard data sets with the Kaldi toolkit. We use the proposed theorem to explain the lower WERs obtained by ASR confidence-weighted ROVER as compared to word frequency-based ROVER. We also quantify the reduction in ROVER WER with increasing diversity of the N-best list. We finally present a simple discriminative framework for jointly training multiple diverse acoustic models (AMs) based on the proposed theorem. Our framework generalizes and provides a theoretical basis for some recent intuitive modifications to well-known discriminative training criterion for training diverse AMs.
Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Analysis and Predictive Modeling of Body Language Behavior in Dyadic Interactions From Multimodal Interlocutor Cues
abstract
During dyadic interactions, participants adjust their behavior and give feedback continuously in response to the behavior of their interlocutors and the interaction context. In this paper, we study how a participant in a dyadic interaction adapts his/her body language to the behavior of the interlocutor, given the interaction goals and context. We apply a variety of psychology-inspired body language features to describe body motion and posture. We first examine the coordination between the dyad's behavior for two interaction stances: friendly and conflictive. The analysis empirically reveals the dyad's behavior coordination, and helps identify informative interlocutor features with respect to the participant's target body language features. The coordination patterns between the dyad's behavior are found to depend on the interaction stances assumed. We apply a Gaussian-Mixture-Model-based (GMM) statistical mapping in combination with a Fisher kernel framework for automatically predicting the body language of an interacting participant from the speech and gesture behavior of an interlocutor. The experimental results show that the Fisher kernel-based approach outperforms methods using only the GMM-based mapping, and using the support vector regression, in terms of correlation coefficient and RMSE. These results suggest a significant level of predictability of body language behavior from interlocutor cues.
Angeliki Metallinou, Shri Narayanan
IEEE Trans. Multim.3
2013 Joint training of interpolated exponential n-gram models
abstract
For many speech recognition tasks, the best language model performance is achieved by collecting text from multiple sources or domains, and interpolating language models built separately on each individual corpus. When multiple corpora are available, it has also been shown that when using a domain adaptation technique such as feature augmentation [1], the performance on each individual domain can be improved by training a joint model across all of the corpora. In this paper, we explore whether improving each domain model via joint training also improves performance when interpolating the models together. We show that the diversity of the individual models is an important consideration, and propose a method for adjusting diversity to optimize overall performance. We present results using word n-gram models and Model M, a class-based n-gram model, and demonstrate improvements in both perplexity and word-error rate relative to state-of-the-art results on a Broadcast News transcription task.
Abhinav Sethy, Stanley F. Chen, Ebru Arisoy, Bhuvana Ramabhadran, Kartik Audhkhasi, Shri Narayanan, Paul Vozila
ASRU6
2013 Using physiology and language cues for modeling verbal response latencies of children with ASD
abstract
Signal-derived measures can provide effective ways towards quantifying human behavior. Verbal Response Latencies (VRLs) of children with Autism Spectrum Disorders (ASD) during conversational interactions are able to convey valuable information about their cognitive and social skills. Motivated by the inherent gap between the external behavior and inner affective state of children with ASD, we study their VRLs in relation to their explicit but also implicit behavioral cues. Explicit cues include the children's language use, while implicit cues are based on physiological signals. Using these cues, we perform classification and regression tasks to predict the duration type (short/long) and value of VRLs of children with ASD while they interacted with an Embodied Conversational Agent (ECA) and their parents. Since parents are active participants in these triadic interactions, we also take into account their linguistic and physiological behaviors. Our results suggest an association between VRLs and these externalized and internalized signal information streams, providing complementary views of the same problem.
Theodora Chaspari, Daniel Bone, James Gibson, Chi-Chun Lee, Shri Narayanan
ICASSP5
2013 On-line genre classification of TV programs using audio content
abstract
Automatic genre classification of TV programs can benefit users in various ways such as allowing for rapid selection of multimedia content. In this paper, we introduce an on-line method that can classify genres of TV programs using audio content. We deploy an acoustic topic model (ATM) which was originally designed to capture contextual information embedded within audio segments. With a dataset based on RAI content, we perform both on-line and off-line classification; we segment audio signals with a fixed length and feed into the system for on-line classification tasks, while we use whole audio signals for off-line tasks. The off-line experimental results suggest that the proposed method using audio content yields competitive performance with conventional methods using audio-visual features and outperforms conventional audio-based approaches. The on-line results show promising results in classifying genre of TV programs with short segments and also suggest that ATM performs better than conventional GMM method if the length of audio segments is longer (>1 second).
Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan
ICASSP3
2013 Spatial and temporal alignment of multimodal human speech production data: Real time imaging, flesh point tracking and audio
abstract
In speech production research, the integration of articulatory data derived from multiple measurement modalities can provide rich description of vocal tract dynamics by overcoming the limited spatio-temporal representations offered by individual modalities. This paper presents a spatial and temporal alignment method between two promising modalities using a corpus of TIMIT sentences obtained from the same speaker: flesh point tracking from Electromagnetic Articulography (EMA) that offers high temporal resolution but sparse spatial information and real time Magnetic Resonance Imaging (MRI) that offers good spatial details but at lower temporal rates. Spatial alignment is done by using palate tracking of EMA, but distortion in MRI audio and articulatory data variability make temporal alignment challenging. This paper proposes a novel alignment technique using joint acoustic-articulatory features which combines dynamic time warping and automatic feature extraction from MRI images. Experimental results show that the temporal alignment obtained using this technique is better (12% relative) than that using acoustic feature only.
Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Shri Narayanan
ICASSP4
2013 Speaker verification using simplified and supervised i-vector modeling
abstract
This paper presents a simplified and supervised i-vector modeling framework that is applied in the task of robust and efficient speaker verification (SRE). First, by concatenating the mean supervector and the i-vector factor loading matrix with respectively the label vector and the linear classifier matrix, the traditional i-vectors are then extended to label-regularized supervised i-vectors. These supervised i-vectors are optimized to not only reconstruct the mean supervectors well but also minimize the mean squared error between the original and the reconstructed label vectors, such that they become more discriminative. Second, factor analysis (FA) can be performed on the pre-normalized centered GMM first order statistics supervector to ensure that the Gaussian statistics sub-vector of each Gaussian component is treated equally in the FA, which reduces the computational cost significantly. Experimental results are reported on the female part of the NIST SRE 2010 task with common condition 5. The proposed supervised i-vector approach outperforms the i-vector baseline by relatively 12% and 7% in terms of equal error rate (EER) and norm old minDCF values, respectively.
Ming Li 0026, Andreas Tsiartas, Maarten Van Segbroeck, Shri Narayanan
ICASSP4
2013 Continuous models of affect from text using n-grams
abstract
We propose a method of affective text analysis and modeling that is capable of generating continuous valence ratings at the sentence level starting from word and multi-word term valence ratings. Motivated from the language modeling literature, a back-off algorithm is employed to efficiently fuse the valence of single-word and multi-word terms. Specifically, a term detection criterion is used to select the appropriate n-gram terms, starting with bigrams and potentially backing off to unigrams. Term affective ratings are generated by a lexicon expansion method, using semantic similarity estimates computed on a large web corpus. The proposed framework provides state-of-the art results in the sentence level SemEval'07 task of news headline polarity detection, reaching an accuracy of 75%.
Nikos Malandrakis, Alexandros Potamianos, Shri Narayanan
ICASSP3
2013 A robust frontend for ASR: Combining denoising, noise masking and feature normalization
abstract
The sensitivity of Automatic Speech Recognition (ASR) systems to the presence of background noises in the speaking environment, still remains a challenging task. Extracting noise robust features to compensate for speech degradations due to the noise, regained popularity in recent years. This paper contributes to this trend by proposing a cost-efficient denoising method that can serve as a preprocessing stage in any feature extraction scheme to boost its ASR performance. Recognition performance on Aurora2 shows that a noise robust frontend is obtained when combined with noise masking and feature normalization. Without the requirement of high computational costs, the method achieves similar recognition results when compared to other state-of-the art noise compensation methods.
Maarten Van Segbroeck, Shri Narayanan
ICASSP2
2013 Combining window predictions efficiently - A new imputation approach for noise robust automatic speech recognition
abstract
This paper introduces a new optimization-based approach to Sparse Imputation/spectral denoising for robust Automatic Speech Recognition (ASR) applications. In particular, we propose an algorithm which couples frame-level optimization and strategic reconciliation of the predictions in a tight manner. We demonstrate that the proposed algorithm outperforms the current state-of-the-art two-step strategy of first optimizing and then averaging across windows, while maintaining the complexity advantages of efficient techniques like the Elastic Net. Our algorithm is also theoretically able to better exploit the properties of a collinear dictionary, which occurs with spectral exemplars from most speech corpora. Through experiments on the Aurora 2.0 noisy digits database, we demonstrate that this new technique achieves significant performance gains (7.67% on average over various SNR levels) over just simply averaging across large number of predictions.
Qun Feng Tan, Shri Narayanan
ICASSP2
2013 A study on the effect of prosodic emphasis transfer on overall speech translation quality
abstract
Despite the increasing interest in Speech-to-speech (S2S) translation, research and development has focused almost exclusively on the lexical aspects of translation. The importance of transferring prosodic and other paralinguistic information through S2S devices and evaluating its impact on the translation quality are yet to be well established. The novelty in this work is a large scale human evaluation study to test the hypothesis that cross-lingual prosodic emphasis transfer is directly related to the perceived quality of speech translation. This hypothesis is validated at the 0.53-0.54 correlation level on the data sets considered with results significant at p-value=0.01. The second contribution of this work is an evaluation methodology based on crowd sourcing using English-Spanish language bilingual data from two distinct domains and evaluated with over 200 bilingual speakers. We also present lessons learned on this type of S2S subjective experiments when using crowd sourcing.
Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan
ICASSP3
2013 Data driven modeling of head motion towards analysis of behaviors in couple interactions
abstract
We propose a data driven approach for modeling head motion behavior in human dyadic interactions, by establishing a structure for unconstrained natural head movement. Using recordings of couples' conversations in real psychotherapy sessions, we first track the head of each subject, compute the head motion and detect active versus non-active intervals. For detected active intervals, we use a sliding window to collect motion sequences. Linear Prediction Coefficients are used to represent the sequence, based on which we train a Gaussian Mixture Model (GMM) such that each mixture would ideally associate with one type of prototypical movement, which we will refer to as a “kineme”. For each complete interaction session, we compute the sum of posterior probabilities of all sequences over the GMM normalized by session length to predict specific “low” versus “high” expert annotated behavior code scores for Acceptance, Blame, Positive and Negative behaviors. We achieved an overall accuracy of about 70% employing these GMMs. This result shows data driven modeling of head motion provides useful information for human behavioral analysis.
Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan
ICASSP4
2013 Toward body language generation in dyadic interaction settings from interlocutor multimodal cues
abstract
During dyadic interactions, participants influence each other's verbal and nonverbal behaviors. In this paper, we examine the coordination between a dyad's body language behavior, such as body motion, posture and relative orientation, given the participants' communication goals, e.g., friendly or conflictive, in improvised interactions. We further describe a Gaussian Mixture Model (GMM) based statistical methodology for automatically generating body language of a listener from speech and gesture cues of a speaker. The experimental results show that automatically generated body language trajectories generally follow the trends of observed trajectories, especially for velocities of body and arms, and that the use of speech information improves prediction performance. These results suggest that there is a significant level of predictability of body language in the examined goal-driven improvisations, which could be exploited for interaction-driven and goal-driven body language generation.
Angeliki Metallinou, Shri Narayanan
ICASSP3
2013 Quantifying atypicality in affective facial expressions of children with autism spectrum disorders
abstract
We focus on the analysis, quantification and visualization of atypicality in affective facial expressions of children with High Functioning Autism (HFA). We examine facial Motion Capture data from typically developing (TD) children and children with HFA, using various statistical methods, including Functional Data Analysis, in order to quantify atypical expression characteristics and uncover patterns of expression evolution in the two populations. Our results show that children with HFA display higher asynchrony of motion between facial regions, more rough facial and head motion, and a larger range of facial region motion. Overall, subjects with HFA consistently display a wider variability in the expressive facial gestures that they employ. Our analysis demonstrates the utility of computational approaches for understanding behavioral data and brings new insights into the autism domain regarding the atypicality that is often associated with facial expressions of subjects with HFA.
Angeliki Metallinou, Ruth B. Grossman, Shri Narayanan
ICME3
2013 Using emotional noise to uncloud audio-visual emotion perceptual evaluation
abstract
Emotion perception underlies communication and social interaction, shaping how we interpret our world. However, there are many aspects of this process that we still do not fully understand. Notably, we have not yet identified how audio and video information are integrated during the perception of emotion. In this work we present an approach to enhance our understanding of this process using the McGurk effect paradigm, a framework in which stimuli composed of mismatched audio and video cues are presented to human evaluators. Our stimuli set contain sentence-level emotional stimuli with either the same emotion on each channel (“matched”) or different emotions on each channel (“mismatched”, for example, an angry face with a happy voice). We obtain dimensional evaluations (valence and activation) of these emotionally consistent and noisy stimuli using crowd sourcing via Amazon Mechanical Turk. We use these data to investigate the audio-visual feature bias that underlies the evaluation process. We demonstrate that both audio and video information individually contribute to the perception of these dimensional properties. We further demonstrate that the change in perception from the emotionally matched to emotionally mismatched stimuli can be modeled using only unimodal feature variation. These results provide insight into the nature of audio-visual feature integration in emotion perception.
Emily Mower Provost, Irene Zhu, Shri Narayanan
ICME3
2013 Head motion synchrony and its correlation to affectivity in dyadic interactions
abstract
Behavioral synchrony, or entrainment, is a phenomenon of great interest to psychologists and a challenging construct to quantify. In this work we study the synchrony behavior of head motion in human dyadic interactions. We model head motion using Gaussian Mixture Model (GMM) of line spectral frequencies extracted from the motion vectors of the head. We quantify interlocutor head motion similarity through the Kullback-Leibler divergence of the GMM posteriors of their respective motion sequences. We use an audiovisual database of distressed couple interactions, extensively annotated by psychologists, to test two hypotheses using the derived similarity measure. We validate the first hypothesis — that people are more likely to increase their degree of synchrony as the interaction progresses — by comparing the first and second halves of the interaction. The second hypothesis tests if the relative change of the similarity measure from these two halves is significantly correlated with the behavioral annotation by the domain experts. This work underscores the importance of head motion as an interaction cue, and the feasibility of using it in a computational model for synchrony behavior.
Bo Xiao 0003, Panayiotis G. Georgiou, Chi-Chun Lee, Brian R. Baucom, Shri Narayanan
ICME5
2013 Empirical link between hypothesis diversity and fusion performance in an ensemble of automatic speech recognition systems
Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH4
2013 Classifying language-related developmental disorders from speech cues: the promise and the potential confounds
abstract
Speech and spoken language cues offer a valuable means to measure and model human behavior. Computational models of speech behavior have the potential to support health care through assistive technologies, informed intervention, and effi-cient long-term monitoring. The Interspeech 2013 Autism Sub-Challenge addresses two developmental disorders that manifest in speech: autism spectrum disorders and specific language im-pairment. We present classification results with an analysis on the development set including a discussion of potential con-founds in the data such as recording condition differences. We hence propose study of features within these domains that may inform realistic separability between groups as well as have the potential to be used for behavioral intervention and monitoring. We investigate template-based prosodic and formant modeling as well as goodness of pronunciation modeling, reporting above chance classification accuracies. Index Terms: autism spectrum disorders, intonation, specific language impairment, goodness of pronunciation
Daniel Bone, Theodora Chaspari, Kartik Audhkhasi, James Gibson, Andreas Tsiartas, Maarten Van Segbroeck, Ming Li 0026, Sungbok Lee, Shri Narayanan
INTERSPEECH9
2013 Acoustic-prosodic, turn-taking, and language cues in child-psychologist interactions for varying social demand
abstract
Impaired social communication and social reciprocity are the primary phenotypic distinctions between autism spectrum dis-orders (ASD) and other developmental disorders. We investi-gate quantitative conversational cues in child-psychologist in-teractions using acoustic-prosodic, turn-taking, and language features. Results indicate the conversational quality degraded for children with higher ASD severity, as the child exhibited difficulties conversing and the psychologist varied her speech and language strategies to engage the child. When interacting with children with increasing ASD severity, the psychologist exhibited higher prosodic variability, increased pausing, more speech, atypical voice quality, and less use of conventional con-versational cue such as assents and non-fluencies. Children with increasing ASD severity spoke less, spoke slower, responded later, had more variable prosody, and used personal pronouns, affect language, and fillers less often. We also investigated the predictive power of features from interaction subtasks with varying social demands placed on the child. We found that acoustic prosodic and turn-taking features were more predictive during higher social demand tasks, and that the most predictive features vary with context of interaction. We also observed that psychologist language features may be robust to the amount of speech in a subtask, showing significance even when the child is participating in minimal-speech, low social-demand tasks. Index Terms: autism spectrum disorders, atypical prosody, so-cial reciprocity, turn-taking, language cues
Daniel Bone, Chi-Chun Lee, Theodora Chaspari, Matthew Black, Marian E. Williams, Sungbok Lee, Pat Levitt, Shri Narayanan
INTERSPEECH8
2013 Analyzing eye-voice coordination in rapid automatized naming
abstract
Rapid Automatized Naming (RAN) is a powerful tool for pre-dicting future reading skill. A person’s ability to quickly name symbols as they scan a table is related to higher-level reading proficiency in adults and is predictive of future literacy gains in children. However, noticeable differences are present in the strategies or patterns within groups having similar task comple-tion times. Thus, a further stratification of RAN dynamics may lead to better characterization and later intervention to support reading skill acquisition. In this work, we analyze the dynamics of the eyes, voice, and the coordination between the two during performance. It is shown that fast performers are more similar to each other than to slow performers in their patterns, but not vice versa. Further insights are provided about the patterns of more proficient subjects. For instance, fast performers tended to exhibit smoother behavior contours, suggesting a more sta-ble perception-production process.
Daniel Bone, Chi-Chun Lee, Vikram Ramanarayanan, Shri Narayanan, Renske S. Hoedemaker, Peter C. Gordon
INTERSPEECH4
2013 On the computation of document frequency statistics from spoken corpora using factor automata
Dogan Can, Shri Narayanan
INTERSPEECH2
2013 Analyzing the structure of parent-moderated narratives from children with ASD using an entity-based approach
abstract
Storytelling is a commonly used technique for rating linguistic and communicative abilities of children with Autism Spectrum Disorders (ASD). It highlights their language use beyond sentence-level production, and their ability to cohesively link events into a plot, including incorporating social context. A key scenario of interest we consider is spoken narrative creation in interactive settings, where confederates such as parents can offer scaffolding to their children’s narratives by eliciting answers with appropriate questions, shaping the structure of the resulting narrative. We analyze the structure of children’s stories narrated with the help of their parents using entity-based feature-level patterns in order to see how there are influenced by the parents’ narrative elicitation techniques. The frequency distribution and evolution of entities -meaning the co-referent people, objects and ideas- can capture the main axis of the story plot. Our results indicate that the type of questions the parents ask can be reflected in the entity-based features of a narrative, affecting its underlying structure and coherence. Index Terms: Narrative Structure, Coherence, Text Entities, Autism Spectrum Disorders
Theodora Chaspari, Emily Mower Provost, Shri Narayanan
INTERSPEECH3
2013 Information theoretic acoustic feature selection for acoustic-to-articulatory inversion
abstract
We use mutual information as the criterion to rank the Mel frequency cepstral coefficients (MFCCs) and their derivatives according to the information they provide about different articulatory features in acoustic-to-articulatory (AtoA) inversion. It is found that just a small subset of the coefficients encodes maximal information about articulatory features and interestingly, this subset is articulatory feature specific. We use these subsets of MFCCs(+derivatives) in AtoA inversion using Gaussian mixture model (GMM) mapping. Inversion experiments with articulatory data support the information theoretic finding that the subsets of MFCCs(+derivatives) as selected by feature ranking method are sufficient to achieve an inversion performance similar to that obtained by a conventional full set of MFCCs(+derivatives). This drastically reduces the modeling complexity of the acoustic-articulatory map using GMM without degrading inversion performance significantly. Index Terms: Acoustic-to-articulatory inversion, mutual information, Gaussian mixture model.
Prasanta Kumar Ghosh, Shri Narayanan
INTERSPEECH2
2013 Spectro-temporal directional derivative features for automatic speech recognition
abstract
We introduce a novel spectro-temporal representation of speech by applying directional derivative filters to the Melspectrogram, with the aim of improving the robustness of automatic speech recognition. Previous studies have shown that two-dimensional wavelet functions, when tuned to appropriate spectral scales and temporal rates, are able to accurately capture the acoustic modulations of speech, even in high noise conditions. Therefore, spectro-temporal features extracted from the wavelet transformation of the spectrogram, offer additional noise robustness to important signal processing tasks, such as voice activity detection and speech recognition. In this paper, we explore the use of the steerable pyramid, a directional wavelet transform that is common in image processing, to derive a spectro-temporal feature representation of speech that can serve as an alternative to cepstral derivatives and Gabor filterbank features. We discuss their application for the task of robust automatic speech recognition. Experiments conducted on the Aurora-2 database demonstrate their competitive robustness to other state-of-the-art speech features, especially in low signalto-noise ratio conditions. Index Terms: spectro-temporal features, automatic speech recognition, directional wavelet transforms
James Gibson, Maarten Van Segbroeck, Antonio Ortega, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH5
2013 Paralinguistic event detection from speech using probabilistic time-series smoothing and masking
Rahul Gupta 0001, Kartik Audhkhasi, Sungbok Lee, Shri Narayanan
INTERSPEECH4
2013 TRAP language identification system for RATS phase II evaluation
abstract
Automatic language identification or detection of audio data has become an important preprocessing step for speech/speaker recognition and audio data mining. In many surveillance applications, language detection has to be performed on highly degraded audio inputs. In this paper, we present our work on language detection in highly degraded radio channel scenarios. We provide a brief description of the Targeted Robust Audio Processing (TRAP) language detection system builtfor the Phase II Evaluationof the RobustAutomatic Transcription of Speech (RATS) program. This system is a combination of 15 systems with different frontends and speech activity decisions. We also analyze the usefulness of multi-layer perceptron (MLP) based non-linear projection of i-vectors before SVM classification. The proposed backend reduces the Equal Error Rate (EER) by 11%–25% relative compared to the baseline PCA-based feature representation for SVM classification, on the RATS test data consisting of data from eight highfrequency radio communication channels. Index Terms: Language identification (detection), highly degraded radio channel, RATS, i-vector, multi-layer perceptron.
Kyu Jeong Han, Sriram Ganapathy, Ming Li 0026, Mohamed Kamal Omar, Shri Narayanan
INTERSPEECH5
2013 Truncation of pharyngeal gesture in English diphthong [aɪ]
abstract
It is well acknowledged that [a] in English diphthongs (e.g. [a] in “pie’d”) has a different formant structure from its closest corresponding monophthong (e.g. [a] in “pod”). The current study proposes that these two sounds share the same cognitive unit, i.e. the pharyngeal constriction gesture that produces [a], and the surface difference can be modeled as a consequence of truncating the same articulatory movement in time by the following palatal glide in the diphthongal environment. Formation of pharyngeal constriction gesture during the production of [a] in a diphthong and in its corresponding monophthong was observed in various timing contexts using Realtime MRI; and the collected production data were quantitatively analyzed using the direct image analysis (DIA) technique, which infers tissue movement by tracking pixel intensity change over time in regions of interest. Results support our truncation account in that: (1) formation time of pharyngeal constriction is significantly longer in monophthongs than in diphthongs; (2) this duration correlates with the resulting constriction degree; and (3) the resulting constriction degree predicts the acoustic difference in the F2 dimension as predicted by our hypothesis. Index Terms: English diphthongs, speech production, rtMRI
Fang-Ying Hsieh, Louis Goldstein, Dani Byrd, Shri Narayanan
INTERSPEECH4
2013 Annotation and classification of Political advertisements
Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH3
2013 Vocal tract cross-distance estimation from real-time MRI using region-of-interest analysis
abstract
Real-Time Magnetic Resonance Imaging affords speech articu-lation data with good spatial and temporal resolution and com-plete midsagittal views of the moving vocal tract, but also brings many challenges in the domain of image processing and analy-sis. Region-of-interest analysis has previously been proposed for simple, efficient and robust extraction of linguistically-meaningful constriction degree information. However, the ac-curacy of such methods has not been rigorously evaluated, and no method has been proposed to calibrate the pixel intensity values or convert them into absolute measurements of length. This work provides such an evaluation, as well as insights into the placement of regions in the image plane and calibration of the resultant pixel intensity measurements. Measurement errors are shown to be generally at or below the spatial resolution of the imaging protocol with a high degree of consistency across time and overall vocal tract configuration, validating the utility of this method of image analysis. Index Terms: speech production data, real-time mri, analysis tools, vocal tract area functions
Adam C. Lammert, Vikram Ramanarayanan, Michael I. Proctor, Shri Narayanan
INTERSPEECH4
2013 Speaker verification based on fusion of acoustic and articulatory information
abstract
We propose a practical, feature-level fusion approach for com-bining acoustic and articulatory information in speaker ver-ification task. We find that concatenating articulation fea-tures obtained from the measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) im-proves the overall speaker verification performance. However, since access to the measured articulatory data is impractical for real world speaker verification applications, we also ex-periment with estimated articulatory features obtained using acoustic-to-articulatory inversion technique. Specifically, we show that augmenting MFCCs with articulatory features ob-tained from subject-independent acoustic-to-articulatory inver-sion technique also significantly enhances the speaker verifi-cation performance. This performance boost could be due to the information about inter-speaker variation present in the es-timated articulatory features, especially at the mean and vari-ance level. Experimental results on the Wisconsin X-Ray Mi-crobeam database show that the proposed acoustic-estimated-articulatory fusion approach significantly outperforms the tra-ditional acoustic-only baseline, providing up to 10 % relative re-duction in Equal Error Rate (EER). We further show that we can achieve an additional 5 % relative reduction in EER after score-level fusion. Index Terms: speech production, speaker verification, articula-tion features, acoustic-to-articulatory inversion, biometrics
Ming Li 0026, Jangwon Kim, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH5
2013 Velic coordination in French nasals: a real-time magnetic resonance imaging study
abstract
Production of nasal vowels in French, and nasal consonants in French and English, was examined using real-time magnetic resonance imaging (rtMRI). The coordination of velic and lin-gual gestures was found to be tightly controlled across differ-ent prosodic contexts in French nasals. Velum lowering in En-glish nasal consonants did not show the same control, although the timing of the corresponding lingual gestures varied with prosodic context in the same way as for French nasals, suggest-ing a coordinative relationship in which oral and velic articula-tors are consistently phased in French nasal production. These findings illustrate the utility of real-time MRI as a method for studying velic activity and articulatory coordination in vocalic and nasal phonology. Index Terms: speech production, velum, nasals, nasal vowels, French, articulation, real-time MRI
Michael I. Proctor, Louis Goldstein, Adam C. Lammert, Dani Byrd, Asterios Toutios, Shri Narayanan
INTERSPEECH6
2013 Articulatory settings facilitate mechanically advantageous motor control of vocal tract articulators
abstract
It was recently shown that vocal tract postures assumed during pauses in read speech are significantly different from those assumed at absolute rest. This paper examines whether the former category of “articulatory settings” are more mechanically advantageous than absolute rest postures with respect to speech articulation. Appropriate task and articulator variables are extracted from real-time Magnetic Resonance Imaging (rtMRI) data of five speakers reading aloud. Locally-weighted regression is then used to calculate Jacobian matrices representing the transformation between articulatory task velocities and postural velocities. A measure of mechanical advantage is proposed based on the obtained Jacobian. Speech-ready postures and postures during inter-speech pauses are observed to be significantly more mechanically advantageous as compared to rest postures. Furthermore, other postures, such as those that occur during the production of different vowels and consonants, are shown to have mechanical advantages that lie in between this continuum. These results could provide insights into understanding postural motor control and other linguistic phenomena, such as sonority hierarchies, in speech production. Index Terms: speech production, real-time MRI, articulatory setting, postural motor control, task dynamics, forward kinematics, vocal tract shaping.
Vikram Ramanarayanan, Adam C. Lammert, Louis Goldstein, Shri Narayanan
INTERSPEECH4
2013 A robust frontend for VAD: exploiting contextual, discriminative and spectral cues of human voice
Maarten Van Segbroeck, Andreas Tsiartas, Shri Narayanan
INTERSPEECH3
2013 Stable articulatory tasks and their variable formation: tamil retroflex consonants
abstract
A real-time MRI examination of retroflex stops and rhotics in Tamil reveals that in some contexts these consonants may in fact be achieved with little or no retroflexion of the tongue tip. Rather, maneuvering and shaping of the tongue in order to achieve post-alveolar contact varies across vowel contexts. Between back vowels /a / and /u/, post-alveolar constriction involves curling back of the tongue tip, but in the context of high front vowel /i/, the same constriction is achieved by bunching of the tongue. It appears that though there is a stable constriction target in the post-alveolar region, its achievement is not fixed but is instead a consequence of the variable state of the vocal tract in different vowel contexts. Articulatory configurations of the tongue across these vowel contexts were examined by comparing measures of Gaussian curvature at evenly spaced points along the vocal tract. The results support the notion that so-called retroflex consonants have a specified target constriction in the post-alveolar region, but that the specific articulations employed to achieve this constriction are not fixed, in keeping with the task dynamic model of speech production. Index Terms: retroflex, Tamil, real-time MRI 1.
Caitlin Smith, Michael I. Proctor, Khalil Iskarous, Louis Goldstein, Shri Narayanan
INTERSPEECH5
2013 Articulatory synthesis of French connected speech from EMA data
abstract
This paper reports an experiment in synthesizing French connected speech using Maeda’s digital simulation of the vocaltract system. The dynamics of the vocal-tract shape are estimated from the dynamics of Electromagnetic Articulograph (EMA) sensors via Maeda’s geometrical articulatory model. Time-varying characteristics of the glottis and the velopharyngeal port are set using empirical rules, while the fundamental frequency pattern is copied from the concurrently recorded audio signal. A subjective experiment was performed online to assess the perceived intelligibility and naturalness of the synthesized speech. Results indicate that a properly driven simulation of the vocal tract has the potential to provide a scientifically grounded alternative to the development of text-to-speech synthesis systems. Index Terms: articulatory synthesis, speech production, electromagnetic articulography, vocal-tract simulation
Asterios Toutios, Shri Narayanan
INTERSPEECH2
2013 Multi-band long-term signal variability features for robust voice activity detection
abstract
In this paper, we propose robust features for the problem of voice activity detection (VAD). In particular, we extend the long term signal variability (LTSV) feature to accommodate multiple spectral bands. The motivation of the multi-band approach stems from the non-uniform frequency scale of speech phonemes and noise characteristics. Our analysis shows that the multi-band approach offers advantages over the single band LTSV for voice activity detection. In terms of classification accuracy, we show 0.3%-61.2% relative improvement over the best accuracy of the baselines considered for 7 out 8 different noisy channels. Experimental results, and error analysis, are reported on the DARPA RATS corpora of noisy speech. Index Terms: noisy speech data, voice activity detection, robust feature extraction
Andreas Tsiartas, Theodora Chaspari, Athanasios Katsamanis, Prasanta Kumar Ghosh, Ming Li 0026, Maarten Van Segbroeck, Alexandros Potamianos, Shri Narayanan
INTERSPEECH8
2013 Toward transfer of acoustic cues of emphasis across languages
Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH3
2013 A two-step technique for MRI audio enhancement using dictionary learning and wavelet packet analysis
abstract
We present a method for speech enhancement of data collected in extremely noisy environments, such as those found during magnetic resonance imaging (MRI) scans. We propose a twostep algorithm to perform this noise suppression. First, we use probabilistic latent component analysis to learn dictionaries of the noise and speech+noise portions of the data and use these to factor the noisy spectrum into estimated speech and noise components. Second, we apply a wavelet packet analysis in conjunction with a wavelet threshold that minimizes the KL divergence between the estimated speech and noise to achieve further noise suppression. Based on both objective and subjective assessments, we find that our algorithm significantly outperforms traditional techniques such as nLMS, while not requiring prior knowledge or periodicity of the noise waveforms that current state-of-the-art algorithms require. Index Terms: rtMRI, noise suppression, wavelets, pLCA, dictionary learning.
Colin Vaz, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH3
2013 Modeling therapist empathy and vocal entrainment in drug addiction counseling
Bo Xiao 0003, Panayiotis G. Georgiou, Zac E. Imel, David C. Atkins, Shri Narayanan
INTERSPEECH5
2013 The effect of word frequency and lexical class on articulatory-acoustic coupling
abstract
Word frequency and lexical class distinction between function and content words have been shown to significantly influence word production. In this paper, we use real-time magnetic resonance imaging to investigate the effect of word frequency and lexical class on articulatory characteristics (the articulator speed) as well as acoustic characteristics (F0 and short-term en-ergy) in word production. Multiple regression analyses showed that word frequency exhibits significantly higher correlation with articulatory and acoustic factors for content words com-pared to function words. A Granger causality analysis uncov-ered a causal relationship from articulatory speed to F0/energy for low-frequency content words. We further observed, us-ing functional canonical correlation analysis, a tight coupling of articulatory and acoustic characteristics for low-frequency content words. These results support the view that word fre-quency distinctly influences the production of function and con-tent words as manifested in their articulation and acoustics, as well as the dynamic coupling of these temporal streams. Index Terms: word frequency, lexical class, articulatory-acoustic coupling, real-time MRI, speech production.
Vikram Ramanarayanan, Dani Byrd, Shri Narayanan
INTERSPEECH4
2013 Faster 3d vocal tract real-time MRI using constrained reconstruction
abstract
Real-time magnetic resonance imaging (rtMRI) is a valuable emerging tool for studying the dynamics of vocal production. Conventional 2D rtMRI typically images the midsagittal plane of the vocal tract, acquiring data from all the important articulators. Dynamic 3D MRI would be a major advance, as it would provide 3D visualization of the vocal tract shaping dynamics, especially for the modeling of complex vocal tract geometries, such as liquid and fricative consonants, which is not easily available from 2D rtMRI. In this paper, we present an approach to directly acquire data in 3D using a highly temporally-undersampled stack-of-spirals imaging sequence, and perform reconstruction using a partially separable model. We demonstrate visualization of vocal tract dynamics from midsagittal and coronal scan planes in the articulation of fricative consonants. Index Terms: speech production, vocal tract shaping, dynamic 3D MRI, spiral imaging, English fricative
Yinghua Zhu, Asterios Toutios, Shri Narayanan, Krishna S. Nayak
INTERSPEECH3
2013 Which ASR should I choose for my dialogue system?
Fabrizio Morbini, Kartik Audhkhasi, Kenji Sagae, Ron Artstein, Dogan Can, Panayiotis G. Georgiou, Shri Narayanan, Anton Leuski, David R. Traum
SIGDIAL Conference7
2013 Unsupervised data processing for classifier-based speech translator
Emil Ettelaie, Panayiotis G. Georgiou, Shri Narayanan
Comput. Speech Lang.3
2013 Automatic speaker age and gender recognition using acoustic and prosodic level information fusion
Ming Li 0026, Kyu Jeong Han, Shri Narayanan
Comput. Speech Lang.3
2013 Paralinguistics in speech and language - State-of-the-art and the challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller 0014, Shri Narayanan
Comput. Speech Lang.7
2013 Enabling effective design of multimodal interfaces for speech-to-speech translation system: An empirical study of longitudinal user behaviors over time and user strategies for coping with errors
JongHo Shin, Panayiotis G. Georgiou, Shri Narayanan
Comput. Speech Lang.3
2013 Enriching machine-mediated speech-to-speech translation using contextual information
Vivek Kumar Rangarajan Sridhar, Srinivas Bangalore, Shri Narayanan
Comput. Speech Lang.3
2013 High-quality bilingual subtitle document alignments with application to spontaneous speech translation
Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan
Comput. Speech Lang.4
2013 Tracking continuous emotional trends of participants during affective dyadic interactions using body language and speech information
Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan
Image Vis. Comput.3
2013 A Globally-Variant Locally-Constant Model for Fusion of Labels from Multiple Diverse Experts without Using Reference Labels
abstract
Researchers have shown that fusion of categorical labels from multiple experts—humans or machine classifiers—improves the accuracy and generalizability of the overall classification system. Simple plurality is a popular technique for performing this fusion, but it gives equal importance to labels from all experts, who may not be equally reliable or consistent across the dataset. Estimation of expert reliability without knowing the reference labels is, however, a challenging problem. Most previous works deal with these challenges by modeling expert reliability as constant over the entire data (feature) space. This paper presents a model based on the consideration that in dealing with real-world data, expert reliability is variable over the complete feature space but constant over local clusters of homogeneous instances. This model jointly learns a classifier and expert reliability parameters without assuming knowledge of the reference labels using the Expectation-Maximization (EM) algorithm. Classification experiments on simulated data, data from the UCI Machine Learning Repository, and two emotional speech classification datasets show the benefits of the proposed model. Using a metric based on the Jensen-Shannon divergence, we empirically show that the proposed model gives greater benefit for datasets where expert reliability is highly variable over the feature space.
Kartik Audhkhasi, Shri Narayanan
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Behavioral Signal Processing: Deriving Human Behavioral Informatics From Speech and Language
abstract
The expression and experience of human behavior are complex and multimodal and characterized by individual and contextual heterogeneity and variability. Speech and spoken language communication cues offer an important means for measuring and modeling human behavior. Observational research and practice across a variety of domains from commerce to healthcare rely on speech- and language-based informatics for crucial assessment and diagnostic information and for planning and tracking response to an intervention. In this paper, we describe some of the opportunities as well as emerging methodologies and applications of human behavioral signal processing (BSP) technology and algorithms for quantitatively understanding and modeling typical, atypical, and distressed human behavior with a specific focus on speech- and language-based communicative, affective, and social behavior. We describe the three important BSP components of acquiring behavioral data in an ecologically valid manner across laboratory to real-world settings, extracting and analyzing behavioral cues from measured data, and developing models offering predictive and decision-making support. We highlight both the foundational speech and language processing building blocks as well as the novel processing and modeling opportunities. Using examples drawn from specific real-world applications ranging from literacy assessment and autism diagnostics to psychotherapy for addiction and marital well being, we illustrate behavioral informatics applications of these signal processing techniques that contribute to quantifying higher level, often subjectively described, human behavior in a domain-sensitive fashion.
Shri Narayanan, Panayiotis G. Georgiou
Proc. IEEE1
2013 An Overview on Perceptually Motivated Audio Indexing and Classification
abstract
An audio indexing system aims at describing audio content by identifying, labeling, or categorizing different acoustic events. Since the resulting audio classification and indexing is meant for direct human consumption, it is highly desirable that it produces perceptually relevant results. This can be obtained by integrating specific knowledge of the human auditory system in the design process to various extent. In this paper, we highlight some of the important concepts used in audio classification and indexing that are perceptually motivated or that exploit some principles of perception. In particular, we discuss several different strategies to integrate human perception, including: 1) the use of generic audition models; 2) the use of perceptually relevant features for the analysis stage that are perceptually justified either as a component of a hearing model or as being correlated with a perceptual dimension of sound similarity; and 3) the involvement of the user in the audio indexing or classification task. In this paper, we also illustrate some of the recent trends in semantic audio retrieval that approximate higher level perceptual processing and cognitive aspects of human audio recognition capabilities, including affect-based audio retrieval.
Gaël Richard, Shiva Sundaram, Shri Narayanan
Proc. IEEE3
2013 Toward automating a human behavioral coding system for married couples' interactions using speech acoustic features
Matthew Black, Athanasios Katsamanis, Brian R. Baucom, Chi-Chun Lee, Adam C. Lammert, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan
Speech Commun.8
2013 Statistical methods for estimation of direct and differential kinematics of the vocal tract
Adam C. Lammert, Louis Goldstein, Shri Narayanan, Khalil Iskarous
Speech Commun.3
2013 Iterative Feature Normalization Scheme for Automatic Emotion Detection from Speech
abstract
The externalization of emotion is intrinsically speaker-dependent. A robust emotion recognition system should be able to compensate for these differences across speakers. A natural approach is to normalize the features before training the classifiers. However, the normalization scheme should not affect the acoustic differences between emotional classes. This study presents the iterative feature normalization (IFN) framework, which is an unsupervised front-end, especially designed for emotion detection. The IFN approach aims to reduce the acoustic differences, between the neutral speech across speakers, while preserving the inter-emotional variability in expressive speech. This goal is achieved by iteratively detecting neutral speech for each speaker, and using this subset to estimate the feature normalization parameters. Then, an affine transformation is applied to both neutral and emotional speech. This process is repeated till the results from the emotion detection system are consistent between consecutive iterations. The IFN approach is exhaustively evaluated using the IEMOCAP database and a data set obtained under free uncontrolled recording conditions with different evaluation configurations. The results show that the systems trained with the IFN approach achieve better performance than systems trained either without normalization or with global normalization.
Carlos Busso, Soroosh Mariooryad, Angeliki Metallinou, Shri Narayanan
IEEE Trans. Affect. Comput.4
2013 Distributional Semantic Models for Affective Text Analysis
abstract
We present an affective text analysis model that can directly estimate and combine affective ratings of multi-word terms, with application to the problem of sentence polarity/semantic orientation detection. Starting from a hierarchical compositional method for generating sentence ratings, we expand the model by adding multi-word terms that can capture non-compositional semantics. The method operates similarly to a bigram language model, using bigram terms or backing off to unigrams based on a (degree of) compositionality criterion. The affective ratings for n-gram terms of different orders are estimated via a corpus-based method using distributional semantic similarity metrics between unseen words and a set of seed words. N-gram ratings are then combined into sentence ratings via simple algebraic formulas. The proposed framework produces state-of-the-art results for word-level tasks in English and German and the sentence-level news headlines classification SemEval'07-Task14 task. The inclusion of bigram terms to the model provides significant performance improvement, even if no term selection is applied.
Nikos Malandrakis, Alexandros Potamianos, Elias Iosif, Shri Narayanan
IEEE Trans. Speech Audio Process.4
2013 Toward the Automatic Extraction of Policy Networks Using Web Links and Documents
abstract
Policy networks are widely used by political scientists and economists to explain various financial and social phenomena, such as the development of partnerships between political entities or institutions from different levels of governance. The analysis of policy networks demands a series of arduous and time-consuming manual steps including interviews and questionnaires. In this paper, we estimate the strength of relations between actors in policy networks using features extracted from data harvested from the web. Features include webpage counts, outlinks, and lexical information extracted from web documents or web snippets. The proposed approach is automatic and does not require any external knowledge source, other than the specification of the word forms that correspond to the political actors. The features are evaluated both in isolation and jointly for both positive and negative (antagonistic) actor relations. The proposed algorithms are evaluated on two EU policy networks from the political science literature. Performance is measured in terms of correlation and mean square error between the human rated and the automatically extracted relations. Correlation of up to 0.74 is achieved for positive relations. The extracted networks are validated by political scientists and useful conclusions about the evolution of the networks over time are drawn.
Theodosis Moschopoulos, Elias Iosif, Leeda Demetropoulou, Alexandros Potamianos, Shri Narayanan
IEEE Trans. Knowl. Data Eng.5
2013 Dynamic 3-D Visualization of Vocal Tract Shaping During Speech
abstract
Noninvasive imaging is widely used in speech research as a means to investigate the shaping and dynamics of the vocal tract during speech production. 3-D dynamic MRI would be a major advance, as it would provide 3-D dynamic visualization of the entire vocal tract. We present a novel method for the creation of 3-D dynamic movies of vocal tract shaping based on the acquisition of 2-D dynamic data from parallel slices and temporal alignment of the image sequences using audio information. Multiple sagittal 2-D real-time movies with synchronized audio recordings are acquired for English vowel-consonant-vowel stimuli /ala/, /a.ιa/, /asa/, and /a∫a/. Audio data are aligned using mel-frequency cepstral coefficients (MFCC) extracted from windowed intervals of the speech signal. Sagittal image sequences acquired from all slices are then aligned using dynamic time warping (DTW). The aligned image sequences enable dynamic 3-D visualization by creating synthesized movies of the moving airway in the coronal planes, visualizing desired tissue surfaces and tube-shaped vocal tract airway after manual segmentation of targeted articulators and smoothing. The resulting volumes allow for dynamic 3-D visualization of salient aspects of lingual articulation, including the formation of tongue grooves and sublingual cavities, with a temporal resolution of 78 ms.
Yinghua Zhu, Yoon-Chul Kim, Michael I. Proctor, Shri Narayanan, Krishna S. Nayak
IEEE Trans. Medical Imaging4
2012 Analyzing quality of crowd-sourced speech transcriptions of noisy audio for acoustic model adaptation
abstract
The accuracy of crowd-sourced speech transcriptions varies depending on a variety of factors. This paper studies the impact of one such factor, namely, the quality of audio. We employed a speech database with babble noise at three SNR levels (clean, 2 dB and -2 dB) and asked workers on Amazon Mechanical Turk to transcribe it. Two interesting observations emerge. First, as expected, the quality of transcripts combined by word frequency based ROVER decreases with decreasing SNR. Further, we demonstrate that the use of some unsupervised reliability scores can improve the transcription quality, with increasing benefits at lower SNR. Second, we do not observe a significant drop in the performance of acoustic models adapted with increasing transcription noise. This highlights the surprising robustness of crowd-sourced transcripts for acoustic model adaptation.
Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan
ICASSP3
2012 Creating ensemble of diverse maximum entropy models
abstract
Diversity of a classifier ensemble has been shown to benefit overall classification performance. But most conventional methods of training ensembles offer no control on the extent of diversity and are meta-learners. We present a method for creating an ensemble of diverse maximum entropy (∂MaxEnt) models, which are popular in speech and language processing. We modify the objective function for conventional training of a MaxEnt model such that its output posterior distribution is diverse with respect to a reference model. Two diversity scores are explored - KL divergence and posterior cross-correlation. Experiments on the CoNLL-2003 Named Entity Recognition task and the IEMOCAP emotion recognition database show the benefits of a ∂MaxEnt ensemble.
Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shri Narayanan
ICASSP4
2012 Improvements in predicting children's overall reading ability by modeling variability in evaluators' subjective judgments
abstract
Automatic literacy assessment is one promising application of speech and language processing research. In our previous work, we showed we could accurately predict children's overall ability to read a list of English words aloud, an integral component of early literacy assessment. In this paper, we improve upon our results by exploiting the fact that evaluators' level of agreement significantly varies, depending on the child being judged. This source of evaluator variability is directly modeled using generalized least squares linear regression. In this framework, the children for which the evaluators were more confident in rating are weighted higher. Performance in predicting the mean evaluator's scores increases from a Pearson's correlation coefficient of 0.946 to 0.952, a relative improvement of 0.63%. This is a significantly higher correlation than the mean inter-evaluator agreement of 0.899 (p <; 0.05). Critically, the mean and maximum absolute errors are significantly reduced.
Matthew Black, Shri Narayanan
ICASSP2
2012 An acoustic analysis of shared enjoyment in ECA interactions of children with autism
abstract
The quality of shared enjoyment in interactions is a key aspect related to Autism Spectrum Disorders (ASD). This paper discusses two types of enjoyment: the first refers to humorous events and is associated with one's positive affective state and the second is used to facilitate social interactions between people. These types of shared enjoyment are objectively specified by their proximity to a voiced and unvoiced laughter instance, respectively. The goal of this work is to study the acoustic differences of areas surrounding the two kinds of shared enjoyment instances, called “social zones”, using data collected from children with autism, and their parents, interacting with an Embodied Conversational Agent (ECA). A classification task was performed to predict whether a “social zone” surrounds a voiced or an unvoiced laughter instance. Our results indicate that humorous events are more easily recognized than events acting as social facilitators and that related speech patterns vary more across children compared to other interlocutors.
Theodora Chaspari, Emily Mower Provost, Athanasios Katsamanis, Shri Narayanan
ICASSP4
2012 Classification of emotional content of sighs in dyadic human interactions
abstract
Emotions are an important part of human communication and are expressed both verbally and non-verbally. Common nonverbal vocalizations such as laughter, cries and sighs carry important emotional content in conversations. Sighs often are associated with negative emotion. In this work, we show that emotional sighs exist along both ends of the valence axis (positive-emotion vs. negative-emotion sighs) in spontaneous affective dialogs and that they have certain distinct multimodal characteristics. Classification results show that it is possible to differentiate between the two types of emotionally valenced sighs, using a combination of acoustic and gestural features with an overall unweighted accuracy of 58.26%.
Rahul Gupta 0001, Chi-Chun Lee, Shri Narayanan
ICASSP3
2012 Object classification in sidescan sonar images with sparse representation techniques
abstract
Most supervised classification approaches try to learn patterns in inter class variabilities using training samples. However in the real world, their discriminative power is often diminished, because data is seldom free from irregularities within a class. Apriori modeling of these intra class variabilities poses a challenge even in underwater sidescan sonar images that we consider for object classification in this work. Sparse representation techniques prove particularly useful in this regard because of their data driven approach to model these variabilities. Results on the NSWC sidescan sonar database suggest that sparse representation classifier with zernike magnitude features is significantly robust in the presence of these non-idealities.
Naveen Kumar 0004, Qun Feng Tan, Shri Narayanan
ICASSP3
2012 Speaker states recognition using latent factor analysis based Eigenchannel factor vector modeling
abstract
This paper presents an automatic speaker state recognition approach which models the factor vectors in the latent factor analysis framework improving upon the Gaussian Mixture Model (GMM) baseline performance. We investigate both intoxicated and affective speaker states. We consider the affective speech signal as the original normal average speech signal being corrupted by the affective channel effects. Rather than reducing the channel variability to enhance the robustness as in the speaker verification task, we directly model the speaker state on the channel factors under the factor analysis framework. In this work, the speaker state factor vectors are extracted and modeled by the latent factor analysis approach in the GMM modeling framework and support vector machine classification method. Experimental results show that the proposed speaker state factor vector modeling system achieved 5.34% and 1.49% unweighted accuracy improvement over the GMM baseline on the intoxicated speech detection task (Alcohol Language Corpus) and the emotion recognition task (IEMOCAP database), respectively.
Ming Li 0026, Angeliki Metallinou, Daniel Bone, Shri Narayanan
ICASSP4
2012 A hierarchical framework for modeling multimodality and emotional evolution in affective dialogs
abstract
Incorporating multimodal information and temporal context from speakers during an emotional dialog can contribute to improving performance of automatic emotion recognition systems. Motivated by these issues, we propose a hierarchical framework which models emotional evolution within and between emotional utterances, i.e., at the utterance and dialog level respectively. Our approach can incorporate a variety of generative or discriminative classifiers at each level and provides flexibility and extensibility in terms of multimodal fusion; facial, vocal, head and hand movement cues can be included and fused according to the modality and the emotion classification task. Our results using the multimodal, multi-speaker IEMOCAP database indicate that this framework is well-suited for cases where emotions are expressed multimodally and in context, as in many real-life situations.
Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan
ICASSP3
2012 Automatic recognition of emotion evoked by general sound events
abstract
Without a doubt there is emotion in sound. So far, however, research efforts have focused on emotion in speech and music despite many applications in emotion-sensitive sound retrieval. This paper is an attempt at automatic emotion recognition of general sounds. We selected sound clips from different areas of the daily human environment and model them using the increasingly popular dimensional approach in the emotional arousal and valence space. To establish a reliable ground truth, we compare mean and median of four annotators with their evaluator weighted estimator. We discuss human labelers' consistency, feature relevance, and automatic regression. Results reach correlation coefficients of .61 (arousal) and .49 (valence).
Björn W. Schuller, Simone Hantke, Felix Weninger, Wenjing Han, Zixing Zhang 0001, Shri Narayanan
ICASSP6
2012 Analyzing the memory of BLSTM Neural Networks for enhanced emotion classification in dyadic spoken interactions
abstract
Recent studies indicate that bidirectional Long Short-Term Memory (BLSTM) recurrent neural networks are well-suited for automatic emotion recognition systems and may lead to better results than systems applying other widely used classifiers such as Support Vector Machines or feedforward Neural Networks. The good performance of BLSTM emotion recognition systems could be attributed to their ability to model and exploit contextual information self-learned via recurrently connected memory blocks which allows them to incorporate information about how emotion evolves over time. However, the actual amount of bidirectional context that a BLSTM classifier takes into account when classifying an observation has not been investigated so far. This paper presents a methodology to systematically investigate the number of past and future utterance-level observations that are considered to generate an emotion prediction for a given utterance, and to examine to what extent this temporal bidirectional context contributes to the overall BLSTM performance.
Martin Wöllmer, Angeliki Metallinou, Athanasios Katsamanis, Björn W. Schuller, Shri Narayanan
ICASSP5
2012 Multimodal detection of salient behaviors of approach-avoidance in dyadic interactions
abstract
Approach-Avoidance (AA) coding is a measure of involvement and immediacy in human dyadic interactions. We focus on analyzing the salient events in interactions that trigger change points in AA code in time, as perceived by domain experts. We employ coarse level visual cues associated with body parts, as well as vocal energy features. Motion vector extraction and body pose estimation techniques are used for extracting visual cues. Functionals of these cues are used as features for SVM based machine learning experiments. We found that the coder's judgments on salient events are related to the short time interval preceding the labeling. We also show that visual cues are the main information source for decision making on salient AA events, and that considering the information from a subset of body parts provides the same information as considering the full set. The mean of absolute value and standard deviation of motion streams are the most effective functionals as feature. We achieve an F-score of 0.55 in detecting salient events using cross-validation with a one-subject-out approach.
Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan
ICMI4
2012 Speaker Personality Classification Using Systems Based on Acoustic-Lexical Cues and an Optimal Tree-Structured Bayesian Network
abstract
Automatic classification of human personality along the Big Five dimensions is an interesting problem with several prac-tical applications. This paper makes some contributions in this regard. First, we propose a few automatically-derived personality-discriminating lexical features which provide infor-mation complementary to the conventional acoustic-prosodic cues. We also design a frame-level Gaussian mixture model based system which adds complimentary information to the sys-tems trained on global statistical functionals. Next, we note that the Big Five dimensions are correlated and thus model the de-pendency between these dimensions in the form of an optimal tree-structured Bayesian network. Our final sub-system con-sists of within class covariance normalization followed by L1-regularized logistic regression. Fusion of all these sub-systems achieves better classification performance than independently trained classifiers using just acoustic features.
Kartik Audhkhasi, Angeliki Metallinou, Ming Li 0026, Shri Narayanan
INTERSPEECH4
2012 Spontaneous-Speech Acoustic-Prosodic Features of Children with Autism and the Interacting Psychologist
abstract
Atypical prosody, often reported in children with Autism Spectrum Disorders, is described by a range of qualitative terms that reflect the eccentricities and variability among persons in the spectrum. We investigate various word-and phonetic-level features from spontaneous speech that may quantify the cues reflecting prosody. Furthermore, we introduce the importance of jointly modeling the psy-chologist’s vocal behavior in this dyadic interaction. We demonstrate that acoustic-prosodic features of both par-ticipants correlate with the children’s rated autism sever-ity. For increasing perceived atypicality, we find chil-dren’s prosodic features that suggest ‘monotonic ’ speech, variable volume, atypical voice quality, and slower rate of speech. Additionally, we find the psychologist’s features inform their perception of a child’s atypical behavior– e.g., the psychologist’s pitch slope and jitter are increas-ingly variable and their speech rate generally decreases. Index Terms: atypical prosody, autism spectrum disor-der, intonation, psychologist, voice quality, ADOS
Daniel Bone, Matthew Black, Chi-Chun Lee, Marian E. Williams, Pat Levitt, Sungbok Lee, Shri Narayanan
INTERSPEECH7
2012 A Robust Unsupervised Arousal Rating Framework using Prosody with Cross-Corpora Evaluation
abstract
This paper presents an unsupervised method for produc-ing a bounded rating of affective arousal from speech. One of the major challenges in such behavioral signal classification is the design of methods that generalize well across domains and datasets. We propose a frame-work that provides robustness across databases by: se-lecting coherent features based on empirical and theoret-ical evidence, fusing activation confidences from mul-tiple features, and effectively weighting the soft-labels without knowing the true labels. Spearman’s rank-correlation (and binary classification accuracy) on four
Daniel Bone, Chi-Chun Lee, Shri Narayanan
INTERSPEECH3
2012 A Case Study: Detecting Counselor Reflections in Psychotherapy for Addictions using Linguistic Features
abstract
Motivational Interviewing (MI) is a goal-oriented psychotherapy, employed in cases such as addiction, which helps clients (i.e., patients) explore and resolve their ambivalence about the problem at hand in a dialog setting. Measuring the counselor’s proficiency with MI has typically been assessed via behavioral coding a time consuming, non-technological approach. This paper examines a computational approach to assessing the quality of MI. Specifically, we focus on a particular aspect of the counselor behavior – reflections – believed to be a critical indicator of MI therapy quality. We automatically tag reflection instances in a maximum entropy Markov modeling framework using several linguistic features with rich contextual information obtained from the session transcripts. We achieve an Fscore of over 80% while gaining insight about the information sources as perceived by the trained annotators.
Dogan Can, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan
INTERSPEECH4
2012 Interplay between verbal response latency and physiology of children with autism during ECA interactions
Theodora Chaspari, Chi-Chun Lee, Shri Narayanan
INTERSPEECH3
2012 Characterizing Covert Articulation in Apraxic Speech Using real-time MRI
abstract
We explore the use of real-time magnetic resonance imaging (rtMRI) as a tool to investigate apraxic speech, in particular, by examining articulatory behavior. Our pilot data reveal that covert (silent) gestural intrusion errors (employing an intrinsically simple 1:1 mode of coupling) are made more frequently by an apraxic subject than by fluent speakers. Covert intrusion errors are also found to be pervasive in non-repetitious apraxic speech. We demonstrate that acoustically silent periods observed before the initiation of apraxic speech oftentimes contain completely covert gestures that occur frequently with multigestural segments. Covert gestures corresponding to entire words are also observed. These data demonstrate that rtMRI can provide important new insights into apraxic speech that are not available using traditional methods of transcription based on acoustic data alone. Index Terms: Apraxia, speech production, covert articulation,
Christina Hagedorn, Michael I. Proctor, Louis Goldstein, Maria Luisa Gorno-Tempini, Shri Narayanan
INTERSPEECH5