EDBT 2026 Demo / reviewers in the wild / expert
Panayiotis G. Georgiou
dblp:80/7046
· DBLP profile ↗
142ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0002-0790-7161ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 109 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 84 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Speech recognition and synthesis · 71% Language models and text generation · 10% Machine translation · 9% | |
| Human-computer interaction and pervasive computing
1 paper |
Accessibility and assistive technology · 100% | |
| Computer graphics and multimedia
5 papers |
Audio and music processing · 74% Multimedia analysis and retrieval · 26% |
Topics — the 20 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.9 | 3 | 2023 | From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition · CHI 2023 Theoretical Analysis of Diversity in an Ensemble of Automatic Speech Recognition Systems · IEEE ACM Trans. Audio Speech Lang. Process. 2014 An Iterative Relative Entropy Minimization-Based Data Selection Approach for n-Gram Model Adaptation · IEEE Trans. Speech Audio Process. 2009 |
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker embedding |
0.4 | 1 | 2019 | Neural Predictive Coding Using Convolutional Neural Networks Toward Unsupervised Learning of Speaker Characteristics · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Natural language and speech › Speech recognition and synthesis
speaker recognition |
0.4 | 1 | 2019 | Neural Predictive Coding Using Convolutional Neural Networks Toward Unsupervised Learning of Speaker Characteristics · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Multimedia analysis and retrieval › video content analysis
human behavior analysis |
0.2 | 1 | 2015 | Head Motion Modeling for Human Behavior Analysis in Dyadic Interaction · IEEE Trans. Multim. 2015 |
Accessibility and assistive technology
augmentative and alternative communication |
0.2 | 1 | 2023 | From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition · CHI 2023 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training |
0.2 | 1 | 2014 | Theoretical Analysis of Diversity in an Ensemble of Automatic Speech Recognition Systems · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Natural language and speech › Machine translation
system combination |
0.2 | 1 | 2014 | Theoretical Analysis of Diversity in an Ensemble of Automatic Speech Recognition Systems · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing › speech recognition
robust speech recognition |
0.1 | 1 | 2011 | Enhanced Sparse Imputation Techniques for a Robust Speech Recognition Front-End · IEEE ACM Trans. Audio Speech Lang. Process. 2011 |
Audio and music processing
speech recognition |
0.1 | 1 | 2011 | Enhanced Sparse Imputation Techniques for a Robust Speech Recognition Front-End · IEEE ACM Trans. Audio Speech Lang. Process. 2011 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.1 | 1 | 2019 | Neural Predictive Coding Using Convolutional Neural Networks Toward Unsupervised Learning of Speaker Characteristics · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.1 | 1 | 2009 | An Iterative Relative Entropy Minimization-Based Data Selection Approach for n-Gram Model Adaptation · IEEE Trans. Speech Audio Process. 2009 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.1 | 1 | 2009 | An Iterative Relative Entropy Minimization-Based Data Selection Approach for n-Gram Model Adaptation · IEEE Trans. Speech Audio Process. 2009 |
Audio and music processing
music information retrieval |
0.1 | 1 | 2008 | Challenging Uncertainty in Query by Humming Systems: A Fingerprinting Approach · IEEE Trans. Speech Audio Process. 2008 |
Audio and music processing › music information retrieval › melody retrieval
query-by-humming |
0.1 | 1 | 2008 | Challenging Uncertainty in Query by Humming Systems: A Fingerprinting Approach · IEEE Trans. Speech Audio Process. 2008 |
Natural language and speech › Language models and text generation › pre-trained language model
domain-specific language model |
0.1 | 1 | 2006 | Text data acquisition for domain-specific language models · EMNLP 2006 |
Audio and music processing
sound source localization |
0.1 | 1 | 2006 | Robust maximum likelihood source localization: the case for sub-Gaussian versus Gaussian · IEEE Trans. Speech Audio Process. 2006 |
Natural language and speech › Machine translation › speech translation
speech-to-speech translation |
0.1 | 1 | 2005 | Transonics: A Practical Speech-to-Speech Translator for English-Farsi Medical Dialogs · ACL 2005 |
Audio and music processing › sound source localization
time delay estimation |
0.0 | 1 | 1999 | Alpha-Stable Modeling of Noise and Robust Time-Delay Estimation in the Presence of Impulsive Noise · IEEE Trans. Multim. 1999 |
Medical and health informatics › clinical text processing
clinical dialogue |
0.0 | 1 | 2005 | Transonics: A Practical Speech-to-Speech Translator for English-Farsi Medical Dialogs · ACL 2005 |
Wireless sensing and localization
acoustic source localization |
0.0 | 1 | 1999 | Alpha-Stable Modeling of Noise and Robust Time-Delay Estimation in the Presence of Impulsive Noise · IEEE Trans. Multim. 1999 |
Methods — techniques the papers use, named apart from their topics
voice assistant evaluation · 1.3survey · 1.3error mitigation · 1.3siamese network · 0.4predictive coding · 0.4convolutional neural network · 0.4signal processing · 0.3predictive modeling · 0.3linear predictive features · 0.2kinesics · 0.2gaussian mixture model · 0.2reversible-jump metropolis-hastings · 0.2ROVER · 0.2least angle regression · 0.1lasso · 0.1elastic net · 0.1hidden markov model · 0.1edit distance · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Multimodal Approach to Device-Directed Speech Detection with Large Language ModelsabstractInteractions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore this task in three ways: First, we train classifiers using only acoustic information obtained from the audio waveform. Second, we take the decoder outputs of an automatic speech recognition (ASR) system, such as 1-best hypotheses, as input features to a large language model (LLM). Finally, we explore a multimodal system that combines acoustic and lexical features, as well as ASR decoder signals in an LLM. Using multimodal information yields relative equal-error-rate improvements over text-only and audio-only models of up to 39% and 61%. Increasing the size of the LLM and training with low-rank adaption leads to further relative EER reductions of up to 18% on our dataset. Dominik Wagner 0002, Alexander W. Churchill, Siddharth Sigtia, Panayiotis G. Georgiou, Matt Mirsamadi, Aarshee Mishra, Erik Marchi |
ICASSP | 4 |
| 2023 | From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionabstractConsumer speech recognition systems do not work as well for many people with speech differences, such as stuttering, relative to the rest of the general population. However, what is not clear is the degree to which these systems do not work, how they can be improved, or how much people want to use them. In this paper, we first address these questions using results from a 61-person survey from people who stutter and find participants want to use speech recognition but are frequently cut off, misunderstood, or speech predictions do not represent intent. In a second study, where 91 people who stutter recorded voice assistant commands and dictation, we quantify how dysfluencies impede performance in a consumer-grade speech recognition system. Through three technical investigations, we demonstrate how many common errors can be prevented, resulting in a system that cuts utterances off 79.1% less often and improves word error rate from 25.4% to 9.9%. Colin Lea, Zifang Huang, Jaya Narain, Lauren Tooley, Dianna Yee, Tien Dung Tran, Panayiotis G. Georgiou, Jeffrey P. Bigham, Leah Findlater |
CHI | 7 |
| 2022 | Multi-Label Multi-Task Deep Learning for Behavioral CodingabstractWe propose a methodology for estimating human behaviors in psychotherapy sessions using multi-label and multi-task learning paradigms. We discuss the problem of behavioral coding in which data of human interactions are annotated with labels to describe relevant human behaviors of interest. We describe two related, yet distinct, corpora consisting of therapist-client interactions in psychotherapy sessions. We experimentally compare the proposed learning approaches for estimating behaviors of interest in these datasets. Specifically, we compare single and multiple label learning approaches, single and multiple task learning approaches, and evaluate the performance of these approaches when incorporating turn context. We demonstrate that the best multi-label, multi-task learning model with turn context achieves 18.9 and 19.5 percent absolute improvements with respect to a logistic regression classifier (for each behavioral coding task respectively) and 6.4 and 6.1 percent absolute improvements with respect to the best single-label, single-task deep neural network models. Lastly, we discuss the insights these modeling paradigms provide into these complex interactions including key commonalities and differences of behaviors within and between the two prevalent psychotherapy approaches–Motivational Interviewing and Cognitive Behavioral Therapy–considered. James Gibson, David C. Atkins, Torrey A. Creed, Zac E. Imel, Panayiotis G. Georgiou, Shri Narayanan |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Modeling Vocal Entrainment in Conversational Speech Using Deep Unsupervised LearningabstractIn interpersonal spoken interactions, individuals tend to adapt to their conversation partner's vocal characteristics to become similar, a phenomenon known as entrainment. A majority of the previous computational approaches are often knowledge driven and linear and fail to capture the inherent nonlinearity of entrainment. In this article, we present an unsupervised deep learning framework to derive a representation from speech features containing information relevant for vocal entrainment. We investigate both an encoding based approach and a more robust triplet network based approach within the proposed framework. We also propose a number of distance measures in the representation space and use them for quantification of entrainment. We first validate the proposed distances by using them to distinguish real conversations from fake ones. Then we also demonstrate their applications in relation to modeling several entrainment-relevant behaviors in observational psychotherapy, namely agreement, blame and emotional bond. Md. Nasir, Brian R. Baucom, Craig J. Bryan, Shri Narayanan, Panayiotis G. Georgiou |
IEEE Trans. Affect. Comput. | 5 |
| 2021 | Analysis and Tuning of a Voice Assistant System for Dysfluent SpeechabstractDysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice operated systems do not work. Current speech recognition systems are trained primarily with data from fluent speakers and as a consequence do not generalize well to speech with dysfluencies such as sound or word repetitions, sound prolongations, or audible blocks. The focus of this work is on quantitative analysis of a consumer speech recognition system on individuals who stutter and production-oriented approaches for improving performance for common voice assistant tasks (i.e., "what is the weather?"). At baseline, this system introduces a significant number of insertion and substitution errors resulting in intended speech Word Error Rates (isWER) that are 13.64\% worse (absolute) for individuals with fluency disorders. We show that by simply tuning the decoding parameters in an existing hybrid speech recognition system one can improve isWER by 24\% (relative) for individuals with fluency disorders. Tuning these parameters translates to 3.6\% better domain recognition and 1.7\% better intent recognition relative to the default setup for the 18 study participants across all stuttering severities. Vikramjit Mitra, Zifang Huang, Colin Lea, Lauren Tooley, Sarah Wu, Darren Botten, Ashwini Palekar, Shrinath Thelapurath, Panayiotis G. Georgiou, Sachin Kajarekar, Jeffrey P. Bigham |
Interspeech | 9 |
| 2021 | RNN Based Incremental Online Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) typically comprises of an automatic speech recognition (ASR) followed by a natural language understanding (NLU) module. The two modules process signals in a blocking sequential fashion, i.e., the NLU often has to wait for the ASR to finish processing on an utterance basis, potentially leading to high latencies that render the spoken interaction less natural. In this paper, we propose recurrent neural network (RNN) based incremental processing towards the SLU task of intent detection. The proposed methodology offers lower latencies than a typical SLU system, without any significant reduction in system accuracy. We introduce and analyze different recurrent neural network architectures for incremental and online processing of the ASR transcripts and compare it to the existing offline systems. A lexical End-of-Sentence (EOS) detector is proposed for segmenting the stream of transcript into sentences for intent classification. Intent detection experiments are conducted on benchmark ATIS, Snips and Facebook's multilingual task oriented dialog datasets modified to emulate a continuous incremental stream of words with no utterance demarcation. We also analyze the prospects of early intent detection, before EOS, with our proposed system. Prashanth Gurunath Shivakumar, Naveen Kumar 0004, Panayiotis G. Georgiou, Shri Narayanan |
SLT | 3 |
| 2021 | An analysis of observation length requirements for machine understanding of human behaviors from spoken language
Sandeep Nallan Chakravarthula, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou |
Comput. Speech Lang. | 4 |
| 2021 | Unsupervised speech representation learning for behavior modeling using triplet enhanced contextualized networks
Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou |
Comput. Speech Lang. | 4 |
| 2021 | Multimodal Embeddings From Language Models for Emotion Recognition in the WildabstractWord embeddings such as ELMo and BERT have been shown to model word usage in language with greater efficacy through contextualized learning on large-scale language corpora, resulting in significant performance improvement across many natural language processing tasks. In this work we integrate acoustic information into contextualized lexical embeddings through the addition of a parallel stream to the bidirectional language model. This multimodal language model is trained on spoken language data that includes both text and audio modalities. We show that embeddings extracted from this model integrate paralinguistic cues into word meanings and can provide vital affective information by applying these multimodal embeddings to the task of speaker emotion recognition. Shao-Yen Tseng, Shri Narayanan, Panayiotis G. Georgiou |
IEEE Signal Process. Lett. | 3 |
| 2020 | Automatic Prediction of Suicidal Risk in Military Couples Using Multimodal Interaction Cues from Couples ConversationsabstractSuicide is a major societal challenge globally, with a wide range of risk factors, from individual health, psychological and behavioral elements to socio-economic aspects. Military personnel, in particular, are at especially high risk. Crisis resources, while helpful, are often constrained by access to clinical visits or therapist availability, especially when needed in a timely manner. There have hence been efforts on identifying whether communication patterns between couples at home can provide preliminary information about potential suicidal behaviors, prior to intervention. In this work, we investigate whether acoustic, lexical, behavior and turn-taking cues from military couples’ conversations can provide meaningful markers of suicidal risk. We test their effectiveness in real-world noisy conditions by extracting these cues through an automatic diarization and speech recognition front-end. Evaluation is performed by classifying 3 degrees of suicidal risk: none, ideation, attempt. Our automatic system performs significantly better than chance in all classification scenarios and we find that behavior and turn-taking cues are the most informative ones. We also observe that conditioning on factors such as speaker gender and topic of discussion tends to improve classification performance. Sandeep Nallan Chakravarthula, Md. Nasir, Shao-Yen Tseng, Tae Jin Park, Brian R. Baucom, Craig J. Bryan, Shri Narayanan, Panayiotis G. Georgiou |
ICASSP | 9 |
| 2020 | Speaker-Invariant Affective Representation Learning via Adversarial TrainingabstractRepresentation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold-standard references. In addition, there is much variability from input speech signals, human subjective perception of the signals and emotion label ambiguity. In this paper, we propose a machine learning framework to obtain speech emotion representations by limiting the effect of speaker variability in the speech signals. Specifically we propose to disentangle the speaker characteristics from emotion through an adversarial training network in order to better represent emotion. Our method combines the gradient reversal technique with an entropy loss function to remove such speaker information. Our approach is evaluated on both IEMOCAP and CMU-MOSEI datasets. We show that our method improves speech emotion classification and increases generalization to unseen speakers. Jing Huang 0019, Shri Narayanan, Panayiotis G. Georgiou |
ICASSP | 5 |
| 2020 | Transfer learning from adult to children for speech recognition: Evaluation, analysis and recommendations
Prashanth Gurunath Shivakumar, Panayiotis G. Georgiou |
Comput. Speech Lang. | 2 |
| 2019 | Improving the Prediction of Therapist Behaviors in Addiction Counseling by Exploiting Class ConfusionsabstractIn this work we address the problem of joint prosodic and lexical behavioral annotation for addiction counseling. We expand on past work that employed Recurrent Neural Networks (RNNs) on multimodal features by grouping and classifying subsets of classes. We propose two implementations: One is hierarchical classification, which uses the behavior confusion matrix to cluster similar classes and makes the prediction based on a tree structure. The second is a graph-based method which uses the result of the original classification just to find a certain subset of the most probable candidate classes, where the candidate sets of different predicted classes are determined by the class confusions. We make a second prediction with simpler classifier to discriminate the candidates. The evaluation shows that the strict hierarchical approach degrades performance, likely due to error propagation, while the graph-based hierarchy provides significant gains. Zhuohao Chen, Karan Singla, James Gibson, Dogan Can, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 7 |
| 2019 | Role Specific Lattice Rescoring for Speaker Role Recognition from Speech Recognition OutputsabstractThe language patterns followed by different speakers who play specific roles in conversational interactions provide valuable cues for the task of Speaker Role Recognition (SRR). Given the speech signal, existing algorithms typically try to find such patterns in the output of an Automatic Speech Recognition (ASR) system. In this work we propose an alternative way of revealing role-specific linguistic characteristics, by making use of role-specific ASR outputs, which are built by suitably rescoring the lattice produced after a first pass of ASR decoding. That way, we avoid pruning the lattice too early, eliminating the potential risk of information loss. Nikolaos Flemotomos, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
ICASSP | 2 |
| 2019 | Hierarchy-aware Loss Function on a Tree Structured Label Space for Audio Event DetectionabstractThe paper introduces a hierarchy-aware loss function in a Deep Neural Network for an audio event detection task that has a bi-level tree structured label space. The goal is not only to improve audio event detection performance at all levels in the label hierarchy, but also to produce better audio embeddings. We exploit the label tree structure to preserve that information in the hierarchy-aware loss function. Two different loss functions are separately employed. First, a triplet loss with probabilistic multi-level batch mining is introduced. Second, a quadruplet learning method is applied, which is a special case of generalized triplet learning for bi-level label taxonomy. The training is performed in a multi-task learning framework by jointly optimizing cross entropy based loss and hierarchy-aware loss function. The proposed method is found to outperform the baseline cross entropy based models at both levels of the hierarchy. The multi-task model is also able to learn better audio representations as observed in our clustering experiments. Moreover, the model is shown to transfer well when an out-of-domain dataset is used for evaluation. Arindam Jati, Naveen Kumar 0004, Ruxin Chen, Panayiotis G. Georgiou |
ICASSP | 4 |
| 2019 | Predicting Behavior in Cancer-Afflicted Patient and Spouse Interactions Using Speech and LanguageabstractCancer impacts the quality of life of those diagnosed as well as their spouse caregivers, in addition to potentially influencing their day-to-day behaviors. There is evidence that effective communication between spouses can improve well-being related to cancer but it is difficult to efficiently evaluate the quality of daily life interactions using manual annotation frameworks. Automated recognition of behaviors based on the interaction cues of speakers can help analyze interactions in such couples and identify behaviors which are beneficial for effective communication. In this paper, we present and detail a dataset of dyadic interactions in 85 real-life cancer-afflicted couples and a set of observational behavior codes pertaining to interpersonal communication attributes. We describe and employ neural network-based systems for classifying these behaviors based on turn-level acoustic and lexical speech patterns. Furthermore, we investigate the effect of controlling for factors such as gender, patient/caregiver role and conversation content on behavior classification. Analysis of our preliminary results indicates the challenges in this task due to the nature of the targeted behaviors and suggests that techniques incorporating contextual processing might be better suited to tackle this problem. Sandeep Nallan Chakravarthula, Shao-Yen Tseng, Maija Reblin, Panayiotis G. Georgiou |
INTERSPEECH | 5 |
| 2019 | Multi-Task Discriminative Training of Hybrid DNN-TVM Model for Speaker Verification with Noisy and Far-Field Speech
Arindam Jati, Raghuveer Peri, Monisankha Pal, Tae Jin Park, Naveen Kumar 0004, Ruchir Travadi, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 7 |
| 2019 | Modeling Interpersonal Linguistic Coordination in Conversations Using Word Mover's Distanceabstractembeddings and extend it to measure the dissimilarity in language used in multiple consecutive speaker turns. To validate our approach, we apply this measure for two case studies in the clinical psychology domain. We find that our proposed measure is correlated with the therapist's empathy towards their patient in Motivational Interviewing and with affective behaviors in Couples Therapy. In both case studies, our proposed metric exhibits higher correlation than previously proposed measures. When applied to the couples with relationship improvement, we also notice a significant decrease in the proposed measure over the course of therapy, indicating higher linguistic coordination. Md. Nasir, Sandeep Nallan Chakravarthula, Brian R. Baucom, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 5 |
| 2019 | The Second DIHARD Challenge: System Description for USC-SAIL Team
Tae Jin Park, Manoj Kumar 0007, Nikolaos Flemotomos, Monisankha Pal, Raghuveer Peri, Rimita Lahiri, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 7 |
| 2019 | Speaker Diarization with Lexical InformationabstractThis work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with speaker embeddings into a speaker clustering process to improve the overall diarization accuracy. To integrate lexical and acoustic information in a comprehensive way during clustering, we introduce an adjacency matrix integration for spectral clustering. Since words and word boundary information for word-level speaker turn probability estimation are provided by a speech recognition system, our proposed method works without any human intervention for manual transcriptions. We show that the proposed method improves diarization performance on various evaluation datasets compared to the baseline diarization system using acoustic information only in speaker embeddings. Tae Jin Park, Kyu Jeong Han, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 6 |
| 2019 | Spoken Language Intent Detection Using Confusion2VecabstractDecoding speaker's intent is a crucial part of spoken language understanding (SLU). The presence of noise or errors in the text transcriptions, in real life scenarios make the task more challenging. In this paper, we address the spoken language intent detection under noisy conditions imposed by automatic speech recognition (ASR) systems. We propose to employ confusion2vec word feature representation to compensate for the errors made by ASR and to increase the robustness of the SLU system. The confusion2vec, motivated from human speech production and perception, models acoustic relationships between words in addition to the semantic and syntactic relations of words in human language. We hypothesize that ASR often makes errors relating to acoustically similar words, and the confusion2vec with inherent model of acoustic relationships between words is able to compensate for the errors. We demonstrate through experiments on the ATIS benchmark dataset, the robustness of the proposed model to achieve state-of-the-art results under noisy ASR conditions. Our system reduces classification error rate (CER) by 20.84% and improves robustness by 37.48% (lower CER degradation) relative to the previous state-of-the-art going from clean to noisy transcripts. Improvements are also demonstrated when training the intent detection models on noisy transcripts. Prashanth Gurunath Shivakumar, Mu Yang, Panayiotis G. Georgiou |
INTERSPEECH | 3 |
| 2019 | Multiview Shared Subspace Learning Across Speakers and Speech Commands
Krishna Somandepalli, Naveen Kumar 0004, Arindam Jati, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 4 |
| 2019 | Neural Predictive Coding Using Convolutional Neural Networks Toward Unsupervised Learning of Speaker CharacteristicsabstractLearning speaker-specific features is vital in many applications like speaker recognition, diarization, and speech recognition. This paper provides a novel approach, we term neural predictive coding (NPC), to learn speaker-specific characteristics in a completely unsupervised manner from large amounts of unlabeled training data that even contain many non-speech events and multi-speaker audio streams. The NPC framework exploits the proposed short-term active-speaker stationarity hypothesis which assumes two temporally close short speech segments belong to the same speaker, and thus a common representation that can encode the commonalities of both the segments, should capture the vocal characteristics of that speaker. We train a convolutional deep siamese network to produce “speaker embeddings” by learning to separate “same” versus “different” speaker pairs which are generated from an unlabeled data of audio streams. Two sets of experiments are done in different scenarios to evaluate the strength of NPC embeddings and compare with state-of-the-art in-domain supervised methods. First, two speaker identification experiments with different context lengths are performed in a scenario with comparatively limited within-speaker channel variability. NPC embeddings are found to perform the best at short duration experiment, and they provide complementary information to i-vectors for full utterance experiments. Second, a large-scale speaker verification task having a wide range of within-speaker channel variability is adopted as an upper-bound experiment where comparisons are drawn with in-domain supervised methods. Arindam Jati, Panayiotis G. Georgiou |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Towards Predicting Physiology from Speech During Stressful Conversations: Heart Rate and Respiratory Sinus ArrhythmiaabstractBeing affected by mental stress during conversations might have a direct or indirect effect on our speech acoustics as well as on our physiological responses. This paper presents a study on finding the relationship between these two modalities, speech acoustics and physiology, during stressful conversations between humans. Heart rate and respiratory sinus arrhythmia have been considered as physiological variables in the present study. Two datasets, one from stress induction sessions and the other one from in-lab discussions of relationship conflicts between couples, have been analyzed. A series of experiments have been performed separately on the two datasets, as well as on the combined dataset. The research finds acoustic features that are significantly correlated with the physiological variables during stressful conversations. It also predicts the physiological signals from speech features through a nonlinear regression analysis. The results take us one step forward towards building an extremely non-intrusive and relatively inexpensive method of predicting physiological responses from speech, and thus detecting the presence and quantifying the intensity of stress during stressful conversations. Arindam Jati, Paula G. Williams, Brian R. Baucom, Panayiotis G. Georgiou |
ICASSP | 4 |
| 2018 | A Deep Reinforcement Learning Framework for Identifying Funny Scenes in MoviesabstractThis paper presents a novel deep Reinforcement Learning (RL) framework for classifying movie scenes based on affect using the face images detected in the video stream as input. Extracting affective information from the video is a challenging task modulating complex visual and temporal representations intertwined with the complex aspects of human perception and information integration. This also makes it difficult to collect a large annotated corpus restricting the use of supervised learning methods. We present an alternative learning framework based on RL that is tolerant to label sparsity and can easily make use of any available ground truth in an online fashion. We employ this modified RL model for the binary classification of whether a scene is funny or not on a dataset of movie scene clips. The results show that our model correctly predicts 72.95% of the time on the 2-3 minute long movie scenes while on shorter scenes the accuracy obtained is 84.13%. Naveen Kumar 0004, Ruxin Chen, Panayiotis G. Georgiou |
ICASSP | 4 |
| 2018 | "Honey, I Learned to Talk": Multimodal Fusion for Behavior AnalysisabstractIn this work we analyze the importance of lexical and acoustic modalities in behavioral expression and perception. We demonstrate that this importance relates to the amount of therapy, and hence communication training, that a person received. It also exhibits some relationship to gender. We proceed to provide an analysis on couple therapy data by splitting the data into clusters based on gender or stage in therapy. Our analysis demonstrates the significant difference between optimal modality weights per cluster and relationship to therapy stage. Given this finding we propose the use of communication-skill aware fusion models to account for these differences in modality importance. The fusion models operate on partitions of the data according to the gender of the speaker or the therapy stage of the couple. We show that while most multimodal fusion methods can improve mean absolute error of behavioral estimates, the best results are given by a model that considers the degree of communication training among the interlocutors. Shao-Yen Tseng, Brian R. Baucom, Panayiotis G. Georgiou |
ICMI | 4 |
| 2018 | Modeling Interpersonal Influence of Verbal Behavior in Couples Therapy Dyadic InteractionsabstractDyadic interactions among humans are marked by speakers continuously influencing and reacting to each other in terms of responses and behaviors, among others. Understanding how interpersonal dynamics affect behavior is important for successful treatment in psychotherapy domains. Traditional schemes that automatically identify behavior for this purpose have often looked at only the target speaker. In this work, we propose a Markov model of how a target speaker's behavior is influenced by their own past behavior as well as their perception of their partner's behavior, based on lexical features. Apart from incorporating additional potentially useful information, our model can also control the degree to which the partner affects the target speaker. We evaluate our proposed model on the task of classifying Negative behavior in Couples Therapy and show that it is more accurate than the single-speaker model. Furthermore, we investigate the degree to which the optimal influence relates to how well a couple does on the long-term, via relating to relationship outcomes Sandeep Nallan Chakravarthula, Brian R. Baucom, Panayiotis G. Georgiou |
INTERSPEECH | 3 |
| 2018 | An Unsupervised Neural Prediction Framework for Learning Speaker Embeddings Using Recurrent Neural Networks
Arindam Jati, Panayiotis G. Georgiou |
INTERSPEECH | 2 |
| 2018 | Towards an Unsupervised Entrainment Distance in Conversational Speech Using Deep Neural NetworksabstractEntrainment is a known adaptation mechanism that causes interaction participants to adapt or synchronize their acoustic characteristics. Understanding how interlocutors tend to adapt to each other's speaking style through entrainment involves measuring a range of acoustic features and comparing those via multiple signal comparison methods. In this work, we present a turn-level distance measure obtained in an unsupervised manner using a Deep Neural Network (DNN) model, which we call Neural Entrainment Distance (NED). This metric establishes a framework that learns an embedding from the population-wide entrainment in an unlabeled training corpus. We use the framework for a set of acoustic features and validate the measure experimentally by showing its efficacy in distinguishing real conversations from fake ones created by randomly shuffling speaker turns. Moreover, we show real world evidence of the validity of the proposed measure. We find that high value of NED is associated with high ratings of emotional bond in suicide assessment interviews, which is consistent with prior studies. Md. Nasir, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou |
INTERSPEECH | 4 |
| 2018 | Multimodal Speaker Segmentation and Diarization Using Lexical and Acoustic Cues via Sequence to Sequence Neural NetworksabstractWhile there has been substantial amount of work in speaker diarization recently, there are few efforts in jointly employing lexical and acoustic information for speaker segmentation. Towards that, we investigate a speaker diarization system using a sequence-to-sequence neural network trained on both lexical and acoustic features. We also propose a loss function that allows for selecting not only the speaker change points but also the best speaker at any time by allowing for different speaker groupings. We incorporate Mel Frequency Cepstral Coefficients (MFCC) as an acoustic feature alongside lexical information that are obtained from conversations from the Fisher dataset. Thus, we show that acoustics provide complementary information to the lexical modality. The experimental results show that sequence-to-sequence system trained on both word sequences and MFCC can improve on speaker diarization result compared to the system that only relies on lexical modality or the baseline MFCC-based system. In addition, we test the performance of our proposed method with Automatic Speech Recognition (ASR) transcripts. While the performance on ASR transcripts drops, the Diarization Error Rate (DER) of our proposed method still outperforms the traditional method based on Bayesian Information Criterion (BIC). Tae Jin Park, Panayiotis G. Georgiou |
INTERSPEECH | 2 |
| 2017 | Exploring sparse representation measures of physiological synchrony for romantic couplesabstractQuantifying the inherent coordination between interacting individuals can afford us new insights into their emotions, communicative intent, and relationship quality. We propose a novel framework to capture the physiological synchrony between romantic partners through sparse representation techniques and appropriately designed parametric dictionaries that take into account the characteristic structure of the considered signals. Physiological synchrony is operationalized as the similarity of co-occurring electrodermal activity (EDA) streams captured through the distance in the corresponding parametric representation space, as well as through the joint signal representation errors. Results indicate that the proposed sparse EDA synchrony measures (SESM)-evaluated on two datasets of couples' interactions-differ across tasks of various emotional intensity and are associated with the partners' attachment style. These results provide a foundation towards designing novel descriptors of interaction and physiological linkage between individuals for emerging affective computing applications. Theodora Chaspari, Adela C. Timmons, Brian R. Baucom, Laura Perrone, Katherine J. W. Baucom, Panayiotis G. Georgiou, Gayla Margolin, Shri Narayanan |
ACII | 6 |
| 2017 | Unsupervised latent behavior manifold learning from acoustic features: Audio2behaviorabstractBehavioral annotation using signal processing and machine learning is highly dependent on training data and manual annotations of behavioral labels. Previous studies have shown that speech information encodes significant behavioral information and be used in a variety of automated behavior recognition tasks. However, extracting behavior information from speech is still a difficult task due to the sparseness of training data coupled with the complex, high-dimensionality of speech, and the complex and multiple information streams it encodes. In this work we exploit the slow varying properties of human behavior. We hypothesize that nearby segments of speech share the same behavioral context and hence share a similar underlying representation in a latent space. Specifically, we propose a Deep Neural Network (DNN) model to connect behavioral context and derive the behavioral manifold in an unsupervised manner. We evaluate the proposed manifold in the couples therapy domain and also provide examples from publicly available data (e.g. stand-up comedy). We further investigate training within the couples' therapy domain and from movie data. The results are extremely encouraging and promise improved behavioral quantification in an unsupervised manner and warrants further investigation in a range of applications. Brian R. Baucom, Panayiotis G. Georgiou |
ICASSP | 3 |
| 2017 | Attention Networks for Modeling Behaviors in Addiction Counseling
James Gibson, Dogan Can, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 3 |
| 2017 | Speaker2Vec: Unsupervised Learning and Adaptation of a Speaker Manifold Using Deep Neural Networks with an Evaluation on Speaker Segmentation
Arindam Jati, Panayiotis G. Georgiou |
INTERSPEECH | 2 |
| 2017 | Exploiting Intra-Annotator Rating Consistency Through Copeland's Method for Estimation of Ground Truth Labels in Couples' Therapy
Karel Mundnich, Md. Nasir, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2017 | Complexity in Speech and its Relation to Emotional Bond in Therapist-Patient Interactions During Suicide Risk Assessment Interviews
Md. Nasir, Brian R. Baucom, Craig J. Bryan, Shri Narayanan, Panayiotis G. Georgiou |
INTERSPEECH | 5 |
| 2017 | Approaching Human Performance in Behavior Estimation in Couples Therapy Using Deep Sentence Embeddings
Shao-Yen Tseng, Brian R. Baucom, Panayiotis G. Georgiou |
INTERSPEECH | 3 |
| 2017 | Multiple Instance Learning for Behavioral CodingabstractWe propose a computational methodology for automatically estimating human behavioral patterns using the multiple instance learning (MIL) paradigm. We describe the incremental diverse density algorithm, a particular formulation of multiple instance learning, and discuss its suitability for behavioral coding. We use a rich multi-modal corpus comprised of chronically distressed married couples having problem-solving discussions as a case study to experimentally evaluate our approach. In the multiple instance learning framework, we treat each discussion as a collection of short-term behavioral expressions which are manifested in the acoustic, lexical, and visual channels. We experimentally demonstrate that this approach successfully learns representations that carry relevant information about the behavioral coding task. Furthermore, we employ this methodology to gain novel insights into human behavioral data, such as the local versus global nature of behavioral constructs as well as the level of ambiguity in the expression of behaviors through each respective modality. Finally, we assess the success of each modality for behavioral classification and compare schemes for multimodal fusion within the proposed framework. James Gibson, Athanasios Katsamanis, Francisco Romero, Bo Xiao 0003, Panayiotis G. Georgiou, Shri Narayanan |
IEEE Trans. Affect. Comput. | 5 |
| 2016 | A Deep Learning Approach to Modeling Empathy in Addiction Counseling
James Gibson, Dogan Can, Bo Xiao 0003, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 6 |
| 2016 | Laughter Valence Prediction in Motivational Interviewing Based on Lexical and Acoustic Cues
Rahul Gupta 0001, Nishant Nath, Taruna Agrawal, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 4 |
| 2016 | Robust Multichannel Gender Classification from Speech in Movie Audio
Naveen Kumar 0004, Md. Nasir, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2016 | Sparsely Connected and Disjointly Trained Deep Neural Networks for Low Resource Behavioral Annotation: Acoustic Classification in Couples' TherapyabstractObservational studies are based on accurate assessment of human state. A behavior recognition system that models interlocutors' state in real-time can significantly aid the mental health domain. However, behavior recognition from speech remains a challenging task since it is difficult to find generalizable and representative features because of noisy and high-dimensional data, especially when data is limited and annotated coarsely and subjectively. Deep Neural Networks (DNN) have shown promise in a wide range of machine learning tasks, but for Behavioral Signal Processing (BSP) tasks their application has been constrained due to limited quantity of data. We propose a Sparsely-Connected and Disjointly-Trained DNN (SD-DNN) framework to deal with limited data. First, we break the acoustic feature set into subsets and train multiple distinct classifiers. Then, the hidden layers of these classifiers become parts of a deeper network that integrates all feature streams. The overall system allows for full connectivity while limiting the number of parameters trained at any time and allows convergence possible with even limited data. We present results on multiple behavior codes in the couples' therapy domain and demonstrate the benefits in behavior classification accuracy. We also show the viability of this system towards live behavior annotations. Brian R. Baucom, Panayiotis G. Georgiou |
INTERSPEECH | 3 |
| 2016 | Complexity in Prosody: A Nonlinear Dynamical Systems Approach for Dyadic Conversations; Behavior and Outcomes in Couples Therapy
Md. Nasir, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou |
INTERSPEECH | 4 |
| 2016 | Multimodal Fusion of Multirate Acoustic, Prosodic, and Lexical Speaker Characteristics for Native Language Identification
Prashanth Gurunath Shivakumar, Sandeep Nallan Chakravarthula, Panayiotis G. Georgiou |
INTERSPEECH | 3 |
| 2016 | Perception Optimized Deep Denoising AutoEncoders for Speech Enhancement
Prashanth Gurunath Shivakumar, Panayiotis G. Georgiou |
INTERSPEECH | 2 |
| 2016 | Couples Behavior Modeling and Annotation Using Low-Resource LSTM Language Models
Shao-Yen Tseng, Sandeep Nallan Chakravarthula, Brian R. Baucom, Panayiotis G. Georgiou |
INTERSPEECH | 4 |
| 2016 | Behavioral Coding of Therapist Language in Addiction Counseling Using Recurrent Neural Networks
Bo Xiao 0003, Dogan Can, James Gibson, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 6 |
| 2015 | Modeling head motion entrainment for prediction of couples' behavioral characteristicsabstractOur work examines the link between head motion entrainment of interacting couples and human expert's judgment on certain overall behavioral characteristics (e.g., Blame patterns). We employ a data-driven model that clusters head motion in an unsupervised manner into elementary types called kinemes. We propose three groups of similarity measures based on Kullback-Leibler divergence to model entrainment. We find that the divergence of the (joint) distribution of kinemes yields consistent and significant correlation with target behavior characteristics. The divergence of the conditional distribution of kinemes is shown to predict the polarity of the behavioral characteristics. We partly explain the strong correlations via associating the conditional distributions with the prominent behavioral implications of their respective associated kinemes. These results show the possibility of inferring human behavioral characteristics through the modeling of dyadic head motion entrainment. Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan |
ACII | 2 |
| 2015 | A language-based generative model framework for behavioral analysis of couples' therapyabstractObservational studies for psychological evaluations rely on careful assessment of multiple behavioral cues. Recent studies have made good progress in automating the psychological evaluation, which often involved tedious manual annotation of a set of behavioral codes. However, the current methods impose strict and often unnatural assumptions for evaluation. In this work, we specifically investigate two goals: (1) Human behavior changes throughout an interaction and better models of this evolution can improve automated behavioral annotation and (2) Human perception of this evolution can be quite complex and non-linear and better techniques than averaging need to be investigated. For this purpose, we propose a Dynamic Behavior Modeling (DBM) scheme, which models a spouse as undergoing changes in behavioral state within a session, and contrast it against a Static Behavior Model (SBM) which allows only a constant session-long behavioral state. We use Negativity in a couples therapy task as our case study. We present results and analysis on both models for capturing the local behavior information and predicting the session level negativity label. Sandeep Nallan Chakravarthula, Rahul Gupta 0001, Brian R. Baucom, Panayiotis G. Georgiou |
ICASSP | 4 |
| 2015 | Quantifying EDA synchrony through joint sparse representation: A case-study of couples' interactionsabstractThe co-variation degree between individuals in their physiological signals can reveal insights about the quality of their interaction as well as their personal characteristics. In an effort to capture the amount of synchrony between Electrodermal Activity (EDA) streams occurring in parallel during dyadic interactions, we propose Sparse EDA Synchrony Measure (SESM), an index derived from the joint sparse representation of EDA ensembles. Sparse decomposition is performed using Simultaneous Orthogonal Matching Pursuit (SOMP) from a knowledge-driven dictionary of tonic and phasic atoms, capturing the slow-varying trends and high-frequency signal fluctuations, respectively. At each iteration the atom having the maximum average correlation with the residuals is selected. We compute SESM as the negative natural logarithm of the joint reconstruction error and evaluate it with data from interactions of married and young dating couples participating in tasks of varying emotional intensity. Through statistical analysis and multiple linear regression experiments, our results indicate that SESM depicts significant differences across tasks in both datasets considered and can be associated to individuals' attachment-related characteristics. Theodora Chaspari, Brian R. Baucom, Adela C. Timmons, Andreas Tsiartas, Larissa Borofsky Del Piero, Katherine J. W. Baucom, Panayiotis G. Georgiou, Gayla Margolin, Shri Narayanan |
ICASSP | 7 |
| 2015 | Redundancy analysis of behavioral coding for couples therapy and improved estimation of behavior from noisy annotationsabstractAssessment and quantification of behavior is an important research objective in the recently developed field of behavioral signal processing. This paper focuses on the estimation of behavior from noisy human assessment. It aims to address the redundancy of behavioral descriptors for couples therapy by introducing a lower-dimensional representation of the behavioral space. We present an improved method for estimating the ground truth of behavioral ratings from assessment by multiple experts or annotators. The results show improved estimation performance using the proposed method and provide an insightful analysis of reconstruction error and decorrelation of annotator bias in the reduced behavioral space. Md. Nasir, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 3 |
| 2015 | Automated evaluation of non-native English pronunciation quality: combining knowledge- and data-driven features at multiple time scalesabstractAutomatically evaluating pronunciation quality of non-native speech has seen tremendous success in both research and com-mercial settings, with applications in L2 learning. In this paper, submitted for the INTERSPEECH 2015 Degree of Nativeness Sub-Challenge, this problem is posed under a challenging cross-corpora setting using speech data drawn from multiple speakers from a variety of language backgrounds (L1) reading different English sentences. Since the perception of non-nativeness is re-alized at the segmental and suprasegmental linguistic levels, we explore a number of acoustic cues at multiple time scales. We experiment with both data-driven and knowledge-inspired fea-tures that capture degree of nativeness from pauses in speech, speaking rate, rhythm/stress, and goodness of phone pronunci-ation. One promising finding is that highly accurate automated assessment can be attained using a small diverse set of intuitive and interpretable features. Performance is further boosted by smoothing scores across utterances from the same speaker; our best system significantly outperforms the challenge baseline. Matthew Black, Daniel Bone, Z.-I. Skordilis, Rahul Gupta 0001, Pavlos Papadopoulos, Sandeep Nallan Chakravarthula, Bo Xiao 0003, Maarten Van Segbroeck, Jangwon Kim, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 11 |
| 2015 | Assessing empathy using static and dynamic behavior models based on therapist's language in addiction counseling
Sandeep Nallan Chakravarthula, Bo Xiao 0003, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou |
INTERSPEECH | 5 |
| 2015 | Analysis and modeling of the role of laughter in motivational interviewing based psychotherapy conversations
Rahul Gupta 0001, Theodora Chaspari, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 3 |
| 2015 | Automatic estimation of parkinson's disease severity from diverse speech tasksabstractThe need for reliable, scalable and efficient diagnosis of Parkin-son’s Disease (PD) is a major clinical need. Automating the diagnosis can lead to more accurate and objective predictions as well as provide insights regarding the nature of Parkinson’s condition. This paper proposes a fully automated system to rate the severity (UPDRS-III scale) of PD from patients ’ speech. Specifically, the system captures atypicalities in an individ-ual’s voice when performing multiple diverse speaking tasks and makes a unified prediction of the PD severity. The perfor-mance is tested in a cross-data setting, with different subjects and dissimilar recording conditions. Results indicate that (i) effective features vary depending on the nature of the specific speech task, (ii) additional novel feature sets to detect distor-tions in Parkinson’s speech significantly improve the prediction accuracy from the Interspeech15 Challenge baseline system and (iii) our fusion system based on an unsupervised clustering tech-nique also improves the accuracy. Our system incorporates i-vector and functionals for segmental features, non-linear time series features, speech rhythm and automatic speech recogni-tion decoding based features. By its application on the Inter-speech15 eating condition challenge, the system also shows its potential for detecting other sources of speech variability. Jangwon Kim, Md. Nasir, Rahul Gupta 0001, Maarten Van Segbroeck, Daniel Bone, Matthew Black, Z.-I. Skordilis, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 9 |
| 2015 | Still together?: the role of acoustic features in predicting marital outcomeabstractThe assessment and prediction of marital outcome in couple therapy has intrigued many clinical psychologists. In this work, we analyze the significance of various acoustic features extracted from couples ’ spoken interaction in predicting the success or failure of their marriage. We also investigate whether speech acoustic features can provide complementary information to behavioral descriptions or codes provided by human experts (e.g., relationship satisfaction, blame patterns, global negativity). We formulate marital outcome prediction as both binary (improvement vs. no improvement) and multiclass (different levels of improvement) classification problem. Our experiments show that acoustic features can predict marital outcome more accurately than those based on behavioral descriptors provided by human experts. We also find that dialog turn-level acoustic features generally perform better than frame-level signal descriptors. This observation supports the notion that the impact of the behavior of one interlocutor on the other is more important than the behavior itself looked in isolation. Finally, acoustic features together with human-derived behavioral codes show the best performance in outcome prediction, suggesting some complementarity in the information captured by these behavioral representations. Md. Nasir, Bo Xiao 0003, Brian R. Baucom, Shri Narayanan, Panayiotis G. Georgiou |
INTERSPEECH | 6 |
| 2015 | A dynamic model for behavioral analysis of couple interactions using acoustic features
James Gibson, Bo Xiao 0003, Brian R. Baucom, Panayiotis G. Georgiou |
INTERSPEECH | 5 |
| 2015 | Analyzing speech rate entrainment and its relation to therapist empathy in drug addiction counseling
Bo Xiao 0003, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 4 |
| 2015 | Head Motion Modeling for Human Behavior Analysis in Dyadic InteractionabstractThis paper presents a computational study of head motion in human interaction, notably of its role in conveying interlocutors' behavioral characteristics. Head motion is physically complex and carries rich information; current modeling approaches based on visual signals, however , are still limited in their ability to adequately capture these important properties. Guided by the methodology of kinesics , we propose a data-driven approach to identify typical head motion patterns. The approach follows the steps of first segmenting motion events, then parametrically representing the motion by linear predictive features, and finally generalizing the motion types using Gaussian mixture models. The proposed approach is experimentally validated using video recordings of communication sessions from real couples involved in a couples therapy study. In particular we use the head motion model to classify binarized expert judgments of the interactants' specific behavioral characteristics where entrainment in head motion is hypothesized to play a role: Acceptance, Blame, Positive, and Negative behavior. We achieve accuracies in the range of 60% to 70% for the various experimental settings and conditions. In addition, we describe a measure of motion similarity between the interaction partners based on the proposed model. We show that the relative change of head motion similarity during the interaction significantly correlates with the expert judgments of the interactants' behavioral characteristics. These findings demonstrate the effectiveness of the proposed head motion model, and underscore the promise of analyzing human behavioral characteristics through signal processing methods. Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan |
IEEE Trans. Multim. | 2 |
| 2014 | Barista: A framework for concurrent speech processing by usc-sailabstractWe present Barista, an open-source framework for concurrent speech processing based on the Kaldi speech recognition toolkit and the libcppa actor library. With Barista, we aim to provide an easy-to-use, extensible framework for constructing highly customizable concurrent (and/or distributed) networks for a variety of speech processing tasks. Each Barista network specifies a flow of data between simple actors, concurrent entities communicating by message passing, modeled after Kaldi tools. Leveraging the fast and reliable concurrency and distribution mechanisms provided by libcppa, Barista lets demanding speech processing tasks, such as real-time speech recognizers and complex training workflows, to be scheduled and executed on parallel (and/or distributed) hardware. Barista is released under the Apache License v2.0. Dogan Can, James Gibson, Colin Vaz, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 4 |
| 2014 | Classification of clean and noisy bilingual movie audio for speech-to-speech translation corpora designabstractIdentifying suitable sources of bilingual audio and text data is a crucial part of statistical Speech to Speech (S2S) research and development. Movies, often dubbed in other languages, offer a good source for this purpose; but not all data are directly usable because of noise and other audio condition differences. Hence, automatically selecting the bilingual audio data that are suitable for analysis, and training S2S systems for specific environments becomes crucial. In this work, we extract bilingual speech segments from movies and aim at classifying segments as clean speech or speech with background noise (i.e. music, babble noise etc.). We examine various features in solving this problem and our best performing method delivers accuracy up to 87% in discriminating clean and noisy speech in bilingual data. Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 3 |
| 2014 | Power-spectral analysis of head motion signal for behavioral modeling in human interactionabstractWe examine whether head motion can be used for predicting human expert's judgments of behavioral characteristics relevant to the couples therapy domain. Specifically we predict “high” or “low” presence of several behavioral characteristics such as “Blame” that are discerned by human experts, through data-driven clustering of the head motion signal based on power-spectral features. We employ the distribution of motion samples in each cluster for behavior judgment prediction. We find clustering horizontal and vertical motion separately is superior to combined clustering in predicting behavior. The performance of gender-specific and gender-independent clustering of head motion is comparable in average while different for each gender. The proposed power-spectral features outperform linear prediction features in average. Using data from a clinical study of distressed couples, we empirically show that the derived clusters quantize head motion into meaningful types that relate to interpretable behavior characteristics. These findings demonstrate the feasibility of inferring behavior characteristics from head motion signals. Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan |
ICASSP | 2 |
| 2014 | Predicting client's inclination towards target behavior change in motivational interviewing and investigating the role of laughter
Rahul Gupta 0001, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 2 |
| 2014 | Unsupervised speaker diarization using riemannian manifold clusteringabstractWe address the problem of speaker clustering for robust unsupervised speaker diarization. We model each speakerhomogeneous segment as one single full multivariate Gaussian probability density function (pdf) and take into consideration the Riemannian property of Gaussian pdfs. By assuming that segments from different speakers lie on different (possibly intersected) sub-manifolds of the manifold of Gaussian pdfs, we formulate the original problem as a Riemannian manifold clustering problem. To apply the computationally simple Riemannian locally linear embedding (LLE) algorithm, we impose a constraint on the length of each segment so as to ensure the fitness of single-Gaussian modeling and to increase the chance that all k-nearest neighbors of a pdf are from the same submanifold (speaker). Experiments on the microphone-recorded conversational interviews from NIST 2010 speaker recognition evaluation set demonstrate promising results of less than 1% Che-Wei Huang, Bo Xiao 0003, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2014 | Modeling therapist empathy through prosody in drug addiction counselingabstractEmpathy measures the capacity of the therapist to experience the same cognitive and emotional dispositions as the patient, and is a key quality factor in counseling. In this work we build computational models to infer the empathy of therapist using prosodic cues. We extract pitch, energy, jitter, shimmer and utterance duration from the speech signal, and normalize and quantize these features in order to estimate the distribution of certain prosodic patterns during each interaction. We find significant correlation between empathy and the distribution of prosodic patterns, and achieve 75% accuracy in classifying therapist empathy levels using this distribution. Experiment results suggest high pitch and energy of the therapist are negatively correlated with empathy. These observations agree with domain literature and human intuition. Bo Xiao 0003, Daniel Bone, Maarten Van Segbroeck, Zac E. Imel, David C. Atkins, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 6 |
| 2014 | Computing vocal entrainment: A signal-derived PCA-based quantification scheme with application to affect analysis in married couple interactions
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 6 |
| 2014 | Theoretical Analysis of Diversity in an Ensemble of Automatic Speech Recognition SystemsabstractDiversity or complementarity of automatic speech recognition (ASR) systems is crucial for achieving a reduction in word error rate (WER) upon fusion using the ROVER algorithm. We present a theoretical proof explaining this often-observed link between ASR system diversity and ROVER performance. This is in contrast to many previous works that have only presented empirical evidence for this link or have focused on designing diverse ASR systems using intuitive algorithmic modifications. We prove that the WER of the ROVER output approximately decomposes into a difference of the average WER of the individual ASR systems and the average WER of the ASR systems with respect to the ROVER output. We refer to the latter quantity as the diversity of the ASR system ensemble because it measures the spread of the ASR hypotheses about the ROVER hypothesis. This result explains the trade-off between the WER of the individual systems and the diversity of the ensemble. We support this result through ROVER experiments using multiple ASR systems trained on standard data sets with the Kaldi toolkit. We use the proposed theorem to explain the lower WERs obtained by ASR confidence-weighted ROVER as compared to word frequency-based ROVER. We also quantify the reduction in ROVER WER with increasing diversity of the N-best list. We finally present a simple discriminative framework for jointly training multiple diverse acoustic models (AMs) based on the proposed theorem. Our framework generalizes and provides a theoretical basis for some recent intuitive modifications to well-known discriminative training criterion for training diverse AMs. Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shri Narayanan |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | On-line genre classification of TV programs using audio contentabstractAutomatic genre classification of TV programs can benefit users in various ways such as allowing for rapid selection of multimedia content. In this paper, we introduce an on-line method that can classify genres of TV programs using audio content. We deploy an acoustic topic model (ATM) which was originally designed to capture contextual information embedded within audio segments. With a dataset based on RAI content, we perform both on-line and off-line classification; we segment audio signals with a fixed length and feed into the system for on-line classification tasks, while we use whole audio signals for off-line tasks. The off-line experimental results suggest that the proposed method using audio content yields competitive performance with conventional methods using audio-visual features and outperforms conventional audio-based approaches. The on-line results show promising results in classifying genre of TV programs with short segments and also suggest that ATM performs better than conventional GMM method if the length of audio segments is longer (>1 second). Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2013 | A study on the effect of prosodic emphasis transfer on overall speech translation qualityabstractDespite the increasing interest in Speech-to-speech (S2S) translation, research and development has focused almost exclusively on the lexical aspects of translation. The importance of transferring prosodic and other paralinguistic information through S2S devices and evaluating its impact on the translation quality are yet to be well established. The novelty in this work is a large scale human evaluation study to test the hypothesis that cross-lingual prosodic emphasis transfer is directly related to the perceived quality of speech translation. This hypothesis is validated at the 0.53-0.54 correlation level on the data sets considered with results significant at p-value=0.01. The second contribution of this work is an evaluation methodology based on crowd sourcing using English-Spanish language bilingual data from two distinct domains and evaluated with over 200 bilingual speakers. We also present lessons learned on this type of S2S subjective experiments when using crowd sourcing. Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2013 | Data driven modeling of head motion towards analysis of behaviors in couple interactionsabstractWe propose a data driven approach for modeling head motion behavior in human dyadic interactions, by establishing a structure for unconstrained natural head movement. Using recordings of couples' conversations in real psychotherapy sessions, we first track the head of each subject, compute the head motion and detect active versus non-active intervals. For detected active intervals, we use a sliding window to collect motion sequences. Linear Prediction Coefficients are used to represent the sequence, based on which we train a Gaussian Mixture Model (GMM) such that each mixture would ideally associate with one type of prototypical movement, which we will refer to as a “kineme”. For each complete interaction session, we compute the sum of posterior probabilities of all sequences over the GMM normalized by session length to predict specific “low” versus “high” expert annotated behavior code scores for Acceptance, Blame, Positive and Negative behaviors. We achieved an overall accuracy of about 70% employing these GMMs. This result shows data driven modeling of head motion provides useful information for human behavioral analysis. Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan |
ICASSP | 2 |
| 2013 | Head motion synchrony and its correlation to affectivity in dyadic interactionsabstractBehavioral synchrony, or entrainment, is a phenomenon of great interest to psychologists and a challenging construct to quantify. In this work we study the synchrony behavior of head motion in human dyadic interactions. We model head motion using Gaussian Mixture Model (GMM) of line spectral frequencies extracted from the motion vectors of the head. We quantify interlocutor head motion similarity through the Kullback-Leibler divergence of the GMM posteriors of their respective motion sequences. We use an audiovisual database of distressed couple interactions, extensively annotated by psychologists, to test two hypotheses using the derived similarity measure. We validate the first hypothesis — that people are more likely to increase their degree of synchrony as the interaction progresses — by comparing the first and second halves of the interaction. The second hypothesis tests if the relative change of the similarity measure from these two halves is significantly correlated with the behavioral annotation by the domain experts. This work underscores the importance of head motion as an interaction cue, and the feasibility of using it in a computational model for synchrony behavior. Bo Xiao 0003, Panayiotis G. Georgiou, Chi-Chun Lee, Brian R. Baucom, Shri Narayanan |
ICME | 2 |
| 2013 | Empirical link between hypothesis diversity and fusion performance in an ensemble of automatic speech recognition systems
Kartik Audhkhasi, Andreas M. Zavou, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Spectro-temporal directional derivative features for automatic speech recognitionabstractWe introduce a novel spectro-temporal representation of speech by applying directional derivative filters to the Melspectrogram, with the aim of improving the robustness of automatic speech recognition. Previous studies have shown that two-dimensional wavelet functions, when tuned to appropriate spectral scales and temporal rates, are able to accurately capture the acoustic modulations of speech, even in high noise conditions. Therefore, spectro-temporal features extracted from the wavelet transformation of the spectrogram, offer additional noise robustness to important signal processing tasks, such as voice activity detection and speech recognition. In this paper, we explore the use of the steerable pyramid, a directional wavelet transform that is common in image processing, to derive a spectro-temporal feature representation of speech that can serve as an alternative to cepstral derivatives and Gabor filterbank features. We discuss their application for the task of robust automatic speech recognition. Experiments conducted on the Aurora-2 database demonstrate their competitive robustness to other state-of-the-art speech features, especially in low signalto-noise ratio conditions. Index Terms: spectro-temporal features, automatic speech recognition, directional wavelet transforms James Gibson, Maarten Van Segbroeck, Antonio Ortega, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 4 |
| 2013 | Annotation and classification of Political advertisements
Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Toward transfer of acoustic cues of emphasis across languages
Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Modeling therapist empathy and vocal entrainment in drug addiction counseling
Bo Xiao 0003, Panayiotis G. Georgiou, Zac E. Imel, David C. Atkins, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Which ASR should I choose for my dialogue system?
Fabrizio Morbini, Kartik Audhkhasi, Kenji Sagae, Ron Artstein, Dogan Can, Panayiotis G. Georgiou, Shri Narayanan, Anton Leuski, David R. Traum |
SIGDIAL Conference | 6 |
| 2013 | Unsupervised data processing for classifier-based speech translator
Emil Ettelaie, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2013 | Enabling effective design of multimodal interfaces for speech-to-speech translation system: An empirical study of longitudinal user behaviors over time and user strategies for coping with errors
JongHo Shin, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2013 | High-quality bilingual subtitle document alignments with application to spontaneous speech translation
Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 3 |
| 2013 | Behavioral Signal Processing: Deriving Human Behavioral Informatics From Speech and LanguageabstractThe expression and experience of human behavior are complex and multimodal and characterized by individual and contextual heterogeneity and variability. Speech and spoken language communication cues offer an important means for measuring and modeling human behavior. Observational research and practice across a variety of domains from commerce to healthcare rely on speech- and language-based informatics for crucial assessment and diagnostic information and for planning and tracking response to an intervention. In this paper, we describe some of the opportunities as well as emerging methodologies and applications of human behavioral signal processing (BSP) technology and algorithms for quantitatively understanding and modeling typical, atypical, and distressed human behavior with a specific focus on speech- and language-based communicative, affective, and social behavior. We describe the three important BSP components of acquiring behavioral data in an ecologically valid manner across laboratory to real-world settings, extracting and analyzing behavioral cues from measured data, and developing models offering predictive and decision-making support. We highlight both the foundational speech and language processing building blocks as well as the novel processing and modeling opportunities. Using examples drawn from specific real-world applications ranging from literacy assessment and autism diagnostics to psychotherapy for addiction and marital well being, we illustrate behavioral informatics applications of these signal processing techniques that contribute to quantifying higher level, often subjectively described, human behavior in a domain-sensitive fashion. Shri Narayanan, Panayiotis G. Georgiou |
Proc. IEEE | 2 |
| 2013 | Toward automating a human behavioral coding system for married couples' interactions using speech acoustic features
Matthew Black, Athanasios Katsamanis, Brian R. Baucom, Chi-Chun Lee, Adam C. Lammert, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
Speech Commun. | 7 |
| 2012 | Analyzing quality of crowd-sourced speech transcriptions of noisy audio for acoustic model adaptationabstractThe accuracy of crowd-sourced speech transcriptions varies depending on a variety of factors. This paper studies the impact of one such factor, namely, the quality of audio. We employed a speech database with babble noise at three SNR levels (clean, 2 dB and -2 dB) and asked workers on Amazon Mechanical Turk to transcribe it. Two interesting observations emerge. First, as expected, the quality of transcripts combined by word frequency based ROVER decreases with decreasing SNR. Further, we demonstrate that the use of some unsupervised reliability scores can improve the transcription quality, with increasing benefits at lower SNR. Second, we do not observe a significant drop in the performance of acoustic models adapted with increasing transcription noise. This highlights the surprising robustness of crowd-sourced transcripts for acoustic model adaptation. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2012 | Multimodal detection of salient behaviors of approach-avoidance in dyadic interactionsabstractApproach-Avoidance (AA) coding is a measure of involvement and immediacy in human dyadic interactions. We focus on analyzing the salient events in interactions that trigger change points in AA code in time, as perceived by domain experts. We employ coarse level visual cues associated with body parts, as well as vocal energy features. Motion vector extraction and body pose estimation techniques are used for extracting visual cues. Functionals of these cues are used as features for SVM based machine learning experiments. We found that the coder's judgments on salient events are related to the short time interval preceding the labeling. We also show that visual cues are the main information source for decision making on salient AA events, and that considering the information from a subset of body parts provides the same information as considering the full set. The mean of absolute value and standard deviation of motion streams are the most effective functionals as feature. We achieve an F-score of 0.55 in detecting salient events using cross-validation with a one-subject-out approach. Bo Xiao 0003, Panayiotis G. Georgiou, Brian R. Baucom, Shri Narayanan |
ICMI | 2 |
| 2012 | A Case Study: Detecting Counselor Reflections in Psychotherapy for Addictions using Linguistic FeaturesabstractMotivational Interviewing (MI) is a goal-oriented psychotherapy, employed in cases such as addiction, which helps clients (i.e., patients) explore and resolve their ambivalence about the problem at hand in a dialog setting. Measuring the counselor’s proficiency with MI has typically been assessed via behavioral coding a time consuming, non-technological approach. This paper examines a computational approach to assessing the quality of MI. Specifically, we focus on a particular aspect of the counselor behavior – reflections – believed to be a critical indicator of MI therapy quality. We automatically tag reflection instances in a maximum entropy Markov modeling framework using several linguistic features with rich contextual information obtained from the session transcripts. We achieve an Fscore of over 80% while gaining insight about the information sources as perceived by the trained annotators. Dogan Can, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 2 |
| 2012 | A Sequential Bayesian Dialog Agent for Computational EthnographyabstractWe present a sequential Bayesian belief update algorithm for an emotional dialog agent’s inference and behavior. This agent’s purpose is to collect usage patterns of natural language description of emotions among a community of speakers, a task which can be seen as a type of computational ethnography. We describe our target application, an emotionally-intelligent agent that can ask questions and learn about emotions through playing the emotion twenty questions (EMO20Q) game. We formalize the agent’s algorithms mathematically and algorithmically and test our model experimentally in an experiment of 45 humancomputer dialogs with a range of emotional words as the independent variable. We found that 44% of these human-computer dialog games are completed successfully, in comparison with earlier work in which human-human dialogs resulted in 85% successful completion on average. Despite being lower than this upper-bound of human performance, especially on difficult emotion words, the subjects rated that the agent’s humanity was 6.1 on a 0 to 10 scale. This indicates that the algorithm we present produces realistic behavior, but that issues of data sparsity may remain. Abe Kazemzadeh, James Gibson, Juanchen Li, Sungbok Lee, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 5 |
| 2012 | Based on Isolated Saliency or Causal Integration? Toward a Better Understanding of Human Annotation Process using Multiple Instance Learning and Sequential Probability Ratio TestabstractHuman perception is capable of integrating local events to generate an overall impression at the global level; this is evident in daily life and is utilized repeatedly in behavioral science studies to bring objective measures into studies of human behavior. In this work, we explore two hypotheses considering whether it is the isolated-saliency or the causal-integration of information that can trigger the global perceptual behavioral ratings as trained annotators engage in tasks of observational coding. We carry out analyses using Multiple Instance Learning and Sequential Probability Ratio Test in a corpus of real and spontaneous distressed couples’ interaction with global sessionlevel abstract behavioral coding done by trained human annotators. We present various analyses based on different behavioral detection schemes demonstrating the potential of utilizing these algorithms in bringing insights into the human annotation process. We further show that while annotating behaviors with more positive impression, annotators gather information throughout the session compared to behaviors with more negative impression, where a single salient instance is enough to trigger the final global decision. Chi-Chun Lee, Athanasios Katsamanis, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2012 | A reranking approach for recognition and classification of speech input in conversational dialogue systemsabstractWe address the challenge of interpreting spoken input in a conversational dialogue system with an approach that aims to exploit the close relationship between the tasks of speech recognition and language understanding through joint modeling of these two tasks. Instead of using a standard pipeline approach where the output of a speech recognizer is the input of a language understanding module, we merge multiple speech recognition and utterance classification hypotheses into one list to be processed by a joint reranking model. We obtain substantially improved performance in language understanding in experiments with thousands of user utterances collected from a deployed spoken dialogue system. Fabrizio Morbini, Kartik Audhkhasi, Ron Artstein, Maarten Van Segbroeck, Kenji Sagae, Panayiotis G. Georgiou, David R. Traum, Shri Narayanan |
SLT | 6 |
| 2011 | "That's Aggravating, Very Aggravating": Is It Possible to Classify Behaviors in Couple Interactions Using Automatically Derived Lexical Features?
Panayiotis G. Georgiou, Matthew Black, Adam C. Lammert, Brian R. Baucom, Shri Narayanan |
ACII (1) | 1 |
| 2011 | EMO20Q Questioner Agent
Abe Kazemzadeh, James Gibson, Panayiotis G. Georgiou, Sungbok Lee, Shri Narayanan |
ACII (2) | 3 |
| 2011 | Emotion Twenty Questions: Toward a Crowd-Sourced Theory of Emotions
Abe Kazemzadeh, Sungbok Lee, Panayiotis G. Georgiou, Shri Narayanan |
ACII (2) | 3 |
| 2011 | Affective State Recognition in Married Couples' Interactions Using PCA-Based Vocal Entrainment Measures with Multiple Instance Learning
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
ACII (2) | 5 |
| 2011 | Accurate transcription of broadcast news speech using multiple noisy transcribers and unsupervised reliability metricsabstractProfessional manual transcription of speech is an expensive and time consuming process. This paper focuses on the problem of combining noisy transcriptions from multiple non-expert transcribers, where the quality of work from each worker varies. Computing transcriber reliability is a difficult task in the absence of gold standard reference transcripts. Three simple metrics for quantifying this reliability without using a gold standard are proposed. We create a database of 1000 Mexican Spanish broadcast news audio clips transcribed by five transcribers each through Amazon Mechanical Turk. Combination of multiple noisy transcripts using these reliability scores improves the word error rate of the combined transcript with respect to the LDC gold standard by 8% relative, and the sentence error rate by 4.1% relative, when compared with a combination without any reliability information. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2011 | Estimation of ordinal approach-avoidance labels in dyadic interactions: Ordinal logistic regression approachabstractBehavioral Signal Processing aims at automating behavioral coding schemes such as those prevalent in psychology and mental health research. This paper describes a method to automatically quantify the approach-and-avoidance (AA) behavior, described by ordinal labels manually assigned by experts using either video-only or video-with-audio. We propose a novel ordinal regression (OR) algorithm and its hidden Markov model (HMM) extension for estimation of AA labels from visual motion capture based and acoustic features. The proposed algorithm transforms the OR to multiple binary classification problems, solves them by independent score-outputting classifiers and fits the cumulative logit logistic regression model with proportional odds (CLLRMP) to vectors of the classifier scores. The time series extension treats labels as states of the HMM with a likelihood function derived from the probabilistic CLLRMP output. We compare performances of the proposed algorithm applying the weighted binary SVMs in the second step (SVM-OLR), its time-series extension (HMM-SVM-OLR) and the baseline multi-class SVM. On the used dyadic interaction dataset the HMM-SVM-OLR achieves the highest estimation accuracies 71.6 % and 65.7 % for AA labels assigned respectively using video-only and video-with-audio. Viktor Rozgic, Bo Xiao 0003, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 5 |
| 2011 | Bilingual audio-subtitle extraction using automatic segmentation of movie audioabstractExtraction of bilingual audio and text data is crucial for designing Speech to Speech (S2S) systems. In this work, we propose an automatic method to segment multilingual audio streams from movies. In addition, the audio streams are aligned with the corresponding subtitles. We found that the proposed method gives 89% perfectly segmented bilingual audio and 6% partially segmented bilingual audio. In addition, the mapping of the audio to the corresponding subtitles has accuracy 91%. Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 3 |
| 2011 | Overlapped speech detection using long-term spectro-temporal similarity in stereo recordingabstractThe problem of detecting overlapped speech in stereo recordings using close-talk microphones is important for a variety of applications including the identification of back-channels, interruptions etc. in a dyadic or multi-party interactions. For detecting overlapped speech, we propose a feature derived using the spectral similarity of two channels over a range of acoustic frames. During overlapped speech frames the proposed spectro-temporal similarity-based feature values decrease and during non-overlapped speech frames the feature values increase due to the presence of cross-talk. Thus the proposed feature helps to discriminate the overlapped speech frames from the non-overlapped ones. Using overlapped speech detection experiments on a dyadic interaction corpus, it is shown that the proposed feature provides a significant improvement ~26% absolute, in the accuracy of detecting the overlapped speech frames when used as an additional feature to the baseline feature obtained from the two channels' intensity profiles. Bo Xiao 0003, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 3 |
| 2011 | Reliability-Weighted Acoustic Model Adaptation Using Crowd Sourced TranscriptionsabstractThis paper focuses on adaptation of acoustic models using speech transcribed by multiple noisy experts. A simple approach involves combining multiple transcripts using word frequency based Recognizer Output Voting Error Reduction (ROVER) followed by adaptation using the combined transcripts. But this assumes that the transcripts being combined are equally reliable. To overcome this assumption, we use two sets of scores to estimate this reliability. The first set is based on answers to some questions given by the transcribers. The second set is derived in an unsupervised way using the word frequency based ROVER transcripts and baseline acoustic models. The overall confidence is a convex combination of these scores and is used to perform a confidence weighted fusion. We adapt the baseline acoustic models using these combined transcripts. Recognition results for a Mexican Spanish ASR system show an absolute improvement of 0.5% in word error rate and 0.9% in sentence error rate. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2011 | "You made me do it": Classification of Blame in Married Couples' Interactions by Fusing Automatically Derived Speech and Language InformationabstractOne of the goals of behavioral signal processing is the automatic prediction of relevant high-level human behaviors from complex, realistic interactions. In this work, we analyze dyadic discussions of married couples and try to classify extreme instances (low/high) of blame expressed from one spouse to another. Since blame can be conveyed through various communicative channels (e.g., speech, language, gestures), we compare two different classification methods in this paper. The first classifier is trained with the conventional static acoustic features and models “how” the spouses spoke. The second is a novel automatic speech recognition-derived classifier, which models “what” the spouses said. We get the best classification performance (82% accuracy) by exploiting the complementarity of these acoustic and lexical information sources through scorelevel fusion of the two classification methods. Index Terms: behavioral signal processing (BSP), couple therapy, blame, acoustic features, lexical features, fusion Matthew Black, Panayiotis G. Georgiou, Athanasios Katsamanis, Brian R. Baucom, Shri Narayanan |
INTERSPEECH | 2 |
| 2011 | Enhancements to the Training Process of Classifier-Based Speech Translator via Topic Modeling
Emil Ettelaie, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2011 | Determining what Questions to Ask, with the Help of Spectral Graph TheoryabstractThis paper considers questions and the objects being asked about to be a graph and formulates the knowledge goal of a question-asking agent in terms of connecting this graph. The game of twenty questions can be thought of as a testbed of such a question-asking agent’s knowledge. If the agent’s knowledge of the domain were completely specified, the goal of questionasking would be to find the answer as quickly as possible and could follow a decision tree approach to narrow down the candidate answers. However, if the agent’s knowledge is incomplete, it must have a secondary goal for the questions it plans: to complete its knowledge. We claim that this secondary goal of a question asking agent can be formulated in terms of spectral graph theory. In particular, disconnected portions of the graph must be connected in a principled way. We show how the eigenvalues of a graph Laplacian of the the question-object adjacency graph can identify whether a set of knowledge contains disconnected components and the zero elements of the powers of the question-object adjacency graph provide a way to identify these questions. We illustrate the approach using an emotion description task. Abe Kazemzadeh, Sungbok Lee, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | An Analysis of PCA-Based Vocal Entrainment Measures in Married Couples' Affective Spoken InteractionsabstractEntrainment has played a crucial role in analyzing marital couples interactions. In this work, we introduce a novel technique for quantifying vocal entrainment based on Principal Component Analysis (PCA). The entrainment measure, as we define in this work, is the amount of preserved variability of one interlocutor’s speaking characteristic when projected onto representing space of the other’s speaking characteristics. Our analysis on real couples interactions shows that when a spouse is rated as having positive emotion, he/she has a higher value of vocal entrainment compared when rated as having negative emotion. We further performed various statistical analyses on the strength and the directionality of vocal entrainment under different affective interaction conditions to bring quantitative insights into the entrainment phenomenon. These analyses along with a baseline prediction model demonstrate the validity and utility of the proposed PCA-based vocal entrainment measure. Index Terms: vocal entrainment, couples therapy, behavioral signal processing, principal component analysis Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 5 |
| 2011 | Acoustic and Visual Cues of Turn-Taking Dynamics in Dyadic InteractionsabstractIn this paper we introduce an empirical study of multimodal cues of turn-taking dynamics in a social interaction context. We first identify pauses, gaps and overlapped speech segments in the dyadic conversation dataset. Second, we define two types of measurements, Mean Equalized Energy (MEE) and Animation Level (AL) on the audio and video channels, respectively. Then, we verify the hypothesis that the speaker with higher MEE or AL is more likely to take the floor after silence or overlapped speech. The results suggest that both the vocal and visual movement energy offer useful cues towards inferring the intention of the interlocutor to grab the floor. Index Terms: turn-taking, cues, equalized energy, motion vector Bo Xiao 0003, Viktor Rozgic, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 5 |
| 2011 | Enhanced Sparse Imputation Techniques for a Robust Speech Recognition Front-EndabstractMissing data techniques (MDTs) have been widely employed and shown to improve speech recognition results under noisy conditions. This paper presents a new technique which improves upon previously proposed sparse imputation techniques relying on the least absolute shrinkage and selection operator (LASSO). LASSO is widely employed in compressive sensing problems. However, the problem with LASSO is that it does not satisfy oracle properties in the event of a highly collinear dictionary, which happens with features extracted from most speech corpora. When we say that a variable selection procedure satisfies the oracle properties, we mean that it enjoys the same performance as though the underlying true model is known. Through experiments on the Aurora 2.0 noisy spoken digits database, we demonstrate that the Least Angle Regression implementation of the Elastic Net (LARS-EN) algorithm is able to better exploit the properties of a collinear dictionary, and thus is significantly more robust in terms of basis selection when compared to LASSO on the continuous digit recognition task with estimated mask. In addition, we investigate the effects and benefits of a good measure of sparsity on speech recognition rates. In particular, we demonstrate that a good measure of sparsity greatly improves speech recognition rates, and that the LARS modification of LASSO and LARS-EN can be terminated early to achieve improved recognition results, even though the estimation error is increased. Qun Feng Tan, Panayiotis G. Georgiou, Shri Narayanan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Using naïve text queries for robust audio information retrievalabstractThe goal of this work is to build an audio information retrieval system which provides users with flexibility in formulating their queries: from audio examples to naïve text. Specifically, the focus of this paper is on using naïve text to create input queries describing the desired information of the users. Using naïve text queries, however, raises interoperability issues between annotation and retrieval processes due to the wide variety of available audio descriptions. In this paper, we propose an intermediate audio description layer (iADL) to solve the interoperability issues between the annotation and retrieval processes. The iADL comprises two axes corresponding to semantic and onomatopoeic descriptions based on human-to-human communication experiments on how humans express sounds verbally. Various text modeling schemes, such as latent semantic analysis (LSA) and latent topic model, are utilized to transform the naïve text onto the proposd iADL. Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan, Shiva Sundaram |
ICASSP | 2 |
| 2010 | Language model adaptation using WWW documents obtained by utterance-based queriesabstractIn this paper, we consider the estimation of topic specific Language Models (LM) by exploiting documents from the World Wide Web (WWW). We focus on the quality of the generated queries and propose a novel query generation method. In contrast to the n-gram based queries used in past works, our approach relies on utterances as queries candidates. The proposed approach does not rely on any language specific information other than the initial in-domain training text. We have conducted experiments with Web texts of size 0-150 million words, and we have shown that despite not using any language specific information, the proposed approach results in up to 1.1% absolute Word Error Rate (WER) improvement as compared to keyword-based approaches. The proposed approach reduces the WER by 6.3% absolute in our experiments, compared to an in-domain LM without considering any Web data. Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2010 | Automatic classification of married couples' behavior using audio featuresabstractIn this work, we analyzed a 96-hour corpus of married couples spontaneously interacting about a problem in their relationship. Each spouse was manually coded with relevant session-level perceptual observations (e.g., level of blame toward other spouse, global positive affect), and our goal was to classify the spouses’ behavior using features derived from the audio signal. Based on automatic segmentation, we extracted prosodic/spectral features to capture global acoustic properties for each spouse. We then trained gender-specific classifiers to predict the behavior of each spouse for six codes. We compare performance for the various factors (across codes, gender, classifier type, and feature type) and discuss future work for this novel and challenging corpus. Index Terms: behavioral signal processing, human behavior analysis, couples therapy, prosody, emotion recognition Matthew Black, Athanasios Katsamanis, Chi-Chun Lee, Adam C. Lammert, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 7 |
| 2010 | Hierarchical classification for speech-to-speech translation
Emil Ettelaie, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2010 | Robust voice activity detection in stereo recording with crosstalkabstractCrosstalk in a stereo recording occurs when the speech from one participant is leaked into the close-talking microphones of the other participants. This crosstalk causes degradation of the voice activity detection (VAD) performance on individual channels, in spite of the strength of the crosstalk signal being lower than that of the participant’s speech. To address this problem, we first detect speech using a standard VAD scheme on the merged signal obtained by adding the signals from two channels and then determine the target channel using a channel selection scheme. Although VAD is performed on a short-term frame basis, we found that the channel selection performance improves with long-term signal information. Experiments using stereo recordings of real conversations demonstrate that the VAD accuracy averaged over both channels improves by 22% (absolute) indicating the robustness of the proposed approach to crosstalk compared to the single channel VAD scheme. Prasanta Kumar Ghosh, Andreas Tsiartas, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | Quantification of prosodic entrainment in affective spontaneous spoken interactions of married couplesabstractInteraction synchrony among interlocutors happens naturally as people adapt their speaking style gradually to promote efficient communication. In this work, we quantify one aspect of interaction synchrony prosodic entrainment, specifically pitch and energy, in married couples’ problem-solving interactions using speech signal-derived measures. Statistical testings demonstrate that some of these measures capture useful information; they show higher values in interactions with couples having high positive attitude compared to high negative attitude. Further, by using quantized entrainment measures employed with statistical symbol sequence matching in a maximum likelihood framework, we obtained 76% accuracy in predicting positive affect vs. negative affect. Chi-Chun Lee, Matthew Black, Athanasios Katsamanis, Adam C. Lammert, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 7 |
| 2010 | A new multichannel multi modal dyadic interaction databaseabstractIn this work we present a new multi-modal database for analysis of participant behaviors in dyadic interactions. This database contains multiple channels with closeand far-field audio, a high definition camera array and motion capture data. Presence of the motion capture allows precise analysis of the body language low-level descriptors and its comparison with similar descriptors derived from video data. Data is manually labeled by multiple human annotators using psychologyinformed guides. This work also presents an initial analysis of approach-avoidance (A-A) behavior. Two sets of annotations are provided, one based on video only and the other obtained by using both the audio and video channels. Additionally, we describe the statistics of interaction descriptors and A-A labels on participants’ roles. Finally we provide an analysis of relations between various non-verbal features and approach/avoidance labels. Viktor Rozgic, Bo Xiao 0003, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 5 |
| 2010 | Automatic speech recognition system channel modelingabstractIn this paper, we present a systems approach for channel mod-eling of an Automatic Speech Recognition (ASR) system. This can have implications in improving speech recognition com-ponents, such as through discriminative language modeling. We simulate the ASR corruption using a phrase-based machine translation system trained between the reference phoneme and output phoneme sequences of a real ASR. We demonstrate that local optimization on the quality of phoneme-to-phoneme map-pings does not directly translate to overall improvement of the entire model. However, we are still able to capitalize on contex-tual information of the phonemes which a simple acoustic dis-tance model is not able to accomplish. Hence we show that the use of longer context results in a significantly improved model of the ASR channel. Qun Feng Tan, Kartik Audhkhasi, Panayiotis G. Georgiou, Emil Ettelaie, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | An N-gram model for unstructured audio signals toward information retrievalabstractAn N-gram modeling approach for unstructured audio signals is introduced with applications to audio information retrieval. The proposed N-gram approach aims to capture local dynamic information in acoustic words within the acoustic topic model framework which assumes an audio signal consists of latent acoustic topics and each topic can be interpreted as a distribution over acoustic words. Experimental results on classifying audio clips from BBC Sound Effects Library according to both semantic and onomatopoeic labels indicate that the proposed N-gram approach performs better than using only a bag-of-words approach by providing complementary local dynamic information. Samuel Kim, Shiva Sundaram, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 3 |
| 2010 | Towards modeling user behavior in interactions mediated through an automated bidirectional speech translation system
JongHo Shin, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2009 | Lattice-based lexical cues for word fragment detection in conversational speechabstractPrevious approaches to the problem of word fragment detection in speech have focussed primarily on acoustic-prosodic features. This paper proposes that the output of a continuous automatic speech recognition (ASR) system can also be used to derive robust lexical features for the task. We hypothesize that the confusion in the word lattice generated by the ASR system can be exploited for detecting word fragments. Two sets of lexical features are proposed -one which is based on the word confusion, and the other based on the pronunciation confusion between the word hypotheses in the lattice. Classification experiments with a support vector machine (SVM) classifier show that these lexical features perform better than the previously proposed acoustic-prosodic features by around 5.20% (relative) on a corpus chosen from the DARPA Transtac Iraqi-English (San Diego) corpus. A combination of both these feature sets improves the word fragment detection accuracy by 11.50% relative to using just the acoustic-prosodic features. Kartik Audhkhasi, Panayiotis G. Georgiou, Shri Narayanan |
ASRU | 2 |
| 2009 | A robust harmony structure modeling scheme for classical music opus identificationabstractA robust algorithm to model the harmony structure of a music piece is proposed. The harmony structure is extracted directly from a music audio signal using a second-order statistic of chroma feature vectors. The method is experimentally shown to be robust against the degradation of chroma feature vectors due to noisy pitch estimation in our classical music opus identification evaluation. To analyze the effects of the noisy pitch estimation, we propose a noise model that describes difference between the oracle chroma feature vectors as obtained from a symbolic representation and those extracted from the rendered audio signal. The results suggest that the harmony structure modeling scheme employing the covariance matrix is more robust than the alternative investigated second-order statistics. The results also show that the proposed method obtains 84.3% accuracy with the symbolic representations and 72.0% with the synthesized audio data, which suggest that the proposed harmony structure modeling method has room for further improvement by addressing the signal processing challenges of pitch extraction, or through employing more robust features. Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 2 |
| 2009 | Robust word boundary detection in spontaneous speech using acoustic and lexical cuesabstractWe consider the problem of word boundary detection in spontaneous speech utterances. Acoustic features have been well explored in the literature in the context of word boundary detection; however, in spontaneous speech of Switchboard-I corpus, we found that the accuracy of word boundary detection using acoustic features is poor (F-score ~ 0.63). We propose a new feature - that captures lexical cues in the context of the word boundary detection problem. We show that including proposed lexical feature along with the usual acoustic features, the accuracy of the word boundary detection improves considerably (F-score ~ 0.81). We also demonstrate the robustness of our proposed feature in presence of different noise levels for additive white and pink noise. Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 3 |
| 2009 | Context-driven automatic bilingual movie subtitle alignmentabstractMovie subtitle alignment is a potentially useful approach for deriving automatically parallel bilingual/multilingual spoken language data for automatic speech translation. In this paper, we consider the movie subtitle alignment task. We propose a distance metric between utterances of different languages based on lexical features derived from bilingual dictionaries. We use the dynamic time warping algorithm to obtain the best alignment. The best F-score of ! 0.713 is obtained using the proposed approach. Index Terms :s ubtitles alignment, dynamic time warping Andreas Tsiartas, Prasanta Kumar Ghosh, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2009 | An Iterative Relative Entropy Minimization-Based Data Selection Approach for n-Gram Model AdaptationabstractPerformance of statistical n-gram language models depends heavily on the amount of training text material and the degree to which the training text matches the domain of interest. The language modeling community is showing a growing interest in using large collections of text (obtainable, for example, from a diverse set of resources on the Internet) to supplement sparse in-domain resources. However, in most cases the style and content of the text harvested from the web differs significantly from the specific nature of these domains. In this paper, we present a relative entropy based method to select subsets of sentences whose n-gram distribution matches the domain of interest. We present results on language model adaptation using two speech recognition tasks: a medium vocabulary medical domain doctor-patient dialog system and a large vocabulary transcription system for European parliamentary plenary speeches (EPPS). We show that the proposed subset selection scheme leads to performance improvements over state of the art speech recognition systems in terms of both speech recognition word error rate (WER) and language model perplexity (PPL). Abhinav Sethy, Panayiotis G. Georgiou, Bhuvana Ramabhadran, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Towards unsupervised training of the classifier-based speech translator
Emil Ettelaie, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2008 | Multimodal Speaker Segmentation in Presence of Overlapped Speech SegmentsabstractWe propose a multimodal speaker segmentation algorithm with two main contributions: First, we suggest a hidden Markov model architecture that performs fusion of the three modalities: a multi-camera system for participant localization, a microphone array for speaker localization, and a speaker identification system; Second, we present a novel method for dealing with overlapped speech segments through a likelihood model of the microphone array observations that uses multiple local maxima of the Steered Power Response Generalized Cross Correlation Phase Transform (SPR-GCC-PHAT) function in the Joint Probabilistic Data Association (JPDA) framework. Results show that the proposed method outperforms standard speaker segmentation systems based on: (a) speaker identification and; (b) microphone array processing, for datasets with the significant portion (27.4%) of overlapped speech, and scores as high as 94.4% on the F-measure scale. Viktor Rozgic, Kyu Jeong Han, Panayiotis G. Georgiou, Shri Narayanan |
ISM | 3 |
| 2008 | The SAIL speaker diarization system for analysis of spontaneous meetingsabstractIn this paper, we propose a novel approach to speaker diarization of spontaneous meetings in our own multimodal SmartRoom environment. The proposed speaker diarization system first applies a sequential clustering concept to segmentation of a given audio data source, and then performs agglomerative hierarchical clustering for speaker-specific classification (or speaker clustering) of speech segments. The speaker clustering algorithm utilizes an incremental Gaussian mixture cluster modeling strategy, and a stopping point estimation method based on information change rate. Through experiments on various meeting conversation data of approximately 200 minutes total length, this system is demonstrated to provide diarization error rate of 18.90% on average. Kyu Jeong Han, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 2 |
| 2008 | Challenging Uncertainty in Query by Humming Systems: A Fingerprinting ApproachabstractRobust data retrieval in the presence of uncertainty is a challenging problem in multimedia information retrieval. In query-by-humming (QBH) systems, uncertainty can arise in query formulation due to user-dependent variability, such as incorrectly hummed notes, and in query transcription due to machine-based errors, such as insertions and deletions. We propose a fingerprinting (FP) algorithm for representing salient melodic information so as to better compare potentially noisy voice queries with target melodies in a database. The FP technique is employed in the QBH system back end; a hidden Markov model (HMM) front end segments and transcribes the hummed audio input into a symbolic representation. The performance of the FP search algorithm is compared to the conventional edit distance (ED) technique. Our retrieval database is built on 1500 MIDI files and evaluated using 400 hummed samples from 80 people with different musical backgrounds. A melody retrieval accuracy of 88% is demonstrated for humming samples from musically trained subjects, and 70% for samples from untrained subjects, for the FP algorithm. In contrast, the widely used ED method achieves 86% and 62% accuracy rates, respectively, for the same samples, thus suggesting that the proposed FP technique is more robust under uncertainty, particularly for queries by musically untrained users. Erdem Ünal, Elaine Chew, Panayiotis G. Georgiou, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Real-Time Monitoring of Participants' Interaction in a Meeting using Audio-Visual SensorsabstractIntelligent environments equipped with audio-visual sensors provide suitable means for automatically monitoring and tracking the behavior, strategies and engagement of the participants in multiperson meetings. In this paper, high-level features are calculated from active speaker segmentations, automatically annotated by our smart room system, to infer the interaction dynamics between the participants. These features include the number and the average duration of each turn, statistics of turn-taking such as time as active speaker, and turn-taking transition patterns between participants. The results show that it is possible to accurately estimate in real-time not only the flow of the interaction, but also how dominant and engaged each participant was during the discussion. These high-level features, which cannot be inferred from any of the individual modalities by themselves, can be useful for summarization, classification, retrieval and (after action) analysis of meetings. Carlos Busso, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP (2) | 2 |
| 2007 | Real-time Emotion Detection System using Speech: Multi-modal Fusion of Different Timescale FeaturesabstractThe goal of this work is to build a real-time emotion detection system which utilizes multi-modal fusion of different timescale features of speech. Conventional spectral and prosody features are used for intra-frame and supra-frame features respectively, and a new information fusion algorithm which takes care of the characteristics of each machine learning algorithm is introduced. In this framework, the proposed system can be associated with additional features, such as lexical or discourse information, in later steps. To verify the realtime system performance, binary decision tasks on angry and neutral emotion are performed using concatenated speech signal simulating realtime conditions. Samuel Kim, Panayiotis G. Georgiou, Sungbok Lee, Shri Narayanan |
MMSP | 2 |
| 2007 | Multimodal Meeting Monitoring: Improvements on Speaker Tracking and Segmentation through a Modified Mixture Particle FilterabstractIn this paper we address improvements to our multimodal system for tracking of meeting participants and speaker segmentation with a focus on the microphone array modality. We propose an algorithm that uses Directions-of-Arrival estimated for each microphone pair as observations and performs tracking of an unknown number of acoustically-active meeting participants and subsequent speaker segmentation. We propose modified mixture particle filter (mMPF) for tracking of acoustic sources in the track-before-detection (TbD) framework. Trajectories of sound sources are reconstructed by the optimal assignment of posterior mixture components produced by mMPF in consecutive frames. Further, we propose a sequential optimal change-point detection algorithm which discovers speech segments in the reconstructed trajectories i.e., performs speaker segmentation. The algorithm is tested on a multi-participant meeting dataset both separately and as a part of the multimodal system. On the task of speaker detection in the multimodal setup we report significant improvement over our previous state of the art implementation. Viktor Rozgic, Carlos Busso, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 3 |
| 2007 | Analyzing the Multimodal Behaviors of Users of a Speech-to-Speech Translation Device by using Concept Matching ScoresabstractWe investigate factors related to interfacing a speech-to-speech translation device with multimodal capabilities. We evaluate the efficacy of the interactions using a measure for meaning transfer, we call concept score. We show that employing a multimodal interface improves translation quality, in this study, by 24%. We also show that while some users require perfect representation of what they said in order to allow transfer, others accept concept degradation to some extent, in median up to 20% in our experiments. An appropriate system strategy is required to recognize this behavior and guide users towards optimum performance points. For example, we show that appropriate feedback is required to guide the users in their choices of translation method, as 13% of the choices users made are worse than the alternatives the system provided. JongHo Shin, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 2 |
| 2007 | Statistical Modeling and Retrieval of Polyphonic MusicabstractIn this article, we propose a solution to the problem of query by example for polyphonic music audio. We first present a generic mid-level representation for audio queries. Unlike previous efforts in the literature, the proposed representation is not dependent on the different spectral characteristics of different musical instruments and the accurate location of note onsets and offsets. This is achieved by first mapping the short term frequency spectrum of consecutive audio frames to the musical space (the spiral array) and defining a tonal identity with respect to center of effect that is generated by the spectral weights of the musical notes. We then use the resulting single dimensional text representations of the audio to create a-gram statistical sequence models to track the tonal characteristics and the behavior of the pieces. After performing appropriate smoothing, we build a collection of melodic n-gram models for testing. Using perplexity-based scoring, we test the likelihood of a sequence of lexical chords (an audio query) given each model in the database collection. Initial results show that, some variations of the input piece appears in the top 5 results 81% of the time for whole melody inputs within a 500 polyphonic melody database. We also tested the retrieval engine for small audio clips. Using 25s segments, variations of the input piece are among the top 5 results 75% of the time. Erdem Ünal, Panayiotis G. Georgiou, Shri Narayanan, Elaine Chew |
MMSP | 2 |
| 2006 | Text data acquisition for domain-specific language models
Abhinav Sethy, Panayiotis G. Georgiou, Shri Narayanan |
EMNLP | 2 |
| 2006 | Speech Recognition Engineering Issues in Speech to Speech Translation System Design for Low Resource Languages and DomainsabstractEngineering automatic speech recognition (ASR) for speech to speech (S2S) translation systems, especially targeting languages and domains that do not have readily available spoken language resources, is immensely challenging due to a number of reasons. In addition to contending with the conventional data-hungry speech acoustic and language modeling needs, these designs have to accommodate varying requirements imposed by the domain needs and characteristics, target device and usage modality (such as phrase-based, or spontaneous free form interactions, with or without visual feedback) and huge spoken language variability arising due to socio-linguistic and cultural differences of the users. This paper, using case studies of creating speech translation systems between English and languages such as Pashto and Farsi, describes some of the practical issues and the solutions that were developed for multilingual ASR development. These include novel acoustic and language modeling strategies such as language adaptive recognition, active-learning based language modeling, class-based language models that can better exploit resource poor language data, efficient search strategies, including N-best and confidence generation to aid multiple hypotheses translation, use of dialog information and clever interface choices to facilitate ASR, and audio interface design for meeting both usability and robustness requirements Shri Narayanan, Panayiotis G. Georgiou, Abhinav Sethy, Dagen Wang, Murtaza Bulut, Shiva Sundaram, Emil Ettelaie, Sankaranarayanan Ananthakrishnan, Horacio Franco, Kristin Precoda, Dimitra Vergyri, Jing Zheng 0001, Wen Wang 0001, Venkata Ramana Rao Gadde, Martin Graciarena, Victor Abrash, Michael W. Frandsen, Colleen Richey |
ICASSP (5) | 2 |
| 2006 | Cross-lingual dialog model for speech to speech translation
Emil Ettelaie, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2006 | How to talk to a hologramabstractThere is a growing need for creating life-like virtual human simulations that can conduct a natural spoken dialog with a human student on a predefined subject. We present an overview of a spoken-dialog system that supports a person interacting with a full-size hologram-like virtual human character in an exhibition kiosk settings. We also give a brief summary of the natural language classification component of the system and describe the experiments we conducted with the system. Anton Leuski, Jarrell Pair, David R. Traum, Peter J. McNerney, Panayiotis G. Georgiou, Ronakkumar Patel |
IUI | 5 |
| 2006 | Selecting relevant text subsets from web-data for building topic specific language models
Abhinav Sethy, Panayiotis G. Georgiou, Shri Narayanan |
HLT-NAACL | 2 |
| 2006 | Maximum likelihood parameter estimation under impulsive conditions, a sub-Gaussian signal approach
Panayiotis G. Georgiou, Chris Kyriakakis |
Signal Process. | 1 |
| 2006 | Robust maximum likelihood source localization: the case for sub-Gaussian versus GaussianabstractIn this paper, we investigate an alternative to the Gaussian density for modeling signals encountered in audio environments. The observation that sound signals are impulsive in nature, combined with the reverberation effects commonly encountered in audio, motivates the use of the sub-Gaussian density. The new sub-Gaussian statistical model and the separable solution of its maximum likelihood estimator are presented. These are used in an array scenario to demonstrate with both simulations and two different microphone arrays the achievable performance gains. The simulations exhibit the robustness of the sub-Gaussian-based method while the real world experiments reveal a significant performance gain, supporting the claim that the sub-Gaussian model is better suited for sound signals Panayiotis G. Georgiou, Chris Kyriakakis |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Transonics: A Practical Speech-to-Speech Translator for English-Farsi Medical Dialogs
Robert S. Belvin, Emil Ettelaie, Sudeep Gandhe, Panayiotis G. Georgiou, Kevin Knight, Daniel Marcu, Scott Millward, Shri Narayanan, Howard Neely, David R. Traum |
ACL | 4 |
| 2005 | Smart room: participant and speaker localization and identificationabstractOur long-term objective is to create smart room technologies that are aware of the users presence and their behavior and can become an active, but not an intrusive, part of the interaction. In this work, we present a multimodal approach for estimating and tracking the location and identity of the participants including the active speaker. Our smart room design contains three user-monitoring systems: four CCD cameras, an omnidirectional camera and a 16 channel microphone array. The various sensory modalities are processed both individually and jointly and it is shown that the multimodal approach results in significantly improved performance in spatial localization, identification and speech activity detection of the participants. Carlos Busso, Sergi Hernanz, Chi-Wei Chu, Soonil Kwon, Sung Lee, Panayiotis G. Georgiou, Isaac Cohen, Shri Narayanan |
ICASSP (2) | 6 |
| 2005 | Building topic specific language models from webdata using competitive models
Abhinav Sethy, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2004 | Speaker identification using supra-segmental pitch pattern dynamicsabstractMost conventional speaker identification systems rely on short-time spectral envelope features. Recent efforts have yielded significant progress by capturing and modeling speaker-specific aspects of long-term information in the spoken language signal such as prosodic, syntactic and other conversational features. Although significant results have been reported, substantial improvements can be made by using detailed models better describing the specific behavior of each feature. We focus on modeling pitch pattern dynamics at different prosodic scale levels. Trends in pitch variation are believed to appear at different time-scales - such as microprosody, accent, phrase and discourse levels - making wavelet analysis of the f0 contour a suitable choice for investigating the corresponding pitch patterns. We then introduce a transform of the f0 contour wavelet coefficients that results in a compact representation and better reveals the spatio-temporal details in the coefficient sequences representation. In turn, the dynamics of the transformed sequence are modeled by a first order Markov chain, at each scale level. Classification is carried out at each level and the scores of the classifiers operating at the different supra-segmental levels are fused together. The proposed method achieves an EER of 4.8% on the NIST 2001 speaker ID extended data task using a 16-conversation subset, based solely on f0-based information. Farhad Farahani, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP (1) | 2 |
| 2004 | Context dependent statistical augmentation of persian transcripts
Panayiotis G. Georgiou, Shri Narayanan, Hooman Shirani Mehr |
INTERSPEECH | 1 |
| 2004 | Creation of a Doctor-Patient Dialogue Corpus Using Standardized Patients
Robert S. Melvin, Win May, Shri Narayanan, Panayiotis G. Georgiou, Shadi Ganjavi |
LREC | 4 |
| 1999 | Alpha-stable robust modeling of background noise for enhanced sound source localizationabstractIn this paper we address the problem of sound source localization in the presence of impulsive noise for application in immersive telepresence and teleconferencing. Traditional Gaussian modeling of noise signals fails when the signals exhibit impulsive behaviour. A new model is used, namely the symmetric /spl alpha/-stable (S/spl alpha/S), which can better account for the outliers that exist in real-world signals. Real data is used to compare the performance of both the Gaussian and the /spl alpha/-stable models. We demonstrate that the astable model gives a much better approximation to the noise signal than the Gaussian model. Furthermore, we study the problem of time delay estimation (TDE) and we demonstrate the shortcomings of TDE techniques based on second-order statistics when the noise is of S/spl alpha/S nature. We propose an alternative to second-order based methods, based on fractional lower-order statistics, and demonstrate the achieved improvement via simulation experiments. Panayiotis G. Georgiou, Panagiotis Tsakalides, Chris Kyriakakis |
ICASSP | 1 |
| 1999 | Alpha-Stable Modeling of Noise and Robust Time-Delay Estimation in the Presence of Impulsive NoiseabstractA new representation of audio noise signals is proposed, based on symmetric /spl alpha/-stable (S/spl alpha/S) distributions in order to better model the outliers that exist in real signals. This representation addresses a shortcoming of the Gaussian model, namely, the fact that it is not well suited for describing signals with impulsive behavior. The /spl alpha/-stable and Gaussian methods are used to model measured noise signals. It is demonstrated that the /spl alpha/-stable distribution, which has heavier tails than the Gaussian distribution, gives a much better approximation to real-world audio signals. The significance of these results is shown by considering the time delay estimation (TDE) problem for source localization in teleimmersion applications. In order to achieve robust sound source localization, a novel time delay estimation approach is proposed. It is based on fractional lower order statistics (FLOS), which mitigate the effects of heavy-tailed noise. An improvement in TDE performance is demonstrated using FLOS that is up to a factor of four better than what can be achieved with second-order statistics. Panayiotis G. Georgiou, Panagiotis Tsakalides, Chris Kyriakakis |
IEEE Trans. Multim. | 1 |