VLDB 2026 Research / reviewers in the wild / expert
Angeliki Metallinou
dblp:87/4411
· DBLP profile ↗
28ranked-venue papers
11as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-authorArtificial intelligence and machine learning · 13 · 4 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Question answering and dialogue systems · 31% Efficient and distributed learning · 24% Information extraction and text analysis · 12% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 64% Collaborative and social computing · 36% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking |
0.6 | 2 | 2020 | MA-DST: Multi-Attention-Based Scalable Dialog State Tracking · AAAI 2020 Discriminative state tracking for spoken dialog systems · ACL (1) 2013 |
Machine learning › Efficient and distributed learning › model compression
embedding compression |
0.4 | 1 | 2019 | Online Embedding Compression for Text Classification Using Low Rank Matrix Factorization · AAAI 2019 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2019 | Online Embedding Compression for Text Classification Using Low Rank Matrix Factorization · AAAI 2019 |
Natural language and speech › Speech recognition and synthesis
spoken language understanding |
0.4 | 1 | 2019 | Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents · AAAI 2019 |
Natural language and speech › Information extraction and text analysis
text classification |
0.4 | 1 | 2019 | Online Embedding Compression for Text Classification Using Low Rank Matrix Factorization · AAAI 2019 |
Machine learning › Transfer learning and domain adaptation › knowledge transfer
unsupervised transfer learning |
0.4 | 1 | 2019 | Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents · AAAI 2019 |
Human-AI interaction
conversational agents |
0.3 | 1 | 2018 | Context Aware Conversational Understanding for Intelligent Agents With a Screen · AAAI 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.2 | 1 | 2014 | Analysis and Predictive Modeling of Body Language Behavior in Dyadic Interactions From Multimodal Interlocutor Cues · IEEE Trans. Multim. 2014 |
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation |
0.1 | 1 | 2019 | Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents · AAAI 2019 |
Algorithms and data structures › numerical linear algebra › matrix factorization
low-rank matrix factorization |
0.1 | 1 | 2019 | Online Embedding Compression for Text Classification Using Low Rank Matrix Factorization · AAAI 2019 |
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems |
0.0 | 1 | 2013 | Discriminative state tracking for spoken dialog systems · ACL (1) 2013 |
Methods — techniques the papers use, named apart from their topics
low-rank matrix factorization · 0.8fixed-point quantization · 0.8cyclically annealed learning rate · 0.8end-to-end training · 0.7deep learning · 0.7transformer · 0.4self-attention · 0.4cross-attention · 0.4pre-training · 0.4ELMo · 0.4support vector regression · 0.2gaussian mixture model · 0.2fisher kernel · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | MA-DST: Multi-Attention-Based Scalable Dialog State TrackingabstractTask oriented dialog agents provide a natural language interface for users to complete their goal. Dialog State Tracking (DST), which is often a core component of these systems, tracks the system's understanding of the user's goal throughout the conversation. To enable accurate multi-domain DST, the model needs to encode dependencies between past utterances and slot semantics and understand the dialog context, including long-range cross-domain references. We introduce a novel architecture for this task to encode the conversation history and slot semantics more robustly by using attention mechanisms at multiple granularities. In particular, we use cross-attention to model relationships between the context and slots at different semantic levels and self-attention to resolve cross-domain coreferences. In addition, our proposed architecture does not rely on knowing the domain ontologies beforehand and can also be used in a zero-shot setting for new domains or unseen slot values. Our model improves the joint goal accuracy by 5% (absolute) in the full-data setting and by up to 2% (absolute) in the zero-shot setting over the present state-of-the-art on the MultiWoZ 2.1 dataset. Adarsh Kumar 0001, Peter Ku, Anuj Kumar Goyal, Angeliki Metallinou, Dilek Hakkani-Tür |
AAAI | 4 |
| 2019 | Online Embedding Compression for Text Classification Using Low Rank Matrix FactorizationabstractDeep learning models have become state of the art for natural language processing (NLP) tasks, however deploying these models in production system poses significant memory constraints. Existing compression methods are either lossy or introduce significant latency. We propose a compression method that leverages low rank matrix factorization during training, to compress the word embedding layer which represents the size bottleneck for most NLP models. Our models are trained, compressed and then further re-trained on the downstream task to recover accuracy while maintaining the reduced size. Empirically, we show that the proposed method can achieve 90% compression with minimal impact in accuracy for sentence classification tasks, and outperforms alternative methods like fixed-point quantization or offline word embedding compression. We also analyze the inference time and storage space for our method through FLOP calculations, showing that we can compress DNN models by a configurable ratio and regain accuracy loss without introducing additional latency compared to fixed point quantization. Finally, we introduce a novel learning rate schedule, the Cyclically Annealed Learning Rate (CALR), which we empirically demonstrate to outperform other popular adaptive learning rate algorithms on a sentence classification benchmark. Anish Acharya, Rahul Goel, Angeliki Metallinou, Inderjit S. Dhillon |
AAAI | 3 |
| 2019 | Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent AgentsabstractUser interaction with voice-powered agents generates large amounts of unlabeled utterances. In this paper, we explore techniques to efficiently transfer the knowledge from these unlabeled utterances to improve model performance on Spoken Language Understanding (SLU) tasks. We use Embeddings from Language Model (ELMo) to take advantage of unlabeled data by learning contextualized word representations. Additionally, we propose ELMo-Light (ELMoL), a faster and simpler unsupervised pre-training method for SLU. Our findings suggest unsupervised pre-training on a large corpora of unlabeled utterances leads to significantly better SLU performance compared to training from scratch and it can even outperform conventional supervised transfer. Additionally, we show that the gains from unsupervised transfer techniques can be further improved by supervised transfer. The improvements are more pronounced in low resource settings and when using only 1000 labeled in-domain samples, our techniques match the performance of training from scratch on 10-15x more labeled in-domain data. Aditya Siddhant, Anuj Kumar Goyal, Angeliki Metallinou |
AAAI | 3 |
| 2018 | Context Aware Conversational Understanding for Intelligent Agents With a ScreenabstractWe describe an intelligent context-aware conversational system that incorporates screen context information to service multimodal user requests. Screen content is used for disambiguation of utterances that refer to screen objects and for enabling the user to act upon screen objects using voice commands. We propose a deep learning architecture that jointly models the user utterance and the screen and incorporates detailed screen content features. Our model is trained to optimize end to end semantic accuracy across contextual and non-contextual functionality, therefore learns the desired behavior directly from the data. We show that this approach outperforms a rule-based alternative, and can be extended in a straightforward manner to new contextual use cases. We perform detailed evaluation of contextual and non-contextual use cases and show that our system displays accurate contextual behavior without degrading the performance of non-contextual user requests. Vishal Ishwar Naik, Angeliki Metallinou, Rahul Goel |
AAAI | 2 |
| 2018 | Human-Habitat for Health (H3): Human-habitat Multimodal Interaction for Promoting Health and Well-being in the Internet of Things EraabstractThis paper presents an introduction to the "Human-Habitat for Health (H3): Human-habitat multimodal interaction for promoting health and well-being in the Internet of Things era" workshop, which was held at the 20th ACM International Conference on Multimodal Interaction on October 16th, 2018, in Boulder, CO, USA. The main theme of the workshop focused on the effect of the physical or virtual environment on individual's behavior, well-being, and health. The H3 workshop included keynote speeches that provided an overview and future directions of the field, as well as presentations including position papers and research contributions. The workshop brought together experts from academia and industry spanning a set of multi-disciplinary fields, including computer science, speech and spoken language understanding, construction science, life-sciences, health sciences, and psychology, to discuss their respective views and identify synergistic and converging research directions and solutions. Theodora Chaspari, Angeliki Metallinou, Leah I. Stein Duker, Amir H. Behzadan |
ICMI | 2 |
| 2018 | Contextual Language Model Adaptation for Conversational AgentsabstractStatistical language models (LM) play a key role in Automatic Speech Recognition (ASR) systems used by conversational agents. These ASR systems should provide a high accuracy under a variety of speaking styles, domains, vocabulary and argots. In this paper, we present a DNN-based method to adapt the LM to each user-agent interaction based on generalized contextual information, by predicting an optimal, context-dependent set of LM interpolation weights. We show that this framework for contextual adaptation provides accuracy improvements under different possible mixture LM partitions that are relevant for both (1) Goal-oriented conversational agents where it's natural to partition the data by the requested application and for (2) Non-goal oriented conversational agents where the data can be partitioned using topic labels that come from predictions of a topic classifier. We obtain a relative WER improvement of 3% with a 1-pass decoding strategy and 6% in a 2-pass decoding framework, over an unadapted model. We also show up to a 15% relative improvement in recognizing named entities which is of significant value for conversational ASR systems. Anirudh Raju, Behnam Hedayatnia, Linda Liu, Ankur Gandhe, Chandra Khatri, Angeliki Metallinou, Anu Venkatesh, Ariya Rastrow |
INTERSPEECH | 6 |
| 2015 | Context-sensitive learning for enhanced audiovisual emotion classification (Extended abstract)abstractHuman emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs) and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations. Angeliki Metallinou, Athanasios Katsamanis, Martin Wöllmer, Florian Eyben, Björn W. Schuller, Shri Narayanan |
ACII | 1 |
| 2015 | Deep neural network acoustic models for spoken assessment applications
Angeliki Metallinou |
Speech Commun. | 3 |
| 2014 | Analysis of interaction attitudes using data-driven hand gesture phrasesabstractHand gesture is one of the most expressive, natural and common types of body language for conveying attitudes and emotions in human interactions. In this paper, we study the role of hand gesture in expressing attitudes of friendliness or conflict towards the interlocutors during interactions. We first employ an unsupervised clustering method using a parallel HMM structure to extract recurring patterns of hand gesture (hand gesture phrases or primitives). We further investigate the validity of the derived hand gesture phrases by examining the correlation of dyad's hand gesture for different interaction types defined by the attitudes of interlocutors. Finally, we model the interaction attitudes with SVM using the dynamics of the derived hand gesture phrases over an interaction. The classification results are promising, suggesting the expressiveness of the derived hand gesture phrases for conveying attitudes and emotions. Angeliki Metallinou, Engin Erzin, Shri Narayanan |
ICASSP | 2 |
| 2014 | Using deep neural networks to improve proficiency assessment for children English language learners
Angeliki Metallinou |
INTERSPEECH | 1 |
| 2014 | Analysis and Predictive Modeling of Body Language Behavior in Dyadic Interactions From Multimodal Interlocutor CuesabstractDuring dyadic interactions, participants adjust their behavior and give feedback continuously in response to the behavior of their interlocutors and the interaction context. In this paper, we study how a participant in a dyadic interaction adapts his/her body language to the behavior of the interlocutor, given the interaction goals and context. We apply a variety of psychology-inspired body language features to describe body motion and posture. We first examine the coordination between the dyad's behavior for two interaction stances: friendly and conflictive. The analysis empirically reveals the dyad's behavior coordination, and helps identify informative interlocutor features with respect to the participant's target body language features. The coordination patterns between the dyad's behavior are found to depend on the interaction stances assumed. We apply a Gaussian-Mixture-Model-based (GMM) statistical mapping in combination with a Fisher kernel framework for automatically predicting the body language of an interacting participant from the speech and gesture behavior of an interlocutor. The experimental results show that the Fisher kernel-based approach outperforms methods using only the GMM-based mapping, and using the support vector regression, in terms of correlation coefficient and RMSE. These results suggest a significant level of predictability of body language behavior from interlocutor cues. Angeliki Metallinou, Shri Narayanan |
IEEE Trans. Multim. | 2 |
| 2013 | Discriminative state tracking for spoken dialog systems
Angeliki Metallinou, Dan Bohus, Jason D. Williams |
ACL (1) | 1 |
| 2013 | Toward body language generation in dyadic interaction settings from interlocutor multimodal cuesabstractDuring dyadic interactions, participants influence each other's verbal and nonverbal behaviors. In this paper, we examine the coordination between a dyad's body language behavior, such as body motion, posture and relative orientation, given the participants' communication goals, e.g., friendly or conflictive, in improvised interactions. We further describe a Gaussian Mixture Model (GMM) based statistical methodology for automatically generating body language of a listener from speech and gesture cues of a speaker. The experimental results show that automatically generated body language trajectories generally follow the trends of observed trajectories, especially for velocities of body and arms, and that the use of speech information improves prediction performance. These results suggest that there is a significant level of predictability of body language in the examined goal-driven improvisations, which could be exploited for interaction-driven and goal-driven body language generation. Angeliki Metallinou, Shri Narayanan |
ICASSP | 2 |
| 2013 | Quantifying atypicality in affective facial expressions of children with autism spectrum disordersabstractWe focus on the analysis, quantification and visualization of atypicality in affective facial expressions of children with High Functioning Autism (HFA). We examine facial Motion Capture data from typically developing (TD) children and children with HFA, using various statistical methods, including Functional Data Analysis, in order to quantify atypical expression characteristics and uncover patterns of expression evolution in the two populations. Our results show that children with HFA display higher asynchrony of motion between facial regions, more rough facial and head motion, and a larger range of facial region motion. Overall, subjects with HFA consistently display a wider variability in the expressive facial gestures that they employ. Our analysis demonstrates the utility of computational approaches for understanding behavioral data and brings new insights into the autism domain regarding the atypicality that is often associated with facial expressions of subjects with HFA. Angeliki Metallinou, Ruth B. Grossman, Shri Narayanan |
ICME | 1 |
| 2013 | Tracking continuous emotional trends of participants during affective dyadic interactions using body language and speech information
Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan |
Image Vis. Comput. | 1 |
| 2013 | Iterative Feature Normalization Scheme for Automatic Emotion Detection from SpeechabstractThe externalization of emotion is intrinsically speaker-dependent. A robust emotion recognition system should be able to compensate for these differences across speakers. A natural approach is to normalize the features before training the classifiers. However, the normalization scheme should not affect the acoustic differences between emotional classes. This study presents the iterative feature normalization (IFN) framework, which is an unsupervised front-end, especially designed for emotion detection. The IFN approach aims to reduce the acoustic differences, between the neutral speech across speakers, while preserving the inter-emotional variability in expressive speech. This goal is achieved by iteratively detecting neutral speech for each speaker, and using this subset to estimate the feature normalization parameters. Then, an affine transformation is applied to both neutral and emotional speech. This process is repeated till the results from the emotion detection system are consistent between consecutive iterations. The IFN approach is exhaustively evaluated using the IEMOCAP database and a data set obtained under free uncontrolled recording conditions with different evaluation configurations. The results show that the systems trained with the IFN approach achieve better performance than systems trained either without normalization or with global normalization. Carlos Busso, Soroosh Mariooryad, Angeliki Metallinou, Shri Narayanan |
IEEE Trans. Affect. Comput. | 3 |
| 2012 | Speaker states recognition using latent factor analysis based Eigenchannel factor vector modelingabstractThis paper presents an automatic speaker state recognition approach which models the factor vectors in the latent factor analysis framework improving upon the Gaussian Mixture Model (GMM) baseline performance. We investigate both intoxicated and affective speaker states. We consider the affective speech signal as the original normal average speech signal being corrupted by the affective channel effects. Rather than reducing the channel variability to enhance the robustness as in the speaker verification task, we directly model the speaker state on the channel factors under the factor analysis framework. In this work, the speaker state factor vectors are extracted and modeled by the latent factor analysis approach in the GMM modeling framework and support vector machine classification method. Experimental results show that the proposed speaker state factor vector modeling system achieved 5.34% and 1.49% unweighted accuracy improvement over the GMM baseline on the intoxicated speech detection task (Alcohol Language Corpus) and the emotion recognition task (IEMOCAP database), respectively. Ming Li 0026, Angeliki Metallinou, Daniel Bone, Shri Narayanan |
ICASSP | 2 |
| 2012 | A hierarchical framework for modeling multimodality and emotional evolution in affective dialogsabstractIncorporating multimodal information and temporal context from speakers during an emotional dialog can contribute to improving performance of automatic emotion recognition systems. Motivated by these issues, we propose a hierarchical framework which models emotional evolution within and between emotional utterances, i.e., at the utterance and dialog level respectively. Our approach can incorporate a variety of generative or discriminative classifiers at each level and provides flexibility and extensibility in terms of multimodal fusion; facial, vocal, head and hand movement cues can be included and fused according to the modality and the emotion classification task. Our results using the multimodal, multi-speaker IEMOCAP database indicate that this framework is well-suited for cases where emotions are expressed multimodally and in context, as in many real-life situations. Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 1 |
| 2012 | Analyzing the memory of BLSTM Neural Networks for enhanced emotion classification in dyadic spoken interactionsabstractRecent studies indicate that bidirectional Long Short-Term Memory (BLSTM) recurrent neural networks are well-suited for automatic emotion recognition systems and may lead to better results than systems applying other widely used classifiers such as Support Vector Machines or feedforward Neural Networks. The good performance of BLSTM emotion recognition systems could be attributed to their ability to model and exploit contextual information self-learned via recurrently connected memory blocks which allows them to incorporate information about how emotion evolves over time. However, the actual amount of bidirectional context that a BLSTM classifier takes into account when classifying an observation has not been investigated so far. This paper presents a methodology to systematically investigate the number of past and future utterance-level observations that are considered to generate an emotion prediction for a given utterance, and to examine to what extent this temporal bidirectional context contributes to the overall BLSTM performance. Martin Wöllmer, Angeliki Metallinou, Athanasios Katsamanis, Björn W. Schuller, Shri Narayanan |
ICASSP | 2 |
| 2012 | Speaker Personality Classification Using Systems Based on Acoustic-Lexical Cues and an Optimal Tree-Structured Bayesian NetworkabstractAutomatic classification of human personality along the Big Five dimensions is an interesting problem with several prac-tical applications. This paper makes some contributions in this regard. First, we propose a few automatically-derived personality-discriminating lexical features which provide infor-mation complementary to the conventional acoustic-prosodic cues. We also design a frame-level Gaussian mixture model based system which adds complimentary information to the sys-tems trained on global statistical functionals. Next, we note that the Big Five dimensions are correlated and thus model the de-pendency between these dimensions in the form of an optimal tree-structured Bayesian network. Our final sub-system con-sists of within class covariance normalization followed by L1-regularized logistic regression. Fusion of all these sub-systems achieves better classification performance than independently trained classifiers using just acoustic features. Kartik Audhkhasi, Angeliki Metallinou, Ming Li 0026, Shri Narayanan |
INTERSPEECH | 2 |
| 2012 | Context-Sensitive Learning for Enhanced Audiovisual Emotion ClassificationabstractHuman emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore, the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs), and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations. Context-sensitive approaches outperform those without context for classification tasks such as discrimination between valence levels or between clusters in the valence-activation space. The analysis of emotional transitions in our database sheds light into the flow of affective expressions, revealing potentially useful patterns. Angeliki Metallinou, Martin Wöllmer, Athanasios Katsamanis, Florian Eyben, Björn W. Schuller, Shri Narayanan |
IEEE Trans. Affect. Comput. | 1 |
| 2011 | Iterative feature normalization for emotional speech detectionabstractContending with signal variability due to source and channel effects is a critical problem in automatic emotion recognition. Any approach in mitigating these effects however has to be done so as to not compromise emotion-relevant information in the signal. A promising approach to this problem has been through feature normalization using features drawn from non-emotional ("neutral") speech samples. This paper considers a scheme for minimizing the inter-speaker differences while still preserving the emotional discrimination of the acoustic features. This can be achieved by estimating the normalization parameters using only neutral speech, and then applying the coefficients to the entire corpus (including emotional set). Specifically, this paper introduces a feature normalization scheme that implements these ideas by iteratively detecting neutral speech and normalizing the features. As the approximation error of the normalization parameters is reduced, the accuracy of the emotion detection system increases. The accuracy of the proposed iterative approach, evaluated across three databases, is only 2.5% lower than the one trained with optimal normalization parameters, and 9.7% higher than the one trained without any normalization scheme. Carlos Busso, Angeliki Metallinou, Shri Narayanan |
ICASSP | 2 |
| 2011 | Tracking changes in continuous emotion states using body language and prosodic cuesabstractHuman expressive interactions are characterized by an ongoing unfolding of verbal and nonverbal cues. Such cues convey the interlocutor's emotional state which is continuous and of variable intensity and clarity over time. In this paper, we examine the emotional content of body language cues describing a participant's posture, relative position and approach/withdraw behaviors during improvised affective interactions, and show that they reflect changes in the participant's activation and dominance levels. Furthermore, we describe a framework for tracking changes in emotional states during an interaction using a statistical mapping between the observed audiovisual cues and the underlying user state. Our approach shows promising results for tracking changes in activation and dominance. Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 1 |
| 2011 | Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM SupervectorsabstractSpeaker state recognition is a challenging problem due to speaker and context variability. Intoxication detection is an important area of paralinguistic speech research with potential real-world applications. In this work, we build upon a base set of various static acoustic features by proposing the combination of several different methods for this learning task. The methods include extracting hierarchical acoustic features, performing iterative speaker normalization, and using a set of GMM supervectors. We obtain an optimal unweighted recall for intoxication recognition using score-level fusion of these subsystems. Unweighted average recall performance is 70.54 % on the test set, an improvement of 4.64 % absolute (7.04 % relative) over the baseline model accuracy of 65.9%. Index Terms: intoxication detection, speaker state, hierarchical features, speaker normalization, GMM supervectors 1. Daniel Bone, Matthew Black, Ming Li 0026, Angeliki Metallinou, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 4 |
| 2010 | Visual emotion recognition using compact facial representations and viseme informationabstractEmotion expression is an essential part of human interaction. Rich emotional information is conveyed through the human face. In this study, we analyze detailed motion-captured facial information of ten speakers of both genders during emotional speech. We derive compact facial representations using methods motivated by Principal Component Analysis and speaker face normalization. Moreover, we model emotional facial movements by conditioning on knowledge of speech-related movements (articulation). We achieve average classification accuracies on the order of 75% for happiness, 50-60% for anger and sadness and 35% for neutrality in speaker independent experiments. We also find that dynamic modeling and the use of viseme information improves recognition accuracy for anger, happiness and sadness, as well as for the overall unweighted performance. Angeliki Metallinou, Carlos Busso, Sungbok Lee, Shri Narayanan |
ICASSP | 1 |
| 2010 | Decision level combination of multiple modalities for recognition and analysis of emotional expressionabstractEmotion is expressed and perceived through multiple modalities. In this work, we model face, voice and head movement cues for emotion recognition and we fuse classifiers using a Bayesian framework. The facial classifier is the best performing followed by the voice and head classifiers and the multiple modalities seem to carry complementary information, especially for happiness. Decision fusion significantly increases the average total unweighted accuracy, from 55% to about 62%. Overall, we achieve average accuracy on the order of 65-75% for emotional states and 30-40% for neutral state using a large multi-speaker, multimodal database. Performance analysis for the case of anger and neutrality suggests a positive correlation between the number of classifiers that performed well and the perceptual salience of the expressed emotion. Angeliki Metallinou, Sungbok Lee, Shri Narayanan |
ICASSP | 1 |
| 2010 | Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional LSTM modelingabstractIn this paper, we apply a context-sensitive technique for multimodal emotion recognition based on feature-level fusion of acoustic and visual cues. We use bidirectional Long Short-Term Memory (BLSTM) networks which, unlike most other emotion recognition approaches, exploit long-range contextual information for modeling the evolution of emotion within a conversation. We focus on recognizing dimensional emotional labels, which enables us to classify both prototypical and nonprototypical emotional expressions contained in a large audiovisual database. Subject-independent experiments on various classification tasks reveal that the BLSTM network approach generally prevails over standard classification techniques such as Hidden Markov Models or Support Vector Machines, and achieves F1-measures of the order of 72 %, 65 %, and 55 % for the discrimination of three clusters in emotional space and the distinction between three levels of valence and activation, respectively. Index Terms: emotion recognition, multimodality, long shortterm memory, hidden markov models, context modeling Martin Wöllmer, Angeliki Metallinou, Florian Eyben, Björn W. Schuller, Shri Narayanan |
INTERSPEECH | 2 |
| 2008 | Audio-Visual Emotion Recognition Using Gaussian Mixture Models for Face and VoiceabstractEmotion expression associated with human communication is known to be a multimodal process. In this work, we investigate the way that emotional information is conveyed by facial and vocal modalities, and how these modalities can be effectively combined to achieve improved emotion recognition accuracy. In particular, the behaviors of different facial regions are studied in detail. We analyze an emotion database recorded from ten speakers (five female, five male), which contains speech and facial marker data. Each individual modality is modeled by Gaussian mixture models (GMMs). Multiple modalities are combined using two different methods: a Bayesian classifier weighting scheme and support vector machines that use post classification accuracies as features. Individual modality recognition performances indicate that anger and sadness have comparable accuracies for facial and vocal modalities, while happiness seems to be more accurately transmitted by facial expressions than voice. The neutral state has the lowest performance, possibly due to the vague definition of neutrality. Cheek regions achieve better emotion recognition accuracy compared to other facial regions. Moreover, classifier combination leads to significantly higher performance, which confirms that training detailed single modality classifiers and combining them at a later stage is an effective approach. Angeliki Metallinou, Sungbok Lee, Shri Narayanan |
ISM | 1 |