Louis-Philippe Morency

dblp:31/739 · DBLP profile ↗
← Back
246ranked-venue papers
19as first author
55since 2021 · last 2026
0000-0001-6376-7696ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 159 · 10 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 83 · 6 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 81 · 9 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 LLA MADRS: Evaluating Open-Source LLMs on Real Clinical Interviews - To Reason or Not to Reason?
abstract
Abstract Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored. We present LLAMADRS, a benchmark for structured clinical assessment from dialogue built on the CAMI corpus of psychiatric interviews, comprising 5,804 expert annotations across 541 sessions. We evaluate 25 open-source models (standard and reasoning-augmented; 0.6B–400B parameters) and generate over 400,000 predictions. Our results demonstrate that strong open-source LLMs achieve item-level accuracy with residual error below clinically substantial thresholds. Additionally, an Item-then-Sum (ITS) strategy, assessing symptoms individually through discrete LLM calls before synthesizing final scores, significantly reduces error relative to Direct Total Score (DTS) prediction across most model architectures and scales, despite reasoning models attempting similar decomposition in the reasoning traces of their DTS predictions. In fact, we find that performance gains attributed to “reasoning” depend fundamentally on prompt design: standard models equipped with structured task definitions and examples match reasoning-augmented counterparts. Among the latter, longer reasoning traces correlate with reduced error; while higher model scale does across both architectures. Our results clarify when and why reasoning helps and offer actionable guidance for deploying LLMs in semi-structured clinical assessment.
Gaoussou Youssouf Kebe, Jeffrey M. Girard, Einat Liebenthal, Justin T. Baker, Fernando De la Torre, Louis-Philippe Morency
Trans. Assoc. Comput. Linguistics6
2025 Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
abstract
Social reasoning abilities are crucial for AI systems to effectively interpret and respond to multimodal human communication and interaction within social contexts.We introduce SOCIAL GENOME, the first benchmark for fine-grained, grounded social reasoning abilities of multimodal models.SOCIAL GENOME contains 272 videos of interactions and 1,486 humanannotated reasoning traces related to inferences about these interactions.These traces contain 5,777 reasoning steps that reference evidence from visual cues, verbal cues, vocal cues, and external knowledge (contextual knowledge external to videos).SOCIAL GENOME is also the first modeling challenge to study external knowledge in social reasoning.SOCIAL GENOME computes metrics to holistically evaluate semantic and structural qualities of modelgenerated social reasoning traces.We demonstrate the utility of SOCIAL GENOME through experiments with state-of-the-art models, identifying performance gaps and opportunities for future research to improve the grounded social reasoning abilities of multimodal models.
Leena Mathur, Marian Qian, Paul Pu Liang, Louis-Philippe Morency
EMNLP4
2025 OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
abstract
In recent years, there has been increasing interest in automatic facial behavior analysis systems from computing communities such as vision, multimodal interaction, robotics, and affective computing. Building upon the widespread utility of prior open-source facial analysis systems, we introduce OpenFace 3.0, an open-source toolkit capable of facial landmark detection, facial action unit detection, eye-gaze estimation, and facial emotion recognition. OpenFace 3.0 contributes a lightweight unified model for facial analysis, trained with a multi-task architecture across diverse populations, head poses, lighting conditions, video resolutions, and facial analysis tasks. By leveraging the benefits of parameter sharing through a unified model and training paradigm, OpenFACE 3.0 exhibits improvements in prediction performance, inference speed, and memory efficiency over similar toolkits and rivals state-of-theart models. Openface 3.0 can be installed and run with a single line of code and operate in real-time without specialized hardware. OpenFace 3.0 code for training models and running the system is freely available for research purposes and supports contributions from the community.
Jiewen Hu, Leena Mathur, Paul Pu Liang, Louis-Philippe Morency
FG4
2025 Instant 3DCG Dance Generation System Based on Music and Dance Composition
abstract
We present a novel system that automatically generates and visualizes 3DCG dance animations based on the user’s preferred music and dance composition. The key technology of the system is a transformer-based diffusion model that produces dance choreographies conditioned on arbitrary inputs of music audio and dance composition. Integrated into a user-friendly GUI, the system allows users to instantly generate and preview multiple dance sequences simply by selecting their desired music and dance composition. This capability supports both creative choreography ideation and effective dance practice.
Ryo Ishii, Shin'ichiro Eitoku, Keigo Fushio, Yoshihide Sato, Louis-Philippe Morency
FG5
2025 CDCGM: Composition-specified Dance Choreography Generation from Music
abstract
Significant research attention has recently been focused on the automatic generation of human dance choreography from music. While several generation models have been proposed, they cannot specify what kind of movements to generate, and as a result, random movements are generated. We therefore propose a generation model called Composition-specified Dance Choreography Generation from Music (CDCGM) that enables creators to specify which dance composition (i.e., type of movement) to take when generating a dance at each time step. We implemented CDCGM by first constructing a new dataset that includes motion captures of breakdancing and time-series annotation data of representative movement types. Evaluation experiments using our corpus showed that CDCGM can generate dances that faithfully reflect the specified dance composition with high quality. Compared to conventional state-of-the-art models, CDCGM is capable of generating quality dances that improve the expressiveness and the degree to which the dance matches the content and timing of the music. We also propose a new application for CDCGM in which users watch newly generated dance choreography simply by entering music and dance composition. The results of a user study evaluation of the application demonstrated that users found the experience of generating dance by specifying any dance composition for any music extremely fun, that it has the potential to greatly contribute to dance choreography and learning, and that there is a strong desire to use this application on a daily basis.
Ryo Ishii, Shin'ichiro Eitoku, Louis-Philippe Morency
FG3
2025 AV-Flow: Transforming Text to Audio-Visual Human-Like Interactions
abstract
We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We demonstrate human-like speech synthesis, synchronized lip motion, lively facial expressions and head pose; all generated from just text characters. The core premise of our approach lies in the architecture of our two parallel diffusion transformers. Intermediate highway connections ensure communication between the audio and visual modalities, and thus, synchronized speech intonation and facial dynamics (e.g., eyebrow motion). Our model is trained with flow matching, leading to expressive results and fast inference. In case of dyadic conversations, AV-Flow produces an always-on avatar, that actively listens and reacts to the audio-visual input of a user. Through extensive experiments, we show that our method outperforms prior work, synthesizing natural-looking 4D talking avatars. Project page: https://aggelinacha.github.io/AV-Flow/
Aggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer, Dimitris Samaras, Alexander Richard
ICCV2
2025 ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
abstract
Recent Large Vision-Language Models (LVLMs) have introduced a new paradigm for understanding and reasoning about image input through textual responses. Although they have achieved remarkable performance across a range of multi-modal tasks, they face the persistent challenge of hallucination, which introduces practical weaknesses and raises concerns about their reliable deployment in real-world applications. Existing work has explored contrastive decoding approaches to mitigate this issue, where the output of the original LVLM is compared and contrasted with that of a perturbed version. However, these methods require two or more queries that slow down LVLM response generation, making them less suitable for real-time applications. To overcome this limitation, we propose ONLY, a training-free decoding approach that requires only a single query and a one-layer intervention during decoding, enabling efficient real-time deployment. Specifically, we enhance textual outputs by selectively amplifying crucial textual information using a text-to-visual entropy ratio for each token. Extensive experimental results demonstrate that our proposed ONLY consistently outperforms state-of-the-art methods across various benchmarks while requiring minimal implementation effort and computational cost. Code is available at https://github.com/zifuwan/ONLY.
Zifu Wan, Ce Zhang 0009, Silong Yong, Martin Q. Ma, Simon Stepputtis, Louis-Philippe Morency, Deva Ramanan, Katia P. Sycara, Yaqi Xie 0001
ICCV6
2025 Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
abstract
While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not align with the given visual input, which restricts their practical applicability in real-world scenarios. In this work, inspired by the observation that the text-to-image generation process is the inverse of image-conditioned response generation in LVLMs, we explore the potential of leveraging text-to-image generative models to assist in mitigating hallucinations in LVLMs. We discover that generative models can offer valuable self-feedback for mitigating hallucinations at both the response and token levels. Building on this insight, we introduce self-correcting Decoding with Generative Feedback (DeGF), a novel training-free algorithm that incorporates feedback from text-to-image generative models into the decoding process to effectively mitigate hallucinations in LVLMs. Specifically, DeGF generates an image from the initial response produced by LVLMs, which acts as an auxiliary visual reference and provides self-feedback to verify and correct the initial response through complementary or contrastive decoding. Extensive experimental results validate the effectiveness of our approach in mitigating diverse types of hallucinations, consistently surpassing state-of-the-art methods across six benchmarks. Code is available at https://github.com/zhangce01/DeGF.
Ce Zhang 0009, Zifu Wan, Zhehan Kan, Martin Q. Ma, Simon Stepputtis, Deva Ramanan, Ruslan Salakhutdinov, Louis-Philippe Morency, Katia P. Sycara, Yaqi Xie 0001
ICLR8
2025 Isolated Causal Effects of Natural Language
abstract
As language technologies become widespread, it is important to understand how changes in language affect reader perceptions and behaviors. These relationships may be formalized as the isolated causal effect of some focal language-encoded intervention (e.g., factual inaccuracies) on an external outcome (e.g., readers’ beliefs). In this paper, we introduce a formal estimation framework for isolated causal effects of language. We show that a core challenge of estimating isolated effects is the need to approximate all non-focal language outside of the intervention. Drawing on the principle of omitted variable bias, we provide measures for evaluating the quality of both non-focal language approximations and isolated effect estimates themselves. We find that poor approximation of non-focal language can lead to bias in the corresponding isolated effect estimates due to omission of relevant variables, and we show how to assess the sensitivity of effect estimates to such bias along the two key axes of fidelity and overlap. In experiments on semi-synthetic and real-world data, we validate the ability of our framework to correctly recover isolated effects and demonstrate the utility of our proposed measures.
Victoria Lin 0001, Louis-Philippe Morency, Eli Ben-Michael
ICML2
2025 Comparative Knowledge Distillation
abstract
In the era of large-scale pretrained models, Knowledge Distillation (KD) serves an important role in transferring the wisdom of computationally-heavy teacher models to lightweight, efficient student models while preserving performance. Yet KD settings often assume readily available access to teacher models capable of performing many in-ferences-a notion increasingly at odds with the realities of costly large-scale models. Addressing this gap, we study an important question: how KD algorithms fare as the number of teacher inferences decreases, a setting we term Reduced-Teacher-Inference Knowledge Distillation (RTI-KD). We observe that the performance of prevalent KD techniques and state-of-the-art data augmentation strategies suffers considerably as the number of teacher inferences is reduced. One class of approaches, termed “relational” knowledge distillation underperforms the rest, yet we hypothesize that they hold promise for reduced dependency on teacher models because they can augment the effective dataset size without additional teacher calls. We find that a simple change - performing high-dimensional comparisons instead of low-dimensional relations, which we term Comparative Knowledge Distillation - vaults performance well over existing KD approaches. We perform empirical evaluation across varied experimental settings and rigorous analysis to understand the learning outcomes of our method. All code is made publicly available.
Alex Tianyi Xu, Alex Wilf, Paul Pu Liang, Alexander Obolenskiv, Daniel Fried, Louis-Philippe Morency
WACV6
2024 Think Twice: Perspective-Taking Improves Large Language Models' Theory-of-Mind Capabilities
abstract
Human interactions are deeply rooted in the interplay of thoughts, beliefs, and desires made possible by Theory of Mind (ToM): our cognitive ability to understand the mental states of ourselves and others.Although ToM may come naturally to us, emulating it presents a challenge to even the most advanced Large Language Models (LLMs).Recent improvements to LLMs' reasoning capabilities from simple yet effective prompting techniques such as Chain-of-Thought (CoT) (Wei et al., 2022) have seen limited applicability to ToM (Gandhi et al., 2023).In this paper, we turn to the prominent cognitive science theory "Simulation Theory" to bridge this gap.We introduce SIMTOM, a novel two-stage prompting framework inspired by Simulation Theory's notion of perspective-taking.To implement this idea on current ToM benchmarks, SIMTOM first filters context based on what the character in question knows before answering a question about their mental state.Our approach, which requires no additional training and minimal prompt-tuning, shows substantial improvement over existing methods, and our analysis reveals the importance of perspective-taking to Theory-of-Mind capabilities.Our findings suggest perspectivetaking as a promising direction for future research into improving LLMs' ToM capabilities.Our code is publicly available.
Alex Wilf, Sihyun Shawn Lee, Paul Pu Liang, Louis-Philippe Morency
ACL (1)4
2024 Global Reward to Local Rewards: Multimodal-Guided Decomposition for Improving Dialogue Agents
abstract
We describe an approach for aligning an LLMbased dialogue agent for long-term social dialogue, where there is only a single global score given by the user at the end of the session.In this paper, we propose the usage of denser naturally-occurring multimodal communicative signals as local implicit feedback to improve the turn-level utterance generation.Therefore, our approach (dubbed GELI) learns a local, turn-level reward model by decomposing the human-provided Global Explicit (GE) sessionlevel reward, using Local Implicit (LI) multimodal reward signals to crossmodally shape the reward decomposition step.This decomposed reward model is then used as part of the RLHF pipeline to improve an LLM-based dialog agent.We run quantitative and qualitative human studies on two large-scale datasets to evaluate the performance of our GELI approach, and find that it shows consistent improvements across various conversational metrics compared to baseline methods.
Dong Won Lee 0007, Hae Park, Cynthia Breazeal, Louis-Philippe Morency
EMNLP5
2024 Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
abstract
Building socially-intelligent AI agents (Social-AI) is a multidisciplinary, multimodal research goal that involves creating agents that can sense, perceive, reason about, learn from, and respond to affect, behavior, and cognition of other agents (human or artificial).Progress towards Social-AI has accelerated in the past decade across several computing communities, including natural language processing, machine learning, robotics, human-machine interaction, computer vision, and speech.Natural language processing, in particular, has been prominent in Social-AI research, as language plays a key role in constructing the social world.In this position paper, we identify a set of underlying technical challenges and open questions for researchers across computing communities to advance Social-AI.We anchor our discussion in the context of social intelligence concepts and prior progress in Social-AI research.
Leena Mathur, Paul Pu Liang, Louis-Philippe Morency
EMNLP3
2024 MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
abstract
Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled.Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like imagetext matching.However, this covers only a subset of real-world interactions.Novel interactions, such as sarcasm expressed through opposing spoken words and gestures or humor expressed through utterances and tone of voice, remain challenging.In this paper, we introduce an approach to enhance multimodal models, which we call Multimodal Mixtures of Experts (MMOE).The key idea in MMOE is to train separate expert models for each type of multimodal interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both modalities are fused.On a sarcasm detection task (MUStARD) and a humor detection task (URFunny), we obtain new state-of-the-art results.MMOE is also able to be applied to various types of models to gain improvement.
Haofei Yu, Zhengyang Qi, Lawrence Jang, Ruslan Salakhutdinov, Louis-Philippe Morency, Paul Pu Liang
EMNLP5
2024 Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
abstract
In many machine learning systems that jointly learn from multiple modalities, a core research question is to understand the nature of multimodal interactions: how modalities combine to provide new task-relevant information that was not present in either alone. We study this challenge of interaction quantification in a semi-supervised setting with only labeled unimodal data and naturally co-occurring multimodal data (e.g., unlabeled images and captions, video and corresponding audio) but when labeling them is time-consuming. Using a precise information-theoretic definition of interactions, our key contribution is the derivation of lower and upper bounds to quantify the amount of multimodal interactions in this semi-supervised setting. We propose two lower bounds: one based on the shared information between modalities and the other based on disagreement between separately trained unimodal classifiers, and derive an upper bound through connections to approximate algorithms for min-entropy couplings. We validate these estimated bounds and show how they accurately track true interactions. Finally, we show how these theoretical results can be used to estimate multimodal model performance, guide data collection, and select appropriate multimodal models for various tasks.
Paul Pu Liang, Chun Kai Ling, Alex Obolenskiy, Rohan Pandey, Alex Wilf, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR8
2024 SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
abstract
*Humans are social beings*; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and *interact* under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
Hao Zhu 0011, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, Maarten Sap
ICLR7
2024 SMURF: Statistical Modality Uniqueness and Redundancy Factorization
abstract
Multimodal late fusion is a well-performing fusion method that sums the outputs of separately processed modalities, so-called modality contributions, to create a prediction; for example, summing contributions from vision, acoustic, and language to predict affective states. In this paper, our primary goal is to improve the interpretability of what modalities contribute to the prediction in late fusion models. More specifically, we want to factorize modality contributions into what is consistently shared by at least two modalities (pairwise redundant contributions) and what the remaining modality-specific contributions are (unique contributions). Our secondary goal is to improve robustness to missing modalities by encouraging the model to learn redundant contributions. To achieve our two goals, we propose SMURF (Statistical Modality Uniqueness and Redundancy Factorization), a late fusion method that factorizes its outputs into a) unique contributions that are uncorrelated with all other modalities and b) pairwise redundant contributions that are maximally correlated between two modalities. For our primary goal, we 1) verify SMURF's factorization on a synthetic dataset, 2) ensure that its factorization does not degrade predictive performance on eight affective datasets, and 3) observe significant relationships between its factorization and human judgments on three datasets. For our secondary goal, we demonstrate that SMURF leads to more robustness to missing modalities at test time compared to three late fusion baselines.
Torsten Wörtwein, Nicholas B. Allen, Jeffrey F. Cohn, Louis-Philippe Morency
ICMI4
2024 GeSTICS: A Multimodal Corpus for Studying Gesture Synthesis in Two-party Interactions with Contextualized Speech
abstract
Generating natural co-speech gestures and facial expressions for effective human-agent interactions requires modeling the intricate interplay between verbal, non-verbal, and contextual cues observed in dyadic human communication. Two types of contextual cues are of particular interest: (1) individual factors of the interlocutors, such as their demographic attributes, and (2) situational factors, like the outcome of a preceding event. To facilitate their study, we introduce the GeSTICS Dataset, a novel multimodal corpus comprising 9,853 questions and 10,460 answers from audiovisual recordings of post-game sports interviews by 147 interviewees. The dataset contains speech data, including textual transcriptions, lexical descriptors, and acoustic features, as well as visual data encompassing the interviewee’s body pose and facial expressions, with an emphasis on capturing these modalities during both the question-listening and answering phases of the interview. Furthermore, GeSTICS incorporates metadata about individual factors, such as the age and cultural background of the interviewees, and situational factors, like the results of the games, which are often overlooked in existing multimodal datasets. Our preliminary analysis of GeSTICS reveals that the effects of speech features, such as loudness and lexical choice, on the production of co-speech gestures in both speaking and listening phases are moderated by situational factors and the interviewee’s individual factors. GeSTICS is designed to enhance the generation of realistic nonverbal behaviors in virtual agents, animated characters, and human-robot interaction systems, thus contributing to more engaging and effective human-agent communication. The analysis code and the dataset are available at https://gestics.github.io.
Gaoussou Youssouf Kebe, Mehmet Deniz Birlikci, Auriane Boudin, Ryo Ishii, Jeffrey M. Girard, Louis-Philippe Morency
IVA6
2024 HEMM: Holistic Evaluation of Multimodal Foundation Models
abstract
Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study progress in multimodal foundation models, given the range of possible modeling decisions, tasks, and domains. In this paper, we introduce Holistic Evaluation of Multimodal Models (HEMM) to systematically evaluate the capabilities of multimodal foundation models across a set of 3 dimensions: basic skills, information flow, and real-world use cases. Basic multimodal skills are internal abilities required to solve problems, such as learning interactions across modalities, fine-grained alignment, multi-step reasoning, and the ability to handle external knowledge. Information flow studies how multimodal content changes during a task through querying, translation, editing, and fusion. Use cases span domain-specific challenges introduced in real-world multimedia, affective computing, natural sciences, healthcare, and human-computer interaction applications. Through comprehensive experiments across the 30 tasks in HEMM, we (1) identify key dataset dimensions (e.g., basic skills, information flows, and use cases) that pose challenges to today’s models, and (2) distill performance trends regarding how different modeling dimensions (e.g., scale, pre-training data, multimodal alignment, pre-training, and instruction tuning objectives) influence performance. Our conclusions regarding challenging multimodal interactions, use cases, and tasks requiring reasoning and external knowledge, the benefits of data and model scale, and the impacts of instruction-tuning yield actionable insights for future work in multimodal foundation models.
Paul Pu Liang, Akshay Goindani, Talha Chafekar, Leena Mathur, Haofei Yu, Ruslan Salakhutdinov, Louis-Philippe Morency
NeurIPS7
2024 Optimizing Language Models for Human Preferences is a Causal Inference Problem
abstract
As large language models (LLMs) see greater use in academic and commercial settings, there is increasing interest in methods that allow language models to generate texts aligned with human preferences. In this paper, we present an initial exploration of language model optimization for human preferences from *direct outcome datasets*, where each sample consists of a text and an associated numerical outcome measuring the reader’s response. We first propose that language model optimization should be viewed as a *causal problem* to ensure that the model correctly learns the relationship between the text and the outcome. We formalize this causal language optimization problem, and we develop a method{—}*causal preference optimization* (CPO){—}that solves an unbiased surrogate objective for the problem. We further extend CPO with *doubly robust* CPO (DR-CPO), which reduces the variance of the surrogate objective while retaining provably strong guarantees on bias. Finally, we empirically demonstrate the effectiveness of (DR-)CPO in optimizing state-of-the-art LLMs for human preferences on direct outcome data, and we validate the robustness of DR-CPO under difficult confounding conditions.
Victoria Lin 0001, Eli Ben-Michael, Louis-Philippe Morency
UAI3
2023 Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
abstract
Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL (1)5
2023 Understanding Masked Autoencoders via Hierarchical Latent Variable Models
abstract
Masked autoencoder (MAE), a simple and effective self-supervised learning framework based on the reconstruction of masked image regions, has recently achieved prominent success in a variety of vision tasks. Despite the emergence of intriguing empirical observations on MAE, a theoretically principled understanding is still lacking. In this work, we formally characterize and justify existing empirical in-sights and provide theoretical guarantees of MAE. We formulate the underlying data-generating process as a hierarchical latent variable model, and show that under reasonable assumptions, MAE provably identifies a set of latent variables in the hierarchical model, explaining why MAE can extract high-level information from pixels. Further, we show how key hyperparameters in MAE (the masking ratio and the patch size) determine which true latent variables to be recovered, therefore influencing the level of semantic information in the representation. Specifically, extremely large or small masking ratios inevitably lead to low-level representations. Our theory offers coherent explanations of existing empirical observations and provides insights for potential empirical improvements and fundamental limitations of the masked-reconstruction paradigm. We conduct extensive experiments to validate our theoretical insights.
Martin Q. Ma, Guangyi Chen 0002, Eric P. Xing, Yuejie Chi, Louis-Philippe Morency, Kun Zhang 0001
CVPR6
2023 Text-Transport: Toward Learning Causal Effects of Natural Language
abstract
As language technologies gain prominence in real-world settings, it is important to understand how changes to language affect reader perceptions.This can be formalized as the causal effect of varying a linguistic attribute (e.g., sentiment) on a reader's response to the text.In this paper, we introduce TEXT-TRANSPORT, a method for estimation of causal effects from natural language under any text distribution.Current approaches for valid causal effect estimation require strong assumptions about the data, meaning the data from which one can estimate valid causal effects often is not representative of the actual target domain of interest.To address this issue, we leverage the notion of distribution shift to describe an estimator that transports causal effects between domains, bypassing the need for strong assumptions in the target domain.We derive statistical guarantees on the uncertainty of this estimator, and we report empirical results and analyses that support the validity of TEXT-TRANSPORT across data settings.Finally, we use TEXT-TRANSPORT to study a realistic setting-hate speech on social media-in which causal effects do shift significantly between text domains, demonstrating the necessity of transport when conducting causal inference on natural language.
Victoria Lin 0001, Louis-Philippe Morency, Eli Ben-Michael
EMNLP2
2023 Multimodal Feature Selection for Detecting Mothers' Depression in Dyadic Interactions with their Adolescent Offspring
abstract
Depression is the most common psychological disorder, a leading cause of disability world-wide, and a major contributor to inter-generational transmission of psychopathology within families. To contribute to our understanding of depression within families and to inform modality selection and feature reduction, it is critical to identify interpretable features in developmentally appropriate contexts. Mothers with and without depression were studied. Depression was defined as history of treatment for depression and elevations in current or recent symptoms. We explored two multimodal feature selection strategies in dyadic interaction tasks of mothers with their adolescent children for depression detection. Modalities included face and head dynamics, facial action units, speech-related behavior, and verbal features. The initial feature space was vast and inter-correlated (collinear). To reduce dimensionality and gain insight into the relative contribution of each modality and feature, we explored feature selection strategies using Variance Inflation Factor (VIF) and Shapley values. On an average collinearity correction through VIF resulted in about 4 times feature reduction across unimodal and multimodal features. Collinearity correction was also found to be an optimal intermediate step prior to Shapley analysis. Shapley feature selection following VIF yielded best performance. The top 15 features obtained through Shapley achieved 78% accuracy. The most informative features came from all four modalities sampled, which supports the importance of multimodal feature selection.
Maneesh Bilalpur, Saurabh Hinduja, Laura A. Cariola, Lisa Sheeber, Nick Alien, László A. Jeni, Louis-Philippe Morency, Jeffrey F. Cohn
FG7
2023 Face-to-Face Contrastive Learning for Social Intelligence Question-Answering
abstract
Creating artificial social intelligence – algorithms that can understand the nuances of multi-person interactions – is an exciting and emerging challenge in processing facial expressions and gestures from multimodal videos. Recent multimodal methods have set the state of the art on many tasks, but have difficulty modeling the complex face-to-face conversational dynamics across speaking turns in social interaction, particularly in a self-supervised setup. In this paper, we propose Face-to-Face Contrastive Learning (F2F-CL), a graph neural network designed to model social interactions using factorization nodes to contextualize the multimodal face-to-face interaction along the boundaries of the speaking turn. With the F2F-CL model, we propose to perform contrastive learning between the factorization nodes of different speaking turns within the same video. We experimentally evaluate our method on the challenging Social-IQ dataset and show state-of-the-art results.
Alex Wilf, Martin Q. Ma, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
FG5
2023 Continual Learning for Personalized Co-Speech Gesture Generation
abstract
Co-speech gestures are a key channel of human communication, making them important for personalized chat agents to generate. In the past, gesture generation models assumed that data for each speaker is available all at once, and in large amounts. However in practical scenarios, speaker data comes sequentially and in small amounts as the agent personalizes with more speakers, akin to a continual learning paradigm. While more recent works have shown progress in adapting to low-resource data, they catastrophically forget the gesture styles of initial speakers they were trained on. Also, prior generative continual learning works are not multimodal, making this space less studied. In this paper, we explore this new paradigm and propose C-DiffGAN: an approach that continually learns new speaker gesture styles with only a few minutes of per-speaker data, while retaining previously learnt styles. Inspired by prior continual learning works, C-DiffGAN encourages knowledge retention by 1) generating reminiscences of previous low-resource speaker data, then 2) crossmodally aligning to them to mitigate catastrophic forgetting. We quantitatively demonstrate improved performance and reduced forgetting over strong baselines through standard continual learning measures, reinforced by a qualitative user study that shows that our method produces more natural, style-preserving gestures. Code and videos can be found at https://chahuja.com/cdiffgan
Chaitanya Ahuja, Pratik Joshi, Ryo Ishii, Louis-Philippe Morency
ICCV4
2023 Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational Videos
abstract
Many educational videos use slide presentations, a sequence of visual pages that contain text and figures accompanied by spoken language, which are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and psychology attribute the effectiveness of lecture presentations to their multimodal nature. As a step toward developing vision-language models to aid in student learning as intelligent teacher assistants, we introduce the Lecture Presentations Multimodal (LPM) Dataset as a large-scale benchmark testing the capabilities of vision-and-language models in multimodal understanding of educational videos. Our dataset contains aligned slides and spoken language, for 180+ hours of video and 9000+ slides, with 10 lecturers from various subjects (e.g., computer science, dentistry, biology). We introduce three research tasks, (1) figure-to-text retrieval, (2) text-to-figure retrieval, and (3) generation of slide explanations, which are grounded in multimedia learning and psychology principles to test a vision-language model’s understanding of multimodal content. We provide manual annotations to help implement these tasks and establish baselines on them. Comparing baselines and human student performances, we find that state-of-the-art vision-language models (zero-shot and fine-tuned) struggle in (1) weak crossmodal alignment between slides and spoken text, (2) learning novel visual mediums, (3) technical language, and (4) long-range sequences. We introduce PolyViLT, a novel multimodal transformer trained with a multi-instance learning loss that is more effective than current approaches for retrieval. We conclude by shedding light on the challenges and opportunities in multimodal understanding of educational presentation videos.
Dong Won Lee 0007, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu, Louis-Philippe Morency
ICCV5
2023 MultiViz: Towards Visualizing and Understanding Multimodal Models
Paul Pu Liang, Yiwei Lyu 0001, Gunjan Chhablani, Nihal Jain, Xingbo Wang 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR7
2023 Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework
abstract
The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interactions that are necessary to solve a multimodal task? Subsequently, what are the most suitable multimodal models to capture these interactions? To answer these questions, we propose an information-theoretic approach to quantify the degree of redundancy, uniqueness, and synergy relating input modalities with an output task. We term these three measures as the PID statistics of a multimodal distribution (or PID for short), and introduce two new estimators for these PID statistics that scale to high-dimensional distributions. To validate PID estimation, we conduct extensive experiments on both synthetic datasets where the PID is known and on large-scale multimodal benchmarks where PID estimations are compared with human annotations. Finally, we demonstrate their usefulness in (1) quantifying interactions within multimodal datasets, (2) quantifying interactions captured by multimodal models, (3) principled approaches for model selection, and (4) three real-world case studies engaging with domain experts in pathology, mood prediction, and robotic perception where our framework helps to recommend strong multimodal models for each application.
Paul Pu Liang, Chun Kai Ling, Suzanne Nie, Richard J. Chen, Nicholas B. Allen, Randy Auerbach, Faisal Mahmood 0001, Ruslan Salakhutdinov, Louis-Philippe Morency
NeurIPS12
2023 Factorized Contrastive Learning: Going Beyond Multi-view Redundancy
abstract
In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FactorCL, a new multimodal representation learning method to go beyond multi-view redundancy. FactorCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FactorCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.
Paul Pu Liang, Martin Q. Ma, James Zou 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
NeurIPS5
2023 MultiZoo and MultiBench: A Standardized Toolkit for Multimodal Deep Learning
abstract
Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robustness, we release MultiZoo, a public toolkit consisting of standardized implementations of >20 core multimodal algorithms and MultiBench, a large-scale benchmark spanning 15 datasets, 10 modalities, 20 prediction tasks, and 6 research areas. Together, these provide an automated end-to-end machine learning pipeline that simplifies and standardizes data loading, experimental setup, and model evaluation. To enable holistic evaluation, we offer a comprehensive methodology to assess (1) generalization, (2) time and space complexity, and (3) modality robustness. MultiBench paves the way towards a better understanding of the capabilities and limitations of multimodal models, while ensuring ease of use, accessibility, and reproducibility. Our toolkits are publicly available, will be regularly updated, and welcome inputs from the community.
Paul Pu Liang, Yiwei Lyu 0001, Arav Agarwal, Louis-Philippe Morency, Ruslan Salakhutdinov
J. Mach. Learn. Res.6
2022 Language Use in Mother-Adolescent Dyadic Interaction: Preliminary Results
abstract
This preliminary study applied a computer-assisted quantitative linguistic analysis to examine the effectiveness of language-based classification models to discriminate between mothers (n = 140) with and without history of treatment for depression (51% and 49%, respectively). Mothers were recorded during a problem-solving interaction with their adolescent child. Transcripts were manually annotated and analyzed using a dictionary-based, natural-language program approach (Linguistic Inquiry and Word Count). To assess the importance of linguistic features to correctly classify history of depression, we used Support Vector Machines (SVM) with interpretable features. Using linguistic features identified in the empirical literature, an initial SVM achieved nearly 63% accuracy. A second SVM using only the top 5 highest ranked SHAP features improved accuracy to 67.15%. The findings extend the existing literature base on understanding language behavior of depressed mood states, with a focus on the linguistic style of mothers with and without a history of treatment for depression and its potential impact on child development and trans-generational transmission of depression.
Laura A. Cariola, Saurabh Hinduja, Maneesh Bilalpur, Lisa Sheeber, Nicholas B. Allen, Louis-Philippe Morency, Jeffrey F. Cohn
ACII6
2022 HOLM: Hallucinating Objects with Language Models for Referring Expression Recognition in Partially-Observed Scenes
abstract
AI systems embodied in the physical world face a fundamental challenge of partial observability; operating with only a limited view and knowledge of the environment.This creates challenges when AI systems try to reason about language and its relationship with the environment: objects referred to through language (e.g.giving many instructions) are not immediately visible.Actions by the AI system may be required to bring these objects in view.A good benchmark to study this challenge is Dynamic Referring Expression Recognition (dRER) task where the goal is to find a target location by dynamically adjusting the field of view (FoV) in a partially observed 360 • scenes.In this paper, we introduce HOLM, Hallucinating Objects with Language Models, to address the challenge of partial observability.HOLM uses large pre-trained language models (LMs) to infer object hallucinations for the unobserved part of the environment.Our core intuition is that if a pair of objects coappear in an environment frequently, our usage of language should reflect this fact about the world.Based on this intuition, we prompt language models to extract knowledge about object affinities which gives us a proxy for spatial relationships of objects.Our experiments show that HOLM performs better than the state-of-the-art approaches on two datasets for dRER; allowing to study generalization for both indoor and outdoor settings.
Volkan Cirik, Louis-Philippe Morency, Taylor Berg-Kirkpatrick
ACL (1)2
2022 DIME: Fine-grained Interpretations of Multimodal Models via Disentangled Local Explanations
abstract
The ability for a human to understand an Artificial Intelligence (AI) model's decision-making process is critical in enabling stakeholders to visualize model behavior, perform model debugging, promote trust in AI models, and assist in collaborative human-AI decision-making. As a result, the research fields of interpretable and explainable AI have gained traction within AI communities as well as interdisciplinary scientists seeking to apply AI in their subject areas. In this paper, we focus on advancing the state-of-the-art in interpreting multimodal models - a class of machine learning methods that tackle core challenges in representing and capturing interactions between heterogeneous data sources such as images, text, audio, and time-series data. Multimodal models have proliferated numerous real-world applications across healthcare, robotics, multimedia, affective computing, and human-computer interaction. By performing model disentanglement into unimodal contributions (UC) and multimodal interactions (MI), our proposed approach, DIME, enables accurate and fine-grained analysis of multimodal models while maintaining generality across arbitrary modalities, model architectures, and tasks. Through a comprehensive suite of experiments on both synthetic and real-world multimodal tasks, we show that DIME generates accurate disentangled explanations, helps users of multimodal models gain a deeper understanding of model behavior, and presents a step towards debugging and improving these models for real-world deployment.
Yiwei Lyu 0001, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
AIES5
2022 Low-Resource Adaptation for Personalized Co-Speech Gesture Generation
abstract
Personalizing an avatar for co-speech gesture generation from spoken language requires learning the idiosyncrasies of a person's gesture style from a small amount of data. Previous methods in gesture generation require large amounts of data for each speaker, which is often infeasible. We propose an approach, named DiffGAN, that efficiently personalizes co-speech gesture generation models of a high-resource source speaker to target speaker with just 2 minutes of target training data. A unique characteristic of DiffGAN is its ability to account for the crossmodal grounding shift, while also addressing the distribution shift in the output domain. We substantiate the effectiveness of our approach a large scale publicly available dataset through quantitative, qualitative and user studies, which show that our proposed methodology significantly outperforms prior approaches for low-resource adaptation of gesture generation. Code and videos can be found at https://chahuja.com/diffgan.
Chaitanya Ahuja, Dong Won Lee 0007, Louis-Philippe Morency
CVPR3
2022 PACS: A Dataset for Physical Audiovisual CommonSense Reasoning
Samuel Yu, Peter Wu, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
ECCV (37)5
2022 Learning Weakly-supervised Contrastive Representations
Yao-Hung Tsai, Tianqin Li, Peiyuan Liao, Ruslan Salakhutdinov, Louis-Philippe Morency
ICLR6
2022 Conditional Contrastive Learning with Kernel
Yao-Hung Tsai, Tianqin Li, Martin Q. Ma, Han Zhao 0002, Kun Zhang 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR6
2022 What is Multimodal?
abstract
Our experience of the world is multimodal: we see objects, hear sounds, feel texture, smell odors, and taste flavors. In recent years, a broad and impactful body of research emerged in artificial intelligence under the umbrella of multimodal, characterized by multiple modalities. As we formalize a long-term research vision for multimodal research, it is important to reflect on its foundational principles and core technical challenges. What is multimodal? Answering this question is complicated by the multi-disciplinary nature of the problem, spread across many domains and research fields. This talk is based on a recent review of 700+ research papers, to study computational and theoretical foundations for multimodal research, with a focus on multimodal machine learning. Two key principles have driven many multimodal innovations: heterogeneity and interconnections from multiple modalities. Historical and recent progress will be synthesized in a research-oriented taxonomy, centered around 6 core technical challenges: representation, alignment, reasoning, generation, transference, and quantification. The talk will conclude with open questions and unsolved challenges essential for a long-term research vision in multimodal research.
Louis-Philippe Morency
ICMI1
2022 Toward Causal Understanding of Therapist-Client Relationships: A Study of Language Modality and Social Entrainment
abstract
The relationship between a therapist and their client is one of the most critical determinants of successful therapy. The working alliance is a multifaceted concept capturing the collaborative aspect of the therapist-client relationship; a strong working alliance has been extensively linked to many positive therapeutic outcomes. Although therapy sessions are decidedly multimodal interactions, the language modality is of particular interest given its recognized relationship to similar dyadic concepts such as rapport, cooperation, and affiliation. Specifically, in this work we study language entrainment, which measures how much the therapist and client adapt toward each other’s use of language over time. Despite the growing body of work in this area, however, relatively few studies examine causal relationships between human behavior and these relationship metrics: does an individual’s perception of their partner affect how they speak, or does how they speak affect their perception? We explore these questions in this work through the use of structural equation modeling (SEM) techniques, which allow for both multilevel and temporal modeling of the relationship between the quality of the therapist-client working alliance and the participants’ language entrainment. In our first experiment, we demonstrate that these techniques perform well in comparison to other common machine learning models, with the added benefits of interpretability and causal analysis. In our second analysis, we interpret the learned models to examine the relationship between working alliance and language entrainment and address our exploratory research questions. The results reveal that a therapist’s language entrainment can have a significant impact on the client’s perception of the working alliance, and that the client’s language entrainment is a strong indicator of their perception of the working alliance. We discuss the implications of these results and consider several directions for future work in multimodality.
Alexandria K. Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay Fournier, Holly Swartz, Louis-Philippe Morency
ICMI7
2022 Paraphrasing Is All You Need for Novel Object Captioning
abstract
Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequence-to-sequence training or CIDEr optimization. As a result, we present Paraphrasing-to-Captioning (P2C), a two-stage learning framework for NOC, which would heuristically optimize the output captions via paraphrasing. With P2C, the captioning model first learns paraphrasing from a language model pre-trained on text-only corpus, allowing expansion of the word bank for improving linguistic fluency. To further enforce the output caption sufficiently describing the visual content of the input image, we perform self-paraphrasing for the captioning model with fidelity and adequacy objectives introduced. Since no ground truth captions are available for novel object images during training, our P2C leverages cross-modality (image-text) association modules to ensure the above caption characteristics can be properly preserved. In the experiments, we not only show that our P2C achieves state-of-the-art performances on nocaps and COCO Caption datasets, we also verify the effectiveness and flexibility of our learning framework by replacing language and cross-modality association models for NOC. Implementation details and code are available in the supplementary materials.
Cheng-Fu Yang, Yao-Hung Tsai, Wan-Cyuan Fan, Ruslan Salakhutdinov, Louis-Philippe Morency, Frank Wang
NeurIPS5
2021 Humor Knowledge Enriched Transformer for Understanding Multimodal Humor
abstract
Recognizing humor from a video utterance requires understanding the verbal and non-verbal components as well as incorporating the appropriate context and external knowledge. In this paper, we propose Humor Knowledge enriched Transformer (HKT) that can capture the gist of a multimodal humorous expression by integrating the preceding context and external knowledge. We incorporate humor centric external knowledge into the model by capturing the ambiguity and sentiment present in the language. We encode all the language, acoustic, vision, and humor centric features separately using Transformer based encoders, followed by a cross attention layer to exchange information among them. Our model achieves 77.36% and 79.41% accuracy in humorous punchline detection on UR-FUNNY and MUStaRD datasets -- achieving a new state-of-the-art on both datasets with the margin of 4.93% and 2.94% respectively. Furthermore, we demonstrate that our model can capture interpretable, humor-inducing patterns from all modalities.
Md. Kamrul Hasan 0003, Sangwu Lee, Wasifur Rahman, Amir Zadeh 0001, Rada Mihalcea, Louis-Philippe Morency, Mohammed E. Hoque 0001
AAAI6
2021 Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile Data
abstract
Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nick Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nicholas B. Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL/IJCNLP (1)10
2021 Goals, Tasks, and Bonds: Toward the Computational Assessment of Therapist Versus Client Perception of Working Alliance
abstract
Early client dropout is one of the most significant challenges facing psychotherapy: recent studies suggest that at least one in five clients will leave treatment prematurely. Clients may terminate therapy for various reasons, but one of the most common causes is the lack of a strong working alliance. The concept of working alliance captures the collaborative relationship between a client and their therapist when working toward the progress and recovery of the client seeking treatment. Unfortunately, clients are often unwilling to directly express dissatisfaction in care until they have already decided to terminate therapy. On the other side, therapists may miss subtle signs of client discontent during treatment before it is too late. In this work, we demonstrate that nonverbal behavior analysis may aid in bridging this gap. The present study focuses primarily on the head gestures of both the client and therapist, contextualized within conversational turn-taking actions between the pair during psychotherapy sessions. We identify multiple behavior patterns suggestive of an individual's perspective on the working alliance; interestingly, these patterns also differ between the client and the therapist. These patterns inform the development of predictive models for self-reported ratings of working alliance, which demonstrate significant predictive power for both client and therapist ratings. Future applications of such models may stimulate preemptive intervention to strengthen a weak working alliance, whether explicitly attempting to repair the existing alliance or establishing a more suitable client-therapist pairing, to ensure that clients encounter fewer barriers to receiving the treatment they need.
Alexandria K. Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay Fournier, Holly Swartz, Louis-Philippe Morency
FG7
2021 Self-supervised Learning from a Multi-view Perspective
Yao-Hung Tsai, Yue Wu 0001, Ruslan Salakhutdinov, Louis-Philippe Morency
ICLR4
2021 Self-supervised Representation Learning with Relative Predictive Coding
Yao-Hung Tsai, Martin Q. Ma, Muqiao Yang, Han Zhao 0002, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR5
2021 M2H2: A Multimodal Multiparty Hindi Dataset For Humor Recognition in Conversations
abstract
Humor recognition in conversations is a challenging task that has recently gained popularity due to its importance in dialogue understanding, including in multimodal settings (i.e., text, acoustics, and visual). The few existing datasets for humor are mostly in English. However, due to the tremendous growth in multilingual content, there is a great demand to build models and systems that support multilingual information access. To this end, we propose a dataset for Multimodal Multiparty Hindi Humor (M2H2) recognition in conversations containing 6,191 utterances from 13 episodes of a very popular TV series ”Shrimaan Shrimati Phir Se”. Each utterance is annotated with humor/non-humor labels and encompasses acoustic, visual, and textual modalities. We propose several strong multimodal baselines and show the importance of contextual and multimodal information for humor recognition in conversations. The empirical results on M2H2 dataset demonstrate that multimodal information complements unimodal information for humor recognition. The dataset and the baselines are available at http://www.iitp.ac.in/~ai-nlp-ml/resources.html and https://github.com/declare-lab/M2H2-dataset.
Dushyant Singh Chauhan, Gopendra Vikram Singh, Navonil Majumder, Amir Zadeh 0001, Asif Ekbal, Pushpak Bhattacharyya, Louis-Philippe Morency, Soujanya Poria
ICMI7
2021 Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis aims to extract and integrate semantic information collected from multiple modalities to recognize the expressed emotions and sentiment in multimodal data. This research area’s major concern lies in developing an extraordinary fusion scheme that can extract and integrate key information from various modalities. However, previous work is restricted by the lack of leveraging dynamics of independence and correlation between modalities to reach top performance. To mitigate this, we propose the Bi-Bimodal Fusion Network (BBFN), a novel end-to-end network that performs fusion (relevance increment) and separation (difference increment) on pairwise modality representations. The two parts are trained simultaneously such that the combat between them is simulated. The model takes two bimodal pairs as input due to the known information imbalance among modalities. In addition, we leverage a gated control mechanism in the Transformer architecture to further improve the final output. Experimental results on three datasets (CMU-MOSI, CMU-MOSEI, and UR-FUNNY) verifies that our model significantly outperforms the SOTA. The implementation of this work is available at https://github.com/declare-lab/multimodal-deep-learning and https://github.com/declare-lab/BBFN.
Wei Han 0002, Hui Chen 0023, Alexander F. Gelbukh, Amir Zadeh 0001, Louis-Philippe Morency, Soujanya Poria
ICMI5
2021 Human-Guided Modality Informativeness for Affective States
abstract
This paper studies the hypothesis that not all modalities are always needed to predict affective states. We explore this hypothesis in the context of recognizing three affective states that have shown a relation to a future onset of depression: positive, aggressive, and dysphoric. In particular, we investigate three important modalities for face-to-face conversations: vision, language, and acoustic modality. We first perform a human study to better understand which subset of modalities people find informative, when recognizing three affective states. As a second contribution, we explore how these human annotations can guide automatic affect recognition systems to be more interpretable while not degrading their predictive performance. Our studies show that humans can reliably annotate modality informativeness. Further, we observe that guided models significantly improve interpretability, i.e., they attend to modalities similarly to how humans rate the modality informativeness, while at the same time showing a slight increase in predictive performance.
Torsten Wörtwein, Lisa Sheeber, Nicholas B. Allen, Jeffrey F. Cohn, Louis-Philippe Morency
ICMI5
2021 Towards Understanding and Mitigating Social Biases in Language Models
abstract
As machine learning methods are deployed in real-world settings such as healthcare, legal systems, and social science, it is crucial to recognize how they shape social biases and stereotypes in these sensitive decision-making processes. Among such real-world deployments are large-scale pretrained language models (LMs) that can be potentially dangerous in manifesting undesirable representational biases - harmful biases resulting from stereotyping that propagate negative generalizations involving gender, race, religion, and other social constructs. As a step towards improving the fairness of LMs, we carefully define several sources of representational biases before proposing new benchmarks and metrics to measure them. With these tools, we propose steps towards mitigating social biases during text generation. Our empirical results and human evaluation demonstrate effectiveness in mitigating bias while retaining crucial contextual information for high-fidelity text generation, thereby pushing forward the performance-fairness Pareto frontier.
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan Salakhutdinov
ICML3
2021 Multimodal and Multitask Approach to Listener's Backchannel Prediction: Can Prediction of Turn-changing and Turn-management Willingness Improve Backchannel Modeling?
abstract
The listener's backchannel has the important function of encouraging a current speaker to hold their turn and continue to speak, which enables smooth conversation. The listener monitors the speaker's turn-management (a.k.a. speaking and listening) willingness and his/her own willingness to display backchannel behavior. Many studies have focused on predicting the appropriate timing of the backchannel so that conversational agents can display backchannel behavior in response to a user who is speaking. To the best of our knowledge, none of them added the prediction of turn-changing and participants' turn-management willingness to the backchannel prediction model in dyad interactions. In this paper, we proposed a novel backchannel prediction model that can jointly predict turn-changing and turn-management willingness. We investigated the impact of modeling turn-changing and willingness to improve backchannel prediction. Our proposed model is based on trimodal inputs, that is, acoustic, linguistic, and visual cues from conversations. Our results suggest that adding turn-management willingness as a prediction task improves the performance of backchannel prediction within the multi-modal multi-task learning approach, while adding turn-changing prediction is not useful for improving the performance of backchannel prediction.
Ryo Ishii, Xutong Ren, Michal Muszynski, Louis-Philippe Morency
IVA4
2021 Social Signals and Multimedia: Past, Present, Future
abstract
The rising popularity of Artificial Intelligence (AI) has brought considerable public interest as well faster and more direct transfer of research ideas into practice. One of the aspects of AI that still trails behind considerably is the role of machines in interpreting, enhancing, modeling, generating, and influencing social behavior. Such behavior is captured as social signals, usually by sensors recording multiple modalities, making it classic multimedia data. Such behavior can also be generated by an AI system when interacting with humans. Using AI techniques in combination with multimedia data can be used to pursue multiple goals, two of which are high-lighted here. First, supporting people during social interactions and helping them to fulfil their social needs either actively or passively.Second, improving our understanding of how people collaborate, build relationships, and process self identity. Despite the rise of fields such as Social Signal Processing, a similar panel organised at ACM Multimedia 2014, and an area on social and emotional signal sat the ACM MM since 2014, we argue that we have yet to truly fulfil the potential of the combining social signals and multimedia. This panel asks where we have come far enough and what remaining challenges there are in light of recent global events.
Hayley Hung, Cathal Gurrin, Martha A. Larson, Hatice Gunes, Fabien Ringeval, Elisabeth André, Louis-Philippe Morency
ACM Multimedia7
2021 Cross-Modal Generalization: Learning in Low Resource Modalities via Meta-Alignment
abstract
How can we generalize to a new prediction task at test time when it also uses a new modality as input? More importantly, how can we do this with as little annotated data as possible? This problem of cross-modal generalization is a new research milestone with concrete impact on real-world applications. For example, can an AI system start understanding spoken language from mostly written text? Or can it learn the visual steps of a new recipe from only text descriptions? In this work, we formalize cross-modal generalization as a learning paradigm to train a model that can (1) quickly perform new tasks (from new domains) while (2) being originally trained on a different input modality. Such a learning paradigm is crucial for generalization to low-resource modalities such as spoken speech in rare languages while utilizing a different high-resource modality such as text. One key technical challenge that makes it different from other learning paradigms such as meta-learning and domain adaptation is the presence of different source and target modalities which will require different encoders. We propose an effective solution based on meta-alignment, a novel method to align representation spaces using strongly and weakly paired cross-modal data while ensuring quick generalization to new tasks across different modalities. This approach uses key ideas from cross-modal learning and meta-learning, and presents strong results on the cross-modal generalization problem. We benchmark several approaches on 3 real-world classification tasks: few-shot recipe classification from text to images of recipes, object classification from images to audio of objects, and language classification from text to spoken speech across 100 languages spanning many rare languages. Our results demonstrate strong performance even when the new target modality has only a few (1-10) labeled samples and in the presence of noisy labels, a scenario particularly prevalent in low-resource modalities.
Paul Pu Liang, Peter Wu, Liu Ziyin 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ACM Multimedia4
2021 StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer
abstract
Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yiwei Lyu 0001, Paul Pu Liang, Hai Pham, Eduard H. Hovy, Barnabás Póczos, Ruslan Salakhutdinov, Louis-Philippe Morency
NAACL-HLT7
2021 MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences
abstract
Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, Louis-Philippe Morency. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ruitao Yi, Yuying Zhu 0004, Azaan Rehman, Amir Zadeh 0001, Soujanya Poria, Louis-Philippe Morency
NAACL-HLT8
2020 Refer360$^\circ$: A Referring Expression Recognition Dataset in 360$^\circ$ Images
abstract
We propose a novel large-scale referring expression recognition dataset, Refer360°, consisting of 17,137 instruction sequences and ground-truth actions for completing these instructions in 360° scenes.Refer360° differs from existing related datasets in three ways.First, we propose a more realistic scenario where instructors and the followers have partial, yet dynamic, views of the scene -followers continuously modify their field-of-view (FoV) while interpreting instructions that specify a final target location.Second, instructions to find the target location consist of multiple steps for followers who will start at random FoVs.As a result, intermediate instructions are strongly grounded in object references and followers must identify intermediate FoVs to find the final target location correctly.Third, the target locations are neither restricted to predefined objects nor chosen by annotators; instead, they are distributed randomly across scenes.This "point anywhere" approach leads to more linguistically complex instructions, as shown in our analyses.Our examination of the dataset shows that Refer360° manifests linguistically rich phenomena in a language grounding task that poses novel challenges for computational modeling of language, vision, and navigation.
Volkan Cirik, Taylor Berg-Kirkpatrick, Louis-Philippe Morency
ACL3
2020 Language to Network: Conditional Parameter Adaptation with Natural Language Descriptions
abstract
Transfer learning using ImageNet pre-trained models has been the de facto approach in a wide range of computer vision tasks. However, fine-tuning still requires task-specific training data. In this paper, we propose N3 (Neural Networks from Natural Language) - a new paradigm of synthesizing task-specific neural networks from language descriptions and a generic pre-trained model. N3 leverages language descriptions to generate parameter adaptations as well as a new task-specific classification layer for a pre-trained neural network, effectively "fine-tuning" the network for a new task using only language descriptions as input. To the best of our knowledge, N3 is the first method to synthesize entire neural networks from natural language. Experimental results show that N3 can out-perform previous natural-language based zero-shot learning methods across 4 different zero-shot image classification benchmarks. We also demonstrate a simple method to help identify keywords in language descriptions leveraged by N3 when synthesizing model parameters.
Zhun Liu, Shengjia Yan, Alexandre E. Eichenberger, Louis-Philippe Morency
ACL5
2020 Towards Debiasing Sentence Representations
abstract
As natural language processing methods are increasingly deployed in real-world scenarios such as healthcare, legal systems, and social science, it becomes necessary to recognize the role they potentially play in shaping social biases and stereotypes.Previous work has revealed the presence of social biases in widely used word embeddings involving gender, race, religion, and other social constructs.While some methods were proposed to debias these word-level embeddings, there is a need to perform debiasing at the sentence-level given the recent shift towards new contextualized sentence representations such as ELMo and BERT.In this paper, we investigate the presence of social biases in sentence-level representations and propose a new method, SENT-DEBIAS, to reduce these biases.We show that SENT-DEBIAS is effective in removing biases, and at the same time, preserves performance on sentence-level downstream tasks such as sentiment analysis, linguistic acceptability, and natural language understanding.We hope that our work will inspire future research on characterizing and removing social biases from widely adopted sentence representations for fairer NLP.
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL6
2020 Integrating Multimodal Information in Large Pretrained Transformers
abstract
Recent Transformer-based contextual word representations, including BERT and XLNet, have shown state-of-the-art performance in multiple disciplines within NLP. Fine-tuning the trained contextual models on task-specific datasets has been the key to achieving superior performance downstream. While fine-tuning these pre-trained models is straight-forward for lexical applications (applications with only language modality), it is not trivial for multimodal language (a growing area in NLP focused on modeling face-to-face communication). Pre-trained models don't have the necessary components to accept two extra modalities of vision and acoustic. In this paper, we proposed an attachment to BERT and XLNet called Multimodal Adaptation Gate (MAG). MAG allows BERT and XLNet to accept multimodal nonverbal data during fine-tuning. It does so by generating a shift to internal representation of BERT and XLNet; a shift that is conditioned on the visual and acoustic modalities. In our experiments, we study the commonly used CMU-MOSI and CMU-MOSEI datasets for multimodal sentiment analysis. Fine-tuning MAG-BERT and MAG-XLNet significantly boosts the sentiment analysis performance over previous baselines as well as language-only fine-tuning of BERT and XLNet. On the CMU-MOSI dataset, MAG-XLNet achieves human-level multimodal sentiment analysis performance for the first time in the NLP community.
Wasifur Rahman, Md. Kamrul Hasan 0003, Sangwu Lee, Amir Zadeh 0001, Chengfeng Mao, Louis-Philippe Morency, Mohammed E. Hoque 0001
ACL6
2020 Style Transfer for Co-speech Gesture Animation: A Multi-speaker Conditional-Mixture Approach
Chaitanya Ahuja, Dong Won Lee 0007, Yukiko I. Nakano, Louis-Philippe Morency
ECCV (18)4
2020 Diverse and Admissible Trajectory Forecasting Through Multimodal Context Understanding
Seong Hyeon Park, Gyubok Lee, Jimin Seo, Manoj Bhat, Jonathan Francis, Ashwin R. Jadhav, Paul Pu Liang, Louis-Philippe Morency
ECCV (11)9
2020 Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis
abstract
, which dynamically adjusts weights between input modalities and output representations differently for each input sample. Multimodal routing can identify relative importance of both individual modalities and cross-modality features. Moreover, the weight assignment by routing allows us to interpret modality-prediction relationships not only globally (i.e. general trends over the whole dataset), but also locally for each single input sample, mean-while keeping competitive performance compared to state-of-the-art methods.
Yao-Hung Tsai, Martin Q. Ma, Muqiao Yang, Ruslan Salakhutdinov, Louis-Philippe Morency
EMNLP (1)5
2020 CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French
abstract
Modeling multimodal language is a core research area in natural language processing. While languages such as English have relatively large multimodal language resources, other widely spoken languages across the globe have few or no large-scale datasets in this area. This disproportionately affects native speakers of languages other than English. As a step towards building more equitable and inclusive multimodal systems, we introduce the first large-scale multimodal language dataset for Spanish, Portuguese, German and French. The proposed dataset, called CMU-MOSEAS (CMU Multimodal Opinion Sentiment, Emotions and Attributes), is the largest of its kind with 40, 000 total labelled sentences. It covers a diverse set topics and speakers, and carries supervision of 20 labels including sentiment (and subjectivity), emotions, and attributes. Our evaluations on a state-of-the-art multimodal model demonstrates that CMU-MOSEAS enables further research for multilingual studies in multimodal language.
Amir Zadeh 0001, Yansheng Cao, Smon Hessner, Paul Pu Liang, Soujanya Poria, Louis-Philippe Morency
EMNLP (1)6
2020 Simple and Effective Approaches for Uncertainty Prediction in Facial Action Unit Intensity Regression
abstract
Knowing how much to trust a prediction is important for many critical applications. We describe two simple approaches to estimate uncertainty in regression prediction tasks and compare their performance and complexity against popular approaches. We operationalize uncertainty in regression as the absolute error between a model's prediction and the ground truth. Our two proposed approaches use a secondary model to predict the uncertainty of a primary predictive model. Our first approach leverages the assumption that similar observations are likely to have similar uncertainty and predicts uncertainty with a non-parametric method. Our second approach trains a secondary model to directly predict the uncertainty of the primary predictive model. Both approaches outperform other established uncertainty estimation approaches on the MNIST, DISFA, and BP4D+ datasets. Furthermore, we observe that approaches that directly predict the uncertainty generally perform better than approaches that indirectly estimate uncertainty.
Torsten Wörtwein, Louis-Philippe Morency
FG2
2020 Toward Multimodal Modeling of Emotional Expressiveness
abstract
Emotional expressiveness captures the extent to which a person tends to outwardly display their emotions through behavior. Due to the close relationship between emotional expressiveness and behavioral health, as well as the crucial role that it plays in social interaction, the ability to automatically predict emotional expressiveness stands to spur advances in science, medicine, and industry. In this paper, we explore three related research questions. First, how well can emotional expressiveness be predicted from visual, linguistic, and multimodal behavioral signals? Second, how important is each behavioral modality to the prediction of emotional expressiveness? Third, which behavioral signals are reliably related to emotional expressiveness? To answer these questions, we add highly reliable transcripts and human ratings of perceived emotional expressiveness to an existing video database and use this data to train, validate, and test predictive models. Our best model shows promising predictive performance on this dataset (RMSE=0.65, R^2=0.45, r=0.74). Multimodal models tend to perform best overall, and models trained on the linguistic modality tend to outperform models trained on the visual modality. Finally, examination of our interpretable models' coefficients reveals a number of visual and linguistic behavioral signals---such as facial action unit intensity, overall word count, and use of words related to social processes---that reliably predict emotional expressiveness.
Victoria Lin 0001, Jeffrey M. Girard, Michael A. Sayette, Louis-Philippe Morency
ICMI4
2020 Depression Severity Assessment for Adolescents at High Risk of Mental Disorders
abstract
Recent progress in artificial intelligence has led to the development of automatic behavioral marker recognition, such as facial and vocal expressions. Those automatic tools have enormous potential to support mental health assessment, clinical decision making, and treatment planning. In this paper, we investigate nonverbal behavioral markers of depression severity assessed during semi-structured medical interviews of adolescent patients. The main goal of our research is two-fold: studying a unique population of adolescents at high risk of mental disorders and differentiating mild depression from moderate or severe depression. We aim to explore computationally inferred facial and vocal behavioral responses elicited by three segments of the semi-structured medical interviews: Distress Assessment Questions, Ubiquitous Questions, and Concept Questions. Our experimental methodology reflects best practise used for analyzing small sample size and unbalanced datasets of unique patients. Our results show a very interesting trend with strongly discriminative behavioral markers from both acoustic and visual modalities. These promising results are likely due to the unique classification task (mild depression vs. moderate and severe depression) and three types of probing questions.
Michal Muszynski, Jamie Zelazny, Jeffrey M. Girard, Louis-Philippe Morency
ICMI4
2020 Impact of Personality on Nonverbal Behavior Generation
abstract
To realize natural-looking virtual agents, one key technical challenge is to automatically generate nonverbal behaviors from spoken language. Since nonverbal behavior varies depending on personality, it is important to generate these nonverbal behaviors to match the expected personality of a virtual agent. In this work, we study how personality traits relate to the process of generating individual nonverbal behaviors from the whole body, including the head, eye gaze, arms, and posture. To study this, we first created a dialogue corpus including transcripts, a broad range of labelled nonverbal behaviors, and the Big Five personality scores of participants in dyad interactions. We constructed models that can predict each nonverbal behavior label given as an input language representation from the participants' spoken sentences. Our experimental results show that personality can help improve the prediction of nonverbal behaviors.
Ryo Ishii, Chaitanya Ahuja, Yukiko I. Nakano, Louis-Philippe Morency
IVA4
2020 Can Prediction of Turn-management Willingness Improve Turn-changing Modeling?
abstract
For smooth conversation, participants must carefully monitor the turn-management (a.k.a. speaking and listening) willingness of other conversational partners and adjust turn-changing behaviors accordingly. Many studies have focused on predicting the actual moments of speaker changes (a.k.a. turn-changing), but to the best of our knowledge, none of them explicitly modeled the turn-management willingness from both speakers and listeners in dyad interactions. We address the problem of building models for predicting this willingness of both. Our models are based on trimodal inputs, including acoustic, linguistic, and visual cues from conversations. We also study the impact of modeling willingness to help improve the task of turn-changing prediction. We introduce a dyadic conversation corpus with annotated scores of speaker/listener turn-management willingness. Our results show that using all of three modalities of speaker and listener is important for predicting turn-management willingness. Furthermore, explicitly adding willingness as a prediction task improves the performance of turn-changing prediction. Also, turn-management willingness prediction becomes more accurate with this multi-task learning approach.
Ryo Ishii, Xutong Ren, Michal Muszynski, Louis-Philippe Morency
IVA4
2020 Neural Methods for Point-wise Dependency Estimation
abstract
Since its inception, the neural estimation of mutual information (MI) has demonstrated the empirical success of modeling expected dependency between high-dimensional random variables. However, MI is an aggregate statistic and cannot be used to measure point-wise dependency between different events. In this work, instead of estimating the expected dependency, we focus on estimating point-wise dependency (PD), which quantitatively measures how likely two outcomes co-occur. We show that we can naturally obtain PD when we are optimizing MI neural variational bounds. However, optimizing these bounds is challenging due to its large variance in practice. To address this issue, we develop two methods (free of optimizing MI variational bounds): Probabilistic Classifier and Density-Ratio Fitting. We demonstrate the effectiveness of our approaches in 1) MI estimation, 2) self-supervised representation learning, and 3) cross-modal retrieval task.
Yao-Hung Tsai, Han Zhao 0002, Makoto Yamada, Louis-Philippe Morency, Ruslan Salakhutdinov
NeurIPS4
2019 Language2Pose: Natural Language Grounded Pose Forecasting
abstract
Generating animations from natural language sentences finds its applications in a a number of domains such as movie script visualization, virtual human animation and, robot motion planning. These sentences can describe different kinds of actions, speeds and direction of these actions, and possibly a target destination. The core modeling challenge in this language-to-pose application is how to map linguistic concepts to motion animations. In this paper, we address this multimodal problem by introducing a neural architecture called Joint Language-to-Pose (or JL2P), which learns a joint embedding of language and pose. This joint embedding space is learned end-to-end using a curriculum learning approach which emphasizes shorter and easier sequences first before moving to longer and harder ones. We evaluate our proposed model on a publicly available corpus of 3D pose data and human-annotated sentences. Both objective metrics and human judgment evaluation confirm that our proposed approach is able to generate more accurate animations and are deemed visually more representative by humans than other data driven approaches.
Chaitanya Ahuja, Louis-Philippe Morency
3DV2
2019 Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities
abstract
Multimodal sentiment analysis is a core research area that studies speaker sentiment expressed from the language, visual, and acoustic modalities. The central challenge in multimodal learning involves inferring joint representations that can process and relate information from these modalities. However, existing work learns joint representations by requiring all modalities as input and as a result, the learned representations may be sensitive to noisy or missing modalities at test time. With the recent success of sequence to sequence (Seq2Seq) models in machine translation, there is an opportunity to explore new ways of learning joint representations that may not require all input modalities at test time. In this paper, we propose a method to learn robust joint representations by translating between modalities. Our method is based on the key insight that translation from a source to a target modality provides a method of learning joint representations using only the source modality as input. We augment modality translations with a cycle consistency loss to ensure that our joint representations retain maximal information from all modalities. Once our translation model is trained with paired multimodal data, we only need data from the source modality at test time for final sentiment prediction. This ensures that our model remains robust from perturbations or missing information in the other modalities. We train our model with a coupled translationprediction objective and it achieves new state-of-the-art results on multimodal sentiment analysis datasets: CMU-MOSI, ICTMMMO, and YouTube. Additional experiments show that our model learns increasingly discriminative joint representations with more input modalities while maintaining robustness to missing or perturbed modalities.
Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, Barnabás Póczos
AAAI4
2019 Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors
abstract
Humans convey their intentions through the usage of both verbal and nonverbal behaviors during face-to-face communication. Speaker intentions often vary dynamically depending on different nonverbal contexts, such as vocal patterns and facial expressions. As a result, when modeling human language, it is essential to not only consider the literal meaning of the words but also the nonverbal contexts in which these words appear. To better model human language, we first model expressive nonverbal representations by analyzing the fine-grained visual and acoustic patterns that occur during word segments. In addition, we seek to capture the dynamic nature of nonverbal intents by shifting word representations based on the accompanying nonverbal behaviors. To this end, we propose the Recurrent Attended Variation Embedding Network (RAVEN) that models the fine-grained structure of nonverbal subword sequences and dynamically shifts word representations based on nonverbal cues. Our proposed model achieves competitive performance on two publicly available datasets for multimodal sentiment analysis and emotion recognition. We also visualize the shifted word representations in different nonverbal contexts and summarize common patterns regarding multimodal variations of word representations.
Yansen Wang, Ying Shen 0006, Zhun Liu, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
AAAI6
2019 Reconsidering the Duchenne Smile: Indicator of Positive Emotion or Artifact of Smile Intensity?
abstract
The Duchenne smile hypothesis is that smiles that include eye constriction (AU6) are the product of genuine positive emotion, whereas smiles that do not are either falsified or related to negative emotion. This hypothesis has become very influential and is often used in scientific and applied settings to justify the inference that a smile is either true or false. However, empirical support for this hypothesis has been equivocal and some researchers have proposed that, rather than being a reliable indicator of positive emotion, AU6 may just be an artifact produced by intense smiles. Initial support for this proposal has been found when comparing smiles related to genuine and feigned positive emotion; however, it has not yet been examined when comparing smiles related to genuine positive and negative emotion. The current study addressed this gap in the literature by examining spontaneous smiles from 136 participants during the elicitation of amusement, embarrassment, fear, and pain (from the BP4D+ dataset). Bayesian multilevel regression models were used to quantify the associations between AU6 and self-reported amusement while controlling for smile intensity. Models were estimated to infer amusement from AU6 and to explain the intensity of AU6 using amusement. In both cases, controlling for smile intensity substantially reduced the hypothesized association, whereas the effect of smile intensity itself was quite large and reliable. These results provide further evidence that the Duchenne smile is likely an artifact of smile intensity rather than a reliable and unique indicator of genuine positive emotion.
Jeffrey M. Girard, Gayatri Shandar, Zhun Liu, Jeffrey F. Cohn, Lijun Yin 0001, Louis-Philippe Morency
ACII6
2019 Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization
abstract
There has been an increased interest in multimodal language processing including multimodal dialog, question answering, sentiment analysis, and speech recognition.However, naturally occurring multimodal data is often imperfect as a result of imperfect modalities, missing entries or noise corruption.To address these concerns, we present a regularization method based on tensor rank minimization.Our method is based on the observation that high-dimensional multimodal time series data often exhibit correlations across time and modalities which leads to low-rank tensor representations.However, the presence of noise or incomplete values breaks these correlations and results in tensor representations of higher rank.We design a model to learn such tensor representations and effectively regularize their rank.Experiments on multimodal language data show that our model achieves good results across various levels of imperfection.
Paul Pu Liang, Zhun Liu, Yao-Hung Tsai, Qibin Zhao, Ruslan Salakhutdinov, Louis-Philippe Morency
ACL (1)6
2019 Multimodal Transformer for Unaligned Multimodal Language Sequences
abstract
Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data non-alignment due to variable sampling rates for the sequences from each modality; and 2) long-range dependencies between elements across modalities. In this paper, we introduce the Multimodal Transformer (MulT) to generically address the above issues in an end-to-end manner without explicitly aligning the data. At the heart of our model is the directional pairwise cross-modal attention, which attends to interactions between multimodal sequences across distinct time steps and latently adapt streams from one modality to another. Comprehensive experiments on both aligned and non-aligned multimodal time-series show that our model outperforms state-of-the-art methods by a large margin. In addition, empirical analysis suggests that correlated crossmodal signals are able to be captured by the proposed crossmodal attention mechanism in MulT.
Yao-Hung Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov
ACL (1)5
2019 Social-IQ: A Question Answering Benchmark for Artificial Social Intelligence
abstract
As intelligent systems increasingly blend into our everyday life, artificial social intelligence becomes a prominent area of research. Intelligent systems must be socially intelligent in order to comprehend human intents and maintain a rich level of interaction with humans. Human language offers a unique unconstrained approach to probe through questions and reason through answers about social situations. This unconstrained approach extends previous attempts to model social intelligence through numeric supervision (e.g. sentiment and emotions labels). In this paper, we introduce Social-IQ, a unconstrained benchmark specifically designed to train and evaluate socially intelligent technologies. By providing a rich source of open-ended questions and answers, Social-IQ opens the door to explainable social intelligence. The dataset contains rigorously annotated and validated videos, questions and answers, as well as annotations for the complexity level of each question and answer. Social-IQ contains 1,250 natural in-the-wild social situations, 7,500 questions and 52,500 correct and incorrect answers. Although humans can reason about social situations with very high accuracy (95.08%), existing state-of-the-art computational models struggle on this task. As a result, Social-IQ brings novel challenges that will spark future research in social intelligence modeling, visual reasoning, and multimodal question answering (QA).
Amir Zadeh 0001, Paul Pu Liang, Edmund Tong, Louis-Philippe Morency
CVPR5
2019 Video Relationship Reasoning Using Gated Spatio-Temporal Energy Graph
abstract
Visual relationship reasoning is a crucial yet challenging task for understanding rich interactions across visual concepts. For example, a relationship {man, open, door} involves a complex relation {open} between concrete entities {man, door}. While much of the existing work has studied this problem in the context of still images, understanding visual relationships in videos has received limited attention. Due to their temporal nature, videos enable us to model and reason about a more comprehensive set of visual relationships, such as those requiring multiple (temporal) observations (e.g., {man, lift up, box} vs. {man, put down, box}), as well as relationships that are often correlated through time (e.g., {woman, pay, money} followed by {woman, buy, coffee}). In this paper, we construct a Conditional Random Field on a fully-connected spatiotemporal graph that exploits the statistical dependency between relational entities spatially and temporally. We introduce a novel gated energy function parametrization that learns adaptive relations conditioned on visual observations. Our model optimization is computationally efficient, and its space computation complexity is significantly amortized through our proposed parameterization. Experimental results on benchmark video datasets (ImageNet Video and Charades) demonstrate state-of-the-art performance across three standard relationship reasoning tasks: Detection, Tagging, and Recognition.
Yao-Hung Tsai, Santosh Kumar Divvala, Louis-Philippe Morency, Ruslan Salakhutdinov, Ali Farhadi
CVPR3
2019 UR-FUNNY: A Multimodal Language Dataset for Understanding Humor
abstract
Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, Mohammed (Ehsan) Hoque. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Md. Kamrul Hasan 0003, Wasifur Rahman, Amir Zadeh 0001, Jianyuan Zhong, Md. Iftekhar Tanveer, Louis-Philippe Morency, Mohammed E. Hoque 0001
EMNLP/IJCNLP (1)6
2019 Transformer Dissection: An Unified Understanding for Transformer's Attention via the Lens of Kernel
abstract
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, Ruslan Salakhutdinov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yao-Hung Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, Ruslan Salakhutdinov
EMNLP/IJCNLP (1)4
2019 Learning Factorized Multimodal Representations
Yao-Hung Tsai, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ICLR (Poster)4
2019 To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations
abstract
Non verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepresence, in form of an avatar, it is important to model these behaviours, especially in dyadic interactions. Creating such personalized avatars not only requires to model intrapersonal dynamics between a avatar’s speech and their body pose, but it also needs to model interpersonal dynamics with the interlocutor present in the conversation. In this paper, we introduce a neural architecture named Dyadic Residual-Attention Model (DRAM), which integrates intrapersonal (monadic) and interpersonal (dyadic) dynamics using selective attention to generate sequences of body pose conditioned on audio and body pose of the interlocutor and audio of the human operating the avatar. We evaluate our proposed model on dyadic conversational data consisting of pose and audio of both participants, confirming the importance of adaptive attention between monadic and dyadic dynamics when predicting avatar pose. We also conduct a user study to analyze judgments of human observers. Our results confirm that the generated body pose is more natural, models intrapersonal dynamics and interpersonal dynamics better than non-adaptive monadic/dyadic models.
Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency, Yaser Sheikh
ICMI3
2019 ElderReact: A Multimodal Dataset for Recognizing Emotional Response in Aging Adults
abstract
Automatic emotion recognition plays a critical role in technologies such as intelligent agents and social robots and is increasingly being deployed in applied settings such as education and healthcare. Most research to date has focused on recognizing the emotional expressions of young and middle-aged adults and, to a lesser extent, children and adolescents. Very few studies have examined automatic emotion recognition in older adults (i.e., elders), which represent a large and growing population worldwide. Given that aging causes many changes in facial shape and appearance and has been found to alter patterns of nonverbal behavior, there is strong reason to believe that automatic emotion recognition systems may need to be developed specifically (or augmented) for the elder population. To promote and support this type of research, we introduce a newly collected multimodal dataset of elders reacting to emotion elicitation stimuli. Specifically, it contains 1323 video clips of 46 unique individuals with human annotations of six discrete emotions: anger, disgust, fear, happiness, sadness, and surprise as well as valence. We present a detailed analysis of the most indicative features for each emotion. We also establish several baselines using unimodal and multimodal features on this dataset. Finally, we show that models trained on dataset of another age group do not generalize well on elders.
Kaixin Ma, Xinru Yang, Jeffrey M. Girard, Louis-Philippe Morency
ICMI6
2019 Multimodal Behavioral Markers Exploring Suicidal Intent in Social Media Videos
abstract
Suicide is one of the leading causes of death in the modern world. In this digital age, individuals are increasingly using social media to express themselves and often use these platforms to express suicidal intent. Various studies have inspected suicidal intent behavioral markers in controlled environments but it is still unexplored if such markers will generalize to suicidal intent expressed on social media. In this work, we set out to study multimodal behavioral markers related to suicidal intent when expressed on social media videos. We explore verbal, acoustic and visual behavioral markers in the context of identifying individuals at higher risk of suicidal attempt. Our analysis reveals that frequent silences, slouched shoulders, rapid hand movements and profanity are predominant multimodal behavioral markers indicative of suicidal intent1.
Ankit Parag Shah, Vasu Sharma, Vaibhav Vaibhav, Mahmoud Alismail, Louis-Philippe Morency
ICMI5
2019 Bag-of-Acoustic-Words for Mental Health Assessment: A Deep Autoencoding Approach
Wenchao Du, Louis-Philippe Morency, Jeffrey F. Cohn, Alan W. Black
INTERSPEECH2
2019 PANEL: Challenges for Multimedia/Multimodal Research in the Next Decade
abstract
The multimedia and multi-modal community is witnessing an explosive transformation in the recent years with major societal impact. With the unprecedented deployment of multimedia devices and systems, multimedia research is critical to our abilities and prospects in advancing state-of-the-art technologies and solving real-world challenges facing the society and the nation. To respond to these challenges and further advance the frontiers of the field of multimedia, this panel will discuss the challenges and visions that may guide future research in the next ten years.
Shih-Fu Chang, Louis-Philippe Morency, Alex Hauptmann 0001, Alberto Del Bimbo, Cathal Gurrin, Hayley Hung, Heng Ji 0001, Alan F. Smeaton
ACM Multimedia2
2019 Deep Gamblers: Learning to Abstain with Portfolio Theory
abstract
We deal with the selective classification problem (supervised-learning problem with a rejection option), where we want to achieve the best performance at a certain level of coverage of the data. We transform the original $m$-class classification problem to (m+1)-class where the (m+1)-th class represents the model abstaining from making a prediction due to disconfidence. Inspired by portfolio theory, we propose a loss function for the selective classification problem based on the doubling rate of gambling. Minimizing this loss function corresponds naturally to maximizing the return of a horse race, where a player aims to balance between betting on an outcome (making a prediction) when confident and reserving one's winnings (abstaining) when not confident. This loss function allows us to train neural networks and characterize the disconfidence of prediction in an end-to-end fashion. In comparison with previous methods, our method requires almost no modification to the model inference algorithm or model architecture. Experiments show that our method can identify uncertainty in data points, and achieves strong results on SVHN and CIFAR10 at various coverages of the data.
Liu Ziyin 0001, Zhikang Wang, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency, Masahito Ueda
NeurIPS5
2019 Multimodal Machine Learning: A Survey and Taxonomy
abstract
Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is characterized as multimodal when it includes multiple such modalities. In order for Artificial Intelligence to make progress in understanding the world around us, it needs to be able to interpret such multimodal signals together. Multimodal machine learning aims to build models that can process and relate information from multiple modalities. It is a vibrant multi-disciplinary field of increasing importance and with extraordinary potential. Instead of focusing on specific multimodal applications, this paper surveys the recent advances in multimodal machine learning itself and presents them in a common taxonomy. We go beyond the typical early and late fusion categorization and identify broader challenges that are faced by multimodal machine learning, namely: representation, translation, alignment, fusion, and co-learning. This new taxonomy will enable researchers to better understand the state of the field and identify directions for future research.
Tadas Baltrusaitis, Chaitanya Ahuja, Louis-Philippe Morency
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Learning Pose-Aware Models for Pose-Invariant Face Recognition in the Wild
abstract
We propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2].
Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.13
2018 Lattice Recurrent Unit: Improving Convergence and Statistical Efficiency for Sequence Modeling
abstract
Recurrent neural networks have shown remarkable success in modeling sequences. However low resource situations still adversely affect the generalizability of these models. We introduce a new family of models, called Lattice Recurrent Units (LRU), to address the challenge of learning deep multi-layer recurrent models with limited resources. LRU models achieve this goal by creating distinct (but coupled) flow of information inside the units: a first flow along time dimension and a second flow along depth dimension. It also offers a symmetry in how information can flow horizontally and vertically. We analyze the effects of decoupling three different components of our LRU model: Reset Gate, Update Gate and Projected State. We evaluate this family of new LRU models on computational convergence rates and statistical efficiency.Our experiments are performed on four publicly-available datasets, comparing with Grid-LSTM and Recurrent Highway networks. Our results show that LRU has better empirical computational convergence rates and statistical efficiency values, along with learning more accurate language models.
Chaitanya Ahuja, Louis-Philippe Morency
AAAI2
2018 Using Syntax to Ground Referring Expressions in Natural Images
abstract
We introduce GroundNet, a neural network for referring expression recognition---the task of localizing (or grounding) in an image the object referred to by a natural language expression. Our approach to this task is the first to rely on a syntactic analysis of the input referring expression in order to inform the structure of the computation graph. Given a parse tree for an input expression, we explicitly map the syntactic constituents and relationships present in the tree to a composed graph of neural modules that defines our architecture for performing localization. This syntax-based approach aids localization of both the target object and auxiliary supporting objects mentioned in the expression. As a result, GroundNet is more interpretable than previous methods: we can (1) determine which phrase of the referring expression points to which object in the image and (2) track how the localization of the target object is determined by the network. We study this property empirically by introducing a new set of annotations on the GoogleRef dataset to evaluate localization of supporting objects. Our experiments show that GroundNet achieves state-of-the-art accuracy in identifying supporting objects, while maintaining comparable performance in the localization of target objects.
Volkan Cirik, Taylor Berg-Kirkpatrick, Louis-Philippe Morency
AAAI3
2018 Memory Fusion Network for Multi-view Sequential Learning
abstract
Multi-view sequential learning is a fundamental problem in machine learning dealing with multi-view sequences. In a multi-view sequence, there exists two forms of interactions between different views: view-specific interactions and cross-view interactions. In this paper, we present a new neural architecture for multi-view sequential learning called the Memory Fusion Network (MFN) that explicitly accounts for both interactions in a neural architecture and continuously models them through time. The first component of the MFN is called the System of LSTMs, where view-specific interactions are learned in isolation through assigning an LSTM function to each view. The cross-view interactions are then identified using a special attention mechanism called the Delta-memory Attention Network (DMAN) and summarized through time with a Multi-view Gated Memory. Through extensive experimentation, MFN is compared to various proposed approaches for multi-view sequential learning on multiple publicly available benchmark datasets. MFN outperforms all the multi-view approaches. Furthermore, MFN outperforms all current state-of-the-art models, setting new state-of-the-art results for all three multi-view datasets.
Amir Zadeh 0001, Paul Pu Liang, Navonil Majumder, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
AAAI6
2018 Multi-attention Recurrent Network for Human Communication Comprehension
abstract
Human face-to-face communication is a complex multimodal signal. We use words (language modality), gestures (vision modality) and changes in tone (acoustic modality) to convey our intentions. Humans easily process and understand face-to-face communication, however, comprehending this form of communication remains a significant challenge for Artificial Intelligence (AI). AI must understand each modality and the interactions between them that shape the communication. In this paper, we present a novel neural architecture for understanding human communication called the Multi-attention Recurrent Network (MARN). The main strength of our model comes from discovering interactions between modalities through time using a neural component called the Multi-attention Block (MAB) and storing them in the hybrid memory of a recurrent component called the Long-short Term Hybrid Memory (LSTHM). We perform extensive comparisons on six publicly available datasets for multimodal sentiment analysis, speaker trait recognition and emotion recognition. MARN shows state-of-the-art results performance in all the datasets.
Amir Zadeh 0001, Paul Pu Liang, Soujanya Poria, Prateek Vij, Erik Cambria, Louis-Philippe Morency
AAAI6
2018 Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
abstract
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Amir Zadeh 0001, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
ACL (1)5
2018 Efficient Low-rank Multimodal Fusion With Modality-Specific Factors
abstract
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, Louis-Philippe Morency. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Zhun Liu, Ying Shen 0006, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
ACL (1)6
2018 Multimodal Language Analysis with Recurrent Multistage Fusion
abstract
Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities.Comprehending multimodal language requires modeling not only the interactions within each modality (intra-modal interactions) but more importantly the interactions between modalities (cross-modal interactions).In this paper, we propose the Recurrent Multistage Fusion Network (RMFN) which decomposes the fusion problem into multiple stages, each of them focused on a subset of multimodal signals for specialized, effective fusion.Crossmodal interactions are modeled using this multistage fusion approach which builds upon intermediate representations of previous stages.Temporal and intra-modal interactions are modeled by integrating our proposed fusion approach with a system of recurrent neural networks.The RMFN displays state-of-the-art performance in modeling human multimodal language across three public datasets relating to multimodal sentiment analysis, emotion recognition, and speaker traits recognition.We provide visualizations to show that each stage of fusion focuses on a different subset of multimodal signals, learning increasingly discriminative multimodal representations.
Paul Pu Liang, Liu Ziyin 0001, Amir Zadeh 0001, Louis-Philippe Morency
EMNLP4
2018 OpenFace 2.0: Facial Behavior Analysis Toolkit
abstract
Over the past few years, there has been an increased interest in automatic facial behavior analysis and understanding. We present OpenFace 2.0 - a tool intended for computer vision and machine learning researchers, affective computing community and people interested in building interactive applications based on facial behavior analysis. OpenFace 2.0 is an extension of OpenFace toolkit and is capable of more accurate facial landmark detection, head pose estimation, facial action unit recognition, and eye-gaze estimation. The computer vision algorithms which represent the core of OpenFace 2.0 demonstrate state-of-the-art results in all of the above mentioned tasks. Furthermore, our tool is capable of real-time performance and is able to run from a simple webcam without any specialist hardware. Finally, unlike a lot of modern approaches or toolkits, OpenFace 2.0 source code for training models and running them is freely available for research purposes.
Tadas Baltrusaitis, Amir Zadeh 0001, Yao Chong Lim, Louis-Philippe Morency
FG4
2018 Toward Visual Behavior Markers of Suicidal Ideation
abstract
Suicide is an increasingly present issue in our society whose eradication could be greatly aided by decision support technologies that can objectively identify behavior markers of suicidal ideation. In this paper, we examine the predictive ability of a variety of smiling and eye gaze behaviors in categorizing hospital patients by mental health status: patients with suicidal ideation, patients with other mental illnesses such as depression, or control group without suicidal ideation or mental illness. We study three main research questions related to suicide behavior markers: (1) Do people with suicidal ideation smile with different dynamics (e.g. genuine vs fake smile)? (2) Do smiles while speaking, listening, and laughing show different levels of occurrence between the three groups? (3) Is gaze aversion (e.g. looking down) also a useful behavior marker? To answer these questions, we created new behavioral annotations on 74 semi-structured interviews from hospital patients, each of them within one of the three mental health conditions. Our data analysis identified behavior markers of mental health status from both smiling and eye gaze behaviors. Using these behavioral features, we created predictive models that show promising results when distinguishing between these three mental health conditions, especially when differentiating suicidal from non-suicidal patients.
Naomi Eigbe, Tadas Baltrusaitis, Louis-Philippe Morency, John Pestian
FG3
2018 Edge Convolutional Network for Facial Action Intensity Estimation
abstract
In this paper, we propose a novel convolutional neural architecture for facial action unit intensity estimation. While Convolutional Neural Networks (CNNs) have shown great promise in a wide range of computer vision tasks, these achievements have not translated as well to facial expression analysis, with hand crafted features (e.g. the Histogram of Orientated Gradient) still being very competitive. We introduce a novel Edge Convolutional Network (ECN) that is able to capture subtle changes in facial appearance. Our model is able to learn edge-like detectors that can capture subtle wrinkles and facial muscle contours at multiple orientations and frequencies. The core novelty of our ECN model is in its first layer which integrates three main components: an edge filter generator, a receptive gate and a filter rotator. All the components are differentiable and our ECN model is end-to-end trainable and learns the important edge detectors for facial expression analysis. Experiments on two facial action unit datasets show that the proposed ECN outperforms state-of-the-art methods for both AU intensity estimation tasks.
Liandong Li, Tadas Baltrusaitis, Bo Sun 0006, Louis-Philippe Morency
FG4
2018 Multimodal Local-Global Ranking Fusion for Emotion Recognition
abstract
Emotion recognition is a core research area at the intersection of artificial intelligence and human communication analysis. It is a significant technical challenge since humans display their emotions through complex idiosyncratic combinations of the language, visual and acoustic modalities. In contrast to traditional multimodal fusion techniques, we approach emotion recognition from both direct person-independent and relative person-dependent perspectives. The direct person-independent perspective follows the conventional emotion recognition approach which directly infers absolute emotion labels from observed multimodal features. The relative person-dependent perspective approaches emotion recognition in a relative manner by comparing partial video segments to determine if there was an increase or decrease in emotional intensity. Our proposed model integrates these direct and relative prediction perspectives by dividing the emotion recognition task into three easier subtasks. The first subtask involves a multimodal local ranking of relative emotion intensities between two short segments of a video. The second subtask uses local rankings to infer global relative emotion ranks with a Bayesian ranking algorithm. The third subtask incorporates both direct predictions from observed multimodal behaviors and relative emotion ranks from local-global rankings for final emotion prediction. Our approach displays excellent performance on an audio-visual emotion recognition benchmark and improves over other algorithms for multimodal fusion.
Paul Pu Liang, Amir Zadeh 0001, Louis-Philippe Morency
ICMI3
2018 Toward Objective, Multifaceted Characterization of Psychotic Disorders: Lexical, Structural, and Disfluency Markers of Spoken Language
abstract
Psychotic disorders are forms of severe mental illness characterized by abnormal social function and a general sense of disconnect with reality. The evaluation of such disorders is often complex, as their multifaceted nature is often difficult to quantify. Multimodal behavior analysis technologies have the potential to help address this need and supply timelier and more objective decision support tools in clinical settings. While written language and nonverbal behaviors have been previously studied, the present analysis takes the novel approach of examining the rarely-studied modality of spoken language of individuals with psychosis as naturally used in social, face-to-face interactions. Our analyses expose a series of language markers associated with psychotic symptom severity, as well as interesting interactions between them. In particular, we examine three facets of spoken language: (1) lexical markers, through a study of the function of words; (2) structural markers, through a study of grammatical fluency; and (3) disfluency markers, through a study of dialogue self-repair. Additionally, we develop predictive models of psychotic symptom severity, which achieve significant predictive power on both positive and negative psychotic symptom scales. These results constitute a significant step toward the design of future multimodal clinical decision support tools for computational phenotyping of mental illness.
Alexandria K. Vail, Elizabeth S. Liebson, Justin T. Baker, Louis-Philippe Morency
ICMI4
2018 Multimodal Polynomial Fusion for Detecting Driver Distraction
abstract
Distracted driving is deadly, claiming 3,477 lives in the U.S. in 2015 alone. Although there has been a considerable amount of research on modeling the distracted behavior of drivers under various conditions, accurate automatic detection using multiple modalities and especially the contribution of using the speech modality to improve accuracy has received little attention. This paper introduces a new multimodal dataset for distracted driving behavior and discusses automatic distraction detection using features from three modalities: facial expression, speech and car signals. Detailed multimodal feature analysis shows that adding more modalities monotonically increases the predictive accuracy of the model. Finally, a simple and effective multimodal fusion technique using a polynomial fusion layer shows superior distraction detection results compared to the baseline SVM and neural network models.
Yulun Du, Alan W. Black, Louis-Philippe Morency, Maxine Eskénazi
INTERSPEECH3
2018 Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos
abstract
Emotion recognition in conversations is crucial for the development of empathetic machines. Present methods mostly ignore the role of inter-speaker dependency relations while classifying emotions in conversations. In this paper, we address recognizing utterance-level emotions in dyadic conversational videos. We propose a deep neural framework, termed conversational memory network, which leverages contextual information from the conversation history. The framework takes a multimodal approach comprising audio, visual and textual features with gated recurrent units to model past utterances of each speaker into memories. Such memories are then merged using attention-based hops to capture inter-speaker dependencies. Experiments show an accuracy improvement of 3-4% over the state of the art.
Devamanyu Hazarika, Soujanya Poria, Amir Zadeh 0001, Erik Cambria, Louis-Philippe Morency, Roger Zimmermann
NAACL-HLT5
2018 Speaker-Follower Models for Vision-and-Language Navigation
abstract
Navigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete low-level motor behaviors; much of the missing information must be inferred based on perceptual context. In machine learning settings, this is doubly challenging: it is difficult to collect enough annotated data to enable learning of this reasoning process from scratch, and also difficult to implement the reasoning process using generic sequence models. Here we describe an approach to vision-and-language navigation that addresses both these issues with an embedded speaker model. We use this speaker model to (1) synthesize new instructions for data augmentation and to (2) implement pragmatic reasoning, which evaluates how well candidate action sequences explain an instruction. Both steps are supported by a panoramic action space that reflects the granularity of human-generated instructions. Experiments show that all three components of this approach---speaker-driven data augmentation, pragmatic reasoning and panoramic action space---dramatically improve the performance of a baseline instruction follower, more than doubling the success rate over the best existing approach on a standard benchmark.
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Daniel Klein 0001, Trevor Darrell
NeurIPS6
2018 Factorized Convolutional Networks: Unsupervised Fine-Tuning for Image Clustering
abstract
Deep convolutional neural networks (CNNs) have recognized promise as universal representations for various image recognition tasks. One of their properties is the ability to transfer knowledge from a large annotated source dataset (e.g., ImageNet) to a (typically smaller) target dataset. This is usually accomplished through supervised fine-tuning on labeled new target data. In this work, we address "unsupervised fine-tuning" that transfers a pre-trained network to target tasks with unlabeled data such as image clustering tasks. To this end, we introduce group-sparse non-negative matrix factorization (GSNMF), a variant of NMF, to identify a rich set of high-level latent variables that are informative on the target task. The resulting "factorized convolutional network" (FCN) can itself be seen as a feed-forward model that combines CNN and two-layer structured NMF. We empirically validate our approach and demonstrate state-of-the-art image clustering performance on challenging scene (MIT-67) and fine-grained (Birds-200, Flowers-102) benchmarks. We further show that, when used as unsupervised initialization, our approach improves image classification performance as well.
Liangyan Gui, Liangke Gui, Yu-Xiong Wang, Louis-Philippe Morency, José M. F. Moura
WACV4
2018 GazeDirector: Fully Articulated Eye Gaze Redirection in Video
abstract
Abstract We present GazeDirector, a new approach for eye gaze redirection that uses model‐fitting. Our method first tracks the eyes by fitting a multi‐part eye region model to video frames using analysis‐by‐synthesis, thereby recovering eye region shape, texture, pose, and gaze simultaneously. It then redirects gaze by 1) warping the eyelids from the original image using a model‐derived flow field, and 2) rendering and compositing synthesized 3D eyeballs onto the output image in a photorealistic manner. GazeDirector allows us to change where people are looking without person‐specific training data, and with full articulation, i.e. we can precisely specify new gaze directions in 3D. Quantitatively, we evaluate both model‐fitting and gaze synthesis, with experiments for gaze estimation and redirection on the Columbia gaze dataset. Qualitatively, we compare GazeDirector against recent work on gaze redirection, showing better results especially for large redirection angles. Finally, we demonstrate gaze redirection on YouTube videos by introducing new 3D gaze targets and by manipulating visual behavior.
Erroll Wood, Tadas Baltrusaitis, Louis-Philippe Morency, Peter Robinson 0001, Andreas Bulling
Comput. Graph. Forum3
2017 Integrating Verbal and Nonvebval Input into a Dynamic Response Spoken Dialogue System
abstract
In this work, we present a dynamic response spoken dialogue system (DRSDS). It is capable of understanding the verbal and nonverbal language of users and making instant, situation-aware response. Incorporating with two external systems, MultiSense and email summarization, we built an email reading agent on mobile device to show the functionality of DRSDS.
Ting-Yao Hu, Chirag Raman, Salvador Medina Maza, Liangke Gui, Tadas Baltrusaitis, Robert E. Frederking, Louis-Philippe Morency, Alan W. Black, Maxine Eskénazi
AAAI7
2017 Local-global ranking for facial expression intensity estimation
abstract
Facial action units provide an objective characterization of facial muscle movements. Automatic estimation of facial action unit intensities is a challenging problem given individual differences in neutral face appearances and the need to generalize across different pose, illumination and datasets. In this paper, we introduce the Local-Global Ranking method as a novel alternative to direct prediction of facial action unit intensities. Our method takes advantage of the additional information present in videos and image collections of the same person (e.g. a photo album). Instead of trying to estimate facial expression intensities independently for each image, our proposed method performs a two-stage ranking: a local pair-wise ranking followed by a global ranking. The local ranking is designed to be accurate and robust by making a simple 3-class comparison (higher, equal, or lower) between randomly sampled pairs of images. We use a Bayesian model to integrate all these pair-wise rankings and construct a global ranking. Our Local-Global Ranking method shows state-of-the-art performance on two publicly-available datasets. Our cross-dataset experiments also show better generalizability.
Tadas Baltrusaitis, Liandong Li, Louis-Philippe Morency
ACII3
2017 Hand2Face: Automatic synthesis and recognition of hand over face occlusions
abstract
A person's face discloses important information about their affective state. Although there has been extensive research on recognition of facial expressions, the performance of existing approaches is challenged by facial occlusions. Facial occlusions are often treated as noise and discarded in recognition of affective states. However, hand over face occlusions can provide additional information for recognition of some affective states such as curiosity, frustration and boredom. One of the reasons that this problem has not gained attention is the lack of naturalistic occluded faces that contain hand over face occlusions as well as other types of occlusions. Traditional approaches for obtaining affective data are time demanding and expensive, which limits researchers in affective computing to work on small datasets. This limitation affects the generalizability of models and deprives researchers from taking advantage of recent advances in deep learning that have shown great success in many fields but require large volumes of data. In this paper, we first introduce a novel framework for synthesizing naturalistic facial occlusions from an initial dataset of non-occluded faces and separate images of hands, reducing the costly process of data collection and annotation. We then propose a model for facial occlusion type recognition to differentiate between hand over face occlusions and other types of occlusions such as scarves, hair, glasses and objects. Finally, we present a model to localize hand over face occlusions and identify the occluded regions of the face.
Behnaz Nojavanasghari, Charles E. Hughes, Tadas Baltrusaitis, Louis-Philippe Morency
ACII4
2017 Visual attention in schizophrenia: Eye contact and gaze aversion during clinical interactions
abstract
Many of the essential clues to the psychiatric condition of an individual lie within the nonverbal and communicative behavior patterns they express during social interactions. Unfortunately, these behaviors are particularly difficult to assess subjectively in a time-constrained environment, to which clinicians are often limited in realistic settings. The present analysis examines quantified patterns of gaze aversion across a set of persons recently admitted to an inpatient psychotic disorder unit at a major psychiatric hospital. These patterns are used to inform the development of discriminative models with the task of predicting schizophrenic symptom severity from both a typological and a dimensional assessment perspective. The results expose a novel set of gaze aversion behaviors distinguishing between positive subtype schizophrenia, characterized by excessive behaviors such as hallucinations and grandiosity, and negative subtype schizophrenia, characterized by diminished behaviors such as blunted affect and emotional withdrawal. The predictive models constitute a significant step toward the development of automated tools to aid medical professionals in the diagnosis of psychotic disorders.
Alexandria K. Vail, Tadas Baltrusaitis, Luciana Pennant, Elizabeth S. Liebson, Justin T. Baker, Louis-Philippe Morency
ACII6
2017 Affect-LM: A Neural Language Model for Customizable Affective Text Generation
abstract
Sayan Ghosh, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Sayan Ghosh 0004, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer
ACL (1)4
2017 Context-Dependent Sentiment Analysis in User-Generated Videos
abstract
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, Louis-Philippe Morency. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh 0001, Louis-Philippe Morency
ACL (1)6
2017 Combating Human Trafficking with Multimodal Deep Models
abstract
Human trafficking is a global epidemic affecting millions of people across the planet.Sex trafficking, the dominant form of human trafficking, has seen a significant rise mostly due to the abundance of escort websites, where human traffickers can openly advertise among at-will escort advertisements.In this paper, we take a major step in the automatic detection of advertisements suspected to pertain to human trafficking.We present a novel dataset called Trafficking-10k, with more than 10,000 advertisements annotated for this task.The dataset contains two sources of information per advertisement: text and images.For the accurate detection of trafficking advertisements, we designed and trained a deep multimodal model called the Human Trafficking Deep Network (HTDN).
Edmund Tong, Amir Zadeh 0001, Cara Jones, Louis-Philippe Morency
ACL (1)4
2017 Temporal Attention-Gated Model for Robust Sequence Classification
abstract
Typical techniques for sequence classification are designed for well-segmented sequences which have been edited to remove noisy or irrelevant parts. Therefore, such methods cannot be easily applied on noisy sequences expected in real-world applications. In this paper, we present the Temporal Attention-Gated Model (TAGM) which integrates ideas from attention models and gated recurrent networks to better deal with noisy or unsegmented sequences. Specifically, we extend the concept of attention model to measure the relevance of each observation (time step) of a sequence. We then use a novel gated recurrent network to learn the hidden representation for the final prediction. An important advantage of our approach is interpretability since the temporal attention weights provide a meaningful value for the salience of each time step in the sequence. We demonstrate the merits of our TAGM approach, both for prediction accuracy and interpretability, on three different tasks: spoken digit recognition, text-based sentiment analysis and visual event recognition.
Wenjie Pei, Tadas Baltrusaitis, David M. J. Tax, Louis-Philippe Morency
CVPR4
2017 Tensor Fusion Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis is an increasingly popular research area, which extends the conventional language-based definition of sentiment analysis to a multimodal setup where other relevant modalities accompany language.In this paper, we pose the problem of multimodal sentiment analysis as modeling intra-modality and inter-modality dynamics.We introduce a novel model, termed Tensor Fusion Network, which learns both such dynamics end-to-end.The proposed approach is tailored for the volatile nature of spoken language in online videos as well as accompanying gestures and voice.In the experiments, our model outperforms state-ofthe-art approaches for both multimodal and unimodal sentiment analysis.
Amir Zadeh 0001, Minghai Chen, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
EMNLP5
2017 Curriculum Learning for Facial Expression Recognition
abstract
Over the past few years, there has been an increased interest in machine understanding and recognition of affective states based on facial expressions. While great progress has been made, there are still a lot of challenges facing automatic emotion recognition, namely: generalizability of models across datasets, accounting for individual differences, and recognition of subtle expressions. While deep learning techniques enabled a large amount of progress in many areas of computer vision, this progress has not yet been fully translated to emotion recognition. Our work attempts to partly address that by presenting a novel learning technique for deep learning methods that leads to better generalization for emotion recognition from facial expressions.
Liangke Gui, Tadas Baltrusaitis, Louis-Philippe Morency
FG3
2017 Local-Global Landmark Confidences for Face Recognition
abstract
A key to successful face recognition is accurate and reliable face alignment using automatically-detected facial landmarks. Given this strong dependency between face recognition and facial landmark detection, robust face recognition requires knowledge of when the facial landmark detection algorithm succeeds and when it fails. Facial landmark confidence represents this measure of success. In this paper, we propose two methods to measure landmark detection confidence: local confidence based on local predictors of each facial landmark, and global confidence based on a 3D rendered face model. A score fusion approach is also introduced to integrate these two confidences effectively. We evaluate both confidence metrics on two datasets for face recognition: JANUS CS2 and IJB-A datasets. Our experiments show up to 9% improvements when face recognition algorithm integrates the local-global confidence metrics.
KangGeon Kim, Feng-Ju Chang, Jongmoo Choi, Louis-Philippe Morency, Ramakant Nevatia, Gérard G. Medioni
FG4
2017 Investigating Facial Behavior Indicators of Suicidal Ideation
abstract
Suicide is the deliberate self-inflicted act with the intent to end one’s own life. It reflects both profound personal suffering and societal failure. While certain suicide risk factors are well understood, predicting suicide attempts remains a very challenging problem. In this paper, we investigate non-verbal facial behaviors to discriminate among control, mentally ill, and suicidal patients. For this task, we used a balanced corpus containing interviews of male and female patients with and without suicide ideation and/or mental health disorders from 3 different hospitals. In our experiments, we explored smiling, frowning, eyebrow raising, and head motion behaviors. We investigated both the occurrence of these behaviors and also how they were conducted. We found that facial behavior descriptors such as the percentage of smiles involving the contraction of the orbicularis oculi muscles (Duchenne smiles) had statistically significant differences between the suicidal and nonsuicidal groups. Our experiments also demonstrated that the stage of the interview in which these facial behaviors occur impacts their discriminative power.
Eugene Laksana, Tadas Baltrusaitis, Louis-Philippe Morency, John Pestian
FG3
2017 Constrained Ensemble Initialization for Facial Landmark Tracking in Video
abstract
Accurate and robust facial landmark tracking is a crucial step for face recognition and affect analysis systems. We often want to not only detect facial landmarks in images but to be able to track them reliably and consistently over time. Recently there has been an increase in research interest in facial landmark detection, especially in cascaded regression based methods such as the Supervised Descent Method (SDM). However, while facial landmark detection in images has improved significantly, comparably very little attention has been given to the task of landmark detection/tracking in videos. In our work we present a novel initialization procedure that can help with cascaded regression based facial landmark detection and tracking. Our initialization technique exploits the fact that cascaded regression is sensitive to initialization noise, especially in the presence of out-of-plane head pose variation, e.g. when a person is looking down when reading or during fast head motion. Our approach allows to learn good candidates for initialization, that we exploit in our tracking framework. We evaluate our technique on 300VW dataset – a large publicly available corpus of in-the-wild videos and demonstrate its effectiveness for a number of cascaded-regression landmark detection approaches.
Christy Li, Tadas Baltrusaitis, Louis-Philippe Morency
FG3
2017 Multi-level Multiple Attentions for Contextual Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis involves identifying sentiment in videos and is a developing field of research. Unlike current works, which model utterances individually, we propose a recurrent model that is able to capture contextual information among utterances. In this paper, we also introduce attentionbased networks for improving both context learning and dynamic feature fusion. Our model shows 6-8% improvement over the state of the art on a benchmark dataset.
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh 0001, Louis-Philippe Morency
ICDM6
2017 Select-additive learning: Improving generalization in multimodal sentiment analysis
abstract
Multimodal sentiment analysis is drawing an increasing amount of attention these days. It enables mining of opinions in video reviews which are now available aplenty on online platforms. However, multimodal sentiment analysis has only a few high-quality data sets annotated for training machine learning algorithms. These limited resources restrict the generalizability of models, where, for example, the unique characteristics of a few speakers (e.g., wearing glasses) may become a confounding factor for the sentiment classification task. In this paper, we propose a Select-Additive Learning (SAL) procedure that improves the generalizability of trained neural networks for multimodal sentiment analysis. In our experiments, we show that our SAL approach improves prediction accuracy significantly in all three modalities (verbal, acoustic, visual), as well as in their fusion. Our results show that SAL, even when trained on one dataset, achieves good generalization across two new test datasets.
Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, Eric P. Xing
ICME3
2017 Automatically predicting human knowledgeability through non-verbal cues
abstract
Humans possess an incredible ability to transmit and decode ``metainformation" through non-verbal actions in daily communication amongst each other. One communicative phenomena that is transmitted through these subtle cues is knowledgeability. In this work, we conduct two experiments. First, we analyze which non-verbal features are important for identifying knowledgeable people when responding to a question. Next, we train a model to predict the knowledgeability of speakers in a game show setting. We achieve results that surpass chance and human performance using a multimodal approach fusing prosodic and visual features. We believe computer systems that can incorporate emotional reasoning of this level can greatly improve human-computer communication and interaction.
Abdelwahab Bourai, Tadas Baltrusaitis, Louis-Philippe Morency
ICMI3
2017 Multimodal sentiment analysis with word-level fusion and reinforcement learning
abstract
With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment analysis which focus on holistic information in speech segments such as bag of words representations and average facial expression intensity, we propose a novel deep architecture for multimodal sentiment analysis that is able to perform modality fusion at the word level. In this paper, we propose the Gated Multimodal Embedding LSTM with Temporal Attention (GME-LSTM(A)) model that is composed of 2 modules. The Gated Multimodal Embedding allows us to alleviate the difficulties of fusion when there are noisy modalities. The LSTM with Temporal Attention can perform word level fusion at a finer fusion resolution between the input modalities and attends to the most important time steps. As a result, the GME-LSTM(A) is able to better model the multimodal structure of speech through time and perform better sentiment comprehension. We demonstrate the effectiveness of this approach on the publicly-available Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis (CMU-MOSI) dataset by achieving state-of-the-art sentiment classification and regression results. Qualitative analysis on our model emphasizes the importance of the Temporal Attention Layer in sentiment prediction because the additional acoustic and visual modalities are noisy. We also demonstrate the effectiveness of the Gated Multimodal Embedding in selectively filtering these noisy modalities out. These results and analysis open new areas in the study of sentiment analysis in human communication and provide new models for multimodal fusion.
Minghai Chen, Paul Pu Liang, Tadas Baltrusaitis, Amir Zadeh 0001, Louis-Philippe Morency
ICMI6
2017 Computational Analysis of Acoustic Descriptors in Psychotic Patients
Torsten Wörtwein, Tadas Baltrusaitis, Eugene Laksana, Luciana Pennant, Elizabeth S. Liebson, Dost Öngür, Justin T. Baker, Louis-Philippe Morency
INTERSPEECH8
2017 Temporally Selective Attention Model for Social and Affective State Recognition in Multimedia Content
abstract
The sheer amount of human-centric multimedia content has led to increased research on human behavior understanding. Most existing methods model behavioral sequences without considering the temporal saliency. This work is motivated by the psychological observation that temporally selective attention enables the human perceptual system to process the most relevant information. In this paper, we introduce a new approach, named Temporally Selective Attention Model (TSAM), designed to selectively attend to salient parts of human-centric video sequences. Our TSAM models learn to recognize affective and social states using a new loss function called speaker-distribution loss. Extensive experiments show that our model achieves the state-of-the-art performance on rapport detection and multimodal sentiment analysis. We also show that our speaker-distribution loss function can generalize to other computational models, improving the prediction performance of deep averaging network and Long Short Term Memory (LSTM).
Liangke Gui, Michael A. Madaio, Amy Ogan, Justine Cassell, Louis-Philippe Morency
ACM Multimedia6
2017 MultiSense - Context-Aware Nonverbal Behavior Analysis Framework: A Psychological Distress Use Case
abstract
During face-to-face interactions, people naturally integrate nonverbal behaviors such as facial expressions and body postures as part of the conversation to infer the communicative intent or emotional state of their interlocutor. The interpretation of these nonverbal behaviors will often be contextualized by interactional cues such as the previous spoken question, the general discussion topic or the physical environment. A critical step in creating computers able to understand or participate in this type of social face-to-face interactions is to develop a computational platform to synchronously recognize nonverbal behaviors as part of the interactional context. In this platform, information for the acoustic and visual modalities should be carefully synchronized and rapidly processed. At the same time, contextual and interactional cues should be remembered and integrated to better interpret nonverbal (and verbal) behaviors. In this article, we introduce a real-time computational framework, MultiSense, which offers flexible and efficient synchronization approaches for context-based nonverbal behavior analysis. MultiSense is designed to utilize interactional cues from both interlocutors (e.g., from the computer and the human participant) and integrate this contextual information when interpreting nonverbal behaviors. MultiSense can also assimilate behaviors over a full interaction and summarize the observed affective states of the user. We demonstrate the capabilities of the new framework with a concrete use case from the mental health domain where MultiSense is used as part of a decision support tool to assess indicators of psychological distress such as depression and post-traumatic stress disorder (PTSD). In this scenario, MultiSense not only infers psychological distress indicators from nonverbal behaviors but also broadcasts the user state in real-time to a virtual agent (i.e., a digital interviewer) designed to conduct semi-structured interviews with human participants. Our experiments show the added value of our multimodal synchronization approaches and also demonstrate the importance of MultiSense contextual interpretation when inferring distress indicators.
Giota Stratou, Louis-Philippe Morency
IEEE Trans. Affect. Comput.2
2017 Adolescent Suicidal Risk Assessment in Clinician-Patient Interaction
abstract
Youth suicide is a major public health problem. It is the third leading cause of death in the United States for ages 13 through 18. Many adolescents that face suicidal thoughts or make a suicide plan never seek professional care or help. Within this work, we evaluate both verbal and nonverbal responses to a five-item ubiquitous questionnaire to identify and assess suicidal risk of adolescents. We utilize a machine learning approach to identify suicidal from non-suicidal speech as well as characterize adolescents that repeatedly attempted suicide in the past. Our findings investigate both verbal and nonverbal behavior information of the face-to-face clinician-patient interaction. We investigate 60 audio-recorded dyadic clinician-patient interviews of 30 suicidal (13 repeaters and 17 non-repeaters) and 30 non-suicidal adolescents. The interaction between clinician and adolescents is statistically analyzed to reveal differences between suicidal versus non-suicidal adolescents and to investigate suicidal repeaters’ behaviors in comparison to suicidal non-repeaters. By using a hierarchical classifier we were able to show that the verbal responses to the ubiquitous questions sections of the interviews were useful to discriminate suicidal and non-suicidal patients. However, to additionally classify suicidal repeaters and suicidal non-repeaters more information especially nonverbal information is required.
Verena Venek, Stefan Scherer, Louis-Philippe Morency, Albert A. Rizzo, John Pestian
IEEE Trans. Affect. Comput.3
2016 Holistically Constrained Local Model: Going Beyond Frontal Poses for Facial Landmark Detection
KangGeon Kim, Tadas Baltrusaitis, Amir Zadeh 0001, Louis-Philippe Morency, Gérard G. Medioni
BMVC4
2016 Extending Long Short-Term Memory for Multi-View Structured Learning
Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrusaitis, Roland Göcke
ECCV (7)2
2016 A 3D Morphable Eye Region Model for Gaze Estimation
Erroll Wood, Tadas Baltrusaitis, Louis-Philippe Morency, Peter Robinson 0001, Andreas Bulling
ECCV (1)3
2016 Riding an emotional roller-coaster: A multimodal study of young child's math problem solving activities
Lujie Chen, Zhuyun Xia, Zhanmei Song, Louis-Philippe Morency, Artur Dubrawski
EDM5
2016 Unsupervised Text Recap Extraction for TV Series
Shikun Zhang, Louis-Philippe Morency
EMNLP3
2016 Learning an appearance-based gaze estimator from one million synthesised images
abstract
Learning-based methods for appearance-based gaze estimation achieve state-of-the-art performance in challenging real-world settings but require large amounts of labelled training data. Learning-by-synthesis was proposed as a promising solution to this problem but current methods are limited with respect to speed, appearance variability, and the head pose and gaze angle distribution they can synthesize. We present UnityEyes, a novel method to rapidly synthesize large amounts of variable eye region images as training data. Our method combines a novel generative 3D model of the human eye region with a real-time rendering framework. The model is based on high-resolution 3D face scans and uses real-time approximations for complex eyeball materials and structures as well as anatomically inspired procedural geometry methods for eyelid animation. We show that these synthesized images can be used to estimate gaze in difficult in-the-wild scenarios, even for extreme gaze angles or in cases in which the pupil is fully occluded. We also demonstrate competitive gaze estimation results on a benchmark in-the-wild dataset, despite only using a light-weight nearest-neighbor algorithm. We are making our UnityEyes synthesis framework available online for the benefit of the research community.
Erroll Wood, Tadas Baltrusaitis, Louis-Philippe Morency, Peter Robinson 0001, Andreas Bulling
ETRA3
2016 EmoReact: a multimodal approach and dataset for recognizing emotional responses in children
abstract
Automatic emotion recognition plays a central role in the technologies underlying social robots, affect-sensitive human computer interaction design and affect-aware tutors. Although there has been a considerable amount of research on automatic emotion recognition in adults, emotion recognition in children has been understudied. This problem is more challenging as children tend to fidget and move around more than adults, leading to more self-occlusions and non-frontal head poses. Also, the lack of publicly available datasets for children with annotated emotion labels leads most researchers to focus on adults. In this paper, we introduce a newly collected multimodal emotion dataset of children between the ages of four and fourteen years old. The dataset contains 1102 audio-visual clips annotated for 17 different emotional states: six basic emotions, neutral, valence and nine complex emotions including curiosity, uncertainty and frustration. Our experiments compare unimodal and multimodal emotion recognition baseline models to enable future research on this topic. Finally, we present a detailed analysis of the most indicative behavioral cues for emotion recognition in children.
Behnaz Nojavanasghari, Tadas Baltrusaitis, Charles E. Hughes, Louis-Philippe Morency
ICMI4
2016 Deep multimodal fusion for persuasiveness prediction
abstract
Persuasiveness is a high-level personality trait that quantifies the influence a speaker has on the beliefs, attitudes, intentions, motivations, and behavior of the audience. With social multimedia becoming an important channel in propagating ideas and opinions, analyzing persuasiveness is very important. In this work, we use the publicly available Persuasive Opinion Multimedia (POM) dataset to study persuasion. One of the challenges associated with this problem is the limited amount of annotated data. To tackle this challenge, we present a deep multimodal fusion architecture which is able to leverage complementary information from individual modalities for predicting persuasiveness. Our methods show significant improvement in performance over previous approaches.
Behnaz Nojavanasghari, Deepak Gopinath, Jayanth Koushik, Tadas Baltrusaitis, Louis-Philippe Morency
ICMI5
2016 Representation Learning for Speech Emotion Recognition
Sayan Ghosh 0004, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer
INTERSPEECH3
2016 Recognizing Human Actions in the Motion Trajectories of Shapes
abstract
People naturally anthropomorphize the movement of nonliving objects, as social psychologists Fritz Heider and Marianne Simmel demonstrated in their influential 1944 research study. When they asked participants to narrate an animated film of two triangles and a circle moving in and around a box, participants described the shapes' movement in terms of human actions. Using a framework for authoring and annotating animations in the style of Heider and Simmel, we established new crowdsourced datasets where the motion trajectories of animated shapes are labeled according to the actions they depict. We applied two machine learning approaches, a spatial-temporal bag-of-words model and a recurrent neural network, to the task of automatically recognizing actions in these datasets. Our best results outperformed a majority baseline and showed similarity to human performance, which encourages further use of these datasets for modeling perception from motion trajectories. Future progress on simulating human-like motion perception will require models that integrate motion information with top-down contextual knowledge.
Melissa Roemmele, Soja-Marie Morgens, Andrew S. Gordon, Louis-Philippe Morency
IUI4
2016 Manipulating the Perception of Virtual Audiences Using Crowdsourced Behaviors
Mathieu Chollet, Nithin Chandrashekhar, Ari Shapiro, Louis-Philippe Morency, Stefan Scherer
IVA4
2016 A Multimodal Corpus for the Assessment of Public Speaking Ability and Anxiety
Mathieu Chollet, Torsten Wörtwein, Louis-Philippe Morency, Stefan Scherer
LREC3
2016 Keynote - Modeling Human Communication Dynamics
abstract
Human face-to-face communication is a little like a dance, in that participants continuously adjust their behaviors based on verbal and nonverbal cues from the social context. Today's computers and interactive devices are still lacking many of these human-like abilities to hold fluid and natural interactions. Leveraging recent advances in machine learning, audio-visual signal processing and computational linguistic, my research focuses on creating computational technologies able to analyze, recognize and predict human subtle communicative behaviors in social context. I formalize this new research endeavor with a Human Communication Dynamics framework, addressing four key computational challenges: behavioral dynamic, multimodal dynamic, interpersonal dynamic and societal dynamic. Central to this research effort is the introduction of new probabilistic models able to learn the temporal and fine-grained latent dependencies across behaviors, modalities and interlocutors. In this talk, I will present some of our recent achievements modeling multiple aspects of human communication dynamics, motivated by applications in healthcare (depression, PTSD, suicide, autism), education (learning analytics), business (negotiation, interpersonal skills) and social multimedia (opinion mining, social influence).
Louis-Philippe Morency
SIGDIAL Conference1
2016 OpenFace: An open source facial behavior analysis toolkit
abstract
Over the past few years, there has been an increased interest in automatic facial behavior analysis and understanding. We present OpenFace - an open source tool intended for computer vision and machine learning researchers, affective computing community and people interested in building interactive applications based on facial behavior analysis. OpenFace is the first open source tool capable of facial landmark detection, head pose estimation, facial action unit recognition, and eye-gaze estimation. The computer vision algorithms which represent the core of OpenFace demonstrate state-of-the-art results in all of the above mentioned tasks. Furthermore, our tool is capable of real-time performance and is able to run from a simple webcam without any specialist hardware. Finally, OpenFace allows for easy integration with other applications and devices through a lightweight messaging system.
Tadas Baltrusaitis, Peter Robinson 0001, Louis-Philippe Morency
WACV3
2016 Self-Reported Symptoms of Depression and PTSD Are Associated with Reduced Vowel Space in Screening Interviews
abstract
Reduced frequency range in vowel production is a well documented speech characteristic of individuals with psychological and neurological disorders. Affective disorders such as depression and post-traumatic stress disorder (PTSD) are known to influence motor control and in particular speech production. The assessment and documentation of reduced vowel space and reduced expressivity often either rely on subjective assessments or on analysis of speech under constrained laboratory conditions (e.g. sustained vowel production, reading tasks). These constraints render the analysis of such measures expensive and impractical. Within this work, we investigate an automatic unsupervised machine learning based approach to assess a speaker's vowel space. Our experiments are based on recordings of 253 individuals. Symptoms of depression and PTSD are assessed using standard self-assessment questionnaires and their cut-off scores. The experiments show a significantly reduced vowel space in subjects that scored positively on the questionnaires. We show the measure's statistical robustness against varying demographics of individuals and articulation rate. The reduced vowel space for subjects with symptoms of depression can be explained by the common condition of psychomotor retardation influencing articulation and motor control. These findings could potentially support treatment of affective disorders, like depression and PTSD in the future.
Stefan Scherer, Gale M. Lucas, Jonathan Gratch, Albert A. Rizzo, Louis-Philippe Morency
IEEE Trans. Affect. Comput.5
2016 Multimodal Analysis and Prediction of Persuasiveness in Online Social Multimedia
Sunghyun Park 0001, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, Louis-Philippe Morency
ACM Trans. Interact. Intell. Syst.5
2015 SimSensei Demonstration: A Perceptive Virtual Human Interviewer for Healthcare Applications
abstract
We present the SimSensei system, a fully automatic virtual agent that conducts interviews to assess indicators of psychological distress. We emphasize on the perception part of the system, a multimodal framework which captures and analyzes user state for both behavioral understanding and interactional purposes.
Louis-Philippe Morency, Giota Stratou, David DeVault, Arno Hartholt, Margot Lhommet, Gale M. Lucas, Fabrizio Morbini, Kallirroi Georgila, Stefan Scherer, Jonathan Gratch, Stacy Marsella, David R. Traum, Albert A. Rizzo
AAAI1
2015 A multi-label convolutional neural network approach to cross-domain action unit detection
abstract
Action Unit (AU) detection from facial images is an important classification task in affective computing. However most existing approaches use carefully engineered feature extractors along with off-the-shelf classifiers. There has also been less focus on how well classifiers generalize when tested on different datasets. In our paper, we propose a multi-label convolutional neural network approach to learn a shared representation between multiple AUs directly from the input image. Experiments on three AU datasets- CK+, DISFA and BP4D indicate that our approach obtains competitive results on all datasets. Cross-dataset experiments also indicate that the network generalizes well to other datasets, even when under different training and testing conditions.
Sayan Ghosh 0004, Eugene Laksana, Stefan Scherer, Louis-Philippe Morency
ACII4
2015 A demonstration of the perception system in SimSensei, a virtual human application for healthcare interviews
abstract
We present the SimSensei system, a fully automatic virtual agent that conducts interviews to assess indicators of psychological distress. With this demo, we focus our attention on the perception part of the system, a multimodal framework which captures and analyzes user state behavior for both behavioral understanding and interactional purposes. We will demonstrate real-time user state sensing as a part of the SimSensei architecture and discuss how this technology enabled automatic analysis of behaviors related to psychological distress.
Giota Stratou, Louis-Philippe Morency, David DeVault, Arno Hartholt, Edward Fast, Margot Lhommet, Gale M. Lucas, Fabrizio Morbini, Kallirroi Georgila, Stefan Scherer, Jonathan Gratch, Stacy Marsella, David R. Traum, Albert A. Rizzo
ACII2
2015 Automatic assessment and analysis of public speaking anxiety: A virtual audience case study
abstract
Public speaking has become an integral part of many professions and is central to career building opportunities. Yet, public speaking anxiety is often referred to as the most common fear in everyday life and can hinder one's ability to speak in public severely. While virtual and real audiences have been successfully utilized to treat public speaking anxiety in the past, little work has been done on identifying behavioral characteristics of speakers suffering from anxiety. In this work, we focus on the characterization of behavioral indicators and the automatic assessment of public speaking anxiety. We identify several indicators for public speaking anxiety, among them are less eye contact with the audience, reduced variability in the voice, and more pauses. We automatically assess the public speaking anxiety as reported by the speakers through a self-assessment questionnaire using a speaker independent paradigm. Our approach using ensemble trees achieves a high correlation between ground truth and our estimation (r=0.825). Complementary to automatic measures of anxiety, we are also interested in speakers' perceptual differences when interacting with a virtual audience based on their level of anxiety in order to improve and further the development of virtual audiences for the training of public speaking and the reduction of anxiety.
Torsten Wörtwein, Louis-Philippe Morency, Stefan Scherer
ACII2
2015 Time-slice Prediction of Dyadic Human Activities
abstract
Recognizing human activities from video data is being leveraged for surveillance and human-computer interaction applications. In this paper, we introduce the problem of time-slice activity recognition which aims to explore human activity at a smaller temporal granularity. Time-slice recognition is able to infer human behaviors from a short temporal window. It has been shown that the temporal slice analysis is helpful for motion characterization and in general for video content representation. These studies motivate us to consider time-slices for activity recognition. We present in Figure 1 an overview of our approach based on timeslice action prediction and contrast it with the conventional approaches which recognize actions based on either the whole video sequence (referred as “holistic” approach) or the first part of it (early recognition). Our time-slice approach studies not only the beginning of the action sequence but generalizes this to any short-term observation anywhere in the video sequence. Another key novelty is in the explicit modeling of the uncertainty occurring when predicting actions based on time-slices. TAP Dataset: We introduce a new dataset, named Time-slice Action Prediction (TAP) dataset, to evaluate our proposed feature descriptors and enable future research on this topic. The dataset was created by extracting time-slices from existing public human action datasets (UT-Interaction, HMDB, TV Interaction, and Hollywood datasets) and perform a perception study with multiple annotators giving continuous ratings for each action. The continuous ratings allow to represent the uncertainty in timeslice action prediction. 3 annotators rated each time-slice on how likely a specific action is occurring. For each time-slice and for each action, the annotator was asked to pick one of 5 likelihoods from “Definitely Not Occurring” to “Definitely Occurring”. Figure 3 illustrates how annotators rated for two example videos. Methodology: Stage 1Discriminative segments: When analyzing an interaction, we can definitely recognize the ongoing activity from specific time slices such as “two people are shaking each other’s hands” slice in handshaking activity. To extract discriminative segments from our dataset, we used Fleiss’ kappa coefficient k [2] to measure the reliability of agreement between annotators. For each interaction video, time-slices where the annotators are in complete agreement, i.e. k=1, on definitely including the interaction of interest, are selected as discriminative segments. Stage 2Predict-STIP: Existing STIP detectors are vulnerable to model the inherent uncertainty in partially observed action recognition Figure 2: Human annotation: This figure shows the average rate of 3 annotators for two video examples: hug and push. The label provided by one annotator is converted to a number on a linear scale from 0 to 1 called the average rate. This average rate will be used to evaluate the performance of our method. Time-slices between dashed lines is the discriminative segment of the interaction.
Maryam Ziaeefard, Robert Bergevin, Louis-Philippe Morency
BMVC3
2015 Exploring feedback strategies to improve public speaking: an interactive virtual audience framework
abstract
Good public speaking skills convey strong and effective communication, which is critical in many professions and used in everyday life. The ability to speak publicly requires a lot of training and practice. Recent technological developments enable new approaches for public speaking training that allow users to practice in a safe and engaging environment. We explore feedback strategies for public speaking training that are based on an interactive virtual audience paradigm. We investigate three study conditions: (1) a non-interactive virtual audience (control condition), (2) direct visual feedback, and (3) nonverbal feedback from an interactive virtual audience. We perform a threefold evaluation based on self-assessment questionnaires, expert assessments, and two objectively annotated measures of eye-contact and avoidance of pause fillers. Our experiments show that the interactive virtual audience brings together the best of both worlds: increased engagement and challenge as well as improved public speaking skills as judged by experts.
Mathieu Chollet, Torsten Wörtwein, Louis-Philippe Morency, Ari Shapiro, Stefan Scherer
UbiComp3
2015 Reduced vowel space is a robust indicator of psychological distress: A cross-corpus analysis
abstract
Reduced frequency range in vowel production is a well documented speech characteristic of individuals' with psychological and neurological disorders. Depression is known to influence motor control and in particular speech production. The assessment and documentation of reduced vowel space and associated perceived hypoarticulation and reduced expressivity often rely on subjective assessments. Within this work, we investigate an automatic unsupervised machine learning approach to assess a speaker's vowel space within three distinct speech corpora and compare observed vowel space measures of subjects with and without psychological conditions associated with psychological distress, namely depression, post-traumatic stress disorder (PTSD), and suicidality. Our experiments are based on recordings of over 300 individuals. The experiments show a significantly reduced vowel space in conversational speech for depression, PTSD, and suicidality. We further observe a similar trend of reduced vowel space for read speech. A possible explanation for a reduced vowel space is psychomotor retardation, a common symptom of depression that influences motor control and speech production.
Stefan Scherer, Louis-Philippe Morency, Jonathan Gratch, John Pestian
ICASSP2
2015 Acoustic and para-verbal indicators of persuasiveness in social multimedia
abstract
Persuasive communication and interaction play an important and pervasive role in many aspects of our lives. With the rapid growth of social multimedia websites such as YouTube, it has become more important and useful to understand persuasiveness in the context of online social multimedia content. In this paper, we present our results of conducting various analyses of persuasiveness in speech with our multimedia corpus of 1,000 movie review videos obtained from ExpoTV.com, a popular social multimedia website. Our experiments firstly show that a speaker's level of persuasiveness can be predicted from acoustic characteristics and para-verbal cues related to speech fluency. Secondly, we show that taking acoustic cues in different time periods of a movie review can improve the performance of predicting a speaker's level of persuasiveness. Lastly, we show that a speaker's positive or negative attitude toward a topic influences the prediction performance as well.
Han Suk Shim, Sunghyun Park 0001, Moitreya Chatterjee, Stefan Scherer, Kenji Sagae, Louis-Philippe Morency
ICASSP6
2015 Combining Two Perspectives on Classifying Multimodal Data for Recognizing Speaker Traits
abstract
Human communication involves conveying messages both through verbal and non-verbal channels (facial expression, gestures, prosody, etc.). Nonetheless, the task of learning these patterns for a computer by combining cues from multiple modalities is challenging because it requires effective representation of the signals and also taking into consideration the complex interactions between them. From the machine learning perspective this presents a two-fold challenge: a) Modeling the intermodal variations and dependencies; b) Representing the data using an apt number of features, such that the necessary patterns are captured but at the same time allaying concerns such as over-fitting. In this work we attempt to address these aspects of multimodal recognition, in the context of recognizing two essential speaker traits, namely passion and credibility of online movie reviewers. We propose a novel ensemble classification approach that combines two different perspectives on classifying multimodal data. Each of these perspectives attempts to independently address the two-fold challenge. In the first, we combine the features from multiple modalities but assume inter-modality conditional independence. In the other one, we explicitly capture the correlation between the modalities but in a space of few dimensions and explore a novel clustering based kernel similarity approach for recognition. Additionally, this work investigates a recent technique for encoding text data that captures semantic similarity of verbal content and preserves word-ordering. The experimental results on a recent public dataset shows significant improvement of our approach over multiple baselines. Finally, we also analyze the most discriminative elements of a speaker's non-verbal behavior that contribute to his/her perceived credibility/passionateness.
Moitreya Chatterjee, Sunghyun Park 0001, Louis-Philippe Morency, Stefan Scherer
ICMI3
2015 Exploring Behavior Representation for Learning Analytics
abstract
Multimodal analysis has long been an integral part of studying learning. Historically multimodal analyses of learning have been extremely laborious and time intensive. However, researchers have recently been exploring ways to use multimodal computational analysis in the service of studying how people learn in complex learning environments. In an effort to advance this research agenda, we present a comparative analysis of four different data segmentation techniques. In particular, we propose affect- and pose-based data segmentation, as alternatives to human-based segmentation, and fixed-window segmentation. In a study of ten dyads working on an open-ended engineering design task, we find that affect- and pose-based segmentation are more effective, than traditional approaches, for drawing correlations between learning-relevant constructs, and multimodal behaviors. We also find that pose-based segmentation outperforms the two more traditional segmentation strategies for predicting student success on the hands-on task. In this paper we discuss the algorithms used, our results, and the implications that this work may have in non-education-related contexts.
Marcelo Worsley, Stefan Scherer, Louis-Philippe Morency, Paulo Blikstein
ICMI3
2015 Multimodal Public Speaking Performance Assessment
abstract
The ability to speak proficiently in public is essential for many professions and in everyday life. Public speaking skills are difficult to master and require extensive training. Recent developments in technology enable new approaches for public speaking training that allow users to practice in engaging and interactive environments. Here, we focus on the automatic assessment of nonverbal behavior and multimodal modeling of public speaking behavior. We automatically identify audiovisual nonverbal behaviors that are correlated to expert judges' opinions of key performance aspects. These automatic assessments enable a virtual audience to provide feedback that is essential for training during a public speaking performance. We utilize multimodal ensemble tree learners to automatically approximate expert judges' evaluations to provide post-hoc performance assessments to the speakers. Our automatic performance evaluation is highly correlated with the experts' opinions with r = 0.745 for the overall performance assessments. We compare multimodal approaches with single modalities and find that the multimodal ensembles consistently outperform single modalities.
Torsten Wörtwein, Mathieu Chollet, Boris Schauerte, Louis-Philippe Morency, Rainer Stiefelhagen, Stefan Scherer
ICMI4
2015 Predicting Co-verbal Gestures: A Deep and Temporal Modeling Approach
Chung-Cheng Chiu, Louis-Philippe Morency, Stacy Marsella
IVA2
2015 NRGsuite: a PyMOL plugin to perform docking simulations in real time using FlexAID
abstract
UNLABELLED: Ligand protein docking simulations play a fundamental role in understanding molecular recognition. Herein we introduce the NRGsuite, a PyMOL plugin that permits the detection of surface cavities in proteins, their refinements, calculation of volume and use, individually or jointly, as target binding-sites for docking simulations with FlexAID. The NRGsuite offers the users control over a large number of important parameters in docking simulations including the assignment of flexible side-chains and definition of geometric constraints. Furthermore, the NRGsuite permits the visualization of the docking simulation in real time. The NRGsuite give access to powerful docking simulations that can be used in structure-guided drug design as well as an educational tool. The NRGsuite is implemented in Python and C/C++ with an easy to use package installer. The NRGsuite is available for Windows, Linux and MacOS. AVAILABILITY AND IMPLEMENTATION: http://bcb.med.usherbrooke.ca/flexaid. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Francis Gaudreault, Louis-Philippe Morency, Rafael Najmanovich
Bioinform.2
2015 Variational Infinite Hidden Conditional Random Fields
abstract
Hidden conditional random fields (HCRFs) are discriminative latent variable models which have been shown to successfully learn the hidden structure of a given classification problem. An Infinite hidden conditional random field is a hidden conditional random field with a countably infinite number of hidden states, which rids us not only of the necessity to specify a priori a fixed number of hidden states available but also of the problem of overfitting. Markov chain Monte Carlo (MCMC) sampling algorithms are often employed for inference in such models. However, convergence of such algorithms is rather difficult to verify, and as the complexity of the task at hand increases the computational cost of such algorithms often becomes prohibitive. These limitations can be overcome by variational techniques. In this paper, we present a generalized framework for infinite HCRF models, and a novel variational inference approach on a model based on coupled Dirichlet Process Mixtures, the HCRF-DPM. We show that the variational HCRF-DPM is able to converge to a correct number of represented hidden states, and performs as well as the best parametric HCRFs-chosen via cross-validation-for the difficult tasks of recognizing instances of agreement, disagreement, and pain in audiovisual sequences.
Konstantinos Bousmalis, Stefanos Zafeiriou, Louis-Philippe Morency, Maja Pantic, Zoubin Ghahramani
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Preface of pattern recognition in human computer interaction
Friedhelm Schwenker, Stefan Scherer, Louis-Philippe Morency
Pattern Recognit. Lett.3
2015 I Can Already Guess Your Answer: Predicting Respondent Reactions during Dyadic Negotiation
abstract
Negotiation is a component deeply ingrained in our daily lives, and it can be challenging for a person to predict the respondent's reaction (acceptance or rejection) to a negotiation offer. In this work, we focus on finding acoustic and visual behavioral cues that are predictive of the respondent's immediate reactions using a face-to-face negotiation dataset, which consists of 42 dyadic interactions in a simulated negotiation setting. We show our results of exploring four different sources of information, namely nonverbal behavior of the proposer, that of the respondent, mutual behavior between the interactants related to behavioral symmetry and asymmetry, and past negotiation history between the interactants. Firstly, we show that considering other sources of information (other than the nonverbal behavior of the respondent) can also have comparable performance in predicting respondent reactions. Secondly, we show that automatically extracted mutual behavioral cues of symmetry and asymmetry are predictive partially due to their capturing information of the nature of the interaction itself, whether it is cooperative or competitive. Lastly, we identify audio-visual behavioral cues that are most predictive of the respondent's immediate reactions.
Sunghyun Park 0001, Stefan Scherer, Jonathan Gratch, Peter J. Carnevale, Louis-Philippe Morency
IEEE Trans. Affect. Comput.5
2014 Continuous Conditional Neural Fields for Structured Regression
Tadas Baltrusaitis, Peter Robinson 0001, Louis-Philippe Morency
ECCV (4)3
2014 Context-based signal descriptors of heart-rate variability for anxiety assessment
abstract
In this paper, we investigate the role of multiple context-based heart-rate variability descriptors for evaluating a person's psychological health, specifically anxiety disorders. The descriptors are extracted from visually sensed heart-rate signals obtained during the course of a semi-structured interview with a virtual human and can potentially integrate question context as well. The proposed descriptors are motivated by prior related work and are constructed based on histogram-based approaches, time and frequency domain analysis of heart-rate variability. In order to contextualize our descriptors, we use information about the polarity and intimacy levels of the questions asked. Our experiments reveal that the descriptors, both with and without context, perform far better than chance in predicting anxiety. Further on, we perform at-a-par with the state-of-the-art in predicting anxiety and other psychological disorders when we integrate the question context information into the descriptors.
Moitreya Chatterjee, Giota Stratou, Stefan Scherer, Louis-Philippe Morency
ICASSP4
2014 A Multimodal Context-based Approach for Distress Assessment
abstract
The increasing prevalence of psychological distress disorders, such as depression and post-traumatic stress, necessitates a serious effort to create new tools and technologies to help with their diagnosis and treatment. In recent years, new computational approaches were proposed to objectively analyze patient non-verbal behaviors over the duration of the entire interaction between the patient and the clinician. In this paper, we go beyond non-verbal behaviors and propose a tri-modal approach which integrates verbal behaviors with acoustic and visual behaviors to analyze psychological distress during the course of the dyadic semi-structured interviews. Our approach exploits the advantages of the dyadic nature of these interactions to contextualize the participant responses based on the affective components (intimacy and polarity levels) of the questions. We validate our approach using one of the largest corpus of semi-structured interviews for distress assessment which consists of 154 multimodal dyadic interactions. Our results show significant improvement on distress prediction performance when integrating verbal behaviors with acoustic and visual behaviors. In addition, our analysis shows that contextualizing the responses improves the prediction performance, most significantly with positive and intimate questions.
Sayan Ghosh 0004, Moitreya Chatterjee, Louis-Philippe Morency
ICMI3
2014 Computational Analysis of Persuasiveness in Social Multimedia: A Novel Dataset and Multimodal Prediction Approach
abstract
Our lives are heavily influenced by persuasive communication, and it is essential in almost any types of social interactions from business negotiation to conversation with our friends and family. With the rapid growth of social multimedia websites, it is becoming ever more important and useful to understand persuasiveness in the context of social multimedia content online. In this paper, we introduce our newly created multimedia corpus of 1,000 movie review videos obtained from a social multimedia website called ExpoTV.com, which will be made freely available to the research community. Our research results presented here revolve around the following 3 main research hypotheses. Firstly, we show that computational descriptors derived from verbal and nonverbal behavior can be predictive of persuasiveness. We further show that combining descriptors from multiple communication modalities (audio, text and visual) improve the prediction performance compared to using those from single modality alone. Secondly, we investigate if having prior knowledge of a speaker expressing a positive or negative opinion helps better predict the speaker's persuasiveness. Lastly, we show that it is possible to make comparable prediction of persuasiveness by only looking at thin slices (shorter time windows) of a speaker's behavior.
Sunghyun Park 0001, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, Louis-Philippe Morency
ICMI5
2014 Dyadic Behavior Analysis in Depression Severity Assessment Interviews
abstract
Previous literature suggests that depression impacts vocal timing of both participants and clinical interviewers but is mixed with respect to acoustic features. To investigate further, 57 middle-aged adults (men and women) with Major Depression Disorder and their clinical interviewers (all women) were studied. Participants were interviewed for depression severity on up to four occasions over a 21 week period using the Hamilton Rating Scale for Depression (HRSD), which is a criterion measure for depression severity in clinical trials. Acoustic features were extracted for both participants and interviewers using COVAREP Toolbox. Missing data occurred due to missed appointments, technical problems, or insufficient vocal samples. Data from 36 participants and their interviewers met criteria and were included for analysis to compare between high and low depression severity. Acoustic features for participants varied between men and women as expected, and failed to vary with depression severity for participants. For interviewers, acoustic characteristics strongly varied with severity of the interviewee's depression. Accommodation - the tendency of interactants to adapt their communicative behavior to each other - between interviewers and interviewees was inversely related to depression severity. These findings suggest that interviewers modify their acoustic features in response to depression severity, and depression severity strongly impacts interpersonal accommodation.
Stefan Scherer, Zakia Hammal, Ying Yang 0007, Louis-Philippe Morency, Jeffrey F. Cohn
ICMI4
2014 Toward crowdsourcing micro-level behavior annotations: the challenges of interface, training, and generalization
abstract
Research that involves human behavior analysis usually requires laborious and costly efforts for obtaining micro-level behavior annotations on a large video corpus. With the emerging paradigm of crowdsourcing however, these efforts can be considerably reduced. We first present OCTAB (Online Crowdsourcing Tool for Annotations of Behaviors), a web-based annotation tool that allows precise and convenient behavior annotations in videos, directly portable to popular crowdsourcing platforms. As part of OCTAB, we introduce a training module with specialized visualizations. The training module's design was inspired by an observational study of local experienced coders, and it enables an iterative procedure for effectively training crowd workers online. Finally, we present an extensive set of experiments that evaluates the feasibility of our crowdsourcing approach for obtaining micro-level behavior annotations in videos, showing the reliability improvement in annotation accuracy when properly training online crowd workers. We also show the generalization of our training approach to a new independent video corpus.
Sunghyun Park 0001, Philippa Shoemark, Louis-Philippe Morency
IUI3
2014 Towards Learning Nonverbal Identities from the Web: Automatically Identifying Visually Accentuated Words
Amir Zadeh 0001, Kenji Sagae, Louis-Philippe Morency
IVA3
2014 The Distress Analysis Interview Corpus of human and computer interviews
Jonathan Gratch, Ron Artstein, Gale M. Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, David R. Traum, Albert A. Rizzo, Louis-Philippe Morency
LREC13
2014 Search Strategies for Pattern Identification in Multimodal Data: Three Case Studies
abstract
The analysis of multimodal data benefits from meaningful search and retrieval. This paper investigates strategies of searching multimodal data for event patterns. Through three longitudinal case studies, we observed researchers exploring and identifying event patterns in multimodal data. The events were extracted from different multimedia signal sources ranging from annotated video transcripts to interaction logs. Each researcher's data has varying temporal characteristics (e.g., sparse, dense, or clustered) that posed several challenges for identifying relevant patterns. We identify unique search strategies and better understand the aspects that contributed to each.
Chreston A. Miller, Francis K. H. Quek, Louis-Philippe Morency
ICMR3
2014 A Demonstration of Dialogue Processing in SimSensei Kiosk
abstract
Fabrizio Morbini, David DeVault, Kallirroi Georgila, Ron Artstein, David Traum, Louis-Philippe Morency. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014.
Fabrizio Morbini, David DeVault, Kallirroi Georgila, Ron Artstein, David R. Traum, Louis-Philippe Morency
SIGDIAL Conference6
2014 Adolescent suicidal risk assessment in clinician-patient interaction: A study of verbal and acoustic behaviors
abstract
Suicide among adolescents is a major public health problem: it is the third leading cause of death in the US for ages 13-18. Up to now, there is no objective ways to assess the suicidal risk, i.e. whether a patient is non-suicidal, suicidal re-attempter (i.e. repeater) or suicidal non-repeater (i.e. individuals with one suicide attempt or showing signs of suicidal gestures or ideation). Therefore, features of the conversation including verbal information and nonverbal acoustic information were investigated from 60 audio-recorded interviews of 30 suicidal (13 repeaters and 17 non-repeaters) and 30 non-suicidal adolescents interviewed by a social worker. The interaction between clinician and patients was statistically analyzed to reveal differences between suicidal vs. non-suicidal adolescents and to investigate suicidal repeaters' behaviors in comparison to suicidal non-repeaters. By using a hierarchical ensemble classifier we were able to successfully discriminate non-suicidal patients, suicidal repeaters and suicidal non-repeaters.
Verena Venek, Stefan Scherer, Louis-Philippe Morency, Albert A. Rizzo, John Pestian
SLT3
2014 Relative facial action unit detection
abstract
This paper presents a subject-independent facial action unit (AU) detection method by introducing the concept of relative AU detection, for scenarios where the neutral face is not provided. We propose a new classification objective function which analyzes the temporal neighborhood of the current frame to decide if the expression recently increased, decreased or showed no change. This approach is a significant change from the conventional absolute method which decides about AU classification using the current frame, without an explicit comparison with its neighboring frames. Our proposed method improves robustness to individual differences such as face scale and shape, age-related wrinkles, and transitions among expressions (e.g., lower intensity of expressions). Our experiments on three publicly available datasets (Extended Cohn-Kanade (CK+), Bosphorus, and DISFA databases) show significant improvement of our approach over conventional absolute techniques.
Mahmoud Khademi, Louis-Philippe Morency
WACV2
2014 Automatic audiovisual behavior descriptors for psychological disorder analysis
Stefan Scherer, Giota Stratou, Gale M. Lucas, Marwa Mahmoud, Jill Boberg, Jonathan Gratch, Albert A. Rizzo, Louis-Philippe Morency
Image Vis. Comput.8
2013 Fifth International Workshop on Affective Interaction in Natural Environments (AFFINE 2013): Interacting with Affective Artefacts in the Wild
abstract
This workshop covers real-time computational techniques for the recognition and interpretation of human affective and social behaviour, and techniques for synthesis of believable social behaviour supporting real-time adaptive human-agent and human-robot interaction in real-world environments.
Ginevra Castellano, Kostas Karpouzis, Jean-Claude Martin, Louis-Philippe Morency, Christopher Peters 0001, Laurel D. Riek
ACII4
2013 Mutual Behaviors during Dyadic Negotiation: Automatic Prediction of Respondent Reactions
abstract
In this paper, we analyze face-to-face negotiation interactions with the goal of predicting the respondent's immediate reaction (i.e., accept or reject) to a negotiation offer. Supported by the theory of social rapport, we focus on mutual behaviors which are defined as nonverbal characteristics that occur due to interactional influence. These patterns include behavioral symmetry (e.g., synchronized smiles) as well as asymmetry (e.g., opposite postures) between the two negotiators. In addition, we put emphasis on finding audio-visual mutual behaviors that can be extracted automatically, with the vision of a real-time decision support tool. We introduce a dyadic negotiation dataset consisting of 42 face-to-face interactions and show experiments confirming the importance of multimodal and mutual behaviors.
Sunghyun Park 0001, Stefan Scherer, Jonathan Gratch, Peter J. Carnevale, Louis-Philippe Morency
ACII5
2013 Automatic Nonverbal Behavior Indicators of Depression and PTSD: Exploring Gender Differences
abstract
In this paper, we show that gender plays an important role in the automatic assessment of psychological conditions such as depression and post-traumatic stress disorder (PTSD). We identify a directly interpretable and intuitive set of predictive indicators, selected from three general categories of nonverbal behaviors: affect, expression variability and motor variability. For the analysis, we introduce a semi-structured virtual human interview dataset which includes 53 video recorded interactions. Our experiments on automatic classification of psychological conditions show that a gender-dependent approach significantly improves the performance over a gender agnostic one.
Giota Stratou, Stefan Scherer, Jonathan Gratch, Louis-Philippe Morency
ACII4
2013 Utterance-Level Multimodal Sentiment Analysis
Verónica Pérez-Rosas, Rada Mihalcea, Louis-Philippe Morency
ACL (1)3
2013 Action Recognition by Hierarchical Sequence Summarization
abstract
Recent progress has shown that learning from hierarchical feature representations leads to improvements in various computer vision tasks. Motivated by the observation that human activity data contains information at various temporal resolutions, we present a hierarchical sequence summarization approach for action recognition that learns multiple layers of discriminative feature representations at different temporal granularities. We build up a hierarchy dynamically and recursively by alternating sequence learning and sequence summarization. For sequence learning we use CRFs with latent variables to learn hidden spatio-temporal dynamics, for sequence summarization we group observations that have similar semantic meaning in the latent space. For each layer we learn an abstract feature representation through non-linear gate functions. This procedure is repeated to obtain a hierarchical sequence summary representation. We develop an efficient learning method to train our model and show that its complexity grows sub linearly with the size of the hierarchy. Experimental results show the effectiveness of our approach, achieving the best published results on the Arm Gesture and Canal9 datasets.
Yale Song, Louis-Philippe Morency, Randall Davis
CVPR2
2013 Speaker and language independent voice quality classification applied to unlabelled corpora of expressive speech
abstract
Voice quality plays a pivotal role in speech style variation. Therefore, control and analysis of voice quality is critical for many areas of speech technology. Until now, most work has focused on small purpose built corpora. In this paper we apply state-of-the-art voice quality analysis to large speech corpora built for expressive speech synthesis. A fuzzy-input fuzzy-output support vector machine classifier is trained and validated using features extracted from these corpora. We then apply this classifier to freely available audiobook data and demonstrate a clustering of the voice qualities that approximates the performance of human perceptual ratings. The ability to detect voice quality variation in these widely available unlabelled audiobook corpora means that the proposed method may be used as a valuable resource in expressive speech synthesis.
John Kane 0002, Stefan Scherer, Matthew P. Aylett, Louis-Philippe Morency, Christer Gobl
ICASSP4
2013 Investigating the speech characteristics of suicidal adolescents
abstract
Suicide is a very serious problem. In the United states it ranks as the second most frequent cause of death among teenagers between the ages of 12 and 17. In this work, we investigate speech characteristics of prosody as well as voice quality in a dyadic interview corpus with suicidal and non-suicidal adolescents. In these interviews the adolescents answer specifically designed questions. Based on this limited dataset, we reveal statistically significant differences in the speech patterns of suicidal adolescents within the investigated interview corpus. Further, we investigate the classification capabilities of machine learning approaches both on an utterance as well as an interview level. The work shows promising results in a speaker-independent classification experiment based on only a dozen speech features. We believe that once the algorithms are refined and integrated with other methods, they may be of value to the clinician.
Stefan Scherer, John Pestian, Louis-Philippe Morency
ICASSP3
2013 Speaker trait characterization in web videos: Uniting speech, language, and facial features
abstract
We present a multi-modal approach to speaker characterization using acoustic, visual and linguistic features. Full realism is provided by evaluation on a database of real-life web videos and automatic feature extraction including face and eye detection, and automatic speech recognition. Different segmentations are evaluated for the audio and video streams, and the statistical relevance of Linguistic Inquiry and Word Count (LIWC) features is confirmed. In the result, late multimodal fusion delivers 73, 92 and 73% average recall in binary age, gender and race classification on unseen test subjects, outperforming the best single modalities for age and race.
Felix Weninger, Claudia Wagner 0001, Martin Wöllmer, Björn W. Schuller, Louis-Philippe Morency
ICASSP5
2013 Speaker-adaptive multimodal prediction model for listener responses
abstract
The goal of this paper is to analyze and model the variability in speaking styles in dyadic interactions and build a predictive algorithm for listener responses that is able to adapt to these different styles. The end result of this research will be a virtual human able to automatically respond to a human speaker with proper listener responses (e.g., head nods). Our novel speaker-adaptive prediction model is created from a corpus of dyadic interactions where speaker variability is analyzed to identify a subset of prototypical speaker styles. During a live interaction our prediction model automatically identifies the closest prototypical speaker style and predicts listener responses based on this ``communicative style". Central to our approach is the idea of ``speaker profile" which uniquely identifies each speaker and enables the matching between prototypical speakers and new speakers. The paper shows the merits of our speaker-adaptive listener response prediction model by showing improvement over a state-of-the-art approach which does not adapt to the speaker. Besides the merits of speaker-adapta-tion, our experiments highlights the importance of using multimodal features when comparing speakers to select the closest prototypical speaker style.
Iwan de Kok, Dirk Heylen, Louis-Philippe Morency
ICMI3
2013 Automatic multimodal descriptors of rhythmic body movement
abstract
Prolonged durations of rhythmic body gestures were proved to be correlated with different types of psychological disorders. To-date, there is no automatic descriptor that can robustly detect those behaviours. In this paper, we propose a cyclic gestures descriptor that can detect and localise rhythmic body movements by taking advantage of both colour and depth modalities. We show experimentally how our rhythmic descriptor can successfully localise the rhythmic gestures as: hands fidgeting, legs fidgeting or rocking, significantly higher than the majority vote classification baseline. Our experiments also demonstrate the importance of fusing both modalities, with a significant increase in performance when compared to individual modalities.
Marwa Mahmoud, Louis-Philippe Morency, Peter Robinson 0001
ICMI2
2013 Interactive relevance search and modeling: support for expert-driven analysis of multimodal data
abstract
In this paper we present the findings of three longitudinal case studies in which a new method for conducting multimodal analysis of human behavior is tested. The focus of this new method is to engage a researcher integrally in the analysis process and allow them to guide the identification and discovery of relevant behavior instances within multimodal data. The case studies resulted in the creation of two analysis strategies: Single-Focus Hypothesis Testing and Multi-Focus Hypothesis Testing. Each were shown to be beneficial to multimodal analysis through supporting either a single focused deep analysis or analysis across multiple angles in unison. These strategies exemplified how challenging questions can be answered for multimodal datasets. The new method is described and the case studies' findings are presented detailing how the new method supports multimodal analysis and opens the door for a new breed of analysis methods. Two of the three case studies resulted in publishable results for the respective participants.
Chreston A. Miller, Francis K. H. Quek, Louis-Philippe Morency
ICMI3
2013 Who is persuasive?: the role of perceived personality and communication modality in social multimedia
abstract
Persuasive communication is part of everyone's daily life. With the emergence of social websites like YouTube, Facebook and Twitter, persuasive communication is now seen online on a daily basis. This paper explores the effect of multi-modality and perceived personality on persuasiveness of social multimedia content. The experiments are performed over a large corpus of movie review clips from Youtube which is presented to online annotators in three different modalities: only text, only audio and video. The annotators evaluated the persuasiveness of each review across different modalities and judged the personality of the speaker. Our detailed analysis confirmed several research hypotheses designed to study the relationships between persuasion, perceived personality and communicative channel, namely modality. Three hypotheses are designed: the first hypothesis studies the effect of communication modality on persuasion, the second hypothesis examines the correlation between persuasion and personality perception and finally the third hypothesis, derived from the first two hypotheses explores how communication modality influence the personality perception.
Gelareh Mohammadi, Sunghyun Park 0001, Kenji Sagae, Alessandro Vinciarelli, Louis-Philippe Morency
ICMI5
2013 ICMI 2013 grand challenge workshop on multimodal learning analytics
abstract
Advances in learning analytics are contributing new empirical findings, theories, methods, and metrics for understanding how students learn. It also contributes to improving pedagogical support for students' learning through assessment of new digital tools, teaching strategies, and curricula. Multimodal learning analytics (MMLA)[1] is an extension of learning analytics and emphasizes the analysis of natural rich modalities of communication across a variety of learning contexts. This MMLA Grand Challenge combines expertise from the learning sciences and machine learning in order to highlight the rich opportunities that exist at the intersection of these disciplines. As part of the Grand Challenge, researchers were asked to predict: (1) which student in a group was the dominant domain expert, and (2) which problems that the group worked on would be solved correctly or not. Analyses were based on a combination of speech, digital pen and video data. This paper describes the motivation for the grand challenge, the publicly available data resources and results reported by the challenge participants. The results demonstrate that multimodal prediction of the challenge goals: (1) is surprisingly reliable using rich multimodal data sources, (2) can be accomplished using any of the three modalities explored, and (3) need not be based on content analysis.
Louis-Philippe Morency, Sharon L. Oviatt, Stefan Scherer, Nadir Weibel, Marcelo Worsley
ICMI1
2013 Audiovisual behavior descriptors for depression assessment
abstract
We investigate audiovisual indicators, in particular measures of reduced emotional expressivity and psycho-motor retardation, for depression within semi-structured virtual human interviews. Based on a standard self-assessment depression scale we investigate the statistical discriminative strength of the audiovisual features on a depression/no-depression basis. Within subject-independent unimodal and multimodal classification experiments we find that early feature-level fusion yields promising results and confirms the statistical findings. We further correlate the behavior descriptors with the assessed depression severity and find considerable correlation. Lastly, a joint multimodal factor analysis reveals two prominent factors within the data that show both statistical discriminative power as well as strong linear correlation with the depression severity score. These preliminary results based on a standard factor analysis are promising and motivate us to investigate this approach further in the future, while incorporating additional modalities.
Stefan Scherer, Giota Stratou, Louis-Philippe Morency
ICMI3
2013 Learning a sparse codebook of facial and body microexpressions for emotion recognition
abstract
Obtaining a compact and discriminative representation of facial and body expressions is a difficult problem in emotion recognition. Part of the difficulty is capturing microexpressions, i.e., short, involuntary expressions that last for only a fraction of a second: at a micro-temporal scale, there are so many other subtle face and body movements that do not convey semantically meaningful information. We present a novel approach to this problem by exploiting the sparsity of the frequent micro-temporal motion patterns. Local space-time features are extracted over the face and body region for a very short time period, e.g., few milliseconds. A codebook of microexpressions is learned from the data and used to encode the features in a sparse manner. This allows us to obtain a representation that captures the most salient motion patterns of the face and body at a micro-temporal scale. Experiments performed on the AVEC 2012 dataset show our approach achieving the best published performance on the arousal dimension based solely on visual features. We also report experimental results on audio-visual emotion recognition, comparing early and late data fusion techniques.
Yale Song, Louis-Philippe Morency, Randall Davis
ICMI2
2013 A comparative study of glottal open quotient estimation techniques
abstract
The robust and efficient extraction of features related to the glottal excitation source has become increasingly important for speech technology. The glottal open quotient (OQ) is one rel-evant measurement which is known to significantly vary with changes in voice quality on a breathy to tense continuum. The extraction of OQ, however, is hampered in the time-domain by the difficulty in consistently locating the point of glottal open-ing as well the computational load of its measurement. De-termining OQ correlates in the frequency domain is an attrac-tive alternative, however the lower frequencies of glottal source spectrum are also affected by other aspects of the glottal pulse shape thereby precluding closed-form solutions and straightfor-ward mappings. The present study provides a comparison of three OQ estimation methods and shows a new method based on spectral features and artificial neural networks to outperform existing methods in terms of discrimination of voice quality, lower error values on a large volume of speech data and dra-matically reduced computation time.
John Kane 0002, Stefan Scherer, Louis-Philippe Morency, Christer Gobl
INTERSPEECH3
2013 Prediction of strategy and outcome as negotiation unfolds by using basic verbal and behavioral features
abstract
Negotiations can be characterized by the strategy participants adopt to achieve their ends (e.g., individualistic strategies are based on self-interest, cooperative strategies are used when participants try to maximize the joint gain, while competitive strategies focus on maximizing each participant’s score against the other) and the outcomes that each participant achieves in the negotiation. This paper investigates the process and the result of predicting the outcome and strategy of participants throughout the progress of the negotiation by using basic, easy to extract, linguistic and acoustic features. We evaluate our approach on a face-to-face negotiation dataset consisting of 41 dyadic interactions and show that it’s possible to significantly improve over a majority-class baseline in tasks of predicting the strategy and outcome of the interaction by analyzing only basic low level features of the negotiation.
Elnaz Nouri, Sunghyun Park 0001, Stefan Scherer, Jonathan Gratch, Peter J. Carnevale, Louis-Philippe Morency, David R. Traum
INTERSPEECH6
2013 Investigating voice quality as a speaker-independent indicator of depression and PTSD
abstract
We seek to investigate voice quality characteristics, in particular on a breathy to tense dimension, as an indicator for psychological distress, i.e. depression and post-traumatic stress disorder (PTSD), within semi-structured virtual human interviews. Our evaluation identifies significant differences between the voice quality of psychologically distressed participants and not-distressed participants within this limited corpus. We investigate the capability of automatic algorithms to classify psychologically distressed speech in speaker-independent experiments. Additionally, we examine the impact of the posed questions’ affective polarity, as motivated by findings in the literature on positive stimulus attenuation and negative stimulus potentiation in emotional reactivity of psychologically distressed participants. The experiments yield promising results using standard machine learning algorithms and solely four distinct features capturing the tenseness of the speaker’s voice.
Stefan Scherer, Giota Stratou, Jonathan Gratch, Louis-Philippe Morency
INTERSPEECH4
2013 Cicero - Towards a Multimodal Virtual Audience Platform for Public Speaking Training
Ligia Maria Batrinca, Giota Stratou, Ari Shapiro, Louis-Philippe Morency, Stefan Scherer
IVA4
2013 All Together Now - Introducing the Virtual Human Toolkit
Arno Hartholt, David R. Traum, Stacy Marsella, Ari Shapiro, Giota Stratou, Anton Leuski, Louis-Philippe Morency, Jonathan Gratch
IVA7
2013 Prediction of Visual Backchannels in the Absence of Visual Context Using Mutual Influence
Derya Ozkan, Louis-Philippe Morency
IVA2
2013 Variational Hidden Conditional Random Fields with Coupled Dirichlet Process Mixtures
Konstantinos Bousmalis, Stefanos Zafeiriou, Louis-Philippe Morency, Maja Pantic, Zoubin Ghahramani
ECML/PKDD (2)3
2013 Verbal indicators of psychological distress in interactive dialogue with a virtual human
David DeVault, Kallirroi Georgila, Ron Artstein, Fabrizio Morbini, David R. Traum, Stefan Scherer, Albert A. Rizzo, Louis-Philippe Morency
SIGDIAL Conference8
2013 Latent Mixture of Discriminative Experts
abstract
In this paper, we introduce a new model called Latent Mixture of Discriminative Experts which can automatically learn the temporal relationship between different modalities. Since, we train separate experts for each modality, LMDE is capable of improving the prediction performance even with limited amount of data. For model interpretation, we present a sparse feature ranking algorithm that exploitsL1regularization. An empirical evaluation is provided on the task of listener backchannel prediction (i.e., head nod). We introduce a new error evaluation metric called User-adaptive Prediction Accuracy that takes into account the difference in people's backchannel responses. Our results confirm the importance of combining five types of multimodal features: lexical, syntactic structure, part-of-speech, visual and prosody. Latent Mixture of Discriminative Experts model outperforms previous approaches.
Derya Ozkan, Louis-Philippe Morency
IEEE Trans. Multim.2
2013 Infinite Hidden Conditional Random Fields for Human Behavior Analysis
abstract
Hidden conditional random fields (HCRFs) are discriminative latent variable models that have been shown to successfully learn the hidden structure of a given classification problem (provided an appropriate validation of the number of hidden states). In this brief, we present the infinite HCRF (iHCRF), which is a nonparametric model based on hierarchical Dirichlet processes and is capable of automatically learning the optimal number of hidden states for a classification task. We show how we learn the model hyperparameters with an effective Markov-chain Monte Carlo sampling technique, and we explain the process that underlines our iHCRF model with the Restaurant Franchise Rating Agencies analogy. We show that the iHCRF is able to converge to a correct number of represented hidden states, and outperforms the best finite HCRFs--chosen via cross-validation--for the difficult tasks of recognizing instances of agreement, disagreement, and pain. Moreover, the iHCRF manages to achieve this performance in significantly less total training, validation, and testing time.
Konstantinos Bousmalis, Stefanos Zafeiriou, Louis-Philippe Morency, Maja Pantic
IEEE Trans. Neural Networks Learn. Syst.3
2012 Gesture-based Object Recognition using Histograms of Guiding Strokes
abstract
Humans perform iconic gestures to refer to entities through embodying their shapes. For instance, people often gesture the outline of an object (e.g. a circle for a ball) when referring to it during communication. In this paper, we present a gesture-based object recognition algorithm that enables natural human-computer interaction involving iconic gestures. Based on our analysis of multiple gesture performances, we propose a new 3D motion description of iconic gestures, called Histograms of Guiding Strokes (HoGS), which successfully summarizes hand dynamic during gestures. Our gesture-based object recognition algorithm compares favorably to human judgment performance and outper-forms most conventional gesture recognition approaches. 1
Amir Sadeghipour, Louis-Philippe Morency, Stefan Kopp
BMVC2
2012 3D Constrained Local Model for rigid and non-rigid facial tracking
abstract
We present 3D Constrained Local Model (CLM-Z) for robust facial feature tracking under varying pose. Our approach integrates both depth and intensity information in a common framework. We show the benefit of our CLM-Z method in both accuracy and convergence rates over regular CLM formulation through experiments on publicly available datasets. Additionally, we demonstrate a way to combine a rigid head pose tracker with CLM-Z that benefits rigid head tracking. We show better performance than the current state-of-the-art approaches in head pose tracking with our extension of the generalised adaptive view-based appearance model (GAVAM).
Tadas Baltrusaitis, Peter Robinson 0001, Louis-Philippe Morency
CVPR3
2012 Multi-view latent variable discriminative models for action recognition
abstract
Many human action recognition tasks involve data that can be factorized into multiple views such as body postures and hand shapes. These views often interact with each other over time, providing important cues to understanding the action. We present multi-view latent variable discriminative models that jointly learn both view-shared and view-specific sub-structures to capture the interaction between views. Knowledge about the underlying structure of the data is formulated as a multi-chain structured latent conditional model, explicitly learning the interaction between multiple views using disjoint sets of hidden variables in a discriminative manner. The chains are tied using a predetermined topology that repeats over time. We present three topologies - linked, coupled, and linked-coupled - that differ in the type of interaction between views that they model. We evaluate our approach on both segmented and unsegmented human action recognition tasks, using the ArmGesture, the NATOPS, and the ArmGesture-Continuous data. Experimental results show that our approach outperforms previous state-of-the-art action recognition models.
Yale Song, Louis-Philippe Morency, Randall Davis
CVPR2
2012 Towards sensing the influence of visual narratives on human affect
abstract
In this paper, we explore a multimodal approach to sensing affective state during exposure to visual narratives. Using four different modalities, consisting of visual facial behaviors, thermal imaging, heart rate measurements, and verbal descriptions, we show that we can effectively predict changes in human affect. Our experiments show that these modalities complement each other, and illustrate the role played by each of the four modalities in detecting human affect.
Mihai Burzo, Daniel McDuff, Rada Mihalcea, Louis-Philippe Morency, Alexis Narvaez, Verónica Pérez-Rosas
ICMI4
2012 Structural and temporal inference search (STIS): pattern identification in multimodal data
abstract
There are a multitude of annotated behavior corpora (manual and automatic annotations) available as research expands in multimodal analysis of human behavior. Despite the rich representations within these datasets, search strategies are limited with respect to the advanced representations and complex structures describing human interaction sequences. The relationships amongst human interactions are structural in nature. Hence, we present Structural and Temporal Inference Search (STIS) to support search for relevant patterns within a multimodal corpus based on the structural and temporal nature of human interactions. The user defines the structure of a behavior of interest driving a search focused on the characteristics of the structure. Occurrences of the structure are returned. We compare against two pattern mining algorithms purposed for pattern identification amongst sequences of symbolic data (e.g., sequence of events such as behavior interactions). The results are promising as STIS performs well with several datasets.
Chreston A. Miller, Louis-Philippe Morency, Francis K. H. Quek
ICMI2
2012 Step-wise emotion recognition using concatenated-HMM
abstract
Human emotion is an important part of human-human communication, since the emotional state of an individual often affects the way that he/she reacts to others. In this paper, we present a method based on concatenated Hidden Markov Model (co-HMM) to infer the dimensional and continuous emotion labels from audio-visual cues. Our method is based on the assumption that continuous emotion levels can be modeled by a set of discrete values. Based on this, we represent each emotional dimension by step-wise label classes, and learn the intrinsic and extrinsic dynamics using our co-HMM model. We evaluate our approach on the Audio-Visual Emotion Challenge (AVEC 2012) dataset. Our results show considerable improvement over the baseline regression model presented with the AVEC 2012.
Derya Ozkan, Stefan Scherer, Louis-Philippe Morency
ICMI3
2012 I already know your answer: using nonverbal behaviors to predict immediate outcomes in a dyadic negotiation
abstract
Be it in our workplace or with our family or friends, negotiation comprises a fundamental fabric of our everyday life, and it is apparent that a system that can automatically predict negotiation outcomes will have substantial implications. In this paper, we focus on finding nonverbal behaviors that are predictive of immediate outcomes (acceptances or rejections of proposals) in a dyadic negotiation. Looking at the nonverbal behaviors of the respondent alone would be inadequate since ample predictive information could also reside in the behaviors of the proposer, as well as the past history between the two parties. With this intuition in mind, we show that a more accurate prediction can be achieved by considering all the three sources (multimodal) of information together. We evaluate our approach on a face-to-face negotiation dataset consisting of 42 dyadic interactions and show that integrating all three sources of information outperforms each individual predictor.
Sunghyun Park 0001, Jonathan Gratch, Louis-Philippe Morency
ICMI3
2012 1st international workshop on multimodal learning analytics: extended abstract
abstract
This summary describes the 1st International Workshop on Multimodal Learning Analytics. This area of study brings together the technologies of multimodal analysis with the learning sciences. The intersection of these domains should enable researchers to foster an improved understanding of student learning, lead to the creation of more natural and enriching learning interfaces, and motivate the development of novel techniques for tackling challenges that are specific of education.
Stefan Scherer, Marcelo Worsley, Louis-Philippe Morency
ICMI3
2012 Multimodal human behavior analysis: learning correlation and interaction across modalities
abstract
Multimodal human behavior analysis is a challenging task due to the presence of complex nonlinear correlations and interactions across modalities. We present a novel approach to this problem based on Kernel Canonical Correlation Analysis (KCCA) and Multi-view Hidden Conditional Random Fields (MV-HCRF). Our approach uses a nonlinear kernel to map multimodal data to a high-dimensional feature space and finds a new projection of the data that maximizes the correlation across modalities. We use a multi-chain structured graphical model with disjoint sets of latent variables, one set per modality, to jointly learn both view-shared and view-specific sub-structures of the projected data, capturing interaction across modalities explicitly. We evaluate our approach on a task of agreement and disagreement recognition from nonverbal audio-visual cues using the Canal 9 dataset. Experimental results show that KCCA makes capturing nonlinear hidden dynamics easier and MV-HCRF helps learning interaction across modalities.
Yale Song, Louis-Philippe Morency, Randall Davis
ICMI2
2012 Perception Markup Language: Towards a Standardized Representation of Perceived Nonverbal Behaviors
Stefan Scherer, Stacy Marsella, Giota Stratou, Yuyu Xu, Fabrizio Morbini, Alesia Egan, Albert A. Rizzo, Louis-Philippe Morency
IVA8
2012 Dialogue Act Recognition using Reweighted Speaker Adaptation
Congkai Sun, Louis-Philippe Morency
SIGDIAL Conference2
2012 Exploring the effect of illumination on automatic expression recognition using the ICT-3DRFE database
Giota Stratou, Abhijeet Ghosh, Paul E. Debevec, Louis-Philippe Morency
Image Vis. Comput.4
2012 Introduction to the special issue on affective interaction in natural environments
abstract
Affect-sensitive systems such as social robots and virtual agents are increasingly being investigated in real-world settings. In order to work effectively in natural environments, these systems require the ability to infer the affective and mental states of humans and to provide appropriate timely output that helps to sustain long-term interactions. This special issue, which appears in two parts, includes articles on the design of socio-emotional behaviors and expressions in robots and virtual agents and on computational approaches for the automatic recognition of social signals and affective states.
Ginevra Castellano, Laurel D. Riek, Christopher Peters 0001, Kostas Karpouzis, Jean-Claude Martin, Louis-Philippe Morency
ACM Trans. Interact. Intell. Syst.6
2011 Machine Learning for Affective Computing
Mohammed E. Hoque 0001, Daniel McDuff, Louis-Philippe Morency, Rosalind W. Picard
ACII (2)3
2011 Are You Friendly or Just Polite? - Analysis of Smiles in Spontaneous Face-to-Face Interactions
Mohammed E. Hoque 0001, Louis-Philippe Morency, Rosalind W. Picard
ACII (1)2
2011 Modeling Latent Discriminative Dynamic of Multi-dimensional Affective Signals
Geovany A. Ramírez, Tadas Baltrusaitis, Louis-Philippe Morency
ACII (2)3
2011 Modeling hidden dynamics of multimodal cues for spontaneous agreement and disagreement recognition
abstract
This paper attempts to recognize spontaneous agreement and disagreement based only on nonverbal multi-modal cues. Related work has mainly used verbal and prosodic cues. We demonstrate that it is possible to correctly recognize agreement and disagreement without the use of verbal context (i.e. words, syntax). We propose to explicitly model the complex hidden dynamics of the multimodal cues using a sequential discriminative model, the Hidden Conditional Random Field (HCRF). In this paper, we show that the HCRF model is able to capture what makes each of these social attitudes unique. We present an efficient technique to analyze the concepts learned by the HCRF model and show that these coincide with the findings from social psychology regarding which cues are most prevalent in agreement and disagreement. Our experiments are performed on a spontaneous dataset of real televised debates. The HCRF model outperforms conventional approaches such as Hidden Markov Models and Support Vector Machines.
Konstantinos Bousmalis, Louis-Philippe Morency, Maja Pantic
FG2
2011 Effect of illumination on automatic expression recognition: A novel 3D relightable facial database
abstract
One of the main challenges in facial expression recognition is illumination invariance. Our long-term goal is to develop a system for automatic facial expression recognition that is robust to light variations. In this paper, we introduce a novel 3D Relightable Facial Expression (ICT-3DRFE) database that enables experimentation in the fields of both computer graphics and computer vision. The database contains 3D models for 23 subjects and 15 expressions, as well as photometric information that allow for photorealistic rendering. It is also facial action units annotated, using FACS standards. Using the ICT-3DRFE database we create an image set of different expressions/illuminations to study the effect of illumination on automatic expression recognition. We compared the output scores from automatic recognition with expert FACS annotations and found that they agree when the illumination is uniform. Our results show that the output distribution of the automatic recognition can change significantly with light variations and sometimes causes the discrimination of two different expressions to be diminished. We propose a ratio-based light transfer method, to factor out unwanted illuminations from given images and show that it reduces the effect of illumination on expression recognition.
Giota Stratou, Abhijeet Ghosh, Paul E. Debevec, Louis-Philippe Morency
FG4
2011 Towards multimodal sentiment analysis: harvesting opinions from the web
abstract
With more than 10,000 new videos posted online every day on social websites such as YouTube and Facebook, the internet is becoming an almost infinite source of information. One crucial challenge for the coming decade is to be able to harvest relevant information from this constant flow of multimodal data. This paper addresses the task of multimodal sentiment analysis, and conducts proof-of-concept experiments that demonstrate that a joint model that integrates visual, audio, and textual features can be effectively used to identify sentiment in Web videos. This paper makes three important contributions. First, it addresses for the first time the task of tri-modal sentiment analysis, and shows that it is a feasible task that can benefit from the joint exploitation of visual, audio and textual modalities. Second, it identifies a subset of audio-visual features relevant to sentiment analysis and present guidelines on how to integrate these features. Finally, it introduces a new dataset consisting of real online data, which will be useful for future research in this area.
Louis-Philippe Morency, Rada Mihalcea, Payal Doshi
ICMI1
2011 Virtual Rapport 2.0
Lixing Huang, Louis-Philippe Morency, Jonathan Gratch
IVA2
2011 Modeling Nonverbal Behavior of a Virtual Counselor during Intimate Self-disclosure
Sin-Hwa Kang, Candace L. Sidner, Jonathan Gratch, Ron Artstein, Lixing Huang, Louis-Philippe Morency
IVA6
2010 Latent Mixture of Discriminative Experts for Multimodal Prediction Modeling
Derya Ozkan, Kenji Sagae, Louis-Philippe Morency
COLING3
2010 Learning Backchannel Prediction Model from Parasocial Consensus Sampling: A Subjective Evaluation
Lixing Huang, Louis-Philippe Morency, Jonathan Gratch
IVA2
2010 3rd international workshop on affective interaction in natural environments (AFFINE)
abstract
The 3rd International Workshop on Affective Interaction in Natural Environments, AFFINE, follows a number of successful AFFINE workshops and events commencing in 2008.A key aim of AFFINE is the identification and investigation of significant open issues in real-time, affect-aware applications 'in the wild' and especially in embodied interaction, for example, with robots or virtual agents. AFFINE seeks to bring together researchers working on the real-time interpretation of user behaviour with those who are concerned with social robot and virtual agent interaction frameworks.
Ginevra Castellano, Kostas Karpouzis, Jean-Claude Martin, Louis-Philippe Morency, Christopher Peters 0001, Laurel D. Riek
ACM Multimedia4
2010 A probabilistic multimodal approach for predicting listener backchannels
Louis-Philippe Morency, Iwan de Kok, Jonathan Gratch
Auton. Agents Multi Agent Syst.1
2010 Monocular head pose estimation using generalized adaptive view-based appearance model
Louis-Philippe Morency, Jacob Whitehill, Javier R. Movellan
Image Vis. Comput.1
2008 Modeling Latent-Dynamic in Shallow Parsing: A Latent Conditional Model with Imrpoved Inference
Xu Sun 0001, Louis-Philippe Morency, Daisuke Okanohara, Yoshimasa Tsuruoka, Jun'ichi Tsujii
COLING2
2008 Generalized adaptive view-based appearance model: Integrated framework for monocular head pose estimation
abstract
Accurately estimating the person's head position and orientation is an important task for a wide range of applications such as driver awareness and human-robot interaction. Over the past two decades, many approaches have been suggested to solve this problem, each with its own advantages and disadvantages. In this paper, we present a probabilistic framework called generalized adaptive viewbased appearance model (GAVAM) which integrates the advantages from three of these approaches: (1) the automatic initialization and stability of static head pose estimation, (2) the relative precision and user-independence of differential registration, and (3) the robustness and bounded drift of keyframe tracking. In our experiments, we show how the GAVAM model can be used to estimate head position and orientation in real-time using a simple monocular camera. Our experiments on two previously published datasets show that the GAVAM framework can accurately track for a long period of time (>2 minutes) with an average accuracy of 3.5deg and 0.75 in with an inertial sensor and a 3D magnetic sensor.
Louis-Philippe Morency, Jacob Whitehill, Javier R. Movellan
FG1
2008 Context-based recognition during human interactions: automatic feature selection and encoding dictionary
abstract
During face-to-face conversation, people use visual feedback such as head nods to communicate relevant information and to synchronize rhythm between participants. In this paper we describe how contextual information from other participants can be used to predict visual feedback and improve recognition of head gestures in human-human interactions. For example, in a dyadic interaction, the speaker contextual cues such as gaze shifts or changes in prosody will influence listener backchannel feedback (e.g., head nod). To automatically learn how to integrate this contextual information into the listener gesture recognition framework, this paper addresses two main challenges: optimal feature representation using an encoding dictionary and automatic selection of optimal feature-encoding pairs. Multimodal integration between context and visual observations is performed using a discriminative sequential model (Latent-Dynamic Conditional Random Fields) trained on previous interactions. In our experiments involving 38 storytelling dyads, our context-based recognizer significantly improved head gesture recognition performance over a vision-only recognizer.
Louis-Philippe Morency, Iwan de Kok, Jonathan Gratch
ICMI1
2008 Predicting Listener Backchannels: A Probabilistic Multimodal Approach
Louis-Philippe Morency, Iwan de Kok, Jonathan Gratch
IVA1
2008 Reducing drift in differential tracking
Louis-Philippe Morency, Trevor Darrell
Comput. Vis. Image Underst.2
2007 Latent-Dynamic Discriminative Models for Continuous Gesture Recognition
abstract
Many problems in vision involve the prediction of a class label for each frame in an unsegmented sequence. In this paper, we develop a discriminative framework for simultaneous sequence segmentation and labeling which can capture both intrinsic and extrinsic class dynamics. Our approach incorporates hidden state variables which model the sub-structure of a class sequence and learn dynamics between class labels. Each class label has a disjoint set of associated hidden states, which enables efficient training and inference in our model. We evaluated our method on the task of recognizing human gestures from unsegmented video streams and performed experiments on three different datasets of head and eye gestures. Our results demonstrate that our model compares favorably to Support Vector Machines, Hidden Markov Models, and Conditional Random Fields on visual gesture recognition tasks.
Louis-Philippe Morency, Ariadna Quattoni, Trevor Darrell
CVPR1
2007 Head gestures for perceptual interfaces: The role of context in improving recognition
Louis-Philippe Morency, Candace L. Sidner, Christopher Lee 0001, Trevor Darrell
Artif. Intell.1
2007 Hidden Conditional Random Fields
abstract
We present a discriminative latent variable model for classification problems in structured domains where inputs can be represented by a graph of local observations. A hidden-state Conditional Random Field framework learns a set of latent variables conditioned on local features. Observations need not be independent and may overlap in space and time.
Ariadna Quattoni, Sy Bor Wang, Louis-Philippe Morency, Michael Collins 0001, Trevor Darrell
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 The Role of Context in Head Gesture Recognition
Louis-Philippe Morency, Candace L. Sidner, Christopher Lee 0001, Trevor Darrell
AAAI1
2006 Hidden Conditional Random Fields for Gesture Recognition
abstract
We introduce a discriminative hidden-state approach for the recognition of human gestures. Gesture sequences often have a complex underlying structure, and models that can incorporate hidden structures have proven to be advantageous for recognition tasks. Most existing approaches to gesture recognition with hidden states employ a Hidden Markov Model or suitable variant (e.g., a factored or coupled state model) to model gesture streams; a significant limitation of these models is the requirement of conditional independence of observations. In addition, hidden states in a generative model are selected to maximize the likelihood of generating all the examples of a given gesture class, which is not necessarily optimal for discriminating the gesture class against other gestures. Previous discriminative approaches to gesture sequence recognition have shown promising results, but have not incorporated hidden states nor addressed the problem of predicting the label of an entire sequence. In this paper, we derive a discriminative sequence model with a hidden state structure, and demonstrate its utility both in a detection and in a multi-way classification formulation. We evaluate our method on the task of recognizing human arm and head gestures, and compare the performance of our method to both generative hidden state and discriminative fully-observable models.
Sy Bor Wang, Ariadna Quattoni, Louis-Philippe Morency, David Demirdjian, Trevor Darrell
CVPR (2)3
2006 The effect of head-nod recognition in human-robot conversation
abstract
This paper reports on a study of human participants with a robot designed to participate in a collaborative conversation with a human. The purpose of the study was to investigate a particular kind of gestural feedback from human to the robot in these conversations: head nods. During these conversations, the robot recognized head nods from the human participant. The conversations between human and robot concern demonstrations of inventions created in a lab. We briefly discuss the robot hardware and architecture and then focus the paper on a study of the effects of understanding head nods in three different conditions. We conclude that conversation itself triggers head nods by people in human-robot conversations and that telling participants that the robot recognizes their nods as well as having the robot provide gestural feedback of its nod recognition is effective in producing more nods.
Candace L. Sidner, Christopher Lee 0001, Louis-Philippe Morency, Clifton Forlines
HRI3
2006 Co-Adaptation of audio-visual speech and gesture classifiers
abstract
The construction of robust multimodal interfaces often requires large amounts of labeled training data to account for cross-user differences and variation in the environment. In this work, we investigate whether unlabeled training data can be leveraged to build more reliable audio-visual classifiers through co-training, a multi-view learning algorithm. Multimodal tasks are good candidates for multi-view learning, since each modality provides a potentially redundant view to the learning algorithm. We apply co-training to two problems: audio-visual speech unit classification, and user agreement recognition using spoken utterances and head gestures. We demonstrate that multimodal co-training can be used to learn from only a few labeled examples in one or both of the audio-visual modalities. We also propose a co-adaptation algorithm, which adapts existing audio-visual classifiers to a particular user or noise condition by leveraging the redundancy in the unlabeled data.
C. Mario Christoudias, Kate Saenko, Louis-Philippe Morency, Trevor Darrell
ICMI3
2006 Recognizing gaze aversion gestures in embodied conversational discourse
abstract
Eye gaze offers several key cues regarding conversational discourse during face-to-face interaction between people. While a large body of research results exist to document the use of gaze in human-to-human interaction, and in animating realistic embodied avatars, recognition of conversational eye gestures - distinct eye movement patterns relevant to discourse - has received less attention. We analyze eye gestures during interaction with an animated embodied agent and propose a non-intrusive vision-based approach to estimate eye gaze and recognize eye gestures. In our user study, human participants avert their gaze (i.e. with "look-away" or "thinking" gestures) during periods of cognitive load. Using our approach, an agent can visually differentiate whether a user is thinking about a response or is waiting for the agent or robot to take its turn.
Louis-Philippe Morency, C. Mario Christoudias, Trevor Darrell
ICMI1
2006 Head gesture recognition in intelligent interfaces: the role of context in improving recognition
abstract
Acknowledging an interruption with a nod of the head is a natural and intuitive communication gesture which can be performed without significantly disturbing a primary interface activity. In this paper we describe vision-based head gesture recognition techniques and their use for common user interface commands. We explore two prototype perceptual interface components which use detected head gestures for dialog box confirmation and document browsing, respectively. Tracking is performed using stereo-based alignment, and recognition proceeds using a trained discriminative classifier. An additional context learning component is described, which exploits interface context to obtain robust performance. User studies with prototype recognition components indicate quantitative and qualitative benefits of gesture-based confirmation over conventional alternatives.
Louis-Philippe Morency, Trevor Darrell
IUI1
2006 Virtual Rapport
Jonathan Gratch, Anya Okhmatovskaia, Francois Lamothe, Stacy Marsella, Mathieu Morales, Rick J. van der Werf, Louis-Philippe Morency
IVA7
2006 Non-parametric and light-field deformable models
C. Mario Christoudias, Louis-Philippe Morency, Trevor Darrell
Comput. Vis. Image Underst.2
2005 Contextual recognition of head gestures
abstract
Head pose and gesture offer several key conversational grounding cues and are used extensively in face-to-face interaction among people. We investigate how dialog context from an embodied conversational agent (ECA) can improve visual recognition of user gestures. We present a recognition framework which (1) extracts contextual features from an ECA's dialog manager, (2) computes a prediction of head nod and head shakes, and (3) integrates the contextual predictions with the visual observation of a vision-based head gesture recognizer. We found a subset of lexical, punctuation and timing features that are easily available in most ECA architectures and can be used to learn how to predict user feedback. Using a discriminative approach to contextual prediction and multi-modal integration, we were able to improve the performance of head gesture detection even when the topic of the test set was significantly different than the training set.
Louis-Philippe Morency, Candace L. Sidner, Christopher Lee 0001, Trevor Darrell
ICMI1
2004 Light Field Appearance Manifolds
C. Mario Christoudias, Louis-Philippe Morency, Trevor Darrell
ECCV (4)2
2004 From conversational tooltips to grounded discourse: head poseTracking in interactive dialog systems
abstract
Head pose and gesture offer several key conversational grounding cues and are used extensively in face-to-face interaction among people. While the machine interpretation of these cues has previously been limited to output modalities, recent advances in face-pose tracking allow for systems which are robust and accurate enough to sense natural grounding gestures. We present the design of a module that detects these cues and show examples of its integration in three different conversational agents with varying degrees of discourse model complexity. Using a scripted discourse model and off-the-shelf animation and speech-recognition components, we demonstrate the use of this module in a novel "conversational tooltip" task, where additional information is spontaneously provided by an animated character when users attendto various physical objects or characters in the environment. We further describe the integration of our module in two systems where animated and robotic characters interact with users based on rich discourse and semantic models.
Louis-Philippe Morency, Trevor Darrell
ICMI1
2003 Adaptive View-Based Appearance Models
abstract
We present a method for online rigid object tracking using an adaptive view-based appearance model. When the object's pose trajectory crosses itself, our tracker has bounded drift and can track objects undergoing large motion for long periods of time. Our tracker registers each incoming frame against the views of the appearance model using a two-frame registration algorithm. Using a linear Gaussian filter, we simultaneously estimate the pose of the object and adjust the view-based model as pose-changes are recovered from the registration algorithm. The adaptive view-based model is populated online with views of the object as it undergoes different orientations in pose space, allowing us to capture non-Lambertian effects. We tested our approach on a real-time rigid object tracking task using stereo cameras and observed an RMS error within the accuracy limit of an attached inertial sensor.
Louis-Philippe Morency, Trevor Darrell
CVPR (1)1
2003 Robust real-time egomotion from stereo images
abstract
In this paper, we present a novel technique for estimating large camera displacement using stereo images. The relative transformation between two stereo image pairs is estimated using a hybrid registration algorithm which combines the robustness of multi-scale feature tracking for large movements and the accuracy of 3D normal flow constraints. Our hybrid technique takes advantage of depth information available from the stereo camera which makes it less sensitive to lighting variations. We tested the accuracy of our hybrid algorithm on real stereo sequences and showed that our technique handles displacements up to 150 cm and rotations up to 20 degrees between images. Our algorithm runs at 6 Hz on a Pentium 4 1.7GHz.
Louis-Philippe Morency, Rakesh Gupta 0001
ICIP (2)1
2003 A multi-modal approach for determining speaker location and focus
abstract
This paper presents a multi-modal approach to locate a speaker in a scene and determine to whom he or she is speaking. We present a simple probabilistic framework that combines multiple cues derived from both audio and video information. A purely visual cue is obtained using a head tracker to identify possible speakers in a scene and provide both their 3-D positions and orientation. In addition, estimates of the audio signal's direction of arrival are obtained with the help of a two-element microphone array. A third cue measures the association between the audio and the tracked regions in the video. Integrating these cues provides a more robust solution than using any single cue alone. The usefulness of our approach is shown in our results for video sequences with two or more people in a prototype interactive kiosk environment.
Michael Siracusa, Louis-Philippe Morency, Kevin W. Wilson, John W. Fisher III, Trevor Darrell
ICMI2
2002 Face-Responsive Interfaces: From Direct Manipulation to Perceptive Presence
Trevor Darrell, Konrad Tollmar, Frank Bentley, Neal Checka, Louis-Philippe Morency, Alice Oh
UbiComp5
2001 Reducing Drift in Parametric Motion Tracking
Louis-Philippe Morency, Trevor Darrell
ICCV2