Jeffrey M. Girard

dblp:00/10781 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0002-7359-3746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2026 LLA MADRS: Evaluating Open-Source LLMs on Real Clinical Interviews - To Reason or Not to Reason?
abstract
Abstract Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored. We present LLAMADRS, a benchmark for structured clinical assessment from dialogue built on the CAMI corpus of psychiatric interviews, comprising 5,804 expert annotations across 541 sessions. We evaluate 25 open-source models (standard and reasoning-augmented; 0.6B–400B parameters) and generate over 400,000 predictions. Our results demonstrate that strong open-source LLMs achieve item-level accuracy with residual error below clinically substantial thresholds. Additionally, an Item-then-Sum (ITS) strategy, assessing symptoms individually through discrete LLM calls before synthesizing final scores, significantly reduces error relative to Direct Total Score (DTS) prediction across most model architectures and scales, despite reasoning models attempting similar decomposition in the reasoning traces of their DTS predictions. In fact, we find that performance gains attributed to “reasoning” depend fundamentally on prompt design: standard models equipped with structured task definitions and examples match reasoning-augmented counterparts. Among the latter, longer reasoning traces correlate with reduced error; while higher model scale does across both architectures. Our results clarify when and why reasoning helps and offer actionable guidance for deploying LLMs in semi-structured clinical assessment.
Gaoussou Youssouf Kebe, Jeffrey M. Girard, Einat Liebenthal, Justin T. Baker, Fernando De la Torre, Louis-Philippe Morency
Trans. Assoc. Comput. Linguistics2
2024 GeSTICS: A Multimodal Corpus for Studying Gesture Synthesis in Two-party Interactions with Contextualized Speech
abstract
Generating natural co-speech gestures and facial expressions for effective human-agent interactions requires modeling the intricate interplay between verbal, non-verbal, and contextual cues observed in dyadic human communication. Two types of contextual cues are of particular interest: (1) individual factors of the interlocutors, such as their demographic attributes, and (2) situational factors, like the outcome of a preceding event. To facilitate their study, we introduce the GeSTICS Dataset, a novel multimodal corpus comprising 9,853 questions and 10,460 answers from audiovisual recordings of post-game sports interviews by 147 interviewees. The dataset contains speech data, including textual transcriptions, lexical descriptors, and acoustic features, as well as visual data encompassing the interviewee’s body pose and facial expressions, with an emphasis on capturing these modalities during both the question-listening and answering phases of the interview. Furthermore, GeSTICS incorporates metadata about individual factors, such as the age and cultural background of the interviewees, and situational factors, like the results of the games, which are often overlooked in existing multimodal datasets. Our preliminary analysis of GeSTICS reveals that the effects of speech features, such as loudness and lexical choice, on the production of co-speech gestures in both speaking and listening phases are moderated by situational factors and the interviewee’s individual factors. GeSTICS is designed to enhance the generation of realistic nonverbal behaviors in virtual agents, animated characters, and human-robot interaction systems, thus contributing to more engaging and effective human-agent communication. The analysis code and the dataset are available at https://gestics.github.io.
Gaoussou Youssouf Kebe, Mehmet Deniz Birlikci, Auriane Boudin, Ryo Ishii, Jeffrey M. Girard, Louis-Philippe Morency
IVA5
2023 DynAMoS: The Dynamic Affective Movie Clip Database for Subjectivity Analysis
abstract
In this paper, we describe the design, collection, and validation of a new video database that includes holistic and dynamic emotion ratings from 83 participants watching 22 affective movie clips. In contrast to previous work in Affective Computing, which pursued a single "ground truth" label for the affective content of each moment of each video (e.g., by averaging the ratings of 2 to 7 trained participants), we embrace the subjectivity inherent to emotional experiences and provide the full distribution of all participants' ratings (with an average of 76.7 raters per video). We argue that this choice represents a paradigm shift with the potential to unlock new research directions, generate new hypotheses, and inspire novel methods in the Affective Computing community. We also describe several interdisciplinary use cases for the database: to provide dynamic norms for emotion elicitation studies (e.g., in psychology, medicine, and neuroscience), to train and test affective content analysis algorithms (e.g., for dynamic emotion recognition, video summarization, and movie recommendation), and to study subjectivity in emotional reactions (e.g., to identify moments of emotional ambiguity or ambivalence within movies, identify predictors of subjectivity, and develop personalized affective content analysis algorithms). The database is made freely available to researchers for noncommercial use at https://dynamos.mgb.org.
Jeffrey M. Girard, Yanmei Tie, Einat Liebenthal
ACII1
2022 Toward Causal Understanding of Therapist-Client Relationships: A Study of Language Modality and Social Entrainment
abstract
The relationship between a therapist and their client is one of the most critical determinants of successful therapy. The working alliance is a multifaceted concept capturing the collaborative aspect of the therapist-client relationship; a strong working alliance has been extensively linked to many positive therapeutic outcomes. Although therapy sessions are decidedly multimodal interactions, the language modality is of particular interest given its recognized relationship to similar dyadic concepts such as rapport, cooperation, and affiliation. Specifically, in this work we study language entrainment, which measures how much the therapist and client adapt toward each other’s use of language over time. Despite the growing body of work in this area, however, relatively few studies examine causal relationships between human behavior and these relationship metrics: does an individual’s perception of their partner affect how they speak, or does how they speak affect their perception? We explore these questions in this work through the use of structural equation modeling (SEM) techniques, which allow for both multilevel and temporal modeling of the relationship between the quality of the therapist-client working alliance and the participants’ language entrainment. In our first experiment, we demonstrate that these techniques perform well in comparison to other common machine learning models, with the added benefits of interpretability and causal analysis. In our second analysis, we interpret the learned models to examine the relationship between working alliance and language entrainment and address our exploratory research questions. The results reveal that a therapist’s language entrainment can have a significant impact on the client’s perception of the working alliance, and that the client’s language entrainment is a strong indicator of their perception of the working alliance. We discuss the implications of these results and consider several directions for future work in multimodality.
Alexandria K. Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay Fournier, Holly Swartz, Louis-Philippe Morency
ICMI2
2021 Goals, Tasks, and Bonds: Toward the Computational Assessment of Therapist Versus Client Perception of Working Alliance
abstract
Early client dropout is one of the most significant challenges facing psychotherapy: recent studies suggest that at least one in five clients will leave treatment prematurely. Clients may terminate therapy for various reasons, but one of the most common causes is the lack of a strong working alliance. The concept of working alliance captures the collaborative relationship between a client and their therapist when working toward the progress and recovery of the client seeking treatment. Unfortunately, clients are often unwilling to directly express dissatisfaction in care until they have already decided to terminate therapy. On the other side, therapists may miss subtle signs of client discontent during treatment before it is too late. In this work, we demonstrate that nonverbal behavior analysis may aid in bridging this gap. The present study focuses primarily on the head gestures of both the client and therapist, contextualized within conversational turn-taking actions between the pair during psychotherapy sessions. We identify multiple behavior patterns suggestive of an individual's perspective on the working alliance; interestingly, these patterns also differ between the client and the therapist. These patterns inform the development of predictive models for self-reported ratings of working alliance, which demonstrate significant predictive power for both client and therapist ratings. Future applications of such models may stimulate preemptive intervention to strengthen a weak working alliance, whether explicitly attempting to repair the existing alliance or establishing a more suitable client-therapist pairing, to ensure that clients encounter fewer barriers to receiving the treatment they need.
Alexandria K. Vail, Jeffrey M. Girard, Lauren M. Bylsma, Jeffrey F. Cohn, Jay Fournier, Holly Swartz, Louis-Philippe Morency
FG2
2021 To Rate or Not To Rate: Investigating Evaluation Methods for Generated Co-Speech Gestures
abstract
While automatic performance metrics are crucial for machine learning of artificial human-like behaviour, the gold standard for evaluation remains human judgement. The subjective evaluation of artificial human-like behaviour in embodied conversational agents is however expensive and little is known about the quality of the data it returns. Two approaches to subjective evaluation can be largely distinguished, one relying on ratings, the other on pairwise comparisons. In this study we use co-speech gestures to compare the two against each other and answer questions about their appropriateness for evaluation of artificial behaviour. We consider their ability to rate quality, but also aspects pertaining to the effort of use and the time required to collect subjective data. We use crowd sourcing to rate the quality of co-speech gestures in avatars, assessing which method picks up more detail in subjective assessments. We compared gestures generated by three different machine learning models with various level of behavioural quality. We found that both approaches were able to rank the videos according to quality and that the ranking significantly correlated, showing that in terms of quality there is no preference of one method over the other. We also found that pairwise comparisons were slightly faster and came with improved inter-rater reliability, suggesting that for small-scale studies pairwise comparisons are to be favoured over ratings.
Pieter Wolfert, Jeffrey M. Girard, Taras Kucherenko, Tony Belpaeme
ICMI2
2020 Toward Multimodal Modeling of Emotional Expressiveness
abstract
Emotional expressiveness captures the extent to which a person tends to outwardly display their emotions through behavior. Due to the close relationship between emotional expressiveness and behavioral health, as well as the crucial role that it plays in social interaction, the ability to automatically predict emotional expressiveness stands to spur advances in science, medicine, and industry. In this paper, we explore three related research questions. First, how well can emotional expressiveness be predicted from visual, linguistic, and multimodal behavioral signals? Second, how important is each behavioral modality to the prediction of emotional expressiveness? Third, which behavioral signals are reliably related to emotional expressiveness? To answer these questions, we add highly reliable transcripts and human ratings of perceived emotional expressiveness to an existing video database and use this data to train, validate, and test predictive models. Our best model shows promising predictive performance on this dataset (RMSE=0.65, R^2=0.45, r=0.74). Multimodal models tend to perform best overall, and models trained on the linguistic modality tend to outperform models trained on the visual modality. Finally, examination of our interpretable models' coefficients reveals a number of visual and linguistic behavioral signals---such as facial action unit intensity, overall word count, and use of words related to social processes---that reliably predict emotional expressiveness.
Victoria Lin 0001, Jeffrey M. Girard, Michael A. Sayette, Louis-Philippe Morency
ICMI2
2020 Depression Severity Assessment for Adolescents at High Risk of Mental Disorders
abstract
Recent progress in artificial intelligence has led to the development of automatic behavioral marker recognition, such as facial and vocal expressions. Those automatic tools have enormous potential to support mental health assessment, clinical decision making, and treatment planning. In this paper, we investigate nonverbal behavioral markers of depression severity assessed during semi-structured medical interviews of adolescent patients. The main goal of our research is two-fold: studying a unique population of adolescents at high risk of mental disorders and differentiating mild depression from moderate or severe depression. We aim to explore computationally inferred facial and vocal behavioral responses elicited by three segments of the semi-structured medical interviews: Distress Assessment Questions, Ubiquitous Questions, and Concept Questions. Our experimental methodology reflects best practise used for analyzing small sample size and unbalanced datasets of unique patients. Our results show a very interesting trend with strongly discriminative behavioral markers from both acoustic and visual modalities. These promising results are likely due to the unique classification task (mild depression vs. moderate and severe depression) and three types of probing questions.
Michal Muszynski, Jamie Zelazny, Jeffrey M. Girard, Louis-Philippe Morency
ICMI3
2019 Reconsidering the Duchenne Smile: Indicator of Positive Emotion or Artifact of Smile Intensity?
abstract
The Duchenne smile hypothesis is that smiles that include eye constriction (AU6) are the product of genuine positive emotion, whereas smiles that do not are either falsified or related to negative emotion. This hypothesis has become very influential and is often used in scientific and applied settings to justify the inference that a smile is either true or false. However, empirical support for this hypothesis has been equivocal and some researchers have proposed that, rather than being a reliable indicator of positive emotion, AU6 may just be an artifact produced by intense smiles. Initial support for this proposal has been found when comparing smiles related to genuine and feigned positive emotion; however, it has not yet been examined when comparing smiles related to genuine positive and negative emotion. The current study addressed this gap in the literature by examining spontaneous smiles from 136 participants during the elicitation of amusement, embarrassment, fear, and pain (from the BP4D+ dataset). Bayesian multilevel regression models were used to quantify the associations between AU6 and self-reported amusement while controlling for smile intensity. Models were estimated to infer amusement from AU6 and to explain the intensity of AU6 using amusement. In both cases, controlling for smile intensity substantially reduced the hypothesized association, whereas the effect of smile intensity itself was quite large and reliable. These results provide further evidence that the Duchenne smile is likely an artifact of smile intensity rather than a reliable and unique indicator of genuine positive emotion.
Jeffrey M. Girard, Gayatri Shandar, Zhun Liu, Jeffrey F. Cohn, Lijun Yin 0001, Louis-Philippe Morency
ACII1
2019 Democratizing Psychological Insights from Analysis of Nonverbal Behavior
abstract
The affective computing community has invested heavily in building automated tools for the analysis of facial behavior and the expression of emotion. These tools present a valuable, but largely untapped, opportunity for social scientists to perform observational analyses of nonverbal behavior at very large scale. Various tech companies are collecting huge corpora of images and videos from around the world that could be used to study important scientific questions. However, privacy restrictions and intellectual property concerns render these data inaccessible to most academics. Unfortunately, this limits the potential for scientific advancement and leads to the consolidation of data and opportunity into the hands of a few powerful institutions. In this paper, we ask whether similar psychological insights can be gained by analyzing smaller, public datasets that are more within reach for academic researchers. As a proof-of-concept for this idea, we gather, analyze, and release a corpus of public images and metadata and use it to replicate recent psychological findings about smiling, gender, and culture. In so doing, we provide evidence that psychological insights can indeed by democratized through the automated analysis of nonverbal behavior.
Daniel McDuff, Jeffrey M. Girard
ACII2
2019 ElderReact: A Multimodal Dataset for Recognizing Emotional Response in Aging Adults
abstract
Automatic emotion recognition plays a critical role in technologies such as intelligent agents and social robots and is increasingly being deployed in applied settings such as education and healthcare. Most research to date has focused on recognizing the emotional expressions of young and middle-aged adults and, to a lesser extent, children and adolescents. Very few studies have examined automatic emotion recognition in older adults (i.e., elders), which represent a large and growing population worldwide. Given that aging causes many changes in facial shape and appearance and has been found to alter patterns of nonverbal behavior, there is strong reason to believe that automatic emotion recognition systems may need to be developed specifically (or augmented) for the elder population. To promote and support this type of research, we introduce a newly collected multimodal dataset of elders reacting to emotion elicitation stimuli. Specifically, it contains 1323 video clips of 46 unique individuals with human annotations of six discrete emotions: anger, disgust, fear, happiness, sadness, and surprise as well as valence. We present a detailed analysis of the most indicative features for each emotion. We also establish several baselines using unimodal and multimodal features on this dataset. Finally, we show that models trained on dataset of another age group do not generalize well on elders.
Kaixin Ma, Xinru Yang, Jeffrey M. Girard, Louis-Philippe Morency
ICMI5
2017 Open-Source Software for Continuous Measurement and Media Annotation
abstract
Full understanding of behavior and experience requires an appreciation of time-dependent patterns. However, traditional methods of observational measurement and self-reporting are ill-suited to capturing such patterns. These methods tend to polarize into either macro-level (gist) analyses of large swaths of time or micro-level (atomic) analyses of discrete segments. Unfortunately, both approaches miss the continuous, dynamic flow of many psychological processes.Specialized methods are needed that can capture such processes as they unfold over time and across dimensions.
Jeffrey M. Girard
FG1
2017 Sayette Group Formation Task (GFT) Spontaneous Facial Expression Database
abstract
Despite the important role that facial expressions play in interpersonal communication and our knowledge that interpersonal behavior is influenced by social context, no currently available facial expression database includes multiple interacting participants. The Sayette Group Formation Task (GFT) database addresses the need for well-annotated video of multiple participants during unscripted interactions. The database includes 172,800 video frames from 96 participants in 32 three-person groups. To aid in the development of automated facial expression analysis systems, GFT includes expert annotations of FACS occurrence and intensity, facial landmark tracking, and baseline results for linear SVM, deep learning, active patch learning, and personalized classification. Baseline performance is quantified and compared using identical partitioning and a variety of metrics (including means and confidence intervals). The highest performance scores were found for the deep learning and active patch learning methods. Learn more at http://osf.io/7wcyz.
Jeffrey M. Girard, Wen-Sheng Chu, László A. Jeni, Jeffrey F. Cohn
FG1
2017 Historical Heterogeneity Predicts Smiling: Evidence from Large-Scale Observational Analyses
abstract
Facial behavior is a valuable source of information about an individual's feelings and intentions. However, many factors combine to influence and moderate facial behavior including personality, gender, context, and culture. Due to the high cost of traditional observational methods, the relationship between culture and facial behavior is not well-understood. In the current study, we explored the sociocultural factors that influence facial behavior using large-scale observational analyses. We developed and implemented an algorithm to automatically analyze the smiling of 866,726 participants across 31 different countries. We found that participants smiled more when from a country that is higher in individualism, has a lower population density, and has a long history of immigration diversity (i.e., historical heterogeneity). Our findings provide the first evidence that historical heterogeneity predicts actual smiling behavior. Furthermore, they converge with previous findings using selfreport methods. Taken together, these findings support the theory that historical heterogeneity explains, and may even contribute to the development of, permissive cultural display rules that encourage the open expression of emotion.
Jeffrey M. Girard, Daniel McDuff
FG1
2017 FERA 2017 - Addressing Head Pose in the Third Facial Expression Recognition and Analysis Challenge
abstract
The field of Automatic Facial Expression Analysis has grown rapidly in recent years. However, despite progress in new approaches as well as benchmarking efforts, most evaluations still focus on either posed expressions, near-frontal recordings, or both. This makes it hard to tell how existing expression recognition approaches perform under conditions where faces appear in a wide range of poses (or camera views), displaying ecologically valid expressions. The main obstacle for assessing this is the availability of suitable data, and the challenge proposed here addresses this limitation. The FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) extends FERA 2015 to the estimation of Action Units occurrence and intensity under different camera views. In this paper we present the third challenge in automatic recognition of facial expressions, to be held in conjunction with the 12th IEEE conference on Face and Gesture Recognition, May 2017, in Washington, United States. Two sub-challenges are defined: the detection of AU occurrence, and the estimation of AU intensity. In this work we outline the evaluation protocol, the data used, and the results of a baseline method for both sub-challenges.
Michel F. Valstar, Enrique Sánchez-Lozano, Jeffrey F. Cohn, László A. Jeni, Jeffrey M. Girard, Zheng Zhang 0023, Lijun Yin 0001, Maja Pantic
FG5
2016 Multimodal Spontaneous Emotion Corpus for Human Behavior Analysis
abstract
Emotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test algorithms. We present a well-annotated, multimodal, multidimensional spontaneous emotion corpus of 140 participants. Emotion inductions were highly varied. Data were acquired from a variety of sensors of the face that included high-resolution 3D dynamic imaging, high-resolution 2D video, and thermal (infrared) sensing, and contact physiological sensors that included electrical conductivity of the skin, respiration, blood pressure, and heart rate. Facial expression was annotated for both the occurrence and intensity of facial action units from 2D video by experts in the Facial Action Coding System (FACS). The corpus further includes derived features from 3D, 2D, and IR (infrared) sensors and baseline results for facial expression and action unit detection. The entire corpus will be made available to the research community.
Zheng Zhang 0023, Jeffrey M. Girard, Yue Wu 0002, Xing Zhang 0012, Peng Liu 0039, Umur A. Ciftci, Shaun J. Canavan, Michael Reale, Andrew Horowitz, Huiyuan Yang, Jeffrey F. Cohn, Lijun Yin 0001
CVPR2
2015 Estimating smile intensity: A better way
Jeffrey M. Girard, Jeffrey F. Cohn, Fernando De la Torre
Pattern Recognit. Lett.1
2014 Perceptions of Interpersonal Behavior are Influenced by Gender, Facial Expression Intensity, and Head Pose
abstract
Across multiple channels, nonverbal behavior communicates information about affective states and interpersonal intentions. Researchers interested in understanding how these nonverbal messages are transmitted and interpreted have examined the relationship between behavior and ratings of interpersonal motives using dimensions such as agency and communion. However, previous work has focused on images of posed behavior and it is unclear how well these results will generalize to more dynamic representations of real-world behavior. The current study proposes to extend the current literature by examining how gender, facial expression intensity, and head pose influence interpersonal ratings in videos of spontaneous nonverbal behavior.
Jeffrey M. Girard
ICMI1
2014 Nonverbal social withdrawal in depression: Evidence from manual and automatic analyses
Jeffrey M. Girard, Jeffrey F. Cohn, Mohammad H. Mahoor, Seyed Mohammad Mavadati, Zakia Hammal, Dean P. Rosenwald
Image Vis. Comput.1
2014 BP4D-Spontaneous: a high-resolution spontaneous 3D dynamic facial expression database
Xing Zhang 0012, Lijun Yin 0001, Jeffrey F. Cohn, Shaun J. Canavan, Michael Reale, Andy Horowitz, Peng Liu 0039, Jeffrey M. Girard
Image Vis. Comput.8