Mohammad Soleymani 0001

dblp:s/MohammadSoleymani · DBLP profile ↗
← Back
77ranked-venue papers
23as first author
34since 2021 · last 2026
0000-0002-5873-1434ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 13 first-author · 21 since 2021Artificial intelligence and machine learning · 29 · 8 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 20 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic Data
abstract
Facial expression analysis is central to social AI and human-computer interaction. However, existing toolkits often struggle to generalize across diverse demographics, largely due to the limited diversity of training data for tasks such as action unit (AU) detection, which typically require costly per-frame annotations. In this work, we introduce LibreFace 2.0, a toolkit that leverages recent advances in face generation and motion retargeting to enrich AU datasets with broader demographic coverage. Specifically, we employ stable diffusion to synthesize a wide range of identities spanning age, gender, race and facial attributes and retarget AU motions from annotated datasets onto these generated identities. Training on this large-scale, demographically diverse dataset yields consistent improvements in benchmark performance and enhances fairness across demographic groups. Beyond AU detection and intensity estimation, LibreFace 2.0 also supports facial expression recognition and gaze estimation through lightweight models that achieve competitive accuracy with substantially fewer parameters, enabling efficient inference. Our work provides a scalable approach to achieving fairer face analysis in real-world applications. The code and the synthetic data will be released publicly at https://github.com/ihp-lab/LibreFace.
Xulang Guan, Ashutosh Chaubey, Maksim Siniukov, Annabelle Hsieh, Zongjian Li, Mohammad Soleymani 0001
FG6
2026 Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
Ashutosh Chaubey, Xulang Guan, Mohammad Soleymani 0001
WACV3
2026 Discrete Facial Encoding: A Framework for Data-driven Facial Display Discovery
abstract
Facial expression analysis is central to understanding human behavior, yet existing coding systems such as the Facial Action Coding System (FACS) are constrained by limited coverage and costly manual annotation. In this work, we introduce Discrete Facial Encoding (DFE), an unsupervised, data-driven alternative of compact and interpretable dictionary of facial expressions from 3D mesh sequences learned through a Residual Vector Quantized Variational Autoencoder (RVQ-VAE). Our approach first extracts identity-invariant expression features from images using a 3D Morphable Model (3DMM), effectively disentangling factors such as head pose and facial geometry. We then encode these features using an RVQ-VAE, producing a sequence of discrete tokens from a shared codebook, where each token captures a specific, reusable facial deformation pattern that contributes to the overall expression. Through extensive experiments, we demonstrate that Discrete Facial Encoding captures more precise facial behaviors than FACS and other facial encoding alternatives. We evaluate the utility of our representation across three high-level psychological tasks: stress detection, personality prediction, and depression detection. Using a simple Bag-of-Words model built on top of the learned tokens, our system consistently outperforms both FACS-based pipelines and strong image and video representation learning models such as Masked Autoencoders. Further analysis reveals that our representation covers a wider variety of facial displays, highlighting its potential as a scalable and effective alternative to FACS for psychological and affective computing applications. We released the source code and model weights at https://github.com/ihp-lab/vqface.
Minh Tran 0004, Maksim Siniukov, Zhangyu Jin, Mohammad Soleymani 0001
WACV4
2025 X-Dyna: Expressive Dynamic Human Image Animation
abstract
We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior approaches centered on human pose control, X-Dyna addresses key shortcomings causing the loss of dynamic details, enhancing the lifelike qualities of human video animations. At the core of our approach is the Dynamics-Adapter, a lightweight module that effectively integrates reference appearance context into the spatial attentions of the diffusion backbone while preserving the capacity of motion modules in synthesizing fluid and intricate dynamic details. Beyond body pose control, we connect a local control module with our model to capture identity-disentangled facial expressions, facilitating accurate expression transfer for enhanced realism in animated scenes. Together, these components form a unified framework capable of learning physical human motion and natural scene dynamics from a diverse blend of human and scene videos. Comprehensive qualitative and quantitative evaluations demonstrate that X-Dyna outperforms state-of- the-art methods, creating highly lifelike and expressive animations. The code is available at https://github.com/bytedance/X-Dyna
Di Chang, You Xie, Yipeng Gao, Zhengfei Kuang, Shengqu Cai, Guoxian Song, Chao Wang 0088, Yichun Shi, Shijie Zhou 0003, Linjie Luo, Gordon Wetzstein, Mohammad Soleymani 0001
CVPR15
2025 Ditailistener: Controllable High Fidelity Listener Video Generation with Diffusion
abstract
Generating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address these challenges, we introduce DiTaiListener, powered by a video diffusion model with multimodal conditions. Our approach first generates short segments of listener responses conditioned on the speaker's speech and facial motions with DiTaiListener-Gen. It then refines the transitional frames via DiTaiListener-Edit for a seamless transition. Specifically, DiTaiListener-Gen adapts a Diffusion Transformer (DiT) for the task of listener head portrait generation by introducing a Causal Temporal Multimodal Adapter (CTM-Adapter) to process speakers' auditory and visual cues. CTM-Adapter integrates speakers' input in a causal manner into the video generation process to ensure temporally coherent listener responses. For long-form video generation, we introduce DiTaiListener-Edit, a transition refinement video-to-video diffusion model. The model fuses video segments into smooth and continuous videos, ensuring temporal consistency in facial expressions and image quality when merging short video segments produced by DiTaiListener-Gen. Quantitatively, DiTaiListener achieves the state-of-the-art performance on benchmark datasets in both photorealism (+73.8% in FID on RealTalk) and motion representation (+6.1% in FD metric on VICO) spaces. User studies confirm the superior performance of DiTaiListener, with the model being the clear preference in terms of feedback, diversity, and smoothness, outperforming competitors by a significant margin.
Maksim Siniukov, Di Chang, Minh Tran 0004, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani 0001
ICCV6
2025 Multimodal Behavioral Characterization of Dyadic Alliance in Support Groups
Kevin Hyekang Joo, Zongjian Li, Yunwen Wang, Yuanfeixue Nan, Mina J. Kian, Shriya Upadhyay, Maja J. Mataric, Lynn C. Miller, Mohammad Soleymani 0001
ICMI9
2025 Speech and Text Foundation Models for Depression Detection: Cross-Task and Cross-Language Evaluation
Lucía Gómez-Zaragozá, Javier Marín-Morales, Mariano Alcañiz Raya, Mohammad Soleymani 0001
INTERSPEECH4
2025 SetPeER: Set-Based Personalized Emotion Recognition With Weak Supervision
abstract
Individual variability of expressive behaviors is a major challenge for emotion recognition systems. Personalized emotion recognition strives to adapt machine learning models to individual behaviors, thereby enhancing emotion recognition performance and overcoming the limitations of generalized emotion recognition systems. However, existing datasets for audiovisual emotion recognition either have a very low number of data points per speaker or include a limited number of speakers. The scarcity of data significantly limits the development and assessment of personalized models, hindering their ability to effectively learn and adapt to individual expressive styles. This paper introduces EmoCeleb: a large-scale, weakly labeled emotion dataset generated via cross-modal labeling. EmoCeleb comprises over 150 hours of audiovisual content from approximately 1,500 speakers, with a median of 50 utterances per speaker. This rich dataset provides a rich resource for developing and benchmarking personalized emotion recognition methods, including those requiring substantial data per individual, such as set learning approaches. We also propose SetPeER: a novel personalized emotion recognition architecture employing set learning. SetPeER effectively captures individual expressive styles by learning representative speaker features from limited data, achieving strong performance with as few as eight utterances per speaker. By leveraging set learning, SetPeER overcomes the limitations of previous approaches that struggle to learn effectively from limited data per individual. Through extensive experiments on EmoCeleb and established benchmarks,i.e, MSP-Podcast and MSP-Improv, we demonstrate the effectiveness of our dataset and the superior performance of SetPeER compared to existing methods for emotion recognition. Our work paves the way for more robust and accurate personalized emotion recognition systems.
Minh Tran 0004, Yufeng Yin 0002, Mohammad Soleymani 0001
IEEE Trans. Affect. Comput.3
2024 Build Your Own Robot Friend: An Open-Source Learning Module for Accessible and Engaging AI Education
abstract
As artificial intelligence (AI) is playing an increasingly important role in our society and global economy, AI education and literacy have become necessary components in college and K-12 education to prepare students for an AI-powered society. However, current AI curricula have not yet been made accessible and engaging enough for students and schools from all socio-economic backgrounds with different educational goals. In this work, we developed an open-source learning module for college and high school students, which allows students to build their own robot companion from the ground up. This open platform can be used to provide hands-on experience and introductory knowledge about various aspects of AI, including robotics, machine learning (ML), software engineering, and mechanical engineering. Because of the social and personal nature of a socially assistive robot companion, this module also puts a special emphasis on human-centered AI, enabling students to develop a better understanding of human-AI interaction and AI ethics through hands-on learning activities. With open-source documentation, assembling manuals and affordable materials, students from different socio-economic backgrounds can personalize their learning experience based on their individual educational goals. To evaluate the student-perceived quality of our module, we conducted a usability testing workshop with 15 college students recruited from a minority-serving institution. Our results indicate that our AI module is effective, easy-to-follow, and engaging, and it increases student interest in studying AI/ML and robotics in the future. We hope that this work will contribute toward accessible and engaging AI education in human-AI interaction for college and high school students.
Zhonghao Shi, Amy O'Connell, Zongjian Li, Siqi Liu 0012, Jennifer Ayissi, Guy Hoffman, Mohammad Soleymani 0001, Maja J. Mataric
AAAI7
2024 DIM: Dyadic Interaction Modeling for Social Behavior Generation
Minh Tran 0004, Di Chang, Maksim Siniukov, Mohammad Soleymani 0001
ECCV (37)4
2024 Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding
Minh Tran 0004, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Min Sun 0001, Mohammad Soleymani 0001
ECCV (80)6
2024 Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
abstract
Human social behaviors are inherently multi-modal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended version of Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE), which is pre-trained on audiovisual social data. Specifically, we modify CAV-MAE to receive a larger number of frames as input and pre-train it on a large dataset of human social interaction (VoxCeleb2) in a self-supervised manner. We demonstrate the effectiveness of this model by fine-tuning and evaluating the model on different social and affective downstream tasks, namely, emotion recognition, laughter detection and apparent personality estimation. The model achieves state-of-the-art results on multimodal emotion recognition and laughter recognition and competitive results for apparent personality estimation, demonstrating the effectiveness of in-domain self-supervised pre-training. Code and model weight are available here https://github.com/HuBohy/SocialMAE.
Hugo Bohy, Minh Tran 0004, Kevin El Haddad, Thierry Dutoit, Mohammad Soleymani 0001
FG5
2024 SEMPI: A Database for Understanding Social Engagement in Video-Mediated Multiparty Interaction
abstract
We present a database for automatic understanding of Social Engagement in MultiParty Interaction (SEMPI). Social engagement is an important social signal characterizing the level of participation of an interlocutor in a conversation. Social engagement involves maintaining attention and establishing connection and rapport. Machine understanding of social engagement can enable an autonomous agent to better understand the state of human participation and involvement to select optimal actions in human-machine social interaction. Recently, video-mediated interaction platforms, e.g., Zoom, have become very popular. The ease of use and increased accessibility of video calls have made them a preferred medium for multiparty conversations, including support groups and group therapy sessions. To create this dataset, we first collected a set of publicly available video calls posted on YouTube. We then segmented the videos by speech turn and cropped the videos to generate single-participant videos. We developed a questionnaire for assessing the level of social engagement by listeners in a conversation probing the relevant nonverbal behaviors for social engagement, including back-channeling, gaze, and expressions. We used Prolific, a crowd-sourcing platform, to annotate 3,505 videos of 76 listeners by three people, reaching a moderate to high inter-rater agreement of 0.693. This resulted in a database with aggregated engagement scores from the annotators. We developed a baseline multimodal pipeline using the state-of-the-art pre-trained models to track the level of engagement achieving the CCC score of 0.454. The results demonstrate the utility of the database for future applications in video-mediated human-machine interaction and human-human social skill assessment. Our dataset and code are available at https://github.com/ihp-lab/SEMPI.
Maksim Siniukov, Yufeng Yin 0002, Eli Fast, Yingshan Qi, Aarav Monga, Audrey Kim, Mohammad Soleymani 0001
ICMI7
2024 MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
abstract
In this work, we propose MagicPose, a diffusion-based model for 2D human pose and facial expression retargeting. Specifically, given a reference image, we aim to generate a person’s new images by controlling the poses and facial expressions while keeping the identity unchanged. To this end, we propose a two-stage training strategy to disentangle human motions and appearance (e.g., facial expressions, skin tone, and dressing), consisting of (1) the pre-training of an appearance-control block and (2) learning appearance-disentangled pose control. Our novel design enables robust appearance control over generated human images, including body, facial attributes, and even background. By leveraging the prior knowledge of image diffusion models, MagicPose generalizes well to unseen human identities and complex poses without the need for additional fine-tuning. Moreover, the proposed model is easy to use and can be considered as a plug-in module/extension to Stable Diffusion. The project website is here. The code is available here.
Di Chang, Yichun Shi, Quankai Gao, Jessica Fu, Guoxian Song, Yizhe Zhu, Mohammad Soleymani 0001
ICML10
2024 FG-Net: Facial Action Unit Detection with Generalizable Pyramidal Features
abstract
Automatic detection of facial Action Units (AUs) allows for objective facial expression analysis. Due to the high cost of AU labeling and the limited size of existing benchmarks, previous AU detection methods tend to overfit the dataset, resulting in a significant performance loss when evaluated across corpora. To address this problem, we propose FG-Net for generalizable facial action unit detection. Specifically, FG-Net extracts feature maps from a Style-GAN2 model pre-trained on a large and diverse face image dataset. Then, these features are used to detect AUs with a Pyramid CNN Interpreter, making the training efficient and capturing essential local features. The proposed FG-Net achieves a strong generalization ability for heatmap-based AU detection thanks to the generalizable and semantic-rich features extracted from the pre-trained generative model. Extensive experiments are conducted to evaluate within- and cross-corpus AU detection with the widely-used DISFA and BP4D datasets. Compared with the state-of-the-art, the proposed method achieves superior cross-domain performance while maintaining competitive within-domain performance. In addition, FG-Net is dataefficient and achieves competitive performance even when trained on 1000 samples. Our code will be released at https://github.com/ihp-lab/FG-Net
Yufeng Yin 0002, Di Chang, Guoxian Song, Shen Sang, Tiancheng Zhi, Linjie Luo, Mohammad Soleymani 0001
WACV8
2024 LibreFace: An Open-Source Toolkit for Deep Facial Expression Analysis
abstract
Facial expression analysis is an important tool for human-computer interaction. In this paper, we introduce LibreFace, an open-source toolkit for facial expression analysis. This open-source toolbox offers real-time and offline analysis of facial behavior through deep learning models, including facial action unit (AU) detection, AU intensity estimation, and facial expression recognition. To accomplish this, we employ several techniques, including the utilization of a large-scale pre-trained network, feature-wise knowledge distillation, and task-specific fine-tuning. These approaches are designed to effectively and accurately analyze facial expressions by leveraging visual information, thereby facilitating the implementation of real-time interactive applications. In terms of Action Unit (AU) intensity estimation, we achieve a Pearson Correlation Coefficient (PCC) of 0.63 on DISFA, which is 7% higher than the performance of OpenFace 2.0 [4] while maintaining highly-efficient inference that runs two times faster than OpenFace 2.0 [4]. Despite being compact, our model also demonstrates competitive performance to state-of-the-art facial expression analysis methods on AffecNet, FFHQ, and RAF-DB. Our code will be released at https://github.com/ihp-lab/LibreFace
Di Chang, Yufeng Yin 0002, Zongjian Li, Minh Tran 0004, Mohammad Soleymani 0001
WACV5
2024 Guest Editorial Best of ACII 2021
abstract
The 9TH AAAC Conference on Affective Computing and Intelligent Interaction 2021 was held in a virtual format in the fall of 2021. It was technically co-sponsored by the IEEE Computer Society and featured the recent work on Affective Computing. The six best papers from this conference were selected by the technical program chairs. They were invited to submit their extended version to be considered for this special section at the IEEE Transactions on Affective Computing. Each submission was reviewed by at least three expert reviewers and was evaluated in terms of overall contribution and the adequacy of the additional content to warrant a new article. This special section features five accepted submissions whose major contributions are summarized below.
Mohammad Soleymani 0001, Shiro Kumano, Emily Mower Provost, Nadia Bianchi-Berthouze, Akane Sano, Kenji Suzuki 0002
IEEE Trans. Affect. Comput.1
2023 Therapist Empathy Assessment in Motivational Interviews
abstract
The quality and effectiveness of psychotherapy sessions are highly influenced by the therapists’ ability to lead the conversation with empathy and acceptance. Manual assessment of the quality of therapy sessions is labor-intensive and difficult to scale. In this paper, we propose a method for estimating session-level therapist empathy ratings for Motivational Interviewing (MI) using therapist language, which has applications in clinical assessment and training. We analyze different stages within therapy sessions to investigate the importance of each stage and its topics of conversation in estimating session-level therapist empathy. We perform experiments on two datasets of MI therapy sessions for alcohol use disorder with session-level empathy scores provided by expert annotators. We achieve average CCC (Concordance Correlation Coefficient) scores of 0.596 and 0.408 for estimating therapist empathy under therapist-dependent and therapist-independent evaluation settings. Our results suggest that therapist responses to client’s discussions on activities and experiences around the problematic behavior (in this case, alcohol abuse) along with the therapist’s usage of in-depth reflections, are the most significant factors in the perception of therapists’ empathy.
Leili Tavabi, Trang Tran 0001, Brian Borsari, Joannalyn Delacruz, Joshua Woolley, Stefan Scherer, Mohammad Soleymani 0001
ACII7
2023 Investigating the Generalizability of Physiological Characteristics of Anxiety
abstract
Recent works have demonstrated the effectiveness of machine learning (ML) techniques in detecting anxiety and stress using physiological signals, but it is unclear whether ML models are learning physiological features specific to stress. To address this ambiguity, we evaluated the generalizability of physiological features that have been shown to be correlated with anxiety and stress to high-arousal emotions. Specifically, we examine features extracted from electrocardiogram (ECG) and electrodermal (EDA) signals from the following three datasets: Anxiety Phases Dataset (APD), Wearable Stress and Affect Detection (WESAD), and the Continuously Annotated Signals of Emotion (CASE) dataset. We aim to understand whether these features are specific to anxiety or general to other high-arousal emotions through a statistical regression analysis, in addition to a within-corpus, cross-corpus, and leave-one-corpus-out cross-validation across instances of stress and arousal. We used the following classifiers: Support Vector Machines, LightGBM, Random Forest, XGBoost, and an ensemble of the aforementioned models. We found that models trained on an arousal dataset perform relatively well on a previously unseen stress dataset, and vice versa. Our experimental results suggest that the evaluated models may be identifying emotional arousal instead of stress. This work is the first cross-corpus evaluation across stress and arousal from ECG and EDA signals, contributing new findings about the generalizability of stress detection.
Emily Zhou, Mohammad Soleymani 0001, Maja J. Mataric
BIBM2
2023 A Speech Representation Anonymization Framework via Selective Noise Perturbation
abstract
Privacy and security are major concerns when communicating speech signals to cloud services such as automatic speech recognition (ASR) and speech emotion recognition (SER). Existing solutions for speech anonymization mainly focus on voice conversion or voice modification to convert a raw utterance into another one with similar content but different, or no, identity-related information. However, an alternative approach to share speech data under the form of privacy-preserving representation has been largely under-explored. In this paper, we propose a speech anonymization framework that achieves privacy via noise perturbation to a selected subset of the high-utility representations extracted using a pre-trained speech encoder. The subset is chosen with a Transformer-based privacy-risk saliency estimator. We validate our framework on four tasks, namely, Automatic Speaker Verification (ASV), ASR, SER and Intent Classification (IC) for privacy and utility assessment. Experimental results show that our approach is able to achieve a competitive, or even superior, utility compared to the speech anonymization baselines from the VoicePrivacy2022 Challenges, while maintaining the same level of privacy. Moreover, the easily-controlled amount of perturbation allows our framework to have a flexible range of privacy-utility trade-offs without re-training any component.
Minh Tran 0004, Mohammad Soleymani 0001
ICASSP2
2023 Personalized Adaptation with Pre-trained Speech Encoders for Continuous Emotion Recognition
Minh Tran 0004, Yufeng Yin 0002, Mohammad Soleymani 0001
INTERSPEECH3
2023 Privacy-preserving Representation Learning for Speech Understanding
Minh Tran 0004, Mohammad Soleymani 0001
INTERSPEECH2
2023 DIVIS: Digital Interactive Victim Intake Simulator
abstract
The Digital Interactive Victim Intake Simulator ("DIVIS") is an interactive, agent-based simulated training tool that has been deployed at the U.S. Army's Sexual Harassment/Assault Response Prevention Program ("SHARP") Academy since May 2021. The system allows student Sexual Assault Response Coordinators ("SARCs") and Victim Advocates ("VAs") to practice critical interpersonal intake skills needed when conducting the initial interview of a survivor of military sexual assault. Currently the system includes two scenarios -- one with a male victim and a second with a female victim -- with two more scenarios under development. Each victim exhibits one of a possible three different emotional vectors, (e.g., angry, ashamed or defensive). Scenarios can run multiple times, giving trainees the ability to navigate through various potential story paths based on how they engage with the victim during each session.
Alesia Gainer, Allison Aptaker, Ron Artstein, David Cobbins, Mark G. Core, Carla Gordon, Anton Leuski, Zongjian Li, Chirag Merchant, David Nelson, Mohammad Soleymani 0001, David R. Traum
IVA11
2023 SAAML: A Framework for Semi-supervised Affective Adaptation via Metric Learning
abstract
Socially intelligent systems such as home robots should be able to perceive emotions and social behaviors. Affect recognition datasets have limited labeled data, and existing large unlabeled datasets, e.g., VoxCeleb2, suitable for pre-training, mostly contain neutral expressions, limiting their application to affective downstream tasks. We introduce a novel Semi-supervised Affective Adaptation framework via Metric Learning (SAAML) to adapt pre-trained audiovisual models (e.g., AV-HuBERT) to expressive behaviors associated with emotions and social communication. The proposed framework automatically retrieves a large number of emotional excerpts (>100 hours) from the VoxCeleb2 dataset via metric learning from two emotion recognition datasets (MSP-IMPROV and CREMA-D), and learns domain-invariant emotion-aware representations. Experimental results show that fine-tuning the proposed affect-aware AV-HuBERT (AW-HuBERT) improves the emotion recognition accuracy by 3-6% compared to fine-tuning the original pre-trained models. We further validate the effectiveness of the AW-HuBERT on human-centered visual understanding tasks, namely, facial expression recognition, video highlight detection, and continuous emotion recognition. The proposed approach consistently outperforms AV-HuBERT and delivers competitive performance compared to the existing methods. With this work, we demonstrate the effectiveness of adaptive pre-training for existing models on domain-specific data to enhance their performance for human-centered tasks.
Minh Tran 0004, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Mohammad Soleymani 0001
ACM Multimedia5
2022 Speech Behavioral Markers Align on Symptom Factors in Psychological Distress
abstract
Automatic detection of psychological disorders has gained significant attention in recent years due to the rise in their prevalence. However, the majority of studies have overlooked the complexity of disorders in favor of a “present/not present” dichotomy in representing disorders. Recent psychological research challenges favors transdiagnostic approaches, moving beyond general disorder classifications to symptom level analysis, as symptoms are often not exclusive to individual disorder classes. In our study, we investigated the link between speech signals and psychological distress symptoms in a corpus of 333 screening interviews from the Distress Analysis Interview Corpus (DAIC). Given the semi-structured organization of interviews, we aggregated speech utterances from responses to shared questions across interviews. We employed deterministic sample selection in classification to rank salient questions for eliciting symptom-specific behaviors in order to predict symptom presence. Some questions include “Do you find therapy helpful?” and “When was the last time you felt happy?”. The prediction results align closely to the factor structure of psychological distress symptoms, linking speech behaviors primarily to somatic and affective alterations in both depression and PTSD. This lends support for the transdiagnostic validity of speech markers for detecting such symptoms. Surprisingly, we did not find a strong link between speech markers and cognitive or psychomotor alterations. This is surprising, given the complexity of motor and cognitive actions required in speech production. The results of our analysis highlight the importance of aligning affective computing research with psychological research to investigate the use of automatic behavioral sensing to assess psychiatric risk.
Larry Zhang, Jacek Kolacz, Albert A. Rizzo, Stefan Scherer, Mohammad Soleymani 0001
ACII5
2022 A Pre-Trained Audio-Visual Transformer for Emotion Recognition
abstract
In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior understanding. The model aims to capture and extract useful information from the interactions between human facial and auditory behaviors, with application in emotion recognition. We evaluate the model performance on two datasets, namely CREMAD-D (emotion classification) and MSP-IMPROV (continuous emotion regression). Experimental results show that fine-tuning the pre-trained model helps improving emotion classification accuracy by 5-7% and Concordance Correlation Coefficients (CCC) in continuous emotion recognition by 0.03-0.09 compared to the same model trained from scratch. We also demonstrate the robustness of finetuning the pre-trained model in a low-resource setting. With only 10% of the original training set provided, finetuning the pre-trained model can lead to at least 10% better emotion recognition accuracy and a CCC score improvement by at least 0.1 for continuous emotion recognition.
Minh Tran 0004, Mohammad Soleymani 0001
ICASSP2
2022 Self-Supervised Learning for Sentiment Analysis via Image-Text Matching
abstract
There is often a resemblance in the sentiment expressed in social media posts (text) and their accompanying images. In this paper, We leverage this sentiment congruence for self-supervised representation learning for sentiment analysis. By teaching the model to pair an image with its corresponding social media post, the model can learn a representation capturing sentiment features from the image and text without supervision. We then use the pre-trained encoder for feature extraction for sentiment analysis in downstream tasks. We show significant improvement and good transferability for sentiment classification in addition to robustness in performance when available data decreases on public datasets (B-T4SA and IMDb Movie Review). With this work, we demonstrate the effectiveness of self-supervised learning through cross-modal matching for sentiment analysis.
Haidong Zhu, Zhaoheng Zheng, Mohammad Soleymani 0001, Ramakant Nevatia
ICASSP3
2022 X-Norm: Exchanging Normalization Parameters for Bimodal Fusion
abstract
Multimodal learning aims to process and relate information from different modalities to enhance the model’s capacity for perception. Current multimodal fusion mechanisms either do not align the feature spaces closely or are expensive for training and inference. In this paper, we present X-Norm, a novel, simple and efficient method for bimodal fusion that generates and exchanges limited but meaningful normalization parameters between the modalities implicitly aligning the feature spaces. We conduct extensive experiments on two tasks of emotion and action recognition with different architectures including Transformer-based and CNN-based models using IEMOCAP and MSP-IMPROV for emotion recognition and EPIC-KITCHENS for action recognition. The experimental results show that X-Norm achieves comparable or superior performance compared to the existing methods including early and late fusion, Gradient-Blending (G-Blend) [44], Tensor Fusion Network, [48] and Multimodal Transformer [40], with a relatively low training cost.
Yufeng Yin 0002, Tianxin Zu, Mohammad Soleymani 0001
ICMI4
2021 Contrastive Learning for Domain Transfer in Cross-Corpus Emotion Recognition
abstract
Automatic emotion recognition methods are sensitive to the variations across humans and datasets and their performance drops when evaluated across corpora. Domain adaptation (DA) techniques such as Domain-Adversarial Neural Network (DANN) can mitigate this problem. However, domain adaptation cannot guarantee to preserve local features necessary for emotion recognition while reducing domain discrepancies in global features. In this paper, we propose Face wArping emoTion rEcognition (FATE) to address this problem. Unlike the traditional DA models in which the base model is first trained with the source data and then fine-tuned with the source and target data, we reverse the training order. Specifically, we employ first-order facial animation warping to generate a synthetic dataset and utilize contrastive learning to pre-train the encoder. Then, we fine-tune the encoder and the classifier with the source data. After fine-tuning, the model achieves superior emotion recognition performance by preserving the subtle facial features. Our experiments on cross-domain emotion recognition with facial behaviors (Aff-Wild2, SEWA, and SEMAINE) indicate that the proposed FATE model substantially outperforms the domain adaptation models, suggesting that FATE has a better domain generalizability for emotion recognition.
Yufeng Yin 0002, Liupei Lu, Zhi Xu 0013, Kaijie Cai, Jonathan Gratch, Mohammad Soleymani 0001
ACII8
2021 Multimodal Phased Transformer for Sentiment Analysis
abstract
Multimodal Transformers achieve superior performance in multimodal learning tasks.However, the quadratic complexity of the selfattention mechanism in Transformers limits their deployment in low-resource devices and makes their inference and training computationally expensive.We propose multimodal Sparse Phased Transformer (SPT) to alleviate the problem of self-attention complexity and memory footprint.SPT uses a sampling function to generate a sparse attention matrix and compress a long sequence to a shorter sequence of hidden states.SPT concurrently captures interactions between the hidden states of different modalities at every layer.To further improve the efficiency of our method, we use Layer-wise parameter sharing and Factorized Co-Attention that share parameters between Cross Attention Blocks, with minimal impact on task performance.We evaluate our model with three sentiment analysis datasets and achieve comparable or superior performance compared with the existing methods, with a 90% reduction in the number of parameters.We conclude that (SPT) along with parameter sharing can capture multimodal interactions with reduced model size and improved sample efficiency.
Junyan Cheng, Iordanis Fostiropoulos, Barry W. Boehm, Mohammad Soleymani 0001
EMNLP (1)4
2021 Modeling Dynamics of Facial Behavior for Mental Health Assessment
abstract
Facial action unit (FAU) intensities are popular descriptors for the analysis of facial behavior. However, FAUs are sparsely represented when only a few are activated at a time. In this study, we explore the possibility of representing the dynamics of facial expressions by adopting algorithms used for word representation in natural language processing. Specifically, we perform clustering on a large dataset of temporal facial expressions with 5.3M frames before applying the Global Vector representation (GloVe) algorithm to learn the embeddings of the facial clusters. We evaluate the usefulness of our learned representations on two downstream tasks: schizophrenia symptom estimation and depression severity regression. These experimental results show the potential of our approach for improving the assessment of mental health symptoms over baseline models that use FAU intensities alone.
Minh Tran 0004, Ellen Bradley, Michelle Matvey, Joshua Woolley, Mohammad Soleymani 0001
FG5
2021 Self-Supervised Patch Localization for Cross-Domain Facial Action Unit Detection
abstract
Automatic detection of Facial Action Units (AUs) is a fundamental block for objective facial expression analysis. Computer vision-based detection of facial action units is susceptible to variations across corpora. To address this problem, we propose a novel architecture that can be jointly trained for self-supervised optical flow estimation, patch localization, supervised action unit detection, and adversarial domain adaptation. Patch localization allows the encoder to learn local features that are critical to detecting subtle changes caused by the presence of AUs in face. Specifically, an encoder-decoder architecture is used to estimate optical flow from every image. The optical flow is simultaneously used for AU detection, patch localization and adversarial domain adaptation. Majority of the existing work on facial expression analysis is evaluated within corpora. In this paper, we develop and evaluate this novel architecture for facial action unit detection across corpora. The experimental results indicate that our framework improves cross-domain performance (5.5 % F1-score on average), suggesting that the proposed patch localization guides the network to learn a more generalizable representation.
Yufeng Yin 0002, Liupei Lu, Yizhen Wu, Mohammad Soleymani 0001
FG4
2021 Subject-Invariant Eeg Representation Learning For Emotion Recognition
abstract
The discrepancies between the distributions of the train and test data, a.k.a., domain shift, result in lower generalization for emotion recognition methods. One of the main factors contributing to these discrepancies is human variability. Domain adaptation methods are developed to alleviate the problem of domain shift, however, these techniques while reducing between database variations fail to reduce between-subject variability. In this paper, we propose an adversarial deep domain adaptation approach for emotion recognition from electroencephalogram (EEG) signals. The method jointly learns a new representation that minimizes emotion recognition loss and maximizes subject confusion loss. We demonstrate that the proposed representation can improve emotion recognition performance within and across databases.
Soheil Rayatdoost, Yufeng Yin 0002, David Rudrauf, Mohammad Soleymani 0001
ICASSP4
2021 A Systematic Cross-Corpus Analysis of Human Reactions to Robot Conversational Failures
abstract
In this paper, we analyze multimodal behavioral responses to robot failures across different tasks. Two multimodal datasets are examined in which humans interact with guided-task robots in task-oriented dialogues. In both datasets, the robots simulated failures of conversational breakdown and miscommunication typically observed in human-robot interactions. We closely examine human reactions to these failures looking at facial and acoustic features. Our analyses identify the significant behavioral features for automatic detection of such failures in interaction. We also examine human responses to different types of robot failures and if failures occurred early or late in the interaction cause variation in the responses. Our findings indicate that several nonverbal behaviors are consistently present in responses to robots’ failures, e.g., gaze and speech prosody, whereas, linguistic features appear to be task-dependent. We discuss how these findings may generalize to other tasks, and how autonomous robots may identify opportunities to detect and recover from failures in interactions with humans.
Dimosthenis Kontogiorgos, Minh Tran 0004, Joakim Gustafson, Mohammad Soleymani 0001
ICMI4
2020 Self-Supervised Learning for Facial Action Unit Recognition through Temporal Consistency
Liupei Lu, Leili Tavabi, Mohammad Soleymani 0001
BMVC3
2020 Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction
abstract
Offensive and abusive language is a pressing problem on social media platforms.In this work, we propose a method for transforming offensive comments, statements containing profanity or offensive language, into non-offensive ones.We design a RETRIEVE, GENERATE and EDIT unsupervised style transfer pipeline to redact the offensive comments in a word-restricted manner while maintaining a high level of fluency and preserving the content of the original text.We extensively evaluate our method's performance and compare it to previous style transfer models using both automatic metrics and human evaluations.Experimental results show that our method outperforms other models on human evaluations and is the only approach that consistently performs well on all automatic evaluation metrics.
Minh Tran 0004, Mohammad Soleymani 0001
COLING3
2020 Expression-Guided EEG Representation Learning for Emotion Recognition
abstract
Learning a joint and coordinated representation between different modalities can improve multimodal emotion recognition. In this paper, we propose a deep representation learning approach for emotion recognition from electroencephalogram (EEG) signals guided by facial electromyogram (EMG) and electrooculogram (EOG) signals. We recorded EEG, EMG and EOG signals from 60 participants who watched 40 short videos and self-reported their emotions. A cross-modal encoder that jointly learns the features extracted from facial and ocular expressions and EEG responses was designed and evaluated on our recorded data and MAHOB-HCI, a publicly available database. We demonstrate that the proposed representation is able to improve emotion recognition performance. We also show that the learned representation can be transferred to a different database without EMG and EOG and achieve superior performance. Methods that fuse behavioral and neural responses can be deployed in wearable emotion recognition solutions, practical in situations in which computer vision expression recognition is not feasible.
Soheil Rayatdoost, David Rudrauf, Mohammad Soleymani 0001
ICASSP3
2020 Punchline Detection using Context-Aware Hierarchical Multimodal Fusion
abstract
Humor has a history as old as humanity. Humor often induces laughter and elicits amusement and engagement. Humorous behavior involves behavior manifested in different modalities including language, voice tone, and gestures. Thus, automatic understanding of humorous behavior requires multimodal behavior analysis. Humor detection is a well-established problem in Natural Language Processing but its multimodal analysis is less explored. In this paper, we present a context-aware hierarchical fusion network for multimodal punchline detection. The proposed neural architecture first fuses the modalities two by two and then fuses all three modalities. The network also models the context of the punchline using Gated Recurrent Unit(s). The model's performance is evaluated on UR-FUNNY database yielding state-of-the-art performance.
Akshat Choube, Mohammad Soleymani 0001
ICMI2
2020 Incorporating Measures of Intermodal Coordination in Automated Analysis of Infant-Mother Interaction
abstract
Interactions between infants and their mothers can provide meaningful insight into the dyad's health and well-being. Previous work has shown that infant-mother coordination, within a single modality, varies significantly with age and interaction quality. However, as infants are still developing their motor, language, and social skills, they may differ from their mothers in the modes they use to communicate. This work examines how infant-mother coordination across modalities can expand researchers' abilities to observe meaningful trends in infant-mother interactions. Using automated feature extraction tools, we analyzed the head position, arm position, and vocal fundamental frequency of mothers and their infants during the Face-to-Face Still-Face (FFSF) procedure. A de-identified dataset including these features was made available online as a contribution of this work. Analysis of infant behavior over the course of the FFSF indicated that the amount and modality of infant behavior change evolves with age. Evaluating the interaction dynamics, we found that infant and mother behavioral signals are coordinated both within and across modalities, and that levels of both intramodal and intermodal coordination vary significantly with age and across stages of the FFSF. These results support the significance of intermodal coordination when assessing changes in infant-mother interaction across conditions.
Lauren Klein, Victor Ardulov, Yuhua Hu, Mohammad Soleymani 0001, Alma Gharib, Barbara Thompson, Pat Levitt, Maja J. Mataric
ICMI4
2020 Multimodal Gated Information Fusion for Emotion Recognition from EEG Signals and Facial Behaviors
abstract
Emotions associated with neural and behavioral responses are detectable through scalp electroencephalogram (EEG) signals and measures of facial expressions. We propose a multimodal deep representation learning approach for emotion recognition from EEG and facial expression signals. The proposed method involves the joint learning of a unimodal representation aligned with the other modality through cosine similarity and a gated fusion for modality fusion. We evaluated our method on two databases: DAI-EF and MAHNOB-HCI. The results show that our deep representation is able to learn mutual and complementary information between EEG signals and face video, captured by action units, head and eye movements from face videos, in a manner that generalizes across databases. It is able to outperform similar fusion methods for the task at hand.
Soheil Rayatdoost, David Rudrauf, Mohammad Soleymani 0001
ICMI3
2020 OpenSense: A Platform for Multimodal Data Acquisition and Behavior Perception
abstract
Automatic multimodal acquisition and understanding of social signals is an essential building block for natural and effective human-machine collaboration and communication. This paper introduces OpenSense, a platform for real-time multimodal acquisition and recognition of social signals. OpenSense enables precisely synchronized and coordinated acquisition and processing of human behavioral signals. Powered by the Microsoft's Platform for Situated Intelligence, OpenSense supports a range of sensor devices and machine learning tools and encourages developers to add new components to the system through straightforward mechanisms for component integration. This platform also offers an intuitive graphical user interface to build application pipelines from existing components. OpenSense is freely available for academic research.
Kalin Stefanov, Baiyu Huang, Zongjian Li, Mohammad Soleymani 0001
ICMI4
2020 Multimodal Automatic Coding of Client Behavior in Motivational Interviewing
abstract
Motivational Interviewing (MI) is defined as a collaborative conversation style that evokes the client's own intrinsic reasons for behavioral change. In MI research, the clients' attitude (willingness or resistance) toward change as expressed through language, has been identified as an important indicator of their subsequent behavior change. Automated coding of these indicators provides systematic and efficient means for the analysis and assessment of MI therapy sessions. In this paper, we study and analyze behavioral cues in client language and speech that bear indications of the client's behavior toward change during a therapy session, using a database of dyadic motivational interviews between therapists and clients with alcohol-related problems. Deep language and voice encoders, \ie BERT and VGGish, trained on large amounts of data are used to extract features from each utterance. We develop a neural network to automatically detect the MI codes using both the clients' and therapists' language and clients' voice, and demonstrate the importance of semantic context in such detection. Additionally, we develop machine learning models for predicting alcohol-use behavioral outcomes of clients through language and voice analysis. Our analysis demonstrates that we are able to estimate MI codes using clients' textual utterances along with preceding textual context from both the therapist and client, reaching an F1-score of 0.72 for a speaker-independent three-class classification. We also report initial results for using the clients' data for predicting behavioral outcomes, which outlines the direction for future work.
Leili Tavabi, Kalin Stefanov, Larry Zhang, Brian Borsari, Joshua Woolley, Stefan Scherer, Mohammad Soleymani 0001
ICMI7
2020 Mitigating Biases in Multimodal Personality Assessment
abstract
As algorithmic decision making systems are increasingly used in high-stake scenarios, concerns have risen about the potential unfairness of these decisions to certain social groups. Despite its importance, the bias and fairness of multimodal systems are not thoroughly studied. In this work, we focus on the multimodal systems designed for apparent personality assessment and hirability prediction. We use the First Impression dataset as a case study to investigate the biases in such systems. We provide detailed analyses on the biases from different modalities and data fusion strategies. Our analyses reveal that different modalities show various patterns of biases and data fusion process also introduces additional biases to the model. To mitigate the biases, we develop and evaluate two different debiasing approaches based on data balancing and adversarial learning. Experimental results show that both approaches can reduce the biases in model outcomes without sacrificing much performance. Our debiasing strategies can be deployed in real-world multimodal systems to provide fairer outcomes.
Shen Yan 0007, Mohammad Soleymani 0001
ICMI3
2020 Speaker-Invariant Adversarial Domain Adaptation for Emotion Recognition
abstract
Automatic emotion recognition methods are sensitive to the variations across different datasets and their performance drops when evaluated across corpora. We can apply domain adaptation techniques e.g., Domain-Adversarial Neural Network (DANN) to mitigate this problem. Though the DANN can detect and remove the bias between corpora, the bias between speakers still remains which results in reduced performance. In this paper, we propose Speaker-Invariant Domain-Adversarial Neural Network (SIDANN) to reduce both the domain bias and the speaker bias. Specifically, based on the DANN, we add a speaker discriminator to unlearn information representing speakers' individual characteristics with a gradient reversal layer (GRL). Our experiments with multimodal data (speech, vision, and text) and the cross-domain evaluation indicate that the proposed SIDANN outperforms (+5.6% and +2.8% on average for detecting arousal and valence) the DANN model, suggesting that the SIDANN has a better domain adaptation ability than the DANN. Besides, the modality contribution analysis shows that the acoustic features are the most informative for arousal detection while the lexical features perform the best for valence detection.
Yufeng Yin 0002, Baiyu Huang, Yizhen Wu, Mohammad Soleymani 0001
ICMI4
2019 Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
abstract
Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single point in the shared space. Unfortunately, injective embedding cannot effectively handle polysemous instances with multiple possible meanings; at best, it would find an average representation of different meanings. This hinders its use in real-world scenarios where individual instances and their cross-modal associations are often ambiguous. In this work, we introduce Polysemous Instance Embedding Networks (PIE-Nets) that compute multiple and diverse representations of an instance by combining global context with locally-guided features via multi-head self-attention and residual learning. To learn visual-semantic embedding, we tie-up two PIE-Nets and optimize them jointly in the multiple instance learning framework. Most existing work on cross-modal retrieval focus on image-text pairs of data. Here, we also tackle a more challenging case of video-text retrieval. To facilitate further research in video-text retrieval, we release a new dataset of 50K video-sentence pairs collected from social media, dubbed MRW (my reaction when). We demonstrate our approach on both image-text and video-text retrieval scenarios using MS-COCO, TGIF, and our new MRW dataset.
Yale Song, Mohammad Soleymani 0001
CVPR2
2019 Multimodal Analysis and Estimation of Intimate Self-Disclosure
abstract
Self-disclosure to others has a proven benefit for one’s mental health. It is shown that disclosure to computers can be similarly beneficial for emotional and psychological well-being. In this paper, we analyzed verbal and nonverbal behavior associated with self-disclosure in two datasets containing structured human-human and human-agent interviews from more than 200 participants. Correlation analysis of verbal and nonverbal behavior revealed that linguistic features such as affective and cognitive content in verbal behavior, and nonverbal behavior such as head gestures are associated with intimate self-disclosure. A multimodal deep neural network was developed to automatically estimate the level of intimate self-disclosure from verbal and nonverbal behavior. Between modalities, verbal behavior was the best modality for estimating self-disclosure within-corpora achieving r = 0.66. However, the cross-corpus evaluation demonstrated that nonverbal behavior can outperform language modality in cross-corpus evaluation. Such automatic models can be deployed in interactive virtual agents or social robots to evaluate rapport and guide their conversational strategy.
Mohammad Soleymani 0001, Kalin Stefanov, Sin-Hwa Kang, Jan Ondras, Jonathan Gratch
ICMI1
2019 Multimodal Learning for Identifying Opportunities for Empathetic Responses
abstract
Embodied interactive agents possessing emotional intelligence and empathy can create natural and engaging social interactions. Providing appropriate responses by interactive virtual agents requires the ability to perceive users’ emotional states. In this paper, we study and analyze behavioral cues that indicate an opportunity to provide an empathetic response. Emotional tone in language in addition to facial expressions are strong indicators of dramatic sentiment in conversation that warrant an empathetic response. To automatically recognize such instances, we develop a multimodal deep neural network for identifying opportunities when the agent should express positive or negative empathetic responses. We train and evaluate our model using audio, video and language from human-agent interactions in a wizard-of-Oz setting, using the wizard’s empathetic responses and annotations collected on Amazon Mechanical Turk as ground-truth labels. Our model outperforms a text-based baseline achieving F1-score of 0.71 on a three-class classification. We further investigate the results and evaluate the capability of such a model to be deployed for real-world human-agent interactions.
Leili Tavabi, Kalin Stefanov, Setareh Nasihati Gilani, David R. Traum, Mohammad Soleymani 0001
ICMI5
2019 Special Section on Multimodal Understanding of Social, Affective, and Subjective Attributes
abstract
Multimedia scientists have largely focused their research on the recognition of tangible properties of data such as objects and scenes. Recently, the field has started evolving toward the modeling of more complex properties. For example, the understanding of social, affective, and subjective attributes of visual data has attracted the attention of many research teams at the crossroads of computer vision, multimedia, and social sciences. These intangible attributes include, for example, visual beauty, video popularity, or user behavior. Multiple, diverse challenges arise when modeling such properties from multimedia data. The sections concern technical aspects such as reliable groundtruth collection, the effective learning of subjective properties, or the impact of context in subjective perception; see Refs. [2] and [3].
Xavier Alameda-Pineda, Miriam Redi, Mohammad Soleymani 0001, Nicu Sebe, Shih-Fu Chang, Samuel D. Gosling
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Large-scale Affective Content Analysis: Combining Media Content Features and Facial Reactions
abstract
We present a novel multimodal fusion model for affective content analysis, combining visual, audio and deep visual-sentiment descriptors from the media content with automated facial action measurements from naturalistic responses to the media. We collected a dataset of 48,867 facial responses to 384 media clips and extracted a rich feature set from the facial responses and media content. The stimulus videos were validated to be informative, inspiring, persuasive, sentimental or amusing. By combining the features, we were able to obtain a classification accuracy of 63% (weighted F1-score: 0.62) for a five-class task. This was a significant improvement over using the media content features alone. By analyzing the feature sets independently, we found that states of informed and persuaded were difficult to differentiate from facial responses alone due to the presence of similar sets of action units in each state (AU 2 occurring frequently in both cases). Facial actions were beneficial in differentiating between amused and informed states whereas media content features alone performed less well due to similarities in the visual and audio make up of the content. We highlight examples of content and reactions from each class. This is the first affective content analysis based on reactions of 10,000s of people.
Daniel McDuff, Mohammad Soleymani 0001
FG2
2017 Multimodal Analysis of Image Search Intent: Intent Recognition in Image Search from User Behavior and Visual Content
abstract
Users search for multimedia content with different underlying motivations or intentions. Study of user search intentions is an emerging topic in information retrieval since understanding why a user is searching for a content is crucial for satisfying the user's need. In this paper, we aimed at automatically recognizing a user's intent for image search in the early stage of a search session. We designed seven different search scenarios under the intent conditions of finding items, re-finding items and entertainment. We collected facial expressions, physiological responses, eye gaze and implicit user interactions from 51 participants who performed seven different search tasks on a custom-built image retrieval platform. We analyzed the users' spontaneous and explicit reactions under different intent conditions. Finally, we trained machine learning models to predict users' search intentions from the visual content of the visited images, the user interactions and the spontaneous responses. After fusing the visual and user interaction features, our system achieved the F-1 score of 0.722 for classifying three classes in a user-independent cross-validation. We found that eye gaze and implicit user interactions, including mouse movements and keystrokes are the most informative features. Given that the most promising results are obtained by modalities that can be captured unobtrusively and online, the results demonstrate the feasibility of deploying such methods for improving multimedia retrieval platforms.
Mohammad Soleymani 0001, Michael Riegler 0001, Pål Halvorsen
ICMR1
2017 MUSA2: First ACM Workshop on Multimodal Understanding of Social, Affective and Subjective Attributes
abstract
Multimedia scientists have largely focused their research on the recognition of tangible properties of data, such as objects and scenes. Recently, the field has started evolving towards the modeling of more complex properties. For example, the understanding of social, affective and subjective attributes of data has attracted the attention of many research teams at the crossroads of computer vision, multimedia, and social sciences. These intangible attributes include, for example, visual beauty, video popularity, or user behavior. Multiple, diverse challenges arise when modeling such properties from multimedia data. Issues concern technical aspects such as reliable groundtruth collection, the effective learning of subjective properties, or the impact of context in subjective perception. The first edition of the ACM MM'17 MUSA2 workshop has gathered together high-quality research works focusing on the computational understanding of intangible properties from multimodal data, including visual emotions, user intent, human relationships, and personality.
Xavier Alameda-Pineda, Miriam Redi, Mohammad Soleymani 0001, Nicu Sebe, Shih-Fu Chang, Samuel D. Gosling
ACM Multimedia3
2017 A survey of multimodal sentiment analysis
Mohammad Soleymani 0001, David García 0001, Brendan Jou, Björn W. Schuller, Shih-Fu Chang, Maja Pantic
Image Vis. Comput.1
2017 Guest editorial: Multimodal sentiment analysis and mining in the wild
Mohammad Soleymani 0001, Björn W. Schuller, Shih-Fu Chang
Image Vis. Comput.1
2016 Analyzing and Predicting GIF Interestingness
abstract
Animated GIFs have regained huge popularity. They are used in instant messaging, online journalism, social media, among others. In this paper, we present an in-depth study on the interestingness of GIFs. We create and annotate a dataset with a set of affective labels, which allows us to investigate the sources of interest. We show that GIFs of pets are considered more interesting that GIFs of people. Furthermore, we study the connection of interest to other features and factors such as popularity. Finally, we build a predictive model and show that it can estimate GIF interestingness with high accuracy. Our model outperforms the existing methods on GIF popularity, as well as a model based on still image interestingness, by a large margin. We envision that the insights and method developed can be used for automatic recognition and generation of interesting GIFs.
Michael Gygli, Mohammad Soleymani 0001
ACM Multimedia2
2016 Analysis of EEG Signals and Facial Expressions for Continuous Emotion Detection
abstract
Emotions are time varying affective phenomena that are elicited as a result of stimuli. Videos and movies in particular are made to elicit emotions in their audiences. Detecting the viewers' emotions instantaneously can be used to find the emotional traces of videos. In this paper, we present our approach in instantaneously detecting the emotions of video viewers' emotions from electroencephalogram (EEG) signals and facial expressions. A set of emotion inducing videos were shown to participants while their facial expressions and physiological responses were recorded. The expressed valence (negative to positive emotions) in the videos of participants' faces were annotated by five annotators. The stimuli videos were also continuously annotated on valence and arousal dimensions. Long-short-term-memory recurrent neural networks (LSTM-RNN) and continuous conditional random fields (CCRF) were utilized in detecting emotions automatically and continuously. We found the results from facial expressions to be superior to the results from EEG signals. We analyzed the effect of the contamination of facial muscle activities on EEG signals and found that most of the emotionally valuable content in EEG features are as a result of this contamination. However, our statistical analysis showed that EEG signals still carry complementary information in presence of facial expressions.
Mohammad Soleymani 0001, Sadjad Asghari-Esfeden, Yun Fu 0001, Maja Pantic
IEEE Trans. Affect. Comput.1
2015 Multimodal emotion recognition in response to videos (Extended abstract)
abstract
We present a user-independent emotion recognition method with the goal of detecting expected emotions or affective tags for videos using electroencephalogram (EEG), pupillary response and gaze distance. We first selected 20 video clips with extrinsic emotional content from movies and online resources. Then EEG responses and eye gaze data were recorded from 24 participants while watching emotional video clips. Ground truth was defined based on the median arousal and valence scores given to clips in a preliminary study. The arousal classes were calm, medium aroused and activated and the valence classes were unpleasant, neutral and pleasant. A one-participant-out cross validation was employed to evaluate the classification performance in a user-independent approach. The best classification accuracy of 68.5% for three labels of valence and 76.4% for three labels of arousal were obtained using a modality fusion strategy and a support vector machine. The results over a population of 24 participants demonstrate that user-independent emotion recognition can outperform individual self-reports for arousal assessments and do not underperform for valence assessments.
Mohammad Soleymani 0001, Maja Pantic, Thierry Pun
ACII1
2015 Content-based music recommendation using underlying music preference structure
abstract
The cold start problem for new users or items is a great challenge for recommender systems. New items can be positioned within the existing items using a similarity metric to estimate their ratings. However, the calculation of similarity varies by domain and available resources. In this paper, we propose a content-based music recommender system which is based on a set of attributes derived from psychological studies of music preference. These five attributes, namely, Mellow, Unpretentious, Sophisticated, Intense and Contemporary (MUSIC), better describe the underlying factors of music preference compared to music genre. Using 249 songs and hundreds of ratings and attribute scores, we first develop an acoustic content-based attribute detection using auditory modulation features and a regression by sparse representation. We then use the estimated attributes in a cold start recommendation scenario. The proposed content-based recommendation significantly outperforms genre-based and user-based recommendation based on the root-mean-square error. The results demonstrate the effectiveness of these attributes in music preference estimation. Such methods will increase the chance of less popular but interesting songs in the long tail to be listened to.
Mohammad Soleymani 0001, Anna Aljanaki, Frans Wiering, Remco C. Veltkamp
ICME1
2015 The Quest for Visual Interest
abstract
In this paper, we report on identifying the underlying factors that contribute to the visual interest in digital photos. A set of 1005 digital photos covering different topics and of different qualities was collected from Flickr. Images were annotated by a pool of diverse participants on a crowdsourcing platform. 12 bipolar ratings were collected for each photo on 7-point semantic differential scale, including dimensions related to interest, emotions and image quality. Every image received 20 annotations from unique participants. The most important appraisals and visual attributes for visual interest in photos was identified. We found that intrinsic pleasantness, arousal, visual quality and coping potential are the most important factors contributing to visual interest in digital photos. We developed a system that automatically detects the important visual attributes from low level visual features and demonstrated their significance in predicting interest at individual level.
Mohammad Soleymani 0001
ACM Multimedia1
2015 ASM'15: The 1st International Workshop on Affect and Sentiment in Multimedia
abstract
No abstract available.
Mohammad Soleymani 0001, Yi-Hsuan Yang, Yu-Gang Jiang 0001, Shih-Fu Chang
ACM Multimedia1
2015 VSD, a public dataset for the detection of violent scenes in movies: design, annotation, analysis and evaluation
Claire-Hélène Demarty, Cédric Penet, Mohammad Soleymani 0001, Guillaume Gravier
Multim. Tools Appl.3
2015 Guest Editorial: Challenges and Perspectives for Affective Analysis in Multimedia
abstract
The articles in this special section focus on new areas of development in the multimedia industry.
Mohammad Soleymani 0001, Yi-Hsuan Yang, Go Irie, Alan Hanjalic
IEEE Trans. Affect. Comput.1
2014 Continuous emotion detection using EEG signals and facial expressions
abstract
Emotions play an important role in how we select and consume multimedia. Recent advances on affect detection are focused on detecting emotions continuously. In this paper, for the first time, we continuously detect valence from electroencephalogram (EEG) signals and facial expressions in response to videos. Multiple annotators provided valence levels continuously by watching the frontal facial videos of participants who watched short emotional videos. Power spectral features from EEG signals as well as facial fiducial points are used as features to detect valence levels for each frame continuously. We study the correlation between features from EEG and facial expressions with continuous valence. We have also verified our model's performance for the emotional highlight detection using emotion recognition from EEG signals. Finally the results of multimodal fusion between facial expression and EEG signals are presented. Having such models we will be able to detect spontaneous and subtle affective responses over time and use them for video highlight detection.
Mohammad Soleymani 0001, Sadjad Asghari-Esfeden, Maja Pantic, Yun Fu 0001
ICME1
2014 Emotional Analysis of Music: A Comparison of Methods
abstract
Music as a form of art is intentionally composed to be emotionally expressive. The emotional features of music are invaluable for music indexing and recommendation. In this paper we present a cross-comparison of automatic emotional analysis of music. We created a public dataset of Creative Commons licensed songs. Using valence and arousal model, the songs were annotated both in terms of the emotions that were expressed by the whole excerpt and dynamically with 1 Hz temporal resolution. Each song received 10 annotations on Amazon Mechanical Turk and the annotations were averaged to form a ground truth. Four different systems from three teams and the organizers were employed to tackle this problem in an open challenge. We compare their performances and discuss the best practices. While the effect of a larger feature set was not very apparent in the static emotion estimation, the combination of a comprehensive feature set and a recurrent neural network that models temporal dependencies has largely outperformed the other proposed methods for dynamic music emotion estimation.
Mohammad Soleymani 0001, Anna Aljanaki, Yi-Hsuan Yang, Michael N. Caro, Florian Eyben, Konstantin Markov, Björn W. Schuller, Remco C. Veltkamp, Felix Weninger, Frans Wiering
ACM Multimedia1
2014 Corpus Development for Affective Video Indexing
abstract
Affective video indexing is the area of research that develops techniques to automatically generate descriptions of video content that encode the emotional reactions which the video content evokes in viewers. This paper provides a set of corpus development guidelines based on state-of-the-art practice intended to support researchers in this field. Affective descriptions can be used for video search and browsing systems offering users affective perspectives. The paper is motivated by the observation that affective video indexing has yet to fully profit from the standard corpora (data sets) that have benefited conventional forms of video indexing. Affective video indexing faces unique challenges, since viewer-reported affective reactions are difficult to assess. Moreover affect assessment efforts must be carefully designed in order to both cover the types of affective responses that video content evokes in viewers and also capture the stable and consistent aspects of these responses. We first present background information on affect and multimedia and related work on affective multimedia indexing, including existing corpora. Three dimensions emerge as critical for affective video corpora, and form the basis for our proposed guidelines: the context of viewer response, personal variation among viewers, and the effectiveness and efficiency of corpus creation. Finally, we present examples of three recent corpora and discuss how these corpora make progressive steps towards fulfilling the guidelines.
Mohammad Soleymani 0001, Martha A. Larson, Thierry Pun, Alan Hanjalic
IEEE Trans. Multim.1
2013 Multimedia implicit tagging using EEG signals
abstract
Electroencephalogram (EEG) signals reflect brain activities associated with emotional and cognitive processes. In this paper, we demonstrate how they can be used to find tags for multimedia content without users' direct input. Alternative methods for multimedia tagging is attracting increasing interest from multimedia community. The new portable EEG helmets are paving the way for employing brain waves in human computer interaction. In this paper, we demonstrate the performance of EEG for tagging purposes using two different scenarios on MAHNOB-HCI database. First, an emotional tagging and classification using a reduced set of electrodes is presented. The emotional responses of 24 participants to short video clips are classified into three classes on arousal and valence. We show how a reduced set of electrodes based on previous studies can preserve and even enhance the emotional classification rate. We then demonstrate the feasibility of using EEG signals for tag relevance tasks. A set of images were shown to participants first, without any tag and then with a relevant or irrelevant tag. The relevance of the tag was assessed based on the EEG responses of the participants in the first second after the tag was depicted. Finally, we demonstrate that by aggregating multiple participants' responses we can significantly improve the tagging accuracy.
Mohammad Soleymani 0001, Maja Pantic
ICME1
2013 Human behavior sensing for tag relevance assessment
abstract
Users react differently to non-relevant and relevant tags associated with content. These spontaneous reactions can be used for labeling large multimedia databases. We present a method to assess tag relevance to images using the non-verbal bodily responses, namely, electroencephalogram (EEG), facial expressions, and eye gaze. We conducted experiments in which 28 images were shown to 28 subjects once with correct and another time with incorrect tags. The goal of our system is to detect the responses to non-relevant tags and consequently filter them out. Therefore, we trained classifiers to detect the tag relevance from bodily responses. We evaluated the performance of our system using a subject independent approach. The precision at top 5% and top 10% detections were calculated and results of different modalities and different classifiers were compared. The results show that eye gaze outperforms the other modalities in tag relevance detection both overall and for top ranked results.
Mohammad Soleymani 0001, Sebastian Kaltwang, Maja Pantic
ACM Multimedia1
2013 Crowdsourcing for multimedia research
abstract
Crowdsourcing techniques make use of intelligent contributions of large number of human crowdmembers. This tutorial introduces researchers to the applications of crowdsourcing to multimedia analysis with the aim of allowing them to understand the potentials and limitations of crowdsourcing tools and techniques. We emphasize the fact that crowdsourcing represents a further development along a pre-existing continuum of techniques, and discuss the added advantages that new developments offer. We provide a basic overview of human computation, with an emphasis on example cases in which crowdsourcing has been applied to generate data sets, to improve automatic multimedia content analysis, and to elicit user needs or multimedia system requirements. Different techniques and considerations in using human computation methods to acquire high-quality data and annotations are discussed and demonstrated.
Mohammad Soleymani 0001, Martha A. Larson
ACM Multimedia1
2012 Expert Talk for Time Machine Session: Affective Multimedia Analysis: Introduction, Background and Perspectives
abstract
The term "affective computing" was coined by Rosalind Picard in 1995. She presented her ideas about how to use affect for interaction with and analysis of multimedia. Her ideas were inspiring to the studies and applications on affective multimedia analysis in the last decade. In this talk, the initial ideas and their development to the current state as well as challenges and perspectives are presented.
Mohammad Soleymani 0001
ICME1
2012 Human-centered implicit tagging: Overview and perspectives
abstract
Tags are an effective form of metadata which help users to locate and browse multimedia content of interest. Tags can be generated by users (user-generated explicit tags), automatically from the content (content-based tags), or assigned automatically based on non-verbal behavioral reactions of users to multimedia content (implicit human-centered tags). This paper discusses the definition and applications of implicit human-centered tagging. Implicit tagging is an effortless process by which content is tagged based on users' spontaneous reactions. It is a novel but growing research topic which is attracting more attention with the growing availability of built-in sensors. This paper discusses the state of the art in this novel field of research and provides an overview of publicly available relevant databases and annotation tools. We finally discuss in detail challenges and opportunities in the field.
Mohammad Soleymani 0001, Maja Pantic
SMC1
2012 DEAP: A Database for Emotion Analysis ;Using Physiological Signals
abstract
We present a multimodal data set for the analysis of human affective states. The electroencephalogram (EEG) and peripheral physiological signals of 32 participants were recorded as each watched 40 one-minute long excerpts of music videos. Participants rated each video in terms of the levels of arousal, valence, like/dislike, dominance, and familiarity. For 22 of the 32 participants, frontal face video was also recorded. A novel method for stimuli selection is proposed using retrieval by affective tags from the last.fm website, video highlight detection, and an online assessment tool. An extensive analysis of the participants' ratings during the experiment is presented. Correlates between the EEG signal frequencies and the participants' ratings are investigated. Methods and results are presented for single-trial classification of arousal, valence, and like/dislike ratings using the modalities of EEG, peripheral physiological signals, and multimedia content analysis. Finally, decision fusion of the classification results from different modalities is performed. The data set is made publicly available and we encourage other researchers to use it for testing their own affective state estimation methods.
Sander Koelstra, Christian Mühl, Mohammad Soleymani 0001, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, Ioannis Patras
IEEE Trans. Affect. Comput.3
2012 A Multimodal Database for Affect Recognition and Implicit Tagging
abstract
MAHNOB-HCI is a multimodal database recorded in response to affective stimuli with the goal of emotion recognition and implicit tagging research. A multimodal setup was arranged for synchronized recording of face videos, audio signals, eye gaze data, and peripheral/central nervous system physiological signals. Twenty-seven participants from both genders and different cultural backgrounds participated in two experiments. In the first experiment, they watched 20 emotional videos and self-reported their felt emotions using arousal, valence, dominance, and predictability as well as emotional keywords. In the second experiment, short videos and images were shown once without any tag and then with correct or incorrect tags. Agreement or disagreement with the displayed tags was assessed by the participants. The recorded videos and bodily responses were segmented and stored in a database. The database is made available to the academic community via a web-based system. The collected data were analyzed and single modality and modality fusion results for both emotion recognition and implicit tagging experiments are reported. These results show the potential uses of the recorded modalities and the significance of the emotion elicitation protocol.
Mohammad Soleymani 0001, Jeroen Lichtenauer, Thierry Pun, Maja Pantic
IEEE Trans. Affect. Comput.1
2012 Multimodal Emotion Recognition in Response to Videos
abstract
This paper presents a user-independent emotion recognition method with the goal of recovering affective tags for videos using electroencephalogram (EEG), pupillary response and gaze distance. We first selected 20 video clips with extrinsic emotional content from movies and online resources. Then, EEG responses and eye gaze data were recorded from 24 participants while watching emotional video clips. Ground truth was defined based on the median arousal and valence scores given to clips in a preliminary study using an online questionnaire. Based on the participants' responses, three classes for each dimension were defined. The arousal classes were calm, medium aroused, and activated and the valence classes were unpleasant, neutral, and pleasant. One of the three affective labels of either valence or arousal was determined by classification of bodily responses. A one-participant-out cross validation was employed to investigate the classification performance in a user-independent approach. The best classification accuracies of 68.5 percent for three labels of valence and 76.4 percent for three labels of arousal were obtained using a modality fusion strategy and a support vector machine. The results over a population of 24 participants demonstrate that user-independent emotion recognition can outperform individual self-reports for arousal assessments and do not underperform for valence assessments.
Mohammad Soleymani 0001, Maja Pantic, Thierry Pun
IEEE Trans. Affect. Comput.1
2011 Continuous emotion detection in response to music videos
abstract
Viewers' preference for multimedia selection depends highly on their emotional experience. In this paper, we present an emotion detection method for music videos using central and peripheral nervous system physiological signals as well as multimedia content analysis. A set of 40 music clips eliciting a broad range of emotions were first selected. After extracting the one minute long emotional highlight of each video, they were shown to 32 participants while their physiological responses were recorded. Participants self-reported their felt emotions after watching each clip by means of arousal, valence, dominance, and liking ratings. The physiological signals included electroencephalogram, galvanic skin response, respiration pattern, skin temperature, electromyograms and blood volume pulse using plethysmograph. Emotional features were extracted from the signals and the multimedia content. The emotional features were used to train a linear ridge regressor to detect emotions for each participant using a leave-one-out cross-validation strategy. The performance of the personalized emotion detection is shown to be significantly superior to a random regressor.
Mohammad Soleymani 0001, Sander Koelstra, Ioannis Patras, Thierry Pun
FG1
2011 Automatic tagging and geotagging in video collections and communities
abstract
Automatically generated tags and geotags hold great promise to improve access to video collections and online communities. We overview three tasks offered in the MediaEval 2010 benchmarking initiative, for each, describing its use scenario, definition and the data set released. For each task, a reference algorithm is presented that was used within MediaEval 2010 and comments are included on lessons learned. The Tagging Task, Professional involves automatically matching episodes in a collection of Dutch television with subject labels drawn from the keyword thesaurus used by the archive staff. The Tagging Task, Wild Wild Web involves automatically predicting the tags that are assigned by users to their online videos. Finally, the Placing Task requires automatically assigning geo-coordinates to videos. The specification of each task admits the use of the full range of available information including user-generated metadata, speech recognition transcripts, audio, and visual features.
Martha A. Larson, Mohammad Soleymani 0001, Pavel Serdyukov, Stevan Rudinac, Christian Wartena, Vanessa Murdock 0001, Gerald Friedland, Roeland Ordelman, Gareth J. F. Jones
ICMR2
2009 Queries and tags in affect-based multimedia retrieval
abstract
An approach for implementing affective information as tags for multimedia content indexing and retrieval is presented. The approach can be used for implicit as well as explicit tags and is presented here using data recorded during the viewing of movie fragments containing annotations and physiological signal recordings. For retrieval based on affective queries, a representation of the query-words is defined in the arousal-valence space in the form of a Gaussian probability distribution and a retrieval method based on this representation is presented. Validation of retrieval accuracy is performed using Precision and Recall parameters. Results show that the use of arousal and valence as affective tags can improve retrieval results.
Joep J. M. Kierkels, Mohammad Soleymani 0001, Thierry Pun
ICME2
2009 Short-term emotion assessment in a recall paradigm
Guillaume Chanel, Joep J. M. Kierkels, Mohammad Soleymani 0001, Thierry Pun
Int. J. Hum. Comput. Stud.3
2008 Affective Characterization of Movie Scenes Based on Multimedia Content Analysis and User's Physiological Emotional Responses
abstract
In this paper, we propose an approach for affective representation of movie scenes based on the emotions that are actually felt by spectators. Such a representation can be used for characterizing the emotional content of video clips for e.g. affective video indexing and retrieval, neuromarketing studies, etc. A dataset of 64 different scenes from eight movies was shown to eight participants. While watching these clips, their physiological responses were recorded. The participants were also asked to self-assess their felt emotional arousal and valence for each scene. In addition, content-based audio- and video-based features were extracted from the movie scenes in order to characterize each one. Degrees of arousal and valence were estimated by a linear combination of features from physiological signals, as well as by a linear combination of content-based features. We showed that a significant correlation exists between arousal/valence provided by the spectator's self-assessments, and affective grades obtained automatically from either physiological responses or from audio-video features. This demonstrates the ability of using multimedia features and physiological responses to predict the expected affect of the user in response to the emotional video content.
Mohammad Soleymani 0001, Guillaume Chanel, Joep J. M. Kierkels, Thierry Pun
ISM1