Yufeng Yin 0002

dblp:15/7783-2 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-5558-2421ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 SetPeER: Set-Based Personalized Emotion Recognition With Weak Supervision
abstract
Individual variability of expressive behaviors is a major challenge for emotion recognition systems. Personalized emotion recognition strives to adapt machine learning models to individual behaviors, thereby enhancing emotion recognition performance and overcoming the limitations of generalized emotion recognition systems. However, existing datasets for audiovisual emotion recognition either have a very low number of data points per speaker or include a limited number of speakers. The scarcity of data significantly limits the development and assessment of personalized models, hindering their ability to effectively learn and adapt to individual expressive styles. This paper introduces EmoCeleb: a large-scale, weakly labeled emotion dataset generated via cross-modal labeling. EmoCeleb comprises over 150 hours of audiovisual content from approximately 1,500 speakers, with a median of 50 utterances per speaker. This rich dataset provides a rich resource for developing and benchmarking personalized emotion recognition methods, including those requiring substantial data per individual, such as set learning approaches. We also propose SetPeER: a novel personalized emotion recognition architecture employing set learning. SetPeER effectively captures individual expressive styles by learning representative speaker features from limited data, achieving strong performance with as few as eight utterances per speaker. By leveraging set learning, SetPeER overcomes the limitations of previous approaches that struggle to learn effectively from limited data per individual. Through extensive experiments on EmoCeleb and established benchmarks,i.e, MSP-Podcast and MSP-Improv, we demonstrate the effectiveness of our dataset and the superior performance of SetPeER compared to existing methods for emotion recognition. Our work paves the way for more robust and accurate personalized emotion recognition systems.
Minh Tran 0004, Yufeng Yin 0002, Mohammad Soleymani 0001
IEEE Trans. Affect. Comput.2
2024 Hearing Loss Detection From Facial Expressions in One-On-One Conversations
abstract
Individuals with impaired hearing experience difficulty in conversations, especially in noisy environments. This difficulty often manifests as a change in behavior and may be captured via facial expressions, such as the expression of discomfort or fatigue. In this work, we build on this idea and introduce the problem of detecting hearing loss from an individual’s facial expressions during a conversation. Building machine learning models that can represent hearing-related facial expression changes is a challenge. In addition, models need to disentangle spurious age-related correlations from hearing-driven expressions. To this end, we propose a self-supervised pre-training strategy tailored for the modeling of expression variations. We also use adversarial representation learning to mitigate the age bias. We evaluate our approach on a large-scale egocentric dataset with realworld conversational scenarios involving subjects with hearing loss and show that our method for hearing loss detection achieves superior performance over baselines.
Yufeng Yin 0002, Ishwarya Ananthabhotla, Vamsi K. Ithapu, Stavros Petridis, Yu-Hsiang Wu, Christi Miller
ICASSP1
2024 SEMPI: A Database for Understanding Social Engagement in Video-Mediated Multiparty Interaction
abstract
We present a database for automatic understanding of Social Engagement in MultiParty Interaction (SEMPI). Social engagement is an important social signal characterizing the level of participation of an interlocutor in a conversation. Social engagement involves maintaining attention and establishing connection and rapport. Machine understanding of social engagement can enable an autonomous agent to better understand the state of human participation and involvement to select optimal actions in human-machine social interaction. Recently, video-mediated interaction platforms, e.g., Zoom, have become very popular. The ease of use and increased accessibility of video calls have made them a preferred medium for multiparty conversations, including support groups and group therapy sessions. To create this dataset, we first collected a set of publicly available video calls posted on YouTube. We then segmented the videos by speech turn and cropped the videos to generate single-participant videos. We developed a questionnaire for assessing the level of social engagement by listeners in a conversation probing the relevant nonverbal behaviors for social engagement, including back-channeling, gaze, and expressions. We used Prolific, a crowd-sourcing platform, to annotate 3,505 videos of 76 listeners by three people, reaching a moderate to high inter-rater agreement of 0.693. This resulted in a database with aggregated engagement scores from the annotators. We developed a baseline multimodal pipeline using the state-of-the-art pre-trained models to track the level of engagement achieving the CCC score of 0.454. The results demonstrate the utility of the database for future applications in video-mediated human-machine interaction and human-human social skill assessment. Our dataset and code are available at https://github.com/ihp-lab/SEMPI.
Maksim Siniukov, Yufeng Yin 0002, Eli Fast, Yingshan Qi, Aarav Monga, Audrey Kim, Mohammad Soleymani 0001
ICMI2
2024 FG-Net: Facial Action Unit Detection with Generalizable Pyramidal Features
abstract
Automatic detection of facial Action Units (AUs) allows for objective facial expression analysis. Due to the high cost of AU labeling and the limited size of existing benchmarks, previous AU detection methods tend to overfit the dataset, resulting in a significant performance loss when evaluated across corpora. To address this problem, we propose FG-Net for generalizable facial action unit detection. Specifically, FG-Net extracts feature maps from a Style-GAN2 model pre-trained on a large and diverse face image dataset. Then, these features are used to detect AUs with a Pyramid CNN Interpreter, making the training efficient and capturing essential local features. The proposed FG-Net achieves a strong generalization ability for heatmap-based AU detection thanks to the generalizable and semantic-rich features extracted from the pre-trained generative model. Extensive experiments are conducted to evaluate within- and cross-corpus AU detection with the widely-used DISFA and BP4D datasets. Compared with the state-of-the-art, the proposed method achieves superior cross-domain performance while maintaining competitive within-domain performance. In addition, FG-Net is dataefficient and achieves competitive performance even when trained on 1000 samples. Our code will be released at https://github.com/ihp-lab/FG-Net
Yufeng Yin 0002, Di Chang, Guoxian Song, Shen Sang, Tiancheng Zhi, Linjie Luo, Mohammad Soleymani 0001
WACV1
2024 LibreFace: An Open-Source Toolkit for Deep Facial Expression Analysis
abstract
Facial expression analysis is an important tool for human-computer interaction. In this paper, we introduce LibreFace, an open-source toolkit for facial expression analysis. This open-source toolbox offers real-time and offline analysis of facial behavior through deep learning models, including facial action unit (AU) detection, AU intensity estimation, and facial expression recognition. To accomplish this, we employ several techniques, including the utilization of a large-scale pre-trained network, feature-wise knowledge distillation, and task-specific fine-tuning. These approaches are designed to effectively and accurately analyze facial expressions by leveraging visual information, thereby facilitating the implementation of real-time interactive applications. In terms of Action Unit (AU) intensity estimation, we achieve a Pearson Correlation Coefficient (PCC) of 0.63 on DISFA, which is 7% higher than the performance of OpenFace 2.0 [4] while maintaining highly-efficient inference that runs two times faster than OpenFace 2.0 [4]. Despite being compact, our model also demonstrates competitive performance to state-of-the-art facial expression analysis methods on AffecNet, FFHQ, and RAF-DB. Our code will be released at https://github.com/ihp-lab/LibreFace
Di Chang, Yufeng Yin 0002, Zongjian Li, Minh Tran 0004, Mohammad Soleymani 0001
WACV2
2023 Personalized Adaptation with Pre-trained Speech Encoders for Continuous Emotion Recognition
Minh Tran 0004, Yufeng Yin 0002, Mohammad Soleymani 0001
INTERSPEECH2
2022 X-Norm: Exchanging Normalization Parameters for Bimodal Fusion
abstract
Multimodal learning aims to process and relate information from different modalities to enhance the model’s capacity for perception. Current multimodal fusion mechanisms either do not align the feature spaces closely or are expensive for training and inference. In this paper, we present X-Norm, a novel, simple and efficient method for bimodal fusion that generates and exchanges limited but meaningful normalization parameters between the modalities implicitly aligning the feature spaces. We conduct extensive experiments on two tasks of emotion and action recognition with different architectures including Transformer-based and CNN-based models using IEMOCAP and MSP-IMPROV for emotion recognition and EPIC-KITCHENS for action recognition. The experimental results show that X-Norm achieves comparable or superior performance compared to the existing methods including early and late fusion, Gradient-Blending (G-Blend) [44], Tensor Fusion Network, [48] and Multimodal Transformer [40], with a relatively low training cost.
Yufeng Yin 0002, Tianxin Zu, Mohammad Soleymani 0001
ICMI1
2021 Contrastive Learning for Domain Transfer in Cross-Corpus Emotion Recognition
abstract
Automatic emotion recognition methods are sensitive to the variations across humans and datasets and their performance drops when evaluated across corpora. Domain adaptation (DA) techniques such as Domain-Adversarial Neural Network (DANN) can mitigate this problem. However, domain adaptation cannot guarantee to preserve local features necessary for emotion recognition while reducing domain discrepancies in global features. In this paper, we propose Face wArping emoTion rEcognition (FATE) to address this problem. Unlike the traditional DA models in which the base model is first trained with the source data and then fine-tuned with the source and target data, we reverse the training order. Specifically, we employ first-order facial animation warping to generate a synthetic dataset and utilize contrastive learning to pre-train the encoder. Then, we fine-tune the encoder and the classifier with the source data. After fine-tuning, the model achieves superior emotion recognition performance by preserving the subtle facial features. Our experiments on cross-domain emotion recognition with facial behaviors (Aff-Wild2, SEWA, and SEMAINE) indicate that the proposed FATE model substantially outperforms the domain adaptation models, suggesting that FATE has a better domain generalizability for emotion recognition.
Yufeng Yin 0002, Liupei Lu, Zhi Xu 0013, Kaijie Cai, Jonathan Gratch, Mohammad Soleymani 0001
ACII1
2021 Self-Supervised Patch Localization for Cross-Domain Facial Action Unit Detection
abstract
Automatic detection of Facial Action Units (AUs) is a fundamental block for objective facial expression analysis. Computer vision-based detection of facial action units is susceptible to variations across corpora. To address this problem, we propose a novel architecture that can be jointly trained for self-supervised optical flow estimation, patch localization, supervised action unit detection, and adversarial domain adaptation. Patch localization allows the encoder to learn local features that are critical to detecting subtle changes caused by the presence of AUs in face. Specifically, an encoder-decoder architecture is used to estimate optical flow from every image. The optical flow is simultaneously used for AU detection, patch localization and adversarial domain adaptation. Majority of the existing work on facial expression analysis is evaluated within corpora. In this paper, we develop and evaluate this novel architecture for facial action unit detection across corpora. The experimental results indicate that our framework improves cross-domain performance (5.5 % F1-score on average), suggesting that the proposed patch localization guides the network to learn a more generalizable representation.
Yufeng Yin 0002, Liupei Lu, Yizhen Wu, Mohammad Soleymani 0001
FG1
2021 Subject-Invariant Eeg Representation Learning For Emotion Recognition
abstract
The discrepancies between the distributions of the train and test data, a.k.a., domain shift, result in lower generalization for emotion recognition methods. One of the main factors contributing to these discrepancies is human variability. Domain adaptation methods are developed to alleviate the problem of domain shift, however, these techniques while reducing between database variations fail to reduce between-subject variability. In this paper, we propose an adversarial deep domain adaptation approach for emotion recognition from electroencephalogram (EEG) signals. The method jointly learns a new representation that minimizes emotion recognition loss and maximizes subject confusion loss. We demonstrate that the proposed representation can improve emotion recognition performance within and across databases.
Soheil Rayatdoost, Yufeng Yin 0002, David Rudrauf, Mohammad Soleymani 0001
ICASSP2
2020 Supporting children's math learning with feedback-augmented narrative technology
abstract
A key challenge in education is effectively engaging children in learning activities. We investigated how a narrative story impacts engagement and learning, as well as how feedback can provide further benefits. To do so, we created an interactive, tablet-based learning platform with a multi-step math task designed using Common Core State Standards. Subjects completed a pretest and then were assigned to a condition, either one of three variations of the system (narratives, narratives with hints, and narratives with a tutoring chatbot using wizard-of-oz techniques) or a control system that has children complete the same learning task without narratives nor feedback, before the subjects completed a post test. 72 children in U.S. grades 3--5 participated. Our results showed that embedding learning activities into narratives boosted children's engagement as evaluated by coding video responses and surveys, and the integration of a tutoring chatbot improved learning outcomes on the assessment. These results provide evidence that a narrative-based tutoring system with chatbot-mediated help may support effective learning experiences for children.
Sherry Ruan, Jiayu He, Rui Ying, Jonathan Burkle, Dunia Hakim, Yufeng Yin 0002, Lily Zhou, Qianyao Xu, Abdallah A. AbuHashem, Griffin Dietz, Elizabeth L. Murnane, Emma Brunskill, James A. Landay
IDC7
2020 Speaker-Invariant Adversarial Domain Adaptation for Emotion Recognition
abstract
Automatic emotion recognition methods are sensitive to the variations across different datasets and their performance drops when evaluated across corpora. We can apply domain adaptation techniques e.g., Domain-Adversarial Neural Network (DANN) to mitigate this problem. Though the DANN can detect and remove the bias between corpora, the bias between speakers still remains which results in reduced performance. In this paper, we propose Speaker-Invariant Domain-Adversarial Neural Network (SIDANN) to reduce both the domain bias and the speaker bias. Specifically, based on the DANN, we add a speaker discriminator to unlearn information representing speakers' individual characteristics with a gradient reversal layer (GRL). Our experiments with multimodal data (speech, vision, and text) and the cross-domain evaluation indicate that the proposed SIDANN outperforms (+5.6% and +2.8% on average for detecting arousal and valence) the DANN model, suggesting that the SIDANN has a better domain adaptation ability than the DANN. Besides, the modality contribution analysis shows that the acoustic features are the most informative for arousal detection while the lexical features perform the best for valence detection.
Yufeng Yin 0002, Baiyu Huang, Yizhen Wu, Mohammad Soleymani 0001
ICMI1
2019 Understanding the Teaching Styles by an Attention based Multi-task Cross-media Dimensional Modeling
abstract
Teaching style plays an influential role in helping students to achieve academic success. In this paper, we explore a new problem of effectively understanding teachers' teaching styles. Specifically, we study 1) how to quantitatively characterize various teachers' teaching styles for various teachers and 2) how to model the subtle relationship between cross-media teaching related data (speech, facial expressions and body motions, content et al.) and teaching styles. Using the adjectives selected from more than 10,000 feedback questionnaires provided by an educational enterprise, a novel concept called Teaching Style Semantic Space (TSSS) is developed based on the pleasure-arousal dimensional theory to describe teaching styles quantitatively and comprehensively. Then a multi-task deep learning based model, Attention-based Multi-path Multi-task Deep Neural Network (AMMDNN), is proposed to accurately and robustly capture the internal correlations between cross-media features and TSSS. Based on the benchmark dataset, we further develop a comprehensive data set including 4,541 full-annotated cross-modality teaching classes. Our experimental results demonstrate that the proposed AMMDNN outperforms (+0.0842% in terms of the concordance correlation coefficient (CCC) on average) baseline methods. To further demonstrate the advantages of the proposed TSSS and our model, several interesting case studies are carried out, such as teaching styles comparison among different teachers and courses, and leveraging the proposed method for teaching quality analysis.
Suping Zhou, Jia Jia 0001, Yufeng Yin 0002, Xiang Li 0105, Zeyang Ye, Kehua Lei, Jialie Shen 0001
ACM Multimedia3
2019 Inferring Emotions From Large-Scale Internet Voice Data
abstract
As voice dialog applications (VDAs, e.g., Siri,11http://www.apple.com/ios/siri/. Cortana,22http://www.microsoft.com/en-us/mobile/campaign-cortana/. Google Now33http://www.google.com/landing/now/.) are increasing in popularity, inferring emotions from the large-scale internet voice data generated from VDAs can help give a more reasonable and humane response. However, the tremendous amounts of users in large-scale internet voice data lead to a great diversity of users accents and expression patterns. Therefore, the traditional speech emotion recognition methods, which mainly target acted corpora, cannot effectively handle the massive and diverse amount of internet voice data. To address this issue, we carry out a series of observations, find suitable emotion categories for large-scale internet voice data, and verify the indicators of the social attributes (query time, query topic, and users location) and emotion inferring. Based on our observations, two different strategies are employed to solve the problem. First, a deep sparse neural network model that uses acoustic information, textual information, and three indicators (a temporal indicator, descriptive indicator, and geo-social indicator) as the input is proposed. Then, to capture the contextual information, we propose a hybrid emotion inference model that includes long short-term memory to capture the acoustic features and a latent dirichlet allocation to extract text features. Experiments on 93 000 utterances collected from the Sogou Voice Assistant44http://yy.sogou.com. (Chinese Siri) validate the effectiveness of the proposed methodologies. Furthermore, we compare the two methodologies and give their advantages and disadvantages.
Jia Jia 0001, Suping Zhou, Yufeng Yin 0002, Boya Wu, Wei Chen 0071
IEEE Trans. Multim.3
2018 Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach
abstract
To give a more humanized response in Voice Dialogue Applications (VDAs), inferring emotion states from users’ queries may play an important role. However, in VDAs, we have tremendous amount of VDA users and massive scale of unlabeled data with high dimension features from multimodal information, which challenge the traditional speech emotion recognition methods. In this paper, to better infer emotion from conversational voice data, we proposed a semi-supervised multi-path generative neural network. Specifically, first, we build a novel supervised multi-path deep neural network framework. To avoid high dimensional input, raw features are trained by groups in local classifiers. Then high-level features of each local classifiers are concatenated as input of a global classifier. These two kinds classifiers are trained simultaneously through a single objective function to achieve a more effective and discriminative emotion inferring. To further solve the labeled-data-scarcity problem, we extend the multi-path deep neural network to a generative model based on semi-supervised variational autoencoder (semi-VAE), which is able to train the labeled and unlabeled data simultaneously. Experiment based on a 24,000 real-world dataset collected from Sogou Voice Assistant (SVAD13) and a benchmark dataset IEMOCAP show that our method significantly outperforms the existing state-of-the-art results.
Suping Zhou, Jia Jia 0001, Yufei Dong, Yufeng Yin 0002, Kehua Lei
AAAI5