Yingjie Chen 0002

dblp:67/1326-2 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0003-2754-8909ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Trend-Aware Supervision: On Learning Invariance for Semi-supervised Facial Action Unit Intensity Estimation
abstract
With the increasing need for facial behavior analysis, semi-supervised AU intensity estimation using only keyframe annotations has emerged as a practical and effective solution to relieve the burden of annotation. However, the lack of annotations makes the spurious correlation problem caused by AU co-occurrences and subject variation much more prominent, leading to non-robust intensity estimation that is entangled among AUs and biased among subjects. We observe that trend information inherent in keyframe annotations could act as extra supervision and raising the awareness of AU-specific facial appearance changing trends during training is the key to learning invariant AU-specific features. To this end, we propose Trend-AwareSupervision (TAS), which pursues three kinds of trend awareness, including intra-trend ranking awareness, intra-trend speed awareness, and inter-trend subject awareness. TAS alleviates the spurious correlation problem by raising trend awareness during training to learn AU-specific features that represent the corresponding facial appearance changes, to achieve intensity estimation invariance. Experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, show the effectiveness of each kind of awareness. And under trend-aware supervision, the performance can be improved without extra computational or storage costs during inference.
Yingjie Chen 0002, Jiarui Zhang 0007, Tao Wang 0004, Yun Liang 0001
AAAI1
2023 EventFormer: AU Event Transformer for Facial Action Unit Event Detection
Yingjie Chen 0002, Jiarui Zhang 0007, Tao Wang 0004, Yun Liang 0001
BMVC1
2023 A Practical, Robust, Accurate Gaze-Based Intention Inference Method for Everyday Human-Robot Interaction
abstract
Gaze estimation is a crucial component of human-robot Interaction (HRI). While previous gaze estimation methods have been widely applied in advertising and gaming with head-mounted devices or complicated camera systems, little research has been conducted in everyday HRI scenarios. During interactions, robots have the potential to infer human intention through static gaze directions and dynamic eye movements, enabling them to behave more intelligently and friendly. This paper combines appearance-based gaze estimation methods with eye movement analysis methods to infer human intentions, particularly in human-robot interaction scenarios. Real interactions were conducted to test the accuracy and robustness of the methods developed. The experiments demonstrate that our methods deliver practical, robust, and accurate results.
Haoyang Xu, Tao Wang 0004, Yingjie Chen 0002, Tianze Shi
SMC3
2022 Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition
abstract
Subject-invariant facial action unit (AU) recognition remains challenging for the reason that the data distribution varies among subjects. In this paper, we propose a causal inference framework for subject-invariant facial action unit recognition. To illustrate the causal effect existing in AU recognition task, we formulate the causalities among facial images, subjects, latent AU semantic relations, and estimated AU occurrence probabilities via a structural causal model. By constructing such a causal diagram, we clarify the causal-effect among variables and propose a plug-in causal intervention module, CIS, to deconfound the confounder Subject in the causal diagram. Extensive experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, show the effectiveness of our CIS, and the model with CIS inserted, CISNet, has achieved state-of-the-art performance.
Yingjie Chen 0002, Diqi Chen, Tao Wang 0004, Yizhou Wang 0001, Yun Liang 0001
AAAI1
2022 On Mitigating Hard Clusters for Face Clustering
Yingjie Chen 0002, Huasong Zhong, Chong Chen 0002, Chen Shen 0003, Jianqiang Huang 0001, Tao Wang 0004, Yun Liang 0001, Qianru Sun
ECCV (12)1
2022 Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image Retrieval
abstract
In this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios, these methods usually either generate an inaccurate similarity matrix without reflection of similarity ranking or suffer from the violation of the underlying assumption in contrastive learning, resulting in limited retrieval performance. To tackle this issue, we propose a novel method termed HAMAN, which explores semantics from a fine-grained view to enhance the ability of multi-label image retrieval. In particular, we reconstruct the pairwise similarity structure by matching fine-grained patch features generated by the pre-trained neural network, serving as reliable guidance for similarity preserving of hash codes. Moreover, a novel conditional contrastive learning on hash codes is proposed to adopt self-supervised learning in multi-label scenarios. According to extensive experiments on three multi-label datasets, the proposed method outperforms a broad range of state-of-the-art methods.
Zeyu Ma 0001, Xiao Luo 0001, Yingjie Chen 0002, Mi-Xiao Hou, Jinxing Li 0003, Minghua Deng, Guangming Lu 0002
IJCAI3
2022 Pursuing Knowledge Consistency: Supervised Hierarchical Contrastive Learning for Facial Action Unit Recognition
abstract
With the increasing need for emotion analysis, facial action unit (AU) recognition has attracted much more attention as a fundamental task for affective computing. Although deep learning has boosted the performance of AU recognition to a new level in recent years, it remains challenging to extract subject-consistent representations since the appearance changes caused by AUs are subtle and ambiguous among subjects. We observe that there are three kinds of inherent relations among AUs, which can be treated as strong prior knowledge, and pursuing the consistency of such knowledge is the key to learning subject-consistent representations. To this end, we propose a supervised hierarchical contrastive learning method (SupHCL) for AU recognition to pursue knowledge consistency among different facial images and different AUs, which is orthogonal to methods focusing on network architecture design. Specifically, SupHCL contains three relation consistency modules, i.e., unary, binary, and multivariate relation consistency modules, which take the corresponding kind of inherent relations as extra supervision to encourage knowledge-consistent distributions of both AU-level and image-level representations. Experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, demonstrate the effectiveness of each relation consistency module and the superiority of SupHCL.
Yingjie Chen 0002, Chong Chen 0002, Xiao Luo 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Tao Wang 0004, Yun Liang 0001
ACM Multimedia1
2021 AUPro: Multi-label Facial Action Unit Proposal Generation for Sequence-Level Analysis
Yingjie Chen 0002, Jiarui Zhang 0007, Diqi Chen, Tao Wang 0004, Yizhou Wang 0001, Yun Liang 0001
ICONIP (3)1
2021 CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit Recognition
abstract
Facial action unit (AU) recognition has attracted increasing attention due to its indispensable role in affective computing, especially in the field of affective human-computer interaction. Due to the subtle and transient nature of AU, it is challenging to capture the delicate and ambiguous motions in local facial regions among consecutive frames. Considering that context is essential to resolve ambiguity in human visual system, modeling context within or among facial images emerges as a promising approach for AU recognition task. To this end, we propose CaFGraph, a novel context-aware facial multi-graph that can model both morphological & muscular-based region-level local context and region-level temporal context. CaFGraph is the first work to construct a universal facial multi-graph structure that is independent of both task settings and dataset statistics for almost all fine-grained facial behavior analysis tasks, including but not limited to AU recognition. To make full use of the context, we then present CaFNet that learns context-aware facial graph representations via CaFGraph from facial images for multi-label AU recognition. Experiments on two widely used benchmark datasets, BP4D and DISFA, demonstrate the superiority of our CaFNet over the state-of-the-art methods.
Yingjie Chen 0002, Diqi Chen, Yizhou Wang 0001, Tao Wang 0004, Yun Liang 0001
ACM Multimedia1