EDBT 2026 Demo / reviewers in the wild / expert
Yante Li
dblp:178/5018
· DBLP profile ↗
21ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-4824-2044ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards consistent and controllable image synthesis for identity-preserving face editingabstractFace editing involves modifying facial attributes like expression, head pose, or lighting, with the goal of preserving the subject’s unique identity features. Diffusion models have recently emerged as the dominant approach in visual generation, driven by their strong generative power. However, challenges persist in the realm of face editing, where independently and correctly editing target attributes while preserving high-fidelity identity information remains a formidable problem. In this paper, we present RigFace, a novel framework that combines controllable signals derived from a 3D Morphable Model (3DMM) with a fine-tuned Stable Diffusion (SD) model. Our basic idea to achieve by leveraging disentangled facial attributes provided by 3DMM and harnessing the strong generative capacity of Stable Diffusion. Specifically, our method contains: 1) A Spatial Attribute Encoder that provides robust and decoupled conditions of background, pose, expression and lighting; 2) A FaceFusion module that transfers identity information at different resolutions from the Identity Encoder to the Denoising UNet of a pre-trained SD model through self-attention, which facilitates detailed identity preservation. Our model achieves superior performance in both identity preservation and photorealism compared to existing face editing models. Mengting Wei, Tuomas Varanka, Yante Li, Xingxun Jiang, Huai-Qian Khor, Guoying Zhao 0001 |
Pattern Recognit. | 3 |
| 2026 | LatentMag: Self-Supervised 3D Magnification for Micro Expressions via Latent ExtrapolationabstractMicro-expressions (MEs) are subtle and brief facial movements that reveal genuine emotional states but are often imperceptible due to their low intensity. While motion magnification has proven effective for enhancing ME visibility in 2D settings, its extension to 3D remains largely unexplored. In this work, we presentLatentMag, the first controllable 3D micro-expression magnification framework. Unlike traditional editing methods that rely on fixed labels or expression targets, our approach models expression intensity as a relative, input-dependent signal. We adopt registered 3D meshes as our representation, enabling vertex-level correspondence and interpretable displacement analysis. To guide magnification, we introduce a geometric prior that models amplification as a spatially adaptive transformation, where the change in pairwise distance between points on the output mesh scales with that observed between the input shapes, ensuring natural, localized deformation. We operationalize this prior in a generative framework by disentangling a latent intensity code, whose extrapolation drives controllable shape amplification. Trained in a self-supervised manner using unlabeled mesh sequences, LatentMag generalizes well to unseen identities and expressions, offering a novel solution that bridges geometric interpretability with realistic 3D expression modeling. Mengting Wei, Xingxun Jiang, Haoyu Chen 0001, Yante Li, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change DetectionabstractIn environmental protection, tree monitoring plays an essential role in maintaining and improving ecosystem health. However, precise monitoring is challenging because existing datasets fail to capture continuous fine-grained changes in trees due to low-resolution images and high acquisition costs. In this paper, we introduce UAVTC, a large-scale, long-term, high-resolution dataset collected using UAVs equipped with cameras, specifically designed to detect individual Tree Changes (TCs). UAVTC includes rich annotations and statistics based on biological knowledge, offering a fine-grained view for tree monitoring. To address environmental influences and effectively model the hierarchical diversity of physiological TCs, we propose a novel Hyperbolic Siamese Network (HSN) for TC detection, enabling compact and hierarchical representations of dynamic tree changes. Extensive experiments show that HSN can effectively capture complex hierarchical changes and provide a robust solution for fine-grained TC detection. In addition, HSN generalizes well to cross-domain face anti-spoofing task, highlighting its broader significance in AI. We believe our work, combining ecological insights and interdisciplinary expertise, will benefit the community by offering a new benchmark and innovative AI technologies. Source code is available on https://github.com/liyantett/Tree-Changes-Detection-with-Siamese-Hyperbolic-network. Yante Li, Hanwen Qi, Haoyu Chen 0001, Xinlian Liang, Guoying Zhao 0001 |
CVPR | 1 |
| 2025 | Behavior-prompted Learning with Tree Attention for Advanced Facial Action Unit DetectionabstractAction Unit (AU) is a systematic coding of facial behaviors that plays a crucial role in facial expression recognition. AU detection faces significant challenges due to the fine-grained categorical differences and coexistence at varying intensities of the AU. To address these challenges, we propose a refined behavior-prompt AU detection model featuring a coarse-to-fine tree attention mechanism. Specifically, we introduce a learnable behavior-prompt approach that utilizes large vision-language models, harnessing their powerful semantic representation capabilities to encompass comprehensive prior knowledge of AU behaviors. Besides, considering AUs’ diverse intensities and interactive nature, a coarse-to-fine tree attention module is customized to capture the fine-grained details of individual AUs and their long-range dependencies. To further mitigate vision-text bias, a feature interaction learning strategy is employed that progressively incorporates context-related visual information into prompts and decouples AU-specific representations. Extensive experiments demonstrate that our proposed method achieves state-of-the-art results on two widely used benchmarks, BP4D and DISFA. Our code is avaliable at https://github.com/ColinHaoZou/FAUD-CLIP. Yante Li |
IJCNN | 3 |
| 2025 | Facial age estimation by group-centric feature learning
Qilu Zhao, Yante Li |
Expert Syst. Appl. | 2 |
| 2024 | Clothing Sampling Based on Active Learning For Cloth-Changing Person Re-identificationabstractCloth-Changing Person Re-Identification (CC-ReID) aims to match the same person with clothing changes. The challenges mainly include two types: same person wearing different clothing and different person wearing similar clothing. The current methods are usually limited by the number and variation of clothing in training data, making it difficult to cope with the latter. To address this issue, this article proposes a clothing sampler (CS) based on active learning. The main idea is actively selecting valuable clothing images, which can ensure both "clothing diversity with the same identity" and "identity diversity with the similar clothing" in batches, forcing the model to learn features that are independent of clothing. In addition, a multi-clothing loss (MC) is also designed to guide the network to learn clothing-independent features. Experiment results on two cloth-changing datasets show the effectiveness of our proposed CS. Jiansen Jing, Yante Li, Guoying Zhao 0001 |
ICME | 4 |
| 2024 | Data Leakage and Evaluation Issues in Micro-Expression AnalysisabstractMicro-expressions have drawn increasing interest lately due to various potential applications. The task is, however, difficult as it incorporates many challenges from the fields of computer vision, machine learning and emotional sciences. Due to the spontaneous and subtle characteristics of micro-expressions, the available training and testing data are limited, which make evaluation complex. We show that data leakage and fragmented evaluation protocols are issues among the micro-expression literature. We find that fixing data leaks can drastically reduce model performance, in some cases even making the models perform similarly to a random classifier. To this end, we go through common pitfalls, propose a new standardized evaluation protocol using facial action units with over 2000 micro-expression samples, and provide an open source library that implements the evaluation protocols in a standardized manner. Code is publicly available inhttps://github.com/tvaranka/meb. Tuomas Varanka, Yante Li, Wei Peng 0009, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Geometric Graph Representation With Learnable Graph Structure and Adaptive AU Constraint for Micro-Expression RecognitionabstractMicro-expression recognition (MER) holds significance in uncovering hidden emotions. Most works take image sequences as input and cannot effectively explore ME information because subtle ME-related motions are easily submerged in unrelated information. Instead, the facial landmark is a lowdimensional and compact modality, which achieves lower computational cost and potentially concentrates on ME-related movement features. However, the discriminability of facial landmarks for MER is unclear. Thus, this paper investigates the contribution of facial landmarks and proposes a novel framework to efficiently recognize MEs with facial landmarks. Firstly, a geometric twostream graph network is constructed to aggregate the low-order and high-order geometric movement information from facial landmarks to obtain discriminative ME representation. Secondly, a self-learning fashion is introduced to automatically model the dynamic relationship between nodes even long-distance nodes. Furthermore, an adaptive action unit loss is proposed to reasonably build a strong correlation between landmarks, facial action units and MEs. Notably, this work provides a novel idea with much higher efficiency to promote MER, only utilizing graphbased geometric features. The experimental results demonstrate that the proposed method achieves competitive performance with a significantly reduced computational cost. Furthermore, facial landmarks significantly contribute to MER and are worth further study for high-efficient ME analysis. Jinsheng Wei, Wei Peng 0009, Guanming Lu, Yante Li, Jingjie Yan, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Interactions for Socially Shared Regulation in Collaborative Learning: An Interdisciplinary Multimodal DatasetabstractSocially shared regulation plays a pivotal role in the success of collaborative learning. However, evaluating socially shared regulation of learning (SSRL) proves challenging due to the dynamic and infrequent cognitive and socio-emotional interactions, which constitute the focal point of SSRL. To address this challenge, this article gathers interdisciplinary researchers to establish a multimodal dataset with cognitive and socio-emotional interactions for SSRL study. Firstly, to induce cognitive and socio-emotional interactions, learning science researchers designed a special collaborative learning task with regulatory trigger events among triadic people for the SSRL study. Secondly, this dataset includes various modalities like video, Kinect data, audio, and physiological data (accelerometer, EDA, heart rate) from 81 high school students in 28 groups, offering a comprehensive view of the SSRL process. Thirdly, three-level verbal interaction annotations and nonverbal interactions including facial expression, eye gaze, gesture, and posture are provided, which could further contribute to interdisciplinary fields such as computer science, sociology, and education. In addition, comprehensive analysis verifies the dataset’s effectiveness. As far as we know, this is the first multimodal dataset for studying SSRL among triadic group members. Yante Li, Yang Liu 0182, Andy Nguyen, Henglin Shi, Eija Vuorenmaa, Sanna Järvelä, Guoying Zhao 0001 |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2023 | Facial Micro-Expressions: An OverviewabstractMicro-expression (ME) is an involuntary, fleeting, and subtle facial expression. It may occur in high-stake situations when people attempt to conceal or suppress their true feelings. Therefore, MEs can provide essential clues to people’s true feelings and have plenty of potential applications, such as national security, clinical diagnosis, and interrogations. In recent years, ME analysis has gained much attention in various fields due to its practical importance, especially automatic ME analysis in computer vision as MEs are difficult to process by naked eyes. In this survey, we provide a comprehensive review of ME development in the field of computer vision, from the ME studies in psychology and early attempts in computer vision to various computational ME analysis methods and future directions. Four main tasks in ME analysis are specifically discussed, including ME spotting, ME recognition, ME action unit detection, and ME generation in terms of the approaches, advance developments, and challenges. Through this survey, readers can understand MEs in both aspects of psychology and computer vision, and apprehend the future research direction in ME analysis. Guoying Zhao 0001, Yante Li, Matti Pietikäinen |
Proc. IEEE | 3 |
| 2023 | 4DME: A Spontaneous 4D Micro-Expression Dataset With MultimodalitiesabstractMicro-expressions (ME) are a special form of facial expressions which may occur when people try to hide their true feelings for some reasons. MEs are important clues to reveal people’s true feelings, but are difficult or impossible to be captured by ordinary persons with naked-eyes as they are very short and subtle. It is expected that robust computer vision methods can be developed to automatically analyze MEs which requires lots of ME data. The current ME datasets are insufficient, and mostly contain only one single form of 2D color videos. Researches on 4D data of ordinary facial expressions have prospered, but so far no 4D data is available in ME study. In the current study, we introduce the 4DME dataset: a new spontaneous ME dataset which includes 4D data along with three other video modalities. Both micro- and macro-expression clips are labeled out in 4DME, and 22 AU labels and five categories of emotion labels are annotated. Experiments are carried out using three 2D-based methods and one 4D-based method to provide baseline results. The results indicate that the 4D data can potentially benefit ME recognition. The 4DME dataset could be used for developing 4D-based approaches, or exploring fusion of multiple video sources (e.g., texture and depth) for the task of ME analysis in future. Besides, we also emphasize the importance of forming a clear and unified criteria of ME annotation for future ME data collection studies. Several key questions related with ME annotation are listed and discussed in depth, especially about the relationship between AUs and ME emotion categories. A preliminary AU-Emo mapping table is proposed with justified explanations and supportive experimental results. Several unsolved issues are also summarized for future work. Shiyang Cheng 0001, Yante Li, Muzammil Behzad, Jie Shen 0008, Stefanos Zafeiriou, Maja Pantic, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Graph-Based Facial Affect Analysis: A ReviewabstractAs one of the most important affective signals, facial affect analysis (FAA) is essential for developing human-computer interaction systems. Early methods focus on extracting appearance and geometry features associated with human affects while ignoring the latent semantic information among individual facial changes, leading to limited performance and generalization. Recent work attempts to establish a graph-based representation to model these semantic relationships and develop frameworks to leverage them for various FAA tasks. This paper provides a comprehensive review of graph-based FAA, including the evolution of algorithms and their applications. First, the FAA background knowledge is introduced, especially on the role of the graph. We then discuss approaches widely used for graph-based affective representation in literature and show a trend towards graph construction. For the relational reasoning in graph-based FAA, existing studies are categorized according to their non-deep or deep learning methods, emphasizing the latest graph neural networks. Performance comparisons of the state-of-the-art graph-based FAA methods are also summarized. Finally, we discuss the challenges and potential directions. As far as we know, this is the first survey of graph-based FAA methods. Our findings can serve as a reference for future research in this field. Yang Liu 0182, Xingming Zhang 0001, Yante Li, Jinzhao Zhou, Xin Li 0116, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | MEGC2022: ACM Multimedia 2022 Micro-Expression Grand ChallengeabstractFacial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. Unfortunately, the small sample problem severely limits the automation of ME analysis. Furthermore, due to the brief and subtle nature of ME, ME spotting is a challenging task, and the performance is still not satisfactory yet. This challenge focuses on two tasks, i.e., the micro- and macro-expression spotting task, and the ME Generation task. Jingting Li 0001, Moi Hoon Yap, Wen-Huang Cheng, John See, Xiaopeng Hong, Adrian K. Davison, Yante Li, Zizhao Dong |
ACM Multimedia | 9 |
| 2022 | Deep Learning for Micro-Expression Recognition: A SurveyabstractMicro-expressions (MEs) are involuntary facial movements revealing people's hidden feelings in high-stake situations and have practical importance in various fields. Early methods for Micro-expression Recognition (MER) are mainly based on traditional features. Recently, with the success of Deep Learning (DL) in various tasks, neural networks have received increasing interest in MER. Different from macro-expressions, MEs are spontaneous, subtle, and rapid facial movements, leading to difficult data collection and annotation, thus publicly available datasets are usually small-scale. Currently, various DL approaches have been proposed to solve the ME issues and improve MER performance. In this survey, we provide a comprehensive review of deep MER and define a new taxonomy for the field encompassing all aspects of MER based on DL, including datasets, each step of the deep MER pipeline, and performance comparisons of the most influential methods. The basic approaches and advanced developments are summarized and discussed for each aspect. Additionally, we conclude the remaining challenges and potential directions for the design of robust MER systems. Finally, ethical considerations in MER are discussed. To the best of our knowledge, this is the first survey of deep MER methods, and this survey can serve as a reference point for future MER research. Yante Li, Jinsheng Wei, Yang Liu 0182, Janne Kauttonen, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Micro-expression Action Unit Detection with Dual-view Attentive Similarity-Preserving Knowledge DistillationabstractEncoding facial expressions via action units (AUs) has been found to be effective in resolving the ambiguity issue among different expressions. Therefore, AU detection plays an important role for emotion analysis. While a number of AU detection methods have been proposed for common facial expressions, there is very limited study for micro-expression AU detection. Micro-expression AU detection is challenging because of the weakness of micro-expression appearance and the spontaneous characteristic leading to difficult collection, thus has small-scale datasets. In this paper, we focus on the micro-expression AU detection and expect to contribute to the community. To address above issues, a novel dual-view attentive similarity-preserving distillation method is proposed for robust micro-expression AU detection by leveraging massive facial expressions in the wild. Through such an attentive similarity-preserving distillation method, we break the domain shift problem and essential AU knowledge from common facial AUs is efficiently distilled. Furthermore, considering that the generalization ability of teacher network is important for knowledge distillation, a semi-supervised co-training approach is developed to construct a generalized teacher network for learning discriminative AU representation. Extensive experiments have demonstrated that our proposed knowledge distillation method can effectively distill and transfer the cross-domain knowledge for robust micro-expression AU detection. Yante Li, Wei Peng 0009, Guoying Zhao 0001 |
FG | 1 |
| 2021 | Intra- and Inter-Contrastive Learning for Micro-expression Action Unit DetectionabstractEncoding facial expressions via Action Units (AUs) has been found effective for resolving the ambiguity issue among different expressions. In the literature, AU detection has extensive researches in macro-expressions. However, there is limited research about AU analysis for micro-expressions (MEs). Micro-expression Action Unit (MEAU) detection becomes a challenging problem because of the subtle facial motion. To alleviate this problem, in this paper, we study the contrastive learning for modeling subtle AUs and propose a novel MEAU detection method by learning the intra- and inter-contrastive information among MEs. Through the intra-contrastive learning module, the difference between the onset and apex frames is enlarged and utilized to obtain the discriminative representation for low-intensity AU detection. In addition, considering the subtle difference between MEAUs, the inter-contrastive learning is designed to automatically explore and enlarge the difference between different AUs to enhance the MEAU detection robustness. Intensive experiments on two widely used ME databases have demonstrated the effectiveness and generalization ability of our proposed method. Yante Li, Guoying Zhao 0001 |
ICMI | 1 |
| 2021 | Micro-expression action unit detection with spatial and channel attentionabstractAction Unit (AU) detection plays an important role in facial behaviour analysis. In the literature, AU detection has extensive researches in macro-expressions. However, to the best of our knowledge, there is limited research about AU analysis for micro-expressions. In this paper, we focus on AU detection in micro-expressions. Due to the small quantity and low intensity of micro-expression databases, micro-expression AU detection becomes challenging. To alleviate these problems, in this work, we propose a novel micro-expression AU detection method by utilizing self high-order statistics of spatio-wise and channel-wise features which can be considered as spatial and channel attentions, respectively. Through such spatial attention module, we expect to utilize rich relationship information of facial regions to increase the AU detection robustness on limited micro-expression samples. In addition, considering the low intensity of micro-expression AUs, we further propose to explore high-order statistics for better capturing subtle regional changes on face to obtain more discriminative AU features. Intensive experiments show that our proposed approach outperforms the basic framework by 0.0859 on CASME II, 0.0485 on CASME, and 0.0644 on SAMM in terms of the average F1-score. Yante Li, Xiaohua Huang 0003, Guoying Zhao 0001 |
Neurocomputing | 1 |
| 2021 | Joint Local and Global Information Learning With Single Apex Frame Detection for Micro-Expression RecognitionabstractMicro-expressions (MEs) are rapid and subtle facial movements that are difficult to detect and recognize. Most recent works have attempted to recognize MEs with spatial and temporal information from video clips. According to psychological studies, the apex frame conveys the most emotional information expressed in facial expressions. However, it is not clear how the single apex frame contributes to micro-expression recognition. To alleviate that problem, this paper firstly proposes a new method to detect the apex frame by estimating pixel-level change rates in the frequency domain. With frequency information, it performs more effectively on apex frame spotting than the currently existing apex frame spotting methods based on the spatio-temporal change information. Secondly, with the apex frame, this paper proposes a joint feature learning architecture coupling local and global information to recognize MEs, because not all regions make the same contribution to ME recognition and some regions do not even contain any emotional information. More specifically, the proposed model involves the local information learned from the facial regions contributing major emotion information, and the global information learned from the whole face. Leveraging the local and global information enables our model to learn discriminative ME representations and suppress the negative influence of unrelated regions to MEs. The proposed method is extensively evaluated using CASME, CASME II, SAMM, SMIC, and composite databases. Experimental results demonstrate that our method with the detected apex frame achieves considerably promising ME recognition performance, compared with the state-of-the-art methods employing the whole ME sequence. Moreover, the results indicate that the apex frame can significantly contribute to micro-expression recognition. Yante Li, Xiaohua Huang 0003, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Can Micro-Expression be Recognized Based on Single Apex Frame?abstractMicro-expressions are rapid and subtle facial movements such that they are difficult to detect and recognize.Most of recent works have attempted to recognize micro-expression by using the spatial and dynamic information from the video clip.Physiological studies have demonstrated that the apex frame can convey the most emotion expressed in facial expression.It may be reasonable to use apex frame for improving micro-expression recognition.However, it is wonder how much apex frames contribute to micro-expression recognition.In this paper, we primarily focus on resolving the contribution-level by using apex frame for micro-expression recognition.Firstly, we propose a new method to detect the apex frame in frequency domain, as it is found that apex frame has very correlated relationship with the amplitude change in frequency domain.Secondly, we propose to use deep convolutional neural network (DCNN) on apex frame to recognize micro-expression.Intensive experimental results on CASME II database shows that our method has achieved considerably improvement compared with the state-of-the-art methods in micro-expression recognition.These results also demonstrate that apex frame can express the major emotion in micro-expression. Yante Li, Xiaohua Huang 0003, Guoying Zhao 0001 |
ICIP | 1 |
| 2016 | Cross-scenario clothing retrieval and fine-grained style recognitionabstractIn this paper, we propose a new approach for cross-scenario clothing retrieval and fine-grained clothing style recognition. The query clothing photos captured by cameras or other mobile devices are filled with noisy background while the product clothing images online for shopping are usually presented in a pure environment. We tackle this problem by two steps. Firstly, a hierarchical super-pixel merging algorithm based on semantic segmentation is proposed to obtain the intact query clothing item. Secondly, aiming at solving the problem of clothing style recognition in different scenarios, we propose sparse coding based on domain-adaptive dictionary learning to improve the accuracy of the classifier and adaptability of the dictionary. In this way, we obtain fine-grained attributes of the clothing items and use the attributes matching score to re-rank the retrieval results further. The experiment results show that our method outperforms the state-of-the-art approaches. Furthermore, we build a well labeled clothing dataset, where the images are selected from 1.5 billion product clothing images. Zongmin Li, Yante Li, Yunping Pang, Yujie Liu 0002 |
ICPR | 2 |
| 2015 | A More Effective Method for Image Representation: Topic Model Based on Latent Dirichlet AllocationabstractNowadays, the Bag-of-words (BoW) representation is well applied to recent state-of-the-art image retrieval works. However, with the rapid growth in the number of images, the dimension of the dictionary increases substantially which leads to great storage and CPU cost. Besides, the local features do not convey any semantic information which is very important in image retrieval. In this paper, we propose to use "topics" instead of "visual words" as the image representation by topic model to reduce the feature dimension and mine more high level semantic information. We call this as Bag-of-Topics (BoT) which is a type of statistical model for discovering the abstract "topics" from the words. We extract the topics by Latent Dirichlet Allocation (LDA) and calculate the similarity between the images using BoT model instead of BoW directly. The results show that the dimension of the image representation has been reduced significantly, while the retrieval performance is improved. Zongmin Li, Yante Li, Zhenzhong Kuang, Yujie Liu 0002 |
CAD/Graphics | 3 |