EDBT 2026 Demo / reviewers in the wild / expert
Tong Chen 0008
dblp:22/1512-8
· DBLP profile ↗
29ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0003-3805-4138ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Look inside nodes: A novel intranode attention mechanism for graph attention networks
Yingjuan Jia, Tong Chen 0008, Xinyu Liu 0007, Hanpu Wang |
Pattern Recognit. | 2 |
| 2026 | Micro-expression recognition based on dataset balance and local connected bi-branch network
Hanpu Wang, Fuyuan Luo, Ju Zhou, Xinyu Liu 0007, Haolin Xia, Tong Chen 0008 |
Signal Process. Image Commun. | 6 |
| 2025 | Weighted Spatiotemporal Feature and Multi-task Learning for Masked Facial Expression Recognition
Shiwei He, Yingjuan Jia, Hanpu Wang, Xinyu Liu 0007, Jianmeng Zhou, Huijie Gu, Tong Chen 0008 |
CVM (1) | 8 |
| 2025 | Emotion-Qwen-VL: A Fully Fine-Tuned Multimodal Large Language Model for Micro-Expression Visual Question AnsweringabstractThis paper presents our solution for the Micro-Expression Visual Question Answering (ME-VQA) task in the 2025 Facial Micro-Expression Grand Challenge (MEGC). To address the limitations of traditional micro-expression recognition (MER) methods in dynamic modeling, semantic interpretation, and natural language interaction, we propose Emotion-Qwen-VL, a fully fine-tuned multimodal large language model tailored for micro-expression understanding. Specifically, we construct a structured, instruction-based QA dataset that reformulates emotion categories and action unit (AU) annotations into natural language QA pairs, covering classification, AU detection, and causal reasoning. We then adopt a full-parameter fine-tuning strategy to guide Qwen2.5-VL in learning fine-grained temporal facial dynamics and their emotional semantics. Experimental results on the MEGC 2025 test set demonstrate that Emotion-Qwen-VL outperforms strong baselines such as Qwen2.5-VL and QVQ across multiple dimensions, including coarse-grained macro-expression classification, fine-grained micro-expression classification, and language generation. Our results highlight the effectiveness, interpretability, and adaptation potential of large models in micro-expression understanding. The code is available at: https://github.com/2308623956/MEGC2025. Ruotong Fang, Zhiyuan Han, Xiaoqing Lin, Yuhao Shan, Tong Chen 0008 |
ACM Multimedia | 7 |
| 2025 | A cross-database micro-expression recognition framework based on meta-learning
Hanpu Wang, Ju Zhou, Xinyu Liu 0007, Yingjuan Jia, Tong Chen 0008 |
Appl. Intell. | 5 |
| 2025 | Facial StO2-based personal identification: dataset construction, feasibility study, and recognition framework
Zheyuan Zhang 0007, Xinyu Liu 0007, Yingjuan Jia, Ju Zhou, Hanpu Wang, Jiaxiu Wang, Tong Chen 0008 |
Appl. Intell. | 7 |
| 2025 | Masked facial expression recognition based on temporal overlap module and action unit graph convolutional network
Zheyuan Zhang 0007, Bingtong Liu, Ju Zhou, Hanpu Wang, Xinyu Liu 0007, Tong Chen 0008 |
J. Vis. Commun. Image Represent. | 7 |
| 2025 | Human emotion and StO2: Dataset, pattern, and recognition of basic emotions
Xinyu Liu 0007, Tong Chen 0008, Ju Zhou, Hanpu Wang, Guangyuan Liu 0005, Xiaolan Fu |
Pattern Recognit. | 2 |
| 2024 | KF-TMIAF: An Efficient Key-Frames-Based Temporal Modeling of Images and AU Features structure for Masked Facial Expression RecognitionabstractMasked Facial Expression refers to facial expression that people deliberately make which do not correspond to their true emotions. The purpose of masked facial expression is to conceal one’s genuine feelings. Precisely identifying masked facial expression aids in revealing the genuine emotions, and finds applications in diagnosis of mental illness. However, the masked facial expression sequences have redundant information, temporal modeling is difficult, and the expressive ability of AU features needs to be improved. To solve the existing problems, we first propose a Key frame Selection and Data Augmentation (KS-DA) method to obtain the key frame in the sequence. In addition, we propose a Global Feature Relation Block (GFR) to aggregate global information and combine it with Temporal Convolutional Networks (TCN) to better realize temporal modeling. Finally, we propose an Adaptive Weight Generation Module (Ada-WGM) to form different weights from adaptation, allowing for better extraction of the temporal changes in AU features. Finally, the experiments have shown that the proposed method has improved the mixed expression classification task (36 task) by 15.23% and the experienced emotions classification task (6E task) by 17.77%, respectively, and obtain sort results. Huijie Gu, Hanpu Wang, Shiwei He, Jianmeng Zhou, Tong Chen 0008 |
BIBM | 6 |
| 2024 | Facial StO2-based Stress Recognition using Automatic Graph Generation and Dual-Stream GNNabstractThe recognition of stress responses in health, resilience, and psychopathology holds significant scientific importance. The biological information carried by the facial tissue oxygen saturation (StO2) is a useful indicator for stress recognition. Although graph-based methods have achieved state-of-the-art (SOTA), there is room for improvement. This study proposes an automatic graph generation method for facial StO2, replacing manual methods and enhancing the information of graphs. To further enhance the capability of graphs’ representation, a dimensionality reduction strategy for nodes is proposed. The strategy is implemented based on an novel objective function that can increase the inter-class distance and reduce the intra-class distance. After obtaining highly representative graphs, a dual-stream network integrating GAT and GCN is designed to extract stress-related features from the graphs, ultimately achieving the SOTA recognition accuracy. Yingjuan Jia, Hanpu Wang, Xinyu Liu 0007, Tong Chen 0008 |
BIBM | 4 |
| 2024 | IBFNet: A Dual Auxiliary Branch Network for Multimodal Hidden Emotion RecognitionabstractMicro-expression (ME) is a crucial cue to reveal hidden emotions and helps to diagnose mental illnesses such as depression. However, their rapid and subtle characteristics make them difficult to recognize, and relying on a single ME feature is insufficient to capture comprehensive information about hidden emotions. This paper proposes an inverted bottleneck fusion network (IBFNet), which combines ME and facial tissue oxygen saturation (StO2) for multimodal feature fusion to improve the recognition of hidden emotions. Specifically, IBFNet learns different representations of each modal feature, namely common and individual features. The modalities are then mapped to different subspaces for feature refinement. We design the recurrent cross-modal attention (RCA) module to explore commonalities across modalities, retaining modality-specific information through auxiliary branches to enhance diversity. Finally, the inverted bottleneck structure is used to fuse the common and individual feature of ME and StO2. Experimental results demonstrate a recognition accuracy of 90.47%, which surpasses the limitations of single-modality recognition. Jianmeng Zhou, Xinyu Liu 0007, Shiwei He, Huijie Gu, Tong Chen 0008 |
BIBM | 6 |
| 2024 | Microexpression to Macroexpression: Facial Expression Magnification by Single InputabstractMicroexpressions are expressions that people inadvertently express, and therefore often represent a person’s true emotion. However, because it has a low intensity and a short duration, it is hard to be recognized correctly. In this paper, we propose a deep learning magnification method to generate macroexpressions from a single microexpression image. In the first stage, we extract the expression information from a single microexpression image. Then, We combine the idea of cyclegan and optical flow consistency to model the extracted expression features as the optical flow field between the neutral face and microexpressions. To extract a reliable optical flow field from the expression information, we design an optical flow refiner. In the second stage, we adopt an encoder-decoder network and let it learn to magnify the optical flow. Finally, the magnified optical flow guided the microexpression images to generate macroexpression images. We compare our single input based network with current two-frames-input based networks. The results show that our method performs better, even in wild images. We fed our magnified images directly into a simple ResNet18 network for recognition, achieving a competitive score under the MEGC2019 standard, compared with recent complex recognition networks. Yaqi Song, Tong Chen 0008, Shigang Li 0001, Jianfeng Li 0003 |
ICRA | 2 |
| 2024 | ULME-GAN: a generative adversarial network for micro-expression sequence generation
Ju Zhou, Sirui Sun, Haolin Xia, Xinyu Liu 0007, Hanpu Wang, Tong Chen 0008 |
Appl. Intell. | 6 |
| 2024 | Seeing Through the Mask: Recognition of Genuine Emotion Through Masked Facial ExpressionabstractThe purpose of facial expression recognition is to recognize the corresponding emotions. However, people tend to hide their emotions by displaying facial expressions that differ from those evoked by emotions. These inconsistent facial expressions are referred to as masked facial expressions (MFEs). The automatic recognition of hidden emotions within an MFE using image data is challenging. In this study, we find distinctive movement patterns in the facial action units (AUs) of MFE sequences through a detailed analysis. Considering our findings, we propose handcrafted features called dynamic AU intensity features (DAIFs) to represent AU movement. Furthermore, we develop a decoupled AU transformer (DAUT) model for recognition, where the decoupled convolution operators ensure that the temporal information in the DAIF is not damaged. To further improve the recognition performance, we design self-supervised clip prediction for pretraining of DAUT. Experimental results demonstrate that our proposed method performs exceptionally well across all tasks in the MFE dataset, particularly improving accuracy by nearly double on the most challenging 36-class task. This suggests that leveraging temporal information from facial AU movements is a reliable and effective technique for recognizing MFEs. Ju Zhou, Xinyu Liu 0007, Hanpu Wang, Zheyuan Zhang 0007, Tong Chen 0008, Xiaolan Fu, Guangyuan Liu 0005 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | Long Video Micro-expression Spotting Based On OCC TheoryabstractThe generation of micro-expression (ME) is unconscious and unaffected by human subjective consciousness. It can reveal hidden emotions in humans, which has important applications in the domains of psychological diagnosis, lie detection, and so on. ME detection, which aims to locate ME apex frames in long videos, is the foundation of ME recognition. In this paper, we propose a ME detection method based on One-Class Classification (OCC) theory for long videos. The method is the first to use neutral expression frames to detect MEs, which addresses the problem of insufficient training samples in ME detection. We design a geometric motion feature based on the prior knowledge from the FACS and utilize it in conjunction with the features extracted by CNN to train separate one-class classifiers for ME detection. Finally, we evaluate the proposed method on publicly available datasets CASME I, CASME II, SAMM, and a mixed database comprising data from the three datasets, to verify the feasibility of the proposed method in both single-database and cross-database ME detection. The experimental results show that our method outperforms three other ME detection methods, demonstrating its ability to more accurately locate ME apex frames in long videos. Bingtong Liu, Zheyuan Zhang 0007, Ju Zhou, Hanpu Wang, Tong Chen 0008 |
BIBM | 5 |
| 2023 | Stress recognition based on graph structure representation of facial StO2abstractStress is an integral state that affects physical and mental health of individuals. Tissue Oxygen Saturation (StO2) is an emerging physiological signal and can reflect different stress states. Previous studies have simply extracted features of StO2 from specific regions of the face without exploring potential associations between them. In this paper, we analyze the facial Region of Interest (ROI) based on StO2, so as to explore the deep relationship between ROIs. Building upon this, we construct, for the first time, a StO2-based facial graph structure for stress recognition by using ROIs as nodes. In addition, we propose a new graph pooling method called FTPool which allows to measure node significance in terms of both node characteristics and graph topology. Finally, we further propose a dual-stream network (GCNet) combining GNN and CNN for the baseline-independent stress recognition. The experimental results show that the proposed GCNet can achieve state-of-the-art (SOTA) results with an accuracy of 78.57% on the original unbalanced database. Jiaxiu Wang, Xinyu Liu 0007, Yingjuan Jia, Tong Chen 0008, Zheyuan Zhang 0007 |
BIBM | 4 |
| 2023 | Baseline-independent stress classification based on facial StO2
Xinyu Liu 0007, Ju Zhou, Tong Chen 0008 |
Appl. Intell. | 4 |
| 2022 | Facial StO2: A New Promising Biometric IdentityabstractIn this paper, we introduce a new biometric identity, facial tissue oxygen saturation (StO2). StO2 is an index of blood oxygen content in tissues and is related to blood vessel distribution pattern and metabolic rate. Experimental results show that classification accuracy can reach 83.33% in 42 participants with different stress states by using StO2 as the only input to the ResNet-50 model. We also proposed a module called StO2Net to eliminate the effects of stress on classification. The highest accuracy can reach up to 90.48% when the module is used. This pilot study shows that facial StO2 can be a promising biometric feature for identity recognition. Dairong Peng, Sirui Sun, Xinyu Liu 0007, Ju Zhou, Tong Chen 0008 |
BIBM | 5 |
| 2022 | Micro-Expression Recognition by Combining Progressive-Learning Intensity Magnification with Self-Attention-Convolution ClassificationabstractHow to detect the subtle intensity change and find the inherent relation of feature maps in a face image is the two key issues for micro-expression recognition. In this study, in order to cope with the subtle intensity change, we design an innovative Cascade Micro-Expression Magnification Network (Cascade-MEMN) to magnify the low amplitude of the intensity change for learning the intensity progressively, which helps suppress overlapping artifacts and produce more realistic magnification. In order to find the inherent relation of facial feature maps, we design a Self-attention-based Convolution Layer (SCL), which introduces self-attention to every pixel when calculating the weights in the sliding window, considering that the variation of micro-expressions is very subtle and local on face. Concretely, the SCLs are used to replace 3 Resblocks from the last stage of ResNet50 with 3 SCL blocks. Finally, a new micro-expression recognition system is realized by combining the progressive-learning-based intensity magnification network with the modified self-attention-convolution classification network. Consequently, the proposed method achieves competitive results for the composite database evaluation (CDE) protocol from MEGC 2019, as shown in the experimental result. Ling Lei 0003, Tong Chen 0008, Shigang Li 0001, Jianfeng Li 0003 |
IJCB | 3 |
| 2021 | Outlier Detection for Spotting Micro-expressionsabstractFacial expression, as a basic communication method, is an important way of emotion expression and cognition. Facial emotional expression impairment seriously affects interpersonal communication and social life. Micro-expressions (MEs) are involuntary and instant facial dynamics that occurs when the subject failed to suppress their genuine emotions, especially in high-stake situations. Psychological research has shown that MEs can reflect people’s true emotions, which is of great help to the treatment of Facial emotional expression impairment. ME spotting aims to locate the apex frame positions of MEs from long videos, which is the first step in ME analysis. Unlike previous researches that used binary classification or maximum feature difference for analysis, in this paper, we apply the idea of outlier detection to spot ME for the first time. MEs are unusual facial dynamics whose movement patterns diverge from others, so they can be regarded as outliers in the feature space of long videos. Our proposed method uses Gaussian model to estimate the probability density function and locates outliers by analyzing the statistical features of long videos to achieve ME spotting. This method was evaluated on CASME I, CASME II and SAMM datasets that only include spontaneous MEs in long videos. The results show that this method can efficiently locate apex frames of ME efficiently in long videos and also provide a new perspective for ME spotting. Ranlei Cao, Xinyu Liu 0007, Ju Zhou, Dairong Peng, Tong Chen 0008 |
BIBM | 6 |
| 2021 | Holoscopic 3D Microgesture Recognition by Deep Neural Network Model Based on Viewpoint Images and Decision FusionabstractFinger microgestures have been widely used in human computer interaction (HCI), particularly for interactive applications, such as virtual reality (VR) and augmented reality (AR) technologies, to provide immersive experience. However, traditional 2D image-based microgesture recognition suffers from low accuracy due to the limitations of 2D imaging sensors, which have no depth information. In this article, we proposed an innovative 3D microgesture recognition system based on a holoscopic 3D imaging sensor. Due to the lack of holoscopic 3D datasets, a comprehensive holoscopic 3D microgesture (HoMG) database is created and used to develop a robust 3D microgesture recognition method. Then, a fast algorithm is proposed to extract multiviewpoint images from one holoscopic image. Furthermore, we applied a CNN model with an attention-based residual block to each viewpoint image to improve the algorithm performance. Finally, bagging classification tree decision-level fusion is applied to combine the predictions. The experimental results demonstrate that the proposed method outperforms state-of-the-art methods and delivers a better accuracy than existing methods. Yi Liu 0061, Min Peng 0004, Mohammad Rafiq Swash, Tong Chen 0008, Hongying Meng |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2020 | Facial Expression Recognition Method Based on a Part-Based Temporal Convolutional Network with a Graph-Structured Representation
Changmin Bai, Jianfeng Li 0003, Tong Chen 0008, Shigang Li 0001 |
ICANN (1) | 4 |
| 2020 | A Novel Graph-TCN with a Graph Structured Representation for Micro-expression RecognitionabstractFacial micro-expressions (MEs) recognition has attracted much attention recently. However, because MEs are spontaneous, subtle and transient, recognizing MEs is a challenge task. In this paper, first, we use transfer learning to apply learning-based video motion magnification to magnify MEs and extract the shape information, aiming to solve the problem of the low muscle movement intensity of MEs. Then, we design a novel graph-temporal convolutional network (Graph-TCN) to extract the features of the local muscle movements of MEs. First, we define a graph structure based on the facial landmarks. Second, the Graph-TCN deals with the graph structure in dual channels with a TCN block. One channel is for node feature extraction, and the other one is for edge feature extraction. Last, the edges and nodes are fused for classification. The Graph-TCN can automatically train the graph representation to distinguish MEs while not using a hand-crafted graph representation. To the best of our knowledge, we are the first to use the learning-based video motion magnification method to extract the features of shape representations from the intermediate layer while magnifying MEs. Furthermore, we are also the first to use deep learning to automatically train the graph representation for MEs. Ling Lei 0003, Jianfeng Li 0003, Tong Chen 0008, Shigang Li 0001 |
ACM Multimedia | 3 |
| 2020 | Micro-attention for micro-expression recognition
Min Peng 0004, Tao Bi, Tong Chen 0008 |
Neurocomputing | 4 |
| 2019 | A Novel Apex-Time Network for Cross-Dataset Micro-Expression RecognitionabstractThe automatic recognition of micro-expression has been boosted ever since the successful introduction of deep learning approaches. As researchers working on such topics are moving to learn from the nature of micro-expression, the practice of using deep learning techniques has evolved from processing the entire video clip of micro-expression to the recognition on apex frame. Using the apex frame is able to get rid of redundant video frames, but the relevant temporal evidence of micro-expression would be thereby left out. This paper proposes a novel Apex-Time Network (ATNet)to recognize micro-expression based on spatial information from the apex frame as well as on temporal information from the respective-adjacent frames. Through extensive experiments on three benchmarks, we demonstrate the improvement achieved by learning such temporal information. Specially, the model with such temporal information is more robust in cross-dataset validations. Min Peng 0004, Tao Bi, Yu Shi 0003, Tong Chen 0008 |
ACII | 6 |
| 2019 | A Graph-Structured Representation with BRNN for Static-based Facial Expression RecognitionabstractFacial expression is controlled by facial muscle and can be considered as appearance and geometric variation of key parts. One key challenging issue of static-based facial expression recognition is to capture effective information from a single facial image. In this paper, we propose a graph representation with Bidirectional RNN (BRNN) for static-based facial expression recognition. Each node on the graph represents appearance information around the facial landmarks. Edges represent the geometric information encoded by the distance between two nodes. A bidirectional recurrent neural network utilized to process the graph extracts the appearance and geometric representation. The final representation from BRNN is fed into a fully connected layer and a Softmax layer to infer expressions. Experimental results show that this method achieves significant improvements over the state-of-art methods on three widely used facial databases (Oulu-CASIA, CK+, and MMI), and our method reduces the error rates of the previous best methods by 42.2%, 35.9% and 18.7%, respectively. Changmin Bai, Jianfeng Li 0003, Tong Chen 0008, Shigang Li 0001, Yiguang Liu |
FG | 4 |
| 2018 | Attention Based Residual Network for Micro-Gesture RecognitionabstractFinger micro-gesture recognition is increasingly become an important part of human-computer interaction (HCI) in applications of augmented reality (AR) and virtual reality (VR) technologies. To push the boundary of microgesture recognition, a novel Holoscopic 3D Micro-Gesture Database (HoMG) was established for research purpose. HoMG has an image subset and a video subset. This paper is to demonstrate the result achieved on the image subset for Holoscopic Micro-Gesture Recognition Challenge 2018 (HoMGR 2018). The proposed method utilized the state-of-the-art residual network with an attention-involved design. In every block of the network, an attention branch is added to the output of the last convolution layer. The attention branch is designed to spotlight the finger micro-gesture and reduce the noise introduced from the wrist and background. With an extensive analysis on HoMG, the proposed model achieved a recognition accuracy of 80.5% on the validation set and 82.1% on the testing set. Min Peng 0004, Tong Chen 0008 |
FG | 3 |
| 2018 | From Macro to Micro Expression Recognition: Deep Learning on Small Datasets Using Transfer LearningabstractThis paper presents the methods used in our submission to 2018 Facial Micro-Expression Grand Challenge (MEGC). The object of the challenge is to recognize micro-expression in two provided databases, including holdout-database recognition and composite database recognition. Considering the small size of the databases, we follow a rout of transfer learning to implement convolutional neural network to recognize the micro-expression. ResNet10 pre-trained on ImageNet dataset was fine-tuned on macro-expression datasets with large size and then on the provided micro-expression datasets. Experimental results show that the method can achieve weighted average recall (WAR) of 0.561 and unweighted average recall (UAR) of 0.389 in Holdout-database Evaluation Task, and F1 Score of 0.64 in Composite Database Evaluation Task, which are much higher than what baseline methods (LBP-TOP, HOOF, HOG3D) can achieve. Min Peng 0004, Zhan Wu, Tong Chen 0008 |
FG | 4 |
| 2014 | Detection of Psychological Stress Using a Hyperspectral Imaging TechniqueabstractThe detection of stress at early stages is beneficial to both individuals and communities. However, traditional stress detection methods that use physiological signals are contact-based and require sensors to be in contact with test subjects for measurement. In this paper, we present a method to detect psychological stress in a non-contact manner using a human physiological response. In particular, we utilize a hyperspectral imaging (HSI) technique to extract the tissue oxygen saturation (StO2) value as a physiological feature for stress detection. Our experimental results indicate that this new feature may be independent from perspiration and ambient temperature. Trier Social Stress Tests (TSSTs) on 21 volunteers demonstrated a significant difference $p\< 0.005$ and a large practical discrimination (d 1/4 1.37) between normalized baseline and stress StO2 levels. The accuracy for stress recognition from baseline using a binary classifier was 76.19 and 88.1 percent for the automatic and manual selections of the classifier threshold, respectively. These results suggest that the StO2 level could serve as a new modality to recognize stress at standoff distances. Tong Chen 0008, Peter W. T. Yuen, Mark A. Richardson 0001, Guangyuan Liu 0005, Zhishun She |
IEEE Trans. Affect. Comput. | 1 |