Jingjie Yan

dblp:89/10303 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Domain-Category Fusion Guided Diffusion Model for cross-dataset facial expression recognition
Jingjie Yan, Yuebo Yue, Jinsheng Wei, Jianguo Hu
Comput. Vis. Image Underst.1
2026 Multimodal emotion recognition based on temporal-spatio bidirectional dilated causal convolution multi-head attention network
Jingjie Yan, Guangkun Shi, Xiaoyang zhou
Inf. Sci.1
2026 Cross-database facial expression recognition based on Spatio-temporal Feature Point Attention Deep Transfer Network
Jingjie Yan, Yuebo Yue, Jingsheng Wei
Multim. Tools Appl.1
2026 Neonatal Pain Speech Emotion Recognition Based on Multiscale Multilevel MAE Network
abstract
Perception of neonatal pain is a critical indicator for early-life health assessment. However, in real-world clinical scenarios, it faces challenges such as poor objectivity and limited monitoring methods. To enhance the modeling capacity and discriminative performance of neonatal pain-related speech emotion recognition, this article proposes a novel multiscale multilevel masked autoencoder network (MMMAEnet). Specifically, the original speech signal is divided into three temporal views: full-segment, first half segment, and second half segment, from which corresponding Mel-spectrograms are extracted as input. These inputs are processed in parallel by structurally consistent masked autoencoders to obtain multiscale contextual features. A fusion module is then employed to integrate the emotional representations across different time segments, enabling joint modeling of local dynamics and global semantics. To further optimize the structure of the feature space, it proposes a segmented alignment contrastive learning module. By constructing positive pairs between the full segment and each of the first half and second half segments, and negative pairs between the first half and second half segments, this module guides the model to better capture semantic consistency and differences among temporal segments. We conduct experimental verification using the neonatal pain speech (NPS) database and three adult speech emotion databases [CASIA Chinese Emotional Corpus (CASIA), Berlin Emotional Speech Database (EMO-DB), and Interactive Emotional Dyadic Motion Capture Database (IEMOCAP)]. The experimental results demonstrate the effectiveness of the proposed MMMAEnet approach in NPS emotion recognition and its good generalization ability on adult speech emotion databases.
Jingjie Yan, Boyan Sun, Guanming Lu, Xianlan Zheng
IEEE Trans. Comput. Soc. Syst.1
2026 Cross-Hierarchical Multi-Head Sparse Vision Transformer Network for Neonatal Pain Facial Expression Recognition
abstract
Automated neonatal pain recognition and assessment based on deep learning is an emerging interdisciplinary topic that combines clinical pediatric medicine and affective computing. To improve the recognition accuracy of neonatal pain facial expressions, this paper proposes a Cross-hierarchical Multi-head Sparse Vision Transformer Network (CMS-ViT). Based on the Transformer architecture, the network presents a multi-head dynamic token sparsification fusion module, which performs dynamic feature selection and information fusion through three stages: selection, pairing, and fusion. This module sparsifies the tokens involved in computation to reduce redundancy. By inserting the sparsification module into different hierarchical layers, the network gradually reduces computational complexity. A cross-hierarchical feature fusion module is then embedded into the backbone to integrate semantic information from different depths, mitigating information loss caused by sparsification and fusion, and ultimately generating more discriminative feature representations by leveraging high-level semantic cues. In addition, the model utilizes pretrained parameters obtained via masked autoencoding on large-scale facial expression datasets, enhancing performance on the neonatal pain facial expression recognition task. Experimental results show that CMS-ViT achieves state-of-the-art (SOTA) performance on the Facial Expression of Neonatal Pain (FENP) dataset and demonstrates good generalization on the AffectNet and RAF-DB datasets.
Jingjie Yan, Jiaming Jiang, Guanming Lu, Xianlan Zheng
IEEE Trans. Circuits Syst. Video Technol.1
2025 Multi-Information Hierarchical Fusion Transformer with Local Alignment and Global Correlation for Micro-Expression Recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan, Dong Zhang 0018
ACM Multimedia4
2025 Cross-database facial expression recognition based on Multi-feature Representation Multi-layer Domain Adaptive Fusion Network
Jingjie Yan, Chengkun Du
Eng. Appl. Artif. Intell.1
2024 Video-based neonatal pain expression recognition with cross-stream attention
Guanming Lu, Haoxia Chen, Jinsheng Wei, Xianlan Zheng, Hongyao Leng, Yimo Lou, Jingjie Yan
Multim. Tools Appl.8
2024 Learning discriminative features for micro-expression recognition
Guanming Lu, Jinsheng Wei, Jingjie Yan
Multim. Tools Appl.4
2024 Geometric Graph Representation With Learnable Graph Structure and Adaptive AU Constraint for Micro-Expression Recognition
abstract
Micro-expression recognition (MER) holds significance in uncovering hidden emotions. Most works take image sequences as input and cannot effectively explore ME information because subtle ME-related motions are easily submerged in unrelated information. Instead, the facial landmark is a lowdimensional and compact modality, which achieves lower computational cost and potentially concentrates on ME-related movement features. However, the discriminability of facial landmarks for MER is unclear. Thus, this paper investigates the contribution of facial landmarks and proposes a novel framework to efficiently recognize MEs with facial landmarks. Firstly, a geometric twostream graph network is constructed to aggregate the low-order and high-order geometric movement information from facial landmarks to obtain discriminative ME representation. Secondly, a self-learning fashion is introduced to automatically model the dynamic relationship between nodes even long-distance nodes. Furthermore, an adaptive action unit loss is proposed to reasonably build a strong correlation between landmarks, facial action units and MEs. Notably, this work provides a novel idea with much higher efficiency to promote MER, only utilizing graphbased geometric features. The experimental results demonstrate that the proposed method achieves competitive performance with a significantly reduced computational cost. Furthermore, facial landmarks significantly contribute to MER and are worth further study for high-efficient ME analysis.
Jinsheng Wei, Wei Peng 0009, Guanming Lu, Yante Li, Jingjie Yan, Guoying Zhao 0001
IEEE Trans. Affect. Comput.5
2023 FENP: A Database of Neonatal Facial Expression for Pain Analysis
abstract
In this article, we introduce a new neonatal facial expression database for pain analysis. This database, called facial expression of neonatal pain (FENP), contains 11,000 neonatal facial expression images associated with 106 Chinese neonates from two children's hospitals, i.e., the Children's Hospital Affiliated to Nanjing Medical University and Second Affiliated Hospital Affiliated to Nanjing Medical University in China. The facial expression images cover four categories of facial expressions, i.e., severe pain expression, mild pain expression, crying expression and calmness expression, where each category contains 2750 neonatal facial expression images. Based on this database, we also investigate the pain facial expression recognition problem using several state-of-the-art facial expression features and expression recognition methods, such as Gabor+SVM, LBP+SVM, HOG+SVM, LBP+HOG+SVM, and several Convolutional Neural Network (CNN) methods (including AlexNet, VGGNet, GoogLeNet, ResNet and DenseNet). The experimental results indicate that the proposed neonatal pain facial expression database is very suitable for the study of both neonatal pain and facial expression recognition. Moreover, the FENP database is publicly available after signing a license agreement (the users can contact Jingjie Yan ([email protected]), Guanming Lu ([email protected])) or Xiaonan Li ([email protected]).
Jingjie Yan, Guanming Lu, Wenming Zheng, Chengwei Huang, Zhen Cui 0001, Yuan Zong, Mengying Chen, Jindu Zhu, Haibo Li 0001
IEEE Trans. Affect. Comput.1
2022 Learning two groups of discriminative features for micro-expression recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan, Yuan Zong
Neurocomputing3
2022 Micro-expression recognition using local binary pattern from five intersecting planes
Jinsheng Wei, Guanming Lu, Jingjie Yan, Huaming Liu
Multim. Tools Appl.3
2021 A comparative study on movement feature in different directions for micro-expression recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan
Neurocomputing3
2016 Sparse Kernel Reduced-Rank Regression for Bimodal Emotion Recognition From Facial Expression and Speech
abstract
A novel bimodal emotion recognition approach from facial expression and speech based on the sparse kernel reduced-rank regression (SKRRR) fusion method is proposed in this paper. In this method, we use the openSMILE feature extractor and the scale invariant feature transform feature descriptor to respectively extract effective features from speech modality and facial expression modality, and then propose the SKRRR fusion approach to fuse the emotion features of two modalities. The proposed SKRRR method is a nonlinear extension of the traditional reduced-rank regression (RRR), where both predictor and response feature vectors in RRR are kernelized by being mapped onto two high-dimensional feature space via two nonlinear mappings, respectively. To solve the SKRRR problem, we propose a sparse representation (SR)-based approach to find the optimal solution of the coefficient matrices of SKRRR, where the introduction of the SR technique aims to fully consider the different contributions of training data samples to the derivation of optimal solution of SKRRR. Finally, we utilize the eNTERFACE '05 and AFEW 4.0 bimodal emotion database to conduct the experiments of monomodal emotion recognition and bimodal emotion recognition, and the results indicate that our presented approach acquires the highest or comparable bimodal emotion recognition rate among some state-of-the-art approaches.
Jingjie Yan, Wenming Zheng, Qinyu Xu, Guanming Lu, Haibo Li 0001
IEEE Trans. Multim.1
2012 Sparse 2-D Canonical Correlation Analysis via Low Rank Matrix Approximation for Feature Extraction
abstract
Although 2-D canonical correlation analysis (2DCCA) has been proposed to reduce the computational complexity while reserving local data structure of image, the learned canonical variables of 2DCCA are the linear combination of all the original variables, which makes it hard to interpret the solutions and might have less generality. In this paper, we propose a sparse 2-D canonical correlation analysis (S2DCCA) to solve the drawbacks of the 2DCCA method and apply it to image feature extraction. The basic idea of S2DCCA is to impose two lasso penalties on the objective function of 2DCCA to obtain two sets of sparse projection directions via low rank matrix approximation. We conduct extensive experiments on both FERET and AR databases to evaluate the performance of the proposed method.
Jingjie Yan, Wenming Zheng
IEEE Signal Process. Lett.1