EDBT 2026 Demo / reviewers in the wild / expert
Guanming Lu
dblp:72/6015
· DBLP profile ↗
21ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-4860-8229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neonatal Pain Speech Emotion Recognition Based on Multiscale Multilevel MAE NetworkabstractPerception of neonatal pain is a critical indicator for early-life health assessment. However, in real-world clinical scenarios, it faces challenges such as poor objectivity and limited monitoring methods. To enhance the modeling capacity and discriminative performance of neonatal pain-related speech emotion recognition, this article proposes a novel multiscale multilevel masked autoencoder network (MMMAEnet). Specifically, the original speech signal is divided into three temporal views: full-segment, first half segment, and second half segment, from which corresponding Mel-spectrograms are extracted as input. These inputs are processed in parallel by structurally consistent masked autoencoders to obtain multiscale contextual features. A fusion module is then employed to integrate the emotional representations across different time segments, enabling joint modeling of local dynamics and global semantics. To further optimize the structure of the feature space, it proposes a segmented alignment contrastive learning module. By constructing positive pairs between the full segment and each of the first half and second half segments, and negative pairs between the first half and second half segments, this module guides the model to better capture semantic consistency and differences among temporal segments. We conduct experimental verification using the neonatal pain speech (NPS) database and three adult speech emotion databases [CASIA Chinese Emotional Corpus (CASIA), Berlin Emotional Speech Database (EMO-DB), and Interactive Emotional Dyadic Motion Capture Database (IEMOCAP)]. The experimental results demonstrate the effectiveness of the proposed MMMAEnet approach in NPS emotion recognition and its good generalization ability on adult speech emotion databases. Jingjie Yan, Boyan Sun, Guanming Lu, Xianlan Zheng |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Cross-Hierarchical Multi-Head Sparse Vision Transformer Network for Neonatal Pain Facial Expression RecognitionabstractAutomated neonatal pain recognition and assessment based on deep learning is an emerging interdisciplinary topic that combines clinical pediatric medicine and affective computing. To improve the recognition accuracy of neonatal pain facial expressions, this paper proposes a Cross-hierarchical Multi-head Sparse Vision Transformer Network (CMS-ViT). Based on the Transformer architecture, the network presents a multi-head dynamic token sparsification fusion module, which performs dynamic feature selection and information fusion through three stages: selection, pairing, and fusion. This module sparsifies the tokens involved in computation to reduce redundancy. By inserting the sparsification module into different hierarchical layers, the network gradually reduces computational complexity. A cross-hierarchical feature fusion module is then embedded into the backbone to integrate semantic information from different depths, mitigating information loss caused by sparsification and fusion, and ultimately generating more discriminative feature representations by leveraging high-level semantic cues. In addition, the model utilizes pretrained parameters obtained via masked autoencoding on large-scale facial expression datasets, enhancing performance on the neonatal pain facial expression recognition task. Experimental results show that CMS-ViT achieves state-of-the-art (SOTA) performance on the Facial Expression of Neonatal Pain (FENP) dataset and demonstrates good generalization on the AffectNet and RAF-DB datasets. Jingjie Yan, Jiaming Jiang, Guanming Lu, Xianlan Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | BeFA: A General Behavior-driven Feature Adapter for Multimedia RecommendationabstractMultimedia recommender systems focus on utilizing behavioral information and content information to model user preferences. Typically, it employs pre-trained feature encoders to extract content features, then fuses them with behavioral features. However, pre-trained feature encoders often extract features from the entire content simultaneously, including excessive preference-irrelevant details.We speculate that it may result in the extracted features not containing sufficient features to accurately reflect user preferences. To verify our hypothesis, we introduce an attribution analysis method for visually and intuitively analyzing the content features. The results indicate that certain items’ content features exhibit the issues of information drift and information omission, reducing the expressive ability of features. Building upon this finding, we propose an effective and efficient general Behaviordriven Feature Adapter (BeFA) to tackle these issues. This adapter reconstructs the content feature with the guidance of behavioral information, enabling content features accurately reflecting user preferences. Extensive experiments demonstrate the effectiveness of the adapter across all multimedia recommendation methods. Qile Fan, Penghang Yu, Zhiyi Tan 0002, Bing-Kun Bao, Guanming Lu |
AAAI | 5 |
| 2025 | Mind Individual Information! Principal Graph Learning for Multimedia RecommendationabstractGraph Neural Network (GNN)-based methods have recently emerged as effective approaches for multimedia recommendation. Typically, these methods employ message passing on the user-item interaction graph, and model user preferences by exploiting co-occurrence patterns. Despite their effectiveness, we argue that they insufficiently exploit the individual information, potentially limiting recommendation performance. To validate our argument, we first analyze existing methods from spectral graph theory. We identify that existing methods focus on capturing global structural features, but underutilize local structural features that convey individual information. Further detailed experiments reveal that such an underutilization leads to overly similar user preferences modeling. Furthermore, we propose a novel Principal Graph Learning (PGL) framework to address this issue. The idea is to enhance user preference modeling by effectively mining and utilizing principal local structural features. PGL first extracts the principal subgraph from the user-item interaction graph using two novel extraction operators: global-aware and local-aware subgraph extraction. It then employs message passing on the principal subgraph to comprehensively model user perference, with the aim of simultaneously capturing co-occurrence patterns and individual information. Compared to existing methods, PGL achieves an average performance improvement of 9%. Penghang Yu, Zhiyi Tan 0002, Guanming Lu, Bing-Kun Bao |
AAAI | 3 |
| 2025 | Multi-Information Hierarchical Fusion Transformer with Local Alignment and Global Correlation for Micro-Expression Recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan, Dong Zhang 0018 |
ACM Multimedia | 3 |
| 2025 | Hypergraph denoising neural network for session-based recommendation
Zhiyi Tan 0002, Guanming Lu, Jinsheng Wei |
Appl. Intell. | 3 |
| 2025 | Adaptive discriminant feature learning for GNN-based session recommendation
Zhiyi Tan 0002, Guanming Lu, Jinsheng Wei |
Multim. Syst. | 3 |
| 2024 | Unsupervised image blind super resolution via real degradation feature learningabstractAbstract In recent years, many methods for image super‐resolution (SR) have relied on pairs of low‐resolution (LR) and high‐resolution (HR) images for training, where the degradation process is predefined by bicubic downsampling. While such approaches perform well in standard benchmark tests, they often fail to accurately replicate the complexity of real‐world image degradation. To address this challenge, researchers have proposed the use of unpaired image training to implicitly model the degradation process. However, there is a significant domain gap between the real‐world LR and the synthetic LR images from HR, which severely degrades the SR performance. A novel unsupervised image‐blind super‐resolution method that exploits degradation feature‐based learning for real‐image super‐resolution reconstruction (RDFL) is proposed. Their approach learns the degradation process from HR to LR using a generative adversarial network (GAN) and constrains the data distribution of the synthetic LR with real degraded images. The authors then encode the degraded features into a Transformer‐based SR network for image super‐resolution reconstruction through degradation representation learning. Extensive experiments on both synthetic and real datasets demonstrate the effectiveness and superiority of the RDFL method, which achieves visually pleasing reconstruction results. Guanming Lu |
IET Comput. Vis. | 2 |
| 2024 | Video-based neonatal pain expression recognition with cross-stream attention
Guanming Lu, Haoxia Chen, Jinsheng Wei, Xianlan Zheng, Hongyao Leng, Yimo Lou, Jingjie Yan |
Multim. Tools Appl. | 1 |
| 2024 | Learning discriminative features for micro-expression recognition
Guanming Lu, Jinsheng Wei, Jingjie Yan |
Multim. Tools Appl. | 1 |
| 2024 | Geometric Graph Representation With Learnable Graph Structure and Adaptive AU Constraint for Micro-Expression RecognitionabstractMicro-expression recognition (MER) holds significance in uncovering hidden emotions. Most works take image sequences as input and cannot effectively explore ME information because subtle ME-related motions are easily submerged in unrelated information. Instead, the facial landmark is a lowdimensional and compact modality, which achieves lower computational cost and potentially concentrates on ME-related movement features. However, the discriminability of facial landmarks for MER is unclear. Thus, this paper investigates the contribution of facial landmarks and proposes a novel framework to efficiently recognize MEs with facial landmarks. Firstly, a geometric twostream graph network is constructed to aggregate the low-order and high-order geometric movement information from facial landmarks to obtain discriminative ME representation. Secondly, a self-learning fashion is introduced to automatically model the dynamic relationship between nodes even long-distance nodes. Furthermore, an adaptive action unit loss is proposed to reasonably build a strong correlation between landmarks, facial action units and MEs. Notably, this work provides a novel idea with much higher efficiency to promote MER, only utilizing graphbased geometric features. The experimental results demonstrate that the proposed method achieves competitive performance with a significantly reduced computational cost. Furthermore, facial landmarks significantly contribute to MER and are worth further study for high-efficient ME analysis. Jinsheng Wei, Wei Peng 0009, Guanming Lu, Yante Li, Jingjie Yan, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Improving Graph Collaborative Filtering with Directional Behavior Enhanced Contrastive LearningabstractGraph Collaborative Filtering is a widely adopted approach for recommendation, which captures similar behavior features through Graph Neural Network (GNN). Recently, Contrastive Learning (CL) has been demonstrated as an effective method to enhance the performance of graph collaborative filtering. Typically, CL-based methods first perturb users’ history behavior data (e.g., drop clicked items), then construct a self-discriminating task for behavior representations under different random perturbations. However, for widely existing inactive users, random perturbation makes their sparse behavior information more incomplete, thereby harming the behavior feature extraction. To tackle the above issue, we design a novel directional perturbation-based CL method to improve the graph collaborative filtering performance. The idea is to perturb node representations through directionally enhancing behavior features. To do so, we propose a simple yet effective feedback mechanism, which fuses the representations of nodes based on behavior similarity. Then, to avoid irrelevant behavior preferences introduced by the feedback mechanism, we construct a behavior self-contrast task before and after feedback, to align the node representations between the final output and the first layer of GNN. Different from the widely adopted self-discriminating task, the behavior self-contrast task avoids complex message propagation on different perturbed graphs, which is more efficient than previous methods. Extensive experiments on three public datasets demonstrate that the proposed method has distinct advantages over other CL methods on recommendation accuracy. Penghang Yu, Bing-Kun Bao, Zhiyi Tan 0002, Guanming Lu |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | Multi-View Graph Convolutional Network for Multimedia RecommendationabstractMultimedia recommendation has received much attention in recent years. It models user preferences based on both behavior information and item multimodal information. Though current GCN-based methods achieve notable success, they suffer from two limitations: (1) Modality noise contamination to the item representations. Existing methods often mix modality features and behavior features in a single view (e.g., user-item view) for propagation, the noise in the modality features may be amplified and coupled with behavior features. In the end, it leads to poor feature discriminability; (2) Incomplete user preference modeling caused by equal treatment of modality features. Users often exhibit distinct modality preferences when purchasing different items. Equally fusing each modality feature ignores the relative importance among different modalities, leading to the suboptimal user preference modeling. Penghang Yu, Zhiyi Tan 0002, Guanming Lu, Bing-Kun Bao |
ACM Multimedia | 3 |
| 2023 | FENP: A Database of Neonatal Facial Expression for Pain AnalysisabstractIn this article, we introduce a new neonatal facial expression database for pain analysis. This database, called facial expression of neonatal pain (FENP), contains 11,000 neonatal facial expression images associated with 106 Chinese neonates from two children's hospitals, i.e., the Children's Hospital Affiliated to Nanjing Medical University and Second Affiliated Hospital Affiliated to Nanjing Medical University in China. The facial expression images cover four categories of facial expressions, i.e., severe pain expression, mild pain expression, crying expression and calmness expression, where each category contains 2750 neonatal facial expression images. Based on this database, we also investigate the pain facial expression recognition problem using several state-of-the-art facial expression features and expression recognition methods, such as Gabor+SVM, LBP+SVM, HOG+SVM, LBP+HOG+SVM, and several Convolutional Neural Network (CNN) methods (including AlexNet, VGGNet, GoogLeNet, ResNet and DenseNet). The experimental results indicate that the proposed neonatal pain facial expression database is very suitable for the study of both neonatal pain and facial expression recognition. Moreover, the FENP database is publicly available after signing a license agreement (the users can contact Jingjie Yan ([email protected]), Guanming Lu ([email protected])) or Xiaonan Li ([email protected]). Jingjie Yan, Guanming Lu, Wenming Zheng, Chengwei Huang, Zhen Cui 0001, Yuan Zong, Mengying Chen, Jindu Zhu, Haibo Li 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Learning two groups of discriminative features for micro-expression recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan, Yuan Zong |
Neurocomputing | 2 |
| 2022 | Micro-expression recognition using local binary pattern from five intersecting planes
Jinsheng Wei, Guanming Lu, Jingjie Yan, Huaming Liu |
Multim. Tools Appl. | 2 |
| 2021 | A comparative study on movement feature in different directions for micro-expression recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan |
Neurocomputing | 2 |
| 2016 | Sparse Kernel Reduced-Rank Regression for Bimodal Emotion Recognition From Facial Expression and SpeechabstractA novel bimodal emotion recognition approach from facial expression and speech based on the sparse kernel reduced-rank regression (SKRRR) fusion method is proposed in this paper. In this method, we use the openSMILE feature extractor and the scale invariant feature transform feature descriptor to respectively extract effective features from speech modality and facial expression modality, and then propose the SKRRR fusion approach to fuse the emotion features of two modalities. The proposed SKRRR method is a nonlinear extension of the traditional reduced-rank regression (RRR), where both predictor and response feature vectors in RRR are kernelized by being mapped onto two high-dimensional feature space via two nonlinear mappings, respectively. To solve the SKRRR problem, we propose a sparse representation (SR)-based approach to find the optimal solution of the coefficient matrices of SKRRR, where the introduction of the SR technique aims to fully consider the different contributions of training data samples to the derivation of optimal solution of SKRRR. Finally, we utilize the eNTERFACE '05 and AFEW 4.0 bimodal emotion database to conduct the experiments of monomodal emotion recognition and bimodal emotion recognition, and the results indicate that our presented approach acquires the highest or comparable bimodal emotion recognition rate among some state-of-the-art approaches. Jingjie Yan, Wenming Zheng, Qinyu Xu, Guanming Lu, Haibo Li 0001 |
IEEE Trans. Multim. | 4 |
| 2008 | Applying a novel combined classifier for pornographic web filtering in a grid computing environmentabstractAs the Web expands exponentially, there are a flood of pornographic Web sites on the Internet. Thus effective and fast web filtering systems are essential. Web filtering based on hypertext classification has become one of the important techniques to handle and filter inappropriate information on the Web. The task involved can be parallelized and distributed in a grid environment. However, how to improve the performance of the hypertext classification under the situation of noisy data is still a challenging problem. In this paper, we propose a new approach for hypertext classification in Web filtering, which uses a novel support vector machine and k-nearest neighbor (KNN-SVM) to remove noisy training examples. The task of text categorization is distributed in several computers. The experimental results show that the generalization performance in the accuracy of classification and the processing time are improved significantly compared to that of the traditional SVM classifier over the grid, and adapt to engineering applications. Zhong Gao, Guanming Lu, Danni Qin, Mei Qin |
CSCWD | 2 |
| 2008 | Recognition of Neonatal Facial Expressions of Acute Pain Using Boosted Gabor FeaturesabstractFacial expressions are considered a critical factor in neonatal pain assessment. This paper proposes a pain expression recognition method using boosted Gabor features. Each neonatal facial image is convoluted with the 2D Gabor filters to extract 412,160 Gabor features. Since the high-dimension Gabor feature vectors are quite redundant, we employs a modified version of AdaBoost algorithm to select and combine the most informative features for classification. The "pain vs. non-pain" problem is treated as two sub-problems by using a coarse-to-fine hierarchical classifier. Experiments with 510 neonatal expression images show that the proposed method is quite effective. Only 30 Gabor features are enough to achieve good classification performance. The recognition rate of pain versus non-pain is up to 88% (i.e. error rate isin=0.12). Compared with one existing algorithm for neonatal facial pain recognition,our approach can reach similar accuracy in much lower time complexity. Forrest Sheng Bao, Guanming Lu |
ICTAI (2) | 3 |
| 2008 | A novel risk assessment system for port state control inspectionabstractPort state control (PSC) inspection is the most important mechanism to ensure world marine safe. Recently, some SVM-based risk assessment systems have been presented in the world. They estimate the risk of each candidate ship based on its generic factors and history inspection factors to select high-risk one before conducting on-board PSC inspection. However, how to improve the performance of the PSC inspection under the situation of noisy data when applying SVM is still a challenging problem. In this paper, we propose a new approach for PSC inspection, which uses a novel support vector machine and k-nearest neighbor (KNN-SVM) to remove noisy training examples and Bag of Words (BW) to extract some new target factors for the PSC inspection database. The experimental results show that the generalization performance and the accuracy of risk assessment are improved significantly compared to that of the traditional SVM classifier, and adapt to engineering applications. Zhong Gao, Guanming Lu, Mengjue Liu |
ISI | 2 |