VLDB 2026 Research / reviewers in the wild / expert
Pei Geng
dblp:328/9161
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Image recognition and object detection · 48% Video understanding and tracking · 40% Graph learning · 13% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
human-object interaction detection |
1.9 | 2 | 2026 | Language-Driven Visual Data Generation for Zero-Shot HOI Detection · IEEE Trans. Image Process. 2026 HORP: Human-Object Relation Priors Guided HOI Detection · CVPR 2025 |
Computer vision › Image recognition and object detection › human-object interaction detection
zero-shot human-object interaction detection |
1.0 | 1 | 2026 | Language-Driven Visual Data Generation for Zero-Shot HOI Detection · IEEE Trans. Image Process. 2026 |
Computer vision › Video understanding and tracking › activity recognition
interaction recognition |
0.9 | 1 | 2025 | HORP: Human-Object Relation Priors Guided HOI Detection · CVPR 2025 |
Computer vision › Video understanding and tracking
action recognition |
0.8 | 1 | 2024 | Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action Recognition · IEEE Trans. Multim. 2024 |
Machine learning › Graph learning
graph neural network |
0.8 | 1 | 2024 | Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action Recognition · IEEE Trans. Multim. 2024 |
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition |
0.8 | 1 | 2024 | Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action Recognition · IEEE Trans. Multim. 2024 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.0transformer · 1.0text-to-vision adapter · 1.0large language model · 1.0relation prior · 0.9multimodal fusion · 0.9multi-scale temporal convolution · 0.8contrastive learning · 0.8attention fusion · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language-Driven Visual Data Generation for Zero-Shot HOI DetectionabstractZero-shot human-object interaction (HOI) detection aims to recognize both seen and unseen interaction categories while detecting humans and objects in an image. However, due to the absence of training samples for unseen categories, existing methods often overfit on seen HOIs and struggle to generalize to unseen ones. To address this issue, we introduce a novel Language-Driven Visual Data Generation (LD-VDG) approach that generates pseudo visual features from textual semantics of unseen HOIs. This provides an innovative solution enabling generalization to unseen HOIs without relying on visual samples. Specifically, we first design a text-to-vision (T-V) adapter to align HOI text and visual features, trained on seen HOIs with paired image-text data. For unseen HOIs, we guide the large language model to produce multiple fine-grained textual descriptions based on HOI labels, which are then encoded by the vision-language model and transformed into pseudo visual features via the T-V adapter. After that, these pseudo features together with real features from seen HOIs are jointly used to train a transformer-based HOI detector. In this way, our method enables effective recognition of unseen HOIs by leveraging language-driven visual representations. Experimental results on standard datasets demonstrate that the proposed LD-VDG outperforms previous methods. In particular, it achieves superior performance on unseen categories under various zero-shot settings. Pei Geng, Shanshan Zhang 0001, Jian Yang 0003 |
IEEE Trans. Image Process. | 1 |
| 2025 | HORP: Human-Object Relation Priors Guided HOI DetectionabstractHuman-Object Interaction (HOI) detection aims to predict thetriplets, where the core challenge lies in recognizing the interaction of each human-object pair. Despite recent progress thanks to more advanced model architectures, HOI performance remains unsatisfactory. In this work, we first perform some failure analysis and find that the accuracy of the no-interaction category is extremely low, largely hindering the improvement of overall performance. We further look into the error types and find the mis-classification between no-interaction and with-interaction ones can be handled by human-object relation priors. Specifically, to better distinguish no-interaction from direct interactions, we propose 3D location prior, which indicates the distance between human and object; as of no-interaction vs. indirect interactions, we propose gaze area prior, which denotes whether human can see the object or not. The above two types of human-object relation priors are represented by text and are combined with the original visual features, generating multi-modal cues for interaction recognition. Experimental results on the HICO-DET and V-COCO datasets demonstrate that our proposed human-object relation priors are effective and our method HORP surpasses previous methods under various settings and scenarios. In particular, the usage of our priors significantly enhances the model’s recognition ability for the no-interaction category. Code is available at https://github.com/namegp/HORP. Pei Geng, Jian Yang 0003, Shanshan Zhang 0001 |
CVPR | 1 |
| 2025 | Sparse and Dense: Learning Confusion Representation Network for 3-D Action RecognitionabstractIn the field of the Internet of Medical Things (IoMT), the demand for Human action recognition (HAR) is growing. Due to the limitations of portability and privacy of traditional sensors, many endeavors have made significant progress in 3-D skeleton-based action recognition. However, existing methods ignore potential higher-order semantic information between joints and fail to perceive rapidly changing dynamic details, resulting in frequent confusion of actions with similar motion trajectories. To alleviate this issue, we propose a novel learning confusion representation network (LCR-Net). Specifically, a progressive feature enhancement module is first designed to utilize self-attention to gradually aggregate lower-order features to higher-order features, emphasizing the relative movement between body parts. Second, we design the enhanced spatio-temporal convolution to explore the potential spatio-temporal dependencies between joints by adding a mask matrix and an attention fusion mechanism. To further perceive the spatio-temporal relationships in subtle changes, we divide the interactive sparse-dense pathways at different spatio-temporal resolutions and enhance the complementary information between the two pathways through feature interaction. Finally, the frequency excitation learning module is proposed to efficiently learn the importance of different frequencies by cross-channel modeling, promoting the compactness of actions within classes and the separability of confusion actions. In addition, the lightweight LCR-Net${}^{\textbf {+}}$is achieved through model compression optimization to meet the deployment requirements of IoT systems. Comprehensive experiments conducted on three public datasets (NTU-RGB+D60&120, NW-UCLA) demonstrate the superior performance of our model. Xinran Hou, Pei Geng, Tianchen Li, Yan Li 0046, Lei Lyu 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Multi-Scale Adaptive Large Kernel Graph Convolutional Network for Skeleton-Based Action Recognition
Yu-Qing Zhang, Chen Pang 0001, Pei Geng, Xuequan Lu, Lei Lyu 0001 |
J. Comput. Sci. Technol. | 3 |
| 2025 | Skeleton-based action recognition through attention guided heterogeneous graph neural network
Tianchen Li, Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Multimodal functional deep learning for multiomics dataabstractWith rapidly evolving high-throughput technologies and consistently decreasing costs, collecting multimodal omics data in large-scale studies has become feasible. Although studying multiomics provides a new comprehensive approach in understanding the complex biological mechanisms of human diseases, the high dimensionality of omics data and the complexity of the interactions among various omics levels in contributing to disease phenotypes present tremendous analytical challenges. There is a great need of novel analytical methods to address these challenges and to facilitate multiomics analyses. In this paper, we propose a multimodal functional deep learning (MFDL) method for the analysis of high-dimensional multiomics data. The MFDL method models the complex relationships between multiomics variants and disease phenotypes through the hierarchical structure of deep neural networks and handles high-dimensional omics data using the functional data analysis technique. Furthermore, MFDL leverages the structure of the multimodal model to capture interactions between different types of omics data. Through simulation studies and real-data applications, we demonstrate the advantages of MFDL in terms of prediction accuracy and its robustness to the high dimensionality and noise within the data. Pei Geng, Feifei Xiao, Guoshuai Cai, Li Chen 0029, Qing Lu 0004 |
Briefings Bioinform. | 2 |
| 2024 | Variation-aware directed graph convolutional networks for skeleton-based action recognition
Tianchen Li, Pei Geng, Guohui Cai, Xinran Hou, Xuequan Lu, Lei Lyu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Functional Neural Networks for High-Dimensional Genetic Data AnalysisabstractArtificial intelligence (AI) is a thriving research field with many successful applications in areas such as computer vision and speech recognition. Machine learning methods, such as artificial neural networks (ANN), play a central role in modern AI technology. While ANN also holds great promise for human genetic research, the high-dimensional genetic data and complex genetic structure bring tremendous challenges. The vast majority of genetic variants on the genome have small or no effects on diseases, and fitting ANN on a large number of variants without considering the underlying genetic structure (e.g., linkage disequilibrium) could bring a serious overfitting issue. Furthermore, while a single disease phenotype is often studied in a classic genetic study, in emerging research fields (e.g., imaging genetics), researchers need to deal with different types of disease phenotypes. To address these challenges, we propose a functional neural networks (FNN) method. FNN uses a series of basis functions to model high-dimensional genetic data and a variety of phenotype data and further builds a multi-layer functional neural network to capture the complex relationships between genetic variants and disease phenotypes. Through simulations, we demonstrate the advantages of FNN for high-dimensional genetic data analysis in terms of robustness and accuracy. The real data applications also showed that FNN attained higher accuracy than the existing methods. Pei Geng, Qing Lu 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Hierarchical Aggregated Graph Neural Network for Skeleton-Based Action RecognitionabstractSupervised human action recognition methods based on skeleton data have achieved impressive performance recently. However, many current works emphasize the design of different contrastive strategies to gain stronger supervised signals, ignoring the crucial role of the model's encoder in encoding fine-grained action representations. Our key insight is that a superior skeleton encoder can effectively exploit the fine-grained dependencies between different skeleton information (e.g., joint, bone, angle) in mining more discriminative fine-grained features. In this paper, we devise an innovative hierarchical aggregated graph neural network (HA-GNN) that involves several core components. In particular, the proposed hierarchical graph convolution (HGC) module learns the complementary semantic information among joint, bone, and angle in a hierarchical manner. The designed pyramid attention fusion mechanism (PAFM) fuses the skeleton features successively to compensate for the action representations obtained by the HGC. We use the multi-scale temporal convolution (MSTC) module to enrich the expression capability of temporal features. In addition, to learn more comprehensive semantic representations of the skeleton, we construct a multi-task learning framework with simple contrastive learning and design the learnable data-enhanced strategy to acquire different data representations. Extensive experiments on NTU RGB+D 60/120, NW-UCLA, Kinetics-400, UAV-Human, and PKUMMD datasets prove that the proposed HA-GNN without contrastive learning achieves state-of-the-art performance in skeleton-based action recognition, and it achieves even better results with contrastive learning. Pei Geng, Xuequan Lu, Wanqing Li 0001, Lei Lyu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Focusing Fine-Grained Action by Self-Attention-Enhanced Graph Neural Networks With Contrastive LearningabstractWith the aid of graph convolution neural network and transformer model, human action recognition has achieved significant performance based on skeleton data. However, the majority of existing works rarely focus on identifying fine-grained motion information (i.e., “read”, “write”, etc.). Furthermore, they tend to explore correlations between joints and bones ignoring the angular information. Consequently, the recognition accuracy for fine-grained actions with most models is still less desired. To address this issue, we first attempt to bring angular information as a complement to familiar joint and bone information, while learning the potential dependencies of the three kinds of information using graph neural networks. Based on this, we propose a self-attention-enhanced graph neural network (SAE-GNN), which consists of a kernel-unified graph convolution (KUGC) module and an enhanced attention graph convolution (EAGC) module. The KUGC module is devised to effectively extract rich features in the skeleton information. The EAGC consisting of a multi-scale enhanced graph convolution block and a multi-headed self-attention block is designed to learn the potential high-level semantic information in the features. Besides, we introduce contrastive learning in the two blocks to enhance feature representation by maximizing their mutual information. We conduct extensive experiments on four publicly available datasets, and results show that our model outperforms state-of-the-art methods in recognizing fine-grained actions. Pei Geng, Xuequan Lu, Chunyu Hu 0001, Hong Liu 0013, Lei Lyu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Adaptive multi-level graph convolution with contrastive learning for skeleton-based action recognition
Pei Geng, Fuyun Wang, Lei Lyu 0001 |
Signal Process. | 1 |