VLDB 2026 Research / reviewers in the wild / expert
Xing Li 0005
dblp:26/379-5
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-0881-0978ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EchoNet: A hierarchical collaborative network for point cloud-based 3D action recognition
Guojia Huang, Zhenjie Hou, Xing Li 0005, Jiuzhen Liang, Xinwen Zhou |
Knowl. Based Syst. | 3 |
| 2026 | Spatial-Temporal Self-Compensating Graph Convolutional Network for Skeleton-Based Action Recognition Under Data ConstraintsabstractSkeleton-based human action recognition has emerged as a prominent research focus in computer vision, with significant progress achieved in recent years. However, existing methods often suffer substantial performance degradation under real-world data constraints, such as body occlusion, missing frames, and noise. These limitations critically undermine the robustness of related techniques in practical applications. To address these challenges, we propose a Spatial Temporal Self-compensating Graph Convolutional Network (STSc-GCN), which skillfully utilizes the systematic and regular nature of human movement to mitigate performance degradation caused by data constraints through a data self-compensation mechanism. Specifically, STSc-GCN comprises two key modules: 1) collaborative motion spatial compensation (CMSC). This module designs multiple distinct topological relationships, primarily including Walk-probability Generality Topology and Self-organizing Particularity Topology, respectively, to deeply explore the universal and personalized collaborative relationships between human joints. These relationships help compensate for the lack of information caused by spatial data constraints and 2) meta-action sharpening temporal Compensation (MSTC). This module introduces a novel motion sharpening mechanism that enhances key dynamic information within the meta-action sequences through cross-attention technology, thereby improving model adaptability to missing-frame scenarios. STSc-GCN achieves state-of-the-art performance on four constrained datasets and shows superior results on three widely used standard datasets, confirming its effectiveness in both constrained and general scenarios. Code will be available at https://github.com/XingLi1012/STSc-GCN.git. Xing Li 0005, Qian Huang 0008, Xin Li 0090, Jinhui Tang 0001, Qiaolin Ye |
IEEE Trans. Image Process. | 1 |
| 2026 | MD-PCSN: Meta-Motion Decoupling Point Cloud Sequence Network for Privacy-Preserving Human Action Recognition in AI MachinesabstractIn next-generation communication networks and Industry 5.0 based applications, ensuring robust security and reliability in human-computer interaction (HCI) constitutes a fundamental prerequisite for safety-critical AI machine systems. Point cloud sequence-based human action recognition demonstrates intrinsic advantages in privacy-preserving HCI, leveraging its non-intrusive sensing modality to mitigate data vulnerability while maintaining high-precision action interpretation in industrial environments. Existing spatio-temporal encoding methods for point cloud sequence-based action recognition suffer from two fundamental limitations: (1) rigid neighborhood constraints impair multi-scale feature extraction for heterogeneous body parts, and (2) independent spatial-temporal decomposition introduces motion representation distortion. We propose a Meta-motion Decoupling Point Cloud Sequence Network (MD-PCSN) that addresses these challenges through: (1) logarithmic spatio-temporal point convolution for hierarchical meta-motion construction at variable granularities, and (2) a novel Gated-KANsformer architecture with differential motion encoding to explicitly model both short-term displacements and long-term spatio-temporal dependencies. The proposed meta-motion decoupling mechanism significantly enhances robustness against sensor perturbations, making the framework particularly suitable for security-critical applications. Extensive experiments on three benchmark datasets demonstrate MD-PCSN’s superior performance. It outperforms classic PST-Transformer by 1.5% on MSR Action3D and 4.14% on UTD-MHAD. Under the NTU RGB+D 60, it achieves 2.9% cross-view gain over the latest PointActionCLIP. Xing Li 0005, Xin Li 0090, Qian Huang 0008 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2025 | Geo-CF2Net: Geometry-Prior Cross-Frequency Interactive Fusion Network for 3D Human Action RecognitionabstractDynamic point cloud-based human action recognition has garnered increasing attention due to its inherent advantages in privacy preservation and structural completeness. Current methods typically rely on nested point spatio-temporal convolutions to understand motion semantics in a bottom-up manner, which is intractable for capturing high-fidelity human dynamics disentangled from spatio-temporal interference. Motivated by this, designing a practical spatio-temporal factorization backbone is essential. However, the repeated coarsening of aggregated features along the spatial dimension often leads to the degradation of intrinsic geometric texture relations within point cloud data. Moreover, discretizing continuous visual data into isolated temporal hyperpoints significantly diminishes temporal continuity, resulting in the fragmentation of human action. To circumvent above limitations, we propose a novel Geometry-Prior Cross-Frequency Interactive Fusion Network (Geo-CF2Net). Specifically, we investigate a Spatial-Geometry Pose Prior (SGPP) module, which compensates for pose information loss during spatial downsampling by explicitly modeling geometric constraints among neighboring points. In addition, we elaborate on a Temporal Motion Unit Interactive Coordination (TMIC) module to track the interactive composite semantics of low-frequency steady-state venations and high-frequency transient-state details within a high-dimensional pose evolution flow. Extensive experiments on three public benchmarks substantiate the superiority of Geo-CF2Net over state-of-the-art methods. Qian Huang 0008, Xing Li 0005, Shihao Han, Yirui Wu, Xin Li 0090, Ziyang Yin |
ACM Multimedia | 3 |
| 2025 | STFE-VC: Spatio-temporal feature enhancement for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Expert Syst. Appl. | 5 |
| 2025 | PRG-Net: Point Relationship-Guided Network for 3D human action recognitionabstractPoint clouds contain rich spatial information, providing important supplementary clues for human action recognition . Recent methods for action recognition based on point cloud sequences primarily rely on complex spatiotemporal local encoding. However, these methods often utilize max-pooling operations to select features when extracting local features , restricting feature updates to local neighborhoods and failing to fully exploit the relationships between regions. Moreover, cross-frame encoding can also lead to the loss of spatiotemporal information. In this study, we propose PRG-Net, a Point Relation Guided Network, to further improve the learning of spatiotemporal features in point clouds. First, we designed two core modules: the Spatial Feature Aggregation (SFA) and the Spatial Feature Descriptor (SFD) modules. The SFA module expands the spatial structure between regions using dynamic aggregation techniques, while the SFD module guides the region aggregation process by Attention-Weighted Descriptors. They enhance the modeling of human spatial structure by expanding the relationships between regions. Second, we introduce inter-frame motion encoding techniques that can obtain the final spatiotemporal representation of the human body through the aggregation of cross-frame vectors, without relying on complex spatiotemporal local encoding. We evaluate PRG-Net on publicly available human action recognition datasets, including NTU RGB+D 60, NTU RGB+D 120, UTD-MHAD, and MSR Action 3D. Experimental results demonstrate that our method outperforms state-of-the-art point-based 3D action recognition methods significantly. Furthermore, we conduct extended experiments on the SHREC 2017 dataset for gesture recognition , and the results show that our method maintains competitive performance on that dataset as well. Zhenjie Hou, En Lin, Xing Li 0005, Jiuzhen Liang, Xinwen Zhou |
Neurocomputing | 4 |
| 2025 | GaitSTAGCN: Spatial-temporal attention graph convolutional networks for gait recognition
Aofei Wang, Zhenjie Hou, En Lin, Xing Li 0005, Jiuzhen Liang, Xinwen Zhou |
Neurocomputing | 4 |
| 2025 | Multiscale motion-aware and spatial-temporal-channel contextual coding network for learned video compression
Yiming Wang 0008, Qian Huang 0008, Bin Tang 0002, Xin Li 0090, Xing Li 0005 |
Knowl. Based Syst. | 5 |
| 2024 | Multi-granular spatial-temporal synchronous graph convolutional network for robust action recognition
Qian Huang 0008, Yingchi Mao, Xing Li 0005, Jie Wu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | PointDMIG: a dynamic motion-informed graph neural network for 3D action recognition
Zhenjie Hou, Xing Li 0005, Jiuzhen Liang, Kaijun You, Xinwen Zhou |
Multim. Syst. | 3 |
| 2023 | Spatial and temporal information fusion for human action recognition via Center Boundary Balancing Multimodal Classifier
Xing Li 0005, Qian Huang 0008, Zhijian Wang 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Real-Time 3-D Human Action Recognition Based on Hyperpoint SequenceabstractReal-time 3-D human action recognition has broad industrial applications, such as surveillance, human–computer interaction, and healthcare monitoring. By relying on complex spatio-temporal local encoding, most existing point cloud sequence networks capture spatio-temporal local structures to recognize 3-D human actions. To simplify the point cloud sequence modeling task, we propose a lightweight and effective point cloud sequence network referred to as SequentialPointNet for real-time 3-D action recognition. Instead of capturing spatio-temporal local structures, SequentialPointNet encodes the temporal evolution of static appearances to recognize human actions. First, we define a novel type of point data, hyperpoint, to better describe the temporally changing human appearances. A theoretical foundation is provided to clarify the information equivalence property for converting point cloud sequences into hyperpoint sequences. Second, the point cloud sequence modeling task is decomposed into a hyperpoint embedding task and a hyperpoint sequence modeling task. Specifically, for hyperpoint embedding, the static point cloud technology is employed to convert point cloud sequences into hyperpoint sequences, which introduces inherent frame-level parallelism; for hyperpoint sequence modeling, a hyperpoint-mixer module is designed as the basic building block to learning the spatio-temporal features of human actions. Extensive experiments on three widely-used 3-D action recognition datasets demonstrate that the proposed SequentialPointNet achieves a competitive classification performance with up to 10× faster than existing approaches. Xing Li 0005, Qian Huang 0008, Zhijian Wang 0002, Tianjin Yang, Zhenjie Hou, Zhuang Miao |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Hyperpointnet for Point Cloud Sequence-Based 3D Human Action RecognitionabstractPoint cloud sequence-based 3D action recognition achieves impressive performance and efficiency. Conventional approaches for modeling point cloud sequences usually perform cross-frame spatio-temporal local encoding, thus resulting in intensive computation and mutual interference between spatial and temporal information extracting. In this work, to avoid spatio-temporal local encoding, we propose a strong parallelized point cloud sequence network referred to as HyperpointNet for 3D action recognition. HyperpointNet is composed of two serial modules, i.e., a hyperpoint sequence embedding module and a hyperpoint sequence encoding module. In the hyperpoint sequence embedding module, employing static point cloud modeling methods, the point cloud sequence is abstracted into a new point data type named hyper-point sequence. In the hyperpoint sequence encoding module, a temporal PointNet (TPN) layer is designed to model the hyperpoint sequence. Extensive experiments conducted on two public datasets show that HyperpointNet outperforms state-of-the-art approaches. Xing Li 0005, Qian Huang 0008, Tianjin Yang, Qianhan Wu |
ICME | 1 |
| 2022 | Dynamic-scale grid structure with weighted-scoring strategy for fast feature matching
Qian Huang 0008, Xing Li 0005 |
Appl. Intell. | 3 |
| 2022 | VirtualActionNet: A strong two-stream point cloud sequence network for human action recognition
Xing Li 0005, Qian Huang 0008, Zhijian Wang 0002, Tianjin Yang |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Human action recognition based on enhanced data guidance and key node spatial temporal graph convolution
Chengyu Zhang 0004, Jiuzhen Liang, Xing Li 0005, Yunfei Xia, Lan Di, Zhenjie Hou, Zhan Huan |
Multim. Tools Appl. | 3 |
| 2020 | A multi-scale human action recognition method based on Laplacian pyramid depth motion imagesabstractHuman action recognition is an active research area in computer vision. Aiming at the lack of spatial muti-scale information for human action recognition, we present a novel framework to recognize human actions from depth video sequences using multi-scale Laplacian pyramid depth motion images (LP-DMI). Each depth frame is projected onto three orthogonal Cartesian planes. Under three views, we generate depth motion images (DMI) and construct Laplacian pyramids as structured multi-scale feature maps which enhances multi-scale dynamic information of motions and reduces redundant static information in human bodies. We further extract the multi-granularity descriptor called LP-DMI-HOG to provide more discriminative features. Finally, we utilize extreme learning machine (ELM) for action classification. Through extensive experiments on the public MSRAction3D datasets, we prove that our method outperforms state-of-the-art benchmarks. Qian Huang 0008, Xing Li 0005, Qianhan Wu |
MMAsia | 3 |
| 2020 | Human action recognition based on 3D body mask and depth spatial-temporal maps
Xing Li 0005, Zhenjie Hou, Jiuzhen Liang, Chen Chen 0001 |
Multim. Tools Appl. | 1 |