VLDB 2026 Research / reviewers in the wild / expert
Zesheng Cheng
dblp:266/8982
· DBLP profile ↗
26ranked-venue papers
1as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved Dense and Sparse Reconstruction with Pixel-Level Noise Mining for Weakly Supervised Salient Object Detection
Chaoqun Fu, Zesheng Cheng |
ICIC (14) | 2 |
| 2026 | Debiased Semi-supervised Classification for Medical Images with Extremely Limited Labels
Xiangdong Meng, Zesheng Cheng |
ICIC (19) | 2 |
| 2026 | Sparse Fine-Tuning for Vision Transformers via Fisher Information
Zesheng Cheng |
ICIC (26) | 2 |
| 2026 | Adaptive Sparse Spatio-Temporal Graph Convolution with Cross-Attention for Hand Gesture RecognitionabstractSkeleton-based hand gesture recognition requires joint modeling of fine-grained spatial dependencies among hand joints and their temporal evolution. Existing spatio-temporal graph convolutional networks (ST-GCNs) and sparse variants have achieved promising results, but most of them still rely on fixed sparsity patterns or static spatial-temporal fusion, which limits adaptability to gesture-specific coordination patterns. To address these issues, we propose an Adaptive Sparse Spatio-Temporal Graph Convolution Network with Cross-Attention Fusion (ASCA-STGCN) for dynamic hand gesture recognition. First, we introduce an input-dependent sparse graph learning mechanism that modulates pairwise joint correlations with a differentiable gating function, allowing the spatial graph topology to adapt to different gestures. Second, we adopt a decoupled spatial-temporal design and further introduce a bidirectional cross-attention fusion module, so that spatial and temporal streams can reweight each other explicitly rather than being combined by static addition. Experiments on three public benchmarks, namely SHREC'17, Briareo, and IPNHand, show that ASCA-STGCN consistently improves over ST-GCN and ST-SGCN baselines while maintaining low model complexity. The results suggest that adaptive sparse graph learning and cross-attention-based fusion provide a lightweight and effective enhancement for skeleton-based hand gesture recognition. Kunjiang Teng, Zesheng Cheng |
ICIC | 2 |
| 2026 | UMG-Net: Universal Multi-view Geometric Network with Hierarchical Dual-Stream Encoding for 3D Medical Image Representation
Hengyi Yuan, Zesheng Cheng |
ICIC (20) | 4 |
| 2026 | Dynamic Confidence-Weighted Specialist-Generalist Collaborative Learning for Semi-supervised 3D Medical Image Segmentation
Zesheng Cheng |
ICIC (20) | 2 |
| 2026 | Traffic Flow Forecasting Based on Spatio-temporal Blurring-Sharpening Differential Equation Networks
Hongyu Yuan, Zesheng Cheng |
ICIC (3) | 2 |
| 2026 | BF-FreqMaIR: Blur-Aware Frequency-Biased Continuity-Preserving Mamba for Image Deblurring
Hongyu Yuan, Zesheng Cheng |
ICIC (21) | 2 |
| 2026 | SpecTr: Spectral-Adversarial Masked Image Modeling for Efficient 3D Medical Pre-Training
Hengyi Yuan, Zesheng Cheng |
ICIC (8) | 4 |
| 2026 | HBPC-GWN: Hierarchical Bayesian Calibration for Graph-Based Traffic Forecasting
Jinhui Zhao, Zesheng Cheng |
ICIC (9) | 2 |
| 2026 | EDG-ODE: Edge-Driven High-Order Graph Neural ODE for Continuous Traffic Flow Forecasting
Jinhui Zhao, Zesheng Cheng |
ICIC | 2 |
| 2025 | A Gaussian Temporal Based Graph Convolutional Network for Traffic Operation Flow ForecastingabstractHigh-Quality traffic flow data may provide planners with a basis for designing road capacity, pavement and intersection control, etc., thus assisting in the construction of a more rational traffic network. This study developed a new Gaussian Temporal Network Module, which is a module based on a Gaussian process that utilizes a kernel Gaussian convolution kernel to assist in the extraction of temporal features. Which is aim to solve the common gradient vanishing and gradient explosion problems in temporal neural networks Further, this study combines the GTNM with the GCN module as a GT-GCN model and tests its performance on a publicly available dataset. Experiment results show that GT-GCN demonstrates superior performance, attaining state-of-the-art or second best outcomes across different test dataset. Several of these results surpass baseline benchmarks by more than 10%, which then well illustrates the effectiveness and stability of GT-GCN model. Shulan Guo, Shangda Xiao, Mingxuan He, Hequn Xian, Zesheng Cheng |
ICCCN | 6 |
| 2025 | STMOM: A Spatiotemporal Traffic Flow Forecasting Model Based on Multi-Order Markov ChainsabstractWith the development of urban traffic, road networks are becoming increasingly complex and traffic congestion is getting worse. Accurately predicting traffic data is one of the most important means to keep urban roads smooth and is one of the key technologies to improve urban traffic management capabilities. Graph Convolutional Networks (GCNs) are a type of deep learning method used to analyze spatial information in traffic networks. There are two problems with traditional GCNs. First, GCNs can only solve first-order models that only consider the influence between adjacent nodes’ temporal data, and stacking multiple layers of graph neural networks can easily lead to over-smoothing phenomena. Second, as the number of layers increases, the number of trainable parameters in the model increases, so does leading to model overfitting. Therefore, this paper proposes a traffic flow prediction model based on multi-order Markov chains. The paper estimates the parameters of this model and refers to the urban road traffic operation evaluation index system. Through prediction experiments, it is proven that the prediction accuracy of this model is higher than other baseline models. Shulan Guo, Zesheng Cheng |
IJCNN | 4 |
| 2025 | HandDiff-GAN: Handwriting Diffusion-Enhanced Generative Adversarial Networks for Character Generation
Zesheng Cheng |
NLPCC (4) | 2 |
| 2025 | S-Track: An SAM-based Model for Real-Time Surgical Video Tracking under High Dynamic ScenariosabstractReal-time surgical video tracking faces critical challenges in highly dynamic endoscopic environments, including rapid instrument-tissue interactions, computational inefficiency, and annotation scarcity. This paper proposes S-Track, a novel SAM-based framework integrating a memory-augmented spatiotemporal fusion module, a dual-path motion prediction mechanism, and a dynamic attention-guided computation strategy. These innovations synergistically address the accuracy-efficiency trade-off, achieving state-of-the-art performance while reducing GPU memory usage by 27% and accelerating inference speed to 27 FPS. Furthermore, we introduce URS-Lithotripsy, the first comprehensive dataset for multi-target surgical tracking, comprising 20 hours of high-resolution videos with 1.2 million expert-annotated frames. Cross-domain validation on EndoVis 2018 demonstrates robust generalization, and ablation studies confirm the necessity of each module. The efficiency, precision, and adaptability to sparse annotations highlight the potential of S-Track for real-time clinical deployment. Sicheng Lu, Zesheng Cheng |
SMC | 2 |
| 2025 | WEMT: Wavelet-Enhanced Multiscale Transformer for Robust Dynamic Gesture RecognitionabstractDynamic gesture recognition faces critical challenges due to complex backgrounds, multi-scale feature variations, and inefficient multi-modal fusion. To address these challenges, this study proposes the Wavelet-Enhanced Multiscale Transformer (WEMT), a unified framework that integrates three key innovations: wavelet-based feature decomposition for noise-robust detail enhancement, hierarchical attention mechanisms for global-local spatiotemporal modeling, and confidence-driven adaptive fusion for heterogeneous modality integration. The wavelet domain module selectively amplifies discriminative gesture features through frequency-domain analysis and multi-scale pooling, while the pyramid-structured attention heads capture both coarse-grained postures and fine-grained motions. The adaptive fusion framework dynamically prioritizes complementary modalities through entropy-based gating, ensuring robustness in diverse environments. Evaluations on NVGesture and Briareo datasets demonstrate state-of-the-art performance across single-modal and multi-modal configurations, with notable advantages in low-light scenarios. The model also maintains computational efficiency, making it suitable for real-world human-computer interaction applications. This work provides a systematic and scalable solution for accurate dynamic gesture recognition, bridging the gap between theoretical advances and practical deployment. Kunjiang Teng, Xinyu An, Zesheng Cheng |
SMC | 3 |
| 2025 | HBRA: Heterogeneous Graph Learning for Bidirectional Marital Recommendations in Aging SocietiesabstractAddressing the inefficiency of marital matching in aging societies, this study proposes HBRA, a bidirectional recommendation framework leveraging heterogeneous graph learning to model asymmetric preferences and enhance compatibility analysis. The framework integrates behavioral and textual data through a unified graph architecture, overcoming limitations of unidirectional recommendation systems and fragmented feature fusion. By introducing spectral graph theory and meta-path-guided message passing, HBRA achieves noise-robust preference modeling while eliminating traditional neural network training through closed-form matrix decomposition, reducing computational complexity by 99.5%. Evaluated on the novel “FCWR” dataset (built from real-world matchmaking records) and the Speed Dating benchmark, HBRA outperforms 24 state-of-the-art models, improving NDCG@2 by 25.8% and training efficiency by three orders of magnitude. Parameter analysis reveals text-based lifestyle keywords as the dominant matching factor, surpassing traditional attributes like age and geography. This work provides a scalable, interpretable solution for intelligent matchmaking platforms, directly addressing demographic challenges posed by population aging through data-driven compatibility optimization. Fenghao Zhang, Zesheng Cheng, Zi Jin |
SMC | 2 |
| 2025 | LGM4TFP: A Large Graph Model for Traffic Flow PredictionabstractUrban traffic flow prediction is a critical task in Intelligent Transportation Systems, yet existing methods often struggle with capturing long-range spatial dependencies and preserving high-frequency signals. To address these challenges, this paper proposes LGM4TFP, a novel framework that incorporates a Graph Attention with High-frequency Enhancement (GAHE) module. GAHE integrates dynamic graph attention for spatial feature extraction and a spectral-domain high-frequency enhancement mechanism to alleviate the over-smoothing problem. Extensive experiments on four real-world datasets demonstrate that LGM4TFP consistently outperforms state-of-the-art models by 5–10% in RMSE, MAE, and MAPE. Ablation studies confirm the effectiveness of the proposed modules, and results show that the model maintains strong robustness across diverse prediction scenarios, highlighting its practical value for dynamic traffic management. Jinhui Zhao, Zesheng Cheng |
SMC | 2 |
| 2025 | STA-Hyper: Hypergraph-Based Spatio-Temporal Attention Network for Next Point-of-Interest Recommendation
Huarui Yu, Zesheng Cheng |
WASA (1) | 2 |
| 2024 | UG-STNN: A Spatial-Temporal Neural Network Based on Unsupervised Graph Representation Module for Traffic Flow PredictionabstractAccurate and efficient traffic flow prediction helps to build an intelligent transportation system and improve the travel experience in daily life. In this study, a new Spatial-Temporal Neural Network Based on Unsupervised Graph Representation Module (UG-STNN) is proposed to improve the graph convolution module, which uses unsupervised learning to extract features in spatial dimensions, and it can learn the structural and feature information in the graph better. Our UG-STNN uses fewer convolutional layers to reduce the number of parameters, decrease the complexity of the model, and improve performance and accuracy. From the experimental results of UG-STNN on different test datasets, the model can approach or even achieve better prediction results compared with other models, which well illustrates the accuracy and stability of the UG-STNN model. Enwei Zhang, Zesheng Cheng, Tiankuan Wang |
SMC | 2 |
| 2024 | REHG: A Recommender Engine Based on Heterogeneous Graph
Xiaoyang Xin, Zesheng Cheng, Tiankuan Wang, Mengqiu Yan, Ruixuan Zhao |
WASA (1) | 2 |
| 2024 | ST-TDCN: A two-channel tree-structure spatial-temporal convolutional network model for traffic velocity prediction
Zhiqiang Lv, Zesheng Cheng, Sisi Jian |
Expert Syst. Appl. | 3 |
| 2024 | TreeCN: Time Series Prediction With the Tree Convolutional Network for Traffic PredictionabstractThe complexity of traffic scenarios, the spatial-temporal feature correlations pose higher challenges for traffic prediction research. Traffic spatial-temporal model is an essential method in this research field, primarily focusing on capturing the spatial-temporal features among nodes and their neighboring nodes. However, existing methods lack comprehensive consideration of directional and hierarchical features among traffic nodes. They are mostly applicable to scenarios with random uniform distribution of nodes, but not suitable for more complex small-scale aggregation distribution scenarios. Therefore, this study proposes the Tree Convolutional Network (TreeCN), a tree-based structure. The data design and model design of TreeCN focus on capturing the directional and hierarchical features among nodes. The directional and hierarchical relationships among nodes are represented by the plane tree matrix and constructed as the spatial tree matrix. The TreeCN, with a full convolution network, performs a bottom-up convolution structure on the tree matrix to complete the task of node feature capturing. In this study, TreeCN is thoroughly compared with statistical, machine learning, and deep learning methods in traffic time series prediction. The experimental results show that TreeCN not only performs well in scenarios with random uniform distribution but also exhibits outstanding effect in more complex small-scale aggregation distribution. Moreover, TreeCN adheres to the design principles of Graph Convolutional Networks (GCN) in capturing the spatial features of traffic nodes and can further capture directional and hierarchical features among them. This is expected to make TreeCN a new method to handle complex traffic scenarios and improve prediction accuracy. Zhiqiang Lv, Zesheng Cheng, Zhihao Xu 0002, Zheng Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | A new approach to COVID-19 data mining: A deep spatial-temporal prediction model based on tree structure for traffic revitalization index
Zhiqiang Lv, Zesheng Cheng, Haoran Li 0021, Zhihao Xu 0002 |
Data Knowl. Eng. | 3 |
| 2022 | A Spatial-Temporal Convolutional Model with Improved Graph Representation
Zesheng Cheng, Zhiqiang Lv |
WASA (1) | 2 |
| 2020 | Integrating Household Travel Survey and Social Media Data to Improve the Quality of OD Matrix: A Comparative Case StudyabstractCollecting effective data is a fundamental step in developing transport networks and related research. Social media have become an emerging source of data for traffic analyses. In this paper, we demonstrate that the function of a city influences the utility of social media data in travel demand models by generating models for eight US cities with different functions. Data from Twitter and Foursquare, as well as other socio-demographic information, are considered as independent variables in Origin-Destination trip regression models generated via a Random Forest regression technique. Model performance with and without use of social media data are compared via 10-fold cross-validation. The results indicate that the accuracy of the models for all eight cities improved when independent variables based on social media data were included. The performance was most improved in metropolitan areas, followed by rural and tourist areas. Inspired by this finding, we conclude that the city function influences the utility of social media data in travel demand models. Meanwhile, we create models based on trip purpose and transport mode to explore other factors that may impact the efficiency of applying social media data in transport research. Zesheng Cheng, Sisi Jian, Taha Hossein Rashidi, Mojtaba Maghrebi, S. Travis Waller |
IEEE Trans. Intell. Transp. Syst. | 1 |