VLDB 2026 Research / reviewers in the wild / expert
Zhanpeng Shao
dblp:135/8444
· DBLP profile ↗
26ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-8130-5230ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Altering Query Prompting With Contrastive Learning for Multimodal Intent RecognitionabstractMultimodal intent recognition utilizes heterogeneous modalities such as visual, auditory, and textual cues to infer user intent, serving as a pivotal component in human-machine interaction. Existing approaches, however, often rely on unimodal paradigms or shallow multimodal fusion, failing to model cross-modal semantic dependencies and struggling to extract discriminative features from non-verbal modalities, limiting their robustness in complex scenarios. To mitigate these limitations, we propose an Altering Query Prompting with Contrastive Learning framework (AQP-CL) that dynamically aligns and refines multimodal representations. Specifically, the Altering Query Prompting (AQP) module introduces a tri-modality rotation attention mechanism, where textual, visual, and acoustic modalities cyclically alternate as queries in cross-attention operations. This approach addresses modality bias while strengthening interdependencies between modalities, ultimately yielding intent-aware fused feature representations that preserve discriminative cues. The Label-semantic Augmented Contrastive Learning (LACL) strategy generates augmented samples through the intent-aware query prompt and enhances feature discrimination via NT-Xent loss on label tokens. By integrating high-confidence textual semantics from intent labels, LACL refines auxiliary modality features through contrastive alignment, ensuring robust cross-modal representation learning. Evaluations on IEMOCAP and MIntRec validate AQP-CL's superiority, achieving state-of-the-art precision of 77.78% on IEMOCAP, a 3.41% improvement over existing methods. Yuxin Jia, Zhanpeng Shao, Min Liu 0008 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Noise-robust re-identification with triple-consistency perception
Zhanpeng Shao, Shixi Luo, Jiazheng Wang 0001, Min Liu 0008, Jianhua Dai 0003 |
Image Vis. Comput. | 2 |
| 2024 | A temporal densely connected recurrent network for event-based human pose estimation
Zhanpeng Shao, Wuzhen Wang, Jianyu Yang 0002, Youfu Li 0001 |
Pattern Recognit. | 1 |
| 2024 | Event Voxel Set Transformer for Spatiotemporal Representation Learning on Event StreamsabstractEvent cameras are neuromorphic vision sensors that record a scene as sparse and asynchronous event streams. Most event-based methods project events into dense frames and process them using conventional vision models, resulting in high computational complexity. A recent trend is to develop point-based networks that achieve efficient event processing by learning sparse representations. However, existing works may lack robust local information aggregators and effective feature interaction operations, thus limiting their modeling capabilities. To this end, we propose an attention-aware model named Event Voxel Set Transformer (EVSTr) for efficient spatiotemporal representation learning on event streams. It first converts the event stream into voxel sets and then hierarchically aggregates voxel features to obtain robust representations. The core of EVSTr is an event voxel transformer encoder that consists of two well-designed components, including the Multi-Scale Neighbor Embedding Layer (MNEL) for local information aggregation and the Voxel Self-Attention Layer (VSAL) for global feature interaction. Enabling the network to incorporate a long-range temporal structure, we introduce a segment modeling strategy (S2TM) to learn motion patterns from a sequence of segmented voxel sets. The proposed model is evaluated on two recognition tasks, including object classification and action recognition. To provide a convincing model evaluation, we present a new event-based action recognition dataset (NeuroHAR) recorded in challenging scenarios. Comprehensive experiments show that EVSTr achieves state-of-the-art performance while maintaining low model complexity. Bochen Xie, Yongjian Deng, Zhanpeng Shao, Qingsong Xu 0002, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | EISNet: A Multi-Modal Fusion Network for Semantic Segmentation With Events and ImagesabstractBio-inspired event cameras record a scene as sparse and asynchronous “events” by detecting per-pixel brightness changes. Such cameras show great potential in challenging scene understanding tasks, benefiting from the imaging advantages of high dynamic range and high temporal resolution. Considering the complementarity between event and standard cameras, we propose a multi-modal fusion network (EISNet) to improve the semantic segmentation performance. The key challenges of this topic lie in (i) how to encode event data to represent accurate scene information and (ii) how to fuse multi-modal complementary features by considering the characteristics of two modalities. To solve the first challenge, we propose an Activity-Aware Event Integration Module (AEIM) to convert event data into frame-based representations with high-confidence details via scene activity modeling. To tackle the second challenge, we introduce the Modality Recalibration and Fusion Module (MRFM) to recalibrate modal-specific representations and then aggregate multi-modal features at multiple stages. MRFM learns to generate modal-oriented masks to guide the merging of complementary features, achieving adaptive fusion. Based on these two core designs, our proposed EISNet adopts an encoder-decoder transformer architecture for accurate semantic segmentation using events and images. Experimental results show that our model outperforms state-of-the-art methods by a large margin on event-based semantic segmentation datasets. The code is publicly available athttps://github.com/bochenxie/EISNet. Bochen Xie, Yongjian Deng, Zhanpeng Shao, Youfu Li 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Towards General and Fast Video Derain via Knowledge DistillationabstractAs a common natural weather condition, rain can obscure video frames and thus affect the performance of the visual system, so video derain receives a lot of attention. In natural environments, rain has a wide variety of streak types, which increases the difficulty of the rain removal task. In this paper, we propose a Rain Review-based General video derain Network via knowledge distillation (named RRGNet) that handles different rain streak types with one pre-training weight. Specifically, we design a frame grouping-based encoder-decoder network that makes full use of the temporal information of the video. Further, we use the old task model to guide the current model in learning new rain streak types while avoiding forgetting. To consolidate the network’s ability to derain, we design a rain review module to play back data from old tasks for the current model. The experimental results show that our developed general method achieves the best results in terms of running speed and derain effect. Defang Cai, Pan Mu, Sixian Chan 0001, Zhanpeng Shao, Cong Bai |
ICME | 4 |
| 2023 | SGPT: The Secondary Path Guides the Primary Path in Transformers for HOI DetectionabstractHOI detection is essential for human-computer interaction, especially in behavior detection and robot manipulation. Existing mainstream transformer methods of HOI detection are focused on single-stream detection only, e.g.,$image \rightarrow HOI(\mathcal{P}_{1})$, or$image \rightarrow HO\rightarrow I(\mathcal{P}_{2})$. Both paths have their own characteristics of concern, so we propose a novel method, using the Secondary path$(\mathcal{P}_{2})$Guides the Primary path$(\mathcal{P}_{1})$in Transformers (SGPT). SGPT contains two core modules: the Dual-Path Consistency (DPC) module and the Instance Interaction Attention (IIA) module. DPC keeps human, object and interaction consistent on the dual-path and lets$\mathcal{P}_{2}$guide$\mathcal{P}_{1}$to learn more meaningful features. IIA fuses human and object to enhance interaction in$\mathcal{P}_{2}$, which allows instance to constrain interaction. Our proposed dual-path are employed during training, and only the$\mathcal{P}_{1}$path is used for inference. Hence, SGPT improves generalization without increasing model capacity in HICO-DET and V-COCO datasets compared to the state-of-the-arts. The code of this work is available at https://github.com/visualVk/sgpt.git. Sixian Chan 0001, Weixiang Wang, Zhanpeng Shao, Cong Bai |
ICRA | 3 |
| 2022 | Automatic and accurate segmentation of peripherally inserted central catheter (PICC) from chest X-rays using multi-stage attention-guided learning
Xiaoyan Wang 0007, Ye Sheng, Chenglu Zhu, Cong Bai, Ming Xia 0005, Zhanpeng Shao, Ruiyi Zhao, Zhenjie Liu |
Neurocomputing | 8 |
| 2022 | Multi-stream feature refinement network for human object interaction detection
Zhanpeng Shao, Zhongyan Hu, Jianyu Yang 0002, Youfu Li 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | A Multi-Level Network for Human Pose EstimationabstractAlthough multi-person human pose estimation has made great progress in recent years, the challenges such as various scales of persons, occluded keypoints, and crowded backgrounds in complex scenes are still remained to be solved. In this paper, we propose a novel multi-level pose estimation network (MLPE) to learn multi-level features that can preserve both the strong semantic clues and spatial resolution for keypoint prediction and location. More specifically, a multi-level prediction network with a feature enhancement strategy is first proposed to learn multi-level features to achieve a good trade-off between the global context information and spatial resolution. We then build a high-resolution fine network to restore high spatial resolution information based on transposed convolutions to accurately locate the keypoints. We have conducted extensive experiments on the challenging MS COCO dataset, which has proved the effectiveness of our proposed method. Code†and the experimental results are publicly online available for further research. Zhanpeng Shao, Youfu Li 0001, Jianyu Yang 0002, Xiaolong Zhou 0001 |
ICRA | 1 |
| 2021 | Learning discriminative motion feature for enhancing multi-modal action recognition
Jianyu Yang 0002, Zhanpeng Shao, Chunping Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Learning Representations From Skeletal Self-Similarities for Cross-View Action RecognitionabstractExisting research attention in vision-based action recognition is generally paid on recognizing actions from the same views seen in the training data. One of the big challenges in action recognition lies in the large variations of action representations as actions are captured from totally different viewpoints. This paper addresses this problem by learning view-invariant representations from skeletal self-similarities of varying scales with a very light multi-stream neural network (MSNN). As human skeletons have been proved to be an effective feature modality used for action recognition and are easy to obtain, we first create a view-invariant action description by formulating skeletal self-similarities at each frame as an image (SSI), which can show a high structural stability under view changes. Accordingly, a MSNN is designed based on 3D CNN and LSTM units to learn representations from SSIs of multiple scales, where the scheme of multiple scales provides our method with a good robustness to view changes. In addition, we integrate the computation of SSIs into the MSNN by wrapping it as a custom learnable layer thanks to its simplicity, instead of normalizing and transforming skeletons using a hand-crafted preprocessing. Extensive experimental evaluations on three challenging cross-view datasets demonstrate the effectiveness of our proposed method, which achieves superior performance to the state-of-the-art algorithms on cross-view recognition. The source code of this work will be released shortly to facilitate future studies in this field. Zhanpeng Shao, Youfu Li 0001, Hong Zhang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Improved itracker combined with bidirectional long short-term memory for 3D gaze estimation using appearance cues
Xiaolong Zhou 0001, Jianing Lin, Zhuo Zhang 0012, Zhanpeng Shao, Shenyong Chen, Honghai Liu 0001 |
Neurocomputing | 4 |
| 2019 | A Hierarchical Model for Human Action Recognition From Body-PartsabstractAs increasing attention is paid to human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in actions for better analysis of human actions in the skeleton data. Considering human actions as simultaneous motions of body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical rotation and relative velocity (HRRV) descriptor. The hierarchical representations encoded by Fisher vectors of the HRRV descriptors are properly formulated into the hierarchical model via the proposed mixed norm, to apply the sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to the state-of-the-art algorithms on datasets with various sizes, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Xiaolong Zhou 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | A Hierarchical Model for Action Recognition Based on Body PartsabstractAs increasing attention is paid on human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in human actions. Considering human actions as simultaneous motions of different body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical RRV (Rotation and Relative Velocity) descriptors. The hierarchical representations encoded by Fisher vectors of the hierarchical RRV descriptors are properly formulated into the hierarchical model via the proposed hierarchical mixed norm, to apply sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to state-of-the-art results on different sizes of datasets, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Jianyu Yang 0002, Zhenhua Wang 0003 |
ICRA | 1 |
| 2018 | Deep CRF-Graph Learning for Semantic Image Segmentation
Fuguang Ding, Zhenhua Wang 0003, Dongyan Guo, Shengyong Chen, Jianhua Zhang 0002, Zhanpeng Shao |
PRICAI | 6 |
| 2018 | Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao |
Neurocomputing | 9 |
| 2018 | DSRF: A flexible trajectory descriptor for articulated human action recognition
Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
Pattern Recognit. | 3 |
| 2018 | RRV: A Spatiotemporal Descriptor for Rigid Body Motion RecognitionabstractThe motion behaviors of a rigid body can be characterized by a six degrees of freedom motion trajectory, which contains the 3-D position vectors of a reference point on the rigid body and 3-D rotations of this rigid body over time. This paper devises a rotation and relative velocity (RRV) descriptor by exploring the local translational and rotational invariants of rigid body motion trajectories, which is insensitive to noise, invariant to rigid transformation and scale. The RRV descriptor is then applied to characterize motions of a human body skeleton modeled as articulated interconnections of multiple rigid bodies. To show the descriptive ability of our RRV descriptor, we explore its potentials and applications in different rigid body motion recognition tasks. The experimental results on benchmark datasets demonstrate that our RRV descriptor learning discriminative motion patterns can achieve superior results for various recognition tasks. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IEEE Trans. Cybern. | 3 |
| 2017 | MSM-HOG: A flexible trajectory descriptor for rigid body motion recognitionabstractThis paper proposes a flexible descriptor for representing 6-D rigid body motion trajectories, which not only shows strong invariances and descriptive ability but also achieves satisfactory results in both recognition accuracy and efficiency. 6-D rigid body motion trajectories are first transformed into the Multi-layer Self-similarity Matrices (MSM) representation. The MSM is the combination of the square similarity matrices in three layers, which captures both local and global spatiotemporal features of the trajectories. Next, the Histogram of Oriented Gradients (HOG) features extracted from the MSM representation are concatenated as the final MSM-HOG trajectory descriptor. Then we train the Support Vector Machine (SVM) classifier with the linear kernel for multicalss motion recognition. Finally, rigid body motion recognition experiments on two public datasets are conducted to verify the effectiveness and efficiency of the proposed method. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IROS | 3 |
| 2017 | On Multiscale Self-Similarities Description for Effective Three-Dimensional/Six-Dimensional Motion Trajectory RecognitionabstractMotion trajectories provide compact informative clues in characterizing motion behaviors of human bodies, robots, and moving objects. This paper devises an invariant and unified descriptor for three-dimensional/six-dimensional (3-D/6-D) motion trajectories recognition by exploring the latent motion patterns in the multiscale self-similarity matrices (MSM) within a motion trajectory and its components. The MSM approach transforms a motion trajectory in Euclidean space into a set of similarity matrices and exhibits strong invariances, in which each matrix can be regarded as a grayscale image. Next, the histograms of oriented gradients features extracted from the MSM representation are concatenated as the final trajectory descriptor. In addition, an improved kernel MSM is raised by calculating the pairwise kernel distances. Finally, extensive 3-D/6-D motion trajectory recognition experiments on three public datasets with a linear support vector machine classifier are conducted to verify the effectiveness and efficiency of the proposed approach. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IEEE Trans. Ind. Informatics | 3 |
| 2016 | On Integral Invariants for Effective 3-D Motion Trajectory Matching and RecognitionabstractMotion trajectories tracked from the motions of human, robots, and moving objects can provide an important clue for motion analysis, classification, and recognition. This paper defines some new integral invariants for a 3-D motion trajectory. Based on two typical kernel functions, we design two integral invariants, the distance and area integral invariants. The area integral invariants are estimated based on the blurred segment of noisy discrete curve to avoid the computation of high-order derivatives. Such integral invariants for a motion trajectory enjoy some desirable properties, such as computational locality, uniqueness of representation, and noise insensitivity. Moreover, our formulation allows the analysis of motion trajectories at a range of scales by varying the scale of kernel function. The features of motion trajectories can thus be perceived at multiscale levels in a coarse-to-fine manner. Finally, we define a distance function to measure the trajectory similarity to find similar trajectories. Through the experiments, we examine the robustness and effectiveness of the proposed integral invariants and find that they can capture the motion cues in trajectory matching and sign recognition satisfactorily. Zhanpeng Shao, Youfu Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2015 | Integral invariants for space motion trajectory matching and recognition
Zhanpeng Shao, Youfu Li 0001 |
Pattern Recognit. | 1 |
| 2014 | Online visual object tracking with supervised sparse representation and learningabstractIn this paper, an online visual object tracking algorithm based on the discriminative sparse representation framework with supervised learning is proposed. Different from the generative sparse representation based tracking algorithms, the proposed method casts the tracking problem into a binary classification task. A linear classifier is embedded into the sparse representation model by incorporating the classification error into the objective function to achieve discriminative classification. The dictionary and the classifier are jointly trained using the online dictionary learning algorithm, thus allow the model can adapt the dynamic variations of target appearance and background environment. The target locations are updated based on the classification score and the greedy search motion model. The proposed method is evaluated using four benchmark datasets and is compared with three state-of-the-art tracking algorithms. The results show that the discriminative sparse representation facilitates the tracking performance. Tianxiang Bai, Youfu Li 0001, Zhanpeng Shao |
ICARCV | 3 |
| 2013 | Learning appearance manifolds with structured sparse representation for robust visual trackingabstractThis paper presents a novel algorithm for robust visual object tracking based on the structured sparse representation framework. Conventional structured sparse representation based tracker models the nonlinear appearance manifold with a single subspace that is difficult to handle significant pose and illumination changes. Different from the afore-mentioned method, the proposed algorithm approximates the nonlinear appearance manifold by multiple low dimensional subspaces computed by an incremental learning scheme based on the merging and insert strategy. In order to enhance the discriminative power of the model, a number of clustered background subspaces are also added into the basis library and updated during tracking. With the Block Orthogonal Matching Pursuit (BOMP) algorithm, we show that the complex nonlinear appearance manifold can effectively represent by a sparse linear combination of structured union of subspaces. Experiments on benchmark video sequences show that the new structured sparse representation model improves the robustness of tracking. Tianxiang Bai, Youfu Li 0001, Zhanpeng Shao |
ICRA | 3 |
| 2013 | A new descriptor for multiple 3D motion trajectories recognitionabstractMotion trajectory gives a meaningful and informative clue in characterizing the motions of human, robots or moving objects. Hence, the descriptor for motion trajectory plays an importance role in motion recognition for many robotic tasks. However, an effective and compact descriptor for multiple 3D motion trajectories under complicated situation is lacking. In this paper, we propose a novel invariant descriptor for multiple motion trajectories based on the kinematic relation among multiple moving parts. There are two kinds of kinematic relation among multiple trajectories: articulated and independent trajectories. Spherical coordinate system is introduced to get a uniform and compact representation for both kinds of trajectories, where the relative trajectory concept are firstly defined based on orientation and distance changes in favor of acquiring relative movement features of each child trajectory with respect to the root trajectory. Then, by incorporating both the differential invariants of root trajectory and orientation, distance variations of each relative trajectory respectively, the new descriptor is constructed. Finally, effectiveness and robustness of our proposed new descriptor for multiple trajectories under complex circumstance are validated by the conducted two experiments for sign language and human action recognition. Zhanpeng Shao, Youfu Li 0001 |
ICRA | 1 |