EDBT 2026 Demo / reviewers in the wild / expert
Yisheng Zhu
dblp:13/6578
· DBLP profile ↗
17ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0001-9742-1984ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Prompting Spatial Temporal Actor Transformer for Fine-Grained Skeleton-Based Action RecognitionabstractThe scarcity of flexibility and effectiveness in skeleton models, combined with the characteristic of limited information in skeleton data, has resulted in fine-grained modeling insufficiently explored in recent skeleton-based action recognition. Confronting these challenges, we propose Dynamic prompting Spatial Temporal Actor transFormer (DSTAFormer), a powerful framework which neatly unifies vision and language. Specifically, we introduce a decoupled vision transformer, which consists of three components: Spatial transFormer (SF), Temporal transFormer (TF), and Actor transFormer (AF), to account for numerous visual aspects of the human body, namely spatial, temporal, and interactive relations. Compared to the vanilla transformer, we reformulate self-attention using Statistically-inspired Attention Reconstruction (SAR) module and Local-specific constraints, thereby enabling a more explicit and interpretable exploration of the action’s fine-grained compositions. The skeleton sequences are processed by this decoupled structure to generate the visual embeddings. To encode environmental interactions that skeletal coordinates inherently lack, we utilize Dynamic Prompting (Dp) strategy to generate visual-based textual prompts. These prompts are transformed into discriminative textual embeddings via a pre-trained large language model (LLM). We also design a Semantic Adapter (SA) to bridge the modality gap. The cross-modality embeddings are projected into a unified feature space for contrastive co-training. This infusion of knowledge into the skeleton data enhances its semantic richness, pushing the boundaries of fine-grained understanding. We evaluate our framework on NTU RGB+D, NTU RGB+D 120, and Toyota Smarthome datasets. DSTAFormer achieves comparable performance against state-of-the-arts. Yisheng Zhu, Guangcan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Decoupled spatio-temporal grouping transformer for skeleton-based action recognition
Shengkun Sun, Zihao Jia, Yisheng Zhu, Guangcan Liu, Zhengtao Yu 0001 |
Vis. Comput. | 3 |
| 2024 | A real-time semi-dense depth-guided depth completion network
JieJie Xu, Yisheng Zhu, Guangcan Liu |
Vis. Comput. | 2 |
| 2023 | Modeling the Relative Visual Tempo for Self-supervised Skeleton-based Action RecognitionabstractVisual tempo characterizes the dynamics and the temporal evolution, which helps describe actions. Recent approaches directly perform visual tempo prediction on skeleton sequences, which may suffer from insufficient feature representation issue. In this paper, we observe that relative visual tempo is more in line with human intuition, and thus providing more effective supervision signals. Based on this, we propose a novel Relative Visual Tempo Contrastive Learning framework for skeleton action Representation (RVTCLR). Specifically, we design a Relative Visual Tempo Learning (RVTL) task to explore the motion information in intra-video clips, and an Appearance-Consistency (AC) task to learn appearance information simultaneously, resulting in more representative spatiotemporal features. Furthermore, skeleton sequence data is much sparser than RGB data, making the network learn shortcuts, and overfit to low-level information such as skeleton scales. To learn high-order semantics, we further design a new Distribution-Consistency (DC) branch, containing three components: Skeleton-specific Data Augmentation (S-DA), Fine-grained Skeleton Encoding Module (FSEM), and Distribution-aware Diversity (DD) Loss. We term our entire method (RVTCLR with DC) as RVTCLR+. Extensive experiments on NTU RGB+D 60 and NTU RGB+D 120 datasets demonstrate that our RVTCLR+ can achieve competitive results over the state-of-the-art methods. Code is available at https://github.com/Zhuysheng/RVTCLR. Yisheng Zhu, Hu Han 0001, Zhengtao Yu 0001, Guangcan Liu |
ICCV | 1 |
| 2023 | Multilevel Spatial-Temporal Excited Graph Network for Skeleton-Based Action RecognitionabstractThe ability to capture joint connections in complicated motion is essential for skeleton-based action recognition. However, earlier approaches may not be able to fully explore this connection in either the spatial or temporal dimension due to fixed or single-level topological structures and insufficient temporal modeling. In this paper, we propose a novel multilevel spatial-temporal excited graph network (ML-STGNet) to address the above problems. In the spatial configuration, we decouple the learning of the human skeleton into general and individual graphs by designing a multilevel graph convolution (ML-GCN) network and a spatial data-driven excitation (SDE) module, respectively. ML-GCN leverages joint-level, part-level, and body-level graphs to comprehensively model the hierarchical relations of a human body. Based on this, SDE is further introduced to handle the diverse joint relations of different samples in a data-dependent way. This decoupling approach not only increases the flexibility of the model for graph construction but also enables the generality to adapt to various data samples. In the temporal configuration, we apply the concept of temporal difference to the human skeleton and design an efficient temporal motion excitation (TME) module to highlight the motion-sensitive features. Furthermore, a simplified multiscale temporal convolution (MS-TCN) network is introduced to enrich the expression ability of temporal features. Extensive experiments on the four popular datasets NTU-RGB+D, NTU-RGB+D 120, Kinetics Skeleton 400, and Toyota Smarthome demonstrate that ML-STGNet gains considerable improvements over the existing state of the art. Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Presswork defect inspection using only defect-free high-resolution images
Zhenyu Guan 0001, Yisheng Zhu, Guangcan Liu |
Vis. Comput. | 3 |
| 2023 | Multi-level feature fusion pyramid network for object detection
Zebin Guo, Hui Shuai, Guangcan Liu, Yisheng Zhu |
Vis. Comput. | 4 |
| 2022 | Self-Supervised Video Representation Learning Using Improved Instance-Wise Contrastive Learning and Deep ClusteringabstractInstance-wise contrastive learning (Instance-CL), which learns to map similar instances closer and different instances farther apart in the embedding space, has achieved considerable progress in self-supervised video representation learning. However, canonical Instance-CL does not handle properly the temporal similarities between different videos, limiting the representation capabilities of learned models. This paper presents a novel two-stage framework that combines Instance-CL and unsupervised clustering to progressively learn desirable temporal representations with high intra-class compactness. Specifically, (a) we first introduce a new consistency-preserving sampling strategy to generate positive/negative pairs. Compared to the traditional sampling methods, our sampling strategy focuses more on motion dynamics, resulting in more temporal-related feature representations. (b) To further explore the temporal similarities between videos so as to encourage intra-class compactness, we set temporal representations extracted from Instance-CL as an initializer, and iteratively use k-means clustering to generate pseudo-labels for training the encoder. We term our method as Improved Instance-CL with Deep Clustering (ICDC) and apply it to two downstream tasks, including action recognition and video retrieval. Extensive experimental results show that ICDC gains considerable improvements compared to the existing self-supervised methods. Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Collaborative Local-Global Learning for Temporal Action ProposalabstractTemporal action proposal generation is an essential and challenging task in video understanding, which aims to locate the temporal intervals that likely contain the actions of interest. Although great progress has been made, the problem is still far from being well solved. In particular, prevalent methods can handle well only the local dependencies (i.e., short-term dependencies) among adjacent frames but are generally powerless in dealing with the global dependencies (i.e., long-term dependencies) between distant frames. To tackle this issue, we propose CLGNet, a novel Collaborative Local-Global Learning Network for temporal action proposal. The majority of CLGNet is an integration of Temporal Convolution Network and Bidirectional Long Short-Term Memory, in which Temporal Convolution Network is responsible for local dependencies while Bidirectional Long Short-Term Memory takes charge of handling the global dependencies. Furthermore, an attention mechanism called the background suppression module is designed to guide our model to focus more on the actions. Extensive experiments on two benchmark datasets, THUMOS’14 and ActivityNet-1.3, show that the proposed method can outperform state-of-the-art methods, demonstrating the strong capability of modeling the actions with varying temporal durations. Yisheng Zhu, Hu Han 0001, Guangcan Liu, Qingshan Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2020 | Fine-grained action recognition using multi-view attentions
Yisheng Zhu, Guangcan Liu |
Vis. Comput. | 1 |
| 2006 | Automatic Speech Segmentation Combining an HMM-Based Approach and Recurrence Trend AnalysisabstractAiming at improving the speech segmentation accuracy acquired from standard HMM-based approach, this paper presents a nonlinear dynamical method for phoneme boundary adjustment by discerning and measuring the nonstationarity of speech dynamics. Dynamical systems of different phones present diversified invariant attractor structures in phase space. Therefore, when analyzing adjacent phones, there may exist a point, at which the underlying dynamics changes. In this study, time-dependent recurrence trend (TDRT) is proposed to describe the local changing degree of the nonstationarity of speech dynamics as time progress and identify the largest paling slop in the windowed recurrence plots (RPs) as the phoneme boundary. The experimental result shows that 9.41% increase in agreement within 20 ms with TDRT correction is obtained on TIMIT database. Run-Qiang Yan, Yiqing Zu, Yisheng Zhu |
ICASSP (1) | 3 |
| 2004 | Data compression with spherical wavelets and wavelets for the image-based relighting
Ze Wang 0018, Andrew Chi-Sing Leung, Yisheng Zhu, Tien-Tsin Wong |
Comput. Vis. Image Underst. | 3 |
| 2004 | Eigen-image based compression for the image-based relighting with cascade recursive least squared networks
Ze Wang 0018, Andrew Chi-Sing Leung, Tien-Tsin Wong, Yisheng Zhu |
Pattern Recognit. | 4 |
| 2004 | Data compression on the illumination adjustable images by PCA and ICA
Ze Wang 0018, Andrew Chi-Sing Leung, Yisheng Zhu, Tien-Tsin Wong |
Signal Process. Image Commun. | 3 |
| 2003 | An improved sequential method for principal component analysis
Ze Wang 0018, Yin Lee, Simone G. O. Fiori, Andrew Chi-Sing Leung, Yisheng Zhu |
Pattern Recognit. Lett. | 5 |
| 2003 | An improved optimal bit allocation method for sub-band coding
Ze Wang 0018, Yin Lee, Andrew Chi-Sing Leung, Tien-Tsin Wong, Yisheng Zhu |
Pattern Recognit. Lett. | 5 |
| 1999 | Fixed-point error analysis and an efficient array processor design of two-dimensional sliding DFT
Yisheng Zhu, Zhizhong Wang |
Signal Process. | 1 |