VLDB 2026 Research / reviewers in the wild / expert
Zhiyang Lu
dblp:260/7699
· DBLP profile ↗
13ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range ConditionsabstractGait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and cross-distance scenarios under real-world conditions. To address this gap, we present LRGait, the first LiDAR-Camera multimodal benchmark designed for robust long-range gait recognition across diverse outdoor distances and environments. We further propose EMGaitNet, an end-to-end framework tailored for long-range multimodal gait recognition. To bridge the modality gap between RGB images and point clouds, we introduce a semantic-guided fusion pipeline. A CLIP-based Semantic Mining (SeMi) module first extracts human body-part-aware semantic cues, which are then employed to align 2D and 3D features via a Semantic-Guided Alignment (SGA) module within a unified embedding space. A Symmetric Cross-Attention Fusion (SCAF) module hierarchically integrates visual contours and 3D geometric features, and a Spatio-Temporal (ST) module captures global gait dynamics. Extensive experiments on various gait datasets validate the effectiveness of our method. Zhiyang Lu, Tianren Wu, Changwang Zhang, Ming Cheng 0002 |
AAAI | 1 |
| 2026 | Language as a Bridge: Semantic-Guided Cross-Modal Gait Recognition via Text Prototype and Feature DecouplingabstractGait recognition aims to identify individuals based on walking patterns in a long-range, contactless manner. While camera-based methods have advanced significantly, their performance deteriorates under poor lighting conditions. LiDAR offers a promising alternative by capturing accurate 3D gait information regardless of illumination. However, effectively integrating heterogeneous data from diverse sensors, such as LiDAR and cameras, remains a key challenge for cross-modal gait recognition. Existing approaches often minimize modality discrepancy directly, which can lead to class collapse and damage to inter-class discriminability. To overcome these limitations, we propose a Semantic-Guided Cross-modal Gait recognition framework, SG-CrossGait, that introduces text features as the prototype space to bridge camera and LiDAR modalities. We design structured Gait Description Factors (GDF) and leverage multimodal large language models (MLLMs) for automatic factor annotation and text generation, enriching existing datasets with textual descriptions, yielding SUSTech1K-Text and FreeGait-Text. A CLIP-based pipeline aligns multi-grained representations from both modalities to the text prototype space. We further propose the Dual-stream Cross-attention Fusion (DCF) module for fine-grained feature integration and the Semantic-Guided Feature Decoupling (SGFD) module to disentangle shared and modality-specific features. A Multi-task Training (MT) scheme incorporating Gait Attribute Recognition (GAR) further enhances intra-class compactness. Extensive experiments validate the effectiveness of our approach. On SUSTech1K-Text, our method achieves 61% accuracy in LiDAR-to-Camera recognition, outperforming the state-of-the-art method by 8.3%. We also release the Gait-Text benchmark to promote future research at the intersection of gait analysis and vision-language learning. Code and datasets are available at: https://github.com/O-VIGIA/SCCG.git. Zhiyang Lu, Wankang Zeng, Ming Cheng 0002, Cheng Wang 0003 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | UniCrossGait: Unified Cross-Modal Gait Recognition Based on Knowledge DistillationabstractGait recognition provides a crucial biometric identification technique for intelligent security since it is long-range and nonintrusive. Previous methods have generally focused on gait recognition within a single modality, but with the constant advancement of multimedia sensors, there is a growing requirement for the development of multimodal and cross-modal gait recognition. The various modalities are aligned directly in a hard manner by current cross-modal approaches, which could result in class collapse caused by imbalanced learning and make learning discriminative representations difficult for the weaker modality. To overcome these limitations, we propose a unified cross-modal gait recognition framework, UniCrossGait. Specifically, we employ the knowledge distillation paradigm to transfer the knowledge of a fused multimodal teacher network, Exist Methods UniCrossGait(Ours) Figure 1: Illustration of comparison between exist methods and which serves as a shared intermediate representation, to unimodal student networks, which enables them to learn modality-invariant representations. To perform knowledge distillation effectively, we design the Direction-level Feature Imitation loss and Intra&Inter sample Correlation loss. Furthermore, we construct the Warm up operator and Distillation Balancing operator to prevent the network from falling into suboptimal solutions due to pure imitation and imbalanced distillation. Our method achieves state of-the-art (SOTA) results on cross-modal gait recognition datasets, and extensive ablation experiments demonstrate the effectiveness of the proposed paradigm and modules. Our code is available at https://github.com/O-VIGIA/UniCrossGait. Zhiyang Lu, Ming Cheng 0002, Cheng Wang 0003 |
IEEE Trans. Multim. | 1 |
| 2025 | SSRFlow: Semantic-Aware Fusion with Spatial Temporal Re-Embedding for Real-World Scene FlowabstractScene flow, which provides the 3D motion field of the first frame from two consecutive point clouds, is vital for dynamic scene perception. However, contemporary scene flow methods face three major challenges. Firstly, they only consider the context of individual point clouds before flow embedding, leading to embedded points struggling to perceive the consistent semantic relationship of another frame. To address this issue, we propose a novel approach called Dual Cross Attentive (DCA) for the latent fusion and alignment between two frames based on semantic contexts. This is then integrated into Global Fusion Flow Embedding (GF) to initialize flow embedding based on global correlations in both contextual and Euclidean spaces. Secondly, deformations exist in non-rigid objects after the warping layer, which distorts the spatiotemporal relation between the consecutive frames. For a more precise estimation of residual flow at next-level, the Spatial Temporal Re-embedding (STR) module is devised to update the point sequence features at current-level. Lastly, poor generalization is often observed due to the significant domain gap between synthetic and LiDAR-scanned datasets. We leverage novel domain adaptive losses to effectively bridge the gap of motion inference from synthetic to real-world. Experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance across various datasets, with particularly outstanding results in real-world LiDAR-scanned situations. Zhiyang Lu, Qinghan Chen, Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
3DV | 1 |
| 2025 | MOJO: MOtion Pattern Learning and JOint-Based Fine-Grained Mining for Person Re-Identification Based on 4D LiDAR Point Clouds
Zhiyang Lu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | An Entropy-Based Pseudo-Label Mixup Method for Source-Free Domain Adaptation
Qinghan Chen, Zhiyang Lu, Ming Cheng 0002 |
PRCV (2) | 2 |
| 2024 | Multilevel Interactive Enhanced Network for Infrared Small-Target DetectionabstractInfrared small target detection (IRSTD) aims to identify small and faint targets amidst cluttered background in infrared images, which is vital for applications like maritime surveillance. Traditional methods struggle due to low signal-to-noise ratio (SNR) and contrast. However, recent CNN-based approaches show promise, leveraging deep learning’s strong modeling capabilities. In this letter, we propose a multilevel interactive enhanced network (MIE-Net). In MIE-Net, we use multiple backbones that have progressively decreasing numbers of blocks. Features transfer and information interaction are carried out between different backbones. We designed an attention mechanism-based feature filter (AFF) to reduce background noise interference by filtering the low-level features with high-level features. Furthermore, we proposed a global information enhancement module (GIEM), through which features are enhanced as they are delivered, while further mitigating the problem of small target loss. Experiments on public datasets validate the effectiveness of our method. MIE-Net outperforms the current state-of-the-art (SOTA) methods by approximately 6% in terms of the intersection over union (IoU). There was also about a 2% increase in average area under the curve (AUC). Youliang Chu, Ming Cheng 0002, Zhiyang Lu, Zhentao Xiong, Cheng Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | GMA3D: Local-Global Attention Learning to Estimate Occluded Motions of Scene Flow
Zhiyang Lu, Ming Cheng 0002 |
PRCV (2) | 1 |
| 2023 | Two-Stage Self-Supervised Cycle-Consistency Transformer Network for Reducing Slice Gap in MR ImagesabstractMagnetic resonance (MR) images are usually acquired with large slice gap in clinical practice, i.e., low resolution (LR) along the through-plane direction. It is feasible to reduce the slice gap and reconstruct high-resolution (HR) images with the deep learning (DL) methods. To this end, the paired LR and HR images are generally required to train a DL model in a popular fully supervised manner. However, since the HR images are hardly acquired in clinical routine, it is difficult to get sufficient paired samples to train a robust model. Moreover, the widely used convolutional Neural Network (CNN) still cannot capture long-range image dependencies to combine useful information of similar contents, which are often spatially far away from each other across neighboring slices. To this end, a Two-stage Self-supervised Cycle-consistency Transformer Network (TSCTNet) is proposed to reduce the slice gap for MR images in this work. A novel self-supervised learning (SSL) strategy is designed with two stages respectively for robust network pre-training and specialized network refinement based on a cycle-consistency constraint. A hybrid Transformer and CNN structure is utilized to build an interpolation model, which explores both local and global slice representations. The experimental results on two public MR image datasets indicate that TSCTNet achieves superior performance over other compared SSL-based algorithms. Zhiyang Lu, Jian Wang 0135, Shihui Ying, Jun Wang 0024, Jun Shi 0004, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | A Channel Attention Based MLP-Mixer Network for Motor Imagery Decoding With EEGabstractConvolutional neural networks (CNNs) and their variants have been successfully applied to the electroencephalogram (EEG) based motor imagery (MI) decoding task. However, these CNN-based algorithms generally have limitations in perceiving global temporal dependencies of EEG signals. Besides, they also ignore the diverse contributions of different EEG channels to the classification task. To address such issues, a novel channel attention based MLP-Mixer network (CAMLP-Net) is proposed for EEG-based MI decoding. Specifically, the MLP-based architecture is applied in this network to capture the temporal and spatial information. The attention mechanism is further embedded into MLP-Mixer to adaptively exploit the importance of different EEG channels. Therefore, the proposed CAMLP-Net can effectively learn more global temporal and spatial information. The experimental results on the newly built MI-2 dataset indicate that our proposed CAMLP-Net achieves superior classification performance over all the compared algorithms. Yanbin He, Zhiyang Lu, Jun Wang 0024, Jun Shi 0004 |
ICASSP | 2 |
| 2022 | A Convolutional Neural Network and Graph Convolutional Network Based Framework for Classification of Breast Histopathological ImagesabstractThe spatial correlation among different tissue components is an essential characteristic for diagnosis of breast cancers based on histopathological images. Graph convolutional network (GCN) can effectively capture this spatial feature representation, and has been successfully applied to the histopathological image based computer-aided diagnosis (CAD). However, the current GCN-based approaches need complicated image preprocessing for graph construction. In this work, we propose a novel CAD framework for classification of breast histopathological images, which integrates both convolutional neural network (CNN) and GCN (named CNN-GCN) into a unified framework, where CNN learns high-level features from histopathological images for further adaptive graph construction, and the generated graph is then fed to GCN to learn the spatial features of histopathological images for the classification task. In particular, a novel clique GCN (cGCN) is proposed to learn more effective graph representation, which can arrange both forward and backward connections between any two graph convolution layers. Moreover, a new group graph convolution is further developed to replace the classical graph convolution of each layer in cGCN, so as to reduce redundant information and implicitly select superior fused feature representation. The proposed clique group GCN (cgGCN) is then embedded in the CNN-GCN framework (named CNN-cgGCN) to promote the learned spatial representation for diagnosis of breast cancers. The experimental results on two public breast histopathological image datasets indicate the effectiveness of the proposed CNN-cgGCN with superior performance to all the compared algorithms. Zhiyang Gao, Zhiyang Lu, Jun Wang 0024, Shihui Ying, Jun Shi 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | S2Q-Net: Mining the High-Pass Filtered Phase Data in Susceptibility Weighted Imaging for Quantitative Susceptibility MappingabstractSusceptibility weighted imaging (SWI) is a routine magnetic resonance imaging (MRI) sequence that combines the magnitude and high-pass filtered phase images to qualitatively enhance the image contrasts related to tissue susceptibility. Tremendous amounts of the high-pass filtered phase data with low signal to noise ratio and incomplete background field removal have thus been collected under default clinical settings. Since SWI cannot quantitatively estimate the susceptibility, it is thus non-trivial to derive quantitative susceptibility mapping (QSM) directly from these redundant phase data, which effectively promotes the mining of the SWI data collected previously. To this end, a novel deep learning based SWI-to-QSM-Net (S2Q-Net) is proposed for QSM reconstruction from SWI high-pass filtered phase data. S2Q-Net firstly estimates the edge maps of QSM to integrate edge prior into features, which benefits the network to reconstruct QSM with realistic and clear tissue boundaries. Furthermore, a novel Second-order Cross Dense Block is proposed in S2Q-Net, which can capture rich inter-region interactions to provide more non-local phase information related to local tissue susceptibility. Experimental results on both simulated and in-vivo data indicate its superiority over all the compared deep learning based QSM reconstruction methods. Zhiyang Lu, Rongjun Ge, Hongjian He, Jun Shi 0004 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Two-Stage Self-supervised Cycle-Consistency Network for Reconstruction of Thin-Slice MR Images
Zhiyang Lu, Jun Wang 0024, Jun Shi 0004, Dinggang Shen |
MICCAI (6) | 1 |