Zhenjun Zhang

dblp:171/6089 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-3087-8239ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Face, body and person analysis · 67% Efficient and distributed learning · 33%
Computer graphics and multimedia
2 papers
Image and video processing · 100%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › gait analysis › gait recognition
event-based gait recognition
1.012026
EdinoGait: Transferring Large Visual Models to Event-Based Vision for Enhancing Gait Recognition · IEEE Trans. Multim. 2026
Computer vision › Face, body and person analysis › gait analysis
gait recognition
1.012026
EdinoGait: Transferring Large Visual Models to Event-Based Vision for Enhancing Gait Recognition · IEEE Trans. Multim. 2026
Machine learning › Efficient and distributed learning › model compression
lightweight neural network
1.012026
BCNet: Butterfly-Shaped Convolutions Network for Lightweight Edge Detection · IEEE Trans. Multim. 2026
Image and video processing
edge detection
1.012026
BCNet: Butterfly-Shaped Convolutions Network for Lightweight Edge Detection · IEEE Trans. Multim. 2026
Image and video processing › motion analysis › human motion analysis
gait recognition
0.812024
EGST: An Efficient Solution for Human Gaits Recognition Using Neuromorphic Vision Sensor · IEEE Trans. Inf. Forensics Secur. 2024
Image and video processing
motion analysis
0.812024
EGST: An Efficient Solution for Human Gaits Recognition Using Neuromorphic Vision Sensor · IEEE Trans. Inf. Forensics Secur. 2024
Wearable and physiological sensing › camera-based sensing
event camera
0.212024
EGST: An Efficient Solution for Human Gaits Recognition Using Neuromorphic Vision Sensor · IEEE Trans. Inf. Forensics Secur. 2024

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 2.0conditional entropy · 2.0butterfly-shaped convolution · 2.0transformer · 1.5spatio-temporal pattern extraction · 1.5event graph sequence · 1.5event encoder · 1.0dual alignment · 1.0contrastive loss · 1.0autoencoder · 1.0
YearPublicationVenuePosition
2026 Visual intelligence-driven hybrid feature learning with efficient dual-stage recurrent attention network for human activity recognition
Tariq Ahmad 0004, Zhenjun Zhang, Asif Rahim, Kamal M. Othman
Pattern Recognit.3
2026 Fusing Events and Frames for Robust Gait Recognition
abstract
Reliable gait recognition under low-light conditions remains challenging for traditional cameras. Event cameras, with their high dynamic range and fine temporal resolution, offer a promising alternative but suffer from sparse signals under small motion amplitudes and vulnerability to noise or irrelevant movements (e.g., shadows). To address these issues, we propose FusionGait, a complementary fusion framework that integrates event with standard frames to achieve robust gait recognition. Specifically, we propose a Self-supervised Hierarchical Feature Extractor (SSL-HFE) built upon DINOv2, which employs learnable prompts to bridge the gap between gray frames and RGB frames, extracts multi-level semantic features, and enhances their discriminability through a self-supervised learning strategy. Then, we introduce the Complementary Fusion Learning Module (CFLM), which employs cross-cost volumes to explicitly model pixel-level correlations between frames and events, enabling effective cross-modal interaction and fusion. Furthermore, we propose EvGSimulator, a sensor-specific data augmentation strategy that simulates diverse illumination conditions based on physical properties. The framework maintains robustness even when real frames are unavailable by reconstructing frame-like representations from events, and it scales to ultra–high-frame-rate scenarios with hundreds of frames per second. We also collect DAVIS346-Gait-RGE, the multi-view semi-indoor & outdoor gait dataset captured with a DAVIS346 event camera, including three modalities: event streams, gray frames, and reconstructed frames. Experiments across multiple datasets show FusionGait achieves state-of-the-art performance, effectively surpassing single-modality methods.
Liaogehao Chen, Zhenjun Zhang, Changchang Li, Yaonan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 EdinoGait: Transferring Large Visual Models to Event-Based Vision for Enhancing Gait Recognition
abstract
Current gait recognition methods heavily rely on various gait representations (e.g., silhouette sequences) generated by task-specific, supervised upstream processes, which inevitably incur high annotation costs and the risk of cumulative errors. Recently, generic knowledge from task–agnostic large visual models (LVMs) has been successfully applied to gait recognition, freeing the field from such dependencies. However, this approach does not address challenges posed by traditional cameras in handling scenarios with low latency, high speed, and high dynamic range. In this paper, we introduce EdinoGait, a novel and effective gait recognition framework that leverages event-based LVMs to overcome the scarcity of large-scale event-based datasets. Specifically, due to the distinct modality gap between image and event data and the lack of large-scale datasets, transferring LVMs to event-based vision is non-trivial. To address this, we introduce a novel event encoder that mitigates the modality gap through event prompts and a$CLS$patch contrastive loss. Subsequently, we design an autoencoder-based dual-alignment module to eliminate background noise brought by LVMs while preserving the motion details provided by event data. Additionally, to promote the application of event cameras in gait recognition, we collect the first semi-indoor, multi-view gait dataset captured by the DAVIS346 event camera. This dataset comprises 6,150 sequences (two modalities: grayscale images and event streams) of 41 subjects captured under two lighting conditions and five view angles ($0^{\circ }$,$45^{\circ }$,$90^{\circ }$,$135^{\circ }$, and$180^{\circ }$). Specifically, for each lighting condition and viewing angle, there are six sequences representing normal walking (NM), three representing walking with a backpack (BG), three with a portable bag (PT), and three with a coat (CL). Comprehensive experiments conducted on our event-based gait dataset and EV-CASIA-B demonstrate that EdinoGait significantly outperforms frame-based LVMs. Notably, under low-light conditions, the recognition accuracy of frame-based LVMs declines sharply, while EdinoGait exhibits robust performance.
Liaogehao Chen, Zhenjun Zhang, Yaonan Wang 0001
IEEE Trans. Multim.2
2026 BCNet: Butterfly-Shaped Convolutions Network for Lightweight Edge Detection
abstract
Aiming at the multi-task optimization conflicts (structure-detail-denoising coupling) and high computational costs caused by existing edge detectors' reliance on complex pre-trained models, this paper proposes BCNet - the first innovative framework that synergistically integrates biological visual mechanisms and information theory to achieve triple decoupling. First, inspired by butterfly-shaped receptive fields in the visual system, we design learnable butterfly-shaped convolution kernels as fundamental operators. These kernels inherently enhance structural perception and suppress noise without requiring deep architectures and pre-training, achieving structure-denoising separation. Second, to address the detail-noise coupling issue in existing methods that reconstruct edge images from multi-scale downsampled features, we propose a conditional entropy-based uncertainty modeling approach. The uncertainty of feature loss during downsampling is quantified via Gaussian distributions, while a self-supervised mechanism dynamically assigns detaillearning weights, enabling the model to learn more details from high-uncertainty regions while avoiding noise introduction. With at most about 2M parameters and real-time inference speeds up to 152 FPS, BCNet demonstrates strong competitiveness across four benchmark datasets, providing a novel, efficient, and lightweight solution for edge detection.https://github.com/StarkLuo/BCNetCode will be available.
Zhengqiao Luo, Zhenjun Zhang, Chuan Lin 0003, Yaonan Wang 0001
IEEE Trans. Multim.2
2025 MCFNet: Multiscale Cross-Modal Fusion Network for Remote Sensing Image Semantic Segmentation
abstract
Multimodal fusion methods have made great advancement in the field of remote sensing image segmentation in recent years. However, the efficient integration of local and global features from multiple modalities to improve the robustness of segmentation model remains a challenging task. In this paper, we propose a multiscale cross-modal fusion network (MCFNet) for semantic segmentation of high-resolution remote sensing images. This model employs a symmetrical dual-CNN encoder and a lightweight MLP decoder to learn comprehensive feature representation across diverse modalities. Specifically, to refine modality features at different stages, we construct a feature fusion module (FFM) that fuses rich complementary information by gradually aggregating local and global features. Additionally, a multiscale cross-modal feature correlation module (MCFCM) to deeply capture and correlate interactive feature information from the fused features. Finally, an efficient multiscale enhancement module (MEM) is leveraged to further enhance multiscale information extraction. Extensive comparison experiments on two benchmark datasets reveal that our method obtains superior performance compared with state-of-the-art methods. The codes will be available athttps://github.com/DrWuHonglin/MCFNet.
Zhaobin Zeng, Zhenjun Zhang
IEEE Signal Process. Lett.3
2024 EGST: An Efficient Solution for Human Gaits Recognition Using Neuromorphic Vision Sensor
abstract
Traditional cameras struggle to perform in challenging scenarios such as low latency, high speed and high dynamic range. In contrast, neuromorphic vision sensors (event cameras) have great potential for robotics and computer vision due to the advantages of high temporal resolution, high dynamic range, and ultra-low resource consumption. Event cameras are novel bio-inspired sensors that monitor the brightness change of each pixel asynchronously and provide a stream of events encoding the time, position and sign of the brightness changes. Hence, traditional computer vision methods cannot be directly applied to the event-stream. Finding event representations that completely maintain event attributes, as well as efficient and accurate learning approaches, is the key to unlocking the potential of event cameras. In this study, we reveal the rigid transfer from event-stream to graph that has been overlooked in previous work and introduce a novel event representation, namely event graph sequence (EGS) considering the local and global temporal clues. Coupled with EGS, we propose a spatio-temporal pattern extracting (STPE) module to capture the spatio-temporal correlation and evolution of EGS. Our novel framework, Event Graph Sequence Transformer (EGST), exploits event properties to provide efficient and accurate recognition. This study focuses on the event-based human gaits recognition task, and EGST is evaluated on three different event-based gait datasets. The evaluation results show better or comparable accuracy than the state-of-the-art, while requiring extremely low computation resources. The code will be available athttps://github.com/C19h/EGST.
Liaogehao Chen, Zhenjun Zhang, Yang Xiao 0007, Yaonan Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2022 Discriminative Multi-View Dynamic Image Fusion for Cross-View 3-D Action Recognition
abstract
Dramatic imaging viewpoint variation is the critical challenge toward action recognition for depth video. To address this, one feasible way is to enhance view-tolerance of visual feature, while still maintaining strong discriminative capacity. Multi-view dynamic image (MVDI) is the most recently proposed 3-D action representation manner that is able to compactly encode human motion information and 3-D visual clue well. However, it is still view-sensitive. To leverage its performance, a discriminative MVDI fusion method is proposed by us via multi-instance learning (MIL). Specifically, the dynamic images (DIs) from different observation viewpoints are regarded as the instances for 3-D action characterization. After being encoded using Fisher vector (FV), they are then aggregated by sum-pooling to yield the representative 3-D action signature. Our insight is that viewpoint aggregation helps to enhance view-tolerance. And, FV can map the raw DI feature to the higher dimensional feature space to promote the discriminative power. Meanwhile, a discriminative viewpoint instance discovery method is also proposed to discard the viewpoint instances unfavorable for action characterization. The wide-range experiments on five data sets demonstrate that our proposition can significantly enhance the performance of cross-view 3-D action recognition. And, it is also applicable to cross-view 3-D object recognition. The source code is available at https://github.com/3huo/ActionView.
Yancheng Wang 0002, Yang Xiao 0007, Zhiguo Cao 0001, Zhenjun Zhang, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.6
2017 Traffic Sign Recognition via Multi-Modal Tree-Structure Embedded Multi-Task Learning
abstract
Traffic sign recognition is a rather challenging task for intelligent transportation systems since signs in different subsets, e.g., speed limit signs, prohibition signs, and mandatory signs, are very different from each other in color or shape, whereas they share some similarities to the ones in the same subset. Therefore, it is important to integrate different modalities of visual features, such as color and shape, and select discriminative features for better sign description; in addition, it benefits to explore the correlations between the classes of traffic signs to learn the classifiers jointly to improve the generalization performance. In this paper, we propose Multi- Modal tree-structure embedded Multi-Task Learning called M2- tMTL to select discriminative visual features both between and within modalities, as well as the correlated features shared by similar classification tasks. Our method simultaneously introduces two structured sparsity-induced norms into a least squares regression. One of the norms can be used not only to select modality of features but also to conduct within-modality feature selection. Moreover, the hierarchical correlations among the classification tasks are well represented by a tree structure, and therefore, the tree-structure sparsity-induced norm is used for learning the regression coefficients jointly to boost the performance of multi-class traffic sign recognition. Alternating direction method of multipliers (ADMM) is used to efficiently solve the proposed model with guaranteed convergence. Extensive experiments on public benchmark data sets demonstrate that the proposed algorithm leads to a quite interpretable model, and it has better or competitive performance with several state-of-the-art methods but with less computational and memory cost.
Xiao Lu 0002, Yaonan Wang 0001, Xuanyu Zhou, Zhenjun Zhang, Zhigang Ling
IEEE Trans. Intell. Transp. Syst.4
2016 Bi-level weighted multi-view clustering via hybrid particle swarm optimization
Bo Jiang 0016, Feiyue Qiu, Zhenjun Zhang
Inf. Process. Manag.4
2015 A Method to Calibrate Vehicle-Mounted Cameras Under Urban Traffic Scenes
abstract
We address the problem of vehicle-mounted camera calibration under urban traffic scenes regarding the fact that the traditional calibration methods are practically restricted, since the internal parameters should be calibrated in the laboratory and it is impossible for recalibration that resulted from the parameters drifting or re-focusing when driving on roads. In this paper, we propose to utilize the manual lines lying in Manhattan directions in the scenes to compute their corresponding vanishing points for camera calibration, as the urban traffic scenes are usually man-made and the important lines and signs for driving are typically lying in the Manhattan directions. For “Manhattan world” scenes, where there are plenty of lines lying in Manhattan directions, the lines in the scene are detected automatically, and the clusters corresponding to Manhattan directions are obtained using RANSAC-like methods. For the more general “quasi-Manhattan world” scenes, where only the lines in two directions can be found naturally, while the lines in the other direction are usually detected trivially or even can be hardly detected, we propose a method to estimate the lines in the third direction to improve the vanishing point estimation accuracy. The method proposed is tested on both two types of scenes, and the accuracy and practicability of this method are demonstrated. Furthermore, calibration experiments on both one image and multiple images are conducted, which show that the results can be more accurate when more images are used.
Yaonan Wang 0001, Xiao Lu 0002, Zhigang Ling, Yimin Yang 0001, Zhenjun Zhang, Kena Wang
IEEE Trans. Intell. Transp. Syst.5