VLDB 2026 Research / reviewers in the wild / expert
Zepeng Wang 0002
dblp:207/1892-2
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-6782-0551ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 10 |
| 2024 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2024abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To advance the algorithm development and provide fair evaluations, the International Competition on Human Identification at a Distance (HID) has been held annually since 2020, with HID 2024 marking the fifth edition. Despite increased difficulty, participants demonstrated remarkable capabilities, surpassing previous accuracy levels. This paper, co-authored by competition organizers and top participants, provides a comprehensive summary of HID 2024, including an overview of the competition, and insights into the methods employed by the top teams. Specifically, inspired by the achievements of the 5 competitions of HID, we also provide the insights for the future directions on gait recognition. Shiqi Yu 0001, Weiming Wu, Jiacong Hu, Zepeng Wang 0002, Runsheng Wang, Yunfei Ni, Yongzhen Huang, Liang Wang 0001, Md. Atiqur Rahman Ahad |
IJCB | 4 |
| 2024 | Compressed Video Action Recognition With Dual-Stream and Dual-Modal TransformerabstractCompressed video action recognition offers the advantage of reducing decoding and inference time compared to the RGB domain. However, the compressed domain poses unique challenges with different types of frames (I-frames and P-frames). I-frames consistent with RGB are rich in frame information, but the redundant information may interfere with the recognition task. There are two modalities in P-frames, residual (R) and motion vector (MV). Although with less information, they can reflect the motion cue. To address these challenges and leverage the independent information from different frames and modalities, we propose a novel approach called Dual-Stream and Dual-Modal Transformer (DSDMT). Our approach consists of two streams: 1) The short-span P-frames stream contains temporal information. We propose the Dual-Modal Attention Module (DAM) to mine different modal variability in P-frames and complement the orthogonal feature vector. Besides, considering the sparsity of P-frames, we extract action features with Frame-level Patch Embedding (FPE) to avoid redundant computation. 2) The long-span I-frames stream extracts the global context feature of the entire video, including content and scene information. By fusing the global video context and local key-frame features, our model represents the action feature in terms of fine-grained and coarse-grained. We evaluated our proposed DSDMT on three public benchmarks with different scales: HMDB-51, UCF-101, and Kinetics-400. Ours achieve better performance with fewer Flops and lower latency. Our analysis shows that the independence and complements of the I-frames and P-frames extracted from the compressed video stream play a crucial role in action recognition. Yuting Mou, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun, Zepeng Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Feature Mixing and Disentangling for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID) has recently attracted lots of attention for its applicability in practical scenarios. However, previous pose-based methods always neglect the non-target pedestrian (NTP) problem. In contrast, we propose a feature mixing and disentangling method to train a robust network for occluded person Re-ID without extra data. Based on ViT, we design our network as follows: 1) A multi-target patch mixing (MPM) module is proposed to generate complex multi-target images with refined labels in the training stage. 2) We propose an identity-based patch realignment (IPR) module in the decoder layer to disentangle local features from the multi-target sample. In contrast to pose-guided methods, our approach overcomes the difficulties of NTP. More importantly, our approach does not bring additional computational costs in the training and testing phases. Experimental results show that our method effectively on occluded person Re-ID. For example, our method performs 3.3%/3.2% better than the baseline on Occluded-Duke in terms of mAP/rank-1 and outperforms the previous state-of-the-art. Zepeng Wang 0002, Ke Xu 0003, Yuting Mou, Xinghao Jiang |
ICME | 1 |
| 2022 | A Transformer-Based Cloth-Irrelevant Patches Feature Extracting Method for Long-Term Cloth-Changing Person Re-identification
Zepeng Wang 0002, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
CGI | 1 |