VLDB 2026 Research / reviewers in the wild / expert
Zhifan Wan
dblp:340/8664
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0009-2742-3613ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 54% Representation and self-supervised learning · 34% Generative modeling · 6% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
1.5 | 2 | 2025 | Collaboratively Self-Supervised Video Representation Learning for Action Recognition · IEEE Trans. Inf. Forensics Secur. 2025 Data-Efficient Masked Video Modeling for Self-supervised Action Recognition · ACM Multimedia 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
pretext task |
0.9 | 1 | 2025 | Collaboratively Self-Supervised Video Representation Learning for Action Recognition · IEEE Trans. Inf. Forensics Secur. 2025 |
Computer vision › Video understanding and tracking › video representation learning
self-supervised video representation learning |
0.9 | 1 | 2025 | Collaboratively Self-Supervised Video Representation Learning for Action Recognition · IEEE Trans. Inf. Forensics Secur. 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked video modeling |
0.7 | 1 | 2023 | Data-Efficient Masked Video Modeling for Self-supervised Action Recognition · ACM Multimedia 2023 |
Machine learning › Generative modeling › generative adversarial network
conditional GAN |
0.3 | 1 | 2025 | Collaboratively Self-Supervised Video Representation Learning for Action Recognition · IEEE Trans. Inf. Forensics Secur. 2025 |
Computer vision › Face, body and person analysis › human pose estimation
human pose forecasting |
0.3 | 1 | 2025 | Collaboratively Self-Supervised Video Representation Learning for Action Recognition · IEEE Trans. Inf. Forensics Secur. 2025 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 0.9contrastive learning · 0.9conditional GAN · 0.9progressive masking · 0.7optical flow · 0.73d tokenizer · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BIMM: Brain-Inspired Masked Modeling for Video Representation LearningabstractThe visual pathway of human brain includes two sub-pathways,i.e., the ventral pathway and the dorsal pathway, which focus on object identification and dynamic information modeling, respectively. Both pathways comprise multi-layer structures, with each layer responsible for processing different aspects of visual information. Inspired by the human visual information processing mechanism, we propose the Brain Inspired Masked Modeling (BIMM) framework, aiming to learn comprehensive representations from videos. Specifically, our approach consists of ventral and dorsal branches, which learn image and video representations, respectively. Both branches employ the Vision Transformer (ViT) as their backbone and are trained through a masked modeling method. To emulate the distinct functions of the visual cortices, we segment the encoder of each branch into three intermediate blocks and reconstruct progressive prediction targets with light weight decoders. Furthermore, drawing inspiration from the information-sharing mechanism in the brain’s visual pathways, we introduce a partial parameter sharing strategy between the branches during training. Extensive experiments demonstrate that BIMM achieves superior performance compared to the state-of-the-art methods. Jie Zhang 0071, Zhifan Wan, Sen Nie, Changzhen Li, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Collaboratively Self-Supervised Video Representation Learning for Action RecognitionabstractConsidering the close connection between action recognition and human pose estimation, we design a Collaboratively Self-supervised Video Representation (CSVR) learning framework specific to action recognition by jointly factoring in generative pose prediction and discriminative context matching as pretext tasks. Specifically, our CSVR consists of three branches: a generative pose prediction branch, a discriminative context matching branch, and a video generating branch. Among them, the first one encodes dynamic motion feature by utilizing Conditional-GAN to predict the human poses of future frames, and the second branch extracts static context features by contrasting positive and negative video feature and I-frame feature pairs. The third branch is designed to generate both current and future video frames, for the purpose of collaboratively improving dynamic motion features and static context features. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple popular video datasets. Jie Zhang 0071, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Data-Efficient Masked Video Modeling for Self-supervised Action RecognitionabstractRecently, self-supervised video representation learning based on Masked Video Modeling (MVM) has demonstrated promising results for action recognition. However, existing methods face two significant challenges: (1) video actions involve a crucial temporal dimension, yet current masking strategies adopt inefficient random approaches that undermine low-density dynamic motion clues in videos; (2) pre-training requires large-scale datasets and significant computing resources (including large batch sizes and enormous iterations). To address these issues, we propose a novel method named Data-Efficient Masked Video Modeling (DEMVM) for self-supervised action recognition. Specifically, a novel masking strategy named Flow-Guided Dense Masking (FGDM) is proposed to facilitate efficient learning by focusing more on the action-related temporal clues, which applies dense masking to dynamic regions based on optical flow priors, while sparse masking to background regions. Furthermore, DEMVM introduces a 3D video tokenizer to enhance the modeling of temporal clues. Finally, Progressive Masking Ratio (PMR) and 2D initialization strategies are presented to enable the model to adapt to the characteristics of the MVM paradigm during different training stages. Extensive experiments on multiple benchmarks, UCF101, HMDB51, and Mimetics, demonstrate that our method achieves state-of-the-art performance in the downstream action recognition task with both efficient data and low computational cost. More interestingly, the few-shot experiment on the Mimetics dataset shows that DEMVM can accurately recognize actions even in the presence of context bias. Qiankun Li 0004, Xiaolong Huang 0001, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang 0071, Shiguang Shan, Zengfu Wang |
ACM Multimedia | 3 |