VLDB 2026 Research / reviewers in the wild / expert
Zhanzhou Feng
dblp:352/8462
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0006-0071-454XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAN: Self-Calibrated AutoregressioN for High-Quality Visual GenerationabstractHuman artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCAN) model capable of self-evaluating and refining generation quality without regenerating the entire image. We unify image token generation and quality evaluation into a single autoregressive model, formulating both tasks as categorical prediction problems. During inference, the model first generates a coarse initial image, then iteratively refines the lowest-quality patches until satisfactory image quality is achieved. Experimental results demonstrate that SCAN effectively handles diverse real-world generation errors and achieves a promising balance between image quality and speed. For example, SCAN-XL achieves an FID of 2.10 and an IS of 326.1, surpassing the LlamaGen-XL by 1.29 (+38%) in FID and 99.0 (+43.6%) in IS, with a 5.6× speedup (19.76s to 3.56s). Compared to recent works, SCAN improves FID and speed by +18.3% and +23% over VAR-d20, and by +7% and +46% over RandAR-XL. Zhanzhou Feng, Qingpei Guo, Jingdong Chen, Ming Yang 0007, Shiliang Zhang |
AAAI | 1 |
| 2026 | VideoAligner: Text-driven feature decomposition for precise video-text alignment
Zhanzhou Feng, Shunan Mao, Yaowei Wang 0001, Shiliang Zhang |
Pattern Recognit. | 1 |
| 2026 | Efficient Human Feature Refinement for Weakly Supervised Group Activity RecognitionabstractWeakly supervised group activity recognition (WSGAR) aims to identify the joint activity of a group of people without relying on hand-annotated human bounding boxes. Existing WSGAR methods typically acquire coarse human-level features by pooling from detected bounding boxes or applying human queries with cross attentions. These approaches focus on learning human relations from the acquired features. However, discriminative person-specific clues might be confused with irrelevant backgrounds, hindering the effectiveness of downstream human relation learning. To address this limitation, we propose a Human Feature Refinement framework that enhances human-level information with graph convolutional networks and self-attention. We define in-box regions as tokens and learn their spatial correspondence through GCN and self-attention. By explicitly extracting in-box details and suppressing irrelevant regions, our method acquires more discriminative human-level features for relation learning and group activity prediction. We further propose a Graph-based Token Merging algorithm to reduce the computation cost of Human Feature Refinement, while minimizing information loss and overfitting risk. Experiments show that our method outperforms previous WSGAR methods on Volleyball, NBA and JRDB-PAR benchmarks, with reduced computation cost. Yihui Zhou, Hao Chen 0061, Zhanzhou Feng, Shiliang Zhang |
IEEE Trans. Multim. | 3 |
| 2025 | Unified Visual Generation via Next-Set Prediction in Continuous Domain
Zhanzhou Feng, Qingpei Guo, Xinyu Xiao, Ruihan Xu 0002, Ming Yang 0007, Shiliang Zhang |
ICCV | 1 |
| 2025 | Evolved Hierarchical Masking for Self-Supervised LearningabstractExisting Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those mask patterns resort to different criteria to depict image contents, sticking to a fixed pattern leads to a limited vision cues modeling capability. This paper introduces an evolved hierarchical masking method to pursue general visual cues modeling in self-supervised learning. The proposed method leverages the vision model being trained to parse the input visual cues into a hierarchy structure, which is hence adopted to generate masks accordingly. The accuracy of hierarchy is on par with the capability of the model being trained, leading to evolved mask patterns at different training stages. Initially, generated masks focus on low-level visual cues to grasp basic textures, then gradually evolve to depict higher-level cues to reinforce the learning of more complicated object semantics and contexts. Our method does not require extra pre-trained models or annotations and ensures training efficiency by evolving the training difficulty. We conduct extensive experiments on seven downstream tasks including partial-duplicate image retrieval relying on low-level details, as well as image classification and semantic segmentation that require semantic parsing capability. Experimental results demonstrate that it substantially boosts performance across these tasks. For instance, it surpasses the recent MAE by 1.1% in imageNet-1K classification and 1.4% in ADE20K segmentation with the same training epochs. We also align the proposed method with the current research focus on LLMs. The proposed approach bridges the gap with large-scale pre-training on semantic demanding tasks and enhances intricate detail perception in tasks requiring low-level feature recognition. Zhanzhou Feng, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Efficient Video Transformers via Spatial-temporal Token Merging for Action RecognitionabstractTransformer has exhibited promising performance in various video recognition tasks but brings a huge computational cost in modeling spatial-temporal cues. This work aims to boost the efficiency of existing video transformers for action recognition through eliminating redundancies in their tokens and efficiently learning motion cues of moving objects. We propose a lightweight and plug-and-play module, namely Spatial-temporal Token Merger (STTM), to merge the tokens belonging to the same object into a more compact representation. STTM first adaptively identifies crucial object clues underlying the video as meta tokens. Similarity scores between input tokens and meta tokens are hence computed and used to guide the fusion of similar tokens in both spatial and temporal domains, respectively. To compensate for motion cues lost in the merging procedure, we compute the linear aggregation of spatial-temporal positions of tokens as motion features. STTM hence outputs a compact set of tokens fusing both appearance and motion features of moving objects. This procedure substantially decreases the number of tokens that need to be processed by each Transformer block and boosts the efficiency. As a general module, STTM can be applied to different layers of various video Transformers. Extensive experiments on the action recognition datasets Kinectics-400 and SSv2 demonstrate its promising performance. For example, it reduces the computation complexity of ViT by 38% while maintaining a similar performance on Kinectics-400. It also brings 1.7% gains of top-1 accuracy on SSv2 under the same computational cost. Zhanzhou Feng, Lei Ma 0008, Shiliang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Evolved Part Masking for Self-Supervised LearningabstractExisting Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those patterns resort to different criteria to mask local regions, sticking to a fixed pattern leads to limited vision cues modeling capability. This paper proposes an evolved part-based masking to pursue more general visual cues modeling in self-supervised learning. Our method is based on an adaptive part partition module, which leverages the vision model being trained to construct a part graph, and partitions parts with graph cut. The accuracy of partitioned parts is on par with the capability of the pretrained model, leading to evolved mask patterns at different training stages. It generates simple patterns at the initial training stage to learn low-level visual cues, which hence evolves to eliminate accurate object parts to reinforce the learning of object semantics and contexts. Our method does not require extra pretrained models or annotations, and effectively ensures the training efficiency by evolving the training difficulty. Experiment results show that it substantially boosts the performance on various tasks including image classification, object detection, and semantic segmentation. For example, it outperforms the recent MAE by 0.69% on imageNet-1K classification and 1.61% on ADE20K segmentation with the same training epochs. Zhanzhou Feng, Shiliang Zhang |
CVPR | 1 |
| 2023 | Efficient Vision Transformer via Token MergerabstractVision Transformers (ViTs) split an image into fixed-size patches as tokens. This strategy has succeeded in computer vision tasks, but introduces considerable tokens similar in semantics and appearances. This work proposes Token Merger to spot redundant tokens and merge them into a compact representation to accelerate ViTs. For each forward inference, the Token Merger first identifies meta tokens to represent meaningful cues of the image content, then adaptively merges similar tokens into a uniform one referring to meta tokens. To pursue a reasonable tradeoff between accuracy and efficiency, we further introduce learnable gates to adaptively decide the token merge ratios of different layers. As a generalizable module, Token Merger can be easily plugged into different layers of ViTs to boost their efficiency. Visualizations show that Token Merger progressively merges tokens and finally learns a compact set of tokens representing clear semantics. Compared with token pruning methods, Token Merger is more effective in preserving meaning contextual cues, thus performs and generalizes substantially better in different vision tasks. Extensive experiments and comparisons with other state-of-the-art downsampling methods also demonstrate its promising performance. For instance, it reduces 95% tokens and accelerates the inference speed by 62%. Meanwhile, the ImageNet classification accuracy only drops by 0.4%. The code will be available. Zhanzhou Feng, Shiliang Zhang |
IEEE Trans. Image Process. | 1 |