VLDB 2026 Research / reviewers in the wild / expert
Wencheng Zhu
dblp:184/0845
· DBLP profile ↗
15ranked-venue papers
10as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point Cloud Quantization Through Multimodal Prompting for 3D UnderstandingabstractVector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current prototype-based approaches relying on trainable vectors or clustered centroids fall short in representativeness and interpretability, even as multimodal alignment demonstrates its promise in vision-language models. To address these limitations, we propose a simple multimodal prompting-driven quantization framework for point cloud analysis. Our methodology is built upon two core insights: 1) Text embeddings from pre-trained models inherently encode visual semantics through many-to-one contrastive alignment, naturally serving as robust prototype priors; and 2) Multimodal prompts enable adaptive refinement of these prototypes, effectively mitigating vision-language semantic gaps. The framework introduces a dual-constrained quantization space, enforced by compactness and separation regularization, which seamlessly integrates visual and prototype features, resulting in hybrid representations that jointly encode geometric and semantic information. Furthermore, we employ Gumbel-Softmax relaxation to achieve differentiable discretization while maintaining quantization sparsity. Extensive experiments on the ModelNet40 and ScanObjectNN datasets clearly demonstrate the superior effectiveness of the proposed method. Wencheng Zhu, Xinzhong Zhu, Pengfei Zhu 0001 |
AAAI | 2 |
| 2026 | VTD-CLIP: Video-to-Text Discretization via Prompting CLIPabstractVision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing methods primarily rely on parameter-efficient fine-tuning of pre-trained image-text models, suffering from limited interpretability and poor generalization due to inadequate temporal modeling. To address these, we propose a simple yet effective video-to-text discretization framework. Our approach leverages the frozen text encoder to build a visual codebook derived from video class labels, exploiting the many-to-one contrastive alignment between visual and textual embeddings in multimodal pretraining. This enables the transformation of temporal visual features into discrete textual tokens via feature lookups, yielding interpretable video representations through explicit video modeling. Then, to improve robustness against noisy or irrelevant frames, we introduce a confidence-aware fusion module that dynamically weights keyframes based on their semantic relevance, as measured by the codebook. Furthermore, we incorporate learnable text prompts to conduct adaptive codebook updates during training. Experiments on four datasets, including HMDB-51, UCF-101, Something-Something-v2, and Kinetics-400, validate the superiority of our approach, achieving competitive improvements over state-of-the-art approaches. Wencheng Zhu, Pengfei Zhu 0001 |
AAAI | 1 |
| 2025 | CKD: Contrastive Knowledge Distillation From a Sample-Wise PerspectiveabstractIn this paper, we propose a simple yet effective contrastive knowledge distillation framework that achieves sample-wise logit alignment while preserving semantic consistency. Conventional knowledge distillation approaches exhibit over-reliance on feature similarity per sample, which risks overfitting, and contrastive approaches focus on inter-class discrimination at the expense of intra-sample semantic relationships. Our approach transfers "dark knowledge" through teacher-student contrastive alignment at the sample level. Specifically, our method first enforces intra-sample alignment by directly minimizing teacher-student logit discrepancies within individual samples. Then, we utilize inter-sample contrasts to preserve semantic dissimilarities across samples. By redefining positive pairs as aligned teacher-student logits from identical samples and negative pairs as cross-sample logit combinations, we reformulate these dual constraints into an InfoNCE loss framework, reducing computational complexity lower than sample squares while eliminating dependencies on temperature parameters and large batch sizes. We conduct comprehensive experiments across three benchmark datasets, including the CIFAR-100, ImageNet-1K, and MS COCO datasets, and experimental results clearly confirm the effectiveness of the proposed method on image classification, object detection, and instance segmentation tasks. Wencheng Zhu, Pengfei Zhu 0001, Yu Wang 0106, Qinghua Hu |
IEEE Trans. Image Process. | 1 |
| 2024 | PRM: A Pixel-Region-Matching Approach for Fast Video Object Segmentation
Wencheng Zhu, Danqing Song |
PRCV (4) | 4 |
| 2022 | Learning multiscale hierarchical attention for video summarization
Wencheng Zhu, Jiwen Lu, Yucheng Han, Jie Zhou 0001 |
Pattern Recognit. | 1 |
| 2022 | Separable Structure Modeling for Semi-Supervised Video Object SegmentationabstractIn this paper, we propose a separable structure modeling approach for semi-supervised video object segmentation. Unlike most existing methods which preclude the semantically structural information of target objects, our method not only captures pixel-level similarity relationships between the reference and target frames but also reveals the separable structure of the specified objects in target frames. Specifically, we first compute a pixel-wise similarity matrix by using representations of reference and target pixels and then select top rank reference pixels for target pixel classification. According to the prior knowledge from these top-rank reference pixels, we further appoint the representative target pixels for object structure modeling. Particularly, in the structure modeling branch, we extract the shared and individual features that can well represent the whole object and its components, respectively. Moreover, the proposed method is a fast algorithm without online fine-tuning and any post-processing. We conduct extensive experiments and ablation studies on the DAVIS-16, DAVIS-17, and YouTube-VOS datasets, and experimental results on three widely-used datasets demonstrate that our method achieves a superior performance, compared with state-of-the-art semi-supervised video object segmentation approaches in terms of speed and accuracy. Wencheng Zhu, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Relational Reasoning Over Spatial-Temporal Graphs for Video SummarizationabstractIn this paper, we propose a dynamic graph modeling approach to learn spatial-temporal representations for video summarization. Most existing video summarization methods extract image-level features with ImageNet pre-trained deep models. Differently, our method exploits object-level and relation-level information to capture spatial-temporal dependencies. Specifically, our method builds spatial graphs on the detected object proposals. Then, we construct a temporal graph by using the aggregated representations of spatial graphs. Afterward, we perform relational reasoning over spatial and temporal graphs with graph convolutional networks and extract spatial-temporal representations for importance score prediction and key shot selection. To eliminate relation clutters caused by densely connected nodes, we further design a self-attention edge pooling module, which disregards meaningless relations of graphs. We conduct extensive experiments on two popular benchmarks, including the SumMe and TVSum datasets. Experimental results demonstrate that the proposed method achieves superior performance against state-of-the-art video summarization methods. Wencheng Zhu, Yucheng Han, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | DSNet: A Flexible Detect-to-Summarize Network for Video SummarizationabstractIn this paper, we propose a Detect-to-Summarize network (DSNet) framework for supervised video summarization. Our DSNet contains anchor-based and anchor-free counterparts. The anchor-based method generates temporal interest proposals to determine and localize the representative contents of video sequences, while the anchor-free method eliminates the pre-defined temporal proposals and directly predicts the importance scores and segment locations. Different from existing supervised video summarization methods which formulate video summarization as a regression problem without temporal consistency and integrity constraints, our interest detection framework is the first attempt to leverage temporal consistency via the temporal interest detection formulation. Specifically, in the anchor-based approach, we first provide a dense sampling of temporal interest proposals with multi-scale intervals that accommodate interest variations in length, and then extract their long-range temporal features for interest proposal location regression and importance prediction. Notably, positive and negative segments are both assigned for the correctness and completeness information of the generated summaries. In the anchor-free approach, we alleviate drawbacks of temporal proposals by directly predicting importance scores of video frames and segment locations. Particularly, the interest detection framework can be flexibly plugged into off-the-shelf supervised video summarization methods. We evaluate the anchor-based and anchor-free approaches on the SumMe and TVSum datasets. Experimental results clearly validate the effectiveness of the anchor-based and anchor-free approaches. Wencheng Zhu, Jiwen Lu, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-label Quadruplet Dictionary Learning
Jiayu Zheng, Wencheng Zhu, Pengfei Zhu 0001 |
ICANN (2) | 2 |
| 2019 | Structured general and specific multi-view subspace clustering
Wencheng Zhu, Jiwen Lu, Jie Zhou 0001 |
Pattern Recognit. | 1 |
| 2018 | Nonlinear subspace clustering for image clustering
Wencheng Zhu, Jiwen Lu, Jie Zhou 0001 |
Pattern Recognit. Lett. | 1 |
| 2017 | Nonlinear subspace clusteringabstractThis paper presents a nonlinear subspace clustering (NSC) method for image clustering. Unlike most existing subspace clustering methods which only exploit the linear relationship of samples to learn the affine matrix, our NSC reveals the multi-cluster nonlinear structure of samples via a nonlinear neural network. While kernel-based clustering methods can also address the nonlinear issue of samples, this type of methods suffers from the scalability issue. Differently, our NSC employs a feed-forward neural network to map samples into a nonlinear space and performs subspace clustering at the top layer of the network, so that the mapping functions and the clustering issues are iteratively learned. Experimental results illustrate that our NSC outperforms the state-of-the-arts. Wencheng Zhu, Jiwen Lu, Jie Zhou 0001 |
ICIP | 1 |
| 2017 | Non-convex regularized self-representation for unsupervised feature selection
Pengfei Zhu 0001, Wencheng Zhu, Weizhi Wang, Wangmeng Zuo, Qinghua Hu |
Image Vis. Comput. | 2 |
| 2017 | Subspace clustering guided unsupervised feature selection
Pengfei Zhu 0001, Wencheng Zhu, Qinghua Hu, Changqing Zhang 0002, Wangmeng Zuo |
Pattern Recognit. | 2 |
| 2016 | Set to Set Visual Tracking
Wencheng Zhu, Pengfei Zhu 0001, Qinghua Hu, Changqing Zhang 0002 |
PRICAI | 1 |