VLDB 2026 Research / reviewers in the wild / expert
Zhilin Zeng
dblp:355/4732
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 47% Segmentation and scene understanding · 43% Vision and language · 8% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
3.3 | 4 | 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025 Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
2.5 | 3 | 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025 Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024 |
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation |
1.7 | 2 | 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025 |
Machine learning › Efficient and distributed learning
inter-module communication |
0.8 | 1 | 2024 | Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.8 | 1 | 2024 | SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space Reconstruction · AAAI 2024 |
Visual content generation and editing › face editing
identity swapping |
0.8 | 1 | 2024 | SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Vision and language › vision-language model › vision-language model adaptation
CLIP fine-tuning |
0.5 | 2 | 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025 |
Computer vision › Vision and language › vision-language model
vision-language foundation model |
0.5 | 2 | 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025 |
Machine learning › Efficient and distributed learning › dynamic neural network
dynamic convolution |
0.2 | 1 | 2024 | SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Transfer learning and domain adaptation
foundation model adaptation |
0.2 | 1 | 2024 | Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024 |
Computer vision › Segmentation and scene understanding › prompt-based segmentation
segment anything model |
0.2 | 1 | 2024 | Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
parameter-efficient fine-tuning · 1.6dynamic convolution · 1.5scaling transformation · 0.9hyperspherical energy · 0.9hyperbolic space embedding · 0.9block-diagonal transformation · 0.9reparameterization · 0.8relation matrix · 0.8matrix decomposition · 0.8low-rank adaptation · 0.8knowledge distillation · 0.8hyper-complex layer · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation seeks to label each pixel in an image with arbitrary text descriptions. Vision-language foundation models, especially CLIP, have recently emerged as powerful tools for acquiring open-vocabulary capabilities. However, fine-tuning CLIP to equip it with pixel-level prediction ability often suffers three issues: 1) high computational cost, 2) misalignment between the two inherent modalities of CLIP, and 3) degraded generalization ability on unseen categories. To address these issues, we propose H-CLIP, a symmetrical parameter-efficient fine-tuning (PEFT) strategy conducted in hyperspherical space for both of the two CLIP modalities. Specifically, the PEFT strategy is achieved by a series of efficient block-diagonal learnable transformation matrices and a dual cross-relation communication module among all learnable matrices. Since the PEFT strategy is conducted symmetrically to the two CLIP modalities, the misalignment between them is mitigated. Furthermore, we apply an additional constraint to PEFT on the CLIP text encoder according to the hyperspherical energy principle, i.e., minimizing hyperspherical energy during fine-tuning preserves the intrinsic structure of the original parameter space, to prevent the destruction of the generalization ability offered by the CLIP text encoder. Extensive evaluations across various benchmarks show that H-CLIP achieves new SOTA open-vocabulary semantic segmentation results while only requiring updating approximately 4% of the total parameters of CLIP. The code is available at: https://github.com/SJTU-DeepVisionLab/H-CLIP. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Wei Shen 0002 |
CVPR | 3 |
| 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic SpaceabstractCLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing the text encoder preserves its powerful embeddings, recent studies show that fine-tuning both the text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchical alignment, since during fine-tuning, the hierarchy level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encoders hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP’s text embeddings decreases, facilitating better alignment with the pixel-level hierarchical structure of visual data. Building on this insight, we propose HyperCLIP, a novel fine-tuning strategy that adjusts the hyperbolic radius of the text embeddings through scaling transformations. By doing so, HyperCLIP equips CLIP with segmentation capability while introducing only a small number of learnable parameters. Our experiments demonstrate that HyperCLIP achieves state-of-the-art performance on open-vocabulary semantic segmentation tasks across three benchmarks, while fine-tuning only approximately 4% of the total parameters of CLIP. More importantly, we observe that after adjustment, CLIP’s text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the granularity required for this segmentation task might be quantified using the hyperbolic radius. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Changsong Wen, Menglin Yang 0001, Wei Shen 0002 |
CVPR | 3 |
| 2025 | BHRAM: a knowledge graph embedding model based on bidirectional and heterogeneous relational attention mechanism
Wanqiu Li, Yuanbin Mo, Weidong Tang, Zhilin Zeng |
Appl. Intell. | 6 |
| 2024 | SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space ReconstructionabstractSegment Anything Model (SAM) has received remarkable attention as it offers a powerful and versatile solution for object segmentation in images. However, fine-tuning SAM for downstream segmentation tasks under different scenarios remains a challenge, as the varied characteristics of different scenarios naturally requires diverse model parameter spaces. Most existing fine-tuning methods attempt to bridge the gaps among different scenarios by introducing a set of new parameters to modify SAM's original parameter space. Unlike these works, in this paper, we propose fine-tuning SAM efficiently by parameter space reconstruction (SAM-PARSER), which introduce nearly zero trainable parameters during fine-tuning. In SAM-PARSER, we assume that SAM's original parameter space is relatively complete, so that its bases are able to reconstruct the parameter space of a new scenario. We obtain the bases by matrix decomposition, and fine-tuning the coefficients to reconstruct the parameter space tailored to the new scenario by an optimal linear combination of the bases. Experimental results show that SAM-PARSER exhibits superior segmentation performance across various scenarios, while reducing the number of trainable parameters by approximately 290 times compared with current parameter-efficient fine-tuning methods. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Xiaokang Yang 0001, Wei Shen 0002 |
AAAI | 3 |
| 2024 | DeCo-Net: Robust Multimodal Brain Tumor Segmentation via Decoupled Complementary Knowledge DistillationabstractAutomated brain tumor segmentation with multimodal magnetic resonance imaging (MRI) plays a pivotal rule in clinical application. However, most existing algorithms require complete image modalities as input, which is often impractical to obtain for every patient in real clinical practice. Therefore, a robust multimodal algorithm that is capable of handling various modality-incomplete data is highly desirable. In this paper, we propose DeCo-Net, a Decoupled Complementary knowledge distillation framework for multimodal brain tumor segmentation with incomplete modalities. Specifically, our approach decouples the feature learning of the modality-incomplete data into two branches: one dedicated to extracting the inherent features from the available modalities and the other focused on inferring the complementary missing modal information. We employ a teacher-student co-training framework where the teacher network is collaboratively trained to dynamically transfer the complementary knowledge to the student model based on the specific type of modality-incomplete data fed to student. To this end, we propose a modality-aware contrastive distillation strategy that guides the student model to distill a discriminative and complementary knowledge representation that acts as supplements to the original modality-incomplete representation. Extensive evaluations on the BraTS2018, BraTS2020 and BraTS2023 datasets demonstrate that our method achieves state-of-the-art performance in multimodal brain tumor segmentation with incomplete modalities. Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002 |
BIBM | 1 |
| 2024 | Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything ModelabstractParameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community, PEFT has shown effectiveness in image classification, but little research has studied its ability for image segmentation. Fine-tuning segmentation models usually requires a heavier adjustment of parameters to align the proper projection directions in the parameter space for new scenarios. This raises a challenge to existing PEFT algorithms, as they often inject a limited number of individual parameters into each block, which prevents substantial adjustment of the projection direction of the parameter space due to the limitation of Hidden Markov Chain along blocks. In this paper, we equip PEFT with a cross-block orchestration mechanism to enable the adaptation of the Segment Anything Model (SAM) to various downstream scenarios. We introduce a novel inter-block communication module, which integrates a learnable relation matrix to facilitate communication among different coefficient sets of each PEFT block's parameter space. Moreover, we propose an intra-block enhancement module, which introduces a linear projection head whose weights are generated from a hyper-complex layer, further enhancing the impact of the adjustment of projection directions on the entire parameter space. Extensive experiments on diverse benchmarks demonstrate that our proposed approach consistently improves the segmentation performance significantly on novel scenarios with only around 1K additional parameters. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Lingxi Xie, Qi Tian 0001, Wei Shen 0002 |
CVPR | 3 |
| 2024 | SLSM: An Efficient Strategy for Lazy Schema Migration on Shared-Nothing Databases
Zhilin Zeng, Hui Li 0005, Xiyue Gao, Hui Zhang 0129, Huiquan Zhang, Jiangtao Cui |
DASFAA (1) | 1 |
| 2024 | Missing as Masking: Arbitrary Cross-Modal Feature Reconstruction for Incomplete Multimodal Brain Tumor Segmentation
Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002 |
MICCAI (8) | 1 |
| 2024 | Seamless group handover authentication protocol for vehicle networks: Services continuity
Ye Bi, Kai Fan 0001, Zhilin Zeng, Kan Yang 0001, Hui Li 0006, Yintang Yang |
Comput. Networks | 3 |
| 2024 | SimSwap++: Towards Faster and High-Quality Identity SwappingabstractFace identity editing (FIE) shows great value in AI content creation. Low-resolution FIE approaches have achieved tremendous progress, but high-quality FIE struggles. Two major challenges hinder higher-resolution and higher-performance development of FIE: lack of high-resolution dataset and unacceptable complexity forbidding for mobile platforms. To address both issues, we establish a novel large-scale, high-quality dataset tailored for FIE. Based on our SimSwap (Chen et al. 2020), we propose an upgraded version named SimSwap++ with significantly boosted model efficiency. SimSwap++ features two major innovations for high-performance model compression. First, a novel computational primitive named Conditional Dynamic Convolution (CD-Conv) is proposed to address the inefficiency of conditional schemes (e.g., AdaIN) in tiny models. CD-Conv achieves anisotropic processing and injection with significantly lower complexity compared to standard conditional operators, e.g., modulated convolution. Second, a Morphable Knowledge Distillation (MKD) is presented to further trim the overall model. Unlike conventional homogeneous teacher-student structures, MKD is designed to be heterogeneous and mutually compensable, endowing the student with the multi-path morphable property; thus, our student maximally inherits the teacher's knowledge after distillation while further reducing its complexity through structure re-parameterization. Extensive experiments demonstrate that our SimSwap++ achieves state-of-the-art performance (97.55% ID accuracy on FaceForensics++) with extremely low complexity (2.5 GFLOPs). Xuanhong Chen, Bingbing Ni, Yutian Liu 0004, Naiyuan Liu, Zhilin Zeng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |