Zhilin Zeng

dblp:355/4732 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 47% Segmentation and scene understanding · 43% Vision and language · 8%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
3.342025
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025
Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
2.532025
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025
Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
1.722025
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025
Machine learning › Efficient and distributed learning
inter-module communication
0.812024
Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Segmentation and scene understanding
object segmentation
0.812024
SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space Reconstruction · AAAI 2024
Visual content generation and editing › face editing
identity swapping
0.812024
SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Vision and language › vision-language model › vision-language model adaptation
CLIP fine-tuning
0.522025
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025
Computer vision › Vision and language › vision-language model
vision-language foundation model
0.522025
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space · CVPR 2025
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation · CVPR 2025
Machine learning › Efficient and distributed learning › dynamic neural network
dynamic convolution
0.212024
SimSwap++: Towards Faster and High-Quality Identity Swapping · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Transfer learning and domain adaptation
foundation model adaptation
0.212024
Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024
Computer vision › Segmentation and scene understanding › prompt-based segmentation
segment anything model
0.212024
Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model · CVPR 2024

Methods — techniques the papers use, named apart from their topics

parameter-efficient fine-tuning · 1.6dynamic convolution · 1.5scaling transformation · 0.9hyperspherical energy · 0.9hyperbolic space embedding · 0.9block-diagonal transformation · 0.9reparameterization · 0.8relation matrix · 0.8matrix decomposition · 0.8low-rank adaptation · 0.8knowledge distillation · 0.8hyper-complex layer · 0.8
YearPublicationVenuePosition
2025 Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation seeks to label each pixel in an image with arbitrary text descriptions. Vision-language foundation models, especially CLIP, have recently emerged as powerful tools for acquiring open-vocabulary capabilities. However, fine-tuning CLIP to equip it with pixel-level prediction ability often suffers three issues: 1) high computational cost, 2) misalignment between the two inherent modalities of CLIP, and 3) degraded generalization ability on unseen categories. To address these issues, we propose H-CLIP, a symmetrical parameter-efficient fine-tuning (PEFT) strategy conducted in hyperspherical space for both of the two CLIP modalities. Specifically, the PEFT strategy is achieved by a series of efficient block-diagonal learnable transformation matrices and a dual cross-relation communication module among all learnable matrices. Since the PEFT strategy is conducted symmetrically to the two CLIP modalities, the misalignment between them is mitigated. Furthermore, we apply an additional constraint to PEFT on the CLIP text encoder according to the hyperspherical energy principle, i.e., minimizing hyperspherical energy during fine-tuning preserves the intrinsic structure of the original parameter space, to prevent the destruction of the generalization ability offered by the CLIP text encoder. Extensive evaluations across various benchmarks show that H-CLIP achieves new SOTA open-vocabulary semantic segmentation results while only requiring updating approximately 4% of the total parameters of CLIP. The code is available at: https://github.com/SJTU-DeepVisionLab/H-CLIP.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Wei Shen 0002
CVPR3
2025 Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space
abstract
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing the text encoder preserves its powerful embeddings, recent studies show that fine-tuning both the text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchical alignment, since during fine-tuning, the hierarchy level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encoders hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP’s text embeddings decreases, facilitating better alignment with the pixel-level hierarchical structure of visual data. Building on this insight, we propose HyperCLIP, a novel fine-tuning strategy that adjusts the hyperbolic radius of the text embeddings through scaling transformations. By doing so, HyperCLIP equips CLIP with segmentation capability while introducing only a small number of learnable parameters. Our experiments demonstrate that HyperCLIP achieves state-of-the-art performance on open-vocabulary semantic segmentation tasks across three benchmarks, while fine-tuning only approximately 4% of the total parameters of CLIP. More importantly, we observe that after adjustment, CLIP’s text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the granularity required for this segmentation task might be quantified using the hyperbolic radius.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Changsong Wen, Menglin Yang 0001, Wei Shen 0002
CVPR3
2025 BHRAM: a knowledge graph embedding model based on bidirectional and heterogeneous relational attention mechanism
Wanqiu Li, Yuanbin Mo, Weidong Tang, Zhilin Zeng
Appl. Intell.6
2024 SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space Reconstruction
abstract
Segment Anything Model (SAM) has received remarkable attention as it offers a powerful and versatile solution for object segmentation in images. However, fine-tuning SAM for downstream segmentation tasks under different scenarios remains a challenge, as the varied characteristics of different scenarios naturally requires diverse model parameter spaces. Most existing fine-tuning methods attempt to bridge the gaps among different scenarios by introducing a set of new parameters to modify SAM's original parameter space. Unlike these works, in this paper, we propose fine-tuning SAM efficiently by parameter space reconstruction (SAM-PARSER), which introduce nearly zero trainable parameters during fine-tuning. In SAM-PARSER, we assume that SAM's original parameter space is relatively complete, so that its bases are able to reconstruct the parameter space of a new scenario. We obtain the bases by matrix decomposition, and fine-tuning the coefficients to reconstruct the parameter space tailored to the new scenario by an optimal linear combination of the bases. Experimental results show that SAM-PARSER exhibits superior segmentation performance across various scenarios, while reducing the number of trainable parameters by approximately 290 times compared with current parameter-efficient fine-tuning methods.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Xiaokang Yang 0001, Wei Shen 0002
AAAI3
2024 DeCo-Net: Robust Multimodal Brain Tumor Segmentation via Decoupled Complementary Knowledge Distillation
abstract
Automated brain tumor segmentation with multimodal magnetic resonance imaging (MRI) plays a pivotal rule in clinical application. However, most existing algorithms require complete image modalities as input, which is often impractical to obtain for every patient in real clinical practice. Therefore, a robust multimodal algorithm that is capable of handling various modality-incomplete data is highly desirable. In this paper, we propose DeCo-Net, a Decoupled Complementary knowledge distillation framework for multimodal brain tumor segmentation with incomplete modalities. Specifically, our approach decouples the feature learning of the modality-incomplete data into two branches: one dedicated to extracting the inherent features from the available modalities and the other focused on inferring the complementary missing modal information. We employ a teacher-student co-training framework where the teacher network is collaboratively trained to dynamically transfer the complementary knowledge to the student model based on the specific type of modality-incomplete data fed to student. To this end, we propose a modality-aware contrastive distillation strategy that guides the student model to distill a discriminative and complementary knowledge representation that acts as supplements to the original modality-incomplete representation. Extensive evaluations on the BraTS2018, BraTS2020 and BraTS2023 datasets demonstrate that our method achieves state-of-the-art performance in multimodal brain tumor segmentation with incomplete modalities.
Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002
BIBM1
2024 Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model
abstract
Parameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community, PEFT has shown effectiveness in image classification, but little research has studied its ability for image segmentation. Fine-tuning segmentation models usually requires a heavier adjustment of parameters to align the proper projection directions in the parameter space for new scenarios. This raises a challenge to existing PEFT algorithms, as they often inject a limited number of individual parameters into each block, which prevents substantial adjustment of the projection direction of the parameter space due to the limitation of Hidden Markov Chain along blocks. In this paper, we equip PEFT with a cross-block orchestration mechanism to enable the adaptation of the Segment Anything Model (SAM) to various downstream scenarios. We introduce a novel inter-block communication module, which integrates a learnable relation matrix to facilitate communication among different coefficient sets of each PEFT block's parameter space. Moreover, we propose an intra-block enhancement module, which introduces a linear projection head whose weights are generated from a hyper-complex layer, further enhancing the impact of the adjustment of projection directions on the entire parameter space. Extensive experiments on diverse benchmarks demonstrate that our proposed approach consistently improves the segmentation performance significantly on novel scenarios with only around 1K additional parameters.
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Lingxi Xie, Qi Tian 0001, Wei Shen 0002
CVPR3
2024 SLSM: An Efficient Strategy for Lazy Schema Migration on Shared-Nothing Databases
Zhilin Zeng, Hui Li 0005, Xiyue Gao, Hui Zhang 0129, Huiquan Zhang, Jiangtao Cui
DASFAA (1)1
2024 Missing as Masking: Arbitrary Cross-Modal Feature Reconstruction for Incomplete Multimodal Brain Tumor Segmentation
Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002
MICCAI (8)1
2024 Seamless group handover authentication protocol for vehicle networks: Services continuity
Ye Bi, Kai Fan 0001, Zhilin Zeng, Kan Yang 0001, Hui Li 0006, Yintang Yang
Comput. Networks3
2024 SimSwap++: Towards Faster and High-Quality Identity Swapping
abstract
Face identity editing (FIE) shows great value in AI content creation. Low-resolution FIE approaches have achieved tremendous progress, but high-quality FIE struggles. Two major challenges hinder higher-resolution and higher-performance development of FIE: lack of high-resolution dataset and unacceptable complexity forbidding for mobile platforms. To address both issues, we establish a novel large-scale, high-quality dataset tailored for FIE. Based on our SimSwap (Chen et al. 2020), we propose an upgraded version named SimSwap++ with significantly boosted model efficiency. SimSwap++ features two major innovations for high-performance model compression. First, a novel computational primitive named Conditional Dynamic Convolution (CD-Conv) is proposed to address the inefficiency of conditional schemes (e.g., AdaIN) in tiny models. CD-Conv achieves anisotropic processing and injection with significantly lower complexity compared to standard conditional operators, e.g., modulated convolution. Second, a Morphable Knowledge Distillation (MKD) is presented to further trim the overall model. Unlike conventional homogeneous teacher-student structures, MKD is designed to be heterogeneous and mutually compensable, endowing the student with the multi-path morphable property; thus, our student maximally inherits the teacher's knowledge after distillation while further reducing its complexity through structure re-parameterization. Extensive experiments demonstrate that our SimSwap++ achieves state-of-the-art performance (97.55% ID accuracy on FaceForensics++) with extremely low complexity (2.5 GFLOPs).
Xuanhong Chen, Bingbing Ni, Yutian Liu 0004, Naiyuan Liu, Zhilin Zeng
IEEE Trans. Pattern Anal. Mach. Intell.5