VLDB 2026 Research / reviewers in the wild / expert
Zelin Peng
dblp:256/7867
· DBLP profile ↗
29ranked-venue papers
10as first author
28since 2021 · last 2026
0009-0002-4066-7929ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 18 · 8 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CP-CLIP: Customized Parameter Generation for Open-vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation aims to assign pixel-level labels to images based on textual descriptions, even for categories beyond predefined closed sets. While vision-language foundation models like CLIP are widely used for this task, fine-tuning them for pixel-level predictions often compromises their generalization capabilities. To address this, we propose a novel fine-tuning strategy, CP-CLIP, which generates customized parameters for CLIP without sacrificing its generalization. Our method employs a customized parameter generator that produces newly added parameters based on random noise, using local visual features from CLIP's image encoder as conditions, enabling generalization to new images from unseen scenarios. Additionally, we introduce an orthogonal adaptation technique to ensure the update direction is orthogonal to the pre-trained weights, largely preserving the initial generalization ability. Extensive experiments demonstrate that CP-CLIP achieves state-of-the-art performance across multiple benchmarks in open-vocabulary semantic segmentation. Zelin Peng, Zhengqin Xu, Wei Shen 0002 |
AAAI | 1 |
| 2026 | Efficient Segmentation with Multimodal Large Language Model via Token RoutingabstractRecent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in addressing open-world segmentation tasks. However, the substantial computational cost of the LLM components presents a significant challenge, especially in segmentation tasks, where efficiency has long been a central concern. Existing efficient MLLM approaches typically reduce computation cost by pruning visual tokens in the early layers, as they account for the majority of the input sequence. Despite their efficiency, this is incompatible with dense prediction tasks such as segmentation, since removing visual tokens leads to the loss of essential object parts and spatial details. To better understand the roles of visual tokens in segmentation, we analyze the attention weights of both image and mask tokens within LLM. We find that image tokens are important throughout all layers, whereas mask tokens only attend to image tokens at deeper layers. Based on the observation, we build an efficient segmentation framework based on MLLMs by introducing a sophisticated token routing strategy. This strategy dynamically determines when and how different tokens participate in computation: For mask tokens, they are only inserted at deeper layers of the LLM to reduce redundant computation, since they rarely attend to image tokens in early layers; For image tokens, only a small number of them, named proxies, are updated via full feedforward network (FFN) computation, while the update of the remaining tokens is guided by these proxies, i.e., efficiently computed through a lightweight projector applied on the difference of the proxies during their update. Our method achieves a 1.5× acceleration over the original LLM process by reducing its FLOPs to 56%, while maintaining the same segmentation performance. Changsong Wen, Zelin Peng, Wei Shen 0002 |
AAAI | 2 |
| 2026 | Bridging Visual Dynamics and Narrative Reasoning: Multimodal Large Language Models for Short Drama Quality Assessment
Qingyang Liu 0008, Jiangtong Li, Zelin Peng, Shaobo Wang 0001, Zhaohe Liao, Shuochen Chang, Bingjie Gao, Mu Liu, Jidong Jiang, Li Niu 0002 |
WWW | 3 |
| 2026 | DMformer: Difficulty-Adapted Masked Transformer for Semi-Supervised Medical Image SegmentationabstractThe shared anatomy among different human bodies can serve as a strong prior for effectively leveraging unlabeled data in semi-supervised medical image segmentation. Inspired by the success of masked image modeling, we notice that this prior can be explicitly realized by incorporating an auxiliary unsupervised gross anatomy reconstruction task into a teacher-student semi-supervised segmentation framework. In this auxiliary task, consistency is maintained between the student's predictions on masked images and the teacher's predictions on the original images. Despite its potential, we observe that the reconstruction difficulties of different organs/tissues can vary significantly and therefore reconstructing them requires tailored learning strategies. To address this issue, we introduce a difficulty-adapted mask mechanism based on the teacher-student framework, wherein the reconstruction difficulty is adapted to facilitate training. Specifically, we control the reconstruction difficulty by modulating two important factors: masked region ratio and masked class ratio. Accordingly, we design two corresponding mask strategies. 1) Region-based masking: randomly masks a fraction of each class according to an automatically computed mask ratio. 2) Class-based masking: masks the entire regions of the specific classes according to the class confidence predicted by the teacher model. During training, a conflict-aware gradient computation strategy is introduced to mitigate potential optimization conflicts arising from modulating the two reconstruction factors simultaneously. By building on vision transformers, we develop an Difficulty-adapted Masked Transformer (DMformer) for semi-supervised medical image segmentation. Extensive experiments demonstrate the superiority of DMformer, which outperforms the previous SOTA by 9.53% and 4.63% in terms of DSC on ACDC dataset with 5% labeled images and Synapse dataset with 30% labeled images, respectively. Zelin Peng, Guanchun Wang, Zhengqin Xu, Xiaokang Yang 0001, Wei Shen 0002 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | FATE: Feature-Adapted Parameter Tuning for Vision-Language ModelsabstractFollowing the recent popularity of vision language models, several attempts, e.g., parameter-efficient fine-tuning (PEFT), have been made to extend them to different downstream tasks. Previous PEFT works motivate their methods from the view of introducing new parameters for adaptation but still need to learn this part of weight from scratch, i.e., random initialization. In this paper, we present a novel strategy that incorporates the potential of prompts, e.g., vision features, to facilitate the initial parameter space adapting to new scenarios. We introduce a Feature-Adapted parameTer Efficient tuning paradigm for vision-language models, dubbed as FATE, which injects informative features from the vision encoder into language encoder's parameters space. Specifically, we extract vision features from the last layer of CLIP's vision encoder and, after projection, treat them as parameters for fine-tuning each layer of CLIP's language encoder. By adjusting these feature-adapted parameters, we can directly enable communication between the vision and language branches, facilitating CLIP's adaptation to different scenarios. Experimental results show that FATE exhibits superior generalization performance on 11 datasets with a very small amount of extra parameters and computation. Zhengqin Xu, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002 |
AAAI | 2 |
| 2025 | PG-SAM: A Fine-Grained Prior-Guided SAM Framework for Prompt-Free Medical Image SegmentationabstractSegment Anything Model (SAM) demonstrates powerful zero-shot capabilities; however, its accuracy and robustness significantly decrease when applied to medical image segmentation. Existing methods address this issue through modality fusion, integrating textual and image information to provide more detailed priors. In this study, we argue that the granularity of text and the domain gap affect the accuracy of the priors. Furthermore, the discrepancy between high-level abstract semantics and pixel-level boundary details in images can introduce noise into the fusion process. To address this, we propose Prior-Guided SAM (PG-SAM), which employs a fine-grained modality prior aligner to leverage specialized medical knowledge for better modality alignment. The core of our method lies in efficiently addressing the domain gap with fine-grained text from a medical large language model (LLM). Meanwhile, it also enhances the priors' quality after modality alignment, ensuring more accurate segmentation. In addition, our decoder enhances the model's expressive capabilities through multi-level feature fusion and iterative mask optimizer operations, supporting unprompted learning. We also propose a unified pipeline that effectively supplies high-quality semantic information to SAM. Extensive experiments on the datasets demonstrate that the proposed PG-SAM achieves state-of-the-art performance. Our anonymous code is released at https://github.com/logan-0623/PG-SAM. Yiheng Zhong, Zihong Luo, Yingzhen Hu, Zelin Peng, Jionglong Su, ZongYuan Ge, Muhammad Imran Razzak |
BIBM | 6 |
| 2025 | Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher DistillationabstractContrastive language–image pretraining models such as CLIP have demonstrated remarkable performance in various text-image alignment tasks. However, the inherent 77-token input limitation and reliance on predominantly short-text training data restrict its ability to handle long-text tasks effectively. To overcome these constraints, we propose LongD-CLIP, a dual-teacher distillation framework designed to enhance long-text representation while mitigating knowledge forgetting. In our approach, a teacher model, fine-tuned on long-text data, distills rich representation knowledge into a student model, while the original CLIP serves as a secondary teacher to help the student retain its foundational knowledge. Extensive experiments reveal that LongD-CLIP significantly outperforms existing models across long-text retrieval, short-text retrieval, and zero-shot image classification tasks. For instance, in the image-to-text retrieval task on the ShareGPT4V test set, LongD-CLIP exceeds Long-CLIP’s performance by 2.5%, achieving an accuracy of 98.3%. Similarly, on the Urban1k dataset, it records a 9.2% improvement, reaching 91.9%, thereby underscoring its robust generalization capabilities. Additionally, the text encoder of LongD-CLIP exhibits reduced latent space drift and improved compatibility with existing generative models, effectively overcoming the 77token input constraint. Yuheng Feng, Changsong Wen, Zelin Peng, Li jiaye |
CVPR | 3 |
| 2025 | Star with Bilinear MappingabstractContextual modeling is crucial for robust visual representation learning, especially in computer vision. Although Transformers have become a leading architecture for vision tasks due to their attention mechanism, the quadratic complexity of full attention operations presents substantial computational challenges. To address this, we introduce Star with Bilinear Mapping (SBM), a Transformer-Like architecture that achieves global contextual modeling with linear complexity. SBM employs a bilinear mapping module (BM) with low-rank decomposition strategy and star operations (element-wise multiplication) to efficiently capture global contextual information. Our model demonstrates competitive performance on image classification and semantic segmentation tasks, delivering significant computational efficiency gains compared to traditional attention-based models. Code is available at https://github.com/SJTU-DeepVisionLab/SBM. Zelin Peng, Zhengqin Xu, Xiaokang Yang 0001, Wei Shen 0002 |
CVPR | 1 |
| 2025 | Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation seeks to label each pixel in an image with arbitrary text descriptions. Vision-language foundation models, especially CLIP, have recently emerged as powerful tools for acquiring open-vocabulary capabilities. However, fine-tuning CLIP to equip it with pixel-level prediction ability often suffers three issues: 1) high computational cost, 2) misalignment between the two inherent modalities of CLIP, and 3) degraded generalization ability on unseen categories. To address these issues, we propose H-CLIP, a symmetrical parameter-efficient fine-tuning (PEFT) strategy conducted in hyperspherical space for both of the two CLIP modalities. Specifically, the PEFT strategy is achieved by a series of efficient block-diagonal learnable transformation matrices and a dual cross-relation communication module among all learnable matrices. Since the PEFT strategy is conducted symmetrically to the two CLIP modalities, the misalignment between them is mitigated. Furthermore, we apply an additional constraint to PEFT on the CLIP text encoder according to the hyperspherical energy principle, i.e., minimizing hyperspherical energy during fine-tuning preserves the intrinsic structure of the original parameter space, to prevent the destruction of the generalization ability offered by the CLIP text encoder. Extensive evaluations across various benchmarks show that H-CLIP achieves new SOTA open-vocabulary semantic segmentation results while only requiring updating approximately 4% of the total parameters of CLIP. The code is available at: https://github.com/SJTU-DeepVisionLab/H-CLIP. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Wei Shen 0002 |
CVPR | 1 |
| 2025 | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic SpaceabstractCLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing the text encoder preserves its powerful embeddings, recent studies show that fine-tuning both the text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchical alignment, since during fine-tuning, the hierarchy level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encoders hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP’s text embeddings decreases, facilitating better alignment with the pixel-level hierarchical structure of visual data. Building on this insight, we propose HyperCLIP, a novel fine-tuning strategy that adjusts the hyperbolic radius of the text embeddings through scaling transformations. By doing so, HyperCLIP equips CLIP with segmentation capability while introducing only a small number of learnable parameters. Our experiments demonstrate that HyperCLIP achieves state-of-the-art performance on open-vocabulary semantic segmentation tasks across three benchmarks, while fine-tuning only approximately 4% of the total parameters of CLIP. More importantly, we observe that after adjustment, CLIP’s text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the granularity required for this segmentation task might be quantified using the hyperbolic radius. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Changsong Wen, Menglin Yang 0001, Wei Shen 0002 |
CVPR | 1 |
| 2025 | Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingabstractRecent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness. Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge |
CVPR | 8 |
| 2025 | Domain Generalization in CLIP via Learning with Diverse Text PromptsabstractDomain generalization (DG) aims to train a model on source domains that can generalize well to unseen domains. Recent advances in Vision-Language Models (VLMs), such as CLIP, exhibit remarkable generalization capabilities across a wide range of data distributions, benefiting tasks like DG. However, CLIP is pre-trained by aligning images with their descriptions, which inevitably captures domain-specific details. Moreover, adapting CLIP to source domains with limited feature diversity introduces bias. These limitations hinder the model’s ability to generalize across domains. In this paper, we propose a new DG approach by learning with diverse text prompts. These text prompts incorporate varied contexts to imitate different domains, enabling DG model to learn domain-invariant features. The text prompts guide DG model learning in three aspects: feature suppression, which uses these prompts to identify domain-sensitive features and suppress them; feature consistency, which ensures the model’s features are robust to domain variations imitated by the diverse prompts; and feature diversification, which diversifies features based on the prompts to mitigate bias. Experimental results show that our approach improves domain generalization performance on five datasets from the DomainBed benchmark, achieving state-of-the-art results. Changsong Wen, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002 |
CVPR | 2 |
| 2025 | OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingabstractSurgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance. Kun Yuan 0004, Yaling Shen, Xiaohao Xu, Wei Li 0320, Zhongxing Xu, Zelin Peng, Siyuan Yan, Vinkle Srivastav, Diping Song, Tianbin Li, Danli Shi, Jin Ye 0002, Nicolas Padoy, Nassir Navab, Junjun He, ZongYuan Ge |
ICCV | 10 |
| 2025 | ACMamba: Fast Unsupervised Anomaly Detection via An Asymmetrical Consensus State Space ModelabstractUnsupervised anomaly detection in hyperspectral images (HSI), aiming to detect unknown targets from backgrounds, is challenging for earth surface monitoring. However, current studies are hindered by steep computational costs due to the high-dimensional property of HSI and dense sampling-based training paradigm, constraining their rapid deployment. Our key observation is that, during training, not all samples within the same homogeneous area are indispensable, whereas ingenious sampling can provide a powerful substitute for reducing costs. Motivated by this, we propose an Asymmetrical Consensus State Space Model (ACMamba) to significantly reduce computational costs without compromising accuracy. Specifically, we design an asymmetrical anomaly detection paradigm that utilizes region-level instances as an efficient alternative to dense pixel-level samples. In this paradigm, a low-cost Mamba-based module is introduced to discover global contextual attributes of regions that are essential for HSI reconstruction. Additionally, we develop a consensus learning strategy from the optimization perspective to simultaneously facilitate background reconstruction and anomaly compression, further alleviating the negative impact of anomaly reconstruction. Theoretical analysis and extensive experiments across eight benchmarks verify the superiority of ACMamba, demonstrating a faster speed and stronger performance over the state-of-the-art. Code is released at https://github.com/PURE-melo/ACMamba. Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Licheng Jiao |
ACM Multimedia | 4 |
| 2025 | HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language ModelsabstractMulti-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computational resources (e.g., thousands of GPUs) for training to achieve cross-modal alignment at multi-granularity levels. We argue that a key source of this inefficiency lies in the vision encoders they widely equip with, e.g., CLIP and SAM, which lack the alignment with language at multi-granularity levels. To address this issue, in this paper, we leverage hyperbolic space, which inherently models hierarchical levels and thus provides a principled framework for bridging the granularity gap between visual and textual modalities at an arbitrary granularity level. Concretely, we propose an efficient training paradigm for MLLMs, dubbed as \blg, which can optimize visual representations to align with their textual counterparts at an arbitrary granularity level through dynamic hyperbolic radius adjustment in hyperbolic space. \alg employs learnable matrices with M\"{o}bius multiplication operations, implemented via three effective configurations: diagonal scaling matrices, block-diagonal matrices, and banded matrices, providing a flexible yet efficient parametrization strategy. Comprehensive experiments across multiple MLLM benchmarks demonstrate that \alg consistently improves both existing pre-training and fine-tuning MLLMs clearly with less than 1\% additional parameters. Code is available at \url{https://github.com/godlin-sjtu/HyperET}. Zelin Peng, Zhengqin Xu, Xiaokang Yang 0001, Wei Shen 0002 |
NeurIPS | 1 |
| 2025 | OraL: An Observational Learning Paradigm for Unsupervised Hyperspectral Change DetectionabstractUnsupervised hyperspectral change detection (UHCD), detecting subtle changes between bi-temporal images without manual annotations, is an essential but challenging task in the earth observation community. The current modus operandi often performs it in a feature comparison manner, which is limited by variations in imaging conditions. We observe that fully supervised paradigms using limited annotations are capable of overcoming this challenge. Based on this, we introduce a novel Observational Learning Paradigm (OraL) for UHCD by mimicking fully supervised paradigms. OraL comprises two sequential stages: Observation, which designs a spatial-temporal observation strategy (STO) that records the learning consistency of pixels under different training steps and views, to obtain reliable pseudo-labels. Reproduction, which retrains the model with these pseudo-labels and introduces a distribution-aware spectral learning strategy (DSL) to adaptively increase their learning difficulty according to spectral distributions, enhancing the robustness and generalization of the model. Extensive experiments on several public hyperspectral image datasets demonstrate its state-of-the-art performance and pluggability for previous unsupervised methods. Code will be made available. Guanchun Wang, Xiangrong Zhang, Zelin Peng, Shunli Tian, Tianyang Zhang 0002, Xu Tang 0004, Licheng Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | S2Mamba: A Spatial-Spectral State Space Model for Hyperspectral Image ClassificationabstractThe land cover analysis using hyperspectral images (HSIs) remains an open problem due to their low spatial resolution and complex spectral information. Recent studies are primarily dedicated to designing Transformer-based architectures for spatial-spectral long-range dependencies modeling, which is computationally expensive with quadratic complexity. Selective structured state space model (SSM; Mamba), which is efficient for modeling long-range dependencies with linear complexity, has recently shown promising progress. However, its potential in HSI processing that requires handling numerous spectral bands has not yet been explored. In this article, we innovatively propose S2Mamba, a spatial-spectral SSM for HSI classification, to excavate spatial-spectral contextual features, resulting in more efficient and accurate land cover analysis. In S2Mamba, two selective structured SSMs through different dimensions are designed for feature extraction, one for spatial, and the other for spectral, along with a spatial-spectral mixture gate (SMG) for optimal fusion. More specifically, S2Mamba first captures spatial contextual relations by interacting each pixel with its adjacent through a patch cross scanning (PCS) module and then explores semantic information from continuous spectral bands through a bidirectional spectral scanning (BSS) module. Considering the distinct expertise of the two attributes in homogenous and complicated texture scenes, we realize the SMG by a group of learnable matrices, allowing for the adaptive incorporation of representations learned across different dimensions. Extensive experiments conducted on HSI classification benchmarks demonstrate the superiority and prospect of S2Mamba. The code will be made available at:https://github.com/PURE-melo/S2Mamba. Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Negative Deterministic Information-Based Multiple Instance Learning for Weakly Supervised Object Detection and SegmentationabstractWeakly supervised object detection (WSOD) and semantic segmentation with image-level annotations have attracted extensive attention due to their high label efficiency. Multiple instance learning (MIL) offers a feasible solution for the two tasks by treating each image as a bag with a series of instances (object regions or pixels) and identifying foreground instances that contribute to bag classification. However, conventional MIL paradigms often suffer from issues, e.g., discriminative instance domination and missing instances. In this article, we observe that negative instances usually contain valuable deterministic information, which is the key to solving the two issues. Motivated by this, we propose a novel MIL paradigm based on negative deterministic information (NDI), termed NDI-MIL, which is based on two core designs with a progressive relation: NDI collection and negative contrastive learning (NCL). In NDI collection, we identify and distill NDI from negative instances online by a dynamic feature bank. The collected NDI is then utilized in a NCL mechanism to locate and punish those discriminative regions, by which the discriminative instance domination and missing instances issues are effectively addressed, leading to improved object- and pixel-level localization accuracy and completeness. In addition, we design an NDI-guided instance selection (NGIS) strategy to further enhance the systematic performance. Experimental results on several public benchmarks, including PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO, show that our method achieves satisfactory performance. The code is available at: https://github.com/GC-WSL/NDI. Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space ReconstructionabstractSegment Anything Model (SAM) has received remarkable attention as it offers a powerful and versatile solution for object segmentation in images. However, fine-tuning SAM for downstream segmentation tasks under different scenarios remains a challenge, as the varied characteristics of different scenarios naturally requires diverse model parameter spaces. Most existing fine-tuning methods attempt to bridge the gaps among different scenarios by introducing a set of new parameters to modify SAM's original parameter space. Unlike these works, in this paper, we propose fine-tuning SAM efficiently by parameter space reconstruction (SAM-PARSER), which introduce nearly zero trainable parameters during fine-tuning. In SAM-PARSER, we assume that SAM's original parameter space is relatively complete, so that its bases are able to reconstruct the parameter space of a new scenario. We obtain the bases by matrix decomposition, and fine-tuning the coefficients to reconstruct the parameter space tailored to the new scenario by an optimal linear combination of the bases. Experimental results show that SAM-PARSER exhibits superior segmentation performance across various scenarios, while reducing the number of trainable parameters by approximately 290 times compared with current parameter-efficient fine-tuning methods. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Xiaokang Yang 0001, Wei Shen 0002 |
AAAI | 1 |
| 2024 | LERE: Learning-Based Low-Rank Matrix Recovery with Rank EstimationabstractA fundamental task in the realms of computer vision, Low-Rank Matrix Recovery (LRMR) focuses on the inherent low-rank structure precise recovery from incomplete data and/or corrupted measurements given that the rank is a known prior or accurately estimated. However, it remains challenging for existing rank estimation methods to accurately estimate the rank of an ill-conditioned matrix. Also, existing LRMR optimization methods are heavily dependent on the chosen parameters, and are therefore difficult to adapt to different situations. Addressing these issues, A novel LEarning-based low-rank matrix recovery with Rank Estimation (LERE) is proposed. More specifically, considering the characteristics of the Gerschgorin disk's center and radius, a new heuristic decision rule in the Gerschgorin Disk Theorem is significantly enhanced and the low-rank boundary can be exactly located, which leads to a marked improvement in the accuracy of rank estimation. According to the estimated rank, we select row and column sub-matrices from the observation matrix by uniformly random sampling. A 17-iteration feedforward-recurrent-mixed neural network is then adapted to learn the parameters in the sub-matrix recovery processing. Finally, by the correlation of the row sub-matrix and column sub-matrix, LERE successfully recovers the underlying low-rank matrix. Overall, LERE is more efficient and robust than existing LRMR methods. Experimental results demonstrate that LERE surpasses state-of-the-art (SOTA) methods. The code for this work is accessible at https://github.com/zhengqinxu/LERE. Zhengqin Xu, Yulun Zhang 0001, Chao Ma 0004, Yichao Yan, Zelin Peng, Shoulie Xie, Shiqian Wu, Xiaokang Yang 0001 |
AAAI | 5 |
| 2024 | DeCo-Net: Robust Multimodal Brain Tumor Segmentation via Decoupled Complementary Knowledge DistillationabstractAutomated brain tumor segmentation with multimodal magnetic resonance imaging (MRI) plays a pivotal rule in clinical application. However, most existing algorithms require complete image modalities as input, which is often impractical to obtain for every patient in real clinical practice. Therefore, a robust multimodal algorithm that is capable of handling various modality-incomplete data is highly desirable. In this paper, we propose DeCo-Net, a Decoupled Complementary knowledge distillation framework for multimodal brain tumor segmentation with incomplete modalities. Specifically, our approach decouples the feature learning of the modality-incomplete data into two branches: one dedicated to extracting the inherent features from the available modalities and the other focused on inferring the complementary missing modal information. We employ a teacher-student co-training framework where the teacher network is collaboratively trained to dynamically transfer the complementary knowledge to the student model based on the specific type of modality-incomplete data fed to student. To this end, we propose a modality-aware contrastive distillation strategy that guides the student model to distill a discriminative and complementary knowledge representation that acts as supplements to the original modality-incomplete representation. Extensive evaluations on the BraTS2018, BraTS2020 and BraTS2023 datasets demonstrate that our method achieves state-of-the-art performance in multimodal brain tumor segmentation with incomplete modalities. Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002 |
BIBM | 2 |
| 2024 | Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything ModelabstractParameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community, PEFT has shown effectiveness in image classification, but little research has studied its ability for image segmentation. Fine-tuning segmentation models usually requires a heavier adjustment of parameters to align the proper projection directions in the parameter space for new scenarios. This raises a challenge to existing PEFT algorithms, as they often inject a limited number of individual parameters into each block, which prevents substantial adjustment of the projection direction of the parameter space due to the limitation of Hidden Markov Chain along blocks. In this paper, we equip PEFT with a cross-block orchestration mechanism to enable the adaptation of the Segment Anything Model (SAM) to various downstream scenarios. We introduce a novel inter-block communication module, which integrates a learnable relation matrix to facilitate communication among different coefficient sets of each PEFT block's parameter space. Moreover, we propose an intra-block enhancement module, which introduces a linear projection head whose weights are generated from a hyper-complex layer, further enhancing the impact of the adjustment of projection directions on the entire parameter space. Extensive experiments on diverse benchmarks demonstrate that our proposed approach consistently improves the segmentation performance significantly on novel scenarios with only around 1K additional parameters. Zelin Peng, Zhengqin Xu, Zhilin Zeng, Lingxi Xie, Qi Tian 0001, Wei Shen 0002 |
CVPR | 1 |
| 2024 | Missing as Masking: Arbitrary Cross-Modal Feature Reconstruction for Incomplete Multimodal Brain Tumor Segmentation
Zhilin Zeng, Zelin Peng, Xiaokang Yang 0001, Wei Shen 0002 |
MICCAI (8) | 2 |
| 2023 | USAGE: A Unified Seed Area Generation Paradigm for Weakly Supervised Semantic SegmentationabstractSeed area generation is usually the starting point of weakly supervised semantic segmentation (WSSS). Computing the Class Activation Map (CAM) from a multi-label classification network is the de facto paradigm for seed area generation, but CAMs generated from Convolutional Neural Networks (CNNs) and Transformers are prone to be under- and over-activated, respectively, which makes the strategies to refine CAMs for CNNs usually inappropriate for Transformers, and vice versa. In this paper, we propose a Unified optimization paradigm for Seed Area GEneration (USAGE) for both types of networks, in which the objective function to be optimized consists of two terms: One is a generation loss, which controls the shape of seed areas by a temperature parameter following a deterministic principle for different types of networks; The other is a regularization loss, which ensures the consistency between the seed areas that are generated by self-adaptive network adjustment from different views, to overturn false activation in seed areas. Experimental results show that USAGE consistently improves seed area generation for both CNNs and Transformers by large margins, e.g., outperforming state-of-the-art methods by a mIoU of 4.1% on PASCAL VOC. Moreover, based on the USAGE-generated seed areas on Transformers, we achieve state-of-the-art WSSS results on both PASCAL VOC and MS COCO. Zelin Peng, Guanchun Wang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001 |
ICCV | 1 |
| 2023 | A Survey on Label-Efficient Deep Image Segmentation: Bridging the Gap Between Weak Supervision and Dense PredictionabstractThe rapid development of deep learning has made a great progress in image segmentation, one of the fundamental tasks of computer vision. However, the current segmentation algorithms mostly rely on the availability of pixel-level annotations, which are often expensive, tedious, and laborious. To alleviate this burden, the past years have witnessed an increasing attention in building label-efficient, deep-learning-based image segmentation algorithms. This paper offers a comprehensive review on label-efficient image segmentation methods. To this end, we first develop a taxonomy to organize these methods according to the supervision provided by different types of weak labels (including no supervision, inexact supervision, incomplete supervision and inaccurate supervision) and supplemented by the types of segmentation problems (including semantic segmentation, instance segmentation and panoptic segmentation). Next, we summarize the existing label-efficient image segmentation methods from a unified perspective that discusses an important question: how to bridge the gap between weak supervision and dense prediction - the current methods are mostly based on heuristic priors, such as cross-pixel similarity, cross-label constraint, cross-view consistency, and cross-image relation. Finally, we share our opinions about the future research directions for label-efficient deep image segmentation. Wei Shen 0002, Zelin Peng, Huayu Wang, Jiazhong Cen, Dongsheng Jiang, Lingxi Xie, Xiaokang Yang 0001, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Absolute Wrong Makes Better: Boosting Weakly Supervised Object Detection via Negative Deterministic InformationabstractWeakly supervised object detection (WSOD) is a challenging task, in which image-level labels (e.g., categories of the instances in the whole image) are used to train an object detector. Many existing methods follow the standard multiple instance learning (MIL) paradigm and have achieved promising performance. However, the lack of deterministic information leads to part domination and missing instances. To address these issues, this paper focuses on identifying and fully exploiting the deterministic information in WSOD. We discover that negative instances (i.e. absolutely wrong instances), ignored in most of the previous studies, normally contain valuable deterministic information. Based on this observation, we here propose a negative deterministic information (NDI) based method for improving WSOD, namely NDI-WSOD. Specifically, our method consists of two stages: NDI collecting and exploiting. In the collecting stage, we design several processes to identify and distill the NDI from negative instances online. In the exploiting stage, we utilize the extracted NDI to construct a novel negative contrastive learning mechanism and a negative guided instance selection strategy for dealing with the issues of part domination and missing instances, respectively. Experimental results on several public benchmarks including VOC 2007, VOC 2012 and MS COCO show that our method achieves satisfactory performance. Guanchun Wang, Xiangrong Zhang, Zelin Peng, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao |
IJCAI | 3 |
| 2021 | Time-aware Neural Collaborative Filtering with Multi-dimensional Features on Academic Paper RecommendationabstractIn modern academic social network, it is very difficult for scholars to find academic papers consistent with their research direction. Time is a critical factor in paper recommendation. As time goes on, the impact of an academic paper would gradually fade. Likewise, the research interests of users may also change. Therefore, we propose a temporal perceptual neural collaborative filtering model that integrates the multi-dimensional features of papers. We conducted our experiments on the dataset from CiteULike, comparing the recommended results by using four time-decay functions and evaluating our model with multiple evaluation indicators. The satisfactory results show that our model is effective in filtering out the expired papers by considering the characteristics of papers and the changes of scholars' interests. Yibo Lu, Yixiang Cai, Zelin Peng, Yong Tang 0001 |
CSCWD | 4 |
| 2021 | Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic SegmentationabstractSemantic segmentation has been continuously investigated in the last ten years, and majority of the established technologies are based on supervised models. In recent years, image-level weakly supervised semantic segmentation (WSSS), including single- and multi-stage process, has attracted large attention due to data labeling efficiency. In this paper, we propose to embed affinity learning of multi-stage approaches in a single-stage model. To be specific, we introduce an adaptive affinity loss to thoroughly learn the local pairwise affinity. As such, a deep neural network is used to deliver comprehensive semantic information in the training phase, whilst improving the performance of the final prediction module. On the other hand, considering the existence of errors in the pseudo labels, we propose a novel label reassign loss to mitigate over-fitting. Extensive experiments are conducted on the PASCAL VOC 2012 dataset to evaluate the effectiveness of our proposed approach that outperforms other standard single-stage methods and achieves comparable performance against several multi-stage methods. Xiangrong Zhang, Zelin Peng, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Huiyu Zhou 0001, Licheng Jiao |
ACM Multimedia | 2 |
| 2020 | A Learnable Blur Kernel for Remote Sensing Image RetrievalabstractWith the explosive increase of remote sensing images, content-based remote sensing image retrieval (CBRSIR) has aroused widespread attention. Convolutional Neural Network (CNN) based methods are widely used in CBRSIR due to the development of deep learning. However, common used CNN models have difficulties in holding shift-invariant property due to the widely used down-sampling method, which means a little shift of input may cause a mutation of feature representation. To mitigate the absence of shift-invariant in down-sampling, we propose the learnable blur kernel (LBK), that can enhance the feature extraction capability by leveraging more context information. We build on this concept without extra cost, which can be simply integrated with modern CNNs architecture. Our method is validated on the public remote sensing dataset and compared with other retrieval methods. The overall experimental results show that the proposed method achieves outstanding performance. Zelin Peng, Guanchun Wang, Xiangrong Zhang, Xu Tang 0004, Licheng Jiao |
IGARSS | 1 |