EDBT 2026 Demo / reviewers in the wild / expert
Wentao Zhang 0005
dblp:41/3249-5
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-7866-6377ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language ModelsabstractVision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs and have achieved strong performance on standard image datasets. However, their effectiveness is often limited when confronted with cross-domain tasks where imaging domains differ from natural images. To address this limitation, we propose Consistency-guided Multi-view Collaborative Optimization (CoMuCo), a novel fine-tuning strategy for VLMs. This strategy employs two functionally complementary expert modules to extract multi-view features, while incorporating prior knowledge-based consistency constraints and information geometry-based consensus mechanisms to enhance the robustness of feature learning. Additionally, a new cross-domain few-shot benchmark is established to help comprehensively evaluate methods on imaging domains distinct from natural images. Extensive empirical evaluations on both existing and newly proposed benchmarks suggest CoMuCo consistently outperforms current methods. Dexia Chen, Wentao Zhang 0005, Qianjie Zhu, Weibing Li, Tong Zhang 0017 |
AAAI | 2 |
| 2026 | Decoupling Continual Semantic SegmentationabstractContinual Semantic Segmentation (CSS) requires learning new classes without forgetting previously acquired knowledge, addressing the fundamental challenge of catastrophic forgetting in dense prediction tasks. However, existing CSS methods typically employ single-stage encoder-decoder architectures where segmentation masks and class labels are tightly coupled, leading to interference between old and new class learning and suboptimal retention-plasticity balance. We introduce DecoupleCSS, a novel two-stage framework for CSS. By decoupling class-aware detection from class-agnostic segmentation, DecoupleCSS enables more effective continual learning, preserving past knowledge while learning new classes. The first stage leverages pre-trained text and image encoders, adapted using LoRA, to encode class-specific information and generate location-aware prompts. In the second stage, the Segment Anything Model (SAM) is employed to produce precise segmentation masks, ensuring that segmentation knowledge is shared across both new and previous classes. This approach improves the balance between retention and adaptability in CSS, achieving state-of-the-art performance across a variety of challenging tasks. Yifu Guo, Yuquan Lu, Wentao Zhang 0005, Zishan Xu, Dexia Chen, Yizhe Zhang 0001 |
AAAI | 3 |
| 2026 | Dual-modality adaptation in vision-language models for continual learning
Jiayang Zeng, Wentao Zhang 0005, Kanghao Chen, Jiantao Tan, Wei-Shi Zheng 0001 |
Neural Networks | 2 |
| 2025 | DAT: Dual-Branch Adapter-Tuning for Few-Shot RecognitionabstractParameter-Efficient Fine-Tuning methods based on vision-language models (such as CLIP) for few-shot learning have recently received considerable attention. However, previous works only fine-tune either the image or text branch, breaking the alignment of the original two branches, meanwhile fine-tuning both branches of the CLIP would inevitably introduce more trainable parameters and likely cause more severe over-fitting due to the limited training data. In this study, we propose a novel Dual-branch Adapter-Tuning framework (DAT), which collaboratively trains the visual adapter and textual adapter added to the two branches of the original CLIP with multiple consistency constraints. By effectively utilizing the semantically detailed class-specific prompts and outputs of the original CLIP to guide the fine-tuning of both branches, our method gains exceptional adaptation ability to the downstream few-shot learning tasks and alleviates the over-fitting issue, meanwhile maximally preserving the generalization ability of the original CLIP model. Our proposed framework has achieved superior performance on diverse datasets under various few-shot learning settings compared to the existing approaches. The source code is available athttps://github.com/SandyXi/DAT. Junxi Chen, Guangxing Wu, Hongxiang Li 0004, Jiankang Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Visual Class Incremental Learning With Textual Priors Guidance Based on an Adapted Vision-Language ModelabstractAn ideal artificial intelligence (AI) system should have the capability to continually learn like humans. However, when learning new knowledge, AI systems often suffer from catastrophic forgetting of old knowledge. Although many continual learning methods have been proposed, they often ignore the issue of misclassifying similar classes and make insufficient use of textual priors of visual classes to improve continual learning performance. In this study, we propose a continual learning framework based on a pre-trained vision-language model (VLM) that does not require storing old class data. This framework utilizes parameter-efficient fine-tuning of the VLM's text encoder for constructing a shared and consistent semantic textual space throughout the continual learning process. The textual priors of visual classes are encoded by the adapted VLM's text encoder to generate discriminative semantic representations, which are then used to guide the learning of visual classes. Additionally, fake out-of-distribution (OOD) images constructed from each training image further assist in the learning of visual classes. Extensive empirical evaluations on three natural datasets and one medical dataset demonstrate the superiority of the proposed framework. Wentao Zhang 0005, Jianhui Xie, Emanuele Trucco, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | TexCIL: Text-Guided Continual Learning of Disease with Vision-Language ModelabstractCurrent intelligent diagnostic systems often catastrophically forget old knowledge when learning new diseases only from the training dataset of the new diseases. Inspired by human learning of visual classes with the effective help of language, we propose a continual learning framework based on a pre-trained visual-language model (VLM) without storing any image of previously learned diseases. In this framework, textual prior knowledge of each new disease can be obtained by utilizing the frozen VLM’s text encoder, and then used to guide the visual learning of the new disease. This framework innovatively utilizes the textual prior knowledge of all previously learned diseases as out-of-distribution (OOD) information to help differentiate currently being-learned diseases from others. Extensive empirical evaluations on both medical and natural image datasets confirm the superiority of the proposed method over existing state-of-the-art methods in continual learning of new visual classes. The source code is available at https://openi.pcl.ac.cn/OpenMedIA/TexCIL. Wentao Zhang 0005, Defeng Zhao, Wei-Shi Zheng 0001 |
BIBM | 1 |
| 2024 | CtF: Mitigating Visual Confusion in Continual Learning Through a Coarse-To-Fine Screening
Zejun Ye, Defeng Zhao, Wentao Zhang 0005 |
ICIC (6) | 3 |
| 2024 | Region Attention Fine-tuning with CLIP for Few-shot ClassificationabstractWith the advancements in visual language models such as CLIP and their strong performance in zero-shot recognition, numerous CLIP-based methods have emerged in the field of few-shot classification. However, many of them do not fully leverage the abundant feature information within the CLIP visual encoder and overlook the issue of varying region-specific importance for image classification across different datasets. To address these limitations, we present an attention pooling-based framework for few-shot fine-tuning. Our framework enables the model to learn task-specific attention weights for image regions, while also incorporating background features and a consistency constraint to enhance training. As a result, our approach outperforms the state-of-the-art approaches on 11 benchmarks, demonstrating its effectiveness. Guangxing Wu, Junxi Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001 |
ICME | 4 |
| 2024 | Expand and Merge: Continual Learning with the Guidance of Fixed Text Embedding SpaceabstractDeep neural networks lack the ability to sequentially learn from new data and adapt to new scenarios. In particular, after learning new data, neural networks will have a significant performance degradation on old knowledge. This phenomenon is known as catastrophic forgetting. To mitigate this issue, we propose expanding and merging additional parameters, encapsulated in a specially designed adapter layer, in the frozen pretrained vision encoder. Along with the newly added parameters, adapter scaling weights in each layer are also introduced to adaptively control the fusion of new and old knowledge. Additionally, the fixed embedding space of a pretrained text encoder is used to guide the continual learning of the vision encoder. Extensive experiments on three datasets demonstrate that the proposed method outperform current state-of-the-art methods. The source code is available at https://github.com/GiantJun/Expand_and_Merge. Yujun Huang, Wentao Zhang 0005 |
IJCNN | 2 |
| 2024 | Enhancing Task Identification Through Pseudo-OOD Features for Class-Incremental Learning
Weizhuo Zhang, Jiankang Chen, Wentao Zhang 0005, Zhijun Tan |
PRCV (3) | 3 |
| 2024 | PAMI: Partition Input and Aggregate Outputs for Model Interpretation
Wentao Zhang 0005, Wei-Shi Zheng 0001 |
Pattern Recognit. | 2 |
| 2024 | Continual Learning of Image Classes With Language Guidance From a Vision-Language ModelabstractCurrent deep learning models often catastrophically forget the knowledge of old classes when continually learning new ones. State-of-the-art approaches to continual learning of image classes often require retaining a small subset of old data to partly alleviate the catastrophic forgetting issue, and their performance would be degraded sharply when no old data can be stored due to privacy or safety concerns. In this study, inspired by human learning of visual knowledge with the effective help of language, we propose a novel continual learning framework based on a pre-trained vision-language model (VLM) without retaining any old data. Rich prior knowledge of each new image class is effectively encoded by the frozen text encoder of the VLM, which is then used to guide the learning of new image classes. The output space of the frozen text encoder is unchanged over the whole process of continual learning, through which image representations of different classes become comparable during model inference even when the image classes are learned at different times. Extensive empirical evaluations on multiple image classification datasets under various settings confirm the superior performance of our method over existing ones. The source code is available athttps://github.com/Fatflower/CIL_LG_VLM/. Wentao Zhang 0005, Yujun Huang, Weizhuo Zhang, Tong Zhang 0017, Qicheng Lao, Yue Yu 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases
Wentao Zhang 0005, Yujun Huang, Tong Zhang 0017, Qingsong Zou, Wei-Shi Zheng 0001 |
MICCAI (2) | 1 |
| 2023 | Feature Adaptation with CLIP for Few-shot ClassificationabstractLarge Vision-Language models such as CLIP have demonstrated impressive capabilities in zero-shot recognition. To apply CLIP to few-shot classification tasks, several methods have been proposed based on CLIP, achieving significant improvements. However, these methods either insufficiently leverage CLIP’s prior knowledge during training or neglect the impact of feature adaptation. In this paper, we propose FAR, a novel approach that balances distribution-altered Feature Adaptation with pRior knowledge of CLIP to further improve the performance of CLIP in few-shot classification tasks. Firstly, we introduce an adapter that enhances the effectiveness of CLIP adaptation by amplifying the differences between the fine-tuned CLIP features and the original CLIP features. Secondly, we leverage the prior knowledge of CLIP to mitigate the risk of overfitting. Through this framework, a good trade-off between feature adaptation and preserving prior knowledge is achieved, enabling effective utilization of both components to enhance performance on downstream tasks. We evaluate our method on over 10 datasets for classification, and our approach consistently outperforms existing methods, demonstrating its effectiveness and robustness. Guangxing Wu, Junxi Chen, Wentao Zhang 0005 |
MMAsia | 3 |