Wei Feng 0015

dblp:17/1152-15 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-7398-6988ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Leveraging Image-Text Pairs for Generalized Category Discovery in Medical Image Classification
abstract
Generalized category discovery aims to identify known medical categories and unknown new medical categories from unlabeled data by migrating knowledge from labeled datasets containing only known categories, which is crucial for disease understanding and precision medicine. Many methods have been proposed and significantly improved the performance of GCD in medical images. However, most of the existing methods discover new categories based on image modalities only, ignoring useful information in the large amount of textual data related to diseases. In this paper, we propose M3GCD (Medical Multi-Modal Generalized Category Discovery), which exploits image– text pairs to jointly recognize known classes and discover novel categories in medical images. To address the varying contribution of different modalities across samples, we develop a Dynamic Expert Fusion module to automatically learn sample-specific modality weights, and further design a Local Experts Balancing mechanism to preserve the discriminative power of individual modalities. By integrating global and local perspectives, our framework adaptively balances modality contributions and enhances multi-modal robustness. Subsequently, to enable the discovery of novel unknown categories during training, we propose a Category Diffusion module grounded in the Metropolis– Hastings framework. This module adaptively merges and splits categories, allowing the model to simultaneously recognize known classes and uncover previously unseen categories during training, without requiring any prior knowledge about the unknown categories. Extensive experiments on two public multi-modal datasets (MIMIC-CXR and PatchGastric), together with a private multi-modal fundus dataset, MM-Retina, demonstrate that our method consistently improves clustering performance on both known and unknown categories compared with existing approaches.
Wei Feng 0015, Sijin Zhou, ZongYuan Ge
IEEE Trans. Medical Imaging1
2025 Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models
abstract
Concept Bottleneck Models (CBMs) decompose image classification into a process governed by interpretable, human-readable concepts. Recent advances in CBMs have used Large Language Models (LLMs) to generate candidate concepts. However, a critical question remains: What is the optimal number of concepts to use? Current concept banks suffer from redundancy or insufficient coverage. To address this issue, we introduce a dynamic, agent-based approach that adjusts the concept bank in response to environmental feedback, optimizing the number of concepts for sufficiency yet concise coverage. Moreover, we propose Conditional Concept Bottleneck Models (CoCoBMs) to overcome the limitations in traditional CBMs’ concept scoring mechanisms. It enhances the accuracy of assessing each concept’s contribution to classification tasks and feature an editable matrix that allows LLMs to correct concept scores that conflict with their internal knowledge. Our evaluations across 6 datasets show that our method not only improves classification accuracy by 6% but also enhances interpretability assessments by 30%.
Yiwen Jiang, Deval Mehta 0001, Wei Feng 0015, ZongYuan Ge
ACL (1)3
2025 Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
abstract
Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness.
Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge
CVPR15
2025 Generalized Category Discovery under Domain Shift: A Frequency Domain Perspective
abstract
Generalized Category Discovery (GCD) aims to leverage labeled samples from known categories to cluster unlabeled data that may include both known and unknown categories. While existing methods have achieved impressive results under standard conditions, their performance often deteriorates in the presence of distribution shifts. In this paper, we explore a more realistic task: Domain-Shifted Generalized Category Discovery (DS\_GCD), where the unlabeled data includes not only unknown categories but also samples from unknown domains. To tackle this challenge, we propose a \textbf{\underline{F}}requency-guided Gene\textbf{\underline{r}}alized Cat\textbf{\underline{e}}gory Discov\textbf{\underline{e}}ry framework (FREE) that enhances the model's ability to discover categories under distributional shift by leveraging frequency-domain information. Specifically, we first propose a frequency-based domain separation strategy that partitions samples into known and unknown domains by measuring their amplitude differences. We then propose two types of frequency-domain perturbation strategies: a cross-domain strategy, which adapts to new distributions by exchanging amplitude components across domains, and an intra-domain strategy, which enhances robustness to intra-domain variations within the unknown domain. Furthermore, we extend the self-supervised contrastive objective and semantic clustering loss to better guide the training process. Finally, we introduce a clustering-difficulty-aware resampling technique to adaptively focus on harder-to-cluster categories, further enhancing model performance. Extensive experiments demonstrate that our method effectively mitigates the impact of distributional shifts across various benchmark datasets and achieves superior performance in discovering both known and unknown categories.
Wei Feng 0015, ZongYuan Ge
NeurIPS1
2025 Neighbor-Guided Unbiased Framework for Generalized Category Discovery in Medical Image Classification
abstract
Generalized category discovery (GCD) utilizes seen category knowledge to automatically discover new semantic categories that are not defined in the training phase. Nevertheless, there has been no research conducted on identifying new classes using medical images and disease categories, which is essential for understanding and diagnosing specific diseases. Moreover, existing methods still produce predictions that are biased towards seen categories since the model is mainly supervised by labeled seen categories, which in turn leads to sub-optimal clustering performance. In this paper, we propose a new neighbor-guided unbiased framework (NGUF) that leverages neighbor information to mitigate prediction bias to address the GCD problem in medical tasks. Specifically, we devise a neighbor-guided cross-pseudo-clustering strategy, which exploits the knowledge of the nearest-neighbor samples to adjust the model predictions thereby generating unbiased pseudo-clustering supervision. Then, based on the unbiased pseudo-clustering supervision, we use a view-invariant learning strategy to assign labels to all samples. In addition, we propose an adaptive weight learning strategy that dynamically determines the degree of adjustment of the predictions of different samples based on the distance density values. Finally, we further propose a cross-batch knowledge distillation module to utilize information from successive iterations to encourage training consistency. Extensive experiments on four medical image datasets show that NGUF is effective in mitigating the model's prediction bias and has superior performance to other state-of-the-art GCD algorithms. Our code will be released soon.
Wei Feng 0015, Sijin Zhou, Yiwen Jiang, ZongYuan Ge
IEEE J. Biomed. Health Informatics1
2024 Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation
abstract
Recent weakly supervised semantic segmentation (WSSS) methods strive to incorporate contextual knowledge to improve the completeness of class activation maps (CAM). In this work, we argue that the knowledge bias between instances and contexts affects the capability of the prototype to sufficiently understand instance semantics. Inspired by prototype learning theory, we propose leveraging prototype awareness to capture diverse and fine-grained feature attributes of instances. The hypothesis is that contextual prototypes might erroneously activate similar and frequently co-occurring object categories due to this knowledge bias. Therefore, we propose to enhance the prototype representation ability by mitigating the bias to better capture spatial coverage in semantic object regions. With this goal, we present a Context Prototype-Aware Learning (CPAL) strategy, which leverages semantic context to enrich instance comprehension. The core of this method is to accurately capture intra-class variations in object features through context-aware prototypes, facilitating the adaptation to the semantic attributes of various instances. We design feature distribution alignment to optimize prototype awareness, aligning instance feature distributions with dense features. In addition, a unified training framework is proposed to combine label-guided classification supervision and prototypes-guided self-supervision. Experimental results on PASCAL VOC 2012 and MS COCO 2014 show that CPAL significantly improves off-the-shelf methods and achieves state-of-the-art performance. The project is available at https://github.com/Barrett-python/CPAL.
Zhongxing Xu, Zhaojun Qu, Wei Feng 0015, Xingjian Jiang, ZongYuan Ge
CVPR4
2024 Universal Semi-supervised Learning for Medical Image Classification
Lie Ju, Yicheng Wu 0001, Wei Feng 0015, Lin Wang 0027, Zhuoting Zhu, ZongYuan Ge
MICCAI (12)3
2023 Unsupervised Domain Adaptation for Medical Image Segmentation by Selective Entropy Constraints and Adaptive Semantic Alignment
abstract
Generalizing a deep learning model to new domains is crucial for computer-aided medical diagnosis systems. Most existing unsupervised domain adaptation methods have made significant progress in reducing the domain distribution gap through adversarial training. However, these methods may still produce overconfident but erroneous results on unseen target images. This paper proposes a new unsupervised domain adaptation framework for cross-modality medical image segmentation. Specifically, We first introduce two data augmentation approaches to generate two sets of semantics-preserving augmented images. Based on the model's predictive consistency on these two sets of augmented images, we identify reliable and unreliable pixels. We then perform a selective entropy constraint: we minimize the entropy of reliable pixels to increase their confidence while maximizing the entropy of unreliable pixels to reduce their confidence. Based on the identified reliable and unreliable pixels, we further propose an adaptive semantic alignment module which performs class-level distribution adaptation by minimizing the distance between same class prototypes between domains, where unreliable pixels are removed to derive more accurate prototypes. We have conducted extensive experiments on the cross-modality cardiac structure segmentation task. The experimental results show that the proposed method significantly outperforms the state-of-the-art comparison algorithms. Our code and data are available at https://github.com/fengweie/SE_ASA.
Wei Feng 0015, Lie Ju, Lin Wang 0027, Kaimin Song, ZongYuan Ge
AAAI1
2023 Towards Novel Class Discovery: A Study in Novel Skin Lesions Clustering
Wei Feng 0015, Lie Ju, Lin Wang 0027, Kaimin Song, ZongYuan Ge
MICCAI (6)1
2023 NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding
abstract
The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education, improve quality control, and enable operational compliance monitoring. However, the development of automatic recognition systems in this field is currently hindered by the scarcity of appropriately labeled datasets. The existing video datasets pose several limitations: 1) these datasets are small-scale in size to support comprehensive investigations of nursing activity; 2) they primarily focus on single procedures, lacking expert-level annotations for various nursing procedures and action steps; and 3) they lack temporally localized annotations, which prevents the effective localization of targeted actions within longer video sequences. To mitigate these limitations, we propose NurViD, a large video dataset with expert-level annotation for nursing procedure activity understanding. NurViD consists of over 1.5k videos totaling 144 hours, making it approximately four times longer than the existing largest nursing activity datasets. Notably, it encompasses 51 distinct nursing procedures and 177 action steps, providing a much more comprehensive coverage compared to existing datasets that primarily focus on limited procedures. To evaluate the efficacy of current deep learning methods on nursing activity understanding, we establish three benchmarks on NurViD: procedure recognition on untrimmed videos, procedure and action recognition on trimmed videos, and action detection. Our benchmark and code will be available at https://github.com/minghu0830/NurViD-benchmark.
Lin Wang 0027, Siyuan Yan, Don Ma, Qingli Ren, Peng Xia 0005, Wei Feng 0015, Peibo Duan, Lie Ju, ZongYuan Ge
NeurIPS7
2022 Unsupervised Domain Adaptive Fundus Image Segmentation with Category-Level Regularization
Wei Feng 0015, Lin Wang 0027, Lie Ju, Xin Wang 0094, ZongYuan Ge
MICCAI (2)1