Wenxue Li 0003

dblp:12/8919-3 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-1301-4933ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 PromptHaze: Prompting Real-world Dehazing via Depth Anything Model
abstract
Real-world image dehazing remains a challenging task due to the diverse nature of haze degradation and the lack of large-scale paired datasets. Existing methods based on hand-crafted priors or generative priors struggle to recover accurate backgrounds and fine details from dense haze regions. In this work, we propose a novel paradigm, PromptHaze, for real-world image dehazing via the depth prompt from the Depth Anything model. By employing a prompt-by-prompt strategy, our method iteratively updates the depth prompt and progressively restores the background through a dehazing network with controllable dehazing strength. Extensive experiments on widely-used real-world dehazing benchmarks demonstrate the superiority of PromptHaze in recovering authentic backgrounds and fine details from various haze scenes, outperforming state-of-the-art methods across multiple quality metrics.
Tian Ye 0001, Sixiang Chen, Haoyu Chen 0003, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003
AAAI7
2025 Towards Realistic Semi-supervised Medical Image Classification
abstract
Existing semi-supervised learning (SSL) approaches follow the idealized closed-world assumption, neglecting the challenges present in realistic medical scenarios, such as open-set distribution and imbalanced class distribution. Although some methods in natural domains attempt to address the open-set problem, they are insufficient for medical domains, where intertwined challenges like class imbalance and small inter-class lesion discrepancies persist. Thus, this paper presents a novel self-recalibrated semantic training framework, which is tailored for SSL in medical imaging by ingeniously harvesting realistic unlabeled samples. Inspired by the observation that certain open-set samples share some similar disease-related representations with in-distribution samples, we first propose an informative sample selection strategy that identifies high-value samples to serve as augmentations, thereby effectively enriching the semantics of known categories. Furthermore, we adopt a compact semantic clustering strategy to address the semantic confusion raised by the above newly introduced open-set semantics. Moreover, to mitigate the interference of class imbalance in open-set SSL, we introduce a less biased dual-balanced classifier with similarity pseudo-label regularization and category-customized regularization. Extensive experiments on a variety of medical image datasets demonstrate the superior performance of our proposed method over state-of-the-art Closed-set and Open-set SSL methods.
Wenxue Li 0003, Lie Ju, Peng Xia 0005, Xinyu Xiong, Lei Zhu 0002, ZongYuan Ge
AAAI1
2025 AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image Enhancement
abstract
Existing low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often impractical and sometimes even unavailable, and 2) accurately modeling complex degradations presents a non-trivial problem. To overcome them, we propose the Attribute Guidance Diffusion framework (AGLLDiff), a training-free method for effective real-world LIE. Instead of specifically defining the degradation process, AGLLDiff shifts the paradigm and models the desired attributes, such as image exposure, structure and color of normal-light images. These attributes are readily available and impose no assumptions about the degradation process, which guides the diffusion sampling process to a reliable high-quality solution space. Extensive experiments demonstrate that our approach outperforms the current leading unsupervised LIE methods across benchmarks in terms of distortion-based and perceptual-based metrics, and it performs well even in sophisticated wild degradation.
Yunlong Lin, Tian Ye 0001, Sixiang Chen, Zhenqi Fu, Yingying Wang 0005, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003, Xinghao Ding
AAAI8
2025 Neighbor Does Matter: Density-Aware Contrastive Learning for Medical Semi-supervised Segmentation
abstract
In medical image analysis, multi-organ semi-supervised segmentation faces challenges such as insufficient labels and low contrast in soft tissues. To address these issues, existing studies typically employ semi-supervised segmentation techniques using pseudo-labeling and consistency regularization. However, these methods mainly rely on individual data samples for training, ignoring the rich neighborhood information present in the feature space. In this work, we argue that supervisory information can be directly extracted from the geometry of the feature space. Inspired by the density-based clustering hypothesis, we propose using feature density to locate sparse regions within feature clusters. Our goal is to increase intra-class compactness by addressing sparsity issues. To achieve this, we propose a Density-Aware Contrastive Learning (DACL) strategy, pushing anchored features in sparse regions towards cluster centers approximated by high-density positive samples, resulting in more compact clusters. Specifically, our method constructs density-aware neighbor graphs using labeled and unlabeled data samples to estimate feature density and locate sparse regions. We also combine label-guided co-training with density-guided geometric regularization to form complementary supervision for unlabeled data. Experiments on the Multi-Organ Segmentation Challenge dataset demonstrate that our proposed method outperforms state-of-the-art methods, highlighting its efficacy in medical image segmentation tasks.
Zhongxing Xu, Wenxue Li 0003, Peng Xia 0005, Yiheng Zhong, Hanjun Wu, Jionglong Su, ZongYuan Ge
AAAI4
2025 Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
abstract
Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness.
Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge
CVPR11
2025 Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data Engine
abstract
Mirror detection is a challenging task because a mirror’s visual appearance varies depending on the reflected content. Due to limited annotated data, current methods failed to generalize well for detecting diverse mirror scenes. Semi-supervised learning with large-scale unlabeled data can improve generalization capabilities on mirror detection, but these methods often suffer from unreliable pseudo-labels due to distribution differences between labeled and unlabeled data, therefore affecting the learning process. To address this issue, we first collect a large-scale dataset of approximately 0.4 million mirror-related images from the internet, significantly expanding the data scale for mirror detection. To effectively exploit this unlabeled dataset, we propose the first semi-supervised framework (namely an iterative data engine) consisting of four steps: (1) mirror detection model training, (2) pseudo label prediction, (3) dual guidance scoring, and (4) selection of highly reliable pseudo labels. In each iteration of the data engine, we employ a geometric accuracy scoring approach to assess pseudo labels based on multiple segmentation metrics, and design a multi-modal agent-driven semantic scoring approach to enhance the semantic perception of pseudo labels. These two scoring approaches can effectively improve the reliability of pseudo labels by selecting unlabeled samples with higher scores. Our method demonstrates promising performance across three mirror detection tasks and exhibits strong generalization on unseen examples. Our code will be publicly available at https://github.com/ge-xing/DAM.
Zhaohu Xing, Hongqiu Wang, Tian Ye 0001, Sixiang Chen, Wenxue Li 0003, Guang Liu 0006, Lei Zhu 0003
CVPR7
2025 GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection
Wenxue Li 0003, Tian Ye 0001, Xinyu Xiong, Jinbin Bai, Wenxuan Song, Zhaohu Xing, Lie Ju, Guanbin Li, Lei Zhu 0003
ICCV1
2025 FSA-Net: Fractal-Driven Synergistic Anatomy-Aware Network for Segmenting White Line of Toldt in Laparoscopic Images
Kecheng Wu, Zhaohu Xing, Zerong Cai, Feng Gao 0023, Wenxue Li 0003, Lei Zhu 0003
MICCAI (9)5
2025 Free Meal: Boosting Semi-Supervised Polyp Segmentation by Harvesting Negative Samples
abstract
Existing semi-supervised polyp segmentation methods assume that unlabeled images are positive, containing lesions to be annotated, while neglecting negative samples that are widely available in practice. This letter reveals that harvesting lesion-free negative samples can effectively boost polyp segmentation performance. Directly extending the labeled set with negative samples is sub-optimal since it introduces potential class imbalance. To overcome this challenge, we first introduce a data augmentation strategy named TypeMix. By fusing unlabeled samples with negative samples, the network can better benefit from diverse features provided by negatives while alleviating the potential side effects. Furthermore, it is observed that the number of negative samples significantly exceeds that of lesion samples. To reduce redundancy and improve training efficiency, we propose a dynamic informativeness-aware sampling strategy, prioritizing the active selection of high-valuable negative samples. Extensive experiments on public datasets demonstrate that our simple but effective strategies are enough to consistently outperform other state-of-the-art methods, offering new possibilities for future work from a data collection perspective.
Xinyu Xiong, Wenxue Li 0003, Duojun Huang
IEEE Signal Process. Lett.2
2024 TP-DRSeg: Improving Diabetic Retinopathy Lesion Segmentation with Explicit Text-Prompts Assisted SAM
Wenxue Li 0003, Xinyu Xiong, Peng Xia 0005, Lie Ju, ZongYuan Ge
MICCAI (8)1
2024 Generalizing to Unseen Domains in Diabetic Retinopathy with Disentangled Representations
Peng Xia 0005, Wenxue Li 0003, Lie Ju, Peibo Duan, Huaxiu Yao, ZongYuan Ge
MICCAI (10)4
2024 HybridVPS: Hybrid-Supervised Video Polyp Segmentation Under Low-Cost Labels
abstract
Deep polyp segmentation methods have shown remarkable potential in boosting diagnostic efficiency. Nevertheless, these methods rely on sufficient pixel-wise annotated data, which is time-consuming and labor-intensive to acquire in clinical practice. This challenge is further escalated under the polyp segmentation scenario due to the massive video frames. To alleviate annotating burden, in this letter, we propose a label-efficient polyp segmentation framework named HybridVPS, which drastically reduces the annotation cost while maintaining satisfactory performance. Our core insight is to take full advantage of the similar semantics between consecutive video frames. Specifically, only a few frames require pixel-wise annotations, while the cheap scribble annotations are enough for the remaining part. To fully leverage the coarse location information provided by scribble annotations, we introduce an adaptive label prompter, which utilizes pixel-wise annotation to provide reliable guidance for scribble-annotated neighboring frames, thus facilitating the overall accuracy of the segmentation. Extensive experiments on the large-scale video polyp dataset SUN-SEG demonstrate the superiority of our approach. HybridVPS achieves comparable performance to the fully supervised scheme while requiring only 2% of the pixel-level annotations.
Wenxue Li 0003, Xinyu Xiong, Fugui Fan
IEEE Signal Process. Lett.1