Yuanchen Wu

dblp:333/8215 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unveiling the complementary synergy of CLIP and diffusion models for weakly supervised semantic segmentation
Hang Yao 0002, Yuanchen Wu, Jide Li, Kequan Yang, Jingxin Han, Xiaoqiang Li 0002
Expert Syst. Appl.2
2026 Disentangling co-occurrence with class-specific banks for Weakly Supervised Semantic Segmentation
Hang Yao 0002, Yuanchen Wu, Kequan Yang, Jide Li, Chao Yin 0001, Xiaoqiang Li 0002
Image Vis. Comput.2
2026 MFDP: Multi-View Feature Integration and Enhanced Disease Prompting for Radiology Report Generation
abstract
Radiology report generation aims to automatically produce diagnostic reports from medical images, reducing radiologists' workload. Most existing models commonly use an encoder-decoder architecture, where the text decoder generates reports based on encoded image tokens. However, these approaches have two major limitations: 1) they always use a single-view feature or simple static fusion multi-view feature, which fails to capture complementary information from multi-view images, and 2) they lack explicit diagnostic information related to the disease during the text decoding process, resulting in reduced clinical accuracy and relevance of the generated report. To deal with the above limitations, this paper proposes a novel framework employing Multi-view Feature Integration and Enhanced Disease Prompting for Radiology Report Generation, called MFDP. Specifically, MFDP introduces two key innovations:1) the Multi-view Feature Fusion (MFF) module is designed to dynamically integrate multi-view images (e.g., frontal and lateral views) through a multi-view attention mechanism that adaptively captures inter-view dependencies, enriching the decoder's input features to generate more comprehensive reports. 2) the Enhanced Disease Prompting (EDP) module is designed to provide explicit diagnostic information by constructing enhanced disease prompts to guide the text decoding process. Experiments on two benchmark datasets, MIMIC-CXR and IU X-Ray, demonstrate that the proposed MFDP is competitive in both Clinical Efficacy (CE) and Natural Language Generation (NLG) metrics. Notably, MFDP achieves a 10% average improvement in CE Recall compared to SOTA models, enabling more precise localization of critical abnormalities while maintaining diagnostic completeness.
Yongxu Zhao, Kequan Yang, Yuanchen Wu, Xiaoqiang Li 0002
IEEE J. Biomed. Health Informatics3
2025 Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
abstract
Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on the models’ response generation, and overlooking the task question itself. This paper discusses the vulnerability of LVLMs in solving counterfactual presupposition questions (CPQs), where the models are prone to accept the presuppositions of counterfactual objects and produce severe hallucinatory responses. To this end, we introduce "Antidote", a unified, synthetic data-driven post-training framework for mitigating both types of hallucination above. It leverages synthetic data to incorporate factual priors into questions to achieve self-correction, and decouple the mitigation process into a preference optimization problem. Furthermore, we construct "CP-Bench", a novel benchmark to evaluate LVLMs’ ability to correctly handle CPQs and produce factual responses. Applied to the LLaVA series, Antidote can simultaneously enhance performance on CP-Bench by over 50%, POPE by 1.8-3.3%, and CHAIR & SHR by 30-50%, all without relying on external supervision from stronger LVLMs or human feedback and introducing noticeable catastrophic forgetting issues.
Yuanchen Wu, Lu Zhang 0060, Hang Yao 0002, Junlong Du, Shouhong Ding, Yunsheng Wu, Xiaoqiang Li 0002
CVPR1
2025 Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
abstract
The rapid development of AI-generated content (AIGC) technology has led to the misuse of highly realistic AI-generated images (AIGI) in spreading misinformation, posing a threat to public information security. Although existing AIGI detection techniques are generally effective, they face two issues: 1) a lack of human-verifiable explanations, and 2) a lack of generalization in the latest generation technology. To address these issues, we introduce a large-scale and comprehensive dataset, Holmes-Set, which includes the Holmes-SFTSet, an instruction-tuning dataset with explanations on whether images are AI-generated, and the Holmes-DPOSet, a human-aligned preference dataset. Our work introduces an efficient data annotation method called the Multi-Expert Jury, enhancing data generation through structured MLLM explanations and quality control via cross-model evaluation, expert defect filtering, and human preference modification. In addition, we propose Holmes Pipeline, a meticulously designed three-stage training framework comprising visual expert pre-training, supervised fine-tuning, and direct preference optimization. Holmes Pipeline adapts multimodal large language models (MLLMs) for AIGI detection while generating human-verifiable and human-aligned explanations, ultimately yielding our model AIGI-Holmes. During the inference stage, we introduce a collaborative decoding strategy that integrates the model perception of the visual expert with the semantic reasoning of MLLMs, further enhancing the generalization capabilities. Extensive experiments on three benchmarks validate the effectiveness of our AIGI-Holmes.
Ziyin Zhou, Yunpeng Luo, Yuanchen Wu, Ke Sun 0016, Jiayi Ji, Shouhong Ding, Xiaoshuai Sun, Yunsheng Wu, Rongrong Ji
ICCV3
2025 ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision Experts
abstract
Vision-language (VL) learning requires extensive visual perception capabilities, such as fine-grained object recognition and spatial perception. Recent works typically rely on training huge models on massive datasets to develop these capabilities. As a more efficient alternative, this paper proposes a new framework that Transfers the knowledge from a hub of Vision Experts (ToVE) for efficient VL learning, leveraging pre-trained vision expert models to promote visual perception capability. Specifically, building on a frozen CLIP image encoder that provides vision tokens for image-conditioned language generation, ToVE introduces a hub of multiple vision experts and a token-aware gating network that dynamically routes expert knowledge to vision tokens. In the transfer phase, we propose a "residual knowledge transfer" strategy, which not only preserves the generalizability of the vision tokens but also allows selective detachment of low-contributing experts to improve inference efficiency. Further, we explore to merge these expert knowledge to a single CLIP encoder, creating a knowledge-merged CLIP that produces more informative vision tokens without expert inference during deployment. Experiment results across various VL tasks demonstrate that the proposed ToVE achieves competitive performance with two orders of magnitude fewer training data.
Yuanchen Wu, Junlong Du, Shouhong Ding, Xiaoqiang Li 0002
ICLR1
2025 Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
abstract
Large Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces Self-Rationale Calibration (SRC) framework to iteratively calibrate the alignment between rationales and answers. SRC begins by employing a lightweight “rationale fine-tuning” approach, which modifies the model’s response format to require a rationale before deriving answer without explicit prompts. Next, SRC searches a diverse set of candidate responses from the fine-tuned LVLMs for each sample, followed by a proposed pairwise scoring strategy using a tailored scoring model, R-Scorer, to evaluate both rationale quality and factual consistency of candidates. Based on a confidence-weighted preference curation process, SRC decouples the alignment calibration into a preference fine-tuning manner, leading to significant improvements of LVLMs in perception, reasoning, and generalization across multiple benchmarks. Our results emphasize the rationale-oriented alignment in exploring the potential of LVLMs.
Yuanchen Wu, Shouhong Ding, Ziyin Zhou, Xiaoqiang Li 0002
ICML1
2025 See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual understanding and multimodal reasoning. However, LVLMs frequently exhibit hallucination phenomena, manifesting as the generated textual responses that demonstrate inconsistencies with the provided visual content. Existing hallucination mitigation methods are predominantly text-centric, the challenges of visual-semantic alignment significantly limit their effectiveness, especially when confronted with fine-grained visual understanding scenarios. To this end, this paper presents ViHallu, a Vision-Centric Hallucination mitigation framework that enhances visual-semantic alignment through Visual Variation Image Generation and Visual Instruction Construction. ViHallu introduces visual variation images with controllable visual alterations while maintaining the overall image structure. These images, combined with carefully constructed visual instructions, enable LVLMs to better understand fine-grained visual content through fine-tuning, allowing models to more precisely capture the correspondence between visual content and text, thereby enhancing visual-semantic alignment. Extensive experiments on multiple benchmarks show that ViHallu effectively enhances models' fine-grained visual understanding while significantly reducing hallucination tendencies. Furthermore, we release ViHallu-Instruction, a visual instruction dataset specifically designed for hallucination mitigation and visual-semantic alignment. Code is available at https://github.com/oliviadzy/ViHallu.
Ziyun Dai, Xiaoqiang Li 0002, Yuanchen Wu, Jide Li
ACM Multimedia4
2025 Mutual learning with discrepancy for weakly supervised object detection
Kequan Yang, Xichen Ye, Yuanchen Wu, Jide Li, Xiaoqiang Li 0002, Pinpin Zhu
Expert Syst. Appl.3
2025 PrioMatch: Semi-supervised learning guided by prior knowledge
Jiaquan Wang, Yuqing Zou, Yuanchen Wu, Xiaoqiang Li 0002
Neurocomputing3
2025 ProxyMatting: Transformer-based image matting via region proxy
Jide Li, Kequan Yang, Yuanchen Wu, Xichen Ye, Hanqi Yang, Xiaoqiang Li 0002
Knowl. Based Syst.3
2025 Pseudo-label enhancement for weakly supervised object detection using self-supervised vision transformer
Kequan Yang, Yuanchen Wu, Jide Li, Chao Yin 0001, Xiaoqiang Li 0002
Knowl. Based Syst.2
2025 MBS: A High-Precision Approximation Method for Softmax and Efficient Hardware Implementation
abstract
The softmax function needs to be frequently used in the multi-head attention layer of Transformer networks. Compared to DNNs and other networks, Transformers have higher computational complexity, requiring higher accuracy and hardware performance for softmax function calculations. Therefore, we propose mixed-base softmax (MBS) for the first time for the approximation of the softmax function. This method combines exponential functions with bases of 2 and 4, which is advantageous for hardware implementation. MBS has a high similarity to the softmax function and demonstrates advanced performance during inference in Transformer network. Through algorithm transformation and hardware optimization, we have designed a low-complexity and highly parallel hardware architecture, which only occupies few additional hardware resources compared to base-2 softmax but achieves higher accuracy. Experimental results show that, under TSMC 90nm CMOS technology at the frequency of 0.5 GHz, our design can achieve the efficiency of 236.18 Gps/(mm2⋅mW) with the area of 4234 μm2. Furthermore, MBS exhibits higher computational accuracy and inference precision compared with base-2 softmax.
Yuanchen Wu, Zhiheng Xie, Hongbing Pan
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
abstract
Recently, One-stage Weakly Supervised Semantic Segmentation (WSSS) with image-level labels has gained increasing interest due to simplification over its cumbersome multi-stage counterpart. Limited by the inherent ambiguity of Class Activation Map (CAM), we observe that one-stage pipelines often encounter confirmation bias caused by incorrect CAM pseudo-labels, impairing their final segmentation performance. Although recent works discard many unreliable pseudo-labels to implicitly alleviate this issue, they fail to exploit sufficient supervision for their models. To this end, we propose a dual student framework with trustworthy progressive learning (DuPL). Specifically, we propose a dual student network with a discrepancy loss to yield diverse CAMs for each sub-net. The two sub-nets generate supervision for each other, mitigating the confirmation bias caused by learning their own incorrect pseudo-labels. In this process, we progressively introduce more trustworthy pseudo-labels to be involved in the supervision through dynamic threshold adjustment with an adaptive noise filtering strategy. Moreover, we believe that every pixel, even discarded from supervision due to its unreliability, is important for WSSS. Thus, we develop consistency regularization on these discarded regions, providing supervision of every pixel. Experiment results demonstrate the superiority of the proposed DuPL over the recent state-of-the-art alternatives on PASCAL VOC 2012 and MS COCO datasets. Code is available at https://github.com/Wu0409/DuPL.
Yuanchen Wu, Xichen Ye, Kequan Yang, Jide Li, Xiaoqiang Li 0002
CVPR1
2024 DINO is Also a Semantic Guider: Exploiting Class-aware Affinity for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) using image-level labels is a challenging task, with relying on Class Activation Map (CAM) to derive segmentation supervision. Although many efficient single-stage solutions have been proposed, their performance is hindered by the inherent ambiguity of CAM. This paper introduces a new approach, dubbed ECA, to Exploit the self-supervised Vision Transformer, DINO, inducing the Class-aware semantic Affinity to overcome this limitation. Specifically, we introduce a Semantic Affinity Exploitation module (SAE). It establishes the class-agnostic affinity graph through the self-attention of DINO. Using the highly activated patches on CAMs as 'seeds', we propagate them across the affinity graph and yield the Class-aware Affinity Region Map (CARM) as supplementary semantic guidance. Moreover, the selection of reliable 'seeds' is crucial to the CARM generation. Inspired by the observed CAM inconsistency between the global and local views, we develop a CAM Correspondence Enhancement module (CCE) to encourage dense local-to-global CAM correspondences, advancing high-fidelity CAM for seed selection in SAE. Our experimental results demonstrate that ECA effectively improves the model's object pattern understanding. Remarkably, it outperforms state-of-the-art alternatives on the PASCAL VOC 2012 and MS COCO 2014 datasets, achieving 90.1% upper bound performance compared to its fully supervised counterpart. Code is available at https://github.com/Wu0409/ECA.
Yuanchen Wu, Xiaoqiang Li 0002, Jide Li, Kequan Yang, Pinpin Zhu
ACM Multimedia1
2024 Robust Offline Active Learning on Graphs
abstract
We consider the problem of active learning on graphs for node-level tasks, which has crucial applications in many real-world networks where labeling node responses is expensive. In this paper, we propose an offline active learning method that selects nodes to query by explicitly incorporating information from both the network structure and node covariates. Building on graph signal recovery theories and the random spectral sparsification technique, the proposed method adopts a two-stage biased sampling strategy that takes both informativeness and representativeness into consideration for node querying. Informativeness refers to the complexity of graph signals that are learnable from the responses of queried nodes, while representativeness refers to the capacity of queried nodes to control generalization errors given noisy node-level information. We establish a theoretical relationship between generalization error and the number of nodes selected by the proposed method. Our theoretical results demonstrate the trade-off between Informativeness and representativeness in active learning. Extensive numerical experiments show that the proposed method is competitive with existing graph-based active learning methods, especially when node covariates and responses contain noises. Additionally, the proposed method is applicable to both regression and classification tasks on graphs.
Yuanchen Wu, Yubai Yuan
NeurIPS1
2024 Uncertainty-aware representation calibration for semi-supervised medical imaging segmentation
Yuanchen Wu, Yue Zhou 0007
Neurocomputing1
2024 Cycle-Consistent Adversarial chest X-rays Domain Adaptation for pneumonia diagnosis
Yue Zhou 0007, Yuanchen Wu
Neurocomputing3
2023 Hierarchical Semantic Contrast for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make CAM fit for WSSS and demonstrate encouraging results. However, they generally exploit single-level semantics, which may hamper the model to learn a comprehensive semantic structure. Motivated by the prior that each image has multiple levels of semantics, we propose hierarchical semantic contrast (HSC) to ameliorate the above problem. It conducts semantic contrast from coarse-grained to fine-grained perspective, including ROI level, class level, and pixel level, making the model learn a better object pattern understanding. To further improve CAM quality, building upon HSC, we explore consistency regularization of cross supervision and develop momentum prototype learning to utilize abundant semantics across different images. Extensive studies manifest that our plug-and-play learning paradigm, HSC, can significantly boost CAM quality on both non-saliency-guided and saliency-guided baselines, and establish new state-of-the-art WSSS performance on PASCAL VOC 2012 dataset. Code is available at https://github.com/Wu0409/HSC_WSSS.
Yuanchen Wu, Xiaoqiang Li 0002, Songmin Dai, Jide Li, Tong Liu 0001, Shaorong Xie
IJCAI1
2023 Semi-supervised medical imaging segmentation with soft pseudo-label fusion
Xiaoqiang Li 0002, Yuanchen Wu, Songmin Dai
Appl. Intell.2
2022 An Attention-Based 3D CNN With Multi-Scale Integration Block for Alzheimer's Disease Classification
abstract
Convolutional Neural Networks (CNNs) have recently been introduced to Alzheimer's Disease (AD) diagnosis. Despite their encouraging prospects, most of the existing models only process AD-related brain atrophy on a single spatial scale, and have high computational complexity. Here, we propose a novel Attention-based 3D Multi-scale CNN model (AMSNet), which can better capture and integrate multiple spatial-scale features of AD, with a concise structure. For the binary classification between 384 AD patients and 389 Cognitively Normal (CN) controls using sMRI scannings, AMSNet achieves remarkable overall performance (91.3% accuracy, 88.3% sensitivity, and 94.2% specificity) with fewer parameters and lower computational load, generally surpassing seven comparative models. Furthermore, AMSNet generalizes well in other AD-related classification tasks, such as the three-way classification (AD-MCI-CN). Our results manifest the feasibility and efficiency of the proposed multi-scale spatial feature integration and attention mechanism used in AMSNet for AD classification, and provide potential biomarkers to explore the neuropathological causes of AD.
Yuanchen Wu, Weiming Zeng, Miao Song 0002
IEEE J. Biomed. Health Informatics1