Chunwei Wu

dblp:302/7835 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
abstract
Fengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang, Chunwei Wu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fengqi Zhu, Rongzhen Wang, Shen Nie, Chunwei Wu, Jun Zhou 0011, Yankai Lin 0001, Ji-Rong Wen, Chongxuan Li
ACL (1)5
2025 Dual Knowledge-Aware Guidance for Source-Free Domain Adaptive Fundus Image Segmentation
Chunwei Wu, Guitao Cao
MICCAI (6)3
2025 Uncertainty guided semi-supervised few-shot segmentation with prototype level fusion
Chunwei Wu, Guitao Cao, Wenming Cao 0001
Neural Networks2
2024 HyperEditor: Achieving Both Authenticity and Cross-Domain Capability in Image Editing via Hypernetworks
abstract
Editing real images authentically while also achieving cross-domain editing remains a challenge. Recent studies have focused on converting real images into latent codes and accomplishing image editing by manipulating these codes. However, merely manipulating the latent codes would constrain the edited images to the generator's image domain, hindering the attainment of diverse editing goals. In response, we propose an innovative image editing method called HyperEditor, which utilizes weight factors generated by hypernetworks to reassign the weights of the pre-trained StyleGAN2's generator. Guided by CLIP's cross-modal image-text semantic alignment, this innovative approach enables us to simultaneously accomplish authentic attribute editing and cross-domain style transfer, a capability not realized in previous methods. Additionally, we ascertain that modifying only the weights of specific layers in the generator can yield an equivalent editing result. Therefore, we introduce an adaptive layer selector, enabling our hypernetworks to autonomously identify the layers requiring output weight factors, which can further improve our hypernetworks' efficiency. Extensive experiments on abundant challenging datasets demonstrate the effectiveness of our method.
Chunwei Wu, Guitao Cao, Wenming Cao 0001
AAAI2
2024 CausVSR: Causality Inspired Visual Sentiment Recognition
Xinyue Zhang 0010, Zhaoxia Wang 0001, Jing Xiang, Chunwei Wu, Guitao Cao
IJCAI5
2024 DLE: Document Illumination Correction with Dynamic Light Estimation
abstract
Document images captured through mobile devices in natural environments are often affected by various types of illumination degradation. The degradation diminishes the clarity and readability of document images, thereby complicating their application to OCR downstream tasks. Existing methods typically address only one or a limited number of degradation types and do not consider the diversity of image degradation types. Additionally, these methods typically involve a pre-trained fixed sub-network to estimate background light or shadows, which lacks flexibility and adaptability. To overcome these challenges, this study proposes a novel framework named DLE, which comprises a two-loop generative adversarial network and a multi-modal discriminator. Specifically, to improve the quality of image representation, a mask extractor is embedded before the image input generator. This forces the model to focus on the distinct features in the image, enhancing the representation of illumination anomalous and degraded regions. The mask extractor generates a luminance mask to evaluate the difference in illumination between the input and target images. Subsequently, the consistency loss computation incorporates a dynamic optimization of the mask extractor, strengthening its ability to estimate the illumination degradation part. Moreover, a pre-trained visual-language model is introduced into the multi-modal discriminator, leveraging its robust cross-modal alignment capability to improve the semantic consistency of the generated images with the preset input text. Extensive experiments demonstrate that our approach achieves the SOTA performance in terms of edit distance (ED) and character error rate (CER).
Jiahao Quan, Chunwei Wu, Guitao Cao
SMC3
2024 A Trustworthy Counterfactual Explanation Method With Latent Space Smoothing
abstract
Despite the large-scale adoption of Artificial Intelligence (AI) models in healthcare, there is an urgent need for trustworthy tools to rigorously backtrack the model decisions so that they behave reliably. Counterfactual explanations take a counter-intuitive approach to allow users to explore "what if" scenarios gradually becoming popular in the trustworthy field. However, most previous work on model's counterfactual explanation cannot generate in-distribution attribution credibly, produces adversarial examples, or fails to give a confidence interval for the explanation. Hence, in this paper, we propose a novel approach that generates counterfactuals in locally smooth directed semantic embedding space, and at the same time gives an uncertainty estimate in the counterfactual generation process. Specifically, we identify low-dimensional directed semantic embedding space based on Principal Component Analysis (PCA) applied in differential generative model. Then, we propose latent space smoothing regularization to rectify counterfactual search within in-distribution, such that visually-imperceptible changes are more robust to adversarial perturbations. Moreover, we put forth an uncertainty estimation framework for evaluating counterfactual uncertainty. Extensive experiments on several challenging realistic Chest X-ray and CelebA datasets show that our approach performs consistently well and better than the existing several state-of-the-art baseline approaches.
Yan Li 0063, Chunwei Wu, Xiao Lin 0012, Guitao Cao
IEEE Trans. Image Process.3
2023 Calibrated Uncertainty-Guided Multi-task Framework for Medical Image Segmentation
abstract
Medical image segmentation is a crucial part of computer-aided diagnosis. Due to the enormous cost of labeling medical images, researchers have turned to exploring semi-supervised learning. However, the lack of supervisory information makes it difficult to accurately segment the fuzzy regions (e.g., complex edges or corners of organs). In this paper, we propose a novel method called Multi-task Consistency Segmentation Network based on Calibrated Uncertainty (CU-MCSNet). This model incorporates calibrated uncertainty to guide the network’s learning process. In addition, the model consists of two tasks: i) semantic segmentation as the primary task and ii) signed distance regression as the auxiliary task. To enhance the accuracy of edge segmentation, we propose the Edge Calibration Network for the primary task. This network integrates essential spatial and channel features, employing gradient complementation to hinder the accumulation of defective information and supply pertinent data to fuzzy regions. We also use the inter-task consistency loss to explore the underlying information of the images. In the multi-task domain, it is tough to balance each task manually, and we note that homoscedastic uncertainty focuses on inter-task variation. However, its numerical estimation may still be subject to bias. Therefore, we propose an adaptive loss-balancing strategy based on calibrated homoscedastic uncertainty. Extensive experiments show that our proposed method achieves state-of-the-art performance.
Chunwei Wu, Guitao Cao
BIBM2
2023 DCNet: Weakly Supervised Saliency Guided Dual Coding Network for Visual Sentiment Recognition
abstract
Visual sentiment recognition is a challenging task with scientific significance in probing vision-processing mechanisms. Recent approaches mainly focused on using overall images or precise annotations to learn emotional representations, yet neglected to capture abstract semantics from regional information, or led to a heavy annotation burden. In this paper, we propose an end-to-end weakly supervised framework, called Dual Coding Network (DCNet), which models a dual coding process for both shallow features and high-level regional information. On the one hand, with the help of the fine-grained module (FG), visual features (e.g. texture features) are utilized to enhance the learning of distinguished representation. On the other hand, the DCNet innovatively leverages saliency information to imitate the neural decoding of perceived visual sentiment contents in human brain activity. Specifically, the saliency information guides the generation of sentiment-specific pseudo affective maps (SAMG), which serve as weak annotations. Then the DCNet couples fine-grained features with pseudo affective maps, and obtains semantic vectors for final sentiment prediction. Extensive experiments show that the proposed DCNet outperforms the state-of-the-art performance on five benchmark datasets.
Xinyue Zhang 0010, Jing Xiang, Hanxiu Zhang, Chunwei Wu, Guitao Cao
ECAI4
2023 Making Adversarial Attack Imperceptible in Frequency Domain: A Watermark-based Framework
abstract
With the development of multimedia communication technology, the image information stored in electronic devices faces increasing privacy risks and requires processing for protection. However, it is found that adversarial perturbations added to images for semantic information protection may corrupt the frequency domain watermarks added for copyright statement. With such challenges, we propose an Adversarial Frequency domain Watermarking (AFW) framework to protect images from both copyright and semantic content. Specifically, the AFW framework constructs the images as adversarial examples by embedding crafted adversarial watermarks in the frequency domain, followed by an optimization algorithm to improve the visual quality. Notably, AFW can generally integrate with existing watermark and attack methods. Extensive experiments on five network models and the ImageNet dataset demonstrate that the AFW framework can achieve information hiding and adversarial attacking goals under visual quality assurance.
Hanxiu Zhang, Guitao Cao, Xinyue Zhang 0010, Jing Xiang, Chunwei Wu
ICME5
2023 Class-Balanced Universal Perturbations for Adversarial Training
abstract
Universal attack generates image-agnostic perturbation called universal adversarial perturbation (UAP), which can be added to all samples in the data distribution to fool the classifier. However, a universal perturbation will likely mislead the classifier to identify most adversarial examples as the same label, resulting in the imbalance of attack strength between classes. In this paper, we propose class-balanced UAPs that enlarge the dispersion of the predicted labels for adversarial examples. To ensure attack strength and balance simultaneously, we design a novel diversity objective containing probability calibration and penalty regularizer, which fully considers the predicted label distribution between samples and the predicted probability distribution within samples. Furthermore, we apply class-balanced attacks in adversarial training to defend against universal perturbations since the class-balanced UAP provides diverse perturbation directions. We correspondingly reformulate adversarial training from the min-max optimization problem into a new two-stage framework. Experiments on several benchmark datasets demonstrate that the class-balanced attack achieves better performance than the universal attack, while adversarial training with class-balanced UAP achieves state-of-the-art results in clean accuracy and robustness to universal perturbations.
Kexue Ma, Guitao Cao, Mengqian Xu, Chunwei Wu, Wenming Cao 0001
IJCNN4
2023 Discriminative Feature Mining and Alignment for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain by reducing the cross-domain distribution discrepancy. Existing UDA methods mainly learn domain-invariant features by directly aligning the marginal distribution of source and target domains. However, they ignore mining the discriminative information of target data and aligning the cross-domain discriminative features, which may lead to performance degradation. To tackle these two issues simultaneously, we propose a Discriminative Feature Mining and Alignment (DFMA) algorithm for UDA. Specifically, DFMA advances a three-stage Instance-Class-Pseudo (ICP) strategy consisting of instance-level contrastive learning, class-level contrastive learning and pseudo-labeling methods to mine the discriminative structure of target data. Then we conduct the cross-domain discriminative feature alignment by integrating the adversarial-based method which designs a novel conditional domain discriminator, and the discrepancy-based method which leverages the first-order and second-order statistical information of features. Furthermore, we build a reconstruction network to enhance the class-level feature alignment. Extensive experiments on several standard UDA benchmark datasets validate the superiority of our proposed DFMA.
Jing Xiang, Guitao Cao, Xinyue Zhang 0010, Hanxiu Zhang, Chunwei Wu
IJCNN5
2023 Chaos to Order: A Label Propagation Perspective on Source-Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA), where only a pre-trained source model is used to adapt to the target distribution, is a more general approach to achieving domain adaptation in the real world. However, it can be challenging to capture the inherent structure of the target features accurately due to the lack of supervised information on the target domain. By analyzing the clustering performance of the target features, we show that they still contain core features related to discriminative attributes but lack the collation of semantic information. Inspired by this insight, we present Chaos to Order (CtO), a novel approach for SFDA that strives to constrain semantic credibility and propagate label information among target subpopulations. CtO divides the target data into inner and outlier samples based on the adaptive threshold of the learning state, customizing the learning strategy to fit the data properties best. Specifically, inner samples are utilized for learning intra-class structure thanks to their relatively well-clustered properties. The low-density outlier samples are regularized by input consistency to achieve high accuracy with respect to the ground truth labels. In CtO, by employing different learning strategies to propagate the labels from the inner local to outlier instances, it clusters the global samples from chaos to order. We further adaptively regulate the neighborhood affinity of the inner samples to constrain the local semantic credibility. In theoretical and empirical analyses, we demonstrate that our algorithm not only propagates from inner to outlier but also prevents local clustering from forming spurious clusters. Empirical evidence demonstrates that CtO outperforms the state of the arts on three public benchmarks: Office-31, Office-Home, and VisDA.
Chunwei Wu, Guitao Cao, Yan Li 0063, Xidong Xi, Wenming Cao 0001
ACM Multimedia1
2023 Geometric Contrastive Learning for Heterogeneous Graphs Encoding
abstract
Heterogeneous graphs can represent many network structures in the real world, and research on heterogeneous graph data has attracted more attention. Most existing approaches require additional label information to obtain meaningful node representations. However, labeling heterogeneous graphs is tedious and time-consuming, and low-quality labels will harm the efficiency of models. In addition, the number of nodes increases exponentially with the distance from the root node, but the linearly expanding Euclidean space is difficult to match this growth rate. In this paper, we propose a unified framework that leverages the label irrelevance of contrastive learning and the unique expressive ability of hyperbolic space to encode heterogeneous graphs. Specifically, we use the contrast mechanism to obtain semantic information in an unsupervised way. Meanwhile, we design hyperbolic encoders that are more suitable for graph structure to learn the latent information from heterogeneous graphs efficiently. Experiments on four real-world heterogeneous graph data sets demonstrate the competitive efficacy of the proposed method.
Siheng Wang, Guitao Cao, Chunwei Wu
SMC3
2022 Mix-up Consistent Cross Representations for Data-Efficient Reinforcement Learning
abstract
Deep reinforcement learning (RL) has achieved re-markable performance in sequential decision-making problems. However, it is a challenge for deep RL methods to extract task-relevant semantic information when interacting with limited data from the environment. In this paper, we propose Mix-up Consistent Cross Representations (MCCR), a novel self-supervised auxiliary task, which aims to improve data efficiency and encourage representation prediction. Specifically, we calculate the contrastive loss between low-dimensional and high-dimensional representations of different state observations to boost the mutual information between states, thus improving data efficiency. Furthermore, we employ a mixed strategy to generate intermediate samples, increasing data diversity and the smoothness of representations prediction in nearby timesteps. Experimental results show that MCCR achieves competitive results over the state-of-the-art approaches for complex control tasks in DeepMind Control Suite, notably improving the ability of pretrained encoders to generalize to unseen tasks.
Guitao Cao, Yan Li 0063, Chunwei Wu, Xidong Xi
IJCNN5
2022 Adversarial Discriminative Feature Separation for Generalization in Reinforcement Learning
abstract
Imporving the generalization ability of an agent is an important and challenging task in deep reinforcement learning (RL). Procedually generated environment is an important benchmark for testing generalization in deep RL. In this benchmark, each game consists of multiple levels, each level is an algorithmically created environment instance with a unique configuration of its factors of variation. Existing methods (e.g., regularization, data augmentation) for improving the generalization of RL agent do not learn well the invariant representation among multiple levels. Besides, existing methods for learning invariant representations in RL using adversarial training can only learn invariant information across two levels. To solve this problem, we propose Adversarial Discriminative Feature Separate (ADFS). First, ADFS design a new discriminator for distinguishing whether two observations belong to the same level. Thus, the policy encoder is encouraged to learn invariant information between multiple levels. Second, it separates the representation of observation into level-invariant features and level-discriminative features, so that correction of the optimization direction of the discriminator. The discriminative features are learned by reducing the similarity of specific features intra-levels and increasing that of inter-levels, respectively. Experimental results demonstrate that our method is quite competitive with existing state-of-the-art methods on Procgen Benchmark.
Chunwei Wu, Xidong Xi, Yan Li 0063, Guitao Cao, Wenming Cao 0001
IJCNN2
2021 DCFG: Discovering Directional CounterFactual Generation for Chest X-rays
abstract
While Deep Neural Networks (DNNs) are achieving state-of-the-art performance on medical domains across a variety of tasks, the need for explainability of model predictions in these high-stakes tasks is still lacking. Current for the explainability in model predictions potentially relies on the supervised counterfactual generation that is time-consuming and direction uncontrollable. Yet, the counterfactual generation needs to be easy to implement and have a controllable direction. In light of this trend, we propose an approach for the unsupervised latent direction search of black-box models that are steerable to the user by enabling the user to effectively explore counterfactual generation in a directional way, without relying on domain- or data-specific assumptions. To identify these explainable directions, we use Principal Component Analysis (PCA), a general manifold learning framework to extract low-dimensional subspaces based on a local noise injection of the pre-trained generative model, so that a small perturbation in the subspaces would provide enough change in the resulting data. With experiments on three real-world CXR datasets involving 6 tasks, we find that our approach is capable of learning explainable predictions that discard unrelated confounding factors. Moreover, our method enables practitioners to edit directions to better understand which features are used for predictions.
Yan Li 0063, Chunwei Wu, Xidong Xi, Guitao Cao, Wenming Cao 0001
BIBM3
2021 Debiased Prototype Network for Adversarial Domain Adaptation
abstract
Domain adaptation is an important and challenging task. Existing adversarial domain adaptation methods explore the relationship between the source and target domains, with the knowledge learned in the source domain supporting the target domain task. The quality of the knowledge will affect the task performance of the transfer, i.e., the higher the quality of the knowledge, the better the transfer task performance. To obtain better domain-invariant knowledge, we extract domain-invariant semantic information over the unit sphere via the prototype network. With the help of geometric constraints from the hypersphere, the features can be more tightly clustered with the estimated prototype (representatives of each class). Adaptation is achieved by adversarial learning to align the domain distribution, which enhances the transferability of the learned features and obtains the basic prototype. Since the basic prototypes dominantly computed from the source domain are biased against the expected domain-invariant prototype, a debiased method is further proposed to obtain the domain-invariant prototypes. Specifically, our method diminishes the intra- and inter- class bias to achieve the class-level alignment. Extensive experiments demonstrate that our model achieves state-of-the-art performance on several domain adaptation benchmark datasets. Our code is available at https://github.com/Chunweiwu-source/DPN.
Chunwei Wu, Guitao Cao, Wenming Cao 0001
IJCNN1