Zicheng Wang 0012

dblp:24/10098-12 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-8351-0329ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels
abstract
Medical Visual Question Answering (Med-VQA) systems benefit the interpretation of medical images containing critical clinical information. However, the challenge of noisy labels and limited high-quality datasets remains underexplored. To address this, we establish the first benchmark for noisy labels in Med-VQA by simulating human mislabeling with semantically designed noise types. More importantly, we introduce the DiN framework, which leverages a diffusion model to handle noisy labels in Med-VQA. Unlike the dominant classification-based VQA approaches that directly predict answers, our Answer Diffuser (AD) module employs a coarse-to-fine process, refining answer candidates with a diffusion model for improved accuracy. The Answer Condition Generator (ACG) further enhances this process by generating task-specific conditional information via integrating answer embeddings with fused image-question features. To address label noise, our Noisy Label Refinement(NLR) module introduces a robust loss function and dynamic answer adjustment to further boost the performance of the AD module. Our DiN framework consistently outperforms existing methods across multiple benchmarks with varying noise levels1.
Erjian Guo, Zhen Zhao 0001, Zicheng Wang 0012, Tong Chen 0011, Yunyi Liu, Luping Zhou
CVPR3
2025 Imbalanced Medical Image Segmentation With Pixel-Dependent Noisy Labels
abstract
Accurate medical image segmentation is often hindered by noisy labels in training data, due to the challenges of annotating medical images. Prior research works addressing noisy labels tend to make class-dependent assumptions, overlooking the pixel-dependent nature of most noisy labels. Furthermore, existing methods typically apply fixed thresholds to filter out noisy labels, risking the removal of minority classes and consequently degrading segmentation performance. To bridge these gaps, our proposed framework, Collaborative Learning with Curriculum Selection (CLCS), addresses pixel-dependent noisy labels with class imbalance. CLCS advances the existing works by i) treating noisy labels as pixel-dependent and addressing them through a collaborative learning framework, and ii) employing a curriculum dynamic thresholding approach adapting to model learning progress to select clean data samples to mitigate the class imbalance issue, and iii) applying a noise balance loss to noisy data samples to improve data utilization instead of discarding them outright. Specifically, our CLCS contains two modules: Curriculum Noisy Label Sample Selection (CNS) and Noise Balance Loss (NBL). In the CNS module, we designed a two-branch network with discrepancy loss for collaborative learning so that different feature representations of the same instance could be extracted from distinct views and used to vote the class probabilities of pixels. Besides, a curriculum dynamic threshold is adopted to select clean-label samples through probability voting. In the NBL module, instead of directly dropping the suspiciously noisy labels, we further adopt a robust loss to leverage such instances to boost the performance. We verify our CLCS on two benchmarks with different types of segmentation noise. Our method can obtain new state-of-the-art performance in different settings, yielding more than 3% Dice and mIoU improvements. Our code is available at https://github.com/Erjian96/CLCS.git.
Erjian Guo, Zicheng Wang 0012, Zhen Zhao 0001, Luping Zhou
IEEE Trans. Medical Imaging2
2025 SOEDiff: Efficient Distillation for Small Object Editing
abstract
In this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion . These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA , which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation , which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID.
Yiming Wu 0005, Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Ronghua Liang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Progressive Classifier and Feature Extractor Adaptation for Unsupervised Domain Adaptation on Point Clouds
Zicheng Wang 0012, Zhen Zhao 0001, Yiming Wu 0005, Luping Zhou, Dong Xu 0001
ECCV (28)1
2024 Alternate Diverse Teaching for Semi-supervised Medical Image Segmentation
Zhen Zhao 0001, Zicheng Wang 0012, Longyue Wang, Dian Yu 0001, Yixuan Yuan, Luping Zhou
ECCV (5)2
2024 Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
abstract
A plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/
Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang
ACM Multimedia3
2023 Conflict-Based Cross-View Consistency for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation (SSS) has recently gained increasing research interest as it can reduce the requirement for large-scale fully-annotated training data. The current methods often suffer from the confirmation bias from the pseudo-labelling process, which can be alleviated by the co-training framework. The current co-training-based SSS methods rely on hand-crafted perturbations to prevent the different sub-nets from collapsing into each other, but these artificial perturbations cannot lead to the optimal solution. In this work, we propose a new conflict-based cross-view consistency (CCVC) method based on a two-branch co-training framework which aims at enforcing the two sub-nets to learn informative features from irrelevant views. In particular, we first propose a new cross-view consistency (CVC) strategy that encourages the two sub-nets to learn distinct features from the same input by introducing a feature discrepancy loss, while these distinct features are expected to generate consistent prediction scores of the input. The CVC strategy helps to prevent the two sub-nets from stepping into the collapse. In addition, we further propose a conflict-based pseudo-labelling (CPL) method to guarantee the model will learn more useful information from conflicting predictions, which will lead to a stable training process. We validate our new CCVC approach on the SSS benchmark datasets where our method achieves new state-of-the-art performance. Our code is available at https://github.com/xiaoyao3302/CCVC.
Zicheng Wang 0012, Zhen Zhao 0001, Xiaoxia Xing, Dong Xu 0001, Luping Zhou
CVPR1
2023 Domain Adaptive Sampling for Cross-Domain Point Cloud Recognition
abstract
Point cloud recognition has recently gained increasing research interest due to the huge potential in real-world applications such as autonomous driving, robotics, etc. However, the point clouds of similar objects often exhibit notable geometric variations due to the difference in capturing devices or environmental changes. This leads to significant performance degradation when the learned point cloud recognition model is applied to a new scenario, which is also known as the domain adaptation issue. In this work, we propose a new unsupervised domain adaptation approach for point cloud recognition via domain adaptive sampling (DAS). In particular, we propose a two-level sampling strategy of point level and instance level to improve the cross-domain recognition ability of the model. First, we propose a domain adaptive point sampling (DAPS) strategy to enhance the domain-invariant representation of point clouds by progressively focusing on representative points in each point cloud based on geometric consistency. Then, we further propose an instance-level domain adaptive cloud sampling (DACS) strategy to learn target-specific information based on a self-paced learning paradigm, where we select a set of pseudo-labeled target point clouds to train our designed light-weighted adapters without modifying the learned domain-invariant representation. We validate our domain adaptive sampling approach on the benchmark datasets PointDA-10 and GraspNetPC-10, where our method achieves new state-of-the-art performance.
Zicheng Wang 0012, Wen Li 0001, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.1