Jiawei Wu 0001

dblp:37/10763-1 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6251-2202ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SparseGS-W: Sparse-View Gaussian Splatting for Unconstrained Image Collections With Diffusion Priors
abstract
Synthesizing novel views from unconstrained image collections is an important but challenging task in computer vision. Existing methods, which optimize per-image appearance and transient occlusion through implicit neural networks from dense training views (approximately 1000 images), struggle to perform effectively with sparse inputs, resulting in noticeable artifacts. In this work, we introduce SparseGS-W, a novel framework designed to boost the reconstruction of unconstrained scenes and novel view synthesis using as few as five training images. Motivated by the observation that diffusion prior constrained by limited sparse inputs can remove artifacts through fast and efficient fine-tuning, we propose a plug-and-play Constrained Novel-View Enhancement module to iteratively enhance the quality of rendered novel views. We further present an Occlusion Handling scheme, which flexibly removes occlusions utilizing the inherent inpainting capability of constrained diffusion priors. Both components are capable of extracting appearance features from any user-provided reference image, enabling flexible modeling of illumination-consistent scenes. Extensive experiments demonstrate that SparseGS-W achieves superior performance not only in full-reference metrics, but also in commonly used non-reference metrics such as FID, ClipIQA and MUSIQ.
Jiawei Wu 0001, Yikun Ma, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.3
2026 FreeDehaze: Towards Training-Free Real-World Image Dehazing via Diffusion Degradation Prior
abstract
Restoring high-quality images from degraded hazy images is a challenging task, particularly in real-world scenarios. Recent investigations seek to address this limitation by exploring advanced methods for synthesizing haze and incorporating real-world hazy images. Due to the inherent diversity and complexity of real-world haze, these methods struggle to accurately model haze representations. Based on our observation that the hazy images generated by advanced text-to-image diffusion models exhibit a remarkable resemblance to real-world haze, it suggests that these diffusion models effectively internalize haze representations. Hence, we propose FreeDehaze, a novel training-free diffusion method for real-world image dehazing. FreeDehaze is a posterior-based framework capable of addressing non-linear dehazing challenges without relying on additional degradation estimation networks. It follows the human cognition for image restoration, beginning with perception and subsequently enhancing the image. The core method initially generates pseudo-clean images based on abstract textual descriptions. Subsequently, optimal transport aligns the denoising network output with the pseudo-clean image within a PCA-based haze subspace, facilitating high-fidelity dehazing. Extensive experiments demonstrate that FreeDehaze outperforms comparative methods in subjective metrics on challenging datasets (e.g., RTTS, URHI, and O-HAZE) and achieves competitive objective metrics, demonstrating strong generalization although without additional training.
Jiawei Wu 0001, Yikun Ma, Wenqi Ren, Zhi Jin 0002, Xiaochun Cao
IEEE Trans. Image Process.1
2026 Virtual Consistency Model for All-in-One Image Restoration
abstract
All-in-one Image Restoration (AIR) seeks to address diverse degradations using a unified model trained only once. Existing methods often rely on degradation-specific guidance, leading to conflicting gradients during training. In contrast, diffusion models offer a promising alternative by operating in a high-noise space where diverse degradations exhibit a homogeneous Gaussian distribution. This characteristic alleviates gradient conflicts associated with task-specific degradations. However, existing diffusion-based AIR methods often suffer from a lack of direct supervision in the image space, leading to error accumulation during the iterative denoising process and image fidelity compromisation. This highlights a fundamental dilemma for AIR: the optimal space for modeling degradations is inherently suboptimal for preserving image fidelity. To address this issue, we propose a Virtual Consistency Model for AIR (VCMAIR), which restores images in the high-noise space while employing a novel consistency function to enforce accurate supervision in the image space. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art methods across a comprehensive benchmark of diverse degradation scenarios, including both standard AIR tasks and challenging real-world image restoration tasks.
Jiawei Wu 0001, Luwei Tu, Zhi Jin 0002, Kaihao Zhang, Wenqi Ren, Xiaochun Cao
IEEE Trans. Image Process.1
2026 Pathology-Aligned Contrastive Representation Learning for Gleason Grading
abstract
Gleason grading, the clinical gold standard for prostate cancer assessment, is based on subjective evaluation of glandular architecture, resulting in inter-observer variability and limited scalability. This highlights the need for automated grading systems. However, their development is hindered by the scarcity of annotated pathology data. Self-supervised learning (SSL) presents a promising solution by utilizing unlabeled data. Pathology images differ significantly in their diagnostic content, often characterized by subtle glandular structures rather than broad visual patterns. Existing SSL methods, primarily designed for natural images, struggle to capture these pathology-specific cues, as they typically rely on random masking or global augmentations that overlook fine-grained morphological characteristics. To address this, we propose Pathology-Aligned Contrastive Representation Learning (PA-CRL), which adaptively aligns representations with diagnostically relevant glandular architecture. The core component is a diagnostic-aware masking strategy that selectively emphasizes pathology-relevant regions by constructing pathology-aligned contrastive pairs. In addition, a stability-based regularization mechanism, termed mask-driven entropic label smoothing (MDELS), leverages entropy differences between masked and unmasked views to regularize contrastive supervision, explicitly promoting representation stability under directed masking perturbations. This distinguishes MDELS from existing uncertainty-aware or soft-label contrastive learning approaches that do not model masked-unmasked representation consistency. Experiments on a clinical multiphoton microscopy dataset and two public H&E-stained datasets demonstrate that PA-CRL learns pathology-relevant representations and achieves 2.89% F1-score improvement in downstream Gleason Grade Group evaluation over state-of-the-art methods. Code is available at https://github.com/phenixsp/PA-CRL.
Maoye Huang, Jiawei Wu 0001, Xiaoqin Zhu
IEEE Trans. Medical Imaging3
2025 MotionDiff: Training-Free Zero-Shot Interactive Motion Editing via Flow-Assisted Multi-View Diffusion
Yikun Ma, Jiawei Wu 0001, Zhi Jin 0002
ICCV3
2025 Gradient-Aware Revitalization of Non-Effective Samples in Medical Image Segmentation
abstract
Deep learning-based medical image segmentation, with its precise lesion localization capabilities, serves as a core component in multimedia medical applications and intelligent diagnostic assistance systems. Recent innovations in network architectures significantly improve segmentation performance. However, the Non-Effective Samples (NES) on model optimization receive little attention. These samples are characterized by minimal gradient variations in loss during training and exhibit a slight contribution to model optimization. They encompass well-segmented samples with near-zero loss values and challenging samples with consistently high loss values. Especially, when NES accumulate, the model will fall into an optimization trap, causing the optimization to stagnate. To address this issue, we propose a lightweight plug-and-play Gradient-Aware Sample Selection and Reactivation Strategy (GA-SRS) that efficiently identifies and revitalizes the training potential of NES. Firstly, GA-SRS filters NES out based on the historical training information of samples and the variations in the loss gradient during the training process. Then, GA-SRS revitalizes the training values of these samples through strong data augmentation. Extensive experiments on four public datasets and three general models demonstrate the effectiveness of GA-SRS. For example, GA-SRS helps improve the IoU metric of U-KAN from 67.20% to 71.24% on the BUSI and from 81.20% to 82.51% on the ISIC dataset, achieving state-of-the-art experimental results.
Shiying Lin, Qinghua Lin, Jiawei Wu 0001, Changqing Zhang 0002
ACM Multimedia5
2025 Leukocyte classification using relative-relationship-guided contrastive learning
Qinghua Lin, Jiawei Wu 0001, Taotao Lai, Rongteng Wu, David Zhang 0001
Expert Syst. Appl.3
2025 Fourier-Based Decoupling Network for Joint Low-Light Image Enhancement and Deblurring
abstract
Nighttime handheld photography is often simultaneously affected by low light and blur degradations due to object motion and camera shake. Previous methods typically design specific modules to restore the degradations in the spatial domain independently. However, the interdependence of low light and blur degradations in the spatial domain makes it difficult for these approaches to effectively decouple the degradations, limiting the performance of the designed modules. In this paper, we observe that in the Fourier domain, low light and blur degradations can be represented independently in the amplitude and phase of the image. Through an in-depth analysis of the underlying physical degradation process, we discover that low light degradation exhibits distinct characteristics across different frequency bands in amplitude, while blur degradation is characterized by phase correlation. Leveraging these insights, we mathematically derive a frequency attention mechanism and a filtering mechanism for learning decoupled representations of these degradations, proposing a Fourier-based Decoupling Network for joint low-light image enhancement and deblurring. Experimental results demonstrate that our method achieves the state-of-the-art performance on both synthetic and real-world datasets and exhibits significantly sharper edges. Code is available at https://github.com/Jabruson/FDN-TIP2025.
Luwei Tu, Jiawei Wu 0001, Deyu Meng, Zhi Jin 0002
IEEE Trans. Image Process.2
2024 ASCL: Accelerating semi-supervised learning via contrastive learning
abstract
Summary SSL (semi‐supervised learning) is widely used in machine learning, which leverages labeled and unlabeled data to improve model performance. SSL aims to optimize class mutual information, but noisy pseudo‐labels introduce false class information due to the scarcity of labels. Therefore, these algorithms often need significant training time to refine pseudo‐labels for performance improvement iteratively. To tackle this challenge, we propose a novel plug‐and‐play method named Accelerating semi‐supervised learning via contrastive learning (ASCL). This method combines contrastive learning with uncertainty‐based selection for performance improvement and accelerates the convergence of SSL algorithms. Contrastive learning initially emphasizes the mutual information between samples as a means to decrease dependence on pseudo‐labels. Subsequently, it gradually turns to maximizing the mutual information between classes, aligning with the objective of semi‐supervised learning. Uncertainty‐based selection provides a robust mechanism for acquiring pseudo‐labels. The combination of the contrastive learning module and the uncertainty‐based selection module forms a virtuous cycle to improve the performance of the proposed model. Extensive experiments demonstrate that ASCL outperforms state‐of‐the‐art methods in terms of both convergence efficiency and performance. In the experimental scenario where only one label is assigned per class in the CIFAR‐10 dataset, the application of ASCL to Pseudo‐label, UDA (unsupervised data augmentation for consistency training), and Fixmatch benefits substantial improvements in classification accuracy. Specifically, the results demonstrate notable improvements in respect of 16.32%, 6.9%, and 24.43% when compared to the original outcomes. Moreover, the required training time is reduced by almost 50%.
Haixiong Liu, Jiawei Wu 0001, Wei Zeng 0003
Concurr. Comput. Pract. Exp.3
2024 Information Transfer in Semi-Supervised Semantic Segmentation
abstract
Enhancing the accuracy of dense classification with limited labeled data and abundant unlabeled data, known as semi-supervised semantic segmentation, is an essential task in vision comprehension. Due to the lack of annotation in unlabeled data, additional pseudo-supervised signals, typically pseudo-labeling, are required to improve the performance. Although effective, these methods fail to consider the internal representation of neural networks and the inherent class-imbalance in dense samples. In this work, we propose an information transfer theory, which establishes a theoretical relationship between shallow and deep representations. We further apply this theory at both the semantic and pixel levels, referred to as IIT-SP, to align different types of information. The proposed IIT-SP optimizes shallow representations to match the target representation required for segmentation. This limits the upper bound of deep representations to enhance segmentation performance. We also propose a momentum-based Cluster-State bar that updates class status online, along with a HardClassMix augmentation and a loss weighting technique to address class imbalance issues based on it. The effectiveness of the proposed method is demonstrated through comparative experiments on PASCAL VOC and Cityscapes benchmarks, where the proposed IIT-SP achieves state-of-the-art performance, reaching mIoU of 68.34% with only 2% labeled data on PASCAL VOC and mIoU of 64.20% with only 12.5% labeled data on Cityscapes.
Jiawei Wu 0001, Haoyi Fan, Guanghai Liu 0001, Shouying Lin
IEEE Trans. Circuits Syst. Video Technol.1
2023 dugMatting: Decomposed-Uncertainty-Guided Matting
abstract
Cutting out an object and estimating its opacity mask, known as image matting, is a key task in image and video editing. Due to the highly ill-posed issue, additional inputs, typically user-defined trimaps or scribbles, are usually needed to reduce the uncertainty. Although effective, it is either time consuming or only suitable for experienced users who know where to place the strokes. In this work, we propose a decomposed-uncertainty-guided matting (dugMatting) algorithm, which explores the explicitly decomposed uncertainties to efficiently and effectively improve the results. Basing on the characteristic of these uncertainties, the epistemic uncertainty is reduced in the process of guiding interaction (which introduces prior knowledge), while the aleatoric uncertainty is reduced in modeling data distribution (which introduces statistics for both data and possible noise). The proposed matting framework relieves the requirement for users to determine the interaction areas by using simple and efficient labeling. Extensively quantitative and qualitative results validate that the proposed method significantly improves the original matting algorithms in terms of both efficiency and efficacy.
Jiawei Wu 0001, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou
ICML1
2022 Foreground-background decoupling matting
abstract
Image matting aims to extract specific objects, deployed in many applications. Generally, the automatic matting methods need an extra before overcome the intricate details and the diverse appearances. Recently, the matting community has paid more attentions to the investigation of trimap-free matting direction to address the dependency of priors. Most trimap-free approaches divide the matting task into global segmentation and detail matting subtasks. Unfortunately, these methods suffer from stagewise modeling, uncorrectable errors, or subtasks bottleneck problems. To address these issues, we propose a new set of matting subtasks, including foreground segmentation, background segmentation, and disambiguation. And we present a novel Foreground–Background Decoupling Matting (FBDM) network motivated by the new subtasks. Specifically, we first design a nested attention mechanism to decouple the backbone features. Then, we utilize two independent progressive semantic decoders by the decoupling features to complete the foreground and background segmentation subtasks. Finally, we utilize multiple of the proposed frequency division local disambiguation modules to achieve the disambiguation subtask. Besides, we establish a challenging potted plant (PPT) benchmark which contains 100 potted plants images in the real world for the matting community. Extensive experiments on several public benchmarks and the PPTs benchmark demonstrate that the proposed FBDM generates the best results compared with the state-of-the-art trimap-free methods.
Jiawei Wu 0001, Guolin Zheng, Haoyi Fan
Int. J. Intell. Syst.1
2021 Semi-Supervised Semantic Segmentation via Entropy Minimization
abstract
In this paper, we propose a novel entropy minimization based semi-supervised method for semantic segmentation. Entropy minimization has proven to be an effective semi-supervised method for realizing the cluster assumption, where the decision boundary should lie in low-density regions. Inspired by the existing consistency training semi-supervised segmentation networks with encoder-decoder architecture, we found that there tend to be more large gradient values at the object edges than other positions in the feature map of the encoder, and therefore propose a feature gradient map regularization to enlarge inter-class distance in the feature space for low-entropy of segmentation prediction. Additionally, we introduce an adaptive sharpening scheme with aleatoric uncertainty, and a class consistency constraint regularization, to alleviate the interference of noise with pseudo labels. Extensive experiments on PASCAL VOC, PASCAL-Context, and Leukocyte datasets show that the proposed method achieves state-of-the-art semi-supervised semantic segmentation performance without almost additional calculations and network structures.
Jiawei Wu 0001, Haoyi Fan, Shouying Lin
ICME1