Yuhui Wu 0001

dblp:41/915-1 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0007-9492-7649ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 InsViE-1M: Effective Instruction-Based Video Editing with Elaborate Dataset Construction
abstract
Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video, instruction) is a challenging task. Existing datasets mostly consist of low-resolution, short duration, and limited amount of source videos with unsatisfactory editing quality, limiting the performance of trained editing models. In this work, we present a high-quality Instruction-based Video Editing dataset with 1M triplets, namely InsViE-1M. We first curate high-resolution and high-quality source videos and images, then design an effective editing-filtering pipeline to construct high-quality editing triplets for model training. For a source video, we generate multiple edited samples of its first frame with different intensities of classifier-free guidance, which are automatically filtered by GPT-4o with carefully crafted guidelines. The edited first frame is propagated to subsequent frames to produce the edited video, followed by another round of filtering for frame quality and motion evaluation. We also generate and filter a variety of video editing triplets from high-quality images. With the InsViE-1M dataset, we propose a multi-stage learning strategy to train our InsViE model, progressively enhancing its instruction following and editing ability. Extensive experiments demonstrate the advantages of our InsViE-1M dataset and the trained model over state-of-the-art works. Codes are available at \href{https://github.com/langmanbusi/InsViE}{InsViE}.
Yuhui Wu 0001, Liyi Chen 0002, Ruibin Li, Lei Zhang 0006
ICCV1
2025 Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu 0009, Rongyuan Wu, Lingchen Sun, Yuhui Wu 0001, Lei Zhang 0006
ICCV5
2025 DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
abstract
Leveraging the powerful generation capability of large-scale pretrained text-to-image models, training-free methods have demonstrated impressive image editing results. Conventional diffusion-based methods, as well as recent rectified flow (RF)-based methods, typically reverse synthesis trajectories by gradually adding noise to clean images, during which the noisy latent at the current timestep is used to approximate that at the next timesteps, introducing accumulated drift and degrading reconstruction accuracy. Considering the fact that in RF the noisy latent is estimated through direct interpolation between Gaussian noises and clean images at each timestep, we propose Direct Noise Alignment (DNA), which directly refines the desired Gaussian noise in the noise domain, significantly reducing the error accumulation in previous methods. Specifically, DNA estimates the velocity field of the interpolated noised latent at each timestep and adjusts the Gaussian noise by computing the difference between the predicted and expected velocity field. We validate the effectiveness of DNA and reveal its relationship with existing RF-based inversion methods. Additionally, we introduce a Mobile Velocity Guidance (MVG) to control the target prompt-guided generation process, balancing image background preservation and target object editability. DNA and MVG collectively constitute our proposed method, namely DNAEdit. Finally, we introduce DNA-Bench, a long-prompt benchmark, to evaluate the performance of advanced image editing models. Experimental results demonstrate that our DNAEdit achieves superior performance to state-of-the-art text-guided editing methods. Our code, model, and benchmark will be made publicly available.
Minghan Li 0001, Shuai Li 0014, Yuhui Wu 0001, Qiaosi Yi, Lei Zhang 0006
NeurIPS4
2025 Toward Generalized and Realistic Unpaired Image Dehazing via Region-Aware Physical Constraints
abstract
Supervised dehazing models, trained on synthetic hazy-clean image pairs, often face a notable decline in performance when applied to real-world scenes. Consequently, CycleGAN-based unpaired dehazing methods are proposed to improve the model’s generalization. One successful approach among these methods involves decomposing the physical properties of the atmospheric scattering model (ASM). However, estimating physical properties individually from input images is difficult without supervised labels, which ignores the semantic consistency between different physical regions. We claim semantic region information can offer additional geometric spatial constraints for estimating physical properties, as natural images can be divided into regions with similar scene depths. Motivated by this, we propose a novel generalized and realistic unpaired image dehazing framework via region-aware physical constraints (RPC-Dehaze). Our approach utilizes fine-grained semantic region maps from the Segment Anything Model (SAM) in a specially designed region prompt enhancement module. This enables the dehazing and hazing cyclic networks to learn region-aware physical constraints, leading to accurate estimation of haze imaging physical properties. In contrast to existing unpaired methods that treat dehazing and hazing networks equally, we incorporate Retinex theory into the hazing network, allowing it to learn diverse illumination effects in different regions. We adaptively refine the Retinex-based illumination component, resulting in more realistic hazy images. To further facilitate unsupervised learning in our framework, we propose a physical consensual contrastive regularization to ensure compact representation constraints in the latent feature space. Extensive experiments on synthetic and real image datasets show our method surpasses state-of-the-art unpaired dehazing methods in both effectiveness and generalization capability.
Kaihao Lin, Guoqing Wang 0001, Tianyu Li 0003, Yuhui Wu 0001, Chongyi Li, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.4
2024 Domain Prompt Learning Framework for Real Image Dehazing
abstract
Supervised dehazing models trained on synthetic datasets exhibit severe performance degradation in real-world scenarios due to the domain gap. Unsupervised methods are proposed to process real hazy images, while they suffer from the complex training procedure. In this paper, we present a universal domain prompt learning framework (DPLF) for boosting the performance of supervised dehazing models in real scenarios by introducing prompt learning and well-designed domain adapters. We train the learnable text prompts by CLIP feature alignment, which can discriminate between real hazy and clean images, and use these prompts as unsupervised text constraints. Notably, the distribution gap between synthetic and real haze can be regarded as the difference of the haze-relevant style domain. Motivated by this, we design the style domain prompt adapter to align features from synthetic and real haze domains. Extensive experiments on real-world datasets demonstrate significant performance improvement of the baseline dehazing models with our DPLF.
Kaihao Lin, Guoqing Wang 0001, Yuhui Wu 0001, Shuhang Gu, Xing Xu 0001, Yang Yang 0002
ICME3
2024 Cascaded Adversarial Attack: Simultaneously Fooling Rain Removal and Semantic Segmentation Networks
abstract
When applying high-level visual algorithms to rainy scenes, it is customary to preprocess the rainy images using low-level rain removal networks, followed by visual networks to achieve the desired objectives. Such a setting has never been explored by adversarial attack methods, which are only limited to attacking one kind of them. Considering the deficiency of multi-functional attacking strategies and the significance for open-world perception scenarios, we are the first to propose a Cascaded Adversarial Attack (CAA) setting, where the adversarial example can simultaneously attack different-level tasks, such as rain removal and semantic segmentation in an integrated system. Specifically, our attack on the rain removal network aims to preserve rain streaks in the output image, while for the semantic segmentation network, we employ powerful existing adversarial attack methods to induce misclassification of the image content. Importantly, CAA innovatively utilizes binary masks to effectively concentrate the aforementioned two significantly disparate perturbation distributions on the input image, enabling attacks on both networks. Additionally, we propose two variants of CAA, which minimize the differences between the two generated perturbations by introducing a carefully designed perturbation interaction mechanism, resulting in enhanced attack performance. Extensive experiments validate the effectiveness of our methods, demonstrating their superior ability to significantly degrade the performance of the downstream task compared to methods that solely attack a single network.
Zhiwen Wang 0004, Yuhui Wu 0001, Zheng Wang 0044, Jiwei Wei, Tianyu Li 0003, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen
ACM Multimedia2
2024 JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image Enhancement
abstract
Low-light image enhancement (LLIE) has achieved promising performance by employing conditional diffusion models. Despite the success of some conditional methods, previous methods may neglect the importance of a sufficient formulation of task-specific condition strategy, resulting in suboptimal visual outcomes. In this study, we propose JoReS-Diff, a novel approach that incorporates Retinex- and semantic-based priors as the additional pre-processing condition to regulate the generating capabilities of the diffusion model. We first leverage pre-trained decomposition network to generate the Retinex prior, which is updated with better quality by an adjustment network and integrated into a refinement network to implement Retinex-based conditional generation at both feature- and image-levels. Moreover, the semantic prior is extracted from the input image with an off-the-shelf semantic segmentation model and incorporated through semantic attention layers. By treating Retinex- and semantic-based priors as the condition, JoReS-Diff presents a unique perspective for establishing an diffusion model for LLIE and similar image enhancement tasks. Extensive experiments validate the rationality and superiority of our approach.
Yuhui Wu 0001, Guoqing Wang 0001, Zhiwen Wang 0004, Yang Yang 0002, Tianyu Li 0003, Malu Zhang, Chongyi Li, Heng Tao Shen
ACM Multimedia1
2024 Towards a Flexible Semantic Guided Model for Single Image Enhancement and Restoration
abstract
Low-light image enhancement (LLIE) investigates how to improve the brightness of an image captured in illumination-insufficient environments. The majority of existing methods enhance low-light images in a global and uniform manner, without taking into account the semantic information of different regions. Consequently, a network may easily deviate from the original color of local regions. To address this issue, we propose a semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that adaptively integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in the LLIE task. We further present a refined framework SKF++ with two new techniques: (a) Extra convolutional branch for intra-class illumination and color recovery through extracting local information and (b) Equalization-based histogram transformation for contrast enhancement and high dynamic range adjustment. Extensive experiments on various benchmarks of LLIE task and other image processing tasks show that models equipped with the SKF/SKF++ significantly outperform the baselines and our SKF/SKF++ generalizes to different models and scenes well. Besides, the potential benefits of our method in face detection and semantic segmentation in low-light conditions are discussed.
Yuhui Wu 0001, Guoqing Wang 0001, Shaochong Liu, Yang Yang 0002, Wei Liu 0005, Xiongxin Tang, Shuhang Gu, Chongyi Li, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Learning Semantic-Aware Knowledge Guidance for Low-Light Image Enhancement
abstract
Low-light image enhancement (LLIE) investigates how to improve illumination and produce normal-light images. The majority of existing methods improve low-light images via a global and uniform manner, without taking into account the semantic information of different regions. Without semantic priors, a network may easily deviate from a region's original color. To address this issue, we propose a novel semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that wisely integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in LLIE task. Extensive experiments show that models equipped with the SKF significantly outperform the baselines on multiple datasets and our SKF generalizes to different models and scenes well. The code is available at Semantic-Aware-Low-Light-Image-Enhancement.
Yuhui Wu 0001, Guoqing Wang 0001, Yang Yang 0002, Jiwei Wei, Chongyi Li, Heng Tao Shen
CVPR1