Yunwei Lan

dblp:311/9108 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 DA-haze: Exploring domain alignment for realistic haze image synthesis
Lanqing Zhang, Yanzhao Su, Zhigao Cui, Nian Wang 0001, Yunwei Lan, Liangyu Zhu
Neurocomputing5
2026 StableV2V: Stabilizing Shape Consistency in Video-to-Video Editing
abstract
Recent advancements in generative artificial intelligence have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. These studies mainly transfer the inherent motion patterns from the source videos to the edited ones, where they often produce inferior results with inconsistency to user intentions, especially when shape changes between the edited and original objects might occur, due to the lack of particular alignments between the delivered motions and edited content. To address this limitation, we present a shape-consistent video editing method, namely StableV2V. Our method decomposes the entire editing pipeline into several sequential procedures, where we first edit the initial video frame, then simulate the shape-aware alignment between the delivered motions and edited sequence, and propagate the edited content to all other frames based on such alignment. Furthermore, we curate a testing benchmark, namely DAVIS-Edit, to offer a comprehensive evaluation of video editing, considering various types of prompts and difficulties. Experimental results and analyses illustrate the superior performance, visual consistency, and inference efficiency of our proposed method compared to existing state-of-the-art video editing studies.
Chang Liu 0165, Kaidong Zhang, Yunwei Lan, Dong Liu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 Weakly Supervised Image Dehazing via Physics-Based Decomposition
abstract
Recent weakly supervised image dehazing (WSID) works have succeeded to improve models’ generalization ability to real scene dehazing by using generative adversarial network (GAN) for unpaired image training. However, it is still difficult for current WSID methods to train one effective dehazing model for various scenes since 1) they always result in residual haze due to insufficient generalization to the feature distribution of real scenes, and 2) they are prone to cause distortions like color shifts, artifacts or halos etc, owing to embedding manual prior or threshold hypothesis for image reconstruction. To solve above problems, in this paper, we propose a novel WSID model via physics-based decomposition (PBD), which estimates atmospheric light, scattering coefficient and scene depth of real haze input to effectively capture the illumination information and haze distribution to recover a preliminary dehazed image by minimizing reconstruction loss. With this constraint, we subtly design a discrete wavelet discriminator (DWD) to effectively improve the generalization to real scene from both spatial and frequency aspect under the supervision of unpaired real clear image. Our PBD is a purely data-driven model freeing from any manual setting or partially correct prior, thus simultaneously ensuring the realness and visibility of dehazed images. Experiments on seven benchmarks verified the strong generalization ability of our PBD, which achieves SOTA dehazing performance with realistic details. Code will be published at https://github.com/NianWang-HJJGCDX/PBD.
Nian Wang 0001, Zhigao Cui, Yanzhao Su, Yunwei Lan, Yuanliang Xue, Aihua Li
IEEE Trans. Circuits Syst. Video Technol.4
2026 LaCon: Late-Constraint Controllable Visual Generation
abstract
Diffusion models have demonstrated impressive abilities in generating photo-realistic and creative images. To offer more controllability for the generation process of diffusion models, previous studies normally adopt extra modules to integrate condition signals by manipulating the intermediate features of the noise predictors, where they often fail in conditions not seen in the training. Although subsequent studies are motivated to handle multi-condition control, they are mostly resource-consuming to implement, where more generalizable and efficient solutions are expected for controllable visual generation. In this paper, we present a late-constraint controllable visual generation method, namely LaCon, which enables generalization across various modalities and granularities for each single-condition control. LaCon establishes an alignment between the external condition and specific diffusion timesteps, and guides diffusion models to produce conditional results based on this built alignment. Experimental results on prevailing benchmark datasets illustrate the promising performance and generalization capability of LaCon under various conditions and settings. Ablation studies analyze different components in LaCon, illustrating its great potential to offer flexible condition controls for different backbones.
Chang Liu 0165, Kaidong Zhang, Yunwei Lan, Dong Liu 0002
IEEE Trans. Image Process.4
2025 Exploiting Diffusion Prior for Real-World Image Dehazing with Unpaired Training
abstract
Unpaired training has been verified as one of the most effective paradigms for real scene dehazing by learning from unpaired real-world hazy and clear images. Although numerous studies have been proposed, current methods demonstrate limited generalization for various real scenes due to limited feature representation and insufficient use of real-world prior. Inspired by the strong generative capabilities of diffusion models in producing both hazy and clear images, we exploit diffusion prior for real-world image dehazing, and propose an unpaired framework named Diff-Dehazer. Specifically, we leverage diffusion prior as bijective mapping learners within the CycleGAN, a classic unpaired learning framework. Considering that physical priors contain pivotal statistics information of real-world data, we further excavate real-world knowledge by integrating physical priors into our framework. Furthermore, we introduce a new perspective for adequately leveraging the representation ability of diffusion models by removing degradation in image and text modalities, so as to improve the dehazing effect. Extensive experiments on multiple real-world datasets demonstrate the superior performance of our method.
Yunwei Lan, Zhigao Cui, Chang Liu 0165, Jialun Peng, Nian Wang 0001, Dong Liu 0002
AAAI1
2025 When Schrödinger Bridge Meets Real-World Image Dehazing with Unpaired Training
abstract
Recent advancements in unpaired dehazing, particularly those using GANs, show promising performance in processing real-world hazy images. However, these methods tend to face limitations due to the generator's limited transport mapping capability, which hinders the full exploitation of their effectiveness in unpaired training paradigms. To address these challenges, we propose DehazeSB, a novel unpaired dehazing framework based on the Schrödinger Bridge. By leveraging optimal transport (OT) theory, DehazeSB directly bridges the distributions between hazy and clear images. This enables optimal transport mappings from hazy to clear images in fewer steps, thereby generating high-quality results. To ensure the consistency of structural information and details in the restored images, we introduce detail-preserving regularization, which enforces pixel-level alignment between hazy inputs and dehazed outputs. Furthermore, we propose a novel prompt learning to leverage pre-trained CLIP models in distinguishing hazy images and clear ones, by learning a haze-aware vision-language alignment. Extensive experiments on multiple real-world datasets demonstrate our method's superiority. Code: https://github.com/ywxjm/DehazeSB.
Yunwei Lan, Zhigao Cui, Chang Liu 0165, Nian Wang 0001, Menglin Zhang, Yanzhao Su, Dong Liu 0002
ICCV1
2025 Efficient Perceptual Video Super-Resolution via One-Step Diffusion Denoising
abstract
This paper addresses the inherent trade-off between perceptual quality and temporal consistency in one-step denoising frameworks for video super-resolution (VSR). We propose a novel diffusion-based approach that integrates SDTurbo as its one-step denoising prior, enhanced with two specialized components for inter-frame interaction: temporal attention adapter and bidirectional recurrent propagation layer operating in the latent domain. We leverage LoRA fine-tuning, enabling significantly reduced computational overhead while maintaining high perceptual quality. This design achieves an optimal balance between perceptual quality and temporal consistency. Evaluated on benchmark datasets including REDS, UDM10 and Vimeo, our method delivers perceptual quality comparable to leading method while requiring significantly fewer parameters and achieving faster inference speeds. Supplementary material and demo videos are available at https://github.com/wym1233/OneStepVSR.
Yunwei Lan
VCIP2
2025 AFE-Dehaze: Image Dehazing Method Based on Adaptive Feature Enhancement Contrastive Learning
abstract
ABSTRACT To address the issue of recovery imbalance caused by the spatial heterogeneity of haze concentration in real scenarios, this paper proposes an adaptive feature enhanced contrastive learning framework (AFE‐Dehaze). The framework achieves breakthroughs through three major collaborative mechanisms: (1) a hierarchical multi‐scale fusion architecture that combines diffusion convolution and channel attention, preserving edge textures (such as leaf veins and building contours) in thin haze areas, while semantically guiding the reconstruction of structural details in dense haze areas, improving texture retention by 18% compared to traditional U‐Net; (2) a concentration‐sensitive contrastive learning paradigm that uses pre‐trained VGG features as semantic anchors, applying pixel‐level constraints in thin haze and feature space constraints in dense haze, which reduces color distortion () by 23%, significantly outperforming methods like refusion; (3) a gradient dynamic balancing strategy that automatically adjusts the optimization direction by analyzing positive and negative sample gradient contributions, enhancing PSNR by 1.2dB and SSIM by 0.05 in non‐uniform haze scenarios. Experiments on a mixed dataset (RESIDE OTS real scenes) demonstrate that AFE‐Dehaze achieves an average PSNR of 28.7dB and SSIM of 0.91, especially improving structural similarity in dense haze areas by 9% compared to Mamba, validating its generalization capability in complex haze environments. This framework provides a solution that balances accuracy and robustness for dehazing in real scenarios such as vehicular vision and remote sensing imaging.
Lanqing Zhang, Zhigao Cui, Yanzhao Su, Nian Wang 0001, Yunwei Lan, Liangyu Zhu
IET Image Process.5
2025 Dual representation modeling and progressive contrastive learning for unsupervised video person re-identification
Yanzhao Su, Nian Wang 0001, Yunwei Lan, Aihua Li
Neurocomputing4
2022 Multi-priors Guided Dehazing Network Based on Knowledge Distillation
Nian Wang 0001, Zhigao Cui, Aihua Li, Yanzhao Su, Yunwei Lan
PRCV (4)5
2021 Prior-guided multiscale network for single-image dehazing
abstract
Abstract Single‐image dehazing is an important problem because it is a key prerequisite for most high‐level computer vision tasks. Traditional prior‐based methods adopt priors generated from clear images to restrain the atmospheric scattering model and then recover haze‐free images. However, these prior‐based methods always encounter over‐enhancement, such as halos and colour distortion. To solve this problem, many works use a convolutional neural network to retrieve original images. However, without priors as guidance, these learning‐based methods dehaze effectively in synthetic datasets but perform poorly in real scenes. Hence, in this paper, we propose a prior‐guided multiscale network for single‐image dehazing named PGMNet. Specifically, prior‐based methods are adopted to acquire dehazed images of the training dataset in advance and then send these dehazed images to a parameter‐shared encoder to form multiscale features. During the decoding process, these multiscale features are adopted to guide the prior‐guided multiscale network to recover more image details. Moreover, considering that these prior‐based dehazed images usually contain some over‐enhanced regions, a spatial attention guided feature aggregation module and squeeze‐and‐excitation module are adopted to alleviate colour distortion. The proposed PGMNet takes the advantage of prior‐based methods in real haze removal and provides superior performance compared with the state‐of‐the‐art methods on both synthetic and real‐world datasets.
Nian Wang 0001, Zhigao Cui, Yanzhao Su, Chuan He 0003, Yunwei Lan, Aihua Li
IET Image Process.5