VLDB 2026 Research / reviewers in the wild / expert
Junkai Fan
dblp:316/6492
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0000-8162-6280ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-Guided Posterior Sampling for Diffusion-Based Real-World Dehazing and Image EnhancementabstractReal-world image dehazing is a challenging task due to the collection of aligned hazy/clear image pairs under unpredictable and complex environments. To address this limitation, we propose a Physical-Guided Posterior Sampling (PGPS) method that designs a dehazing reconstruction posterior to sample an RGB and depth from pre-trained unconditional diffusion generation process. First, we introduce a Hybrid Degradation Atmospheric Scattering Model (HD-ASM) to adapt the diffusion model, enabling the generation of high-fidelity dehazed images from posterior samples without relying on the aligned hazy/clear image pairs. Second, we propose a two-stage sampling strategy with piecewise loss to improve sampling quality and stability, along with a post-processing technique to remove JPEG compression artifacts amplified by dehazing. Extensive experiments show that our method outperforms state-of-the-art techniques in image dehazing, and in the RTTS dataset’s complex human-vehicle environment. Additionally, our approach also surpasses other benchmarks in object detection, exhibiting superior generalization performance. Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Jianjun Qian, Heyou Chang, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | DVDPEC: Driving-Video Dehazing via Position Embedding-Based CodebookabstractDespite significant progress in real-world image dehazing, efficiently generating high-fidelity, haze-free videos (especially in driving scenarios) remains challenging. Existing methods generally extend image dehazing techniques to videos by employing pre-trained single image dehazing models for preprocessing followed by refinement stages. However, this disjointed two-stage process often leads to unrealistic textures and loss of detail, as it fails to leverage large amounts of high-quality images for prior learning and the subsequent refinement struggles to correct temporal inconsistencies across frames introduced in the first stage. To address these issues, we propose DVDPEC: a Driving Video Dehazing framework utilizing a Position Embedding-based (PE-based) Codebook and a novel Flow Selective Block (FSB). The PE-based codebook stores fine-grained, spatially aware textural information specific to driving videos and leverages implicit positional embeddings for precise, position-aware codebook matching. This enables accurate prior retrieval and improves dehazing results. The FSB aggregates information from adjacent frames by dynamically combining both image flow and prior flow, effectively mitigating flow estimation ambiguities caused by haze. It enhances information fusion across frames, leading to more coherent and visually appealing dehazed videos. Extensive experiments demonstrate that DVDPEC achieves state-of-the-art performance on real-world driving video dehazing tasks, significantly enhancing texture preservation and visual fidelity. Yu Zheng 0036, Wenxuan Fang 0001, Xiantao Hu, Junkai Fan, Jiangwei Weng, Jun Li 0027, Kai Zhang 0008, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving VideoabstractIn this paper, we study the challenging problem of simultaneously removing haze and estimating depth from real monocular hazy videos. These tasks are inherently complementary: enhanced depth estimation improves dehazing via the atmospheric scattering model (ASM), while superior dehazing contributes to more accurate depth estimation through the brightness consistency constraint (BCC). To tackle these intertwined tasks, we propose a novel depth-centric learning framework that integrates the ASM model with the BCC constraint. Our key idea is that both ASM and BCC rely on a shared depth estimation network. This network simultaneously exploits adjacent dehazed frames to enhance depth estimation via BCC and uses the refined depth cues to more effectively remove haze through ASM. Additionally, we leverage a non-aligned clear video and its estimated depth to independently regularize the dehazing and depth estimation networks. This is achieved by designing two discriminator networks: D_MFIR enhances high-frequency details in dehazed videos, and D_MDR reduces the occurrence of black holes in low-texture regions. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art techniques in both video dehazing and depth estimation tasks, especially in real-world hazy scenes. Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Xiang Chen 0015, Shangbing Gao, Jun Li 0027, Jian Yang 0003 |
AAAI | 1 |
| 2025 | Guided Real Image Dehazing Using YCbCr Color SpaceabstractImage dehazing, particularly with learning-based methods, has gained significant attention due to its importance in real-world applications. However, relying solely on the RGB color space often fall short, frequently leaving residual haze. This arises from two main issues: the difficulty in obtaining clear textural features from hazy RGB images and the complexity of acquiring real haze/clean image pairs outside controlled environments like smoke-filled scenes. To address these issues, we first propose a novel Structure Guided Dehazing Network (SGDN) that leverages the superior structural properties of YCbCr features over RGB. It comprises two key modules: Bi-Color Guidance Bridge (BGB) and Color Enhancement Module (CEM). BGB integrates a phase integration module and an interactive attention module, utilizing the rich texture features of the YCbCr space to guide the RGB space, thereby recovering clearer features in both frequency and spatial domains. To maintain tonal consistency, CEM further enhances the color perception of RGB features by aggregating YCbCr channel information. Furthermore, for effective supervised learning, we introduce a Real-World Well-Aligned Haze dataset, which includes a diverse range of scenes from various geographical regions and climate conditions. Experimental results demonstrate that our method surpasses existing state-of-the-art methods across multiple real-world smoke/haze datasets. Wenxuan Fang 0001, Junkai Fan, Yu Zheng 0036, Jiangwei Weng, Ying Tai, Jun Li 0027 |
AAAI | 2 |
| 2025 | Non-Aligned Supervision for Real Image DehazingabstractRemoving haze from real-world images is challenging due to unpredictable weather conditions, resulting in the misalignment of hazy and clear image pairs. In this paper, we propose an innovative dehazing framework that operates under non-aligned supervision. This framework is grounded in the atmospheric scattering model, and consists of three interconnected networks: dehazing, airlight, and transmission networks. In particular, we explore a non-alignment scenario that a clear reference image, unaligned with the input hazy image, is utilized to supervise the dehazing network. To implement this, we present a multi-scale reference loss that compares the feature representations between the referred image and the dehazed output. Our scenario makes it easier to collect hazy/clear image pairs in real-world environments, even under conditions of misalignment and shift views. To showcase the effectiveness of our scenario, we have collected a new hazy dataset including 415 image pairs captured by mobile Phone in both rural and urban areas, called "Phone-Hazy". Furthermore, we introduce a self-attention network based on mean and variance for modeling real infinite airlight, using the dark channel prior as positional guidance. Experimental results demonstrate the superior performance of our framework over existing state-of-the-art techniques in the real-world image dehazing task. Phone-Hazy and code will be available at https://fanjunkai1.github.io/projectpage/NSDNet/index.html. Junkai Fan, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Daytime-Mixed Non-Aligned Learning for Real Nighttime Image EnhancementabstractEnhancing real nighttime images is a significant challenge due to the deterioration of visual quality caused by limited perceptibility under adverse illumination conditions, leading to loss of details and color deviation. In this paper, we propose a novel nighttime image enhancement framework using daytime-mixed non-aligned supervision. It aims to couple the information between non-aligned daytime and nighttime image pairs. Specifically, our framework consists of a simple yet effective daytime-mixed supervised learning phase and a Retinex-based reconstruction phase. In the first phase, we employ a multi-instance with adaptive information fusion (AIF) module integrated within a UNet enhancement network called MIFUNet, which is trained via a daytime-mixed supervised loss. In the second phase, the Retinex-based reconstruction employs both a light-effect estimation network and an illumination adjustment network to restore the nighttime image, guided by physical principles. To evaluate the effectiveness of our approach, we collect a real non-aligned day-night dataset named the NANE dataset, which contains 748 non-aligned image pairs and 100 nighttime images solely for testing. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art image enhancement methods. Jiangwei Weng, Junkai Fan, Jianjun Qian, Haiyang Zou, Ying Tai, Jian Yang 0003, Jun Li 0027 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Driving-Video Dehazing with Non-Aligned Regularization for Safety AssistanceabstractReal driving-video dehazing poses a significant challenge due to the inherent difficulty in acquiring precisely aligned hazy/clear video pairs for effective model training, especially in dynamic driving scenarios with unpredictable weather conditions. In this paper, we propose a pioneering approach that addresses this challenge through a nonaligned regularization strategy. Our core concept involves identifying clear frames that closely match hazy frames, serving as references to supervise a video dehazing network. Our approach comprises two key components: reference matching and video dehazing. Firstly, we introduce a non-aligned reference frame matching module, leveraging an adaptive sliding window to match high-quality reference frames from clear videos. Video dehazing incorporates flow-guided cosine attention sampler and deformable cosine attention fusion modules to enhance spatial multi-frame alignment and fuse their improved information. To validate our approach, we collect a GoProHazy dataset captured effortlessly with GoPro cameras in diverse rural and urban road environments. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art methods in the challenging task of real driving-video dehazing. Project page. Junkai Fan, Jiangwei Weng, Kun Wang 0042, Jianjun Qian, Jun Li 0027, Jian Yang 0003 |
CVPR | 1 |
| 2024 | DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine DomainabstractIn this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cosine domain. This unique formulation allows for the modeling of local depth correlations within each patch. Crucially, the frequency transformation segregates the depth information into various frequency components, with low-frequency components encapsulating the core scene structure and high-frequency components detailing the finer aspects. This decomposition forms the basis of our progressive strategy, which begins with the prediction of low-frequency components to establish a global scene context, followed by successive refinement of local details through the prediction of higher-frequency components. We conduct comprehensive experiments on NYU-Depth-V2, TOFDC, and KITTI datasets, and demonstrate the state-of-the-art performance of DCDepth. Code is available at https://github.com/w2kun/DCDepth. Kun Wang 0042, Zhiqiang Yan 0001, Junkai Fan, Wanlu Zhu, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
NeurIPS | 3 |
| 2023 | Enhanced Frequency Information for Image Dehazing
Junkai Fan, Jun Li 0027, Jian Yang 0003 |
ICIG (1) | 2 |