Hai Jiang 0006

dblp:15/5983-6 · also Jiang Hai 0006 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-7087-6775ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021
YearPublicationVenuePosition
2026 RAW-Flow: Advancing RGB-to-RAW Image Reconstruction with Deterministic Latent Flow Matching
abstract
RGB-to-RAW reconstruction, or the reverse modeling of a camera Image Signal Processing (ISP) pipeline, aims to recover high-fidelity RAW data from RGB images. Despite notable progress, existing learning-based methods typically treat this task as a direct regression objective and still struggle with detail inconsistency and color deviation, due to the ill-posed nature of inverse ISP and the inherent information loss in quantized RGB images. To address these limitations, we pioneer a generative perspective by reformulating RGB-to-RAW reconstruction as a deterministic latent transport problem and introduce a novel framework named RAW-Flow, which leverages flow matching to learn a deterministic vector field in latent space, to effectively bridge the gap between RGB and RAW representations and enable accurate reconstruction of structural details and color information. To further enhance latent transport, we introduce a cross-scale context guidance module that injects hierarchical RGB features into the flow estimation process. Moreover, we design a Dual-domain Latent Autoencoder (DLAE) with a feature alignment constraint to support the proposed latent transport framework, which jointly encodes RGB and RAW inputs while promoting stable training and high-fidelity reconstruction. Extensive experiments demonstrate that RAW-Flow outperforms state-of-the-art approaches both quantitatively and visually.
Zhen Liu 0022, Diedong Feng, Hai Jiang 0006, Liaoyuan Zeng, Hao Wang 0073, Chaoyu Feng, Bing Zeng 0001, Shuaicheng Liu
AAAI3
2026 Supervised Small-Baseline and Large-Baseline Homography Learning With Diffusion-Based Data Generation
abstract
In this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data for supervised small-baseline and large-baseline homography learning and yield a state-of-the-art homography estimation network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane masks and homography of the pair, along with another sampled homography that serves as ground truth to generate a new labeled training pair with realistic motion. In the training phase, the generated data is used to train the supervised homography network, in which the training data is refined via a content refinement diffusion model. Once an iteration is finished, the trained network is used in the next data generation phase to update the pre-estimated homography. Through such an iterative strategy, the quality of the dataset and the performance of the network can be gradually and simultaneously improved. Experimental results show that our method outperforms existing competitors and previous supervised methods can also be improved based on the generated dataset.
Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Bing Zeng 0001, Shuaicheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Solving ILL-posed Regions in High Dynamic Range Reconstruction With Uncertainty-Aware Diffusion Models
abstract
Learning-based approaches have achieved promising progress in High Dynamic Range (HDR) image reconstruction, particularly in ghost removal. However, they often struggle in ill-posed regions, such as areas with occlusion or saturation, where insufficient or unreliable information leads to persistent residual ghosting artifacts and structural distortions. In this paper, we present UA-Diff, an uncertainty-aware diffusion framework designed to generate visually coherent, ghost-free HDR images. Specifically, our approach introduces an Uncertainty Generation Module (UGM) that estimates pixel-wise reconstruction confidence via a probabilistic Laplacian loss, producing an uncertainty map that explicitly highlights challenging ill-posed regions. To address these regions effectively, we develop an Uncertainty-Aware Diffusion Module (UADM) that operates selectively on the average-coefficient component of a 2D discrete wavelet transform, where dominant artifacts tend to concentrate. This enables reduced computational overhead while preserving high-quality details. Moreover, we propose an Uncertainty-Guided Sampling (UGS) strategy that leverages the uncertainty map to guide the denoising process, ensuring faithful reconstruction in reliable regions and targeted refinement in uncertain areas. Extensive experiments on three public HDR benchmarks demonstrate that UA-Diff surpasses state-of-the-art methods both quantitatively and perceptually, especially in challenging ill-posed scenarios.
Zhen Liu 0022, Hai Jiang 0006, Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 ISPDiffuser: Learning RAW-to-sRGB Mappings with Texture-Aware Diffusion Models and Histogram-Guided Color Consistency
abstract
RAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from raw data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing learning-based methods still struggle with detail disparity and color distortion. In this paper, we present ISPDiffuser, a diffusion-based decoupled framework that separates the RAW-to-sRGB mapping into detail reconstruction in grayscale space and color consistency mapping from grayscale to sRGB. Specifically, we propose a texture-aware diffusion model that leverages the generative ability of diffusion models to focus on local detail recovery, in which a texture enrichment loss is further proposed to prompt the diffusion model to generate more intricate texture details. Subsequently, we introduce a histogram-guided color consistency module that utilizes color histogram as guidance to learn precise color information for grayscale to sRGB color consistency mapping, with a color consistency loss designed to constrain the learned color information. Extensive experimental results show that the proposed ISPDiffuser outperforms state-of-the-art competitors both quantitatively and visually.
Yang Ren 0001, Hai Jiang 0006, Menglong Yang, Wei Li 0075, Shuaicheng Liu
AAAI2
2025 Learning to See in the Extremely Dark
Hai Jiang 0006, Binhao Guan, Zhen Liu 0022, Xiaohong Liu 0001, Songchen Han, Shuaicheng Liu
ICCV1
2025 Learning Arbitrary-Scale RAW Image Downscaling with Wavelet-based Recurrent Reconstruction
abstract
Image downscaling is critical for efficient storage and transmission of high-resolution (HR) images. Existing learning-based methods focus on performing downscaling within the sRGB domain, which typically suffers from blurred details and unexpected artifacts. RAW images, with their unprocessed photonic information, offer greater flexibility but lack specialized downscaling frameworks. In this paper, we propose a wavelet-based recurrent reconstruction framework that leverages the information lossless attribute of wavelet transformation to fulfill the arbitrary-scale RAW image downscaling in a coarse-to-fine manner, in which the Low-Frequency Arbitrary-Scale Downscaling Module (LASDM) and the High-Frequency Prediction Module (HFPM) are proposed to preserve structural and textural integrity of the reconstructed low-resolution (LR) RAW images, alongside an energy-maximization loss to align high-frequency energy between HR and LR domain. Furthermore, we introduce the Realistic Non-Integer RAW Downscaling (Real-NIRD) dataset, featuring a non-integer downscaling factor of 1.3×, and incorporate it with publicly available datasets with integer factors (2×, 3×, 4×) for comprehensive benchmarking arbitrary-scale image downscaling purposes. Extensive experiments demonstrate that our method outperforms existing state-of-the-art competitors both quantitatively and visually. The code and dataset will be released at https://github.com/RenYangSCU/ASRD.
Yang Ren 0001, Hai Jiang 0006, Wei Li 0075, Menglong Yang, Heng Zhang 0042, Zehua Sheng, Qingsheng Ye, Shuaicheng Liu
ACM Multimedia2
2025 ISPFormer: Learning RAW-to-sRGB mappings with wavelet-based self-attention
Yang Ren 0001, Xinhan Niu, Hai Jiang 0006, Zhen Liu 0022, Ting Jiang 0005, Guanghui Liu 0001, Shuaicheng Liu
Neurocomputing3
2025 Incorporating Fourier Transformation With Diffusion Models for Low-Light Image Enhancement
abstract
In this letter, we propose a diffusion-based framework that leverages the generative ability of diffusion models and the advantages of the physically explainable Fourier transformation for visually satisfactory low-light image enhancement. Specifically, we first employ an encoder to convert the paired low-light and normal-light images into latent features and transform the features into the frequency domain through Fourier transformation, resulting in amplitude components that contain illumination information and phase components that represent details information. Subsequently, we present the latent-Fourier diffusion model which performs diffusion operations on the phase components for details reconstruction. Furthermore, we propose a lightness boost module to reconstruct amplitude aiming to improve the contrast in the frequency domain, and the restored feature obtained by performing inverse Fourier transformation on the reconstructed phase and amplitude components is further refined by the proposed latent feature fusion module to achieve better visual perception. Finally, the refined feature is taken as input to a decoder to produce the final restored image. Extensive experiments on publicly available benchmarks demonstrate our proposed method outperforms state-of-the-art competitors.
Ailin Ma, Hai Jiang 0006, Binbin Liang, Songchen Han
IEEE Signal Process. Lett.2
2024 LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models
Hai Jiang 0006, Ao Luo, Xiaohong Liu 0001, Songchen Han, Shuaicheng Liu
ECCV (48)1
2024 Revisiting coarse-to-fine strategy for low-light image enhancement with deep decomposition guided training
Hai Jiang 0006, Yang Ren 0001, Songchen Han
Comput. Vis. Image Underst.1
2024 DMHomo: Learning Homography with Diffusion Models
abstract
Supervised homography estimation methods face a challenge due to the lack of adequate labeled training data. To address this issue, we propose DMHomo , a diffusion model-based framework for supervised homography learning. This framework generates image pairs with accurate labels, realistic image content, and realistic interval motion, ensuring that they satisfy adequate pairs. We utilize unlabeled image pairs with pseudo labels such as homography and dominant plane masks, computed from existing methods, to train a diffusion model that generates a supervised training dataset. To further enhance performance, we introduce a new probabilistic mask loss, which identifies outlier regions through supervised training, and an iterative mechanism to optimize the generative and homography models successively. Our experimental results demonstrate that DMHomo effectively overcomes the scarcity of qualified datasets in supervised homography learning and improves generalization to real-world scenes. The code and dataset are available at GitHub ( https://github.com/lhaippp/DMHomo ).
Haipeng Li 0001, Hai Jiang 0006, Ao Luo, Ping Tan 0002, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu
ACM Trans. Graph.2
2023 Semi-supervised Deep Large-Baseline Homography Estimation with Progressive Equivalence Constraint
abstract
Homography estimation is erroneous in the case of large-baseline due to the low image overlay and limited receptive field. To address it, we propose a progressive estimation strategy by converting large-baseline homography into multiple intermediate ones, cumulatively multiplying these intermediate items can reconstruct the initial homography. Meanwhile, a semi-supervised homography identity loss, which consists of two components: a supervised objective and an unsupervised objective, is introduced. The first supervised loss is acting to optimize intermediate homographies, while the second unsupervised one helps to estimate a large-baseline homography without photometric losses. To validate our method, we propose a large-scale dataset that covers regular and challenging scenes. Experiments show that our method achieves state-of-the-art performance in large-baseline scenes while keeping competitive performance in small-baseline scenes. Code and dataset are available at https://github.com/megvii-research/LBHomo.
Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Shuaicheng Liu
AAAI1
2023 Supervised Homography Learning with Realistic Dataset Generation
abstract
In this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data and yield a supervised homography network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane masks and homography of the pair, along with another sampled homography that serves as ground truth to generate a new labeled training pair with realistic motion. In the training phase, the generated data is used to train the supervised homography network, in which the training data is refined via a content consistency module and a quality assessment module. Once an iteration is finished, the trained network is used in the next data generation phase to update the pre-estimated homography. Through such an iterative strategy, the quality of the dataset and the performance of the network can be gradually and simultaneously improved. Experimental results show that our method achieves state-of-the-art performance and existing supervised methods can be also improved based on the generated dataset. Code and dataset are available at https://github.com/JianghaiSCU/RealSH.
Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu
ICCV1
2023 R2RNet: Low-light image enhancement via Real-low to Real-normal Network
Hai Jiang 0006, Zhu Xuan, Yang Ren 0001, Yutong Hao, Fengzhu Zou, Songchen Han
J. Vis. Commun. Image Represent.1
2023 Unsupervised Global and Local Homography Estimation With Motion Basis Learning
abstract
In this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homography flow bases. Second, considering a homography contains 8 Degree-of-Freedoms (DOFs) that is much less than the rank of the network features, we propose a Low Rank Representation (LRR) block that reduces the feature rank, so that features corresponding to the dominant motions are retained while others are rejected. Last, we propose a Feature Identity Loss (FIL) to enforce the learned image feature warp-equivariant, meaning that the result should be identical if the order of warp operation and feature extraction is swapped. With this constraint, the unsupervised optimization can be more effective and the learned features are more stable. With global-to-local homography flow refinement, we also naturally generalize the proposed method to local mesh-grid homography estimation, which can go beyond the constraint of a single homography. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the state-of-the-art on the homography benchmark dataset both qualitatively and quantitatively. Code is available at https://github.com/megvii-research/BasesHomo.
Shuaicheng Liu, Hai Jiang 0006, Nianjin Ye, Chuan Wang 0001, Bing Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Advanced RetinexNet: A fully convolutional network for low-light image enhancement
Hai Jiang 0006, Yutong Hao, Fengzhu Zou, Songchen Han
Signal Process. Image Commun.1
2023 Low-Light Image Enhancement with Wavelet-Based Diffusion Models
abstract
Diffusion models have achieved promising results in image restoration tasks, yet suffer from time-consuming, excessive computational resource consumption, and unstable restoration. To address these issues, we propose a robust and efficient Diffusion-based Low-Light image enhancement approach, dubbed DiffLL. Specifically, we present a wavelet-based conditional diffusion model (WCDM) that leverages the generative power of diffusion models to produce results with satisfactory perceptual fidelity. Additionally, it also takes advantage of the strengths of wavelet transformation to greatly accelerate inference and reduce computational resource usage without sacrificing information. To avoid chaotic content and diversity, we perform both forward diffusion and denoising in the training phase of WCDM, enabling the model to achieve stable denoising and reduce randomness during inference. Moreover, we further design a high-frequency restoration module (HFRM) that utilizes the vertical and horizontal details of the image to complement the diagonal information for better fine-grained restoration. Extensive experiments on publicly available real-world benchmarks demonstrate that our method outperforms the existing state-of-the-art methods both quantitatively and visually, and it achieves remarkable improvements in efficiency compared to previous diffusion-based methods. In addition, we empirically show that the application for low-light face detection also reveals the latent practical values of our method. Code is available at https://github.com/JianghaiSCU/Diffusion-Low-Light.
Hai Jiang 0006, Ao Luo, Haoqiang Fan, Songchen Han, Shuaicheng Liu
ACM Trans. Graph.1
2022 Combining Spatial and Frequency Information for Image Deblurring
abstract
This paper aims to combine spatial and frequency information for single image deblurring. Although some methods have tried to use frequency information to perform deblurring, they only simply process the different frequencies information separately or concatenate the real part and imaginary part of frequency features but ignore the strong correlation between them. To address this problem, we propose a simple but effective frequency interaction pipeline to realize the mutual conversion of the real part and the imaginary part. Then, we construct a spatial-frequency conversion module (SFCM) to promote the mutual conversion between the frequency information and the spatial information. Based on the proposed components, we build a multi-scale deblurring network, dubbed SFDNet, which can fully exploit coarse and middle-level information in spatial and frequency domains for finer scale image deblurring. Extensive experiments on the GoPro and HIDE datasets demonstrate that the proposed network outperforms the state-of-the-art methods both quantitatively and visually.
Hai Jiang 0006, Yang Ren 0001, Yaqi Yu, Songchen Han
IEEE Signal Process. Lett.1