EDBT 2026 Demo / reviewers in the wild / expert
Chongyi Li
dblp:169/5585
· DBLP profile ↗
134ranked-venue papers
18as first author
112since 2021 · last 2026
0000-0003-2609-2460ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 9 first-author · 71 since 2021Artificial intelligence and machine learning · 77 · 9 first-author · 69 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 7 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PerTouch: VLM-Driven Agent for Personalized and Semantic Image RetouchingabstractImage retouching aims to enhance visual quality while aligning with users' personalized aesthetic preferences. To address the challenge of balancing controllability and subjectivity, we propose a unified diffusion-based image retouching framework called PerTouch. Our method supports semantic-level image retouching while maintaining global aesthetics. Using parameter maps containing attribute values in specific semantic regions as input, PerTouch constructs an explicit parameter-to-image mapping for fine-grained image retouching. To improve semantic boundary perception, we introduce semantic replacement and parameter perturbation mechanisms during training. To connect natural language instructions with visual control, we develop a VLM-driven agent to handle both strong and weak user instructions. Equipped with mechanisms of feedback-driven rethinking and scene-aware memory, PerTouch better aligns with user intent and captures long-term preferences. Extensive experiments demonstrate each component’s effectiveness and the superior performance of PerTouch in personalized image retouching. Zewei Chang, Zheng-Peng Duan, Jianxing Zhang, Chunle Guo, Hyungju Chun, Hyunhee Park, Zikun Liu 0001, Chongyi Li |
AAAI | 9 |
| 2026 | EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment EvaluationabstractText-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated methods emerge to evaluate the image-text alignment capabilities of generative models. However, the performance comparison among these automated methods is constrained by the limited scale of existing datasets. Additionally, existing datasets lack the capacity to assess the performance of automated methods at a fine-grained level. In this study, we contribute an EvalMuse-40K dataset, gathering 40K image-text pairs with fine-grained human annotations for image-text alignment-related tasks. In the construction process, we employ various strategies such as balanced prompt sampling and data re-annotation to ensure the diversity and reliability of our dataset. This allows us to comprehensively evaluate the performance of image-text alignment methods for T2I models. Based on this dataset, we introduce an efficient automated evaluation method termed FGA-BLIP2, which enables Fine-Grained Alignment evaluation solely by inputting images and text leveraging BLIP2, without visual question answering for each fine-grained element. Experimental results show the proposed FGA-BLIP2 efficiently achieves good performance on multiple image-text alignment datasets. Meanwhile, benefiting from the high efficiency and fine-grained evaluation capability of FGA-BLIP2, we apply it as a reward model to improve text-to-image models, which effectively enhances the image-text alignment ability of text-to-image models. Shuhao Han, Haotian Fan, Jiachen Fu, Tao Li 0008, Junhui Cui, Yunqiu Wang, Yang Tai, Chunle Guo, Chongyi Li |
AAAI | 11 |
| 2026 | VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame InterpolationabstractDue to large pixel movement and high computational cost, estimating the motion of high-resolution frames is challenging. Thus, most flow-based Video Frame Interpolation (VFI) methods first predict bidirectional flows at low resolution and then use high-magnification upsampling (e.g., bilinear) to obtain the high-resolution ones. However, this kind of upsampling strategy may cause blur or mosaic at the flows' edges. Additionally, the motion of fine pixels at high resolution cannot be adequately captured in motion estimation at low resolution, which leads to the misalignment of task-oriented flows. With such inaccurate flows, input frames are warped and combined pixel-by-pixel, resulting in ghosting and discontinuities in the interpolated frame. In this study, we propose a novel VFI pipeline, VTinker, which consists of two core components: guided flow upsampling (GFU) and Texture Mapping. After motion estimation at low resolution, GFU introduces input frames as guidance to alleviate the blurring details in bilinear upsampling flows, which makes flows' edges clearer. Subsequently, to avoid pixel-level ghosting and discontinuities, Texture Mapping generates an initial interpolated frame, referred to as the intermediate proxy. The proxy serves as a cue for selecting clear texture blocks from the input frames, which are then mapped onto the proxy to facilitate producing the final interpolated frame via a reconstruction module. Extensive experiments demonstrate that VTinker achieves state-of-the-art performance in VFI. Jiayi Fu, Chunle Guo, Shuhao Han, Chongyi Li |
AAAI | 5 |
| 2026 | Control-Lit: Illumination Controllable Backlit Image EnhancementabstractBacklit image enhancement (BIE) aims to address image degradation caused by challenging lighting conditions. By enhancing the illumination of underexposed areas and restoring image details while avoiding overexposure, it achieves an overall harmonious luminance. In contrast to traditional BIE methods that apply global enhancements with limited effectiveness, we propose a controllable Mamba-based enhancement method, termed Control-Lit. Our method not only delivers effective backlit image enhancement but also allows users to interactively adjust the illumination of specific regions. Control-Lit achieves superior global enhancement performance by employing the Dark Channel Prior (DCP)-based Finite Scalar Quantization (DFSQ) module that provides a high-quality image prior. Additionally, the Dark Channel Prior Enhancement (DCPE) module is designed to guide the network in pixel-wise illumination adjustment at the feature level, thereby achieving enhancement of backlit regions. Furthermore, to address the loss of content and details in backlit regions, the Global-Local Vision State Space (GLVSS) module is incorporated to extract both global and local features for BIE. To enable customizable, controllable enhancement, we introduce an illumination control vector. By adjusting the coefficient map elements that are multiplied by this vector, users can achieve precise regional illumination adjustments. To further validate the generalization of different light enhancement methods, we contribute a synthetic backlit image dataset using a relighting generative model. Along with current widely used datasets, the experimental results demonstrate that our method achieves state-of-the-art performance on all datasets while enabling high-quality and controllable image illumination adjustment. The code is available at https://github.com/wuhj43/Control-Lit. Hongjun Wu 0003, Yi Tang 0008, Chongyi Li, Zhi Jin 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Expose Camouflage in the Water: Underwater Camouflaged Instance Segmentation and DatasetabstractWith the development of underwater exploration and marine protection, underwater vision tasks are widespread. Due to the degraded underwater environment, characterized by color distortion, low contrast, and blurring, camouflaged instance segmentation (CIS) faces greater challenges in accurately segmenting objects that blend closely with their surroundings. Traditional camouflaged instance segmentation methods, trained on terrestrial-dominated datasets with limited underwater samples, may exhibit inadequate performance in underwater scenes. To address these issues, we introduce the first underwater camouflaged instance segmentation (UCIS) dataset, abbreviated as UCIS4K, which comprises 3,953 images of camouflaged marine organisms with instance-level annotations. In addition, we propose an Underwater Camouflaged Instance Segmentation network based on Segment Anything Model (UCIS-SAM). Our UCIS-SAM includes three key modules. First, the Channel Balance Optimization Module (CBOM) enhances channel characteristics to improve underwater feature learning, effectively addressing the model's limited understanding of underwater environments. Second, the Frequency Domain True Integration Module (FDTIM) is proposed to emphasize intrinsic object features and reduce interference from camouflage patterns, enhancing the segmentation performance of camouflaged objects blending with their surroundings. Finally, the Multi-scale Feature Frequency Aggregation Module (MFFAM) is designed to strengthen the boundaries of low-contrast camouflaged instances across multiple frequency bands, improving the model's ability to achieve more precise segmentation of camouflaged objects. Extensive experiments on the proposed UCIS4K and public benchmarks show that our UCIS-SAM outperforms state-of-the-art approaches. The code and dataset are released at https://github.com/wchchw/UCIS4K. Chuhong Wang, Hua Li 0012, Chongyi Li, Huazhong Liu, Xiongxin Tang, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2026 | Incorporating Uncertainty-Guided and Top-k Codebook Matching for Real-World Blind Image Super-ResolutionabstractRecent advancements in codebook-based real image super-resolution (SR) have shown promising results in real-world applications. The core idea involves matching high-quality image features from a codebook based on low-resolution (LR) image features. However, existing methods face two major challenges: inaccurate feature matching with the codebook and poor texture detail reconstruction. To address these issues, we propose a novel Uncertainty-Guided and Top-k Codebook Matching SR (UGTSR) framework, which incorporates three key components: 1) an uncertainty learning mechanism that guides the model to focus on texture-rich regions, 2) a Top-k feature matching strategy that enhances feature matching accuracy by fusing multiple candidate features, and 3) an Align-Attention module that enhances the alignment of information between LR and HR features. Experimental results demonstrate significant improvements in texture realism and reconstruction fidelity compared to existing methods. The source code can be found at https://github.com/wwlCape/UGTSR-main. Weilei Wen, Zhaohui Zheng 0003, Chunle Guo, Xiuli Shao, Chongyi Li |
IEEE Trans. Image Process. | 7 |
| 2026 | 3D-UIR: 3D Gaussian for Underwater 3D Scene Reconstruction via Physics-Based Appearance-Medium DecouplingabstractNovel view synthesis for underwater scene reconstruction presents unique challenges due to complex light-media interactions. Optical scattering and absorption in water body bring inhomogeneous medium attenuation interference that disrupts conventional volume rendering assumptions of uniform propagation medium. While 3D Gaussian Splatting (3DGS) offers real-time rendering capabilities, it struggles with underwater inhomogeneous environments where scattering media introduces artifacts and inconsistent appearance. In this study, we propose a physics-based framework that disentangles object appearance from water medium effects through tailored Gaussian modeling. Our approach introduces appearance embeddings, which are explicit medium representations for backscatter and attenuation, enhancing scene consistency. In addition, we propose a depth-guided optimization strategy that leverages pseudo-depth maps as supervision with depth regularization and scale penalty terms to improve geometric fidelity. By integrating the proposed appearance and medium modeling components via an underwater imaging model, our approach achieves both high-quality novel view synthesis and physically accurate scene restoration. Experiments demonstrate our significant improvements in rendering quality and restoration accuracy over existing methods. The project page is available at https://bilityniu.github.io/3D-UIR. Jieyu Yuan, Yuanlin Zhang 0009, Chunle Guo, Xiongxin Tang, Ruixing Wang, Chongyi Li |
IEEE Trans. Image Process. | 7 |
| 2025 | A Diffusion-Based Framework for Occluded Object MovementabstractSeamlessly moving objects within a scene is a common requirement for image editing, but it is still a challenge for existing editing methods. Especially for real-world images, the occlusion situation further increases the difficulty. The main difficulty is that the occluded portion needs to be completed before movement can proceed. To leverage the real-world knowledge embedded in the pre-trained diffusion models, we propose a Diffusion-based framework specifically designed for Occluded Object Movement, named DiffOOM. The proposed DiffOOM consists of two parallel branches that perform object de-occlusion and movement simultaneously. The de-occlusion branch utilizes a background color-fill strategy and a continuously updated object mask to focus the diffusion process on completing the obscured portion of the target object. Concurrently, the movement branch employs latent optimization to place the completed object in the target location and adopts local text-conditioned guidance to integrate the object into new surroundings appropriately. Extensive evaluations across various metrics demonstrate the superior performance of our method, which is further validated by a comprehensive user study. Zheng-Peng Duan, Jiawei Zhang 0002, Zheng Lin 0005, Chunle Guo, Dongqing Zou, Jimmy S. J. Ren, Chongyi Li |
AAAI | 8 |
| 2025 | DiffRetouch: Using Diffusion to Retouch on the Shoulder of ExpertsabstractImage retouching aims to enhance the visual quality of photos. Considering the different aesthetic preferences of users, the target of retouching is subjective. However, current retouching methods mostly adopt deterministic models, which not only neglects the style diversity in the expert-retouched results and tends to learn an average style during training, but also lacks sample diversity during inference. In this paper, we propose a diffusion-based method, named DiffRetouch. Thanks to the excellent distribution modeling ability of diffusion, our method can capture the complex fine-retouched distribution covering various visual-pleasing styles in the training data. Moreover, four image attributes are made adjustable to provide a user-friendly editing mechanism. By adjusting these attributes in specified ranges, users are allowed to customize preferred styles within the learned fine-retouched distribution. Additionally, the affine bilateral grid and contrastive learning scheme are introduced to handle the problem of texture distortion and control insensitivity respectively. Extensive experiments have demonstrated the superior performance of our method on visually appealing and sample diversity. Zheng-Peng Duan, Jiawei Zhang 0002, Zheng Lin 0005, Xin Jin 0005, Xundong Wang, Dongqing Zou, Chunle Guo, Chongyi Li |
AAAI | 8 |
| 2025 | FaceMe: Robust Blind Face Restoration with Personal IdentificationabstractBlind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a personalized face restoration method, FaceMe, based on a diffusion model. Given a single or a few reference images, we use an identity encoder to extract identity-related features, which serve as prompts to guide the diffusion model in restoring high-quality and identity-consistent facial images. By simply combining identity-related features, we effectively minimize the impact of identity-irrelevant features during training and support any number of reference image inputs during inference. Additionally, thanks to the robustness of the identity encoder, synthesized images can be used as reference images during training, and identity changing during inference does not require fine-tuning the model. We also propose a pipeline for constructing a reference image training pool that simulates the poses and expressions that may appear in real-world scenarios. Experimental results demonstrate that our FaceMe can restore high-quality facial images while maintaining identity consistency, achieving excellent performance and robustness. Zheng-Peng Duan, Jia Ouyang, Jiayi Fu, Hyunhee Park, Zikun Liu 0001, Chunle Guo, Chongyi Li |
AAAI | 8 |
| 2025 | IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model
Yifan Chen 0001, Chunle Guo, Chongyi Li, Yujiu Yang 0001 |
CGI (3) | 4 |
| 2025 | Classic Video Denoising in a Machine Learning World: Robust, Fast, and ControllableabstractDenoising is a crucial step in many video processing pipelines such as in interactive editing, where high quality, speed, and user control are essential. While recent approaches achieve significant improvements in denoising quality by leveraging deep learning, they are prone to unexpected failures due to discrepancies between training data distributions and the wide variety of noise patterns found in real-world videos. These methods also tend to be slow and lack user control. In contrast, traditional denoising methods perform reliably on in-the-wild videos and run relatively quickly on modern hardware. However, they require manually tuning parameters for each input video, which is not only tedious but also requires skill. We bridge the gap between these two paradigms by proposing a differentiable denoising pipeline based on traditional methods. A neural network is then trained to predict the optimal denoising parameters for each specific input, resulting in a robust and efficient approach that also supports user control. Xin Jin 0005, Simon Niklaus, Zhoutong Zhang, Zhihao Xia, Chunle Guo, Jiawen Chen 0001, Chongyi Li |
CVPR | 8 |
| 2025 | Iterative Predictor-Critic Code Decoding for Real-World Image DehazingabstractWe propose a novel Iterative Predictor-Critic Code Decoding framework for real-world image dehazing, abbreviated as IPC-Dehaze, which leverages the high-quality codebook prior encapsulated in a pre-trained VQGAN. Apart from previous codebook-based methods that rely on oneshot decoding, our method utilizes high-quality codes obtained in the previous iteration to guide the prediction of the Code-Predictor in the subsequent iteration, improving code prediction accuracy and ensuring stable dehazing performance. Our idea stems from the observations that 1) the degradation of hazy images varies with haze density and scene depth, and 2) clear regions play crucial cues in restoring dense haze regions. However, it is nontrivial to progressively refine the obtained codes in subsequent iterations, owing to the difficulty in determining which codes should be retained or replaced at each iteration. Another key insight of our study is to propose CodeCritic to capture interrelations among codes. The CodeCritic is used to evaluate code correlations and then resample a set of codes with the highest mask scores, i.e., a higher score indicates that the code is more likely to be rejected, which helps retain more accurate codes and predict difficult ones. Extensive experiments demonstrate the superiority of our method over state-of-the-art methods in real-world dehazing. Our project page can be found at https://github.com/Jiayi-Fu/IPC-Dehaze. Jiayi Fu, Zikun Liu 0001, Chunle Guo, Hyunhee Park, Guoqing Wang 0001, Chongyi Li |
CVPR | 8 |
| 2025 | DiT4SR: Taming Diffusion Transformer for Real-World Image Super-ResolutionabstractLarge-scale pre-trained diffusion models are becoming increasingly popular in solving the Real-World Image Super-Resolution (Real-ISR) problem because of their rich generative priors. The recent development of diffusion transformer (DiT) has witnessed overwhelming performance over the traditional UNet-based architecture in image generation, which also raises the question: Can we adopt the advanced DiT-based diffusion model for Real-ISR? To this end, we propose our DiT4SR, one of the pioneering works to tame the large-scale DiT model for Real-ISR. Instead of directly injecting embeddings extracted from low-resolution (LR) images like ControlNet, we integrate the LR embeddings into the original attention mechanism of DiT, allowing for the bidirectional flow of information between the LR latent and the generated latent. The sufficient interaction of these two streams allows the LR stream to evolve with the diffusion process, producing progressively refined guidance that better aligns with the generated latent at each diffusion step. Additionally, the LR guidance is injected into the generated latent via a cross-stream convolution layer, compensating for DiT's limited ability to capture local information. These simple but effective designs endow the DiT model with superior performance in Real-ISR, which is demonstrated by extensive experiments. Project Page: https://adam-duan.github.io/projects/dit4sr/. Zheng-Peng Duan, Jiawei Zhang 0002, Xin Jin 0005, Zheng Xiong, Dongqing Zou, Jimmy S. J. Ren, Chunle Guo, Chongyi Li |
ICCV | 9 |
| 2025 | Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local Encoding
Jingming He, Chongyi Li, Shiqi Wang 0001, Sam Kwong |
ICCV | 2 |
| 2025 | MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data Generation
Fu-Zhao Ou, Chongyi Li, Shiqi Wang 0001, Sam Kwong |
ICCV | 2 |
| 2025 | DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
Jiachen Fu, Chunle Guo, Chongyi Li |
ACM Multimedia | 3 |
| 2025 | UltraLED: Learning to See Everything in Ultra-High Dynamic Range ScenesabstractUltra-high dynamic range (UHDR) scenes exhibit pronounced exposure disparities between bright and dark regions. Such conditions are Ultra-high dynamic range (UHDR) scenes exhibit significant exposure disparities between bright and dark regions. Such conditions are commonly encountered in nighttime scenes with light sources. Even with standard exposure settings, a bimodal intensity distribution with boundary peaks often emerges, making it difficult to preserve both highlight and shadow details simultaneously. RGB-based bracketing methods can capture details at both ends using short-long exposure pairs, but are susceptible to misalignment and ghosting artifacts. We found that a short-exposure image already retains sufficient highlight detail. The main challenge of UHDR reconstruction lies in denoising and recovering information in dark regions. In comparison to the RGB images, RAW images, thanks to their higher bit depth and more predictable noise characteristics, offer greater potential for addressing this challenge. This raises a key question: can we learn to see everything in UHDR scenes using only a single short-exposure RAW image? In this study, we rely solely on a single short-exposure frame, which inherently avoids ghosting and motion blur, making it particularly robust in dynamic scenes. To achieve that, we introduce UltraLED, a two-stage framework that performs exposure correction via a ratio map to balance dynamic range, followed by a brightness-aware RAW denoiser to enhance detail recovery in dark regions. To support this setting, we design a 9-stop bracketing pipeline to synthesize realistic UHDR images and contribute a corresponding dataset based on diverse scenes, using only the shortest exposure as input for reconstruction. Extensive experiments show that UltraLED significantly outperforms existing single-frame approaches. Our code and dataset are made publicly available at https://srameo.github.io/projects/ultraled. Yuang Meng, Xin Jin 0005, Lina Lei, Chunle Guo, Chongyi Li |
NeurIPS | 5 |
| 2025 | DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse DataabstractWe present **DIPO**, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state.
Compared to the single-image approach, our dual-image input imposes only a modest overhead for data collection, but at the same time provides important motion information, which is a reliable guide for predicting kinematic relationships between parts.
Specifically, we propose a dual-image diffusion model that captures relationships between the image pair to generate part layouts and joint parameters. In addition, we introduce a Chain-of-Thought (CoT) based **graph reasoner** that explicitly infers part connectivity relationships.
To further improve robustness and generalization on complex articulated objects, we develop a fully automated dataset expansion pipeline, name **LEGO-Art**, that enriches the diversity and complexity of PartNet-Mobility dataset. We propose **PM-X**, a large-scale dataset of complex articulated 3D objects, accompanied by rendered images, URDF annotations, and textual descriptions.
Extensive experiments demonstrate that DIPO significantly outperforms existing baselines in both the resting state and the articulated state, while the proposed PM-X dataset further enhances generalization to diverse and structurally complex articulated objects.
Our code and dataset are available at https://github.com/RQ-Wu/DIPO. Chunle Guo, Jiaxiong Qiu, Chongyi Li, Lichao Huang, Zhizhong Su, Ming-Ming Cheng |
NeurIPS | 6 |
| 2025 | FCNet: A feature complementary network for nighttime flare removal
Kejing Qi, Chongyi Li |
Comput. Vis. Image Underst. | 3 |
| 2025 | Control Color: Multimodal Diffusion-Based Interactive Image Colorization
Zhexin Liang, Zhaochen Li, Shangchen Zhou, Chongyi Li, Chen Change Loy |
Int. J. Comput. Vis. | 4 |
| 2025 | Semantic segmentation in adverse scenes with fewer labeled images
Guanhua An, Jichang Guo, Chunle Guo, Yudong Wang 0002, Chongyi Li |
Neural Networks | 5 |
| 2025 | A General Spatial-Frequency Learning Framework for Multimodal Image FusionabstractMultimodal image fusion involves tasks like pan-sharpening and depth super-resolution. Both tasks aim to generate high-resolution target images by fusing the complementary information from the texture-rich guidance and low-resolution target counterparts. They are inborn with reconstructing high-frequency information. Despite their inherent frequency domain connection, most existing methods only operate solely in the spatial domain and rarely explore the solutions in the frequency domain. This study addresses this limitation by proposing solutions in both the spatial and frequency domains. We devise a Spatial-Frequency Information Integration Network, abbreviated as SFINet for this purpose. The SFINet includes a core module tailored for image fusion. This module consists of three key components: a spatial-domain information branch, a frequency-domain information branch, and a dual-domain interaction. The spatial-domain information branch employs the spatial convolution-equipped invertible neural operators to integrate local information from different modalities in the spatial domain. Meanwhile, the frequency-domain information branch adopts a modality-aware deep Fourier transformation to capture the image-wide receptive field for exploring global contextual information. In addition, the dual-domain interaction facilitates information flow and the learning of complementary representations. We further present an improved version of SFINet, SFINet++, that enhances the representation of spatial information by replacing the basic convolution unit in the original spatial domain branch with the information-lossless invertible neural operator. We conduct extensive experiments to validate the effectiveness of the proposed networks and demonstrate their outstanding performance against state-of-the-art methods in two representative multimodal image fusion tasks: pan-sharpening and depth super-resolution. Man Zhou 0003, Jie Huang 0017, Danfeng Hong, Xiuping Jia, Jocelyn Chanussot, Chongyi Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Data-driven gradient priors integrated into blind image deblurring
Qing Qi, Jichang Guo, Chongyi Li |
Signal Process. Image Commun. | 3 |
| 2025 | Toward Generalized and Realistic Unpaired Image Dehazing via Region-Aware Physical ConstraintsabstractSupervised dehazing models, trained on synthetic hazy-clean image pairs, often face a notable decline in performance when applied to real-world scenes. Consequently, CycleGAN-based unpaired dehazing methods are proposed to improve the model’s generalization. One successful approach among these methods involves decomposing the physical properties of the atmospheric scattering model (ASM). However, estimating physical properties individually from input images is difficult without supervised labels, which ignores the semantic consistency between different physical regions. We claim semantic region information can offer additional geometric spatial constraints for estimating physical properties, as natural images can be divided into regions with similar scene depths. Motivated by this, we propose a novel generalized and realistic unpaired image dehazing framework via region-aware physical constraints (RPC-Dehaze). Our approach utilizes fine-grained semantic region maps from the Segment Anything Model (SAM) in a specially designed region prompt enhancement module. This enables the dehazing and hazing cyclic networks to learn region-aware physical constraints, leading to accurate estimation of haze imaging physical properties. In contrast to existing unpaired methods that treat dehazing and hazing networks equally, we incorporate Retinex theory into the hazing network, allowing it to learn diverse illumination effects in different regions. We adaptively refine the Retinex-based illumination component, resulting in more realistic hazy images. To further facilitate unsupervised learning in our framework, we propose a physical consensual contrastive regularization to ensure compact representation constraints in the latent feature space. Extensive experiments on synthetic and real image datasets show our method surpasses state-of-the-art unpaired dehazing methods in both effectiveness and generalization capability. Kaihao Lin, Guoqing Wang 0001, Tianyu Li 0003, Yuhui Wu 0001, Chongyi Li, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Improving Robustness of Point Cloud Analysis Through Perturbation Simulation and Distortion-Guided Feature AugmentationabstractRobust analysis of 3D point cloud data is essential for high-precision applications such as autonomous driving and industrial automation, where models need to consistently perform under complex and unpredictable real-world conditions. Current strategies, including data augmentation techniques and robust network designs, often struggle to effectively capture dynamic disturbances or accommodate the spatial variations of point clouds, thus lacking the required flexibility across diverse environments. To overcome these limitations, we propose a novel methodology to improve the robustness of 3D point cloud processing systems. Our approach simulates generalized corrupted input samples during training, using Radial Basis Functions (RBF) to model smooth deformations based on the control points. These deformations are applied selectively to different regions of the point cloud, adapting to spatial heterogeneity based on local density and geometric complexity. While generating these samples, we employ a combined adversarial loss that simultaneously induces model errors and maximizes the difference in internal feature distributions between the original and perturbed data. Additionally, we introduce a sub-network for distortion-guided feature augmentation to enhance important features while suppressing unreliable ones. This sub-network estimates distortion levels by compressing features and identifying discrepancies, then adjusts feature extraction process accordingly. Experimental results demonstrate that our method outperforms existing approaches on both Computer-Aided Design (CAD) models and real-world LiDAR datasets, enhancing model resilience and accuracy in handling diverse 3D scenarios. Jingming He, Chongyi Li, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2025 | Deep Underwater Image Quality Assessment With Explicit Degradation Awareness EmbeddingabstractUnderwater Image Quality Assessment (UIQA) is currently an area of intensive research interest. Existing deep learning-based UIQA models always learn a deep neural network to directly map the input degraded underwater image into a final quality score via end-to-end training. However, a wide variety of image contents or distortion types may correspond to the same quality score, making it challenging to train such a deep model merely with a single subjective quality score as supervision. An intuitive idea to solve this problem is to exploit more detailed degradation-aware information as supplementary guidance to facilitate model learning. In this paper, we devise a novel deep UIQA model with Explicit Degradation Awareness embedding, i.e., EDANet. To train the EDANet, a two-stage training strategy is adopted. First, a tailored Degradation Information Discovery subnetwork (DIDNet) is pre-trained to infer a residual map between the input degraded underwater image and its pseudoreference counterpart. The inferred residual map explicitly characterizes the local degradation of the input underwater image. The intermediate feature representations on the decoder side of DIDNet are then embedded into the Degradation-guided Quality Evaluation subnetwork (DQENet), which significantly enhances the feature characterization capability with higher degradation awareness for quality prediction. The superiority of our EDANet against 18 state-of-the-art methods has been well demonstrated by extensive comparisons on two benchmark datasets. The source code of our EDANet is available at https://github.com/yia-yuese/EDANet. Qiuping Jiang, Yuese Gu, Zongwei Wu, Chongyi Li, Huan Xiong, Feng Shao 0001, Zhihua Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2025 | Heterogeneous Experts and Hierarchical Perception for Underwater Salient Object DetectionabstractExisting underwater salient object detection (USOD) methods design fusion strategies to integrate multimodal information, but lack exploration of modal characteristics. To address this, we separately leverage the RGB and depth branches to learn disentangled representations, formulating the heterogeneous experts and hierarchical perception network (HEHP). Specifically, to reduce modal discrepancies, we propose the hierarchical prototype guided interaction (HPI), which achieves fine-grained alignment guided by the semantic prototypes, and then refines with complementary modalities. We further design the mixture of frequency experts (MoFE), where experts focus on modeling high- and low-frequency respectively, collaborating to explicitly obtain hierarchical representations. To efficiently integrate diverse spatial and frequency information, we formulate the four-way fusion experts (FFE), which dynamically selects optimal experts for fusion while being sensitive to scale and orientation. Since depth maps with poor quality inevitably introduce noises, we design the uncertainty injection (UI) to explore high uncertainty regions by establishing pixel-level probability distributions. We further formulate the holistic prototype contrastive (HPC) loss based on semantics and patches to learn compact and general representations across modalities and images. Finally, we employ varying supervision based on branch distinctions to implicitly construct difference modeling. Extensive experiments on two USOD datasets and four relevant underwater scene benchmarks validate the effect of the proposed method, surpassing state-of-the-art binary detection models. Impressive results on seven natural scene benchmarks further demonstrate the scalability. Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Chongyi Li, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 6 |
| 2024 | Synergistic Multiscale Detail Refinement via Intrinsic Supervision for Underwater Image EnhancementabstractVisually restoring underwater scenes primarily involves mitigating interference from underwater media. Existing methods ignore the inherent scale-related characteristics in underwater scenes. Therefore, we present the synergistic multi-scale detail refinement via intrinsic supervision (SMDR-IS) for enhancing underwater scene details, which contain multi-stages. The low-degradation stage from the original images furnishes the original stage with multi-scale details, achieved through feature propagation using the Adaptive Selective Intrinsic Supervised Feature (ASISF) module. By using intrinsic supervision, the ASISF module can precisely control and guide feature transmission across multi-degradation stages, enhancing multi-scale detail refinement and minimizing the interference from irrelevant information in the low-degradation stage. In multi-degradation encoder-decoder framework of SMDR-IS, we introduce the Bifocal Intrinsic-Context Attention Module (BICA). Based on the intrinsic supervision principles, BICA efficiently exploits multi-scale scene information in images. BICA directs higher-resolution spaces by tapping into the insights of lower-resolution ones, underscoring the pivotal role of spatial contextual relationships in underwater image restoration. Throughout training, the inclusion of a multi-degradation loss function can enhance the network, allowing it to adeptly extract information across diverse scales. When benchmarked against state-of-the-art methods, SMDR-IS consistently showcases superior performance. Our code is available at https://github.com/zhoujingchun03/SMDR-IS Dehuan Zhang, Jingchun Zhou, Chunle Guo, Weishi Zhang, Chongyi Li |
AAAI | 5 |
| 2024 | AMSP-UOD: When Vortex Convolution and Stochastic Perturbation Meet Underwater Object DetectionabstractIn this paper, we present a novel Amplitude-Modulated Stochastic Perturbation and Vortex Convolutional Network, AMSP-UOD, designed for underwater object detection. AMSP-UOD specifically addresses the impact of non-ideal imaging factors on detection accuracy in complex underwater environments. To mitigate the influence of noise on object detection performance, we propose AMSP Vortex Convolution (AMSP-VConv) to disrupt the noise distribution, enhance feature extraction capabilities, effectively reduce parameters, and improve network robustness. We design the Feature Association Decoupling Cross Stage Partial (FAD-CSP) module, which strengthens the association of long and short range features, improving the network performance in complex underwater environments. Additionally, our sophisticated post-processing method, based on non-maximum suppression with aspect-ratio similarity thresholds, optimizes detection in dense scenes, such as waterweed and schools of fish, improving object detection accuracy. Extensive experiments on the URPC and RUOD datasets demonstrate that our method outperforms existing state-of-the-art methods in terms of accuracy and noise immunity. AMSP-UOD proposes an innovative solution with the potential for real-world applications. Our code is available at https://github.com/zhoujingchun03/AMSP-UOD. Jingchun Zhou, Zongxin He, Kin-Man Lam 0001, Yudong Wang 0002, Weishi Zhang, Chunle Guo, Chongyi Li |
AAAI | 7 |
| 2024 | Learning Inclusion Matching for Animation Paint Bucket ColorizationabstractColorizing line art is a pivotal task in the production of hand-drawn cel animation. This typically involves digital painters using a paint bucket tool to manually color each segment enclosed by lines, based on RGB values predeter-mined by a color designer. This frame-by-frame process is both arduous and time-intensive. Current automated meth-ods mainly focus on segment matching. This technique mi-grates colors from a reference to the target frame by aligning features within line-enclosed segments across frames. However, issues like occlusion and wrinkles in animations often disrupt these direct correspondences, leading to mis-matches. In this work, we introduce a new learning-based inclusion matching pipeline, which directs the network to comprehend the inclusion relationships between segments rather than relying solely on direct visual correspondences. Our method features a two-stage pipeline that integrates a coarse color warping module with an inclusion matching module, enabling more nuanced and accurate colorization. To facilitate the training of our network, we also develope a unique dataset, referred to as PaintBucket-Character. This dataset includes rendered line arts alongside their colorized counterparts, featuring various 3D characters. Extensive experiments demonstrate the effectiveness and superiority of our method over existing techniques. Yuekun Dai, Shangchen Zhou, Qinyue Li, Chongyi Li, Chen Change Loy |
CVPR | 4 |
| 2024 | Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and DeblurringabstractExisting joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data, which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models, they typically rely on the assumed degradation process and cannot handle unknown real-world degradations well. To address these problems, we propose a novel zero-shot framework, FourierDiff, which embeds Fourier priors into a pre-trained diffusion model to harmoniously handle the joint degradation of luminance and structures. FourierDiff is appealing in its relaxed requirements on paired training data and degradation assumptions. The key zero-shot insight is motivated by image characteristics in the Fourier domain: most luminance information concentrates on amplitudes while structure and content information are closely related to phases. Based on this observation, we decompose the sampled results of the reverse diffusion process in the Fourier domain and take advantage of the amplitude of the generative prior to align the enhanced brightness with the distribution of natural images. To yield a sharp and content-consistent enhanced result, we further design a spatial-frequency alternating optimization strategy to progressively refine the phase of the input. Extensive experiments demonstrate the superior effectiveness of the proposed method, especially in real-world scenes. The code is available at https://github.com/aipixel/FourierDiff. Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Yichen Zheng, Bineng Zhong 0001, Chongyi Li, Liqiang Nie |
CVPR | 6 |
| 2024 | CLIB-FIQA: Face Image Quality Assessment with Confidence CalibrationabstractFace Image Quality Assessment (FIQA) is pivotal for guaranteeing the accuracy of face recognition in unconstrained environments. Recent progress in deep quality-fitting-based methods that train models to align with quality anchors, has shown promise in FIQA. However, these methods heavily depend on a recognition model to yield quality anchors and indiscriminately treat the confidence of inaccurate anchors as equivalent to that of accurate ones during the FIQA model training, leading to a fitting bottleneck issue. This paper seeks a solution by putting forward the Confidence-Calibrated Face Image Quality Assessment (CLIB-FIQA) approach, underpinned by the synergistic interplay between the quality anchors and objective quality factors such as blur, pose, expression, occlusion, and illumination. Specifically, we devise a joint learning framework built upon the vision-language alignment model, which leverages the joint distribution with multiple quality factors to facilitate the quality fitting of the FIQA model. Furthermore, to alleviate the issue of the model placing excessive trust in inaccurate quality anchors, we propose a confidence calibration method to correct the quality distribution by exploiting to the fullest extent of these objective quality factors characterized as the merged-factor distribution during training. Experimental results on eight datasets reveal the superior performance of the proposed method. Fu-Zhao Ou, Chongyi Li, Shiqi Wang 0001, Sam Kwong |
CVPR | 2 |
| 2024 | LAMP: Learn A Motion Pattern for Few-Shot Video GenerationabstractIn this paper, we present a few-shot text-to-video frame-work, LAMP, which enables a text-to-image diffusion model to Learn A specific Motion Pattern with 8 ~16 videos on a single GPU. Unlike existing methods, which re-quire a large number of training resources or learn motions that are precisely aligned with template videos, it achieves a trade-off between the degree of generation freedom and the resource costs for model training. Specifically, we design a motion-content decoupled pipeline that uses an off-the-shelf text-to-image model for content generation so that our tuned video diffusion model mainly focuses on motion learning. The well-developed text-to-image techniques can provide visually pleasing and diverse content as generation conditions, which highly improves video quality and gen-eration freedom. To capture the features of temporal di-mension, we expand the pre-trained 2D convolution lay-ers of the T2I model to our novel temporal-spatial motion learning layers and modify the attention blocks to the temporal level. Additionally, we develop an effective in-ference trick, shared-noise sampling, which can improve the stability of videos without computational costs. Our method can also be flexibly applied to other tasks, e.g. real-world image animation and video editing. Extensive ex-periments demonstrate that LAMP can effectively learn the motion pattern on limited data and generate high-quality videos. The code and models are available at https://rq-wu.github.io/projects/LAMP. Chunle Guo, Chongyi Li |
CVPR | 5 |
| 2024 | Kalman-Inspired Feature Propagation for Video Face Super-Resolution
Ruicheng Feng, Chongyi Li, Chen Change Loy |
ECCV (26) | 2 |
| 2024 | Restore Anything with Masks: Leveraging Mask Image Modeling for Blind All-in-One Image Restoration
Chu-Jie Qin, Zikun Liu 0001, Chunle Guo, Hyun Hee Park, Chongyi Li |
ECCV (45) | 7 |
| 2024 | Adaptive Window Pruning for Efficient Local Motion DeblurringabstractLocal motion blur commonly occurs in real-world photography due to the mixing between moving objects and stationary backgrounds during exposure. Existing image deblurring methods predominantly focus on global deblurring, inadvertently affecting the sharpness of backgrounds in locally blurred images and wasting unnecessary computation on sharp pixels, especially for high-resolution images.
This paper aims to adaptively and efficiently restore high-resolution locally blurred images. We propose a local motion deblurring vision Transformer (LMD-ViT) built on adaptive window pruning Transformer blocks (AdaWPT). To focus deblurring on local regions and reduce computation, AdaWPT prunes unnecessary windows, only allowing the active windows to be involved in the deblurring processes. The pruning operation relies on the blurriness confidence predicted by a confidence predictor that is trained end-to-end using a reconstruction loss with Gumbel-Softmax re-parameterization and a pruning loss guided by annotated blur masks. Our method removes local motion blur effectively without distorting sharp regions, demonstrated by its exceptional perceptual and quantitative improvements (+0.28dB) compared to state-of-the-art methods. In addition, our approach substantially reduces FLOPs by 66% and achieves more than a twofold increase in inference speed compared to Transformer-based deblurring methods. We will make our code and annotated blur masks publicly available. Haoying Li, Jixin Zhao, Shangchen Zhou, Huajun Feng, Chongyi Li, Chen Change Loy |
ICLR | 5 |
| 2024 | JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image EnhancementabstractLow-light image enhancement (LLIE) has achieved promising performance by employing conditional diffusion models. Despite the success of some conditional methods, previous methods may neglect the importance of a sufficient formulation of task-specific condition strategy, resulting in suboptimal visual outcomes. In this study, we propose JoReS-Diff, a novel approach that incorporates Retinex- and semantic-based priors as the additional pre-processing condition to regulate the generating capabilities of the diffusion model. We first leverage pre-trained decomposition network to generate the Retinex prior, which is updated with better quality by an adjustment network and integrated into a refinement network to implement Retinex-based conditional generation at both feature- and image-levels. Moreover, the semantic prior is extracted from the input image with an off-the-shelf semantic segmentation model and incorporated through semantic attention layers. By treating Retinex- and semantic-based priors as the condition, JoReS-Diff presents a unique perspective for establishing an diffusion model for LLIE and similar image enhancement tasks. Extensive experiments validate the rationality and superiority of our approach. Yuhui Wu 0001, Guoqing Wang 0001, Zhiwen Wang 0004, Yang Yang 0002, Tianyu Li 0003, Malu Zhang, Chongyi Li, Heng Tao Shen |
ACM Multimedia | 7 |
| 2024 | Generalizing ISP Model by Unsupervised Raw-to-raw MappingabstractISP (Image Signal Processor) serves as a pipeline converting unprocessed raw images to sRGB images, positioned before nearly all visual tasks. Due to the varying spectral sensitivities of cameras, raw images captured by different cameras exist in different color spaces, making it challenging to deploy ISP across cameras with consistent performance. To address this challenge, it is intuitively to incorporate a raw-to-raw mapping (mapping raw images across camera color spaces) module into the ISP. However, the lack of paired data (i.e., images of the same scene captured by different cameras) makes it difficult to train a raw-to-raw model using supervised learning methods. In this paper, we aim to achieve ISP generalization by proposing the first unsupervised raw-to-raw model. To be specific, we propose a CSTPP (Color Space Transformation Parameters Predictor) module to predict the space transformation parameters in a patch-wise manner, which can accurately perform color space transformation and flexibly manage complex lighting conditions. Additionally, we design a CycleGAN-style training framework to realize unsupervised learning, overcoming the deficiency of paired data. Our proposed unsupervised model achieved performance comparable to that of the state-of-the-art semi-supervised method in raw-to-raw task. Furthermore, to assess its ability to generalize the ISP model across different cameras, we for the first formulated cross-camera ISP task and demonstrated the performance of our method through extensive experiments. The codes are released at https://github.com/ydxxxx/Unsupervised-Raw-to-raw-Mapping. Dongyu Xie, Chaofan Qiao, Lanyue Liang, Zhiwen Wang 0004, Tianyu Li 0003, Qiao Liu 0003, Chongyi Li, Guoqing Wang 0001, Yang Yang 0002 |
ACM Multimedia | 7 |
| 2024 | Lighting Every Darkness with 3DGS: Fast Training and Real-Time Rendering for HDR View SynthesisabstractVolumetric rendering-based methods, like NeRF, excel in HDR view synthesis from RAW images, especially for nighttime scenes. They suffer from long training times and cannot perform real-time rendering due to dense sampling requirements. The advent of 3D Gaussian Splatting (3DGS) enables real-time rendering and faster training. However, implementing RAW image-based view synthesis directly using 3DGS is challenging due to its inherent drawbacks: 1) in nighttime scenes, extremely low SNR leads to poor structure-from-motion (SfM) estimation in dis- tant views; 2) the limited representation capacity of the spherical harmonics (SH) function is unsuitable for RAW linear color space; and 3) inaccurate scene structure hampers downstream tasks such as refocusing. To address these issues, we propose LE3D (Lighting Every darkness with 3DGS). Our method proposes Cone Scatter Initialization to enrich the estimation of SfM and replaces SH with a Color MLP to represent the RAW linear color space. Additionally, we introduce depth distortion and near-far regularizations to improve the accuracy of scene structure for down- stream tasks. These designs enable LE3D to perform real-time novel view synthesis, HDR rendering, refocusing, and tone-mapping changes. Compared to previous vol- umetric rendering-based methods, LE3D reduces training time to 1% and improves rendering speed by up to 4,000 times for 2K resolution images in terms of FPS. Code and viewer can be found in https://srameo.github.io/projects/le3d. Xin Jin 0005, Pengyi Jiao, Zheng-Peng Duan, Xingchao Yang, Chongyi Li, Chunle Guo, Bo Ren 0003 |
NeurIPS | 5 |
| 2024 | Transformer with large convolution kernel decoder network for salient object detection in optical remote sensing images
Pengwei Dong, Bo Wang 0070, Runmin Cong, Hai-Han Sun, Chongyi Li |
Comput. Vis. Image Underst. | 5 |
| 2024 | HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image EnhancementabstractUnderwater image enhancement presents a significant challenge due to the complex and diverse underwater environments that result in severe degradation phenomena such as light absorption, scattering, and color distortion. More importantly, obtaining paired training data for these scenarios is a challenging task, which further hinders the generalization performance of enhancement models. To address these issues, we propose a novel approach, the Hybrid Contrastive Learning Regularization (HCLR-Net). Our method is built upon a distinctive hybrid contrastive learning regularization strategy that incorporates a unique methodology for constructing negative samples. This approach enables the network to develop a more robust sample distribution. Notably, we utilize non-paired data for both positive and negative samples, with negative samples are innovatively reconstructed using local patch perturbations. This strategy overcomes the constraints of relying solely on paired data, boosting the model’s potential for generalization. The HCLR-Net also incorporates an Adaptive Hybrid Attention module and a Detail Repair Branch for effective feature extraction and texture detail restoration, respectively. Comprehensive experiments demonstrate the superiority of our method, which shows substantial improvements over several state-of-the-art methods in terms of quantitative metrics, significantly enhances the visual quality of underwater images, establishing its innovative and practical applicability. Our code is available at: https://github.com/zhoujingchun03/HCLR-Net . Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 3 |
| 2024 | Correction: HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image Enhancement
Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 3 |
| 2024 | CVANet: Cascaded visual attention network for single image super-resolution
Weidong Zhang 0007, Wenyi Zhao, Jia Li 0019, Peixian Zhuang, Hai-Han Sun, Chongyi Li |
Neural Networks | 7 |
| 2024 | Flare7K++: Mixing Synthetic and Real Datasets for Nighttime Flare Removal and BeyondabstractArtificial lights commonly leave strong lens flare artifacts on the images captured at night, degrading both the visual quality and performance of vision algorithms. Existing flare removal approaches mainly focus on removing daytime flares and fail in nighttime cases. Nighttime flare removal is challenging due to the unique luminance and spectrum of artificial lights, as well as the diverse patterns and image degradation of the flares. The scarcity of the nighttime flare removal dataset constrains the research on this crucial task. In this paper, we introduce Flare7K++, the first comprehensive nighttime flare removal dataset, consisting of 962 real-captured flare images (Flare-R) and 7000 synthetic flares (Flare7K). Compared to Flare7K, Flare7K++ is particularly effective in eliminating complicated degradation around the light source, which is intractable by using synthetic flares alone. Besides, the previous flare removal pipeline relies on the manual threshold and blur kernel settings to extract light sources, which may fail when the light sources are tiny or not overexposed. To address this issue, we additionally provide the annotations of light sources in Flare7K++ and propose a new end-to-end pipeline to preserve the light source while removing lens flares. Our dataset and pipeline offer a valuable foundation and benchmark for future investigations into nighttime flare removal studies. Extensive experiments demonstrate that Flare7K++ supplements the diversity of existing flare datasets and pushes the frontier of nighttime flare removal toward real-world scenarios. Yuekun Dai, Chongyi Li, Shangchen Zhou, Ruicheng Feng, Yihang Luo, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Towards a Flexible Semantic Guided Model for Single Image Enhancement and RestorationabstractLow-light image enhancement (LLIE) investigates how to improve the brightness of an image captured in illumination-insufficient environments. The majority of existing methods enhance low-light images in a global and uniform manner, without taking into account the semantic information of different regions. Consequently, a network may easily deviate from the original color of local regions. To address this issue, we propose a semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that adaptively integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in the LLIE task. We further present a refined framework SKF++ with two new techniques: (a) Extra convolutional branch for intra-class illumination and color recovery through extracting local information and (b) Equalization-based histogram transformation for contrast enhancement and high dynamic range adjustment. Extensive experiments on various benchmarks of LLIE task and other image processing tasks show that models equipped with the SKF/SKF++ significantly outperform the baselines and our SKF/SKF++ generalizes to different models and scenes well. Besides, the potential benefits of our method in face detection and semantic segmentation in low-light conditions are discussed. Yuhui Wu 0001, Guoqing Wang 0001, Shaochong Liu, Yang Yang 0002, Wei Liu 0005, Xiongxin Tang, Shuhang Gu, Chongyi Li, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Non-Uniform Illumination Underwater Image Restoration via Illumination Channel Sparsity PriorabstractUnderwater image quality is seriously degraded due to the insufficient light in water. Although artificial illumination can assist imaging, it often brings non-uniform illumination phenomenon. To this end, we develop an illumination channel sparsity prior (ICSP) guided variational framework for non-uniform illumination underwater image restoration. Technically, the illumination channel sparsity prior is built on the observation that the illumination channel of a uniform-light underwater image in HSI color space contains few pixels whose intensity is very low. Then according to the Retinex theory, we design a variational model with L0 norm term, constraint term, and gradient term, by integrating the proposed ICSP into an extended underwater image formation model. Such three regularizations are effective in enhancing the brightness, correcting color distortion, and revealing structures and fine-scale details. Meanwhile, we exploit a fast numerical algorithm on the base of the alternating direction method of multipliers (ADMM) to accelerate solving this optimization problem. We also collect a benchmark dataset, namely NUID that contains 925 real underwater images of different non-uniform illumination. Extensive experiments demonstrate that our proposed method is effective in terms of qualitative and quantitative comparisons, ablation studies, convergence analysis, and applications. The code and dataset are available athttps://github.com/Hou-Guojia/ICSP. Guojia Hou, Peixian Zhuang, Kunqian Li, Hai-Han Sun, Chongyi Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Underwater Image Quality Improvement via Color, Detail, and Contrast RestorationabstractDue to the complex imaging mechanism, underwater images often suffer from multiple degradation issues, such as color cast, blurry detail, and low contrast, which affect the extraction of valuable information. To deal with these degradation issues, a simple yet effective underwater image quality improvement method based on color, detail and contrast restoration (CDCR) is developed, which consists of three key modules: a well-preserved finding-driven color balance module (CBM), a linear saturation transformation-based discriminant function-based detail restoration module (DRM), and a transmission minimization-oriented contrast restoration module (CRM). First, the CBM explores a well-preserved channel finding and employs a channel compensation strategy to balance the color differences among three color channels. Second, the DRM uses a piecewise underwater image saturation estimation strategy, which takes the various spectral properties of water into account and designs an additional linear saturation transformation-based discriminant function to prevent the transmission from being under-estimated. At last, the CRM estimates a global backscatter light based on transmission minimization and further improves the contrast by locally removing the backscatter light of the base layer. Our restored image is appealing in its natural color, fine details, and high contrast. Extensive experiments on three underwater image enhancement datasets show that our CDCR achieves better results than state-of-the-art methods, i.e., compared with the second-best method, the average PCQI and UIQM values of our method increase by 5.7% and 0.2%, and the average Blur and DFAD values of our method decrease by 8.0% and 5.3%. Meanwhile, experiments further suggest that the rate of new visible edges and the quality of contrast restoration of our CDCR at least increase by 7.7% and 51.2% in most tested sandstorm and foggy images, respectively, which demonstrates that our method has a good generalization capability for sandstorm and foggy image restoration. Zheng Liang 0001, Weidong Zhang 0007, Rui Ruan, Peixian Zhuang, Xiwang Xie, Chongyi Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Salient Object Detection Toward Single-Pixel ImagingabstractReplacing CCD and CMOS image sensors in conventional cameras with digital micromirror devices (DMD), single-pixel cameras low-costly shot images by capturing compressed measurements and computation. However, the compressed measurements lack explicit spatial information, causing difficulties for high-level tasks such as salient object detection (SOD) that are usually designed to have visual inputs. To address the issue, we propose a single-pixel imaging-based SOD network called SPISODNet that enables predicting saliency maps directly from compressed measurements with high accuracy. Specifically, we first design an underlying feature inversion module (UFIM) to capture the underlying scene information, and then develop a context-aware flow (CAF) consisting of a feature focus module (FFM), three bidirectional attention modules (BAMs), and a spatial information-induced attention module (SIAM) to acquire and polish saliency predictions. Extensive experiments demonstrate that our method achieves superior performance for single-pixel imaging-based SOD. HuiHui Yue, Jichang Guo, Xiangjun Yin, Yi Zhang 0107, Bihan Wen, Chongyi Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Underwater Image Enhancement via Weighted Wavelet Visual Perception FusionabstractUnderwater images typically suffer from various quality degradation issues due to the scattering and absorption of light, but these degraded-quality underwater images are unbeneficial for analysis and applications. To effectively solve these quality degradation issues, an underwater image enhancement method via weighted wavelet visual perception fusion is introduced, called WWPF. Concretely, we first present an attenuation-map-guided color correction strategy to correct the color distortion of an underwater image. Subsequently, we employ the maximum information entropy optimized global contrast strategy to the color-corrected image to obtain a global contrast-enhanced image. Meanwhile, we apply a fast integration optimized local contrast strategy to the color-corrected image to get a local contrast-enhanced image. To exploit the complementary of the global contrast-enhanced image and the local contrast-enhanced image, we introduce a weighted wavelet visual perception fusion strategy to obtain a high-quality underwater image by fusing the high-frequency and low-frequency components of images at different scales. Our extensive experiments on three benchmarks validate that our WWPF outperforms the state-of-the-art methods in qualitative and quantitative. Besides, the underwater images processed by our WWPF also benefit practical underwater applications. The code is availablehttps://github.com/Li-Chongyi/WWPF_code. Weidong Zhang 0007, Ling Zhou 0003, Peixian Zhuang, Guohou Li, Xipeng Pan, Wenyi Zhao, Chongyi Li |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | GACNet: Generate Adversarial-Driven Cross-Aware Network for Hyperspectral Wheat Variety IdentificationabstractWheat variety identification from hyperspectral images holds significant importance in both fine breeding and intelligent agriculture. However, the discriminatory accuracy of some techniques is limited due to insufficient datasets, data redundancy, and noise interference. To address these issues, we propose a wheat variety identification framework called generate adversarial-driven cross-aware network (GACNet), comprising a semi-supervised generative adversarial network (GAN) for data augmentation and a cross-aware attention network (CAANet) for variety identification. First, the semi-supervised GAN (SSGAN) alleviates data scarcity by generating fake hyperspectral images as realistically as possible through learning the distribution hypothesis of real hyperspectral images, while the discriminator distinguishes between real and fake hyperspectral images. Subsequently, the CAANet is employed for wheat variety identification, which leverages a cascading cross-learning of 3-D and 2-D convolutions to fully exploit spectral, spatial, and texture features and refines the features through an embedded attention mechanism in the cross-convolutional module. Additionally, we constructed a hyperspectral wheat variety dataset (HWVD) comprising 4560 samples of 19 categories. Extensive experiments on our dataset demonstrate that our GACNet outperforms state-of-the-art methods for wheat variety identification. The HWVD will be made available. Weidong Zhang 0007, Guohou Li, Peixian Zhuang, Guojia Hou, Qiang Zhang 0011, Chongyi Li |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Perception-Driven Deep Underwater Image Enhancement Without Paired SupervisionabstractUnderwater image enhancement (UIE) aims to improve the visual quality of raw underwater images. Current UIE algorithms primarily train a deep neural network (DNN) on synthetic datasets or datasets with pseudo labels by minimizing the reconstruction loss between enhanced images and ground truth images. However, there is a domain gap between synthetic and real-world underwater images, and the widely used$\ell _{1}$or$\ell _{2}$loss tends to overlook the importance of human perception, resulting in unsatisfactory perceptual quality of the final enhanced results. In this paper, we propose an unsupervised perception-driven DNN called PDD-Net for generalizable UIE. Instead of relying on paired images for training, we resort to an unsupervised generative adversarial network (GAN) with a large-scale set of easily available natural images as the target domain. This enables training on larger image sets collected from various domains while avoiding over-fitted to any specific data generation protocol. Additionally, to make the visual quality of enhanced underwater images more in line with human perception, we pre-train a DNN-based pairwise quality ranking (PQR) model based on which a PQR loss is formulated to progressively guides the enhancement of raw underwater image toward the higher quality direction. In addition, we introduce a global attention module (GAM) that integrates modulation and attention mechanisms to enable capturing rich global and local information, leading to improvements in both brightness and contrast. Extensive experiments demonstrate that our proposed PDD-Net exhibits excellent generalization capabilities and outperforms existing methods in terms of both visual perception quality and quantitative indicators across different datasets. Qiuping Jiang, Yaozu Kang, Zhihua Wang 0002, Wenqi Ren, Chongyi Li |
IEEE Trans. Multim. | 5 |
| 2024 | Rethinking Pan-Sharpening in Closed-Loop RegularizationabstractIt is generally known that pan-sharpening is fundamentally a PAN-guided multispectral (MS) image super-resolution problem that involves learning the nonlinear mapping from low-resolution (LR) to high-resolution (HR) MS images. Since an infinite number of HR-MS images can be downsampled to produce the same corresponding LR-MS image, learning the mapping from LR-MS to HR-MS image is typically ill-posed and the space of the possible pan-sharpening functions can be extremely large, making it difficult to estimate the optimal mapping solution. To address the above issue, we propose a closed-loop scheme that learns the two opposite mapping including the pan-sharpening and its corresponding degradation process simultaneously to regularize the solution space in a single pipeline. More specifically, an invertible neural network (INN) is introduced to perform a bidirectional closed-loop: the forward operation for LR-MS pan-sharpening and the backward operation for learning the corresponding HR-MS image degradation process. In addition, given the vital importance of high-frequency textures for the Pan-sharpened MS images, we further strengthen the INN by designing a specified multiscale high-frequency texture extraction module. Extensive experimental results demonstrate that the proposed algorithm performs favorably against state-of-the-art methods qualitatively and quantitatively with fewer parameters. Ablation studies also verify the effectiveness of the closed-loop mechanism in pan-sharpening. The source code is made publicly available at https://github.com/manman1995/pan-sharpening-Team-zhouman/. Man Zhou 0003, Jie Huang 0017, Danfeng Hong, Feng Zhao 0004, Chongyi Li, Jocelyn Chanussot |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Underwater Ranker: Learn Which Is Better and How to Be BetterabstractIn this paper, we present a ranking-based underwater image quality assessment (UIQA) method, abbreviated as URanker. The URanker is built on the efficient conv-attentional image Transformer. In terms of underwater images, we specially devise (1) the histogram prior that embeds the color distribution of an underwater image as histogram token to attend global degradation and (2) the dynamic cross-scale correspondence to model local degradation. The final prediction depends on the class tokens from different scales, which comprehensively considers multi-scale dependencies. With the margin ranking loss, our URanker can accurately rank the order of underwater images of the same scene enhanced by different underwater image enhancement (UIE) algorithms according to their visual quality. To achieve that, we also contribute a dataset, URankerSet, containing sufficient results enhanced by different UIE algorithms and the corresponding perceptual rankings, to train our URanker. Apart from the good performance of URanker, we found that a simple U-shape UIE network can obtain promising performance when it is coupled with our pre-trained URanker as additional supervision. In addition, we also propose a normalization tail that can significantly improve the performance of UIE networks. Extensive experiments demonstrate the state-of-the-art performance of our method. The key designs of our method are discussed. Our code and dataset are available at https://li-chongyi.github.io/URanker_files/. Chunle Guo, Xin Jin 0005, Linghao Han, Weidong Zhang 0007, Chongyi Li |
AAAI | 7 |
| 2023 | Nighttime Smartphone Reflective Flare Removal Using Optical Center Symmetry PriorabstractReflective flare is a phenomenon that occurs when light reflects inside lenses, causing bright spots or a “ghosting effect” in photos, which can impact their quality. Eliminating reflective flare is highly desirable but challenging. Many existing methods rely on manually designed features to detect these bright spots, but they often fail to identify reflective flares created by various types of light and may even mistakenly remove the light sources in scenarios with multiple light sources. To address these challenges, we propose an optical center symmetry prior, which suggests that the reflective flare and light source are always symmetrical around the lens's optical center. This prior helps to locate the reflective flare's proposal region more accurately and can be applied to most smartphone cameras. Building on this prior, we create the first reflective flare removal dataset called BracketFlare, which contains diverse and realistic reflective flare patterns. We use continuous bracketing to capture the reflective flare pattern in the underexposed image and combine it with a normally exposed image to synthesize a pair of flare-corrupted and flare-free images. With the dataset, neural networks can be trained to remove the reflective flares effectively. Extensive experiments demonstrate the effectiveness of our method on both synthetic and real-world datasets. Yuekun Dai, Yihang Luo, Shangchen Zhou, Chongyi Li, Chen Change Loy |
CVPR | 4 |
| 2023 | Generating Aligned Pseudo-Supervision from Non-Aligned Data for Image Restoration in Under-Display CameraabstractDue to the difficulty in collecting large-scale and perfectly aligned paired training data for Under-Display Camera (UDC) image restoration, previous methods resort to monitor-based image systems or simulation-based methods, sacrificing the realness of the data and introducing domain gaps. In this work, we revisit the classic stereo setup for training data collection – capturing two images of the same scene with one UDC and one standard camera. The key idea is to “copy” details from a high-quality reference image and “paste” them on the UDC image. While being able to generate real training pairs, this setting is susceptible to spatial misalignment due to perspective and depth of field changes. The problem is further compounded by the large domain discrepancy between the UDC and normal images, which is unique to UDC restoration. In this paper, we mitigate the non-trivial domain discrepancy and spatial misalignment through a novel Transformer-based framework that generates well-aligned yet high-quality target data for the corresponding UDC input. This is made possible through two carefully designed components, namely, the Domain Alignment Module (DAM) and Geometric Alignment Module (GAM), which encourage robust and accurate discovery of correspondence between the UDC and normal views. Extensive experiments show that high-quality and well-aligned pseudo UDC training pairs are beneficial for training a robust restoration network. Code and the dataset are available at https://github.com/jnjaby/AlignFormer. Ruicheng Feng, Chongyi Li, Huaijin G. Chen, Jinwei Gu, Chen Change Loy |
CVPR | 2 |
| 2023 | DNF: Decouple and Feedback Network for Seeing in the DarkabstractThe exclusive properties of RAW data have shown great potential for low-light image enhancement. Nevertheless, the performance is bottlenecked by the inherent limitations of existing architectures in both single-stage and multi-stage methods. Mixed mapping across two different domains, noise-to-clean and RAW-to-sRGB, misleads the single-stage methods due to the domain ambiguity. The multi-stage methods propagate the information merely through the resulting image of each stage, neglecting the abundant features in the lossy image-level dataflow. In this paper, we probe a generalized solution to these bottlenecks and propose a Decouple aNd Feedback framework, abbreviated as DNF. To mitigate the domain ambiguity, domain-specific subtasks are decoupled, along with fully utilizing the unique properties in RAW and sRGB domains. The feature propagation across stages with a feedback mechanism avoids the information loss caused by image-level dataflow. The two key insights of our method resolve the inherent limitations of RAW data-based low-light image enhancement satisfactorily, empowering our method to outperform the previous state-of-the-art method by a large margin with only 19% parameters, achieving 0.97dB and 1.30dB PSNR improvements on the Sony and Fuji subsets of SID. Xin Jin 0005, Linghao Han, Zhen Li 0031, Chunle Guo, Chongyi Li |
CVPR | 6 |
| 2023 | RIDCP: Revitalizing Real Image Dehazing via High-Quality Codebook PriorsabstractExisting dehazing approaches struggle to process real-world hazy images owing to the lack of paired real data and robust priors. In this work, we present a new paradigm for real image dehazing from the perspectives of synthesizing more realistic hazy data and introducing more robust priors into the network. Specifically, (1) instead of adopting the de facto physical scattering model, we rethink the degradation of real hazy images and propose a phenomenological pipeline considering diverse degradation types. (2) We propose a Real Image Dehazing network via high-quality Codebook Priors (RIDCP). Firstly, a VQGAN is pre-trained on a large-scale high-quality dataset to obtain the discrete codebook, encapsulating high-quality priors (HQPs). After replacing the negative effects brought by haze with HQPs, the decoder equipped with a novel normalized feature alignment module can effectively utilize high-quality features and produce clean results. However, although our degradation pipeline drastically mitigates the domain gap between synthetic and real data, it is still intractable to avoid it, which challenges HQPs matching in the wild. Thus, we recalculate the distance when matching the features to the HQPs by a controllable matching operation, which facilitates finding better counterparts. We provide a recommendation to control the matching based on an explainable solution. Users can also flexibly adjust the enhancement degree as per their preference. Extensive experiments verify the effectiveness of our data synthesis pipeline and the superior performance of RIDCP in real image dehazing. Code and data are released at https://rqwu.github.io/projects/RIDCP. Zheng-Peng Duan, Chunle Guo, Chongyi Li |
CVPR | 5 |
| 2023 | Learning Semantic-Aware Knowledge Guidance for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) investigates how to improve illumination and produce normal-light images. The majority of existing methods improve low-light images via a global and uniform manner, without taking into account the semantic information of different regions. Without semantic priors, a network may easily deviate from a region's original color. To address this issue, we propose a novel semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that wisely integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in LLIE task. Extensive experiments show that models equipped with the SKF significantly outperform the baselines on multiple datasets and our SKF generalizes to different models and scenes well. The code is available at Semantic-Aware-Low-Light-Image-Enhancement. Yuhui Wu 0001, Guoqing Wang 0001, Yang Yang 0002, Jiwei Wei, Chongyi Li, Heng Tao Shen |
CVPR | 6 |
| 2023 | Lighting Every Darkness in Two Pairs : A Calibration-Free Pipeline for RAW DenoisingabstractCalibration-based methods have dominated RAW image denoising under extremely low-light environments. However, these methods suffer from several main deficiencies: 1) the calibration procedure is laborious and time-consuming, 2) denoisers for different cameras are difficult to transfer, and 3) the discrepancy between synthetic noise and real noise is enlarged by high digital gain. To overcome the above shortcomings, we propose a calibration-free pipeline for Lighting Every Drakness (LED), regardless of the digital gain or camera sensor. Instead of calibrating the noise parameters and training repeatedly, our method could adapt to a target camera only with few-shot paired data and fine-tuning. In addition, well-designed structural modification during both stages alleviates the domain gap between synthetic and real noise without any extra computational cost. With 2 pairs for each additional digital gain (in total 6 pairs) and 0.5% iterations, our method achieves superior performance over other calibration-based methods. Xin Jin 0005, Jia-Wen Xiao, Linghao Han, Chunle Guo, Ruixun Zhang, Xialei Liu, Chongyi Li |
ICCV | 7 |
| 2023 | Iterative Prompt Learning for Unsupervised Backlit Image EnhancementabstractWe propose a novel unsupervised backlit image enhancement method, abbreviated as CLIP-LIT, by exploring the potential of Contrastive Language-Image Pre-Training (CLIP) for pixel-level image enhancement. We show that the openworld CLIP prior not only aids in distinguishing between backlit and well-lit images, but also in perceiving heterogeneous regions with different luminance, facilitating the optimization of the enhancement network. Unlike high-level and image manipulation tasks, directly applying CLIP to enhancement tasks is non-trivial, owing to the difficulty in finding accurate prompts. To solve this issue, we devise a prompt learning framework that first learns an initial prompt pair by constraining the text-image similarity between the prompt (negative/positive sample) and the corresponding image (backlit image/well-lit image) in the CLIP latent space. Then, we train the enhancement network based on the textimage similarity between the enhanced result and the initial prompt pair. To further improve the accuracy of the initial prompt pair, we iteratively fine-tune the prompt learning framework to reduce the distribution gaps between the backlit images, enhanced results, and well-lit images via rank learning, boosting the enhancement performance. Our method alternates between updating the prompt learning framework and enhancement network until visually pleasing results are achieved. Extensive experiments demonstrate that our method outperforms state-of-the-art methods in terms of visual quality and generalization ability, without requiring any paired data. Zhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng, Chen Change Loy |
ICCV | 2 |
| 2023 | Troubleshooting Ethnic Quality Bias with Curriculum Domain Adaptation for Face Image Quality AssessmentabstractFace Image Quality Assessment (FIQA) lays the foundation for ensuring the stability and accuracy of face recognition systems. However, existing FIQA methods mainly formulate quality relationships within the training set to yield quality scores, ignoring the generalization problem caused by ethnic quality bias between the training and test sets. Domain adaptation presents a potential solution to mitigate the bias, but if FIQA is treated essentially as a regression task, it will be limited by the challenge of feature scaling in transfer learning. Additionally, how to guarantee source risk is also an issue due to the lack of ground-truth labels of the source domain for FIQA. This paper presents the first attempt in the field of FIQA to address these challenges with a novel Ethnic-Quality-Bias Mitigating (EQBM) framework. Specifically, to eliminate the restriction of scalar regression, we first compute the Likert-scale quality probability distributions as source domain annotations. Furthermore, we design an easy-to-hard training scheduler based on the inter-domain uncertainty and intra-domain quality margin as well as the ranking-based domain adversarial network to enhance the effectiveness of transfer learning and further reduce the source risk in domain adaptation. Extensive experiments demonstrate that the EQBM significantly mitigates the quality bias and improves the generalization capability of FIQA across races on different datasets. Fu-Zhao Ou, Baoliang Chen, Chongyi Li, Shiqi Wang 0001, Sam Kwong |
ICCV | 3 |
| 2023 | Empowering Low-Light Image Enhancer through Customized Learnable PriorsabstractDeep neural networks have achieved remarkable progress in enhancing low-light images by improving their brightness and eliminating noise. However, most existing methods construct end-to-end mapping networks heuristically, neglecting the intrinsic prior of image enhancement task and lacking transparency and interpretability. Although some unfolding solutions have been proposed to relieve these issues, they rely on proximal operator networks that deliver ambiguous and implicit priors. In this work, we propose a paradigm for low-light image enhancement that explores the potential of customized learnable priors to improve the transparency of the deep unfolding paradigm. Motivated by the powerful feature representation capability of Masked Autoencoder (MAE), we customize MAE-based illumination and noise priors and redevelop them from two perspectives: 1) structure flow: we train the MAE from a normal-light image to its illumination properties and then embed it into the proximal operator design of the unfolding architecture; and 2) optimization flow: we train MAE from a normal-light image to its gradient representation and then employ it as a regularization term to constrain noise in the model output. These designs improve the interpretability and representation capability of the model. Extensive experiments on multiple low-light image enhancement datasets demonstrate the superiority of our proposed paradigm over state-of-the-art methods. Code is available at https://github.com/zheng980629/CUE. Naishan Zheng, Man Zhou 0003, Yanmeng Dong, Xiangyu Rui, Jie Huang 0017, Chongyi Li, Feng Zhao 0004 |
ICCV | 6 |
| 2023 | Learned Image Reasoning Prior Penetrates Deep Unfolding Network for Panchromatic and Multi-Spectral Image FusionabstractThe success of deep neural networks for pan-sharpening is commonly in a form of black box, lacking transparency and interpretability. To alleviate this issue, we propose a novel model-driven deep unfolding framework with image reasoning prior tailored for the pan-sharpening task. Different from existing unfolding solutions that deliver the proximal operator networks as the uncertain and vague priors, our framework is motivated by the content reasoning ability of masked autoencoders (MAE) with insightful designs. Specifically, the pre-trained MAE with spatial masking strategy, acting as intrinsic reasoning prior, is embedded into unfolding architecture. Meanwhile, the pre-trained MAE with spatial-spectral masking strategy is treated as the regularization term within loss function to constrain the spatial-spectral consistency. Such designs penetrate the image reasoning prior into deep unfolding networks while improving its interpretability and representation capability. The uniqueness of our framework is that the holistic learning process is explicitly integrated with the inherent physical mechanism underlying the pan-sharpening task. Extensive experiments on multiple satellite datasets demonstrate the superiority of our method over the existing state-of-the-art approaches. Code will be released at https://manman1995.github.io/. Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Chongyi Li |
ICCV | 4 |
| 2023 | Improving Lens Flare Removal with General-Purpose Pipeline and Multiple Light Sources RecoveryabstractWhen taking images against strong light sources, the resulting images often contain heterogeneous flare artifacts. These artifacts can importantly affect image visual quality and downstream computer vision tasks. While collecting real data pairs of flare-corrupted/flare-free images for training flare removal models is challenging, current methods utilize the direct-add approach to synthesize data. However, these methods do not consider automatic exposure and tone mapping in image signal processing pipeline (ISP), leading to the limited generalization capability of deep models training using such data. Besides, existing methods struggle to handle multiple light sources due to the different sizes, shapes and illuminance of various light sources. In this paper, we propose a solution to improve the performance of lens flare removal by revisiting the ISP and remodeling the principle of automatic exposure in the synthesis pipeline and design a more reliable light sources recovery strategy. The new pipeline approaches realistic imaging by discriminating the local and global illumination through convex combination, avoiding global illumination shifting and local over-saturation. Our strategy for recovering multiple light sources convexly averages the input and output of the neural network based on illuminance levels, thereby avoiding the need for a hard threshold in identifying light sources. We also contribute a new flare removal testing dataset containing the flare-corrupted images captured by ten types of consumer electronics. The dataset facilitates the verification of the generalization capability of flare removal methods. Extensive experiments show that our solution can effectively improve the performance of lens flare removal and push the frontier toward more general situations. Yuyan Zhou, Dong Liang 0008, Songcan Chen, Sheng-Jun Huang, Chongyi Li |
ICCV | 6 |
| 2023 | ProPainter: Improving Propagation and Transformer for Video InpaintingabstractFlow-based propagation and spatiotemporal Transformer are two mainstream mechanisms in video inpainting (VI). Despite the effectiveness of these components, they still suffer from some limitations that affect their performance. Previous propagation-based approaches are performed separately either in the image or feature domain. Global image propagation isolated from learning may cause spatial misalignment due to inaccurate optical flow. Moreover, memory or computational constraints limit the temporal range of feature propagation and video Transformer, preventing exploration of correspondence information from distant frames. To address these issues, we propose an improved framework, called ProPainter, which involves enhanced ProPagation and an efficient Transformer. Specifically, we introduce dual-domain propagation that combines the advantages of image and feature warping, exploiting global correspondences reliably. We also propose a mask-guided sparse video Transformer, which achieves high efficiency by discarding unnecessary and redundant tokens. With these components, ProPainter outperforms prior arts by a large margin of 1.46 dB in PSNR while maintaining appealing efficiency. Shangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change Loy |
ICCV | 2 |
| 2023 | Exploring Temporal Frequency Spectrum in Deep Video DeblurringabstractVideo deblurring aims to restore the latent video frames from their blurred counterparts. Despite the remarkable progress, most promising video deblurring methods only investigate the temporal priors in the spatial domain and rarely explore their its potential in the frequency domain. In this paper, we revisit the blurred sequence in the Fourier space and figure out some intrinsic frequency-temporal priors that imply the temporal blur degradation can be accessibly decoupled in the potential frequency domain. Based on these priors, we propose a novel Fourier-based frequency-temporal video deblurring solution, where the core design accommodates the temporal spectrum to a popular video deblurring pipeline of feature extraction, alignment, aggregation, and optimization. Specifically, we design a Spectrum Prior-guided Alignment module by leveraging enlarged blur information in the potential spectrum to mitigate the blur effects on the alignment. Then, Temporal Energy prior-driven Aggregation is implemented to replenish the original local features by estimating the temporal spectrum energy as the global sharpness guidance. In addition, the customized frequency loss is devised to optimize the proposed method for decent spectral distribution. Extensive experiments demonstrate that our model performs favorably against other state-of-the-art methods, thus confirming the effectiveness of frequency-temporal prior modeling. Qi Zhu 0010, Man Zhou 0003, Naishan Zheng, Chongyi Li, Jie Huang 0017, Feng Zhao 0004 |
ICCV | 4 |
| 2023 | Embedding Fourier for Ultra-High-Definition Low-Light Image Enhancement
Chongyi Li, Chunle Guo, Man Zhou 0003, Zhexin Liang, Shangchen Zhou, Ruicheng Feng, Chen Change Loy |
ICLR | 1 |
| 2023 | Fourmer: An Efficient Global Modeling Paradigm for Image RestorationabstractGlobal modeling-based image restoration frameworks have become popular. However, they often require a high memory footprint and do not consider task-specific degradation. Our work presents an alternative approach to global modeling that is more efficient for image restoration. The key insights which motivate our study are two-fold: 1) Fourier transform is capable of disentangling image degradation and content component to a certain extent, serving as the image degradation prior, and 2) Fourier domain innately embraces global properties, where each pixel in the Fourier space is involved with all spatial pixels. While adhering to the ``spatial interaction + channel evolution'' rule of previous studies, we customize the core designs with Fourier spatial interaction modeling and Fourier channel evolution. Our paradigm, Fourmer, achieves competitive performance on common image restoration tasks such as image de-raining, image enhancement, image dehazing, and guided image super-resolution, while requiring fewer computational resources. The code for Fourmer will be made publicly available. Man Zhou 0003, Jie Huang 0017, Chunle Guo, Chongyi Li |
ICML | 4 |
| 2023 | Transition-constant Normalization for Image EnhancementabstractNormalization techniques that capture image style by statistical representation have become a popular component in deep neural networks.
Although image enhancement can be considered as a form of style transformation, there has been little exploration of how normalization affect the enhancement performance.
To fully leverage the potential of normalization, we present a novel Transition-Constant Normalization (TCN) for various image enhancement tasks.
Specifically, it consists of two streams of normalization operations arranged under an invertible constraint, along with a feature sub-sampling operation that satisfies the normalization constraint.
TCN enjoys several merits, including being parameter-free, plug-and-play, and incurring no additional computational costs.
We provide various formats to utilize TCN for image enhancement, including seamless integration with enhancement networks, incorporation into encoder-decoder architectures for downsampling, and implementation of efficient architectures.
Through extensive experiments on multiple image enhancement tasks, like low-light enhancement, exposure correction, SDR2HDR translation, and image dehazing, our TCN consistently demonstrates performance improvements.
Besides, it showcases extensive ability in other tasks including pan-sharpening and medical segmentation.
The code is available at \textit{\textcolor{blue}{https://github.com/huangkevinj/TCNorm}}. Jie Huang 0017, Man Zhou 0003, Mingde Yao, Chongyi Li, Zhiwei Xiong, Feng Zhao 0004 |
NeurIPS | 6 |
| 2023 | Training Your Image Restoration Network Better with Random Weight Network as Optimization FunctionabstractThe blooming progress made in deep learning-based image restoration has been largely attributed to the availability of high-quality, large-scale datasets and advanced network structures. However, optimization functions such as L_1 and L_2 are still de facto. In this study, we propose to investigate new optimization functions to improve image restoration performance. Our key insight is that ``random weight network can be acted as a constraint for training better image restoration networks''. However, not all random weight networks are suitable as constraints. We draw inspiration from Functional theory and show that alternative random weight networks should be represented in the form of a strict mathematical manifold. We explore the potential of our random weight network prototypes that satisfy this requirement: Taylor's unfolding network, invertible neural network, central difference convolution, and zero-order filtering. We investigate these prototypes from four aspects: 1) random weight strategies, 2) network architectures, 3) network depths, and 4) combinations of random weight networks. Furthermore, we devise the random weight in two variants: the weights are randomly initialized only once during the entire training procedure, and the weights are randomly initialized in each training epoch. Our approach can be directly integrated into existing networks without incurring additional training and testing computational costs. We perform extensive experiments across multiple image restoration tasks, including image denoising, low-light image enhancement, and guided image super-resolution to demonstrate the consistent performance gains achieved by our method. Upon acceptance of this paper, we will release the code. Man Zhou 0003, Naishan Zheng, Chunle Guo, Chongyi Li |
NeurIPS | 5 |
| 2023 | FouriDown: Factoring Down-Sampling into Shuffling and SuperposingabstractSpatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employed in previous approaches. To overcome this limitation, we propose a novel down-sampling paradigm in the Fourier domain, abbreviated as FouriDown, which unifies existing down-sampling techniques. Drawing inspiration from the signal sampling theorem, we parameterize the non-parameter static weighting down-sampling operator as a learnable and context-adaptive operator within a unified Fourier function. Specifically, we organize the corresponding frequency positions of the 2D plane in a physically-closed manner within a single channel dimension. We then perform point-wise channel shuffling based on an indicator that determines whether a channel's signal frequency bin is susceptible to aliasing, ensuring the consistency of the weighting parameter learning. FouriDown, as a generic operator, comprises four key components: 2D discrete Fourier transform, context shuffling rules, Fourier weighting-adaptively superposing rules, and 2D inverse Fourier transform. These components can be easily integrated into existing image restoration networks. To demonstrate the efficacy of FouriDown, we conduct extensive experiments on image de-blurring and low-light image enhancement. The results consistently show that FouriDown can provide significant performance improvements. We will make the code publicly available to facilitate further exploration and application of FouriDown. Qi Zhu 0010, Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Hongzhi Gao, Chongyi Li, Feng Zhao 0004 |
NeurIPS | 6 |
| 2023 | Underwater Image Enhancement via Piecewise Color Correction and Dual Prior Optimized Contrast EnhancementabstractDue to the absorption and scattering of light, underwater captured images often face serious quality degradation issues. In this letter, we propose to cope with the aforementioned issues via piecewise color correction and dual prior optimized contrast enhancement. Specifically, we first present the piecewise color correction method using the maximum mean and two gain factors to correct the color cast of each color channel. Then, we propose a dual prior optimized contrast enhancement method, which relies on the spatial and texture priors to decompose the base layer and detail layer of the V channel in HSV color space. Meanwhile, we employ different enhancement strategies in different layers to enhance the contrast and texture detail of underwater images. Our extensive experiments on several benchmark datasets show that our method outperforms eleven compared state-of-the-art methods. Moreover, our method has good generalization capability for fog and low-light images. The code is available athttps://github.com/Li-Chongyi/PCDE. Weidong Zhang 0007, Songlin Jin, Peixian Zhuang, Zheng Liang 0001, Chongyi Li |
IEEE Signal Process. Lett. | 5 |
| 2023 | A Perception-Aware Decomposition and Fusion Framework for Underwater Image EnhancementabstractThis paper presents a perception-aware decomposition and fusion framework for underwater image enhancement (UIE). Specifically, a general structural patch decomposition and fusion (SPDF) approach is introduced. SPDF is built upon the fusion of two complementary pre-processed inputs in a perception-aware and conceptually independent image space. First, a raw underwater image is pre-processed to produce two complementary versions including a contrast-corrected image and a detail-sharpened image. Then, each of them is decomposed into three conceptually independent components, i.e., mean intensity, contrast, and structure, via structural patch decomposition (SPD). Afterwards, the corresponding components are fused using tailored strategies. The three components after fusion are finally integrated via inverting the decomposition to reconstruct a final enhanced underwater image. The main advantage of SPDF is that two complementary pre-processed images are fused in a perception-aware and conceptually independent image space and the fusions of different components can be performed separately without any interactions and information loss. Comprehensive comparisons on two benchmark datasets demonstrate that SPDF outperforms several state-of-the-art UIE algorithms qualitatively and quantitatively. Moreover, the effectiveness of SPDF is also verified on another two relevant tasks, i.e., low-light image enhancement and single image dehazing. The code will be made available soon. Yaozu Kang, Qiuping Jiang, Chongyi Li, Wenqi Ren, Hantao Liu, Pengjun Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Embedding Global Contrastive and Local Location in Self-Supervised LearningabstractSelf-supervised representation learning (SSL) typically suffers from inadequate data utilization and feature-specificity due to the suboptimal sampling strategy and the monotonous optimization method. Existing contrastive-based methods alleviate these issues through exceedingly long training time and large batch size, resulting in non-negligible computational consumption and memory usage. In this paper, we present an efficient self-supervised framework, called GLNet. The key insights of this work are the novel sampling and ensemble learning strategies embedded in the self-supervised framework. We first propose a location-based sampling strategy to integrate the complementary advantages of semantic and spatial characteristics. Whereafter, a Siamese network with momentum update is introduced to generate representative vectors, which are used to optimize the feature extractor. Finally, we particularly embed global contrastive and local location tasks in the framework, which aims to leverage the complementarity between the high-level semantic features and low-level texture features. Such complementarity is significant for mitigating the feature-specificity and improving the generalizability, thus effectively improving the performance of downstream tasks. Extensive experiments on representative benchmark datasets demonstrate that GLNet performs favorably against the state-of-the-art SSL methods. Specifically, GLNet improves MoCo-v3 by 2.4% accuracy on ImageNet dataset, while improves 2% accuracy and consumes only 75% training time on the ImageNet-100 dataset. In addition, GLNet is appealing in its compatibility with popular SSL frameworks. Code is available at GLNet. Wenyi Zhao, Chongyi Li, Weidong Zhang 0007, Lu Yang 0006, Peixian Zhuang, Lingqiao Li, Kefeng Fan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Global-and-Local Collaborative Learning for Co-Salient Object DetectionabstractThe goal of co-salient object detection (CoSOD) is to discover salient objects that commonly appear in a query group containing two or more relevant images. Therefore, how to effectively extract interimage correspondence is crucial for the CoSOD task. In this article, we propose a global-and-local collaborative learning (GLNet) architecture, which includes a global correspondence modeling (GCM) and a local correspondence modeling (LCM) to capture the comprehensive interimage corresponding relationship among different images from the global and local perspectives. First, we treat different images as different time slices and use 3-D convolution to integrate all intrafeatures intuitively, which can more fully extract the global group semantics. Second, we design a pairwise correlation transformation (PCT) to explore similarity correspondence between pairwise images and combine the multiple local pairwise correspondences to generate the local interimage relationship. Third, the interimage relationships of the GCM and LCM are integrated through a global-and-local correspondence aggregation (GLA) module to explore more comprehensive interimage collaboration cues. Finally, the intra and inter features are adaptively integrated by an intra-and-inter weighting fusion (AEWF) module to learn co-saliency features and predict the co-saliency map. The proposed GLNet is evaluated on three prevailing CoSOD benchmark datasets, demonstrating that our model trained on a small dataset (about 3k images) still outperforms 11 state-of-the-art competitors trained on some large datasets (about 8k-200k images). Runmin Cong, Ning Yang 0008, Chongyi Li, Huazhu Fu, Yao Zhao 0001, Qingming Huang, Sam Kwong |
IEEE Trans. Cybern. | 3 |
| 2023 | PUGAN: Physical Model-Guided Underwater Image Enhancement Using GAN With Dual-DiscriminatorsabstractDue to the light absorption and scattering induced by the water medium, underwater images usually suffer from some degradation problems, such as low contrast, color distortion, and blurring details, which aggravate the difficulty of downstream underwater understanding tasks. Therefore, how to obtain clear and visually pleasant images has become a common concern of people, and the task of underwater image enhancement (UIE) has also emerged as the times require. Among existing UIE methods, Generative Adversarial Networks (GANs) based methods perform well in visual aesthetics, while the physical model-based methods have better scene adaptability. Inheriting the advantages of the above two types of models, we propose a physical model-guided GAN model for UIE in this paper, referred to as PUGAN. The entire network is under the GAN architecture. On the one hand, we design a Parameters Estimation subnetwork (Par-subnet) to learn the parameters for physical model inversion, and use the generated color enhancement image as auxiliary information for the Two-Stream Interaction Enhancement sub-network (TSIE-subnet). Meanwhile, we design a Degradation Quantization (DQ) module in TSIE-subnet to quantize scene degradation, thereby achieving reinforcing enhancement of key regions. On the other hand, we design the Dual-Discriminators for the style-content adversarial constraint, promoting the authenticity and visual aesthetics of the results. Extensive experiments on three benchmark datasets demonstrate that our PUGAN outperforms state-of-the-art methods in both qualitative and quantitative metrics. The code and results can be found from the link of https://rmcong.github.io/proj_PUGAN.html. Runmin Cong, Wei Zhang 0021, Chongyi Li, Chunle Guo, Qingming Huang, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2023 | Learning to remove sandstorm for image enhancement
Pengwei Liang, Pengwei Dong, Fan Wang 0005, Peng Ma, Jiajing Bai, Bo Wang 0070, Chongyi Li |
Vis. Comput. | 7 |
| 2022 | Image Dehazing Transformer with Transmission-Aware 3D Position EmbeddingabstractDespite single image dehazing has been made promising progress with Convolutional Neural Networks (CNNs), the inherent equivariance and locality of convolution still bottleneck deharing performance. Though Transformer has occupied various computer vision tasks, directly leveraging Transformer for image dehazing is challenging: 1) it tends to result in ambiguous and coarse details that are undesired for image reconstruction; 2) previous position embedding of Transformer is provided in logic or spatial position order that neglects the variational haze densities, which results in the sub-optimal dehazlng performance. The key insight of this study is to investigate how to combine CNN and Transformer for image dehazing. To solve the feature inconsistency issue between Transformer and CNN, we propose to modulate CNN features via learning modulation matrices (i.e., coefficient matrix and bias matrix) conditioned on Transformer features instead of simple feature addition or concatenation. The feature modulation naturally inherits the global context modeling capability of Transformer and the local representation capability of CNN. We bring a haze density-related prior into Trans-former via a novel transmission-aware 3D position embedding module, which not only provides the relative position but also suggests the haze density of different spatial regions. Extensive experiments demonstrate that our method, DeHamer, attains state-of-the-art performance on several image dehazing benchmarks. Chunle Guo, Qixin Yan, Saeed Anwar, Runmin Cong, Wenqi Ren, Chongyi Li |
CVPR | 6 |
| 2022 | LEDNet: Joint Low-Light Enhancement and Deblurring in the Dark
Shangchen Zhou, Chongyi Li, Chen Change Loy |
ECCV (6) | 2 |
| 2022 | Adaptively Learning Low-high Frequency Information Integration for Pan-sharpeningabstractPan-sharpening aims to generate high-spatial resolution multi-spectral (MS) image by fusing high-spatial resolution panchromatic (PAN) image and its corresponding low-spatial resolution MS image. Despite the remarkable progress, most existing pan-sharpening methods only work in the spatial domain and rarely explore the potential solutions in the frequency domain. In this paper, we propose a novel pan-sharpening framework by adaptively learning low-high frequency information integration in the spatial and frequency dual domains. It consists of three key designs: mask prediction sub-network, low-frequency learning sub-network and high-frequency learning sub-network. Specifically, the first is responsible for measuring the modality-aware frequency information difference of PAN and MS images and further predicting the low-high frequency boundary in the form of a two-dimensional mask. In view of the mask, the second adaptively picks out the corresponding low-frequency components of different modalities and then restores the expected low-frequency one by spatial and frequency dual domains information integration while the third combines the above refined low-frequency and the original high-frequency for the latent high-frequency reconstruction. In this way, the low-high frequency information is adaptively learned, thus leading to the pleasing results. Extensive experiments validate the effectiveness of the proposed network and demonstrate the favorable performance against other state-of-the-art methods. The source code will be released at https://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Chongyi Li, Hu Yu 0001, Naishan Zheng, Feng Zhao 0004 |
ACM Multimedia | 3 |
| 2022 | Normalization-based Feature Selection and Restitution for Pan-sharpeningabstractPan-sharpening is essentially a panchromatic (PAN) image-guided low-spatial resolution MS image super-resolution problem. The commonly challenging issue of pan-sharpening is how to correctly select consistent features and propagate them, and properly handle inconsistent ones between PAN and MS modalities. To solve this issue, we propose a Normalization-based Feature Selection and Restitution mechanism, which is capable of filtering out the inconsistent features and promoting to learn the consistent ones. Specifically, we first modulate the PAN feature as the MS style in feature space by AdaIN operation \citeAdaIN. However, such operation inevitably removes the favorable features. We thus propose to distill the effective information from the removed part and restitute it back to the modulated part. To better distillation, we enforce a contrastive learning constraint to close the distance between the restituted feature and the ground truth, and push the removed part away from the ground truth. In this way, the consistent features of PAN images are correctly selected and the inconsistent ones are filtered out, thus relieving the over-transferred artifacts in the process of PAN-guided MS super-resolution. Extensive experiments validate the effectiveness of the proposed network and demonstrate its favorable performance against other state-of-the-art methods. The source code will be released at https://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Aiping Liu, Chongyi Li, Feng Zhao 0004 |
ACM Multimedia | 6 |
| 2022 | Flare7K: A Phenomenological Nighttime Flare Removal DatasetabstractArtificial lights commonly leave strong lens flare artifacts on images captured at night. Nighttime flare not only affects the visual quality but also degrades the performance of vision algorithms. Existing flare removal methods mainly focus on removing daytime flares and fail in nighttime. Nighttime flare removal is challenging because of the unique luminance and spectrum of artificial lights and the diverse patterns and image degradation of the flares captured at night. The scarcity of nighttime flare removal datasets limits the research on this crucial task. In this paper, we introduce, Flare7K, the first nighttime flare removal dataset, which is generated based on the observation and statistics of real-world nighttime lens flares. It offers 5,000 scattering and 2,000 reflective flare images, consisting of 25 types of scattering flares and 10 types of reflective flares. The 7,000 flare patterns can be randomly added to flare-free images, forming the flare-corrupted and flare-free image pairs. With the paired data, we can train deep models to restore flare-corrupted images taken in the real world effectively. Apart from abundant flare patterns, we also provide rich annotations, including the labeling of light source, glare with shimmer, reflective flare, and streak, which are commonly absent from existing datasets. Hence, our dataset can facilitate new work in nighttime flare removal and more fine-grained analysis of flare patterns. Extensive experiments show that our dataset adds diversity to existing flare datasets and pushes the frontier of nighttime flare removal. Yuekun Dai, Chongyi Li, Shangchen Zhou, Ruicheng Feng, Chen Change Loy |
NeurIPS | 2 |
| 2022 | Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering NetworkabstractPanchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to generate the expected high-resolution (HR) MS images, conditioning on the corresponding high-resolution PAN images. In this paper, we present a simple yet effective alternating reverse filtering network for pan-sharpening. Inspired by the classical reverse filtering that reverses images to the status before filtering, we formulate pan-sharpening as an alternately iterative reverse filtering process, which fuses LR MS and HR MS in an interpretable manner. Different from existing model-driven methods that require well-designed priors and degradation assumptions, the reverse filtering process avoids the dependency on pre-defined exact priors. To guarantee the stability and convergence of the iterative process via contraction mapping on a metric space, we develop the learnable multi-scale Gaussian kernel module, instead of using specific filters. We demonstrate the theoretical feasibility of such formulations. Extensive experiments on diverse scenes to thoroughly verify the performance of our method, significantly outperforming the state of the arts. Man Zhou 0003, Jie Huang 0017, Feng Zhao 0004, Chengjun Xie, Chongyi Li, Danfeng Hong |
NeurIPS | 6 |
| 2022 | Towards Robust Blind Face Restoration with Codebook Lookup TransformerabstractBlind face restoration is a highly ill-posed problem that often requires auxiliary guidance to 1) improve the mapping from degraded inputs to desired outputs, or 2) complement high-quality details lost in the inputs. In this paper, we demonstrate that a learned discrete codebook prior in a small proxy space largely reduces the uncertainty and ambiguity of restoration mapping by casting \textit{blind face restoration} as a \textit{code prediction} task, while providing rich visual atoms for generating high-quality faces. Under this paradigm, we propose a Transformer-based prediction network, named \textit{CodeFormer}, to model the global composition and context of the low-quality faces for code prediction, enabling the discovery of natural faces that closely approximate the target faces even when the inputs are severely degraded. To enhance the adaptiveness for different degradation, we also propose a controllable feature transformation module that allows a flexible trade-off between fidelity and quality. Thanks to the expressive codebook prior and global modeling, \textit{CodeFormer} outperforms the state of the arts in both quality and fidelity, showing superior robustness to degradation. Extensive experimental results on synthetic and real-world datasets verify the effectiveness of our method. Shangchen Zhou, Kelvin C. K. Chan, Chongyi Li, Chen Change Loy |
NeurIPS | 3 |
| 2022 | Deep Fourier Up-SamplingabstractExisting convolutional neural networks widely adopt spatial down-/up-sampling for multi-scale modeling. However, spatial up-sampling operators (e.g., interpolation, transposed convolution, and un-pooling) heavily depend on local pixel attention, incapably exploring the global dependency. In contrast, the Fourier domain is in accordance with the nature of global modeling according to the spectral convolution theorem. Unlike the spatial domain that easily performs up-sampling with the property of local similarity, up-sampling in the Fourier domain is more challenging as it does not follow such a local property. In this study, we propose a theoretically feasible Deep Fourier Up-Sampling (FourierUp) to solve these issues. We revisit the relationships between spatial and Fourier domains and reveal the transform rules on the features of different resolutions in the Fourier domain, which provide key insights for FourierUp's designs. FourierUp as a generic operator consists of three key components: 2D discrete Fourier transform, Fourier dimension increase rules, and 2D inverse Fourier transform, which can be directly integrated with existing networks. Extensive experiments across multiple computer vision tasks, including object detection, image segmentation, image de-raining, image dehazing, and guided image super-resolution, demonstrate the consistent performance gains obtained by introducing our FourierUp. Code will be publicly available. Man Zhou 0003, Hu Yu 0001, Jie Huang 0017, Feng Zhao 0004, Jinwei Gu, Chen Change Loy, Deyu Meng, Chongyi Li |
NeurIPS | 8 |
| 2022 | Conditional mutual information-based feature selection algorithm for maximal relevance minimal redundancy
Xiangyuan Gu, Jichang Guo, Chongyi Li |
Appl. Intell. | 4 |
| 2022 | Nighttime image dehazing using color cast removal and dual path multi-scale fusion strategy
Bo Wang 0070, Bowen Wei, Zitong Kang, Chongyi Li |
Frontiers Comput. Sci. | 5 |
| 2022 | Underwater image enhancement by maximum-likelihood based adaptive color correction and robust scattering removal
Bo Wang 0070, Zitong Kang, Pengwei Dong, Fan Wang 0005, Peng Ma, Jiajing Bai, Pengwei Liang, Chongyi Li |
Frontiers Comput. Sci. | 8 |
| 2022 | Multi-scale and multi-patch transformer for sandstorm image enhancement
Pengwei Liang, Wenyu Ding, Zihong Li, Bo Wang 0070, Chongyi Li |
J. Vis. Commun. Image Represent. | 8 |
| 2022 | Salient object detection in low-light images via functional optimization-inspired feature polishing
HuiHui Yue, Jichang Guo, Xiangjun Yin, Yi Zhang 0107, Sida Zheng, Zenan Zhang, Chongyi Li |
Knowl. Based Syst. | 7 |
| 2022 | The Orientation Estimation of Elongated Underground Objects via Multipolarization Aggregation and Selection Neural NetworkabstractThe horizontal orientation angle and the vertical inclination angle of an elongated subsurface object are key parameters for object identification and imaging in ground-penetrating radar (GPR) applications. Conventional methods can only extract the horizontal orientation angle or estimate both angles in narrow ranges due to limited polarimetric information and detection capability. To address these issues, this letter, for the first time, explores the possibility of leveraging neural networks with multipolarimetric GPR data to estimate both angles of an elongated subsurface object in the entire spatial range. Based on the polarization-sensitive characteristic of an elongated object, we propose a multipolarization aggregation and selection network (MASNet), which takes the multipolarimetric radargrams as inputs, integrates their characteristics in the feature space, and selects discriminative features of reflected signal patterns for accurate orientation estimation. Numerical results show that our proposed MASNet achieves high estimation accuracy with an angle estimation error of less than 5°. The promising results obtained by the proposed method encourage one to think of new solutions for GPR-related tasks by integrating multipolarization information with deep learning techniques. The data and code implemented in the letter can be found athttps://haihan-sun.github.io/GPR.html. Hai-Han Sun, Yee Hui Lee, Chongyi Li, Genevieve Lai Fern Ow, Mohamed Lokman Mohd Yusof, Abdulkadir C. Yucel |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | SSTNet: Spatial, Spectral, and Texture Aware Attention Network Using Hyperspectral Image for Corn Variety IdentificationabstractCurrently, most existing methods using hyperspectral image to assist seed identification only consider the spectral information but ignore the spatial information resulting in unsatisfactory classification results. To cope with this issue, we propose a spatial, spectral, and texture-aware attention network to identify corn varieties, called SSTNet. Specifically, we first employ 3D convolution to extract the spatial and inter-spectral features. Subsequently, we utilize 2D convolution to extract the spatial and texture features. Meanwhile, we embed an attention mechanism into the 2D convolution module to further refine the spatial and texture features. The advantageous complementary properties of 3D and 2D convolutions allow the spatial and textural features of hyperspectral images to be fully exploited. Besides, we construct a hyperspectral image dataset including 1200 samples of 10 corn varieties. Experiments on our proposed dataset demonstrate that our SSTNet outperforms the state-of-the-art methods for identifying corn varieties. Weidong Zhang 0007, Hai-Han Sun, Qiang Zhang 0011, Peixian Zhuang, Chongyi Li |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | A Feature Selection Algorithm Based on Equal Interval Division and Conditional Mutual Information
Xiangyuan Gu, Jichang Guo, Tao Ming, Chongyi Li |
Neural Process. Lett. | 5 |
| 2022 | Low-Light Image and Video Enhancement Using Deep Learning: A SurveyabstractLow-light image enhancement (LLIE) aims at improving the perception or interpretability of an image captured in an environment with poor illumination. Recent advances in this area are dominated by deep learning-based solutions, where many learning strategies, network structures, loss functions, training data, etc. have been employed. In this paper, we provide a comprehensive survey to cover various aspects ranging from algorithm taxonomy to unsolved open issues. To examine the generalization of existing methods, we propose a low-light image and video dataset, in which the images and videos are taken by different mobile phones' cameras under diverse illumination conditions. Besides, for the first time, we provide a unified online platform that covers many popular LLIE methods, of which the results can be produced through a user-friendly web interface. In addition to qualitative and quantitative evaluation of existing methods on publicly available and our proposed datasets, we also validate their performance in face detection in the dark. This survey together with the proposed dataset and online platform could serve as a reference source for future study and promote the development of this research field. The proposed platform and dataset as well as the collected methods, datasets, and evaluation metrics are publicly available and will be regularly updated. Project page: https://www.mmlab-ntu.com/project/lliv_survey/index.html. Chongyi Li, Chunle Guo, Linghao Han, Ming-Ming Cheng, Jinwei Gu, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Learning to Enhance Low-Light Image via Zero-Reference Deep Curve EstimationabstractThis paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network. Our method trains a lightweight deep network, DCE-Net, to estimate pixel-wise and high-order curves for dynamic range adjustment of a given image. The curve estimation is specially designed, considering pixel value range, monotonicity, and differentiability. Zero-DCE is appealing in its relaxed assumption on reference images, i.e., it does not require any paired or even unpaired data during training. This is achieved through a set of carefully formulated non-reference loss functions, which implicitly measure the enhancement quality and drive the learning of the network. Despite its simplicity, we show that it generalizes well to diverse lighting conditions. Our method is efficient as image enhancement can be achieved by an intuitive and simple nonlinear curve mapping. We further present an accelerated and light version of Zero-DCE, called Zero-DCE++, that takes advantage of a tiny network with just 10K parameters. Zero-DCE++ has a fast inference speed (1000/11 FPS on a single GPU/CPU for an image of size 1200×900×3) while keeping the enhancement performance of Zero-DCE. Extensive experiments on various benchmarks demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively. Furthermore, the potential benefits of our method to face detection in the dark are discussed. The source code is made publicly available at https://li-chongyi.github.io/Proj_Zero-DCE++.html. Chongyi Li, Chunle Guo, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective MetricabstractDue to the attenuation and scattering of light by water, there are many quality defects in raw underwater images such as color casts, decreased visibility, reduced contrast,et al.. Many different underwater image enhancement (UIE) algorithms have been proposed to enhance underwater image quality. However, how to fairly compare the performance among UIE algorithms remains a challenging problem. So far, the lack of comprehensive human subjective user study with large-scale benchmark dataset and reliable objective image quality assessment (IQA) metric makes it difficult to fully understand the true performance of UIE algorithms. We in this paper make efforts in both subjective and objective aspects to fill these gaps. Firstly, we construct a new Subjectively-Annotated UIE benchmark Dataset (SAUD) which simultaneously provides real-world raw underwater images, readily available enhanced results by representative UIE algorithms, and subjective ranking scores of each enhanced result. Secondly, we propose an effective No-reference (NR) Underwater Image Quality metric (NUIQ) to automatically evaluate the visual quality of enhanced underwater images. Experiments on the constructed SAUD dataset demonstrate the superiority of our proposed NUIQ metric, achieving higher consistency with subjective rankings than 22 mainstream NR-IQA metrics. The dataset and source code will be made available athttps://github.com/yia-yuese/SAUD-Dataset. Qiuping Jiang, Yuese Gu, Chongyi Li, Runmin Cong, Feng Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | GIFM: An Image Restoration Method With Generalized Image Formation Model for Poor Visible ConditionsabstractRecently, image restoration has attracted considerable attention from researchers, and these methods generally restore degraded images based on the atmospheric scattering model (ATSM) and retinex model (RM). The two models only take into the single attenuation process during imaging, thereby introducing undesirable results. To deal with this issue, we propose an image restoration method based on a generalized image formation model (GIFM). First, unlike the existing image restoration methods, we rebuild a novel image formation model, which describes the light attenuation process that includes the light source-scene path and scene-sensor path. Second, we construct an objective optimization function to decompose a degraded image into a color distorted component and color corrected component, and an augmented Lagrange multiplier-based alternating direction minimization algorithm is provided to solve the optimization problem. Finally, we fully consider the advantages of the small-scale neighborhood and large-scale neighborhood in image restoration, and an image itself brightness-based weighted fusion strategy is proposed to balance brightness enhancement and contrast improvement. Extensive experiments on three image enhancement datasets show that our GIFM achieves better results than state-of-the-art methods. Experiments further suggest that our GIFM performs well for image restoration of extreme scenes, keypoint detection, object detection, and image segmentation. Zheng Liang 0001, Weidong Zhang 0007, Rui Ruan, Peixian Zhuang, Chongyi Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Estimating Parameters of the Tree Root in Heterogeneous Soil Environments via Mask-Guided Multi-Polarimetric Integration Neural NetworkabstractGround-penetrating radar (GPR) has been used as a nondestructive tool for tree root inspection. Estimating root-related parameters from GPR radargrams greatly facilitates root health monitoring and imaging. However, the task of estimating root-related parameters is challenging as the root reflection is a complex function of multiple root parameters and root orientations. Existing methods can only estimate a single root parameter at a time without considering the influence of other parameters and root orientations, resulting in limited estimation accuracy under different root conditions. In addition, soil heterogeneity introduces clutter in GPR radargrams, making the data processing and interpretation even harder. To address these issues, a novel neural network architecture, called mask-guided multi-polarimetric integration neural network (MMI-Net), is proposed to automatically and simultaneously estimate multiple root-related parameters in heterogeneous soil environments. The MMI-Net includes two subnetworks: a MaskNet that predicts a mask to highlight the root reflection area to eliminate interfering environmental clutter and a parameter estimation subnetwork (ParaNet) that uses the predicted mask as guidance to integrate, extract, and emphasize informative features in multi-polarimetric radargrams for accurate estimation of five key root-related parameters. The parameters include the root depth, diameter, relative permittivity, and horizontal and vertical orientation angles. Experimental results demonstrate that the proposed MMI-Net achieves high estimation accuracy in these root-related parameters. This is the first work that takes the combined contributions of root parameters and spatial orientations into account and simultaneously estimates multiple root-related parameters. The data and code implemented in this article can be found athttps://haihan-sun.github.io/GPR.html. Hai-Han Sun, Yee Hui Lee, Qiqi Dai, Chongyi Li, Genevieve Lai Fern Ow, Mohamed Lokman Mohd Yusof, Abdulkadir C. Yucel |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | CIR-Net: Cross-Modality Interaction and Refinement for RGB-D Salient Object DetectionabstractFocusing on the issue of how to effectively capture and utilize cross-modality information in RGB-D salient object detection (SOD) task, we present a convolutional neural network (CNN) model, named CIR-Net, based on the novel cross-modality interaction and refinement. For the cross-modality interaction, 1) a progressive attention guided integration unit is proposed to sufficiently integrate RGB-D feature representations in the encoder stage, and 2) a convergence aggregation structure is proposed, which flows the RGB and depth decoding features into the corresponding RGB-D decoding streams via an importance gated fusion unit in the decoder stage. For the cross-modality refinement, we insert a refinement middleware structure between the encoder and the decoder, in which the RGB, depth, and RGB-D encoder features are further refined by successively using a self-modality attention refinement unit and a cross-modality weighting refinement unit. At last, with the gradually refined features, we predict the saliency map in the decoder stage. Extensive experiments on six popular RGB-D SOD benchmarks demonstrate that our network outperforms the state-of-the-art saliency detectors both qualitatively and quantitatively. The code and results can be found from the link of https://rmcong.github.io/proj_CIRNet.html. Runmin Cong, Qinwei Lin, Chen Zhang 0013, Chongyi Li, Xiaochun Cao, Qingming Huang, Yao Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Under-Display Camera Image Enhancement via Cascaded Curve EstimationabstractThe new trend of full-screen devices encourages manufacturers to position a camera behind a screen, i.e., the newly-defined Under-Display Camera (UDC). Therefore, UDC image restoration has been a new realistic single image enhancement problem. In this work, we propose a curve estimation network operating on the hue (H) and saturation (S) channels to perform adaptive enhancement for degraded images captured by UDCs. The proposed network aims to match the complicated relationship between the images captured by under-display and display-free cameras. To extract effective features, we cascade the proposed curve estimation network with sharing weights, and we introduce a spatial and channel attention module in each curve estimation network to exploit attention-aware features. In addition, we learn the curve estimation network in a semi-supervised manner to alleviate the restriction of the requirement for amounts of labeled images and improve the generalization ability for unseen degraded images in various realistic scenes. The semi-supervised network consists of a supervised branch trained on labeled data and an unsupervised branch trained on unlabeled data. To train the proposed model, we build a new dataset comprised of real-world labeled and unlabeled images. Extensive experiments demonstrate that our proposed algorithm performs favorably against state-of-the-art image enhancement methods for UDC images in terms of accuracy and speed, especially on ultra-high-definition (UHD) images. Jun Luo 0012, Wenqi Ren, Tao Wang 0074, Chongyi Li, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2022 | Underwater Image Enhancement via Minimal Color Loss and Locally Adaptive Contrast EnhancementabstractUnderwater images typically suffer from color deviations and low visibility due to the wavelength-dependent light absorption and scattering. To deal with these degradation issues, we propose an efficient and robust underwater image enhancement method, called MLLE. Specifically, we first locally adjust the color and details of an input image according to a minimum color loss principle and a maximum attenuation map-guided fusion strategy. Afterward, we employ the integral and squared integral maps to compute the mean and variance of local image blocks, which are used to adaptively adjust the contrast of the input image. Meanwhile, a color balance strategy is introduced to balance the color differences between channel a and channel b in the CIELAB color space. Our enhanced results are characterized by vivid color, improved contrast, and enhanced details. Extensive experiments on three underwater image enhancement datasets demonstrate that our method outperforms the state-of-the-art methods. Our method is also appealing in its fast processing speed within 1s for processing an image of size 1024×1024×3 on a single CPU. Experiments further suggest that our method can effectively improve the performance of underwater image segmentation, keypoint detection, and saliency detection. The project page is available at https://li-chongyi.github.io/proj_MMLE.html. Weidong Zhang 0007, Peixian Zhuang, Hai-Han Sun, Guohou Li, Sam Kwong, Chongyi Li |
IEEE Trans. Image Process. | 6 |
| 2022 | Underwater Image Enhancement With Hyper-Laplacian Reflectance PriorsabstractUnderwater image enhancement aims at improving the visibility and eliminating color distortions of underwater images degraded by light absorption and scattering in water. Recently, retinex variational models show remarkable capacity of enhancing images by estimating reflectance and illumination in a retinex decomposition course. However, ambiguous details and unnatural color still challenge the performance of retinex variational models on underwater image enhancement. To overcome these limitations, we propose a hyper-laplacian reflectance priors inspired retinex variational model to enhance underwater images. Specifically, the hyper-laplacian reflectance priors are established with thel1/2-norm penalty on first-order and second-order gradients of the reflectance. Such priors exploit sparsity-promoting and complete-comprehensive reflectance that is used to enhance both salient structures and fine-scale details and recover the naturalness of authentic colors. Besides, thel2norm is found to be suitable for accurately estimating the illumination. As a result, we turn a complex underwater image enhancement issue into simple subproblems that separately and simultaneously estimate the reflection and the illumination that are harnessed to enhance underwater images in a retinex variational model. We mathematically analyze and solve the optimal solution of each subproblem. In the optimization course, we develop an alternating minimization algorithm that is efficient on element-wise operations and independent of additional prior knowledge of underwater conditions. Extensive experiments demonstrate the superiority of the proposed method in both subjective results and objective assessments over existing methods. Peixian Zhuang, Fatih Porikli, Chongyi Li |
IEEE Trans. Image Process. | 4 |
| 2021 | Investigating Attention Mechanism in 3D Point Cloud Object DetectionabstractObject detection in three-dimensional (3D) space attracts much interest from academia and industry since it is an essential task in AI-driven applications such as robotics, autonomous driving, and augmented reality. As the basic format of 3D data, the point cloud can provide detailed geometric information about the objects in the original 3D space. However, due to 3D data s sparsity and unorderedness, specially designed networks and modules are needed to process this type of data. Attention mechanism has achieved impressive performance in diverse computer vision tasks; however, it is unclear how attention modules would affect the performance of 3D point cloud object detection and what sort of attention modules could fit with the inherent properties of 3D data. This work investigates the role of the attention mechanism in 3D point cloud object detection and provides insights into the potential of different attention modules. To achieve that, we comprehensively investigate classical 2D attentions, novel 3D attentions, including the latest point cloud transformers on SUN RGB-D and ScanNetV2 datasets. Based on the detailed experiments and analysis, we conclude the effects of different attention modules. This paper is expected to serve as a reference source for benefiting attention-embedded 3D point cloud object detection. The code and trained models are available at: https://qithub.com/SkiQiu0419/attentions_in_3D_detection. Shi Qiu 0001, Saeed Anwar, Chongyi Li |
3DV | 4 |
| 2021 | Removing Diffraction Image Artifacts in Under-Display Camera via Dynamic Skip Connection NetworkabstractDevelopment of Under-Display Camera (UDC) systems provides a true bezel-less and notch-free viewing experience on smartphones (and TV, laptops, tablets), while allowing images to be captured from the selfie camera embedded underneath. In a typical UDC system, the microstructure of the semi-transparent organic light-emitting diode (OLED) pixel array attenuates and diffracts the incident light on the camera, resulting in significant image quality degradation. Oftentimes, noise, flare, haze, and blur can be observed in UDC images. In this work, we aim to analyze and tackle the aforementioned degradation problems. We define a physics-based image formation model to better understand the degradation. In addition, we utilize one of the world’s first commodity UDC smartphone prototypes to measure the real-world Point Spread Function (PSF) of the UDC system, and provide a model-based data synthesis pipeline to generate realistically degraded images. We specially design a new domain knowledge-enabled Dynamic Skip Connection Network (DISCNet) to restore the UDC images. We demonstrate the effectiveness of our method through extensive experiments on both synthetic and real UDC data. Our physics-based image formation model and proposed DISCNet can provide foundations for further exploration in UDC image restoration, and even for general diffraction artifact removal in a broader sense.1 Ruicheng Feng, Chongyi Li, Huaijin G. Chen, Chen Change Loy, Jinwei Gu |
CVPR | 2 |
| 2021 | A feature selection algorithm based on redundancy analysis and interaction weight
Xiangyuan Gu, Jichang Guo, Chongyi Li |
Appl. Intell. | 3 |
| 2021 | Bayesian retinex underwater image enhancement
Peixian Zhuang, Chongyi Li |
Eng. Appl. Artif. Intell. | 2 |
| 2021 | Stereo superpixel: An iterative framework based on parallax consistency and collaborative optimization
Hua Li 0012, Runmin Cong, Sam Kwong, Chuanbo Chen, Qianqian Xu 0001, Chongyi Li |
Inf. Sci. | 6 |
| 2021 | Blind face images deblurring with enhancement
Qing Qi, Jichang Guo, Chongyi Li |
Multim. Tools Appl. | 3 |
| 2021 | ASIF-Net: Attention Steered Interweave Fusion Network for RGB-D Salient Object DetectionabstractSalient object detection from RGB-D images is an important yet challenging vision task, which aims at detecting the most distinctive objects in a scene by combining color information and depth constraints. Unlike prior fusion manners, we propose an attention steered interweave fusion network (ASIF-Net) to detect salient objects, which progressively integrates cross-modal and cross-level complementarity from the RGB image and corresponding depth map via steering of an attention mechanism. Specifically, the complementary features from RGB-D images are jointly extracted and hierarchically fused in a dense and interweaved manner. Such a manner breaks down the barriers of inconsistency existing in the cross-modal data and also sufficiently captures the complementarity. Meanwhile, an attention mechanism is introduced to locate the potential salient regions in an attention-weighted fashion, which advances in highlighting the salient objects and suppressing the cluttered background regions. Instead of focusing only on pixelwise saliency, we also ensure that the detected salient objects have the objectness characteristics (e.g., complete structure and sharp boundary) by incorporating the adversarial learning that provides a global semantic constraint for RGB-D salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably against 17 state-of-the-art saliency detectors on four publicly available RGB-D salient object detection datasets. The code and results of our method are available at https://github.com/Li-Chongyi/ASIF-Net. Chongyi Li, Runmin Cong, Sam Kwong, Junhui Hou, Huazhu Fu, Guopu Zhu, Dingwen Zhang, Qingming Huang |
IEEE Trans. Cybern. | 1 |
| 2021 | Underwater Image Enhancement via Medium Transmission-Guided Multi-Color Space EmbeddingabstractUnderwater images suffer from color casts and low contrast due to wavelength- and distance-dependent attenuation and scattering. To solve these two degradation issues, we present an underwater image enhancement network via medium transmission-guided multi-color space embedding, called Ucolor. Concretely, we first propose a multi-color space encoder network, which enriches the diversity of feature representations by incorporating the characteristics of different color spaces into a unified structure. Coupled with an attention mechanism, the most discriminative features extracted from multiple color spaces are adaptively integrated and highlighted. Inspired by underwater imaging physical models, we design a medium transmission (indicating the percentage of the scene radiance reaching the camera)-guided decoder network to enhance the response of network towards quality-degraded regions. As a result, our network can effectively improve the visual quality of underwater images by exploiting multiple color spaces embedding and the advantages of both physical model-based and learning-based methods. Extensive experiments demonstrate that our Ucolor achieves superior performance against state-of-the-art methods in terms of both visual quality and quantitative metrics. The code is publicly available at: https://li-chongyi.github.io/Proj_Ucolor.html. Chongyi Li, Saeed Anwar, Junhui Hou, Runmin Cong, Chunle Guo, Wenqi Ren |
IEEE Trans. Image Process. | 1 |
| 2021 | Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing ImagesabstractDespite the remarkable advances in visual saliency analysis for natural scene images (NSIs), salient object detection (SOD) for optical remote sensing images (RSIs) still remains an open and challenging problem. In this paper, we propose an end-to-end Dense Attention Fluid Network (DAFNet) for SOD in optical RSIs. A Global Context-aware Attention (GCA) module is proposed to adaptively capture long-range semantic context relationships, and is further embedded in a Dense Attention Fluid (DAF) structure that enables shallow attention cues flow into deep layers to guide the generation of high-level feature attention maps. Specifically, the GCA module is composed of two key components, where the global feature aggregation module achieves mutual reinforcement of salient feature embeddings from any two spatial locations, and the cascaded pyramid attention module tackles the scale variation issue by building up a cascaded pyramid framework to progressively refine the attention map in a coarse-to-fine manner. In addition, we construct a new and challenging optical RSI dataset for SOD that contains 2,000 images with pixel-wise saliency annotations, which is currently the largest publicly available benchmark. Extensive experiments demonstrate that our proposed DAFNet significantly outperforms the existing state-of-the-art SOD competitors. https://github.com/rmcong/DAFNet_TIP20. Qijian Zhang, Runmin Cong, Chongyi Li, Ming-Ming Cheng, Yuming Fang 0001, Xiaochun Cao, Yao Zhao 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2020 | Zero-Reference Deep Curve Estimation for Low-Light Image EnhancementabstractThe paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network. Our method trains a lightweight deep network, DCE-Net, to estimate pixel-wise and high-order curves for dynamic range adjustment of a given image. The curve estimation is specially designed, considering pixel value range, monotonicity, and differentiability. Zero-DCE is appealing in its relaxed assumption on reference images, i.e., it does not require any paired or unpaired data during training. This is achieved through a set of carefully formulated non-reference loss functions, which implicitly measure the enhancement quality and drive the learning of the network. Our method is efficient as image enhancement can be achieved by an intuitive and simple nonlinear curve mapping. Despite its simplicity, we show that it generalizes well to diverse lighting conditions. Extensive experiments on various benchmarks demonstrate the advantages of our method over state-of-the-art methods qualitatively and quantitatively. Furthermore, the potential benefits of our Zero-DCE to face detection in the dark are discussed. Chunle Guo, Chongyi Li, Jichang Guo, Chen Change Loy, Junhui Hou, Sam Kwong, Runmin Cong |
CVPR | 2 |
| 2020 | RGB-D Salient Object Detection with Cross-Modality Modulation and Selection
Chongyi Li, Runmin Cong, Yongri Piao, Qianqian Xu 0001, Chen Change Loy |
ECCV (8) | 1 |
| 2020 | NuI-Go: Recursive Non-Local Encoder-Decoder Network for Retinal Image Non-Uniform Illumination RemovalabstractRetinal images have been widely used by clinicians for early diagnosis of ocular diseases. However, the quality of retinal images is often clinically unsatisfactory due to eye lesions and imperfect imaging process. One of the most challenging quality degradation issues in retinal images is non-uniform which hinders the pathological information and further impairs the diagnosis of ophthalmologists and computer-aided analysis. To address this issue, we propose a non-uniform illumination removal network for retinal image, called NuI-Go, which consists of three Recursive Non-local Encoder-Decoder Residual Blocks (NEDRBs) for enhancing the degraded retinal images in a progressive manner. Each NEDRB contains a feature encoder module that captures the hierarchical feature representations, a non-local context module that models the context information, and a feature decoder module that recovers the details and spatial dimension. Additionally, the symmetric skip-connections between the encoder module and the decoder module provide long-range information compensation and reuse. Extensive experiments demonstrate that the proposed method can effectively remove the non-uniform illumination on retinal images while well preserving the image details and color. We further demonstrate the advantages of the proposed method for improving the accuracy of retinal vessel segmentation. Chongyi Li, Huazhu Fu, Runmin Cong, Zechao Li, Qianqian Xu 0001 |
ACM Multimedia | 1 |
| 2020 | CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object DetectionabstractCo-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue is how to effectively capture co-saliency cues by modeling and exploiting inter-image relationships. In this paper, we present an end-to-end collaborative aggregation-and-distribution network (CoADNet) to capture both salient and repetitive visual patterns from multiple images. First, we integrate saliency priors into the backbone features to suppress the redundant background information through an online intra-saliency guidance structure. After that, we design a two-stage aggregate-and-distribute architecture to explore group-wise semantic interactions and produce the co-saliency features. In the first stage, we propose a group-attentional semantic aggregation module that models inter-image relationships to generate the group-wise semantic representations. In the second stage, we propose a gated group distribution module that adaptively distributes the learned group semantics to different individuals in a dynamic gating mechanism. Finally, we develop a group consistency preserving decoder tailored for the CoSOD task, which maintains group constraints during feature decoding to predict more consistent full-resolution co-saliency maps. The proposed CoADNet is evaluated on four prevailing CoSOD benchmark datasets, which demonstrates the remarkable performance improvement over ten state-of-the-art competitors. Qijian Zhang, Runmin Cong, Junhui Hou, Chongyi Li, Yao Zhao 0001 |
NeurIPS | 4 |
| 2020 | A parallel down-up fusion network for salient object detection in optical remote sensing images
Chongyi Li, Runmin Cong, Chunle Guo, Hua Li 0012, Chunjie Zhang 0001, Feng Zheng 0001, Yao Zhao 0001 |
Neurocomputing | 1 |
| 2020 | A Feature Selection Algorithm Based on Equal Interval Division and Minimal-Redundancy-Maximal-Relevance
Xiangyuan Gu, Jichang Guo, Tao Ming, Chongyi Li |
Neural Process. Lett. | 5 |
| 2020 | Underwater scene prior inspired deep underwater image and video enhancement
Chongyi Li, Saeed Anwar, Fatih Porikli |
Pattern Recognit. | 1 |
| 2020 | Diving deeper into underwater image enhancement: A survey
Saeed Anwar, Chongyi Li |
Signal Process. Image Commun. | 2 |
| 2020 | An Underwater Image Enhancement Benchmark Dataset and BeyondabstractUnderwater image enhancement has been attracting much attention due to its significance in marine engineering and aquatic robotics. Numerous underwater image enhancement algorithms have been proposed in the last few years. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real-world images. It is thus unclear how these algorithms would perform on images acquired in the wild and how we could gauge the progress in the field. To bridge this gap, we present the first comprehensive perceptual study and analysis of underwater image enhancement using large-scale real-world images. In this paper, we construct an Underwater Image Enhancement Benchmark (UIEB) including 950 real-world underwater images, 890 of which have the corresponding reference images. We treat the rest 60 underwater images which cannot obtain satisfactory reference images as challenging data. Using this dataset, we conduct a comprehensive study of the state-of-the-art underwater image enhancement algorithms qualitatively and quantitatively. In addition, we propose an underwater image enhancement network (called Water-Net) trained on this benchmark as a baseline, which indicates the generalization of the proposed UIEB for training Convolutional Neural Networks (CNNs). The benchmark evaluations and the proposed Water-Net demonstrate the performance and limitations of state-of-the-art algorithms, which shed light on future research in underwater image enhancement. The dataset and code are available at. Chongyi Li, Chunle Guo, Wenqi Ren, Runmin Cong, Junhui Hou, Sam Kwong, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2020 | PDR-Net: Perception-Inspired Single Image Dehazing Network With RefinementabstractDuring recent years, we have witnessed a rapid development of wireless network technologies which have revolutionized the way people take and share multimedia content. However, images captured in the outdoor scenes usually suffer from limited visibility due to suspended atmospheric particles, which directly affects the quality of photos. Despite the recent progress of image dehazing methods, the visual quality of dehazed results still needs further improvement. In this paper, we propose a deep convolutional neural network (CNN) for single image dehazing called PDR-Net, which includes a perception-inspired haze removal subnetwork that reconstructs the latent dehazed image and a refinement subnetwork that further enhances the contrast and color properties of the dehazed result by joint multi-term loss optimization. Compared to the previous methods, our method combines the advantages of existing indoor and outdoor image dehazing training data, which makes the proposed PDR-Net generalized to various hazy images and effective for improving the visual quality of the dehazed results. Extensive experiments demonstrate that the proposed method achieves comparable and even better performance on both real and synthetic images in qualitative and quantitative metrics. Additionally, the potential usage of our method in high-level vision tasks is discussed. Chongyi Li, Chunle Guo, Jichang Guo, Ping Han, Huazhu Fu, Runmin Cong |
IEEE Trans. Multim. | 1 |
| 2019 | A visual hierarchical framework based model for underwater image enhancement
Bo Wang 0070, Chongyi Li |
Frontiers Comput. Sci. | 2 |
| 2019 | Nested Network With Two-Stream Pyramid for Salient Object Detection in Optical Remote Sensing ImagesabstractArising from the various object types and scales, diverse imaging orientations, and cluttered backgrounds in optical remote sensing image (RSI), it is difficult to directly extend the success of salient object detection for nature scene image to the optical RSI. In this paper, we propose an end-to-end deep network called LV-Net based on the shape of network architecture, which detects salient objects from optical RSIs in a purely data-driven fashion. The proposed LV-Net consists of two key modules, i.e., a two-stream pyramid module (L-shaped module) and an encoder-decoder module with nested connections (V-shaped module). Specifically, the L-shaped module extracts a set of complementary information hierarchically by using a two-stream pyramid structure, which is beneficial to perceiving the diverse scales and local details of salient objects. The V-shaped module gradually integrates encoder detail features with decoder semantic features through nested connections, which aims at suppressing the cluttered backgrounds and highlighting the salient objects. In addition, we construct the first publicly available optical RSI data set for salient object detection, including 800 images with varying spatial resolutions, diverse saliency types, and pixel-wise ground truth. Experiments on this benchmark data set demonstrate that the proposed method outperforms the state-of-the-art salient object detection methods both qualitatively and quantitatively. Chongyi Li, Runmin Cong, Junhui Hou, Sanyi Zhang, Sam Kwong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Hierarchical Features Driven Residual Learning for Depth Map Super-ResolutionabstractRapid development of affordable and portable consumer depth cameras facilitates the use of depth information in many computer vision tasks such as intelligent vehicles and 3D reconstruction. However, depth map captured by low-cost depth sensors (e.g., Kinect) usually suffers from low spatial resolution, which limits its potential applications. In this paper, we propose a novel deep network for depth map super-resolution (SR), called DepthSR-Net. The proposed DepthSR-Net automatically infers a high resolution (HR) depth map from its low resolution (LR) version by hierarchical features driven residual learning. Specifically, DepthSR-Net is built on a residual U-Net deep network architecture. Given LR depth map, we first obtain the desired HR by bicubic interpolation upsampling, and then construct an input pyramid to achieve multiple level receptive fields. Next, we extract hierarchical features from the input pyramid, intensity image, and encoder-decoder structure of UNet. Finally, we learn the residual between the interpolated depth map and the corresponding HR one using the rich hierarchical features. The final HR depth map is achieved by adding the learned residual to the interpolated depth map. We conduct an ablation study to demonstrate the effectiveness of each component in the proposed network. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art methods. Additionally, the potential usage of the proposed network in other low-level vision problems is discussed. Chunle Guo, Chongyi Li, Jichang Guo, Runmin Cong, Huazhu Fu, Ping Han |
IEEE Trans. Image Process. | 2 |
| 2018 | Image compressed sensing based on non-convex low-rank approximation
Jichang Guo, Chongyi Li |
Multim. Tools Appl. | 3 |
| 2018 | LightenNet: A Convolutional Neural Network for weakly illuminated image enhancement
Chongyi Li, Jichang Guo, Fatih Porikli, Yanwei Pang |
Pattern Recognit. Lett. | 1 |
| 2018 | Emerging From Water: Underwater Image Color Correction Based on Weakly Supervised Color TransferabstractUnderwater vision suffers from severe effects due to selective attenuation and scattering when light propagates through water. Such degradation not only affects the quality of underwater images, but limits the ability of vision tasks. Different from existing methods that either ignore the wavelength dependence on the attenuation or assume a specific spectral profile, we tackle color distortion problem of underwater images from a new view. In this letter, we propose a weakly supervised color transfer method to correct color distortion. The proposed method relaxes the need for paired underwater images for training and allows the underwater images being taken in unknown locations. Inspired by cycle-consistent adversarial networks, we design a multiterm loss function including adversarial loss, cycle consistency loss, and structural similarity index measure loss, which makes the content and structure of the outputs same as the inputs, meanwhile the color is similar to the images that were taken without the water. Experiments on underwater images captured under diverse scenes show that our method produces visually pleasing results, even outperforms the state-of-the-art methods. Besides, our method can improve the performance of vision tasks. Chongyi Li, Jichang Guo, Chunle Guo |
IEEE Signal Process. Lett. | 1 |
| 2017 | Compressive sensing in wireless multimedia sensor networks based on low-rank approximationabstractFor the large number of the image data produced by the sensor nodes in wireless multimedia sensor networks (WMSNs), the reduction of the sensor data and energy efficient transmission of this data are the most challenging problem. Compressed sensing based image compression provides the dramatic reduction of image sampling rates, energy consumption for WMSNs data collection and transmission. For better suitable for real-time sensing of image and reduce the sensing matrix size, block compressed sensing method acquire and process image in a block-by-block manner, by the same operator. Recently, nonlocal sparsity has been evidenced to improve the reconstruction of image details in various compressed sensing studies. In this paper, based on block compressed sensing, both the local sparsity of the whole image and nonlocal sparsity of similar patches are integrated into a compressed sensing recovery framework. And the nonlocal sparsity of images is characterized by the low-rank approximation. For better approximation of the rank function, the non-convex low-rank regularization namely Schatten p-norm minimization is applied for block compressed sensing recovery. The experiments show that the proposed algorithm can reduce the reconstruction error and improve the quality of the reconstruction image. Moreover, the proposed algorithm is suitable for WMSNs with few memory and low transmission bandwidth. Jichang Guo, Chongyi Li |
ICC | 3 |
| 2017 | A hybrid method for underwater image correction
Chongyi Li, Jichang Guo, Chunle Guo, Runmin Cong, Jiachang Gong |
Pattern Recognit. Lett. | 1 |
| 2017 | Hierarchical feature concatenation-based kernel sparse representations for image categorization
Bo Wang 0070, Jichang Guo, Chongyi Li |
Vis. Comput. | 4 |
| 2016 | Single underwater image restoration by blue-green channels dehazing and red channel correctionabstractRestoring underwater image from a single image is know to be ill-posed, and some assumptions made in previous methods are not suitable for many situations. In this paper, we propose a method based on blue-green channels dehazing and red channel correction for underwater image restoration. Firstly, blue-green channels are recovered via dehazing algorithm based on an extension and modification of Dark Channel Prior algorithm. Then, red channel is corrected following the Gray-World assumption theory. Finally, in order to resolve the problem which some recovered image regions may look too dim or too bright, an adaptive exposure map is built. Qualitative analysis demonstrates that our method significantly improves visibility and contrast, and reduces the effects of light absorption and scattering. For quantitative analysis, our results obtain best values in terms of entropy, local feature points and average gradient, which outperform three existing physical model available methods. Chongyi Li, Jichang Quo, Yanwei Pang, Shanji Chen, Jian Wang 0087 |
ICASSP | 1 |
| 2016 | Underwater image restoration based on minimum information loss principle and optical properties of underwater imagingabstractRestoring underwater image from a single image is known to be an ill-posed problem. Some assumptions made in previous methods are not suitable in many situations. In this paper, an effective method is proposed to restore underwater images. Using the quad-tree subdivision and graph-based segmentation, the global background light can be robustly estimated. The medium transmission map is estimated based on minimum information loss principle and optical properties of underwater imaging. Qualitative experiments show that our results are characterized by relatively genuine color, natural appearance, and improved contrast and visibility. Quantitative comparisons demonstrate that the proposed method can achieve better quality of underwater images when compared with several other methods. Chongyi Li, Jichang Guo, Shanji Chen, Yibin Tang, Yanwei Pang, Jian Wang 0087 |
ICIP | 1 |
| 2016 | Underwater Image Enhancement by Dehazing With Minimum Information Loss and Histogram Distribution PriorabstractImages captured under water are usually degraded due to the effects of absorption and scattering. Degraded underwater images show some limitations when they are used for display and analysis. For example, underwater images with low contrast and color cast decrease the accuracy rate of underwater object detection and marine biology recognition. To overcome those limitations, a systematic underwater image enhancement method, which includes an underwater image dehazing algorithm and a contrast enhancement algorithm, is proposed. Built on a minimum information loss principle, an effective underwater image dehazing algorithm is proposed to restore the visibility, color, and natural appearance of underwater images. A simple yet effective contrast enhancement algorithm is proposed based on a kind of histogram distribution prior, which increases the contrast and brightness of underwater images. The proposed method can yield two versions of enhanced output. One version with relatively genuine color and natural appearance is suitable for display. The other version with high contrast and brightness can be used for extracting more valuable information and unveiling more details. Simulation experiment, qualitative and quantitative comparisons, as well as color accuracy and application tests are conducted to evaluate the performance of the proposed method. Extensive experiments demonstrate that the proposed method achieves better visual quality, more valuable information, and more accurate color restoration than several state-of-the-art methods, even for underwater images taken under several challenging scenes. Chongyi Li, Jichang Guo, Runmin Cong, Yanwei Pang, Bo Wang 0070 |
IEEE Trans. Image Process. | 1 |