Xiang Chen 0015

dblp:64/3062-15 · DBLP profile ↗
← Back
34ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0002-8966-8159ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 20 · 8 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Rethinking Rainy 3D Scene Reconstruction via Perspective Transforming and Brightness Tuning
abstract
Rain degrades the visual quality of multi-view images, which are essential for 3D scene reconstruction, resulting in inaccurate and incomplete reconstruction results. Existing datasets often overlook two critical characteristics of real rainy 3D scenes: the viewpoint-dependent variation in the appearance of rain streaks caused by their projection onto 2D images, and the reduction in ambient brightness resulting from cloud coverage during rainfall. To improve data realism, we construct a new dataset named OmniRain3D that incorporates perspective heterogeneity and brightness dynamicity, enabling more faithful simulation of rain degradation in 3D scenes. Based on this dataset, we propose an end-to-end reconstruction framework named REVR-GSNet (Rain Elimination and Visibility Recovery for 3D Gaussian Splatting). Specifically, REVR-GSNet integrates recursive brightness enhancement, Gaussian primitive optimization, and GS-guided rain elimination into a unified architecture through joint alternating optimization, achieving high-fidelity reconstruction of clean 3D scenes from rain-degraded inputs. Extensive experiments show the effectiveness of our dataset and method. Our dataset and method provide a foundation for future research on multi-view image deraining and rainy 3D scene reconstruction.
Qianfeng Yang, Xiang Chen 0015, Pengpeng Li 0001, Qiyuan Guan, Guiyue Jin, Jiyu Jin
AAAI2
2026 Convergence-aware task scheduling with position-constrained semantic Mamba for all-in-one adverse weather image restoration
Xianhao Wu, Guili Xu, Xiang Chen 0015, Qianfeng Yang, Qiyuan Guan
Neurocomputing3
2026 Dual prompts guided cross-domain transformer for unified day-night image dehazing
Jianlei Liu, Jiaming Niu, Xiang Chen 0015, Yuting Pang, Shilong Wang 0005
Knowl. Based Syst.3
2026 Collaborative Feedback Discriminative Propagation for Video Super-Resolution
abstract
The key success of existing video super-resolution (VSR) methods stems mainly from exploring spatial and temporal information that is usually achieved by a temporal propagation with alignment strategies. However, inaccurate alignment usually leads to significant artifacts that will be accumulated during propagation and thus affect video restoration. Moreover, only propagating the same timestep features forward or backward does not handle the videos with complex motion or occlusion. To address these issues, we propose a collaborative feedback discriminative (CFD) method to correct inaccurate aligned features and better model spatial and temporal information for VSR. Specifically, we first develop a discriminative alignment correction (DAC) method to reduce the influences of the artifacts caused by inaccurate alignment. Then, we propose a collaborative feedback propagation (CFP) module based on feedback and gating mechanisms to explore spatial and temporal information of different timestep features from forward and backward propagation simultaneously. Finally, we embed the proposed DAC and CFP into commonly used VSR networks to verify the effectiveness of our method. Experimental results demonstrate that our method improves the performance of existing VSR models while maintaining a lower model complexity.
Hao Li 0058, Xiang Chen 0015, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Haze has many faces: Multi-domain haze style transfer for diverse haze removal
Cunchuan Huang, Shuai Li 0005, Xiang Chen 0015, Jianlei Liu, Dengwang Li
Pattern Recognit.3
2026 Towards Ultra-High-Definition Image Deraining: A Benchmark and an Efficient Method
abstract
Despite significant advancements in image deraining, most existing methods are carried out on low-resolution images, leaving their effectiveness on high-resolution images uncertain. This limitation becomes even more pronounced with the rise of ultra-high-definition (UHD) imaging. In this paper, we tackle the challenge of UHD image deraining and introduce 4K-Rain13 k, the first large-scale UHD image deraining dataset, featuring 13,000 paired images at 4 K resolution. Leveraging this dataset, we conduct a benchmark study on existing methods for processing UHD images. To better address this task, we propose UDR-Mixer, an efficient and effective architecture tailored for UHD image deraining. Our model comprises two key components: a spatial feature rearrangement layer, which captures long-range dependencies in UHD images, and a frequency feature modulation layer, which enhances high-fidelity image reconstruction. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while maintaining lower model complexity. The source code and proposed dataset are available athttps://github.com/cschenxiang/UDR-Mixer.
Hongming Chen 0004, Xiang Chen 0015, Chen Wu 0006, Zhuoran Zheng, Jinshan Pan, Xianping Fu
IEEE Trans. Multim.2
2025 Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving Video
abstract
In this paper, we study the challenging problem of simultaneously removing haze and estimating depth from real monocular hazy videos. These tasks are inherently complementary: enhanced depth estimation improves dehazing via the atmospheric scattering model (ASM), while superior dehazing contributes to more accurate depth estimation through the brightness consistency constraint (BCC). To tackle these intertwined tasks, we propose a novel depth-centric learning framework that integrates the ASM model with the BCC constraint. Our key idea is that both ASM and BCC rely on a shared depth estimation network. This network simultaneously exploits adjacent dehazed frames to enhance depth estimation via BCC and uses the refined depth cues to more effectively remove haze through ASM. Additionally, we leverage a non-aligned clear video and its estimated depth to independently regularize the dehazing and depth estimation networks. This is achieved by designing two discriminator networks: D_MFIR enhances high-frequency details in dehazed videos, and D_MDR reduces the occurrence of black holes in low-texture regions. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art techniques in both video dehazing and depth estimation tasks, especially in real-world hazy scenes.
Junkai Fan, Kun Wang 0042, Zhiqiang Yan 0001, Xiang Chen 0015, Shangbing Gao, Jun Li 0027, Jian Yang 0003
AAAI4
2025 DeRainGS: Gaussian Splatting for Enhanced Scene Reconstruction in Rainy Environments
abstract
Reconstruction under adverse rainy conditions poses significant challenges due to reduced visibility and the distortion of visual perception. These conditions can severely impair the quality of geometric maps, which is essential for applications ranging from autonomous planning to environmental monitoring. In response to these challenges, this study introduces the novel task of 3D Reconstruction in Rainy Environments (3DRRE), specifically designed to address the complexities of reconstructing 3D scenes under rainy conditions. To benchmark this task, we construct the HydroViews dataset that comprises a diverse collection of both synthesized and real-world scene images characterized by various intensities of rain streaks and raindrops. Furthermore, we propose DeRainGS, the first 3DGS method tailored for reconstruction in adverse rainy environments. Extensive experiments across a wide range of rain scenarios demonstrate that our method delivers state-of-the-art performance, remarkably outperforming existing occlusion-free methods by a large margin.
Shuhong Liu, Xiang Chen 0015, Hongming Chen 0004, Quanfeng Xu
AAAI2
2025 FoundIR: Unleashing Million-Scale Training Data to Advance Foundation Models for Image Restoration
abstract
Despite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with limited degradations. Therefore, large-scale high-quality real-world training data is urgently needed to facilitate the emergence of foundational models for image restoration. To advance this field, we spare no effort in contributing a million-scale dataset with two notable advantages over existing training data: real-world samples with larger-scale, and degradation types with higher diversity. By adjusting internal camera settings and external imaging conditions, we can capture aligned image pairs using our well-designed data acquisition system over multiple rounds and our data alignment criterion. Moreover, we propose a robust model, FoundIR, to better address a broader range of restoration tasks in real-world scenarios, taking a further step toward foundation models. Specifically, we first utilize a diffusion-based generalist model to remove degradations by learning the degradation-agnostic common representations from diverse inputs, where incremental learning strategy is adopted to better guide model training. To refine the model's restoration capability in complex scenarios, we introduce degradation-aware specialist models for achieving final high-quality results. Extensive experiments show the value of our dataset and the effectiveness of our method.
Hao Li 0058, Xiang Chen 0015, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan
ICCV2
2025 K-Buffers: A Plug-in Method for Enhancing Neural Fields with Multiple Buffers
abstract
Neural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose a plug-in method named K-Buffers that leverages multiple buffers to improve the rendering performance. Our method first renders K buffers from scene representations and constructs K pixel-wise feature maps. Then, We introduce a K-Feature Fusion Network (KFN) to merge the K pixel-wise feature maps. Finally, we adopt a feature decoder to generate the rendering image. We also introduce an acceleration strategy to improve rendering speed and quality. We apply our method to well-known radiance field baselines, including neural point fields and 3D Gaussian Splatting (3DGS). Extensive experiments demonstrate that our method effectively enhances the rendering performance of neural point fields and 3DGS.
Haofan Ren, Zunjie Zhu, Xiang Chen 0015, Ming Lu 0002, Rongfeng Lu, Chenggang Yan 0001
IJCAI3
2025 WeatherBench: A Real-World Benchmark Dataset for All-in-One Adverse Weather Image Restoration
abstract
Existing all-in-one image restoration approaches, which aim to handle multiple weather degradations within a single framework, are predominantly trained and evaluated using mixed single-weather synthetic datasets. However, these datasets often differ significantly in resolution, style, and domain characteristics, leading to substantial domain gaps that hinder the development and fair evaluation of unified models. Furthermore, the lack of a large-scale, real-world all-in-one weather restoration dataset remains a critical bottleneck in advancing this field. To address these limitations, we present a real-world all-in-one adverse weather image restoration benchmark dataset, which contains image pairs captured under various weather conditions, including rain, snow, and haze, as well as diverse outdoor scenes and illumination settings. The resulting dataset provides precisely aligned degraded and clean images, enabling supervised learning and rigorous evaluation. We conduct comprehensive experiments by benchmarking a variety of task-specific, task-general, and all-in-one restoration methods on our dataset. Our dataset offers a valuable foundation for advancing robust and practical all-in-one image restoration in real-world scenarios. The dataset has been publicly released and is available at https://github.com/guanqiyuan/WeatherBench.
Qiyuan Guan, Qianfeng Yang, Xiang Chen 0015, Tianyu Song 0003, Guiyue Jin, Jiyu Jin
ACM Multimedia3
2025 SmokeBench: A Real-World Dataset for Surveillance Image Desmoking in Early-Stage Fire Scenes
abstract
Early-stage fire scenes (0-15 minutes after ignition) represent a crucial temporal window for emergency interventions. During this stage, the smoke produced by combustion significantly reduces the visibility of surveillance systems, severely impairing situational awareness and hindering effective emergency response and rescue operations. Consequently, there is an urgent need to remove smoke from images to obtain clear scene information. However, the development of smoke removal algorithms remains limited due to the lack of large-scale, real-world datasets comprising paired smoke-free and smoke-degraded images. To address these limitations, we present a real-world surveillance image desmoking benchmark dataset named SmokeBench, which contains image pairs captured under diverse scenes setup and smoke concentration. The curated dataset provides precisely aligned degraded and clean images, enabling supervised learning and rigorous evaluation. We conduct comprehensive experiments by benchmarking a variety of desmoking methods on our dataset. Our dataset provides a valuable foundation for advancing robust and practical image desmoking in real-world fire scenes. This dataset has been released to the public and can be downloaded from https://github.com/ncfjd/SmokeBench.
Wenzhuo Jin, Qianfeng Yang, Xianhao Wu, Hongming Chen 0004, Pengpeng Li 0001, Xiang Chen 0015
ACM Multimedia6
2025 Rethinking Nighttime Image Deraining via Learnable Color Space Transformation
abstract
Compared to daytime image deraining, nighttime image deraining poses significant challenges due to inherent complexities of nighttime scenarios and the lack of high-quality datasets that accurately represent the coupling effect between rain and illumination. In this paper, we rethink the task of nighttime image deraining and contribute a new high-quality benchmark, HQ-NightRain, which offers higher harmony and realism compared to existing datasets. In addition, we develop an effective Color Space Transformation Network (CST-Net) for better removing complex rain from nighttime scenes. Specifically, we propose a learnable color space converter (CSC) to better facilitate rain removal in the Y channel, as nighttime rain is more pronounced in the Y channel compared to the RGB color space. To capture illumination information for guiding nighttime deraining, implicit illumination guidance is introduced enabling the learned features to improve the model's robustness in complex scenarios. Extensive experiments show the value of our dataset and the effectiveness of our method. The source code and datasets are available at https://github.com/guanqiyuan/CST-Net.
Qiyuan Guan, Xiang Chen 0015, Guiyue Jin, Jiyu Jin, Shumin Fan, Tianyu Song 0003, Jinshan Pan
NeurIPS2
2025 Towards Unified Deep Image Deraining: A Survey and a New Benchmark
abstract
Recent years have witnessed significant advances in image deraining due to the progress of effective image priors and deep learning models. As each deraining approach has individual settings (e.g., training and test datasets, evaluation criteria), how to fairly evaluate existing approaches comprehensively is not a trivial task. Although existing surveys aim to thoroughly review image deraining approaches, few of them focus on unifying evaluation settings to examine the deraining capability and practicality evaluation. In this paper, we provide a comprehensive review of existing image deraining methods and provide a unified evaluation setting to evaluate their performance. Furthermore, we construct a new high-quality benchmark named HQ-RAIN to conduct extensive evaluations, consisting of 5,000 paired high-resolution synthetic images with high harmony and realism. We also discuss existing challenges and highlight several future research opportunities worth exploring. To facilitate the reproduction and tracking of the latest deraining technologies for general users, we build an online platform to provide the off-the-shelf toolkit, involving the large-scale performance evaluation.
Xiang Chen 0015, Jinshan Pan, Jiangxin Dong, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Rethinking Multi-Scale Representations in Deep Deraining Transformer
abstract
Existing Transformer-based image deraining methods depend mostly on fixed single-input single-output U-Net architecture. In fact, this not only neglects the potentially explicit information from multiple image scales, but also lacks the capability of exploring the complementary implicit information across different scales. In this work, we rethink the multi-scale representations and design an effective multi-input multi-output framework that constructs intra- and inter-scale hierarchical modulation to better facilitate rain removal and help image restoration. We observe that rain levels reduce dramatically in coarser image scales, thus proposing to restore rain-free results from the coarsest scale to the finest scale in image pyramid inputs, which also alleviates the difficulty of model learning. Specifically, we integrate a sparsity-compensated Transformer block and a frequency-enhanced convolutional block into a coupled representation module, in order to jointly learn the intra-scale content-aware features. To facilitate representations learned at different scales to communicate with each other, we leverage a gated fusion module to adaptively aggregate the inter-scale spatial-aware features, which are rich in correlated information of rain appearances, leading to high-quality results. Extensive experiments demonstrate that our model achieves consistent gains on five benchmarks.
Hongming Chen 0004, Xiang Chen 0015, Jiyang Lu, Yufeng Li 0001
AAAI2
2024 Bidirectional Multi-Scale Implicit Neural Representations for Image Deraining
abstract
How to effectively explore multi-scale representations of rain streaks is important for image deraining. In contrast to existing Transformer-based methods that depend mostly on single-scale rain appearance, we develop an end-to-end multi-scale Transformer that leverages the potentially useful features in various scales to facilitate high-quality image reconstruction. To better explore the common degradation representations from spatially-varying rain streaks, we incorporate intra-scale implicit neural representations based on pixel coordinates with the degraded inputs in a closed-loop design, enabling the learned features to facilitate rain removal and improve the robustness of the model in complex scenarios. To ensure richer collaborative representation from different scales, we embed a simple yet effective inter-scale bidirectional feedback operation into our multi-scale Transformer by performing coarse-to-fine and fine-to-coarse information communication. Extensive experiments demonstrate that our approach, named as NeRD-Rain, performs favorably against the state-of-the-art ones on both synthetic and real-world benchmark datasets. The source code and trained models are available at https://github.com/cschenxiang/NeRD-Rain.
Xiang Chen 0015, Jinshan Pan, Jiangxin Dong
CVPR1
2024 Learning a Spiking Neural Network for Efficient Image Deraining
Tianyu Song 0003, Guiyue Jin, Pengpeng Li 0001, Kui Jiang, Xiang Chen 0015, Jiyu Jin
IJCAI5
2024 SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) is an emerging technology, that generates a multienergy attenuation map for the interior of an object and extends the traditional image volume into a 4-D form. Compared with traditional CT based on energy-integrating detectors, spectral CT can make full use of spectral information, resulting in high resolution and providing accurate material quantification. Numerous model-based iterative reconstruction methods have been proposed for spectral CT reconstruction. However, these methods usually suffer from difficulties such as laborious parameter selection and expensive computational costs. In addition, due to the image similarity of different energy bins, spectral CT usually implies a strong low-rank prior, which has been widely adopted in current iterative reconstruction models. Singular value thresholding (SVT) is an effective algorithm to solve the low-rank constrained model. However, the SVT method requires a manual selection of thresholds, which may lead to suboptimal results. To relieve these problems, in this article, we propose a sparse and low-rank unrolling network (SOUL-Net) for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner. Furthermore, a Taylor expansion-based neural network backpropagation method is introduced to improve the numerical stability. The qualitative and quantitative results demonstrate that the proposed method outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction.
Xiang Chen 0015, Wenjun Xia, Ziyuan Yang 0001, Hu Chen 0002, Yan Liu 0052, Jiliu Zhou, Yang Chen 0008, Bihan Wen, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.1
2023 Hybrid CNN-Transformer Feature Fusion for Single Image Deraining
abstract
Since rain streaks exhibit diverse geometric appearances and irregular overlapped phenomena, these complex characteristics challenge the design of an effective single image deraining model. To this end, rich local-global information representations are increasingly indispensable for better satisfying rain removal. In this paper, we propose a lightweight Hybrid CNN-Transformer Feature Fusion Network (dubbed as HCT-FFN) in a stage-by-stage progressive manner, which can harmonize these two architectures to help image restoration by leveraging their individual learning strengths. Specifically, we stack a sequence of the degradation-aware mixture of experts (DaMoE) modules in the CNN-based stage, where appropriate local experts adaptively enable the model to emphasize spatially-varying rain distribution features. As for the Transformer-based stage, a background-aware vision Transformer (BaViT) module is employed to complement spatially-long feature dependencies of images, so as to achieve global texture recovery while preserving the required structure. Considering the indeterminate knowledge discrepancy among CNN features and Transformer features, we introduce an interactive fusion branch at adjacent stages to further facilitate the reconstruction of high-quality deraining results. Extensive evaluations show the effectiveness and extensibility of our developed HCT-FFN. The source code is available at https://github.com/cschenxiang/HCT-FFN.
Xiang Chen 0015, Jinshan Pan, Jiyang Lu, Zhentao Fan, Hao Li 0058
AAAI1
2023 Learning A Sparse Transformer Network for Effective Image Deraining
abstract
Transformers-based methods have achieved significant performance in image deraining as they can model the non-local information which is vital for high-quality image reconstruction. In this paper, we find that most existing Transformers usually use all similarities of the tokens from the query-key pairs for the feature aggregation. However, if the tokens from the query are different from those of the key, the self-attention values estimated from these tokens also involve in feature aggregation, which accordingly interferes with the clear image restoration. To overcome this problem, we propose an effective DeRaining network, Sparse Transformer (DRSformer) that can adaptively keep the most useful self-attention values for feature aggregation so that the aggregated features better facilitate high-quality image reconstruction. Specifically, we develop a learnable top-k selection operator to adaptively retain the most crucial attention scores from the keys for each query for better feature aggregation. Simultaneously, as the naive feed-forward network in Transformers does not model the multi-scale information that is important for latent clear image restoration, we develop an effective mixed-scale feed-forward network to generate better features for image deraining. To learn an enriched set of hybrid features, which combines local context from CNN operators, we equip our model with mixture of experts feature compensator to present a cooperation refinement deraining scheme. Extensive experimental results on the commonly used benchmarks demonstrate that the proposed method achieves favorable performance against state-of-the-art approaches. The source code and trained models are available at https://github.com/cschenxiang/DRSformer.
Xiang Chen 0015, Hao Li 0058, Mingqiang Li, Jinshan Pan
CVPR1
2023 Image Deraining Transformer with Sparsity and Frequency Guidance
abstract
In recent years, Transformer has witnessed significant progress in the single image deraining field. However, most existing methods do not consider the latent sparse representation and distinguished frequency information. To this end, this paper proposes an effective Image Deraining Transformer with Sparsity and Frequency Guidance, called SFG-IDT. To achieve such guidance, the proposed method consists two key designs: sparsity-compensated multi-head attention (SCMA) and frequency-enhanced multi-scale operator (FEMO). Specifically, the SCMA enhances the concentration of attention while explicitly retaining non-local connectivity with Locality Sensitive Hashing (LSH), to facilitate rain removal better and help image restoration. Simultaneously, the FEMO integrates the frequency information into the multi-scale convolution operators with Fast Fourier Transform (FFT) to obtain a more accurate representation for achieving high-quality derained results. Extensive experimental results show that our developed SFG-IDT outperforms the state-of-the-art approach (Restormer) by 0.27 dB on average, but saves 50.3% parameters and 46.7% computational cost.
Tianyu Song 0003, Pengpeng Li 0001, Guiyue Jin, Jiyu Jin, Shumin Fan, Xiang Chen 0015
ICME6
2022 Unpaired Deep Image Deraining Using Dual Contrastive Learning
abstract
Learning single image deraining (SID) networks from an unpaired set of clean and rainy images is practical and valuable as acquiring paired real-world data is almost infeasible. However, without the paired data as the supervision, learning a SID network is challenging. Moreover, simply using existing unpaired learning methods (e.g., unpaired adversarial learning and cycle-consistency constraints) in the SID task is insufficient to learn the underlying relationship from rainy inputs to clean outputs as there exists significant domain gap between the rainy and clean images. In this paper, we develop an effective unpaired SID adversarial framework which explores mutual properties of the unpaired exemplars by a dual contrastive learning manner in a deep feature space, named as DCD-GAN. The proposed method mainly consists of two cooperative branches: Bidirectional Translation Branch (BTB) and Contrastive Guidance Branch (CGB). Specifically, BTB exploits full advantage of the circulatory architecture of adversarial consistency to generate abundant exemplar pairs and excavates latent feature distributions between two domains by equipping it with bidirectional mapping. Simultaneously, CGB implicitly constrains the embeddings of different exemplars in the deep feature space by encouraging the similar feature distributions closer while pushing the dissimilar further away, in order to better facilitate rain removal and help image restoration. Extensive experiments demonstrate that our method performs favorably against existing unpaired deraining approaches on both synthetic and real-world datasets, and generates comparable results against several fully-supervised or semi-supervised models.
Xiang Chen 0015, Jinshan Pan, Kui Jiang, Yufeng Li 0001, Caihua Kong, Longgang Dai, Zhentao Fan
CVPR1
2022 Unpaired Deep Image Dehazing Using Contrastive Disentanglement Learning
Xiang Chen 0015, Zhentao Fan, Pengpeng Li 0001, Longgang Dai, Caihua Kong, Zhuoran Zheng, Yufeng Li 0001
ECCV (17)1
2022 HD-Net: Hierarchical Distillation Network for High-Efficiency Single Image Deraining
abstract
Rain streaks usually result in severe image visual degradation and foreground occlusion, affecting the quality of computer tasks in outdoor scenes. Currently, the mainstream methods in single-image deraining are based on data-driven. However, the deep learning network could be imperfect, with limited power for learning the global information from rain streaks all over the map. In order to solve this problem, we proposed a novel Hierarchical Distillation Network (HD-Net). In this network, Hierarchical Feature Extraction Block (HFEB) can fully utilize the Transformer's learning ability in high-level features, integrate local detail extraction and global structure representation, and compensate for the weakness of the Convolutional Neural Network (CNN), which is overattentive to the underlying image features. Furthermore, the Distillation-Calibration Block (DCB) are adopted to avoid feature redundancy during model training and calibrate the channel and spatial information through the feature transmission, which could significantly improve the learning efficiency. Finally, the experiment results show that our model performs better than traditional CNN models and state-of-the-art methods.
Kejian Hu, Zhichen Zhang, Xiang Chen 0015, Nanfeng Jiang, Yu Zhou 0048, Tiesong Zhao
MMSP4
2022 Memory-Oriented Unpaired Learning for Single Remote Sensing Image Dehazing
abstract
Remote sensing image dehazing (RSID) is an extremely challenging problem due to the irregular and nonuniform distribution of haze. The existing RSID methods achieve excellent performance using deep learning; however, relying on paired synthetic data is limited to their generality in various haze distribution. In this letter, we present a memory-oriented generative adversarial network (MO-GAN), which tries to capture the desired hazy features in an unpaired learning manner toward single RSID. For better extracting the haze-relevant features, a novel multistage attentive-recurrent memory module is designed to guide an autoencoder neural network, which can record the various appearances of haze distribution at different stages. To well differentiate fake images from real ones, a dual region discriminator is constructed to handle spatially varying haze densities in global and local regions. Extensive experiments demonstrate that our designed MO-GAN outperforms the recent comparing approaches on the various frequently used datasets, especially in real world nonuniform haze conditions. The source code is released inhttps://github.com/cxtalk/MO-GAN.
Xiang Chen 0015
IEEE Geosci. Remote. Sens. Lett.1
2022 Hybrid High-Resolution Learning for Single Remote Sensing Satellite Image Dehazing
abstract
Recently, deep learning models have shown convincing performance in removing a single satellite image haze, which arouses increasing attention in the field of remote sensing (RS). Unfortunately, these models still suffer from an insufficient ability to recover the desired fine spatial details from the hazy image. In this letter, we first attempt to explore an end-to-end hybrid high-resolution learning network framework termed H2RL-Net to address this issue due to its novel feature extraction architecture, where spatially precise outputs are guaranteed by the main high-resolution branch and semantically richer features are collected by the complementary set of multiresolution convolution streams. To improve representation learning, H2RL-Net is constructed primarily by exploiting the parallel cross-scale fusion (PCF) module, thereby increasingly aggregating information from the multiple scales at the respective resolution level, which allows both top-down and bottom-up information exchanging processes. Simultaneously, we also introduce the channel feature refinement (CFR) block to our model, aiming to perform dynamic feature recalibration among the channelwise features and produce better dehazed results. The experimental analysis illustrates that the designed framework can deliver significant improvements over other baseline methods in the synthetic and real-world hazy RS images under various scenes.
Xiang Chen 0015, Yufeng Li 0001, Longgang Dai, Caihua Kong
IEEE Geosci. Remote. Sens. Lett.1
2022 Single-Stage Detector With Dual Feature Alignment for Remote Sensing Object Detection
abstract
As a fundamental vision-based task in the remote sensing filed, object detection has achieved significant progress. However, remote sensing object detection is still an urgent challenge owing to dense distribution, large aspect ratio, and arbitrary orientations. To address this issue, we develop an end-to-end dual align single-stage rotation detector (DA-Net) consisting of two main components: a Rotation Feature Selection (RFS) module and a Rotation Feature Align (RFA) module. Specifically, RFS module can empower neurons with the capability to adjust receptive fields, which achieves the first stage of feature alignment on the image level. Furthermore, RFA module is employed to adaptively align the feature based on the size, shapes, orientations of its corresponding anchors, realizing the second stage of instance-level feature alignment. Extensive experiments have shown that our DA-Net can significantly improve remote sensing detection performance against several start-of-the-art algorithms on two benchmark datasets.
Yufeng Li 0001, Caihua Kong, Longgang Dai, Xiang Chen 0015
IEEE Geosci. Remote. Sens. Lett.4
2022 A deep hourglass-structured fusion model for efficient single image dehazing
Yufeng Li 0001, Xiang Chen 0015, Caihua Kong, Longgang Dai
Multim. Tools Appl.2
2022 Unpaired Image Dehazing With Physical-Guided Restoration and Depth-Guided Refinement
abstract
Most existing single image dehazing methods aim to learn supervised models from paired synthetic data, which often limits their generalization ability in real-world applications. Besides, due to ignoring the merits of physical model in visibility restoration and the properties of depth features in clarity improvement, we observe that only relying on the transfer capability of unpaired adversarial learning will suffer from low-quality recovery. To this end, we develop an effective end-to-end unpaired image dehazing method by integrating a physical-guided restoration stage and a depth-guided refinement stage in a GAN framework, named as PDR-GAN. Specifically, the dark channel prior is embedded in the restoration stage to provide constraints for the network, and the preliminary dehazed image is first generated. For the refinement stage, we excavate the potential relationship between the depth and transmission map to better refine the results of the previous stage and further recover the distant area details. Our framework benefits from the stage-wise learning strategy of model-based restoration and feature-based reconstruction, which is especially helpful for image dehazing when paired data is not available. Experimental results show that our method is superior to the current unpaired dehazing approaches in terms of both quantitative and qualitative.
Xiang Chen 0015, Yufeng Li 0001, Caihua Kong, Longgang Dai
IEEE Signal Process. Lett.1
2022 TAO-Net: Task-Adaptive Operation Network for Image Restoration and Enhancement
abstract
Recently, deep learning based models have been extensively applied to image restoration and enhancement tasks. However, existing approaches neglect the potential correlation among these tasks, and do not fully consider the different intensity factors of degradation which affect image quality. To this end, we propose Task-Adaptive Operation Network (TAO-Net) that stacks a series of operation blocks, guided by the supervised attention mechanism to adaptively enable the model to assign corresponding intensity weights for different degradation factors depending on the input signals. To better recover and preserve the accurate details during the repeated operations, we make full use of the image prior to intruduce structure refinement block as auxiliary information for reconstructing high-quality results. Furthermore, feature aggregation block is also adopted to avoid feature interference caused by the direct concatenation between the backbone operation block and the auxiliary refinement block together. Extensive experiments on multiple benchmark datasets demonstrate that our proposed method outperforms previous baselines in image deraining and low-light image enhancement.
Yufeng Li 0001, Zhentao Fan, Jiyang Lu, Xiang Chen 0015
IEEE Signal Process. Lett.4
2022 FONT-SIR: Fourth-Order Nonlocal Tensor Decomposition Model for Spectral CT Image Reconstruction
abstract
Spectral computed tomography (CT) reconstructs images from different spectral data through photon counting detectors (PCDs). However, due to the limited number of photons and the counting rate in the corresponding spectral segment, the reconstructed spectral images are usually affected by severe noise. In this paper, we propose a fourth-order nonlocal tensor decomposition model for spectral CT image reconstruction (FONT-SIR). To maintain the original spatial relationships among similar patches and improve the imaging quality, similar patches without vectorization are grouped in both spectral and spatial domains simultaneously to form the fourth-order processing tensor unit. The similarity of different patches is measured with the cosine similarity of latent features extracted using principal component analysis (PCA). By imposing the constraints of the weighted nuclear and total variation (TV) norms, each fourth-order tensor unit is decomposed into a low-rank component and a sparse component, which can efficiently remove noise and artifacts while preserving the structural details. Moreover, the alternating direction method of multipliers (ADMM) is employed to solve the decomposition model. Extensive experimental results on both simulated and real data sets demonstrate that the proposed FONT-SIR achieves superior qualitative and quantitative performance compared with several state-of-the-art methods.
Xiang Chen 0015, Wenjun Xia, Yan Liu 0052, Hu Chen 0002, Jiliu Zhou, Zhiyuan Zha, Bihan Wen, Yi Zhang 0018
IEEE Trans. Medical Imaging1
2021 Single Remote Sensing Image Dehazing Using a Dual-Step Cascaded Residual Dense Network
abstract
Remote sensing (RS) dehazing is an extremely challenging task since the non-uniform distribution of haze and fog severely degrade the images and difficult to extract features. To address these issues, we propose an end to end Dual-step Cascaded Residual Dense Network called DCRD-Net, which can exactly remove haze from the hazy RS image and precisely restoring the details. The architecture of network contains two cascaded task-driven subnetworks, in order to deal with the coarse and fine haze-relevant features separately. Besides, the Residual Dense Enhancement Block (RDEB) is involved to guide the feature extraction, so that multi-scale information can be used to estimate the local and global features. Further, the Squeeze and Excitation (SE) Block is employed to optimize the RDEB, for the purpose to get the contextual feature and reduce computation complex. Quantitative and qualitative results illustrate that the designed framework outperforms the recent outstanding dehazing methods on promoting haze remove and restoring detailed in the synthetic and real-world RS images under various scenes. To encourage more comparisons, we release our codes on GitHub https://github.com/cxtalk/DCRD-Net.
Xiang Chen 0015
ICIP2
2021 A Coarse-to-Fine Two-Stage Attentive Network for Haze Removal of Remote Sensing Images
abstract
In many remote sensing (RS) applications, haze seriously degrades the quality of optical RS images and even brings inconvenience to the following high-level visual tasks such as RS detection. In this letter, we address this challenge by designing a first-coarse-then-fine two-stage dehazing neural network, named FCTF-Net. The structure is simple but effective: the first stage of image dehazing extracts multiscale features through the encoder–decoder architecture and, therefore, allows the second stage of dehazing for better refining the results of the previous stage. In addition, we combine the channel attention mechanism with the basic convolution block, considering that different channel characteristics contain entirely different weighting information, to effectively deal with irregular distribution of haze in RS images. Owing to the scarcity of various and quality hazy RS data sets, we adopt two different synthesis methods to generate large-scale image pairs for uniform and nonuniform hazy images. This two-stage network, when trained in an end-to-end fashion, yields the state-of-the-art performances on both the synthetic data sets and real-world images with more visually pleasing dehazed results. Both the synthetic data set and the code are publicly available athttps://github.com/cxtalk/FCTF-Net.
Yufeng Li 0001, Xiang Chen 0015
IEEE Geosci. Remote. Sens. Lett.2
2020 Multi-scale Attentive Residual Dense Network for Single Image Rain Removal
Xiang Chen 0015
ACCV (2)1