EDBT 2026 Demo / reviewers in the wild / expert
Han Xu 0001
dblp:32/34-1
· DBLP profile ↗
26ranked-venue papers
13as first author
21since 2021 · last 2026
0000-0002-6291-2924ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking Low-Light Fusion Barrier: Unsupervised Darkness and Noise-Aware Visible and Infrared Image Fusion NetworkabstractInfrared and visible image fusion aims to integrate complementary information but suffers from severe residual degradations under low-light conditions. Existing methods face two main limitations: darkness-aware fusion focuses on illumination enhancement while neglecting noise suppression, and degradation-aware supervised fusion relies on paired data and struggles with enhancement-denoising imbalance due to multi-task optimization conflicts. To address these issues, we propose BLFusion, an unsupervised darkness- and noise-aware fusion framework that performs illumination enhancement and noise suppression in two dedicated stages while fusing complementary information without high-quality references. First, a Retinex-guided state space model-based decomposition network models illumination degradation to brighten dark visible images. Then, an unsupervised denoising fusion network jointly performs fusion and denoising, where noise correlation is disrupted by shuffling and a blind-spot network with dilated convolutions estimates clean representations from surrounding pixels. Finally, noise-free features from both modalities are fused to generate the final image. Moreover, we construct the MRLL dataset with 500 well-aligned infrared-visible image pairs, filling the gap for real-world noise-degraded nighttime scenarios. Experiments demonstrate that BLFusion outperforms state-of-the-art methods and generalizes robustly across diverse low-light and noisy conditions. The MRLL dataset and code are publicly available at https://github.com/ChenDoubleJ/BLFusion-MRLL. Han Xu 0001, Guangcan Liu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Diff-MEF: Cross-Modal Diffusion Framework With Text Prompts and Semantic Perception for Multi-Exposure Image FusionabstractThe absence of real-world ground truth (GT) remains a challenge in multi-exposure image fusion (MEF). Benchmarks synthesizing pseudo GT through algorithm ensembles. Existing methods, hampered by inherent imperfections of pseudo GT and fixed mapping relationships, show limited performance and robustness. To address the limitations, we propose a novel cross-modal diffusion framework that synergizes text prompts and semantic perception for MEF, termed as Diff-MEF. First, it reformulates MEF as a probabilistic estimation task with conditional diffusion model for progressive transition and fusion. Then, we explicitly infer semantic and exposure priors as text prompts and semantic perception to improve performance and robustness. The priors are synergized through multi-modal prior embedding and optimization guidance. On the one hand, regarding cross-modal interaction, multi-modal priors, including segmentation masks, and exposure- and content-aware text prompts, are embedded into diffusion process by dedicated encoders and refine visual features through a text-segmentation refinement module. On the other hand, a semantic-level contrastive loss builds a regularization between cross-modal features in the semantic space of CLIP to mitigate degradations introduced by pseudo GT and fusion distortions. Experiments demonstrate that Diff-MEF outperforms SOTA methods and pseudo GT with superior fusion performance and robustness across diverse exposure scenarios. Code is available at https://github.com/hanna-xu/Diff-MEF. Han Xu 0001, Yunfei Huang, Linfeng Tang, Jiayi Ma 0001, Guangcan Liu |
IEEE Trans. Image Process. | 1 |
| 2025 | Cross-Modal Stealth: A Coarse-to-Fine Attack Framework for RGB-T TrackerabstractCurrent research on adversarial attacks mainly focuses on RGB trackers, with no existing methods for attacking RGB-T cross-modal trackers. To fill this gap and overcome its challenges, we propose a progressive adversarial patch generation framework and achieve cross-modal stealth. On the one hand, we design a coarse-to-fine architecture grounded in the latent space to progressively and precisely uncover the vulnerabilities of RGB-T trackers. On the other hand, we introduce a correlation-breaking loss that disrupts the modal coupling within trackers, spanning from the pixel to the semantic level. These two design elements ensure that the proposed method can overcome the obstacles posed by cross-modal information complementarity in implementing attacks. Furthermore, to enhance the reliable application of the adversarial patches in real world, we develop a point tracking-based reprojection strategy that effectively mitigates performance degradation caused by multi-angle distortion during imaging. Extensive experiments demonstrate the superiority of our method. Xinyu Xiang, Qinglong Yan, Hao Zhang 0073, Jianfeng Ding, Han Xu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
AAAI | 5 |
| 2025 | LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesabstractCurrent advanced research on infrared and visible image fusion primarily focuses on improving fusion performance, often neglecting the applicability on real-time fusion devices. In this paper, we propose a novel approach that towards extremely fast fusion via distillation to learnable lookup tables specifically designed for image fusion, termed as LUT-Fuse. Firstly, we develop a look-up table structure that utilizing low-order approximation encoding and high-level joint contextual scene encoding, which is well-suited for multi-modal fusion. Moreover, given the lack of ground truth in multi-modal image fusion, we naturally proposed the efficient LUT distillation strategy instead of traditional quantization LUT methods. By integrating the performance of the multi-modal fusion network (MM-Net) into the MM-LUT model, our method achieves significant breakthroughs in efficiency and performance. It typically requires less than one-tenth of the time compared to the current lightweight SOTA fusion algorithms, ensuring high operational speed across various scenarios, even in low-power mobile devices. Extensive experiments validate the superiority, reliability, and stability of our fusion approach. The code is available at https://github.com/zyb5/LUT-Fuse. Xunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan, Han Xu 0001, Jiayi Ma 0001 |
ICCV | 5 |
| 2025 | Towards Perfection: Building Inter-component Mutual Correction for Retinex-based Low-light Image EnhancementabstractIn low-light image enhancement, Retinex-based deep learning methods have garnered significant attention due to their exceptional interpretability. These methods decompose images into mutually independent illumination and reflectance components, allows each component to be enhanced separately. In fact, achieving perfect decomposition of illumination and reflectance components proves to be quite challenging, with some residuals still existing after decomposition. In this paper, we formally name these residuals as inter-component residuals (ICR), which has been largely underestimated by previous methods. In our investigation, ICR not only affects the accuracy of the decomposition but also causes enhanced components to deviate from the ideal outcome, ultimately reducing the final synthesized image quality. To address this issue, we propose a novel Inter-correction Retinex model (IRetinex) to alleviate ICR during the decomposition and enhancement stage. In the decomposition stage, we leverage inter-component residual reduction module to reduce the feature similarity between illumination and reflectance components. In the enhancement stage, we utilize the feature similarity between the two components to detect and mitigate the impact of ICR within each enhancement unit. Extensive experiments on three low-light benchmark datasets demonstrated that by reducing ICR, our method outperforms state-of-the-art approaches both qualitatively and quantitatively. Our code is available at: https://github.com/caoluyang0830/IRetinex.git. Luyang Cao, Han Xu 0001, Jian Zhang 0090, Lei Qi 0001, Jiayi Ma 0001, Yinghuan Shi, Yang Gao 0001 |
ACM Multimedia | 2 |
| 2025 | Deno-IF: Unsupervised Noisy Visible and Infrared Image Fusion MethodabstractMost image fusion methods are designed for ideal scenarios and struggle to handle noise. Existing noise-aware fusion methods are supervised and heavily rely on constructed paired data, limiting performance and generalization. This paper proposes a novel unsupervised noisy visible and infrared image fusion method, comprising two key modules. First, when only noisy source images are available, a convolutional low-rank optimization module decomposes clean components based on convolutional low-rank priors, guiding subsequent optimization. The unsupervised approach eliminates data dependency and enhances generalization across various and variable noise. Second, a unified network jointly realizes denoising and fusion. It consists of both intra-modal recovery and inter-modal recovery and fusion, also with a convolutional low-rankness loss for regularization. By exploiting the commonalities of denoising and fusion, the joint framework significantly reduces network complexity while expanding functionality. Extensive experiments validate the effectiveness and generalization of the proposed method for image fusion under various and variable noise conditions. The code is publicly available at https://github.com/hanna-xu/Deno-IF. Han Xu 0001, Yuyang Li 0005, Yunfei Deng, Jiayi Ma 0001, Guangcan Liu |
NeurIPS | 1 |
| 2025 | Diff-Retinex++: Retinex-Driven Reinforced Diffusion Model for Low-Light Image EnhancementabstractThis paper proposes a Retinex-driven reinforced diffusion model for low-light image enhancement, termed Diff-Retinex++, to address various degradations caused by low light. Our main approach integrates the diffusion model with Retinex-driven restoration to achieve physically-inspired generative enhancement, making it a pioneering effort. To be detailed, Diff-Retinex++ consists of two-stage view modules, including the Denoising Diffusion Model (DDM), and the Retinex-Driven Mixture of Experts Model (RMoE). First, DDM treats low-light image enhancement as one type of image generation task, benefiting from the powerful generation ability of diffusion model to handle the enhancement. Second, we design the Retinex theory into the plug-and-play supervision attention module. It leverages the latent features in the backbone and knowledge distillation to learn Retinex rules, and further regulates these latent features through the attention mechanism. In this way, it couples the relationship between Retinex decomposition and image enhancement in a new view, achieving dual improvement. In addition, the Low-Light Mixture of Experts preserves the vividness of the diffusion model and fidelity of the Retinex-driven restoration to the greatest extent. Ultimately, the iteration of DDM and RMoE achieves the goal of Retinex-driven reinforced diffusion model. Extensive experiments conducted on real-world low-light datasets qualitatively and quantitatively demonstrate the effectiveness, superiority, and generalization of the proposed method. Xunpeng Yi, Han Xu 0001, Hao Zhang 0073, Linfeng Tang, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | URFusion: Unsupervised Unified Degradation-Robust Image Fusion NetworkabstractWhen dealing with low-quality source images, existing image fusion methods either fail to handle degradations or are restricted to specific degradations. This study proposes an unsupervised unified degradation-robust image fusion network, termed as URFusion, in which various types of degradations can be uniformly eliminated during the fusion process, leading to high-quality fused images. URFusion is composed of three core modules: intrinsic content extraction, intrinsic content fusion, and appearance representation learning and assignment. It first extracts degradation-free intrinsic content features from images affected by various degradations. These content features then provide feature-level rather than image-level fusion constraints for optimizing the fusion network, effectively eliminating degradation residues and reliance on ground truth. Finally, URFusion learns the appearance representation of images and assigns the statistical appearance representation of high-quality images to the content-fused result, producing the final high-quality fused image. Extensive experiments on multi-exposure image fusion and multi-modal image fusion tasks demonstrate the advantages of URFusion in fusion performance and suppression of multiple types of degradations. The code is available at https://github.com/hanna-xu/URFusion. Han Xu 0001, Xunpeng Yi, Chen Lu 0004, Guangcan Liu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionabstractImage fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF. Xunpeng Yi, Han Xu 0001, Hao Zhang 0073, Linfeng Tang, Jiayi Ma 0001 |
CVPR | 2 |
| 2024 | CRetinex: A Progressive Color-Shift Aware Retinex Model for Low-Light Image Enhancement
Han Xu 0001, Hao Zhang 0073, Xunpeng Yi, Jiayi Ma 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | CTCFNet: CNN-Transformer Complementary and Fusion Network for High-Resolution Remote Sensing Image Semantic SegmentationabstractSemantic segmentation of high-resolution remote sensing images poses challenges such as scale variability, diverse objects, and obstruction by surface elements. These factors often lead existing methods to suffer from issues like missed and false detections, as well as coarse segmentation boundaries. To tackle these challenges, this article proposes a CNN-transformer complementary and fusion network, termed as CTCFNet. It aims to enhance segmentation accuracy and robustness by extracting and integrating the complementary global and local information from high-resolution remote sensing images. The CTCFNet operates through two primary stages: feature extraction and fusion. In the feature extraction stage, a feature extractor employs convolutional neural network (CNN) and pyramid vision transformer (PVT) blocks to extract both local and global features. A boundary loss is also proposed to improve the segmentation performance for object textures and boundaries. In the feature fusion stage, a feature aggregation module (FAM) is first designed to effectively fuse local and global features at the same scale, facilitating the feature extractor to obtain more comprehensive representations. On this basis, a bi-directional decoder (BiDecoder) reconstructs multiscale features through both top-down and bottom-up directions, resulting in more precise segmentation outputs. Experiments on several high-resolution remote sensing image datasets demonstrate that the proposed method outperforms the state-of-the-art methods in terms of segmentation accuracy and generalization. The code is available athttps://github.com/ChenLu0000/CTCFNet. Chen Lu 0004, Kaile Du, Han Xu 0001, Guangcan Liu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Unsupervised Multi-Exposure Image Fusion Breaking Exposure Limits via Contrastive LearningabstractThis paper proposes an unsupervised multi-exposure image fusion (MEF) method via contrastive learning, termed as MEF-CL. It breaks exposure limits and performance bottleneck faced by existing methods. MEF-CL firstly designs similarity constraints to preserve contents in source images. It eliminates the need for ground truth (actually not exist and created artificially) and thus avoids negative impacts of inappropriate ground truth on performance and generalization. Moreover, we explore a latent feature space and apply contrastive learning in this space to guide fused image to approximate normal-light samples and stay away from inappropriately exposed ones. In this way, characteristics of fused images (e.g., illumination, colors) can be further improved without being subject to source images. Therefore, MEF-CL is applicable to image pairs of any multiple exposures rather than a pair of under-exposed and over-exposed images mandated by existing methods. By alleviating dependence on source images, MEF-CL shows better generalization for various scenes. Consequently, our results exhibit appropriate illumination, detailed textures, and saturated colors. Qualitative, quantitative, and ablation experiments validate the superiority and generalization of MEF-CL. Our code is publicly available at https://github.com/hanna-xu/MEF-CL. Han Xu 0001, Liang Haochen, Jiayi Ma 0001 |
AAAI | 1 |
| 2023 | Diff-Retinex: Rethinking Low-light Image Enhancement with A Generative Diffusion ModelabstractIn this paper, we rethink the low-light image enhancement task and propose a physically explainable and generative diffusion model for low-light image enhancement, termed as Diff-Retinex. We aim to integrate the advantages of the physical model and the generative network. Furthermore, we hope to supplement and even deduce the information missing in the low-light image through the generative network. Therefore, Diff-Retinex formulates the lowlight image enhancement problem into Retinex decomposition and conditional image generation. In the Retinex decomposition, we integrate the superiority of attention in Transformer and meticulously design a Retinex Transformer decomposition network (TDN) to decompose the image into illumination and reflectance maps. Then, we design multi-path generative diffusion networks to reconstruct the normal-light Retinex probability distribution and solve the various degradations in these components respectively, including dark illumination, noise, color deviation, loss of scene contents, etc. Owing to generative diffusion model, Diff-Retinex puts the restoration of low-light subtle detail into practice. Extensive experiments conducted on real-world low-light datasets qualitatively and quantitatively demonstrate the effectiveness, superiority, and generalization of the proposed method. Xunpeng Yi, Han Xu 0001, Hao Zhang 0073, Linfeng Tang, Jiayi Ma 0001 |
ICCV | 2 |
| 2023 | MURF: Mutually Reinforcing Multi-Modal Image Registration and FusionabstractExisting image fusion methods are typically limited to aligned source images and have to "tolerate" parallaxes when images are unaligned. Simultaneously, the large variances between different modalities pose a significant challenge for multi-modal image registration. This study proposes a novel method called MURF, where for the first time, image registration and fusion are mutually reinforced rather than being treated as separate issues. MURF leverages three modules: shared information extraction module (SIEM), multi-scale coarse registration module (MCRM), and fine registration and fusion module (F2M). The registration is carried out in a coarse-to-fine manner. During coarse registration, SIEM first transforms multi-modal images into mono-modal shared information to eliminate the modal variances. Then, MCRM progressively corrects the global rigid parallaxes. Subsequently, fine registration to repair local non-rigid offsets and image fusion are uniformly implemented in F2M. The fused image provides feedback to improve registration accuracy, and the improved registration result further improves the fusion result. For image fusion, rather than solely preserving the original source information in existing methods, we attempt to incorporate texture enhancement into image fusion. We test on four types of multi-modal data (RGB-IR, RGB-NIR, PET-MRI, and CT-MRI). Extensive registration and fusion results validate the superiority and universality of MURF. Han Xu 0001, Jiteng Yuan, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Multipatch Progressive Pansharpening With Knowledge DistillationabstractIn this paper, we propose a novel multi-patch and multi-stage pansharpening method with knowledge distillation, termed as PSDNet. Different from existing pansharpening methods that typically input single-size patches to the network and implement pansharpening in an overall stage, we design multi-patch inputs and a multi-stage network for more accurate and finer learning. First, multi-patch inputs allow the network to learn more accurate spatial and spectral information by reducing the number of object types. We employ small patches in the early part to learn accurate local information, as small patches contain fewer object types. Then, the later part exploits large patches to fine-tune it for the overall information. Second, the multi-stage network is designed to reduce the difficulty of the previous single-step pansharpening and progressively generate elaborate results. In addition, instead of the traditional perceptual loss, which hardly relates to the specific task or the designed network, we introduce distillation loss to reinforce the guidance of the ground truth. Extensive experiments are conducted to demonstrate the superior performance of our proposed PSDNet to existing state-of-the-art methods. Our code is available at https://github.com/Meiqi-Gong/PSDNet. Meiqi Gong, Hao Zhang 0073, Han Xu 0001, Xin Tian 0006, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Dual Spatial-Spectral Pyramid Network With Transformer for Hyperspectral Image FusionabstractMultispectral image (MSI) and hyperspectral image (HSI) fusion can combine the best of both worlds to produce images with both high spatial and spectral resolution. In this paper, we have designed a network for fusing MSIs and HSIs, called DSPNet. On the one hand, in order to ensure the accuracy of the spectral dimension, i.e. spectral fidelity, we designed the spectral pyramid (SpePy) module and the multiscale spectral information fusion (MLSIF) module. The former extracts the multiscale local spectral information that captures the subtle spectral details and variations between different spectra. The latter establishes long-range dependency in the spectral dimension through the spectral-wise multi-head hybrid-attention (S-MHA) mechanism, thus enabling the network to focus on the local spectral information needed to recover the spectral details. On the other hand, to address the spatial information of MSIs, we designed the spatial pyramid (SpaPy) module. The SpaPy module can extract the non-local spatial information of MSIs at different scales, which enables the network to adapt to different remote-sensing scenes. Experiments performed on simulated and real data demonstrate the superiority of our method over the state-of-the-art methods both qualitatively and quantitatively. Han Xu 0001, Yong Ma 0001, Minghui Wu 0007, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | RFNet: Unsupervised Network for Mutually Reinforcing Multi-modal Image Registration and FusionabstractIn this paper, we propose a novel method to realize multimodal image registration and fusion in a mutually reinforcing framework, termed as RFNet. We handle the registration in a coarse-to-fine fashion. For the first time, we exploit the feedback of image fusion to promote the registration accuracy rather than treating them as two separate issues. The fine-registered results also improve the fusion performance. Specifically, for image registration, we solve the bottlenecks of defining registration metrics applicable for multi-modal images and facilitating the network convergence. The metrics are defined based on image translation and image fusion respectively in the coarse and fine stages. The convergence is facilitated by the designed metrics and a deformable convolution-based network. For image fusion, we focus on texture preservation, which not only increases the information amount and quality of fusion results but also improves the feedback of fusion results. The proposed method is evaluated on multi-modal images with large global parallaxes, images with local misalignments and aligned images to validate the performances of registration and fusion. The results in these cases demonstrate the effectiveness of our method. Han Xu 0001, Jiayi Ma 0001, Jiteng Yuan, Zhuliang Le, Wei Liu 0005 |
CVPR | 1 |
| 2022 | CUFD: An encoder-decoder network for visible and infrared image fusion based on common and unique feature decomposition
Han Xu 0001, Meiqi Gong, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 1 |
| 2022 | U2Fusion: A Unified Unsupervised Image Fusion NetworkabstractThis study proposes a novel unified and unsupervised end-to-end image fusion network, termed as U2Fusion, which is capable of solving different fusion problems, including multi-modal, multi-exposure, and multi-focus cases. Using feature extraction and information measurement, U2Fusion automatically estimates the importance of corresponding source images and comes up with adaptive information preservation degrees. Hence, different fusion tasks are unified in the same framework. Based on the adaptive degrees, a network is trained to preserve the adaptive similarity between the fusion result and source images. Therefore, the stumbling blocks in applying deep learning for image fusion, e.g., the requirement of ground-truth and specifically designed metrics, are greatly mitigated. By avoiding the loss of previous fusion capabilities when training a single model for different tasks sequentially, we obtain a unified model that is applicable to multiple fusion tasks. Moreover, a new aligned infrared and visible image dataset, RoadScene (available at https://github.com/hanna-xu/RoadScene), is released to provide a new option for benchmark evaluation. Qualitative and quantitative experimental results on three typical image fusion tasks validate the effectiveness and universality of U2Fusion. Our code is publicly available at https://github.com/hanna-xu/U2Fusion. Han Xu 0001, Jiayi Ma 0001, Junjun Jiang, Xiaojie Guo 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | D2TNet: A ConvLSTM Network With Dual-Direction Transfer for Pan-SharpeningabstractIn this article, we propose an efficient convolutional long short-term memory (ConvLSTM) network with dual-direction transfer for pan-sharpening, termed D2TNet. We design a specially structured ConvLSTM network that allows for dual-directional communication, including multiscale information and multilevel information. On the one hand, due to the sensitivity of spatial information to scales and the sensitivity of spectral information to levels, multiscale and multilevel information is extracted to facilitate the fuller use of source images. On the other hand, ConvLSTM is employed to capture the strong dependencies between multiscale information and multilevel information. Besides, we introduce a multiscale loss to enable different scales contributing to each other to generate high-resolution multispectral images that are closer to the ground truth. Extensive experiments, including qualitative evaluation, quantitative evaluation, and efficiency comparison, are implemented to verify that our D2TNet outperforms state-of-the-art methods indeed. Meiqi Gong, Jiayi Ma 0001, Han Xu 0001, Xin Tian 0006, Xiao-Ping Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SDPNet: A Deep Network for Pan-Sharpening With Enhanced Information RepresentationabstractIn this article, we propose a surface- and deep-level constraint-based pan-sharpening network, termed SDPNet, to address the pan-sharpening problem. Focusing on the two primary goals of pan-sharpening, i.e., spatial and spectral information preservations, we first design two encoder-decoder networks to extract deep-level features from two types of source images, in addition to surface-level characteristics, as the enhanced information representation. The unique feature maps that characterize the unique information in source images can be obtained through the deep-level feature extraction. We further design a pan-sharpening network with densely connected blocks to strengthen feature propagation and reduce parameter number, where the unique feature maps are utilized to efficiently constrain the similarity between the pan-sharpened result and the ground truth, thus avoiding information distortion. Both qualitative and quantitative comparisons on the reduced-resolution and full-resolution source images demonstrate the advantages of our method over state-of-the-art methods. Our code is publicly available at https://github.com/hanna-xu/SDPNet. Han Xu 0001, Jiayi Ma 0001, Hao Zhang 0073, Junjun Jiang, Xiaojie Guo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | FusionDN: A Unified Densely Connected Network for Image FusionabstractIn this paper, we present a new unsupervised and unified densely connected network for different types of image fusion tasks, termed as FusionDN. In our method, the densely connected network is trained to generate the fused image conditioned on source images. Meanwhile, a weight block is applied to obtain two data-driven weights as the retention degrees of features in different source images, which are the measurement of the quality and the amount of information in them. Losses of similarities based on these weights are applied for unsupervised learning. In addition, we obtain a single model applicable to multiple fusion tasks by applying elastic weight consolidation to avoid forgetting what has been learned from previous tasks when training multiple tasks sequentially, rather than train individual models for every fusion task or jointly train tasks roughly. Qualitative and quantitative results demonstrate the advantages of FusionDN compared with state-of-the-art methods in different fusion tasks. Han Xu 0001, Jiayi Ma 0001, Zhuliang Le, Junjun Jiang, Xiaojie Guo 0001 |
AAAI | 1 |
| 2020 | Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and IntensityabstractIn this paper, we propose a fast unified image fusion network based on proportional maintenance of gradient and intensity (PMGI), which can end-to-end realize a variety of image fusion tasks, including infrared and visible image fusion, multi-exposure image fusion, medical image fusion, multi-focus image fusion and pan-sharpening. We unify the image fusion problem into the texture and intensity proportional maintenance problem of the source images. On the one hand, the network is divided into gradient path and intensity path for information extraction. We perform feature reuse in the same path to avoid loss of information due to convolution. At the same time, we introduce the pathwise transfer block to exchange information between different paths, which can not only pre-fuse the gradient information and intensity information, but also enhance the information to be processed later. On the other hand, we define a uniform form of loss function based on these two kinds of information, which can adapt to different fusion tasks. Experiments on publicly available datasets demonstrate the superiority of our PMGI over the state-of-the-art in terms of both visual effect and quantitative metric in a variety of fusion tasks. In addition, our method is faster compared with the state-of-the-art. Hao Zhang 0073, Han Xu 0001, Xiaojie Guo 0001, Jiayi Ma 0001 |
AAAI | 2 |
| 2020 | DDcGAN: A Dual-Discriminator Conditional Generative Adversarial Network for Multi-Resolution Image FusionabstractIn this paper, we proposed a new end-to-end model, termed as dual-discriminator conditional generative adversarial network (DDcGAN), for fusing infrared and visible images of different resolutions. Our method establishes an adversarial game between a generator and two discriminators. The generator aims to generate a real-like fused image based on a specifically designed content loss to fool the two discriminators, while the two discriminators aim to distinguish the structure differences between the fused image and two source images, respectively, in addition to the content loss. Consequently, the fused image is forced to simultaneously keep the thermal radiation in the infrared image and the texture details in the visible image. Moreover, to fuse source images of different resolutions, e.g., a low-resolution infrared image and a high-resolution visible image, our DDcGAN constrains the downsampled fused image to have similar property with the infrared image. This can avoid causing thermal radiation information blurring or visible texture detail loss, which typically happens in traditional methods. In addition, we also apply our DDcGAN to fusing multi-modality medical images of different resolutions, e.g., a low-resolution positron emission tomography image and a high-resolution magnetic resonance image. The qualitative and quantitative experiments on publicly available datasets demonstrate the superiority of our DDcGAN over the state-of-the-art, in terms of both visual effect and quantitative metrics. Jiayi Ma 0001, Han Xu 0001, Junjun Jiang, Xiaoguang Mei, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 2 |
| 2020 | MEF-GAN: Multi-Exposure Image Fusion via Generative Adversarial NetworksabstractIn this paper, we present an end-to-end architecture for multi-exposure image fusion based on generative adversarial networks, termed as MEF-GAN. In our architecture, a generator network and a discriminator network are trained simultaneously to form an adversarial relationship. The generator is trained to generate a real-like fused image based on the given source images which is expected to fool the discriminator. Correspondingly, the discriminator is trained to distinguish the generated fused images from the ground truth. The adversarial relationship makes the fused image not limited to the restriction of the content loss. Therefore, the fused images are closer to the ground truth in terms of probability distribution, which can compensate for the insufficiency of single content loss. Moreover, aiming at the problem that the luminance of multi-exposure images varies greatly with spatial location, the self-attention mechanism is employed in our architecture to allow for attention-driven and long-range dependency. Thus, local distortion, confusing results, or inappropriate representation can be corrected in the fused image. Qualitative and quantitative experiments are performed on publicly available datasets, where the results demonstrate that MEF-GAN outperforms the state-of-the-art, in terms of both visual effect and objective evaluation metrics. Our code is publicly available at https://github.com/jiayi-ma/MEF-GAN. Han Xu 0001, Jiayi Ma 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 1 |
| 2019 | Learning a Generative Model for Fusing Infrared and Visible Images via Conditional Generative Adversarial Network with Dual DiscriminatorsabstractIn this paper, we propose a new end-to-end model, called dual-discriminator conditional generative adversarial network (DDcGAN), for fusing infrared and visible images of different resolutions. Unlike the pixel-level methods and existing deep learning-based methods, the fusion task is accomplished through the adversarial process between a generator and two discriminators, in addition to the specially designed content loss. The generator is trained to generate real-like fused images to fool discriminators. The two discriminators are trained to calculate the JS divergence between the probability distribution of downsampled fused images and infrared images, and the JS divergence between the probability distribution of gradients of fused images and gradients of visible images, respectively. Thus, the fused images can compensate for the features that are not constrained by the single content loss. Consequently, the prominence of thermal targets in the infrared image and the texture details in the visible image can be preserved or even enhanced in the fused image simultaneously. Moreover, by constraining and distinguishing between the downsampled fused image and the low-resolution infrared image, DDcGAN can be preferably applied to the fusion of different resolution images. Qualitative and quantitative experiments on publicly available datasets demonstrate the superiority of our method over the state-of-the-art. Han Xu 0001, Pengwei Liang, Wei Yu 0018, Junjun Jiang, Jiayi Ma 0001 |
IJCAI | 1 |