VLDB 2026 Research / reviewers in the wild / expert
Kangjian He
dblp:177/5328
· DBLP profile ↗
44ranked-venue papers
5as first author
39since 2021 · last 2026
0000-0001-6207-9728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 26 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 11 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSKFuse: Passive-active distillation learning for multi-modal image fusion via dynamic sparse kansformerabstractAn effective knowledge learning strategy combined with a lightweight network architecture is crucial for the practical deployment of multi-modal image fusion. While existing methods have made significant progress in the visual perception of fused results, their model complexity and generalization capabilities still require further optimization. In this paper, we propose a novel passive-active distillation learning framework for multi-modal image fusion, termed DSKFuse, which integrates the Dynamic Sparse Transformer and the latent Kolmogorov-Arnold Network (KAN). Specifically, we design an efficient fusion architecture trained via a two-stage knowledge distillation strategy, seamlessly integrating passive and active learning methodologies. In the first stage, passive distillation learning enhances the fusion network by extracting valuable knowledge from complex fusion models. In the second stage, an active knowledge distillation approach is implemented, enabling the model to autonomously capture discriminative features from source images, thereby improving the robustness and generalization of DSKFuse. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance in both image fusion and downstream tasks, including detection and segmentation. The code will be released at https://github.com/DZSYUNNAN/DSKFuse . Zhaisheng Ding, Ruichao Hou, Yunzhe Men, Shengyang Luan, Yanyu Liu, Kangjian He, Shidong Xie |
Expert Syst. Appl. | 6 |
| 2026 | Cross-model and attribute-driven dual-stage knowledge distillation for multimodal medical image fusion
Yanyu Liu, Chunxue Liu, Ruichao Hou, Zhaisheng Ding, Kangjian He, Dongming Zhou 0001 |
Multim. Syst. | 5 |
| 2026 | Enhancing infrared-visible image fusion via text-guided adaptive feature integration
Jundong Zhang, Kangjian He, Dan Xu 0001, Songhan Zheng, Wencheng Mei |
Multim. Syst. | 3 |
| 2026 | CUDiff: Consistency and uncertainty guided conditional diffusion for infrared and visible image fusion
Yueying Luo, Kangjian He, Dan Xu 0001 |
Pattern Recognit. | 2 |
| 2026 | Prior knowledge driven dynamic fusion network for infrared and visible images
Yueying Luo, Kangjian He, Dan Xu 0001, Yiqiao Zhou |
Pattern Recognit. | 2 |
| 2026 | Test-Time Domain-Agnostic Meta-Prompt Learning for Multi-Source Few-Shot Domain Adaptation
Kuanghong Liu, Jin Wang 0008, Kangjian He, Dan Xu 0001, Xuejie Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain AdaptationabstractConventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper introduces an uploadable multi-source few-shot domain adaptation (UMFDA) schema. It belongs to a decentralized edge collaborative learning in the edge-side models that must maintain a low computational load. And only a limited amount of annotations in source domain data is provided, with most of the data being unannotated. Further, this paper proposes a vision-aware multimodal prompt tuning framework (VAMP) under the decentralized schema, where the vision-aware prompt guides the text domain-specific prompt to maintain semantic discriminability and perceive the domain information. The cross-modal semantic and domain distribution alignment losses optimize each edge-side model, while text classifier consistency and semantic diversity losses promote collaborative learning among edge-side models. Extensive experiments were conducted on OfficeHome and DomainNet datasets to demonstrate the effectiveness of the proposed VAMP in the UMFDA, which outperformed the previous prompt tuning methods. Kuanghong Liu, Jin Wang 0008, Kangjian He, Dan Xu 0001, Xuejie Zhang 0002 |
AAAI | 3 |
| 2025 | PCM-Net: A Hierarchical Medical Image Registration Framework Integrating Channel Adaptability and Multi-scale Awareness
Zihang Sun, Dan Xu 0001, Kangjian He, Zilong Xue, Yijie He |
CGI (1) | 3 |
| 2025 | Construct a Powerful Discriminative Relationship for Few-Shot Action RecognitionabstractLearning discriminative features from very few labeled samples has gradually become a hot issue in the task of human skeleton action recognition. Most of the existing works follow the paradigms of meta-learning or contrastive learning. However, we argue that the discriminative relationship established in this way is rather simple, because the temporal and spatial features are complex and it is difficult to find the relationships of the same category. In this paper, we propose a Multi-Level Semantic Prompting Joint Contrastive Learning Head (MSJCL-Head), which consists of Joint Hard-Soft Contrastive Learning module (JCL) and Action Semantic Prompts (ASP), to obtain the discriminative representations of texts and skeletons with effect strength and discover and calibrate ambiguous samples in the feature space. A large number of experiments have been conducted on the NTU-T, NTU-S and Kinetics datasets, and the results show that our model has achieved competitive results in few-shot tasks. Qianhan Tang, Ningxin Wang, Kangjian He, Hao Zhang 0110, Dan Xu 0001 |
ICME | 4 |
| 2025 | FCReg: Medical Image Registration Network with Image-Text Feature Coupling
Yijie He, Dan Xu 0001, Yueying Luo, Zihang Sun, Kangjian He |
PRCV (14) | 5 |
| 2025 | DPA-SAM: Enhancing Medical Image Segmentation with 3D-DCAF and PGAttention
Liye Li, Kangjian He, Gaifang Luo, Hao Zhang 0110, Yijie He, Dan Xu 0001 |
PRCV (14) | 2 |
| 2025 | Bridging the Degradation Gap in Real Super-Resolution: A Transfer-Based Paired Dataset Construction
Yinghui Zhu, Congcong Zeng, Dan Xu 0001, Jiangang Pan, Kangjian He, Hongzhen Shi |
PRCV (9) | 5 |
| 2025 | TSSA-Net: Transposed Sparse Self-Attention-based network for image super-resolution
Guanhao Chen, Dan Xu 0001, Kangjian He, Hongzhen Shi, Hao Zhang 0110 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Medical Image Registration via Spatial Feature Extraction Mamba and Substrate Iterative RefinementabstractABSTRACT One of the major challenges in medical image registration is balancing computational efficiency with the ability to capture large deformations in complex anatomical structures. Existing methods often struggle with high computational costs due to the need for extensive feature extraction and attention computations at various levels of the network. Moreover, some methods do not take into account the spatial relationships of the feature images during registration, and the loss of these spatial relationships leads to suboptimal results for these methods. To this end, we introduce a novel medical image registration network, PSMamba‐Net, which leverages optimized iteration and the Mamba framework within a dual‐stream pyramid architecture. The network reduces the computational burden by narrowing attention computations at each decoding level, while an optimized iterative registration module at the bottom of the pyramid captures large deformations. This approach eliminates the need for repeated feature extraction, significantly accelerating the registration process. Additionally, the SMB module is incorporated as a decoder to enhance spatial relationship modelling and leverage Mamba's strengths in long‐sequence processing. PSMamba‐Net balances efficiency and accuracy, surpassing state‐of‐the‐art methods across LPBA40, Mindboggle, and Abdomen CT datasets. Our source code is available at: https://github.com/VCMHE/PSMamba . Zilong Xue, Kangjian He, Dan Xu 0001 |
IET Image Process. | 2 |
| 2025 | Infrared and visible image fusion based on hybrid multi-scale decomposition and adaptive contrast enhancement
Yueying Luo, Kangjian He, Dan Xu 0001, Hongzhen Shi, Wenxia Yin |
Signal Process. Image Commun. | 2 |
| 2025 | MDH-Net: advancing 3D brain MRI registration with multi-stage transformer and dual-stream feature refinement hybrid network
Chenou Liu, Kangjian He, Dan Xu 0001, Hongzhen Shi |
J. Supercomput. | 2 |
| 2025 | Multi-modality medical image fusion by edge supervising and multi-scale attention features extraction
Wencheng Mei, Kangjian He, Dan Xu 0001, Siqi Xie, Yiqiao Zhou |
J. Supercomput. | 2 |
| 2025 | $\hbox {KD}^{3}$mt: knowledge distillation-driven dynamic mixer transformer for medical image fusion
Zhaijuan Ding, Yanyu Liu, Kangjian He, Dongming Zhou 0001 |
Vis. Comput. | 4 |
| 2024 | Multimodal Medical Image Registration Using Optimized Phase Consistency Within Joint Frequency-Space Domain
Dan Xu 0001, Kangjian He |
PRCV (5) | 3 |
| 2024 | Hyperspectral Image Super-Resolution Based on Dual-Domain Gated Attention Network
Songhan Zheng, Dan Xu 0001, Kangjian He |
PRCV (13) | 3 |
| 2024 | ASFusion: Adaptive visual enhancement and structural patch decomposition for infrared and visible image fusion
Yiqiao Zhou, Kangjian He, Dan Xu 0001, Dapeng Tao, Chengzhou Li |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Fidelity based visual compensation and salient information rectification for infrared and visible image fusion
Yueying Luo, Dan Xu 0001, Kangjian He, Hongzhen Shi |
Knowl. Based Syst. | 3 |
| 2024 | Affective image recognition with multi-attribute knowledge in deep neural networks
Hao Zhang 0110, Gaifang Luo, Yingying Yue, Kangjian He, Dan Xu 0001 |
Multim. Tools Appl. | 4 |
| 2024 | A multi-weight fusion framework for infrared and visible image fusion
Yiqiao Zhou, Kangjian He, Dan Xu 0001, Hongzhen Shi, Hao Zhang 0110 |
Multim. Tools Appl. | 2 |
| 2024 | RegFSC-Net: Medical Image Registration via Fourier Transform With Spatial Reorganization and Channel Refinement NetworkabstractMedical image registration is crucial in medical image analysis applications. Recently, U-Net-style networks have been commonly used for unsupervised image registration, predicting dense displacement fields in full-resolution space. However, this process is resource-intensive and time-consuming for high-resolution volumetric image data. To address this challenge, this paper proposes a novel model named RegFSC-Net, which utilizes Fourier transform with spatial reorganization (SR) and channel refinement (CR) network for registration. We embed efficient feature extraction modules SR and CR modules into the encoder, and adopt a parameter-free model to drive the decoder to improve the U-shaped network. Precisely, RegFSC-Net does not directly predict the full-resolution displacement field in space but learns the low-dimensional representation of the displacement field in the bandlimited Fourier domain, which is beneficial in reducing network parameters, memory usage, and computational costs. Experimental results show that RegFSC-Net outperforms various state-of-the-art methods. Specifically, in comparison to the widely recognized Transformer-based method TransMorph, RegFSC-Net utilizes only around 8.2% of its parameters, resulting in a 1.95% higher Dice score and significantly faster inference speeds of 126.67% and 419.99% on GPU and CPU, respectively. Furthermore, we also designed three variants of RegFSC-Net and demonstrated their potential applications in computer-aided diagnosis. Chenou Liu, Kangjian He, Dan Xu 0001, Hongzhen Shi, Hao Zhang 0110, Kunyuan Zhao |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | MVSFusion: infrared and visible image fusion method for multiple visual scenarios
Chengzhou Li, Kangjian He, Dan Xu 0001, Yueying Luo, Yiqiao Zhou |
Vis. Comput. | 2 |
| 2023 | Low-light image enhancement for infrared and visible image fusionabstractAbstract Infrared and visible image fusion (IVIF) is an essential branch of image fusion, and enhancing the visible image of IVIF can significantly improve the fusion performance. However, many existing low‐light enhancement methods are unsuitable for the visible image enhancement of IVIF. In order to solve this problem, this paper proposes a new visible image enhancement method for IVIF. Firstly, the colour balance and contrast enhancement‐based self‐calibrated illumination estimation (CCSCE) is proposed to improve the input image's brightness, contrast, and colour information. Then, the method based on Mutually Guided Image Filtering (muGIF) is adopted to design a strategy to extract details adaptively from the original visible image, which can keep details without introducing additional noise effectively. Finally, the proposed visible image enhancement technique is used for IVIF tasks. In addition, the proposed method can be used for the visible image enhancement of IVIF and other low‐light images. Experiment results on different public datasets and IVIF demonstrate the authors’ method's superiority from both qualitative and quantitative comparisons. The authors’ code will be publicly available at https://github.com/yiqiao666/low‐light‐enhancement‐for‐IVIF/tree/master . Yiqiao Zhou, Lisiqi Xie, Kangjian He, Dan Xu 0001, Dapeng Tao |
IET Image Process. | 3 |
| 2023 | Superpixel-based adaptive salient region analysis for infrared and visible image fusion
Chengzhou Li, Kangjian He, Dan Xu 0001, Dapeng Tao, Hongzhen Shi, Wenxia Yin |
Neural Comput. Appl. | 2 |
| 2023 | Fidelity-driven Optimization Reconstruction and Details Preserving Guided Fusion for Multi-Modality Medical ImageabstractBy integrating effective features of multi-modality medical images to provide richer information, multi-modality medical image fusion has been substantially used in computer-aided diagnosis applications. However, many existing fusion schemes do not consider how to eliminate the effects of the noise in source medical images and cannot provide enough details and textures for disease diagnosis. To address the problems above, we propose a new fidelity-driven optimization (FDO) reconstruction and details preserving guided-based fusion method for multi-modality medical images. To overcome the influence of noise in multi-modality medical images, a rank coefficient optimization method of low-rank approximation based on weighted mean curvature is proposed to reconstruct multi-modality medical image. Moreover, we propose an iterative detail preserving guided fusion (DPGF) method to integrate more textures and detail information of source multi-modality medical images, while ensuring high signal-to noise ratios. The experimental results show that the proposed method outperforms some of the state-of-the-art fusion methods. Specifically, the extensive experiments prove that our method has high robustness for noisy medical images, which also indicates the application prospects in diagnosis applications. Kangjian He, Xuejie Zhang 0002, Dan Xu 0001, Lisiqi Xie |
IEEE Trans. Multim. | 1 |
| 2023 | Skeleton-based Human Action Recognition via Large-kernel Attention Graph Convolutional NetworkabstractThe skeleton-based human action recognition has broad application prospects in the field of virtual reality, as skeleton data is more resistant to data noise such as background interference and camera angle changes. Notably, recent works treat the human skeleton as a non-grid representation, e.g., skeleton graph, then learns the spatio-temporal pattern via graph convolution operators. Still, the stacked graph convolution plays a marginal role in modeling long-range dependences that may contain crucial action semantic cues. In this work, we introduce a skeleton large kernel attention operator (SLKA), which can enlarge the receptive field and improve channel adaptability without increasing too much computational burden. Then a spatiotemporal SLKA module (ST-SLKA) is integrated, which can aggregate long-range spatial features and learn long-distance temporal correlations. Further, we have designed a novel skeleton-based action recognition network architecture called the spatiotemporal large-kernel attention graph convolution network (LKA-GCN). In addition, large-movement frames may carry significant action information. This work proposes a joint movement modeling strategy (JMM) to focus on valuable temporal interactions. Ultimately, on the NTU-RGBD 60, NTU-RGBD 120 and Kinetics-Skeleton 400 action datasets, the performance of our LKA-GCN has achieved a state-of-the-art level. Hao Zhang 0110, Kangjian He, Dan Xu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Cascade connection-based channel attention network for bidirectional medical image registration
Lingxiang Kong, Lisiqi Xie, Dan Xu 0001, Kangjian He |
Vis. Comput. | 5 |
| 2023 | Adaptive low light visual enhancement and high-significant target detection for infrared and visible image fusion
Wenxia Yin, Kangjian He, Dan Xu 0001, Yingying Yue, Yueying Luo |
Vis. Comput. | 2 |
| 2022 | OsaMOT: Occlusion and scale-aware multi-object tracking algorithm for low viewpointabstractAbstract Multi‐object tracking (MOT), which uses the context information of image sequences to locate, maintain identities and generate trajectories of multiple targets in each frame, is key technology in the field of computer vision. To address the problems of occlusion and scale variation in low‐viewpoint MOT, OsaMOT is proposed here. First, according to the global occlusion state of each frame, OsaMOT proposes the adaptive anti‐occlusion feature to enhance the awareness and adaptability for occlusion. At the same time, OsaMOT uses the cascade screening mechanism to reduce the “virtual new target” phenomenon due to the dramatic change in target features caused by scale variation and occlusion. Finally, considering that the occluded templates will affect the tracking performance, OsaMOT proposes an adaptive anti‐noise template update mechanism according to the partial occlusion state of the target, which improves the purity of the template library and further enhances the applicability to occlusion. The experimental results show that OsaMOT can weaken the influence of scale variation, partial occlusion, short‐term full occlusion and long‐term full occlusion in the low‐viewpoint tracking scenes. Most evaluation indexes of OsaMOT under low‐viewpoint tracking scenario are superior to those of some typical algorithms proposed in recent years, and the tracking robustness is improved. Yingying Yue, Dan Xu 0001, Kangjian He, Hongzhen Shi, Hao Zhang 0110 |
IET Image Process. | 3 |
| 2022 | Graph transformer network with temporal kernel attention for skeleton-based action recognitionabstractSkeleton-based human action recognition has caused wide concern, as skeleton data can robustly adapt to dynamic circumstances such as camera view changes and background interference thus allowing recognition methods to focus on robust features. In recent studies, the human body is modeled as a topological graph, and the graph convolution network (GCN) is used to extract features of actions. Although GCN has a strong ability to learn spatial modes, it ignores the varying degrees of higher-order dependencies that are captured by message passing. Moreover, the joints represented by vertices are interdependent, and hence incorporating an attention mechanism to weigh dependencies is beneficial. In this work, we propose a kernel attention adaptive graph transformer network (KA-AGTN), which models the higher-order spatial dependencies between joints by the graph transformer operator based on multihead self-attention. In addition, the Temporal Kernel Attention (TKA) block in KA-AGTN generates a channel-level attention score using temporal features, which can enhance temporal motion correlation. After combining the two-stream framework and adaptive graph strategy, KA-AGTN outperforms the baseline 2s-AGCN by 1.9% and by 1% under X-Sub and X-View on the NTU-RGBD 60 dataset, by 3.2% and 3.1% under X-Sub and X-Set on the NTU-RGBD 120 dataset, and by 2% and 2.3% under Top-1 and Top-5 and achieves the state-of-the-art performance on the Kinetics-Skeleton 400 dataset. Hao Zhang 0110, Dan Xu 0001, Kangjian He |
Knowl. Based Syst. | 4 |
| 2022 | Focus-pixel estimation and optimization for multi-focus image fusionabstractAbstract To integrate the effective information and improve the quality of multi-source images, many spatial or transform domain-based image fusion methods have been proposed in the field of information fusion. The key purpose of multi-focus image fusion is to integrate the focused pixels and remove redundant information of each source image. Theoretically, if the focused pixels and complementary information of different images are detected completely, the fusion image with best quality can be obtained. For this goal, we propose a focus-pixel estimation and optimization based multi-focus image fusion framework in this paper. Because the focused pixels of an image are in the same depth of field (DOF), we propose a multi-scale focus-measure algorithm for the focused pixels matting to integrate the focused region firstly. Then, the boundaries of focused and defocused regions are obtained accurately by the proposed optimizing strategy. And the boundaries are also fused to reduce the influence of insufficient boundary precision. The experimental results demonstrate that the proposed method outperforms some previous typical methods in both objective evaluations and visual perception. Kangjian He, Dan Xu 0001 |
Multim. Tools Appl. | 1 |
| 2022 | Adaptive enhanced infrared and visible image fusion using hybrid decomposition and coupled dictionary
Wenxia Yin, Kangjian He, Dan Xu 0001, Yueying Luo |
Neural Comput. Appl. | 2 |
| 2022 | Learning multi-level representations for affective image recognitionabstractAbstract Images can convey intense affective experiences and affect people on an affective level. With the prevalence of online pictures and videos, evaluating emotions from visual content has attracted considerable attention. Affective image recognition aims to classify the emotions conveyed by digital images automatically. The existing studies using manual features or deep networks mainly focus on low-level visual features or high-level semantic representation without considering all factors. To better understand how deep networks are working for affective recognition tasks, we investigate the convolutional features by visualization them in this work. Our research shows that the hierarchical CNN model mainly relies on deep semantic information while ignoring the shallow visual details, which are essential to evoke emotions. To form a more general and discriminative representation, we propose a multi-level hybrid model that learns and integrates the deep semantics and shallow visual representations for sentiment classification. In addition, this study shows that class imbalance would affect performance as the main category of the affective dataset will overwhelm training and degenerate the deep networks. Therefore, a new loss function is introduced to optimize the deep affective model. Experimental results on several affective image recognition datasets show that our model outperforms various existing studies. The source code is publicly available. Hao Zhang 0110, Dan Xu 0001, Gaifang Luo, Kangjian He |
Neural Comput. Appl. | 4 |
| 2021 | Adaptive colour restoration and detail retention for image enhancementabstractAbstract Computer vision‐based crowd understanding and analysis technology has been widely used in public safety due to the rapid growth of population and the frequent occurrence of various accidents. Improving imaging quality is the key to improve the performance of crowd analysis, density estimation, target recognition, segmentation, and detection in computer vision tasks. Due to the complex imaging environment such as fog and low illumination, some images taken in outdoor environment often have the problems of colour distortion, lack of details, and the poor imaging quality, which affect the subsequent visual tasks. To improve the imaging quality and visual effect, an adaptive colour restoration and detail retention‐based method is proposed for image enhancement. First, to overcome the problem of colour distortion caused by low illumination and fog, a multi‐channel fusion based adaptive image colour restoration method is proposed. To make the enhancement result more consistent with human observation, the detail retention‐based method is applied to enhance the details. Experimental results demonstrate that the authors' results are effective and outperform the compared methods both in visual and objective evaluations. Kangjian He, Dapeng Tao, Dan Xu 0001 |
IET Image Process. | 1 |
| 2021 | Contrastive learning for a single historical painting's blind super-resolutionabstractMost of the existing blind super-resolution(SR) methods explicitly estimate the kernel in pixel space, which usually has a large deviation and results in poor SR performance. As a seminal work, DASR learns abstract representations to distinguish various degradations in the feature space, which effectively reduces degradation estimation bias. Therefore, we also employ the feature space to extract degradation representations for an ancient painting. However, most of the blind SR mehods, including DASR, are committed to removing degradations introduced by kernels, downsampling and additive noise. Among them, downsampling degradation is often accompanied by unpleasant artifacts. To address this issue, the paper designs a high-resolution(HR) representation encoder EHR based on contrastive learning to distinguish artifacts introduced by downsampling. Moreover, to optimize the ill-posed nature of blind SR, we propose a contrastive regularization(CR) to minimize the contrastive loss based on VGG-19. With the help of CR, the SR images are pulled closer to the HR images and pushed far away from bicubic LR observations. Benefiting from these improvements, our method consistently achieves higher quantitative performance and better visual quality with more natural textures than state-of-the-art approaches on a specialized painting dataset. Hongzhen Shi, Dan Xu 0001, Kangjian He, Hao Zhang 0110, Yingying Yue |
Vis. Informatics | 3 |
| 2019 | Multi-focus image fusion combining focus-region-level partition and pulse-coupled neural network
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie, Xin Jin 0005 |
Soft Comput. | 1 |
| 2019 | FuseGAN: Learning to Fuse Multi-Focus Image via Conditional Generative Adversarial NetworkabstractWe study the problem of multi-focus image fusion, where the key challenge is detecting the focused regions accurately among multiple partially focused source images. Inspired by the conditional generative adversarial network (cGAN) to image-to-image task, we propose a novel FuseGAN to fulfill the images-to-image for multi-focus image fusion. To satisfy the requirement of dual input-to-one output, the encoder of the generator in FuseGAN is designed as a Siamese network. The least square GAN objective is employed to enhance the training stability of FuseGAN, resulting in an accurate confidence map for focus region detection. Also, we exploit the convolutional conditional random fields technique on the confidence map to reach a refined final decision map for better focus region detection. Moreover, due to the lack of a large-scale standard dataset, we synthesize a large enough multi-focus image dataset based on a public natural image dataset PASCAL VOC 2012, where we utilize a normalized disk point spread function to simulate the defocus and separate the background and foreground in the synthesis for each image. We conduct extensive experiments on two public datasets to verify the effectiveness of the proposed method. Results demonstrate that the proposed method presents accurate decision maps for focus regions in multi-focus images, such that the fused images are superior to 11 recent state-of-the-art algorithms, not only in visual perception, but also in quantitative analysis in terms of five metrics. Xiaopeng Guo 0001, Rencan Nie, Jinde Cao, Dongming Zhou 0001, Liye Mei, Kangjian He |
IEEE Trans. Multim. | 6 |
| 2018 | Multi-focus: Focused region finding and multi-scale transform for image fusion
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie |
Neurocomputing | 1 |
| 2018 | A lightweight scheme for multi-focus image fusion
Xin Jin 0005, Jingyu Hou 0001, Rencan Nie, Shaowen Yao 0001, Dongming Zhou 0001, Kangjian He |
Multim. Tools Appl. | 7 |
| 2018 | Multi-focus image fusion method using S-PCNN optimized by particle swarm optimization
Xin Jin 0005, Dongming Zhou 0001, Shaowen Yao 0001, Rencan Nie, Kangjian He |
Soft Comput. | 6 |