Jun Chen 0019

dblp:85/5901-19 · DBLP profile ↗
← Back
46ranked-venue papers
24as first author
30since 2021 · last 2026
0000-0001-9005-6849ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 13 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TOFusion: Text-guided and object-aware infrared and visible image fusion
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jiayi Ma 0001
Pattern Recognit.1
2026 MDbFusion++: A Visible and Infrared Image Fusion Framework Capable for Motion Deblurring
abstract
Existing image fusion methods focus on containing more complementary information, but source images always suffer from motion blur owing to object motion, which results in distorted details in fused images and further deteriorates performance on high-level tasks. This paper proposes a novel visible and infrared image fusion framework capable for motion deblurring (MDbFusion++), which can simultaneously perform image fusion and deblurring within a mutually reinforcing framework. MDbFusion++ employs a coarse-to-fine image restoration strategy and comprises two key components: a coarse deblurring part (CDP) and a fine deblurring and fusion part (FDFP). Firstly, CDP transfers multi-modal images into features corresponding to spatial locations and creatively leverages infrared features to coarsely compensate motion blurred visible ones through adaptive weights module (AWM). Subsequently, FDFP further restores fine visible features and achieves multi-modal images fusion in spatial and frequency domains with the help of multi-domain enhancement module (MEM). The deblurred visible features provide clear information to improve fusion results, and the improved fused images, in turn, provide gradient feedback to further improve deblurring effects. We evaluate our network in terms of both image deblurring and fusion, and extensive comparative experiments demonstrate the superior performance and distinct advantages of MDbFusion++.
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001
IEEE Trans. Image Process.1
2025 RAAG:Redundancy-adaptive and attention-guided token pruning for efficient video action detection
Jun Chen 0019, Sailong Deng, Wei Yu 0018, Longsheng Wei
Neurocomputing1
2025 CCAF-Net: Cascade Complementarity-Aware Fusion Network for traffic accident prediction in dashcam videos
Wei Liu 0005, Yixiang Gao, Longsheng Wei, Jun Chen 0019
Neurocomputing6
2025 UAV image stitching method based on dual feature guidance and optimal seam
Jun Chen 0019, Haikuan Gao, Beibei Yu, Xin Tian 0006
Knowl. Based Syst.1
2025 Kernel-guided injection deep network for blind fusion of multispectral and panchromatic images
Chengjie Ke, Wei Zhang 0259, Jun Chen 0019, Xin Tian 0006
Pattern Recognit.4
2025 SDSFusion: A Semantic-Aware Infrared and Visible Image Fusion Network for Degraded Scenes
abstract
A single-modal infrared or visible image offers limited representation in scenes with lighting degradation or extreme weather. We propose a multi-modal fusion framework, named SDSFusion, for all-day and all-weather infrared and visible image fusion. SDSFusion exploits the commonality in image processing to achieve enhancement, fusion, and semantic task interaction in a unified framework guided by semantic awareness and multi-scale features and losses. To address the disparity between infrared and visible images in degraded scenes, we differentiate modal features in a unified fusion model. Unlike existing joint fusion methods, we propose an adversarial generative network that refines the reconstruction of low-light images by embedding fused features. It provides feature-level brightness supplementation and image reconstruction to refine brightness and contrast. Extensive experiments in degraded scenes confirm that our approach is superior to state-of-the-art approaches in visual quality and performance, demonstrating the effectiveness of interaction improvement. The code will be posted at: https://github.com/Liling-yang/SDSFusion.
Jun Chen 0019, Liling Yang, Wei Yu 0018, Wenping Gong, Zhanchuan Cai, Jiayi Ma 0001
IEEE Trans. Image Process.1
2025 VRTNet: Vector Rectifier Transformer for Two-View Correspondence Learning
abstract
Finding reliable correspondences in two-view image and recovering the camera poses are key problems in photogrammetry and image signal processing. Multilayer perceptron (MLP) has a wide application in two-view correspondence learning for which is good at learning disordered sparse correspondences, but it is susceptible to the dominant outliers and requires additional functional blocks to capture context information. CNN can naturally extract local context information, but it cannot handle disordered data and extract global context and channel information. In order to overcome the shortcomings of MLP and CNN, we design a correspondence learning network based on Transformer, named Vector Rectifier Transformer (VRTNet). Transformer is an encoder-decoder structure which can handle disordered sparse correspondences and output sequences of arbitrary length. Therefore, we design two sub-Transformers in VRTNet to achieve the mutual conversion between disordered and ordered correspondences. The self-attention and cross-attention mechanisms in them allow VRTNet to focus on the global context relations of all correspondences. To capture local context and channel information, we propose rectifier network (including CNN and channel attention block) as the backbone of VRTNet, which avoids the complex design of additional blocks. Rectifier network can correct the errors of ordered correspondences to obtain rectified correspondences. Finally, outliers are removed by comparing original and rectified correspondences. VRTNet performs better than the state-of-the-art methods in the tasks of relative pose estimation, outlier removal and image registration.
Meng Yang 0031, Jun Chen 0019, Xin Tian 0006, Longsheng Wei, Jiayi Ma 0001
IEEE Trans. Multim.2
2025 Deep Variational Network for Blind Pansharpening
abstract
Deep-learning-based methods play an important role in pansharpening that uses panchromatic images to enhance the spatial resolution of multispectral images while maintaining spectral features. However, most existing methods mainly consider only one fixed degradation in the training process. Therefore, their performance may drop significantly when the degradation of testing data is unknown (blind) and different from the training data, which is common in real-world applications. To address this issue, we proposed a deep variational network for blind pansharpening, named VBPN, which integrates degradation estimation and image fusion into a whole Bayesian framework. First, by taking the noise and blurring parameters of the multispectral image with the noise parameters of the panchromatic image as hidden variables, we parameterize the approximate posterior distribution for the fusion problem using neural networks. Since all parameters in this posterior distribution are explicitly modeled, the degradation parameters of the multispectral image and the panchromatic image are easily estimated. Furthermore, we designed VPBN composed of degradation estimation and image fusion subnetworks, which can optimize the fusion results guided by the variational inference according to the testing data. As a result, the blind pansharpening performance can be improved. In general, VPBN has good interpretability and generalization ability by combining the advantages of model-based and deep-learning-based approaches. Experiments on simulated and real datasets prove that VPBN can achieve state-of-the-art fusion results.
Chengjie Ke, Jun Chen 0019, Xin Tian 0006
IEEE Trans. Neural Networks Learn. Syst.4
2024 Mdbfusion: A Visible And Infrared Image Fusion Framework Capable For Motion Deblurring
abstract
Existing image fusion methods focus on containing more complementary information, but source images always suffer from motion blur owing to object motion, which results in distorted details in fused images and further deteriorates performance on high-level tasks. This paper proposes a novel visible and infrared image fusion framework capable for motion deblurring (MDbFusion++), which can simultaneously perform image fusion and deblurring within a mutually reinforcing framework. MDbFusion++ employs a coarse-to-fine image restoration strategy and comprises two key components: a coarse deblurring part (CDP) and a fine deblurring and fusion part (FDFP). Firstly, CDP transfers multi-modal images into features corresponding to spatial locations and creatively leverages infrared features to coarsely compensate motion blurred visible ones through adaptive weights module (AWM). Subsequently, FDFP further restores fine visible features and achieves multi-modal images fusion in spatial and frequency domains with the help of multi-domain enhancement module (MEM). The deblurred visible features provide clear information to improve fusion results, and the improved fused images, in turn, provide gradient feedback to further improve deblurring effects. We evaluate our network in terms of both image deblurring and fusion, and extensive comparative experiments demonstrate the superior performance and distinct advantages of MDbFusion++.
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001
ICIP1
2024 Cross-Scale Domain Adaptation with Comprehensive Information for Pansharpening
Meiqi Gong, Hao Zhang 0073, Hebaixu Wang, Jun Chen 0019, Jun Huang 0008, Xin Tian 0006, Jiayi Ma 0001
IJCAI4
2024 Infrared and visible image fusion via gradientlet filter and salience-combined map
Jun Chen 0019, Cai Lei, Liu Wei
Multim. Tools Appl.1
2024 Various Degradation: Dual Cross-Refinement Transformer for Blind Sonar Image Super-Resolution
abstract
Deep learning-based methods have achieved remarkable results in super-resolution (SR) of sonar images. However, most existing methods only consider simple bicubic downsampling degradation, and SR networks suitable for natural images may not be suitable for sonar images. Therefore, they perform poorly on sonar images with unknown degradation parameters in real-world scenarios (i.e.,blindscenario). To address these issues, we propose a dual cross-refinement transformer (DCRT) forblindSR of sonar images. DCRT first constructs a large-scale degradation space based on the sonar image imaging mechanism. More importantly, we randomly sample the task-level training information to make DCRT robust on different SR tasks, thereby enhancing theblindSR capability of the network. Then, DCRT focuses on image features than domain features through spatial-channel self-attention cross-fusion block (S-C-SACFB), so the domain gap between the training and testing data can be reduced. Meanwhile, S-C-SACFB effectively combines inter-attention and high-frequency enhancement residual block to enhance the network’s ability to extract high-frequency features while suppressing speckle noise in sonar images. Finally, DCRT uses global residual connections to generate high-resolution sonar images. A large number of experiments at different SR scale show that DCRT outperforms the state–of–the–art methods in both quantitative and qualitative aspects.
Jiahao Rao, Yini Peng, Jun Chen 0019, Xin Tian 0006
IEEE Trans. Geosci. Remote. Sens.3
2024 HitFusion: Infrared and Visible Image Fusion for High-Level Vision Tasks Using Transformer
abstract
This study proposes an innovative network to fuse infrared and visible images, called HitFusion, which uses the cross-feature transformer module and is compatible with high-level vision tasks. Firstly, existing image fusion approaches primarily concentrate on optimizing human visual perception and image metrics. To enhance the performance of the fusion network in subsequent high-level vision tasks, a segmentation network and a corresponding loss are introduced into the fusion network training process. Specifically, we devise a three-stage training strategy to render the fusion network more suitable for high-level vision tasks, guided by the segmentation network and broadening the fusion network's training set to boost its generalization capability. Secondly, current transformer-based image fusion methods neglect the interaction between visible texture features and infrared contrast features. To tackle this, the cross-feature transformer module is proposed, allowing the fusion network to learn the cross-feature correlation and long-range dependencies between source images, thus achieving fusion results with good complementarity. Finally, a dual-branch fusion network is proposed, based on the distinct characteristics of different images, that targets the extraction of deep features from source images utilizing contrast residual and texture enhancement modules to achieve improved fusion results. Extensive experimental results reveal that our HitFusion method excels in both qualitative and quantitative assessments, while also demonstrating superior performance in addressing high-level vision tasks.
Jun Chen 0019, Jianfeng Ding, Jiayi Ma 0001
IEEE Trans. Multim.1
2023 THFuse: An infrared and visible image fusion network using transformer and hybrid feature extractor
Jun Chen 0019, Jianfeng Ding, Yang Yu 0045, Wenping Gong
Neurocomputing1
2023 THAT-Net: Two-layer hidden state aggregation based two-stream network for traffic accident prediction
Wei Liu 0005, Yisheng Lu, Jun Chen 0019, Longsheng Wei
Inf. Sci.4
2023 Fusion of near-infrared and visible images based on saliency-map-guided multi-scale transformation decomposition
Jun Chen 0019, Cai Lei, Liu Wei
Multim. Tools Appl.1
2023 Multi-Neighborhood Guided Kendall Rank Correlation Coefficient for Feature Matching
abstract
Seeking feature correspondences among two or more images is an important problem in computer vision and image processing. The putative matches constructed by the similarity of feature descriptors are often contaminated by many false matches. Typically, the local neighborhood points of a true match point have a rank order, which will be maintained in the corresponding image, and we call it rank consistency. In this paper, we design a number of sorting plans to obtain the neighborhood rank lists by taking full advantage of the local neighborhood geometry structure. In order to measure the differences between rank lists, we adopt the statistically famous Kendall rank correlation coefficient and generalize its definition for matching problem. We design a neighborhood common element guidance strategy and a multi-neighborhood strategy to improve the universality and robustness of our method. Our method has linear complexity and it has superiority over state-of-the-art methods on several challenging data sets. It also performs well in image registration and loop-closure detection tasks. The source code of our method is publicly available athttps://github.com/MnYangs/mGKRCC.
Jun Chen 0019, Meng Yang 0031, Wenping Gong, Yang Yu 0045
IEEE Trans. Multim.1
2023 DMEF: Multi-Exposure Image Fusion Based on a Novel Deep Decomposition Method
abstract
In this paper, we propose a novel deep decomposition approach based on Retinex theory for multi-exposure image fusion, termed as DMEF. According to the assumption of Retinex theory, we firstly decompose the source images into illumination and reflection maps by the data-driven decomposition network, among which we introduce the pathwise interaction block that reactivates the deep features lost in one path and embeds them into another path. Therefore, loss of illumination and reflection features during decomposition can be effectively suppressed. And then the high dynamic range illumination map could be obtained by fusing the separated illumination maps in the fusion network. Thus, the reconstructed details in under-exposed and over-exposed regions will be clearer with the help of the fused reflection map which contains complete high-frequency scene information. Finally, the fused illumination and reflection maps are multiplied pixel-by-pixel to obtain the final fused image. Moreover, to retain the discontinuity in the illumination map where gradient of reflection map changes steeply, we introduce the structure-preservation smoothness loss function to retain the structure information and eliminate visual artifacts in these regions. The superiority of our proposed network is demonstrated by applying extensive experiments compared with other state-of-the-art fusion methods subjectively and objectively.
Kangle Wu, Jun Chen 0019, Jiayi Ma 0001
IEEE Trans. Multim.2
2023 ACE-MEF: Adaptive Clarity Evaluation-Guided Network With Illumination Correction for Multi-Exposure Image Fusion
abstract
For a natural scene with nonuniform environment light, the captured visible images are always under- or over-exposed because of the limited dynamic range of digital imaging devices. Multi-exposure image fusion (MEF) is a mainstream and effective solution. For a local region that has friendly visual effect in one exposure setting but extremely bad-exposed in another, most existing MEF methods have the ability to transfer the scene detail information to the fused images. However, they will be affected by the over-high or -low light inevitably thus resulting in local visibility reduction. To address this issue, we propose an adaptive clarity evaluation-guided network with illumination correction for MEF in a coarse-to-fine manner, which is termed as ACE-MEF. To be specific, our ACE-MEF is mainly composed of two modules: clarity preservation network (CPN) and illumination adjustment network (IAN). Based on the adaptive clarity evaluation, CPN could be trained to coarsely preserve the environment light and texture details of the clearer regions in source images. Therefore, the need for labeled reference images that are time-consuming to obtain could be mitigated. By measuring the parameter maps of gamma function, IAN is able to refine and correct the local bad-exposed regions so that more details could be further revealed. Extensive experiments demonstrate that our method outperforms multiple state-of-the-art algorithms qualitatively and quantitatively.
Kangle Wu, Jun Chen 0019, Yang Yu 0045, Jiayi Ma 0001
IEEE Trans. Multim.2
2022 UAV Image Stitching Using Shape-Preserving Warp Combined With Global Alignment
abstract
In this letter, we propose a strategy for unmanned aerial vehicle (UAV) image stitching to generate natural-looking panoramas. Traditional methods using homography to perform alignment cannot account for images with parallax, so they require that the input images should be taken from the same viewpoint or the scene should be near the planar. However, remote sensing images obtained by UAVs usually do not satisfy such an ideal situation, and the stitching results always suffer from artifacts. To overcome these challenges and obtain natural-looking panoramas, a global alignment strategy is proposed to better align the input images. Combined with a shape-preserving warp, the stitching results can achieve better alignment accuracy while maintaining the shape. Meanwhile, locality preserving matching (LPM) is used to eliminate mismatches during feature detection and matching for accurate alignment. In addition, to make the stitching results more natural-looking, we also use multiband blending to eliminate artifacts that may exist in the results due to unmodeled effects. Experiments show that our stitching strategy can effectively improve alignment accuracy and obtain natural-looking results compared to other state-of-the-art methods.
Donghai Guo, Jun Chen 0019, Linbo Luo 0002, Wenping Gong, Longsheng Wei
IEEE Geosci. Remote. Sens. Lett.2
2022 Two-view correspondence learning via complex information extraction
Jun Chen 0019, Linbo Luo 0002, Wenping Gong, Yong Wang 0036
Multim. Tools Appl.1
2022 ASF-Net: Adaptive Screening Feature Network for Building Footprint Extraction From Remote-Sensing Images
abstract
Building footprint extraction plays an important role in many remote-sensing (RS) applications such as urban planning and disaster monitoring. Mainly, the exploitation of contextual information in a fixed receptive field is the focus of previous research, which makes it difficult to generically extract buildings that vary greatly in size and shape, especially when isolated large buildings are surrounded by dense small buildings. To improve this problem, we attempt to teach the network to adjust the receptive field and enhance useful feature information adaptively. In this article, we propose a novel adaptive screening feature network (ASF-Net), which can independently screen and enhance effective feature information from two aspects. On the one hand, we propose a deepened space up-sampling block to screen useful information and help establish boundaries. On the other hand, we propose an Adaptive Information Utilization Block (AIUB) to enlarge the receptive field of feature maps and refine the incomplete building footprint. As a result, the more accurate multiscale building footprint is inferred from the enhanced features. Experimental results on the popular aerial image segmentation datasets show that ASF-Net obtains competitive results [80.2% intersection over union (IoU) on the Inria aerial image labeling dataset and 74.2% IoU on the Massachusetts buildings dataset] in comparison with several state-of-the-art models. The TensorFlow implementation is available athttps://github.com/jyx0516/ASF-Net.
Jun Chen 0019, Yuxuan Jiang 0007, Linbo Luo 0002, Wenping Gong
IEEE Trans. Geosci. Remote. Sens.1
2022 Robust Feature Matching via Local Consensus
abstract
Feature matching is the foundation and key task of remote sensing image registration, which is to establish a reliable point corresponding relationship between the feature points of two images. In this article, a simple and effective local consensus method for rigid and nonrigid feature matching is proposed and applied to solve the problem of high outliers ratio caused by nonrigid transformation, nonlinear radiation difference, and speckle noise in the remote sensing image registration task. We first establish the putative feature correspondences according to the similarity between local descriptors and then use local consensus constraints (including neighborhood consensus and motion vector consensus) to remove outliers. The specific steps are given as follows. First, we use the neighborhood consensus constraint of feature points to carry out preliminary filtering to remove outliers with obvious errors and retain a large number of inliers, so as to obtain a clean reliable set. Then, the reliable set space is grided into several nonoverlapping cells, and the estimated motion vector is calculated for each cell. By taking the comprehensive deviation between the ordinary motion vectors and estimated motion vectors, we transform the matching problem into a mathematical optimization model and derive a closed-form solution with linear time and linear space complexities. In this way, our method can also significantly increase the speed of operation without sacrificing accuracy. A large number of feature matching experiments on remote sensing prove that our method is superior to existing methods and also has good results in the general scene.
Jun Chen 0019, Meng Yang 0031, Chengli Peng, Linbo Luo 0002, Wenping Gong
IEEE Trans. Geosci. Remote. Sens.1
2022 Multi-Focus Image Fusion Based on Multi-Scale Gradients and Image Matting
abstract
Multi-focus image fusion technology is to extract different focused regions of the same scene among partially focused images and merge them together to generate a composite image where all objects are clear. Two crucial points to multi-focus image fusion are the effective focus measurement method to evaluate the sharpness of the source images and the accurate segmentation method to extract the focused regions. In conventional multi-focus image fusion methods, the decision map obtained according to the focus measurement is sensitive to mis-registration, or produces an uneven boundary lines. In this paper, the maximum value in the top-hat transform and the bottom-hat transform is used as the gradient measurement value, and the complementary features between multiple scales are used to achieve accurate focus measurement for initial segmentation. In order to obtain a better fusion decision map, a robust image matting algorithm is used to refine the trimap generated by the initial segmentation. Then, make full use of the strong correlation between the source images to optimize the edge regions of the decision map to improve the image fusion quality. Finally, a fusion image is constructed based on the fusion decision map and the source images. We perform qualitative and quantitative experiments on publicly available databases to verify the effectiveness of the method. The results show that compared with several state-of-the-art algorithms, the proposed fusion method can obtain accurate decision maps and achieve better performance in visual perception and quantitative analysis.
Jun Chen 0019, Linbo Luo 0002, Jiayi Ma 0001
IEEE Trans. Multim.1
2021 Building Footprint Generation by Integrating U-Net with Deepened Space Module
abstract
In this paper, we propose a novel and practical convolutional neural network method for building footprint generation in remote sensing images, in order to deal with the problem that the detailed information and geometric structure of ground objects in high-resolution images become more abundant, which leads to a large increase in the calculation amount. So we introduce a deepened space module, which can ignore the channels with weak target features and emphasize the effective features. It is embedded in each splicing layer in the upsampling process of U-net to achieve the effect of feature selection. By means of clipping and data enhancement, we carry out iterative training and model optimization learning on Inria aerial image label dataset, and realize the automatic generation of building footprint. Compared with FCN8s, Unet, SegNet, PSPNet, Deeplabv3 + and GLNet, experimental results show that the method we use to generate building footprint is more accurate, and in IoU, mPA, PA three indicators are better than the comparison algorithms.
Jun Chen 0019, Yuxuan Jiang 0007, Linbo Luo 0002, Kangle Wu
ICIP1
2021 Effective Feature Fusion Network in BIFPN for Small Object Detection
abstract
In view of the difficulty and low accuracy of small object detection in remote sensing images, this paper proposes a bidirectional cross-scale connection feature fusion network with an information direct connection layer and a shallow information fusion layer. Aiming at the problem that the detection targets in remote sensing images are mainly small and medium-sized targets, we fuse the shallow feature maps with rich spatial information in the bidirectional cross-scale connection feature fusion network instead of directly using the shallow feature maps for regression and classification. While ensuring the model inference speed, the detection accuracy of small objects is improved. At the same time, we use the information direct connection layer to perform feature fusion with the initial information in each iteration of the bidirectional cross-scale connection feature fusion pyramid to prevent the loss of small object information. Experimental results show that the algorithm proposed in this paper can obtain good accuracy and real-time performance on the NWPU VHR-10 dataset.
Jun Chen 0019, HongSheng Mai, Linbo Luo 0002, Kangle Wu
ICIP1
2021 Building Area Estimation in Drone Aerial Images Based on Mask R-CNN
abstract
In rural areas where disasters occur frequently, the calculation of building areas is crucial in property assessment. In the segmentation algorithm, Mask R-CNN can distinguish the adjacent objects and extract the outline of an object. Based on this observation, we propose a novel method to calculate the building areas based on Mask R-CNN and adopt the concept of transfer learning to train our model, which can achieve good results with a small number of drone aerial images as training samples. The proposed method involves three main steps: 1) pretraining using open-source satellite remote sensing images; 2) fine-tuning with a small number of drone aerial images; and 3) testing with new images and area calculation based on the number of building pixels. The experiments show that the proposed method can achieve good results in terms of F1 score and intersection over union.
Jun Chen 0019, Ganbei Wang, Linbo Luo 0002, Wenping Gong
IEEE Geosci. Remote. Sens. Lett.1
2021 A saliency-based multiscale approach for infrared and visible image fusion
Jun Chen 0019, Kangle Wu, Linbo Luo 0002
Signal Process.1
2021 Drone Image Stitching Using Local Mesh-Based Bundle Adjustment and Shape-Preserving Transform
abstract
This article proposes a strategy for drone image stitching using local mesh-based bundle adjustment and shape-preserving transform, which aims to effectively stitch multiple overlapping drone images into a natural panoramic image. Existing traditional methods using a simple homography cannot handle the situation that the input drone images have parallax effect, and the image mosaic result always suffers from artifacts. In order to achieve natural-looking stitching results without the above limitation, we divide the proposed method into the following steps. Starting from initial feature sets obtained by off-the-shelf feature extraction methods, we incorporate the parallax errors into an energy minimum framework and construct a robust alignment energy. This energy can be minimized efficiently based on local bundle adjustment and robust$3\sigma $principle, which could eliminate parallax effects and achieve accurate alignment. Then the seamless panoramic image is obtained by warping the target image and the source images onto the mesh plane directly. An image patch can be transformed by projective transformation (e.g., homography), which provides good alignment but may cause distortions. Consequently, combined with mesh-based shape-preserving transform, our proposed strategy can improve the naturalness of the results flexibly. Experiments show that our stitching strategy can eliminate parallax effects more effectively and achieve natural-looking results compared to other state-of-the-art methods.
Qi Wan, Jun Chen 0019, Linbo Luo 0002, Wenping Gong, Longsheng Wei
IEEE Trans. Geosci. Remote. Sens.2
2020 Multiscale Infrared and Visible Image Fusion Based on Phase Congruency and Saliency
abstract
In this paper, in order to enhance the infrared target in infrared image and retain the edge and detail information in visible image, we propose a multi-scale decomposition fusion method based on phase congruency and saliency. In this method, the Laplacian pyramid is first used to decompose the source image into detail layers and base layers. Secondly, we use a method based on phase congruency for the fusion of detail layers. Thirdly, for the base layer, we decompose it into saliency map and residual map. The “max absolute” rule and “averag” rule are adopted for the fusion of saliency map and residual map, then the fused saliency map and residual map are added to attain the fused base image. Finally, we use the inverse transform of Laplacian pyramid to reconstruct the fused image. The experimental results show that the proposed method have better fusion effect than other methods. What's outstanding is that the infrared targets in the fused image are enhanced and abundant edges are preserved.
Jun Chen 0019, Kangle Wu, Linbo Luo 0002, Xin Tian 0006
IGARSS1
2020 Drone Image Stitching Using Local Least Square Alignment
abstract
This paper proposes a strategy for drone image stitching using local least square alignment, which aims to effectively stitch multiple overlapping drone images into a natural panoramic image. Existing traditional methods using simple homography cannot handle the situation that the input drone images have parallax effect, and the mosaic result always suffers from artifacts. In order to achieve natural-looking stitching results without the above limitation, we divide the proposed method into the following two steps, namely, local least square alignment and global similarity constraint. Starting from initial feature sets obtained by traditional feature extraction methods, we construct a robust alignment energy based on parallax errors to adaptively eliminate parallax effects. The energy can be efficiently minimized used least square estimate. Combined with global similarity constraint, our proposed strategy can flexibly improve the naturalness of the results. Experiments show that our stitching strategy can more effectively eliminate parallax effects and achieve natural-looking results compared to other state-of-the-art methods.
Qi Wan, Linbo Luo 0002, Jun Chen 0019, Yong Wang 0036, Donghai Guo
IGARSS3
2020 UAV Image Mosaicing Based Multi-Region Local Projection Deformation
abstract
The goal of unmanned aerial vehicle (UAV) image mosaicing is to create natural-looking mosaics free of artifacts due to the parallax of the image and relative camera motion. UAV remote sensing is a low-altitude technology and the UAV imaged scene is not effectively planar, yielding parallax on images. In this paper, we apply local homography to match UAV images, which can reduce misalignment artifacts or “ghosting” in the results compared with 2D projective transforms or global homography. In addition, when an object in three dimensions is mapped to an image plane, different surfaces have different projections. These projections vary with the viewpoint in a sequence of UAV images, which still causes artifacts near some tall buildings if we only use local homography. We propose a novel stitching method based multi-region local projection deformation, that divides the overlapping regions of input images into several regions, then meshes image to calculate local projections by partitioned regions. Specifically, we use a strategy where multiple regions have different weights for calculating local projections, which can significantly reduce ghosting due to these projections vary with the viewpoint and parallax. The benefits of the proposed approach are demonstrated using a variety of challenging cases.
Linbo Luo 0002, Jun Chen 0019, Wenping Gong, Donghai Guo
IGARSS3
2020 Infrared and visible image fusion based on target-enhanced multiscale transform decomposition
Jun Chen 0019, Linbo Luo 0002, Xiaoguang Mei, Jiayi Ma 0001
Inf. Sci.1
2019 Progressive Filtering for Feature Matching
abstract
In this paper, we propose a simple yet efficient method termed as Progressive Filtering for Feature Matching, which is able to establish accurate correspondences between two images of common or similar scenes. Our algorithm first grids the correspondence space and calculates a typical motion vector for each cell, and then removes false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive matching strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. The density estimation is utilized to address the island samples and accelerate the convergency of the mismatch removal procedure. In addition, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Extensive experiments on several representative real images involving different types of geometric transformations demonstrate the superiority of our approach over the state-of-the-art.
Xingyu Jiang 0005, Jiayi Ma 0001, Jun Chen 0019
ICASSP3
2019 Remote Sensing Image Matching using TPS Transformation and Local Geometrical Constraint
abstract
Focusing on the characteristics of remote sensing images, this study proposes a new algorithm for feature matching of remote sensing images to eliminate mismatch. The algorithm utilizes feature descriptors, such as scale-invariant feature transform, for rough correspondence and the thin-plate spline for non-rigid transformation. Under the Bayesian framework, correspondence and transformation are alternately optimized by the expectation-maximization algorithm to automatically eliminate mismatched points. We also introduce a local geometrical constraint to maintain the internal structure of adjacent feature points. We apply this method to a large number of remote sensing images, and the experimental results reveal the method's superiority over the state-of-the-art.
Jun Chen 0019, Linbo Luo 0002, Wenping Gong
IGARSS1
2019 Drone Image Stitching Guided by Robust Elastic Warping and Locality Preserving Matching
abstract
Image stitching stitches multiple overlapping images into a seamless image according to the corresponding geometric relationship between the reference and source images. In this study, the parallax-tolerant image stitching method based on robust elastic warping is applied to the stitching of drone images, and locality-preserving feature matching is used to effectively remove outliers from the drone images. The method can be divided into three stages, namely, locality-preserving feature matching, robust elastic warping, and global projectivity preservation. First, a set of high- precision point matching is provided for a drone image, and local matching is used. Second, the robust elastic warping function eliminates the parallax error, and the input image is distorted according to the calculated deformation on the grid plane. Finally, the global projectivity-preserving method is applied to obtain high-precision result panoramas. Experiments on several sets of drone images demonstrate that our method can generate better panoramas over the competitors.
Linbo Luo 0002, Qi Wan, Jun Chen 0019, Yongtao Wang, Xiaoguang Mei
IGARSS3
2019 Uav Image Mosaic Based on Non-Rigid Matching and Bundle Adjustment
abstract
This study introduces a robust method for panoramic unmanned aerial vehicle (UAV) image mosaic. The traditional automatic panoramic image stitching method requires the camera to carefully rotate the optical center to obtain an image, but the image used for mosaic in reality cannot easily achieve this ideal state. In particular, remote sensing images obtained by UAVs do not satisfy such a situation. The images may not be on a plane yet, and several of them may even have non-rigid changes. Therefore, the classical method of UAV image stitching is expected to produce poor results. To this end, we improve the traditional stitching method to overcome the abovementioned challenges. Specifically, a non-rigid matching algorithm is introduced to the system to make it suitable for remote sensing images. We perform bundle adjustments using a new strategy to make the mosaic system suitable for UAV images. Experimental results show that our method is more robust than the traditional method.
Linbo Luo 0002, Jun Chen 0019, Tao Lu 0001, Yong Wang 0036
IGARSS3
2019 Gaussian field estimator with manifold regularization for retinal image registration
Jiahao Wang 0001, Jun Chen 0019, Shuaibin Zhang, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001
Signal Process.2
2016 Robust image matching via feature guided Gaussian mixture model
abstract
In this paper, we propose a novel feature guided Gaussian mixture model (FG-GMM) for image matching, which typically requires matching two sets of feature points extracted from the given images. We formulate the problem as estimation of a feature guided mixture of densities: a GMM is fitted to one point set, such that both the centers and local features of the Gaussian densities are constrained to coincide with the other point set. The problem is solved under a unified maximum-likelihood framework together with an iterative semi-supervised Expectation-Maximization (EM) algorithm initialized by the confident feature correspondences. The image transformation is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on various real images show the robustness of our approach, which consistently outperforms other state-of-the-art methods.
Jiayi Ma 0001, Junjun Jiang, Yuan Gao 0015, Jun Chen 0019, Chengyin Liu
ICME4
2016 Registration of remote sensing images with non-rigid distortions
abstract
In this paper, we propose a novel formulation for building accurate pixel-wise alignments between remote sensing images under non-rigid distortions. Our formulation involves two variables: the first is a discrete displacement flow field similar to optical flow which controls the pixel-wise correspondence and allows piecewise smoothness, while the second is a continuous spatial transformation which fits for a few confidential sparse feature correspondences. An additional term is introduced to ensure the coherence between the two variables, and the continuous spatial transformation plays a role of anchor for optimizing the discrete displacement flow field. Experiments on real remote sensing images demonstrate that our approach greatly outperforms state-of-the-art methods.
Jiayi Ma 0001, Jun Chen 0019, Yong Ma 0001
IGARSS2
2016 Multimodal retinal image registration using edge map and feature guided Gaussian mixture model
abstract
In this paper, we propose a method for multimodal retinal image registration based on feature guided Gaussian mixture model (GMM) and edge map. We extract two sets of feature points from the edge maps of two images, and formulate image registration as the estimation of a feature guided mixture of densities: a GMM is fitted to one point set, such that both the centers and local features of the Gaussian densities are constrained to coincide with the other point set. The problem is solved under a maximum-likelihood framework together with an iterative EM algorithm initialized by confident feature matches, where the image transformation is modeled by an affine function. Extensive experiments on various retinal images show the robustness of our method, which consistently outperforms other state-of-the-arts, especially when the data is badly degraded.
Jiayi Ma 0001, Junjun Jiang, Jun Chen 0019, Chengyin Liu, Chang Li 0001
VCIP3
2016 Infrared and visible image fusion using total variation model
Yong Ma 0001, Jun Chen 0019, Chen Chen 0003, Fan Fan 0001, Jiayi Ma 0001
Neurocomputing2
2016 Image retrieval based on image-to-class similarity
Jun Chen 0019, Yong Wang 0036, Linbo Luo 0002, Jin-Gang Yu, Jiayi Ma 0001
Pattern Recognit. Lett.1
2015 Non-rigid point set registration via coherent spatial mapping
Jun Chen 0019, Jiayi Ma 0001, Changcai Yang
Signal Process.1
2013 On contrast combinations for visual saliency detection
abstract
Saliency detection is an important task in computer vision and image processing. The most influential factor in bottom-up visual saliency is contrast operation. In this paper, we propose a unified model to combine widely used contrast measurements, namely, center-surround, corner-surround and global contrast to detect visual saliency. The proposed model benefits from the advantages of each individual contrast operation, and thus produces more robust and accurate saliency maps. Extensive experimental results on natural images show the effectiveness of the proposed model for visual saliency detection task, and demonstrate the combination is superior than individual subcomponent.
Quan Zhou 0004, Shiwei Ren, Yu Zhou 0016, Jun Chen 0019, Wenyu Liu 0001
ICIP5