Xin Tian 0006

dblp:16/4804-6 · DBLP profile ↗
← Back
47ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0003-1993-2708ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TOFusion: Text-guided and object-aware infrared and visible image fusion
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jiayi Ma 0001
Pattern Recognit.4
2026 Intermodal correlation modeling for incomplete multi-modal learning in land use and land cover classification
Yini Peng, Rui Liu 0041, Yishan Wang, Xin Tian 0006
Pattern Recognit.6
2026 MDbFusion++: A Visible and Infrared Image Fusion Framework Capable for Motion Deblurring
abstract
Existing image fusion methods focus on containing more complementary information, but source images always suffer from motion blur owing to object motion, which results in distorted details in fused images and further deteriorates performance on high-level tasks. This paper proposes a novel visible and infrared image fusion framework capable for motion deblurring (MDbFusion++), which can simultaneously perform image fusion and deblurring within a mutually reinforcing framework. MDbFusion++ employs a coarse-to-fine image restoration strategy and comprises two key components: a coarse deblurring part (CDP) and a fine deblurring and fusion part (FDFP). Firstly, CDP transfers multi-modal images into features corresponding to spatial locations and creatively leverages infrared features to coarsely compensate motion blurred visible ones through adaptive weights module (AWM). Subsequently, FDFP further restores fine visible features and achieves multi-modal images fusion in spatial and frequency domains with the help of multi-domain enhancement module (MEM). The deblurred visible features provide clear information to improve fusion results, and the improved fused images, in turn, provide gradient feedback to further improve deblurring effects. We evaluate our network in terms of both image deblurring and fusion, and extensive comparative experiments demonstrate the superior performance and distinct advantages of MDbFusion++.
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001
IEEE Trans. Image Process.3
2025 Multi-receptive field interaction network for shape from polarization
Yini Peng, Rui Liu 0041, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006
Sci. China Inf. Sci.6
2025 UAV image stitching method based on dual feature guidance and optimal seam
Jun Chen 0019, Haikuan Gao, Beibei Yu, Xin Tian 0006
Knowl. Based Syst.4
2025 Kernel-guided injection deep network for blind fusion of multispectral and panchromatic images
Chengjie Ke, Wei Zhang 0259, Jun Chen 0019, Xin Tian 0006
Pattern Recognit.5
2025 VRTNet: Vector Rectifier Transformer for Two-View Correspondence Learning
abstract
Finding reliable correspondences in two-view image and recovering the camera poses are key problems in photogrammetry and image signal processing. Multilayer perceptron (MLP) has a wide application in two-view correspondence learning for which is good at learning disordered sparse correspondences, but it is susceptible to the dominant outliers and requires additional functional blocks to capture context information. CNN can naturally extract local context information, but it cannot handle disordered data and extract global context and channel information. In order to overcome the shortcomings of MLP and CNN, we design a correspondence learning network based on Transformer, named Vector Rectifier Transformer (VRTNet). Transformer is an encoder-decoder structure which can handle disordered sparse correspondences and output sequences of arbitrary length. Therefore, we design two sub-Transformers in VRTNet to achieve the mutual conversion between disordered and ordered correspondences. The self-attention and cross-attention mechanisms in them allow VRTNet to focus on the global context relations of all correspondences. To capture local context and channel information, we propose rectifier network (including CNN and channel attention block) as the backbone of VRTNet, which avoids the complex design of additional blocks. Rectifier network can correct the errors of ordered correspondences to obtain rectified correspondences. Finally, outliers are removed by comparing original and rectified correspondences. VRTNet performs better than the state-of-the-art methods in the tasks of relative pose estimation, outlier removal and image registration.
Meng Yang 0031, Jun Chen 0019, Xin Tian 0006, Longsheng Wei, Jiayi Ma 0001
IEEE Trans. Multim.3
2025 Deep Variational Network for Blind Pansharpening
abstract
Deep-learning-based methods play an important role in pansharpening that uses panchromatic images to enhance the spatial resolution of multispectral images while maintaining spectral features. However, most existing methods mainly consider only one fixed degradation in the training process. Therefore, their performance may drop significantly when the degradation of testing data is unknown (blind) and different from the training data, which is common in real-world applications. To address this issue, we proposed a deep variational network for blind pansharpening, named VBPN, which integrates degradation estimation and image fusion into a whole Bayesian framework. First, by taking the noise and blurring parameters of the multispectral image with the noise parameters of the panchromatic image as hidden variables, we parameterize the approximate posterior distribution for the fusion problem using neural networks. Since all parameters in this posterior distribution are explicitly modeled, the degradation parameters of the multispectral image and the panchromatic image are easily estimated. Furthermore, we designed VPBN composed of degradation estimation and image fusion subnetworks, which can optimize the fusion results guided by the variational inference according to the testing data. As a result, the blind pansharpening performance can be improved. In general, VPBN has good interpretability and generalization ability by combining the advantages of model-based and deep-learning-based approaches. Experiments on simulated and real datasets prove that VPBN can achieve state-of-the-art fusion results.
Chengjie Ke, Jun Chen 0019, Xin Tian 0006
IEEE Trans. Neural Networks Learn. Syst.5
2024 Mdbfusion: A Visible And Infrared Image Fusion Framework Capable For Motion Deblurring
abstract
Existing image fusion methods focus on containing more complementary information, but source images always suffer from motion blur owing to object motion, which results in distorted details in fused images and further deteriorates performance on high-level tasks. This paper proposes a novel visible and infrared image fusion framework capable for motion deblurring (MDbFusion++), which can simultaneously perform image fusion and deblurring within a mutually reinforcing framework. MDbFusion++ employs a coarse-to-fine image restoration strategy and comprises two key components: a coarse deblurring part (CDP) and a fine deblurring and fusion part (FDFP). Firstly, CDP transfers multi-modal images into features corresponding to spatial locations and creatively leverages infrared features to coarsely compensate motion blurred visible ones through adaptive weights module (AWM). Subsequently, FDFP further restores fine visible features and achieves multi-modal images fusion in spatial and frequency domains with the help of multi-domain enhancement module (MEM). The deblurred visible features provide clear information to improve fusion results, and the improved fused images, in turn, provide gradient feedback to further improve deblurring effects. We evaluate our network in terms of both image deblurring and fusion, and extensive comparative experiments demonstrate the superior performance and distinct advantages of MDbFusion++.
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001
ICIP3
2024 Cross-Scale Domain Adaptation with Comprehensive Information for Pansharpening
Meiqi Gong, Hao Zhang 0073, Hebaixu Wang, Jun Chen 0019, Jun Huang 0008, Xin Tian 0006, Jiayi Ma 0001
IJCAI6
2024 Re-decoupling the classification branch in object detectors for few-class scenes
Jie Hua 0005, Zhongyuan Wang 0001, Qin Zou 0001, Jinsheng Xiao, Xin Tian 0006
Pattern Recognit.5
2024 TransFusion-CR: Two-Phase SAR-to-Optical Translation and Deep Feature Fusion for Cloud Removal
abstract
Multispectral image (MSI) is essential for remote sensing (RS) tasks. Unfortunately, the presence of clouds will lead to a reduction in the accuracy and availability of ground information. Synthetic aperture radar (SAR) is unaffected by cloud cover, which can serve as a supplemental information source to alleviate the difficulty of cloud removal in MSI. However, SAR images suffer from the problem of speckle noise and contain information only in a single spatial dimension. To achieve cloud removal effectively by fusing MSI and SAR, we propose the TransFusion-CR in this article. It is a novel two-phase SAR and MSI translation-fusion cloud removal network and designed to enhance the information gains provided by SAR. We use a convolutional neural network (CNN) with a large kernel and design special modules to process the global and local information so that TransFusion-CR can reconstruct clear spatial details and accurate spectral information without artifacts. In the translation phase, we establish an SAR frequency extraction (SFE) module based on the discrete wavelet transform (DWT) to separate and process different frequency information of the SAR image. This assists in dealing with the high-frequency speckle noise while generating translated cloud-free MSI with clear spatial information and reliable spectral information. Subsequently, in the fusion phase, a specially designed feature fusion module realizes a weighted triple-source image fusion guided by the cloud mask score (CMS) image, which high-quality cloud-free MSI. Finally, a new loss function is proposed to balance the learning focus and achieve joint optimization. Extensive experiments demonstrate that TransFusion-CR outperforms the state-of-the-art by about 1.3 dB in terms of PSNR when the percentage of thick cloud cover is more than 60% and can construct reliable terrestrial information even in the complete absence of surface data.
Rui Liu 0041, Yini Peng, Xin Tian 0006
IEEE Trans. Geosci. Remote. Sens.4
2024 Various Degradation: Dual Cross-Refinement Transformer for Blind Sonar Image Super-Resolution
abstract
Deep learning-based methods have achieved remarkable results in super-resolution (SR) of sonar images. However, most existing methods only consider simple bicubic downsampling degradation, and SR networks suitable for natural images may not be suitable for sonar images. Therefore, they perform poorly on sonar images with unknown degradation parameters in real-world scenarios (i.e.,blindscenario). To address these issues, we propose a dual cross-refinement transformer (DCRT) forblindSR of sonar images. DCRT first constructs a large-scale degradation space based on the sonar image imaging mechanism. More importantly, we randomly sample the task-level training information to make DCRT robust on different SR tasks, thereby enhancing theblindSR capability of the network. Then, DCRT focuses on image features than domain features through spatial-channel self-attention cross-fusion block (S-C-SACFB), so the domain gap between the training and testing data can be reduced. Meanwhile, S-C-SACFB effectively combines inter-attention and high-frequency enhancement residual block to enhance the network’s ability to extract high-frequency features while suppressing speckle noise in sonar images. Finally, DCRT uses global residual connections to generate high-resolution sonar images. A large number of experiments at different SR scale show that DCRT outperforms the state–of–the–art methods in both quantitative and qualitative aspects.
Jiahao Rao, Yini Peng, Jun Chen 0019, Xin Tian 0006
IEEE Trans. Geosci. Remote. Sens.4
2024 Interpretable Model-Driven Deep Network for Hyperspectral, Multispectral, and Panchromatic Image Fusion
abstract
Simultaneously fusing hyperspectral (HS), multispectral (MS), and panchromatic (PAN) images brings a new paradigm to generate a high-resolution HS (HRHS) image. In this study, we propose an interpretable model-driven deep network for HS, MS, and PAN image fusion, called HMPNet. We first propose a new fusion model that utilizes a deep before describing the complicated relationship between the HRHS and PAN images owing to their large resolution difference. Consequently, the difficulty of traditional model-based approaches in designing suitable hand-crafted priors can be alleviated because this deep prior is learned from data. We further solve the optimization problem of this fusion model based on the proximal gradient descent (PGD) algorithm, achieved by a series of iterative steps. By unrolling these iterative steps into several network modules, we finally obtain the HMPNet. Therefore, all parameters besides the deep prior are learned in the deep network, simplifying the selection of optimal parameters in the fusion and achieving a favorable equilibrium between the spatial and spectral qualities. Meanwhile, all modules contained in the HMPNet have explainable physical meanings, which can improve its generalization capability. In the experiment, we exhibit the advantages of the HMPNet over other state-of-the-art methods from the aspects of visual comparison and quantitative analysis, where a series of simulated as well as real datasets are utilized for validation.
Xin Tian 0006, Kun Li 0025, Wei Zhang 0259, Zhongyuan Wang 0001, Jiayi Ma 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Robust and Scalable Gaussian Process Regression and Its Applications
abstract
This paper introduces a robust and scalable Gaussian process regression (GPR) model via variational learning. This enables the application of Gaussian processes to a wide range of real data, which are often large-scale and contaminated by outliers. Towards this end, we employ a mixture likelihood model where outliers are assumed to be sampled from a uniform distribution. We next derive a variational formulation that jointly infers the mode of data, i.e., inlier or outlier, as well as hyperparameters by maximizing a lower bound of the true log marginal likelihood. Compared to previous robust GPR, our formulation approximates the exact posterior distribution. The inducing variable approximation and stochastic variational inference are further introduced to our variational framework, extending our model to large-scale data. We apply our model to two challenging real-world applications, namely feature matching and dense gene expression imputation. Extensive experiments demonstrate the superiority of our model in terms of robustness and speed. Notably, when matching 4k feature points, its inference is completed in milliseconds with almost no false matches. The code is at github.com/YifanLu2000/Robust-Scalable-GPR.
Jiayi Ma 0001, Leyuan Fang, Xin Tian 0006, Junjun Jiang
CVPR4
2023 High-Frequency Transformer Network Based on Window Cross-Attention for Pansharpening
abstract
Inspired by the powerful ability to capture long-distance dependencies in the vision transformer, we propose a novel high-frequency transformer network based on window cross-attention to fuse panchromatic (PAN) and multispectral (MS) images for a high-resolution MS image. To overcome the problem brought by shallow feature extraction in the previous transformer-based fusion network, we combine high-pass filtering and deep feature extraction to explore more texture information. As a result, the obtained relationship between MS and PAN images according to feature similarity is more accurate. In particular, we build the cross-modality correlation by a window cross-attention mechanism at pixel-level between MS and PAN images’ local window. Compared with patch-level, pixel-level helps to preserve fine-grained features. Therefore, more spatial details from a PAN image are transferred to an MS image, leading to a clearer fused MS image with good preservation of spectral information. Experimental results demonstrate that the proposed method outperforms the comparison methods in terms of visual and quantitative qualities.
Chengjie Ke, Hao Liang 0008, Duidui Li, Xin Tian 0006
ICASSP4
2023 Variational Bayesian deep network for blind Poisson denoising
Hao Liang 0008, Rui Liu 0041, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006
Pattern Recognit.5
2023 Multipatch Progressive Pansharpening With Knowledge Distillation
abstract
In this paper, we propose a novel multi-patch and multi-stage pansharpening method with knowledge distillation, termed as PSDNet. Different from existing pansharpening methods that typically input single-size patches to the network and implement pansharpening in an overall stage, we design multi-patch inputs and a multi-stage network for more accurate and finer learning. First, multi-patch inputs allow the network to learn more accurate spatial and spectral information by reducing the number of object types. We employ small patches in the early part to learn accurate local information, as small patches contain fewer object types. Then, the later part exploits large patches to fine-tune it for the overall information. Second, the multi-stage network is designed to reduce the difficulty of the previous single-step pansharpening and progressively generate elaborate results. In addition, instead of the traditional perceptual loss, which hardly relates to the specific task or the designed network, we introduce distillation loss to reinforce the guidance of the ground truth. Extensive experiments are conducted to demonstrate the superior performance of our proposed PSDNet to existing state-of-the-art methods. Our code is available at https://github.com/Meiqi-Gong/PSDNet.
Meiqi Gong, Hao Zhang 0073, Han Xu 0001, Xin Tian 0006, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Coarse-to-Fine Cross-Domain Learning Fusion Network for Pansharpening
abstract
Deep learning (DL) based pansharpening methods have shown great advantages in fusing multispectral (MS) and panchromatic (PAN) images to obtain a high-resolution MS image in remote sensing applications. However, most DL methods have low generalization capability that will cause severe spatial or spectral distortions, especially when a large distribution gap exists between training data from a source domain and testing data from another target domain. To overcome this problem, we propose a coarse-to-fine adaption learning fusion network for pansharpening. We first learn the priori mapping relationships between MS and PAN images in the source domain through the coarse-fusion network, which combines the advantages of UNet and Transformer architectures that helps to explore texture information of different characteristics. To generate a clear fusion result with good preservation of spatial and spectral information in the target domain, the fine-fusion network is further proposed to adjust the spatial and spectral information of the coarse-fusion image in an unsupervised learning manner based on the target-specific knowledge. Therefore, the generalization capability can be effectively improved because the general mapping relationship from the source domain and the specific target knowledge from the target domain are both considered. Experiments on simulated and real datasets are conducted to demonstrate the superiority of our proposed method over other state-of-the-art DL methods in terms of visual quality and quantitative analysis.
Chengjie Ke, Wei Zhang 0259, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006
IEEE Trans. Geosci. Remote. Sens.5
2023 Joint Segmentation and Identification Feature Learning for Occlusion Face Recognition
abstract
The existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification.
Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002
IEEE Trans. Neural Networks Learn. Syst.5
2022 Coherent Point Drift Revisited for Non-rigid Shape Matching and Registration
abstract
In this paper, we explore a new type of extrinsic method to directly align two geometric shapes with point-to-point correspondences in ambient space by recovering a deformation, which allows more continuous and smooth maps to be obtained. Specifically, the classic coherent point drift is revisited and generalizations have been proposed. First, by observing that the deformation model is essentially defined with respect to Euclidean space, we generalize the kernel method to non-Euclidean domains. This generally leads to better results for processing shapes, which are known as two-dimensional manifolds. Second, a generalized probabilistic model is proposed to address the sensibility of coherent point drift method to local optima. Instead of directly optimizing over the objective of coherent point drift, the new model allows to focus on a group of most confident ones, thus improves the robustness of the registration system. Experiments are conducted on multiple public datasets with comparison to state-of-the-art competitors, demonstrating the superiority of our method which is both flexible and efficient to improve the matching accuracy due to our extrinsic alignment objective in ambient space.
Aoxiang Fan, Jiayi Ma 0001, Xin Tian 0006, Xiaoguang Mei
CVPR3
2022 CUFD: An encoder-decoder network for visible and infrared image fusion based on common and unique feature decomposition
Han Xu 0001, Meiqi Gong, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001
Comput. Vis. Image Underst.3
2022 Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006
Inf. Sci.8
2022 Realistic frontal face reconstruction using coupled complementarity of far-near-sighted face images
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Baojin Huang, Zhen Han 0002, Xin Tian 0006
Pattern Recognit.7
2022 D2TNet: A ConvLSTM Network With Dual-Direction Transfer for Pan-Sharpening
abstract
In this article, we propose an efficient convolutional long short-term memory (ConvLSTM) network with dual-direction transfer for pan-sharpening, termed D2TNet. We design a specially structured ConvLSTM network that allows for dual-directional communication, including multiscale information and multilevel information. On the one hand, due to the sensitivity of spatial information to scales and the sensitivity of spectral information to levels, multiscale and multilevel information is extracted to facilitate the fuller use of source images. On the other hand, ConvLSTM is employed to capture the strong dependencies between multiscale information and multilevel information. Besides, we introduce a multiscale loss to enable different scales contributing to each other to generate high-resolution multispectral images that are closer to the ground truth. Extensive experiments, including qualitative evaluation, quantitative evaluation, and efficiency comparison, are implemented to verify that our D2TNet outperforms state-of-the-art methods indeed.
Meiqi Gong, Jiayi Ma 0001, Han Xu 0001, Xin Tian 0006, Xiao-Ping Zhang 0002
IEEE Trans. Geosci. Remote. Sens.4
2022 Variational Pansharpening by Exploiting Cartoon-Texture Similarities
abstract
Pansharpening aims to fuse a multispectral (MS) image with low spatial resolution and a panchromatic (PAN) image with a high-spatial resolution to produce an image with both high spectral and high spatial resolution. In this study, we propose a variational pansharpening method by exploiting cartoon-texture similarities. After decomposition of the PAN image, the cartoon component always contains the global structure information, while the texture component includes the locally patterned information. This enables that the fused high-spatial resolution MS image can preserve the global and local spatial details (e.g., high-order information) well after leveraging the similarities of cartoon and texture components from PAN and MS images. To explore such cartoon-texture similarities, we describe cartoon similarity as gradient sparsity, formulated as a reweighted total variation term. Meanwhile, we use group low-rank constraint for texture similarity that is presented as repetitive texture patterns. By incorporating a data fidelity term for preserving the spectral information on the basis that the down-sampled fused MS image is consistent with the MS image, we further formulate pansharpening as an optimization problem and solve it efficiently using the alternative direction multiplier method. Extensive experiments have been conducted on a series of satellite data sets, and we also carry out a simulated vegetation coverage change experiment to verify the efficiency of the proposed method in remote sensing. The qualitative and quantitative results demonstrate that our method outperforms the state-of-the-art pansharpening methods in terms of both visual effect and objective metrics.
Xin Tian 0006, Yuerong Chen, Changcai Yang, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 VP-Net: An Interpretable Deep Network for Variational Pansharpening
abstract
In this study, we propose an interpretable deep network for variational pansharpening (VP), named VP-Net. Different from traditional priors using linear operators, such as the gradient, we construct a prior based on the similarity between panchromatic (PAN) and high-resolution multispectral (HRMS) images by a nonlinear operator that can be learned through a deep network. Considering the spectral difference of various satellite multispectral (MS) imaging platforms, we specifically seek the aforementioned similarity from the PAN image and the intensity of the HRMS image to reduce the spectral distortion. Based on this prior, we propose a novel VP model by further incorporating a data fidelity term from the low-resolution MS image. Specifically, we build the VP-Net by unrolling the variable splitting method for an optimal solution to this model. Consequently, all modules in VP-Net have clear physical meanings and strong generalization capabilities. Meanwhile, all parameters and the aforementioned nonlinear operator are learned in VP-Net, avoiding the difficulty of selecting optimal handcrafted parameters in traditional methods. Therefore, VP-Net not only achieves an optimal balance between spatial and spectral qualities but also has a strong generalization capability across different types of training and testing data. In the experiment, we first demonstrate the superiority of the proposed method over the current state of the arts in terms of both visual effect and quantitative analysis on different satellite datasets. Moreover, we carry out an normalized difference vegetation index (NDVI) experiment to demonstrate its potential in remote sensing.
Xin Tian 0006, Kun Li 0025, Zhongyuan Wang 0001, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 HyperFusion: A Computational Approach for Hyperspectral, Multispectral, and Panchromatic Image Fusion
abstract
Fusing hyperspectral image (HSI) and multispectral image (MSI) of high spatial resolution is typically utilized to obtain HSIs of high spatial resolution. However, the spatial quality of most existing methods is unsatisfactory due to the limited spatial resolution of an MSI. To further improve the spatial resolution of the fused HSI while keeping the spectral information well, we propose a new computational paradigm, named HyperFusion, which simultaneously fuses HSI, MSI, and panchromatic (PAN) image. To achieve this goal, we first establish two data fidelity terms based on a physical observation that HSI and MSI can be treated as degraded versions of the fused HSI. Consequently, the spatial and spectral information from HSI and MSI can be well preserved. To efficiently transfer the spatial details of PAN into the fused HSI while keeping the spectral information well, we further construct a prior constraint from PAN based on the structural similarity. Meanwhile, we impose another low-rank prior constraint on the coefficient matrix to accurately describe the latent characteristics of the HSI with high spatial resolution. By incorporating the aforementioned data fidelity terms and prior constraints, we finally formulate the objective as an optimization problem and utilize the alternative direction multiplier method to solve it efficiently. Comprehensive experiments on simulated and real datasets are carried out to demonstrate the superiority of HyperFusion over other state of the arts in terms of visual quality and quantitative analysis. We also adopt a simulated experiment of vegetation coverage index analysis to verify the effectiveness of HyperFusion in remote sensing applications.
Xin Tian 0006, Wei Zhang 0259, Yuerong Chen, Zhongyuan Wang 0001, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 DBDnet: A Deep Boosting Strategy for Image Denoising
abstract
In this paper, we propose a new deep network architecture named deep boosting denoising net (DBDnet) for image denoising. It is a residual learning network that can generate a noise map from a noisy observation. In detail, it first generates a coarse noise map via a simple structure, and then updates the noise map gradually via a boosting function. The motivation of our DBDnet stems from the observation that the noise map recovered by any algorithm cannot ideally equal the ground-truth noise map, which typically contains noise. We call this noise NoN,i.e., noise of noise map. Based on this observation, we formulate the denoising as a process of reducing NoN, and the role of DBDnet is to eliminate the NoN from the coarse noise map. In particular, we analyze the process of reducing NoN theoretically, and propose an NoN eliminating module to simulate it accordingly. We evaluate the proposed DBDnet on images polluted by different levels of additive white Gaussian noise and real noise. Experiment results demonstrate that our DBDnet can attain better denoising performance compared with state-of-the-art methods on several kinds of image denoising tasks. In particular, for the Gaussian denoising and real image denoising tasks, the average improvements of the PSNR values brought by our DBDnet are about 0.25 dB and 1.01 dB, respectively. In addition, we find and verify that the deep boosting insight can be easily introduced into the state-of-the-art image denoising network, and promotes its denoising performance. Our code is publicly available athttps://github.com/jiayi-ma/DBDNet.
Jiayi Ma 0001, Chengli Peng, Xin Tian 0006, Junjun Jiang
IEEE Trans. Multim.3
2021 When Face Recognition Meets Occlusion: A New Benchmark
abstract
The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC.
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001
ICASSP7
2021 Omniscient Video Super-Resolution
abstract
Most recent video super-resolution (SR) methods either adopt an iterative manner to deal with low-resolution (LR) frames from a temporally sliding window, or leverage the previously estimated SR output to help reconstruct the current frame recurrently. A few studies try to combine these two structures to form a hybrid framework but have failed to give full play to it. In this paper, we propose an omniscient framework to not only utilize the preceding SR output, but also leverage the SR outputs from the present and future. The omniscient framework is more generic because the iterative, recurrent and hybrid frameworks can be regarded as its special cases. The proposed omniscient framework enables a generator to behave better than its counterparts under other frameworks. Abundant experiments on public datasets show that our method is superior to the state-of-the-art methods in objective metrics, subjective visual effects and complexity.
Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Xin Tian 0006, Jiayi Ma 0001
ICCV6
2021 Variation-Net: Interpretable Variation-Inspired Deep Network for Pansharpening
abstract
In this study, we propose Variation-net, an interpretable variation-inspired deep network for pansharpening, which aims to fuse panchromatic (PAN) and multispectral (MS) images for a high-resolution MS image. We first construct a novel variational pan-sharpening model with clear physical meanings. As the relationship between the PAN and MS images in the real situation is complex and nonlinear, we explore the similarity between PAN and MS images from the sparsity of nonlinear transforms in this variational pansharpening model. As a result, spatial details can be accurately transferred from PAN image to MS image. Furthermore, we build the Variation-net by unrolling the iterative shrinkage-thresholding algorithm to solve the proposed variational pansharpening model. Therefore, all modules in Variation-net have clear physical meanings and are easily observed, leading to good generalization capability. Meanwhile, nonlinear transforms and other parameters in the variational pansharpening model are learned end–to–end. The experiments demonstrate that Variation-net outperforms the state-of-the-art methods from the aspects of visual effect and objective quality analysis.
Kun Li 0025, Wei Zhang 0259, Xin Tian 0006, Jiayi Ma 0001, Huabing Zhou, Zhongyuan Wang 0001
ICME3
2021 Robust Feature Matching for Remote Sensing Image Registration via Linear Adaptive Filtering
abstract
As a fundamental and critical task in feature-based remote sensing image registration, feature matching refers to establishing reliable point correspondences from two images of the same scene. In this article, we propose a simple yet efficient method termed linear adaptive filtering (LAF) for both rigid and nonrigid feature matching of remote sensing images and apply it to the image registration task. Our algorithm starts with establishing putative feature correspondences based on local descriptors and then focuses on removing outliers using geometrical consistency priori together with filtering and denoising theory. Specifically, we first grid the correspondence space into several nonoverlapping cells and calculate a typical motion vector for each one. Subsequently, we remove false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a Gaussian kernel convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. In addition, an adaptive parameter setting strategy and posterior probability estimation based on the expectation-maximization algorithm enhance the robustness of our method to different data. Most importantly, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Consequently, some sparse point-based tasks may inspire from our method when they are achieved by deep learning techniques. Extensive feature matching and image registration experiments on several remote sensing data sets demonstrate the superiority of our approach over the state of the art.
Xingyu Jiang 0005, Jiayi Ma 0001, Aoxiang Fan, Haiping Xu, Geng Lin, Tao Lu 0001, Xin Tian 0006
IEEE Trans. Geosci. Remote. Sens.7
2021 FusionNDVI: A Computational Fusion Approach for High-Resolution Normalized Difference Vegetation Index
abstract
Normalized difference vegetation index (NDVI), derived from the near-infrared and red bands of a multispectral (MS) image, has been widely used in remote sensing. To obtain a high-resolution (HR) NDVI, existing attempts typically first generate an HR-MS image using pansharpening and then calculate the HR NDVI accordingly. However, some inaccurate spatial information will be simultaneously introduced into NDVIs, influencing their spatial quality seriously. To overcome this challenge, we investigate a computational fusion approach from a novel perspective for HR NDVI in this study. Rather than pansharpening an HR-MS image, we define an HR vegetation index calculated based on an available HR panchromatic image and an estimated HR red band (VIPR) and fuse the low-resolution (LR) NDVI and HR VIPR directly to acquire an HR NDVI. In particular, we adopt a nonlocal gradient sparsity constraint to force a similar nonlocal spatial structure in the fused NDVI and VIPR, where the VIPR is dynamically updated by adding a constraint to reconstruct the HR red band. We further integrate a data fidelity term to constrain the relationship between the fused NDVI and its LR version, and an efficient strategy based on the alternative direction multiplier method is developed to solve the nonconvex optimization problem. The extensive experimental results demonstrate that the proposed method achieves superior fusion performance over the state of the art, exhibiting its wide application aspect in remote sensing.
Xin Tian 0006, Mengliang Zhang, Changcai Yang, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Cartoon-Texture Decomposition-Based Variational Pansharpening
abstract
Pansharpening is widely used to increase the spatial resolution of a multispectral (MS) image by fusing with a panchromatic (PAN) image that has high-spatial resolution and the same scene. In this paper, the similarities of MS and PAN images in cartoon-texture space are exploited. The cartoon and texture components of an image always contain the global structure information and the locally-patterned information, respectively. Therefore, the global and local spatial details (i.e., high-order information) could be preserved well in the fused high-spatial resolution MS image after leveraging the similarities of these images. Pansharpening is formulated as the optimization problem with respect to the cartoon-texture similarities between the MS and the PAN images in this work based on the aforementioned observation. Specifically, cartoon similarity is determined through gradient sparsity and formulated as a total variation term, whereas texture similarity is described according to the low-rank property. The alternative direction multiplier method is used to solve the optimization problem. In the experiment, the Gaofen-1 satellite dataset is used to compare the proposed method with other classical pansharpening methods. Experimental results demonstrate that our method outperforms the comparison methods in terms of visual and quantitative qualities.
Yuerong Chen, Mengliang Zhang, Zhongyuan Wang 0001, Xin Tian 0006
ICASSP5
2020 Attention-Guided Deraining Network Via Stage-Wise Learning
abstract
Due to diverse rain shapes, directions, densities as well as different distances to cameras, rain streaks in the air are interweaved and overlapped. However, most existing deraining methods are inherently oblivious this phenomenon and tend to learn a single rain streak layer to simulate this complex distribution, consequently failing to restore high-quality rain-free images. To solve this problem, along with the stage-wise learning, we propose a novel attention-guided deraining network (ADN) for rain streak removal. Specially, we decompose the rain streaks into multiple rain streak layers, and individually model them along the stages of the network to match the increasing abstracts. Moreover, the attention mechanism is utilized to guide the fusion of these rain streak layers by handling the overlaps between them. Extensive experiments on several benchmark datasets and real-world scenarios show substantial improvements both on quantitative indicators and visual effects over the current top-performing methods.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Yuhong Yang 0001, Xin Tian 0006, Junjun Jiang
ICASSP6
2020 Fusionndvi: A Novel Fusion Method for NDVI in Remote Sensing
abstract
Normalized difference vegetation index (NDVI) is widely utilized to examine vegetation coverage and estimate crop yield. To obtain a high-resolution (HR) NDVI, fusion techniques, which first generates a HR multispectral (MS) image by fusing a low-resolution (LR) MS image and a HR panchromatic image, and then calculates the HR NDVI based on the fused HR MS image, are utilized in previous studies. A HR vegetation index calculated on the basis of HR panchromatic image could provide HR spatial resolution, and this vegetation index has a spatial structure that is similar to that of NDVI. Therefore, this similarity is investigated to construct a novel method called FusionNDVI to improve the fusion performance in this study. The fusion problem is formulated to minimize a least square fitting error term and a nonlocal gradient sparsity regularization term. The fitting term is used to limit the difference between the fused HR NDVI and the LR NDVI, whereas the regularizer enforces a similar nonlocal spatial structure in the fused NDVI and the HR vegetation index. An efficient solving algorithm based on the augmented Lagrangian method of multipliers is derived. The superiority of the proposed FusionNDVI method over the state-of-the-art ones is verified via simulations.
Mengliang Zhang, Ziping Zhao 0002, Yuerong Chen, Zhongyuan Wang 0001, Xin Tian 0006
ICASSP5
2020 A Generative Adversarial Network For Medical Image Fusion
abstract
In this paper, a novel end-to-end model for fusing medical images characterizing structural information, i.e., IS, and images characterizing functional information, i.e., IF, of different resolutions is proposed, which is achieved by using a conditional generative adversarial network with multiple generators and multiple discriminators (MGMDcGAN). In the first cGAN, a real-like fused image is generated by a generator, simultaneously fooling two discriminators. While the discriminators are to distinguish the fused image from source images. Besides, to prevent the functional information from being weakened in the final fused image when enhancing the dense structure information, we employ the second cGAN with a mask calculated. Meanwhile, the structural information in ISand the functional information in IFthe final fused image can be concurrently kept. Furthermore, our MGMDcGAN is a unified method, which is applicable to different kinds of medical image fusion, including MRI-PET, MRISPECT, and CT-SPECT. Extensive experiments on publicly available datasets substantiate the superiority of our MGMDcGAN over the current state-of-the-art.
Zhuliang Le, Jun Huang 0008, Fan Fan 0001, Xin Tian 0006, Jiayi Ma 0001
ICIP4
2020 Robust CBCT Reconstruction Based On Low-Rank Tensor Decomposition And Total Variation Regularization
abstract
Cone-beam computerized tomography (CBCT) has been widely used in numerous clinical applications. To reduce the effects of X-ray on patients, a low radiation dose is always recommended in CBCT. However, noise will seriously degrade image quality under a low dose condition because the intensity of the signal is relatively low. In this study, we propose to use the Huber loss function as a data fidelity term in CBCT reconstruction, making the reconstruction robust to impulse noise under low radiation dose condition. Furthermore, a low-rank tensor property is adopted as the prior term. Such property is helpful in recovering the missing structure information caused by impulse noise. The proposed CBCT reconstruction model is formulated by further integrating a 3D total variation term for reducing Gaussian noise. An alternative direction multiplier method is adopted to solve the optimization problem. Experiments on simulated and real data show that the proposed model outperforms existing CBCT reconstruction algorithms.
Xin Tian 0006, Zhongyuan Wang 0001
ICIP1
2020 Masked Face Recognition with Identification Association
abstract
In the crime scene, criminals often consciously conceal their facial identity through face-masked disguise, which poses a huge challenge to identity recognition. Existing disguised face recognition techniques aiming for light even slight occlusions are completely invalid for face-masked identification. To this end, this paper proposes a masked face recognition method based on person re-identification association, which converts the masked face recognition problem into an association uncovering problem between the masked face and the appearing faces of the same person. Based on the characteristics that person re-identification technique does not rely solely on facial information, it first takes advantages of re-identification to establish the association between face-masked pedestrians and face-unveiled pedestrians. It further provides an effective face image quality assessment to select the most identifiable faces for subsequent recognition from a variety of appearing candidate faces. Finally, the selected high-quality recognizable faces are used to replace masked faces for identification. The comparison experiments with the existing disguise face recognition methods show its superiority in terms of accuracy.
Zhongyuan Wang 0001, Zheng He 0001, Nanxi Wang, Xin Tian 0006, Tao Lu 0001
ICTAI5
2020 Multiscale Infrared and Visible Image Fusion Based on Phase Congruency and Saliency
abstract
In this paper, in order to enhance the infrared target in infrared image and retain the edge and detail information in visible image, we propose a multi-scale decomposition fusion method based on phase congruency and saliency. In this method, the Laplacian pyramid is first used to decompose the source image into detail layers and base layers. Secondly, we use a method based on phase congruency for the fusion of detail layers. Thirdly, for the base layer, we decompose it into saliency map and residual map. The “max absolute” rule and “averag” rule are adopted for the fusion of saliency map and residual map, then the fused saliency map and residual map are added to attain the fused base image. Finally, we use the inverse transform of Laplacian pyramid to reconstruct the fused image. The experimental results show that the proposed method have better fusion effect than other methods. What's outstanding is that the infrared targets in the fused image are enhanced and abundant edges are preserved.
Jun Chen 0019, Kangle Wu, Linbo Luo 0002, Xin Tian 0006
IGARSS6
2020 Sparse Representation-Based Image Fusion for Multi-Source NDVI Change Detection
abstract
The normalized differential vegetation index (NDVI) is a useful index for change detection in remote sensing vegetation analysis. Multi-source NDVI change detection, which utilizes the NDVI information at different time from multiple satellites, can solve the problem of long-revisiting periods for a single source (satellite). However, the spatial resolution of NDVI calculated from the multispectral images of different satellites is different. A sparse representation-based image fusion method is proposed to improve the spatial resolution of NDVI. First, a high spatial-resolution vegetation index (HRVI) is utilized. The proposed method is based on the assumption that NDVI and HRVI with different resolutions will have the same sparse coefficients under some specific dictionaries. In the experiment, the proposed method is compared with several state-of-the-art methods to demonstrate its efficiency. Furthermore, its application in multi-source NDVI change detection verified by datasets from GF-1 and GF-2 satellites.
Mengliang Zhang, Yuerong Chen, Xin Tian 0006
IGARSS4
2020 Image fusion employing adaptive spectral-spatial gradient sparse regularization in UAV remote sensing
Mengliang Zhang, Xin Tian 0006
Signal Process.4
2020 A Variational Pansharpening Method Based on Gradient Sparse Representation
abstract
By exploiting the gradient similarity between multispectral (MS) and panchromatic (PAN) images, a variational pansharpening method based on gradient sparse representation is proposed, based on the observation that the gradients of corresponding MS and PAN images with different resolutions have the similar sparse coefficients under certain specific dictionaries. By adding a data fidelity term to preserve the spectral information, an optimization model is constructed as a minimization problem of an energy function. The problem can be solved by the gradient descent method efficiently. Experiments on different satellite data reveal that the proposed method outperforms the state-of-the-art methods in terms of visual effect and objective quality analysis.
Xin Tian 0006, Yuerong Chen, Changcai Yang, Jiayi Ma 0001
IEEE Signal Process. Lett.1
2019 Spectral-Spatial Classification of Hyperspectral Image based on a Joint Attention Network
abstract
Deep neural networks have been successfully applied to extracting deep features for many hyperspectral tasks. Attention mechanism has been widely used in computer vision, inspired by this, we have designed a joint attention network for spectral-spatial classification of hyperspectral image. In our method, recurrent neural network (RNN) with attention can learn inner spectral correlations within a continuous spectrum, convolutional neural network (CNN) with attention is designed to focus on saliency features and spatial dependency in the neighbor regions. Experimental results demonstrate that our method can fully utilize spectral and spatial information to obtain competitive performance.
Erting Pan, Yong Ma 0001, Xiaoguang Mei, Xiaobing Dai, Fan Fan 0001, Xin Tian 0006, Jiayi Ma 0001
IGARSS6
2011 Efficient Multi-Input/Multi-Output VLSI Architecture for Two-Dimensional Lifting-Based Discrete Wavelet Transform
abstract
This brief paper proposes an efficient multi-input/multi-output VLSI architecture (MIMOA) for two-dimensional lifting-based discrete wavelet transform (DWT). The novelty is the simplicity and generality to construct the MIMOA, which is a high-speed architecture with computing time as low as N2/M for an N × N image with controlled increase of hardware cost. M is the throughput rate.
Xin Tian 0006, Yihua Tan, Jin-Wen Tian
IEEE Trans. Computers1
2010 Underwater Object Detection Based on Gravity Gradient
abstract
A novel method of underwater object detection based on gravity gradient is presented, which can be used on autonomous underwater vehicles (AUVs) to detect abnormal objects underwater. Gravity gradient anomalies of partial area, which are caused by the object, can be measured by a gravity gradiometer on an AUV. Then, anomalies can be inversed with a gravity gradient inversion algorithm, so the mass and barycenter of an object can be estimated. Simulation results show that approximate information of an object can be provided by the proposed method.
Xin Tian 0006, Jie Ma 0003, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.2