Jian Kang 0005

dblp:56/6072-5 · DBLP profile ↗
← Back
45ranked-venue papers
19as first author
34since 2021 · last 2025
0000-0001-6284-3044ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 17 first-author · 27 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Multipass InSAR Phase Linking for Non-Gaussian Scatterers
abstract
Multipass synthetic aperture radar interferometry (InSAR) is a powerful tool for deformation monitoring, leveraging multiple SAR images to extract surface displacement information over time. However, its application in complex environments, such as scatterers, exhibits a mixture of different scattering behaviors within a single resolution cell or in close proximity, remains challenging due to the errors in phase linking (PL) caused by deviations from the standard statistical model, i.e., the complex circular Gaussian (CCG) distribution. This article proposes an advanced approach for addressing such issues in InSAR time-series analysis. We introduce a novel pixel selection algorithm, HetPS, based on the$\mathcal {KM}$distribution, which models both spatial and temporal heterogeneity of scatterers. This method enhances the robustness of the pixel selection process. Additionally, we present a new PL algorithm, HetPL, which builds on the pixel selection strategy and conducts the phase history reconstruction for those heterogeneous scatterers following the$\mathcal {KM}$distribution. The methods’ effectiveness is demonstrated using both synthetic and real TerraSAR-X data. Results show that the proposed methods significantly improve phase history stability and deformation estimation accuracy, outperforming traditional techniques in regions with strong heterogeneity.
Yusong Bai, Shengkai Liu, Chunping Shi, Jian Kang 0005
IEEE Trans. Geosci. Remote. Sens.7
2025 Align and Complete Samples in Remote Sensing Fine-Grained Rigid Object Detection
abstract
Currently, the remote sensing fine-grained rigid object detectors mainly face two challenges: fuzzy localization and inaccurate classification (including misclassifying and multi-classifying). Firstly, mainstream detectors based on the “dense prediction” paradigm suffer from a misalignment between anchor points (APs) and ground truths (GTs). Commonly, they adopt multi-stage regression as a solution, which brings an ambiguous sample definition problem and redundant computational costs. Secondly, two factors seriously affect the classification. They are the insufficient sample learning caused by the long-tail distribution of categories, and the difficulty in extracting discriminative features caused by slight inter-class variance. To address the issues above, we propose an efficient aligning and completing detector (ACDet) based on a single-stage structure. Firstly, the Adaptive Anchor Alignment Mechanism decouples the centripetal sampling bias in the categorical features and leverages it to learn the APs aligned with GTs. It is plug-and-play, requiring no additional supervision annotations. Secondly, a novel Online Tail-sample Supplementation algorithm is proposed. It dynamically maintains class balance during training and can be easily added as post-processing behind existing sample assignment strategies. Thirdly, an Adaptive Group Perceptron is designed to effectively enrich the diversity of features and enhance the model’s ability to extract discriminative features. Experiments on four public datasets demonstrate that the proposed ACDet achieves the state-of-the-art (SOTA) level, even surpassing competitive multi-stage detectors, with fewer computing resources and a faster inference speed. The code will be public after the paper is published.
Zicong Zhu, Jian Kang 0005, Wenhui Diao, Bing Wang 0015, Jingen Ni
IEEE Trans. Geosci. Remote. Sens.2
2024 Shadow detection using a cross-attentional dual-decoder network with self-supervised image reconstruction features
abstract
Shadow detection is a challenging problem in computer vision due to the high variability in lighting conditions, object shapes, and scene layouts. Despite the positive results achieved by some existing technologies, the problem becomes particularly challenging with complex and heterogeneous images where shadow-casting objects coexist and shadows can have different depths, scales, and morphologies. As a result, more advanced and accurate solutions are still needed to deal with this type of complexities. To address these challenges, this paper proposes a novel deep learning model, called the Cross-Attentional Dual Decoder Network (CADDN), to improve shadow detection by using fine-grained image reconstruction features. Unlike other existing methods, the CADDN uses an innovative encoder-decoder architecture with two decoder segments that work together to reconstruct the input images and their corresponding shadow masks. In this way, the features used to reconstruct the original input image can be used to support the shadow detection process itself. The proposed model also incorporates a cross-attention mechanism to weight the most relevant features for detecting shadows and skip connections with noise to improve the quality of the transferred features. The experimental results, including several benchmark image datasets and state-of-the-art detection methods, demonstrate the suitability of the presented approach for detecting shadows in computer vision applications.
Rubén Fernández-Beltran, Angélica Guzmán-Ponce, Rafael Fernandez, Jian Kang 0005, Ginés García-Mateos
Image Vis. Comput.4
2024 LGNet: Local and global point dependency network for 3D object detection
Yan Huang 0018, Jian Kang 0005, Hui Zhang 0071, Wei Hong 0002
Pattern Recognit.4
2024 TADCG: A Novel Gridless Tomographic SAR Imaging Approach Based on the Alternate Descent Conditional Gradient Algorithm With Robustness and Efficiency
abstract
Sparse signal processing techniques, such as compressed sensing (CS), are commonly used in tomographic synthetic aperture radar (TomoSAR) imaging due to the sparsity present in the elevation direction. However, classical CS methods, such as$\ell _{1}$-norm regularization and orthogonal matching pursuit (OMP), suffer from the off-grid effect. Specifically, they discretize the elevation axis into multiple grids and assume that scatterers are located precisely on the grids, leading to reconstruction results that deviate from the true heights of the scatterers. Although gridless CS methods, such as atomic norm minimization (ANM), have achieved gridless reconstruction in specific scenarios, they face challenges such as the requirement of uniformly distributed baselines and large computational cost. In this article, we propose a gridless CS method based on the alternate descent conditional gradient (ADCG) kernel for TomoSAR inversion and compare it with ANM, iterative soft thresholding (IST), and OMP. We show through numerical simulations and experimental results that our proposed method is applicable not only to scenarios with uniformly distributed baselines but also to scenarios with nonuniformly distributed baselines, and it resolves the off-grid effect in both cases. Finally, we demonstrate the effectiveness of our proposed method by implementing gridless reconstruction using an actual dataset of urban buildings in Yuncheng.
Mingxiao Shao, Zhe Zhang 0026, Jie Li 0065, Jian Kang 0005, Bingchen Zhang
IEEE Trans. Geosci. Remote. Sens.4
2024 Cross-Modal Hashing With Feature Semi-Interaction and Semantic Ranking for Remote Sensing Ship Image Retrieval
abstract
Cross-modal hashing plays a pivotal role in large-scale remote sensing (RS) ship image retrieval. RS ship images often exhibit similar overall appearance with subtle differences. Existing hashing methods typically employ feature non-interaction strategies to generate common hash codes, which may not effectively capture the correlations between cross-modal ship images to reduce intermodality discrepancies. To address this issue, we propose a novel cross-modal hashing approach based on feature semi-interaction and semantic ranking (FSISR) for RS ship image retrieval. Our FSISR approach not only captures intricate correlations between different ship image modalities, but also enables the construction of hash tables for large-scale retrieval. FSISR comprises a feature semi-interaction module and a semantic ranking objective function. The semi-interaction module utilizes clustering centers from one modality to learn the correlations between two modalities and generate robust shared representations. The objective function optimizes these representations in a common Hamming space, consisting of a shared semantic alignment loss and a margin-free ranking loss. The alignment loss employs a shared semantic layer to preserve label-level similarity, while the ranking loss incorporates hard examples to establish a margin-free loss that captures similarity ranking relationships. We evaluate the performance of our method on benchmark datasets and demonstrate its effectiveness for cross-modal RS ship image retrieval.https://github.com/sunyuxi/FSISR.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Sebastian Hafner, Xutao Li 0003, Chuyao Luo, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2024 Extracting Building Footprints in SAR Images via Distilling Boundary Information From Optical Images
abstract
Buildings represent pivotal entities in remote sensing imagery for various applications like urban planning and land resource management. Predominantly, methods for building footprint extraction in the literature focus on optical imagery with visual attributes that faithfully mirror the physical world. Nevertheless, the acquisition of high-quality optical images presents formidable challenges due to the susceptibility to illumination conditions and scene visibility. In contrast, synthetic aperture radar (SAR) images can be acquired in all-weather and all-time situations, unburdened by the aforementioned constraints. However, the coherent imaging mechanism engenders intricate complexities for building footprint extraction SAR images. To address this issue, this paper introduces the Boundary Information Distillation Network (BIDNet) to improve the prediction accuracy in SAR images by distilling knowledge from optical images. The proposed approach adopts a teacher-student framework, featuring two customized components: the Explicit Distillation Module (EDM) and the Latent Distillation Module (LDM). Different from the conventional practice of directly aligning feature maps, BIDNet focuses on leveraging the more conspicuous boundary information in optical images. The EDM operates by simultaneously yielding a boundary map to emphasize the boundary area and assimilating the explicit low-level features of two modalities. The LDM represents the structural attributes within the high-level latent feature space and aligns the representations of the two modalities. Within this module, intrinsic self-correlations among features originating from boundary regions are encoded, and so are the cross-correlations established between features from boundary regions and alternative areas. The two modules also serve as the conduit for knowledge distillation from the teacher network to the student network, enabling the utilization of optical imagery for enhancing the building footprint extraction in SAR imagery. Extensive experiments demonstrate that our BIDNet achieves state-of-the-art performance on the Multi-Sensor All Weather Mapping (MSAW) dataset, outperforming the strong baseline by 4.3-7.2 points in f1-score and 4.9-8.0 points in IoU. The source code and trained models will be publicly available.
Lanxin Zeng, Wen Yang 0001, Jian Kang 0005, Huai Yu, Mihai Datcu, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.4
2024 A Robust Nonlocal Tensor Decomposition Method for InSAR Phase Denoising
abstract
Interferometric synthetic aperture radar (InSAR) images are severely corrupted by noise in both magnitude and phase. It is significantly essential to recover the true interferometric phase during InSAR signal processing. Usually, traditional phase denoising methods are to find homogeneous samples for filtering with the need to balance noise reduction and phase preservation, which may be a problem in dealing with topography scenes. In this article, a novel algorithm of robust nonlocal tensor decomposition (RNLTD) for InSAR phase denoising is proposed. In the scheme, a nonlocal tensor (NLT) model of the interferogram is constructed by selecting and stacking similar image patches in a nonlocal region. Benefiting from the simultaneous use of nonlocal and tensor tools, superior low-rank properties of this NLT can be acquired, which is also confirmed by numerical analysis. Then, a robust tensor decomposition algorithm is proposed to formulate the low-rank recovery of the interferogram and constrain the sparse outliers for noise reduction. Next, an alternating direction method of multipliers (ADMM) solution is applied to robustly and accurately restore the noise-reduced interferometric phase. As a result, the proposed RNLTD algorithm takes advantage of effectively capturing the phase structure in a high-dimension manner, which is helpful in phase preservation with the achievement of excellent noise reduction. Lastly, the experimental analysis using one set of simulated and two sets of measured InSAR data is performed to show the promising performance of the proposed algorithm.
Gang Xu 0002, Fangzheng Xu, Xiang-Gen Xia 0001, Hanwen Yu, Honghao Zhou, Jian Kang 0005, Mengdao Xing
IEEE Trans. Geosci. Remote. Sens.6
2024 SIRS: Multitask Joint Learning for Remote Sensing Foreground-Entity Image-Text Retrieval
abstract
The essence of improving the effect of cross-modal image-text retrieval (CIR) lies in the finer-grained modeling of homogeneous features between modalities. However, in remote sensing (RS) scenarios, existing methods usually apply the image-sentence granular feature alignment paradigm, bringing significant difficulties to the fine-grained representation of homogeneous features between modalities. Besides, more complex background noise and extreme scale ranges of foreground targets are hard to distinguish, causing the feature mottle problem. To address the above issues, we propose a novel Semantic-guided Image-text Retrieval framework with Segmentation (SIRS). It is a multi-task joint learning framework for plug-and-play and end-to-end training RS CIR models efficiently, including Semantic-guided Spatial Attention (SSA) and Adaptive Multi-scale Weighting (AMW) modules. First, SSA introduces a background reconstruction branch based on noise perception and a semantic segmentation branch based on pixel-level prediction. It explores a joint learning strategy that concisely filters background noise and refines foreground features considerably. Secondly, AMW performs multi-scale weighting on various layers of feature map output by the encoder, effectively improving the learning efficiency of foreground targets at different scales. It is worth mentioning that SIRS outputs combination results with image and segmentation mask, which is not available in other methods. Based on the RSITMD dataset, we complete the semantic segmentation annotation RSITMD-SS to verify the performance of the proposed method. Sufficient and complete experiments verify the effectiveness of the proposed method. With SIRS, the mainstream SVP and CLIP-based methods improve about 7 mR and derive segmentation prediction with acceptable computational cost optionally. The code and associated dataset will be available on https://github.com/StarBurstStream0/SIRS.
Zicong Zhu, Jian Kang 0005, Wenhui Diao, Yingchao Feng, Junxi Li, Jingen Ni
IEEE Trans. Geosci. Remote. Sens.2
2023 W-NetPan: Double-U network for inter-sensor self-supervised pan-sharpening
abstract
The increasing availability of remote sensing data allows dealing with spatial-spectral limitations by means of pan-sharpening methods. However, fusing inter-sensor data poses important challenges, in terms of resolution differences, sensor-dependent deformations and ground-truth data availability, that demand more accurate pan-sharpening solutions. In response, this paper proposes a novel deep learning-based pan-sharpening model which is termed as the double-U network for self-supervised pan-sharpening (W-NetPan). In more details, the proposed architecture adopts an innovative W-shape that integrates two U-Net segments which sequentially work for spatially matching and fusing inter-sensor multi-modal data. In this way, a synergic effect is produced where the first segment resolves inter-sensor deviations while stimulating the second one to achieve a more accurate data fusion. Additionally, a joint loss formulation is proposed for effectively training the proposed model without external data supervision. The experimental comparison, conducted over four coupled Sentinel-2 and Sentinel-3 datasets, reveals the advantages of W-NetPan with respect to several of the most important state-of-the-art pan-sharpening methods available in the literature. The codes related to this paper will be available at https://github.com/rufernan/WNetPan.
Rubén Fernández-Beltran, Rafael Fernandez, Jian Kang 0005, Filiberto Pla
Neurocomputing3
2023 CVGG-Net: Ship Recognition for SAR Images Based on Complex-Valued Convolutional Neural Network
abstract
Ship target recognition is a vital task in synthetic aperture radar (SAR) imaging applications. Although convolutional neural networks have been successfully employed for SAR image target recognition, surpassing traditional algorithms, most existing research concentrates on the amplitude domain and neglects the essential phase information. Furthermore, several complex-valued neural networks utilize average pooling to achieve full complex values, resulting in suboptimal performance. To address these concerns, this paper introduces a Complex-valued Convolutional Neural Network (CVGG-Net) specifically designed for SAR image ship recognition. CVGG-Net effectively leverages both the amplitude and phase information in complex-valued SAR data. Additionally, this study examines the impact of various widely-used complex activation functions on network performance and presents a novel complex max-pooling method, called Complex Area Max-Pooling. Experimental results from two measured SAR datasets demonstrate that the proposed algorithm outperforms conventional real-valued convolutional neural networks. The proposed framework is validated on several SAR datasets.
Dandan Zhao 0001, Zhe Zhang 0026, Dongdong Lu, Jian Kang 0005, Xiaolan Qiu, Yirong Wu
IEEE Geosci. Remote. Sens. Lett.4
2023 SAR time series despeckling via nonlocal matrix decomposition in logarithm domain
Jian Kang 0005, Teng-Yu Ji, Zhe Zhang 0026, Rubén Fernández-Beltran
Signal Process.1
2023 LaMIE: Large-Dimensional Multipass InSAR Phase Estimation for Distributed Scatterers
abstract
State-of-the-art phase linking (PL) methods for distributed scatterer interferometry (DSI) retrieve consistent phase histories from the sample coherence matrix or the one whose magnitudes are calibrated. To unify them, we first propose a framework consisting of sample coherence matrix estimation and Kullback–Leibler (KL) divergence minimization. Within such framework, we observe that the current state-of-the-art PL methods mainly focus on calibrating the magnitudes of sample coherence matrix while ignoring the errors caused by it exploited in the complex domain, especially when the PL problem is large-dimensional. In this paper, “large-dimensional” refers to the case where the temporal dimensionNof coherence matrices and the numberPof statistically homogeneous pixels (SHP) are at the same level. To solve this issue, we further propose a PL method, termed LaMIE, which is aimed at precise phase history retrieval from large-dimensional coherence matrices for DSI. It includes two steps: 1) sample coherence matrix shrinkage to calibrate the matrix in complex and real domains and 2) phase history retrieval via the flat coherence metric. Both simulated and real data experiments validate the effectiveness of the proposed method by comparing it with other PL methods. Through LaMIE, the densities of the selected points with stable phases can be significantly improved, and the displacement velocities for more regions can be obtained than with state-of-the-art methods.
Yusong Bai, Jian Kang 0005, Anping Zhang, Zhe Zhang 0026, Naoto Yokoya
IEEE Trans. Geosci. Remote. Sens.2
2023 Cross-View Object Geo-Localization in a Local Region With Satellite Imagery
abstract
Cross-view geo-localization is a critical task in various applications, such as smart city management and disaster monitoring. Current methods typically divide a satellite image into patches and use these patches to identify the geographic location of a query image. However, these methods can only provide the location of an image rather than the location of a specific object of interest. This makes it difficult to link these methods to GeoDatabases to obtain detailed information about a target object, such as its name and construction time. To overcome this limitation, we propose a novel problem of cross-view object geo-localization in a local region with high-resolution satellite images. This problem includes two main challenges: accurately identifying the location of an object and distinguishing the target object from others in satellite images. To address these challenges, we present a new Detection-based Geo-localization method called DetGeo, which consists of an object detection-based framework with a two-branch encoder and a query-aware cross-view fusion module. DetGeo uses cross-view images as input to the detector to provide object-level geo-localization. The fusion module employs cross-view spatial attention to focus on relevant areas of target objects during cross-view feature fusion. To evaluate our method, we constructed a new Cross-View Object Geo-Localization dataset called CVOGL, which comprises ground-view or drone-view images as query images and satellite-view images as geo-tagged reference images. Comprehensive experiments are conducted to demonstrate the effectiveness of our method on CVOGL. https://github.com/sunyuxi/DetGeo.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Shanshan Feng 0001, Xutao Li 0003, Chuyao Luo, Puzhao Zhang, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2023 Consistency Center-Based Deep Cross-Modal Hashing for Multisource Remote Sensing Image Retrieval
abstract
Cross-modal hashing aims to retrieve similar images from large-scale Earth Observation (EO) data archives, which typically contain multiple satellite sources of remote sensing (RS) images. However, existing cross-modal hashing methods primarily focus on dual-source RS images and often face two main limitations when retrieving multi-source RS images. Firstly, these methods exhibit significant redundancy as they require handling all possible dual-source combinations in multi-source RS images. Secondly, they often rely on pairwise or triplet image sources to construct objective functions, which are not significantly effective in reducing the discrepancies among multiple RS image sources. To address these limitations, we propose a novel Consistency Center-based deep cross-modal Hashing method called C2Hash for multi-source RS image retrieval. Our C2Hash employs a multi-branch hashing network to directly encode multi-source RS images into unified hash codes, thereby offering higher processing efficiency. Furthermore, C2Hash introduces consistency centers to construct a novel objective function. The consistency center represents the shared semantic features among similar multi-source RS images and is generated by a label hashing network. The objective function encourages similar multi-source RS images to approach the same consistency center to align all image sources in a unified Hamming space. Our method can effectively reduce the discrepancies across multiple image sources and generate unified hash codes. To evaluate its effectiveness, we construct a new Multi-Source RS Image dataset called MSRSI, comprising five different types of image sources. We conduct comprehensive experiments to demonstrate the superior performance of our method on the MSRSI dataset. https://github.com/sunyuxi/C2Hash.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Xutao Li 0003, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2022 Visual Grounding in Remote Sensing Images
abstract
Ground object retrieval from a large-scale remote sensing image is very important for lots of applications. We present a novel problem of visual grounding in remote sensing images. Visual grounding aims to locate the particular objects (in the form of the bounding box or segmentation mask) in an image by a natural language expression. The task already exists in the computer vision community. However, existing benchmark datasets and methods mainly focus on natural images rather than remote sensing images. Compared with natural images, remote sensing images contain large-scale scenes and the geographical spatial information of ground objects (e.g., longitude, latitude). The existing method cannot deal with these challenges. In this paper, we collect a new visual grounding dataset, called RSVG, and design a new method, namely GeoVG. In particular, the proposed method consists of a language encoder, image encoder, and fusion module. The language encoder is used to learn numerical geospatial relations and represent a complex expression as a geospatial relation graph. The image encoder is applied to learn large-scale remote sensing scenes with adaptive region attention. The fusion module is used to fuse the text and image feature for visual grounding. We evaluate the proposed method by comparing it to the state-of-the-art methods on RSVG. Experiments show that our method outperforms the previous methods on the proposed datasets. https://sunyuxi.github.io/publication/GeoVG
Yuxi Sun 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Jian Kang 0005
ACM Multimedia5
2022 Unsupervised deep hashing through learning soft pseudo label for remote sensing image retrieval
Yuxi Sun 0002, Yunming Ye, Xutao Li 0003, Shanshan Feng 0001, Bowen Zhang 0005, Jian Kang 0005, Kuai Dai
Knowl. Based Syst.6
2022 Toward Tightness of Scalable Neighborhood Component Analysis for Remote-Sensing Image Characterization
abstract
Deep metric learning methods have recently drawn significant attention in the field of remote sensing (RS), owing to their prominent capabilities for modeling relations among RS images based on their semantic contents. In the context of scene classification and large-scale image retrieval, one of the most prominent deep metric learning methods is the scalable neighborhood component analysis (SNCA), which has demonstrated excellent performance on the locality neighborhood structure in the metric space. However, the standard SNCA has important constraints on separating the hard positive and other negative images in the metric space, and this may become a major limitation when dealing with the large-scale variance problem inherent to RS data. To address this issue, we propose a novel deep metric learning formulation that introduces a new margin parameter to enforce the compactness of the within-class feature embeddings. Based on this innovative scheme, we propose two novel loss functions: 1) T-SNCA-c, where the parameter is based on the cosine similarity, and 2) T-SNCA-a, where the parameter is based on the angular distance. Besides, we exploit memory bank optimization to further enhance the semantic diversity during training. Our experimental results, conducted using three downstream applications ($K$-NN classification, clustering, and image retrieval) and two large-scale RS benchmark datasets, demonstrate that the proposed approach can achieve superior performance when compared to current state-of-the-art deep metric learning methods. The codes of this work will be made available online (https://github.com/jiankang1991/GRSL_TSNCA).
Jian Kang 0005, Rubén Fernández-Beltran, Sicong Liu 0001, Antonio Plaza
IEEE Geosci. Remote. Sens. Lett.1
2022 Time-Resolved Sentinel-3 Vegetation Indices Via Inter-Sensor 3-D Convolutional Regression Networks
abstract
Sentinel missions provide widespread opportunities of exploiting inter-sensor synergies to improve the operational monitoring of terrestrial photosynthetic activity and canopy structural variations using vegetation indices (VI). In this context, continuous and consistent temporal data are logically required to rapidly detect vegetation changes across sensors. Nonetheless, the existing temporal limitations inherent to satellite orbits, cloud occlusions, data degradation, and many other factors may severely constrain the availability of data involving multiple satellites. In response, this letter proposes a novel deep 3-D convolutional regression network (3CRN) for temporally enhancing Sentinel-3 (S3) VI by taking advantage of inter-sensor Sentinel-2 (S2) observations. Unlike existing regression and deep learning-based methods, the proposed approach allows convolutional kernels to slide across the temporal dimension to exploit not only the higher spatial resolution of the S2 instrument but also its own temporal evolution to better estimate time-resolved VI in S3. To validate the proposed approach, we built a database made of multiple day-synchronized S2 and S3 operational products from a study area in Extremadura (Spain). The conducted experimental comparison, including multiple state-of-the-art regression and deep learning models, shows the statistically significant advantages of the presented framework. The codes of this work will be made available athttps://github.com/rufernan/3CRN.
Rubén Fernández-Beltran, Damian Ibañez, Jian Kang 0005, Filiberto Pla
IEEE Geosci. Remote. Sens. Lett.3
2022 Deep Learning-Based Building Footprint Extraction With Missing Annotations
abstract
Most state-of-the-art deep learning-based methods for extraction of building footprints are aimed at designing proper convolutional neural network (CNN) architectures or loss functions able to effectively predict building masks from remote sensing (RS) images. To properly train such CNN models, large-scale and pixel-level building annotations are required. One common approach to obtain scalable benchmark data sets for the segmentation of buildings is to register RS images with auxiliary geospatial information data, such as those available from OpenStreetMaps (OSM). However, due to land-cover changes, urban construction, and delayed geospatial information updating, some building annotations may be missing in the corresponding ground-truth building mask layers. This will likely introduce confusion in the training of CNN models for discriminating between background and building pixels. To solve this important issue, we first formulate the problem as a long-tailed classification one. Then, we introduce a new joint loss function based on three terms: 1) logit adjusted cross entropy (LACE) loss, aimed at discriminating between building and background pixels from a long-tailed label distribution; 2) weighted dice loss, aimed at increasing the$F_{1}$scores of the predicted building masks; and 3) boundary (BD) alignment loss, which is optimized for preserving the fine-grained structure of building boundaries. Our experiments, conducted on two benchmark building segmentation data sets, validate the effectiveness of our newly proposed loss with respect to other state-of-the-art losses commonly used for extracting building footprints. The codes of this letter will be publicly available fromhttps://github.com/jiankang1991/GRSL_BFE_MA.
Jian Kang 0005, Rubén Fernández-Beltran, Xian Sun 0001, Jingen Ni, Antonio Plaza
IEEE Geosci. Remote. Sens. Lett.1
2022 SAR Time-Series Despeckling via Nonlocal Total Variation Regularized Robust PCA
abstract
Through the development of Synthetic Aperture Radar (SAR) technology, it is now possible to observe dynamic processes on the earth with fine temporal resolution by forming SAR time series. Nonetheless, such sequential images remain difficult to interpret due to the speckle effect. Despeckling them is further complicated by outliers caused by abrupt changes in weather conditions or the appearance of objects. In spite of the fact that many state-of-the-art methods can achieve excellent filtering performances over stable areas, they often result in artifacts in those areas where outliers existed at the time of acquisition. To simultaneously mitigate the speckle noise and extract outliers, we propose a novel SAR time series despeckling method based on nonlocal total variation regularized robust principle component analysis, which is termed SAR-NL-TVRPCA. By comparing it to other state-of-the-art methods, the effectiveness of its despeckling has been validated in real data experiments. Furthermore, the extracted outliers can provide insight into abrupt changes occurring throughout the observation period, which provides byproducts for further analysis.
Zhanyu Zhu, Jian Kang 0005, Teng-Yu Ji, Zhe Zhang 0026, Rubén Fernández-Beltran
IEEE Geosci. Remote. Sens. Lett.2
2022 Multisensor Fusion and Explicit Semantic Preserving-Based Deep Hashing for Cross-Modal Remote Sensing Image Retrieval
abstract
Cross-modal hashing is an important tool for retrieving useful information from very-high-resolution (VHR) optical images and synthetic aperture radar (SAR) images. Dealing with the intermodal discrepancies, including both spatial–spectral and visual semantic aspects, between VHR and SAR images is extremely vital to generate high-quality common hash codes in the Hamming space. However, existing cross-modal hashing methods ignore the spatial–spectral discrepancy when representing VHR and SAR images. Moreover, existing methods employ derived supervised signals, such as pairwise training images, to implicitly guide hashing learning, which fails to effectively deal with the visual semantic discrepancy, i.e., cannot adequately preserve the intraclass similarity and interclass discrimination between VHR and SAR images. To address these drawbacks, this article proposes a multisensor fusion and explicit semantic preserving-based deep Hashing method, termed as MsEspH, which can effectively deal with the discrepancies. Specifically, we design a novel cross-modal hashing network to eliminate the spatial–spectral discrepancies by fusing extra multispectral images (MSIs), which are generated in real time by a generative adversarial network. Then, we propose an explicit semantic preserving-based objective function by analyzing the connection between classification and hash learning. The objective function can preserve the intraclass similarity and interclass discrimination with class labels directly. Moreover, we theoretically verify that hash learning and classification can be unified into a learning framework under certain conditions. To evaluate our method, we construct and release a large-scale VHR-SAR image dataset. Extensive experiments on the dataset demonstrate that our method outperforms various state-of-the-art cross-modal hashing methods.
Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Jian Kang 0005, Zhichao Huang 0001, Chuyao Luo
IEEE Trans. Geosci. Remote. Sens.5
2022 Coherence-Guided Complex Convolutional Sparse Coding for Interferometric Phase Restoration
abstract
Interferometric phase restoration is a crucial step in retrieving large-scale geophysical parameters from Synthetic Aperture Radar (SAR) images. Existing noise impacts the accuracy of parameter retrieval as a result of decorrelation effects. Most state-of-the-art filtering methods belong to the group of nonlocal filters. In this paper, we propose a novel convolutional sparse coding method in complex domain with the prior knowledge of coherence integrated into the optimization model, which is termed as CoComCSC. CoComCSC is not only capable of reducing noise in regions with continuous phase changes, but also of preserving the phase details prominently. The experiments results on simulated and real data demonstrate the effectiveness of CoComCSC by comparing with other state-of-the-art methods. Moreover, the obtained Digital Elevation Model (DEM) product by CoComCSC from RADARSAT-2 data indicates its superior filtering performance over regions with heterogeneous land-covers, which shows its great potential for generating high-resolution DEM products.
Jian Kang 0005, Zhe Zhang 0026, Yan Huang 0018, Jialin Liu 0003, Naoto Yokoya
IEEE Trans. Geosci. Remote. Sens.2
2022 Rotation-Invariant Deep Embedding for Remote Sensing Images
abstract
Endowing convolutional neural networks (CNNs) with the rotation-invariant capability is important for characterizing the semantic contents of remote sensing (RS) images since they do not have typical orientations. Most of the existing deep methods for learning rotation-invariant CNN models are based on the design of proper convolutional or pooling layers, which aims at predicting the correct category labels of the rotated RS images equivalently. However, a few works have focused on learning rotation-invariant embeddings in the framework of deep metric learning for modeling the fine-grained semantic relationships among RS images in the embedding space. To fill this gap, we first propose a rule that the deep embeddings of rotated images should be closer to each other than those of any other images (including the images belonging to the same class). Then, we propose to maximize the joint probability of the leave-one-out image classification and rotational image identification. With the assumption of independence, such optimization leads to the minimization of a novel loss function composed of two terms: 1) a class-discrimination term and 2) a rotation-invariant term. Furthermore, we introduce a penalty parameter that balances these two terms and further propose a final loss to Rotation-invariant Deep embedding for RS images, termed RiDe. Extensive experiments conducted on two benchmark RS datasets validate the effectiveness of the proposed approach and demonstrate its superior performance when compared to other state-of-the-art methods. The codes of this article will be publicly available athttps://github.com/jiankang1991/TGRS_RiDe.
Jian Kang 0005, Rubén Fernández-Beltran, Zhirui Wang 0003, Xian Sun 0001, Jingen Ni, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2022 DisOptNet: Distilling Semantic Knowledge From Optical Images for Weather-Independent Building Segmentation
abstract
Synthetic aperture radar (SAR) images provide all-weather and all-time capabilities for Earth observation, which becomes highly beneficial in the field of intelligent remote sensing (RS) image interpretation. Due to these advantages, SAR images have been widely exploited in automatic building segmentation tasks under poor weather conditions, especially when disasters happen. However, compared to optical images, the semantics inherent to SAR images are less rich and interpretable due to factors such as speckle noise and imaging geometry. In this scenario, most state-of-the-art methods are focused on designing advanced network architectures or loss functions for building footprint extraction. However, few works have been oriented toward improving segmentation performance through knowledge transfer from optical images. In this article, we propose a novel method based on theDisOptNetnetwork, which can distill the useful semantic knowledge from optical images into a network only trained with SAR data. Specifically, we first analyze the multilevel feature discrepancies between multiple stages of the networks pretrained on the two image modalities. We observe that feature discrepancies start to increase as the encoding stage gradually changes from low level to high level. Based on such observation, we reuse the early stage features and construct parallel convolutional neural network (CNN) branches that are responsible for capturing high-level domain-specific knowledge for each image modality. The optical branch is aimed at mimicking feature generation at the optical pretrained network given the input SAR images. Then, an aggregation module is introduced to calibrate and fuse the features from different modalities while generating the building segments. Extensive experiments were conducted on a large-scale multisensor all-weather building segmentation dataset with state-of-the-art methods used for comparison. Our experimental results validate the effectiveness ofDisOptNet, which demonstrates great potential in the task of weather-independent building footprint generation under real scenarios. The codes of this article will be made publicly available athttps://github.com/jiankang1991/TGRS_DisOptNet.
Jian Kang 0005, Zhirui Wang 0003, Ruoxin Zhu, Junshi Xia, Xian Sun 0001, Rubén Fernández-Beltran, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2022 Multisource Data Reconstruction-Based Deep Unsupervised Hashing for Unisource Remote Sensing Image Retrieval
abstract
Unsupervised hashing for remote sensing (RS) image retrieval first extracts image features and then use these features to construct supervised information (e.g., pseudo-labels) to train hashing networks. Existing methods usually regard RS images as natural images to extract unisource features. However, these features only contain partial information about ground objects and cannot produce reliable pseudo-labels. In addition, existing methods only generate a pseudo single-label to annotate each RS image, which cannot accurately represent multiple scenes in a RS image. To address these drawbacks, this paper proposes a new Multisource data reconstruction-based deep unsupervised Hashing method, called MrHash, which explores the characteristics of RS images to construct reliable pseudo-labels. In particular, we first use geographic coordinates to obtain different satellite images and develop a novel autoencoder network to extract multisource features from these images. Then pseudo multi-labels are designed to deal with the coexistence of multiple scenes in a single image. These labels are generated by a custom probability function with extracted multisource features. Finally, we propose a novel multi-semantic hash loss by using the Kull-back–Leibler (KL) divergence to preserve the semantic similarity of these pseudo multi-labels in Hamming space. Our newly developed MrHash only uses multisource images to construct supervised information, and hash code generation still relies on a unisource input image. Experiments on benchmark datasets clearly show the superiority of the proposed method over state-of-the-art baselines. https://github.com/sunyuxi/MrHash.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Xutao Li 0003, Bowen Zhang 0005, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2022 Normal Assisted Pixel-Visibility Learning With Cost Aggregation for Multiview Stereo
abstract
Multiple-View Stereo (MVS) aims to reconstruct the dense 3D representations of scenes. MVS has potential applications in the fields of autonomous driving (unstructured environment construction) and robotic navigation (visual-inertial navigation). To mitigate the error of depth estimation in low-textured or occluded regions, this work proposes a two-stage multi-view stereo network for fast and accurate depth estimation. The improvements of this work over the state of the art are as follows: 1) Sparse costs are constructed to jointly predict the initial depth map and surface normal by cost regularization, which proves that the surface normals can be estimated in this way with low memory consumption. 2) A new edge refinement block is developed to refine the coarse surface normal to obtain a fine-grained surface normal map. 3) Instead of using the general variance-based metric to equally aggregate cost, a new content-adaptive cost aggregation mechanism based on the similarity of the neighboring surface normal is designed for reliable cost aggregation. To the best of our knowledge, the proposed work is the first trainable network that leverages surface normal as guidance to capture neighboring pixel-visibility, which is an effective supplement to existing depth/normal estimation frameworks. Experimental results indicate that our method can not only achieve accurate depth estimation for scene perception but also make no concession to the real-time performance and limited memory bottleblock. Multiple-view stereo (MVS) aims to reconstruct the dense 3D representations of scenes. It is widely used in the fields of industrial measurement, autonomous driving, and robotic navigation. To mitigate the error of depth estimation in challenging scenarios, this work proposes a two-stage multi-view stereo network for fast and accurate depth estimation. Our method is the first trainable network that leverages surface normal as pixel-visibility guidance to aggregate reliable cost, which could achieve accurate depth estimation and provide the perception ability for the robot. The proposed method has great potential in the fields of 3D reconstruction, industrial measurement, and robotic navigation to estimate real-time and accurate depth with limited memory consumption.
Kevin W. Tong, Xiaorong Guan, Jian Kang 0005, Zhao-Hui Sun, Rob Law 0001, Pedram Ghamisi, Qi Wu 0003
IEEE Trans. Intell. Transp. Syst.3
2021 Robust Deep Metric Learning for Remote Sensing Images with Noisy Annotations
abstract
Manual and automatic annotation of Remote Sensing (RS) scenes are rather complex tasks which may unavoidably introduce some degree of mislabeled data in large-scale archives. In this regard, noisy annotations become an important constraint for deep metric learning-based RS characterization methods since most of them are trained in a supervised way. To address this problem, here we investigate the use of deep metric learning for characterizing RS scenes with noisy labels. Specifically, we consider the Normalized Softmax Loss and develop a robust extension, i.e., the Robust Normalized Softmax Loss (RNSL), in order to effectively capture the semantic relationships among RS scenes with mislabeled ground-truth information. The conducted experiments, using the K-NN classifier and two benchmark RS image archives, show the potential of the proposed approach with respect to other state-of-the-art methods.
Jian Kang 0005, Rubén Fernández-Beltran, Puhong Duan, Xudong Kang, Antonio Plaza
IGARSS1
2021 Generalized Scalable Neighborhood Component Analysis for Single and Multi-Label Remote Sensing Image Characterization
abstract
Deep metric learning has recently become a prominent technology for the semantic understanding of remote sensing (RS) scenes due to its great potential for characterizing visual semantics. However, state-of-the-art deep metric learning models are often constrained in RS by the use of single-label annotations, which eventually reduce their capacity to characterize complex aerial scenes. Additionally, many of the existing works are specialized in particular RS applications which constrains the study of their associated metric spaces from a multi-task perspective. In this paper, we propose a new unified deep metric learning approach for both single- and multi-label RS scene characterization while also taking into account different downstream RS applications. Specifically, we extend the Scalable Neighborhood Component Analysis (SNCA) to the multi-label case and propose its generalized version, i.e., GSNCA. Extensive experiments on single- and multi-label RS benchmark datasets have been conducted to evaluate the effectiveness of the proposed method for RS image classification, clustering and retrieval.
Jian Kang 0005, Rubén Fernández-Beltran, Antonio Plaza
IGARSS1
2021 Graph-Induced Aligned Learning on Subspaces for Hyperspectral and Multispectral Data
abstract
In this article, we have great interest in investigating a common but practical issue in remote sensing (RS)-can a limited amount of one information-rich (or high-quality) data, e.g., hyperspectral (HS) image, improve the performance of a classification task using a large amount of another information-poor (low-quality) data, e.g., multispectral (MS) image? This question leads to a typical cross-modality feature learning. However, classic cross-modality representation learning approaches, e.g., manifold alignment, remain limited in effectively and efficiently handling such problems that the data from high-quality modality are largely absent. For this reason, we propose a novel graph-induced aligned learning (GiAL) framework by 1) adaptively learning a unified graph (further yielding a Laplacian matrix) from the data in order to align multimodality data (MS-HS data) into a latent shared subspace; 2) simultaneously modeling two regression behaviors with respect to labels and pseudo-labels under a multitask learning paradigm; and 3) dramatically updating the pseudo-labels according to the learned graph and refeeding the latest pseudo-labels into model learning of the next round. In addition, an optimization framework based on the alternating direction method of multipliers (ADMMs) is devised to solve the proposed GiAL model. Extensive experiments are conducted on two MS-HS RS data sets, demonstrating the superiority of the proposed GiAL compared with several state-of-the-art methods.
Danfeng Hong, Jian Kang 0005, Naoto Yokoya, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.2
2021 Robust Normalized Softmax Loss for Deep Metric Learning-Based Characterization of Remote Sensing Images With Label Noise
abstract
Most deep metric learning-based image characterization methods exploit supervised information to model the semantic relations among the remote sensing (RS) scenes. Nonetheless, the unprecedented availability of large-scale RS data makes the annotation of such images very challenging, requiring automated supportive processes. Whether the annotation is assisted by aggregation or crowd-sourcing, the RS large-variance problem, together with other important factors [e.g., geo-location/registration errors, land-cover changes, even low-quality Volunteered Geographic Information (VGI), etc.] often introduce the so-called label noise, i.e., semantic annotation errors. In this article, we first investigate the deep metric learning-based characterization of RS images with label noise and propose a novel loss formulation, named robust normalized softmax loss (RNSL), for robustly learning the metrics among RS scenes. Specifically, our RNSL improves the robustness of the normalized softmax loss (NSL), commonly utilized for deep metric learning, by replacing its logarithmic function with the negative Box–Cox transformation in order to down-weight the contributions from noisy images on the learning of the corresponding class prototypes. Moreover, by truncating the loss with a certain threshold, we also propose a truncated robust normalized softmax loss (t-RNSL) which can further enforce the learning of class prototypes based on the image features with high similarities between them, so that the intraclass features can be well grouped and interclass features can be well separated. Our experiments, conducted on two benchmark RS data sets, validate the effectiveness of the proposed approach with respect to different state-of-the-art methods in three different downstream applications (classification, clustering, and retrieval). The codes of this article will be publicly available fromhttps://github.com/jiankang1991.
Jian Kang 0005, Rubén Fernández-Beltran, Puhong Duan, Xudong Kang, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2021 Deep Unsupervised Embedding for Remotely Sensed Images Based on Spatially Augmented Momentum Contrast
abstract
Convolutional neural networks (CNNs) have achieved great success when characterizing remote sensing (RS) images. However, the lack of sufficient annotated data (together with the high complexity of the RS image domain) often makes supervised and transfer learning schemes limited from an operational perspective. Despite the fact that unsupervised methods can potentially relieve these limitations, they are frequently unable to effectively exploit relevant prior knowledge about the RS domain, which may eventually constrain their final performance. In order to address these challenges, this article presents a new unsupervised deep metric learning model, called spatially augmented momentum contrast (SauMoCo), which has been specially designed to characterize unlabeled RS scenes. Based on the first law of geography, the proposed approach defines spatial augmentation criteria to uncover semantic relationships among land cover tiles. Then, a queue of deep embeddings is constructed to enhance the semantic variety of RS tiles within the considered contrastive learning process, where an auxiliary CNN model serves as an updating mechanism. Our experimental comparison, including different state-of-the-art techniques and benchmark RS image archives, reveals that the proposed approach obtains remarkable performance gains when characterizing unlabeled scenes since it is able to substantially enhance the discrimination ability among complex land cover categories. The source codes of this article will be made available to the RS community for reproducible research.
Jian Kang 0005, Rubén Fernández-Beltran, Puhong Duan, Sicong Liu 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2021 Graph Relation Network: Modeling Relations Between Scenes for Multilabel Remote-Sensing Image Classification and Retrieval
abstract
Due to the proliferation of large-scale remote-sensing (RS) archives with multiple annotations, multilabel RS scene classification and retrieval are becoming increasingly popular. Although some recent deep learning-based methods are able to achieve promising results in this context, the lack of research on how to learn embedding spaces under the multilabel assumption often makes these models unable to preserve complex semantic relations pervading aerial scenes, which is an important limitation in RS applications. To fill this gap, we propose a new graph relation network (GRN) for multilabel RS scene categorization. Our GRN is able to model the relations between samples (or scenes) by making use of a graph structure which is fed into network learning. For this purpose, we define a new loss function called scalable neighbor discriminative loss with binary cross entropy (SNDL-BCE) that is able to embed the graph structures through the networks more effectively. The proposed approach can guide deep learning techniques (such as convolutional neural networks) to a more discriminative metric space, where semantically similar RS scenes are closely embedded and dissimilar images are separated from a novel multilabel viewpoint. To achieve this goal, our GRN jointly maximizes a weighted leave-one-out K-nearest neighbors ( KNN) score in the training set, where the weight matrix describes the contributions of the nearest neighbors associated with each RS image on its class decision, and the likelihood of the class discrimination in the multilabel scenario. An extensive experimental comparison, conducted on three multilabel RS scene data archives, validates the effectiveness of the proposed GRN in terms of KNN classification and image retrieval. The codes of this article will be made publicly available for reproducible research in the community.
Jian Kang 0005, Rubén Fernández-Beltran, Danfeng Hong, Jocelyn Chanussot, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2021 Learning Convolutional Sparse Coding on Complex Domain for Interferometric Phase Restoration
abstract
Interferometric phase restoration has been investigated for decades and most of the state-of-the-art methods have achieved promising performances for InSAR phase restoration. These methods generally follow the nonlocal filtering processing chain, aiming at circumventing the staircase effect and preserving the details of phase variations. In this article, we propose an alternative approach for InSAR phase restoration, that is, Complex Convolutional Sparse Coding (ComCSC) and its gradient regularized version. To the best of the authors' knowledge, this is the first time that we solve the InSAR phase restoration problem in a deconvolutional fashion. The proposed methods can not only suppress interferometric phase noise, but also avoid the staircase effect and preserve the details. Furthermore, they provide an insight into the elementary phase components for the interferometric phases. The experimental results on synthetic and realistic high- and medium-resolution data sets from TerraSAR-X StripMap and Sentinel-1 interferometric wide swath mode, respectively, show that our method outperforms those previous state-of-the-art methods based on nonlocal InSAR filters, particularly the state-of-the-art method: InSAR-BM3D. The source code of this article will be made publicly available for reproducible research inside the community.
Jian Kang 0005, Danfeng Hong, Jialin Liu 0003, Gerald Baier, Naoto Yokoya, Begüm Demir
IEEE Trans. Neural Networks Learn. Syst.1
2020 Band-Wise Multi-Scale CNN Architecture for Remote Sensing Image Scene Classification
abstract
Most of the existing convolutional neural network (CNN) architectures in the framework of image scene classification problems are designed for modeling RGB image bands. Direct application of these architectures to the high-dimensional remote sensing (RS) scene classification can be insufficient to accurately describe the spectral content. To address this issue, we propose a novel CNN architecture for the feature embedding of high-dimensional RS images. The proposed architecture aims at: 1) decoupling the spectral and spatial feature extraction for sufficiently describing the complex information content of images; and 2) taking advantage of multi-scale representations of different land-use and land-cover classes present in the images. To this end, the proposed architecture is mainly composed of: 1) a convolutional layer for band-wise extraction of multi-scale spatial features; 2) a convolutional layer for pixel-wise extraction of spectral features; and 3) standard 2D convolution and residual blocks for further feature learning. Experiments on BigEarthNet validate the effectiveness of the proposed method, when compared to the state-of-the-art CNN architectures.
Jian Kang 0005, Begüm Demir
IGARSS1
2020 Sun Glint Removal of Hyperspectral Images via Texture-Aware Total Variation
abstract
Sun glint, as the spectral reflection of solar radiation on non-flat water surfaces, is a serious confounding factor for coastal shallow-water environments. When the coastal areas are observed with a hyperspectral sensor, the existing sun glint in the produced images can seriously influence the quality of the image interpretation. To solve this issue, in this paper, we propose a novel sun glint removal method based on a variation model for hyperspectral images (HSIs). The proposed method aims to decompose the original HSI into a desired clean image and a sun glint image. To achieve this, we exploit a texture-aware total variation to remove the sun glint in HSIs, where the texture information is imposed on the total variation regularization to highlight sun glint. Experiments on simulated and real datasets demonstrate that our method can obtain outstanding performance with respect to other state-of-the-art approaches.
Puhong Duan, Jian Kang 0005, Xudong Kang, Pedram Ghamisi, Shutao Li 0001
IGARSS2
2020 Learning-Shared Cross-Modality Representation Using Multispectral-LiDAR and Hyperspectral Data
abstract
Due to the ever-growing diversity of the data source, multimodality feature learning has attracted more and more attention. However, most of these methods are designed by jointly learning feature representation from multimodalities that exist in both training and test sets, yet they are less investigated in the absence of certain modality in the test phase. To this end, in this letter, we propose to learn a shared feature space across multimodalities in the training process. By this way, the out-of-sample from any of multimodalities can be directly projected onto the learned space for a more effective cross-modality representation. More significantly, the shared space is regarded as a latent subspace in our proposed method, which connects the original multimodal samples with label information to further improve the feature discrimination. Experiments are conducted on the multispectral-Light Detection and Ranging (LIDAR) and hyperspectral data set provided by the 2018 IEEE GRSS Data Fusion Contest to demonstrate the effectiveness and superiority of the proposed method in comparison with several popular baselines.
Danfeng Hong, Jocelyn Chanussot, Naoto Yokoya, Jian Kang 0005, Xiao Xiang Zhu 0001
IEEE Geosci. Remote. Sens. Lett.4
2020 Deep Metric Learning Based on Scalable Neighborhood Components for Remote Sensing Scene Characterization
abstract
With the development of convolutional neural networks (CNNs), the semantic understanding of remote sensing (RS) scenes has been significantly improved based on their prominent feature encoding capabilities. While many existing deep-learning models focus on designing different architectures, only a few works in the RS field have focused on investigating the performance of the learned feature embeddings and the associated metric space. In particular, two main loss functions have been exploited: the contrastive and the triplet loss. However, the straightforward application of these techniques to RS images may not be optimal in order to capture their neighborhood structures in the metric space due to the insufficient sampling of image pairs or triplets during the training stage and to the inherent semantic complexity of remotely sensed data. To solve these problems, we propose a new deep metric learning approach, which overcomes the limitation on the class discrimination by means of two different components: 1) scalable neighborhood component analysis (SNCA) that aims at discovering the neighborhood structure in the metric space and 2) the cross-entropy loss that aims at preserving the class discrimination capability based on the learned class prototypes. Moreover, in order to preserve feature consistency among all the minibatches during training, a novel optimization mechanism based on momentum update is introduced for minimizing the proposed loss. An extensive experimental comparison (using several state-of-the-art models and two different benchmark data sets) has been conducted to validate the effectiveness of the proposed method from different perspectives, including: 1) classification; 2) clustering; and 3) image retrieval. The related codes of this article will be made publicly available for reproducible research by the community.
Jian Kang 0005, Rubén Fernández-Beltran, Zhen Ye 0009, Xiaohua Tong, Pedram Ghamisi, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2020 Multipass SAR Interferometry Based on Total Variation Regularized Robust Low Rank Tensor Decomposition
abstract
Multipass SAR interferometry (InSAR) techniques based on meter-resolution spaceborne SAR satellites, such as TerraSAR-X or COSMO-SkyMed, provide 3D reconstruction and the measurement of ground displacement over large urban areas. Conventional methods such as persistent scatterer interferometry (PSI) usually requires a fairly large SAR image stack (usually in the order of tens) to achieve reliable estimates of these parameters. Recently, low rank property in multipass InSAR data stack was explored and investigated in our previous work (J. Kang et al., “Object-based multipass InSAR via robust low-rank tensor decomposition,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 6, 2018). By exploiting this low rank prior, a more accurate estimation of the geophysical parameters can be achieved, which in turn can effectively reduce the number of interferograms required for a reliable estimation. Based on that, this article proposes a novel tensor decomposition method in a complex domain, which jointly exploits low rank and variational prior of the interferometric phase in InSAR data stacks. Specifically, a total variation (TV) regularized robust low rank tensor decomposition method is exploited for recovering outlier-free InSAR stacks. We demonstrate that the filtered InSAR data stacks can greatly improve the accuracy of geophysical parameters estimated from real data. Moreover, this article demonstrates for the first time in the community that tensor-decomposition-based methods can be beneficial for large-scale urban mapping problems using multipass InSAR. Two TerraSAR-X data stacks with large spatial areas demonstrate the promising performance of the proposed method.
Jian Kang 0005, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 Fusing Spaceborne SAR Interferometry and Street View Images for 4D Urban Modeling
abstract
Obtaining city models in a large scale is usually achieved by means of remote sensing techniques, such as synthetic aperture radar (SAR) interferometry and optical image stereogrammetry. Despite the controlled quality of these products, such observation is restricted by the characteristics of their sensor platform, such as revisit time and spatial resolution. Over the last decade, the rapid development of online geographic information systems, such as Google map, has accumulated vast amount of online images. Despite their uncontrolled quality, these images constitute a set of redundant spatial-temporal observations of our dynamic 3D urban environment. These images contain useful information that can complement the remote sensing data, especially the SAR images. This paper presents a one of the first studies of fusing online street view images and spaceborne SAR images, for the reconstruction of spatial-temporal (hence 4D) city models. We describe a general approach to geometrically combine the information of these two types of images that are nearly impossible to even coregister without a precise 3D city model due to their distinct imaging geometry. It is demonstrated that, one can obtain a new kind of city model that includes high resolution optical texture for better scene understanding and the dynamics of individual buildings up to the precision of millimeter retrieved from SAR interferometry.
Yuanyuan Wang 0002, Jian Kang 0005, Xiao Xiang Zhu 0001
FUSION2
2018 Multi-Pass SAR Interferometry for 3D Reconstruction of Complex Mountainous Areas Based on Robust Low Rank Tensor Decomposition
abstract
During the past decades, multi-pass SAR interferometry (In-SAR) techniques have been developed for retrieving geophysical parameters such as elevation, over large areas. Conventional method such as periodogram usually requires a fairly large SAR image stack (usually in the order of tens), in order to achieve reliable estimates of these parameters. However, when it comes to large-area processing, it is time-consuming and luxury to obtain a sufficient number of SAR images for the reconstruction. In this paper, we demonstrate a novel multi-pass InSAR method for 3D reconstruction using low rank tensor decomposition. By exploiting the low rank prior knowledge in the multi-pass InSAR stack, simulations show that the proposed method can improve the accuracy of elevation estimates by a factor of two, compared to the state-of-the-art InSAR filtering methods, such as SqueeSAR. The capability of the proposed algorithm is also demonstrated on real data using one TanDEM-X InSAR stack of a complex mountainous area.
Jian Kang 0005, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS1
2018 Object-Based Multipass InSAR via Robust Low-Rank Tensor Decomposition
abstract
The most unique advantage of multipass synthetic aperture radar interferometry (InSAR) is the retrieval of long-term geophysical parameters, e.g., linear deformation rates, over large areas. Recently, an object-based multipass InSAR framework has been proposed by Kang, as an alternative to the typical single-pixel methods, e.g., persistent scatterer interferometry (PSI), or pixel-cluster-based methods, e.g., SqueeSAR. This enables the exploitation of inherent properties of InSAR phase stacks on an object level. As a follow-on, this paper investigates the inherent low rank property of such phase tensors and proposes a Robust Multipass InSAR technique via Object-based low rank tensor decomposition. We demonstrate that the filtered InSAR phase stacks can improve the accuracy of geophysical parameters estimated via conventional multipass InSAR techniques, e.g., PSI, by a factor of 10-30 in typical settings. The proposed method is particularly effective against outliers, such as pixels with unmodeled phases. These merits, in turn, can effectively reduce the number of images required for a reliable estimation. The promising performance of the proposed method is demonstrated using high-resolution TerraSAR-X image stacks.
Jian Kang 0005, Yuanyuan Wang 0002, Michael Schmitt 0003, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 Improve multi-baseline InSAR parameter retrieval by semantic information from optical images
abstract
One of the most unique benefits of multi-baseline synthetic aperture radar interferometry (InSAR) is the long-term monitoring of subtle ground deformation over large areas. Most state-of-the-art algorithms for retrieving such parameter are based on single pixels, e.g. Permanent Scatterer InSAR [1] or clusters of ergodic pixels with stationary phases e.g. SqueeSAR [2]. None of the studies has addressed the joint inversion in an object level, where the true interferometric phase may be varying subject to topography and deformation. Recently, one study has investigated SAR and optical data fusion in order to make use of the rich semantic information from optical images [3]. Based on that work, we seek to investigate the possibility of an object-level multi-baseline InSAR deformation reconstruction given the semantic information from the corresponding optical images. In this paper, we introduced the tensor model for the multi-baseline InSAR inversion and proposed a maximum a posteriori estimator of the deformation parameters by including a spatial prior function in the objective function. Substantial improvement in the deformation estimation is observed in the experiments using both simulated and the real SAR data.
Jian Kang 0005, Yuanyuan Wang 0002, Marco Körner 0001, Xiao Xiang Zhu 0001
IGARSS1
2017 Robust Object-Based Multipass InSAR Deformation Reconstruction
abstract
Deformation monitoring by multipass synthetic aperture radar (SAR) interferometry (InSAR) is, so far, the only imaging-based method to assess millimeter-level deformation over large areas from space. Past research mostly focused on the optimal retrieval of deformation parameters on the basis of a single pixel or a pixel cluster. Only until recently, the first demonstration of object-based urban infrastructure monitoring by fusing InSAR and the semantic classification labels derived from optical images was presented by Wanget al.Given such classification labels in the SAR image, we propose a general framework for object-based InSAR parameter retrieval, where the parameters of the whole object are jointly estimated by the inversion of a regularized tensor model instead of pixelwise. Our approach does not assume the stationarity of each sample in the object, which is usually assumed in other pixel cluster-based methods, such as SqueeSAR. In addition, to handle outliers in real data, a robust phase recovery step prior to parameter retrieval is also introduced. In typical settings, the proposed method outperforms the current pixelwise estimators, e.g., periodogram, by a factor of several tens in the accuracy of the linear deformation estimates. Last but not least, for a practical demonstration on bridge monitoring, we present a full workflow of long-term bridge monitoring using the proposed approach.
Jian Kang 0005, Yuanyuan Wang 0002, Marco Körner 0001, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2016 Object-based InSAR deformation reconstruction with application to bridge monitoring
abstract
Deformation monitoring by multi-baseline synthetic aperture radar (SAR) interferometry is so far the only imaging-based method to assess millimeter-level deformation over large areas from space. Past research mostly focused on optimal deformation parameters retrieval on a pixel-basis. Only until recently, the first demonstration of object-based urban infrastructures monitoring by fusing SAR interferometry (InSAR) and the semantic classification labels derived from optical images was presented in [1]–[3]. This paper proposes an algorithm for object-based joint InSAR deformation reconstruction using these classification labels. We derive an object-based multi-baseline InSAR reconstruction model, and propose an efficient algorithm for bridge detection in optical images.
Jian Kang 0005, Yuanyuan Wang 0002, Marco Körner 0001, Xiao Xiang Zhu 0001
IGARSS1