EDBT 2026 Demo / reviewers in the wild / expert
Rubén Fernández-Beltran
dblp:130/1803
· DBLP profile ↗
48ranked-venue papers
13as first author
28since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-modal consistent loss diffusion model for Sentinel-3 single image super resolutionabstractAbstract In the context of Earth observation, the trade-off between spatial, spectral, and temporal resolution often limits the versatility of remote sensing images in many important applications. In response, this paper introduces a novel deep learning diffusion model, specifically tailored to improve the spatial resolution of the optical products acquired by the Sentinel-3 (S3) satellite. Our framework employs a diffusion probabilistic model, benefiting from the higher spatial resolution of the Sentinel-2 satellite during training via a new multi-modal loss formulation. This ensures consistency with the original S3 images while enhancing the spatial details. Two distinct conditional low-resolution encoders were experimented with, providing insights into their respective contributions to the diffusion process. The efficacy of the proposed model is demonstrated through extensive ablation studies and comparisons with state-of-the-art methods, using both synthetic and real S3 products. The findings indicate that our model successfully improves spatial resolution while maintaining the integrity of the spectral information, contributing to the field of remote sensing single-image super-resolution. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla, Naoto Yokoya, Junshi Xia |
Neural Comput. Appl. | 2 |
| 2025 | Inter-Sensor High-Resolution and Multi-Temporal Image Fusion for Unsupervised Domain Adaptation in Remote SensingabstractMotivated by the increasing demand for robust segmentation in unlabeled remote sensing data, we propose DAM-Former, a novel UDA model that fuses high-resolution multimodal imagery with multi-temporal multispectral data. Current UDA approaches in remote sensing rarely exploit the complementary strengths of spatial and temporal features. To address this gap, our framework integrates two interconnected branches: a transformer-based network for high-resolution multimodal data and a lightweight convolutional network with temporal attention for multi-temporal imagery. To improve segmentation accuracy and lower noise, the extracted features are robustly combined through a deep temporal fusion module and a new mixed loss with an ensemble pseudo-label strategy. Extensive experiments and an ablation study on the FLAIR-2 dataset demonstrate that DAM-Former outperforms state-of-the-art methods, marking the first in-depth study of temporal information fusion in UDA segmentation for remote sensing data. Damian Ibañez, Junshi Xia, Naoto Yokoya, Filiberto Pla, Rubén Fernández-Beltran |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Shadow detection using a cross-attentional dual-decoder network with self-supervised image reconstruction featuresabstractShadow detection is a challenging problem in computer vision due to the high variability in lighting conditions, object shapes, and scene layouts. Despite the positive results achieved by some existing technologies, the problem becomes particularly challenging with complex and heterogeneous images where shadow-casting objects coexist and shadows can have different depths, scales, and morphologies. As a result, more advanced and accurate solutions are still needed to deal with this type of complexities. To address these challenges, this paper proposes a novel deep learning model, called the Cross-Attentional Dual Decoder Network (CADDN), to improve shadow detection by using fine-grained image reconstruction features. Unlike other existing methods, the CADDN uses an innovative encoder-decoder architecture with two decoder segments that work together to reconstruct the input images and their corresponding shadow masks. In this way, the features used to reconstruct the original input image can be used to support the shadow detection process itself. The proposed model also incorporates a cross-attention mechanism to weight the most relevant features for detecting shadows and skip connections with noise to improve the quality of the transferred features. The experimental results, including several benchmark image datasets and state-of-the-art detection methods, demonstrate the suitability of the presented approach for detecting shadows in computer vision applications. Rubén Fernández-Beltran, Angélica Guzmán-Ponce, Rafael Fernandez, Jian Kang 0005, Ginés García-Mateos |
Image Vis. Comput. | 1 |
| 2024 | Semi- and Self-Supervised Metric Learning for Remote Sensing ApplicationsabstractEarth data collection from satellites and aircraft has exponentially grown, but a substantial portion of it remains unlabeled. This has prompted the remote sensing community to explore effective methods for leveraging unlabeled data. In our prior investigation [1], we evaluated various deep semi-supervised learning algorithms on two very high-resolution (VHR) optical datasets (UCM [2] and AID [3]). Notably, the CoMatch [4] algorithm demonstrated the highest accuracy, motivating further exploration. This letter extends our earlier work by integrating the established Class-Aware Contrastive Semi-Supervised Learning framework (Comatch+CCSSL) [5] into CoMatch and introducing a new triplet metric learning loss (CoMatch+Triplet). CoMatch+Triplet excelled with 93.2% accuracy on UCM, while CoMatch led with 92.19% on AID. The addition of the triplet loss can produce a clearer separation of the samples from different classes in the embedding space at very early learning stages, being able to learn faster and getting maximum performance with few iterations. The exploration of diverse semi and self-supervised training methodologies presented in this work sheds light on the strengths and limitations of these approaches, enhancing our understanding of their applicability in remote sensing applications. Itza Hernandez-Sequeira, Rubén Fernández-Beltran, Filiberto Pla |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Cross-Modal Hashing With Feature Semi-Interaction and Semantic Ranking for Remote Sensing Ship Image RetrievalabstractCross-modal hashing plays a pivotal role in large-scale remote sensing (RS) ship image retrieval. RS ship images often exhibit similar overall appearance with subtle differences. Existing hashing methods typically employ feature non-interaction strategies to generate common hash codes, which may not effectively capture the correlations between cross-modal ship images to reduce intermodality discrepancies. To address this issue, we propose a novel cross-modal hashing approach based on feature semi-interaction and semantic ranking (FSISR) for RS ship image retrieval. Our FSISR approach not only captures intricate correlations between different ship image modalities, but also enables the construction of hash tables for large-scale retrieval. FSISR comprises a feature semi-interaction module and a semantic ranking objective function. The semi-interaction module utilizes clustering centers from one modality to learn the correlations between two modalities and generate robust shared representations. The objective function optimizes these representations in a common Hamming space, consisting of a shared semantic alignment loss and a margin-free ranking loss. The alignment loss employs a shared semantic layer to preserve label-level similarity, while the ranking loss incorporates hard examples to establish a margin-free loss that captures similarity ranking relationships. We evaluate the performance of our method on benchmark datasets and demonstrate its effectiveness for cross-modal RS ship image retrieval.https://github.com/sunyuxi/FSISR. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Sebastian Hafner, Xutao Li 0003, Chuyao Luo, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | W-NetPan: Double-U network for inter-sensor self-supervised pan-sharpeningabstractThe increasing availability of remote sensing data allows dealing with spatial-spectral limitations by means of pan-sharpening methods. However, fusing inter-sensor data poses important challenges, in terms of resolution differences, sensor-dependent deformations and ground-truth data availability, that demand more accurate pan-sharpening solutions. In response, this paper proposes a novel deep learning-based pan-sharpening model which is termed as the double-U network for self-supervised pan-sharpening (W-NetPan). In more details, the proposed architecture adopts an innovative W-shape that integrates two U-Net segments which sequentially work for spatially matching and fusing inter-sensor multi-modal data. In this way, a synergic effect is produced where the first segment resolves inter-sensor deviations while stimulating the second one to achieve a more accurate data fusion. Additionally, a joint loss formulation is proposed for effectively training the proposed model without external data supervision. The experimental comparison, conducted over four coupled Sentinel-2 and Sentinel-3 datasets, reveals the advantages of W-NetPan with respect to several of the most important state-of-the-art pan-sharpening methods available in the literature. The codes related to this paper will be available at https://github.com/rufernan/WNetPan. Rubén Fernández-Beltran, Rafael Fernandez, Jian Kang 0005, Filiberto Pla |
Neurocomputing | 1 |
| 2023 | FloU-Net: An Optical Flow Network for Multimodal Self-Supervised Image RegistrationabstractImage registration is an essential task in image processing, where the final objective is to geometrically align two or more images. In remote sensing, this process allows comparing, fusing, or analyzing data, especially when multimodal images are used. In addition, multimodal image registration becomes fairly challenging when the images have a significant difference in scale and resolution, together with local small image deformations. For this purpose, this letter presents a novel optical flow (OF)-based image registration network, named the FloU-Net, which tries to further exploit intersensor synergies by means of deep learning. The proposed method is able to extract spatial information from resolution differences and through a U-Net backbone generate an OF field estimation to accurately register small local deformations of multimodal images in a self-supervised fashion. For instance, the registration between Sentinel-2 (S2) and Sentinel-3 (S3) optical data is not trivial, as there are considerable spectral–spatial differences among their sensors. In this case, the higher spatial resolution of S2 results in S2 data being a convenient reference to spatially improve S3 products, as well as those of the forthcoming Fluorescence Explorer (FLEX) mission, since image registration is the initial requirement to obtain higher data processing level products. To validate our method, we compare the proposed FloU-Net with other state-of-the-art techniques using 21 coupled S2/S3 optical images from different locations of interest across Europe. The comparison is performed through different performance measures. Results show that the proposed FloU-Net can outperform the compared methods. The code and dataset are available inhttps://github.com/ibanezfd/FloU-Net. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | SAR time series despeckling via nonlocal matrix decomposition in logarithm domain
Jian Kang 0005, Teng-Yu Ji, Zhe Zhang 0026, Rubén Fernández-Beltran |
Signal Process. | 4 |
| 2023 | Cross-View Object Geo-Localization in a Local Region With Satellite ImageryabstractCross-view geo-localization is a critical task in various applications, such as smart city management and disaster monitoring. Current methods typically divide a satellite image into patches and use these patches to identify the geographic location of a query image. However, these methods can only provide the location of an image rather than the location of a specific object of interest. This makes it difficult to link these methods to GeoDatabases to obtain detailed information about a target object, such as its name and construction time. To overcome this limitation, we propose a novel problem of cross-view object geo-localization in a local region with high-resolution satellite images. This problem includes two main challenges: accurately identifying the location of an object and distinguishing the target object from others in satellite images. To address these challenges, we present a new Detection-based Geo-localization method called DetGeo, which consists of an object detection-based framework with a two-branch encoder and a query-aware cross-view fusion module. DetGeo uses cross-view images as input to the detector to provide object-level geo-localization. The fusion module employs cross-view spatial attention to focus on relevant areas of target objects during cross-view feature fusion. To evaluate our method, we constructed a new Cross-View Object Geo-Localization dataset called CVOGL, which comprises ground-view or drone-view images as query images and satellite-view images as geo-tagged reference images. Comprehensive experiments are conducted to demonstrate the effectiveness of our method on CVOGL. https://github.com/sunyuxi/DetGeo. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Shanshan Feng 0001, Xutao Li 0003, Chuyao Luo, Puzhao Zhang, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Consistency Center-Based Deep Cross-Modal Hashing for Multisource Remote Sensing Image RetrievalabstractCross-modal hashing aims to retrieve similar images from large-scale Earth Observation (EO) data archives, which typically contain multiple satellite sources of remote sensing (RS) images. However, existing cross-modal hashing methods primarily focus on dual-source RS images and often face two main limitations when retrieving multi-source RS images. Firstly, these methods exhibit significant redundancy as they require handling all possible dual-source combinations in multi-source RS images. Secondly, they often rely on pairwise or triplet image sources to construct objective functions, which are not significantly effective in reducing the discrepancies among multiple RS image sources. To address these limitations, we propose a novel Consistency Center-based deep cross-modal Hashing method called C2Hash for multi-source RS image retrieval. Our C2Hash employs a multi-branch hashing network to directly encode multi-source RS images into unified hash codes, thereby offering higher processing efficiency. Furthermore, C2Hash introduces consistency centers to construct a novel objective function. The consistency center represents the shared semantic features among similar multi-source RS images and is generated by a label hashing network. The objective function encourages similar multi-source RS images to approach the same consistency center to align all image sources in a unified Hamming space. Our method can effectively reduce the discrepancies across multiple image sources and generate unified hash codes. To evaluate its effectiveness, we construct a new Multi-Source RS Image dataset called MSRSI, comprising five different types of image sources. We conduct comprehensive experiments to demonstrate the superior performance of our method on the MSRSI dataset. https://github.com/sunyuxi/C2Hash. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Xutao Li 0003, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Transfer Deep Learning for Remote Sensing Datasets: A Comparison StudyabstractRemote sensing is also benefiting from the quick development of deep learning algorithms for image analysis and classification tasks. In this paper, we evaluate the classification performance of a well-known Convolutional Neural Network (CNN) models, such as ResNet50, using a transfer learning approach. We compare the performance when using vector-features acquired from general purpose data, such as the ImageNet [1], versus remote sensing data like BigEarthNet [2], UCMerced [3], RESISC45 [4] and So2Sat [5]. The results show that the model pre-trained on RESISC-45 data achieved the highest accuracy when classifying the Eurosat [6] testing dataset. This was followed by the model pre-trained on Imagenet with 95.94% and BigEarthNet with 95.93%. When presented with diverse remote sensing data, the classification improved in regards to large quantities of general-purpose data. The experiments carried out also show, that multi modal (co-registered synthetic aperture radar and multispectral) did not increase the classification rate with respect to using only multispectral data. The source codes of this work are available for reproducible research at https://github.com/itzahs/CNN-RS. Itza Hernandez-Sequeira, Rubén Fernández-Beltran, Filiberto Pla |
IGARSS | 2 |
| 2022 | SEN23E: A Cloudless Geo-Referenced Multi-Spectral Sentinel-2/Sentinel-3 Dataset for Data Fusion AnalysisabstractThe availability of geo-referenced coupled data of dif-ferent platforms is essential to train remote sensing (RS) multi-modal classification and bio-phyiscal parameter esti-mation learning methods. To properly develop a general-izing model different scenes and topographies are required. For this purpose, different multi-modal datasets have been published for the last years. Nevertheless, to our knowl-edge there is not any dataset composed of Sentinel-2 (S2) and Sentinel-3 (S3) geo-referenced images. In this paper we present SEN23, a dataset composed of 100 complete multi-spectral S2 and S3 paired images of different locations along Europe from the 2021 summer. The coupled images were obtained with a time difference of three or less days, containing less than a 1 % of cloud coverage and have a resolution difference of × 15. SEN23E is expected to help with the development of new multi-spectral, multi-resolution and multi-modal models for complex tasks which need con-text and complete images. SEN23E will be available at https://github.com/ibanezdf/SEN23E. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla |
IGARSS | 2 |
| 2022 | Toward Tightness of Scalable Neighborhood Component Analysis for Remote-Sensing Image CharacterizationabstractDeep metric learning methods have recently drawn significant attention in the field of remote sensing (RS), owing to their prominent capabilities for modeling relations among RS images based on their semantic contents. In the context of scene classification and large-scale image retrieval, one of the most prominent deep metric learning methods is the scalable neighborhood component analysis (SNCA), which has demonstrated excellent performance on the locality neighborhood structure in the metric space. However, the standard SNCA has important constraints on separating the hard positive and other negative images in the metric space, and this may become a major limitation when dealing with the large-scale variance problem inherent to RS data. To address this issue, we propose a novel deep metric learning formulation that introduces a new margin parameter to enforce the compactness of the within-class feature embeddings. Based on this innovative scheme, we propose two novel loss functions: 1) T-SNCA-c, where the parameter is based on the cosine similarity, and 2) T-SNCA-a, where the parameter is based on the angular distance. Besides, we exploit memory bank optimization to further enhance the semantic diversity during training. Our experimental results, conducted using three downstream applications ($K$-NN classification, clustering, and image retrieval) and two large-scale RS benchmark datasets, demonstrate that the proposed approach can achieve superior performance when compared to current state-of-the-art deep metric learning methods. The codes of this work will be made available online (https://github.com/jiankang1991/GRSL_TSNCA). Jian Kang 0005, Rubén Fernández-Beltran, Sicong Liu 0001, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Time-Resolved Sentinel-3 Vegetation Indices Via Inter-Sensor 3-D Convolutional Regression NetworksabstractSentinel missions provide widespread opportunities of exploiting inter-sensor synergies to improve the operational monitoring of terrestrial photosynthetic activity and canopy structural variations using vegetation indices (VI). In this context, continuous and consistent temporal data are logically required to rapidly detect vegetation changes across sensors. Nonetheless, the existing temporal limitations inherent to satellite orbits, cloud occlusions, data degradation, and many other factors may severely constrain the availability of data involving multiple satellites. In response, this letter proposes a novel deep 3-D convolutional regression network (3CRN) for temporally enhancing Sentinel-3 (S3) VI by taking advantage of inter-sensor Sentinel-2 (S2) observations. Unlike existing regression and deep learning-based methods, the proposed approach allows convolutional kernels to slide across the temporal dimension to exploit not only the higher spatial resolution of the S2 instrument but also its own temporal evolution to better estimate time-resolved VI in S3. To validate the proposed approach, we built a database made of multiple day-synchronized S2 and S3 operational products from a study area in Extremadura (Spain). The conducted experimental comparison, including multiple state-of-the-art regression and deep learning models, shows the statistically significant advantages of the presented framework. The codes of this work will be made available athttps://github.com/rufernan/3CRN. Rubén Fernández-Beltran, Damian Ibañez, Jian Kang 0005, Filiberto Pla |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Deep Learning-Based Building Footprint Extraction With Missing AnnotationsabstractMost state-of-the-art deep learning-based methods for extraction of building footprints are aimed at designing proper convolutional neural network (CNN) architectures or loss functions able to effectively predict building masks from remote sensing (RS) images. To properly train such CNN models, large-scale and pixel-level building annotations are required. One common approach to obtain scalable benchmark data sets for the segmentation of buildings is to register RS images with auxiliary geospatial information data, such as those available from OpenStreetMaps (OSM). However, due to land-cover changes, urban construction, and delayed geospatial information updating, some building annotations may be missing in the corresponding ground-truth building mask layers. This will likely introduce confusion in the training of CNN models for discriminating between background and building pixels. To solve this important issue, we first formulate the problem as a long-tailed classification one. Then, we introduce a new joint loss function based on three terms: 1) logit adjusted cross entropy (LACE) loss, aimed at discriminating between building and background pixels from a long-tailed label distribution; 2) weighted dice loss, aimed at increasing the$F_{1}$scores of the predicted building masks; and 3) boundary (BD) alignment loss, which is optimized for preserving the fine-grained structure of building boundaries. Our experiments, conducted on two benchmark building segmentation data sets, validate the effectiveness of our newly proposed loss with respect to other state-of-the-art losses commonly used for extracting building footprints. The codes of this letter will be publicly available fromhttps://github.com/jiankang1991/GRSL_BFE_MA. Jian Kang 0005, Rubén Fernández-Beltran, Xian Sun 0001, Jingen Ni, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | SAR Time-Series Despeckling via Nonlocal Total Variation Regularized Robust PCAabstractThrough the development of Synthetic Aperture Radar (SAR) technology, it is now possible to observe dynamic processes on the earth with fine temporal resolution by forming SAR time series. Nonetheless, such sequential images remain difficult to interpret due to the speckle effect. Despeckling them is further complicated by outliers caused by abrupt changes in weather conditions or the appearance of objects. In spite of the fact that many state-of-the-art methods can achieve excellent filtering performances over stable areas, they often result in artifacts in those areas where outliers existed at the time of acquisition. To simultaneously mitigate the speckle noise and extract outliers, we propose a novel SAR time series despeckling method based on nonlocal total variation regularized robust principle component analysis, which is termed SAR-NL-TVRPCA. By comparing it to other state-of-the-art methods, the effectiveness of its despeckling has been validated in real data experiments. Furthermore, the extracted outliers can provide insight into abrupt changes occurring throughout the observation period, which provides byproducts for further analysis. Zhanyu Zhu, Jian Kang 0005, Teng-Yu Ji, Zhe Zhang 0026, Rubén Fernández-Beltran |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Masked Auto-Encoding Spectral-Spatial Transformer for Hyperspectral Image ClassificationabstractDeep learning has certainly become the dominant trend in hyper-spectral (HS) remote sensing image classification owing to its excellent capabilities to extract highly discriminating spatial-spectral features. In this context, transformer networks have recently shown prominent results in distinguishing even the most subtle spectral differences because of their potential to characterize sequential spectral data. Nonetheless, many complexities affecting HS remote sensing data (e.g. atmospheric effects, thermal noise, quantization noise, etc.) may severely undermine such potential since no mode of relieving noisy feature patterns has still been developed within transformer networks. To address the problem, this paper presents a novel masked auto-encoding spectral-spatial transformer (MAEST), which gathers two different collaborative branches: (i) a reconstruction path, which dynamically uncovers the most robust encoding features based on a masking auto-encoding strategy; and (ii) a classification path, which embeds these features onto a transformer network to classify the data focusing on the features that better reconstruct the input. Unlike other existing models, this novel design pursues to learn refined transformer features considering the aforementioned complexities of the HS remote sensing image domain. The experimental comparison, including several state-of-the-art methods and benchmark datasets, shows the superior results obtained by MAEST. The codes of this paper will be available at https://github.com/ibanezfd/MAEST. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Rotation-Invariant Deep Embedding for Remote Sensing ImagesabstractEndowing convolutional neural networks (CNNs) with the rotation-invariant capability is important for characterizing the semantic contents of remote sensing (RS) images since they do not have typical orientations. Most of the existing deep methods for learning rotation-invariant CNN models are based on the design of proper convolutional or pooling layers, which aims at predicting the correct category labels of the rotated RS images equivalently. However, a few works have focused on learning rotation-invariant embeddings in the framework of deep metric learning for modeling the fine-grained semantic relationships among RS images in the embedding space. To fill this gap, we first propose a rule that the deep embeddings of rotated images should be closer to each other than those of any other images (including the images belonging to the same class). Then, we propose to maximize the joint probability of the leave-one-out image classification and rotational image identification. With the assumption of independence, such optimization leads to the minimization of a novel loss function composed of two terms: 1) a class-discrimination term and 2) a rotation-invariant term. Furthermore, we introduce a penalty parameter that balances these two terms and further propose a final loss to Rotation-invariant Deep embedding for RS images, termed RiDe. Extensive experiments conducted on two benchmark RS datasets validate the effectiveness of the proposed approach and demonstrate its superior performance when compared to other state-of-the-art methods. The codes of this article will be publicly available athttps://github.com/jiankang1991/TGRS_RiDe. Jian Kang 0005, Rubén Fernández-Beltran, Zhirui Wang 0003, Xian Sun 0001, Jingen Ni, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | DisOptNet: Distilling Semantic Knowledge From Optical Images for Weather-Independent Building SegmentationabstractSynthetic aperture radar (SAR) images provide all-weather and all-time capabilities for Earth observation, which becomes highly beneficial in the field of intelligent remote sensing (RS) image interpretation. Due to these advantages, SAR images have been widely exploited in automatic building segmentation tasks under poor weather conditions, especially when disasters happen. However, compared to optical images, the semantics inherent to SAR images are less rich and interpretable due to factors such as speckle noise and imaging geometry. In this scenario, most state-of-the-art methods are focused on designing advanced network architectures or loss functions for building footprint extraction. However, few works have been oriented toward improving segmentation performance through knowledge transfer from optical images. In this article, we propose a novel method based on theDisOptNetnetwork, which can distill the useful semantic knowledge from optical images into a network only trained with SAR data. Specifically, we first analyze the multilevel feature discrepancies between multiple stages of the networks pretrained on the two image modalities. We observe that feature discrepancies start to increase as the encoding stage gradually changes from low level to high level. Based on such observation, we reuse the early stage features and construct parallel convolutional neural network (CNN) branches that are responsible for capturing high-level domain-specific knowledge for each image modality. The optical branch is aimed at mimicking feature generation at the optical pretrained network given the input SAR images. Then, an aggregation module is introduced to calibrate and fuse the features from different modalities while generating the building segments. Extensive experiments were conducted on a large-scale multisensor all-weather building segmentation dataset with state-of-the-art methods used for comparison. Our experimental results validate the effectiveness ofDisOptNet, which demonstrates great potential in the task of weather-independent building footprint generation under real scenarios. The codes of this article will be made publicly available athttps://github.com/jiankang1991/TGRS_DisOptNet. Jian Kang 0005, Zhirui Wang 0003, Ruoxin Zhu, Junshi Xia, Xian Sun 0001, Rubén Fernández-Beltran, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Multisource Data Reconstruction-Based Deep Unsupervised Hashing for Unisource Remote Sensing Image RetrievalabstractUnsupervised hashing for remote sensing (RS) image retrieval first extracts image features and then use these features to construct supervised information (e.g., pseudo-labels) to train hashing networks. Existing methods usually regard RS images as natural images to extract unisource features. However, these features only contain partial information about ground objects and cannot produce reliable pseudo-labels. In addition, existing methods only generate a pseudo single-label to annotate each RS image, which cannot accurately represent multiple scenes in a RS image. To address these drawbacks, this paper proposes a new Multisource data reconstruction-based deep unsupervised Hashing method, called MrHash, which explores the characteristics of RS images to construct reliable pseudo-labels. In particular, we first use geographic coordinates to obtain different satellite images and develop a novel autoencoder network to extract multisource features from these images. Then pseudo multi-labels are designed to deal with the coexistence of multiple scenes in a single image. These labels are generated by a custom probability function with extracted multisource features. Finally, we propose a novel multi-semantic hash loss by using the Kull-back–Leibler (KL) divergence to preserve the semantic similarity of these pseudo multi-labels in Hamming space. Our newly developed MrHash only uses multisource images to construct supervised information, and hash code generation still relies on a unisource input image. Experiments on benchmark datasets clearly show the superiority of the proposed method over state-of-the-art baselines. https://github.com/sunyuxi/MrHash. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Xutao Li 0003, Bowen Zhang 0005, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Robust Deep Metric Learning for Remote Sensing Images with Noisy AnnotationsabstractManual and automatic annotation of Remote Sensing (RS) scenes are rather complex tasks which may unavoidably introduce some degree of mislabeled data in large-scale archives. In this regard, noisy annotations become an important constraint for deep metric learning-based RS characterization methods since most of them are trained in a supervised way. To address this problem, here we investigate the use of deep metric learning for characterizing RS scenes with noisy labels. Specifically, we consider the Normalized Softmax Loss and develop a robust extension, i.e., the Robust Normalized Softmax Loss (RNSL), in order to effectively capture the semantic relationships among RS scenes with mislabeled ground-truth information. The conducted experiments, using the K-NN classifier and two benchmark RS image archives, show the potential of the proposed approach with respect to other state-of-the-art methods. Jian Kang 0005, Rubén Fernández-Beltran, Puhong Duan, Xudong Kang, Antonio Plaza |
IGARSS | 2 |
| 2021 | Generalized Scalable Neighborhood Component Analysis for Single and Multi-Label Remote Sensing Image CharacterizationabstractDeep metric learning has recently become a prominent technology for the semantic understanding of remote sensing (RS) scenes due to its great potential for characterizing visual semantics. However, state-of-the-art deep metric learning models are often constrained in RS by the use of single-label annotations, which eventually reduce their capacity to characterize complex aerial scenes. Additionally, many of the existing works are specialized in particular RS applications which constrains the study of their associated metric spaces from a multi-task perspective. In this paper, we propose a new unified deep metric learning approach for both single- and multi-label RS scene characterization while also taking into account different downstream RS applications. Specifically, we extend the Scalable Neighborhood Component Analysis (SNCA) to the multi-label case and propose its generalized version, i.e., GSNCA. Extensive experiments on single- and multi-label RS benchmark datasets have been conducted to evaluate the effectiveness of the proposed method for RS image classification, clustering and retrieval. Jian Kang 0005, Rubén Fernández-Beltran, Antonio Plaza |
IGARSS | 2 |
| 2021 | Sentinel-3 Image Super-Resolution Using Data Fusion and Convolutional Neural NetworksabstractWith the increasing availability of Sentinel-2 (S2) and Sentinel-3 (S3) data, developing higher-level data products becomes a very attractive option to relieve the spatial limitations of the Ocean and Land Colour Instrument (OLCI) of S3. In this context, this paper investigates the suitability of super-resolving operational OLCI products using the Multi-Spectral Instrument (MSI) of S2 as an offline spatial reference. Specifically, the proposed approach assembles a multi -spectral data fusion scheme together with a convolutinal neural network (CNN) mapping function to project the OLCI sensor onto its corresponding spatial reference which is synthetically generated by the OLCI/MSI fusion. In this way, the trained model is able to super-resolve operational OLCI products under demand without the need of using MSI data. The experimental part of the work shows the suitability of the proposed approach in the context of the Copernicus programme. Rafael Fernandez, Rubén Fernández-Beltran, Filiberto Pla |
IGARSS | 2 |
| 2021 | A Remote Sensing Image Registration Benchmark for Operational Sentinel-2 and Sentinel-3 ProductsabstractImage registration is an essential task in image processing, where the final objective is to align geometrically two or more images. In Remote Sensing this process allows to compare, fusion or analyse data. For this purpose, different methods and techniques have been proposed. In this paper a selection of image registration methods has been compared performing inter-sensor registration between Sentinel-2 (S2) and Sentinel-3 (S3) operational data. Registration between S2 and S3 data is not trivial, as there are considerable spectral-spatial differences among them. Nevertheless, the resolution difference results in S2 products being a convenient reference to improve S3 products spatially. The experimentation has been done using four sample pairs of S2 and S3 operational data and representatives of the main registration algorithms used in the last years. Performance measures and results are shown and discussed to check the accuracy and quality of the selected methods. Damian Ibañez, Rubén Fernández-Beltran, Filiberto Pla |
IGARSS | 2 |
| 2021 | Unsupervised Remote Sensing Image Retrieval Using Probabilistic Latent Semantic HashingabstractUnsupervised hashing methods have attracted considerable attention in large-scale remote sensing (RS) image retrieval, due to their capability for massive data processing with significantly reduced storage and computation. Although existing unsupervised hashing methods are suitable for operational applications, they exhibit limitations when accurately modeling the complex semantic content present in RS images using binary codes (in an unsupervised manner). To address this problem, in this letter, we introduce a novel unsupervised hashing method that takes advantage of the generative nature of probabilistic topic models to encapsulate the hidden semantic patterns of the data into the final binary representation. Specifically, we introduce a new probabilistic latent semantic hashing (pLSH) model to effectively learn the hash codes using three main steps: 1) data grouping, where the input RS archive is clustered into several groups; 2) topic computation, where the pLSH model is used to uncover highly descriptive hidden patterns from each group; and 3) hash code generation, where the data probability distributions are thresholded to generate the final binary codes. Our experimental results, obtained on two benchmark archives, reveal that the proposed method significantly outperforms state-of-the-art unsupervised hashing methods. Rubén Fernández-Beltran, Begüm Demir, Filiberto Pla, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Robust Normalized Softmax Loss for Deep Metric Learning-Based Characterization of Remote Sensing Images With Label NoiseabstractMost deep metric learning-based image characterization methods exploit supervised information to model the semantic relations among the remote sensing (RS) scenes. Nonetheless, the unprecedented availability of large-scale RS data makes the annotation of such images very challenging, requiring automated supportive processes. Whether the annotation is assisted by aggregation or crowd-sourcing, the RS large-variance problem, together with other important factors [e.g., geo-location/registration errors, land-cover changes, even low-quality Volunteered Geographic Information (VGI), etc.] often introduce the so-called label noise, i.e., semantic annotation errors. In this article, we first investigate the deep metric learning-based characterization of RS images with label noise and propose a novel loss formulation, named robust normalized softmax loss (RNSL), for robustly learning the metrics among RS scenes. Specifically, our RNSL improves the robustness of the normalized softmax loss (NSL), commonly utilized for deep metric learning, by replacing its logarithmic function with the negative Box–Cox transformation in order to down-weight the contributions from noisy images on the learning of the corresponding class prototypes. Moreover, by truncating the loss with a certain threshold, we also propose a truncated robust normalized softmax loss (t-RNSL) which can further enforce the learning of class prototypes based on the image features with high similarities between them, so that the intraclass features can be well grouped and interclass features can be well separated. Our experiments, conducted on two benchmark RS data sets, validate the effectiveness of the proposed approach with respect to different state-of-the-art methods in three different downstream applications (classification, clustering, and retrieval). The codes of this article will be publicly available fromhttps://github.com/jiankang1991. Jian Kang 0005, Rubén Fernández-Beltran, Puhong Duan, Xudong Kang, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Deep Unsupervised Embedding for Remotely Sensed Images Based on Spatially Augmented Momentum ContrastabstractConvolutional neural networks (CNNs) have achieved great success when characterizing remote sensing (RS) images. However, the lack of sufficient annotated data (together with the high complexity of the RS image domain) often makes supervised and transfer learning schemes limited from an operational perspective. Despite the fact that unsupervised methods can potentially relieve these limitations, they are frequently unable to effectively exploit relevant prior knowledge about the RS domain, which may eventually constrain their final performance. In order to address these challenges, this article presents a new unsupervised deep metric learning model, called spatially augmented momentum contrast (SauMoCo), which has been specially designed to characterize unlabeled RS scenes. Based on the first law of geography, the proposed approach defines spatial augmentation criteria to uncover semantic relationships among land cover tiles. Then, a queue of deep embeddings is constructed to enhance the semantic variety of RS tiles within the considered contrastive learning process, where an auxiliary CNN model serves as an updating mechanism. Our experimental comparison, including different state-of-the-art techniques and benchmark RS image archives, reveals that the proposed approach obtains remarkable performance gains when characterizing unlabeled scenes since it is able to substantially enhance the discrimination ability among complex land cover categories. The source codes of this article will be made available to the RS community for reproducible research. Jian Kang 0005, Rubén Fernández-Beltran, Puhong Duan, Sicong Liu 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Graph Relation Network: Modeling Relations Between Scenes for Multilabel Remote-Sensing Image Classification and RetrievalabstractDue to the proliferation of large-scale remote-sensing (RS) archives with multiple annotations, multilabel RS scene classification and retrieval are becoming increasingly popular. Although some recent deep learning-based methods are able to achieve promising results in this context, the lack of research on how to learn embedding spaces under the multilabel assumption often makes these models unable to preserve complex semantic relations pervading aerial scenes, which is an important limitation in RS applications. To fill this gap, we propose a new graph relation network (GRN) for multilabel RS scene categorization. Our GRN is able to model the relations between samples (or scenes) by making use of a graph structure which is fed into network learning. For this purpose, we define a new loss function called scalable neighbor discriminative loss with binary cross entropy (SNDL-BCE) that is able to embed the graph structures through the networks more effectively. The proposed approach can guide deep learning techniques (such as convolutional neural networks) to a more discriminative metric space, where semantically similar RS scenes are closely embedded and dissimilar images are separated from a novel multilabel viewpoint. To achieve this goal, our GRN jointly maximizes a weighted leave-one-out K-nearest neighbors ( KNN) score in the training set, where the weight matrix describes the contributions of the nearest neighbors associated with each RS image on its class decision, and the likelihood of the class discrimination in the multilabel scenario. An extensive experimental comparison, conducted on three multilabel RS scene data archives, validates the effectiveness of the proposed GRN in terms of KNN classification and image retrieval. The codes of this article will be made publicly available for reproducible research in the community. Jian Kang 0005, Rubén Fernández-Beltran, Danfeng Hong, Jocelyn Chanussot, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Sentinel-2 Multi-Temporal Data for Rice Crop Classification in NepalabstractThe global coverage of Sentinel-2 provides widespread opportunities for accurately mapping and monitoring key crops in emerging countries, like in the case of Nepal's rice production. While previous studies based on other satellites show some important spatial and temporal limitations, the use of operational Sentinel-2 data still remains unexplored in this regard. As a result, this work investigates the viability of using the Sentinel-2 instrument for a precise rice crop classification in Nepal. Initially, we define a dataset made of multi-temporal Sentinel-2 data from the Terai region of Nepal. Then, we conduct several classification experiments to provide empirical evidences about the suitability of different classification models when identifying rice crops in developing countries, where only limited ground-truth data could be available. The experiments reveal the suitability of using Sentinel-2 for accurately mapping rice crops in Nepal with a CNN-based classification model. Tina Baidar, Rubén Fernández-Beltran, Filiberto Pla |
IGARSS | 2 |
| 2020 | Inter-Sensor Remote Sensing Image Enhancement for Operational Sentinel-2 and Sentinel-3 Data ProductsabstractThe recent availability of operational data from the Sentinel-2 and Sentinel-3 missions provides widespread opportunities to generate diverse high-level remote sensing products. However, the synergies between both multi-spectral instruments are often difficult to exploit from an operational perspective. Standard pansharpening algorithms may encounter important disadvantages due to the limited intersensor data availability in actual production environments. Moreover, the lack of a real high-resolution ground-truth for super-resolution techniques may affect the radiometric quality of the final result. In this scenario, this work investigates the viability of using the Multi-Spectral Instrument of Sentinel-2 for super-resolving data products acquired by the Ocean and Land Colour Instrument of Sentinel-3. Specifically, we define an inter-sensor image enhancement framework which combines a PCA-based component substitution pansharpening scheme with a CNN-based spatial enhancing super-resolution mapping. The conducted experiments reveal the suitability of the proposed approach for generating Level-4 data products within the Copernicus programme context. Rafael Fernandez, Rubén Fernández-Beltran, Filiberto Pla |
IGARSS | 2 |
| 2020 | Endmember Extraction From Hyperspectral Imagery Based on Probabilistic Tensor MomentsabstractThis letter presents a novel hyperspectral endmember extraction approach that integrates a tensor-based decomposition scheme with a probabilistic framework in order to take advantage of both technologies when uncovering the signatures of pure spectral constituents in the scene. On the one hand, statistical unmixing models are generally able to provide accurate endmember estimates by means of rather complex optimization algorithms. On the other hand, tensor decomposition techniques are very effective factorization tools which are often constrained by the lack of physical interpretation within the remote sensing field. In this context, this letter develops a new hybrid endmember extraction approach based on the decomposition of the probabilistic tensor moments of the hyperspectral data. Initially, the input image reflectance values are modeled as a collection of multinomial distributions provided by a family of Dirichlet generalized functions. Then, the unmixing process is effectively conducted by the tensor decomposition of the third-order probabilistic tensor moments of the multivariate data. Our experiments, conducted over four hyperspectral data sets, reveal that the proposed approach is able to provide efficient and competitive results when compared to different state-of-the-art endmember extraction methods. Rubén Fernández-Beltran, Filiberto Pla, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Deep Metric Learning Based on Scalable Neighborhood Components for Remote Sensing Scene CharacterizationabstractWith the development of convolutional neural networks (CNNs), the semantic understanding of remote sensing (RS) scenes has been significantly improved based on their prominent feature encoding capabilities. While many existing deep-learning models focus on designing different architectures, only a few works in the RS field have focused on investigating the performance of the learned feature embeddings and the associated metric space. In particular, two main loss functions have been exploited: the contrastive and the triplet loss. However, the straightforward application of these techniques to RS images may not be optimal in order to capture their neighborhood structures in the metric space due to the insufficient sampling of image pairs or triplets during the training stage and to the inherent semantic complexity of remotely sensed data. To solve these problems, we propose a new deep metric learning approach, which overcomes the limitation on the class discrimination by means of two different components: 1) scalable neighborhood component analysis (SNCA) that aims at discovering the neighborhood structure in the metric space and 2) the cross-entropy loss that aims at preserving the class discrimination capability based on the learned class prototypes. Moreover, in order to preserve feature consistency among all the minibatches during training, a novel optimization mechanism based on momentum update is introduced for minimizing the proposed loss. An extensive experimental comparison (using several state-of-the-art models and two different benchmark data sets) has been conducted to validate the effectiveness of the proposed method from different perspectives, including: 1) classification; 2) clustering; and 3) image retrieval. The related codes of this article will be made publicly available for reproducible research by the community. Jian Kang 0005, Rubén Fernández-Beltran, Zhen Ye 0009, Xiaohua Tong, Pedram Ghamisi, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Open Multi-Processing Acceleration for Unsupervised Land Cover Categorization Using Probabilistic Latent Semantic AnalysisabstractThe probabilistic Latent Semantic Analysis (pLSA) model has recently shown a great potential to uncover highly descriptive semantic features from limited amounts of remote sensing data. Nonetheless, the high computational cost of this algorithm often constraints its operational application for land cover categorization tasks. In this scenario, this paper presents an Open Multi-Processing (OpenMP) implementation of the pLSA algorithm for unsupervised Synthetic Aperture Radar (SAR) and Multi-Spectral Imaging (MSI) image categorization. The experimental results suggest that multi-core systems are an important architecture for the efficient processing of both SAR and MSI datasets. Specifically, the proposed approach is able to cover a real scenario exhibiting good results in both accuracy and performance terms. Sergio Bernabé, Carlos García 0001, Rubén Fernández-Beltran, Mercedes Eugenia Paoletti, Juan Mario Haut, Javier Plaza, Antonio Plaza |
IGARSS | 3 |
| 2019 | Sentinel-2 and Sentinel-3 Intersensor Vegetation Estimation via Constrained Topic ModelingabstractThis letter presents a novel intersensor vegetation estimation framework, which aims at combining Sentinel-2 (S2) spatial resolution with Sentinel-3 (S3) spectral characteristics in order to generate fused vegetation maps. On the one hand, the multispectral instrument (MSI), carried by S2, provides high spatial resolution images. On the other hand, the Ocean and Land Color Instrument (OLCI), one of the instruments of S3, captures the Earth's surface at a substantially coarser spatial resolution but using smaller spectral bandwidths, which makes the OLCI data more convenient to highlight specific spectral features and motivates the development of synergetic fusion products. In this scenario, the approach presented here takes advantage of the proposed constrained probabilistic latent semantic analysis (CpLSA) model to produce intersensor vegetation estimations, which aim at synergically exploiting MSI's spatial resolution and OLCI's spectral characteristics. Initially, CpLSA is used to uncover the MSI reflectance patterns, which are able to represent the OLCI-derived vegetation. Then, the original MSI data are projected onto this higher abstraction-level representation space in order to generate a high-resolution version of the vegetation captured in the OLCI domain. Our experimental comparison, conducted using four data sets, three different regression algorithms, and two vegetation indices, reveals that the proposed framework is able to provide a competitive advantage in terms of quantitative and qualitative vegetation estimation results. Rubén Fernández-Beltran, Filiberto Pla, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Intersensor Remote Sensing Image Registration Using Multispectral Semantic EmbeddingsabstractThis letter presents a novel intersensor registration framework specially designed to register Sentinel-3 (S3) operational data using the Sentinel-2 (S2) instrument as a reference. The substantially higher resolution of the Multispectral Instrument (MSI), on-board S2, with respect to the Ocean and Land Color Instrument (OLCI), carried by S3, makes the former sensor a suitable spatial reference to finely adjust OLCI products. Nonetheless, the important spectral-spatial differences between both instruments may constrain traditional registration mechanisms to effectively align data of such different nature. In this context, the proposed registration scheme advocates the use of a topic model-based embedding approach to conduct the intersensor registration task within a common multispectral semantic space, where the input imagery is represented according to their corresponding spectral feature patterns instead of the low-level attributes. Thus, the OLCI products can be effectively registered to the MSI reference data by aligning those hidden patterns that fundamentally express the same visual concepts across the sensors. The experiments, conducted over four different S2 and S3 operational data collections, reveal that the proposed approach provides performance advantages over six different intersensor registration counterparts. Rubén Fernández-Beltran, Filiberto Pla, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Low-High-Power Consumption Architectures for Deep-Learning Models Applied to Hyperspectral Image ClassificationabstractConvolutional neural networks have emerged as an excellent tool for remotely sensed hyperspectral image (HSI) classification. Nonetheless, the high computational complexity and energy requirements of these models typically limit their application in on-board remote sensing scenarios. In this context, low-power consumption architectures are promising platforms that may provide acceptable on-board computing capabilities to achieve satisfactory classification results with reduced energy demand. For instance, the new NVIDIA Jetson Tegra TX2 device is an efficient solution for on-board processing applications using deep-learning (DL) approaches. So far, very few efforts have been devoted to exploiting this or other similar computing platforms in on-board remote sensing procedures. This letter explores the use of low-power consumption architectures and DL algorithms for HSI classification. The conducted experimental study reveals that the NVIDIA Jetson Tegra TX2 device offers a good choice in terms of performance, cost, and energy consumption for on-board HSI classification tasks. Juan Mario Haut, Sergio Bernabé, Mercedes Eugenia Paoletti, Rubén Fernández-Beltran, Antonio Plaza, Javier Plaza |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Remote Sensing Single-Image Superresolution Based on a Deep Compendium ModelabstractThis letter introduces a novel remote sensing singleimage superresolution (SR) architecture based on a deep efficient compendium model. The current deep learning-based SR trend stands for using deeper networks to improve the performance. However, this practice often results in the degradation of visual results. To address this issue, the proposed approach harmonizes several different improvements on the network design to achieve state-of-the-art performance when superresolving remote sensing imagery. On the one hand, the proposal combines residual units and skip connections to extract more informative features on both local and global image areas. On the other hand, it makes use of parallelized 1×1 convolutional filters (network in network) to reconstruct the superresolved result while reducing the information loss through the network. Our experiments, conducted using seven different SR methods over the well-known UC Merced remote sensing data set, and two additional GaoFen-2 test images, show that the proposed model is able to provide competitive advantages. Juan Mario Haut, Mercedes Eugenia Paoletti, Rubén Fernández-Beltran, Javier Plaza, Antonio Plaza, Jun Li 0009 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Remote Sensing Image Superresolution Using Deep Residual Channel AttentionabstractThe current trend in remote sensing image superresolution (SR) is to use supervised deep learning models to effectively enhance the spatial resolution of airborne and satellite-based optical imagery. Nonetheless, the inherent complexity of these architectures/data often makes these methods very difficult to train. Despite these recent advances, the huge amount of network parameters that must be fine-tuned and the lack of suitable high-resolution remotely sensed imagery in actual operational scenarios still raise some important challenges that may become relevant limitations in the existent earth observation data production environments. To address these problems, we propose a new remote sensing SR approach that integrates a visual attention mechanism within a residual-based network design in order to allow the SR process to focus on those features extracted from land-cover components that require more computations to be superresolved. As a result, the network training process is significantly improved because it aims at learning the most relevant high-frequency information while the proposed architecture allows neglecting the low-frequency features extracted from spatially uninformative earth surface areas by means of several levels of skip connections. Our experimental assessment, conducted using the University of California at Merced and GaoFen-2 remote sensing image collections, three scaling factors, and eight different SR methods, demonstrates that our newly proposed approach exhibits competitive performance in the task of superresolving remotely sensed imagery. Juan Mario Haut, Rubén Fernández-Beltran, Mercedes Eugenia Paoletti, Javier Plaza, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Capsule Networks for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have recently exhibited an excellent performance in hyperspectral image classification tasks. However, the straightforward CNN-based network architecture still finds obstacles when effectively exploiting the relationships between hyperspectral imaging (HSI) features in the spectral-spatial domain, which is a key factor to deal with the high level of complexity present in remotely sensed HSI data. Despite the fact that deeper architectures try to mitigate these limitations, they also find challenges with the convergence of the network parameters, which eventually limit the classification performance under highly demanding scenarios. In this paper, we propose a new CNN architecture based on spectral-spatial capsule networks in order to achieve a highly accurate classification of HSIs while significantly reducing the network design complexity. Specifically, based on Hinton's capsule networks, we develop a CNN model extension that redefines the concept of capsule units to become spectral-spatial units specialized in classifying remotely sensed HSI data. The proposed model is composed by several building blocks, called spectral-spatial capsules, which are able to learn HSI spectral-spatial features considering their corresponding spatial positions in the scene, their associated spectral signatures, and also their possible transformations. Our experiments, conducted using five well-known HSI data sets and several state-of-the-art classification methods, reveal that our HSI classification approach based on spectral-spatial capsules is able to provide competitive advantages in terms of both classification accuracy and computational time. Mercedes Eugenia Paoletti, Juan Mario Haut, Rubén Fernández-Beltran, Javier Plaza, Antonio Plaza, Jun Li 0009, Filiberto Pla |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Deep Pyramidal Residual Networks for Spectral-Spatial Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) exhibit good performance in image processing tasks, pointing themselves as the current state-of-the-art of deep learning methods. However, the intrinsic complexity of remotely sensed hyperspectral images still limits the performance of many CNN models. The high dimensionality of the HSI data, together with the underlying redundancy and noise, often makes the standard CNN approaches unable to generalize discriminative spectral-spatial features. Moreover, deeper CNN architectures also find challenges when additional layers are added, which hampers the network convergence and produces low classification accuracies. In order to mitigate these issues, this paper presents a new deep CNN architecture specially designed for the HSI data. Our new model pursues to improve the spectral-spatial features uncovered by the convolutional filters of the network. Specifically, the proposed residual-based approach gradually increases the feature map dimension at all convolutional layers, grouped in pyramidal bottleneck residual blocks, in order to involve more locations as the network depth increases while balancing the workload among all units, preserving the time complexity per layer. It can be seen as a pyramid, where the deeper the blocks, the more feature maps can be extracted. Therefore, the diversity of high-level spectral-spatial attributes can be gradually increased across layers to enhance the performance of the proposed network with the HSI data. Our experiments, conducted using four well-known HSI data sets and 10 different classification techniques, reveal that our newly developed HSI pyramidal residual model is able to provide competitive advantages (in terms of both classification accuracy and computational time) over the state-of-the-art HSI classification methods. Mercedes Eugenia Paoletti, Juan Mario Haut, Rubén Fernández-Beltran, Javier Plaza, Antonio Plaza, Filiberto Pla |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Inter-Sensor Regression Analysis for Operational Sentinel-2 and Sentinel-3 Data ProductsabstractThe relatively recent availability of operational products from Sentinel-2 and Sentinel-3 missions gives widespread opportunities to combine data collected from different sensors in order to provide products of a higher processing level. Nonetheless, the availability of these products may be affected by multiple factors, such as cloud occlusions, band saturation, geolocation errors or even misaligned detectors. All these anomalies affecting remote sensing data may eventually limit the accessibility to fused products because some of the required information may become partially unavailable for specific areas of interest. In this scenario, the work presented here aims at analyzing the effectiveness of several state-of-the-art regression models in order to restore Sentinel-3 products with partial anomalies from Sentinel-2 integral data. In particular this work investigates three regression methods, two linear-regression method and a non-linear artificial neural networks based method. Obtained results prove that the nonlinear approach and linear RIDGE method are able to carry out a good estimation of S3 from S2 data. Juan Mario Haut, Rubén Fernández-Beltran, Mercedes Eugenia Paoletti, Javier Plaza, Antonio Plaza, Filiberto Pla |
IGARSS | 2 |
| 2018 | Multimodal Probabilistic Latent Semantic Analysis for Sentinel-1 and Sentinel-2 Image FusionabstractProbabilistic topic models have recently shown a great potential in the remote sensing image fusion field, which is particularly helpful in land-cover categorization tasks. This letter first studies the application of probabilistic latent semantic analysis (pLSA) and latent Dirichlet allocation to remote sensing synthetic aperture radar (SAR) and multispectral imaging (MSI) unsupervised land-cover categorization. Then, a novel pLSA-based image fusion approach is presented, which pursues to uncover multimodal feature patterns from SAR and MSI data in order to effectively fuse and categorize Sentinel-1 and Sentinel-2 remotely sensed data. Experiments conducted over two different data sets reveal the advantages of the proposed approach for unsupervised land-cover categorization tasks. Rubén Fernández-Beltran, Juan Mario Haut, Mercedes Eugenia Paoletti, Javier Plaza, Antonio Plaza, Filiberto Pla |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Prior-based probabilistic latent semantic analysis for multimedia retrieval
Rubén Fernández-Beltran, Filiberto Pla |
Multim. Tools Appl. | 1 |
| 2018 | Sparse multi-modal probabilistic latent semantic analysis for single-image super-resolution
Rubén Fernández-Beltran, Filiberto Pla |
Signal Process. | 1 |
| 2018 | Hyperspectral Unmixing Based on Dual-Depth Sparse Probabilistic Latent Semantic AnalysisabstractThis paper presents a novel approach for spectral unmixing of remotely sensed hyperspectral data. It exploits probabilistic latent topics in order to take advantage of the semantics pervading the latent topic space when identifying spectral signatures and estimating fractional abundances from hyperspectral images. Despite the contrasted potential of topic models to uncover image semantics, they have been merely used in hyperspectral unmixing as a straightforward data decomposition process. This limits their actual capabilities to provide semantic representations of the spectral data. The proposed model, called dual-depth sparse probabilistic latent semantic analysis (DEpLSA), makes use of two different levels of topics to exploit the semantic patterns extracted from the initial spectral space in order to relieve the ill-posed nature of the unmixing problem. In other words, DEpLSA defines a first level of deep topics to capture the semantic representations of the spectra, and a second level of restricted topics to estimate endmembers and abundances over this semantic space. An experimental comparison in conducted using the two standard topic models and the seven state-of-the-art unmixing methods available in the literature. Our experiments, conducted using four different hyperspectral images, reveal that the proposed approach is able to provide competitive advantages over available unmixing approaches. Rubén Fernández-Beltran, Antonio Plaza, Javier Plaza, Filiberto Pla |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | A New Deep Generative Network for Unsupervised Remote Sensing Single-Image Super-ResolutionabstractSuper-resolution (SR) brings an excellent opportunity to improve a wide range of different remote sensing applications. SR techniques are concerned about increasing the image resolution while providing finer spatial details than those captured by the original acquisition instrument. Therefore, SR techniques are particularly useful to cope with the increasing demand remote sensing imaging applications requiring fine spatial resolution. Even though different machine learning paradigms have been successfully applied in SR, more research is required to improve the SR process without the need of external high-resolution (HR) training examples. This paper proposes a new convolutional generator model to super-resolve low-resolution (LR) remote sensing data from an unsupervised perspective. That is, the proposed generative network is able to initially learn relationships between the LR and HR domains throughout several convolutional, downsampling, batch normalization, and activation layers. Then, the data are symmetrically projected to the target resolution while guaranteeing a reconstruction constraint over the LR input image. An experimental comparison is conducted using 12 different unsupervised SR methods over different test images. Our experiments reveal the potential of the proposed approach to improve the resolution of remote sensing imagery. Juan Mario Haut, Rubén Fernández-Beltran, Mercedes Eugenia Paoletti, Javier Plaza, Antonio Plaza, Filiberto Pla |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Latent topics-based relevance feedback for video retrieval
Rubén Fernández-Beltran, Filiberto Pla |
Pattern Recognit. | 1 |
| 2015 | Incremental probabilistic Latent Semantic Analysis for video retrieval
Rubén Fernández-Beltran, Filiberto Pla |
Image Vis. Comput. | 1 |