Claudio Persello

dblp:42/8952 · DBLP profile ↗
← Back
46ranked-venue papers
15as first author
13since 2021 · last 2025
0000-0003-3742-5398ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 46 · 15 first-author · 13 since 2021
YearPublicationVenuePosition
2025 Deep Merge: Deep-Learning-Based Region Merging for Remote Sensing Image Segmentation
abstract
Image segmentation represents a fundamental step in analyzing very high-spatial-resolution (VHR) remote sensing imagery. Its objective is to partition an image into segments that best match with geo-objects. However, the diverse appearances of geospatial objects often lead to interobject homogeneity and intraobject heterogeneity. Existing segmentation methods often struggle to accurately segment geo-objects with varying shapes and scales. To address these challenges, we propose DeepMerge, a novel method that integrates deep learning and region adjacency graphs (RAGs) to accurately segment complete geo-objects in large VHR images. DeepMerge begins with an initial over-segmentation of the image and then iteratively merges similar regions to achieve complete geo-object segmentation. A deep learning model is employed to learn the similarity between adjacent superpixel pairs. This approach only requires labels indicating whether adjacent superpixels belong to the same geo-object eliminating the need for object-level annotations, enabling weakly supervised segmentation. A cross-scale module is incorporated to capture multiscale information, enhancing the representation of superpixels. In addition, the feature distances between neighboring super-pixels are deemed as scale parameters (thresholds) to control the merging procedure, thus yielding an interpretable, predictable, stable, and optimal scale parameter 0.5. DeepMerge can achieve high segmentation accuracy in a weakly supervised manner, which is validated on large-scale remote sensing images of 0.55-m resolution covering an area of 5660 km2. The experimental results demonstrate that DeepMerge achieves the highest F value (0.9552) and the lowest total error (TE) (0.0827), accurately segmenting geo-objects of varying sizes and outperforming all competing methods.
Xianwei Lv 0002, Claudio Persello, Wangbin Li, Xiao Huang 0003, Dongping Ming, Alfred Stein
IEEE Trans. Geosci. Remote. Sens.2
2024 Glacier Mapping from Sentinel-1 SAR Time Series with Deep Learning in Svalbard
abstract
Glaciers are one of the essential climate variables. Tracking their areal changes over time is of high importance for monitoring the impacts of climate change and designing adaptation strategies. Mapping glaciers from optical remote sensing data might result in a very limited temporal resolution due to the absence of cloud-free imagery at the end of the ablation season. Synthetic aperture radar (SAR) solves this problem as it can operate in almost all weather conditions. Here, we present a deep learning strategy for glacier mapping based solely on Sentinel-1 SAR data in Svalbard. We test two options for integrating SAR image time series into deep learning models, namely, 3D convolutions and long short-term memory (LSTM) cells. Both proposed models achieve an intersection over union (IoU) of 0.964 on the test subset. Our results highlight the applicability of SAR data in glacier mapping with the potential to obtain glacier inventories with higher temporal resolution. We shared our dataset, code-base and pretrained models at https://github.com/konstantin-a-maslov/icemapper.
Konstantin A. Maslov, Thomas Schellenberger, Claudio Persello, Alfred Stein
IGARSS3
2024 User and Data-Centric Artificial Intelligence for Mapping Urban Deprivation in Multiple Cities Across the Globe
abstract
The rapid urbanization in many regions worldwide results in the proliferation of deprived urban areas, also known as slums or informal settlements. Our study addresses the pressing need for accurate information by investigating User and Data-centric Artificial Intelligence (AI)-based methods for mapping deprived urban areas and extracting information supporting the Sustainable Development Goals (SDG) Indicator 11.1.1. In collaboration with local communities and several (inter)national stakehlders, we co-designed AI strategies based on free or low-cost Earth Observation (EO) and geospatial data to map informal settlements in eight cties across the globe. The AI methods design, data collection, and validation strategies follow an iterative and agile process consisting of progressive refinement stages necessary to collect reliable labeled data and take user requirements into their centre. Our findings indicate that the combination of Sentinel-2 and morphometric features yields the most accurate results.
Bedru Tareke, Paulo Silva Filho, Claudio Persello, Monika Kuffer, Raian Vargas Maretto, Jon Wang, Ángela Abascal, Priam V. Pillai, Binti Singh, Juan Manuel D'Attoli, Caroline Kabaria, Julio Cesar Pedrassoli, Patricia Lustosa Brito, Peter Elias 0002, Elio Atenógenes, Andrea Ramírez Santiago
IGARSS3
2024 Building Extraction From Very High-Resolution Remote Sensing Image With Few Data
abstract
Deep neural networks have shown remarkable progress in building extraction. However, their effectiveness is limited due to their dependence on data and dense annotations. It remains challenging to extract buildings with few annotated samples when dealing with variable shapes, sizes, and textures amid an insufficient number of labels. Motivated by the aforementioned challenge, this letter proposes a novel approach called meta rectification networks (MRNs) to extract completely new types of buildings using minimal data accurately. We employed a learning strategy for each specific building extraction task where few annotated building images are used to generate feature representations for differentiating between building and nonbuilding pixels. Our methodology involves matching each pixel to the learned features to detect unknown buildings effectively. Since the feature representations generated from few annotated building images have a large inductive bias, we design a pseudolabel rectification (PLR) mechanism to reduce this inductive bias and enhance the representational power of the feature representations. We evaluate the effectiveness of our model in five different regions. Our experiments show that our proposed model achieves mean-IoU values of 65.54% and overall accuracy of 74.46% while using only 1% of the training data. Our findings suggest that MRN can accurately detect buildings even when trained with limited data, thus providing an effective solution to overcome the constraint of insufficient annotations present in building-extraction problems.
Zhenqi Cui, Pei Nie, Claudio Persello
IEEE Geosci. Remote. Sens. Lett.3
2024 Recent Advances in Machine Learning for Remote Sensing Toward the Sustainable Development Goals
Ujjwal Verma, Dalton D. Lunga, Ronny Hänsch, Claudio Persello, Silvia Liberata Ullo
IEEE Geosci. Remote. Sens. Lett.4
2023 Investigating Sar-Optical Deep Learning Data Fusion to Map the Brazilian Cerrado Vegetation with Sentinel Data
abstract
Despite its environmental and societal importance, accurately mapping the Brazilian Cerrado’s vegetation is still an open challenge. Its diverse but spectrally similar physiognomies are difficult to be identified and mapped by state-of-the-art methods from only medium-to high-resolution optical images. This work investigates the fusion of Synthetic Aperture Radar (SAR) and optical data in convolutional neural network architectures to map the Cerrado according to a 2-level class hierarchy. Additionally, the proposed model is designed to deal with uncertainties that are brought by the difference in resolution between the input images (at 10m) and the reference data (at 30m). We tested four data fusion strategies and showed that the position for the data combination is important for the network to learn better features.
Paulo Silva Filho, Claudio Persello, Raian Vargas Maretto, Renato Machado
IGARSS2
2023 GLAVITU: A Hybrid CNN-Transformer for Multi-Regional Glacier Mapping from Multi-Source Data
abstract
Glacier mapping is essential for studying and monitoring the impacts of climate change. However, several challenges such as debris-covered ice and highly variable landscapes across glacierized regions worldwide complicate large-scale glacier mapping in a fully-automated manner. This work presents a novel hybrid CNN-transformer model (GlaViTU) for multi-regional glacier mapping. Our model outperforms three baseline models—SETR-B/16, ResU-Net and TransU-Net—achieving a higher mean IoU of 0.875 and demonstrates better generalization ability. The proposed model is also parameter-efficient, with approximately 10 and 3 times fewer parameters than SETR-B/16 and ResU-Net, respectively. Our results provide a solid foundation for future studies on the application of deep learning methods for global glacier mapping. To facilitate reproducibility, we have shared our data set, codebase and pretrained models on GitHub at https://github.com/konstantin-a-maslov/GlaViTU-IGARSS2023.
Konstantin A. Maslov, Claudio Persello, Thomas Schellenberger, Alfred Stein
IGARSS2
2023 Extracting Polygons of Visible Cadastral Boundaries Using Deep Learning
abstract
Formal land registration systems are out of reach for most of the world’s population. Conventional mapping methods, such as high-precision ground surveys, are costly, making them inaccessible, especially in low- and middle-income countries. With the introduction of fit-for-purpose land administration, automatic feature extraction techniques have been actively investigated to accelerate the land rights mapping process. Therefore, in our research, we assessed the potential of deep learning to extract cadastral boundaries from very high-resolution images. Our study adopts a multitask learning strategy, which utilizes state of the art U-Net model for the segmentation task, whereas the frame field learning method provides structural information for the subsequent active contour model to produce regularized vector polygons. The experimental results show that the combined U-Net model and frame field information produced polygons with higher accuracy compared to a segmentation method on its own.
Bedru Tareke, Mila Koeva, Claudio Persello
IGARSS3
2023 Vectorizing Planar Roof Structure From Very High Resolution Remote Sensing Images Using Transformers
abstract
Grasping the roof structure of a building is a key part of building reconstruction. Directly predicting the geometric structure of the roof from a raster image to a vectorized representation, however, remains challenging. This paper introduces an efficient and accurate parsing method based upon a vision Transformer we dubbed Roof-Former. Our method consists of three steps: 1) Image encoder and edge node initialization, 2) Image feature fusion with an enhanced segmentation refinement branch, and 3) Edge filtering and structural reasoning. The vertex and edge heat map F1-scores have increased by 2.0% and 1.9% on the VWB dataset when compared to HEAT. Additionally, qualitative evaluations suggest that our method is superior to the current state-of-the-art. It indicates effectiveness for extracting global image information and maintaining the consistency and topological validity of the roof structure.
Wufan Zhao, Claudio Persello, Xianwei Lv 0002, Alfred Stein
IGARSS2
2023 AI4SmallFarms: A Dataset for Crop Field Delineation in Southeast Asian Smallholder Farms
abstract
Agricultural field polygons within smallholder farming systems are essential to facilitate the collection of geo-spatial data useful for farmers, managers, and policymakers. However, the limited availability of training labels poses a challenge in developing supervised methods to accurately delineate field boundaries using Earth Observation (EO) data. This letter introduces an open data set for training and benchmarking machine learning methods to delineate agricultural field boundaries in polygon format. The large-scale data set consists of 439,001 field polygons divided into 62 tiles of approximately 5×5 km distributed across Vietnam and Cambodia, covering a range of fields and diverse landscape types. The field polygons have been meticulously digitized from satellite images, following a rigorous multi-step quality control process and topological consistency checks. Multi-temporal composites of Sentinel-2 (S2) images are provided to ensure cloud-free data. We conducted an experimental analysis testing a state-of-the-art Deep Learning (DL) workflow based on fully convolutional networks, contour closing, and polygonization. We anticipate that this large-scale data set will enable researchers to further enhance the delineation of agricultural fields in smallholder farms and to support the achievement of the Sustainable Development Goals (SDG). The data set can be downloaded from https://doi.org/10.17026/dans-xy6-ngg6.
Claudio Persello, Jeroen Grift, Xinyan Fan, Claudia Paris, Ronny Hänsch, Mila Koeva, Andrew Nelson 0003
IEEE Geosci. Remote. Sens. Lett.1
2022 Despeckling Polarimetric SAR Data Using a Multistream Complex-Valued Fully Convolutional Network
abstract
A polarimetric synthetic aperture radar (PolSAR) sensor is able to collect images in different polarization states, making it a rich source of information for target characterization. PolSAR images are inherently affected by speckle. Therefore, before derivingad hocproducts from the data, the polarimetric covariance matrix needs to be estimated by reducing speckle. In recent years, deep learning-based despeckling methods have started to evolve from single-channel SAR images to PolSAR images. To this aim, deep learning-based approaches separate the real and imaginary components of the complex-valued covariance matrix and use them as independent channels in standard convolutional neural networks (CNNs). However, this approach neglects the mathematical relationship that exists between the real and imaginary components, resulting in suboptimal output. Here, we propose a multistream complex-valued fully convolutional network (FCN) (CV-deSpeckNethttps://github.com/adugnag/CV-deSpeckNet) to reduce speckle and effectively estimate the PolSAR covariance matrix. To evaluate the performance of CV-deSpeckNet, we used Sentinel-1 dual polarimetric SAR images to compare against its real-valued counterpart that separates the real and imaginary parts of the complex covariance matrix. CV-deSpeckNet was also compared against the state of the art PolSAR despeckling methods. The results show that CV-deSpeckNet was able to be trained with a fewer number of samples, has a higher generalization capability, and resulted in higher accuracy than its real-valued counterpart and state-of-the-art PolSAR despeckling methods. These results showcase the potential of complex-valued deep learning for PolSAR despeckling.
Adugna G. Mullissa, Claudio Persello, Johannes Reiche
IEEE Geosci. Remote. Sens. Lett.2
2021 EO-Based Low-Cost Frameworks to Address Global Urban Data GAPS on Deprivation and Multiple Hazards
abstract
A continuously growing number of urban inhabitants in Low- and Middle-income Country (LMIC) cities live in deprived areas. Such areas are under-serviced and characterized by poor living and environmental conditions, where the unplanned morphology interacts with physical hazards. While such areas proliferate, climate change has increasingly severe impacts on them. Deprived areas are often located in high-risk zones, e.g., flood zones. However, the absence of global databases on such areas is an obstacle to locate and prioritize hotspots of deprived communities exposed to climate change. Consequently, quantifying the numbers of exposed inhabitants is not possible, though it is required in support of the Sustainable Development Goals (SDGs) (e.g., 11, 13) and local adaptation strategies. Earth observation (EO) that allows producing such data fall short of providing city-level information due to computational constraints and unsolved methodological challenges related to scalability and transferability. This paper presents a framework to combine EO data with data on hazards (e.g., storms, floods) that impact urban areas and, in particular, deprived communities. First results in pilot cities show that deprived communities are systematically more exposed to physical hazards as compared to formal built-up areas. These problems are expected to intensify in the context of climate change, as most hazards will increase in their severity and frequency.
Monika Kuffer, Dana R. Thomson, Andrew Maki, Sabine Vanhuysse, Stefanos Georganos, Richard Sliuzas, Claudio Persello
IGARSS7
2021 End-to-End Roofline Extraction from Very-High-Resolution Remote Sensing Images
abstract
Roof shape information is essential for creating 3D building models. However, the automated extracting of roof structures from Earth observation data is a difficult task involving significant uncertainties caused by scene complexity and limited multi-source data coverage. This paper introduces the integrally-attracted wireframe parsing (IAWP) framework to reconstruct building rooflines as a planar graph from remotely sensed images with a single forward pass. We add global geometric line priors through the Hough transform into deep networks to better extract the linear geometric features. We perform experiments on the vectorizing world building (VWB) dataset. The investigated method improves the F-score metrics of corner points/edges by 0.1%/7.7% and 0.6%/1.1%, respectively. Visual comparison results also indicate that the HT-IHT block gives consistent improvements in terms of geometric regularity.
Wufan Zhao, Claudio Persello, Alfred Stein
IGARSS2
2020 Towards Uncovering Socio-Economic Inequalities Using VHR Satellite Images and Deep Learning
abstract
In many cities of the Global South, informal and deprived neighborhoods, also commonly called slums, continue to proliferate, but their locations and dwellers' socio-economic status are often invisible in official statistics and maps. Very high resolution (VHR) satellite images coupled with deep learning allow us to efficiently map these areas and study their socio-economic and spatio-temporal variability to support interventions. This paper investigates a deep transfer learning approach based on convolutional neural networks (CNN) to identify the socio-economic variability of poor neighborhoods in Bangalore, India. Our deep network, pre-trained on a slum classification data set, is tuned towards the prediction of a continuous-valued socio-economic index capturing multiple levels of deprivation. Experimental results show that the CNN-based regression model can explain the socio-economic variability with an R2 of 0.75. The use of additional publicly available geographic information layers allow us to spatially extend the analysis beyond the surveyed deprived area data samples to uncover city-wide patterns of socio-economic inequalities.
Claudio Persello, Monika Kuffer
IGARSS1
2020 Building Instance Segmentation and Boundary Regularization from High-Resolution Remote Sensing Images
abstract
Building extraction from remote sensing images using convolutional neural networks (CNNs) has been an active research topic in recent years. Most results obtained by CNN-based algorithms, however, still have common issues with the precision of the delineation of building outlines and the separation of different buildings. Recently, efforts have been made towards the automation of building outline regularization. This paper employs a new instance segmentation framework named Hybrid Task Cascade (HTC) as baseline model, integrating detection and segmentation as a joint multi-stage processing. We further integrate regularization methods such as convex hull and Douglas-Peucker algorithm to obtain accurately segmented edges. The method is tested on the crowdAI benchmark dataset by comparing with alternative state-of-the-art models (i.e., Mask R-CNN). The results show that our method achieves better instance segmentation results and improves the results in terms of geometric regularity of building segments.
Wufan Zhao, Claudio Persello, Alfred Stein
IGARSS2
2019 Towards Automated Delineation of Smallholder Farm Fields From VHR Images Using Convolutional Networks
abstract
Automated delineation of smallholder farm fields is difficult because of their small size, irregular shape and the use of mixed-cropping systems. Edges between smallholder plots are often indistinct in satellite imagery and contours have to be identified by considering the transition of the complex textural patterns of the fields. We introduce a strategy to delineate field boundaries using a fully convolutional network in combination with a globalization and grouping algorithm to produce a hierarchical segmentation of the fields. We carry out an experimental analysis in a study area in Kofa, Nigeria, using a WorldView-3 image, comparing several state-of-the-art contour detection algorithms. The proposed strategy outperforms state-of-the-art computer vision methods and shows promising results by automatically delineating field boundaries with an accuracy close to human level photo-interpretation.
Claudio Persello, Valentyn A. Tolpekin, John Ray Bergado, Rolf A. de By
IGARSS1
2019 Deep Convolutional Networks for Cloud Detection Using Resourcesat-2 Data
abstract
Cloud cover creates obstruction in Earth Observation studies. The obstruction is harder to distinguish from features having similar reflectance on the ground, such as snow. To distinguish clouds from snow in a VNIR image, we use an additional SWIR band. The images were fed into a deep Fully Convolutional Network that can fuse the multiresolution SWIR and VNIR bands together, in order to produce pixelwise classification. The accuracy obtained by the model on the test image was 93.35%. We compare the performance of this model with a more commonly used technique, Random Forests. To analyze the effect of SWIR, we use another deep learning model, trained only on the VNIR image, and compare the accuracies obtained.
Debvrat Varshney, Prasun Kumar Gupta, Claudio Persello, Bhaskar Ramachandra Nikam
IGARSS3
2019 Extracting Cadastral Boundaries from UAV Images Using Fully Convolutional Networks
abstract
The cadastre is the foundation of land management. However, it is estimated that 70% of the land rights in the world remain unregistered. Traditional approaches are costly and labor intensive, therefore, recently the use of remotely sensed images has been investigated. The delineation of cadastral boundaries from such data is challenging since not all boundaries are demarcated by visible physical objects. In this paper, we introduce a technique based on deep Fully Convolutional Networks (FCNs), which can automatically learn high-level spatial features from images, to extract cadastral boundaries. Our strategy combines FCN and a grouping algorithm using the Oriented Watershed Transform (OWT) to generate connected contours. We carried out an experimental analysis in a real case study in Busogo, Rwanda, using images acquired by Unmanned Aerial Vehicles (UAV) in 2018. Our investigation shows promising results in automatically extracting visible boundaries, which can contribute to the current mapping and updating practices in Rwanda.
Xue Xia 0001, Mila Koeva, Claudio Persello
IGARSS3
2018 Fusenet: End- to-End Multispectral Vhr Image Fusion and Classification
abstract
Classification of very high resolution (VHR) satellite images faces two major challenges: 1) inherent low intra-class and high inter-class spectral similarities and 2) mismatching resolution of available bands. Conventional methods have addressed these challenges by adopting separate stages of image fusion and spatial feature extraction steps. These steps, however, are not jointly optimizing the classification task at hand. We propose a single-stage framework embedding these processing stages in a multiresolution convolutional network. The network, called FuseNet, aims to match the resolution of the panchromatic and multispectral bands in a VHR image using convolutional layers with corresponding downsampling and upsampling operations. We compared FuseNet against the use of separate processing steps for image fusion, such as pansharpening and resampling through interpolation. We also analyzed the sensitivity of the classification performance of FuseNet to a selected number of its hyperparameters. Results show that FuseNet surpasses conventional methods.
John Ray Bergado, Claudio Persello, Alfred Stein
IGARSS2
2018 Fully Convolutional Networks for Multi-Temporal SAR Image Classification
abstract
Classification of crop types from multi-temporal SAR data is a complex task because of the need to extract spatial and temporal features from images affected by speckle. Previous methods applied speckle filtering and then classification in two separate processing steps. This paper introduces fully convolutional networks (FCN) for pixel-wise classification of crops from multi-temporal SAR data. It applies speckle filtering and classification in a single framework. Furthermore, it also uses dilated kernels to increase the capability to learn long distance spatial dependencies. The proposed FCN was compared with patch-based convolutional neural network (CNN) and support vector machine (SVM) classifiers. The proposed method performed better when compared with the patch-based CNN and SVM.
Adugna G. Mullissa, Claudio Persello, Valentyn A. Tolpekin
IGARSS2
2018 Recurrent Multiresolution Convolutional Networks for VHR Image Classification
abstract
Classification of very high-resolution (VHR) satellite images has three major challenges: 1) inherent low intraclass and high interclass spectral similarities; 2) mismatching resolution of available bands; and 3) the need to regularize noisy classification maps. Conventional methods have addressed these challenges by adopting separate stages of image fusion, feature extraction, and postclassification map regularization. These processing stages, however, are not jointly optimizing the classification task at hand. In this paper, we propose a single-stage framework embedding the processing stages in a recurrent multiresolution convolutional network trained in an end-to-end manner. The feedforward version of the network, called FuseNet, aims to match the resolution of the panchromatic and multispectral bands in a VHR image using convolutional layers with corresponding downsampling and upsampling operations. Contextual label information is incorporated into FuseNet by means of a recurrent version called ReuseNet. We compared FuseNet and ReuseNet against the use of separate processing steps for both image fusions, e.g., pansharpening and resampling through interpolation and map regularization such as conditional random fields. We carried out our experiments on a land-cover classification task using a Worldview-03 image of Quezon City, Philippines, and the International Society for Photogrammetry and Remote Sensing 2-D semantic labeling benchmark data set of Vaihingen, Germany. FuseNet and ReuseNet surpass the baseline approaches in both the quantitative and qualitative results.
John Ray Bergado, Claudio Persello, Alfred Stein
IEEE Trans. Geosci. Remote. Sens.2
2017 Classification of multitemporal SAR images using convolutional neural networks and Markov random fields
abstract
Classification of Synthetic Aperture Radar (SAR) images is a complex task because of the presence of speckle, which affects images in a way similar to a strong noise. In this study, we investigate the use of Convolutional Neural Networks (CNNs) which can effectively learn a bank of spatial filters to simultaneously 1) reduce speckle noise, and 2) extract spatial-contextual features to characterize texture and scattering mechanism. Moreover, we combine CNN with Markov Random Fields (MRFs) for post-classification label smoothing to further reduce the effect of speckle on the land-cover map and to improve classification accuracy. We applied the proposed classification system to the analysis of a multitemporal series of Sentinel-1 images for mapping agricultural fields in Flevoland, The Netherlands. Experimental results confirm the effectiveness of the investigated approach, which outperforms standard methods.
Carolyne Danilla, Claudio Persello, Valentyn A. Tolpekin, John Ray Bergado
IGARSS2
2017 Detection of informal settlements from VHR satellite images using convolutional neural networks
abstract
Convolutional neural networks (CNNs), widely studied in the domain of computer vision, are more recently finding application in the analysis of high-resolution aerial and satellite imagery. In this paper, we investigate a deep feature learning approach based on CNNs for the detection of informal settlements in Dar es Salaam, Tanzania. This information is vital for decision making and planning of upgrading processes. Distinguishing the different urban structure types is challenging because of the abstract semantic definition of the classes as opposed to the separation of standard land-cover classes. This task requires the extraction of complex spatial-contextual features. To this aim, we trained a CNN in an end-to-end fashion and used it to classify informal and formal settlements. Our experimental results show that CNNs outperform state of the art methods using hand-crafted features. We conclude that CNNs are able to effectively learn the spatial-contextual features for accurately discriminating formal and informal settlements.
Nicholus Mboga, Claudio Persello, John Ray Bergado, Alfred Stein
IGARSS2
2017 Deep Fully Convolutional Networks for the Detection of Informal Settlements in VHR Images
abstract
This letter investigates fully convolutional networks (FCNs) for the detection of informal settlements in very high resolution (VHR) satellite images. Informal settlements or slums are proliferating in developing countries and their detection and classification provides vital information for decision making and planning urban upgrading processes. Distinguishing different urban structures in VHR images is challenging because of the abstract semantic definition of the classes as opposed to the separation of standard land-cover classes. This task requires extraction of texture and spatial features. To this aim, we introduce deep FCNs to perform pixel-wise image labeling by automatically learning a higher level representation of the data. Deep FCNs can learn a hierarchy of features associated to increasing levels of abstraction, from raw pixel values to edges and corners up to complex spatial patterns. We present a deep FCN using dilated convolutions of increasing spatial support. It is capable of learning informative features capturing long-range pixel dependencies while keeping a limited number of network parameters. Experiments carried out on a Quickbird image acquired over the city of Dar es Salaam, Tanzania, show that the proposed FCN outperforms state-of-the-art convolutional networks. Moreover, the computational cost of the proposed technique is significantly lower than standard patch-based architectures.
Claudio Persello, Alfred Stein
IEEE Geosci. Remote. Sens. Lett.1
2016 A deep learning approach to the classification of sub-decimetre resolution aerial images
abstract
Spatial-contextual features play a vital role in the classification of very high resolution aerial images characterized by sub-decimetre resolution. However, manually extracting relevant contextual features is difficult and time-consuming in the analysis of sub-decimetre resolution images, where the objects of interest are significantly larger than the pixel size. Deep learning methods allow us to replace hand-crafted features by automatically learning contextual features from the image. In this paper, we investigate the use of convolutional neural networks (CNN) for the classification of urban areas using high resolution airborne images. We also analyse the sensitivity of network hyperparameters providing an interpretation of their effect on the extraction of spatial-contextual features. Experimental results show the effectiveness of CNN in learning discriminative contextual features leading to accurate classified maps and outperforming traditional classification methods based on the extraction of textural features.
John Ray Bergado, Claudio Persello, Caroline Gevaert
IGARSS2
2016 Integration of 2D and 3D features from UAV imagery for informal settlement classification using Multiple Kernel Learning
abstract
Informal settlement upgrading projects require high-resolution and up-to-date thematic maps in order to plan and design effective interventions. To this end, Unmanned Aerial Vehicles (UAVs) provide the opportunity to obtain very high resolution 2D orthomosaics and 3D point clouds where and when needed. The heterogeneous, dense structures which typically make up an informal settlement motivate the importance of integrating complex 2D and 3D features obtained from UAV data into a single classification problem. Multiple Kernel Learning (MKL) Support Vector Machines (SVMs) maintain the distinct characteristics of the different feature spaces by optimizing individual kernels for specific feature groups which are later combined into a single kernel used for classification. Both the kernel parameters and kernel weights can be optimized by considering the alignment between the kernel and an ideal kernel which would perfectly classify the samples. This paper demonstrates how extracting high-level features from both the 2D orthomosaic as well as the 3D point cloud (obtained by an UAV), and integrating them through a MKL approach, can obtain an Overall Accuracy of 90.29%, a 4% increase over the results obtained using single kernel methods.
Caroline Gevaert, Claudio Persello, Richard Sliuzas, George Vosselman
IGARSS2
2016 Two-level active learning method for debris detection using VHR satellite imagery and local aerial surveys
abstract
This paper presents a two-level Active Learning (AL) classification method for the interactive detection of earthquake-induced debris via the synergetic use of post-disaster Very High Resolution (VHR) satellite and local decimeter-resolution aerial images. The proposed method is performed by interactively guiding the human expert in the collection of labeled training samples from aerial images and optimally planning further aerial surveys. A label propagation mechanism is adopted to reliably propagate the labeled pixels annotated by the user on the aerial image by visual photointerpretation to the (lower resolution) satellite image. The experimental analysis is carried out using post-disaster images of the 2010 Yushu earthquake in China. The obtained results confirm that the proposed method can significantly reduce the user annotation effort and the cost for acquiring additional aerial images leading to more accurate and timely debris detection maps.
Zhihua Xu, Claudio Persello, Mengmeng Li 0002, Lixin Wu, Pengtianhao Wu
IGARSS2
2016 Kernel-Based Domain-Invariant Feature Selection in Hyperspectral Images for Transfer Learning
abstract
This paper presents a kernel-based feature selection method for the classification of hyperspectral images. The proposed method aims at selecting a subset of the original features that are both 1) relevant (discriminant) for the considered classification problem, i.e., preserve the functional relationship between input and output variables, and 2) invariant (stable) across different domains, i.e., minimize the data-set shift between the source and the target domains. Domains can be associated with hyperspectral images collected either on different geographical areas or on the same area at different times. We propose a novel measure of data-set shift for evaluating the domain stability, which computes the distance of the conditional distributions between the source and target domains in a reproducing kernel Hilbert space. Such a measure is defined on the basis of the kernel embeddings of the conditional distributions resulting in a nonparametric approach that does not require estimating the distribution of the classes. The adopted search strategy is based on a multiobjective optimization algorithm, which optimizes the two terms of the criterion function for the estimation of the Pareto-optimal solutions. This results in an effective approach of performing feature selection in a transfer learning setting. The experimental results obtained on two hyperspectral images show the effectiveness of the proposed method in selecting features with high generalization capabilities.
Claudio Persello, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2014 Relevant and invariant feature selection of hyperspectral images for domain generalization
abstract
This paper presents a novel feature selection method for the analysis of hyperspectral images. The proposed method aims at selecting a subset of the original features that are both 1) relevant for the considered problem (i.e., preserve the functional relationship between input and output variables), and 2) invariant (stable) across different domains (i.e., minimize the data set shift among different domains). Domains can be associated with images collected on different areas or on the same area at different times. We propose a novel measure of domain stability, which evaluates the distance of the conditional distributions between the source and target domain. Such a measure is defined on the basis of kernel embeddings of conditional distributions and can be applied to both classification and regression problems. Experimental results show the effectiveness of the proposed method in selecting features with high generalization capabilities on the target domain.
Claudio Persello, Lorenzo Bruzzone
IGARSS1
2014 Active and Semisupervised Learning for the Classification of Remote Sensing Images
abstract
This paper aims at analyzing and comparing active learning (AL) and semisupervised learning (SSL) methods for the classification of remote sensing (RS) images. We present a literature review of the two learning paradigms and compare them theoretically and experimentally when addressing classification problems characterized by few training samples (w.r.t. the number of features) and affected by sample selection bias. Commonalities and differences are highlighted in the context of a conceptual framework used to describe the workflow of the two approaches. We point out advantages and disadvantages of the two approaches, delineating the boundary conditions on the applicability of the two paradigms with respect to both the amount and the quality of available training samples. Moreover, we investigate the integration of concepts that are in common between the two learning paradigms for improving state-of-the-art techniques and combining AL and SSL in order to jointly leverage the advantages of both approaches. In this framework, we propose a novel SSL algorithm that improves the progressive semisupervised support vector machine by integrating concepts that are usually considered in AL methods. We performed several experiments considering both synthetic and real multispectral and hyperspectral RS data, defining different classification problems starting from different initial training sets. The experiments are carried out considering classification methods based on support vector machines.
Claudio Persello, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2014 Cost-Sensitive Active Learning With Lookahead: Optimizing Field Surveys for Remote Sensing Data Classification
abstract
Active learning typically aims at minimizing the number of labeled samples to be included in the training set to reach a certain level of classification accuracy. Standard methods do not usually take into account the real annotation procedures and implicitly assume that all samples require the same effort to be labeled. Here, we consider the case where the cost associated with the annotation of a given sample depends on the previously labeled samples. In general, this is the case when annotating a queried sample is an action that changes the state of a dynamic system, and the cost is a function of the state of the system. In order to minimize the total annotation cost, the active sample selection problem is addressed in the framework of a Markov decision process, which allows one to plan the next labeling action on the basis of an expected long-term cumulative reward. This framework allows us to address the problem of optimizing the collection of labeled samples by field surveys for the classification of remote sensing data. The proposed method is applied to the ground sample collection for tree species classification using airborne hyperspectral images. Experiments carried out in the context of a real case study on forest inventory show the effectiveness of the proposed method.
Claudio Persello, Abdeslam Boularias, Michele Dalponte, Terje Gobakken, Erik Næsset, Bernhard Schölkopf
IEEE Trans. Geosci. Remote. Sens.1
2013 Optimizing the ground sample collection with cost-sensitive active learning for tree species classification using hyperspectral images
abstract
This study presents a cost-sensitive active learning method for optimizing the field surveys by a human expert in the classification of single tree species using hyperspectral images. The goal of the proposed method is to guide the human expert in the collection of labeled samples in order to maximize the ratio between the classification accuracy with respect to the travelling costs. Experiments carried out in the context of a real study on forest inventory show the effectiveness of the proposed method.
Claudio Persello, Michele Dalponte, Terje Gobakken, Erik Næsset
IGARSS1
2013 Interactive Domain Adaptation for the Classification of Remote Sensing Images Using Active Learning
abstract
This letter presents a novel interactive domain-adaptation technique based on active learning for the classification of remote sensing (RS) images. The proposed method aims at adapting the supervised classifier trained on a given RS source image to make it suitable for classifying a different but related target image. The two images can be acquired in different locations and/or at different times. The proposed approach iteratively selects the most informative samples of the target image to be labeled by the user and included in the training set, whereas the source image samples are reweighted or possibly removed from the training set on the basis of their disagreement with the target image classification problem. This way, the consistent information available from the source image can be effectively exploited for the classification of the target image and for guiding the selection of new samples to be labeled, whereas the inconsistent information is automatically detected and removed. This approach can significantly reduce the number of new labeled samples to be collected from the target image. Experimental results on both a multispectral very high resolution and a hyperspectral data set confirm the effectiveness of the proposed method.
Claudio Persello
IEEE Geosci. Remote. Sens. Lett.1
2012 Interactive domain adaptation technique for the classification of remote sensing images
abstract
This paper presents a novel interactive domain-adaptation technique based on active learning for the classification of remote sensing (RS) images. The proposed method aims at adapting the supervised classifier trained on a given RS source image to make it suitable for classifying a different but related target image. The two images can be acquired in different locations and/or at different times, but present the same set of land-cover classes. The proposed approach iteratively selects the most informative samples of the target image to be labeled by the user and included in the training set, while the source-image samples are re-weighted or possibly removed from the training set on the basis of their disagreement with the target image classification problem. In this way, the consistent information available from the source image can be effectively exploited for the classification of a target image and for guiding the user in the selection of the new samples to be labeled, whereas the inconsistent information is automatically detected and removed. Experimental results on a Very High Resolution (VHR) multispectral dataset confirm the effectiveness of the proposed method.
Claudio Persello, Francesco Dinuzzo
IGARSS1
2012 Active Learning for Domain Adaptation in the Supervised Classification of Remote Sensing Images
abstract
This paper presents a novel technique for addressing domain adaptation (DA) problems with active learning (AL) in the classification of remote sensing images. DA models the important problem of adapting a supervised classifier trained on a given image (source domain) to the classification of another similar but not identical image (target domain) acquired on a different area. The main idea of the proposed approach is iteratively labeling and adding to the training set the minimum number of the most informative samples from the target domain, while removing the source-domain samples that do not fit with the distributions of the classes in the target domain. In this way, the classification system exploits already available information, i.e., the labeled samples of source domain, in order to minimize the number of target domain samples to be labeled, thus reducing the cost associated to the definition of the training set for the classification of the target domain. In addition, we define a convergence criterion that allows the technique to stop the iterative AL process on the target domain without relying on the availability of a test set for it. This is an important contribution, as in operational applications, it is not realistic to assume that a test set for the target domain is available. Experimental results obtained in the classification of very high resolution and hyperspectral images confirm the effectiveness of the proposed technique.
Claudio Persello, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2011 A novel active learning strategy for domain adaptation in the classification of remote sensing images
abstract
We present a novel technique for addressing domain adaptation problems in the classification of remote sensing images with active learning. Domain adaptation is the important problem of adapting a supervised classifier trained on a given image (source domain) to the classification of another similar (but not identical) image (target domain) acquired on a different area, or on the same area at a different time. The main idea of the proposed approach is to iteratively labeling and adding to the training set the minimum number of the most informative samples from target domain, while removing the source-domain samples that does not fit with the distributions of the classes in the target domain. In this way, the classification system exploits already available information, i.e., the labeled samples of source domain, in order to minimize the number of target domain samples to be labeled, thus reducing the cost associated to the definition of the training set for the classification of the target domain. Experimental results obtained in the classification of a hyperspectral image confirm the effectiveness of the proposed technique.
Claudio Persello, Lorenzo Bruzzone
IGARSS1
2011 Batch-Mode Active-Learning Methods for the Interactive Classification of Remote Sensing Images
abstract
This paper investigates different batch-mode active-learning (AL) techniques for the classification of remote sensing (RS) images with support vector machines. This is done by generalizing to multiclass problem techniques defined for binary classifiers. The investigated techniques exploit different query functions, which are based on the evaluation of two criteria: uncertainty and diversity. The uncertainty criterion is associated to the confidence of the supervised algorithm in correctly classifying the considered sample, while the diversity criterion aims at selecting a set of unlabeled samples that are as more diverse (distant one another) as possible, thus reducing the redundancy among the selected samples. The combination of the two criteria results in the selection of the potentially most informative set of samples at each iteration of the AL process. Moreover, we propose a novel query function that is based on a kernel-clustering technique for assessing the diversity of samples and a new strategy for selecting the most informative representative sample from each cluster. The investigated and proposed techniques are theoretically and experimentally compared with state-of-the-art methods adopted for RS applications. This is accomplished by considering very high resolution multispectral and hyperspectral images. By this comparison, we observed that the proposed method resulted in better accuracy with respect to other investigated and state-of-the art methods on both the considered data sets. Furthermore, we derived some guidelines on the design of AL systems for the classification of different types of RS images.
Begüm Demir, Claudio Persello, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.2
2010 Recent trends in classification of remote sensing data: active and semisupervised machine learning paradigms
abstract
This paper addresses the recent trends in machine learning methods for the automatic classification of remote sensing (RS) images. In particular, we focus on two new paradigms: semisupervised and active learning. These two paradigms allow one to address classification problems in the critical conditions where the available labeled training samples are limited. These operational conditions are very usual in RS problems, due to the high cost and time associated with the collection of labeled samples. Semisupervised and active learning techniques allow one to enrich the initial training set information and to improve classification accuracy by exploiting unlabeled samples or requiring additional labeling phases from the user, respectively. The two aforementioned strategies are theoretically and experimentally analyzed considering SVM-based techniques in order to highlight advantages and disadvantages of both strategies.
Lorenzo Bruzzone, Claudio Persello
IGARSS2
2010 A Novel Protocol for Accuracy Assessment in Classification of Very High Resolution Images
abstract
This paper presents a novel protocol for the accuracy assessment of the thematic maps obtained by the classification of very high resolution images. As the thematic accuracy alone is not sufficient to adequately characterize the geometrical properties of high-resolution classification maps, we propose a protocol that is based on the analysis of two families of indices: 1) the traditional thematic accuracy indices and 2) a set of novel geometric indices that model different geometric properties of the objects recognized in the map. In this context, we present a set of indices that characterize five different types of geometric errors in the classification map: 1) oversegmentation; 2) undersegmentation; 3) edge location; 4) shape distortion; and 5) fragmentation. Moreover, we propose a new approach for tuning the free parameters of supervised classifiers on the basis of a multiobjective criterion function that aims at selecting the parameter values that result in the classification map that jointly optimize thematic and geometric error indices. Experimental results obtained on QuickBird images show the effectiveness of the proposed protocol in selecting classification maps characterized by a better tradeoff between thematic and geometric accuracies than standard procedures based only on thematic accuracy measures. In addition, results obtained with support vector machine classifiers confirm the effectiveness of the proposed multiobjective technique for the selection of free-parameter values for the classification algorithm.
Claudio Persello, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2009 Active Learning for Classification of Remote Sensing Images
abstract
This paper presents an analysis of active learning techniques for the classification of remote sensing images and proposes a novel active learning method based on support vector machines (SVMs). The proposed method exploits a query function for the inclusion of batches of unlabeled samples in the training set, which is based on the evaluation of two criteria: uncertainty and diversity. This query function adopts a stochastic approach to the selection of unlabeled samples, which is based on a function of uncertainty estimated from the distribution of errors on the validation set (which is assumed available for the model selection of the SVM classifier). Experimental results carried out on a very high resolution image confirm the effectiveness of the proposed active learning technique, which results more accurate than standard methods.
Lorenzo Bruzzone, Claudio Persello
IGARSS (3)2
2009 A Novel Approach to the Selection of Spatially Invariant Features for Classification of Hyperspectral Images
abstract
This paper presents a novel approach to feature selection for the classification of hyperspectral images. The proposed approach aims at selecting a subset of the original set of features that exhibits two main properties: i) high capability to discriminate among the considered classes, ii) high invariance in the spatial domain of the investigated scene. This approach results in a more robust classification system with improved generalization properties with respect to standard feature-selection methods. The feature selection is accomplished by defining a multi-objective criterion function made up of two terms: i) a term that measures the class separability, ii) a term that evaluates the spatial invariance of the selected features. In order to assess the spatial invariance of the feature subset we propose both a supervised method and a semisupervised method (which choice depends on the available reference data). The multi-objective problem is solved by an evolutionary algorithm that estimates the set of Pareto-optimal solutions. Experiments carried out on a hyperspectral image acquired by the Hyperion sensor on a complex area confirmed the effectiveness of the proposed approach.
Claudio Persello, Lorenzo Bruzzone
IGARSS (2)1
2009 A Novel Context-Sensitive Semisupervised SVM Classifier Robust to Mislabeled Training Samples
abstract
This paper presents a novel context-sensitive semisupervised support vector machine (CS4VM) classifier, which is aimed at addressing classification problems where the available training set is not fully reliable, i.e., some labeled samples may be associated to the wrong information class (mislabeled patterns). Unlike standard context-sensitive methods, the proposed CS4VM classifier exploits the contextual information of the pixels belonging to the neighborhood system of each training sample in the learning phase to improve the robustness to possible mislabeled training patterns. This is achieved according to both the design of a semisupervised procedure and the definition of a novel contextual term in the cost function associated with the learning of the classifier. In order to assess the effectiveness of the proposed CS4VM and to understand the impact of the addressed problem in real applications, we also present an extensive experimental analysis carried out on training sets that include different percentages of mislabeled patterns having different distributions on the classes. In the analysis, we also study the robustness to mislabeled training patterns of some widely used supervised and semisupervised classification algorithms (i.e., conventional support vector machine (SVM), progressive semisupervised SVM, maximum likelihood, andk-nearest neighbor). Results obtained on a very high resolution image and on a medium resolution image confirm both the robustness and the effectiveness of the proposed CS4VM with respect to standard classification algorithms and allow us to derive interesting conclusions on the effects of mislabeled patterns on different classifiers.
Lorenzo Bruzzone, Claudio Persello
IEEE Trans. Geosci. Remote. Sens.2
2009 A Novel Approach to the Selection of Spatially Invariant Features for the Classification of Hyperspectral Images With Improved Generalization Capability
abstract
This paper presents a novel approach to feature selection for the classification of hyperspectral images. The proposed approach aims at selecting a subset of the original set of features that exhibits at the same time high capability to discriminate among the considered classes and high invariance in the spatial domain of the investigated scene. This approach results in a more robust classification system with improved generalization properties with respect to standard feature-selection methods. The feature selection is accomplished by defining a multiobjective criterion function made up of two terms: (1) a term that measures the class separability and (2) a term that evaluates the spatial invariance of the selected features. In order to assess the spatial invariance of the feature subset, we propose both a supervised method (which assumes that training samples acquired in two or more spatially disjoint areas are available) and a semisupervised method (which requires only a standard training set acquired in a single area of the scene and takes advantage of unlabeled samples selected in portions of the scene spatially disjoint from the training set). The choice for the supervised or semisupervised method depends on the available reference data. The multiobjective problem is solved by an evolutionary algorithm that estimates the set of Pareto-optimal solutions. Experiments carried out on a hyperspectral image acquired by the Hyperion sensor on a complex area confirmed the effectiveness of the proposed approach.
Lorenzo Bruzzone, Claudio Persello
IEEE Trans. Geosci. Remote. Sens.2
2008 A Novel Protocol for Accuracy Assessment in Classification of Very High Resolution Multispectral and SAR Images
abstract
This paper presents a novel protocol for the accuracy assessment of thematic maps obtained by the classification of very high resolution images. As the thematic accuracy alone is not sufficient to adequately characterize the geometrical properties of classification maps, we propose a novel protocol that is based on the analysis of two families of indexes: (i) the traditional thematic accuracy indexes, and (ii) a set of geometric indexes that characterize different geometric properties of the objects recognized in the map. These indexes can be used in the training phase of a classifier for identifying the parameters values that optimize classification results on the basis of a multi-objective criterion. Experimental results obtained on Quickbird images show the effectiveness of the proposed protocol in selecting classification maps characterized by better tradeoff between thematic and geometric accuracy with respect to standard accuracy measures.
Lorenzo Bruzzone, Claudio Persello
IGARSS (2)2
2008 A Novel Approach to the Selection of Robust and Invariant Features for Classification of Hyperspectral Images
abstract
This paper presents a novel approach to feature selection for the classification of hyperspectral images. The proposed approach aims at selecting a subset of the original set of features that exhibits two main properties:( i) high capability to discriminate among the considered classes, (ii) high invariance (stationarity) in the spatial domain of the investigated scene. The feature selection is accomplished by defining a multi-objective criterion that considers two terms: (i) a term that assesses the class separability, (ii) a term that evaluates the spatial invariance of the selected features. The multi-objective problem is solved by an evolutionary algorithm that estimates the Pareto-optimal solutions. Experiments carried out on a hyperspectral image acquired by the Hyperion sensor confirmed the effectiveness of the proposed technique.
Lorenzo Bruzzone, Claudio Persello
IGARSS (1)2
2007 Fusion of spectral and spatial information by a novel SVM classification technique
abstract
A novel context-sensitive semisupervised classification technique based on support vector machines is proposed. This technique aims at exploiting the SVM method for image classification by properly fusing spectral information with spatial- context information. This results in: i) an increased robustness to noisy training sets in the learning phase of the classifier; ii) a higher and more stable classification accuracy with respect to the specific patterns included in the training set; and iii) a regularized classification map. The main property of the proposed context sensitive semisupervised SVM (CS4VM) is to adaptively exploit the contextual information in the training phase of the classifier, without any critical assumption on the expected labels of the pixels included in the same neighborhood system. This is done by defining a novel context-sensitive term in the objective function used in the learning of the classifier. In addition, the proposed CS4VM can be integrated with a Markov random field (MRF) approach for exploiting the contextual information also to regularize the classification map. Experiments carried out on very high geometrical resolution images confirmed the effectiveness of the proposed technique.
Lorenzo Bruzzone, Mattia Marconcini, Claudio Persello
IGARSS3