Devis Tuia

dblp:99/606 · DBLP profile ↗
← Back
141ranked-venue papers
33as first author
36since 2021 · last 2026
0000-0003-0374-2459ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 100 · 27 first-author · 21 since 2021Artificial intelligence and machine learning · 36 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021
YearPublicationVenuePosition
2026 TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
abstract
Recent unsupervised domain adaptation (UDA) methods have shown great success in addressing classical domain shifts (e.g., synthetic-to-real), but they still suffer under complex shifts (e.g. geographical shift), where both the background and object appearances differ significantly across domains. Prior works showed that the language modality can help in the adaptation process, exhibiting more robustness to such complex shifts. In this paper, we introduce TRUST, a novel UDA approach that exploits the robustness of the language modality to guide the adaptation of a vision model. TRUST generates pseudo-labels for target samples from their captions and introduces a novel uncertainty estimation strategy that uses normalised CLIP similarity scores to estimate the uncertainty of the generated pseudo-labels. Such estimated uncertainty is then used to reweight the classification loss, mitigating the adverse effects of wrong pseudo-labels obtained from low-quality captions. To further increase the robustness of the vision model, we propose a multimodal soft-contrastive learning loss that aligns the vision and language feature spaces, by leveraging captions to guide the contrastive training of the vision model on target images. In our contrastive loss, each pair of images acts as both a positive and a negative pair and their feature representations are attracted and repulsed with a strength proportional to the similarity of their captions. This solution avoids the need for hardly determining positive and negative pairs, which is critical in the UDA setting. Our approach outperforms previous methods, setting the new state-of-the-art on classical (DomainNet) and complex (GeoNet) domain shifts. The code is available at https://github.com/MattiaLitrico/TRUST-Leveraging-Text-Robustness-for-Unsupervised-Domain-Adaptation.
Mattia Litrico, Mario Valerio Giuffrida, Sebastiano Battiato, Devis Tuia
AAAI4
2026 Checkmate: Interpretable and Explainable RSVQA is the Endgame
Lucrezia Tosato, Christel Tartini-Chappuis, Syrielle Montariol, Flora Weissgerber, Sylvain Lobry, Devis Tuia
ICPR (7)6
2025 MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps
abstract
Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors enabling the study of wildlife populations at scale with minimal disturbance. However, the lack of annotated video datasets limits the development of powerful video understanding models needed to process the vast amount of fieldwork data collected. To advance research in wild animal behavior monitoring we present MammAlps, a multi-modal and multi-view dataset of wildlife behavior monitoring from 9 camera-traps in the Swiss National Park. Mam-mAlps contains over 14 hours of video with audio, 2D segmentation maps and 8.5 hours of individual tracks densely labeled for species and behavior. Based on 6‘135 single animal clips, we propose the first hierarchical and multi-modal animal behavior recognition benchmark using audio, video and reference scene segmentation maps as inputs. Furthermore, we also propose a second ecology-oriented benchmark aiming at identifying activities, species, number of individuals and meteorological conditions from 397 multi-view and long-term ecological events, including false positive triggers. We advocate that both tasks are complementary and contribute to bridging the gap between machine learning and ecology. Code and data are available at https://github.com/eceo-epfl/MammAlps.
Valentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul, Alexander Mathis, Devis Tuia
CVPR6
2025 GeoExplorer: Active Geo-Localization with Curiosity-Driven Exploration
abstract
Active Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a goal-reaching reinforcement learning (RL) problem with a distance-based reward. They localize the goal by implicitly learning to minimize the relative distance from it. However, when distance estimation becomes challenging or when encountering unseen targets and environments, the agent exhibits reduced robustness and generalization ability due to the less reliable exploration strategy learned during training. In this paper, we propose GeoExplorer, an AGL agent that incorporates curiosity-driven exploration through intrinsic rewards. Unlike distance-based rewards, our curiosity-driven reward is goal-agnostic, enabling robust, diverse, and contextually relevant exploration based on effective environment modeling. These capabilities have been proven through extensive experiments across four AGL benchmarks, demonstrating the effectiveness and generalization ability of GeoExplorer in diverse settings, particularly in localizing unfamiliar targets and environments.
Li Mi, Manon Béchaz, Zeming Chen 0001, Antoine Bosselut, Devis Tuia
ICCV5
2025 SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
Gencer Sumbul, Chang Xu 0027, Emanuele Dalsasso, Devis Tuia
ICCV4
2025 What to align in multimodal contrastive learning?
abstract
Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive Multimodal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on seven multimodal tasks.
Benoit Dufumier, Javiera Castillo-Navarro, Devis Tuia, Jean-Philippe Thiran
ICLR3
2025 Multi-Scale Grouped Prototypes for Interpretable Semantic Segmentation
abstract
Prototypical part learning is emerging as a promising approach for making semantic segmentation interpretable. The model selects real patches seen during training as prototypes and constructs the dense prediction map based on the similarity between parts of the test image and the prototypes. This improves interpretability since the user can inspect the link between the predicted output and the patterns learned by the model in terms of prototypical information. In this paper, we propose a method for inter-pretable semantic segmentation that leverages multi-scale image representation for prototypical part learning. First, we introduce a prototype layer that explicitly learns diverse prototypical parts at several scales, leading to multi-scale representations in the prototype activation output. Then, we propose a sparse grouping mechanism that produces multi-scale sparse groups of these scale-specific prototypical parts. This provides a deeper understanding of the interactions between multi-scale object representations while enhancing the interpretability of the segmentation model. The experiments conducted on Pascal VOC, Cityscapes, and ADE20K demonstrate that the proposed method increases model sparsity, improves interpretability over existing prototype-based methods, and narrows the performance gap with the non-interpretable counterpart models. Code is available at github.com/eceo-epfl/ScaleProtoSeg.
Hugo Porta, Emanuele Dalsasso, Diego Marcos, Devis Tuia
WACV4
2024 ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
abstract
Asking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected answers or knowledge triplets, to generate focused questions. These constraints allow VQG systems to specify the question content or leverage external commonsense knowledge that can not be obtained from the image content only. However, generating focused questions using textual constraints while enforcing a high relevance to the image content remains a challenge, as VQG systems often ignore one or both forms of grounding. In this work, we propose Contrastive Visual Question Generation (ConVQG), a method using a dual contrastive objective to discriminate questions generated using both modalities from those based on a single one. Experiments on both knowledge-aware and standard VQG benchmarks demonstrate that ConVQG outperforms the state-of-the-art methods and generates image-grounded, text-guided, and knowledge-rich questions. Our human evaluation results also show preference for ConVQG questions compared to non-contrastive baselines.
Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia
AAAI6
2024 ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia
ECCV (14)7
2024 Self-Supervised Underwater Caustics Removal and Descattering via Deep Monocular SLAM
Jonathan Sauder, Devis Tuia
ECCV (84)2
2024 Multilingual Vision-Language Pre-training for the Remote Sensing Domain
abstract
Methods based on Contrastive Language-Image Pre-training (CLIP) are nowadays extensively used in support of vision-and-language tasks involving remote sensing data, such as cross-modal retrieval. The adaptation of CLIP to this specific domain has relied on model fine-tuning with the standard contrastive objective, using existing human-labeled image-caption datasets, or using synthetic data corresponding to image-caption pairs derived from other annotations over remote sensing images (e.g., object classes). The use of different pre-training mechanisms has received less attention, and only a few exceptions have considered multilingual inputs. This work proposes a novel vision-and-language model for the remote sensing domain, exploring the fine-tuning of a multilingual CLIP model and testing the use of a self-supervised method based on aligning local and global representations from individual input images, together with the standard CLIP objective. Model training relied on assembling pre-existing datasets of remote sensing images paired with English captions, followed by the use of automated machine translation into nine additional languages. We show that translated data is indeed helpful, e.g. improving performance also on English. Our resulting model, which we named Remote Sensing Multilingual CLIP (RS-M-CLIP), obtains state-of-the-art results in a variety of vision-and-language tasks, including cross-modal and multilingual image-text retrieval, or zero-shot image classification.
João Daniel Silva, João Magalhães, Devis Tuia, Bruno Martins 0001
SIGSPATIAL/GIS3
2024 Geographic Location Encoding with Spherical Harmonics and Sinusoidal Representation Networks
abstract
Learning representations of geographical space is vital for any machine learning model that integrates geolocated data, spanning application domains such as remote sensing, ecology, or epidemiology. Recent work embeds coordinates using sine and cosine projections based on Double Fourier Sphere (DFS) features. These embeddings assume a rectangular data domain even on global data, which can lead to artifacts, especially at the poles. At the same time, little attention has been paid to the exact design of the neural network architectures with which these functional embeddings are combined. This work proposes a novel location encoder for globally distributed geographic data that combines spherical harmonic basis functions, natively defined on spherical surfaces, with sinusoidal representation networks (SirenNets) that can be interpreted as learned Double Fourier Sphere embedding. We systematically evaluate positional embeddings and neural network architectures across various benchmarks and synthetic evaluation datasets. In contrast to previous approaches that require the combination of both positional encoding and neural networks to learn meaningful representations, we show that both spherical harmonics and sinusoidal representation networks are competitive on their own but set state-of-the-art performances across tasks when combined. The model code and experiments are available at https://github.com/marccoru/locationencoder.
Marc Rußwurm, Konstantin Klemmer, Esther Rolf, Robin Zbinden, Devis Tuia
ICLR5
2024 Knowledge-Aware Visual Question Generation for Remote Sensing Images
abstract
With the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific information or performing image retrieval. However, automatically generated image-based questions tend to be simplistic and template-based, which hinders the real deployment of question answering or visual dialogue systems. To enrich and diversify the questions, we propose a knowledge-aware remote sensing visual question generation model, KRSVQG, that incorporates external knowledge related to the image content to improve the quality and contextual understanding of the generated questions. The model takes an image and a related knowledge triplet from external knowledge sources as inputs and leverages image captioning as an intermediary representation to enhance the image grounding of the generated questions. To assess the performance of KRSVQG, we utilized two datasets that we manually annotated: NWPU-300 and TextRS-300. Results on these two datasets demonstrate that KRSVQG outperforms existing methods and leads to knowledge-enriched questions, grounded in both image and domain knowledge.
Li Mi, Javiera Castillo-Navarro, Devis Tuia
IGARSS4
2024 Semantic Segmentation of Coffee Plantations from Sentinel-2 Time Series
abstract
Coffee production serves as an important source of income for millions of farmers across many tropical countries. Given the scale of the production, accurate and up-to-date maps of coffee plantations are needed to detect, monitor, and mitigate potential negative environmental impacts. Such maps can be produced based on satellite imagery in a cost-effective and replicable way. However, this comes with certain challenges, including the dynamic spectral signature of coffee, complex topographies in which it is often cultivated, and the high cloud coverage common in tropical countries. In this work, we train a deep learning model to detect coffee plantations in Brazil directly from time series of Sentinel-2 images, alleviating the need for manual extraction of features. We show that the model outperforms significantly models trained on single images, is more robust against clouds and various seasonal patterns, and generalizes better to new regions.
Jan Pisl, Gaston Lenczner, Devis Tuia, Frank de Morsier
IGARSS3
2024 Training Visual Language Models with Object Detection: Grounded Change Descriptions in Satellite Images
abstract
Recently, generalist Vision Language Models (VLMs) have shown exceptional progress in tasks previously dominated by specialized computer vision models. This becomes more prevalent when visual grounding capabilities, such as the ability to reason over input text and image to generate bounding boxes around objects, are required. However, how these capabilities transfer to specialized domains such as remote sensing remains understudied, despite the recent increase in specialized models for Earth observation. In this work, we evaluate how grounding visual entities – by generating bounding-box coordinates – affects VLM performance in satellite imagery. To this end, we create two instruction-following tasks sourced from the xBD dataset, describing changes due to natural disasters observed in satellite images. We fine-tune several instances of MiniGPTv2, an open-source VLM with grounding capabilities, and evaluate their performance under the "grounded" vs. "not grounded" settings. We find that generating bounding boxes to refer to visual entities increases performance in tasks related to objects in the image, but only when the number of entities in the image is limited.
João Luis Prado, Syrielle Montariol, Javiera Castillo-Navarro, Devis Tuia, Antoine Bosselut
IGARSS4
2024 Estimating Fine-Grained Population Growth Rates from Coarse Census Data
abstract
Fine-grained population estimation is important for several domains, such as urban planning, public health, and humanitarian action. Due to limited resources, population maps with sufficient spatial resolution and temporal frequency are not available for many developing countries. Population estimates are often based on statistics available at the national or provincial level, which moreover are updated infrequently. The United Nations produce estimates of population growth rates, which are widely used to project population numbers to a target year, starting from the most recent census data. However, they use a simplified model of population growth with uniform growth rates for all urban, respectively rural areas within a country. This neglects the complex dynamics of population growth (e.g., growth rates in big cities are usually larger than in smaller urban areas) and leads to significant errors in population projections. In this work, we propose a methodology to estimate fine-grained population growth rates and present experimental results for Mozambique.
John E. Vargas-Munoz, Nando Metzger, Rodrigo Caye Daudt, Konrad Schindler, Devis Tuia
IGARSS5
2024 WildCLIP: Scene and Animal Attribute Retrieval from Camera Trap Data with Domain-Adapted Vision-Language Models
abstract
Abstract Wildlife observation with camera traps has great potential for ethology and ecology, as it gathers data non-invasively in an automated way. However, camera traps produce large amounts of uncurated data, which is time-consuming to annotate. Existing methods to label these data automatically commonly use a fixed pre-defined set of distinctive classes and require many labeled examples per class to be trained. Moreover, the attributes of interest are sometimes rare and difficult to find in large data collections. Large pretrained vision-language models, such as contrastive language image pretraining (CLIP), offer great promises to facilitate the annotation process of camera-trap data. Images can be described with greater detail, the set of classes is not fixed and can be extensible on demand and pretrained models can help to retrieve rare samples. In this work, we explore the potential of CLIP to retrieve images according to environmental and ecological attributes. We create WildCLIP by fine-tuning CLIP on wildlife camera-trap images and to further increase its flexibility, we add an adapter module to better expand to novel attributes in a few-shot manner. We quantify WildCLIP’s performance and show that it can retrieve novel attributes in the Snapshot Serengeti dataset. Our findings outline new opportunities to facilitate annotation processes with complex and multi-attribute captions. The code is available at https://github.com/amathislab/wildclip .
Valentin Gabeff, Marc Rußwurm, Devis Tuia, Alexander Mathis
Int. J. Comput. Vis.3
2024 Knowledge-Aware Text-Image Retrieval for Remote Sensing Images
abstract
Image-based retrieval in large Earth observation archives is challenging because one needs to navigate across thousands of candidate matches only with the query image as a guide. By using text as information supporting the visual query, the retrieval system gains in usability, but at the same time faces difficulties due to the diversity of visual signals that cannot be summarized by a short caption only. For this reason, as a matching-based task, cross-modal text–image retrieval often suffers from information asymmetry between text and images. To address this challenge, we propose a Knowledge-aware Text–Image Retrieval (KTIR) method for remote sensing images. By mining relevant information from an external knowledge graph, KTIR enriches the text scope available in the search query and alleviates the information gaps between text and images for better matching. Moreover, by integrating domain-specific knowledge, KTIR also enhances the adaptation of pretrained vision–language models to remote sensing applications. Experimental results on three commonly used remote sensing text–image retrieval benchmarks show that the proposed knowledge-aware method leads to varied and consistent retrievals, outperforming state-of-the-art retrieval methods.
Li Mi, Xianjie Dai, Javiera Castillo-Navarro, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.4
2023 Semi-Supervised Deep Learning Representations in Earth Observation Based Forest Management
abstract
In this study, we examine the potential of several self-supervised deep learning models in predicting forest attributes and detecting forest changes using ESA Sentinel-1 and Sentinel-2 images. The performance of the proposed deep learning models is compared to established conventional machine learning approaches. Studied use-cases include mapping of forest disturbance (windthrown forests, snowload damages) using deep change vector analysis, forest height mapping using UNet+ based models, Momentum contrast and regression modeling. Study areas were represented by several boreal forest sites in Finland. Our results indicate that developed methods allow to achieve superior classification and prediction accuracies compared to traditional methodologies and mimimize the amount of necessary in-situ forestry data.
Oleg Antropov, Matthieu Molinier, Ridvan Salih Kuzu, Lloyd H. Hughes, Marc Rußwurm, Devis Tuia, Corneliu Octavian Dumitru, Shaojia Ge, Sudipan Saha, Xiao Xiang Zhu 0001
IGARSS6
2023 Improving Few-Shot Object Detection with Object Part Proposals
abstract
Few-Shot Object Detection (FSOD) allows fast adaptation of an object detection model to new classes of objects using few examples per class. This has many applications, in particular in satellite and aerial observation, as it allows learning from experts who can only annotate a few examples for new classes and helps migrate models across tasks. In this work, we present a technique to improve the performance of FSOD in remote sensing by defining a contrastive loss that utilizes parts of objects. For this, we generate, what we call, Object Parts Proposals (OPPs) on the fly for each novel class, and use them to learn more robust features with an additional contrastive objective. We observe that training with OPPs brings a consistent improvement over the state-of-the-art when evaluating on the DIOR dataset.The code is available at https://github.com/arthurchevalley/Improving-FSOD-on-RSI-using-Sub-Parts.
Arthur Chevalley, Ciprian Tomoiaga, Marcin Detyniecki, Marc Rußwurm, Devis Tuia
IGARSS5
2023 Classification of Tropical Deforestation Drivers with Machine Learning and Satellite Image Time Series
abstract
Tropical deforestation is a major environmental problem with severe consequences such as carbon emissions or biodiversity loss. While much research focuses on monitoring and mapping deforestation, less attention is paid to understanding the various reasons and motivations behind it, known as deforestation drivers. Drivers can typically be identified from optical satellite imagery, but it is often necessary to view the deforested site at multiple points in time to determine the driver, making manual annotation of drivers laborious. In this work, we propose a deep learning model that classifies drivers from time series of Sentinel-2 images. The model combines convolutional, LSTM, and attention layers. To train the model, we use a large crowd-sourced dataset spanning across the tropics. We compare its results to other architectures and show that using time series can bring significant improvement in accuracy compared to single images, especially if a suitable architecture is used. Additionally, we analyze the attention scores produced by our model and show that it learns different strategies for different classes.
Jan Pisl, Lloyd H. Hughes, Marc Rußwurm, Devis Tuia
IGARSS4
2023 Detection of Settlements in Tanzania and Mozambique by Many Regional Few-Shot Models
abstract
In this work, we propose an approach to aid in mapping small settlements, which are often misclassified by models trained on a large-scale context (global or regional). We leverage pre-trained land cover models and few-shot learning to enhance the detection of these settlements. The backbone models are trained globally, but their application is localized through a spatial sampling strategy to address the challenge of detecting missed or unlabelled settlements. The proposed sampling strategy is based on the distance around a test patch and allows for the sampling of both backgrounds (non-settlements) points and settlements. Following this strategy results in a balanced dataset for model fine-tuning and ensures that the model is well-adapted to the local context. The idea is that nearby settlements share more similar properties, which is leveraged in our approach. We evaluate these transferred models by measuring the number of previously unmapped settlements detected by the fine-tuned classifier. For this, we manually annotated over two thousand buildings across two regions of Tanzania, previously unmapped in the original urban landcover product. Our results indicate the potential of the sampling approach, particularly when combined with a model pretrained with Momentum Contrast (MoCo). However, we also highlight the limitations in terms of spatial resolution of Sentinel-2 data for the detection of small settlements.
Marc Rußwurm, Lloyd H. Hughes, Giorgio Pasquali, Corneliu Octavian Dumitru, Devis Tuia
IGARSS5
2023 Text as a Richer Source of Supervision in Semantic Segmentation Tasks
abstract
This paper introduces TACOSS a text-image alignment approach that allows explainable land cover semantic segmentation by directly integrating semantic concepts encoded from texts. TACOSS combines convolutional neural networks for visual feature extraction with semantic embeddings provided by a language model. By leveraging contrastive learning approaches, we learn an alignment between the visual and the (fixed) textual representations. In addition to producing standard semantic segmentation outputs, our model enables interactive queries with RS images using natural language prompts. The experimental results obtained on 50cm resolution aerial data from Switzerland show that TACOSS performs similarly to a standard semantic segmentation model while allowing the flexible usage of in- and out-of-vocabulary terms for the interactions with the image.
Valérie Zermatten, Javiera Castillo-Navarro, Lloyd Hughes, Tobias Kellenberger, Devis Tuia
IGARSS5
2023 Revisiting Evaluation Metrics for Semantic Segmentation: Optimization and Evaluation of Fine-grained Intersection over Union
abstract
Semantic segmentation datasets often exhibit two types of imbalance: \textit{class imbalance}, where some classes appear more frequently than others and \textit{size imbalance}, where some objects occupy more pixels than others. This causes traditional evaluation metrics to be biased towards \textit{majority classes} (e.g. overall pixel-wise accuracy) and \textit{large objects} (e.g. mean pixel-wise accuracy and per-dataset mean intersection over union). To address these shortcomings, we propose the use of fine-grained mIoUs along with corresponding worst-case metrics, thereby offering a more holistic evaluation of segmentation techniques. These fine-grained metrics offer less bias towards large objects, richer statistical information, and valuable insights into model and dataset auditing. Furthermore, we undertake an extensive benchmark study, where we train and evaluate 15 modern neural networks with the proposed metrics on 12 diverse natural and aerial segmentation datasets. Our benchmark study highlights the necessity of not basing evaluations on a single metric and confirms that fine-grained mIoUs reduce the bias towards large objects. Moreover, we identify the crucial role played by architecture designs and loss functions, which lead to best practices in optimizing fine-grained metrics. The code is available at \href{https://github.com/zifuwanggg/JDTLosses}{https://github.com/zifuwanggg/JDTLosses}.
Zifu Wang, Maxim Berman, Amal Rannen Triki, Philip Torr 0001, Devis Tuia, Tinne Tuytelaars, Luc Van Gool, Jiaqian Yu, Matthew B. Blaschko
NeurIPS5
2022 Language Transformers for Remote Sensing Visual Question Answering
abstract
Remote sensing visual question answering (RSVQA) opens new avenues to promote the use of satellites data, by interfacing satellite image analysis with natural language processing. Capitalizing on the remarkable advances in natural language processing and computer vision, RSVQA aims at finding an answer to a question formulated by a human user about a remote sensing image. This is achieved by extracting representations from images and questions, and then fusing them in a joint representation. Focusing on the language part of the architecture, this study compares and evaluates the adequacy to the RSVQA task of two language models, a traditional recurrent neural network (Skip-thoughts) and a recent attentionbased Transformer (BERT). We study whether large transformer models are beneficial to the task and whether fine-tuning is needed for these models to perform at their best. Our findings show that the models benefit from fine-tuning language models and that RSVQA with BERT is slightly but consistently better when properly fine-tuned.
Christel Chappuis, Vincent Mendez, Eliot Walt, Sylvain Lobry, Bertrand Le Saux, Devis Tuia
IGARSS6
2022 Training Techniques for Presence-Only Habitat Suitability Mapping with Deep Learning
abstract
The goal of habitat suitability mapping is to predict the lo-cations in which a given species could be present. This is typically accomplished by statistical models which use envi-ronmental variables to predict species observation data. The relationship between the environmental characteristics of a location and the species that live there is likely to be quite complex, so deep learning models would seem natural to use. In practice, there are biases in the training data which present obstacles to standard deep learning approaches. First, large-scale species observation collections typically consist of presence-only data, which means we only have locations where a species has been observed (not where it has been confirmed to be absent). Second, the class distribution tends to be long-tailed. In this work we examine training tech-niques to mitigate these challenges: (i) a method for sharing species information between nearby observations and (ii) a curriculum learning strategy to reduce class imbalance early in training. These methods enable us to outperform state-of-the-art results on the GeoLifeCLEF 2020 dataset and suggest fruitful directions for future work.
Benjamin Kellenberger, Elijah Cole, Diego Marcos, Devis Tuia
IGARSS4
2022 Humans are Poor Few-Shot Classifiers for Sentinel-2 Land Cover
abstract
Learning to predict accurately from a few data samples is a central challenge in modern data-hungry machine learning. On natural images, human vision typically outperforms deep learning approaches on few-shot learning. However, we hypothesize that aerial and satellite images are more challenging to the human eye. This applies particularly when the image resolution is comparatively low, as with the 10m ground sampling distance of Sentinel-2. In this study, we benchmark model-agnostic meta-learning (MAML) algorithms against human participants on few-shot land cover classification with Sentinel-2 imagery on the Sen12MS dataset. We find that categorization of land cover from globally distributed regions is a difficult task for the participants, who classified the given images less accurately than the MAML-trained model and with a highly variable success rate. This suggest that hand-labeling land cover directly on Sentinel-2 imagery is not optimal when tackling a new land cover classification problem. Labeling only a few images and employing a trained meta-learning model to this task may lead to more accurate and consistent solutions compared to hand labeling by multiple individuals.
Marc Rußwurm, Sherrie Wang, Devis Tuia
IGARSS3
2022 Towards Efficient Correction of Coconut Tree Detection Errors
abstract
Coconut tree plantations are one of the main sources of income in several South Pacific countries. Thus, keeping track of the location of coconut trees is important for monitoring and post-disaster assessment. Although deep learning based object detectors can attain considerably accurate results, it is inevitable that errors will remain in the predictions obtained for a large test set. Since every mistake counts, in this work we propose a methodology to efficiently use the time of human annotators to find and correct a large part of erroneous coconut tree detections. We propose to use a Random forest classifer that finds detection errors to sort the regions of the image (tiles) in decreasing order of likeliness to have detection errors. In our experiments involving UAV images in the Kingdom of Tonga, the user could analyze only 24% of the tiles and correct approximatively 71% of the errors thanks to the sorting.
John E. Vargas-Munoz, Diego Schibli, Devis Tuia
IGARSS3
2022 Semantic Segmentation of Remote Sensing Images With Sparse Annotations
abstract
Training convolutional neural networks (CNNs) for very high-resolution images requires a large quantity of high-quality pixel-level annotations, which is extremely labor-intensive and time-consuming to produce. Moreover, professional photograph interpreters might have to be involved in guaranteeing the correctness of annotations. To alleviate such a burden, we propose a framework for semantic segmentation of aerial images based on incomplete annotations, where annotators are asked to label a few pixels with easy-to-draw scribbles. To exploit these sparse scribbled annotations, we propose the FEature and Spatial relaTional regulArization (FESTA) method to complement the supervised task with an unsupervised learning signal that accounts for neighborhood structures both in spatial and feature terms. For the evaluation of our framework, we perform experiments on two remote sensing image segmentation data sets involving aerial and satellite imagery, respectively. Experimental results demonstrate that the exploitation of sparse annotations can significantly reduce labeling costs, while the proposed method can help improve the performance of semantic segmentation when training on such annotations. The sparse labels and codes are publicly available for reproducibility purposes.https://github.com/Hua-YS/Semantic-Segmentation-with-Sparse-Labels
Yuansheng Hua, Diego Marcos, Lichao Mou, Xiao Xiang Zhu 0001, Devis Tuia
IEEE Geosci. Remote. Sens. Lett.5
2022 Wasserstein Adversarial Regularization for Learning With Label Noise
abstract
Noisy labels often occur in vision datasets, especially when they are obtained from crowdsourcing or Web scraping. We propose a new regularization method, which enables learning robust classifiers in presence of noisy data. To achieve this goal, we propose a new adversarial regularization scheme based on the Wasserstein distance. Using this distance allows taking into account specific relations between classes by leveraging the geometric properties of the labels space. Our Wasserstein Adversarial Regularization (WAR) encodes a selective regularization, which promotes smoothness of the classifier between some classes, while preserving sufficient complexity of the decision boundary between others. We first discuss how and why adversarial regularization can be used in the context of noise and then show the effectiveness of our method on five datasets corrupted with noisy labels: in both benchmarks and real datasets, WAR outperforms the state-of-the-art competitors.
Kilian Fatras, Bharath Bhushan Damodaran, Sylvain Lobry, Rémi Flamary, Devis Tuia, Nicolas Courty
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 A Semisupervised CRF Model for CNN-Based Semantic Segmentation With Sparse Ground Truth
abstract
Convolutional neural networks (CNNs) represent the new reference approach for semantic segmentation of very-high-resolution (VHR) images, due to their ability to automatically capture semantic information while learning relevant features. However, as for most supervised methods, the map accuracy depends on the quantity and quality of ground truth (GT) used to train them. The use of densely annotated data (i.e., a detailed, exhaustive, pixel-level GT) allows to obtain effective CNN models but normally implies high efforts in annotation. Such ground truth is often available in benchmark datasets on which new methods are tested, but not on real data for land-cover applications, where only sparse annotations might be sufficiently cost effective. A CNN model trained with such incomplete GT maps has the tendency to smooth object boundaries because they are never precisely delineated in the GT. To cope with those shortcomings, we propose to exploit the intermediate activation maps of the CNN and to deploy a semisupervised fully connected conditional random field (CRF). In comparison with competitors using the same sparse annotations, the proposed method is able to better fill part of the performance gap compared to a CNN trained on the densely annotated, but generally unavailable, GTs.
Luca Maggiolo, Diego Marcos, Gabriele Moser, Sebastiano B. Serpico, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.5
2022 deSpeckNet: Generalizing Deep Learning-Based SAR Image Despeckling
abstract
Deep learning (DL) has proven to be a suitable approach for despeckling synthetic aperture radar (SAR) images. So far, most DL models are trained to reduce speckle that follows a particular distribution, either using simulated noise or a specific set of real SAR images, limiting the applicability of these methods for real SAR images with unknown noise statistics. In this article, we present a DL method, deSpeckNet,1that estimates the speckle noise distribution and the despeckled image simultaneously. Since it does not depend on a specific noise model, deSpeckNet generalizes well across SAR acquisitions in a variety of landcover conditions. We evaluated the performance of deSpeckNet on single polarized Sentinel-1 images acquired in Indonesia, The Democratic Republic of Congo, and The Netherlands, a single polarized ALOS-2/PALSAR-2 image acquired in Japan and an Iceye X2 image acquired in Germany. In all cases, deSpeckNet was able to effectively reduce speckle and restore the images in high quality with respect to the state of the art.
Adugna G. Mullissa, Diego Marcos, Devis Tuia, Martin Herold 0001, Johannes Reiche
IEEE Trans. Geosci. Remote. Sens.3
2021 Geo-Data for Mapping Scenic Beauty: Exploring the Potential of Remote Sensing and Social Media
abstract
Scenic beauty is an important contributing factor to peoples' well-being. Modelling scenic beauty has been made possible at large scales with the availability of open-source remote sensing products. At the same time, the metadata available through social media, including tags and descriptions, offer a novel modelling alternative with a personalised view from the ground. This is especially relevant to policy applications. Using a crowdsourced landscape aesthetics dataset called ScenicOrNot as ground truth, we develop and test models to predict scenic beauty based on remotely sensed indicators and image metadata from social media (Flickr). Initial results show that both model types generate strong predictions of scenic beauty and model accuracy is maximised when the two are combined. Our research shows that both a top-view measurement using remote sensing and a social media-based measurement from the ground can be used to model landscape aesthetics in support of sustainable policy goals.
Ilan Havinga, Diego Marcos, Patrick W. Bogaart, Lars Hein, Devis Tuia
IGARSS5
2021 Liveability from Above: Understanding Quality of Life with Overhead Imagery and Deep Neural Networks
abstract
Urban planners are increasingly interested in understanding what makes a neighbourhood pleasant and liveable. In this paper, we use the overhead perspective as a new way to describe and understand liveability of city neighborhoods. We predict building quality scores from aerial images using deep neural networks and demonstrate that liveability can be predicted from overhead aerial images of a neighbourhood. We make our model interpretable by adding the intermediate task of predicting a list of housing factors, but found this to substantially degrade the results. This suggests that the unconstrained model used visual cues that are unrelated to the housing variables, and shows the difficulty of housing variable prediction from above due to the absence of visual cues such as facades.
Alex Levering, Diego Marcos, Devis Tuia
IGARSS3
2021 RSVQA Meets Bigearthnet: A New, Large-Scale, Visual Question Answering Dataset for Remote Sensing
abstract
Visual Question Answering is a new task that can facilitate the extraction of information from images through textual queries: it aims at answering an open-ended question formulated in natural language about a given image. In this work, we introduce a new dataset to tackle the task of visual question answering on remote sensing images: this large-scale, open access dataset extracts image/question/answer triplets from the BigEarthNet dataset. This new dataset contains close to 15 millions samples and is openly available. We present the dataset construction procedure, its characteristics and first results using a deep-learning based methodology. These first results show that the task of visual question answering is challenging and opens new interesting research avenues at the interface of remote sensing and natural language processing. The dataset and the code to create and process it are open and freely available on https://rsvqa.sylvainlobry.com/
Sylvain Lobry, Begüm Demir, Devis Tuia
IGARSS3
2021 Deploying machine learning to assist digital humanitarians: making image annotation in OpenStreetMap more efficient
abstract
John E. Vargas Muñoza*, Devis Tuiab & Alexandre X. Falcãoaa Laboratory of Image Data Science, Institute of Computing University of Campinas, Campinas, Brazilb Laboratory of Geo-information Science and Remote Sensing, Wageningen University & Research, Wageningen, The NetherlandsJohn E. Vargas Muñoz received the B.Sc. degree in informatics engineering from the National University of San Antonio Abad in Cusco, Cusco, Peru, in 2010, and the master's degree in computer science from the University of Campinas, Campinas, Brazil, in 2015. During 2017-2018, he worked on applications of machine learning to open geographical data, as a visiting Ph.D. student, in the Laboratory of Geo-information Science and Remote Sensing at Wageningen University, the Netherlands. In 2019, he received a Ph.D. in computer science from the University of Campinas, Campinas, Brazil. His research interests include machine learning, image processing, remote sensing image classification, and crowdsourced geographic information analysis.Devis Tuia (S'07, M’09, SM’15) received the Ph.D in environmental sciences at the University of Lausanne, Switzerland, in 2009. He was a Postdoc at the University of Valencia, the University of Colorado, Boulder, CO and EPFL Lausanne. Between 2014 and 2017, he was Assistant Professor at the University of Zurich. He is now Full Professor at the Geo-Information Science and Remote Sensing Laboratory at Wageningen University, the Netherlands. He is interested in algorithms for information extraction and data fusion of geospatial data (including remote sensing) using machine learning and computer vision. He serves as Associate Editor for IEEE TGRS and the Journal of the ISPRS. More info on http://devis.tuia.googlepages.com/Alexandre X. Falcão is full professor at the Institute of Computing, University of Campinas, Campinas, SP, Brazil. He received a B.Sc. in Electrical Engineering from the Federal University of Pernambuco, Recife, PE, Brazil, in 1988. He has worked in biomedical image processing, visualization and analysis since 1991. In 1993, he received a M.Sc. in Electrical Engineering from the University of Campinas, Campinas, SP, Brazil. During 1994-1996, he worked with the Medical Image Processing Group at the Department of Radiology, University of Pennsylvania, PA, USA, on interactive image segmentation for his doctorate. He got his doctorate in Electrical Engineering from the University of Campinas in 1996. In 1997, he worked in a project for Globo TV at a research center, CPqD-TELEBRAS in Campinas, developing methods for video quality assessment. His experience as professor of Computer Science and Engineering started in 1998 at the University of Campinas. His main research interests include image/video processing, visualization, and analysis; graph algorithms and dynamic programming; image annotation, organization, and retrieval; machine learning and pattern recognition; and image analysis applications in Biology, Medicine, Biometrics, Geology, and Agriculture.CONTACT John E. Vargas Muñoz [email protected] populations in rural areas of developing countries has attracted the attention of humanitarian mapping projects since it is important to plan actions that affect vulnerable areas. Recent efforts have tackled this problem as the detection of buildings in aerial images. However, the quality and the amount of rural building annotated data in open mapping services like OpenStreetMap (OSM) is not sufficient for training accurate models for such detection. Although these methods have the potential of aiding in the update of rural building information, they are not accurate enough to automatically update the rural building maps. In this paper, we explore a human-computer interaction approach and propose an interactive method to support and optimize the work of volunteers in OSM. The user is asked to verify/correct the annotation of selected tiles during several iterations and therefore improving the model with the new annotated data. The experimental results, with simulated and real user annotation corrections, show that the proposed method greatly reduces the amount of data that the volunteers of OSM need to verify/correct. The proposed methodology could benefit humanitarian mapping projects, not only by making more efficient the process of annotation but also by improving the engagement of volunteers.
John E. Vargas-Munoz, Devis Tuia, Alexandre X. Falcão
Int. J. Geogr. Inf. Sci.2
2020 Contextual Semantic Interpretability
Diego Marcos, Ruth Fong, Sylvain Lobry, Rémi Flamary, Nicolas Courty, Devis Tuia
ACCV (4)6
2020 ADVANCING DEEP LEARNING FOR EARTH SCIENCES: FROM HYBRID MODELING TO INTERPRETABILITY
abstract
Machine learning and deep learning in particular have made a huge impact in many fields of science and engineering. In the last decade, advanced deep learning methods have been developed and applied to remote sensing and geoscientific data problems extensively. Applications on classification and parameter retrieval are making a difference: methods are very accurate, can handle large amounts of data, and can deal with spatial and temporal data structures efficiently. Nevertheless, several important challenges need still to be addressed. First, current standard deep architectures cannot deal with long-range dependencies so distant driving processes (in space or time) are not captured, and they cannot cope with non-Euclidean spaces efficiently. Second, as other data-driven techniques, deep learning models do not necessarily respect physical or causal relations. Finally, deep learning models are still obscure and resistant to interpretability. Advances are needed to cope with arbitrary signal structures and data relations, physical plausibility and interpretability. This paper discusses about ways forward to develop new DL methods for the Earth sciences in all three directions.
Gustau Camps-Valls, Markus Reichstein, Xiao Xiang Zhu 0001, Devis Tuia
IGARSS4
2020 Learning Multi-Label Aerial Image Classification Under Label Noise: A Regularization Approach Using Word Embeddings
abstract
Training deep neural networks requires well-annotated datasets. However, real world datasets are often noisy, especially in a multi-label scenario, i.e. where each data point can be attributed to more than one class. To this end, we propose a regularization method to learn multi-label classification networks from noisy data. This regularization is based on the assumption that semantically close classes are more likely to appear together in a given image. Hereby, we encode label correlations with prior knowledge and regularize noisy network predictions using label correlations. To evaluate its effectiveness, we perform experiments on a mutli-label aerial image dataset contaminated with controlled levels of label noise. Results indicate that networks trained using the proposed method outperform those directly learned from noisy labels and that the benefits increase proportionally to the amount of noise present.
Yuansheng Hua, Sylvain Lobry, Lichao Mou, Devis Tuia, Xiao Xiang Zhu 0001
IGARSS4
2020 Interpretable Scenicness from Sentinel-2 Imagery
abstract
Landscape aesthetics, or scenicness, has been identified as an important ecosystem service that contribute to human health and well-being. Currently there are no methods to inventorize landscape scenicness on a large scale. In this paper we study how to upscale local assessments of scenicness provided by human observers, and we do so by using satellite images. Moreover, we develop an explicitly interpretable CNN model that allows assessing the connections between landscape scenicness and the presence of specific landcover types. To generate the landscape scenicness ground truth, we use the ScenicOrNot crowdsourcing database, which provides geo-referenced, human-based scenicness estimates for ground based photos in Great Britain. Our results show that it is feasible to predict landscape scenicness based on satellite imagery. The interpretable model performs comparably to an unconstrained model, suggesting that it is possible to learn a semantic bottleneck that represents well the present landcover classes and still contains enough information to accurately predict the location's scenicness.
Alex Levering, Diego Marcos, Sylvain Lobry, Devis Tuia
IGARSS4
2020 Fine-grained landuse characterization using ground-based pictures: a deep learning solution based on globally available data
abstract
We study the problem of landuse characterization at the urban-object level using deep learning algorithms. Traditionally, this task is performed by surveys or manual photo interpretation, which are expensive and difficult to update regularly. We seek to characterize usages at the single object level and to differentiate classes such as educational institutes, hospitals and religious places by visual cues contained in side-view pictures from Google Street View (GSV). These pictures provide geo-referenced information not only about the material composition of the objects but also about their actual usage, which otherwise is difficult to capture using other classical sources of data such as aerial imagery. Since the GSV database is regularly updated, this allows to consequently update the landuse maps, at lower costs than those of authoritative surveys. Because every urban-object is imaged from a number of viewpoints with street-level pictures, we propose a deep-learning based architecture that accepts arbitrary number of GSV pictures to predict the fine-grained landuse classes at the object level. These classes are taken from OpenStreetMap. A quantitative evaluation of the area of Île-de-France, France shows that our model outperforms other deep learning-based methods, making it a suitable alternative to manual landuse characterization.
Shivangi Srivastava, John E. Vargas-Munoz, Sylvain Lobry, Devis Tuia
Int. J. Geogr. Inf. Sci.4
2020 RSVQA: Visual Question Answering for Remote Sensing Data
abstract
This article introduces the task of visual question answering for remote sensing data (RSVQA). Remote sensing images contain a wealth of information, which can be useful for a wide range of tasks, including land cover classification, object counting, or detection. However, most of the available methodologies are task-specific, thus inhibiting generic and easy access to the information contained in remote sensing data. As a consequence, accurate remote sensing product generation still requires expert knowledge. With RSVQA, we propose a system to extract information from remote sensing data that is accessible to every user: we use questions formulated in natural language and use them to interact with the images. With the system, images can be queried to obtain high-level information specific to the image content or relational dependencies between objects visible in the images. Using an automatic method introduced in this article, we built two data sets (using low- and high-resolution data) of image/question/answer triplets. The information required to build the questions and answers is queried from OpenStreetMap (OSM). The data sets can be used to train (when using supervised methods) and evaluate models to solve the RSVQA task. We report the results obtained by applying a model based on convolutional neural networks (CNNs) for the visual part and a recurrent neural network (RNN) for the natural language part of this task. The model is trained on the two data sets, yielding promising results in both cases.
Sylvain Lobry, Diego Marcos, Jesse Murray, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.4
2019 Optimal Transport for Multi-source Domain Adaptation under Target Shift
abstract
In this paper, we tackle the problem of reducing discrepancies between multiple domains, i.e. multi-source domain adaptation, and consider it under the target shift assumption: in all domains we aim to solve a classification problem with the same output classes, but with different labels proportions. This problem, generally ignored in the vast majority of domain adaptation papers, is nevertheless critical in real-world applications, and we theoretically show its impact on the success of the adaptation. Our proposed method is based on optimal transport, a theory that has been successfully used to tackle adaptation problems in machine learning. The introduced approach, Joint Class Proportion and Optimal Transport (JCPOT), performs multi-source adaptation and target shift correction simultaneously by learning the class probabilities of the unlabeled target sample and the coupling allowing to align two (or more) probability distributions. Experiments on both synthetic and real-world data (satellite image pixel classification) task show the superiority of the proposed method over the state-of-the-art.
Ievgen Redko, Nicolas Courty, Rémi Flamary, Devis Tuia
AISTATS4
2019 Adaptive Compression-based Lifelong Learning
Shivangi Srivastava, Maxim Berman, Matthew B. Blaschko, Devis Tuia
BMVC4
2019 Visual Question Answering From Remote Sensing Images
abstract
Remote sensing images carry wide amounts of information beyond land cover or land use. Images contain visual and structural information that can be queried to obtain high level information about specific image content or relational dependencies between the objects sensed. This paper explores the possibility to use questions formulated in natural language as a generic and accessible way to extract this type of information from remote sensing images, i.e. visual question answering. We introduce an automatic way to create a dataset using OpenStreetMap1data and present some preliminary results. Our proposed approach is based on deep learning, and is trained using our new dataset.
Sylvain Lobry, Jesse Murray, Diego Marcos, Devis Tuia
IGARSS4
2019 Zoom In, Zoom Out: Injecting Scale Invariance into Landuse Classification CNNs
abstract
We propose a Convolutional Neural Network (CNN), which encodes local scale invariance and equivariance in a multiresolution, multi-sensor image classification task. We show that the locally scale invariant model achieves results that are in line with state-of-the-art. The scale invariant and equivariant models also prove to be more robust to reductions in training data and number of filters used in each convolutional layer. These results demonstrate the benefit of disentangling scale within the learned features of CNNs, in particular when processing multi-resolution imagery. This is beneficial in the two studied cases: when training data is limited, or when the number of model parameters must be kept to a minimum.
Jesse Murray, Diego Marcos, Devis Tuia
IGARSS3
2019 Interactive Coconut Tree Annotation Using Feature Space Projections
abstract
The detection and counting of coconut trees in aerial images are important tasks for environment monitoring and post-disaster assessment. Recent deep-learning-based methods can attain accurate results, but they require a reasonably high number of annotated training samples. In order to obtain such large training sets with considerably reduced human effort, we present a semi-automatic sample annotation method based on the 2D t-SNE projection of the sample feature space. The proposed approach can facilitate the construction of effective training sets more efficiently than using the traditional manual annotation, as shown in our experimental results with VHR images from the Kingdom of Tonga.
John E. Vargas-Munoz, Alexandre X. Falcão, Devis Tuia
IGARSS4
2019 Nonlinear Feature Normalization for Hyperspectral Domain Adaptation and Mitigation of Nonlinear Effects
abstract
Domain adaptation in remote sensing aims at the automatic knowledge transfer between a set of multitemporal and multisource images. This process is often impaired by nonlinear effects in the data, e.g., varying illumination conditions, different viewing angles, and geometry-dependent reflection. In this paper, we introduce the Nonlinear Feature Normalization (NFN), a fast and robust way to align the spectral characteristics of multiple hyperspectral data sets. NFN employs labeled training spectra for the different classes in an image to describe the corresponding underlying low-dimensional manifold structure. A linear basis for data representation is defined by arbitrary class reference vectors, and the image is aligned to the new basis in the same space. This results in samples of the same class being pulled closer together and samples of different classes pushed apart. NFN transforms the data in its original domain, preserving physical interpretability. We use the continuous invertibility of NFN to derive the NFN Alignment (NFNalign) transformation, which can be used for domain adaptation, by transforming one data set to the domain of a chosen reference. The evaluation is performed on multiple hyperspectral data sets as well as our new benchmark for multitemporal hyperspectral data. In a first step, we show that the NFN transformation successfully mitigates nonlinear effects by comparing classification of the linear Spectral Angle Mapper on original and transformed data. Finally, we demonstrate successful domain adaptation with NFNalign by applying it to the task of hyperspectral data preprocessing. The evaluation shows that our approach for alignment of multitemporal data produces high-spectral similarity and successfully allows knowledge transfer, e.g., of classifier models and training data.
Wolfgang Groß, Devis Tuia, Uwe Sörgel, Wolfgang Middelmann
IEEE Trans. Geosci. Remote. Sens.2
2019 Half a Percent of Labels is Enough: Efficient Animal Detection in UAV Imagery Using Deep CNNs and Active Learning
abstract
We present an Active Learning (AL) strategy for reusing a deep Convolutional Neural Network (CNN)-based object detector on a new data set. This is of particular interest for wildlife conservation: given a set of images acquired with an Unmanned Aerial Vehicle (UAV) and manually labeled ground truth, our goal is to train an animal detector that can be reused for repeated acquisitions, e.g., in follow-up years. Domain shifts between data sets typically prevent such a direct model application. We thus propose to bridge this gap using AL and introduce a new criterion called Transfer Sampling (TS). TS uses Optimal Transport (OT) to find corresponding regions between the source and the target data sets in the space of CNN activations. The CNN scores in the source data set are used to rank the samples according to their likelihood of being animals, and this ranking is transferred to the target data set. Unlike conventional AL criteria that exploit model uncertainty, TS focuses on very confident samples, thus allowing quick retrieval of true positives in the target data set, where positives are typically extremely rare and difficult to find by visual inspection. We extend TS with a new window cropping strategy that further accelerates sample retrieval. Our experiments show that with both strategies combined, less than half a percent of oracle-provided labels are enough to find almost 80% of the animals in challenging sets of UAV images, beating all baselines by a margin.
Benjamin Kellenberger, Diego Marcos, Sylvain Lobry, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.4
2018 Learning Deep Structured Active Contours End-to-End
abstract
The world is covered with millions of buildings, and precisely knowing each instance's position and extents is vital to a multitude of applications. Recently, automated building footprint segmentation models have shown superior detection accuracy thanks to the usage of Convolutional Neural Networks (CNN). However, even the latest evolutions struggle to precisely delineating borders, which often leads to geometric distortions and inadvertent fusion of adjacent building instances. We propose to overcome this issue by exploiting the distinct geometric properties of buildings. To this end, we present Deep Structured Active Contours (DSAC), a novel framework that integrates priors and constraints into the segmentation process, such as continuous boundaries, smooth edges, and sharp corners. To do so, DSAC employs Active Contour Models (ACM), a family of constraint- and prior-based polygonal models. We learn ACM parameterizations per instance using a CNN, and show how to incorporate all components in a structured output model, making DSAC trainable end-to-end. We evaluate DSAC on three challenging building instance segmentation datasets, where it compares favorably against state-of-the-art. Code will be made available on https://github.com/dmarcosg/DSAC.
Diego Marcos, Devis Tuia, Benjamin Kellenberger, Lisa Zhang 0003, Min Bai, Renjie Liao 0001, Raquel Urtasun
CVPR2
2018 DeepJDOT: Deep Joint Distribution Optimal Transport for Unsupervised Domain Adaptation
Bharath Bhushan Damodaran, Benjamin Kellenberger, Rémi Flamary, Devis Tuia, Nicolas Courty
ECCV (4)4
2018 Detecting Animals in Repeated UAV Image Acquisitions by Matching CNN Activations with Optimal Transport
abstract
Repeated animal censuses are crucial for wildlife parks to ensure ecological equilibriums. They are increasingly conducted using images generated by Unmanned Aerial Vehicles (UAVs), often coupled to semi-automatic object detection methods. Such methods have shown great progress also thanks to the employment of Convolutional Neural Networks (CNNs), but even the best models trained on the data acquired in one year struggle predicting animal abundances in subsequent campaigns due to the inherent shift between the datasets. In this paper we adapt a CNN-based animal detector to a follow-up UAV dataset by employing an unsupervised domain adaptation method based on Optimal Transport. We show how to infer updated labels from the source dataset by means of an ensemble of bootstraps. Our method increases the precision compared to the unmodified CNN, while not requiring additional labels from the target set.
Benjamin Kellenberger, Diego Marcos, Nicolas Courty, Devis Tuia
IGARSS4
2018 Improving Maps from CNNs Trained with Sparse, Scribbled Ground Truths Using Fully Connected CRFs
abstract
Convolutional Neural Networks (CNNs) have become the new standard for semantic segmentation of very high resolution images. But as for other methods, the map accuracy depends on the quantity and quality of ground truth used to train them. Having densely annotated data, i.e. a detailed, pixel-level ground truth (GT), allows obtaining effective models, but requires high efforts in annotation. For this reason, it is more common and efficient to work with point or scribbled annotations rather than with dense ones. A CNN model trained with such incomplete ground truths tends to mischaracterize the shapes of the objects and to be inaccurate near their boundaries. We propose to use an approximation of a fully connected Conditional Random Field (CRF) to solve these issues, in which long range connections are accounted for through auxiliary nodes based on clustering of CNN activation features. Experiments on the ISPRS Vaihingen benchmark, where a CNN is trained only with a non-dense, scribbled ground truth, show that the proposed method can fill part of the performance gap with respect to models trained on the densely annotated, but unrealistic, ground truth.
Luca Maggiolo, Diego Marcos, Gabriele Moser, Devis Tuia
IGARSS4
2018 Discovering Temporal Patterns of Air Quality in Different Parts of Europe with Data Driven Feature Extraction
abstract
Air quality is strongly affecting human lifestyle all over the world, and its impact is apparent on healthcare, sustainable development, welfare and public administration policies. Accurate understanding of the polluting processes requires to analyze huge volumes of records, so that significant patterns and regularities can be detected. In this paper, we introduce a framework to explore the air pollution dynamics over all Europe by means of a data driven feature extraction approach. Taking advantage of MODIS records, we are able to investigate daily trends of air quality from 2003 to 2016. By means of an automatic learning scheme based on mutual information maximization, we extract the most significant patterns in the dataset. Experimental results show that the proposed approach is able to identify relevant air pollution trends that can be associated with specific physical phenomena on ground.
Andrea Marinoni, Paolo Gamba, Daniele De Vecchi, Devis Tuia
IGARSS4
2018 A Deep Network Approach to Multitemporal Cloud Detection
abstract
We present a deep learning model with temporal memory to detect clouds in image time series acquired by the Seviri imager mounted on the Meteosat Second Generation (MSG) satellite. The model provides pixel-level cloud maps with related confidence and propagates information in time via a recurrent neural network structure. With a single model, we are able to outline clouds along all year and during day and night with high accuracy.
Devis Tuia, Benjamin Kellenberger, Adrián Pérez-Suay, Gustau Camps-Valls
IGARSS1
2018 Correcting Misaligned Rural Building Annotations in Open Street Map Using Convolutional Neural Networks Evidence
abstract
Mapping rural buildings in developing countries is crucial to monitor and plan in those vulnerable areas. Despite the existence of some rural building annotations in OpenStreetMap (OSM), those are of insufficient quantity and quality to train models able to map large areas accurately. In particular, these annotations are very often misaligned with respect to the buildings that are present in updated aerial imagery. We propose a Markov Random Field (MRF) method to correct misaligned rural building annotations. To do so, our method uses i) the correlation between candidate aligned OSM annotations and buildings roughly detected on aerial images and ii) the local consistency of the alignment vectors.
John E. Vargas-Munoz, Diego Marcos, Sylvain Lobry, Jefersson A. dos Santos, Alexandre X. Falcão, Devis Tuia
IGARSS6
2018 Best Practices to Train Deep Models on Imbalanced Datasets - A Case Study on Animal Detection in Aerial Imagery
Benjamin Kellenberger, Diego Marcos, Devis Tuia
ECML/PKDD (3)3
2018 Decision Fusion With Multiple Spatial Supports by Conditional Random Fields
abstract
Classification of remotely sensed images into land cover or land use is highly dependent on geographical information at least at two levels. First, land cover classes are observed in a spatially smooth domain separated by sharp region boundaries. Second, land classes and observation scale are also tightly intertwined: they tend to be consistent within areas of homogeneous appearance, or regions, in the sense that all pixels within a roof should be classified as roof, independently on the spatial support used for the classification. In this paper, we follow these two observations and encode them as priors in an energy minimization framework based on conditional random fields (CRFs), where classification results obtained at pixel and region levels are probabilistically fused. The aim is to enforce the final maps to be consistent not only in their own spatial supports (pixel and region) but also across supports, i.e., by getting the predictions on the pixel lattice and on the set of regions to agree. To this end, we define an energy function with three terms: 1) a data term for the individual elements in each support (support-specific nodes); 2) spatial regularization terms in a neighborhood for each of the supports (support-specific edges); and 3) a regularization term between individual pixels and the region containing each of them (intersupports edges). We utilize these priors in a unified energy minimization problem that can be optimized by standard solvers. The proposed 2LCRF model consists of a CRF defined over a bipartite graph, i.e., two interconnected layers within a single graph accounting for interlattice connections. 2LCRF is tested on two very high-resolution data sets involving submetric satellite and subdecimeter aerial data. In all cases, 2LCRF improves the result obtained by the independent base model (either random forests or convolutional neural networks) and by standard CRF models enforcing smoothness in the spatial domain.
Devis Tuia, Michele Volpi, Gabriele Moser
IEEE Trans. Geosci. Remote. Sens.1
2017 Rotation Equivariant Vector Field Networks
abstract
In many computer vision tasks, we expect a particular behavior of the output with respect to rotations of the input image. If this relationship is explicitly encoded, instead of treated as any other variation, the complexity of the problem is decreased, leading to a reduction in the size of the required model. In this paper, we propose the Rotation Equivariant Vector Field Networks (RotEqNet), a Convolutional Neural Network (CNN) architecture encoding rotation equivariance, invariance and covariance. Each convolutional filter is applied at multiple orientations and returns a vector field representing magnitude and angle of the highest scoring orientation at every spatial location. We develop a modified convolution operator relying on this representation to obtain deep architectures. We test RotEqNet on several problems requiring different responses with respect to the inputs' rotation: image classification, biomedical image segmentation, orientation estimation and patch matching. In all cases, we show that RotEqNet offers extremely compact models in terms of number of parameters and provides results in line to those of networks orders of magnitude larger.
Diego Marcos, Michele Volpi, Nikos Komodakis, Devis Tuia
ICCV4
2017 Fast animal detection in UAV images using convolutional neural networks
abstract
Illegal wildlife poaching poses one severe threat to the environment. Measures to stem poaching have only been with limited success, mainly due to efforts required to keep track of wildlife stock and animal tracking. Recent developments in remote sensing have led to low-cost Unmanned Aerial Vehicles (UAVs), facilitating quick and repeated image acquisitions over vast areas. In parallel, progress in object detection in computer vision yielded unprecedented performance improvements, partially attributable to algorithms like Convolutional Neural Networks (CNNs). We present an object detection method tailored to detect large animals in UAV images. We achieve a substantial increase in precision over a robust state-of-the-art model on a dataset acquired over the Kuzikus wildlife reserve park in Namibia. Furthermore, our model processes data at over 72 images per second, as opposed 3 for the baseline, allowing for real-time applications.
Benjamin Kellenberger, Michele Volpi, Devis Tuia
IGARSS3
2017 Joint height estimation and semantic labeling of monocular aerial images with CNNS
abstract
We aim to jointly estimate height and semantically label monocular aerial images. These two tasks are traditionally addressed separately in remote sensing, despite their strong correlation. Therefore, a model learning both height and classes jointly seems advantageous and so, we propose a multitask Convolutional Neural Network (CNN) architecture with two losses: one performing semantic labeling, and another predicting normalized Digital Surface Model (nDSM) from the pixel values. Since the nDSM/height information is used only in the second loss, there is no need to have a nDSM map at test time, and the model can estimate height automatically on new images. We test our proposed method on a set of sub-decimeter resolution images and show that our model equals the performances of two separate models, but at the cost of a single one.
Shivangi Srivastava, Michele Volpi, Devis Tuia
IGARSS3
2017 Post classification smoothing in sub-decimeter resolution images with semi-supervised label propagation
abstract
In this paper, we propose a post classification smoothing method aimed at improving the accuracy and visual appearance of sub-decimeter image classification results. Starting from the class confidence maps of a supervised classifier, we find a set of high confidence markers and propagate labels on an extended region adjacency graph. We apply the proposed method on a challenging 5cm resolution dataset over Potsdam, Germany. The proposed algorithm outperforms state-of-the-art post classification smoothing algorithms both when the classifier is trained specifically on the image and when it is trained and tested in different set of images.
John E. Vargas-Munoz, Devis Tuia, Jefersson A. dos Santos, Alexandre X. Falcão
IGARSS2
2017 Optimal Transport for Domain Adaptation
abstract
Domain adaptation is one of the most challenging tasks of modern data analytics. If the adaptation is done correctly, models built on a specific data representation become more robust when confronted to data depicting the same classes, but described by another observation system. Among the many strategies proposed, finding domain-invariant representations has shown excellent properties, in particular since it allows to train a unique classifier effective in all domains. In this paper, we propose a regularized unsupervised optimal transportation model to perform the alignment of the representations in the source and target domains. We learn a transportation plan matching both PDFs, which constrains labeled samples of the same class in the source domain to remain close during transport. This way, we exploit at the same time the labeled samples in the source and the distributions observed in both domains. Experiments on toy and challenging real visual adaptation examples show the interest of the method, that consistently outperforms state of the art approaches. In addition, numerical experiments show that our approach leads to better performances on domain invariant deep learning features and can be easily adapted to the semi-supervised case where few labeled samples are available in the target domain.
Nicolas Courty, Rémi Flamary, Devis Tuia, Alain Rakotomamonjy
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Toward Seamless Multiview Scene Analysis From Satellite to Street Level
abstract
In this paper, we discuss and review how combined multiview imagery from satellite to street level can benefit scene analysis. Numerous works exist that merge information from remote sensing and images acquired from the ground for tasks such as object detection, robots guidance, or scene understanding. What makes the combination of overhead and street-level images challenging are the strongly varying viewpoints, the different scales of the images, their illuminations and sensor modality, and time of acquisition. Direct (dense) matching of images on a per-pixel basis is thus often impossible, and one has to resort to alternative strategies that will be discussed in this paper. For such purpose, we review recent works that attempt to combine images taken from the ground and overhead views for purposes like scene registration, reconstruction, or classification. After the theoretical review, we present three recent methods to showcase the interest and potential impact of such fusion on real applications (change detection, image orientation, and tree cataloging), whose logic can then be reused to extend the use of ground-based images in remote sensing andvice versa. Through this review, we advocate that cross fertilization between remote sensing, computer vision, and machine learning is very valuable to make the best of geographic data available from Earth observation sensors and ground imagery. Despite its challenges, we believe that integrating these complementary data sources will lead to major breakthroughs in Big GeoData. It will open new perspectives for this exciting and emerging field.
Sébastien Lefèvre, Devis Tuia, Jan Dirk Wegner, Timothée Produit, Ahmed Samy Nassar
Proc. IEEE2
2017 Dense Semantic Labeling of Subdecimeter Resolution Images With Convolutional Neural Networks
abstract
Semantic labeling (or pixel-level land-cover classification) in ultrahigh-resolution imagery (<;10 cm) requires statistical models able to learn high-level concepts from spatial data, with large appearance variations. Convolutional neural networks (CNNs) achieve this goal by learning discriminatively a hierarchy of representations of increasing abstraction. In this paper, we present a CNN-based system relying on a downsample-then-upsample architecture. Specifically, it first learns a rough spatial map of high-level representations by means of convolutions and then learns to upsample them back to the original resolution by deconvolutions. By doing so, the CNN learns to densely label every pixel at the original resolution of the image. This results in many advantages, including: 1) the state-of-the-art numerical accuracy; 2) the improved geometric accuracy of predictions; and 3) high efficiency at inference time. We test the proposed system on the Vaihingen and Potsdam subdecimeter resolution data sets, involving the semantic labeling of aerial images of 9- and 5-cm resolution, respectively. These data sets are composed by many large and fully annotated tiles, allowing an unbiased evaluation of models making use of spatial information. We do so by comparing two standard CNN architectures with the proposed one: standard patch classification, prediction of local label patches by employing only convolutions, and full patch labeling by employing deconvolutions. All the systems compare favorably or outperform a state-of-the-art baseline relying on superpixels and powerful appearance descriptors. The proposed full patch labeling CNN outperforms these models by a large margin, also showing a very appealing inference time.
Michele Volpi, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.2
2016 Geospatial Correspondences for Multimodal Registration
abstract
The growing availability of very high resolution (<;1 m/pixel) satellite and aerial images has opened up unprecedented opportunities to monitor and analyze the evolution of land-cover and land-use across the world. To do so, images of the same geographical areas acquired at different times and, potentially, with different sensors must be efficiently parsed to update maps and detect land-cover changes. However, a naϊve transfer of ground truth labels from one location in the source image to the corresponding location in the target image is generally not feasible, as these images are often only loosely registered (with up to ± 50m of non-uniform errors). Furthermore, land-cover changes in an area over time must be taken into account for an accurate ground truth transfer. To tackle these challenges, we propose a mid-level sensor-invariant representation that encodes image regions in terms of the spatial distribution of their spectral neighbors. We incorporate this representation in a Markov Random Field to simultaneously account for nonlinear mis-registrations and enforce locality priors to find matches between multi-sensor images. We show how our approach can be used to assist in several multimodal land-cover update and change detection problems.
Diego Marcos, Raffay Hamid, Devis Tuia
CVPR3
2016 Learning rotation invariant convolutional filters for texture classification
abstract
We present a method for learning discriminative filters using a shallow Convolutional Neural Network (CNN). We encode rotation invariance directly in the model by tying the weights of groups of filters to several rotated versions of the canonical filter in the group. These filters can be used to extract rotation invariant features well-suited for image classification. We test this learning procedure on a texture classification benchmark, where the orientations of the training images differ from those of the test images. We obtain results comparable to the state-of-the-art. Compared to standard shallow CNNs, the proposed method obtains higher classification performance while reducing by an order of magnitude the number of parameters to be learned.
Diego Marcos, Michele Volpi, Devis Tuia
ICPR3
2016 Optimal transport for data fusion in remote sensing
abstract
One of the main objective of data fusion is the integration of several acquisition of the same physical object, in order to build a new consistent representation that embeds all the information from the different modalities. In this paper, we propose the use of optimal transport theory as a powerful mean of establishing correspondences between the modalities. After reviewing important properties and computational aspects, we showcase its application to three remote sensing fusion problems: domain adaptation, time series averaging and change detection in LIDAR data.
Nicolas Courty, Rémi Flamary, Devis Tuia, Thomas Corpetti
IGARSS3
2016 Solving structured segmentation of aerial images as puzzles
abstract
Traditional approaches to structured semantic segmentation employ appearance-based classifiers to provide a class-likelihood at each spatial location and then post-process it with Markov Random Fields (MRF) to enforce label smoothness and structure in the output space. The spatial support for such techniques is usually a patch of pixels, which makes the prediction over-smoothed because the borders of objects are not explicitly taken into account. This is further exacerbated by MRF post-processing employing the standard Potts model, which tends to further over-smooth predictions at boundaries. In this paper, we propose a different but related approach: we optimize an energy function finding the optimal combination of small ground truth (GT) tiles from training data over predictions at test time, effectively solving a puzzle. We optimize over a first configuration given by a Convolutional Neural Network (CNN) output.
Diego Marcos, Michele Volpi, Devis Tuia
IGARSS3
2016 Getting pixels and regions to agree with conditional random fields
abstract
Land cover / land use classification of remotely sensed images is inherently geographical. The use of spatial information, accounting for neighborhood relationship and spatial smoothness of geographical objects, made its proofs in countless occasions and, especially when considering very high resolution images, methods ignoring spatial context do not perform well. In this paper, we propose a hybrid dual-layer conditional random field model that enforces spatial smoothness and consistency between the pixel and region-based maps. We formulate these intuitions as a standard energy minimization problem, and we show that finding a joint solution over both output spaces leads to strong improvements in the numerical and visual senses.
Devis Tuia, Michele Volpi, Gabriele Moser
IGARSS1
2016 Semantic labeling of aerial images by learning class-specific object proposals
abstract
Land-cover and land-use semantic labeling in centimeter resolution imagery (ultra-high resolution) is mostly performed by supervised classification of informative descriptors extracted from spatially coherent but small objects (e.g. superpixels or patches). In this paper, we propose an extension of this reasoning by proposing a class-specific, multi-scale and bottom-up object proposal strategy to perform semantic labeling. Specifically, we rely on a fully trainable boundary (edge) detector, allowing us to extract class-specific object-proposals. Such proposals enable training rich appearance and object models as well as enhanced spatial reasoning. We evaluate the proposed strategy on the Vaihingen dataset with promising results.
Michele Volpi, Devis Tuia
IGARSS2
2016 Kernel Low-Rank and Sparse Graph for Unsupervised and Semi-Supervised Classification of Hyperspectral Images
abstract
In this paper, we present a graph representation that is based on the assumption that data live on a union of manifolds. Such a representation is based on sample proximities in reproducing kernel Hilbert spaces and is thus linear in the feature space and nonlinear in the original space. Moreover, it also expresses sample relationships under sparse and low-rank constraints, meaning that the resulting graph will have limited connectivity (sparseness) and that samples belonging to the same group will be likely to be connected together and not with those from other groups (low rankness). We present this graph representation as a general representation that can be then applied to any graph-based method. In the experiments, we consider the clustering of hyperspectral images and semi-supervised classification (one class and multiclass).
Frank de Morsier, Maurice Borgeaud, Volker Gass, Jean-Philippe Thiran, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.5
2016 Nonconvex Regularization in Remote Sensing
abstract
In this paper, we study the effect of different regularizers and their implications in high-dimensional image classification and sparse linear unmixing. Although kernelization or sparse methods are globally accepted solutions for processing data in high dimensions, we present here a study on the impact of the form of regularization used and its parameterization. We consider regularization via traditional squared (ℓ2) and sparsity-promoting (ℓ1) norms, as well as more unconventional nonconvex regularizers (ℓpand log sum penalty). We compare their properties and advantages on several classification and linear unmixing tasks and provide advices on the choice of the best regularizer for the problem at hand. Finally, we also provide a fully functional toolbox for the community.
Devis Tuia, Rémi Flamary, Michel Barlaud
IEEE Trans. Geosci. Remote. Sens.1
2016 Discriminative Multiple Kernel Learning for Hyperspectral Image Classification
abstract
In this paper, we propose a discriminative multiple kernel learning (DMKL) method for spectral image classification. The core idea of the proposed method is to learn an optimal combined kernel from predefined basic kernels by maximizing separability in reproduction kernel Hilbert space. DMKL achieves the maximum separability via finding an optimal projective direction according to statistical significance, which leads to the minimum within-class scatter and maximum between-class scatter instead of a time-consuming search for the optimal kernel combination. Fisher criterion (FC) and maximum margin criterion (MMC) are used to find the optimal projective direction, thus leading to two variants of the proposed method, DMKL-FC and DMKL-MMC, respectively. After learning the projective direction, all basic kernels are projected to generate a discriminative combined kernel. Three merits are realized by DMKL. First, DMKL can achieve a substantial improvement in classification performance without strict limitation for selection of basic kernels. Second, the discriminating scales of a Gaussian kernel, the useful bands for classification, and the competitive sizes of spatial filters can be selected by ranking the corresponding weights, where the large weights correspond to the most relevant. Third, DMKL reduces the computational burden by requiring fewer support vectors. Experiments are conducted on two hyperspectral data sets and one multispectral data set. The corresponding experimental results demonstrate that the proposed algorithms can achieve the best performance with satisfactory computational efficiency for spectral image classification, compared with several state-of-the-art algorithms.
Qingwang Wang, Yanfeng Gu, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.3
2015 Weakly supervised alignment of multisensor images
abstract
Manifold alignment has become very popular in recent literature. Aligning data distributions prior to product generation is an appealing strategy, since it allows to provide data spaces that are more similar to each other, regardless of the subsequent use of the transformed data. We propose a methodology that finds a common representation among data spaces from different sensors using geographic image correspondences, or semantic ties. To cope with the strong deformations between the data spaces considered, we propose to add nonlinearities by expanding the input space with Gaussian Radial Basis Function (RBF) features with respect to the centroids of a partitioning of the data. Such features allow us to cope with nonlinear transformations, while keeping a simple and efficient linear formulation. The proposed method is multi-domain and does not require co-registration, rather only a partial degree of spatial overlap. We test it on a challenging problem of multisensor classification transferring a model trained on a WorldView 2 image to predict land cover of a 3-bands orthophoto and show that we can transfer the model with an accuracy comparable to the one that would have been obtained by a model trained on the target image with an image-specific ground truth.
Diego Marcos, Gustau Camps-Valls, Devis Tuia
IGARSS3
2015 Large-scale random features for kernel regression
abstract
Kernel methods constitute a family of powerful machine learning algorithms, which have found wide use in remote sensing and geosciences. However, kernel methods are still not widely adopted because of the high computational cost when dealing with large scale problems, such as the inversion of radiative transfer models. This paper introduces the method of random kitchen sinks (RKS) for fast statistical retrieval of bio-geo-physical parameters. The RKS method allows to approximate a kernel matrix with a set of random bases sampled from the Fourier domain. We extend their use to other bases, such as wavelets, stumps, and Walsh expansions. We show that kernel regression is now possible for datasets with millions of examples and high dimensionality. Examples on atmospheric parameter retrieval from infrared sounders and biophysical parameter retrieval by inverting PROSAIL radiative transfer models with simulated Sentinel-2 data show the effectiveness of the technique.
Valero Laparra, Diego Marcos, Devis Tuia, Gustau Camps-Valls
IGARSS3
2015 Individual tree segmentation in deciduous forests using geodesic voting
abstract
Airborne Laser Scanning (ALS) has been widely used to survey forest areas. The extraction (segmentation) of individual trees from ALS point clouds is a prerequisite step for tree biophysical parameter estimation. For this purpose, we develop and evaluate a graph based segmentation algorithm adapted to deciduous forests scanned with high density LiDAR (~50 points / m2) in leaf-off conditions. The algorithm is applied to a 1 ha deciduous forest plot in western Switzerland and the accuracy of individual trunk locations is evaluated in terms of recall, precision and F-score. The results indicate that the algorithm performs satisfactorily within the experimental setup conditions.
Matthew Parkan, Devis Tuia
IGARSS2
2015 To be or not to be convex? A study on regularization in hyperspectral image classification
abstract
Hyperspectral image classification has long been dominated by convex models, which provide accurate decision functions exploiting all the features in the input space. However, the need for high geometrical details, which are often satisfied by using spatial filters, and the need for compact models (i.e. relying on models issued form reduced input spaces) has pushed research to study alternatives such as sparsity inducing regularization, which promotes models using only a subset of the input features. Although successful in reducing the number of active inputs, these models can be biased and sometimes offer sparsity at the cost of reduced accuracy. In this paper, we study the possibility of using non-convex regularization, which limits the bias induced by the regularization. We present and compare four regularizers, and then apply them to hyperspectral classification with different cost functions.
Devis Tuia, Rémi Flamary, Michel Barlaud
IGARSS1
2015 Semisupervised Classification of Remote Sensing Images With Hierarchical Spatial Similarity
abstract
A semisupervised kernel deformation function, including spatial similarity, is proposed for the classification of remote sensing (RS) images. The method exploits the characteristic of these images, in which spatially nearby points are likely to belong to the same class. To fulfill this assumption, a kernel encoding both spatial and spectral proximity using unlabeled samples is proposed. In this letter, two similarity functions for constructing a spatial kernel are proposed. Experimental tests are performed on very high-resolution multispectral and hyperspectral data. With respect to state-of-the-art semisupervised methods for RS images, the proposed method incorporating spatial similarity obtains higher classification accuracy values and smoother classification maps.
Zheng Zhang 0021, Devis Tuia
IEEE Geosci. Remote. Sens. Lett.4
2015 Multimodal Classification of Remote Sensing Images: A Review and Future Directions
abstract
Earth observation through remote sensing images allows the accurate characterization and identification of materials on the surface from space and airborne platforms. Multiple and heterogeneous image sources can be available for the same geographical region: multispectral, hyperspectral, radar, multitemporal, and multiangular images can today be acquired over a given scene. These sources can be combined/fused to improve classification of the materials on the surface. Even if this type of systems is generally accurate, the field is about to face new challenges: the upcoming constellations of satellite sensors will acquire large amounts of images of different spatial, spectral, angular, and temporal resolutions. In this scenario, multimodal image fusion stands out as the appropriate framework to address these problems. In this paper, we provide a taxonomical view of the field and review the current methodologies for multimodal classification of remote sensing images. We also highlight the most recent advances, which exploit synergies with machine learning and signal processing: sparse methods, kernel-based fusion, Markov modeling, and manifold alignment. Then, we illustrate the different approaches in seven challenging remote sensing applications: 1) multiresolution fusion for multispectral image classification; 2) image downscaling as a form of multitemporal image fusion and multidimensional interpolation among sensors of different spatial, spectral, and temporal resolutions; 3) multiangular image classification; 4) multisensor image fusion exploiting physically-based feature extractions; 5) multitemporal image classification of land covers in incomplete, inconsistent, and vague image sources; 6) spatiospectral multisensor fusion of optical and radar images for change detection; and 7) cross-sensor adaptation of classifiers. The adoption of these techniques in operational settings will help to monitor our planet from space in the very near future.
Luis Gómez-Chova, Devis Tuia, Gabriele Moser, Gustau Camps-Valls
Proc. IEEE2
2015 Cluster validity measure and merging system for hierarchical clustering considering outliers
Frank de Morsier, Devis Tuia, Maurice Borgeaud, Volker Gass, Jean-Philippe Thiran
Pattern Recognit.2
2015 Semisupervised Transfer Component Analysis for Domain Adaptation in Remote Sensing Image Classification
abstract
In this paper, we study the problem of feature extraction for knowledge transfer between multiple remotely sensed images in the context of land-cover classification. Several factors such as illumination, atmospheric, and ground conditions cause radiometric differences between images of similar scenes acquired on different geographical areas or over the same scene but at different time instants. Accordingly, a change in the probability distributions of the classes is observed. The purpose of this work is to statistically align in the feature space an image of interest that still has to be classified (the target image) to another image whose ground truth is already available (the source image). Following a specifically designed feature extraction step applied to both images, we show that classifiers trained on the source image can successfully predict the classes of the target image despite the shift that has occurred. In this context, we analyze a recently proposed domain adaptation method aiming at reducing the distance between domains, Transfer Component Analysis, and assess the potential of its unsupervised and semisupervised implementations. In particular, with a dedicated study of its key additional objectives, namely the alignment of the projection with the labels and the preservation of the local data structures, we demonstrate the advantages of Semisupervised Transfer Component Analysis. We compare this approach with other both linear and kernel-based feature extraction techniques. Experiments on multi- and hyperspectral acquisitions show remarkable cross- image classification performances for the considered strategy, thus confirming its suitability when applied to remotely sensed images.
Giona Matasci, Michele Volpi, Mikhail F. Kanevski, Lorenzo Bruzzone, Devis Tuia
IEEE Trans. Geosci. Remote. Sens.5
2014 Network-Based Correlated Correspondence for Unsupervised Domain Adaptation of Hyperspectral Satellite Images
abstract
Adapting a model to changes in the data distribution is a relevant problem in machine learning and pattern recognition since such changes degrade the performances of classifiers trained on undistorted samples. This paper tackles the problem of domain adaptation in the context of hyper spectral satellite image analysis. We propose a new correlated correspondence algorithm based on network analysis. The algorithm finds a matching between two distributions, which preserves the geometrical and topological information of the corresponding graphs. We evaluate the performance of the algorithm on a shadow compensation problem in hyper spectral image analysis: the land use classification obtained with the compensated data is improved.
Julien Rebetez, Devis Tuia, Nicolas Courty
ICPR2
2014 Unsupervised Alignment of Image Manifolds with Centrality Measures
abstract
The re-use of available labeled samples to classify newly acquired data is a hot topic in pattern analysis and machine learning. Classification algorithms developed with data from one domain cannot be directly used in another related domain, unless the data representation or the classifier have been adapted to the new data distribution. This is crucial in satellite/airborne image analysis: when confronted to domain shifts issued from changes in acquisition or illumination conditions, image classifiers tend to become inaccurate. In this paper, we introduce a method to align data manifolds that represent the same land cover classes, but have undergone spectral distortions. The proposed method relies on a semi-supervised manifold alignment technique and relaxes the requirement of labeled data in all domains by exploiting centrality measures over graphs to match the manifolds. Experiments on multispectral pixel classification at very high spatial resolution show the potential of the method.
Devis Tuia, Michele Volpi, Gustau Camps-Valls
ICPR1
2014 Spectral adaptation of hyperspectral flight lines using VHR contextual information
abstract
Due to technological constraints, hyperspectral earth observation imagery are often a mosaic of overlapping flight lines collected in different passes over the area of interest. This causes variations in aqcuisition conditions such that the reflected spectrum can vary significantly between these flight lines. Partly, this problem is solved by atmospherical correction, but residual spectral differences often remain. A probabilistic domain adaptation framework based on graph matching using Hidden Markov Random Fields was recently proposed for transforming hyperspectral data from one image to better correspond to the other. This paper investigates the use of scale and angle invariant textural features for improving the performance of the used Hidden Markov Random Field matching framework in the case of hyperspectral flight lines. These textural features are derived from the filtering of VHR optical imagery with a bank of Gabor filters with varying orientation, scale and frequency and subsequently rendering them invariant to scale and frequency by applying the 2D DFT on the filter responses in the scale and frequency space.
Jan-Pieter Jacobs, Guy Thoonen, Devis Tuia, Gustau Camps-Valls, Pieter Kempeneers, Paul Scheunders
IGARSS3
2014 Domain adaptation in remote sensing through cross-image synthesis with dictionaries
abstract
This contribution studies an approach based on dictionary learning which enables the alignment of the sparse representations of two images. Set in a domain adaptation context, the purpose of this work is to re-synthesize the pixels of a remote sensing image so that, for a given land-cover class, the new values of the samples are comparable across acquisitions. Consequently, the data space of a given source image can be converted to that of a related target image, or vice-versa. After the mentioned transformation, the performance of a classifier trained on the source image and used to predict the thematic classes on the target image is expected to be more robust. A linear transformation is derived thanks to an algorithm simultaneously learning the image-specific dictionaries and the mapping function bridging them via their respective sparse codes. Experiments on knowledge transfer among two co-registered VHR images acquired with different off-nadir angles show promising results. An appropriate cross-image synthesis yields an increased land-cover model portability from one acquisition to another.
Giona Matasci, Frank de Morsier, Mikhail F. Kanevski, Devis Tuia
IGARSS4
2014 Non-linear low-rank and sparse representation for hyperspectral image analysis
abstract
In this paper, we tackle the problem of unsupervised classification of hyperspectral images. We propose a clustering method based on graphs representing the data structure, which is assumed to be an union of multiple manifolds. The method constraints the pixels to be expressed as a low-rank and sparse combination of the others in a reproducing kernel Hilbert spaces (RKHS). This captures the global (low-rank) and local (sparse) structures. Spectral clustering is applied on the graph to assign the pixels to the different manifolds. A large scale approach is proposed, in which the optimization is first performed on a subset of the data and then it is applied to the whole image using a non-linear collaborative representation respecting the manifolds structure. Experiments on two hyperspectral images show very good unsupervised classification results compared to competitive approaches.
Frank de Morsier, Devis Tuia, Maurice Borgeaucft, Volker Gass, Jean-Philippe Thiran
IGARSS2
2014 Crop backscatter modeling and soil moisture estimation with support vector regression
abstract
In this paper, we used an improved version of the Tor Vergata radiative transfer model to simulate the backscattering coefficient for the L-band SAR signals over areas covered with vegetation. Fields of winter wheat, maize and sugar beet observed during the AgriSAR2006 campaign were investigated. For maize field, the presence of periodic soil surface profiles played an important role in determining the total backscattering. Soil moisture was also estimated using an inverse algorithm based on a supervised, non-parametric learning technique, v-SVR. v-SVR proved good generalization properties even with a limited number of training samples available. Dependence to the origin of training samples, as well as the influence of different features, was thoroughly considered.
Jelena Stamenkovic, Paolo Ferrazzoli, Leila Guerriero, Devis Tuia, Jean-Philippe Thiran, Maurice Borgeaud
IGARSS4
2014 Weakly supervised alignment of image manifolds with semantic ties
abstract
Aligning data distributions that underwent spectral distortions related to acquisition conditions is a key issue to improve the performance of classifiers applied to multi-temporal and multi-angular images. In this paper, we propose a feature extraction methodology, which aligns data manifolds based on their internal geometric structure and on a series of object correspondences highlighted on each image, or tie points. The weakly supervised manifold alignment (WeSMA) is a feature extractor that allows to define a common latent space, in which the images can be projected and processed by the same classifier. WeSMA relaxes the need for labeled pixels in all acquisitions of previous manifold alignment methods, an heavy constraint for remote sensing applications. Experiments on a set of World-View II images acquired at different viewing angles show the interest of the method that can compensate the spectral shift generated by the angular distortion without labels issued from the off-nadir image.
Devis Tuia
IGARSS1
2014 Domain Adaptation with Regularized Optimal Transport
Nicolas Courty, Rémi Flamary, Devis Tuia
ECML/PKDD (1)3
2014 Principal Polynomial Analysis
abstract
This paper presents a new framework for manifold learning based on a sequence of principal polynomials that capture the possibly nonlinear nature of the data. The proposed Principal Polynomial Analysis (PPA) generalizes PCA by modeling the directions of maximal variance by means of curves, instead of straight lines. Contrarily to previous approaches, PPA reduces to performing simple univariate regressions, which makes it computationally feasible and robust. Moreover, PPA shows a number of interesting analytical properties. First, PPA is a volume-preserving map, which in turn guarantees the existence of the inverse. Second, such an inverse can be obtained in closed form. Invertibility is an important advantage over other learning methods, because it permits to understand the identified features in the input domain where the data has physical meaning. Moreover, it allows to evaluate the performance of dimensionality reduction in sensible (input-domain) units. Volume preservation also allows an easy computation of information theoretic quantities, such as the reduction in multi-information after the transform. Third, the analytical nature of PPA leads to a clear geometrical interpretation of the manifold: it allows the computation of Frenet-Serret frames (local features) and of generalized curvatures at any point of the space. And fourth, the analytical Jacobian allows the computation of the metric induced by the data, thus generalizing the Mahalanobis distance. These properties are demonstrated theoretically and illustrated experimentally. The performance of PPA is evaluated in dimensionality and redundancy reduction, in both synthetic and real datasets from the UCI repository.
Valero Laparra, Sandra Jiménez, Devis Tuia, Gustau Camps-Valls, Jesús Malo
Int. J. Neural Syst.3
2014 Semi-supervised multiview embedding for hyperspectral data classification
Michele Volpi, Giona Matasci, Mikhail F. Kanevski, Devis Tuia
Neurocomputing4
2014 SVM Active Learning Approach for Image Classification Using Spatial Information
abstract
In the last few years, active learning has been gaining growing interest in the remote sensing community in optimizing the process of training sample collection for supervised image classification. Current strategies formulate the active learning problem in the spectral domain only. However, remote sensing images are intrinsically defined both in the spectral and spatial domains. In this paper, we explore this fact by proposing a new active learning approach for support vector machine classification. In particular, we suggest combining spectral and spatial information directly in the iterative process of sample selection. For this purpose, three criteria are proposed to favor the selection of samples distant from the samples already composing the current training set. In the first strategy, the Euclidean distances in the spatial domain from the training samples are explicitly computed, whereas the second one is based on the Parzen window method in the spatial domain. Finally, the last criterion involves the concept of spatial entropy. Experiments on two very high resolution images show the effectiveness of regularization in spatial domain for active learning purposes.
Edoardo Pasolli, Farid Melgani, Devis Tuia, Fabio Pacifici, William J. Emery
IEEE Trans. Geosci. Remote. Sens.3
2014 Automatic Feature Learning for Spatio-Spectral Image Classification With Sparse SVM
abstract
Including spatial information is a key step for successful remote sensing image classification. In particular, when dealing with high spatial resolution, if local variability is strongly reduced by spatial filtering, the classification performance results are boosted. In this paper, we consider the triple objective of designing a spatial/spectral classifier, which is compact (uses as few features as possible), discriminative (enhances class separation), and robust (works well in small sample situations). We achieve this triple objective by discovering the relevant features in the (possibly infinite) space of spatial filters by optimizing a margin-maximization criterion. Instead of imposing a filter bank with predefined filter types and parameters, we let the model figure out which set of filters is optimal for class separation. To do so, we randomly generate spatial filter banks and use an active-set criterion to rank the candidate features according to their benefits to margin maximization (and, thus, to generalization) if added to the model. Experiments on multispectral very high spatial resolution (VHR) and hyperspectral VHR data show that the proposed algorithm, which is sparse and linear, finds discriminative features and achieves at least the same performances as models using a large filter bank defined in advance by prior knowledge.
Devis Tuia, Michele Volpi, Mauro Dalla Mura, Alain Rakotomamonjy, Rémi Flamary
IEEE Trans. Geosci. Remote. Sens.1
2014 Semisupervised Manifold Alignment of Multimodal Remote Sensing Images
abstract
We introduce a method for manifold alignment of different modalities (or domains) of remote sensing images. The problem is recurrent when a set of multitemporal, multisource, multisensor, and multiangular images is available. In these situations, images should ideally be spatially coregistered, corrected, and compensated for differences in the image domains. Such procedures require massive interaction of the user, involve tuning of many parameters and heuristics, and are usually applied separately. Changes of sensors and acquisition conditions translate into shifts, twists, warps, and foldings of the (typically nonlinear) manifolds where images lie. The proposed semisupervised manifold alignment (SS-MA) method aligns the images working directly on their manifolds and is thus not restricted to images of the same resolutions, either spectral or spatial. SS-MA pulls close together samples of the same class while pushing those of different classes apart. At the same time, it preserves the geometry of each manifold along the transformation. The method builds a linear invertible transformation to a latent space where all images are alike and reduces to solving a generalized eigenproblem of moderate size. We study the performance of SS-MA in toy examples and in real multiangular, multitemporal, and multisource image classification problems. The method performs well for strong deformations and leads to accurate classification for all domains. A MATLAB implementation of the proposed method is provided at http://isp. uv.es/code/ssma.htm.
Devis Tuia, Michele Volpi, Maxime Trolliet, Gustau Camps-Valls
IEEE Trans. Geosci. Remote. Sens.1
2014 Explicit Recursive and Adaptive Filtering in Reproducing Kernel Hilbert Spaces
abstract
This brief presents a methodology to develop recursive filters in reproducing kernel Hilbert spaces. Unlike previous approaches that exploit the kernel trick on filtered and then mapped samples, we explicitly define the model recursivity in the Hilbert space. For that, we exploit some properties of functional analysis and recursive computation of dot products without the need of preimaging or a training dataset. We illustrate the feasibility of the methodology in the particular case of the γ-filter, which is an infinite impulse response filter with controlled stability and memory depth. Different algorithmic formulations emerge from the signal model. Experiments in chaotic and electroencephalographic time series prediction, complex nonlinear system identification, and adaptive antenna array processing demonstrate the potential of the approach for scenarios where recursivity and nonlinearity have to be readily combined.
Devis Tuia, Jordi Muñoz-Marí, José Luis Rojo-Álvarez, Manel Martínez-Ramón, Gustau Camps-Valls
IEEE Trans. Neural Networks Learn. Syst.1
2013 Multi-view feature extraction for hyperspectral image classification
Michele Volpi, Giona Matasci, Mikhail F. Kanevski, Devis Tuia
ESANN4
2013 Investigating Feature Extraction for Domain Adaptation in Remote Sensing Image Classification
Giona Matasci, Lorenzo Bruzzone, Michele Volpi, Devis Tuia, Mikhail F. Kanevski
ICPRAM4
2013 Domain adaptation with Hidden Markov Random Fields
abstract
In this paper, we propose a method to match multitemporal sequences of hyperspectral images using Hidden Markov Random Fields. Based on the matching of the data manifold, the algorithm matches the reflectance spectra of the classes, thus allowing the reuse of labeled examples acquired on one image to classify the other. This allows valorization of spectra collected in situ to other acquisitions than the one they were acquired for, without user supervision, prior knowledge of the class reflectance in the new domain or global information about atmospheric conditions.
Jan-Pieter Jacobs, Guy Thoonen, Devis Tuia, Gustau Camps-Valls, Birgen Haest, Paul Scheunders
IGARSS3
2013 Statistical assessment of dataset shift and model portability in multi-angle in-track image acquisitions
abstract
In this study we propose an evaluation of the angular effects altering the spectral response of the land-cover over multi-angle remote sensing image acquisitions. The shift in the statistical distribution of the pixels observed in an in-track sequence of WorldView-2 images is analyzed by means of a kernel-based measure of distance between probability distributions. Afterwards, the portability of supervised classifiers across the sequence is investigated by looking at the evolution of the classification accuracy with respect to the changing observation angle. In this context, the efficiency of various physically and statistically based preprocessing methods in obtaining angle-invariant data spaces is compared and possible synergies are discussed.
Giona Matasci, Nathan Longbotham, Fabio Pacifici, Mikhail F. Kanevski, Devis Tuia
IGARSS5
2013 Multisensor alignment of image manifolds
abstract
The access to many sources of satellite information is nowadays a reality. However, few methods allow to consider simultaneously data coming from different sensors, due to the differences in numbers of bands, spatial resolution and changes in the acquisition conditions. In this paper, we propose a methodology to align the data structures (also called manifolds) of two (or more) images and to exploit them simultaneously in a joint latent space. The method being invertible, it also have the interesting property to allow to project the image pixels from one sensor to another, thus allowing to synthesize the bands of one sensor using the pixels of the other through the projection learned. Experiments using QuickBird and World-View II images show the properties of the method and open new opportunities for multisensor remote-sensing.
Devis Tuia, Maxime Trolliet, Michele Volpi
IGARSS1
2013 Create the relevant spatial filterbank in the hyperspectral jungle
abstract
Inclusion of spatial information is known to be beneficial to the classification of hyperspectral images. However, given the high dimensionality of the data, it is difficult to know before hand which are the bands to filter or what are the filters to be applied. In this paper, we propose an active set algorithm based on a l1 support vector machine that explores the (possibily infinite) space of spatial filters and retrieves automatically the filters that maximize class separation. Experiments on hyperspectral imagery confirms the power of the method, that reaches state of the art performance with small feature sets generated automatically and without prior knowledge.
Devis Tuia, Michele Volpi, Mauro Dalla Mura, Alain Rakotomamonjy, Rémi Flamary
IGARSS1
2013 Multi-sensor change detection based on nonlinear canonical correlations
abstract
The analysis of multi-modal and multi-sensor images is nowadays of paramount importance for Earth Observation (EO) applications. There exist a variety of methods that aim at fusing the different sources of information to obtain a compact representation of such datasets. However, for change detection existing methods are often unable to deal with heterogeneous image sources and very few consider possible nonlinearities in the data. Additionally, the availability of labeled information is very limited in change detection applications. For these reasons, we present the use of a semi-supervised kernel-based feature extraction technique. It incorporates a manifold regularization accounting for the geometric distribution and jointly addressing the small sample problem. An exhaustive example using Landsat 5 data illustrates the potential of the method for multi-sensor change detection.
Michele Volpi, Frank de Morsier, Gustau Camps-Valls, Mikhail F. Kanevski, Devis Tuia
IGARSS5
2013 Active Learning: Any Value for Classification of Remotely Sensed Data?
abstract
Active learning, which has a strong impact on processing data prior to the classification phase, is an active research area within the machine learning community, and is now being extended for remote sensing applications. To be effective, classification must rely on the most informative pixels, while the training set should be as compact as possible. Active learning heuristics provide capability to select unlabeled data that are the “most informative” and to obtain the respective labels, contributing to both goals. Characteristics of remotely sensed image data provide both challenges and opportunities to exploit the potential advantages of active learning. We present an overview of active learning methods, then review the latest techniques proposed to cope with the problem of interactive sampling of training pixels for classification of remotely sensed data with support vector machines (SVMs). We discuss remote sensing specific approaches dealing with multisource and spatially and time-varying data, and provide examples for high-dimensional hyperspectral imagery.
Melba M. Crawford, Devis Tuia, Hsiuhan Lexie Yang
Proc. IEEE2
2013 Semi-Supervised Novelty Detection Using SVM Entire Solution Path
abstract
Very often, the only reliable information available to perform change detection is the description of some “unchanged” regions. Since, sometimes, these regions do not contain all the relevant information to identify their counterpart (the changes), we consider the use of unlabeled data to perform semi-supervised novelty detection (SSND). SSND can be seen as an unbalanced classification problem solved using the cost-sensitive support vector machine (CS-SVM), but this requires a heavy parameter search. Here, we propose the use of entire solution path algorithms for the CS-SVM in order to facilitate and accelerate parameter selection for SSND. Two algorithms are considered and evaluated. The first algorithm is an extension of the CS-SVM algorithm that returns the entire solution path in a single optimization. This way, optimization of a separate model for each hyperparameter set is avoided. The second algorithm forces the solution to be coherent through the solution path, thus producing classification boundaries that are nested (included in each other). We also present a low-density (LD) criterion for selecting optimal classification boundaries, thus avoiding recourse to cross validation (CV) that usually requires information about the “change” class. Experiments are performed on two multitemporal change detection data sets (flood and fire detection). Both algorithms tracing the solution path provide similar performances than the standard CS-SVM while being significantly faster. The proposed LD criterion achieves results that are close to the ones obtained by CV but without using information about the changes.
Frank de Morsier, Devis Tuia, Maurice Borgeaud, Volker Gass, Jean-Philippe Thiran
IEEE Trans. Geosci. Remote. Sens.2
2013 Learning User's Confidence for Active Learning
abstract
In this paper, we study the applicability of active learning (AL) in operative scenarios. More particularly, we consider the well-known contradiction between the AL heuristics, which rank the pixels according to their uncertainty, and the user's confidence in labeling, which is related to both the homogeneity of the pixel context and user's knowledge of the scene. We propose a filtering scheme based on a classifier that learns the confidence of the user in labeling, thus minimizing the queries where the user would not be able to provide a class for the pixel. The capacity of a model to learn the user's confidence is studied in detail, also showing that the effect of resolution in such a learning task. Experiments on two QuickBird images of different resolutions (with and without pansharpening) and considering committees of users prove the efficiency of the filtering scheme proposed, which maximizes the number of useful queries with respect to traditional AL.
Devis Tuia, Jordi Muñoz-Marí
IEEE Trans. Geosci. Remote. Sens.1
2013 Graph Matching for Adaptation in Remote Sensing
abstract
We present an adaptation algorithm focused on the description of the data changes under different acquisition conditions. When considering a source and a destination domain, the adaptation is carried out by transforming one data set to the other using an appropriate nonlinear deformation. The eventually nonlinear transform is based on vector quantization and graph matching. The transfer learning mapping is defined in an unsupervised manner. Once this mapping has been defined, the samples in one domain are projected onto the other, thus allowing the application of any classifier or regressor in the transformed domain. Experiments on challenging remote sensing scenarios, such as multitemporal very high resolution image classification and angular effects compensation, show the validity of the proposed method to match-related domains and enhance the application of cross-domains image processing techniques.
Devis Tuia, Jordi Muñoz-Marí, Luis Gómez-Chova, Jesús Malo
IEEE Trans. Geosci. Remote. Sens.1
2012 Discovering relevant spatial filterbanks for VHR image classification
Devis Tuia, Mauro Dalla Mura, Michele Volpi, Rémi Flamary, Alain Rakotomamonjy
ICPR1
2012 Discovering single classes in remote sensing images with active learning
abstract
When dealing with supervised target detection, the acquisition of labeled samples is one of the most critical phases: the samples must be yet representative of the class of interest, but must also be found among a vast majority of non-target examples. Moreover, the efficiency of the search is also an issue, since the samples labeled as background are not used by target detectors such as the support vector data description (SVDD). In this work we propose a competitive and effective approach to identify the most relevant training samples for one-class classification based on the use of an active learning strategy. The SVDD classifier is first trained with insufficient target examples. It is then used to detect the most informative samples to be labeled by a user through active learning techniques. By selecting unlabeled samples in a smart way and by adopting a diversity criterion, it is possible to obtain an accurate description of the class of interest with a relatively small number of training samples. The performance of the proposed method is illustrated in a change detection scenario and is validated by comparison with state-of-art active learning techniques originally developed for multiclass problems.
Mirco Furlani, Devis Tuia, Jordi Muñoz-Marí, Francesca Bovolo, Gustau Camps-Valls, Lorenzo Bruzzone
IGARSS2
2012 Putting the user into the active learning loop: Towards realistic but efficient photointerpretation
abstract
In recent years, several studies have been published about the smart definition of training set using active learning algorithms. However, none of these works consider the contradiction between the active learning methods, which rank the pixels according to their uncertainty, and the confidence of the user in labeling, which is related both to the homogeneity of the pixel context and to the knowledge of the user of the scene. In this paper, we propose a two-steps procedure based on a filtering scheme to learn the confidence of the user in labeling. This way, candidate training pixels are ranked according both to their uncertainty and to the chances of being labeled correctly by the user. In this way, we avoid the queries where the user would not be able to provide a class for the pixel. We consider the capacity of a model in learning the user's confidence and report experiments on a QuickBird image: the filtering scheme proposed maximizes the number of useful queries with respect to traditional active learning.
Devis Tuia, Jordi Muñoz-Marí
IGARSS1
2012 Enhanced change detection using nonlinear feature extraction
abstract
This paper presents an application of the kernel principal component analysis aiming at spectrally aligning optical images before the application of change detection techniques. The approach relies on the extraction of nonlinear features from a selected subset of pixels representing unchanged areas in the bi-temporal images. Both images are then projected into the new space defined by the eigenvectors associated to largest variance (eigenvalues). In the transformed space, unchanged pixels are mapped next to each other, thus reducing within-class variance. The difference image that results from subtracting the projected datasets is likely to provide a more suitable representation for detecting changes. A subset of two Landsat TM scenes validates the proposed approach. The new representation is studied thanks to the change vector analysis and to the support vector domain description.
Michele Volpi, Giona Matasci, Devis Tuia, Mikhail F. Kanevski
IGARSS3
2012 Unsupervised Change Detection With Kernels
abstract
In this letter, an unsupervised kernel-based approach to change detection is introduced. Nonlinear clustering is utilized to partition in two a selected subset of pixels representing both changed and unchanged areas. Once the optimal clustering is obtained, the learned representatives of each group are exploited to assign all the pixels composing the multitemporal scenes to the two classes of interest. Two approaches based on different assumptions of the difference image are proposed. The first accounts for the difference image in the original space, while the second defines a mapping describing the difference image directly in feature spaces. To optimize the parameters of the kernels, a novel unsupervised cost function is proposed. An evidence of the correctness, stability, and superiority of the proposed solution is provided through the analysis of two challenging change-detection problems.
Michele Volpi, Devis Tuia, Gustau Camps-Valls, Mikhail F. Kanevski
IEEE Geosci. Remote. Sens. Lett.2
2012 Remote sensing image segmentation by active queries
Devis Tuia, Jordi Muñoz-Marí, Gustau Camps-Valls
Pattern Recognit.1
2012 Semisupervised Classification of Remote Sensing Images With Active Queries
abstract
We propose a semiautomatic procedure to generate land cover maps from remote sensing images. The proposed algorithm starts by building a hierarchical clustering tree, and exploits the most coherent pixels with respect to the available class information. For a given amount of labeled pixels, the algorithm returns both classification and confidence maps. Since the quality of the map depends of the number and informativeness of the labeled pixels, active learning methods are used to select the most informative samples to increase confidence in class membership. Experiments on four different data sets, accounting for hyperspectral and multispectral images at different spatial resolutions, confirm the effectiveness of the proposed approach, and how active learning techniques reduce the uncertainty of the classification maps. Specifically, more accurate results with fewer labeled samples are obtained. Inclusion of spatial information in the classifiers drastically improves the classification accuracy, leading to faster convergence curves and tighter confidence intervals. In conclusion, the presented algorithm provides efficient image classification and, at the same time, yields a confidence map that may be very useful in many Earth observation applications.
Jordi Muñoz-Marí, Devis Tuia, Gustau Camps-Valls
IEEE Trans. Geosci. Remote. Sens.2
2012 Memory-Based Cluster Sampling for Remote Sensing Image Classification
abstract
In this paper, we address the problem of semi-automatic definition of training sets for the classification of remotely sensed images. We propose two approaches based on active learning aiming at removing both the proximal (low diversity) and the dense (low exploration during iterations) sampling redundancies. The first is encountered when several samples carrying similar spectral information are selected by the algorithm, while the second occurs when the heuristic is unable to explore undiscovered parts of the feature space during iterations. For this purpose, kernel$k$-means is used to cluster a set of uncertain candidates in the same space spanned by the kernel function defined in the SVM classification step. Two heuristics are proposed to maximize the speed of convergence to high classification accuracies: The first is based on binary hierarchical partitioning of the set of selected uncertain samples, while the second extends this approach by considering memory in the selection and thus dynamically adapts to the problem throughout the iterations. Experiments on both VHR and hyperspectral imagery confirm fast convergence of the algorithm, that outperforms state-of-the-art sampling schemes.
Michele Volpi, Devis Tuia, Mikhail F. Kanevski
IEEE Trans. Geosci. Remote. Sens.2
2011 Explicit recursivity into reproducing kernel Hilbert spaces
abstract
This paper presents a methodology to develop recursive filters in reproducing kernel Hilbert spaces (RKHS). Unlike previous approaches that exploit the kernel trick on filtered and then mapped samples, we explicitly define model recursivity in the Hilbert space. The method exploits some properties of functional analysis and recursive computation of dot products without the need of pre-imaging. We illustrate the feasibility of the methodology in the particular case of the gamma filter, an infinite impulse response (IIR) filter with controlled stability and memory depth. Different algorithmic formulations emerge from the signal model. Experiments in chaotic and electroencephalographic time series prediction scenarios demonstrate the potentiality of the approach.
Devis Tuia, Gustau Camps-Valls, Manel Martínez-Ramón
ICASSP1
2011 Principal polynomial analysis for remote sensing data processing
abstract
Inspired by the concept of Principal Curves, in this paper, we define Principal Polynomials as a non-linear generalization of Principal Components to overcome the conditional mean independence restriction of PCA. Principal Polynomials deform the straight Principal Components by minimizing the regression error (or variance) in the corresponding orthogonal subspaces. We propose to use a projection on a series of these polynomials to set a new nonlinear data representation: the Principal Polynomial Analysis (PPA). We prove that the dimensionality reduction error in PPA is always lower than in PCA. Lower truncation error and increased independence suggest that unsupervised PPA features can be better suited to image classification than those identified by other unsupervised techniques. We analyze the performance of Linear Discriminant Analysis in the feature space after dimensionality reduction using the proposed PPA, the classical PCA, and locally linear embedding (LLE). Experiments on very high resolution data confirm the suitability of PPA to describe nonlinear manifolds found in remote sensing data.
Valero Laparra, Devis Tuia, Sandra Jiménez, Gustau Camps-Valls, Jesús Malo
IGARSS2
2011 Domain separation for efficient adaptive active learning
abstract
This paper proposes a procedure aimed at efficiently adapting a classifier trained on a source image to a similar target image. The adaptation is carried out through active queries in the target domain following a strategy particularly designed for the case where class distributions have shifted between the two images. We first suggest a pre-selection of candidate pixels issued from the target image by keeping only those samples appearing to be lying in a region of the input space not yet covered by the existing ground truth (source domain pixels). Then, exploiting a classifier integrating instance weights, active queries are performed on the target image. As the inclusion to the training set of the samples progresses, the weights associated with the training pixels are updated using different criteria according to their origin (source or target domain). Experiments on a pair of QuickBird images of urban scenes prove the validity of the proposed approach if compared to existing benchmark methods.
Giona Matasci, Devis Tuia, Mikhail F. Kanevski
IGARSS2
2011 Improving active learning methods using spatial information
abstract
Active learning process represents an interesting solution to the problem of training sample collection for the classification of remote sensing images. In this work, we propose a criterion based on the spatial information that can be used in combination with a spectral criterion in order to improve the selection of training samples. Experimental results obtained on a very high resolution image show the effectiveness of regularization in spatial domain and open challenging perspectives for terrain campaigns planning.
Edoardo Pasolli, Farid Melgani, Devis Tuia, Fabio Pacifici, William J. Emery
IGARSS3
2011 Large scale semi-supervised image segmentation with active queries
abstract
A semiautomatic procedure to generate classification maps of remote sensing images is proposed. Starting from a hierarchical unsupervised classification, the algorithm exploits the few available labeled pixels to assign each cluster to the most probable class. For a given amount of labeled pixels, the algorithm returns a classified segmentation map, along with confidence levels of class membership for each pixel. Active learning methods are used to select the most informative samples to increase confidence in the class membership. Experiments on a AVIRIS hyperspectral image confirm the effectiveness of the method, especially when used with active learning query functions and spatial regularization.
Devis Tuia, Jordi Muñoz-Marí, Gustau Camps-Valls
IGARSS1
2011 Graph matching for efficient classifiers adaptation
abstract
In this work we present an adaptation algorithm focused on the description of the measurement changes under different acquisition conditions. The adaptation is carried out by transforming the manifold in the first observation conditions into the corresponding manifold in the second. The eventually non-linear transform is based on vector quantization and graph matching. The transfer learning mapping is defined in an unsupervised manner. Once this mapping has been defined, the labeled samples in the first are projected into the second domain, thus allowing the application of any classifier in the transformed domain. Experiments on VHR series of images show the validity of the proposed method to adapt the classifiers to related domains.
Devis Tuia, Jordi Muñoz-Marí, Jesús Malo
IGARSS1
2011 Unsupervised change detection in the feature space using kernels
abstract
In this paper we propose an unsupervised approach to change detection by computing the difference image directly in the feature spaces. The resulting difference kernel, that is a combination of kernels computed on the coregistered and radiometrically matched input images, is used to train a nonlinear partitioning algorithm. In order to apply the kernel k-means, issues related to the initialization and to the tuning of parameters (e.g. the Gaussian RBF bandwidth) are considered. To validate the proposed unsupervised algorithm, two multitemporal VHR remote sensing images are used.
Michele Volpi, Devis Tuia, Gustau Camps-Valls, Mikhail F. Kanevski
IGARSS2
2011 Multioutput Support Vector Regression for Remote Sensing Biophysical Parameter Estimation
abstract
This letter proposes a multioutput support vector regression (M-SVR) method for the simultaneous estimation of different biophysical parameters from remote sensing images. General retrieval problems require multioutput (and potentially nonlinear) regression methods. M-SVR extends the single-output SVR to multiple outputs maintaining the advantages of a sparse and compact solution by using an$\varepsilon$-insensitive cost function. The proposed M-SVR is evaluated in the estimation of chlorophyll content, leaf area index and fractional vegetation cover from a hyperspectral compact high-resolution imaging spectrometer images. The achieved improvement with respect to the single-output regression approach suggests that M-SVR can be considered a convenient alternative for nonparametric biophysical parameter estimation and model inversion.
Devis Tuia, Jochem Verrelst, Luis Alonso 0002, Fernando Pérez-Cruz, Gustau Camps-Valls
IEEE Geosci. Remote. Sens. Lett.1
2010 Time series input selection using multiple kernel learning
Loris Foresti, Devis Tuia, Vadim Timonin, Mikhail F. Kanevski
ESANN2
2010 Estimating biophysical variable dependences with kernels
abstract
This paper introduces a nonlinear measure of dependence between random variables in the context of remote sensing data analysis. The Hilbert-Schmidt Independence Criterion (HSIC) is a kernel method for evaluating statistical dependence. HSIC is based on computing the Hilbert-Schmidt norm of the cross-covariance operator of mapped samples in the corresponding Hilbert spaces. The HSIC empirical estimator is very easy to compute and has good theoretical and practical properties. We exploit the capabilities of HSIC to explain nonlinear dependences in two remote sensing problems: temperature estimation and chlorophyll concentration prediction from spectra. Results show that, when the relationship between random variables is nonlinear or when few data are available, the HSIC criterion outperforms other standard methods, such as the linear correlation or mutual information.
Gustau Camps-Valls, Devis Tuia, Valero Laparra, Jesús Malo
IGARSS2
2010 Cluster-based active learning for compact image classification
abstract
In this paper, we consider active sampling to label pixels grouped with hierarchical clustering. The objective of the method is to match the data relationships discovered by the clustering algorithm with the user's desired class semantics. The first is represented as a complete tree to be pruned and the second is iteratively provided by the user. The active learning algorithm proposed searches the pruning of the tree that best matches the labels of the sampled points. By choosing the part of the tree to sample from according to current pruning's uncertainty, sampling is focused on most uncertain clusters. This way, large clusters for which the class membership is already fixed are no longer queried and sampling is focused on division of clusters showing mixed labels. The model is tested on a VHR image in a multiclass classification setting. The method clearly outperforms random sampling in a transductive setting, but cannot generalize to unseen data, since it aims at optimizing the classification of a given cluster structure.
Devis Tuia, Mikhail F. Kanevski, Jordi Muñoz-Marí, Gustau Camps-Valls
IGARSS1
2010 Advanced active sampling for remote sensing image classification
abstract
A novel approach to active sampling is proposed for the semi automatic selection of training patterns in a given pool of candidates. In the method proposed, each candidate is ranked with a double criterion: first, informativeness of the candidate is assessed using the Support Vector Machine (SVM) real valued decision function. Then, diverse sampling is ensured by considering the relative position of the candidate in the SVM feature space. Such position is evaluated by using partitioning of the feature space via nonlinear clustering. Among all candidates belonging to the same cluster, the winner is the pixel minimizing the weighted combination of the distances from both the SVM hyperplane and the cluster center. This way, the proposed approach provides a way to i) account for and minimize the redundancy of the sampled pixels and ii) maximize the speed of convergence to an optimal classification accuracy. In order to discover clusters and to evaluate the distance between the cluster center and the samples, the kernel k-means algorithm is used in a hierarchical way. By its kernel nature, the algorithm partitions data in the SVM induced space, thus ensuring coherent diversity with respect to the linear model therein. The reliability of the proposed heuristic is evaluated on a QuickBird VHR image of the city of Zurich, Switzerland.
Michele Volpi, Devis Tuia, Mikhail F. Kanevski
IGARSS2
2010 Multisource Composite Kernels for Urban-Image Classification
abstract
This letter presents advanced classification methods for very high resolution images. Efficient multisource information, both spectral and spatial, is exploited through the use of composite kernels in support vector machines. Weighted summations of kernels accounting for separate sources of spectral and spatial information are analyzed and compared to classical approaches such as pure spectral classification or stacked approaches using all the features in a single vector. Model selection problems are addressed, as well as the importance of the different kernels in the weighted summation.
Devis Tuia, Frédéric Ratle, Alexei Pozdnoukhov, Gustau Camps-Valls
IEEE Geosci. Remote. Sens. Lett.1
2010 Learning Relevant Image Features With Multiple-Kernel Classification
abstract
The increase in spatial and spectral resolution of the satellite sensors, along with the shortening of the time-revisiting periods, has provided high-quality data for remote sensing image classification. However, the high-dimensional feature space induced by using many heterogeneous information sources precludes the use of simple classifiers: thus, a proper feature selection is required for discarding irrelevant features and adapting the model to the specific problem. This paper proposes to classify the images and simultaneously to learn the relevant features in such high-dimensional scenarios. The proposed method is based on the automatic optimization of a linear combination of kernels dedicated to different meaningful sets of features. Such sets can be groups of bands, contextual or textural features, or bands acquired by different sensors. The combination of kernels is optimized through gradient descent on the support vector machine objective function. Even though the combination is linear, the ranked relevance takes into account the intrinsic nonlinearity of the data through kernels. Since a naive selection of the free parameters of the multiple-kernel method is computationally demanding, we propose an efficient model selection procedure based on the kernel alignment. The result is a weight (learned from the data) for each kernel where both relevant and meaningless image features automatically emerge after training the model. Experiments carried out in multi- and hyperspectral, contextual, and multisource remote sensing data classification confirm the capability of the method in ranking the relevant features and show the computational efficience of the proposed strategy.
Devis Tuia, Gustau Camps-Valls, Giona Matasci, Mikhail F. Kanevski
IEEE Trans. Geosci. Remote. Sens.1
2010 Correction to "Active Learning Methods for Remote Sensing Image Classification" [Jul 09 2218-2232]
abstract
In the above titled paper (ibid., vol. 47, no. 7, pp. 2218-2232, Jul. 09), three lines from Algorithm 2 were inadvertently omitted during the paper's typesetting. The corrected algorithm is presented here.
Devis Tuia, Frédéric Ratle, Fabio Pacifici, Mikhail F. Kanevski, William J. Emery
IEEE Trans. Geosci. Remote. Sens.1
2009 Multiple Kernel Learning of Environmental Data. Case Study: Analysis and Mapping of Wind Fields
Loris Foresti, Devis Tuia, Alexei Pozdnoukhov, Mikhail F. Kanevski
ICANN (2)2
2009 Recent advances in remote sensing image processing
abstract
Remote sensing image processing is nowadays a mature research area. The techniques developed in the field allow many real-life applications with great societal value. For instance, urban monitoring, fire detection or flood prediction can have a great impact on economical and environmental issues. To attain such objectives, the remote sensing community has turned into a multidisciplinary field of science that embraces physics, signal theory, computer science, electronics, and communications. From a machine learning and signal/image processing point of view, all the applications are tackled under specific formalisms, such as classification and clustering, regression and function approximation, image coding, restoration and enhancement, source unmixing, data fusion or feature selection and extraction. This paper serves as a survey of methods and applications, and reviews the last methodological advances in remote sensing image processing.
Devis Tuia, Gustau Camps-Valls
ICIP1
2009 Learning the Relevant Image Features with Multiple Kernels
abstract
This paper proposes to learn the relevant features of remote sensing images for automatic spatio-spectral classification with the automatic optimization of multiple kernels. The method consists of building dedicated kernels for different sets of bands, contextual or textural features. The optimal linear combination of kernels is optimized through gradient descent on the support vector machine (SVM) objective function. Since a naive implementation is computationally demanding, we propose an efficient model selection procedure based on kernel alignment. The result is a weight - learned from the data - for each kernel where both relevant and meaningless image features emerge after training. Excellent results are observed in both multi and hyperspectral image classification, improving standard SVM and other spatio-spectral formulations.
Devis Tuia, Giona Matasci, Gustau Camps-Valls, Mikhail F. Kanevski
IGARSS (2)1
2009 Semisupervised Remote Sensing Image Classification With Cluster Kernels
abstract
A semisupervised support vector machine is presented for the classification of remote sensing images. The method exploits the wealth of unlabeled samples for regularizing the training kernel representation locally by means of cluster kernels. The method learns a suitable kernel directly from the image and thus avoids assuming a priori signal relations by using a predefined kernel structure. Good results are obtained in image classification examples when few labeled samples are available. The method scales almost linearly with the number of unlabeled samples and provides out-of-sample predictions.
Devis Tuia, Gustau Camps-Valls
IEEE Geosci. Remote. Sens. Lett.1
2009 Decision Fusion for the Classification of Hyperspectral Data: Outcome of the 2008 GRS-S Data Fusion Contest
abstract
The 2008 Data Fusion Contest organized by the IEEE Geoscience and Remote Sensing Data Fusion Technical Committee deals with the classification of high-resolution hyperspectral data from an urban area. Unlike in the previous issues of the contest, the goal was not only to identify the best algorithm but also to provide a collaborative effort: The decision fusion of the best individual algorithms was aiming at further improving the classification performances, and the best algorithms were ranked according to their relative contribution to the decision fusion. This paper presents the five awarded algorithms and the conclusions of the contest, stressing the importance of decision fusion, dimension reduction, and supervised classification methods, such as neural networks and support vector machines.
Giorgio Licciardi, Fabio Pacifici, Devis Tuia, Saurabh Prasad, Terrance West, Ferdinando Giacco, Christian Thiel 0002, Jordi Inglada, Emmanuel Christophe, Jocelyn Chanussot, Paolo Gamba
IEEE Trans. Geosci. Remote. Sens.3
2009 Classification of Very High Spatial Resolution Imagery Using Mathematical Morphology and Support Vector Machines
abstract
We investigate the relevance of morphological operators for the classification of land use in urban scenes using sub-metric panchromatic imagery. A support vector machine is used for the classification. Six types of filters have been employed: opening and closing, opening and closing by reconstruction, and opening and closing top hat. The type and scale of the filters are discussed, and a feature selection algorithm called recursive feature elimination is applied to decrease the dimensionality of the input data. The analysis performed on two QuickBird panchromatic images showed that simple opening and closing operators are the most relevant for classification at such a high spatial resolution. Moreover, mixed sets combining simple and reconstruction filters provided the best performance. Tests performed on both images, having areas characterized by different architectural styles, yielded similar results for both feature selection and classification accuracy, suggesting the generalization of the feature sets highlighted.
Devis Tuia, Fabio Pacifici, Mikhail F. Kanevski, William J. Emery
IEEE Trans. Geosci. Remote. Sens.1
2009 Active Learning Methods for Remote Sensing Image Classification
abstract
In this paper, we propose two active learning algorithms for semiautomatic definition of training samples in remote sensing image classification. Based on predefined heuristics, the classifier ranks the unlabeled pixels and automatically chooses those that are considered the most valuable for its improvement. Once the pixels have been selected, the analyst labels them manually and the process is iterated. Starting with a small and nonoptimal training set, the model itself builds the optimal set of samples which minimizes the classification error. We have applied the proposed algorithms to a variety of remote sensing data, including very high resolution and hyperspectral images, using support vector machines. Experimental results confirm the consistency of the methods. The required number of training samples can be reduced to 10% using the methods proposed, reaching the same level of accuracy as larger data sets. A comparison with a state-of-the-art active learning method, margin sampling, is provided, highlighting advantages of the methods proposed. The effect of spatial resolution and separability of the classes on the quality of the selection of pixels is also discussed.
Devis Tuia, Frédéric Ratle, Fabio Pacifici, Mikhail F. Kanevski, William J. Emery
IEEE Trans. Geosci. Remote. Sens.1
2008 Socio-economic Data Analysis with Scan Statistics and Self-organizing Maps
Devis Tuia, Christian Kaiser, Antonio Da Cunha, Mikhail F. Kanevski
ICCSA (1)1
2008 Very-High Resolution Image Classification using Morphological Operators and SVM
abstract
An extensive analysis based on the use of different morphological filters for the classification of very-high resolution panchromatic images is presented. Feature selection on high-dimensional input space is performed using recursive feature elimination, a support vector machines specific method performing backward elimination based on margin-estimation criterion. Experimental results on an eight-classes image of Las Vegas (USA) confirmed the effectiveness of the analysis pointing out the relevancy of the most contributing morphological features which resulted in high classification accuracy using panchromatic imagery.
Devis Tuia, Fabio Pacifici, Alexei Pozdnoukhov, Christian Kaiser, Domenico Solimini, William J. Emery
IGARSS (4)1
2008 Active Learning of Very-High Resolution Optical Imagery with SVM: Entropy vs Margin Sampling
abstract
An active learning method is proposed for the semi-automatic selection of training sets in remote sensing image classification. The method adds iteratively to the current training set the unlabeled pixels for which the prediction of an ensemble of classifiers based on bagged training sets show maximum entropy. This way, the algorithm selects the pixels that are the most uncertain and that will improve the model if added in the training set. The user is asked to label such pixels at each iteration. Experiments using support vector machines (SVM) on an 8 classes QuickBird image show the excellent performances of the methods, that equals accuracies of both a model trained with ten times more pixels and a model whose training set has been built using a state-of-the-art SVM specific active learning method.
Devis Tuia, Frédéric Ratle, Fabio Pacifici, Alexei Pozdnoukhov, Mikhail F. Kanevski, Fabio Del Frate, Domenico Solimini, William J. Emery
IGARSS (4)1
2008 Support-Based Implementation of Bayesian Data Fusion for Spatial Enhancement: Applications to ASTER Thermal Images
abstract
In this letter, a general Bayesian data fusion (BDF) approach is proposed and applied to the spatial enhancement of ASTER thermal images. This method fuses information coming from the visible or near-infrared bands (15 times 15 m pixels) with the thermal infrared bands (90 times 90 m pixels) by explicitly accounting for the change of support. By relying on linear multivariate regression assumptions, differences of support size for input images can be explicitly accounted for. Due to the use of locally varying variances, it also avoids producing artifacts on the fused images. Based on a set of ASTER images over the region of Lausanne, Switzerland, the advantages of this support-based approach are assessed and compared to the downscaling cokriging approach recently proposed in the literature. Results show that improvements are substantial with respect to both visual and quantitative criteria. Although the method is illustrated here with a specific case study, it is versatile enough to be applied to the spatial enhancement problem in general. It thus opens new avenues in the context of remotely sensed images.
Dominique Fasbender, Devis Tuia, Patrick Bogaert, Mikhail F. Kanevski
IEEE Geosci. Remote. Sens. Lett.2